Business Voice Intent Presentation Methods, Devices, Media and Electronic Equipment

By performing speech transcription, verification, and intent recognition on outbound call audio data, and utilizing a feature vector extraction module and the XGBOOST model, the problem of low efficiency in speech data intent analysis in the marketing field is solved, achieving efficient and accurate intent recognition and marketing assistance.

CN115733925BActive Publication Date: 2026-03-10CHINA TELECOM CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-08-26
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies for analyzing intent from voice data in the marketing field are inefficient and rely heavily on human experience, resulting in inaccurate analysis results and failing to effectively support marketing campaigns.

Method used

Data processing is performed using a feature vector extraction module and an XGBOOST model, including an attention mechanism layer. Combining the feature vector extraction module and the XGBOOST model, a multi-channel speech recognition engine is used to verify outbound call audio data. The final text transcription result of the speech transcription is also verified. Finally, the data is obtained and output through a business-related keyword recognition model at the application layer.

Benefits of technology

It achieves automated processing and intelligent intent recognition of outbound call audio data, converts speech to text and verifies it, and uses a business-related keyword recognition model for intent analysis, thereby improving the accuracy and efficiency of voice data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115733925B_ABST
    Figure CN115733925B_ABST
Patent Text Reader

Abstract

This application relates to the field of intent recognition, disclosing a method, apparatus, medium, and electronic device for presenting business voice intent. The method includes: acquiring outbound call audio data; transcribing the outbound call audio data using at least two speech-to-text engines to obtain corresponding text transcription results; verifying the text transcription results of each speech-to-text engine and determining the final text transcription result of the outbound call audio data based on the verification results; acquiring a first business-related keyword from the final text transcription result based on a preset business-related keyword recognition model, wherein the business-related keyword recognition model includes a feature vector extraction module and an XGBOOST model connected in sequence; determining the intent information corresponding to the first business-related keyword; and outputting the intent information. This method improves the accuracy and efficiency of intent recognition and enhances the effectiveness of marketing assistance; furthermore, the entire method can be completed automatically, significantly reducing labor costs.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intent recognition, and in particular relates to a business voice intent presentation method and device, a medium and an electronic device. BACKGROUND

[0002] With the development of the Internet and information technology, the amount of data generated by humans is growing exponentially, and the era of big data has arrived. In the marketing field, big data technology has been widely used.

[0003] For structured data, big data technology can process to some extent; however, for non-structured data such as voice, big data technology cannot effectively process. Voice data is the key data in the marketing field, which leads to the fact that in the marketing field, the intention behind the voice data is mainly analyzed based on manual methods, which is low in efficiency, high in cost, and the analysis result is heavily dependent on the experience of manual work, and the auxiliary effect on marketing is limited. SUMMARY

[0004] In the technical field of intent recognition, in order to solve the above technical problems, the purpose of the present application is to provide a business voice intent presentation method, device, medium and electronic device.

[0005] According to an aspect of the present application, a business voice intent presentation method is provided, the method comprising:

[0006] obtaining outbound audio data;

[0007] transcribing the outbound audio data according to at least two voice transcription engines to obtain corresponding text transcription results;

[0008] verifying the text transcription results of each voice transcription engine, and determining the final text transcription result of the outbound audio data based on the verification result;

[0009] obtaining a first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model, wherein the business-related keyword recognition model comprises a feature vector extraction module and an XGBOOST model connected in sequence, the feature vector extraction module adds an attention mechanism layer, the feature vector extraction module is used to extract a discrete feature vector of the keyword, and the XGBOOST model is used to output a classification result of the keyword;

[0010] determining intent information corresponding to the first business-related keyword according to the first business-related keyword;

[0011] outputting the intent information.

[0012] According to another aspect of the present application, a business voice intent presentation device is provided, the device comprising:

[0013] An acquisition module configured to acquire outbound call audio data;

[0014] A transcription module configured to transcribe the outbound call audio data according to at least two voice transcription engines to obtain corresponding text transcription results;

[0015] A verification module configured to verify the text transcription results of each voice transcription engine, and determine a final text transcription result of the outbound call audio data based on the verification result;

[0016] A keyword acquisition module configured to acquire a first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model, wherein the business-related keyword recognition model comprises a feature vector extraction module and an XGBOOST model connected in sequence, the feature vector extraction module is added with an attention mechanism layer, the feature vector extraction module is used to extract discrete feature vectors of keywords, and the XGBOOST model is used to output classification results of keywords;

[0017] A determination module configured to determine intent information corresponding to the first business-related keyword according to the first business-related keyword;

[0018] An output module configured to output the intent information.

[0019] According to another aspect of the present application, a computer readable program medium is provided, which stores computer program instructions, when the computer program instructions are executed by a computer, the computer executes the method as described above.

[0020] According to another aspect of the present application, an electronic device is provided, the electronic device comprising:

[0021] A processor;

[0022] A memory, wherein computer readable instructions are stored on the memory, and when the computer readable instructions are executed by the processor, the method as described above is implemented.

[0023] The technical solutions provided by the embodiments of the present application can include the following beneficial effects:

[0024] The business voice intent presentation method provided in the application comprises the following steps: acquiring outbound audio data; transcribing the outbound audio data according to at least two voice transcription engines to obtain corresponding text transcription results; checking the text transcription results of each voice transcription engine, and determining the final text transcription result of the outbound audio data based on the checking result; obtaining a first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model, wherein the business-related keyword recognition model comprises a feature vector extraction module and an XGBOOST model connected in sequence, the feature vector extraction module adds an attention mechanism layer, the feature vector extraction module is used to extract a discrete feature vector of a keyword, and the XGBOOST model is used to output a classification result of the keyword; determining the intent information corresponding to the first business-related keyword according to the first business-related keyword; and outputting the intent information.

[0025] In this method, the outbound audio data is first acquired, then the voice transcription engine is used to transcribe the outbound audio data to obtain the text transcription result, then the business-related keyword recognition model is used to determine the business-related keyword, finally the intent information is determined using the business-related keyword, and the intent information is output, thereby realizing the processing and value mining of the unstructured data of the outbound audio data, obtaining the intent information based on the outbound audio data, assisting the marketing activities, improving the marketing effect, realizing the effective use of the outbound audio data, improving the intent recognition efficiency of the audio data, realizing the accurate recognition of the business-related keyword and the self-growth of the business-related keyword through the feature vector extraction module and the XGBOOST model with the added attention mechanism layer, thereby improving the accuracy of intent recognition and the effect of marketing assistance, checking the text transcription results of the multiple voice transcription engines, and determining the final text transcription result based on the checking result, thereby improving the accuracy of text transcription and the accuracy of intent recognition to a certain extent. In addition, the entire method can be automatically completed, the massive unstructured data can be quickly mined, and the labor cost is greatly reduced.

[0026] It should be understood that the above general description and the following detailed description are only exemplary and cannot limit the application. BRIEF DESCRIPTION OF DRAWINGS

[0027] The accompanying drawings, which are incorporated into and form part of the specification, illustrate an embodiment consistent with the application and, together with the specification, serve to explain the principles of the application.

[0028] Figure 1 is a system architecture schematic diagram of a business voice intent presentation method according to an exemplary embodiment;

[0029] Figure 2 FIG. 6 is a flowchart of a business voice intent presentation method according to an example embodiment;

[0030] Figure 3 FIG. 7 is a schematic diagram of a storage manner of a data and voice transcription engine according to an example embodiment;

[0031] Figure 4 FIG. 8 is a process schematic diagram of transcribing outbound audio data based on a two-way voice transcription engine according to an example embodiment;

[0032] Figure 5 FIG. 9 is a structural schematic diagram of a business-related keyword recognition model according to an example embodiment;

[0033] Figure 6 FIG. 10 is a data processing process schematic diagram when training a business-related keyword recognition model according to an example embodiment;

[0034] Figure 7 FIG. 11 is a schematic diagram of role recognition on a final text transcription result according to an example embodiment;

[0035] Figure 8 FIG. 12 is a schematic diagram of intent recognition by role according to an example embodiment;

[0036] Figure 9 FIG. 13 is a schematic diagram of an interface for displaying excellent scripts according to an example embodiment;

[0037] Figure 10 FIG. 14 is a schematic diagram of an interface for displaying statistical results of failure reasons according to an example embodiment;

[0038] Figure 11 FIG. 15 is a block diagram of a business voice intent presentation device according to an example embodiment;

[0039] Figure 12 FIG. 16 is an example block diagram of an electronic device for implementing the above business voice intent presentation method according to an example embodiment;

[0040] Figure 13 FIG. 17 is a program product for implementing the above business voice intent presentation method according to an example embodiment. DETAILED DESCRIPTION

[0041] The exemplary embodiments will be described in detail herein with reference to the attached drawings. The following description is made with reference to the accompanying drawings in which like reference numerals refer to like elements, and the term "exemplary" is used herein to mean "serving as an example, instance, or illustration." The following description is not intended to limit the scope of the present application, but is merely intended to describe the exemplary embodiments of the present application.

[0042] In addition, the accompanying drawings are included to provide a thorough understanding of the present application and are not intended to be exhaustive or to limit the present application to the precise outline described herein. The same or similar components are denoted by the same or similar reference numerals throughout the drawings, and repeated description thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, and do not necessarily correspond to physical or logical entities.

[0043] The present application first provides a business voice intention presentation method. The business voice is voice data generated by a marketing person and a customer in a conversation process in a marketing process. The marketing person herein can be an outbound marketing person or an offline marketing person, and the customer is a user in need of a product or a service. The intention of the business voice is the connotative information reflected by the conversation behavior of the conversation person in the voice data. The embodiment of the present application can efficiently and accurately identify the intention information of the conversation person in the voice data, and can display the intention information in multiple ways. The scheme in the embodiment of the present application can be applied to the field of marketing and customer service, such as the service industry with outbound marketing scenarios, such as communication, catering, real estate, insurance, etc., to improve the level of service and marketing. When the scheme of the embodiment of the present application is applied to the communication field, it can be specifically applied to the 5G business outbound marketing assistance.

[0044] The terminal of the present application can be any device with computing function, which can be connected with external devices for receiving or sending data. Specifically, it can be a portable mobile device, such as a smart phone, a tablet computer, a notebook computer, a PDA (Personal Digital Assistant), etc., or a fixed device, such as a computer device, a field terminal, a desktop computer, a server, a workstation, etc., or a collection of multiple devices, such as a physical infrastructure of cloud computing or a server cluster.

[0045] Optionally, the terminal of the present application can be a server or a physical infrastructure of cloud computing.

[0046] Figure 1 is a schematic diagram of a system architecture of a business voice intention presentation method according to an exemplary embodiment. As shown in Figure 1As shown, the system architecture includes a data layer, an algorithm layer, and an application layer. Specifically, the data layer is used to obtain outbound voice files, which can include voice files for marketing activities and customer service business. The algorithm layer includes a speech transcription engine and an NLP module. The outbound audio is transcribed into text through the speech transcription engine. After the transcribed text is processed for data cleaning, classification, and arrangement, the processed data is input into the NLP module, which can perform operations such as business dictionary construction, agent identification, similarity calculation, dynamic window detection, and finally intent recognition. The intent recognition result is output through the application layer. The application layer can include business analysis, marketing assistance, business opportunity insight, outbound reminder, excellent marketing technique, and business dynamic following modules. The outbound reminder module can feed back the customer's intent information to the outbound personnel in real time, so that the outbound personnel can better understand the user's intent and help the outbound personnel make better marketing decisions. The excellent marketing technique module can arrange and output excellent marketing techniques. For example, if a certain outbound personnel successfully completes marketing through voice dialogue, the technique information can be extracted from the voice data as an excellent marketing technique, so that the marketing personnel can learn the technique and continuously improve the marketing ability of the marketing personnel.

[0047] Figure 2 is a flowchart of a business voice intent presentation method according to an exemplary embodiment. The business voice intent presentation method provided in this embodiment can be executed by a server, such as Figure 2 As shown, the method includes the following steps:

[0048] Step 210: Obtain outbound audio data.

[0049] The outbound audio data is audio data generated by a customer service or agent through a dialogue with a customer in an outbound scenario. The outbound audio data can be outbound audio data generated by various channels.

[0050] The outbound audio data can be audio data accumulated and saved before the current time, or can be real-time acquired.

[0051] In one embodiment, the outbound audio data is obtained by acquiring real-time received outbound audio data.

[0052] Specifically, a certain agent is currently talking to a customer and recommending a product to the customer. At this time, the outbound audio data is acquired in real time, and the outbound audio data is subjected to intent recognition.

[0053] In the embodiment of the present application, by acquiring the outbound audio data when the agent communicates with the customer in real time, the intention recognition can be performed in real time based on the outbound audio data, and the agent can feed back the intention information of the customer in real time, so that the agent can adjust the dialogue strategy according to the intention information, thereby improving the success rate of marketing.

[0054] In step 220, the outbound audio data is transcribed according to at least two speech transcription engines to obtain corresponding text transcription results.

[0055] The speech transcription engine can transcribe the outbound audio data into text data, that is, the speech transcription engine includes a speech recognition model.

[0056] The at least two speech transcription engines are a plurality of different speech transcription engines, and each speech transcription engine can output the transcription result of the outbound audio data independently. These speech transcription engines can include the same or different speech recognition models. When the speech recognition models are different, these speech recognition models can use the same or different algorithms, and when the same algorithm is used, the parameters in each model can be different. The speech recognition model can use a long short-term memory (LSTM) network, a DFSMN model, etc.

[0057] In one embodiment, the outbound audio data and the text transcription result are stored in a cloud platform, and the speech transcription engine is deployed on the cloud platform.

[0058] Specifically, the speech transcription engine can be deployed in a private cloud platform in advance by using a local development server of the system. The outbound audio data can be directly uploaded to the cloud platform by the outbound system, and the text transcription result obtained after the speech transcription engine in the cloud platform transcribes the outbound audio data is also saved in the cloud platform.

[0059] In the embodiment of the present application, by storing the outbound audio data, the text transcription result and the speech transcription engine in the cloud platform, the outbound audio data and the speech transcription engine cannot be accessed and obtained by the terminal, thereby effectively ensuring the security of the data.

[0060] Figure 3 is a schematic diagram of a storage method of a data and speech transcription engine according to an exemplary embodiment. Please refer to Figure 3As shown, the private cloud platform includes outbound call cloud data, a speech-to-text engine, and outbound call transcription data. The outbound call cloud data includes campaign marketing outbound call cloud data and customer service outbound call cloud data. Campaign marketing outbound call cloud data can be the audio data generated when agents proactively call customers, while customer service outbound call cloud data can be the audio data generated when agents receive inquiries from customers. The outbound call transcription data includes campaign marketing outbound call transcription data and customer service outbound call transcription data; these are the text transcription results corresponding to the campaign marketing outbound call cloud data and customer service outbound call cloud data, respectively. The aforementioned outbound call audio data can belong to any type of outbound call cloud data.

[0061] Figure 4 This is a schematic diagram illustrating the process of transcribing outbound call audio data based on a two-channel speech-to-text engine, according to an exemplary embodiment. Figure 4 As shown, when two speech-to-text engines are deployed, a large amount of outbound call audio data is transcribed using the X and Y transcription engines in the dual-channel speech-to-text engine. The resulting transcription results can be divided into two categories: activity marketing outbound call transcription data and 10,000-number outbound call transcription data. The 10,000-number outbound call transcription data mentioned above is similar to the customer service outbound call transcription data mentioned earlier.

[0062] Figure 3 and Figure 4 The illustrated embodiment actually involves first transcribing pre-obtained outbound call audio data, and then constructing a business keyword library after obtaining the transcribed data. The business keyword library will be described in the following embodiments.

[0063] Step 230: Verify the text transcription results of each speech transcription engine, and determine the final text transcription result of the outbound call audio data based on the verification results.

[0064] Verifying the text transcription results of various speech-to-text engines is the process of determining the differences in the text transcription results. The verification results are whether the text transcription results are inconsistent and the specific circumstances of the inconsistency.

[0065] In one embodiment, the step of transcribing the outbound call audio data using at least two speech-to-text engines to obtain corresponding text transcription results includes: transcribing the outbound call audio data using two speech-to-text engines to obtain a first text transcription result and a second text transcription result; the step of verifying the text transcription results of each speech-to-text engine and determining the final text transcription result of the outbound call audio data based on the verification results includes: discarding the first and second transcribed statements if the number of consecutive inconsistent characters in the first transcribed statement in the first text transcription result and the second transcribed statement corresponding to the first transcribed statement in the second text transcription result reaches a predetermined number; retaining any one of the first and second transcribed statements if the number of consecutive inconsistent characters in the first transcribed statement in the first text transcription result and the second transcribed statement corresponding to the first transcribed statement in the second text transcription result is less than a predetermined number; and using all retained transcribed statements as the final text transcription result of the outbound call audio data.

[0066] Specifically, the first transcription statement corresponds to the second transcription statement. That is to say, the first transcription statement and the second transcription statement are transcription statements in the same position, or transcription statements corresponding to the same segment of audio data in the outbound call audio data. The predetermined number can be set arbitrarily as needed. For example, the predetermined number can be 10. When the number of consecutive inconsistent characters in the first and second transcribed statements is ≥10, both the first and second transcribed statements are discarded. When the number of consecutive inconsistent characters in the first and second transcribed statements is <10, there are two cases: If the first and second transcribed statements are consistent, that is, there are no inconsistent characters in the first and second transcribed statements, then either the first or the second transcribed statement is retained. If the first and second transcribed statements are inconsistent and the number of consecutive inconsistent characters is <10, in addition to retaining either the first or the second transcribed statement, the differing words can be extracted from the inconsistent parts of the first and second transcribed statements, and then a difference word dictionary can be constructed. The differing words are often homophones in the same position in the first and second transcribed statements. Therefore, the difference word dictionary can be used for text repair.

[0067] In one embodiment, retaining either the first or the second transcribing statement based on the fact that the number of consecutive inconsistent characters in the first transcribing statement in the first text transcribing result and the second transcribing statement corresponding to the first transcribing statement in the second text transcribing result is less than a predetermined number includes: retaining either the first or the second transcribing statement based on the fact that the first transcribing statement in the first text transcribing result and the second transcribing statement corresponding to the first transcribing statement in the second text transcribing result are consistent.

[0068] All retained transcription statements can be stored as the final text transcription result.

[0069] Please continue reading Figure 4 As shown, after obtaining the outbound call transcription data for marketing activities and the outbound call transcription data for Wanhao, the outbound call transcription data is transcribed out. Then, the transcription results corresponding to the two speech transcription engines are compared, and errors are corrected or erroneous sentences are discarded. Finally, the transcribed text is stored. The original words and verification words that do not match can be re-inputted into the two speech transcription engines to train them again and improve their performance.

[0070] Step 240: Based on the preset business-related keyword recognition model, obtain the first business-related keyword in the final text transcription result.

[0071] The business-related keyword recognition model includes a feature vector extraction module and an XGBOOST model connected in sequence. An attention mechanism layer is added to the feature vector extraction module. The feature vector extraction module is used to extract discrete feature vectors of keywords, and the XGBOOST model is used to output the classification results of keywords.

[0072] In one embodiment, the feature vector extraction module includes a bidirectional long short-term memory network module and a deep semantic matching model.

[0073] Long Short-Term Memory (LSTM) is a type of recurrent neural network. RNN and LSTM algorithms compress all past information into a single vector and pass it forward. Therefore, when sentences are long, a small vector dimension can easily lead to information compression or even loss. A bidirectional LSTM network module is essentially a bidirectional LSTM. Deep semantic matching models (DSSMs) can output low-dimensional vector representations. Therefore, a bidirectional LSTM network module is essentially a bidirectional LSTM-DSSM with an added attention mechanism layer. Since the LSTM-DSSM algorithm suffers from the vanishing gradient problem and long-distance dependency problem inherent in neural networks, adding an attention mechanism layer expands the external storage mechanism, avoiding information compression caused by simply passing a single hidden state vector forward. This improves prediction accuracy and operational efficiency.

[0074] The XGBOOST model, or eXtreme Gradient Boosting model, is a strong classifier model composed of multiple weak classifiers; it is a boosting tree model.

[0075] The business-related keyword recognition model can directly output the corresponding first business-related keyword based on the input of the final text transcription result; the business-related keyword recognition model can also output the corresponding first business-related keyword based on the input of a keyword sequence.

[0076] In one embodiment, before obtaining the first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model, the method further includes: extracting keywords from the final text transcription result to obtain a keyword sequence corresponding to the final text transcription result; obtaining the first business-related keyword in the final text transcription result based on the preset business-related keyword recognition model includes: inputting the keyword sequence into the business-related keyword recognition model to obtain the first business-related keyword in the final text transcription result.

[0077] Figure 5 This is a schematic diagram illustrating the structure of a business-related keyword recognition model according to an exemplary embodiment. For example... Figure 5 As shown, the business-related keyword recognition model includes not only bidirectional LSTM-DSSM and XGBOOST, but also an N-gram layer and a Max-pooling layer located between the bidirectional LSTM-DSSM and XGBOOST. XGBOOST consists of multiple decision trees constructed through tree splits.

[0078] N-gram is an algorithm based on statistical language models. Its basic idea is to process the text content into a sliding window of size N bytes, forming a sequence of byte segments of length N. Each byte segment is called a gram. The frequency of all grams is statistically analyzed, and then filtered according to a pre-defined threshold to form a list of key grams, which is the vector feature space of the text. Each gram in the list represents a feature vector dimension.

[0079] After obtaining a keyword sequence composed of keywords such as "5G" and "broadband", it is input into a business-related keyword recognition model. The model outputs word vectors through an N-gram layer, and then uses a weighted summation of the features of each feature in a bidirectional LTMN-DSSM model to output a hidden layer feature vector. The model automatically filters and combines features using the memory property of the neural network to generate a new discrete feature vector. This discrete feature vector is then used as the input to the XGBOOST model. The XGBOOST model can output the predicted probability of each keyword for each category, thus obtaining the model classification result for the keywords.

[0080] Figure 6 This is a schematic diagram illustrating the data processing procedure for training a business-related keyword recognition model according to an exemplary embodiment. The training and use of the business-related keyword recognition model can employ Spark 2.0 tools and the PySpark data framework. The data processing procedure is as follows: First, the modeling samples are organized, and the keywords in the modeling samples are vectorized to obtain word vectors. The modeling samples include keywords and corresponding business keyword tags. The modeling samples can be divided into a training set and a test set, with the training set containing 7562 training samples and the test set containing 1687 test samples. Next, the word vectors are used as input to the bidirectional LSTMN-DSSM model. The word vectors are processed by the first layer of the bidirectional LSTMN-DSSM model, and then converted into 1028-dimensional word vectors through the hidden layer of the first layer. Finally, the word vectors generated by the first layer are... The 1028-dimensional word vectors output from the model are input into the second layer of the bidirectional LSTMN-DSSM model. After processing, they are mapped to a 300-dimensional vector space to obtain 300-dimensional word vectors. Next, the 300-dimensional vectors are processed to obtain discrete feature vectors, which are then input into the XGBOOST model to output business keyword probabilities. Finally, based on the keyword probabilities, a threshold is set as a filtering condition to achieve business keyword recognition. Specifically, based on the predicted probability, a probability value greater than 0.5 is defined as a telecommunications professional term, and vice versa. Throughout the training process, the model parameters are adjusted based on the keyword recognition results, the corresponding business keyword labels, and the loss function. The model training is iteratively executed until the model meets the training termination condition.

[0081] After training is complete, the model's performance can be tested using a test set.

[0082] Based on the test set performance comparison, the traditional LSTM-DSSM algorithm achieved a classification accuracy of 0.67 and a recall of 0.87 for telecom service terms and non-service terms; the improved bidirectional LSTMN-DSSM-XGBOOST algorithm achieved a classification accuracy of 0.81 and a recall of 0.84 for telecom service terms and non-service terms, showing a significant improvement in performance.

[0083] Step 250: Determine the intent information corresponding to the first business-related keywords based on the first business-related keywords.

[0084] In one embodiment, the outbound call audio data is generated during the business marketing process, and determining the intent information corresponding to the first business-related keywords based on the first business-related keywords includes: querying the marketing result reason information that matches the first business-related keywords from the business keyword library based on the first business-related keywords, and using the marketing result reason information as the intent information corresponding to the first business-related keywords.

[0085] Specifically, marketing outcome information can include reasons for marketing success and reasons for marketing failure.

[0086] In one embodiment, before querying marketing result reason information matching the first business-related keywords from the business keyword library based on the first business-related keywords, the method further includes: acquiring pre-stored target outbound call audio data and marketing result information corresponding to the target outbound call audio data; transcribing the target outbound call audio data using at least two speech-to-text engines to obtain corresponding text transcription results; verifying the text transcription results of each speech-to-text engine and determining the final text transcription result of the target outbound call audio data based on the verification results; acquiring second business-related keywords from the final text transcription result based on a preset business-related keyword recognition model; and pushing the second business-related keywords and the marketing result information to a reason summarization end, so that users at the reason summarization end can summarize the marketing result reason information based on the second business-related keywords and the marketing result information, and add the second business-related keywords and the marketing result reason information to the business keyword library.

[0087] In other embodiments, after obtaining the second business-related keywords, they can be directly added to the business keyword library.

[0088] In this embodiment of the application, business keywords are self-generated by continuously performing text transcription and business-related keyword recognition.

[0089] The second business-related keywords can be the same as or different from the first business-related keywords. The difference between the second business-related keywords and the first business-related keywords is that the second business-related keywords are keywords established during the business keyword database building phase.

[0090] In one embodiment, the method further includes: if the intent information corresponding to the first business-related keyword cannot be determined, then the first business-related keyword is pushed to the reason summarization end, so that after the user of the reason summarization end summarizes the marketing result reason information corresponding to the first business-related keyword, the first business-related keyword and the marketing result reason information are added to the business keyword library.

[0091] The reasoning terminal can be a tool used by business experts. When the corresponding marketing result is a successful marketing campaign, the business expert can define the reasons for the success based on second-level business-related keywords; conversely, when the corresponding marketing result is a failed marketing campaign, the business expert can define the reasons for the failure based on second-level business-related keywords. Therefore, by searching the business keyword database, the reasons for marketing success or failure summarized by experts can be obtained, thereby achieving intent recognition.

[0092] Step 260: Output the intent information.

[0093] Intent information can be output in multiple forms of visualization, and the analysis results of intent information can also be presented in multiple forms of visualization on the web front end.

[0094] In one embodiment, determining the intent information corresponding to the first business-related keywords based on the first business-related keywords further includes: performing role recognition on the final text transcription result based on a logistic regression model to obtain text transcription results corresponding to the customer role and the agent role respectively; determining the intent information corresponding to the first business-related keywords based on the text transcription result corresponding to the customer role; and outputting the intent information includes: returning the intent information to the terminal where the agent role is located.

[0095] Role identification based on logistic regression model in the final text transcription result is actually the process of locating the roles of the two parties in the dialogue (agent and customer) in the final text transcription result. Figure 7 This is a schematic diagram illustrating role recognition of the final text transcription result according to an exemplary embodiment. For example... Figure 7As shown, the operation for role recognition in the final text transcription result is as follows: First, the dialogue content of the two roles, Spk0 and Spk1, is segmented and vectorized, where Spk0 and Spk1 are the labeled data corresponding to the dialogue content; next, the frequency of business words appearing in the dialogue content is added to the word vector to obtain the text vector; finally, the text vector and the corresponding labeled data are used as a manually labeled dataset and input into the Logistic Regression (LR) model for training. The Logistic Regression model can output the category corresponding to each sentence in the text transcription result. For example, for a sentence, if the Logistic Regression model predicts that the probability of it being an agent is 88% and the probability of it being a customer is 12%, then the Logistic Regression model will predict the sentence as Spk0: Agent; for another sentence, if the Logistic Regression model predicts that the probability of it being a customer is 88% and the probability of it being an agent is 12%, then the Logistic Regression model will predict the sentence as Spk1: Customer.

[0096] In other embodiments of this application, role recognition can also be performed based on other classification models.

[0097] In one embodiment, the method further includes: when marketing result information corresponding to the outbound call audio data is obtained, and marketing success is determined based on the marketing result information, extracting marketing script information from the text transcription result corresponding to the agent role, and saving the marketing script information.

[0098] Marketing script information is equivalent to the intentions of the agents; therefore, extracting marketing script information is equivalent to identifying the intentions of the agents.

[0099] Figure 8 This is a schematic diagram illustrating role-based intent recognition according to an exemplary embodiment. For example... Figure 8 As shown, on the one hand, based on the keyword self-growth model algorithm, business experts summarize the lexicon to define the key reasons for marketing success / failure. The customer's original language is matched with the keyword lexicon, ultimately outputting the expert-summarized reasons for success / failure, thus achieving the recognition of the customer's stated intent. On the other hand, based on marketing results and combined with the fluency of the translated dialogue, key marketing scripts for customer service from different businesses are selected to achieve intelligent extraction of excellent customer service scripts. Figure 8 In the process, business-related keywords are continuously extracted from the final text transcription result through a sliding window, achieving customer keyword capture and obtaining marketing success keywords for customer intent identification. Simultaneously, based on the agent semantic recognition model, excellent agent dialogue is intelligently extracted. Both customer intent identification marketing success keywords and excellent agent dialogue can be output through the page. Figure 8You can also see a button or hyperlink on the page to listen to the original recording. When the agent clicks the button or hyperlink, they can hear the original recording of the excellent sales script. This allows them to learn the tone and style of the script and further improve their marketing skills.

[0100] Since outbound calls are a form of dialogue between agents and customers, their words are intertwined in the transcribed text. Therefore, this application embodiment is based on agent semantic recognition and differentially captures the keyword content of the dialogue text between agents and customers, realizing intelligent extraction of customer service scripts and interpretation of customer language intent.

[0101] In one embodiment, the method further includes: summarizing the intent information and the marketing script information, and outputting the summary results in a visual manner.

[0102] Visualization can be achieved using Tableau, and the summarized results can be sent via outbound calls to assist in analysis reports.

[0103] After intelligently acquiring intent information and marketing script information through the above methods, this information can be displayed in various forms. Furthermore, the intent information and marketing script information can be further processed for display, such as using multi-form echarts graphical tools to present it in the form of word cloud dot plots, Pareto charts, and circular scale charts. Figure 9 This is a schematic diagram illustrating an interface for displaying excellent conversational skills, according to an exemplary embodiment. Figure 9 As shown, the page not only displays successful and excellent sales scripts, but also successful keywords. Each successful keyword is displayed in the form of a word cloud dot plot. The size of the word cloud dot plot represents the frequency of the keyword's occurrence, so that the successful keywords that can ultimately improve marketing results can be understood by the agents more quickly.

[0104] Figure 10 This is a schematic diagram illustrating an interface displaying statistical results of failure reasons, according to an exemplary embodiment. Please refer to [link / reference]. Figure 10 As shown, the page displays statistical results indicating the reasons for failure. These results are presented in the form of a pie chart, detailing the percentage of reasons for each failed marketing campaign. For example, the percentage for "the package is sufficient" could be 30.31%, etc. Figure 10 The document also displays the top 10 failed keywords, showcasing the 10 keywords most frequently mentioned by customers in voice data related to marketing failures; the top 10 failed keywords can be dynamically presented in real time according to keywords / content.

[0105] In summary, this application's embodiments, based on outbound call data, a speech-to-text engine, NLP algorithms, and application software programming, are deployed on a general-purpose server and have been applied on the company's big data capability platform, successfully achieving intent recognition of outbound marketing recordings. By leveraging outbound call script parsing capabilities, highly acceptable marketing scripts are extracted and output. In Hefei, this has resulted in a more than 150% increase in successful orders from proactive outbound marketing compared to ordinary scripts, with an estimated revenue increase exceeding 100 million yuan. Furthermore, by performing intent recognition on customer speech-to-text data, this application can pinpoint the root causes of marketing failures and identify customers with 5G needs for remarketing.

[0106] This application also provides a business voice intent presentation device, and the following are device embodiments of this application.

[0107] Figure 11 This is a block diagram illustrating a business voice intent presentation device according to an exemplary embodiment. Figure 11 As shown, the device 1100 includes:

[0108] Module 1110 is configured to acquire outbound call audio data;

[0109] The transcription module 1120 is configured to transcribe the outbound call audio data according to at least two speech transcription engines to obtain the corresponding text transcription result.

[0110] The verification module 1130 is configured to verify the text transcription results of each speech transcription engine and determine the final text transcription result of the outbound call audio data based on the verification results.

[0111] The keyword acquisition module 1140 is configured to acquire the first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model. The business-related keyword recognition model includes a feature vector extraction module and an XGBOOST model connected in sequence. An attention mechanism layer is added to the feature vector extraction module. The feature vector extraction module is used to extract discrete feature vectors of the keywords. The XGBOOST model is used to output the classification result of the keywords.

[0112] The determination module 1150 is configured to determine the intent information corresponding to the first service-related keywords based on the first service-related keywords;

[0113] Output module 1160 is configured to output the intent information.

[0114] According to a third aspect of this application, an electronic device capable of implementing the above-described method is also provided.

[0115] Those skilled in the art will understand that various aspects of this application can be implemented as a system, method, or program product. Therefore, various aspects of this application can be specifically implemented in the following forms: a completely hardware implementation, a completely software implementation (including firmware, microcode, etc.), or a combination of hardware and software implementations, collectively referred to herein as a "circuit," "module," or "system."

[0116] The following reference Figure 12 To describe an electronic device 1200 according to this embodiment of the present application. Figure 12 The electronic device 1200 shown is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.

[0117] like Figure 12 As shown, the electronic device 1200 is manifested in the form of a general-purpose computing device. The components of the electronic device 1200 may include, but are not limited to: at least one processing unit 1210, at least one storage unit 1220, and a bus 1230 connecting different system components (including storage unit 1220 and processing unit 1210).

[0118] The storage unit stores program code that can be executed by the processing unit 1210, causing the processing unit 1210 to perform the steps described in the "Embodiment Methods" section above according to various exemplary embodiments of this application.

[0119] Storage unit 1220 may include a readable medium in the form of a volatile storage unit, such as random access memory (RAM) 1221 and / or cache memory 1222, and may further include a read-only memory (ROM) 1223.

[0120] Storage unit 1220 may also include a program / utility 1224 having a set (at least one) of program modules 1225, such program modules 1225 including but not limited to: operating system, one or more application programs, other program modules and program data, each or some combination of these examples may include an implementation of a network environment.

[0121] Bus 1230 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any of the various bus structures.

[0122] Electronic device 1200 can also communicate with one or more external devices 1400 (e.g., keyboard, pointing device, Bluetooth device, etc.), and with one or more devices that enable a user to interact with electronic device 1200, and / or with any device that enables electronic device 1200 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be made via input / output (I / O) interface 1250, such as communication with display unit 1240. Furthermore, electronic device 1200 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via network adapter 1260. As shown, network adapter 1260 communicates with other modules of electronic device 1200 via bus 1230. It should be understood that, although not shown in the figures, other hardware and / or software modules can be used in conjunction with electronic device 1200, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.

[0123] Through the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the method according to the embodiments of this application.

[0124] According to a fourth aspect of this application, a computer-readable storage medium is also provided, on which a program product capable of implementing the methods described above is stored. In some possible embodiments, various aspects of this application may also be implemented as a program product comprising program code that, when the program product is run on a terminal device, causes the terminal device to perform the steps of the various exemplary embodiments of this application described in the "Exemplary Methods" section above.

[0125] refer to Figure 13 As shown, a program product 1300 for implementing the above-described method according to an embodiment of this application is described. It may employ a portable compact disc read-only memory (CD-ROM) and include program code, and can run on a terminal device, such as a personal computer. However, the program product of this application is not limited thereto. In this document, a readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0126] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0127] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium, capable of sending, propagating, or transmitting programs for use by or in conjunction with an instruction execution system, apparatus, or device.

[0128] The program code contained on the readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.

[0129] Program code for performing the operations of this application can be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, and conventional procedural programming languages ​​such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).

[0130] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this application, and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0131] It should be understood that this application is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A business voice intent presentation method, characterized by, The method comprises: acquiring outbound call audio data; transcribing the outbound call audio data according to two-way speech transcription engines to obtain a first text transcription result and a second text transcription result; discarding a first transcription sentence in the first text transcription result and a second transcription sentence corresponding to the first transcription sentence in the second text transcription result when the number of consecutive inconsistent characters in the first transcription sentence and the second transcription sentence reaches a predetermined number; retaining either the first transcription sentence or the second transcription sentence when the number of consecutive inconsistent characters in the first transcription sentence and the second transcription sentence is less than the predetermined number; taking all retained transcription sentences as a final text transcription result of the outbound call audio data; obtaining a first business-related keyword in the final text transcription result based on a preset business-related keyword recognition model, wherein the business-related keyword recognition model comprises a feature vector extraction module and an XGBOOST model connected in sequence, the feature vector extraction module is added with an attention mechanism layer, the feature vector extraction module is used to extract a discrete feature vector of a keyword, and the XGBOOST model is used to output a classification result of the keyword; determining intent information corresponding to the first business-related keyword according to the first business-related keyword; outputting the intent information.

2. The method of claim 1, wherein, The outbound call audio data and the text transcription result are stored in a cloud platform, and the speech transcription engine is deployed on the cloud platform.

3. The method of claim 1, wherein, The outbound call audio data is generated in a business marketing process, and the determination of the intent information corresponding to the first business-related keyword according to the first business-related keyword comprises: querying marketing result reason information matching the first business-related keyword from a business keyword library according to the first business-related keyword, and taking the marketing result reason information as the intent information corresponding to the first business-related keyword.

4. The method of claim 3, wherein, Before querying the marketing result reason information matching the first business-related keyword from the business keyword library according to the first business-related keyword, the method further comprises: acquiring pre-stored target outbound call audio data and marketing result information corresponding to the target outbound call audio data; transcribing the target outbound call audio data according to at least two-way speech transcription engines to obtain corresponding text transcription results; verifying the text transcription results of each way of speech transcription engine, and determining a final text transcription result of the target outbound call audio data based on a verification result; obtaining a second business-related keyword in the final text transcription result based on a preset business-related keyword recognition model; pushing the second business-related keyword and the marketing result information to a reason induction terminal, so that a user of the reason induction terminal induces marketing result reason information according to the second business-related keyword and the marketing result information, and adds the second business-related keyword and the marketing result reason information to the business keyword library.

5. The method of claim 3, wherein, The determining, according to the first service-related keyword, the intention information corresponding to the first service-related keyword further includes: Performing role recognition on the final text transcription result based on a logistic regression model to obtain text transcription results corresponding to a customer role and an agent role, respectively; Determining, according to the text transcription result corresponding to the customer role of the first service-related keyword, the intention information corresponding to the first service-related keyword; The outputting of the intention information includes: Returning the intention information to a terminal where the agent role is located.

6. The method of claim 5, wherein, The method further includes: When the marketing result information corresponding to the outbound audio data is acquired, extracting marketing script information from the text transcription result corresponding to the agent role according to the marketing result information, and saving the marketing script information, if the marketing is successful.

7. A service voice intent presentation apparatus, characterized by, The device includes: An acquisition module configured to acquire outbound audio data; A transcription module configured to transcribe the outbound audio data according to two speech transcription engines to obtain a first text transcription result and a second text transcription result; A verification module configured to discard a first transcription sentence in the first text transcription result and a second transcription sentence corresponding to the first transcription sentence in the second text transcription result if the number of continuously inconsistent characters in the first transcription sentence and the second transcription sentence reaches a predetermined number, and to retain any one of the first transcription sentence and the second transcription sentence if the number of continuously inconsistent characters in the first transcription sentence and the second transcription sentence is less than the predetermined number, and to take all retained transcription sentences as a final text transcription result of the outbound audio data; A keyword acquisition module configured to acquire a first service-related keyword in the final text transcription result based on a preset service-related keyword recognition model, wherein the service-related keyword recognition model includes a feature vector extraction module and an XGBOOST model connected in sequence, the feature vector extraction module is added with an attention mechanism layer, the feature vector extraction module is used to extract a discrete feature vector of a keyword, and the XGBOOST model is used to output a classification result of the keyword; A determination module configured to determine, according to the first service-related keyword, the intention information corresponding to the first service-related keyword; An output module configured to output the intention information.

8. A computer readable program medium characterized in that, The computer program instructions stored therein cause the computer to execute the method according to any one of claims 1 to 6 when the computer program instructions are executed by the computer.

9. An electronic device, comprising: The electronic device includes: A processor; A memory having computer-readable instructions stored thereon, the computer-readable instructions being executed by the processor to implement the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Annotation method and device of voice data, equipment and computer storage medium

    CN109599095A

  • Voice data intention determination method and device, computer equipment and storage medium

    CN110162633A