Analysis and recommendation method and device for dialogue analysis and gold-brand verbal skill mining

Through dialogue preprocessing, coding and semantic analysis, combining cluster analysis and business indicator association, gold medal speech is identified and recommended, and the problems of inefficiency of dialogue analysis and speech recommendation and lack of semantic matching layers in the existing technology are solved, and more efficient and accurate dialogue analysis and speech recommendation are achieved.

CN119988550APending Publication Date: 2025-05-13BEIJING ZHUGE YUNYOU TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510084337.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art has problems with inefficiency, noise and redundant data in dialogue analysis and speech recommendation, making it difficult to capture the correlation between complex dialogue emotions and multiple rounds of dialogue, and the recommendation algorithm lacks a deep semantic matching layer.

Method used

By preprocessing the conversation content provided by the user, structured conversation data is extracted, and it is encoded and semantic analysis is performed to generate semantic coded vectors. Based on these vectors, cluster analysis and business indicator correlation are carried out, positive related speeches are identified, and matched with the standard speech library through the matching layer, and the gold medal speech with the highest similarity is recommended.

Benefits of technology

It improves the efficiency and accuracy of dialogue analysis, can effectively capture the correlation between complex dialogue emotions and multiple rounds of dialogue, improves the accuracy and effectiveness of speech recommendations, and adapts to the needs of different business scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988550A_ABST
    Figure CN119988550A_ABST
Patent Text Reader

Abstract

The invention discloses an analysis and recommendation method for dialogue analysis and gold-brand verbal skill mining, which comprises the following steps: preprocessing dialogue content provided by a user, and extracting structured dialogue data; encoding the dialogue data to generate semantic representation data of the dialogue; performing semantic analysis on the semantic representation data to generate a semantic coding vector of the dialogue; performing clustering analysis on different dialogues based on semantic coding vectors of the dialogues, and identifying forward correlation verbal skills in combination with business indexes; matching the forward correlation verbal skill with a predefined standard verbal skill library through a matching layer, and selecting the verbal skill with the highest similarity as a recommended golden verbal skill; and displaying the analyzed and mined recommended gold-brand verbal skill at a user side.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recommendation, and in particular to an analysis and recommendation method and device for dialogue analysis and gold medal speech mining. Background Art

[0002] In the field of modern customer service, sales, and business communication, companies are increasingly focusing on improving customer conversion rates, customer satisfaction, and employee service quality through conversation content analysis. At present, conversation analysis and speech recommendation technology are gradually becoming key tools in enterprise management and operations. Through the analysis of historical conversation data and speech mining, companies can effectively identify high-quality speech, improve employee service performance, and enhance the overall customer experience. However, existing technologies still have some shortcomings, which are mainly reflected in the following aspects:

[0003] Traditional conversation parsing methods mostly rely on manual annotation or simple rule matching. This method is acceptable when the amount of data is small, but it is inefficient when faced with large-scale and diverse customer conversation data, and is prone to problems such as noise and redundant data. The lack of automated preprocessing means makes it difficult to efficiently and accurately extract structured information from conversation data. In existing conversation parsing technologies, most systems only understand the semantics of conversations on the surface, and it is difficult to effectively capture complex conversation emotions and the correlation between multiple rounds of conversations. The existing technology usually encodes the content of the conversation in a simple way, which makes it difficult to generate deep semantic representations, affecting the accuracy of conversation clustering and speech recommendation.

[0004] Many current speech mining systems are often based on keyword matching or simple rule definitions in conversations, and are unable to automatically identify positively related speech, i.e., "golden speech", from a large amount of conversation data. This method is not only time-consuming and labor-intensive, but also difficult to adapt to the needs of different business scenarios and unable to quickly respond to changes in business needs. Existing conversation speech recommendation systems often fail to accurately match the most suitable standard speech, and the accuracy and effectiveness of recommended speech are low. The recommendation algorithm lacks a deep semantic matching layer, which leads to the situation that when recommending golden speech, the recommended speech is easily inconsistent with the actual business needs, affecting the actual application effect.

[0005] Therefore, there is an urgent need for an analysis and recommendation method and device for dialogue analysis and gold medal speech mining. Summary of the invention

[0006] The present invention provides an analysis and recommendation method and device for dialogue analysis and gold medal speech mining to solve the above-mentioned problems existing in the prior art.

[0007] In order to achieve the above object, the present invention provides the following technical solutions:

[0008] Analysis and recommendation methods for conversation analysis and gold medal speech mining, including:

[0009] S101: pre-processing the conversation content provided by the user to extract structured conversation data;

[0010] S102: Encoding the dialogue data to generate semantic representation data of the dialogue;

[0011] S103: performing semantic analysis on the semantic representation data to generate a semantic encoding vector of the conversation;

[0012] S104: Based on the semantic coding vectors of the conversations, cluster analysis is performed on different conversations, and positively related conversations are identified in combination with business indicators;

[0013] S105: Match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words;

[0014] S106: The recommended golden words analyzed and mined are displayed on the user side.

[0015] Among them, S101 includes:

[0016] S1011: Clean the conversation content provided by the user, remove redundant information and noise data, and obtain the pre-processed conversation content, where the conversation content includes customer service, business conversations, and multiple rounds of conversations in sales;

[0017] S1012: labeling the preprocessed conversation content, and marking attribute information of each round of conversation according to the characteristics of the conversation content. The attribute information includes basic information of users mentioned in the conversation, whether the order is successful, emotional tendency of the conversation, and key question types;

[0018] S1013: extracting structured conversation data from the labeled conversation data, and converting the conversation data into a unified data format.

[0019] The semantic representation data of the generated dialogue includes:

[0020] Get multiple encoding strategies corresponding to the conversation data;

[0021] Traverse the encoding strategies in sequence, and obtain the feature-semantic representation library corresponding to the traversed encoding strategy each time;

[0022] Splitting the conversation data into multiple conversation data items;

[0023] Perform feature extraction on the conversation data items to obtain multiple features;

[0024] Based on the feature-semantic representation library, determine the semantic representation corresponding to the feature and associate it with the corresponding conversation data item;

[0025] Accumulating and calculating the semantic representations associated with the second conversation data item to obtain a semantic representation sum;

[0026] Obtaining the semantic representation and threshold corresponding to the traversed first encoding strategy, and if the semantic representation and threshold are greater than or equal to the semantic representation and threshold, taking the corresponding second conversation data item as the third conversation data item;

[0027] Obtaining an adaptation semantic model corresponding to the traversed first coding strategy, performing adaptation semantic representation on the traversed first coding strategy based on the adaptation semantic model and according to the dialogue data, obtaining an adaptation value, and associating the adaptation value with the traversed coding strategy;

[0028] When the traversal of the coding strategies is finished, the adaptation values ​​associated with the coding strategies are accumulated and calculated to obtain the adaptation value sum;

[0029] The maximum adaptation value and the corresponding coding strategy are used as the final coding strategy;

[0030] Based on the final coding strategy, the conversation data is encoded to obtain the coding result.

[0031] The semantic encoding vector of the generated dialogue includes:

[0032] Acquire a preset semantic representation data set, the semantic representation data set including: a plurality of semantic nodes;

[0033] Get the semantic feature value corresponding to the semantic node;

[0034] If the semantic feature value meets the preset semantic feature value threshold, the corresponding semantic node is used as the target node;

[0035] Acquire at least one semantic feature item corresponding to the conversation content through the target node;

[0036] Integrate various semantic feature items to generate the semantic encoding vector of the conversation content and complete semantic parsing.

[0037] Among them, cluster analysis is performed on different conversations, including:

[0038] Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content.

[0039] Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded;

[0040] Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators.

[0041] Mark the conversation content that is positively correlated with business indicators as positively correlated words, and positively correlated words are gold medal words.

[0042] Among them, the matching layer matches the positive related words with the predefined standard words library, including:

[0043] Obtain a predefined standard script library, which includes: multiple script samples;

[0044] Obtain positive related words through the matching layer;

[0045] Calculate the similarity between the positive related speech and each speech sample in the standard speech library;

[0046] Calculate the similarity value between each speech sample and the semantic vector of the positively related speech;

[0047] Use the Softmax layer to convert the similarity value into a probability value to obtain the matching probability value of each speech sample;

[0048] Select the speech sample with the highest matching probability value as the recommended golden speech;

[0049] Complete the selection of recommended golden words.

[0050] Among them, the recommended golden words analyzed and mined are displayed on the user side, including:

[0051] Display recommended golden words on the user side and provide dynamic update function so that the recommended words can be adjusted and optimized according to real-time analysis data;

[0052] The recommended golden words are displayed in grades, and the efficient words and optimization suggestions are marked separately;

[0053] Based on user feedback or business personnel's choice, adjust the display priority and display method of recommended golden words to ensure the best display effect;

[0054] The recommended golden scripts will be presented to the user in a form that is easy to understand and apply, including the type of scripts, usage scenarios, and application effect evaluation information.

[0055] Among them, the analysis and recommendation devices used for dialogue analysis and gold medal speech mining include:

[0056] A dialogue extraction unit, used to pre-process the dialogue content provided by the user and extract structured dialogue data;

[0057] A speech coding unit is used to encode the dialogue data and generate semantic representation data of the dialogue;

[0058] A semantic encoding unit, which is used to perform semantic analysis on the semantic representation data and generate a semantic encoding vector of the dialogue;

[0059] The speech mining unit is used to perform cluster analysis on different conversations based on the semantic coding vectors of the conversations, and identify positively related speech based on business indicators;

[0060] The golden words recommendation unit is used to match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words;

[0061] The speech display unit is used to display the recommended golden speech that has been analyzed and mined on the user side.

[0062] The dialogue extraction unit includes:

[0063] The first dialogue extraction module is used to clean the dialogue content provided by the user, remove redundant information and noise data, and obtain the pre-processed dialogue content, where the dialogue content includes customer service, business dialogue and multiple rounds of dialogue in sales;

[0064] The second dialogue extraction module is used to label the preprocessed dialogue content. According to the characteristics of the dialogue content, the attribute information of each round of dialogue is marked. The attribute information includes the basic information of the user mentioned in the dialogue, whether the order is completed, the emotional tendency of the dialogue, and the type of key questions.

[0065] The third module of conversation extraction is used to extract structured conversation data from the labeled conversation data and convert the conversation data into a unified data format.

[0066] Among them, cluster analysis is performed on different conversations, including:

[0067] Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content.

[0068] Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded;

[0069] Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators.

[0070] Mark the conversation content that is positively correlated with business indicators as positively correlated words, and positively correlated words are gold medal words.

[0071] Compared with the prior art, the present invention has the following advantages:

[0072] The analysis and recommendation method for conversation analysis and gold medal speech mining includes: preprocessing the conversation content provided by the user to extract structured conversation data; encoding the conversation data to generate the semantic representation data of the conversation; semantically parsing the semantic representation data to generate the semantic encoding vector of the conversation; clustering analysis of different conversations based on the semantic encoding vector of the conversation, and identifying positive related speech in combination with business indicators; matching the positive related speech with the predefined standard speech library through the matching layer, and selecting the speech with the highest similarity as the recommended gold medal speech; displaying the analyzed and mined recommended gold medal speech on the user side. The system can continuously optimize the analysis capability of the semantic model according to the new conversation data, so that the analysis accuracy is continuously improved to adapt to the diversified needs in different business scenarios.

[0073] Other features and advantages of the present invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of the present invention.

[0074] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0076] Figure 1 It is a flow chart of an analysis and recommendation method for conversation analysis and gold medal speech mining in an embodiment of the present invention;

[0077] Figure 2 A flowchart of extracting structured conversation data in an embodiment of the present invention;

[0078] Figure 3 The structure diagram of the analysis and recommendation device for dialogue analysis and gold medal speech mining in an embodiment of the present invention. DETAILED DESCRIPTION

[0079] The preferred embodiments of the present invention are described below in conjunction with the accompanying drawings. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present invention, and are not used to limit the present invention.

[0080] The embodiment of the present invention provides an analysis and recommendation method for dialogue analysis and gold medal speech mining, including:

[0081] S101: pre-processing the conversation content provided by the user to extract structured conversation data;

[0082] S102: Encoding the dialogue data to generate semantic representation data of the dialogue;

[0083] S103: performing semantic analysis on the semantic representation data to generate a semantic encoding vector of the conversation;

[0084] S104: Based on the semantic coding vectors of the conversations, cluster analysis is performed on different conversations, and positively related conversations are identified in combination with business indicators;

[0085] S105: Match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words;

[0086] S106: The recommended golden words analyzed and mined are displayed on the user side.

[0087] The working principle of the above technical solution is as follows: the system first pre-processes the conversation data provided by the user and extracts structured conversation content from the raw data. These conversation data usually come from scenarios such as customer service, business conversations, sales, etc., including multiple rounds of conversations between customers and service personnel. The content of each round of conversation is sorted and classified through labels (such as whether the order is completed), historical conversation records, user attributes, questions and replies, etc., to form a standardized input data set.

[0088] In the speech coding stage, the system encodes each conversation data and generates a semantic representation of the conversation. The coding process extracts key information in the conversation (such as customer questions, customer service replies, historical conversations, etc.) and represents it as structured information units (such as i, q, a, u). Here, i represents the order mark of the conversation, q represents the question raised by the customer, a is the customer service reply content, and u is the user's related attribute information. The encoded data serves as the basic input for large model analysis.

[0089] The semantic encoding unit is mainly composed of a token-level encoder module, a sentence-level encoder module, and a user attribute extraction module. The system uses a large model to perform semantic analysis and deep parsing on the encoded conversation data. The large model generates a semantic vector representation (such as h1, h2, ..., hM) for each conversation through a deep understanding of the semantic information in the conversation. These vector representations not only capture the explicit information in the conversation (such as keywords and sentiment tendencies), but also identify the complex semantic relationships and contextual connections implied in the conversation, thereby providing high-quality semantic information for subsequent analysis.

[0090] Based on the vector representation generated by semantic coding, the system performs cluster analysis on different conversations and automatically identifies positively correlated speech strategies (i.e., gold medal speech strategies) and negatively correlated speech strategies by associating them with business indicators (such as whether an order is placed). The system can filter out outstanding speech strategies, such as those communication methods that significantly improve customer satisfaction or conversion rate during customer service. Positively correlated speech strategies are gold medal speech strategies that are automatically identified by the system.

[0091] The system matches the generated semantic encoding vector h with the predefined standard speech library through the matching layer to obtain the recommended speech that best fits the current context. The matching layer calculates the similarity between each speech d and the semantic vector h, and converts the similarity into a probability value p through the Softmax layer. The system finally selects the speech with the highest similarity as the recommendation result and outputs it to the user, that is, finding the speech that best fits the current situation.

[0092] Finally, the system will display the results of the gold medal speech found through analysis and mining, and recommend the most suitable speech in different conversation scenarios to help enterprises make optimization adjustments in business strategy formulation. These analysis results are presented in an easy-to-understand form, including high-frequency customer problems and their corresponding best solutions, and excellent speech categories, for reference and application by business personnel.

[0093] The beneficial effects of the above technical solution are: by combining large models with automated coding and analysis methods, the system can process massive conversation data without human intervention, significantly improving analysis efficiency. The system uses the semantic understanding ability of the large model to dig out detailed communication strategies from complex multi-round conversations, including positive (efficient) and negative (inefficient) speech, and automatically cluster and classify them to generate analysis results with high business value. The system can continuously optimize the analytical capabilities of the semantic model based on new conversation data, so that the analysis accuracy continues to improve and adapt to the diverse needs in different business scenarios. Help companies identify "gold medal speech" and guide customer service and sales teams to optimize strategies in actual applications, thereby improving customer service quality, conversion rate and overall business performance.

[0094] In another embodiment, S101 includes:

[0095] S1011: Clean the conversation content provided by the user, remove redundant information and noise data, and obtain the pre-processed conversation content, where the conversation content includes customer service, business conversations, and multiple rounds of conversations in sales;

[0096] S1012: labeling the preprocessed conversation content, and marking attribute information of each round of conversation according to the characteristics of the conversation content. The attribute information includes basic information of users mentioned in the conversation, whether the order is successful, emotional tendency of the conversation, and key question types;

[0097] S1013: extracting structured conversation data from the labeled conversation data, and converting the conversation data into a unified data format.

[0098] Among them, labeling processing is performed on the preprocessed conversation content, including:

[0099] Determine characteristic information of the conversation content;

[0100] Calculate the order probability, sentiment tendency and key question type of each conversation content to determine the overall attribute information of the conversation content;

[0101] If the order probability is lower than the preset threshold, the number of units with negative sentiment in the dialogue unit is counted to determine the impact of negative sentiment on the dialogue content and whether the impact exceeds the positive guidance ability of the dialogue content;

[0102] If not, use the ability of positive guidance to adjust negative emotions;

[0103] If so, the preset golden strategy is enabled and applied to the conversation content to improve the emotional tendency of the conversation. The goal of the golden strategy is to increase the probability of closing a deal and optimize the emotional tendency.

[0104] Calculation of the order probability, sentiment tendency, and key question types for each dialogue unit includes:

[0105] Get the text content and context information of each dialogue unit;

[0106] Compare the text content with the preset order-making model to determine whether the conversation content has the potential to make an order, and analyze the frequency and type of key issues;

[0107] Determine the emotional tendency of the conversation;

[0108] The attribute information of the dialogue unit is marked according to the potential for order formation, key question types and emotional tendencies.

[0109] The working principle of the above technical solution is as follows: First, the conversation content provided by the user needs to be cleaned to remove redundant information and noise data. Redundant information includes repeated conversation content, irrelevant information or unnecessary sentences, while noise data refers to grammatical errors, spelling errors or irrelevant emotional expressions. The goal of cleaning is to make the conversation content more concise and easier to analyze.

[0110] Suppose there is a conversation like this:

[0111] User A: “Hello, I would like to know the price of this product.”

[0112] Customer Service B: "Hello, thank you very much for your inquiry! We are happy to help you."

[0113] User A: "Then please tell me the price. Why is it so slow? What's going on?"

[0114] Customer Service B: "I'm sorry that my response was a little slow. The price of this product is 500 yuan."

[0115] User A: “Okay, thanks.”

[0116] The cleaned conversation will remove some unnecessary parts, such as customer service B's polite expressions ("Thank you very much for your inquiry", "I'm sorry"), and retain key information:

[0117] User A: “Hello, I would like to know the price of this product.”

[0118] Customer Service B: “500 yuan.”

[0119] User A: “Okay, thanks.”

[0120] After cleaning the conversation data, labeling is performed to extract the key information of the conversation and assign corresponding attribute labels to each round of conversation. Attribute information may include but is not limited to:

[0121] User basic information: user's identity information (such as user ID, user preferences, etc.), if any.

[0122] Whether a deal is concluded: whether there is an intention to close a deal or an actual deal is concluded in the conversation, such as whether the buyer clearly indicates that the product will be purchased.

[0123] Conversation sentiment tendency: Analyze the user's sentiment tendency in the conversation, whether it is positive, negative or neutral.

[0124] Key question types: Analyze the main categories of questions users ask in conversations, such as product price, features, after-sales service, etc.

[0125] Suppose the following is a conversation:

[0126] User A: “How many colors does this phone come in?”

[0127] Customer Service B: “It comes in three colors: black, white, and blue.”

[0128] The labeled information is as follows:

[0129] User basic information: User A (identified by the system as a first-time buyer);

[0130] Whether the order is concluded: It is not clear whether the order is concluded;

[0131] Emotional tendency of the conversation: neutral (no obvious emotional tendency);

[0132] Key question type: Product color.

[0133] After labeling, the conversation data is converted into a structured format, usually a standard data format (such as JSON or a database table). The structured form makes it easier to further analyze and mine. The core of this process is to extract key information and standardize it so that the data can be quickly used in subsequent analysis, machine learning, or recommendation systems.

[0134] Convert the above conversations and labeled information into structured data:

[0135] {

[0136] "user_id":"A",

[0137] "dialog_rounds":[

[0138] {

[0139] "question":"How many colors does this phone come in?",

[0140] "response":"It comes in three colors: black, white and blue.",

[0141] "emotion":"neutral",

[0142] "key_issue":"Product color",

[0143] "transaction":"no"

[0144] } ]

[0146] }

[0147] The beneficial effects of the above technical solution are as follows: through the dialogue cleaning process, redundant information and noise data are removed, making the dialogue more accurate and concise. This helps customer service personnel or intelligent customer service systems understand user needs more quickly and improve response efficiency. For users, concise and clear dialogues can also improve user experience and avoid interference from irrelevant information. Labeling processing gives specific attributes to each round of dialogue, which facilitates the automatic classification and analysis of dialogue content. Through structured data formats, personalized dialogue history archives can be established for each user to help customer service better understand user needs. By analyzing a large amount of labeled dialogue data, efficient dialogue strategies and words can be identified. These golden words are usually those expressions that can guide customers to quickly reach transactions, solve problems or improve customer satisfaction. By comparing the effects of different words, companies can develop more targeted customer service training plans and standardized words libraries. Labeled data can help companies identify users' real needs, especially some potential needs.

[0148] In another embodiment, generating semantic representation data of a conversation includes:

[0149] Get multiple encoding strategies corresponding to the conversation data;

[0150] Traverse the encoding strategies in sequence, and obtain the feature-semantic representation library corresponding to the traversed encoding strategy each time;

[0151] Splitting the conversation data into multiple conversation data items;

[0152] Perform feature extraction on the conversation data items to obtain multiple features;

[0153] Based on the feature-semantic representation library, determine the semantic representation corresponding to the feature and associate it with the corresponding conversation data item;

[0154] Accumulating and calculating the semantic representations associated with the second conversation data item to obtain a semantic representation sum;

[0155] Obtaining the semantic representation and threshold corresponding to the traversed first encoding strategy, and if the semantic representation and threshold are greater than or equal to the semantic representation and threshold, taking the corresponding second conversation data item as the third conversation data item;

[0156] Obtaining an adaptation semantic model corresponding to the traversed first coding strategy, performing adaptation semantic representation on the traversed first coding strategy based on the adaptation semantic model and according to the dialogue data, obtaining an adaptation value, and associating the adaptation value with the traversed coding strategy;

[0157] When the traversal of the coding strategies is finished, the adaptation values ​​associated with the coding strategies are accumulated and calculated to obtain the adaptation value sum;

[0158] The maximum adaptation value and the corresponding coding strategy are used as the final coding strategy;

[0159] Based on the final coding strategy, the conversation data is encoded to obtain the coding result.

[0160] Among them, for each dialogue data item D, extract the feature set F:

[0161] F(D)={f1(D),f2(D),…,f n (D)}

[0162] Use the feature-semantic representation library S to map features to semantic representations:

[0163] S(F)={s1,s2,…,s n}

[0164] s i Represented as a semantic representation;

[0165] For the second dialogue data item D2, its semantic representation is accumulated:

[0166]

[0167] ω i represents the weight of each semantic representation, based on its importance or confidence.

[0168] The adaptation semantic model M adapts the semantic representation and calculates the adaptation value A:

[0169]

[0170] C represents the current encoding strategy, α i , β i , γ i , δ i Represents the parameters of the model, which need to be optimized through training. Tanh is the hyperbolic tangent function, which is used to introduce nonlinearity.

[0171] The working principle of the above technical solution is as follows: the encoding strategy is a processing scheme for conversation data, which aims to extract key information from the data and generate semantic representations for subsequent analysis. Assume that the conversation data is a customer service conversation, in which the customer asks: "How do I apply for a refund?" The customer service replies: "You can apply for a refund through our official website, and the refund will be processed within 7 working days." The encoding strategies corresponding to this conversation data include "keyword extraction strategy" (extracting keywords such as "refund") and "sentiment analysis strategy" (analyzing the customer's emotional needs, such as whether they are anxious).

[0172] The feature-semantic representation library is a database built based on different encoding strategies to map conversation features to their semantic representations. Each feature (such as keywords, emotions, intentions, etc.) corresponds to a semantic expression. For the "keyword extraction strategy", the feature-semantic representation library maps "refund" to the semantic representation of "refund request". For the "sentiment analysis strategy", it maps "anxiety" to the semantic representation of "high priority customer".

[0173] A dialog data item is the basic unit in a conversation, usually a question or a sentence. Split the above conversation into two dialog data items:

[0174] The first data item: "How to apply for a refund?" The second data item: "You can apply for a refund through our official website, and the refund will be processed within 7 working days."

[0175] Extract semantic features from each conversation data item, such as keywords, emotions, intentions, etc. For the first data item "How to apply for a refund?", extract the features: "refund", "apply". For the second data item "You can apply for a refund through our official website, and the refund will be processed within 7 working days." Extract the features: "official website", "apply for a refund", "processed".

[0176] Map the extracted features into semantic representations so that the machine can understand them. The feature "refund" is mapped to the semantic representation "refund request". The feature "official website" is mapped to "application path".

[0177] The semantic representation of each conversation data item is accumulated to form an overall semantic representation. The semantic representation of the first data item "How to apply for a refund?" is "Refund request". The semantic representation of the second data item "You can apply for a refund through our official website, and the refund will be processed within 7 working days." is "Application method" + "Refund processing time". The semantic representation sum is the sum of these two representations to get an overall semantic representation.

[0178] Check whether the calculated semantic representation and reaches the predetermined threshold, and decide whether to use the second conversation data item as the third data item. Set the threshold as the semantic representation and of the combination of "refund request" and "application path". If the combined semantic representation reaches the threshold, the second data item can be included in the analysis as the third data item.

[0179] The adaptive semantic model is a model used to further optimize and adjust the semantic representation according to different encoding strategies. For example, the adaptive semantic model can be used to adjust the semantic representation through sentiment analysis to make it more suitable for customers with high sentiment needs in customer service scenarios, thereby improving the accuracy of analysis.

[0180] According to the semantic representation adjusted by the adaptation model, the adaptation values ​​of different encoding strategies are calculated, and the encoding strategy with the largest adaptation value is finally selected. If a certain encoding strategy best meets the emotional requirements of the dialogue data and has a high adaptation value during the gold medal speech mining process, then this strategy is selected.

[0181] The final selected coding strategy is used for the final coding processing of the dialogue data to generate the final output results. If the final coding strategy of the golden speech mining is the "emotion priority" strategy, the dialogue will be optimized according to the results of the sentiment analysis to ensure that the needs of high-priority customers are responded to quickly.

[0182] The beneficial effect of the above technical solution is that by decomposing the conversation data into multiple data items and performing feature extraction and semantic mapping on them, more accurate semantic understanding can be achieved. This not only enables the system to understand the content of the conversation, but also captures the deep user needs. Through the combination and adaptation of multiple coding strategies, golden words can be extracted from different dimensions (such as emotions, keywords, and user intentions). These words can effectively improve customer satisfaction and service efficiency. Through the "emotion priority" strategy, customer service can give priority to responding to anxious customers and provide more humane services. The adaptive ability of coding strategies and adaptive semantic models enables the system to continuously optimize as the conversation data changes. Through automated coding strategies and semantic analysis, not only the need for manual intervention is reduced, but also efficient real-time feedback and recommendations can be achieved. This is particularly important for large-scale online customer service systems, which can significantly improve processing speed and accuracy. Through detailed semantic representation and accumulation of adaptation values, the system can provide more valuable support for subsequent data analysis and decision-making. Merchants can optimize marketing decisions based on different golden words strategies to improve conversion rates and customer retention rates.

[0183] In another embodiment, generating a semantic encoding vector of a conversation includes:

[0184] Acquire a preset semantic representation data set, the semantic representation data set including: a plurality of semantic nodes;

[0185] Get the semantic feature value corresponding to the semantic node;

[0186] If the semantic feature value meets the preset semantic feature value threshold, the corresponding semantic node is used as the target node;

[0187] Acquire at least one semantic feature item corresponding to the conversation content through the target node;

[0188] Integrate various semantic feature items to generate the semantic encoding vector of the conversation content and complete semantic parsing.

[0189] The working principle of the above technical solution is as follows: Semantic node: In a semantic network, a semantic node represents the core theme or concept of a conversation content. For example, for a customer service scenario, "price inquiry", "product quality", "after-sales service", etc. can be semantic nodes. Semantic feature value: It is a value used to measure whether a semantic node is related to the current conversation content. For example, if the customer mentions "price discount" in the conversation, the feature value of the semantic node "price inquiry" is 0.85, while the feature value related to "after-sales service" is only 0.2. Semantic feature value threshold: A preset value used to determine whether a semantic node is important enough. For example, the threshold is set to 0.7, and only semantic nodes with feature values ​​higher than 0.7 can be selected as target nodes. Target node: It is a node that is closely related to the current conversation content and is selected from multiple semantic nodes. For example, if a customer asks "Is there any discount for this product?", "price inquiry" is selected as the target node. Semantic feature item: It is the specific detailed information further extracted under the target node. For example, "discount", "discount", "price reduction", etc. are semantic feature items of the target node "price inquiry". Semantic encoding vector: All semantic feature items are integrated into a mathematical representation for further analysis and recommendation by machine learning models. Vectors such as [0.9, 0.75, 0.85] represent the weights of "discount", "price reduction", and "offer".

[0190] Scenario: Customer service conversation analysis and golden words recommendation;

[0191] Customer input: "Is there any discount for this product? How is the after-sales service?"

[0192] The system parses the semantics and extracts semantic nodes: "price inquiry", "after-sales service", and "product information".

[0193] The system calculates the semantic feature value: the feature value of "price inquiry" is 0.85; the feature value of "after-sales service" is 0.75; the feature value of "product information" is 0.5. According to the preset threshold (0.7), the target nodes are filtered out: "price inquiry" and "after-sales service". Extract semantic feature items: For the target node "price inquiry", the extracted semantic feature items include: "discount" and "discount". For the target node "after-sales service", the extracted semantic feature items include: "warranty" and "repair support".

[0194] Integration to generate semantic encoding vector: The feature items and their correlation values ​​are integrated into a semantic encoding vector. For example: "Price inquiry": [0.9 (discount), 0.8 (discount)]; "After-sales service": [0.85 (warranty), 0.75 (repair support)]. The final generated conversation semantic encoding vector is: [0.9, 0.8, 0.85, 0.75].

[0195] Recommended golden words: According to the semantic coding vector, match the words database. Recommend corresponding words: "This product is currently 20% off, I can tell you the details." "In addition, our products have a one-year warranty, and the maintenance service is also very complete."

[0196] The beneficial effects of the above technical solution are: through the hierarchical analysis of semantic nodes, semantic feature values, and semantic feature value thresholds, the system can accurately identify the core needs of user conversations. The generation and matching of semantic coding vectors can quickly lock in the appropriate words and improve customer communication efficiency. The system automatically recommends gold medal words based on the results of semantic analysis, which can quickly and accurately respond to customer questions and improve customer experience. By adding semantic nodes and feature items, the system can adapt to different scenarios, such as e-commerce customer service, insurance sales, etc.

[0197] In another embodiment, cluster analysis is performed on different conversations, including:

[0198] Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content.

[0199] Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded;

[0200] Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators.

[0201] Mark the conversation content that is positively correlated with business indicators as positively correlated words, and positively correlated words are gold medal words.

[0202] The generated semantic coding vector is clustered and analyzed, including:

[0203] Obtain a preset conversation content library, and generate a corresponding semantic encoding vector based on the conversation content;

[0204] Preprocessing the semantic coding vector to obtain a processed first semantic vector set;

[0205] Based on a preset similarity measure, similarity calculation is performed on the first semantic vector set to determine a similarity relationship between the semantic vectors;

[0206] Based on a preset cluster analysis model, the semantic vectors in the first semantic vector set are clustered to obtain a plurality of initial conversation clusters;

[0207] Obtain the cluster center point vector corresponding to the initial conversation cluster, perform cluster consistency analysis on the cluster center point vector, and if the consistency meets the preset threshold, use the initial conversation cluster as the final conversation cluster; otherwise, re-cluster the initial conversation clusters that do not meet the consistency, and obtain the re-clustered conversation clusters until all clusters meet the consistency requirements, and obtain multiple final conversation clusters;

[0208] Obtain the representative conversation content of each cluster in the final conversation cluster, and determine the category label of each conversation cluster based on the representative conversation content to facilitate semantic classification and retrieval of the conversation content;

[0209] Extract content features from the final conversation clusters, obtain multiple features corresponding to each conversation cluster, and build a feature library;

[0210] By calculating the matching degree between each conversation cluster feature and the target feature in the preset feature library, if the match is met, the corresponding conversation cluster is marked as an associated conversation cluster, and at least one inspection item corresponding to the matching feature is obtained, and the inspection item includes an inspection standard, an inspection strategy and a similarity range;

[0211] Based on the verification strategy, the associated dialog cluster is verified, and if the verification passes, the first similarity score corresponding to the verification strategy is obtained; if the verification fails, the error value corresponding to the verification strategy is obtained;

[0212] The first similarity score and error value of each dialog cluster are cumulatively calculated to obtain the final clustering accuracy score, thus completing the quality evaluation of the clustering analysis.

[0213] The working principle of the above technical solution is: obtain the conversation content library and generate the semantic encoding vector

[0214] The system first obtains a large amount of conversation data, which can come from customer service, sales and other scenarios. Each conversation is converted into a "semantic encoding vector" through natural language processing technology (such as BERT, GPT and other models), which is an information representation that can capture the semantics of the conversation.

[0215] Dialogue content: "Hello, thank you for your purchase. Do you need help finding other products?"

[0216] After being converted into a semantic encoding vector, the vector represents the semantic features of the conversation, regardless of the specific vocabulary. The system can understand that it is a helpful, service-oriented conversation.

[0217] The generated semantic encoding vectors are preprocessed, such as normalization and denoising, to ensure that each vector can accurately reflect the core content of the conversation.

[0218] These processed semantic encoding vectors are clustered using similarity-based clustering algorithms (such as K-means, DBSCAN, etc.), and each cluster represents a class of similar conversation content.

[0219] Suppose the conversation content is divided into three categories: Category 1: "Customer service helps answer product usage questions"; Category 2: "Customers ask about product prices"; Category 3: "Customers express interest in purchasing." The clustering result will assign each conversation content to the corresponding cluster to help identify different types of conversations.

[0220] Get the center point vector of each cluster and judge the quality of each cluster through cluster consistency analysis. If the internal consistency of a cluster does not meet the preset standard, it will be re-clustered until the quality of all clusters meets the standard.

[0221] The system identifies which conversation clusters are positively correlated with business results by correlating business data such as customer conversion rate and customer satisfaction. For example, some service-oriented conversation clusters have a significant impact on improving customer satisfaction.

[0222] Based on the correlation analysis of business indicators, the system selects the content in the conversation cluster that is positively correlated with the business indicators and marks it as "gold medal words". These gold medal words are usually those that are effective in promoting sales, increasing customer loyalty or improving customer satisfaction. For example: the conversation content: "Thank you for your patience, I will provide you with a more favorable plan." This wording shows a positive effect in improving customer conversion rate, so it is marked as a gold medal wording.

[0223] Content features are extracted for each conversation cluster to obtain multiple features of the cluster (such as tone, care, product recommendations, etc.), and matched with the preset feature library. Conversation clusters with high matching degrees are marked as associated conversation clusters. Associated conversation clusters are verified according to preset verification strategies. Verification strategies include evaluation criteria, similarity ranges, etc., such as the relationship between semantic similarity and customer satisfaction. After the verification is passed, the system will give a "similarity score" and sum up these values ​​to calculate the final accuracy score of the cluster.

[0224] The beneficial effects of the above technical solution are: through automated cluster analysis and golden words recognition, the customer service team or sales team can quickly identify which words have a significant positive effect on improving customer conversion rate, customer satisfaction, etc. These golden words can be quickly shared with team members to improve the overall word quality and work efficiency. By combining business indicators (such as customer conversion rate, customer satisfaction, etc.) with the content of the conversation, the system can accurately identify which conversation content is related to business goals. In this way, enterprises can optimize words and service strategies according to actual business needs. With the continuous updating of conversation content and business data, the system can automatically optimize the word library according to new conversations and business indicators, and continue to dig out more efficient words to help enterprises maintain competitiveness. The system provides decision support based on data and algorithm analysis to help management understand which service words or sales words perform best, and then formulate optimized training and word guidance strategies.

[0225] In another embodiment, the forward related speech is matched with a predefined standard speech library through a matching layer, including:

[0226] Obtain a predefined standard script library, which includes: multiple script samples;

[0227] Obtain positive related words through the matching layer;

[0228] Calculate the similarity between the positive related speech and each speech sample in the standard speech library;

[0229] Calculate the similarity value between each speech sample and the semantic vector of the positively related speech;

[0230] Use the Softmax layer to convert the similarity value into a probability value to obtain the matching probability value of each speech sample;

[0231] Select the speech sample with the highest matching probability value as the recommended golden speech;

[0232] Complete the selection of recommended golden words.

[0233] The working principle of the above technical solution is as follows: the standard speech library is a library composed of multiple speech samples, which contains standard response dialogues for different scenarios. Each speech sample is a carefully designed sentence or a group of sentences designed to respond to specific user needs or questions. For example, a speech sample in the standard speech library is: "Hello, I am a customer service representative of XX Company. How can I help you?" "Thank you for your feedback. We will handle your problem as soon as possible." These speech samples have clear purposes and are very common in customer service systems, voice assistants and other scenarios.

[0234] When a user has a conversation with the system, the system will analyze the user's input and filter out a set of "positively related words" based on the input content (for example, asking questions, making demands, etc.). These words are highly semantically matched with the user input. For example, if the user enters: "I want to know the functions of the product", the system can match standard words related to the product introduction, such as: "Our product functions include XX, YY, ZZ, etc. You can choose to use them as needed." "Regarding the product functions, we provide a detailed user manual for your reference."

[0235] After obtaining the positively related speech, the system will calculate the semantic similarity of these speech with all speech samples in the standard speech library. This step is usually achieved by comparing the "semantic vector" of the dialogue sentence. Each speech sample is converted into a vector to represent its semantic content, and the positively related speech is also converted into a vector. Suppose the user inputs: "I want to know the function of the product", the system converts it into a semantic vector (through natural language processing technology, such as BERT, GPT, etc.). Then, the system calculates the similarity between the vector and all speech samples in the library, and selects the sample that is closest to the semantics of the user's statement.

[0236] In order to more accurately recommend the most appropriate speech, the system uses the Softmax layer to convert the calculated similarity values ​​into probability values. Softmax is a function commonly used in multi-classification problems. It normalizes multiple similarity values ​​into a probability distribution so that the matching probability value of each speech sample is between 0 and 1. Suppose the similarities of the three speech samples obtained by similarity calculation are 0.7, 0.9, and 0.85 respectively. The Softmax layer converts these similarity values ​​into probabilities, for example: the matching probability of speech 1: 0.25, the matching probability of speech 2: 0.50, and the matching probability of speech 3: 0.25. In the end, the system will select the speech sample with the highest probability (speech 2 in this example) as the recommended gold medal speech.

[0237] After being processed by the Softmax layer, the system selects the words with the highest matching probability as the recommended "golden words". This word is the standard word sample that best meets the current user needs and can usually provide the best response in a specific scenario. If the user asks "how to apply for a refund", and the system calculates and finds that a certain word sample (such as: "For refund applications, you need to provide the order number and the reason for the refund, and we will process it within 24 hours.") has the highest matching probability, then this word will be recommended as the "golden word".

[0238] The beneficial effects of the above technical solution are: through semantic matching and similarity calculation, the system can accurately understand the needs of users and recommend the most relevant standard words. It avoids users from encountering irrelevant or unclear answers during the communication process, thereby improving the service quality. By recommending gold medal words, the system can give the most accurate answers required by users in the shortest time, reducing user waiting time and improving the fluency of interaction. Especially in customer service or online service scenarios, users can quickly get satisfactory replies, which improves the overall user experience. The process of automated recommendation of words reduces the burden of manual customer service, especially in high-concurrency scenarios, the system can intelligently respond to a large number of user requests without manual intervention. This not only improves efficiency, but also allows manual customer service to focus more on handling complex problems. By continuously accumulating user feedback and optimizing the word library, the system can continuously optimize the word recommendation strategy. As the amount of data increases, the accuracy and intelligence of the word library will gradually increase.

[0239] In another embodiment, the recommended golden words analyzed and mined are displayed on the user end, including:

[0240] Display recommended golden words on the user side and provide dynamic update function so that the recommended words can be adjusted and optimized according to real-time analysis data;

[0241] The recommended golden words are displayed in grades, and the efficient words and optimization suggestions are marked separately;

[0242] Based on user feedback or business personnel's choice, adjust the display priority and display method of recommended golden words to ensure the best display effect;

[0243] The recommended golden scripts will be presented to the user in a form that is easy to understand and apply, including the type of scripts, usage scenarios, and application effect evaluation information.

[0244] The working principle of the above technical solution is as follows: the recommended golden words will be adjusted and optimized according to real-time data analysis. By collecting and analyzing the user's behavior data, feedback information and the salesperson's operating habits during use, the system can dynamically update the recommended golden words. For example, if a certain word has received good user feedback (such as improving the transaction rate or user satisfaction) after multiple uses, the system will automatically mark it as an efficient word and increase its display priority. Suppose a salesperson uses the following words in a conversation with a customer: Word A: "According to your needs, we can provide a more personalized solution, which can not only save time but also improve efficiency." The system will collect customer feedback data (such as whether the user continues to communicate, transaction rate, etc.) and analyze the effect of this word. If word A performs well among most customers (high transaction rate, high customer satisfaction), then the word will be marked as "efficient word" and will be recommended to other salespeople first.

[0245] When a user / salesperson uses a certain line of speech, the system will collect customer feedback (such as whether the purchase is completed, whether the conversation is over, customer sentiment, etc.). The system will evaluate the effectiveness of the line of speech through data analysis and compare it with other lines of speech. If a line of speech is effective, the system will update it in real time and increase its display priority.

[0246] The recommended golden scripts will be divided into different levels according to their effects and usage scenarios. Highly effective scripts will be marked as priority recommendations, and there will be corresponding optimization suggestions. For scripts with average or less than ideal performance, the system will give optimization suggestions to help salesmen adjust the content of the script to improve the effect. Suppose in multiple sales conversations, script B: "Your current needs are the most important, and we can definitely help you solve this problem." seems to be relatively general and has average effects. The system will provide improvement suggestions for this script: Optimization suggestion: "Add more specific demand analysis, such as 'Based on the needs you mentioned last time, we have specially designed XX plan for you'". These optimization suggestions will be automatically displayed to salesmen, helping them adjust the content of the script to make it more effective in future conversations.

[0247] After a salesperson uses a certain line of speech, the system will collect feedback and evaluate the effect. If the effect of the line of speech is not ideal, the system will provide optimization suggestions based on data analysis and assign grades according to certain rules. Efficient lines of speech are recommended first, and inefficient lines of speech are displayed with optimization suggestions.

[0248] The system will automatically adjust the display priority and display method of the scripts based on the user's operational feedback or the salesperson's choice. For example, one salesperson prefers to use personalized recommended scripts, while another salesperson prefers to use standardized, universal scripts. The system will adjust the order of the script recommendations based on their choices and preferences. Suppose salesperson C prefers to use personalized gold-medal scripts, while salesperson D prefers universal scripts. The system will prioritize the types that suit them in their interfaces based on their preferences:

[0249] For salesperson C, the system recommends "personalized sales talk based on customer characteristics";

[0250] For salesperson D, the system recommends "standardized, conventional speech".

[0251] Salespeople can choose to display specific categories of recommended scripts (such as personalized, standardized). The system adjusts the priority of script display based on the selection to ensure that each salesperson sees the content that best suits them.

[0252] On the user side, the recommended golden scripts need to be presented in a concise and easy-to-understand form so that salespeople can quickly understand and apply them flexibly. Each script will be accompanied by a description of the usage scenario, expected results, and application evaluation information. Suppose there is a recommended script: "You can choose our product because it will help you save a lot of time and improve work efficiency." Description of usage scenario: Applicable to situations where customers are concerned about efficiency and time management. Expected effect: Improve customer recognition of product value and increase transaction opportunities. Application effect evaluation: Based on the transaction data after using the script, the system will display its contribution to improving the transaction rate (for example: after using the script, the transaction rate increased by 15%).

[0253] Each recommended script will be accompanied by clear scenario description, expected results and evaluation data to ensure that sales staff can understand and apply it effectively.

[0254] The beneficial effects of the above technical solution are: by updating the words and optimizing suggestions in real time, the salesperson can quickly adopt the most effective communication strategy and improve the success rate of communicating with customers. The system adjusts the words recommendation based on actual feedback, so that the recommendation is more in line with the needs of the salesperson and the needs of the customer. The system can help the salesperson choose the best words for the current sales situation through words analysis and recommendation, thereby improving the transaction rate. The system can automatically adjust the display priority according to the operating habits and customer feedback of different salespeople, and personally recommend the communication words that are most suitable for them. This makes the words more adaptable and can effectively help different salespeople deal with different customers and sales situations.

[0255] In another embodiment, an analysis and recommendation device for conversation analysis and gold medal speech mining includes:

[0256] A dialogue extraction unit, used to pre-process the dialogue content provided by the user and extract structured dialogue data;

[0257] A speech coding unit is used to encode the dialogue data and generate semantic representation data of the dialogue;

[0258] A semantic encoding unit, which is used to perform semantic analysis on the semantic representation data and generate a semantic encoding vector of the dialogue;

[0259] The speech mining unit is used to perform cluster analysis on different conversations based on the semantic coding vectors of the conversations, and identify positively related speech based on business indicators;

[0260] The golden words recommendation unit is used to match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words;

[0261] The speech display unit is used to display the recommended golden speech that has been analyzed and mined on the user side.

[0262] The working principle of the above technical solution is: the role of the dialogue extraction unit is to extract structured data from the original dialogue for subsequent analysis and processing. Assume that the dialogue between the user and the customer service is as follows:

[0263] User: I would like to know how much your mobile phones cost?

[0264] Customer Service: Hello, we have many types of mobile phones with prices ranging from 2,000 yuan to 5,000 yuan. Which type of mobile phone do you need?

[0265] At this point, the dialogue extraction unit recognizes the following structured information:

[0266] User intention: to inquire about the price of the mobile phone; Customer service response: provided mobile phone options with a variety of price ranges; Key information: mobile phone, price, 2,000 yuan to 5,000 yuan.

[0267] By extracting this information, the system can convert conversations into structured data that is easy to analyze and process.

[0268] The role of the speech encoding unit is to convert structured data into semantic representation data, that is, to generate a digital representation for the conversation content. Here, we can use natural language processing technology (such as Word2Vec, BERT, etc.) to achieve this. Suppose the system encodes the two keywords "mobile phone" and "price" into vector representations: "mobile phone" will be encoded as [0.12, 0.47, -0.33, ...]; "price" will be encoded as [0.09, 0.15, 0.32, ...]. These digital vector representations can help the system better understand the semantics of the conversation and lay the foundation for subsequent semantic analysis.

[0269] The task of the semantic encoding unit is to further parse the dialogue and generate a more abstract semantic vector for semantic matching and clustering analysis. The system converts each dialogue into a higher-level vector representation based on the context and semantic relationship of the dialogue. For example, the system generates a vector [0.25, 0.63, -0.10, ...] containing semantic information by analyzing the relationship between "mobile phone price" and "2,000 to 5,000 yuan". This vector is no longer just a representation of a single word, but combines the semantic structure and contextual information of the entire dialogue.

[0270] The speech mining unit clusters different conversations based on semantic coding vectors, identifies conversation groups with similar semantics, and selects the most effective and positively related speech based on business indicators (such as conversion rate, customer satisfaction, etc.). For example, the system finds that conversations about "asking about the price of a mobile phone" are usually clustered in a specific semantic area, and based on historical data, it finds that such conversations can often be successfully converted into purchases. The system clusters the speech with other similar conversations, finds the best speech pattern, and marks it as a positively related speech.

[0271] The role of the gold medal speech recommendation unit is to compare the positive related speech with the speech in the standard speech library through the matching layer based on the results of cluster analysis, and select the "gold medal speech" that best matches the semantics of the user's current conversation for recommendation. Suppose there is a speech in the standard speech library: "According to your needs, we have suitable mobile phones with prices between 2,000 yuan and 5,000 yuan. You can choose according to your budget." If the system recognizes that the user is asking about the price of the mobile phone and the range is between 2,000 yuan and 5,000 yuan, then this speech will be matched as a gold medal speech and recommended to the customer service.

[0272] The role of the script display unit is to display the recommended golden script to the customer service staff or the automatic customer service system for their direct use. In this step, the system not only provides the script text, but also provides some suggestions and contextual information to help customer service staff interact with customers more efficiently. For example, the customer service interface will display: "Recommended golden script: 'According to your needs, we have suitable mobile phones with prices ranging from 2,000 yuan to 5,000 yuan. You can choose according to your budget.'" At the same time, the system will also display reference data such as the success rate and customer satisfaction of the script to help customer service make the most appropriate response.

[0273] The beneficial effects of the above technical solution are: by automating the traditional manual speech recommendation system, customer service personnel can quickly obtain optimized "golden speech" that meets user needs, thereby reducing the time of repeated thinking and querying standard speech, and improving work efficiency. By analyzing user intentions and historical data, the system can recommend the most relevant and most conversion-rate speech to help customer service better understand and solve user problems. Better service quality helps to improve customer satisfaction and loyalty. The speech mining unit can continuously mine new and effective speech through cluster analysis and semantic coding, thereby optimizing the speech library. As the amount of data increases, the system can continue to learn, and the recommended "golden speech" will become more accurate and personalized. Through intelligent semantic analysis, the system can ensure that the recommended speech is highly consistent with user needs, avoiding errors and inaccuracies in the manual speech selection process.

[0274] In another embodiment, the dialogue extraction unit includes:

[0275] The first dialogue extraction module is used to clean the dialogue content provided by the user, remove redundant information and noise data, and obtain the pre-processed dialogue content, where the dialogue content includes customer service, business dialogue and multiple rounds of dialogue in sales;

[0276] The second dialogue extraction module is used to label the preprocessed dialogue content. According to the characteristics of the dialogue content, the attribute information of each round of dialogue is marked. The attribute information includes the basic information of the user mentioned in the dialogue, whether the order is completed, the emotional tendency of the dialogue, and the type of key questions.

[0277] The third module of conversation extraction is used to extract structured conversation data from the labeled conversation data and convert the conversation data into a unified data format.

[0278] The working principle of the above technical solution is as follows: In the first module, the goal is to clean the original conversation data and remove redundant information and noise data so that subsequent processing is more accurate. Conversation cleaning mainly screens and filters out valuable information based on some predefined rules. The processing process can be completed through the following steps:

[0279] Remove irrelevant noise information: such as advertisements, system notifications, meaningless catchphrases, etc., which are usually not of substantial help to the analysis. If a customer says, "I don't have time now, let me watch the advertisement first", the part "watch the advertisement first" can be removed after cleaning.

[0280] Remove duplicate content: In multiple rounds of conversations, customers and customer service may repeat similar questions or statements, such as "How is your product?", "Is your product good?". After cleaning, the repeated parts will be merged or filtered. The cleaned content will become: "The customer asked about the product quality."

[0281] Merge the core questions in long conversations: In multiple rounds of conversations, users repeatedly express the same question or demand. By extracting key information, the conversation content can be simplified. For example, if a user repeatedly asks "What is the battery capacity of this phone?", it can be merged into "The user asks about the battery capacity of the phone."

[0282] Original conversation: User: "I saw your ad, what's the price?" Customer service: "Hello, thank you for your attention. Our prices vary according to the model. Which one do you need?" User: "The ad seems to say it's the 3,000 yuan model?" Customer service: "Yes, the 3,000 yuan model is the standard version. Do you have any other questions?"

[0283] Cleaned conversation: User asks about price, customer service responds with price-related information.

[0284] In the second module, the goal is to label each round of dialogue based on the characteristics of the dialogue content, so as to facilitate subsequent data analysis. Labeled attributes may include: Basic user information: Mark the user's identity information, demand areas, etc., such as whether the user has registered, whether the user's demand is for a certain function of the product or the price, etc. If the user mentions: "I registered on your website", the system can mark "the user has registered". Whether the order is completed: Mark whether a sales or service agreement has been reached in the dialogue, that is, whether the order is successfully signed. If the customer confirms to purchase the product at the end of the dialogue, it can be marked as "the order has been completed". Emotional tendency: Analyze the user's emotional tendency (such as positive, negative, neutral, etc.) based on the tone and expression of the dialogue content. If the user says: "The price is too high, I don't plan to buy it", it can be marked as "negative emotion". Key question type: According to the theme of the dialogue content, mark the question type in the dialogue, such as "product consultation", "technical support", "after-sales service", "payment problem", etc. For example, "Does this phone support wireless charging?" can be marked as "product consultation".

[0285] For example: Original conversation: User: "How is your phone?"

[0286] Customer service: "Hello, this phone has powerful functions and excellent performance, and is loved by customers."

[0287] User: "What about the battery? I'm more concerned about that."

[0288] Customer service: "The battery capacity is 5000 mAh, and the battery life is very strong."

[0289] User: "I want to place an order, how do I pay?"

[0290] Customer Service: "You can choose to pay by Alipay, WeChat Pay or credit card."

[0291] Labeled conversations: user asks about product features (key question type), user is concerned about batteries (key question type), user decides to buy (whether the order is completed).

[0292] In the third module, the goal is to convert the labeled conversation content into a structured data format. Structured data facilitates subsequent analysis, statistics, and model training. At this stage, the data will be stored in a standardized format, common formats include JSON, CSV, etc. Data is converted into a standard format: each round of conversation and its labels (such as sentiment, question type, etc.) are stored in a structured field. For example, each conversation round corresponds to a record, including a timestamp, what the user said, the customer service response, and related labels (such as sentiment, question type, whether the order was completed, etc.).

[0293] The beneficial effects of the above technical solution are: through structured data, valuable information can be quickly extracted, such as which questions are most frequently asked, which words are most effective, and which users with emotional tendencies are more likely to be converted into customers. This provides data support for subsequent optimization. By analyzing the effects of different words, the team can refine the golden words. Through label analysis of customer emotional tendencies, personalized dialogue strategies can be customized. Structured dialogue data can provide decision support for management, help analyze each link of the sales funnel, optimize customer service processes, and improve conversion rates.

[0294] In another embodiment, cluster analysis is performed on different conversations, including:

[0295] Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content.

[0296] Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded;

[0297] Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators.

[0298] Mark the conversation content that is positively correlated with business indicators as positively correlated words, and positively correlated words are gold medal words.

[0299] The working principle of the above technical solution is as follows: First, a large amount of conversation data is collected. For example, the conversation records of an online customer service platform. Each conversation includes different customer questions and customer service answers.

[0300] Each conversation content is converted into a high-dimensional semantic encoding vector through natural language processing technology (such as BERT, GPT and other models). This vector represents the semantic information of the conversation, so that the similarity between different conversations can be calculated by mathematical methods. For example: Customer Conversation 1: "I want to know the price of your products." Customer Conversation 2: "What price discounts do you have for your products?" Customer service answers: "We have a variety of discounts, you can choose according to your needs." After converting these conversations into semantic vectors, it can be seen that the semantics of "Customer Conversation 1" and "Customer Conversation 2" are very similar and belong to the same type of conversation content.

[0301] Clustering algorithms (such as K-means, hierarchical clustering, etc.) are used to cluster the semantic vectors of all conversations to obtain multiple "conversation clusters". Each conversation cluster represents a class of conversation content with similar semantics or intentions. For example, all conversations asking about product prices may be clustered in the same cluster, while conversations asking about after-sales service are in another cluster. Conversation Cluster A: Questions about prices, discounts, etc. Conversation Cluster B: Questions about after-sales, return and exchange policies, etc. Conversation Cluster C: Technical questions about product functions and performance.

[0302] Next, we need to associate business indicators to evaluate which conversation clusters can bring better business results. These business indicators include: Customer conversion rate: for example, the ratio from conversation to actual purchase. Customer satisfaction: for example, satisfaction based on customer service ratings or questionnaires. Whether the order is completed: for example, whether the customer finally completes the purchase. Suppose we have these business data and associate them with each conversation cluster. Through association analysis, we can identify which conversation clusters have conversation content that shows a trend of positive correlation with business indicators.

[0303] By analyzing the relationship between customer conversion rate, customer satisfaction, etc. and conversation clusters, we can identify which conversation clusters have a positive impact on business results. For example, we found that all conversation clusters asking about prices and discounts are strongly correlated with increased customer conversion rates. Through analysis, we found that the conversation cluster containing "asking about prices" (conversation cluster A) is usually strongly correlated with customer purchase intentions, while the conversation clusters containing "asking about after-sales" or "technical support" (conversation clusters B and C) rarely lead to transactions. Finally, according to the content in the conversation cluster that is positively correlated with business indicators (such as customer conversion rate, satisfaction, etc.), it is marked as "gold medal speech".

[0304] In the conversation cluster, select those conversation contents with high conversion rate or customer satisfaction and mark them as "gold medal words". These words can be customer service answers, strategies to guide customers to buy, skills to deal with objections, etc. In conversation cluster A, if some customer service answers help customers quickly understand the discount and motivate purchases, such as "Buy now and there is a 50% discount, it is recommended that you seize the opportunity", it is considered a gold medal word. In conversation cluster C, if the customer service can effectively answer the customer's technical questions, make the customer trust the product and promote the purchase, the answer can also be marked as a gold medal word.

[0305] Recommend these golden words to the customer service team as training materials or templates for automatic replies to help customer service improve their effectiveness in responding to customer questions. The customer service system can automatically recommend golden words: when a customer asks about prices, the system can automatically display words like "Seize the opportunity, buy today for the biggest discount", helping customer service improve customer conversion rates.

[0306] The beneficial effects of the above technical solution are as follows: Golden Medal Talk can improve customers' purchasing desire and trust by optimizing customer communication, directly promote customer conversion, and improve sales performance. Through association analysis, it is found which conversation content can effectively improve customer satisfaction. After identifying and recommending these words, customer service can respond to customers more targetedly and provide personalized services, thereby improving customer satisfaction and loyalty. The mining of golden words provides the customer service team with specific training content and templates to help them improve the effectiveness of communication with customers. This systematic training method helps to promote excellent conversation skills among all employees and improve the overall performance of the team. By automatically recommending golden words, the customer service system can reduce manual intervention, allowing customer service staff to quickly find the best way to answer common problems. This not only improves the response speed, but also allows customer service staff to focus on dealing with complex problems, thereby improving work efficiency.

[0307] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.

Claims

1. An analysis and recommendation method for conversation analysis and gold medal speech mining, characterized in that: include: S101: pre-processing the conversation content provided by the user to extract structured conversation data; S102: Encoding the dialogue data to generate semantic representation data of the dialogue; S103: performing semantic analysis on the semantic representation data to generate a semantic encoding vector of the conversation; S104: Based on the semantic coding vectors of the conversations, cluster analysis is performed on different conversations, and positively related conversations are identified in combination with business indicators; S105: Match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words; S106: The recommended golden words analyzed and mined are displayed on the user side.

2. The analysis and recommendation method for dialogue analysis and gold medal speech mining according to claim 1 is characterized in that: S101 includes: S1011: Clean the conversation content provided by the user, remove redundant information and noise data, and obtain the pre-processed conversation content, where the conversation content includes customer service, business conversations, and multiple rounds of conversations in sales; S1012: labeling the pre-processed conversation content, and marking attribute information of each round of conversation according to the characteristics of the conversation content. The attribute information includes basic information of the user, whether the order is established, emotional tendency of the conversation, and key question type; S1013: extracting structured conversation data from the labeled conversation data, and converting the conversation data into a unified data format.

3. The analysis and recommendation method for dialogue analysis and gold medal speech mining according to claim 1 is characterized in that: Generate semantic representation data of the conversation, including: Get multiple encoding strategies corresponding to the conversation data; Traverse the encoding strategies in sequence, and obtain the feature-semantic representation library corresponding to the traversed encoding strategy each time; Splitting the conversation data into multiple conversation data items; Perform feature extraction on the conversation data items to obtain multiple features; Based on the feature-semantic representation library, determine the semantic representation corresponding to the feature and associate it with the corresponding conversation data item; Accumulating and calculating the semantic representations associated with the second conversation data item to obtain a semantic representation sum; Obtaining the semantic representation and threshold corresponding to the traversed first encoding strategy, and if the semantic representation and threshold are greater than or equal to the semantic representation and threshold, taking the corresponding second conversation data item as the third conversation data item; Obtaining an adaptation semantic model corresponding to the traversed first coding strategy, performing adaptation semantic representation on the traversed first coding strategy based on the adaptation semantic model and according to the dialogue data, obtaining an adaptation value, and associating the adaptation value with the traversed coding strategy; When the traversal of the coding strategies is finished, the adaptation values ​​associated with the coding strategies are accumulated and calculated to obtain the adaptation value sum; The maximum adaptation value and the corresponding coding strategy are used as the final coding strategy; Based on the final coding strategy, the conversation data is encoded to obtain the coding result.

4. The analysis and recommendation method for dialogue analysis and gold medal speech mining according to claim 1 is characterized in that: Generate a semantic encoding vector of the conversation, including: Acquire a preset semantic representation data set, the semantic representation data set including: a plurality of semantic nodes; Get the semantic feature value corresponding to the semantic node; If the semantic feature value meets the preset semantic feature value threshold, the corresponding semantic node is used as the target node; Acquire at least one semantic feature item corresponding to the conversation content through the target node; Integrate various semantic feature items to generate the semantic encoding vector of the conversation content and complete semantic parsing.

5. The analysis and recommendation method for dialogue analysis and gold medal speech mining according to claim 1 is characterized in that: Cluster analysis of different conversations, including: Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content. Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded; Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators. Mark the conversation content that is positively correlated with business indicators as positively correlated words, and positively correlated words are gold medal words.

6. The analysis and recommendation method for dialogue analysis and golden words mining according to claim 1 is characterized in that: The matching layer matches positive related words with the predefined standard words library, including: Obtain a predefined standard script library, which includes: multiple script samples; Obtain positive related words through the matching layer; Calculate the similarity between the positive related speech and each speech sample in the standard speech library; Calculate the similarity value between each speech sample and the semantic vector of the positively related speech; Use the Softmax layer to convert the similarity value into a probability value to obtain the matching probability value of each speech sample; Select the speech sample with the highest matching probability value as the recommended golden speech; Complete the selection of recommended golden words.

7. The analysis and recommendation method for dialogue analysis and gold medal speech mining according to claim 1 is characterized in that: The recommended golden words analyzed and mined are displayed on the user side, including: Display recommended golden words on the user side and provide dynamic update function so that the recommended words can be adjusted and optimized according to real-time analysis data; The recommended golden words are displayed in grades, and the efficient words and optimization suggestions are marked separately; Based on user feedback or business personnel's choice, adjust the display priority and display method of recommended golden words to ensure the best display effect; The recommended golden scripts will be presented to the user in a form that is easy to understand and apply, including the type of scripts, usage scenarios, and application effect evaluation information.

8. An analysis and recommendation device for conversation analysis and gold medal speech mining, characterized in that: include: A dialogue extraction unit, used to pre-process the dialogue content provided by the user and extract structured dialogue data; A speech coding unit is used to encode the dialogue data and generate semantic representation data of the dialogue; A semantic encoding unit, which is used to perform semantic analysis on the semantic representation data and generate a semantic encoding vector of the dialogue; The speech mining unit is used to perform cluster analysis on different conversations based on the semantic coding vectors of the conversations, and identify positively related speech based on business indicators; The golden words recommendation unit is used to match the positive related words with the predefined standard words library through the matching layer, and select the words with the highest similarity as the recommended golden words; The speech display unit is used to display the recommended golden speech that has been analyzed and mined on the user side.

9. The analysis and recommendation device for dialogue analysis and gold medal speech mining according to claim 8 is characterized in that: The dialogue extraction unit includes: The first dialogue extraction module is used to clean the dialogue content provided by the user, remove redundant information and noise data, and obtain the pre-processed dialogue content, where the dialogue content includes customer service, business dialogue and multiple rounds of dialogue in sales; The second dialogue extraction module is used to label the preprocessed dialogue content. According to the characteristics of the dialogue content, the attribute information of each round of dialogue is marked. The attribute information includes basic user information, whether the order is completed, dialogue sentiment tendency, and key question type. The third module of conversation extraction is used to extract structured conversation data from the labeled conversation data and convert the conversation data into a unified data format.

10. The analysis and recommendation device for dialogue analysis and gold medal speech mining according to claim 8, characterized in that: Cluster analysis of different conversations, including: Perform cluster analysis on the generated semantic encoding vectors to obtain multiple conversation clusters, each of which represents a class of similar conversation content. Obtaining business indicators associated with the conversation data, where the business indicators include at least one indicator related to customer conversion rate, customer satisfaction, or whether a deal is concluded; Based on business indicators, correlation analysis is performed on each of the multiple conversation clusters to identify conversation content that is positively correlated with the business indicators. Mark the conversation content that is positively related to business indicators as positively related words, and positively related words are the gold medal words.