Consultation on call transfer methods, devices, computer equipment, and storage media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-08-14
AI Technical Summary
[0004]本申请实施例提供一种咨询通话转接方法、装置、计算机设备及存储介质,以解决因现有技术无法在通话过程中实时识别购买意向并立即转接坐席,导致错失最佳服务时机、影响通话咨询的服务转化率的技术问题
[0009]上述咨询通话转接方法、装置、计算机设备及存储介质所实现的方案中,可以应用于金融科技和医疗健康的业务咨询通话场景,首先获取目标咨询对象与智能客服之间的当前咨询通话的实时通话语音数据,完整捕获咨询全过程的语音信息。接着,对实时通话语音数据进行说话人分离,得到目标咨询对象的原始咨询文本数据,有效过滤智能客服侧语音干扰,确保分析对象的准确性。进一步地,对原始咨询文本数据进行上下文提取,得到上下文融合向量,将分散语句关联为完整语义表征,避免意图识别的片面性;根据上下文融合向量生成实时通话摘要,便于从语义层面更准确地掌握目标咨询对象的咨询要点;根据上下文融合向量进行意图识别,得到实时意图数据,能够更精准地捕捉用户真实需求。最后,基于实时通话摘要和实时意图数据将当前咨询通话转接至预设的人工客服,使得人工客服可以直接响应,显著提升转接效率与服务体验,解决了因现有技术无法在通话过程中实时识别购买意向并立即转接坐席,导致错失最佳服务时机、影响通话咨询的服务转化率的技术问题。
Smart Images

Figure CN122578770A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the fields of artificial intelligence technology and natural language processing technology, and is applicable to the fields of financial technology and medical technology. In particular, it relates to a consultation call transfer method, device, computer equipment and storage medium. Background Technology
[0002] Call transfer methods can be used to analyze customer purchase intentions during real-time calls and transfer calls with high purchase intentions to human agents. These methods can be applied in various scenarios, such as in the fintech sector for telephone sales and intelligent customer service consultations of financial products like car insurance, property insurance, and wealth management; and in the medical technology sector for telephone sales and intelligent customer service consultations of medical products like medical insurance, medical devices, and rehabilitation / nursing equipment. By analyzing call content in real time, customer purchase intentions can be identified and transferred to human agents, thereby improving service conversion efficiency.
[0003] Currently, the entire call recording is usually transcribed into text after the call ends, and then a summary is extracted and the intent is categorized. This method cannot identify the purchase intent in real time during the call and immediately transfer the call to an agent, which can easily lead to missing the best service opportunity and affect the service conversion rate of call consultations. Summary of the Invention
[0004] This application provides a consultation call transfer method, device, computer equipment, and storage medium to solve the technical problem that the existing technology cannot identify purchase intentions in real time during a call and immediately transfer the call to an agent, resulting in missed best service opportunities and affecting the service conversion rate of call consultations.
[0005] Firstly, a method for transferring consultation calls is provided, including: Acquire real-time voice data of the current consultation call; wherein, the current consultation call is a call between the target consultation object and the intelligent customer service; Speaker separation is performed on the real-time call voice data to obtain the original consultation text data of the target consultation object; Context extraction is performed on the original consultation text data to obtain a context fusion vector; A real-time call summary is generated based on the context fusion vector; Intent recognition is performed based on the context fusion vector to obtain the real-time intent data; Based on the real-time call summary and the real-time intent data, the current inquiry call is transferred to a preset human customer service representative.
[0006] Secondly, a consultation call transfer device is provided, comprising: The call voice acquisition module is used to acquire real-time call voice data of the current consultation call; wherein, the current consultation call is a call between the target consultation object and the intelligent customer service; The speaker separation module is used to separate the speaker from the real-time call voice data to obtain the original consultation text data of the target consultation object. The context extraction module is used to extract the context from the original consultation text data to obtain a context fusion vector; A call summary generation module is used to generate a real-time call summary based on the context fusion vector. The intent recognition module is used to perform intent recognition based on the context fusion vector to obtain the real-time intent data; The call transfer module is used to transfer the current inquiry call to a preset human customer service representative based on the real-time call summary and the real-time intent data.
[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described consultation call transfer method.
[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described consultation call transfer method.
[0009] The aforementioned consultation call transfer method, device, computer equipment, and storage medium can be applied to business consultation call scenarios in fintech and healthcare. First, it acquires real-time voice data of the current consultation call between the target customer and the intelligent customer service representative, completely capturing the voice information of the entire consultation process. Next, it performs speaker separation on the real-time voice data to obtain the original consultation text data of the target customer, effectively filtering out voice interference from the intelligent customer service side and ensuring the accuracy of the analysis. Further, it extracts the context from the original consultation text data to obtain a context fusion vector, linking scattered statements into a complete semantic representation, avoiding the one-sidedness of intent recognition. A real-time call summary is generated based on the context fusion vector, facilitating a more accurate understanding of the target customer's consultation points from a semantic perspective. Intent recognition is then performed based on the context fusion vector to obtain real-time intent data, enabling a more precise capture of the user's true needs. Finally, based on real-time call summaries and real-time intent data, the current consultation call is transferred to a preset human customer service representative, enabling the human customer service representative to respond directly. This significantly improves transfer efficiency and service experience, and solves the technical problem that existing technology cannot identify purchase intentions in real time during the call and immediately transfer the call to an agent, resulting in missed best service opportunities and affecting the service conversion rate of the call consultation. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a schematic diagram of an application environment for a consultation call transfer method according to one embodiment of this application; Figure 2 This is a flowchart illustrating a consultation call transfer method in one embodiment of this application; Figure 3 yes Figure 2 A schematic diagram of a specific implementation method for step S20; Figure 4 yes Figure 2 A schematic diagram of a specific implementation method for step S30; Figure 5 yes Figure 4 A flowchart illustrating a specific implementation of step S32; Figure 6 yes Figure 2 A schematic diagram of a specific implementation of step S40; Figure 7 yes Figure 6 A schematic diagram of a specific implementation method for step S42; Figure 8 yes Figure 2 A schematic diagram of a specific implementation method for step S60; Figure 9 This is a schematic diagram of a consultation call transfer device in one embodiment of this application; Figure 10 This is a schematic diagram of the structure of a computer device according to one embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.
[0014] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.
[0015] Information extraction is a text processing technique that extracts factual information such as entities, relationships, and events from natural language text and outputs it as structured data. Information extraction is a technique for extracting specific information from text data. Text data is composed of specific units, such as sentences, paragraphs, and chapters. Text information is composed of smaller, specific units, such as characters, words, phrases, sentences, paragraphs, or combinations of these units. Extracting noun phrases, names of people, and place names from text data is an example of text information extraction. Of course, text information extraction techniques can extract information of various types.
[0016] Call transfer methods can be used to analyze customer purchase intentions during real-time calls and transfer calls with high purchase intentions to human agents. These methods can be applied in various scenarios, such as in the fintech sector for telephone sales and intelligent customer service consultations of financial products like car insurance, property insurance, and wealth management; and in the medical technology sector for telephone sales and intelligent customer service consultations of medical products like medical insurance, medical devices, and rehabilitation / nursing equipment. By analyzing call content in real time, customer purchase intentions can be identified and transferred to human agents, thereby improving service conversion efficiency.
[0017] Currently, the entire call recording is usually transcribed into text after the call ends, and then a summary is extracted and the intent is categorized. This method cannot identify the purchase intent in real time during the call and immediately transfer the call to an agent, which can easily lead to missing the best service opportunity and affect the service conversion rate of call consultations.
[0018] In addition, some algorithmic models attempt to perform streaming processing (sentence-by-sentence or segment-by-segment analysis), but in the process, they mostly analyze isolated sentences. For example, they only classify intent based on the current sentence, ignoring the context established by previous dialogue history. In actual service dialogues, a customer's purchase intention is built up gradually. For example, a customer might first ask, "Is there a discount on this product?", then ask, "How do I pay?", and finally say, "I'll think about it." If we look at the sentence "I'll think about it" in isolation, the intent might be negative, but in the context of "asking about discounts" and "payment methods," it is likely to be a high-intent signal. Because some algorithmic models lack effective long-distance context modeling capabilities, the misjudgment rate of intent recognition is high.
[0019] Furthermore, existing summary generation models and intent recognition models are typically two independent modules, or even implemented by two separate algorithms. They process the same speech stream, but their outputs are independent, lacking information exchange and coordination. This can easily lead to the generated summaries failing to include crucial information essential for intent judgment, and the criteria for intent recognition failing to be highlighted in the summaries. This disconnect results in subsequent decisions (such as transferring call centers) lacking a unified and comprehensive information view to support them, leading to insufficient decision-making basis.
[0020] All of the above situations can lead to missing the best service opportunity and affect the service conversion rate of call consultations.
[0021] Based on this, embodiments of this application provide a consultation call transfer method, apparatus, computer equipment, and storage medium to solve the technical problem that the existing technology cannot identify purchase intentions in real time during a call and immediately transfer the call to an agent, resulting in missed optimal service opportunities and affecting the service conversion rate of call consultations.
[0022] The consultation call transfer method, apparatus, computer equipment, and storage medium provided in this application are specifically described through the following embodiments. First, the consultation call transfer method in this application embodiment is described.
[0023] The consultation call transfer method provided in this application embodiment can be applied to, for example, Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain real-time voice data of the current consultation call from the client; the current consultation call is a call between the target inquirer and the intelligent customer service; speaker separation is performed on the real-time voice data to obtain the original consultation text data of the target inquirer; context extraction is performed on the original consultation text data to obtain a context fusion vector; a real-time call summary is generated based on the context fusion vector; intent recognition is performed based on the context fusion vector to obtain real-time intent data; based on the real-time call summary and real-time intent data, the current consultation call is transferred to a preset human customer service representative, and the call transfer result is fed back to the client.
[0024] In this application, targeting business consultation call scenarios in fintech and healthcare, the real-time voice data of the current consultation call between the target customer and the intelligent customer service is first acquired to fully capture the voice information of the entire consultation process. Next, speaker separation is performed on the real-time voice data to obtain the original consultation text data of the target customer, effectively filtering out voice interference from the intelligent customer service side and ensuring the accuracy of the analysis. Furthermore, context extraction is performed on the original consultation text data to obtain a context fusion vector, linking scattered statements into a complete semantic representation, avoiding the one-sidedness of intent recognition. A real-time call summary is generated based on the context fusion vector, facilitating a more accurate understanding of the target customer's consultation points from a semantic perspective. Intent recognition is performed based on the context fusion vector to obtain real-time intent data, enabling more precise capture of the user's true needs. Finally, based on the real-time call summary and real-time intent data, the current consultation call is transferred to a pre-set human customer service representative, allowing for direct response and significantly improving transfer efficiency and service experience. This solves the technical problem that existing technologies cannot identify purchase intentions in real-time during the call and immediately transfer the customer to an agent, leading to missed optimal service opportunities and impacting the service conversion rate of the consultation.
[0025] The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The following detailed description uses specific embodiments to illustrate this application.
[0026] Please see Figure 2 As shown, Figure 2A flowchart illustrating the consultation call transfer method provided in this application embodiment includes the following steps: S10. Obtain the real-time voice data of the current consultation call; wherein, the current consultation call is the call between the target consultation object and the intelligent customer service; The current consultation call refers to the ongoing call, involving both the target customer and the intelligent customer service system. In fintech scenarios, the target customer could be a bank customer or someone about to purchase insurance products; in healthcare technology scenarios, the target customer could be a patient or their family member. Intelligent customer service refers to an AI system that automatically responds based on voice interaction technology, such as intelligent voice assistants in banks and hospitals.
[0027] Real-time call voice data refers to the unfiltered voice signal stream continuously collected during the current consultation call, including the voices of both the target inquirer and the intelligent customer service representative.
[0028] It should be noted that the real-time call voice data is automatically collected after obtaining the consent of the target consultation recipient. The voice data stream is continuously captured in real time after the call is established through the voice acquisition interface of the communication system. The voice information of the current conversation is completely saved without any preprocessing, providing a complete and timely raw data foundation for subsequent speaker separation and intent analysis.
[0029] S20. Speaker separation is performed on the real-time call voice data to obtain the original consultation text data of the target consultation object; It's important to understand that speaker separation here refers to separating the real-time call audio data, which includes the voices of both the target customer and the AI customer service representative, from the speaker. This process identifies which audio segments belong to the target customer, and then only those segments are converted to text. The final output is clean, original consultation text data, effectively eliminating audio interference from the AI customer service representative and ensuring that all subsequent analyses are based on the target customer's authentic expression.
[0030] Among them, such as Figure 3 As shown, step S20, which involves speaker separation of the real-time call voice data to obtain the original consultation text data of the target consultation recipient, includes the following steps: S21. Extract features from real-time call voice data to obtain real-time call voice features; S22. Perform speaker identification on the real-time call voice features to obtain speaker identification data; S23. Based on the speaker identification data, audio is extracted from the real-time call voice data to obtain the original audio data of the target consultation object; S24. Perform text conversion on the original audio data to obtain the original consultation text data of the target consultation recipient.
[0031] First, signal processing algorithms (such as short-time Fourier transform, Mel filter bank, etc.) are used to convert the waveform of high-dimensional real-time call speech data into low-dimensional, high-discrimination feature vectors, thereby obtaining real-time call speech features. These real-time call speech features refer to feature vectors extracted from the waveform of real-time call speech data that can effectively represent the essential attributes of speech, preserving key information such as speaker identity and speech content.
[0032] Next, using pre-trained machine learning or deep learning models (such as x-vector, ECAPA-TDNN, etc.), the speaker identification data is obtained by determining which specific speaker (target inquiry recipient or intelligent customer service representative) a particular segment of speech belongs to based on speech features. This speaker identification data is used to indicate which speaker a speech segment in real-time call speech data belongs to, avoiding misidentification of intelligent customer service representative's speech as the target inquiry recipient's speech, thus providing a reliable identity basis for subsequent accurate separation.
[0033] After clarifying the speaker identification data, the pure voice segment corresponding to the target inquirer is separated from the real-time call voice data using the speaker identification data as a guiding condition. The original audio data of the target inquirer is obtained. Through identity-based audio extraction, interference from irrelevant information such as customer service voice can be effectively removed, and the pure voice of the target inquirer can be obtained, ensuring the accuracy of subsequent text conversion and avoiding misjudgment of intent due to voice mixing.
[0034] Finally, a pre-trained speech recognition model (ASR model, such as Conformer, Whisper, etc.) is used to decode the raw audio data frame by frame into corresponding text sequences, obtaining the original consultation text data of the target consultant. The original consultation text data completely preserves the textual data of the content expressed by the target consultant during the call. By converting speech into structured text, the target consultant's requests can be understood and analyzed by the computer in subsequent processing, facilitating intent recognition, summary generation, and intelligent transfer decisions, thus helping to improve the automation level and service accuracy of consultation transfer.
[0035] S30. Extract the context from the original consultation text data to obtain the context fusion vector; It is important to understand that context extraction here refers to associating all statements in the original consultation text data and compressing and fusing the multi-turn dialogue into a context fusion vector. The context fusion vector retains complete semantic information and avoids misunderstandings caused by taking things out of context.
[0036] Among them, such as Figure 4 As shown, step S30, which involves extracting the context from the original consultation text data to obtain the context fusion vector, includes the following steps: S31. Perform context encoding on the original consultation text data to obtain text context encoding features; Specifically, a pre-trained context encoder is used to perform context encoding on the original consultation text data, so that the representation of each word in the original consultation text data incorporates its contextual information, resulting in text context encoding features. Among them, the text context encoding features are summary-oriented context vectors, used to accurately represent the factual narrative flow in the dialogue, providing a basis for subsequent summary generation and service decision-making.
[0037] In one embodiment, the pre-trained context encoder specifically employs a bidirectional GRU network, which includes a forward GRU network (reading text from front to back) and a backward GRU network (reading text from back to front). The forward GRU network generates a forward hidden state sequence in word order, and the backward GRU network generates a backward hidden state sequence in reverse order. Then, the forward and backward hidden states at each position are weighted (e.g., averaged or summed of learnable weights) to output the final text context encoding features. Here, the "hidden state" is the internal memory vector output by the GRU network at each time step, carrying the semantic information up to that position.
[0038] Understandably, because bidirectional GRU networks can capture both preceding and following contextual information, the generated text context encoding features can comprehensively preserve the overall narrative flow and factual information of the original consultation text data, making them suitable for generating accurate dialogue summaries. For example, when it is necessary to show the transfer agent "what the user said," this feature can fully present the chain of facts, avoiding the loss of key information.
[0039] S32. Perform intent encoding on the original consultation text data to obtain intent context encoding features; It is important to understand that intent encoding here refers to extracting feature representations from the original consultation text data that are directly related to the core needs of the target consultation recipient (such as purchase, redemption, registration, consultation, etc.), which can accurately capture the user's core needs and decision-making intent.
[0040] Among them, such as Figure 5 As shown, step S32, which involves encoding the intent of the original consultation text data to obtain intent context encoding features, includes the following steps: S321. Perform word segmentation on the original consultation text data to obtain the call segmentation data; It is important to understand that word segmentation here refers to using pre-defined word segmentation tools (such as jieba, HanLP, etc.) to divide continuous raw consultation text data into independent word sequences according to semantic boundaries. This transforms strings that cannot be directly calculated into word sequences that can be analyzed one by one, enabling subsequent operations to be performed at the word level rather than the character level, thus significantly improving the efficiency and accuracy of intent pattern capture.
[0041] For example, in a fintech scenario, suppose the original consultation text data is "I want to redeem my fixed-income wealth management product now". After word segmentation, we get ["I", "now", "want", "redeem", "fixed-income wealth management product"]. Among them, "fixed-income wealth management product" and "redeem" are correctly segmented into independent words, which facilitates subsequent recognition.
[0042] In a medical technology scenario, assuming the original consultation text data is "I want to buy a neck massager", after word segmentation, we get ["I", "want", "buy", "one", "neck massager"], where "neck massager" is correctly segmented into independent words, which facilitates subsequent recognition.
[0043] S322. Extract intent keywords from the segmented call data to obtain intent keyword data; After obtaining the call segmentation data, we first acquire the pre-defined intent keyword dictionary for fintech or medical technology scenarios. Then, we match the call segmentation data with the intent keyword dictionary and output a concise sequence containing only intent-related keywords. This process filters out a large number of descriptive and modifier words that are irrelevant to the intent, allowing subsequent encoding to focus only on the keywords that truly carry the decision-making intent, thus improving the accuracy of intent encoding.
[0044] S323. Encode the intent keyword data to obtain intent context encoding features.
[0045] It's important to note that the encoding process here refers to mapping the extracted intent keyword data into a low-dimensional dense vector through an embedding layer, then inputting it into an intent-aware convolutional pooling layer (like the EVE convolutional layer) for local pattern capture. Finally, a max pooling layer extracts the most salient features from all convolutional results, aggregating them into a fixed-dimensional vector representation, which is the intent context encoding feature. This intent context encoding feature is an intent-driven context vector, a compact vector that integrates the semantics of all intent keywords and their combinations, used to represent the core decision-making intent of the target consultation recipient.
[0046] S33. Calculate the weights based on the text context encoding features and the intent context encoding features to obtain the fused weight features; S34. Based on the fusion weight features, the text context encoding features and the intent context encoding features are weighted and fused to obtain the context fusion vector.
[0047] Specifically, the text context encoding features and intent context encoding features are concatenated to obtain concatenated encoding features. These concatenated encoding features are then input into a pre-trained small fully connected neural network (containing one linear transform layer and a sigmoid layer) to dynamically balance the importance of the text context encoding features and intent context encoding features, resulting in a fusion weight feature with a value ranging from 0 to 1. This achieves an adaptive dynamic fusion strategy, avoiding the "one-size-fits-all" problem caused by fixed weights. When the user provides detailed facts, the system automatically prioritizes summary accuracy to ensure the agent receives complete background information, with the fusion weight feature approaching 1, indicating greater reliance on the text context encoding features. When the user's intent signal is clear, the system automatically prioritizes intent recognition, with the fusion weight feature approaching 0, indicating greater reliance on the intent context encoding features. This dynamic fusion strategy significantly improves the adaptability of context representation to different types of inquiries.
[0048] After clarifying the fusion weight features, the text context encoding features and intent context encoding features are weighted and fused based on the fusion weight features to obtain the context fusion vector. The context fusion vector can simultaneously take into account the factual completeness and intent accuracy of the dialogue, providing the optimal contextual basis for subsequent call transfer decisions. The specific calculation formula is: Context fusion vector = Fusion weight feature * Text context encoding feature + (1 - Fusion weight feature) * Intent context encoding feature.
[0049] S40. Generate a real-time call summary based on the context fusion vector; It is important to understand that generating real-time call summaries based on context fusion vectors here refers to inputting the context fusion vectors into a pre-trained summary generation model (such as a sequence-to-sequence generation model based on Transformer). The summary generation model decodes and outputs a well-structured summary text that is updated in real time and reflects the key points of the user's inquiry up to the current moment. This allows for a quick grasp of the core content of the user's inquiry, significantly shortens the information acquisition time for human customer service representatives, and improves the response efficiency of customer service personnel.
[0050] Among them, such as Figure 6 As shown, step S40, which involves generating a real-time call summary based on the context fusion vector, includes the following steps: S41. Obtain the current call summary; It should be noted that the current call summary is a summary text generated in real time based on the previously completed call content at the current moment of the current consultation call. The current call summary is a condensed expression of the historical dialogue content. That is, in step S40, only the summary needs to be generated in real time at the current moment, instead of generating it from scratch. This can ensure the continuity and consistency of the summary, avoid repeated calculations, and greatly reduce the computational overhead of real-time generation.
[0051] S42. For each original statement in the original consultation text data, calculate the contribution score of the original statement based on the current call summary and context fusion vector to obtain the contribution score of the current statement. It should be noted that since the statements made before the current moment have all generated the current call summary, the original statement in step S42 refers to the latest content said by the target consultant at the current moment.
[0052] The contribution score calculation here refers to comparing the semantics of the original statement with the context fusion vector to quantify the degree of contribution of the original statement to the summary. This allows for the accurate identification of which statements truly provide new and effective information for the summary and which statements are repetitive or low-value content, thus providing a basis for subsequent compression decisions.
[0053] Among them, such as Figure 7 As shown, step S42, which involves calculating the contribution score of the original statement based on the current call summary and context fusion vector, includes the following steps: S421. Encode the original statement to obtain the statement encoding features; Specifically, pre-trained language models (such as BERT, Transformer, etc.) are used to map the original sentence into a fixed-dimensional high-dimensional vector representation, thus obtaining the sentence encoding features. The sentence encoding features are vectors that can represent the core semantic information of the original sentence.
[0054] S422. Concatenate the sentence encoding features, context fusion vector, and current call summary to obtain the summary concatenation features; It should be noted that the feature concatenation here refers to concatenating the sentence encoding features, context fusion vector, and current call summary along the dimensional direction to obtain the summary concatenation feature. This summary concatenation feature is a comprehensive feature vector that simultaneously contains sentence semantics, contextual information, and existing summary information, thereby avoiding misjudgment of contribution due to incomplete information and making the generated summary more comprehensive and three-dimensional.
[0055] S423. Perform bias calculation on the abstract splicing features to obtain the abstract bias features; The bias calculation here refers to the element-wise addition or linear transformation of a learnable bias term in the summary splicing features to obtain the summary bias features. The summary bias features reflect the weighted projection of the splicing features into the contribution judgment space, while the bias term is a learnable bias vector or bias matrix used to adjust the importance of different dimensions in the summary splicing features. This amplifies certain feature dimensions that are more critical to contribution judgment while suppressing noisy dimensions, thereby making the contribution assessment more focused on key information and improving the accuracy of contribution score calculation.
[0056] S424. Normalize the abstract bias features to obtain the contribution score of the current sentence.
[0057] Finally, the summary bias features are mapped to a uniform range (such as 0 to 1) by a normalization algorithm (such as the Sigmoid function, Softmax function, or Min-Max normalization) to obtain the current statement contribution score. This current statement contribution score is used to quantify the degree of information contribution of the original statement to the current call summary. The closer the score is to 1, the more worthy the original statement is to be included in the current call summary.
[0058] S43. If the contribution score of the current statement exceeds the preset contribution threshold, the original statement is compressed to obtain the target statement. It should be noted that the preset contribution threshold is a pre-set score threshold (such as 0.5) based on the actual application scenario, used to determine whether the original statement has sufficient information value to be included in the current call summary.
[0059] If the contribution score of the current statement exceeds the preset contribution threshold, the original statement is compressed to obtain the target statement, which is a refined statement obtained after compression.
[0060] Compression here refers to the process of removing redundancy and simplifying the expression of the original sentences without losing the core semantics. For example, removing vague words such as "approximately", "around", and "more or less", and breaking long sentences down into short core sentences, thereby ensuring the conciseness of the summary.
[0061] For example, in a fintech scenario, suppose the original statement is "Hello, I would like to learn about your newly launched critical illness insurance," the target statement after compression can be "Customer inquiring about critical illness insurance."
[0062] S44. Update the current call summary based on the target statement to obtain the real-time call summary.
[0063] Finally, the compressed target statement is incrementally merged into the current call summary to generate the latest version of the summary, which is the real-time call summary.
[0064] It's understandable that the current call summary is the call summary generated up to the current moment, while the real-time call summary is the latest call summary generated at the current moment. By using incremental updates instead of full regeneration, the real-time nature of the summary is ensured, and the summary always focuses on the latest call content, avoiding redundancy and distortion caused by the accumulation of invalid information. This significantly improves the conciseness, accuracy, and real-time nature of the call summary. Simultaneously, it provides accurate and concise decision-making basis for transferring the current inquiry call to a human agent.
[0065] S50. Perform intent recognition based on the context fusion vector to obtain real-time intent data; Specifically, the context fusion vector is input into the pre-trained intent classification model. Based on the complete semantic information in the context fusion vector, the intent classification model outputs the most likely intent category of the target consultant and its confidence level, forming real-time intent data. Compared to intent recognition based on single sentences, recognition based on context fusion vectors can accurately capture the true needs of the target consultant.
[0066] In some embodiments, real-time intent data includes three dimensions: intent categories and their confidence levels, namely: A strong intention corresponds to explicit expressions such as "buy now" or "order now"; Conditional intent corresponds to the intent that includes conditions such as "I'll buy it if it's cheaper" or "Are there any free gifts?" Implicit intentions refer to the subtle intentions behind repeatedly asking for product details or requesting information.
[0067] Among them, strong intent, conditional intent, and implicit intent together constitute real-time intent data.
[0068] S60: Transfer the current inquiry call to a preset human customer service representative based on real-time call summary and real-time intent data.
[0069] In step S60 of some embodiments, the real-time call summary and real-time intent data are combined to determine whether the current consultation call needs to be transferred to a preset human customer service representative. If the transfer to a human customer service representative is required, the summary and intent data are simultaneously pushed to the human customer service representative's interface.
[0070] Specifically, such as Figure 8 As shown, step S60, which involves transferring the current inquiry call to a preset live agent based on real-time call summary and real-time intent data, includes the following steps: S61. Perform keyword matching on the real-time call summary based on preset high-intent keywords to obtain intent keyword matching data; It should be noted that the preset high-intent keywords refer to pre-configured high-intent keywords. These high-intent keywords are a set of words selected based on business experience that can clearly indicate that users have a strong intention to process or inquire about services. For example, in the fintech scenario, high-intent keywords may include keywords such as "early repayment," "want to purchase," "loan amount," and "interest rate discount"; in the medical technology scenario, they may include keywords such as "medical appointment," "medical insurance reimbursement," "medical insurance purchase," and "medical device purchase."
[0071] The text in the real-time call summary is compared with preset high-intent keywords to determine whether the text in the real-time call summary contains any number of links from the high-intent keywords. The matching results are then statistically analyzed to obtain intent keyword matching data. Understandably, keyword matching can quickly and interpretably capture explicit high-intent signals from target inquiries, offering greater transparency and controllability compared to a purely black-box model, thus providing data support for subsequent fusion decisions.
[0072] S62. Calculate the purchase intent strength based on the real-time call summary to obtain purchase intent strength data; Next, a pre-trained large model can be used to perform semantic understanding on the real-time call summary, calculate the quantitative strength of the target consultant's current purchase / processing intention, and obtain purchase intention strength data. Purchase intention strength data is used to measure the intensity of the target consultant's current intention to process business or consultation services. The value range is usually from 0 to 1. The higher the value, the stronger the intention. This solves the problem that "conditional intention" may eventually turn into strong intention over time.
[0073] For example, in a fintech scenario, if a target customer says, "Let me look at the information first and then consider it," the semantic analysis shows a purchase intent strength of 0.5, indicating a medium intention. However, if the target customer says, "I want to apply now, just tell me how," the intent strength is 0.95, indicating a high intention. In a healthcare technology scenario, if a target customer says, "I'm just asking around, just to learn more," the intent strength is 0.3, indicating a low intention. However, if the target customer says, "I want to buy this medical insurance product now," the intent strength is 0.88, indicating a high intention.
[0074] S63. Weighted fusion of real-time intent data, intent keyword matching data, and purchase intent intensity data is performed to obtain a fusion decision score; After obtaining the intent keyword matching data and purchase intent intensity data, the real-time intent data, intent keyword matching data, and purchase intent intensity data are weighted and fused according to preset weight parameters to obtain the final fusion decision score. The fusion decision score integrates three dimensions: current intent, explicit keyword signal, and implicit intensity signal. The three complement each other to make the decision more robust and significantly reduce the probability of mis-transfer and missed transfer. At the same time, the configurable weight design allows the system to flexibly adapt to the focus of different business scenarios.
[0075] It should be noted that the preset weight parameters are set according to different business scenarios. For example, in the fintech scenario, more emphasis is placed on real-time intent data and purchase intent intensity data, so higher weights can be assigned to real-time intent data and purchase intent intensity data.
[0076] S64. If the fusion decision score exceeds the preset score threshold, the current consultation call and real-time call summary will be transferred to a human customer service representative.
[0077] It should be noted that the preset scoring threshold is a pre-set score threshold for triggering call transfer based on the actual application scenario, such as 0.5. If the fusion decision score exceeds the preset scoring threshold, the current consultation call, along with the real-time call summary, will be transferred to a human customer service representative. The human customer service representative can see the complete summary and clear intent the moment they answer the call, eliminating the need for the target inquirer to repeat their question, thus achieving a seamless call transition. At the same time, accurate matching based on intent ensures that the user is transferred to the most professional personnel, avoiding multiple layers of transfers and significantly improving transfer efficiency and user satisfaction.
[0078] The pre-set human customer service refers to human agents with corresponding professional skills, such as loan specialists or medical professionals. It should be noted that the capabilities of the human customer service representatives matched to different application scenarios vary, and no specific limitations are made here.
[0079] For example, in a fintech scenario, a fusion decision score of 0.68 exceeds 0.5. When a human customer service representative receives a current inquiry with a summary stating, "The customer is inquiring about mortgage interest rates, mentioning repayment methods, and has a strong desire to apply," they can quickly respond and provide relevant mortgage application services. In a healthcare technology scenario, a fusion decision score of 0.71 exceeds 0.5. When a human customer service representative receives a current inquiry with a summary stating, "The patient's family wants to purchase medical insurance and has a strong desire to do so," they can quickly respond, recommend medical insurance to the patient's family, and provide relevant purchase services.
[0080] By transferring calls with a high purchase intention to human customer service, the workload of human customer service resources can be reduced, the invalid transfer rate can be lowered, and the real-time call summary can be attached so that human customer service can respond quickly without repeating the questions, which significantly improves the accuracy of transfer, the efficiency of human service and the customer experience.
[0081] In some embodiments, in addition to transferring the current inquiry call and the real-time call summary to a human agent, corresponding reminder information can be generated based on real-time intent data to remind the human agent.
[0082] After step S63 in some embodiments, if the fusion decision score does not exceed a preset score threshold, the current consultation call between the target consultant and the intelligent customer service is maintained without transferring to a human customer service representative, thereby reducing the occupation of human customer service resources and lowering the invalid transfer rate.
[0083] As can be seen, in the above solution, for business consultation call scenarios in fintech and healthcare, the first step is to acquire real-time call voice data of the current consultation call between the target customer and the intelligent customer service, completely capturing the voice information of the entire consultation process. Next, speaker separation is performed on the real-time call voice data to obtain the original consultation text data of the target customer, effectively filtering out voice interference from the intelligent customer service side and ensuring the accuracy of the analysis. Furthermore, context extraction is performed on the original consultation text data to obtain a context fusion vector, linking scattered statements into a complete semantic representation, avoiding the one-sidedness of intent recognition; a real-time call summary is generated based on the context fusion vector, facilitating a more accurate understanding of the target customer's consultation points from a semantic perspective; intent recognition is performed based on the context fusion vector to obtain real-time intent data, enabling more precise capture of the user's true needs. Finally, based on the real-time call summary and real-time intent data, the current consultation call is transferred to a preset human customer service representative, allowing for direct response from the human representative, significantly improving transfer efficiency and service experience. This solves the technical problem that existing technologies cannot identify purchase intentions in real-time during the call and immediately transfer the call to an agent, leading to missed optimal service opportunities and impacting the service conversion rate of the consultation.
[0084] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0085] In one embodiment, a consultation call transfer device is provided, which corresponds one-to-one with the consultation call transfer method described in the above embodiments. For example... Figure 9 As shown, the consultation call transfer device includes a call voice acquisition module 101, a speaker separation module 102, a context extraction module 103, a call summary generation module 104, an intent recognition module 105, and a call transfer module 106. Detailed descriptions of each functional module are as follows: The call voice acquisition module 101 is used to acquire real-time call voice data of the current consultation call; wherein, the current consultation call is a call between the target consultation object and the intelligent customer service; Speaker separation module 102 is used to separate the speaker from the real-time call voice data to obtain the original consultation text data of the target consultation object; The context extraction module 103 is used to extract the context from the original consultation text data to obtain a context fusion vector. The call summary generation module 104 is used to generate a real-time call summary based on the context fusion vector. The intent recognition module 105 is used to perform intent recognition based on the context fusion vector to obtain real-time intent data; The call transfer module 106 is used to transfer the current inquiry call to a preset human customer service representative based on real-time call summary and real-time intent data.
[0086] In one embodiment, the speaker separation module 102 is specifically used for: Feature extraction is performed on real-time call voice data to obtain real-time call voice features; Speaker identification data is obtained by performing speaker recognition on real-time call voice features. Audio is extracted from real-time call voice data based on speaker identification data to obtain the original audio data of the target consultation object; The original audio data is converted into text to obtain the original consultation text data of the target consultation recipient.
[0087] In one embodiment, the context extraction module 103 is specifically used for: Context encoding is performed on the original consultation text data to obtain text context encoding features; Intent-based encoding is performed on the original consultation text data to obtain intent-context encoding features; Weights are calculated based on text context encoding features and intent context encoding features to obtain fused weight features; The text context encoding features and intent context encoding features are weighted and fused based on the fusion weight features to obtain the context fusion vector.
[0088] In one embodiment, the context extraction module 103 is specifically used for: The original consultation text data is segmented into words to obtain the call segmentation data. Intent keyword data is obtained by extracting intent keywords from the segmented call data; The intent keyword data is encoded to obtain intent context encoding features.
[0089] In one embodiment, the call summary generation module 104 is specifically used for: Get the current call summary; For each original statement in the original consultation text data, a contribution score is calculated based on the current call summary and context fusion vector to obtain the current statement contribution score; If the contribution score of the current statement exceeds the preset contribution threshold, the original statement is compressed to obtain the target statement. The current call summary is updated based on the target statement to obtain the real-time call summary.
[0090] In one embodiment, the call summary generation module 104 is specifically used for: The original statement is encoded to obtain its encoding features; The sentence encoding features, context fusion vector, and current call summary are concatenated to obtain the summary concatenation features; Bias calculation is performed on the summary splicing features to obtain the summary bias features; The abstract bias features are normalized to obtain the contribution score of the current sentence.
[0091] In one embodiment, the call transfer module 106 is specifically used for: Based on preset high-intent keywords, keyword matching is performed on the real-time call summary to obtain intent keyword matching data; Purchase intent strength data is obtained by calculating the purchase intent strength based on real-time call summaries. The real-time intent data, intent keyword matching data, and purchase intent intensity data are weighted and fused to obtain a fusion decision score; If the fusion decision score exceeds the preset score threshold, the current consultation call and the real-time call summary will be transferred to a human customer service representative.
[0092] This application provides a consultation call transfer device for business consultation call scenarios in fintech and healthcare. First, it acquires real-time voice data of the current consultation call between the target inquirer and the intelligent customer service representative, completely capturing the voice information of the entire consultation process. Next, it performs speaker separation on the real-time voice data to obtain the original consultation text data of the target inquirer, effectively filtering out voice interference from the intelligent customer service representative and ensuring the accuracy of the analysis. Further, it extracts the context from the original consultation text data to obtain a context fusion vector, associating scattered statements into a complete semantic representation, avoiding the one-sidedness of intent recognition. A real-time call summary is generated based on the context fusion vector, facilitating a more accurate understanding of the target inquirer's consultation points from a semantic perspective. Intent recognition is performed based on the context fusion vector to obtain real-time intent data, enabling more precise capture of the user's true needs. Finally, based on the real-time call summary and real-time intent data, the current consultation call is transferred to a preset human customer service representative, allowing for direct response and significantly improving transfer efficiency and service experience. This solves the technical problem that existing technologies cannot identify purchase intentions in real-time during a call and immediately transfer the call to an agent, leading to missed optimal service opportunities and impacting service conversion rates.
[0093] Specific limitations regarding the consultation call transfer device can be found in the limitations of the consultation call transfer method above, and will not be repeated here. Each module in the aforementioned consultation call transfer device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0094] Please see Figure 10 , Figure 10 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 1001 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 1002 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 1002 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1002 and is called and executed by the processor 1001 to execute the consultation call transfer method of the embodiments of this application, including: Acquire real-time voice data of the current consultation call; where the current consultation call is the call between the target inquiry recipient and the intelligent customer service; Speaker separation is performed on real-time call voice data to obtain the original consultation text data of the target consultation recipient; Context extraction is performed on the original consultation text data to obtain a context fusion vector; Generate a real-time call summary based on the context fusion vector; Intent recognition is performed based on the context fusion vector to obtain real-time intent data; Based on real-time call summaries and real-time intent data, the current inquiry call is transferred to a pre-set human customer service representative.
[0095] Input / output interface 1003 is used to implement information input and output; The communication interface 1004 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 1005 transmits information between various components of the device (e.g., processor 1001, memory 1002, input / output interface 1003, and communication interface 1004); The processor 1001, memory 1002, input / output interface 1003 and communication interface 1004 are connected to each other within the device via bus 1005.
[0096] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire real-time voice data of the current consultation call; where the current consultation call is the call between the target inquiry recipient and the intelligent customer service; Speaker separation is performed on real-time call voice data to obtain the original consultation text data of the target consultation recipient; Context extraction is performed on the original consultation text data to obtain a context fusion vector; Generate a real-time call summary based on the context fusion vector; Intent recognition is performed based on the context fusion vector to obtain real-time intent data; Based on real-time call summaries and real-time intent data, the current inquiry call is transferred to a pre-set human customer service representative.
[0097] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0098] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0099] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0100] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0101] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for transferring consultation calls, characterized in that, The method includes: Acquire real-time voice data of the current consultation call; wherein, the current consultation call is a call between the target consultation object and the intelligent customer service; Speaker separation is performed on the real-time call voice data to obtain the original consultation text data of the target consultation object; Context extraction is performed on the original consultation text data to obtain a context fusion vector; A real-time call summary is generated based on the context fusion vector; Intent recognition is performed based on the context fusion vector to obtain the real-time intent data; Based on the real-time call summary and the real-time intent data, the current inquiry call is transferred to a preset human customer service representative.
2. The consultation call transfer method according to claim 1, characterized in that, The step of extracting context from the original consultation text data to obtain a context fusion vector includes: The original consultation text data is subjected to context encoding to obtain text context encoding features; The original consultation text data is subjected to intent encoding to obtain intent context encoding features; Weights are calculated based on the text context encoding features and the intent context encoding features to obtain fused weight features; The text context encoding features and the intent context encoding features are weighted and fused based on the fusion weight features to obtain the context fusion vector.
3. The consultation call transfer method according to claim 2, characterized in that, The process of performing intent encoding on the original consultation text data to obtain intent context encoding features includes: The original consultation text data is segmented into words to obtain call segmentation data; Intent keyword data is obtained by extracting intent keywords from the segmented call data; The intent keyword data is encoded to obtain the intent context encoding features.
4. The consultation call transfer method according to claim 1, characterized in that, The step of generating a real-time call summary based on the context fusion vector includes: Get the current call summary; For each original statement in the original consultation text data, a contribution score is calculated for the original statement based on the current call summary and the context fusion vector to obtain the current statement contribution score; If the contribution score of the current statement exceeds the preset contribution threshold, the original statement is compressed to obtain the target statement; The current call summary is updated based on the target statement to obtain the real-time call summary.
5. The consultation call transfer method according to claim 4, characterized in that, The calculation of the contribution score of the original statement based on the current call summary and the context fusion vector to obtain the current statement contribution score includes: The original statement is encoded to obtain statement encoding features; The sentence encoding features, the context fusion vector, and the current call summary are concatenated to obtain the summary concatenation features; The summary splicing features are subjected to bias calculation to obtain the summary bias features; The summary bias features are normalized to obtain the contribution score of the current statement.
6. The consultation call transfer method according to any one of claims 1 to 5, characterized in that, The step of transferring the current inquiry call to a preset human customer service representative based on the real-time call summary and the real-time intent data includes: Based on preset high-intent keywords, keyword matching is performed on the real-time call summary to obtain intent keyword matching data; Purchase intent strength is calculated based on the real-time call summary to obtain purchase intent strength data; The real-time intent data, the intent keyword matching data, and the purchase intent intensity data are weighted and fused to obtain a fusion decision score; If the fusion decision score exceeds a preset score threshold, the current consultation call and the real-time call summary will be transferred to the human customer service representative.
7. The consultation call transfer method according to any one of claims 1 to 5, characterized in that, The step of performing speaker separation on the real-time call voice data to obtain the original consultation text data of the target consultation recipient includes: Feature extraction is performed on the real-time call voice data to obtain real-time call voice features; Speaker identification data is obtained by performing speaker recognition on the real-time call voice features; Based on the speaker identification data, audio is extracted from the real-time call voice data to obtain the original audio data of the target consultation object; The original audio data is converted into text to obtain the original consultation text data of the target consultation recipient.
8. A consultation call transfer device, characterized in that, The device includes: The call voice acquisition module is used to acquire real-time call voice data of the current consultation call; wherein, the current consultation call is a call between the target consultation object and the intelligent customer service; The speaker separation module is used to separate the speaker from the real-time call voice data to obtain the original consultation text data of the target consultation object. The context extraction module is used to extract the context from the original consultation text data to obtain a context fusion vector; A call summary generation module is used to generate a real-time call summary based on the context fusion vector. The intent recognition module is used to perform intent recognition based on the context fusion vector to obtain the real-time intent data; The call transfer module is used to transfer the current inquiry call to a preset human customer service representative based on the real-time call summary and the real-time intent data.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the consultation call transfer method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the consultation call transfer method as described in any one of claims 1 to 7.