Outbound response optimization method and device, computer equipment and storage medium

By combining automatic speech recognition and multi-model processing with response quality analysis, the response mechanism of intelligent customer service or intelligent assistant is optimized, solving the response latency problem of large language models in highly interactive scenarios, and achieving more natural interaction and higher response accuracy.

CN120895037APending Publication Date: 2025-11-04PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511200452.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-04

AI Technical Summary

Technical Problem

Existing intelligent customer service or intelligent assistants driven by large language models suffer from significant response and inference delays in highly interactive and real-time scenarios, resulting in a poor user experience.

Method used

Automatic speech recognition methods are used to obtain speech recognition text, and the speech recognition text is processed by a large language model, a lightweight micro-response prediction model and a semantic cache search and matching method. Combined with response quality analysis, an optimization mechanism is used to obtain the optimal response.

Benefits of technology

It reduces reliance on large language models, avoids cloud fluctuation risks, improves response efficiency, shortens response inference latency, enhances robustness, adapts to high-concurrency scenarios, supports large-scale deployment, improves response accuracy and stability, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895037A_ABST
    Figure CN120895037A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, is applied to intelligent medical and financial scenes, and discloses an outbound response optimization method and device, computer equipment and a storage medium, and the method comprises the steps: receiving user voice, employing an automatic voice recognition method to carry out the voice recognition of the user voice, and obtaining a voice recognition text; processing the obtained speech recognition text by using a large language model, a lightweight micro-response prediction model and a semantic cache search matching method to obtain a large model response, a prediction response and a cache matching response; and performing response quality analysis on the large model response, the prediction response and the cache matching response, acquiring the response with the optimal response quality as the final response by adopting a preferential mechanism, and outputting the final response. Dependence on a large language model is reduced, response efficiency is improved, response reasoning delay is shortened, interaction is more natural, and user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and natural language processing, and in particular to an outbound call response optimization method and device, a computer device and a storage medium. BACKGROUND

[0002] Large language models (LLM) are widely used in intelligent customer service or intelligent assistants, but in scenarios requiring strong interaction and strong real-time, the existing large language model-driven intelligent customer service or intelligent assistants mostly adopt a "single round request-response" mode, which is time-consuming and has reasoning delay, and there is a problem of obvious response reasoning delay, resulting in poor user experience. SUMMARY

[0003] The present application provides an outbound call response optimization method, device, computer device and medium to solve the technical problem of poor user experience caused by obvious response reasoning delay in the existing large language model-driven intelligent customer service or intelligent assistant.

[0004] In a first aspect, an outbound call response optimization method is provided, comprising:

[0005] receiving user speech, using an automatic speech recognition method to perform speech recognition on the user speech, and obtaining speech recognition text;

[0006] respectively using a large language model, a lightweight micro-response prediction model and a semantic cache search matching method to process the obtained speech recognition text, and obtaining a large model response, a predicted response and a cache matching response;

[0007] performing response quality analysis on the large model response, the predicted response and the cache matching response, using an optimization mechanism to obtain the response with the best response quality as the final response, and outputting the final response.

[0008] In a second aspect, an outbound call response optimization device is provided, comprising:

[0009] a speech recognition module configured to receive user speech, use an automatic speech recognition method to perform speech recognition on the user speech, and obtain speech recognition text;

[0010] a multi-response generation module configured to respectively use a large language model, a lightweight micro-response prediction model and a semantic cache search matching method to process the obtained speech recognition text, and obtain a large model response, a predicted response and a cache matching response;

[0011] a response quality analysis module configured to perform response quality analysis on the large model response, the predicted response and the cache matching response, use an optimization mechanism to obtain the response with the best response quality as the final response, and output the final response.

[0012] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the outbound call response optimization method described above.

[0013] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the aforementioned outbound call response optimization method.

[0014] In the aforementioned outbound call response optimization method, apparatus, computer equipment, and storage medium, the solution involves receiving user voice through a client, performing automatic speech recognition on the user voice to obtain speech recognition text, processing the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain a large model response, a predicted response, and a cache matching response, analyzing the response quality of the large model response, predicted response, and cache matching response, using an optimization mechanism to obtain the response with the best response quality as the final response, outputting the final response, and feeding the final response back to the client. In this invention, for intelligent customer service in the medical field, or for intelligent customer service in the financial field, outbound call response optimization can be utilized. This approach involves obtaining speech-recognized text through speech recognition of user speech. It then processes the obtained speech-recognized text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method. This yields the large model response, the predicted response, and the cached matching response, allowing for simultaneous asynchronous parallel processing of the speech-recognized text. This reduces reliance on the large language model, mitigates cloud fluctuation risks, enhances robustness, improves response efficiency, shortens response inference latency, and makes interactions more natural. Furthermore, it reduces the average processing time per turn of dialogue, adapts to high-concurrency scenarios, supports large-scale deployment, and has strong scalability. By using an optimization mechanism based on response quality analysis to select the response with the best quality as the final response and output it, it improves response accuracy and stability, thereby enhancing the user experience. Attached Figure Description

[0015] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of an application environment for an outbound call response optimization method according to an embodiment of the present invention;

[0017] Figure 2This is a flowchart illustrating an outbound call response optimization method according to an embodiment of the present invention;

[0018] Figure 3 This is a schematic diagram of an outbound call response optimization device according to an embodiment of the present invention;

[0019] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0020] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] The outbound call response optimization method provided in this embodiment of the invention can be applied to, for example, Figure 1In application environments such as healthcare and finance, intelligent customer service or intelligent assistants are typically implemented through a server-side architecture, where the client communicates with the server via a network. The server receives user voice from the client, performs automatic speech recognition to obtain speech-recognized text, and processes the text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, predicted response, and cached matching response. Response quality analysis is performed on these responses, and a selection mechanism is used to obtain the response with the best quality as the final response, which is then output and fed back to the client. In this invention, for intelligent customer service in the healthcare or financial fields, an outbound call response optimization scheme can be used to optimize user voice through speech recognition. The speech recognition method obtains speech-recognized text, which is then processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method. This yields the large model response, predicted response, and cached matching response, allowing for simultaneous asynchronous parallel processing of the speech-recognized text. This reduces reliance on the large language model, mitigates cloud fluctuation risks, enhances robustness, improves response efficiency, shortens response inference latency, and makes interactions more natural. It also reduces the average processing time per turn of dialogue, adapts to high-concurrency scenarios, supports large-scale deployment, and has strong scalability. By using a selection mechanism based on response quality analysis to obtain the response with the best quality as the final response and output it, the method improves response accuracy and stability, and enhances user experience. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.

[0023] Please see Figure 2 As shown, Figure 2 A flowchart illustrating the outbound call response optimization method provided in this embodiment of the invention includes the following steps:

[0024] S10: Receive user voice, use automatic speech recognition method to perform speech recognition on user voice, and obtain speech recognition text.

[0025] Automatic Speech Recognition (ASR) methods are used to perform speech recognition, converting user speech into text to obtain speech recognition text, which is then analyzed and processed to generate appropriate responses. The outbound call response optimization method provided in this invention can be applied to intelligent customer service or intelligent assistants in various application scenarios such as healthcare, finance, and insurance. It is typically implemented through a server-side mechanism that can receive user voice messages in real time. For example, in the medical field, medical device products can be promoted to users via telephone, often requiring intelligent customer service or intelligent assistants to answer user inquiries about these products to improve user experience. For instance, users can inquire about the use or price of related medical device products via voice call. An automatic speech recognition method can be used to recognize the user's speech, obtaining speech recognition text, which is then processed to generate a response.

[0026] Alternatively, for example, in the field of financial applications, financial products can be promoted to users via telephone, and intelligent customer service or intelligent assistants can be used to reply with relevant information about financial products. Automatic speech recognition methods can be used to recognize user voice and obtain speech-recognized text, which can be processed and responded to in a subsequent manner.

[0027] S20: The obtained speech recognition text is processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method, respectively, to obtain the large model response, the prediction response, and the cache matching response.

[0028] Step S20, which involves processing the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, includes the following steps:

[0029] Semantic embedding clustering method is used to perform semantic clustering on historical dialogues to obtain representative questions and their corresponding responses;

[0030] The search and matching are performed based on the representative questions obtained by combining the speech recognition text. The representative questions that semantically match the speech recognition text are obtained based on semantic similarity, and the response corresponding to the representative question is obtained as the cached matching response.

[0031] Among these features, existing responses can be directly reused by searching and matching representative questions, thus improving response speed. Representative questions are those that are considered representative, meaning they occur more frequently than a preset frequency threshold, exhibit high response stability, and do not require strong contextual dependence. Representative questions also serve as semantic cluster centers. High response stability means that the differences between multiple answers are small, and the lack of strong contextual dependence indicates a generic answer. Semantic cluster centers refer to those with high cluster similarity to other questions. For example, typical high-frequency questions such as "How much does this cost?" and "How long will it take to deliver?" are representative questions. For responses to questions containing variables, a response template combined with real-time variable filling can be used to generate corresponding responses, giving the responses personalized capabilities. For example, "The price is" and "The delivery date is" can be used as response templates, which, after being filled with real-time variables, can become the corresponding responses.

[0032] The response can include opening greetings and closing invitations. These response fragments can be extracted separately and concatenated with the response from the large model, which can reduce the length of the response generated by the large language model and improve speed.

[0033] Specifically, step S20 involves processing the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, including:

[0034] Based on historical intent tags, the obtained speech recognition text is subjected to intent classification prediction and matching to obtain the matching intent category. Based on the matching intent category, a predefined predictive response fragment template is called to generate the predictive response.

[0035] Specifically, combining speech recognition text with historical intent tags for intent classification and prediction matching can determine the approximate question intent of the speech recognition text, generate predicted responses, accelerate response speed, effectively compensate for the inference latency of large language models, and ensure natural dialogue flow. Historical intent tags can be obtained from manually annotated historical dialogue intent datasets, or automatically classified through intent recognition models in historical dialogues. Historical intent tags can also be obtained by defining intent tags in historical dialogues using semantic embedding clustering methods, or by synchronizing tags from historical dialogues through a customer relationship management system. Historical intent tags include prices and shipping dates, among others.

[0036] For example, when promoting medical device products in the medical field or financial products in the financial field, if a user asks "What is the unit price of the product?", the system performs intent classification and prediction matching based on historical intent tags. If the matched intent category is price, a predefined predictive response fragment template is called, and the predictive response is generated as "The current unit price of this product is".

[0037] Specifically, in some embodiments, the step of performing intent classification prediction and matching on the obtained speech recognition text based on historical intent tags includes:

[0038] The speech recognition text input is used to perform intent classification, prediction, and matching based on historical intent labels. This open-source text modeling tool can be fastText.

[0039] Specifically, in some embodiments, the step of performing intent classification prediction and matching on the obtained speech recognition text based on historical intent tags includes:

[0040] The speech recognition text is encoded into BERT (Bidirectional Encoder Representation from Transformer) vectors. Based on historical intent labels, the MLP (Multilayer Perceptron) model is used to perform intent classification prediction and matching on the BERT vectors.

[0041] Specifically, in some embodiments, the step of performing intent classification prediction and matching on the obtained speech recognition text based on historical intent tags includes:

[0042] A small semantic allocator is used to perform intent classification prediction and matching on speech recognition text based on historical intent labels.

[0043] Among them, the small semantic allocator can be a DSSM (Deep Structured Semantic Model) model or a dual-tower semantic allocator (dual-tower model).

[0044] Specifically, step S20, which involves processing the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, includes the following steps:

[0045] Obtain historical dialogues, perform semantic compression on the historical dialogues, and obtain a summary of the historical dialogues;

[0046] By using historical dialogue summaries as cues for the large language model, the obtained speech recognition text is processed to obtain the large model response.

[0047] Specifically, the number of rounds of historical dialogues acquired should be no more than 10 to ensure the accuracy of the extracted historical dialogue summaries. Using these summaries as cues for the large language model reduces redundant input, accelerates the generation of large model responses, and reduces response inference latency.

[0048] Specifically, in some embodiments, the semantic compression of historical dialogues includes:

[0049] The large language model or extraction compression model is used to extract and compress information based on keywords in user-asked questions from historical dialogues, information related to keywords in responses, and user-focused feedback.

[0050] For example, if a user's question is "the price of the product," then the keyword for the question is "price." A response could be "the price of the product is..." or "this product is also participating in a discount promotion." The product price and discount information in the response would then be information related to the keyword in the question. The user's focus based on the response might be "is the product genuine?" or "how can I verify its authenticity?" Therefore, the user's focus based on the response would be on product authenticity assurance. Semantic compression of historical dialogues yields a historical dialogue summary, retaining only key instructions and information points. This historical dialogue summary, combined with the obtained speech recognition text, is then processed using a large language model, which can shorten the processing time of the large language model and improve efficiency.

[0051] Preferably, the user's question can be a product brand history introduction. The brand history introduction in the large model response can be generated using a fill-in-the-blank mode, shortening the response generation time by combining templates with sample prompts and input variables. The fill-in-the-blank mode can be a Few-shot fill-in-the-blank mode. For example, the predefined template for the brand history introduction is: "{Product Name} is a {Product Type} product launched by our {Brand Name}, {Product Related Information Description}, suitable for {Usage Scenarios}". The sample prompt for the brand history introduction is: "Jiangxiaobai sorghum liquor is a light-aroma product launched by our Jiangxiaobai, with a smooth taste, suitable for gatherings with friends". The input variables are: "{Product Name} = Langjiu Honghualang 10, {Brand Name} = Langjiu, {Type} = Sauce-aroma type, {Product Related Information Description} = Smooth taste". The final large model response is: "Langjiu Honghualang 10 is a sauce-aroma product launched by our Langjiu, with a smooth taste, suitable for gatherings with friends".

[0052] Preferably, the large model response can be implemented using asynchronous decoding, which allows the generated token to be output in real time without waiting for the complete statement to be generated, thus shortening the response inference latency.

[0053] Specifically, speech recognition and response generation are distributed across different asynchronous task queues, and communication is coordinated using an event-driven and status notification approach. This ensures that the inference latency is determined by the longest step rather than the sum of all steps, effectively shortening the response inference latency.

[0054] S30: Perform response quality analysis on the large model response, predicted response, and cache-matched response, and use an optimization mechanism to obtain the response with the best response quality as the final response, and output the final response.

[0055] Specifically, step S30, which involves performing response quality analysis on the large model response, predicted response, and cache-matched response, using an optimization mechanism to obtain the response with the best response quality as the final response, and outputting the final response, includes:

[0056] The response quality of the large model response, predicted response, and cache matching response is evaluated based on semantic coverage, context consistency, and response time. The weighted score of each response is calculated by combining the weight coefficients corresponding to each response. The response with the highest weighted score is selected as the response with the best response quality and is used as the final response. The final response is then output.

[0057] Semantic coverage is used to determine whether the response answers the core question; contextual consistency is used to determine whether the response continues the main dialogue; and response time is used to determine whether the response is better—a shorter response time indicates a better response. A language model can also be used to score the fluency and politeness of the response. A weighted score is calculated based on a comprehensive evaluation of semantic coverage, contextual consistency, and response time, combined with corresponding response weight coefficients. This achieves a holistic consideration of response quality. By dynamically selecting the best response based on its quality, both efficiency and naturalness can be guaranteed. The weight coefficients of each response are ranked as follows: the largest weight coefficient is for responses from large models, followed by cached matching responses, and the smallest weight coefficient is for predicted responses.

[0058] Preferably, in some embodiments, the step of performing response quality analysis on the large model response, predicted response, and cache-matched response, and using an optimization mechanism to obtain the response with the best response quality as the final response, further includes:

[0059] Calculate the cache hit confidence for the cache-matched response and determine whether the cache hit confidence is higher than the preset confidence threshold. If so, the cache-matched response is the response with the best response quality and is used as the final response.

[0060] In some embodiments, response quality analysis is performed on the large model response, predicted response, and cache-matched response, and an optimization mechanism is used to obtain the response with the best response quality as the final response. This also includes:

[0061] A prediction bias analysis is performed on the predicted response, and the semantic similarity score between the predicted response and the large model response is calculated. When the semantic similarity score is not lower than the preset similarity threshold, the response quality of the combined predicted response and the large model response is optimal, and this response is taken as the final response.

[0062] Multi-path response improves the efficiency of interactive responses. The parallel inference and streaming synthesis structure reduces the average single-turn dialogue processing time, adapting to high-concurrency scenarios. The semantic similarity score can be either the BERT vector score or the ROUGE (Recall-Oriented Understudy for Gisting Evaluation) score, with a preset similarity threshold of 0.6. During the playback of the predicted response, the speech progress can be dynamically monitored. When the large model response is ready and semantically consistent, the playback of the second half of the large model response is seamlessly switched. When the large model response is ready but the semantics of the predicted response are inconsistent with the large model response, the playback of the complete large model response is switched after the predicted response has finished.

[0063] When the semantic similarity score is lower than the preset similarity threshold, it is determined that the predicted response is seriously biased. The predicted response is replaced by the large model response or the cached matching response, and the replay process is automatically triggered.

[0064] In some embodiments, after step S30, that is, after performing response quality analysis on the large model response, predicted response, and cache matching response, and using an optimization mechanism to obtain the response with the best response quality as the final response, and after outputting the final response, the outbound call response optimization method further includes:

[0065] The final response is converted into speech using a text-to-speech method, and the converted speech is then played.

[0066] During the text-to-speech process, the speech synthesis status and response specifications can be monitored. If either the speech synthesis status or the response specifications are abnormal, an interruption prompt will be played. The interruption prompt may be something like, "Please wait while we verify this." Abnormal speech synthesis statuses include synthesis failure, the presence of illegal characters in the text, and text length exceeding a threshold. Synthesis failure can be caused by insufficient memory or a corrupted audio synthesizer. Illegal characters can be special symbols or emoticons. Abnormal response specifications include semantically incoherent responses, grammatically incorrect responses, conflicts between the response content and the context, and repetitions, sentence skipping, or abrupt changes in language style within the response content.

[0067] When the replay process is triggered, the frequency of the end note of the original response segment and the beginning note of the new response segment must be aligned using a pitch smoothing method, and the replacement large model response or cached matching response must be played seamlessly.

[0068] As can be seen, in the above solution, for intelligent customer service in the medical field or the financial field, speech recognition text is obtained by performing speech recognition on the user's voice. The obtained speech recognition text is then processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, predicted response, and cached matching response. This allows for simultaneous asynchronous parallel processing of speech recognition text, reducing reliance on the large language model, mitigating cloud fluctuation risks, enhancing robustness, improving response efficiency, shortening response inference latency, making the interaction more natural, reducing the average single-turn dialogue processing time, adapting to high-concurrency scenarios, supporting large-scale deployment, and exhibiting strong scalability. By using an optimization mechanism based on response quality analysis to obtain the response with the best response quality as the final response and output it, the response accuracy and stability are improved, thus enhancing the user experience.

[0069] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0070] In one embodiment, an outbound call response optimization device is provided, which corresponds one-to-one with the outbound call response optimization method in the above embodiments. For example... Figure 3 As shown, the outbound call response optimization device includes a voice recognition module 101, a multi-response generation module 102, and a response quality analysis module 103. Detailed descriptions of each functional module are as follows:

[0071] The speech recognition module 101 is used to receive user speech, perform speech recognition on user speech using an automatic speech recognition method, and obtain speech recognition text;

[0072] The multi-response generation module 102 is used to process the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method respectively, to obtain the large model response, the predicted response, and the cache matching response.

[0073] The response quality analysis module 103 is used to perform response quality analysis on the large model response, predicted response and cached matching response, and adopts the best selection mechanism to obtain the response with the best response quality as the final response and output the final response.

[0074] In one embodiment, the multi-response generation module 102 is specifically used for:

[0075] Semantic embedding clustering method is used to perform semantic clustering on historical dialogues to obtain representative questions and their corresponding responses;

[0076] The search and matching are performed based on the representative questions obtained by combining the speech recognition text. The representative questions that semantically match the speech recognition text are obtained based on semantic similarity, and the response corresponding to the representative question is obtained as the cached matching response.

[0077] In one embodiment, the multi-response generation module 102 is specifically used for:

[0078] Based on historical intent tags, the obtained speech recognition text is subjected to intent classification prediction and matching to obtain the matching intent category. Based on the matching intent category, a predefined predictive response fragment template is called to generate the predictive response.

[0079] In one embodiment, the multi-response generation module 102 is specifically used for:

[0080] Obtain historical dialogues, perform semantic compression on the historical dialogues, and obtain a summary of the historical dialogues;

[0081] By using historical dialogue summaries as cues for the large language model, the obtained speech recognition text is processed to obtain the large model response.

[0082] In one embodiment, the multi-response generation module 102 is further configured to:

[0083] The large language model or extraction compression model is used to extract and compress information based on keywords in user-asked questions from historical dialogues, information related to keywords in responses, and user-focused feedback.

[0084] In one embodiment, the response quality analysis module 103 is specifically used for:

[0085] The response quality of the large model response, predicted response, and cache matching response is evaluated based on semantic coverage, context consistency, and response time. The weighted score of each response is calculated by combining the weight coefficients corresponding to each response. The response with the highest weighted score is selected as the response with the best response quality and is used as the final response. The final response is then output.

[0086] In one embodiment, the response quality analysis module 103 is further configured to:

[0087] The final response is converted into speech using a text-to-speech method, and the converted speech is then played.

[0088] This invention provides an outbound call response optimization device. It obtains speech recognition text by performing speech recognition on user speech, and processes the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method. This yields a large model response, a predicted response, and a cached matching response, enabling simultaneous asynchronous parallel processing of the speech recognition text. This reduces reliance on the large language model, mitigates cloud fluctuation risks, enhances robustness, improves response efficiency, shortens response inference latency, and makes interactions more natural. It also reduces the average processing time per turn of dialogue, adapts to high-concurrency scenarios, supports large-scale deployment, and has strong scalability. By using an optimization mechanism based on response quality analysis to obtain the response with the best quality as the final response and output it, it improves response accuracy and stability, and enhances user experience.

[0089] Specific limitations regarding the outbound call response optimization device can be found in the limitations of the outbound call response optimization method described above, and will not be repeated here. Each module in the aforementioned outbound call response optimization device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0090] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the server-side functions or steps of an outbound call response optimization method.

[0091] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements client-side functions or steps of an outbound call response optimization method.

[0092] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0093] Receive user voice, use automatic speech recognition method to recognize user voice, and obtain speech recognition text;

[0094] The obtained speech recognition text was processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method, respectively, to obtain the large model response, the prediction response, and the cache matching response.

[0095] Response quality analysis is performed on the large model response, predicted response, and cached matching response. The best response is selected as the final response using an optimization mechanism, and the final response is output.

[0096] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0097] Receive user voice, use automatic speech recognition method to recognize user voice, and obtain speech recognition text;

[0098] The obtained speech recognition text was processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method, respectively, to obtain the large model response, the prediction response, and the cache matching response.

[0099] Response quality analysis is performed on the large model response, predicted response, and cached matching response. The best response is selected as the final response using an optimization mechanism, and the final response is output.

[0100] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0101] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0102] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0103] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for optimizing outbound call response, characterized in that, include: Receive user voice, use automatic speech recognition method to recognize user voice, and obtain speech recognition text; The obtained speech recognition text was processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method, respectively, to obtain the large model response, the prediction response, and the cache matching response. Response quality analysis is performed on the large model response, predicted response, and cached matching response. The best response is selected as the final response using an optimization mechanism, and the final response is output.

2. The outbound call response optimization method as described in claim 1, characterized in that, The obtained speech recognition text is processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, including: Semantic embedding clustering method is used to perform semantic clustering on historical dialogues to obtain representative questions and their corresponding responses; The search and matching are performed based on the representative questions obtained by combining the speech recognition text. The representative questions that semantically match the speech recognition text are obtained based on semantic similarity, and the response corresponding to the representative question is obtained as the cached matching response.

3. The outbound call response optimization method as described in claim 1, characterized in that, The obtained speech recognition text is processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, including: Based on historical intent tags, the obtained speech recognition text is subjected to intent classification prediction and matching to obtain the matching intent category. Based on the matching intent category, a predefined predictive response fragment template is called to generate the predictive response.

4. The outbound call response optimization method as described in claim 1, characterized in that, The obtained speech recognition text is processed using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method to obtain the large model response, the predicted response, and the cache matching response, including: Obtain historical dialogues, perform semantic compression on the historical dialogues, and obtain a summary of the historical dialogues; By using historical dialogue summaries as cues for the large language model, the obtained speech recognition text is processed to obtain the large model response.

5. The outbound call response optimization method as described in claim 4, characterized in that, The semantic compression of historical dialogues includes: The large language model or extraction compression model is used to extract and compress information based on keywords in user-asked questions from historical dialogues, information related to keywords in responses, and user-focused points based on responses.

6. The outbound call response optimization method as described in claim 1, characterized in that, The process involves analyzing the response quality of the large model response, predicted response, and cache-matched response, employing a selection mechanism to obtain the response with the best quality as the final response, and outputting the final response, including: The response quality of the large model response, predicted response, and cache matching response is evaluated based on semantic coverage, context consistency, and response time. The weighted score of each response is calculated by combining the weight coefficients corresponding to each response. The response with the highest weighted score is selected as the response with the best response quality and is used as the final response. The final response is then output.

7. The outbound call response optimization method as described in claim 1, characterized in that, After the final output response, the outbound call response optimization method further includes: The final response is converted into speech using a text-to-speech method, and the converted speech is then played.

8. An outbound call response optimization device, characterized in that, include: The speech recognition module is used to receive user speech, perform speech recognition on user speech using automatic speech recognition methods, and obtain speech recognition text. The multi-response generation module is used to process the obtained speech recognition text using a large language model, a lightweight micro-response prediction model, and a semantic cache search and matching method respectively, to obtain the large model response, the predicted response, and the cache matching response. The response quality analysis module is used to perform response quality analysis on the large model response, predicted response, and cached matching response. It uses an optimization mechanism to select the response with the best response quality as the final response and outputs the final response.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the outbound call response optimization method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the outbound call response optimization method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Session semantic extraction and intelligent combination method for audio semantic recognition

    CN121862095A