An intelligent outbound call dialog method and system based on a multi-modal pipeline arrangement
Patent Information
- Application Number
- CN202611036127.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-13
- Publication Date
- 2026-09-15
Smart Images

Figure CN122765130A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent voice interaction technology, and in particular to an intelligent outbound call dialogue method and system based on multimodal pipeline orchestration. Background Technology
[0002] With the rapid development of artificial intelligence, large language models, real-time speech recognition, and call center technologies, intelligent outbound calling has been widely applied in various business scenarios such as telesales, customer follow-up, notification reminders, financial collection, insurance follow-up, medical follow-up, and government notifications. Compared with the traditional manual outbound calling mode, intelligent outbound calling systems can automatically complete customer list import, batch dialing, voice interaction, intent recognition, call record saving, and result statistics, which reduces the cost of human agents to a certain extent and improves the efficiency of outbound call reach and business processing.
[0003] Existing intelligent outbound calling platforms typically include functions such as outbound call task management, customer list management, SIP trunk access, automatic dialing, ASR speech recognition, NLP or large language model dialogue, TTS speech synthesis, recording retention, call record query, and human agent transfer. They can support basic automatic outbound calling and voice interaction processes, but still have many shortcomings in practical applications:
[0004] The voice interaction process is serial, resulting in high end-to-end response latency. Existing systems generally adopt a fully serial processing mode of "complete ASR transcription - complete response generation by the dialogue engine - complete TTS synthesis - overall playback". After the user finishes speaking, the system needs to wait for the complete transcription, response generation, and synthesis of the speech before it can be played to the user. This is acceptable in simple notification outbound calls, but in multi-turn natural language interaction scenarios, it causes obvious pauses in speech, a strong sense of waiting for the user, and poor dialogue naturalness. Especially in scenarios such as continuous question and answer and objection handling, it reduces the customer's willingness to answer and the overall interaction experience.
[0005] Dialogue engines suffer from single-point dependencies and insufficient service continuity. Existing outbound call systems often rely on a single-path dialogue engine, such as a fixed script, a traditional NLP dialogue platform, or a single large language model service. When the main dialogue service experiences network anomalies, interface timeouts, model unavailability, or knowledge base access failures, the system typically lacks automatic fallback and backup dialogue paths. This can easily lead to failed responses to ongoing calls, prolonged periods of unresponsiveness, or even direct interruptions, impacting system stability under high concurrency and long-term operation scenarios.
[0006] The lack of idempotent handling for call events leads to poor data consistency. Intelligent outbound calling generates numerous call events, including call initiation, ringing, connection, hanging up, agent transfer, recording generation, and transcription completion. Due to network retransmissions, out-of-order events, and system restarts between the softswitch platform, business systems, and databases, the same event may be repeatedly pushed or written. The absence of a reliable idempotent event mechanism easily results in duplicate call records, incorrect status updates, inaccurate statistical results, and abnormal subsequent business processes.
[0007] The existing outbound call scheduling solutions suffer from a lack of adaptability and self-healing capabilities due to their limited dialing modes. Some solutions only support a single dialing mode or focus on progressive dialing in human agent scenarios, failing to simultaneously meet the demands of high-concurrency automated outbound calls and stable connection control for AI robots. Furthermore, when SIP trunks malfunction, gateways become unavailable, or call failure rates increase, the system often relies on manual maintenance, lacking automatic monitoring, recovery, and scheduling slowdown mechanisms, thus impacting the continuous execution of batch outbound call tasks.
[0008] Insufficient utilization of call results and lack of a complete business loop. Existing intelligent outbound calling systems typically focus more on the call process itself, with insufficient support for semantic analysis and business follow-up after the call ends. While some systems can save recordings, transcribe text, or generate summaries, call results often remain at the recording level, lacking an automatic flow mechanism from call content to structured intent, intent level, recommended actions, and follow-up task pools. High-intent customers cannot be identified and prioritized for follow-up in a timely manner, reducing the conversion efficiency of outbound calling business.
[0009] Therefore, there is an urgent need to design an intelligent outbound call dialogue solution based on multimodal pipeline orchestration, which will uniformly orchestrate the entire outbound call link and optimize it from multiple dimensions such as interactive experience, service stability, data reliability, scheduling efficiency and business closed loop, so as to solve the above-mentioned problems of existing technologies. Summary of the Invention
[0010] The purpose of this invention is to at least solve one of the technical problems existing in the prior art, and to provide an intelligent outbound call dialogue method and system based on multimodal pipeline orchestration, which can solve the problems in the background art mentioned above.
[0011] To achieve the above objectives, the present invention provides the following technical solution: an intelligent outbound call dialogue method based on multimodal pipeline orchestration, comprising the following steps:
[0012] S1. Create an outbound call task and configure the task parameters. The task scheduling engine starts the batch dialing process according to the task configuration.
[0013] S2. Initiate an outbound call request through the SIP trunk gateway, receive call events, and perform an exact one-time write of the event based on the idempotent key;
[0014] S3. After the call is connected, a two-way audio channel is established, and the user's voice stream is sent to the ASR service in real time for sentence-by-sentence transcription.
[0015] S4. Perform keyword detection on the ASR transcribed text. If the keyword trigger condition for manual transfer is met, execute the manual transfer process.
[0016] If the call is not successful, the transcribed text and the call context will be sent to the dialogue engine to generate a response.
[0017] S5. Detect the availability of the main dialogue engine. If the main dialogue engine is available, generate a streaming response text using the main dialogue engine.
[0018] If the primary dialogue engine is unavailable, it will automatically switch to the degraded dialogue engine to generate the response text.
[0019] S6. The convection reply text is segmented in real time according to the sentence boundaries. Each time a complete sentence unit is generated, TTS speech synthesis is triggered immediately and the synthesized speech is pushed to the call media stream for playback.
[0020] S7. After each round of dialogue, update the user's intention level and call context, and determine whether the call termination conditions are met. If not, continue to the next round of dialogue.
[0021] S8. After the call ends, a structured call summary is generated. Customers are automatically transferred to the follow-up pool based on their intent level and call events, and the end-to-end monitoring data is updated.
[0022] Preferably, in step S6, the real-time segmentation of the streaming response text according to sentence boundaries specifically involves using periods, question marks, exclamation marks, semicolons, and line breaks as sentence boundary identification markers. During the streaming output process of the dialogue engine, whenever a complete sentence or playable semantic segment is identified, the text segment is immediately submitted to the TTS service for speech synthesis, without waiting for the complete response to be generated.
[0023] ASR transcription, dialogue generation, and TTS synthesis are performed in a pipeline-like parallel collaboration at the sentence level.
[0024] Preferably, in step S5, the main dialogue engine uses a knowledge-based dialogue agent or a cloud-based large language model service to carry the business knowledge base, multi-turn context management, and natural language response generation.
[0025] The degraded dialogue engine uses a local large language model combined with a rule-based dialogue library and a retrieval enhancement generation module to provide basic business response capabilities when the main dialogue engine experiences interface timeouts, network anomalies, service unavailability, or knowledge base access failures.
[0026] Preferably, during the switching process between the main dialogue engine and the degraded dialogue engine, call data is transmitted through a unified call context object. The call context object records user information, task information, dialogue rounds, historical transcribed text, historical reply text, keyword hit status, intention level, transfer to human agent status, and current processing stage.
[0027] Once the main conversation engine is restored to availability, the current call context can be resynchronized to the main conversation engine.
[0028] Preferably, in step S2, performing an exact one-time write based on the idempotency key specifically involves generating a unique idempotency key idempotency_key for each call event, wherein the idempotency key is generated by combining the call identifier, event type, event time, and softswitch event number;
[0029] Using idempotent keys as the unique constraint field in the database, events that arrive for the first time are inserted into a new record, while events that arrive repeatedly are only updated with timestamps or ignored, without generating new business records.
[0030] Preferably, in step S1, the task scheduling engine supports two dialing modes: power and progressive.
[0031] In power mode, multiple outbound call requests are initiated simultaneously based on the maximum number of concurrent calls, trunk capacity, and task queue, making it suitable for high-concurrency notification and marketing outreach scenarios.
[0032] In progressive mode, the next call is initiated only after the previous call has been processed or agent resources have been released. This mode is suitable for scenarios where human agents handle calls and require refined communication.
[0033] The system limits the maximum number of concurrent calls using semaphores, and also sets a daily maximum retry limit for unanswered numbers.
[0034] Preferably, the task scheduling engine periodically monitors the registration status, call failure rate, response timeout count, and abnormal return codes of the SIP trunk;
[0035] When a relay anomaly is detected, a self-healing process is automatically executed, which includes pausing dialing, releasing the abnormal channel, resetting the gateway status, rescanning the relay, and reloading the configuration. After recovery, the dialing task is automatically restarted. If self-healing fails, an alarm is generated and pushed to the monitoring panel.
[0036] Preferably, in step S8, generating a structured call summary specifically involves: calling a summary generator to perform semantic analysis on the complete call transcript and outputting a structured result containing the call summary, detected intent, intent level, and recommended action;
[0037] The testing intent covers categories such as purchase, consultation, refusal, complaint, transfer to human agent, appointment, and contact later. The intent level is divided into four levels: high intent, medium intent, low intent, and no intent.
[0038] Preferably, the automatic transfer of customers to the follow-up pool specifically means: when a customer's final intention level is high or medium, or when a call is transferred to a human operator, the customer is automatically added to the follow-up pool;
[0039] Set high follow-up priority for high-interest customers and normal follow-up priority for medium-interest customers;
[0040] Customers with low or no interest can choose not to enter the follow-up pool or enter the low-priority follow-up pool according to business rules.
[0041] An intelligent outbound call dialogue system based on multimodal pipeline orchestration includes an external interaction layer, a core business layer, an AI capability layer, and a data storage layer.
[0042] The external interaction layer includes a SIP trunk gateway, a management console, and a real-time monitoring panel, which are used to access the public telephone network, configure outbound call tasks, and display the system's operating status, respectively.
[0043] The core business layer includes a call orchestration module, a task scheduling engine, an event processing service, and an audio bridging module.
[0044] The call orchestration module is used to uniformly schedule the entire lifecycle of a single call;
[0045] The task scheduling engine is used to control the batch dialing rhythm, concurrency limits, and relay self-healing.
[0046] The event handling service is used to receive call events and perform idempotent writes;
[0047] The audio bridging module is used to forward call media streams and AI service audio data;
[0048] The AI capability layer includes ASR service, main dialogue engine, degraded dialogue engine, TTS service, intent classifier and summary generator, which are used for speech transcription, main path dialogue generation, degraded path dialogue generation, speech synthesis, user intent classification and structured summary generation, respectively.
[0049] The data storage layer includes a business database, a recording file storage, a caching service, and a follow-up pool, which are used to store business data, call recordings, session context, and follow-up tasks, respectively.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration significantly reduces interaction latency and improves the naturalness of dialogue. This invention breaks the traditional fully serial processing mode of ASR, dialogue and TTS through sentence-level streaming segmentation and real-time TTS triggering mechanism. The three links are pipelined and coordinated at the sentence level, which can reduce end-to-end response latency by 40%-60%, greatly reduce user waiting time, and make human-computer dialogue closer to natural communication.
[0052] 2. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration provides dual-path fault tolerance and improves service continuity. Through the dual-path architecture of the main dialogue engine and the degraded dialogue engine, combined with seamless transmission of call context, the main service can automatically degrade and maintain the call when it is abnormal, avoiding call interruption caused by single point of failure, and greatly improving the system availability in high-concurrency and complex network environments.
[0053] 3. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration features a lightweight idempotent mechanism to ensure data consistency. Based on the idempotent key scheme with unique database constraints, it can achieve accurate one-time writing of call events without complex state machines, effectively solving the problems of data duplication and state errors caused by network retransmission and repeated event push. It has low implementation cost and convenient event type expansion.
[0054] 4. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration features dual-mode intelligent scheduling, which improves outbound call efficiency and adaptability. The power and progressive dual dialing modes can flexibly adapt to different business scenarios. Combined with concurrency control and retry limit mechanism, it balances dialing efficiency and system stability. The automatic self-healing capability of the trunk reduces the reliance on manual operation and maintenance, and ensures the continuous operation of batch tasks.
[0055] 5. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration has a closed-loop business process, which improves the business conversion value. After the call, a structured semantic summary is automatically generated. Combined with intention classification and automatic flow of the follow-up pool, the call results are directly converted into executable follow-up tasks. High-value customers can be identified and processed first, which effectively improves the conversion efficiency and operational value of outbound call business.
[0056] 6. This intelligent outbound call dialogue method and system based on multimodal pipeline orchestration features a modular and layered architecture with good scalability. The modules at each layer of the system are decoupled, and it supports the access of ASR, TTS, and dialogue engines from different vendors through adapters. It can flexibly adapt to the outbound call business needs of different industries and scales, and facilitates function iteration and system migration. Attached Figure Description
[0057] The present invention will be further described below with reference to the accompanying drawings and embodiments:
[0058] Figure 1 System overall architecture diagram;
[0059] Figure 2 Flowchart of intelligent outbound call dialogue method;
[0060] Figure 3 Schematic diagram of sentence-level pipelined flow and dual-path dialogue degradation mechanism;
[0061] Figure 4 A closed-loop diagram illustrating event idempotent writing, batch dialing scheduling, and post-call follow-up. Detailed Implementation
[0062] This section will describe in detail specific embodiments of the present invention. Preferred embodiments of the present invention are shown in the accompanying drawings. The purpose of the drawings is to supplement the textual description with graphics, so that people can intuitively and vividly understand each technical feature and overall technical solution of the present invention, but they should not be construed as limiting the scope of protection of the present invention.
[0063] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.
[0064] In the description of this invention, terms such as greater than, less than, and exceeding are understood to exclude the stated number, while terms such as above, below, and within are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.
[0065] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.
[0066] Please see Figure 1-4 This invention provides a technical solution: an intelligent outbound call dialogue method based on multimodal pipeline orchestration, comprising the following steps:
[0067] S1. Task Configuration and Startup: Operations personnel create outbound call tasks through the management console, and configure customer lists, dialing modes, maximum concurrency, opening remarks, script rules, maximum number of dialogue rounds, and follow-up rules.
[0068] The task scheduling engine periodically detects tasks to be executed, updates tasks that have reached their start time to running status, and initiates batch dialing according to the configured dialing mode.
[0069] The task scheduling engine supports two dialing modes: power and progressive.
[0070] Power mode: Suitable for high-concurrency scenarios such as notification reminders, marketing outreach, and satisfaction surveys. The system initiates multiple calls in batches based on the maximum concurrent call limit and trunk resource availability to maximize dialing throughput.
[0071] Progressive mode: Suitable for scenarios that require human agents to handle calls and provide detailed communication. The system follows the rule of "the next call is initiated only after the previous call is completed and resources are released" to control the call connection rhythm and avoid resource congestion.
[0072] The system also sets a daily maximum number of retries for numbers that are not connected, busy, or unanswered. Once the limit is reached, no more calls will be made that day to avoid excessively disturbing users.
[0073] S2. Outbound call initiation and idempotent event processing: The task scheduling engine controls the SIP trunk gateway to initiate outbound call requests to the target number and listens for call events such as ringing, connection, busy line, missed call, and failure.
[0074] The event processing service generates a unique idempotent key for each call event and performs an exact-once write based on database unique constraints to avoid duplicate records.
[0075] The idempotent key is generated by concatenating the call unique identifier, event type code, event occurrence timestamp, and softswitch event serial number to ensure the global uniqueness of each event;
[0076] The database creates a unique index for the idempotent key field in the event table. When a duplicate event arrives, the database triggers a unique constraint conflict. The system either updates the latest timestamp of the event or discards it directly, without generating a new business record.
[0077] S3. Call Establishment and Audio Bridging: When the called user answers, the audio bridging module establishes a two-way audio channel between the call media stream and the AI capability layer. The system plays an opening statement and sends the user's real-time voice stream to the ASR service for real-time transcription.
[0078] S4. Keyword detection and human intervention judgment: The call orchestration module receives the sentence-by-sentence recognition text output by ASR and first performs keyword detection and human intervention intention judgment.
[0079] If the preset keywords for transferring to a human agent are matched or the conditions for transferring to a human agent are met, the automatic dialogue process will be stopped, a prompt to transfer to a human agent will be played, and a transfer or follow-up action will be triggered.
[0080] If the request to switch to human assistance is not triggered, the dialogue generation process will begin.
[0081] S5, Dual-channel dialogue engine scheduling: The call orchestration module sends the user's transcribed text and call context to the main dialogue engine, while simultaneously detecting the availability status of the main dialogue engine.
[0082] If the main dialogue engine is functioning normally, it will generate the streaming reply text.
[0083] If the main dialogue engine interface times out, network error, or service unavailability is detected, it will automatically switch to the degraded dialogue engine and generate a response based on the local knowledge base and search-enhanced generation technology.
[0084] During the switching process between the main dialogue engine and the degraded dialogue engine, data transfer is achieved through a unified call context object;
[0085] The call context object stores basic user information, task attributes, dialogue history, keyword hit records, intent level, abnormal status and current processing stage for each call;
[0086] The degradation engine continues to generate responses based on the context, ensuring the continuity of multi-turn dialogues;
[0087] Once the main dialogue engine is restored, the context can be synchronized back to the main engine according to the strategy, restoring full-featured dialogue capabilities.
[0088] S6. Sentence-level streaming synthesis and playback: The call orchestration module receives the streaming response text from the dialogue engine and segments the text into fragments in real time according to sentence boundaries.
[0089] Each time a complete sentence unit is identified, it is immediately submitted to the TTS service for speech synthesis, and the synthesized speech is pushed to the call media stream for playback through the audio bridging module, realizing sentence-level pipeline collaboration between ASR, dialogue generation and TTS;
[0090] The specific rules for sentence-level streaming segmentation are as follows: punctuation marks such as periods, question marks, exclamation marks, semicolons, and newlines are used as sentence boundary markers. During the streaming output of the dialogue engine, each character is detected. Each time a complete semantic sentence is formed, a TTS synthesis is triggered, without waiting for the complete response to be generated.
[0091] Through this mechanism, the three stages of ASR transcription, dialogue text generation, and speech synthesis proceed in parallel at the sentence level, significantly reducing the perceived waiting delay for users.
[0092] S7. Intent update and loop judgment: After each round of dialogue, the intent classifier classifies the user input into intent levels and updates the intent state in the call context.
[0093] The system determines whether the maximum number of dialogue rounds has been reached, the user has hung up, the business objective has been achieved, or other termination conditions have been triggered. If none of these conditions are met, the system continues to execute the next dialogue round.
[0094] S8. Post-call processing and business loop: After the call ends, the system summarizes the complete transcribed text, call metadata and dialogue status, and calls the summary generator to generate a structured result containing call summary, detection intent, intent level and recommended action.
[0095] Determine whether to add customers to the follow-up pool and set priorities based on their level of interest and status of being transferred to human agents;
[0096] All call data and follow-up tasks are persistently stored, and various metrics on the real-time monitoring panel are updated simultaneously.
[0097] Furthermore, the system periodically performs health monitoring on SIP trunks, with monitoring indicators including trunk registration status, call failure rate, response timeout rate, percentage of abnormal return codes, and channel occupancy rate.
[0098] When a relay anomaly is detected, an automatic self-healing process is executed: first, new dialing requests are paused, and abnormal channel resources are released; then, gateway reset, relay rescanning, and configuration reloading are performed. After completion, dialing is attempted to be restored and the status is continuously monitored.
[0099] If the self-healing process fails multiple times, an alarm message will be generated and pushed to the operations and maintenance personnel.
[0100] Furthermore, the structured summary is generated by a large language model based on the complete call transcription. The call summary summarizes the core content of the call and the user's needs, detects intent to identify the user's core behavioral tendencies, and divides the intent level into four levels: high, medium, low, and none, based on the intensity of the user's attitude and needs. The recommended action provides specific suggestions for follow-up, such as a human callback, sending an SMS, making another outbound call, or ending the task.
[0101] Furthermore, the automatic follow-up pool rules are as follows: high-interest and medium-interest customers are automatically entered into the follow-up pool, with high-interest customers assigned the highest priority and followed up by agents first;
[0102] Customers who trigger a transfer to a live agent during a call are immediately placed in a high-priority follow-up pool.
[0103] Customers with low or no interest can choose not to enter the follow-up pool, or enter the low-priority pool for further nurturing and outreach, according to business rules.
[0104] On the other hand, this invention provides an intelligent outbound call dialogue system based on multimodal pipeline orchestration, which adopts a layered and decoupled architecture, including an external interaction layer, a core business layer, an AI capability layer, and a data storage layer.
[0105] The external interaction layer includes a SIP trunk gateway, a management console, and a real-time monitoring panel. The SIP trunk gateway is responsible for connecting to the public switched telephone network and completing outbound call signaling interaction, call maintenance, and call termination.
[0106] The management console provides a visual interface that supports outbound call task configuration, customer list import, call script editing, call record query, and follow-up pool management.
[0107] The real-time monitoring panel displays task progress, active call count, connection rate, failure rate, relay status, and abnormal alarms through real-time communication protocols.
[0108] The core business layer includes a call orchestration module, a task scheduling engine, an event processing service, and an audio bridging module;
[0109] As the core scheduling unit of the system, the call orchestration module manages the entire lifecycle of a single call from initiation, connection, multi-round dialogue, exception handling to post-hang-up processing, and uniformly schedules various components of the AI capability layer and maintains the call context.
[0110] The task scheduling engine is responsible for the status transition of batch outbound call tasks, dialing rhythm control, concurrency limit, dual-mode dialing management, and trunk health monitoring and self-healing.
[0111] The event processing service is responsible for receiving various call events pushed by the softswitch platform, generating idempotent keys and performing persistent writing to ensure the consistency of event processing;
[0112] The audio bridging module is responsible for bidirectional forwarding between the call media stream and the AI service audio stream, enabling audio path support for real-time voice interaction.
[0113] The AI capability layer includes ASR service, main dialogue engine, degraded dialogue engine, TTS service, intent classifier and summary generator;
[0114] The ASR service provides real-time speech-to-text capabilities, outputting time-stamped, sentence-by-sentence recognized text;
[0115] The main dialogue engine serves as the main dialogue path, integrating a business knowledge base and multi-turn dialogue management capabilities to generate natural and fluent business responses.
[0116] The fallback dialogue engine serves as a backup dialogue path, consisting of a local lightweight large model, a rule-based dialogue library, and a RAG retrieval module, providing basic responses when the main engine is unavailable.
[0117] The TTS service provides text-to-speech capabilities, supporting streaming synthesis and real-time playback;
[0118] The intent classifier is responsible for identifying user intent and determining intent level in each round of dialogue;
[0119] The summary generator is responsible for the full semantic analysis of the text after the call ends and outputs a structured summary result;
[0120] The data storage layer includes a business database, audio file storage, caching services, and a follow-up pool;
[0121] The business database stores structured data such as customer information, task configurations, call records, event logs, and summary results;
[0122] Recording file storage: Saves call recording files;
[0123] The caching service stores session context, task status, active call information, and temporary monitoring metrics to improve access efficiency.
[0124] The follow-up pool stores information on customers to be followed up, their priorities, and recommended actions to support subsequent business processes.
[0125] Example 1:
[0126] This embodiment provides an intelligent outbound call dialogue system based on multimodal pipeline orchestration. It adopts a layered modular architecture and is designed for intelligent outbound call scenarios such as telephone sales, customer follow-up, government notifications, financial collection, and medical follow-up, to achieve low latency, high reliability, and fully closed-loop intelligent outbound call dialogue services.
[0127] System overall architecture;
[0128] like Figure 1 As shown, the system is divided into a data storage layer, an AI capability layer, a core business layer, and an external interaction layer from bottom to top. Each layer interacts through standardized interfaces, and the modules are decoupled to facilitate independent iteration and replacement.
[0129] The external interaction layer serves as the entry point for the system to interact with the external environment and operational personnel, and comprises three core components:
[0130] SIP trunk gateways use the standard SIP protocol to connect to operator lines or IP voice gateways, and support signaling establishment, media negotiation, call maintenance and disconnection release for bulk outbound calls.
[0131] It can connect to multiple trunk lines simultaneously, supporting line load balancing and fault switching.
[0132] The management console is a web-based visual management backend that provides functions such as task management, list management, script configuration, record query, follow-up pool management, and system settings.
[0133] Operations personnel can upload customer lists, set outbound call times, select dialing modes, configure opening remarks and business scripts, and set follow-up rules through the console.
[0134] The real-time monitoring panel establishes a long connection with the backend service via WebSocket to display key indicators such as the number of active calls, task completion progress, connection rate, call failure rate, number of calls transferred to human operators, new additions to the follow-up pool, and trunk line status in real time.
[0135] It supports displaying abnormal alarm pop-ups, making it easy for operations and maintenance personnel to monitor the system's operating status in real time.
[0136] The core business layer is the system's scheduling and control center, responsible for processing the business logic of the entire outbound call process, and includes four core modules:
[0137] The call orchestration module, as the core scheduling unit for a single call, is responsible for the state management and component scheduling of the entire lifecycle of the call.
[0138] Each call corresponds to a call orchestration instance, which is responsible for triggering ASR recognition, calling the dialogue engine, controlling TTS playback, maintaining call context, handling human-to-human transfer events, and triggering post-processing procedures.
[0139] This module adopts an event-driven architecture, receiving various messages such as audio, text, and events through a message queue, and scheduling various AI capability components according to business logic order.
[0140] The task scheduling engine is responsible for the full lifecycle management of batch outbound call tasks, including task start, pause, resume, and termination, as well as dialing rhythm control, concurrency limit, dialing mode switching, trunk health monitoring and self-healing.
[0141] The dialing dispatch supports both power and progressive modes, and uses distributed semaphores to control the global maximum number of concurrent calls, thus avoiding overload of trunk and AI services.
[0142] Retry management records the number of retries for each number in a day. Numbers that are not connected, busy, or unanswered are retried at the configured intervals. Once the daily maximum retry limit is reached, dialing is stopped for the day, and the count is automatically reset the next day.
[0143] The relay is self-healing and has a built-in health detection thread that checks the registration status and call success rate of each relay line every 10 seconds.
[0144] When the failure rate of three consecutive detections exceeds the threshold, the self-healing process is automatically triggered, including suspending dialing on the line, performing gateway registration reset, reloading line configuration, initiating a test call for verification, and resuming dialing after successful verification.
[0145] If the self-healing fails three times in a row, the line will be marked as faulty and an alarm will be sent.
[0146] The event handling service is specifically designed to receive call events pushed by the softswitch platform, including call initiation, ringing, connection, user hang-up, agent transfer, call end, recording completion, transcription completion, etc.
[0147] Event reception adopts an asynchronous message queue processing mode to avoid system blockage caused by peak event push periods;
[0148] For each event, the service generates an idempotent key according to preset rules, and then performs a database write operation, using the database unique constraint to achieve idempotency.
[0149] The audio bridging module is responsible for bidirectional real-time forwarding between the call media stream and the AI service audio stream;
[0150] After the call is connected, the module establishes an independent audio processing channel for the call, forwards the RTP audio stream from the user side to the ASR service, and pushes the TTS synthesized audio stream to the other end of the call.
[0151] The module supports audio format conversion, volume adjustment, and silence detection, ensuring the sound quality and smoothness of voice interaction.
[0152] The AI capability layer provides the system with various intelligent voice and language processing capabilities, comprising six core components:
[0153] The ASR service uses a real-time speech recognition engine, supports streaming audio input, and outputs time-stamped, sentence-by-sentence recognized text.
[0154] Intermediate results are returned in real time during the recognition process, and a stable full sentence recognition result is finally output. It supports optimized recognition of dialects, numbers and proper nouns.
[0155] The main dialogue engine is built using a cloud-based large language model combined with retrieval-enhanced generation technology. It has a built-in business knowledge base and multi-turn dialogue management capabilities, and can understand user intent, maintain dialogue context, and generate natural language responses that meet business requirements.
[0156] The main engine supports streaming output, with text continuously returned at the granularity of characters or tokens, supporting sentence-level streaming processing;
[0157] The dialogue engine is downgraded, and a lightweight large language model is deployed locally, combined with a rule-based speech library and a local RAG retrieval module, without relying on an external network;
[0158] When the main engine is unavailable, the fallback engine generates basic business responses based on the transmitted call context and the local knowledge base, ensuring that the call is not interrupted and core business issues can be answered normally.
[0159] The TTS service supports streaming text input and speech synthesis. It can receive sentence-level text segments and quickly synthesize speech with a synthesis latency controlled within the hundreds of milliseconds.
[0160] It supports multiple timbre, speech rate and volume configurations, and can select the appropriate pronunciation style according to different business scenarios;
[0161] The intent classifier, based on a text classification model, performs intent recognition and intent level judgment on each round of user input. It supports multiple classifications such as purchase intent, consultation intent, rejection intent, complaint intent, and transfer to human agent intent, and outputs four intent levels: high, medium, low, and none, which are used to update the call context.
[0162] The summary generator, based on the summarizing capabilities of a large language model, takes the complete call transcription text and call metadata as input and outputs structured analysis results, including the following fields:
[0163] Call summary: A 100-300 word summary of the core content of the call;
[0164] Detection intent: User's core intent tags, multiple selections are allowed;
[0165] Intent level: High / Medium / Low / None (four-level rating);
[0166] Recommended actions: Follow-up suggestions, such as prioritizing manual callback, sending product text messages, calling out again after 3 days, and marking the task as unintended termination.
[0167] The data storage layer is responsible for the persistence and caching of all system data, and includes four types of storage:
[0168] The business database uses a relational database (such as MySQL / PostgreSQL) to store customer information tables, outbound call task tables, call record tables, call event tables, transcription segmentation tables, summary result tables, follow-up pool tables, system configuration tables, etc.
[0169] The call event table uses a unique index on the idempotency_key field to enable idempotency writing of events;
[0170] For recording file storage, object storage services (such as MinIO and OSS) are used to save call recording files. The file names are associated with a unique call identifier and online playback and download are supported.
[0171] The caching service uses Redis as the caching middleware to store active call context, task status, concurrent semaphores, temporary monitoring metrics, hotspot configuration data, etc., thereby improving system response speed and reducing database pressure.
[0172] The follow-up pool is implemented in the form of a database table, storing fields such as customer information to be followed up, associated call ID, intention level, priority, recommended action, assigned agent, and follow-up status. It supports agents to retrieve, backfill, and update follow-up records.
[0173] Method execution flow;
[0174] like Figure 2 As shown in the figure, the intelligent outbound call dialogue method based on multimodal pipeline orchestration described in this embodiment has the following specific execution steps:
[0175] S1. Task creation and startup;
[0176] Operations personnel log in to the management console, create new outbound call tasks, upload customer number lists, and set task names, outbound call time periods, dialing modes, maximum concurrent calls, daily single-number retry limit, opening remarks, business script rules, maximum number of conversation rounds, and follow-up pool allocation rules.
[0177] After configuration, save the task; the task status will be "Pending Start".
[0178] The task scheduling engine polls the task list at fixed intervals. When the start time of a task is reached, the task status is updated to "running", the customer number queue for that task is loaded, and dialing scheduling begins.
[0179] S2, Batch dialing and event idempotent processing;
[0180] The task scheduling engine controls the dialing rhythm based on the selected dialing mode:
[0181] If power mode is selected: the system retrieves the corresponding number of numbers from the number queue in batches according to the current remaining concurrent quota, and simultaneously sends outbound call requests to the SIP trunk gateway;
[0182] Each initiated call occupies one concurrent semaphore, and the semaphore is released after the call ends;
[0183] If progressive mode is selected: the system waits for the existing call to end and release resources before taking the next number from the queue and initiating a call;
[0184] If human agents are available, the dialing speed will be controlled based on the number of available agents.
[0185] After a SIP trunk gateway initiates a call, it pushes various events such as call initiation, ringing, connection, busy line, missed call, failure, and hang-up to the event processing service through a message queue.
[0186] After receiving an event, the event handling service performs idempotent processing:
[0187] Extract the call unique identifier (call_id), event type (event_type), event timestamp (event_ts), and softswitch event number (event_no) from the event;
[0188] Generate an idempotency key by concatenating the following format: "call_id:event_type:event_ts:event_no";
[0189] Perform a database INSERT operation. If the insertion is successful, it indicates that this is the first event, and update the status of the corresponding call record.
[0190] If insertion fails due to a unique constraint conflict, it is determined to be a duplicate event. Only the latest reception time of the event is updated or it is discarded directly, and no new business record is generated.
[0191] S3. Call connection and voice interaction initialization;
[0192] Once the event handling service receives the "connection" event and successfully writes it, it notifies the call orchestration module to start the call interaction process.
[0193] The audio bridging module establishes a two-way audio channel for the call, and the system first plays a preset opening speech.
[0194] After the opening remarks are played, the audio bridging module begins forwarding the real-time audio stream from the user side to the ASR service and starts speech recognition.
[0195] S4, ASR transcription and manual judgment;
[0196] The ASR service performs streaming recognition on real-time voice streams. After each complete sentence is recognized, the timestamped recognized text is pushed to the call orchestration module.
[0197] After receiving each sentence of recognized text, the call orchestration module first performs keyword matching detection, matching the preset keyword library for transferring to human agents, including "transfer to human agent", "find customer service", "human service", "need a real person" and "do you have anyone available";
[0198] If the keyword for transferring to a human agent is hit, or if the system identifies the user as requesting human assistance for two consecutive rounds, the current dialogue generation process will be immediately stopped, and a prompt message will be played saying "Transferring you to a human agent, please wait" will be played, while the agent transfer process will be triggered.
[0199] If there are no available agents, mark the customer as pending a callback and add them to the high-priority follow-up pool;
[0200] The entire process of transferring the call to a human operator is simultaneously written to the event log and call record.
[0201] If the keyword for switching to human intervention is not matched, the dialogue generation process will begin.
[0202] S5, dual-channel dialogue engine scheduling and streaming response;
[0203] The call orchestration module assembles the current user-identified text and call context object into a dialogue request and sends it to the main dialogue engine, while simultaneously starting a timeout detection timer;
[0204] If the main dialogue engine responds normally within the timeout period and starts returning streaming text, then the main dialogue engine will provide the dialogue service.
[0205] If the request times out, returns an error code, or the network connection fails, the main dialogue engine is deemed unavailable, and the system automatically switches to the degraded dialogue engine. The same user text and call context are sent to the degraded engine, which then generates the response text.
[0206] During engine switching, the call context object is completely passed, which includes: user name, customer tag, task type, historical dialogue rounds, confirmed business information, current intent level, and matched keywords.
[0207] The degradation engine continues to generate responses based on historical context, avoiding issues such as duplicate questions and contradictions.
[0208] If the main dialogue engine is subsequently detected to have recovered, you can choose to continue using the downgraded engine until the call ends, or smoothly switch back to the main engine after the current round ends, depending on the configuration.
[0209] S6, sentence-level streaming segmentation and real-time TTS playback;
[0210] like Figure 3 As shown, the call orchestration module receives the streaming text returned by the dialogue engine and initiates the sentence boundary detection logic:
[0211] Receives streaming text character by character and caches text fragments in real time;
[0212] Detect whether sentence end markers such as periods, question marks, exclamation marks, semicolons, and line breaks appear in the cached text;
[0213] Once the sentence end marker is detected, the complete sentence is immediately extracted and submitted to the TTS service for speech synthesis. At the same time, the cache for that part is cleared and subsequent text is received.
[0214] Once the TTS synthesis is complete, the voice stream is immediately pushed to the call line for playback via the audio bridging module.
[0215] Through this mechanism, while the dialogue engine is generating the second half of the response, the first half of the generated sentence has already been synthesized and played to the user. The three stages of ASR transcription, dialogue generation, and TTS synthesis are processed in a pipeline-like parallel manner.
[0216] Compared to the traditional serial mode of "sentence generation - complete synthesis - overall playback", this mechanism can significantly reduce end-to-end response latency, significantly reduce silences and pauses during calls, and improve the naturalness of the conversation.
[0217] S7, Intention Update and Dialogue Loop;
[0218] After each round of dialogue is completed, the intent classifier performs intent recognition and intent level judgment on the user input in this round, and updates the results to the call context object;
[0219] The system determines whether the call termination conditions are met. Termination conditions include: the user hangs up voluntarily, the preset maximum number of dialogue rounds is reached, business objectives are achieved (such as completing notifications or information confirmations), the user explicitly indicates that the call is to end, and the call is transferred to a human agent.
[0220] If the termination condition is not met, return to step 4, continue listening to the user's voice, and start the next round of dialogue.
[0221] S8. Call termination and business loop processing;
[0222] Upon detecting a call end event, the call orchestration module triggers the post-call processing flow:
[0223] Summarize the complete transcript of this call, call metadata (duration, connection time, hang-up time, whether to transfer to a human operator, etc.), dialogue history, and records of changes in intent;
[0224] Call the summary generator, input the above data, and generate structured call analysis results, including call summary, core intent, final intent level, and recommended follow-up actions;
[0225] The follow-up pool will be executed based on the final intention level and the status of transitioning to manual intervention:
[0226] If the final intention is high, or if the call is transferred to a human agent but not connected, the call will be automatically added to the follow-up pool with a priority of "high" and assigned to the corresponding agent group for priority handling.
[0227] Those whose final intention is "moderate" are automatically added to the follow-up pool and their priority is set to "normal".
[0228] If the final intention is low or no intention, depending on the task configuration, choose not to enter the follow-up pool, or enter the low-priority nurturing pool and arrange to reach them again later.
[0229] All data, including call logs, transcription segments, event logs, summary results, and follow-up tasks, are persisted to the business database, while audio recordings are uploaded to object storage.
[0230] Meanwhile, the real-time monitoring panel updates the task's completion progress, connection rate, failure rate, number of people transferred to human assistants, and number of new additions to the follow-up pool. If any abnormal indicators are detected, an alarm will be triggered.
[0231] Event and scheduling closed-loop description;
[0232] like Figure 4 As shown, the left side is the batch dialing scheduling and trunk self-healing mechanism, the middle side is the event idempotent processing mechanism, and the right side is the call follow-up pool transfer mechanism. Together, these three constitute a closed loop of system reliability, efficiency, and business value.
[0233] The idempotency mechanism ensures the accuracy of the underlying data, providing a reliable data foundation for scheduling decisions and business analysis;
[0234] Dual-mode scheduling and relay self-healing ensure the efficient and stable operation of batch outbound call tasks;
[0235] The automatic follow-up pool transforms call results into business actions, achieving a closed-loop process from outbound calls to follow-ups.
[0236] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those skilled in the art, various changes can be made without departing from the spirit of the present invention.
Claims
1. An intelligent outbound call dialog method based on a multi-modal pipeline orchestration, characterized in that, Includes the following steps: S1. Create an outbound call task and configure the task parameters. The task scheduling engine starts the batch dialing process according to the task configuration. S2. Initiate an outbound call request through the SIP trunk gateway, receive call events, and perform an exact one-time write of the event based on the idempotent key; S3. After the call is connected, a two-way audio channel is established, and the user's voice stream is sent to the ASR service in real time for sentence-by-sentence transcription. S4. Perform keyword detection on the ASR transcribed text. If the keyword trigger condition for manual transfer is met, execute the manual transfer process. If the call is not successful, the transcribed text and the call context will be sent to the dialogue engine to generate a response. S5. Detect the availability of the main dialogue engine. If the main dialogue engine is available, generate a streaming response text using the main dialogue engine. If the primary dialogue engine is unavailable, it will automatically switch to the degraded dialogue engine to generate the response text. S6. The convection reply text is segmented in real time according to the sentence boundaries. Each time a complete sentence unit is generated, TTS speech synthesis is triggered immediately and the synthesized speech is pushed to the call media stream for playback. S7. After each round of dialogue, update the user's intention level and call context, and determine whether the call termination conditions are met. If not, continue to the next round of dialogue. S8. After the call ends, a structured call summary is generated. Customers are automatically transferred to the follow-up pool based on their intent level and call events, and the end-to-end monitoring data is updated.
2. The intelligent outbound conversation method based on multi-modal pipeline orchestration according to claim 1, characterized in that: In step S6, the real-time segmentation of the streaming response text according to sentence boundaries is specifically as follows: using periods, question marks, exclamation marks, semicolons, and line breaks as sentence boundary identification markers, each time a complete sentence or playable semantic segment is identified during the streaming output of the dialogue engine, the text segment is immediately submitted to the TTS service for speech synthesis, without waiting for the complete response to be generated. ASR transcription, dialogue generation, and TTS synthesis are performed in a pipeline-like parallel collaboration at the sentence level. 3.The intelligent outbound dialog method based on multi-modal pipeline orchestration of claim 1, wherein: In step S5, the main dialogue engine adopts a knowledge-based dialogue agent or a cloud-based large language model service to carry the business knowledge base, multi-turn context management and natural language response generation. The degraded dialogue engine uses a local large language model combined with a rule-based dialogue library and a retrieval enhancement generation module to provide basic business response capabilities when the main dialogue engine experiences interface timeouts, network anomalies, service unavailability, or knowledge base access failures.
4. The intelligent outbound dialog method based on multi-modal pipeline orchestration according to claim 3, characterized in that: During the switching process between the main dialogue engine and the degraded dialogue engine, call data is transmitted through a unified call context object. The call context object records user information, task information, dialogue rounds, historical transcribed text, historical reply text, keyword hit status, intention level, transfer to human agent status, and current processing stage. Once the main conversation engine is restored to availability, the current call context can be resynchronized to the main conversation engine.
5. The intelligent outbound conversation method based on multi-modal pipeline orchestration according to claim 1, characterized in that: In step S2, the precise one-time write of the event based on the idempotency key is specifically as follows: a unique idempotency key idempotency_key is generated for each call event, and the idempotency key is generated by combining the call identifier, event type, event time, and softswitch event number; Using idempotent keys as the unique constraint field in the database, events that arrive for the first time are inserted into a new record, while events that arrive repeatedly are only updated with timestamps or ignored, without generating new business records.
6. The intelligent outbound dialog method based on multi-modal pipeline orchestration according to claim 1, characterized in that: In step S1, the task scheduling engine supports two dialing modes: power and progressive. In power mode, multiple outbound call requests are initiated simultaneously based on the maximum number of concurrent calls, trunk capacity, and task queue, making it suitable for high-concurrency notification and marketing outreach scenarios. In progressive mode, the next call is initiated only after the previous call has been processed or agent resources have been released. This mode is suitable for scenarios where human agents handle calls and require refined communication. The system limits the maximum number of concurrent calls using semaphores, and also sets a daily maximum retry limit for unanswered numbers.
7. The intelligent outbound dialog method based on multi-modal pipeline orchestration according to claim 6, characterized in that: The task scheduling engine periodically monitors the registration status, call failure rate, response timeout count, and abnormal return codes of the SIP trunk. When a relay anomaly is detected, a self-healing process is automatically executed, which includes pausing dialing, releasing the abnormal channel, resetting the gateway status, rescanning the relay, and reloading the configuration. After recovery, the dialing task is automatically restarted. If self-healing fails, an alarm will be generated and pushed to the monitoring panel. 8.The intelligent outbound dialog method and system based on multi-modal pipeline scheduling of claim 1, wherein: In step S8, generating a structured call summary specifically involves calling a summary generator to perform semantic analysis on the complete call transcript and outputting a structured result containing the call summary, detection intent, intent level, and recommended action. The testing intent covers categories such as purchase, consultation, refusal, complaint, transfer to human agent, appointment, and contact later. The intent level is divided into four levels: high intent, medium intent, low intent, and no intent. 9.The intelligent outbound dialog method based on multi-modal pipeline orchestration of claim 8, wherein: The automatic transfer of customers to the follow-up pool specifically means: when a customer's final intention level is high or medium, or when a call is transferred to a human agent, the customer is automatically added to the follow-up pool; Set high follow-up priority for high-interest customers and normal follow-up priority for medium-interest customers; Customers with low or no interest can choose not to enter the follow-up pool or enter the low-priority follow-up pool according to business rules.
10. An intelligent outbound call dialogue system based on multimodal pipeline orchestration, characterized in that, It includes an external interaction layer, a core business layer, an AI capability layer, and a data storage layer; The external interaction layer includes a SIP trunk gateway, a management console, and a real-time monitoring panel, which are used to access the public telephone network, configure outbound call tasks, and display the system's operating status, respectively. The core business layer includes a call orchestration module, a task scheduling engine, an event processing service, and an audio bridging module. The call orchestration module is used to uniformly schedule the entire lifecycle of a single call; The task scheduling engine is used to control the batch dialing rhythm, concurrency limits, and relay self-healing. The event handling service is used to receive call events and perform idempotent writes; The audio bridging module is used to forward call media streams and AI service audio data; The AI capability layer includes ASR service, main dialogue engine, degraded dialogue engine, TTS service, intent classifier and summary generator, which are used for speech transcription, main path dialogue generation, degraded path dialogue generation, speech synthesis, user intent classification and structured summary generation, respectively. The data storage layer includes a business database, a recording file storage, a caching service, and a follow-up pool, which are used to store business data, call recordings, session context, and follow-up tasks, respectively.