An AI outbound calling system and its control method and medium

By building a physical link construction module and an interaction status recognition module in the AI ​​outbound calling system, the problems of cross-terminal hardware compatibility and human intervention delay were solved, achieving time consistency and hardware compatibility of cross-terminal audio data transmission, and improving the system's universality and response efficiency.

CN122093504APending Publication Date: 2026-05-26SHENZHEN DAZHI SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610526154.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-21
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The existing AI outbound calling system lacks a universal conversion interface for hardware collaboration mechanisms across terminals, resulting in poor system versatility and high deployment costs. The voice stream transmission lacks a physical link construction mechanism, and there are significant delays in the process of manual monitoring and intervention. Furthermore, the system architecture, which relies on the softswitch platform, lacks compatibility.

Method used

By constructing a physical link building module to perform signal gain compensation and impedance matching, missing level points in the signal transmission path are identified and filled in, generating a synchronous and continuous uplink and downlink voice signal sequence; combined with the interaction state recognition module, audio segments with abrupt energy changes are extracted, abnormal segments of routing switching frequency are identified and screened, the trend of changes in manual monitoring signals is tracked, the location of manual intervention is located, and a synchronization flag number for manual intervention response is generated.

Benefits of technology

It improves the time consistency and hardware compatibility of cross-terminal audio data transmission, identifies and processes the distribution structure of voice flow paths, establishes a control lag comparison relationship across hardware terminals, and outputs AI outbound call interaction recognition results with routing partitions, switching boundaries and response features, reducing human intervention delays.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122093504A_ABST
    Figure CN122093504A_ABST
Patent Text Reader

Abstract

This invention relates to the field of voice interface control technology, specifically to an AI outbound calling system and its control method and medium, comprising modules for physical link construction, interaction state recognition, flow path control, and intervention and collaborative response. By smoothly connecting analog signal gaps through impedance matching and gain compensation, a synchronous and continuous uplink and downlink voice signal sequence is generated; audio energy jumps are identified to pinpoint call state transition periods; frequency distribution characteristics of AI and human mode switching are analyzed to extract abnormal frequency segments of routing switching; and response synchronization flags are located by combining human intervention trend positioning to ultimately determine the core interactive control result. This invention improves the time consistency and compatibility of audio transmission across hardware terminals, and achieves efficient collaborative response and closed-loop control of interaction states between AI and human modes by accurately identifying flow paths and intervention response locations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice interface control technology, and in particular to an AI outbound calling system and its control method and medium. Background Technology

[0002] The field of voice interface control technology encompasses research directions such as voice acquisition interface connection, voice signal transmission path control, and call audio encoding and decoding processing. Its core content involves acquiring, converting, and routing multi-source audio signals during a call to reflect the voice flow and system coordination status under different modes, including human-computer interaction and manual intervention. It broadly involves analog voice signal processing, hardware-level sound card conversion, dynamic physical link switching, and human-computer voice interaction control mechanisms. Relying on outbound call terminals, chip-based processing modules, and real-time communication technology, it conducts multi-channel real-time processing of outbound call voice streams, AI-generated voice, and manually monitored voice, and manages the interaction process using corresponding switching strategies. This field emphasizes hardware compatibility, real-time audio transmission, and link switching stability, serving multiple scenarios such as customer follow-up, marketing notification, fault warning, and intelligent customer service.

[0003] The AI ​​outbound calling system and control method refers to a specific technical solution that establishes a physical hardware connection between the outbound calling mobile phone, AI processing device, and voice box, and combines this with signal flow logic for voice interaction. The technical aspects addressed include using the outbound calling mobile phone to perform dialing and call functions, using a chip-equipped voice box to complete the sound card conversion and return of analog voice signals, using headphones to achieve real-time monitoring of call content, and dynamically controlling the audio transmission path during the interaction process through analysis and switching logic based on an AI model system. A search reveals that patent CN115866141A discloses a human-machine coupled call flow control scheme, using deep learning to predict outbound call concurrency to dynamically adjust flow allocation; patent CN111131638A discloses an outbound calling method based on a softswitch platform and AI connector, improving service access capabilities through resource allocation. Generally, a flow scheduling method based on the backend logic layer or a signaling control method based on a VoIP architecture is used to complete the access and logical layer management of outbound calling services.

[0004] Existing hardware collaboration mechanisms lack universal conversion interfaces across terminals. Different brands and models of mobile phones typically require authorization and corresponding customized solutions to access the AI ​​calling function, resulting in poor system versatility and high deployment costs. In terms of voice stream transmission, algorithm-based flow control schemes lack physical link construction mechanisms. Especially in instantaneous switching scenarios requiring manual intervention, it is difficult to achieve hardware-level acquisition of real-time call audio from outgoing mobile phones and instantaneous flow of physical transmission paths, leading to significant delays in manual monitoring and intervention. Furthermore, the system architecture relying on softswitch platforms is too complex and lacks compatibility with specific outgoing calling hardware (such as mobile terminals). It cannot utilize simple hardware modules to achieve real-time conversion of analog voice streams, resulting in poor collaboration and delayed interactive responses among multiple devices. Summary of the Invention

[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing an AI outbound calling system and its control method and medium.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: an AI outbound calling system, the system comprising:

[0007] The physical link construction module acquires analog voice signals from the outbound terminal and the chip-equipped voice box, identifies level gaps in the signal transmission path, smoothly connects the two ends of the signal by adjusting the gain compensation and impedance matching of the physical link, and performs bidirectional stream path alignment after arranging them in time order to generate a synchronous and continuous uplink and downlink voice signal sequence.

[0008] The interaction state recognition module observes whether the pitch frequency rises continuously based on the audio content in the synchronous and continuous uplink and downlink voice signal sequence, identifies audio segments with energy jump changes, extracts the time index of the jump segment as the target segment for interaction state analysis, and generates the call state transition time period number.

[0009] The flow path control module checks whether the voice segments marked by the call state transition time period number fluctuate repeatedly between AI and human modes, extracts the interval path between each mode switch, determines whether the switching actions are concentrated, completes the filtering and classification of switching frequencies, and generates the abnormal segment number of the routing switching frequency.

[0010] The intervention and coordination response module calls the monitoring audio sequence corresponding to the abnormal segment number of the route switching frequency, tracks the change trend of the manual monitoring signal, finds the location of the first manual intervention after the AI ​​voice ends, uses the access point as a response reference, and if a delay in the switching action is found, executes the control command replacement operation and generates a manual intervention response synchronization flag number.

[0011] As a further aspect of the present invention, the synchronous and continuous uplink and downlink voice signal sequence includes the time correspondence of audio signals, signal gain compensation consistency, and transmission link continuity and integrity; the call state transition time period number includes the frequency of voice change occurrence, the distribution range of transition time periods, and the transition segment identifier number; the route switching frequency abnormal segment number includes the switching interval distribution density, the mode proportion level, and the switching abnormal segment location identifier; and the manual intervention response synchronization flag number includes the intervention response delay identifier, the signal flow trend offset, and the replacement target control interval.

[0012] As a further aspect of the present invention, the physical link construction module includes:

[0013] The physical access submodule acquires the analog voice signals collected by the outbound terminal and the voice box, extracts the time points and level amplitudes of each hardware channel according to the timestamp sequence, processes them according to time progression after unifying the sampling frequency, inserts fitting values ​​consistent with the signal trends before and after the physical link disconnection position, so that the audio has a consistent time distribution and waveform continuity, and obtains a signal sequence group with link alignment.

[0014] The impedance matching and completion submodule extracts signal segments with discontinuous level spans based on the signal sequence group aligned with the link, performs amplitude trend analysis on the continuous waveforms on both sides of the breakpoint, fills in matching segments according to the impedance change direction of adjacent segments, and embeds continuous waveform point sequences in each signal blank area to obtain a continuous and complete voice channel dataset.

[0015] The bidirectional flow alignment submodule calls up the uplink and downlink voice signals in the continuous and complete voice channel dataset, locates each time point according to the uplink data time sequence, selects adjacent time periods in the downlink data, fills in the corresponding values ​​according to the energy trend of the adjacent segments, and performs corresponding processing on the two types of voice streams at all time points to obtain a synchronous and continuous uplink and downlink voice signal sequence.

[0016] As a further aspect of the present invention, the interaction state recognition module includes:

[0017] The time frame division submodule obtains the audio content in the synchronous continuous uplink and downlink voice signal sequence, extracts continuous audio segments according to the signal timestamp, indexes and locates the starting position in each continuous data segment, extracts the audio amplitude sequence between adjacent time frames, and numbers the segments that do not have silence interruptions to obtain a continuous voice time segment number list.

[0018] The feature change recognition submodule calls each numbered segment in the continuous speech time period number list, extracts all audio amplitude points within it, compares the amplitudes of adjacent points in all time steps in sequence, determines whether there is a continuous rise in sound energy based on the direction of change between adjacent amplitudes, classifies and labels all segments with continuous amplitude increment features, and obtains a continuous rising tone trend segment index set.

[0019] The energy jump extraction submodule extracts signal intervals with short-term surges in amplitude changes in each segment based on the continuous pitch-rising trend segment index set. It extracts the amplitude difference between the time points before and after each jump point and performs amplification judgment. Among all jump points, it filters signal positions where the energy growth exceeds the set audio jump threshold and extracts their corresponding numbers in chronological order to obtain the call state transition time period number.

[0020] As a further aspect of the present invention, the flow path control module includes:

[0021] The flow structure recognition submodule obtains the voice segments marked in the call state transition time period sequence number, extracts the time sequence and flow trend corresponding to each voice sequence, examines whether the alternation order of AI-generated voice and human-monitored voice appears continuously, extracts adjacent position segments between control power transition points, and uses them to subsequently define the routing switching mode to obtain the flow alternation segment index set.

[0022] The switching interval extraction submodule extracts the mode switching interval segments in each segment according to the time sequence based on the flow alternation segment index set. It uses the distance between adjacent route turning points as the switching interval basis, and performs a classification operation on each switching interval segment in the current sequence in sequence. It is then classified into a unified index according to the relative time range to obtain the mode switching interval distribution interval set.

[0023] The path proportion judgment submodule calls the time index sequence of each segment in the mode switching interval distribution set, compares the change in the proportion of switching times in the segment length within the same time period, filters the segment numbers whose proportion value is higher than the set switching density threshold, and summarizes all segment numbers that meet the judgment conditions to obtain the sequence number of abnormal route switching frequency segment.

[0024] As a further aspect of the present invention, the intervention and collaborative response module includes:

[0025] The monitoring trend extraction submodule obtains the time segment corresponding to the abnormal segment number of the routing switching frequency, extracts the monitoring sequence data in each time segment, reads all audio energy values ​​in chronological order, continuously judges the direction of level change at adjacent time points, sorts out the time periods of continuous increase or continuous decrease, and obtains a set of monitoring signal change trend sequences.

[0026] The access point positioning submodule, based on the monitoring signal change trend sequence set, calls the end time point of each AI voice segment as the time reference point, scans the intervention direction backward in the corresponding monitoring sequence, and when the energy shift first appears in the human intervention trend, extracts the time point and corresponding level value at that position to construct the subsequent synchronization reference segment and obtains the human intervention time index table.

[0027] The response synchronization annotation submodule, based on the manual access time index table, compares the end time of each AI voice segment with the interval length between the start time of manual intervention and the end time of AI voice in each segment, extracts the sequence number of the time interval that is higher than the intervention delay threshold, and performs a logical reset operation on the corresponding segment to obtain the manual intervention response synchronization flag number.

[0028] As a further aspect of the present invention, the system further includes:

[0029] The interaction result output module searches for the time segment identified by the synchronization flag number of the artificial intervention response, extracts the AI ​​voice and artificial voice sequences and compares the synchronization trend, determines whether the two types of signals have completed the flow reversal within the preset intervention window, classifies the matching stage as the core interaction segment, and generates the current interaction control result.

[0030] The interactive control results include the signal coordination flow interval, the representative interactive stage number, and the start and end times of the core control segment.

[0031] As a further aspect of the present invention, the interactive result output module includes:

[0032] The segment extraction submodule obtains all time segment numbers listed in the synchronization flag number of the manual intervention response, finds the corresponding start and end time points in the complete call record according to the number, extracts the AI ​​sequence and manual sequence in each time segment in turn, and organizes them into independent signal groups according to the time order to obtain the target segment signal set.

[0033] The control and coordination judgment submodule, based on the target segment signal set, sequentially reads the continuous amplitude sequence of AI and artificial signals in each segment, identifies the continuous occupied segment and continuous silent segment in each sequence, compares whether the control logic of the two types of signals is consistent in direction according to the time step, extracts the continuous range of the synchronous switching trend in the same segment, and obtains the collaborative control coverage interval set.

[0034] The interactive segment labeling submodule calls the continuous signal segments in the coordinated control coverage area, filters them according to the coverage time length within each segment, extracts signal segments with complete control logic and strong continuity within the time range, labels them as continuous segments corresponding to representative interactive stages, and uniformly summarizes the numbering information to obtain the current interactive control result.

[0035] A control method for an AI outbound calling system includes the following steps:

[0036] S1: Obtain the signal sequences of the outbound terminal and the voice box, identify the level gaps in the voice waveform and locate the two ends of the boundary, call the impedance trend extension of the signal before and after the gap to connect them, complete the time sequence arrangement of the two types of sequences, establish the correspondence between the uplink and downlink audio according to the time axis, and generate a synchronous and continuous uplink and downlink voice signal sequence.

[0037] S2: Based on the audio content in the synchronous and continuous uplink and downlink voice signal sequence, extract complete continuous time frames and mark the starting position, identify the segments where the sound energy continuously changes and jumps, extract the relevant time index, classify all call transition segments, and generate call state transition time period number.

[0038] S3: Based on the voice segments contained in the call state transition time period sequence number, identify the fluctuation structure of mode switching within each segment, locate the interval position between adjacent switching actions, determine whether the structure has a continuous alternation trend, extract segments with specific frequency characteristics, mark the classified routing segment number, and generate the routing switching frequency abnormal segment sequence number.

[0039] S4: Call the monitoring audio sequence corresponding to the abnormal segment number of the route switching frequency, read the continuous change trend of the monitoring energy in each segment, find the first intervention point of the artificial trend based on the end time of the AI ​​voice segment, identify the response position of the signal flow to the artificial end, locate the segment of delayed intervention and set a mark, and generate the artificial intervention response synchronization flag number.

[0040] S5: Based on the time segment identified in the synchronization flag number of the human intervention response, extract the corresponding signal sequences of AI and human intervention, observe whether the two types of signals show a synchronous switching or mutually exclusive decline trend, identify continuous segments with collaborative characteristics, set them as key intervals for interactive control, and generate the current interactive control result.

[0041] A computer-readable medium having a control program for an AI outbound calling system stored thereon, wherein the control program for the AI ​​outbound calling system, when executed by a processor, implements a control method for the AI ​​outbound calling system.

[0042] Compared with the prior art, the advantages and positive effects of the present invention are as follows:

[0043] In this invention, a hardware-level bidirectional flow path alignment structure is constructed, and impedance matching and level compensation of analog signals are combined to fill physical link gaps, thereby improving the time consistency and hardware compatibility of cross-terminal audio data transmission. By extracting voice energy jump segments and identifying the distribution structure of the flow path, the frequency range and switching characteristics of AI and manual mode switching are divided. Combined with the response position of the intervention inflection point in the listening sequence, a control lag comparison relationship across hardware ends is established. Based on the collaborative switching trend, the core segments of interactive control are extracted, forming a multi-level analysis system from physical access, feature recognition to collaborative response, and outputting AI outbound call interaction recognition results with routing partitions, switching boundaries and response characteristics. Attached Figure Description

[0044] Figure 1 This is a flowchart illustrating the overall system flow of the present invention.

[0045] Figure 2 This is a flowchart of the system modules of the present invention;

[0046] Figure 3 This is a flowchart of the method of the present invention. Detailed Implementation

[0047] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0048] In embodiments of the present invention, words such as "exemplarily," "for example," etc., are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the word "exemplary" is intended to present the concept in a concrete manner. Furthermore, in embodiments of the present invention, the meaning expressed by "and / or" can be both, or either one.

[0049] In the embodiments of this invention, the terms "image" and "picture" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning. Similarly, the terms "of," "corresponding (relevant)," and "corresponding" may sometimes be used interchangeably. It should be noted that, without emphasizing the distinction between them, they convey the same meaning.

[0050] In this embodiment of the invention, sometimes a subscript such as W1 may be written in a non-subscript form such as W1. When the difference is not emphasized, the meaning they express is the same.

[0051] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0052] Please see Figure 1 This invention provides a technical solution: an AI outbound calling system, the system comprising:

[0053] The physical link construction module acquires analog voice signals from the outbound terminal and the chip-equipped voice box, identifies level gaps in the signal transmission path, smoothly connects the two ends of the signal by adjusting the gain compensation and impedance matching of the physical link, and performs bidirectional stream path alignment after arranging them in time order to generate a synchronous and continuous uplink and downlink voice signal sequence.

[0054] The interaction state recognition module observes whether the pitch frequency rises continuously based on the audio content in the synchronous and continuous uplink and downlink voice signal sequence, identifies audio segments with energy jump changes, extracts the time index of the jump segment as the target segment for interaction state analysis, and generates the call state transition time period number.

[0055] The flow path control module checks whether the voice segments marked by the call state transition time period number fluctuate repeatedly between AI and human modes, extracts the interval path between each mode switch, determines whether the switching actions are concentrated, completes the filtering and classification of switching frequencies, and generates the abnormal segment number of the routing switching frequency.

[0056] The intervention and coordination response module calls the monitoring audio sequence corresponding to the abnormal segment number of the route switching frequency, tracks the change trend of the manual monitoring signal, finds the location of the first manual intervention after the AI ​​voice ends, uses the access point as a response reference, and if a delay in the switching action is found, executes the control command replacement operation and generates a manual intervention response synchronization flag number.

[0057] The interaction result output module searches for the time segment identified by the synchronization flag number of the artificial intervention response, extracts the AI ​​voice and artificial voice sequences and compares the synchronization trend, determines whether the two types of signals have completed the flow reversal within the preset intervention window, classifies the matching stage as the core interaction segment, and generates the current interaction control result.

[0058] The synchronous and continuous uplink and downlink voice signal sequence includes the time correspondence of audio signals, the consistency of signal gain compensation, and the continuity and integrity of the transmission link. The call state transition time period number includes the frequency of voice change occurrence, the distribution range of transition time periods, and the transition segment identification number. The route switching frequency abnormal segment number includes the switching interval distribution density, the mode proportion level, and the location identifier of the abnormal switching segment. The manual intervention response synchronization flag number includes the intervention response delay flag, the signal flow trend offset, and the replacement target control interval. The interactive control result includes the signal coordinated flow interval, the representative interactive stage number, and the start and end time of the core control segment.

[0059] Please see Figure 2 The physical link construction module includes:

[0060] The physical access submodule acquires the analog voice signals collected by the outbound terminal and the voice box, extracts the time points and level amplitudes of each hardware channel according to the timestamp sequence, processes them according to time progression after unifying the sampling frequency, inserts fitting values ​​consistent with the signal trends before and after the physical link disconnection position, so that the audio has a consistent time distribution and waveform continuity, and obtains a signal sequence group with link alignment.

[0061] The system connects the audio output port of the outbound calling mobile terminal and the voice box with a DSP chip via I2S or PCM interface protocols. It reads the original outbound voice stream and the returned analog monitoring signal, respectively. Internally, it has a dual-channel ring buffer. Based on nanosecond-level timestamps generated by hardware clock interrupts, each frame of voice data entering the buffer is tagged to form a raw audio data packet with an absolute time index. During extraction, a timestamp-based sorting operation is performed, traversing the buffer queue, checking the monotonically increasing nature of the timestamps, and removing out-of-order data caused by signal cable jitter. For high-frequency sampled audio, a decimation operation is performed, retaining one sample value at fixed intervals; for low-frequency or intermittent signals, a linear interpolation operation is performed, calculating the level slope between two adjacent raw sample points, based on the time advance step size. The interpolated amplitude is generated point-by-point at intermediate moments. During processing, the physical connection status is continuously monitored. Once the level amplitude is detected to be lower than the preset mute threshold and the duration exceeds a single sampling period (e.g., 20 milliseconds), it is determined that the link is momentarily disconnected or the signal is lost. At this time, the interpolation compensation logic is activated to read the effective level of the last frame before the breakpoint. The level of the first frame after recovery Calculate the overall change gradient A series of filling points with the same gradient direction are generated in the fracture zone according to a uniform sampling rate to ensure that the output signal sequence is strictly equidistant on the time axis and the waveform is continuous. Finally, two sets of signal sequence groups with time axis alignment are generated, including the outbound call terminal and the voice box.

[0062] The impedance matching and completion submodule extracts signal segments with discontinuous level spans based on the signal sequence group aligned with the link, performs amplitude trend analysis on the continuous waveforms on both sides of the breakpoint, fills in matching segments according to the impedance change direction of adjacent segments, and embeds continuous waveform point sequences in each signal blank area to obtain a continuous and complete voice channel dataset.

[0063] The above signal sequence group is primarily used to repair level jumps or signal loss caused by impedance mismatches in different hardware. First, the sliding window scanning logic is run, setting the window length to correspond to 100 sampling points, and calculating the stability of the effective level values ​​within the window. When a rate of change in level is detected... When the set hardware fluctuation threshold is exceeded, the area is marked as a level loss or impedance mismatch point. For the identified breakpoint, 150 samples forward and 150 samples backward are extracted as references. The impedance analysis unit is called to perform least-squares linear fitting on the waveforms on both sides to obtain the envelope slopes of the forward and backward waveforms. If the slopes on both sides have the same polarity, a smooth compensation curve (such as a cubic spline curve) is constructed to connect the breakpoint boundary; if the slopes have opposite polarities, the predicted intersection point is calculated as a virtual phase inflection point, and two connecting segments are constructed. During the filling process, impedance matching values ​​for the corresponding time steps are generated one by one and embedded into the blank positions of the original sequence. A marker is set to record that this segment is hardware compensation data, eliminating physical link breaks caused by different terminal accesses.

[0064] The bidirectional flow alignment submodule calls up the uplink and downlink voice signals in the continuous and complete voice channel dataset, locates each time point according to the uplink data time sequence, selects adjacent time periods in the downlink data, fills in the corresponding values ​​according to the energy trend of the adjacent segments, and performs corresponding processing on the two types of voice streams at all time points to obtain a synchronous and continuous uplink and downlink voice signal sequence.

[0065] Using the timeline of the uplink outbound call as a reference, the timestamps of each frame of uplink data are traversed, and the corresponding time is indexed in the downlink monitoring signal sequence. Due to the hardware delay of the analog conversion circuit, a small search window centered on that timestamp is selected in the downlink data (e.g., ...). (milliseconds), calculate the average energy intensity of the signal within this window. The trend of change is used as the corresponding downlink value at that moment. After indexing is completed at all time points, phase synchronization verification is performed. If the downlink signal index of a certain node is found to be out of bounds, the nearest neighbor copying method is used to fill in the trend value of the previous valid time point. Finally, the processed uplink and downlink signals are merged to output a composite signal sequence containing the synchronization timestamp, outbound voice amplitude, and monitoring voice amplitude.

[0066] The interaction state recognition module includes:

[0067] The time frame division submodule obtains the audio content in the synchronous continuous uplink and downlink voice signal sequence, extracts continuous audio segments according to the signal timestamp, indexes and locates the starting position in each continuous data segment, extracts the audio amplitude sequence between adjacent time frames, and numbers the segments that do not have silence interruptions to obtain a continuous voice time segment number list.

[0068] During the scanning process, the continuity of audio energy is monitored in real time. When the signal amplitude is detected to be continuously higher than the set silence threshold and not manually filled in, the start timestamp of that segment is recorded. Internally, a valid length logic is maintained. A segment is only considered a valid speech segment and assigned a unique identification number if its duration exceeds a set interaction length threshold (e.g., 0.5 seconds). For each valid segment, the amplitude of all adjacent time steps is extracted to construct a sequence array. The scanning pointer automatically cuts off when it encounters a long silence or a link interruption flag, restarting the accumulation for the next segment. The final output is a list of time segments containing multiple independent numbers, ensuring that subsequent analysis is based on real voice interaction content.

[0069] The feature change recognition submodule calls each numbered segment in the continuous speech time period number list, extracts all audio amplitude points within it, compares the amplitudes of adjacent points in all time steps in sequence, determines whether there is a continuous rise in sound energy based on the direction of change between adjacent amplitudes, classifies and labels all segments with continuous amplitude increment features, and obtains a continuous rising tone trend segment index set.

[0070] For each segment of speech data in the numbered list, microscopic spectral analysis is performed, and the difference calculation unit is invoked to perform a first-order forward difference operation on the sequence. Specifically, the data is read... Time Amplitude and Time Amplitude ,calculate By setting an up-pitch counter ,when The time counter increments; when The time counter is reset. An internally preset threshold for determining continuous increase in pitch (e.g.) This represents a continuous rise in the frequency of the speech (a segment where the frequency increases continuously). Once the criteria are met, the current segment is locked and marked as a "continuous rise event." All segments are traversed, and all rising intervals that meet the criteria are counted, generating a set containing start and end indices. This process eliminates minor fluctuations caused by background noise, focusing on identifying intonation rise features in outbound speech that represent questions or emphasis.

[0071] The energy jump extraction submodule extracts signal intervals with short-term sudden increases in amplitude changes in each segment based on the continuous pitch-rising trend segment index set. It extracts the amplitude difference corresponding to the time points before and after each jump point and judges the increase. It filters the signal positions where the energy increase exceeds the set audio jump threshold from all jump points and extracts their corresponding numbers in chronological order to obtain the call state transition time period number.

[0072] The system accesses a rising tone trend index set and performs a second-order gradient scan within each rising segment. A short sliding window of 100 milliseconds is set to calculate the cumulative increase in audio energy within the window. Simultaneously, an "audio jump threshold" is introduced as a key discrimination parameter, set based on acoustic statistical analysis of the actual AI dialogue ending point. The setting process is as follows: The system retrieves the outbound call script library, selects 200 sets of tone feature data at the end of AI dialogues, calculates the energy increase within 100 milliseconds immediately after the end, and obtains a baseline mean of ekcilVWF9b2tXufl.answerText5. Based on this, the jump threshold is set to 80% of this baseline. During operation, if the energy difference between the start and end of the window is greater than ekcilVWF9b2tXufl.answerText2$, it is determined that there is an energy jump at that position (i.e., the transition point where the AI ​​dialogue ends and awaits human feedback). All jump point time positions are recorded and filtered by density, outputting a list of call state transition time period numbers.

[0073] Table 1 shows typical audio signal changes during the processing of this module.

[0074] Table 1: Example Table of Speech Signal Transition Detection Data

[0075]

[0076] The flow path control module includes:

[0077] The flow structure recognition submodule obtains the voice segments marked in the call state transition time period sequence number, extracts the time sequence and flow trend corresponding to each voice sequence, examines whether the alternation order of AI-generated voice and human-monitored voice appears continuously, extracts adjacent position segments between control power transition points, and uses them to subsequently define the routing switching mode to obtain the flow alternation segment index set.

[0078] The system reads audio segments from the corresponding time period, performs an extreme value search operation, and compares the energy weights of uplink (AI) and downlink (human / monitoring) frame by frame. If the uplink energy is greater than the downlink energy, it is marked as "AI busy"; if the downlink energy is greater than the uplink energy, it is marked as "human busy". The order of the busy signal markers is then checked to verify whether an alternating flow pattern of "AI-human-AI" is observed. If permission conflicts or continuous unidirectional busy signals occur, the dominant position is retained based on energy significance. After confirming the alternating flow, adjacent permission transition points (i.e., the gap between the end of the AI ​​dialogue and the intervention of human monitoring) are extracted, and the start and end positions of each segment are recorded. These transition segments constitute the flow structure of the call, and after being summarized, they form an index set of alternating flow segments.

[0079] The switching interval extraction submodule extracts the mode switching interval segments in each segment according to the time sequence based on the flow alternation segment index set. It uses the distance between adjacent route turning points as the switching interval basis, and performs a classification operation on each switching interval segment in the current sequence in sequence. It is then classified into a unified index according to the relative time range to obtain the mode switching interval distribution interval set.

[0080] Based on the index set of alternating flow segments, the frequency characteristics of switching are quantified. The time difference between two adjacent permission inflection points (such as two consecutive AI end points) is calculated and defined as the switching interval. The process iterates through the conversion time periods, generating a series of interval values. To categorize these intervals, interval bucketing logic is introduced, defining standard time range buckets (e.g., ...). s represents instantaneous high-frequency switching. s represents normal interactive adjustment. (The above is for long interactive recovery). Each segment The data is compared with the bucket range and assigned to the corresponding index. This process constructs a set of mode switching interval distribution intervals, recording the distribution of different switching frequencies in the current call, providing a basis for evaluating interaction stability.

[0081] The path proportion judgment submodule calls the time index sequence of each segment in the mode switching interval distribution range set, compares the change in the proportion of switching times in the segment length within the same time period, filters the segment numbers whose proportion value is higher than the set switching density threshold, and summarizes all segment numbers that meet the judgment conditions to obtain the abnormal segment number of routing switching frequency.

[0082] First, acquire data for each interval set, focusing on the "instantaneous high-frequency switching" segment, and calculate the total duration of the current call segment. Statistical analysis of the total handover duration falling within the high-frequency range Perform percentage calculation A "handover density threshold" is introduced for anomaly detection. The threshold setting process is as follows: the system presets the handover density threshold to 0.35. Assuming the total duration of the current transition period is 40 seconds, and the cumulative duration of high-frequency handovers is statistically 18 seconds, the calculation result... .because If the switching frequency of this segment is determined to be abnormal, it may be due to frequent competition or logical conflicts between AI and manual modes. The corresponding sequence number is then included in the list of abnormal segments.

[0083] The intervention and collaborative response module includes:

[0084] The monitoring trend extraction submodule obtains the time segment corresponding to the abnormal segment number of the routing switching frequency, extracts the monitoring sequence data in each time segment, reads all audio energy values ​​in chronological order, continuously judges the direction of level change at adjacent time points, sorts out the time periods of continuous increase or continuous decrease, and obtains a set of monitoring signal change trend sequences.

[0085] For segments marked as switching anomalies, an index mapping mechanism is used to locate the corresponding time interval in the monitoring channel, extending the observation window by 3 seconds. All monitoring sample values ​​within the window are read, and moving average smoothing is performed to filter out hardware noise. Subsequently, monotonicity testing is performed on the sequence; if the level difference is positive for 50 consecutive milliseconds, it is marked as a rising segment of human intervention; if negative, it is marked as a silent monitoring segment. This logic helps to analyze the macro-level trend of human monitoring and preparing for intervention at the headset, generating a set of monitoring signal change trend sequences, with a focus on the first response of the downlink signal after the AI's dialogue ends.

[0086] The access point positioning submodule, based on the monitoring signal change trend sequence set, calls the end time point of each AI voice segment as the time reference point, scans the intervention direction backward in the corresponding monitoring sequence, and when the energy shift first appears in the human intervention trend, extracts the time point and corresponding level value at that position to construct the subsequent synchronization reference segment and obtains the human intervention time index table.

[0087] Using the end point of the AI-generated dialogue as the reference zero point Scan the trend sequence backward from this point. Look for the critical point where the listening signal transitions from a "silent" state to a "human intervention" state. When the trend indicator first changes from stable or declining to a sustained increase, lock this reversal time position and define it as the human intervention point. Iterate through all abnormal segments, establish a corresponding relationship for each segment, and generate a list containing... and The index table. This step aims to determine the hardware-level lag point in the transfer of system control.

[0088] The response synchronization annotation submodule, according to the manual access time index table, compares the end time of each AI voice segment with the interval length between the start time of manual intervention and the end time of AI voice in each segment, extracts the sequence number of the time interval that is higher than the intervention delay threshold, and performs a logical reset operation on the corresponding segment to obtain the manual intervention response synchronization flag number.

[0089] Calculate lag time An "intervention delay threshold" is introduced as a criterion, based on the normal flow response time of the outbound call system. Threshold calculation example: The system sets the intervention delay threshold to 800 milliseconds. In actual operation, the AI ​​end point for a certain segment is 120.500s, and the human intervention point is 121.800s, resulting in a calculated lag time of 1300 milliseconds. Comparing the result of 1300 milliseconds with the threshold of 800 milliseconds, since ekcilVWF9b2tXufl.answerText300>800$, a significant interaction delay is determined. A logic reset is then executed, marking this segment as a "delayed response zone" and adding it to the list of human intervention response synchronization flags. Table 2 shows the response determination for typical interaction segments.

[0090] Table 2: Test Data on Delay in Manual Intervention Response

[0091]

[0092] The interactive result output module includes:

[0093] The segment extraction submodule obtains all time segment numbers listed in the synchronization flag number of the manual intervention response, finds the corresponding start and end time points in the complete call record according to the number, extracts the AI ​​sequence and manual sequence in each time segment in turn, and organizes them into independent signal groups according to the time order to obtain the target segment signal set.

[0094] Tracing back to the call database, the voice segments are precisely cut according to the calibrated start and end timestamps associated with the numbers. Dual-channel data containing the complete AI script and subsequent human intervention responses are extracted and encapsulated into objects containing "AI arrays," "human arrays," and "time indices," resulting in the target segment signal set. Each set of data represents a complete "AI-guided - human intervention" interaction loop.

[0095] The control and coordination judgment submodule, based on the target segment signal set, sequentially reads the continuous amplitude sequence of AI and artificial signals in each segment, identifies the continuous occupied segment and continuous silent segment in each sequence, compares whether the control logic of the two types of signals is consistent in direction according to the time step, extracts the continuous range of the synchronous switching trend in the same segment, and obtains the collaborative control coverage interval set.

[0096] Directional comparison is performed on the dual-channel sequences within the target segment, marking each time step as either "online" or "silent." A logical AND operation is performed: when the uplink AI is "silent" and the downlink manual is "online," it is determined as a "valid permission transfer point"; if both are online simultaneously, it is determined as an "interaction conflict point." Continuous valid transfer points are connected into a block to form a synchronous change interval, eliminating interference from unilateral noise or hardware crosstalk, and preserving the true collaborative control range.

[0097] The interactive segment labeling submodule calls the continuous signal segments in the coordinated control coverage area, filters them according to the coverage time length within each segment, extracts signal segments with complete control logic and strong continuity within the time range, labels them as continuous segments corresponding to representative interactive stages, and uniformly summarizes the numbering information to obtain the current interactive control result.

[0098] The duration of the collaborative interval is calculated, and an "effective interaction threshold" (e.g., 1.5 seconds) is introduced. If a collaborative interval lasts 3.2 seconds, meeting the screening criteria, it is marked as a representative interaction segment. A list is generated that includes the interaction type (AI to human / human intervention), start and end times, and quality control assessment. This result directly reflects the closeness of AI and human collaboration during outbound calls. Table 3 shows the final output.

[0099] Table 3: Outbound Call Interaction Control Results Table

[0100]

[0101] Please see Figure 3 A control method for an AI outbound calling system includes the following steps:

[0102] S1: Obtain the signal sequences of the outbound terminal and the voice box, identify the level gaps in the voice waveform and locate the two ends of the boundary, call the impedance trend extension of the signal before and after the gap to connect them, complete the time sequence arrangement of the two types of sequences, establish the correspondence between the uplink and downlink audio according to the time axis, and generate a synchronous and continuous uplink and downlink voice signal sequence.

[0103] S2: Based on the audio content in the synchronous and continuous uplink and downlink voice signal sequence, extract complete continuous time frames and mark the starting position, identify the segments where the sound energy continuously changes and jumps, extract the relevant time index, classify all call transition segments, and generate call state transition time period number.

[0104] S3: Based on the voice segments contained in the call state transition time period sequence number, identify the fluctuation structure of mode switching within each segment, locate the interval position between adjacent switching actions, determine whether the structure has a continuous alternation trend, extract segments with specific frequency characteristics, mark the classified routing segment number, and generate the routing switching frequency abnormal segment sequence number.

[0105] S4: Call the monitoring audio sequence corresponding to the abnormal segment number of the route switching frequency, read the continuous change trend of the monitoring energy in each segment, find the first intervention point of the artificial trend based on the end time of the AI ​​voice segment, identify the response position of the signal flow to the artificial end, locate the segment of delayed intervention and set a mark, and generate the artificial intervention response synchronization flag number.

[0106] S5: Based on the time segment identified in the synchronization flag number of the human intervention response, extract the corresponding signal sequences of AI and human intervention, observe whether the two types of signals show a synchronous switching or mutually exclusive decline trend, identify continuous segments with collaborative characteristics, set them as key intervals for interactive control, and generate the current interactive control result.

[0107] A computer-readable medium having a control program for an AI outbound calling system stored thereon, wherein the control program for the AI ​​outbound calling system, when executed by a processor, implements a control method for the AI ​​outbound calling system.

[0108] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. An AI outbound call system, characterized by, The system comprises: A physical link construction module acquires analog voice signals of an outbound terminal and a voice box with a chip, identifies a level missing point in a signal transmission path, compensates and impedance matches both ends of the signal by adjusting the gain of the physical link, arranges in time sequence, performs bidirectional flow path alignment, and generates a synchronous continuous uplink and downlink voice signal sequence; An interactive state recognition module observes whether the tone frequency is continuously rising according to the audio content in the synchronous continuous uplink and downlink voice signal sequence, identifies an audio segment with energy jump change, extracts the time index of the jump segment as an interactive state analysis target section, and generates a call state conversion time period sequence number; A flow path control module checks whether the internal mode repeatedly fluctuates between AI and manual mode according to the voice segment marked by the call state conversion time period sequence number, extracts the interval path between each mode switching, judges whether the switching action is concentratedly distributed, completes the screening and classification of switching frequency, and generates a routing switching frequency abnormal segment sequence number; An intervention collaborative response module calls the listening audio sequence corresponding to the routing switching frequency abnormal segment sequence number, tracks the trend of manual listening signal change, finds the first manual intervention position after the AI voice ends, takes the access point as a response reference, executes a control instruction replacement operation if the switching action is found to be delayed, and generates a manual intervention response synchronization flag number. 2.The AI outbound call system of claim 1, wherein: The synchronous continuous uplink and downlink voice signal sequence comprises a time corresponding relationship of an audio signal, signal gain compensation consistency, and transmission link continuity, the call state conversion time period sequence number comprises a voice mutation appearance frequency, a conversion time period distribution range, and a conversion section identification number, the routing switching frequency abnormal segment sequence number comprises a switching interval distribution density, a mode proportion level, and a switching abnormal section positioning identification, and the manual intervention response synchronization flag number comprises an intervention response delay identification, a signal flow trend offset, and a replacement target control interval. 3.The AI outbound call system of claim 1, wherein, The physical link construction module comprises: A physical access submodule acquires analog voice signals collected by an outbound terminal and a voice box, extracts time points and level amplitudes of each hardware channel according to a timestamp sequence, processes in time sequence after uniform sampling frequency, inserts a fitting value consistent with the trend of the signals before and after the physical link is disconnected, makes the audio have consistent time distribution and waveform continuity, and obtains a link-aligned signal sequence group; An impedance matching completion submodule extracts a signal segment with discontinuous level span based on the link-aligned signal sequence group, analyzes the amplitude trend of continuous waveforms on both sides of the breakpoint, supplements a matching segment according to the impedance change direction of adjacent segments, embeds a continuous waveform point column in each signal blank area, and obtains a continuous and complete voice channel data set; A bidirectional flow path alignment submodule calls uplink and downlink voice signals in the continuous and complete voice channel data set, positions each time point in the uplink data time sequence, selects adjacent time segments in the downlink data, fills in corresponding values according to the energy trend of adjacent segments, and performs corresponding processing on the two types of voice streams at all time nodes to obtain a synchronous continuous uplink and downlink voice signal sequence. 4.The AI outbound call system of claim 1, wherein, The interactive state recognition module comprises: The time frame division sub-module acquires audio content in the synchronous and continuous uplink / downlink voice signal sequence, sequentially extracts continuous audio segments according to signal timestamps, indexes the starting position in each continuous data segment, extracts the audio amplitude sequence between adjacent time frames, and performs numbering processing on the paragraphs without a silent interruption, to obtain a continuous voice time period numbering list. The feature change recognition sub-module calls each numbered segment in the continuous voice time period numbering list, extracts all audio amplitude points in the segment, sequentially compares the amplitudes of adjacent points in all time steps, determines whether there is a continuous amplitude increment feature according to the change direction between adjacent amplitudes, classifies and labels all paragraphs with the continuous amplitude increment feature, and obtains a continuous rising trend paragraph index set. The energy jump extraction sub-module extracts signal intervals with a short-time sudden increase in amplitude change in each paragraph according to the continuous rising trend paragraph index set, extracts the amplitude difference between time points before and after each jump point and discriminates the amplitude increase, filters signal positions with an energy growth amplitude exceeding a set audio jump threshold among all jump points, sequentially extracts the corresponding number according to time sequence, and obtains a call state conversion period sequence number. 5.The AI outbound call system of claim 1, wherein, The flow transfer path control module includes: The flow direction structure recognition sub-module acquires the voice segments marked in the call state conversion period sequence number, sequentially extracts the time sequence and flow direction trend corresponding to each voice sequence, checks whether the alternating order of AI-generated voice and artificial monitoring voice appears continuously, extracts adjacent position sections between control right conversion points, and uses the sections to subsequently determine the routing switching mode, to obtain a flow transfer alternating section index set. The switching interval extraction sub-module extracts mode switching interval sections in each section according to the flow transfer alternating section index set, takes the distance between adjacent routing turning points as the switching interval basis, sequentially classifies each switching interval in the current sequence, and classifies the intervals into a unified index according to the relative time range, to obtain a mode switching interval distribution interval set. The path proportion judgment sub-module calls the time index sequence of each section in the mode switching interval distribution interval set, compares the proportion change of the switching times in the paragraph length in the same time period, filters section numbers with a proportion value higher than a set switching density threshold, and performs summary processing on all section numbers that meet the determination condition, to obtain a routing switching frequency abnormal segment sequence number. 6.The AI outbound call system of claim 1, wherein, The intervention coordination response module includes: The monitoring trend extraction sub-module acquires the time segments corresponding to the routing switching frequency abnormal segment sequence number, extracts the monitoring sequence data in each time segment, reads all audio energy values in time sequence, continuously judges the level change direction of adjacent time points, processes the time periods with a continuous rise or a continuous fall, and obtains a monitoring signal change trend sequence set. The access point positioning sub-module calls the end time point of each AI voice segment as a time reference point based on the set of listening signal change trend sequences, scans the intervention direction in the corresponding listening sequence backward, extracts the time point and the corresponding level value when the energy direction is first changed when the artificial intervention trend occurs, and uses the time point and the corresponding level value to construct a subsequent synchronization reference section to obtain an artificial access time index table; The response synchronization marking sub-module compares the artificial intervention start time and the AI voice termination time in each segment according to the artificial access time index table, extracts the sequence number of the time interval that is higher than the intervention delay threshold, and performs a logical reset operation on the corresponding paragraph to obtain an artificial intervention response synchronization marker number. 7.The AI outbound call system of claim 1, wherein, The system further comprises: The interactive result output module finds the time segment identified by the artificial intervention response synchronization marker number, extracts the AI voice and artificial voice sequences, compares the synchronization trend, judges whether the two types of signals complete the flow direction conversion within the preset intervention window, classifies the consistent stage as a core interactive segment, and generates a current interactive control result; The interactive control result includes a signal coordination flow conversion interval, a representative interactive stage number, and a core control segment start and end time. 8.The AI outbound call system of claim 1, wherein, The interactive result output module comprises: The segment extraction sub-module obtains all the time segment numbers listed in the artificial intervention response synchronization marker number, finds the corresponding start and end time points in the complete call record according to the numbers, extracts the AI sequence and the artificial sequence in each time segment in turn, organizes the AI sequence and the artificial sequence into independent signal groups in time sequence respectively, and obtains a target segment signal set; The control coordination judgment sub-module reads the continuous amplitude sequence of the AI and artificial signals in each segment in turn based on the target segment signal set, identifies the continuous busy section and the continuous silent section in each sequence, compares the control logic of the two types of signals in time steps, extracts the continuous range of the synchronization switching trend in the same section, and obtains a coordination control coverage interval set; The interactive section marking sub-module calls the continuous signal section in the coordination control coverage interval set, filters the signal sections with complete control logic and strong continuity in each section according to the coverage time length, marks the continuous sections corresponding to the representative interactive stages, uniformly collects the number information, and obtains the current interactive control result. 9.A control method of an AI outbound call system, the method comprising: The method is used in the AI outbound call system of any one of claims 1-8, and comprises the following steps: S1: Obtain the signal sequences of the outbound terminal and the voice box, identify the level missing points in the voice waveform and locate the two end boundaries, call the impedance trend extension of the signals before and after the gap to connect them, complete the time sequence arrangement of the two types of sequences, establish the corresponding relationship between the uplink and downlink audio signals on the time axis, and generate the uplink and downlink voice signal sequences in synchronization and continuity; S2: Based on the audio content in the uplink and downlink voice signal sequences in synchronization and continuity, extract complete continuous time frames and mark the start positions, identify the sections with continuous sound energy changes and jumps, extract the related time indexes, classify all call conversion segments, and generate call state conversion period numbers. S3: Based on the voice segments contained in the call state transition time period sequence number, identify the fluctuation structure of mode switching within each segment, locate the interval position between adjacent switching actions, determine whether the structure has a continuous alternation trend, extract segments with specific frequency characteristics, mark the classified routing segment number, and generate the routing switching frequency abnormal segment sequence number. S4: Call the monitoring audio sequence corresponding to the abnormal segment number of the route switching frequency, read the continuous change trend of the monitoring energy in each segment, find the first intervention point of the artificial trend based on the end time of the AI ​​voice segment, identify the response position of the signal flow to the artificial end, locate the segment of delayed intervention and set a mark, and generate the artificial intervention response synchronization flag number. S5: Based on the time segment identified in the synchronization flag number of the human intervention response, extract the corresponding signal sequences of AI and human intervention, observe whether the two types of signals show a synchronous switching or mutually exclusive decline trend, identify continuous segments with collaborative characteristics, set them as key intervals for interactive control, and generate the current interactive control result.

10. A computer readable medium characterized by It stores the control program of the AI ​​outbound calling system, which, when executed by the processor, implements the control method of the AI ​​outbound calling system as described in claim 9.

Citation Information

Patent Citations

  • Intelligent outbound voice robot system and outbound method

    CN111131638A

  • Flow control method and system in man-machine coupled call, and storage medium

    CN115866141A