Intelligent collection method and device based on audio analysis

By using audio analysis technology to collect and process debtor call audio in real time, and combining a multimodal feature fusion engine and a strategy decision engine, the scripts are dynamically adjusted and compliant documents are generated. This solves the problems of low efficiency and poor compliance in existing intelligent debt collection technologies, and realizes an efficient and accurate intelligent debt collection process.

CN121486498APending Publication Date: 2026-02-06ZHANGZHOU SEETEC OPTOELECTRONICS TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610020860.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing intelligent debt collection technologies rely on human communication, which is inefficient, lack multi-dimensional data analysis capabilities, and struggle to accurately capture debtors' emotions and information. Furthermore, the compliance document generation process is unsystematic, resulting in low collection efficiency and insufficient compliance.

Method used

By using audio analysis technology to collect debtors' call audio in real time, extracting audio frame data and voice activity detection results, and combining a multimodal feature fusion engine and a strategy decision engine, the system can dynamically adjust the script and generate compliant documents to achieve intelligent debt collection throughout the entire process.

Benefits of technology

It improves collection efficiency, reduces reliance on manual labor, accurately adapts to communication scenarios, reduces resistance, ensures accurate and compliant information, optimizes the interactive experience, and achieves an efficient closed loop in the collection process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121486498A_ABST
    Figure CN121486498A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent collection method and device based on audio analysis, and is applied to the technical field of data processing.The method comprises the steps that multi-mode audio analysis serves as the core, and basic information such as debtor credit data and legal document templates is collected and stored in a structured mode; call audio is transferred into a text stream with a timestamp in real time through an ASR technology, and audio framing, voice activity detection and other features are synchronously extracted. Through semantic, acoustic feature and dialogue rhythm three-dimensional analysis, a communication strategy and a switching threshold are determined by a multi-modal fusion and strategy decision engine, an adaptive verbal skill is generated, and a conversation link is established through a third-party platform. The debtor feedback is analyzed in real time, the verbal skill is dynamically adjusted until an effective result is obtained, finally key information such as committed repayment is extracted, a compliance legal document is automatically generated and sent out through one key after manual auditing, and whole-process intelligent collection is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to an intelligent debt collection method and device based on audio analysis. Background Technology

[0002] Currently, the field of intelligent debt collection mainly relies on human communication or basic automated script systems. Existing technologies have significant shortcomings in terms of efficiency, accuracy, and compliance, specifically as follows: High reliance on human intervention and low efficiency: Traditional debt collection relies heavily on human agents to communicate with debtors. This requires human judgment of the debtor's emotions and intentions and adjustment of the script. The daily call volume that a single person can handle is limited, making it difficult to cope with a large number of debt cases. Moreover, human communication is easily affected by subjective emotions, resulting in poor consistency of the script and unstable debt collection success rate.

[0003] Lack of multi-dimensional data analysis capabilities: Existing automated debt collection systems can only conduct fixed-script communication based on text scripts, and cannot analyze the acoustic characteristics and dialogue rhythm in the call audio in real time. It is difficult to accurately capture the debtor's emotional fluctuations and changes in participation, and it is impossible to adjust communication strategies in a timely manner, which can easily cause debtor resistance and reduce the effectiveness of communication.

[0004] Insufficient capture and verification of key information: During the communication process, existing technology is unable to automatically extract key information such as the repayment time and amount promised by the debtor, and lacks a real-time verification mechanism for the accuracy of information, which is prone to information recording deviations; when the debtor raises disputes about debt information, it is impossible to quickly trigger a strategy switch, resulting in a communication deadlock and difficulty in obtaining a valid repayment commitment.

[0005] The generation of compliance documents is disconnected from the process: In the collection process, key information needs to be sorted out manually based on the communication results, legal document templates need to be matched and compliance documents need to be generated. Manual operation is not only time-consuming, but also prone to non-compliance due to information omissions or format errors. In addition, the document review and sending process lacks systematic connection, which affects the efficiency of the collection process. Summary of the Invention

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: An intelligent debt collection method based on audio analysis includes: acquiring basic data and parameter information related to debt collection, including debtor credit data, legal document templates, and type IDs, associated debt IDs, priority IDs, collection node IDs, and communication patterns for each data type; cleaning, verifying, and structurally storing the basic data through a data processing module; real-time acquisition of debtor call audio streams using ASR technology, extracting audio frame data, voice activity detection results, and speaker separation timestamps, and synchronously converting them into timestamped text streams; parsing the audio and text data, classifying them according to semantic, acoustic features, and dialogue rhythm dimensions, and using a multimodal feature fusion engine based on audio volume, speech rate, pause features, and text keywords and intent information. The system consults with the strategy decision engine to determine the direction of communication strategy adjustments, and determines the pace of script adjustments and strategy switching thresholds based on debtor emotion scores and participation indicators. It processes scripts according to the target communication strategy templates to generate voice scripts that align with the tone and communication objectives. Based on the real-time interactive rhythm, it transmits the voice scripts and establishes a call link with the debtor through a third-party telephone platform. It receives feedback information, analyzes the feedback data in real time, and verifies the accuracy of key information. If there is any dispute, it triggers a strategy switching mechanism to adjust the scripts until a valid communication result is obtained. It extracts the promised repayment time and amount information from the call, matches it with the corresponding legal document template to automatically generate compliant documents, and issues them with one click after manual review, completing the entire intelligent collection process.

[0007] An intelligent debt collection device based on audio analysis, the device includes: an acquisition module, used to acquire basic data and parameter information related to debt collection, including debtor credit data, legal document templates, and type ID, associated debt ID, priority ID, collection node ID and communication mode of each data, and a data processing module to clean, verify and store the basic data in a structured manner. The feature extraction module is used to collect the debtor's call audio stream in real time using ASR technology, extract audio frame data, voice activity detection results, speaker separation timestamps, and synchronously convert them into a time-stamped text stream; The voice information generation module is used to parse audio and text data, classify and process them according to semantic, acoustic features and dialogue rhythm dimensions, and determine the direction of communication strategy adjustment by the multimodal feature fusion engine and the strategy decision engine based on the volume, speech rate and pause features of the audio, and the keywords and intent information of the text. It also determines the rhythm of the speech adjustment and the threshold for strategy switching based on the debtor's emotion score and participation index. The speech is processed according to the target communication strategy dialogue template to form voice speech information that conforms to the attitude tone and communication goals. The voice communication module is used to transmit voice scripts based on real-time interactive rhythm, establish a call link with debtors through a third-party telephone platform, receive feedback information, and analyze feedback data and verify the accuracy of key information in real time. If there is a dispute, a strategy switching mechanism is triggered to adjust the script until a valid communication result is obtained. The module extracts the promised repayment time and amount information from the call, matches the corresponding legal document template to automatically generate compliant documents, and sends them out with one click after manual review, completing the entire intelligent collection process.

[0008] Its beneficial effects are as follows: This invention provides an intelligent debt collection method based on audio analysis. It collects and structures basic information such as debtor credit data and legal document templates; it uses ASR technology to transcribe call audio into a timestamped text stream in real time, and simultaneously extracts features such as audio frame segmentation and voice activity detection; through semantic, acoustic features, and dialogue rhythm analysis, a dual-engine system determines the communication strategy and switching threshold, generates appropriate scripts, and establishes a call link; it analyzes debtor feedback in real time, dynamically adjusts the scripts, and finally extracts key repayment information, automatically generates compliant legal documents, and issues them with one click after manual review, thus realizing an intelligent debt collection closed loop.

[0009] This invention significantly improves debt collection efficiency. The automated process reduces reliance on manual labor, allowing a single person to handle more cases while maintaining high consistency in communication. It precisely adapts to communication scenarios, analyzes debtor status from multiple dimensions, dynamically adjusts strategies, reduces resistance, and increases the rate of obtaining repayment commitments. It ensures accurate and compliant information by verifying key information in real time and automatically generating standardized legal documents, reducing human error and compliance risks. The optimized user experience and customized communication scripts, tailored to debtor emotions and pace, enhance communication effectiveness and facilitate an efficient closed-loop debt collection process. Attached Figure Description

[0010] Figure 1 A flowchart illustrating an intelligent debt collection method based on audio analysis provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of a smart debt collection device based on audio analysis, provided as an embodiment of the present invention. Detailed Implementation

[0011] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention. Figure 1 This application describes an audio analysis-based intelligent debt collection method according to exemplary embodiments thereof.

[0012] In this application embodiment, an intelligent debt collection method based on audio analysis is described, such as... Figure 1 As shown: S101, obtain basic data and parameter information related to debt collection.

[0013] In one implementation, the basic data and parameter information related to debt collection includes debtor credit data, legal document templates, and type ID, associated debt ID, priority ID, collection node ID, and communication mode for each data. The basic data is cleaned, verified, and stored in a structured manner through a data processing module. The specific information is as follows.

[0014] The system receives credit data files in CSV and Excel formats uploaded by users via a file upload interface, and extracts core data such as debtor identity information (name, ID number), contact information (mobile phone number, landline number), and basic debt information (debt amount, loan date, number of overdue days). The system includes pre-set templates for various standardized legal documents such as "Collection Notice," "Repayment Commitment Letter," and "Settlement Agreement." These templates contain fixed modules such as party information, debt details, commitment clauses, and explanations of legal validity, while also reserving fields for dynamic filling.

[0015] Assign unique type IDs to various types of data, such as "DATA_CREDIT" for credit data types and "TEMPLATE_LEGAL" for legal document template types; bind associated debt IDs according to debt attribution to ensure accurate association between data and corresponding debts; set priority IDs based on the degree and amount of debt delinquency, such as "PRIO_HIGH" for debts with high delinquency amounts; assign collection node IDs according to the collection process nodes (such as "NODE_INIT" for the initial collection node and "NODE_NEGO" for the negotiation stage node); determine the communication mode based on the debtor's historical communication feedback, such as using a "friendly negotiation mode" for debtors with high cooperation and a "compliance warning mode" for debtors who have repeatedly delayed.

[0016] Check the debtor's credit data for missing required fields such as identity information, debt amount, and contact information. If an empty ID number or an unfilled debt amount is found, mark it as incomplete data and return a completion prompt. Verify that the debt amount format conforms to numerical specifications, and check for anomalies such as non-numeric characters or negative amounts. Verify the legality of the ID number and mobile phone number formats, such as whether the ID number is 18 digits long and whether the mobile phone number conforms to the national number segment rules. Compare the extracted debtor's promised repayment amount and other subsequently collected information with the original debt amount, and mark any discrepancies. Ensure that the debt information corresponding to the associated debt ID is consistent throughout the system to avoid data conflicts such as the amount or number of overdue days for the same debt. Verify the matching of the collection node ID with the current collection process stage to prevent the node ID from becoming disconnected from the actual process.

[0017] The cleaned and verified debtor credit data is categorized and organized into "identity information - debt information - communication records" and transformed into a standardized data format. For example, debt information is stored in the field structure of "debt ID - debt amount - loan date - overdue days". Legal document templates are structured according to template type, applicable scenarios, and field mapping rules, and the data source corresponding to each dynamic field is clearly defined.

[0018] All structured data is stored in a database, and an index is established to link debtor credit data, legal document templates, and various identification parameters. This ensures that corresponding data can be quickly retrieved through associated debt IDs, collection node IDs, etc., during the subsequent collection process, while also guaranteeing the integrity of data storage and access efficiency.

[0019] S102 uses ASR technology to collect the debtor's call audio stream in real time, extract audio frame data, voice activity detection results, and speaker separation timestamps, and synchronously convert them into a time-stamped text stream.

[0020] In one implementation, the audio of a debtor's conversation is processed based on the needs of audio analysis and real-time interaction. An ASR (Automatic Speech Recognition) mechanism is used to extract audio frame data, voice activity detection results, and core audio features such as speaker separation timestamps. The audio signal preprocessing module encapsulates the data at the PCM (Public-Content Management) level according to PCM specifications. Based on multimodal analysis rules, the audio feature type, semantic extraction accuracy, and timestamp synchronization format are extracted and processed to generate a unified format audio frame data packet, a speech-to-text stream, feature encoding mapping information, and audio-text synchronization verification information. Based on the needs of audio analysis and real-time interaction, the original audio stream of the debtor's conversation is processed to provide basic data support for subsequent multimodal analysis and real-time decision-making. The ASR mechanism extracts three types of core features from the conversation audio. First, audio frame data, which divides the continuous audio stream into independent audio frames at fixed time intervals, making it easy to analyze frame by frame; second, voice activity detection results, which identify segments of audio containing valid speech and silent segments, accurately distinguishing the voice interaction periods in a call; and third, speaker separation timestamps, which clarify the start and end times of different speakers (debtors and AI debt collectors) in the audio stream, achieving accurate division of the voices of both parties.

[0021] The audio signal preprocessing module encapsulates the extracted audio frames according to PCM data specifications, unifying the audio data storage format and transmission standards to ensure compatibility and consistency in subsequent processing. Based on preset multimodal analysis rules, audio feature types are clearly classified and labeled to ensure that different features correspond to precise analysis dimensions. Semantic extraction accuracy standards are set to ensure that the ASR-transcribed text accurately reflects the core meaning expressed by the debtor. A unified timestamp synchronization format ensures that the timestamps of audio frames, speech activity segments, and speaker speech segments adhere to consistent encoding rules.

[0022] After the above processing, four types of standardized information are generated, including audio frame data packets in a unified format, complete speech-to-text streams, feature encoding mapping information that records the correspondence between features and analysis dimensions, and audio-text synchronization verification information used to verify the correlation between audio and text.

[0023] Based on audio-text synchronization verification information, bidirectional verification is performed on various generated information. A real-time supplementary acquisition mechanism is initiated for missing audio frames. For abnormal transcribed text, a semantic correction service is invoked for secondary transcription. For data with outdated timestamps, a synchronization calibration process is triggered, generating an optimized audio processing solution that includes anomaly handling strategies and supplementary acquisition time windows. Based on the audio-text synchronization verification information, bidirectional verification is performed on the generated audio frame data packets, speech-to-text streams, and other information. Verification content includes the semantic consistency between the speech segments corresponding to the audio frames and the transcribed text, and the accurate correspondence of timestamps, ensuring accurate matching of audio and text information.

[0024] If a missing audio frame is detected during the verification process, a real-time data acquisition mechanism is immediately activated to re-acquire the audio data for the missing period through the call link, ensuring the integrity of the audio data and preventing data loss from affecting subsequent analysis results. When anomalies such as semantic inconsistencies or missing key information are detected in the transcribed text, the semantic correction service is invoked to perform a secondary transcription of the abnormal text. Combining the contextual semantics and call scenario, transcription errors are corrected, key information is supplemented, and the accuracy and usability of the text stream are ensured.

[0025] If a discrepancy is detected between the timestamps of the audio features and the text stream, a synchronization calibration process is triggered. Using the call start time as a baseline, the timestamps of audio frames, voice activity detection results, and transcribed text are adjusted to ensure complete alignment of all information in the time dimension. These anomaly handling measures are integrated to form an optimized audio processing scheme that includes specific anomaly handling strategies and audio frame re-acquisition time windows, providing clear guidance for subsequent processing.

[0026] The optimized processing scheme is integrated and executed, responding to real-time call analysis commands. Simultaneously, the call start time is bound to the time synchronization module, generating a text stream with a unified timestamp and associated interval information with audio features. The optimized audio processing scheme is comprehensively integrated, performing audio data supplementation, text correction, and timestamp calibration operations according to the scheme requirements. This ensures all information meets the requirements of multimodal analysis and real-time interaction, while rapidly responding to real-time call analysis commands to guarantee the system's timeliness. The time synchronization module binds the call start time to the system time, establishing a unified time benchmark and providing a consistent basis for timestamp annotation of various information.

[0027] Based on the bound call start time, a speech-to-text stream with a unified timestamp is generated to ensure that each sentence in the text stream accurately corresponds to a specific time point in the call; at the same time, audio feature association interval information is generated to clarify the time interval relationship between different audio features (such as audio frames and speech activity segments), laying the foundation for subsequent multimodal feature fusion analysis.

[0028] S103 analyzes audio and text data, classifies and processes them according to semantic, acoustic features, and dialogue rhythm dimensions, and determines the direction of communication strategy adjustment by the multimodal feature fusion engine and the strategy decision engine based on the audio volume, speech rate, and pause features, and the text keywords and intent information. It also determines the rhythm of speech adjustment and the threshold for strategy switching based on the debtor's emotion score and participation index.

[0029] In one implementation, the audio data is parsed and processed to extract acoustic features of volume, speech rate, and pauses, generating an acoustic feature vector. First, preprocessing operations such as framing and windowing are performed on the preprocessed PCM format audio frame data to eliminate inter-frame interference and improve the accuracy of feature extraction. Then, three core acoustic features are extracted: First, volume features, which quantify the volume by detecting the amplitude intensity of the audio signal and calculating the root mean square (RMS) value of each audio frame, reflecting the changes in the strength of the debtor's voice. Second, speech rate features, which combine the ASR-transcribed text stream with the timestamps of the corresponding audio frames to count the number of syllables or effective words pronounced by the debtor per unit time, while also recording the syllable interval duration, comprehensively judging the speed and fluency of their speech. Third, pause features, based on the results of speech activity detection (VAD), identify periods in the audio stream without effective speech, extracting information such as the duration, frequency, and distribution location of pauses. The frequency feature is calculated by counting the number of pauses per unit time, and the distribution location feature marks the specific location of the pause in the sentence (beginning, middle, or end). Finally, the extracted volume feature sequences, speech rate quantization values, and pause-related parameters are sorted and arranged according to preset dimensions (time dimension and feature type dimension) to form a high-dimensional acoustic feature vector that can comprehensively represent the acoustic properties of audio, providing accurate audio data support for subsequent multimodal feature fusion.

[0030] Semantic analysis is performed on text data to extract keywords and intent recognition results, generating semantic understanding features. Based on timestamped text streams generated by ASR transcription, deep semantic analysis is conducted using a locally deployed, finely tuned ChatGLM model to ensure the accuracy and scenario adaptability of the analysis results. Through keyword extraction technology, based on a debt collection scenario corpus, core words strongly related to debt collection are identified in real time, including categories related to repayment ability (e.g., "paycheck," "no money," "cash flow"), time-related (e.g., "next week," "end of the month," "next Friday"), and dispute-related (e.g., "disputed," "disagreement," "calculation error"), clarifying the key information and core demands expressed by the debtor.

[0031] Using an intent classification model based on labeled debt collection case training data, the debtor's current communication intent is accurately categorized into four types: "delay," "rejection," "negotiation," and "commitment." "Delay" refers to the debtor postponing repayment for various reasons; "rejection" refers to explicitly denying the debt or refusing to repay; "negotiation" refers to actively exploring repayment plans; and "commitment" refers to explicitly stating that the debtor will repay on time. This allows for a precise understanding of the debtor's willingness to repay. Simultaneously, combining sentiment analysis technology, the debtor's emotional state is determined through the use of interjections, sentence structure, and semantic tendencies in the text. Specifically, it is categorized into four types: "calm," "excited," "angry," and "frustrated." "Calm" indicates a calm expression without strong emotional tendencies; "excited" indicates an urgent tone and obvious emotional fluctuations; "angry" indicates negative emotional expressions such as accusation and dissatisfaction; and "frustrated" indicates a low mood indicating an inability to repay.

[0032] Finally, the extracted classification keywords, clear intent recognition results, and accurate sentiment analysis results are integrated in multiple dimensions to generate semantic understanding features that include semantic core, intention tendency, and emotional state. This comprehensively and deeply reflects the semantic connotation behind the text and the debtor's subjective state, providing solid semantic support for subsequent strategic decisions.

[0033] This study analyzes the rhythm of call interaction data, extracting engagement metrics and interaction behavior patterns to generate dialogue rhythm features. It also quantifies acoustic feature vectors to generate standardized acoustic evaluation features. Based on speech activity detection results and speaker separation timestamps, the study analyzes the rhythm of call interaction data. Engagement metrics are extracted, and the debtor's participation level in the call is assessed by statistically analyzing data such as the debtor's speaking time percentage and response timeliness. Interaction behavior patterns are identified, such as whether the debtor frequently interrupts, remains silent for extended periods, or initiates questions, clarifying the interaction methods between the two parties. The engagement metrics and interaction behavior patterns are integrated to generate dialogue rhythm features, visually presenting the interaction rhythm and debtor's behavior during the call.

[0034] The generated acoustic feature vectors are quantized using a unified quantization standard, converting the raw data of features such as volume, speech rate, and pauses into standardized values. This process eliminates dimensional differences between different features, making various acoustic features comparable and generating standardized acoustic evaluation features, laying the foundation for subsequent multi-feature fusion.

[0035] The semantic understanding features are structured to generate structured semantic decision features. Unstructured information within the semantic understanding features is systematically organized and standardized according to pre-defined structured rules. First, the core framework of the structured rules is defined, including fixed fields and logical association specifications. Fixed fields cover keyword classification fields, intent identifier fields, sentiment state fields, key entity fields, and semantic confidence fields, while logical association specifications clarify the mapping relationships and priority ranking rules between each field.

[0036] Next, various types of unstructured information are processed in a targeted manner: the extracted keywords are classified according to semantic attributes (such as repayment-related, dispute-related, and delay-related) and filled into the keyword classification field; the intent recognition results ("delay", "refusal", "negotiation", "commitment", etc.) are labeled into the intent identifier field; the sentiment analysis results ("calm", "excited", "angry", "frustrated", etc.) are entered into the sentiment state field; key entity information (such as debt amount, repayment time, disputed matters, etc.) in the text are extracted and filled into the corresponding fields; the credibility of semantic recognition is calculated by algorithm, a semantic confidence value is generated and filled into the corresponding field.

[0037] Finally, based on the preset logical association specifications, the correspondence between keyword classification and intent identifier, the influence weight of emotional state on intent priority, etc., are clarified, forming a structured semantic decision feature with a fixed format, complete fields, and clear logical associations. This feature eliminates the fragmentation of unstructured information and presents it in a standardized data format, which facilitates rapid reading, parsing, and invocation by the multimodal feature fusion engine and the strategy decision engine, providing efficient and accurate semantic data support for determining the direction of subsequent communication strategy adjustments.

[0038] Based on the fusion processing of acoustic feature vectors, semantic understanding features, dialogue rhythm features, standardized acoustic assessment features, and structured semantic decision features, a multimodal feature fusion engine and a strategy decision engine negotiate to determine the direction of communication strategy adjustments, the rhythm of speech adjustments, and the strategy switching threshold. Acoustic feature vectors, semantic understanding features, dialogue rhythm features, standardized acoustic assessment features, and structured semantic decision features are input into the multimodal feature fusion engine. This engine comprehensively integrates information from audio, text, and interaction dimensions to form a unified comprehensive assessment result of the debtor's status. The multimodal feature fusion engine transmits the comprehensive assessment result to the strategy decision engine. The two collaborate to determine the direction of communication strategy adjustments based on the debtor's emotional state, repayment willingness, level of participation, and other comprehensive factors, such as changing from "rational disclosure" to "appeasement + guided evidence presentation" or "friendly but firm confirmation." Simultaneously, based on the debtor's emotional score and participation index, the rhythm of speech adjustments (such as speech speed and sentence length) and the strategy switching threshold (such as triggering strategy switching when the emotional score reaches a certain value) are determined.

[0039] S104, process according to the target communication strategy dialogue template to form voice script information that conforms to the attitude tone and communication goals.

[0040] In one implementation, the core information of the preset script template, key entity variables, semantic adaptation rules, and tone adaptation parameters are categorized and integrated according to the attitude tone, content dimension, and target dimension characteristics of the target communication strategy to generate basic script unit information that meets the needs of the scenario. The categorization and integration are carried out based on the core characteristics of the target communication strategy: First, the specific settings of the three core dimensions are clarified. The attitude tone covers three categories: hard, neutral, and friendly. The content dimension includes directions such as legal notification, explanation of credit impact, and provision of installment or reduction solutions. The target dimension focuses on core demands such as obtaining repayment commitments, confirming key information, and resolving debt disputes. Second, the core communication framework in the preset script template is extracted, and key entity variables (such as the promised repayment date, repayment amount, debtor's name, debt contract number, etc.), semantic adaptation rules (i.e., the adaptation mechanism that adjusts the script expression logic according to the semantic scenario expressed by the debtor to ensure semantic consistency), and tone adaptation parameters (corresponding to the tone intensity level, speech rate fluctuation range, pause interval duration, etc. for different attitude tones) are simultaneously sorted out. Finally, based on the specific communication scenario requirements, the three types of dimensional characteristics are systematically classified and integrated with the script templates, key entity variables, semantic adaptation rules, and tone adaptation parameters to form basic script unit information that can be directly called. Each unit is precisely matched to a specific communication scenario (such as a scenario of comforting a debtor when they are angry, or a scenario of confirming when someone intends to delay), ensuring a high degree of adaptation between the script and the scenario.

[0041] Based on the need for effective communication, the content layout of the basic dialogue units was designed, clarifying the core statements, variable placement, and the proportion of tone adjustment parameters, thus generating dialogue unit content design information. Based on the principle of effective communication, the content layout of the basic dialogue units was optimized: the core statements were clearly identified, prioritizing key information such as legal reminders, core repayment plan terms, and credit impact consequences at the beginning of the dialogue to ensure debtors can quickly grasp the core content; the placement of key entity variables was precisely marked to ensure logical coherence and natural connection between variables and the context of the dialogue, avoiding gaps in expression; and the proportion of tone adjustment parameters was reasonably allocated to ensure that their application aligns with the current tone without interfering with the clear delivery of core information, balancing the effectiveness of tone expression and information delivery. Through the above design, the generated dialogue unit content design information ensures that the overall logic of the dialogue is clear and the key points are highlighted, significantly improving the debtor's efficiency in receiving and understanding key information.

[0042] Combining the speech synthesis features of the TTS engine with the need to enhance naturalness, synthesis rules were established to match timbre to strategic tone and adjust speech rate according to semantic emphasis, ensuring that the speech is adapted to the debtor's state. Specifically, regarding timbre matching, appropriate timbres were precisely selected according to strategic tone: a firm and steady tone corresponded to a strong tone to convey a compliance warning attitude; a friendly tone corresponded to a gentle and approachable tone to shorten the communication distance; and a neutral tone corresponded to a calm and objective tone to clearly state the facts. Regarding speech rate adjustment, the speech rate was flexibly adjusted according to semantic emphasis. A slightly slower speech rate was used for key information such as legal clauses, repayment amounts, and promised deadlines to extend the audio presentation time of key information and enhance the delivery effect; a normal speech rate was used for transitional statements and polite greetings to ensure the overall fluency of communication. Simultaneously, leveraging the characteristics of speech synthesis models (such as WaveNet or Tacotron2), a sound optimization mechanism is incorporated. This involves adaptive noise cancellation via the Wiener filter, dynamic range compression to balance volume, and audio equalization to adjust frequency response, further enhancing speech clarity and naturalness. This rule ensures that the synthesized speech closely matches the debtor's current state (emotional state, willingness to repay, etc.), strengthening the approachability and persuasiveness of voice communication.

[0043] The process integrates the basic dialogue units, content design information, and synthesis rules to generate speech script information that includes unit structure, content arrangement, and synthesis specifications. The integration of basic dialogue unit information, content design information, and synthesis rules involves a full-process fusion process: First, based on the structure of the basic dialogue units, the core modules of the script (such as opening greetings, core request delivery, solution explanation, and closing confirmation) and the composition logic of each module are clarified. Second, according to the content arrangement requirements, the sentence structure of the dialogue is organized, key entity variables are accurately filled, and tone adjustment parameters are reasonably allocated to ensure that the speech expression conforms to logical norms and scenario requirements. Finally, following the synthesis rules and combining the characteristics of the TTS engine, the appropriate timbre type and speech rate level are determined, and the execution order of sound optimization processes such as noise cancellation and dynamic range compression is clarified. Through the above full-process integration, complete speech script information containing unit structure, content arrangement, and synthesis specifications is generated. This information not only meets the requirements of the communication strategy but also has a natural and fluent speech presentation effect, providing standardized and highly adaptable script support for subsequent real-time voice interaction.

[0044] S105 transmits voice script information based on real-time interactive rhythm and establishes a call link with the debtor through a third-party telephone platform.

[0045] In one implementation, the system first determines the transmission rhythm of the voice messages based on the results of preliminary multimodal analysis. For debtors who are calm and highly engaged, a normal interaction rhythm is adopted to ensure smooth and efficient communication. For debtors who are emotionally agitated or less engaged, the transmission rhythm is appropriately slowed down, and the interval between key messages is extended to give debtors sufficient time to receive and respond, avoiding resistance caused by an overly fast pace. Simultaneously, the transmission rhythm is dynamically adapted based on the interactive behavior patterns in the dialogue rhythm characteristics to ensure that the voice message transmission matches the debtor's response habits.

[0046] The system establishes a stable connection with a third-party telephone service platform (such as Twilio) via API, completing communication protocol adaptation and authorization authentication to ensure the security and stability of data transmission. Based on the contact information (mobile phone number, landline number) in the debtor's credit data, the system automatically sends a call initiation request to the third-party telephone platform, clearly identifying the two parties (AI collection agent number, debtor number) and the purpose of the call. After receiving the request, the third-party telephone platform completes the call link setup, initiates a call invitation to the debtor, and establishes a real-time voice call link between the AI ​​collection agent and the debtor once the debtor answers.

[0047] Once the call link is established, the system transmits the generated voice script information (including unit structure, content layout, and synthesis specifications) to the third-party telephone platform in real time, following the adapted real-time interaction rhythm. During transmission, the system simultaneously monitors the transmission status to ensure no loss or delay of voice data, guaranteeing the complete presentation of the script. Simultaneously, taking into account the voice transmission characteristics of the third-party telephone platform, the system performs adaptation processing on the voice data to ensure the voice script maintains a clear and natural presentation during transmission, laying the foundation for subsequent real-time interaction between the two parties.

[0048] S106 receives feedback information and analyzes the feedback data and verifies the accuracy of key information in real time. If there is any dispute about the information, a strategy switching mechanism is triggered to adjust the wording until a valid communication result is obtained.

[0049] In one implementation, the debtor's feedback data, along with key entity information, dispute point identifiers, and response timestamps, are categorized and integrated according to the semantic type, emotional characteristics, and interaction pattern characteristics of the feedback information to generate standardized feedback analysis unit information. Emotional characteristics include debtor emotional states such as calmness, excitement, anger, and frustration; interaction pattern characteristics involve communication behavior patterns such as proactive response, passive reply, frequent interruptions, and prolonged silence. Next, key entity information is precisely extracted from the debtor's feedback data, including the specific repayment time, repayment amount, disputed matters, and reasons for debt objections; dispute point identifiers are used to clearly mark content such as disputes over principal, disputes over term interpretation, and disputes over repayment ability in the feedback; response timestamps are recorded to accurately correspond to the specific time point of the feedback information in the call process, achieving a link between feedback and call progress. Finally, according to the classification criteria of semantic type, emotional characteristics, and interaction pattern characteristics, the extracted key entity information, dispute point identifiers, and response timestamps are systematically categorized and integrated to generate standardized feedback analysis unit information. Each unit fully covers the core content and core attributes of a single feedback, ensuring the standardization and completeness of subsequent analysis.

[0050] Based on the need for effective communication, the content layout of the standardized feedback analysis unit is designed, clarifying the weight of core feedback semantics and points of contention, as well as the proportion of information accuracy verification indicators, and generating feedback unit content design information. Core feedback semantics include key expressions that directly reflect the debtor's demands or attitudes, such as "pay 5000 yuan next Friday," "deny the debt," and "hope for installment payments," which are prioritized to help quickly capture the debtor's core intentions. The weight of points of contention is reasonably allocated, using a quantitative approach to assign weights to these points. Core points of contention affecting collection progress, such as disputes over principal and whether the debt is acknowledged, are given higher weights (e.g., weight values ​​of 0.8-1.0), while secondary points of contention, such as details of repayment methods and notification methods, are given lower weights (e.g., weight values ​​of 0.2-0.5), facilitating the focus on key conflicts. The proportion of information accuracy verification indicators is clearly defined. Verification indicators include the consistency between feedback information and original debt data, the completeness of key entity information, and the clarity of semantic expression. Among these, data consistency and information completeness verification account for no less than 70%, ensuring that the verification focus is prominent and comprehensive, avoiding redundant verification. The above design generates feedback unit content design information, making feedback analysis more targeted and efficient, and improving the decision-making efficiency of strategy adjustment.

[0051] Leveraging the dynamic adjustment capabilities of the strategy decision engine and the communication objectives (obtaining repayment commitments, resolving disputes, and confirming information) to achieve the desired outcome, specific operational rules were established: Clearly defining the conditions that trigger strategy switching based on disputed information. When core points of contention emerge in the feedback and the emotional characteristic is "anger," an immediate strategy switch from "rational notification" to "appeasement + guidance on evidence collection" is triggered. When the semantic type is "delayed expression" and the interaction mode is "proactive response," a "friendly but firm confirmation" strategy switch is triggered. Rules for multi-round verification of valid results were established. Key entity information provided by the debtor (such as repayment time and amount) was verified through repeated inquiries, comparison with original debt contract data, and cross-validation with historical communication records to ensure the accuracy of the information. The rules take into account both the controllability and flexibility of the strategy. All strategy switching is carried out within the preset strategy tree framework. The strategy tree nodes cover compliant communication strategies such as reminders, pressure, reassurance, and providing solutions, avoiding deviation from compliant collection requirements and ensuring that the wording adjustments are adapted to the debtor's current state (emotions, feedback semantics, interaction patterns) in real time, thus ensuring the effectiveness and compliance of communication.

[0052] Integrate standardized feedback analysis unit information, feedback unit content design information, and strategy switching rules for full-process fusion processing: First, based on the composition of standardized feedback analysis units, clarify the core modules of feedback analysis, including core semantic extraction, dispute point identification, accuracy verification, interaction pattern analysis, and their logical connections, and construct a clear analysis framework.

[0053] Secondly, following the requirements for the content arrangement of feedback units, the core feedback semantics, the weighted ranking of disputed points, and the accuracy verification results are analyzed, prioritizing the presentation of high-weighted disputed points and key verification conclusions. Finally, based on the strategy switching rules, the direction of the corresponding wording adjustments for the current feedback is clarified, such as changing from "legal notification" to "solution provision," the timing of the adjustment (e.g., adjusting immediately after the current dialogue round ends, adjusting after confirming disputed points), and the verification standards (e.g., confirming whether the debtor accepts the new strategy and whether the key information is accurate). Through the above integrated processing, effective communication result confirmation information, including unit structure, content arrangement, and adjustment specifications, is generated, providing a clear and accurate feedback analysis basis for the strategy decision engine, supporting the dynamic adjustment of subsequent wording and the optimization of communication strategies.

[0054] S107 extracts promised repayment time and amount information from calls, matches corresponding legal document templates to automatically generate compliant documents, and issues them with one click after manual review, completing the entire intelligent collection process.

[0055] In one implementation, core repayment information is accurately extracted from the speech-to-text stream and structured semantic decision features of the call. First, the promised repayment time information is captured using named entity recognition technology to identify the specific repayment date and time period mentioned by the debtor, such as "next Friday" or "the 10th of next month," and automatically converted to a standard date format. Second, the promised repayment amount information is extracted, showing the debtor's stated repayment amount, such as "pay 1,000 yuan upfront" or "pay in three installments of 5,000 yuan each," and compared with the original debt contract amount for consistency verification. If the extracted repayment amount differs from the total outstanding amount, the repayment time is unclear, or key elements are missing, the system automatically marks it as abnormal information and prompts further confirmation to ensure the accuracy and completeness of the extracted information.

[0056] The system pre-sets various standardized legal document templates, including "Repayment Commitment Letter," "Installment Repayment Agreement," and "Collection Notice." Each template contains fixed modules such as party information, debt details, commitment clauses, and explanations of legal validity, as well as dynamic variable fields. The system automatically matches templates based on extracted key repayment information and the call context: if the debtor explicitly commits to repayment, the "Repayment Commitment Letter" template is matched; if an installment repayment agreement is reached, the "Installment Repayment Agreement" template is matched. Subsequently, the system automatically populates the corresponding dynamic variable fields in the templates with extracted structured information such as the debtor's name, ID number, committed repayment time, amount, and debt contract number. For descriptive paragraphs such as "Facts and Reasons," the system uses NLG technology combined with key call information (such as communication time and the agreed repayment terms) to generate fluent and compliant text, thus automatically generating personalized compliant documents.

[0057] After generating the initial draft of the document, the system pushes it to the review portal of the collection agent or legal personnel, simultaneously linking it to corresponding call recordings, timestamps of key statements, and information verification records for reviewers' reference. Reviewers verify the core content of the document, including party information, repayment agreements, and legal clauses, and can make minor adjustments to address issues such as non-standard wording or missing information. All modification records are automatically saved for tracking and auditing. Once approved, staff can trigger the document issuance operation with a single click through the system. The system supports multiple methods of sending, including email, SMS, and traditional postal services, and automatically records the document's sending time, method, and receipt status, forming a closed loop and completing the entire intelligent collection process.

[0058] like Figure 2 As shown, an intelligent debt collection device based on audio analysis includes: The acquisition module 201 is used to acquire basic data and parameter information related to debt collection, including debtor credit data, legal document templates, as well as the type ID, associated debt ID, priority ID, collection node ID and communication mode of each data. The data processing module cleans, verifies and stores the basic data in a structured manner. The feature extraction module 202 is used to collect the debtor's call audio stream in real time through ASR technology, extract audio frame data, voice activity detection results, speaker separation timestamps, and synchronously convert them into a time-stamped text stream; The voice information generation module 203 is used to parse audio and text data, classify and process them according to semantic, acoustic features and dialogue rhythm dimensions, and determine the direction of communication strategy adjustment by the multimodal feature fusion engine and the strategy decision engine based on the volume, speech rate and pause features of the audio, and the keywords and intent information of the text. It also determines the rhythm of the speech adjustment and the threshold for strategy switching based on the debtor's emotion score and participation index. The speech is processed according to the target communication strategy dialogue template to form voice speech information that conforms to the attitude tone and communication goals. The voice communication module 204 is used to transmit voice script information based on real-time interactive rhythm, establish a call link with the debtor through a third-party telephone platform; receive feedback information, and analyze feedback data and verify the accuracy of key information in real time. If there is a dispute over the information, a strategy switching mechanism is triggered to adjust the script until a valid communication result is obtained; extract the promised repayment time and amount information in the call, match the corresponding legal document template to automatically generate compliant documents, and send them out with one click after manual review, completing the entire process of intelligent collection.

[0059] A computing device includes a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the device is triggered to execute any intelligent collection method based on audio analysis.

[0060] The methods and / or embodiments in this application can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowchart. When the computer program is executed by a processing unit, it performs the functions defined in the methods of this application.

[0061] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0062] It will be apparent to those skilled in the art that this application is not limited to the details of the exemplary embodiments described above, and that this application can be implemented in other specific forms without departing from the spirit or essential characteristics of this application. Therefore, the embodiments should be regarded as exemplary and non-limiting in all respects, and the scope of this application is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within this application.

Claims

1. A smart debt collection method based on audio analysis, characterized in that, include: Acquire basic data and parameter information related to debt collection, including debtor credit data, legal document templates, as well as the type ID, associated debt ID, priority ID, collection node ID, and communication mode of each data. The data processing module cleans, verifies, and stores the basic data in a structured manner. The ASR technology is used to collect the debtor's call audio stream in real time, extract audio frame data, voice activity detection results, and speaker separation timestamps, and convert them into a time-stamped text stream simultaneously; The audio and text data are analyzed and classified according to semantic, acoustic features and dialogue rhythm dimensions. Based on the volume, speech rate and pause features of the audio, and the keywords and intent information of the text, the multimodal feature fusion engine and the strategy decision engine negotiate to determine the direction of communication strategy adjustment. The rhythm of speech adjustment and the threshold of strategy switching are determined based on the debtor's emotional score and participation index. Process the dialogue script template according to the target communication strategy to form voice script information that conforms to the attitude tone and communication goals; Based on real-time interactive rhythm, voice script information is transmitted, and a call link is established with the debtor through a third-party telephone platform; Receive feedback information and analyze the feedback data and verify the accuracy of key information in real time. If there is any dispute about the information, trigger the strategy switching mechanism to adjust the wording until an effective communication result is obtained. Extract promised repayment time and amount information from the call, match the corresponding legal document template to automatically generate compliant documents, and send them out with one click after manual review, completing the entire intelligent collection process.

2. The intelligent debt collection method based on audio analysis according to claim 1, characterized in that, The system uses ASR technology to collect the debtor's call audio stream in real time, extracts audio frame data, voice activity detection results, and speaker separation timestamps, and synchronously converts them into a timestamped text stream, including: Based on the needs of audio analysis and real-time interaction, the audio of debtors' calls is processed. The ASR speech-to-text mechanism is used to extract audio frame data, speech activity detection results, and core audio features of speaker separation timestamp. The speech signal preprocessing module encapsulates the data at the PCM level according to the PCM data specification. Based on the multimodal analysis rules, the audio feature type, semantic extraction accuracy, and timestamp synchronization format are extracted and processed to generate audio frame data packets, speech-to-text streams, feature encoding mapping information, and audio-text synchronization verification information in a unified format. Based on the audio-text synchronization verification information, the generated information is verified bidirectionally. A real-time supplementary acquisition mechanism is initiated for missing audio frames. For abnormal transcribed text, the semantic correction service is called for secondary transcription. For data with out-of-sync timestamps, a synchronization calibration process is triggered. An optimized audio processing scheme including anomaly handling strategies and supplementary acquisition time windows is generated. The optimized processing scheme is integrated and executed, responding to real-time call analysis commands. At the same time, the call start time is bound through the time synchronization module, generating a text stream with a unified timestamp and audio feature association interval information.

3. The intelligent debt collection method based on audio analysis according to claim 2, characterized in that, The audio and text data are analyzed and classified according to semantic, acoustic features, and dialogue rhythm dimensions. Based on the audio's volume, speech rate, and pause features, and the text's keywords and intent information, a multimodal feature fusion engine and a strategy decision engine negotiate to determine the direction of communication strategy adjustments. Furthermore, the pace of dialogue adjustments and strategy switching thresholds are determined based on the debtor's emotional score and participation indicators, including: The audio data is parsed and processed to extract acoustic features such as volume, speech rate, and pauses, and to generate acoustic feature vectors. Perform semantic analysis on text data, extract keywords and intent recognition results, and generate semantic understanding features; Rhythm analysis is performed on call interaction data to extract participation indicators and interaction behavior patterns, generating dialogue rhythm features; acoustic feature vectors are quantified to generate standardized acoustic evaluation features. The semantic understanding features are processed in a structured manner to generate structured semantic decision features; Based on the fusion processing of acoustic feature vectors, semantic understanding features, dialogue rhythm features, standardized acoustic evaluation features, and structured semantic decision features, the multimodal feature fusion engine and the strategy decision engine negotiate to determine the direction of communication strategy adjustment, the rhythm of speech adjustment, and the strategy switching threshold.

4. The intelligent debt collection method based on audio analysis according to claim 1, characterized in that, Process the dialogue script template according to the target communication strategy to form voice script information that conforms to the attitude tone and communication objectives, including: Based on the attitude tone, content dimension, and target dimension characteristics of the target communication strategy, the core information of the preset script templates, key entity variables, semantic adaptation rules, and tone adaptation parameters are classified and integrated to generate basic script unit information that meets the needs of the scenario. Based on the needs of communication effectiveness, the content layout of the basic units of the dialogue script is designed, clarifying the core dialogue expressions and variable filling positions, the proportion of tone adjustment parameters, and generating dialogue unit content design information. Combining the speech synthesis characteristics of the TTS engine with the need to improve naturalness, we set synthesis rules to match the timbre according to the strategy tone and adjust the speech rate according to the semantic focus to ensure that the speech is adapted to the debtor's state. The results of integrating basic dialogue units, content design information, and synthesis rules are processed to generate speech dialogue information that includes unit composition, content arrangement, and synthesis specifications.

5. The intelligent debt collection method based on audio analysis according to claim 1, characterized in that, Receive feedback information and analyze the feedback data and verify the accuracy of key information in real time. If there is any dispute regarding the information, a strategy switching mechanism is triggered to adjust the wording until a valid communication result is obtained, including: Based on the semantic type, emotional characteristics, and interaction mode of the feedback information, the debtor's feedback data, key entity information, dispute point identifiers, and response timestamps are classified and integrated to generate standardized feedback analysis unit information. Based on the need for effective communication, the content layout of the standardized feedback analysis unit is designed, clarifying the weight of core feedback semantics and points of contention, the proportion of information accuracy verification indicators, and generating feedback unit content design information. By combining the dynamic adjustment features of the strategy decision engine with the needs of achieving communication goals, operational rules are set to trigger strategy switching based on disputed information and verify and confirm effective results through multiple rounds of checks, ensuring that the wording adjustments are adapted to the debtor's status in real time. The feedback analysis unit integrates the results, content design information, and strategy switching rules, and generates effective communication result confirmation information that includes unit structure, content layout, and adjustment specifications.

6. A smart debt collection device based on audio analysis, characterized in that, The device includes: The acquisition module is used to acquire basic data and parameter information related to debt collection, including debtor credit data, legal document templates, as well as the type ID, associated debt ID, priority ID, collection node ID and communication mode of each data. The data processing module cleans, verifies and stores the basic data in a structured manner. The feature extraction module is used to collect the debtor's call audio stream in real time using ASR technology, extract audio frame data, voice activity detection results, speaker separation timestamps, and synchronously convert them into a time-stamped text stream; The voice information generation module is used to parse audio and text data, classify and process them according to semantic, acoustic features and dialogue rhythm dimensions, and determine the direction of communication strategy adjustment by the multimodal feature fusion engine and the strategy decision engine based on the volume, speech rate and pause features of the audio, and the keywords and intent information of the text. It also determines the rhythm of the speech adjustment and the threshold for strategy switching based on the debtor's emotion score and participation index. The speech is processed according to the target communication strategy dialogue template to form voice speech information that conforms to the attitude tone and communication goals. The voice communication module is used to transmit voice scripts based on real-time interactive rhythm, establish a call link with debtors through a third-party telephone platform, receive feedback information, and analyze feedback data and verify the accuracy of key information in real time. If there is a dispute, a strategy switching mechanism is triggered to adjust the script until a valid communication result is obtained. The module extracts the promised repayment time and amount information from the call, matches the corresponding legal document template to automatically generate compliant documents, and sends them out with one click after manual review, completing the entire intelligent collection process.

7. An electronic device, characterized in that, include: First processor; and memory for storing executable instructions of the first processor; The first processor is configured to execute the intelligent collection method based on audio analysis according to any one of claims 1 to 5 by executing the executable instructions.

8. A computing device, the device comprising a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein, When the computer program instructions are executed by the processor, the device is triggered to execute the intelligent collection method based on audio analysis as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent collecting robot and method based on intention recognition and finite state automata

    CN109949805A

  • Multi-mode voice call information extraction method and system

    CN114974294A

  • Legal document generation system based on intelligent voice recognition

    CN118095226A

  • Collection strategy generation and execution system and method based on debtor portrait driving

    CN120833213A