Call service quality auditing method and device, equipment and medium

By collecting and analyzing call audio data, and using intent, emotion and service response models to generate audit reports, the problem of insufficient identification of customer intention and emotion status in the existing technology is solved, and the service quality is automated and multi-dimensional audit is realized, and audit efficiency and accuracy are improved.

CN120499313APending Publication Date: 2025-08-15CHINA PING AN LIFE INSURANCE CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510937816.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

The existing technology cannot effectively identify customer intentions, emotional states and service response behaviors during calls, resulting in a lack of depth and accuracy in service quality audits, and it is difficult to support the intelligent, standardized and compliant development of financial and medical telephone service processes.

Method used

Call audio data is collected and transliterated into text data, and in-depth analysis is used to generate customer intention results, emotion recognition results and service response analysis results, and finally an audit report containing service quality analysis is generated.

Benefits of technology

It realizes automated, multi-dimensional, in-depth audit analysis of service quality, improves audit efficiency and accuracy, and reduces the risk of manual burden and subjective deviation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120499313A_ABST
    Figure CN120499313A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a call service quality auditing method, device, equipment and medium, and the method comprises the steps: collecting call audio data, and translating the call audio data into call text data; identifying a client intention in the call text data by using an intention identification model, and generating a client intention result; using an emotion recognition model to recognize customer emotion, and generating an emotion recognition result; using the service response analysis model to identify service behavior performance, and generating a service response analysis result; and generating an audit report containing service quality analysis content based on the client intention result, the emotion recognition result and the service response analysis result. According to the method, the client intention, the emotional state and the service response behavior in the call content are processed in a unified manner, the unstructured call information is converted into the structured evaluation data, automatic, multi-dimensional and deep auditing analysis of the service quality is realized, and the auditing efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a call service quality audit method, device, equipment and storage medium. Background Art

[0002] In the fintech sector, the process management of in-person appointment services often relies on telephone channels for customer guidance and information confirmation. However, existing technical methods often rely on manual listening and recording, resulting in low overall service efficiency. Customers' true intentions and feedback expressed over the phone often require operators to listen to and organize each individual customer, which is not only labor-intensive and costly, but also prone to recording errors due to subjective understanding or lack of attention, affecting subsequent online service processing and quality assessment.

[0003] In the healthcare sector, patients often face a lack of structured documentation of service responses when registering appointments, receiving medical guidance, or making online follow-up appointments by phone. This is especially true for follow-up appointments for chronic conditions and special examinations. Standardized service delivery and accurate understanding of patient intent are crucial for subsequent service arrangements. However, the current lack of effective intent recognition and response recording mechanisms in these telephone service processes makes it difficult to support comprehensive service quality audits and standardized management.

[0004] Current, widely used telephone service audit mechanisms, whether in finance or healthcare, lack the ability to deeply understand customer intent, emotional state, and service response quality during calls. Existing solutions often rely solely on keyword triggering or fuzzy rule matching, making it difficult to identify hidden service issues, customer dissatisfaction, or failed guidance during calls. Furthermore, call data is typically stored in audio format, lacking the ability to be processed and structured in text. This prevents audit systems from efficiently extracting and analyzing key information from the service process.

[0005] More seriously, current systems generally lack quantitative evaluation criteria for service quality, failing to effectively integrate multi-dimensional factors such as customer intent, emotional reactions, and service personnel response behavior to form a unified basis for quality analysis. This not only restricts the feasibility of service optimization but also makes it difficult for related audit work to adapt to the needs of regulatory compliance and refined internal management.

[0006] Therefore, the inadequacy of existing technologies in areas such as voice data transcription, customer intent recognition, emotional state assessment, and service response process auditing has become a major bottleneck hindering the intelligent, standardized, and compliant development of financial and medical telephone service processes. More efficient technologies are urgently needed to achieve structured understanding and in-depth analysis of call content, thereby enhancing the comprehensiveness and intelligence of service quality audits. Summary of the Invention

[0007] The main purpose of the present invention is to provide a call service quality audit method, device, equipment and storage medium, aiming to solve the technical problem that the existing technology cannot uniformly extract and analyze customer intentions, emotional states and service response behaviors during calls, resulting in a lack of depth and accuracy in service quality audits.

[0008] To achieve the above object, the present invention provides a call service quality audit method, comprising:

[0009] Collecting call audio data and performing text transcription on the call audio data through a speech transcription module to generate call text data;

[0010] Performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result;

[0011] Performing emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result;

[0012] Performing service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0013] Based on the customer intention results, emotion recognition results and service response analysis results, an audit report including service quality analysis results is generated.

[0014] Furthermore, to achieve the above-mentioned purpose, the present invention provides a call service quality auditing device, comprising:

[0015] A speech transcription module is used to collect call audio data and perform text transcription on the call audio data through the speech transcription module to generate call text data;

[0016] An intent analysis module, configured to perform intent recognition processing on the call text data based on an intent recognition model and generate a customer intent result;

[0017] An emotion recognition module, configured to perform emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result;

[0018] a response analysis module, configured to perform service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0019] The audit generation module is used to generate an audit report including service quality analysis results based on the customer intention results, emotion recognition results and service response analysis results.

[0020] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a call service quality audit program stored in the memory and runnable on the processor. When the call service quality audit program is executed by the processor, the steps of the call service quality audit method described above are implemented.

[0021] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a call service quality audit program is stored. When the call service quality audit program is executed by a processor, the steps of the call service quality audit method described above are implemented.

[0022] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. A call service quality audit method, device, equipment and medium are disclosed, including: collecting call audio data and transcribing it into call text data; using an intention recognition model to identify customer intentions in call text data and generate customer intention results; using an emotion recognition model to identify customer emotions and generate emotion recognition results; using a service response analysis model to identify service behavior performance and generate service response analysis results; based on the customer intention results, emotion recognition results and service response analysis results, an audit report containing service quality analysis content is generated. The present invention uniformly processes customer intentions, emotional states and service response behaviors in call content, converts unstructured call information into structured evaluation data, realizes automated, multi-dimensional and in-depth audit analysis of service quality, improves audit efficiency and accuracy, and reduces manual burden and subjective bias risks. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0024] Figure 1 A schematic diagram of an application environment of a call service quality audit method according to an embodiment of the present invention;

[0025] Figure 2 This is a flow chart of an embodiment of a method for auditing call service quality according to the present invention;

[0026] Figure 3 A schematic diagram of the functional modules of a preferred embodiment of a call service quality auditing device of the present invention;

[0027] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0028] Figure 5 FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0029] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0030] The call service quality audit method provided by the embodiment of the present invention can be applied in the following situations: Figure 1 In an application environment, the user end communicates with the server end through a network. The server end can collect call audio data through the user end and transcribe it into call text data; use the intention recognition model to identify the customer intention in the call text data and generate a customer intention result; use the emotion recognition model to identify the customer emotion and generate an emotion recognition result; use the service response analysis model to identify the service behavior performance and generate a service response analysis result; based on the customer intention result, the emotion recognition result and the service response analysis result, generate an audit report containing service quality analysis content. The present invention converts unstructured call information into structured evaluation data by uniformly processing the customer intention, emotional state and service response behavior in the call content, thereby realizing automated, multi-dimensional and in-depth audit analysis of service quality, improving audit efficiency and accuracy, and reducing manual burden and subjective bias risk. Among them, the user end can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server end can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0031] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of a call service quality audit method provided by the present invention. It should be noted that although a logical order is shown in the flow chart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0032] like Figure 2 As shown, the call service quality audit method proposed by the present invention includes the following steps:

[0033] S10, collecting call audio data, and performing text transcription on the call audio data through a voice transcription module to generate call text data;

[0034] In this embodiment, obtaining high-quality voice information is the foundation for subsequent semantic analysis and behavioral assessment. To effectively collect voice data, an audio data capture mechanism must be established within the telephone communication link. This mechanism can intercept bidirectional audio streams—that is, both the customer-side and service-side voice—while maintaining the integrity and time synchronization of the original data. This audio data can be collected through an embedded recording interface, a digital trunk signal parsing module, or integrated through a front-end telephone gateway. Audio data not only contains speech content but may also contain background noise, echo feedback, and silence. Therefore, after data collection, it must undergo a preprocessing phase.

[0035] During the purification process, the audio data undergoes noise suppression to eliminate non-target speech elements. Background noise suppression can be achieved using spectral subtraction, beamforming, or deep learning noise reduction models. Echo cancellation relies on return path modeling algorithms, such as adaptive filtering or full-duplex echo cancellation architectures, to eliminate system echo components. This processed audio data achieves higher speech intelligibility and semantic integrity, contributing to the accuracy of subsequent text generation.

[0036] To address the issue of multiple speakers alternating during a call, speaker separation is required to improve the accuracy of subsequent intent and emotion recognition. This involves extracting the audio from the service staff and the customer separately. Speaker separation can be achieved using methods such as time-frequency masking, nested voiceprint clustering, and speech segmentation models. Speech embedding vectors combined with clustering algorithms can effectively identify segments from different speakers and decouple multi-channel speech data.

[0037] The single speaker's speech that has undergone the above processing is input into different speech recognition engines respectively. The service staff's speech can be recognized by speech recognition model A, and the customer's speech can be recognized by model B. The recognition engine is usually based on a deep neural network framework, combining acoustic modeling and language modeling modules to perform character-level or word-level decoding. During the decoding process, it is necessary to combine a domain-adaptive language model to improve the accuracy of professional terminology recognition. The two segments of recognized text data are respectively annotated with role labels, and the final transcription text is constructed with reference to the alignment of the call timeline. The text sequence should contain timestamps, role identifiers, and traceable indexes to facilitate the subsequent analysis and processing modules to locate key segments and behavior segments.

[0038] Voice collection mechanisms deployed in telephone communication systems can automatically activate recording stream routing upon call initiation by connecting to a PBX (Private Branch Exchange) system. Alternatively, they can be deployed in VOIP (Voice over IP) systems to intercept voice streams using protocols such as SIP or RTP. During audio purification, dual-channel microphone signals can be introduced as reference channels to enhance echo cancellation capabilities, or deep learning models (such as the Wavenet noise reduction architecture) can be deployed to adapt to high-noise environments.

[0039] For speaker separation, when both the customer and the agent are speaking continuously, time-window-based VAD (Voice Activity Detection) preprocessing can be used to reduce speech overlap. The voiceprint recognition module uses a pretrained ECAPA-TDNN embedding model combined with K-means clustering for real-time speaker separation. The speech recognition engine can load customized corpora, such as service line terminology lists, sentiment dictionaries, and common Q&A templates, for model fine-tuning.

[0040] For text alignment, a dynamic time warping (DTW) algorithm can be used to synchronize the timelines of text across different channels to avoid semantic misalignment. A turn-based detection mechanism can also be introduced to complete the conversation structure. The resulting text should have a unique identifier field to facilitate integration with emotion recognition, intent recognition, and behavior analysis processes.

[0041] Example: In a healthcare scenario, a patient calls a hotline to discuss their treatment needs and medical history. The conversation includes a variety of medical terminology and emotional fluctuations. The system uses high-precision speech recognition to convert the call content into text data with a timeline. This allows subsequent analysis modules to identify whether the patient clearly expressed their intentions, whether they expressed anxiety or fear, and whether the service staff followed the consultation process.

[0042] In financial service scenarios, customers inquire about loan progress or report account issues over the phone. The system performs voice separation and text transcription on the collected multiple rounds of conversations to ensure that each round of responses is accurately restored to structured content. This facilitates subsequent analysis of whether the service response is timely, whether the language is compliant, and whether the customer has high-risk emotions, thereby effectively supporting compliance audits and service quality assessments.

[0043] This embodiment effectively converts raw audio data into structured text data through real-time collection, voice purification, speaker separation, and high-precision speech transcription. This text data, serving as the unified input for downstream analysis modules, not only ensures semantic integrity and temporal accuracy but also significantly improves the accuracy and consistency of subsequent recognition of customer intent, emotional state, and service response behavior, reducing the burden of manual analysis and increasing the efficiency of automated processing.

[0044] S20, performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result;

[0045] In this embodiment, after the call text data is aligned with the role identification and timeline, further semantic intent information must be extracted to assist in understanding the service needs, questions, or operational goals expressed by the customer during the call. This process first requires extracting the text segments corresponding to the customer role from the call text data. These segments are the natural language sentences uttered by the customer throughout the call. This content can be filtered using the role identification tags and assembled into the customer conversation text. Customer conversation text refers to continuous or discontinuous semantic units uttered by the customer, with temporal order and contextual dependencies. Its completeness is crucial for the accuracy of intent recognition.

[0046] Performing semantic feature extraction on customer conversation text converts natural language content into a vector representation that can be processed by the model. Semantic feature extraction methods may include but are not limited to text encoding models based on the Transformer architecture, context-based word embedding technologies (such as BERT and RoBERTa), or structural embedding mechanisms combined with intent knowledge graphs. During the feature extraction process, in addition to basic semantic embedding vectors, enhanced features such as language style, temporal position information, frequency of negative words, and sentiment trends can also be introduced to improve the intent classification model's ability to distinguish different intent types.

[0047] The generated intent feature vector, as a high-dimensional representation, is input into the intent recognition model for classification. This model, typically based on a multi-classification architecture, outputs probability values for multiple pre-set intent categories, representing the degree to which the text fits each intent label. The initial intent classification result is the confidence output for each category, which is used to characterize the service type, customer request, or operational purpose that the current semantic unit may correspond to. To avoid semantic ambiguity in short texts or biased judgment of isolated utterances, further temporal analysis is required in conjunction with the conversation context. By modeling the preceding and following utterances in the context window, the coherence and consistency of the intent can be calculated, thereby adjusting the confidence score of the initial classification result.

[0048] The final customer intent result is based on the initial intent classification results, excluding low-confidence categories below a preset threshold and retaining only high-confidence outputs, ensuring the stability and relevance of the final recognition results to actual business needs. This confidence threshold can be dynamically adjusted based on business fault tolerance requirements or intent type differences, such as setting a higher confidence threshold for high-risk intents while allowing a certain amount of ambiguity for common business intents. The final output is a clear intent label or set of labels, which can be further accompanied by a probability score, corresponding sentence index, and recognition timestamp.

[0049] During customer text extraction, a rule engine or model segmentation strategy can be used to distinguish between customer and service personnel speech. Furthermore, a short sentence aggregation mechanism can be used to stitch discontinuous speech fragments into complete semantic blocks to reduce information fragmentation. Semantic feature extraction can load a pre-trained BERT model and fine-tune it based on business corpus to enhance semantic understanding capabilities within specific service areas. Alternatively, cross-modal knowledge distillation can be used to map existing business rules into feature vectors for model training.

[0050] Intent recognition models can employ fully connected neural networks, enhance models with attention mechanisms, or incorporate conditional random fields (CRF) structures to enhance semantic boundary perception. When identifying multiple intents, a sigmoid activation function can be used to achieve multi-label output, and methods such as beam search can be used to generate candidate intent paths. Confidence calculations can be based on the output distribution of the cross-entropy loss or weighted corrections based on historical call similarity to improve robustness in scenarios with weak semantics or ambiguous expressions.

[0051] Example: In healthcare service scenarios, when patients make appointments or seek consultations over the phone, their intent may involve different service needs, such as registration application, fee inquiry, medical record copying, and doctor change. The system can automatically identify the type of intent expressed by the patient in different sentences and, based on the context of the preceding and following conversations, determine whether the patient has completed the appointment path.

[0052] In financial services scenarios, customers may express various requests during phone calls, such as loan applications, credit limit adjustments, repayment plan changes, and service complaints. The system can automatically identify the actual business intent from the customer's statements, such as identifying risk concerns or service objections implied in the customer's questions. This can then prompt service personnel to provide targeted responses and audit records, effectively improving the accuracy and assessability of service responses.

[0053] This embodiment automatically identifies the customer's true intention during a call by extracting and converting their sentences into semantic feature vectors, then combining them with contextual analysis and intent classification model processing. This effectively replaces the manual analysis process and improves the efficiency and accuracy of intent recognition. In multi-round conversation scenarios, combining temporal sequence and contextual continuity judgment can reduce recognition errors and improve model stability and applicability. The output of customer intent provides a structured semantic foundation for subsequent behavior evaluation, service adaptation, and anomaly monitoring.

[0054] S30, performing emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result;

[0055] In this embodiment, the call text data contains information about the customer's emotional state during the service process. This information reflects the customer's subjective response to the service process, the service staff's performance, and the overall experience. To extract effective emotional state features from the call text data, the text content corresponding to the customer's role must first be filtered out to form the customer emotion text. The customer emotion text is a collection of continuous speeches with the customer as the role subject. Its chronological order, contextual coherence, and word choice play a decisive role in the accuracy of emotion recognition. This screening can be accomplished through role tagging or text aggregation combined with speech frequency and contextual consistency strategies.

[0056] Based on the customer's emotional text, further linguistic and semantic features expressing emotional states need to be extracted. This emotional feature extraction can include, but is not limited to, linguistic features such as part-of-speech distribution, emotional word frequency, sentence complexity, intonation markers, interrogative and negative structures, and the frequency of hyperbole. Semantic features such as contextual word weights, emotional color dictionary matching, and the degree of emotional tension in the language can also be incorporated. The extracted results are encoded into a multidimensional emotional feature vector, where each dimension represents a different type of emotional signal, including but not limited to anger, anxiety, disappointment, satisfaction, confusion, and other states.

[0057] This emotion feature vector is input into a pre-set emotion recognition model. The model architecture can be a multi-layer neural network based on emotion label classification or a text sequence-level Transformer architecture. The recognition output is a set of initial emotion classification results, reflecting the dominant emotion type favored by the customer in the current text. Furthermore, to avoid the compression of subtle emotional variations by a single classification label, the identified emotion labels are further quantified. Emotional intensity quantification can be calculated based on maximum probability values, entropy distribution, language emotion density, historical customer behavior models, and other methods to produce an emotion intensity score representing the amplitude of emotional fluctuations or the degree of negative emotional deviation.

[0058] Finally, the initial emotion classification results are combined with the emotion intensity scores to form a structured emotion recognition result. This recognition result can be used to label the emotion category and emotional fluctuation level corresponding to each customer speech paragraph, supporting subsequent work such as service sensitivity analysis, potential complaint risk prediction, and emotional response adaptation of service personnel.

[0059] During the text screening phase, customer paragraphs can be extracted using a role recognition model combined with call segmentation results. Alternatively, customer speech text can be directly aggregated using structured transcription data with speaker labels. The sentiment feature extraction module can load multilingual sentiment dictionaries, universal sentiment templates, and syntax tree matching rules to generate multi-layered sentiment signals. For sentiment assessments involving industry-specific terminology or cultural differences in expression, transfer learning can be used to retrain the sentiment recognition model on industry conversation data.

[0060] Emotion recognition models can adopt a dual-channel architecture: one channel processes the sentiment word embeddings, while the other processes the context state vector. The outputs of these channels are then fused in an attention layer to improve recognition accuracy. Emotion intensity scores can be mapped to the original categorical distribution using a sigmoid function, or clustering can be used to learn a scale for emotion magnitude based on historical sentiment annotation data. The final results can be formatted as a JSON structure, annotated with the emotion type, confidence level, intensity level, and corresponding speech text interval.

[0061] Example: In healthcare services, patients often express anxiety, worry, and urgency during phone consultations or appointments. The system can automatically identify and quantify the intensity of negative emotions expressed by patients when they mention difficulties registering, long wait times, or poor service attitudes. This can then prompt customer service or doctors to intervene promptly, improving the patient experience.

[0062] In financial services, when customers discuss account anomalies, financial losses, interest rate changes, and other issues over the phone, they often express doubt, disappointment, or distrust. The system can automatically detect these emotional states and use them as customer satisfaction risk signals, providing a basis for subsequent complaint prevention and service quality tracking, helping financial institutions strengthen customer relationship management and risk prevention.

[0063] This embodiment integrates natural language expressions from customer calls with contextual features into a model and feeds this into an emotion recognition model. This allows for efficient extraction of customer emotional states without relying on emotion labeling, significantly improving the automation of customer feedback analysis. Combined with emotion intensity calculation methods, this method can also identify the magnitude of emotional fluctuations and the potential accumulation of negative emotions, playing a key role in improving the accuracy of service quality monitoring and risk early warning capabilities. The emotion recognition results provide a high-quality data foundation for subsequent service strategy evaluation and personnel performance management.

[0064] S40, performing service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0065] In this example, the call text data contains key information such as the service personnel's language expression, response speed, and problem-solving strategies during interactions with customers. To extract elements for evaluating service quality from this unstructured text, the sentences corresponding to the service personnel's roles must first be identified and aggregated into service personnel response text. This process typically relies on the speaking role labels carried in the transcription results or a segmented aggregation strategy based on contextual judgment, ensuring that subsequent analysis focuses solely on the service provider's language behavior.

[0066] The primary analysis target for agent response text is its response time. Extracting these features involves calculating the agent's first response time, the answer interval between each conversation turn, and the average response delay for the entire call based on the transcribed timestamp information. This information reflects the timeliness of the service response and helps determine the efficiency of the service response.

[0067] Furthermore, the structure and completeness of the service personnel's responses reflect their adherence to process standards. Using predefined standard script templates or keyword sets, we can identify whether the service personnel have used required service process elements such as greetings, confirmations, introductions, and summaries. Based on their coverage and the degree of alignment with the sequential structure, we calculate a script completeness score, reflecting service standards.

[0068] In service interactions, problem resolution is a key measure of service effectiveness. To achieve this, we identify semantic segments within the text that include problem identification, solution suggestions, and status confirmation. We extract and construct a sequence of problem-solving steps, combining these steps with their respective timestamps to generate a structured problem-solving path. We then count the frequency of logical connectives within this path, such as "therefore," "next," "next," and "finally," to assess whether the service process possesses a clear and coherent logical structure, which we then use to calculate a problem-solving efficiency index.

[0069] The above-mentioned multiple-dimensional information includes response delay time data, speech completeness score and problem-solving efficiency index. The three together constitute the service response analysis results, providing basic support for service audit, performance evaluation and process optimization.

[0070] Agent response text can be obtained by merging all text tagged with the agent role in the speech transcription results, or by using semantic analysis of each conversational turn to identify the dominant speaker. Response time can be extracted by calculating the difference between end-to-end timestamps, or by constructing a timeline of agent speech starts and ends to calculate average latency.

[0071] Script integrity analysis can be implemented using a template matching algorithm, defining a set of standardized script templates with different script nodes corresponding to each service scenario. For example, in financial customer service, account changes must include identity confirmation, risk warnings, and operation confirmation. In medical services, initial appointments must include symptom confirmation, registration instructions, and follow-up appointment reminders. The system can deduct points for completeness based on missing nodes.

[0072] Problem-solving efficiency can be assessed by combining contextual semantic rules to determine the logical order of identified semantic action label segments, and then comprehensively evaluating them using total path duration and connective density. This process can be constructed as a temporal logic network model to quantify service continuity and timeliness.

[0073] Example: In healthcare services, the system can automatically identify whether the doctor's assistant completes identity verification, registration time reminders, and risk explanations according to standard procedures during the appointment call, and evaluate whether their response is timely and their suggestions are clear, making it easier for hospital managers to identify process gaps or service bottlenecks.

[0074] In telephone service scenarios in the financial field, when customer service handles credit card freezes or large-value transaction verification requests, whether the customer service response includes key nodes such as risk warnings, customer confirmation, and operating instructions will be automatically extracted by the system and a complete problem-solving path map will be generated. The response delay and speech integrity scoring model will assist in discovering service behaviors with a high probability of complaints, providing structured support for service improvement and increased customer satisfaction.

[0075] By structuring service personnel responses into multi-dimensional features, this embodiment accurately assesses the timeliness of responses, the accuracy of their wording, and the effectiveness of problem resolution. Response latency characterizes the immediacy of service, wording integrity reflects the consistency of the service process, and problem-solving efficiency measures the effectiveness of service outcomes. The combined output of these three factors eliminates the need for audits of service behavior from relying solely on subjective judgment. Instead, it relies on a data-driven quantitative indicator system, improving analytical accuracy and automation, and providing a scientific basis for service process optimization and performance management.

[0076] S50, generating an audit report including service quality analysis results based on the customer intention results, emotion recognition results and service response analysis results.

[0077] In this example, customer intent, emotion recognition, and service response analysis represent understanding of customer needs, perception of emotional states, and assessment of service execution behaviors, respectively. Together, these three components provide a structured understanding of the entire call service process. Before generating an audit report, these three types of information must be converted into metrics that facilitate quantitative evaluation.

[0078] Customer intent results typically include information such as intent type, intent fulfillment status, and the number of conversation turns. The intent fulfillment status can be quantified into an intent fulfillment indicator through methods such as semantic consistency determination and intent transfer path tracking, indicating whether the customer's needs were successfully received and responded to.

[0079] Emotion recognition results include different emotion category labels and corresponding emotion intensity values, particularly the frequency and duration of negative emotions, which are crucial for assessing service satisfaction and potential risks. Aggregating negative emotion intensities above a set threshold creates a negative emotion anomaly indicator.

[0080] The service response analysis results provide scores for multiple dimensions, including response delay, completeness of speech, and problem-solving efficiency. These three values need to be standardized and weighted together to generate a service performance indicator, which is used to comprehensively characterize the service's response quality and task completion capabilities.

[0081] After generating these three types of indicators, they can be integrated by constructing a service quality evaluation matrix. This matrix uses indicator type as a dimension and call unique identifiers as an index, forming a multidimensional analysis structure that allows for horizontal comparison and vertical trend tracking. Furthermore, the evaluation matrix can be mapped into a visual report format, supporting intuitive presentation through charts, scores, and rating labels, providing a visual basis for manual audits or system monitoring.

[0082] Methods for extracting intent fulfillment metrics include tracking whether the intent is confirmed and responded to by the agent during the conversation process based on intent classification results, or indirectly confirming it by combining keyword response matching rates. For the same customer call, multiple sub-intents can be extracted and their fulfillment determined separately, ultimately outputting a fulfillment rate indicator as a ratio.

[0083] The negative emotion anomaly indicator can be used to determine whether the cumulative weight of strong negative labels (such as "anger", "impatience" and "anxiety") in the emotion recognition results exceeds the warning threshold. The emotion volatility rate can also be introduced as an auxiliary criterion to identify calls with extreme emotional fluctuations.

[0084] Service performance indicators can be generated using a weighted scoring mechanism, with weights adjustable for different application scenarios. For example, in healthcare consultations, response time and problem-solving efficiency are prioritized, while in financial services, standardization of conversational language and the completeness of risk warnings are prioritized. The system can adaptively adjust the weighting coefficients of various indicators based on call tags or service type.

[0085] After the service quality assessment matrix is constructed, it can be graphically presented using various methods such as heat maps, score radar charts, trend line charts, etc., and an automatic labeling mechanism can be set to trigger abnormal scores, allowing auditors to quickly locate problematic call records.

[0086] Example: In healthcare scenarios, the evaluation matrix can help hospitals identify whether service staff may be slow to respond or the risk of patient complaints is increased during peak hours. It can also identify service interactions that may lead to medical disputes in advance through negative emotion abnormality indicators.

[0087] In financial services, the assessment matrix can reveal whether customers show difficulty understanding, questioning emotions, or disconnected service language during high-risk operational processes. By linking intention achievement indicators with service effectiveness indicators, it provides a quantitative reference for business compliance management and process optimization, while improving customer experience and risk control levels.

[0088] This embodiment jointly models customer intent, emotional state, and service response behavior and converts them into quantitative indicators, constructing a unified service quality assessment matrix. This allows for refined auditing and comprehensive scoring of the entire telephone service process. This approach not only improves the objectivity and consistency of audits but also provides data support for automated system alerts, service process optimization, and personnel performance evaluations. Visual presentation further enhances information readability and application efficiency, transforming service quality management from static sampling to dynamic, comprehensive analysis.

[0089] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. A call service quality audit method, device, equipment and medium are disclosed, including: collecting call audio data and transcribing it into call text data; using an intention recognition model to identify customer intentions in the call text data and generate customer intention results; using an emotion recognition model to identify customer emotions and generate emotion recognition results; using a service response analysis model to identify service behavior performance and generate service response analysis results; based on the customer intention results, emotion recognition results and service response analysis results, an audit report containing service quality analysis content is generated. The present invention uniformly processes customer intentions, emotional states and service response behaviors in call content, converts unstructured call information into structured evaluation data, realizes automated, multi-dimensional and in-depth audit analysis of service quality, improves audit efficiency and accuracy, and reduces manual burden and subjective bias risks.

[0090] In one embodiment, the above step S10 includes:

[0091] S101, collecting the original two-way audio stream through the telephone communication interface to obtain initial call audio data;

[0092] S102, performing background noise suppression and echo cancellation processing on the initial call audio data to obtain purified audio data;

[0093] S103, performing speaker separation processing on the purified audio data to obtain a service staff audio stream and a customer audio stream;

[0094] S104, inputting the service personnel audio stream into a first speech transcription engine to generate a service personnel text sequence;

[0095] S105, inputting the customer audio stream into a second speech transcription engine to generate a customer text sequence;

[0096] S106: Align the service personnel text sequence and the customer text sequence along a time axis to generate call text data including a role identifier and a timestamp.

[0097] In this embodiment, capturing the original two-way audio stream through the telephone communication interface is the foundational operation for digitizing the entire call process. This is typically accomplished by interfacing with a telephone exchange platform or softphone system, enabling real-time interception of the audio stream upon call establishment. The captured two-way audio stream is typically a dual-channel or mono mixed audio signal, containing the voice interaction between the service representative and the customer.

[0098] The process of generating the initial call audio data includes encapsulating and caching the audio signal. The specific format can use a standardized audio container such as WAV or FLAC. The sampling rate is recommended to be set to 16kHz or higher to ensure speech recognition accuracy.

[0099] Background noise suppression and echo cancellation are performed on the initial call audio data to improve voice quality and subsequent recognition accuracy. Background noise suppression can be implemented using spectral subtraction, Wiener filtering, or neural network denoising models to effectively filter out non-speech frequency components such as telephone line noise and current interference. Echo cancellation, on the other hand, targets the speaker echo signal picked up by the microphone during two-way communication. It often uses adaptive filters or spectral comparison methods to suppress it in real time, improving speech clarity.

[0100] Speaker separation separates the voice data from different speakers within the cleaned audio, creating a service agent audio stream and a customer audio stream. This process can be performed using a sound source separation model, speaker embedding alignment, or a deep learning model based on time-frequency masking to ensure accurate segmentation of voice segments corresponding to different speakers, even in the context of interleaved calls.

[0101] The service agent's audio stream is fed into the first speech transcription engine, while the customer's audio stream is fed into the second speech transcription engine. This allows for role-independent speech recognition processes. The two transcription engines can be copies of the same model or heterogeneous models. The former simplifies the system architecture, while the latter optimizes recognition for different speaking styles. Each engine processes the input audio signal into continuous text output for the corresponding role.

[0102] Timeline alignment of agent and customer text sequences is crucial for creating readable and structured call transcripts. Timeline alignment synchronizes audio sample timestamps or the start and end frame numbers of speech segments, creating a transcript of the interaction that includes conversation turns, role identifiers, and specific speaking times. This allows subsequent analysis models to recover semantic context by role.

[0103] Telephone communication interfaces can be implemented by connecting to an IP phone system via the SIP protocol or by connecting to a traditional PBX gateway to capture and digitize TDM signals. The audio signal sampling module can be configured with a dynamic bitrate adjustment strategy, automatically reducing sampling accuracy when bandwidth is limited.

[0104] The background noise suppression model utilizes a dual-channel noise reduction network trained on audio datasets from multiple industries, dynamically adapting to the current ambient noise spectrum during a call. The echo cancellation module incorporates a spectral hole detection mechanism to automatically identify echo playback intervals and reduce their amplitude.

[0105] Speaker separation processing can adopt a speaker separation network based on BLSTM or Conformer structure, and preset the service staff's timbre reference vector to enhance the separation effect when the customer and service staff's speech overlap.

[0106] The first and second speech transcription engines can model male and female vocal channels respectively, and can also introduce industry-specific adaptation dictionaries, such as adding a medical terminology recognition module to healthcare-related transcription engines, and adding language model extensions for securities and credit terms in financial scenarios.

[0107] In addition to using acoustic timestamps, the text timeline alignment module can also use the VAD (Voice Activity Detection) algorithm to identify natural sentence breaks, improve the semantic clarity of the transcription structure, and simultaneously annotate information such as emotional markers and speaking speed.

[0108] Example: At a healthcare service center, a patient dials a number to schedule a consultation. The system automatically initiates audio capture and character separation. Background noise originates from the ambient sound of the outpatient hall, and noise suppression preserves clear speech. As the patient expresses their needs, their speech accelerates and their interactions with the service staff frequently overlap. Using speaker separation technology, the system accurately reproduces the content of both conversations and transcribes them into a traceable text record for the doctor's subsequent analysis of the patient's concerns.

[0109] In financial services, when a customer calls the hotline regarding credit card installment payments, the background noise of televisions and echoes can interfere with the call. The purification module effectively removes non-human voice signals. The service representative uses professional terminology to explain the rate structure to the customer, and the speech transcription engine accurately restores terms such as "billing cycle" and "early repayment" and identifies the speaker, creating a clearly structured transcript that provides a basis for subsequent risk audits and compliance checks.

[0110] This embodiment achieves high-fidelity, highly time-accurate, and highly structured call text data generation through precise two-way audio acquisition, noise and echo suppression, speaker separation, and independent engine-based text transcription. This processing flow ensures the basic data quality for subsequent intent recognition, sentiment analysis, and service response evaluation, effectively avoiding issues such as role confusion, recognition errors, and context loss. This overall improvement in call data processing accuracy, completeness, and semantic restoration establishes a reliable data foundation for subsequent audits and service quality assessments.

[0111] In one embodiment, the above step S20 includes:

[0112] S201, extracting the text content of the customer role from the call text data to generate a customer conversation text;

[0113] S202, extracting semantic features from the customer conversation text to generate an intention feature vector;

[0114] S203, inputting the intention feature vector into an intention recognition model to generate an initial intention classification result;

[0115] S204, analyzing the initial intent classification result based on the conversation temporal context to generate an intent confidence score;

[0116] S205 , based on the intention confidence score, exclude classification items with a confidence score lower than a preset threshold from the initial intention classification result to generate a final customer intention result.

[0117] In this embodiment, extracting the text content of the customer role from the call text data to generate the customer conversation transcript relies on the role identification and text separation information completed in the previous call audio processing stage. This process typically performs a filtering operation based on the role label, combining all text segments marked as belonging to the customer in chronological order to form a continuous text sequence. The accuracy of role identification directly affects the integrity of the conversation transcript. Therefore, high-confidence role annotation information should be generated simultaneously during the transcription stage, for example, by combining the speaker embedding vector and predefined speech patterns to determine role attribution.

[0118] Extracting semantic features from customer conversation text and generating intent feature vectors is a deep semantic encoding process for the text. This step can generate context-sensitive embedding vectors based on pre-trained language models (such as BERT, ERNIE, RoBERTa, etc.), or it can be combined with domain semantic feature maps to enhance feature representation. For example, in the financial sector, business classification label embedding can be added, and in medical scenarios, consultation target type embedding can be introduced to form composite semantic features that can express intent directions such as requests, inquiries, and complaints.

[0119] Inputting the intent feature vector into the intent recognition model generates an initial intent classification result. This involves using a trained multi-class classifier to determine the customer's intent. The intent recognition model can be a shallow neural network, support vector machine, graph neural network, or a deep model based on a transformer architecture. Different network architectures can be used to balance accuracy and computational complexity based on the business density and semantic dimensionality of the training corpus. The initial intent classification result is typically a multi-class output, containing multiple candidate intents and their predicted probabilities.

[0120] Based on the conversational context, the initial intent classification results are analyzed to generate an intent confidence score. This approach embeds local semantic judgments within the interaction logic over a longer time span for dynamic correction. Methods for analyzing conversational context can include window-based sliding context modeling or using RNN, GRU, and other structures to model context sequences. This analysis not only relies on the current speech content but also considers semantic trends and historical speech behavior, thereby improving the accuracy of intent judgments for ambiguous expressions, semantic shifts, or multi-turn statements. The generated intent confidence score can be determined by integrating the model's predicted probability, context consistency score, and domain knowledge weighting.

[0121] Based on the intent confidence score, classification items below a preset confidence threshold are excluded from the initial intent classification results. This is a confidence filtering operation to ensure the quality of intent output and decision usability. The confidence threshold can be flexibly set based on the business scenario. For example, in high-risk business processing such as customer complaint identification, a higher threshold can be set to avoid false positives, while in low-risk business guidance, a more relaxed threshold can be set to improve recognition coverage. The final customer intent results retain intent labels and their semantic labels with sufficient confidence for subsequent use in modules such as service response analysis and audit modeling.

[0122] This embodiment constructs an intent recognition process based on customer role text, combining semantic feature extraction, context correction, and confidence filtering to effectively identify the customer's true intent during a call. Compared to traditional keyword-based intent matching methods, this structured recognition method can more accurately handle complex expressions such as ambiguous semantics, context transfer, and the coexistence of multiple intents. By performing structured semantic understanding on call text data through the intent recognition model, the accuracy of customer expression content and the reliability of business judgments are effectively improved, providing clear semantic support for subsequent service quality analysis and audit result generation.

[0123] In one embodiment, the above step S30 includes:

[0124] S301, separating the text content of the customer role from the call text data to generate a customer emotion text;

[0125] S302, performing an emotional feature extraction operation on the customer emotional text to generate a multi-dimensional emotional feature vector;

[0126] S303, inputting the multi-dimensional emotion feature vector into an emotion recognition model to generate an initial emotion classification result containing a negative emotion marker;

[0127] S304, performing emotion intensity quantification processing on the initial emotion classification result to generate an emotion intensity score;

[0128] S305 , associating the initial emotion classification result with the emotion intensity score to generate an emotion recognition result.

[0129] In this embodiment, separating the text content of customer roles from call text data to generate customer sentiment text is an operation that extracts customer speech based on the premise that the call content has been fully labeled with roles and aligned with the time sequence. This operation relies on the accurate assignment of service personnel and customer role labels. By associating the role-separated labels with the speech time sequence, the conversation content solely originating from the customer can be extracted from the overall text. Customer sentiment text is the sole source of corpus for emotion recognition analysis, so its accuracy directly determines the subsequent recognition quality. Particular attention should be paid to the complete preservation of boundary content such as short sentences, interjections, and emotional outbursts.

[0130] Performing sentiment feature extraction on customer sentiment texts to generate multi-dimensional sentiment feature vectors is a natural language processing method that converts the emotional signals in customer speech into vector representations that can be processed by recognition models. Feature extraction methods can include a fusion of word vector encoding (such as Word2Vec, FastText), context representation (such as embedding of pre-trained language models such as BERT), syntactic dependency tree structure, intonation and rhetoric pattern encoding, and other forms. In specific business scenarios, dictionaries can also be introduced to enhance sentiment polarity, such as introducing a negative financial vocabulary list, a medical anxiety dictionary, etc., to improve the feature coverage of the sentiment dimension. The multi-dimensional sentiment feature vector finally generated should include multiple dimensions such as sentiment tendency, sentiment intensity, expression form, semantic reversal, etc., to support the comprehensiveness of subsequent model judgments.

[0131] Inputting multidimensional emotion feature vectors into an emotion recognition model to generate an initial emotion classification result with a negative emotion label is the process of making a preliminary qualitative judgment on emotional tendencies. Emotion recognition models can be sentiment classification neural networks, hierarchical emotion label tree structures, or sequence models incorporating attention mechanisms. Training objectives typically include identifying basic emotion categories (such as anger, anxiety, dissatisfaction, disappointment, satisfaction, and calmness) and determining the emotion polarity (positive or negative). The output of the initial emotion classification result should clearly mark the negative emotion label and its type, providing a basic label reference for subsequent service quality analysis.

[0132] Quantifying the intensity of the initial emotion classification results to generate an emotion intensity score is the process of converting qualitative emotion labels into quantitative reference indicators. Emotion intensity can be calculated based on multiple dimensions, including the original model's predicted probability, the number of semantically prominent words, and the density of negative semantics within the text. Quantification methods can include weighted averaging, principal component projection, and fuzzy logic reasoning. This score should reflect the intensity, stability, and concentration of emotional expression, and be comparable to facilitate subsequent trend monitoring and anomaly identification.

[0133] Associating the initial emotion classification results with the emotion intensity score to generate emotion recognition results is the process of merging qualitative labels with quantitative indicators into a structured output. This output should include information such as the emotion label name, emotion polarity type, emotion intensity score, and a description of its source indicator. This facilitates subsequent use in scenarios such as service response quality assessment, service interruption warnings, and potential complaint risk alerts. This structured representation supports comparative analysis of emotion scores and labels under a unified evaluation logic, providing effective emotion data input for the audit system.

[0134] This embodiment establishes an emotion recognition process based on the text content of customer personas, combining multi-dimensional feature extraction, classification, and intensity quantification to more accurately identify the type and intensity of emotions expressed by customers during calls. This recognition mechanism deeply analyzes the emotional changes underlying semantic expressions and converts qualitative results into quantitative scores, providing a more comprehensive and reliable emotional data foundation for subsequent service quality assessments and risk warnings.

[0135] In one embodiment, the above step S40 includes:

[0136] S401, separating the text content of the service personnel role from the call text data to obtain the service personnel response text;

[0137] S402, extracting response time features from the service personnel response text to obtain response delay time data;

[0138] S403, detecting service speech integrity features in the service personnel's response text and generating a speech integrity score;

[0139] S404, identifying a problem-solving step sequence and its time stamp in the service personnel's response text to obtain the problem-solving step sequence and time stamp;

[0140] S405, determining the frequency of occurrence of logical connectives in the problem-solving step sequence;

[0141] S406, determining the total duration of the problem-solving path based on the time stamp;

[0142] S407, generating a problem-solving efficiency index based on the frequency of occurrence of the logical connectives and the total duration of the problem-solving path;

[0143] S408: Combine the response delay time data, the speech integrity score, and the problem-solving efficiency index to generate a multi-dimensional service response analysis result.

[0144] In this embodiment, the service agent role text is separated from the call text data to generate the service agent response text. This is done by identifying a set of sentences in the call log that contain the service agent role identifier and extracting all of the service agent's utterances to form a complete text sequence. This process relies on the role annotation information from the previous processing to accurately separate sentences that appear interspersed with the customer's speech and avoid interfering information. The service agent response text is the core data source for subsequent analysis and must maintain sentence order integrity and time tag relevance to ensure reproducibility in subsequent processing.

[0145] We extract response time features from agent response texts to generate response delay data. This data analyzes the time difference between the agent's initial response and the customer's previous statement to extract a measure of service response timeliness. Response delay can be accurately calculated using the timestamp field included in call text data. This can be expanded to include a collection of response delays for each conversation round, used to assess the agent's average response timeliness and fluctuation throughout the conversation, reflecting service efficiency and interaction quality.

[0146] Detecting the completeness of service language in service personnel responses and generating a completeness score for the language uses predefined service process specifications or script templates to analyze the response content for structural alignment, determining whether the service personnel fully cover the key points of the predefined script. Completeness features can include components such as the use of greetings, question confirmation, introductory statements, information disclosure, and closing remarks. The analysis incorporates techniques such as keyword matching, syntactic structure analysis, and semantic proximity recognition to compare the actual response text with a set of standard scripts. A comprehensive score is generated based on the hit ratio, resulting in a completeness score.

[0147] Identifying the sequence of problem-solving steps and their time stamps in service personnel response texts involves extracting the solution path reflected in the service process in stages and locating its temporal sequence. By building a service task knowledge base, we match and annotate the implicit action statements in the service personnel's texts, extracting semantic units with instructional properties such as "Please log in," "Submit an application," and "Verify your identity." We then generate corresponding time stamps based on their occurrence, forming a structure of step-time pairs. This sequence reflects whether the service process follows the correct path and whether it maintains continuity and integrity.

[0148] Determining the frequency of logical connectives within a problem-solving step sequence is a measure of semantic coherence within the process. Connectives such as "next," "then," "while," and "therefore" can demonstrate whether service personnel maintain a clear structure and task flow. Counting their frequency within a step sequence can serve as an indirect indicator of structural quality. A higher frequency generally indicates a more logical service explanation and lowers customer comprehension costs.

[0149] Determining the total duration of a problem-solving path based on time stamps evaluates the time taken by the entire problem-solving process by analyzing the time difference between the beginning and end of the problem-solving step sequence. This duration not only reflects the complexity of the process but also serves as a core parameter for efficiency evaluation, making it particularly suitable for evaluating responsiveness to complex problem solving.

[0150] The problem-solving efficiency index is generated based on the frequency of logical connectives and the total duration of the problem-solving path. This quantifies the level of service efficiency by establishing a multi-factor fusion model that comprehensively considers the clarity of expression structure and the time cost of operation. The final index can be generated using methods such as linear weighted scoring, standardized scoring, or a trained clustering model. This index should reflect whether the service completes the process in a short time under clear guidance.

[0151] Combining response delay data, script integrity scores, and problem-solving efficiency metrics generates a multi-dimensional service response analysis, providing a quantitative and structured representation of service personnel's overall performance. These three metrics cover the dimensions of response timeliness, script compliance, and problem-solving efficiency. The generated results provide concrete data support for service quality scoring, process audits, and personnel training, and feature structured, visual output for further analysis and comparison.

[0152] This example builds a role-separated response text analysis process to systematically extract three key metrics from service personnel during telephone interactions: response delay, speech completeness, and problem-solving efficiency. This enables a comprehensive and refined, data-driven assessment of service quality. This processing chain not only overcomes the difficulty of objective quantification in traditional manual monitoring but also improves the adaptability of service behavior to the audit system through structured indicator output, effectively supporting accurate decision-making for subsequent service optimization and risk monitoring operations.

[0153] In one embodiment, the above step S50 includes:

[0154] S501, aggregating intention achievement status data in the customer intention result to generate an intention achievement indicator;

[0155] S502, extracting negative emotion intensity data from the emotion recognition result to generate a negative emotion abnormality index;

[0156] S503, integrating the response delay time data, the speech completeness score, and the problem-solving efficiency index in the service response analysis result to generate a service effectiveness index;

[0157] S504, correlating the intention achievement index, negative emotion abnormality index, and service effectiveness index to generate a service quality evaluation matrix;

[0158] S505: Convert the service quality evaluation matrix into a visual report format to generate an audit report containing service quality analysis results.

[0159] In this embodiment, the intention fulfillment status data in the customer intention results are aggregated to generate an intention fulfillment index. This is done by summarizing the state sets formed by status fields identified as "completed," "uncompleted," or "interrupted" in multiple customer intention results, and performing frequency statistics or proportion calculations on these states to obtain an indicator system for evaluating the completeness of customer demand processing. The intention fulfillment status data typically comes from the structural fields of the final customer intention result in the preceding intention recognition processing, such as result tags such as whether the appointment was successful or whether the guidance was completed. The aggregation processing supports single call dimensions or batch session dimensions, and has the ability to analyze trends across data cycles.

[0160] Extracting negative emotion intensity data from emotion recognition results to generate a negative emotion anomaly index involves filtering and thresholding the emotion intensity values corresponding to labels such as "anger," "anxiety," and "dissatisfaction" in the emotion recognition results during a call. Negative emotion intensity data is typically presented as a standardized emotion intensity score. When generating the index, anomaly identification boundaries can be set. For example, if the score exceeds a certain threshold, it will be counted as an abnormal emotion event. Further statistics are collected on its frequency, intensity mean, or degree of variation to reflect the degree of customer emotional fluctuations or potential negative experiences during the service process.

[0161] Integrating response delay data, the speech integrity score, and the problem-solving efficiency metric from service response analysis results to generate a service effectiveness index involves normalizing or merging service performance data from multiple dimensions with uniform weights, creating a comprehensive scoring system that can be directly used for horizontal comparison or vertical trend analysis. Response delay data reflects service responsiveness, the speech integrity score measures service execution compliance, and the problem-solving efficiency metric reflects substantive processing efficiency. These three metrics can be weighted to create a customized comprehensive score tailored to specific business scenarios, ensuring the metric is sensitive and representative of the overall service process.

[0162] Correlating intent fulfillment indicators, negative emotion anomaly indicators, and service effectiveness indicators to generate a service quality assessment matrix maps the assessment results from three sources—customer intent processing results, emotion change trends, and service execution performance—into a unified multidimensional indicator space, expressing each dimension of service quality within a unified structure. This matrix is typically constructed in a two- or three-dimensional form, with each dimension corresponding to a specific indicator type. Each matrix cell represents the quality score result for a specific call sample or call period, thus supporting further statistical aggregation, quality trend analysis, service anomaly identification, and strategic intervention.

[0163] Converting a service quality assessment matrix into a visual report format and generating an audit report containing a service quality analysis involves converting and presenting structured assessment results in various formats, such as charts, tables, and text summaries, and outputting them as standardized data report files. Visual reports can include multi-dimensional radar charts, trend line graphs, heat matrix charts, or indicator rankings, and can be output in various formats, such as HTML, PDF, and Excel, for archiving and analysis. This audit report is not only used for service monitoring but also serves as a crucial basis for audit retention and compliance checks.

[0164] This embodiment aggregates and indexes the results of customer intent recognition, emotion recognition, and service response analysis at the structural level, constructs a complete service quality assessment system, and outputs audit reports in a visual form, which can significantly improve the automation level of service analysis and the comprehensiveness of assessment dimensions. This approach avoids the problems of one-sided manual analysis and scattered information in traditional methods, and realizes accurate audit capabilities based on structured indicators. At the same time, through the collaborative analysis of emotion intensity, service behavior efficiency, and customer intent matching, it can systematically identify potential service anomalies and customer dissatisfaction risks, and effectively support the intelligent management goals of financial technology and healthcare businesses in compliance auditing, customer experience improvement, and process optimization.

[0165] In one embodiment, after the above step S40, the method further includes:

[0166] S601, obtaining the customer intention result, emotion recognition result, and service response analysis result of the current call;

[0167] S602, combining the customer intention result, emotion recognition result, and service response analysis result into real-time monitoring indicator data;

[0168] S603, matching the real-time monitoring indicator data with a preset anomaly strategy library to generate an anomaly matching result;

[0169] S604: When the abnormal matching result meets the preset warning condition, a real-time warning message is generated.

[0170] In this embodiment, obtaining the customer intent results, emotion recognition results, and service response analysis results for the current call refers to retrieving the structured analysis results related to the current conversation from the completed processing flow and using them as the basic input for subsequent monitoring processing. Customer intent results typically include the behavioral goals, intent classification, and achievement status identified during the call; emotion recognition results include emotion category labels and emotion intensity scores; and service response analysis results include data fields such as response delay time, speech completeness, and problem-solving efficiency. This step typically implements the aggregation of multi-dimensional analysis results through a data bus or session state manager to ensure that analysis model results from different sources can be organized and managed on a unified timeline.

[0171] Combining customer intent results, emotion recognition results, and service response analysis results into real-time monitoring indicator data involves field-level mapping and structural fusion of the three dimensions to generate a set of indicators with monitoring properties. This indicator data structure supports standardized fields such as intent matching failure flags, negative emotion intensity values, response delay durations, and processing logic jump flags, and is categorized using a tag system or event dimension. This indicator structure is designed to balance real-time processing performance with contextual semantic expression capabilities, supporting the reflection of call session service quality status and abnormal characteristics within a unified indicator space.

[0172] Matching real-time monitoring indicator data with a pre-set anomaly policy library to generate anomaly matching results involves invoking the anomaly policy matching engine to perform rule comparison or model inference based on the structured monitoring indicator data generated by the current call, determining whether the current service process has triggered a certain type of service anomaly. The anomaly policy library can be composed of predefined rule expressions, contextual condition models, or historically trained anomaly cases. The matching process supports both static rule matching and dynamic risk model calculation. Matching results are typically output as anomaly trigger type, matching rule number, and confidence level, generating event information that can be used for real-time early warning responses.

[0173] When anomaly matching results meet preset warning conditions, real-time warning information is generated. This means that after the exception policy is matched, the pre-set warning trigger mechanism encapsulates the anomaly matching results that meet the trigger conditions into structured warning instructions or prompt signals for downstream systems to respond to. Warning conditions can include rule-level trigger thresholds, cumulative anomaly frequency, and combined triggering of similar indicators. Real-time warning information includes the anomaly type, trigger cause, severity level, and response recommendations. It can be used by the real-time monitoring console, agent assistance system, or business risk control module to implement service interruption control, customer sentiment intervention, or weighted quality inspection processing.

[0174] Example: In a remote banking customer service system in the financial sector, customers call the hotline to inquire about loan applications, credit card limit adjustments, or disputed bills. The system first collects the original two-way audio stream between the customer and the service representative through the telephone communication interface and performs background noise suppression and echo cancellation on the audio data to produce purified audio data. Speaker separation is then used to separate the customer and service representative audio streams, which are then fed into two separate speech transcription engines to generate a customer text sequence and a service representative text sequence. The system then aligns the two along the timeline to generate call text data containing role identifiers and timestamps.

[0175] Call text data is further used as input to extract the text content of the customer persona. After semantic feature extraction and intent recognition model inference, the customer intent results are generated. The intent recognition process combines temporal context judgment to filter out low-confidence classification items and retain only valid intents, reflecting the customer's actual business needs or operational goals during the call.

[0176] Based on the same call text data, the system extracts the customer's emotional text and generates a multi-dimensional emotion feature vector through emotional feature extraction. This vector is then input into the emotion recognition model to generate an initial emotion classification result. The system then quantifies the emotion intensity of this result, generating an emotion intensity score and generating an emotion recognition result that matches the category, identifying the customer's emotional tendencies and the intensity of their emotional fluctuations during the call.

[0177] The service personnel's transcripts were used for service response analysis, including extracting response delays, testing service script completeness, identifying problem-solving steps and their logical structure, and calculating the total duration of the resolution path and the frequency of connectives. This ultimately resulted in a multi-dimensional service response analysis that included response delay data, script completeness scores, and problem-solving efficiency indicators.

[0178] After the call, the system aggregates customer intent fulfillment status, customer sentiment, and service response efficiency to form a service quality assessment matrix. It also generates a visual audit report, which serves as a crucial basis for customer service quality inspection, regulatory audits, and compliance analysis. The system also supports a real-time early warning mechanism. When anomalies such as abnormal customer sentiment, delayed service responses, or unsatisfied intent are triggered, it generates real-time warning information for business personnel to reference, enabling timely intervention and handling of sensitive calls.

[0179] In remote healthcare appointment services, patients often inquire by phone about appointment progress, doctor appointments, medical insurance reimbursement, or test results. The service system accesses the phone interface to collect raw call audio, purifies it, and uses speaker separation technology to extract the patient and customer service audio. A dual-speech transcription engine then generates time-aligned call text data, ensuring clear character roles and complete context.

[0180] The system uses call text data to identify the intent of patient text messages, determining whether they are requests for appointment changes, examination appointments, fee inquiries, or cancellations. It then filters out invalid intents based on recognition confidence. The emotion recognition model also analyzes patient expressions for anxiety, anger, or restlessness, helping to identify possible sources of dissatisfaction or urgent requests.

[0181] The service staff's responses are used for service response analysis, including the response delay time, whether key scripts are covered (such as whether the medical treatment process is informed, whether the patient's identity is confirmed), the logical rationality and efficiency of the problem-solving steps, etc., so as to quantify the service staff's response quality and service closed-loop capabilities.

[0182] The system aggregates identified patient intent fulfillment, emotional volatility, and service response efficiency to form a structured service quality assessment matrix and generates audit reports in the form of charts or tables. This information can be used for internal service supervision and can also be fed back to the doctor scheduling system, intelligent guidance system, or patient safety management module.

[0183] If the patient fails to achieve their intended purpose, experiences strong negative emotions during the call, or the service staff responds slowly or misses steps, the system will trigger an early warning mechanism in real time, generating an early warning message to remind the agent supervisor or medical service coordinator to manually access or reconnect, ensuring the patient's service experience and avoiding interruptions or misleading of the medical process.

[0184] This embodiment builds a service quality monitoring mechanism with dynamic judgment capabilities by instantly aggregating customer intent, emotional state, and service response behavior analysis results during a call and matching them with a structured anomaly policy library in real time. This processing method breaks through the limitations of traditional post-audit and static rule-based judgment, and realizes the linkage between state perception and anomaly identification during the call. When the intention is not achieved during the call, the customer's emotions fluctuate strongly, or the service response efficiency decreases significantly, the system can immediately issue an early warning, significantly improving service response sensitivity, risk prevention and control initiative, and customer satisfaction adjustment capabilities.

[0185] In one embodiment, a call service quality audit device is provided, which corresponds to the call service quality audit method in the above embodiment. Figure 3 , Figure 3 This is a functional module diagram of a preferred embodiment of the call service quality audit device of the present invention. It includes a speech transcription module 10, an intent analysis module 20, an emotion recognition module 30, a response analysis module 40, and an audit generation module 50. Each functional module is described in detail below:

[0186] The speech transcription module 10 is used to collect call audio data and perform text transcription on the call audio data through the speech transcription module to generate call text data;

[0187] The intention analysis module 20 is used to perform intention recognition processing on the call text data based on the intention recognition model to generate a customer intention result;

[0188] The emotion recognition module 30 is used to perform emotion recognition processing on the call text data based on the emotion recognition model to generate an emotion recognition result;

[0189] a response analysis module 40 for performing a service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0190] The audit generation module 50 is used to generate an audit report including service quality analysis results based on the customer intention results, emotion recognition results and service response analysis results.

[0191] In one embodiment, the speech transcription module 10 is specifically configured to:

[0192] Collect the original two-way audio stream through the telephone communication interface to obtain the initial call audio data;

[0193] Performing background noise suppression and echo cancellation processing on the initial call audio data to obtain purified audio data;

[0194] Performing speaker separation processing on the purified audio data to obtain a service staff audio stream and a customer audio stream;

[0195] Inputting the service personnel audio stream into a first speech transcription engine to generate a service personnel text sequence;

[0196] Inputting the customer audio stream into a second speech transcription engine to generate a customer text sequence;

[0197] The service personnel text sequence and the customer text sequence are aligned along a time axis to generate call text data including a role identifier and a timestamp.

[0198] In one embodiment, the intention analysis module 20 is specifically configured to:

[0199] Extracting the text content of the customer role from the call text data to generate the customer conversation text;

[0200] Extracting semantic features from the customer conversation text to generate an intent feature vector;

[0201] Inputting the intent feature vector into an intent recognition model to generate an initial intent classification result;

[0202] Analyzing the initial intent classification result based on the conversation temporal context to generate an intent confidence score;

[0203] Based on the intent confidence score, classification items with a confidence score below a preset threshold are excluded from the initial intent classification results to generate a final customer intent result.

[0204] In one embodiment, the emotion recognition module 30 is specifically configured to:

[0205] Separating the text content of the customer role from the call text data to generate a customer emotion text;

[0206] Performing an emotional feature extraction operation on the customer's emotional text to generate a multi-dimensional emotional feature vector;

[0207] Inputting the multi-dimensional emotion feature vector into an emotion recognition model to generate an initial emotion classification result containing a negative emotion marker;

[0208] Performing emotion intensity quantification processing on the initial emotion classification result to generate an emotion intensity score;

[0209] The initial emotion classification result is associated with the emotion intensity score to generate an emotion recognition result.

[0210] In one embodiment, the response analysis module 40 is specifically configured to:

[0211] Separating the text content of the service personnel role from the call text data to obtain the service personnel response text;

[0212] Extracting response time features from the service personnel's response text to obtain response delay time data;

[0213] Detecting service speech integrity features in the service personnel's response text and generating a speech integrity score;

[0214] Identifying a problem-solving step sequence and its time stamp in the service personnel's response text to obtain the problem-solving step sequence and the time stamp;

[0215] determining the frequency of occurrence of logical connectives in the sequence of problem-solving steps;

[0216] determining a total duration of the problem-solving path based on the time stamp;

[0217] generating a problem-solving efficiency index based on the frequency of occurrence of the logical connectives and the total duration of the problem-solving path;

[0218] The response delay time data, the speech completeness score and the problem-solving efficiency index are combined to generate a multi-dimensional service response analysis result.

[0219] In one embodiment, the audit generation module 50 is specifically configured to:

[0220] Aggregating intention achievement status data in the customer intention results to generate an intention achievement indicator;

[0221] Extracting negative emotion intensity data from the emotion recognition results to generate a negative emotion abnormality indicator;

[0222] Integrate the response delay time data, the speech completeness score and the problem-solving efficiency index in the service response analysis results to generate a service effectiveness index;

[0223] Correlating the intention achievement index, negative emotion abnormality index, and service effectiveness index to generate a service quality evaluation matrix;

[0224] The service quality assessment matrix is converted into a visual report format to generate an audit report containing service quality analysis results.

[0225] In one embodiment, the response analysis module 40 is specifically configured to:

[0226] Obtain customer intent results, emotion recognition results, and service response analysis results for the current call;

[0227] Combining the customer intention results, emotion recognition results, and service response analysis results into real-time monitoring indicator data;

[0228] Matching the real-time monitoring indicator data with a preset abnormal strategy library to generate an abnormal matching result;

[0229] When the abnormal matching result meets the preset warning condition, real-time warning information is generated.

[0230] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the service side of a call service quality audit method.

[0231] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide determination and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a user-side method for auditing call service quality.

[0232] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0233] Collecting call audio data and performing text transcription on the call audio data through a speech transcription module to generate call text data;

[0234] Performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result;

[0235] Performing emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result;

[0236] Performing service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0237] Based on the customer intention results, emotion recognition results and service response analysis results, an audit report including service quality analysis results is generated.

[0238] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0239] Collecting call audio data and performing text transcription on the call audio data through a speech transcription module to generate call text data;

[0240] Performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result;

[0241] Performing emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result;

[0242] Performing service response analysis on the call text data based on a service response analysis model to generate a service response analysis result;

[0243] Based on the customer intention results, emotion recognition results and service response analysis results, an audit report including service quality analysis results is generated.

[0244] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0245] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0246] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0247] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A method for auditing call service quality, characterized in that: The following steps are involved: Collecting call audio data and performing text transcription on the call audio data through a speech transcription module to generate call text data; Performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result; Performing emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result; Performing service response analysis on the call text data based on a service response analysis model to generate a service response analysis result; Based on the customer intention results, emotion recognition results and service response analysis results, an audit report including service quality analysis results is generated.

2. The call service quality audit method according to claim 1, wherein: Collecting call audio data and performing text transcription on the call audio data through a voice transcription module to generate call text data, including: Collect the original two-way audio stream through the telephone communication interface to obtain the initial call audio data; Performing background noise suppression and echo cancellation processing on the initial call audio data to obtain purified audio data; Performing speaker separation processing on the purified audio data to obtain a service staff audio stream and a customer audio stream; Inputting the service personnel audio stream into a first speech transcription engine to generate a service personnel text sequence; Inputting the customer audio stream into a second speech transcription engine to generate a customer text sequence; The service personnel text sequence and the customer text sequence are aligned along a time axis to generate call text data including a role identifier and a timestamp.

3. The call service quality audit method according to claim 1, wherein: Performing intent recognition processing on the call text data based on the intent recognition model to generate a customer intent result, including: Extracting the text content of the customer role from the call text data to generate the customer conversation text; Extracting semantic features from the customer conversation text to generate an intent feature vector; Inputting the intent feature vector into an intent recognition model to generate an initial intent classification result; Analyzing the initial intent classification result based on the conversation temporal context to generate an intent confidence score; Based on the intent confidence score, classification items with a confidence score below a preset threshold are excluded from the initial intent classification results to generate a final customer intent result.

4. The call service quality audit method according to claim 1, wherein: Performing emotion recognition processing on the call text data based on the emotion recognition model to generate an emotion recognition result includes: Separating the text content of the customer role from the call text data to generate a customer emotion text; Performing an emotional feature extraction operation on the customer's emotional text to generate a multi-dimensional emotional feature vector; Inputting the multi-dimensional emotion feature vector into an emotion recognition model to generate an initial emotion classification result containing a negative emotion marker; Performing emotion intensity quantification processing on the initial emotion classification result to generate an emotion intensity score; The initial emotion classification result is associated with the emotion intensity score to generate an emotion recognition result.

5. The call service quality audit method according to claim 1, wherein: Performing service response analysis on the call text data based on the service response analysis model to generate a service response analysis result, including: Separating the text content of the service personnel role from the call text data to obtain the service personnel response text; Extracting response time features from the service personnel's response text to obtain response delay time data; Detecting service speech integrity features in the service personnel's response text and generating a speech integrity score; Identifying a problem-solving step sequence and its time stamp in the service personnel's response text to obtain the problem-solving step sequence and the time stamp; determining the frequency of occurrence of logical connectives in the sequence of problem-solving steps; determining a total duration of the problem-solving path based on the time stamp; generating a problem-solving efficiency index based on the frequency of occurrence of the logical connectives and the total duration of the problem-solving path; The response delay time data, the speech completeness score and the problem-solving efficiency index are combined to generate a multi-dimensional service response analysis result.

6. The call service quality audit method according to claim 1, wherein: Based on the customer intent results, emotion recognition results, and service response analysis results, an audit report containing service quality analysis results is generated, including: Aggregating intention achievement status data in the customer intention results to generate an intention achievement indicator; Extracting negative emotion intensity data from the emotion recognition results to generate a negative emotion abnormality indicator; Integrate the response delay time data, the speech completeness score and the problem-solving efficiency index in the service response analysis results to generate a service effectiveness index; Correlating the intention achievement index, negative emotion abnormality index, and service effectiveness index to generate a service quality evaluation matrix; The service quality assessment matrix is converted into a visual report format to generate an audit report containing service quality analysis results.

7. The call service quality audit method according to claim 1, wherein: After performing service response analysis on the call text data based on the service response analysis model and generating a service response analysis result, the method further includes: Obtain customer intent results, emotion recognition results, and service response analysis results for the current call; Combining the customer intention results, emotion recognition results, and service response analysis results into real-time monitoring indicator data; Matching the real-time monitoring indicator data with a preset abnormal strategy library to generate an abnormal matching result; When the abnormal matching result meets the preset warning condition, real-time warning information is generated.

8. A call service quality audit device, characterized in that: The call service quality audit device includes: A speech transcription module is used to collect call audio data and perform text transcription on the call audio data through the speech transcription module to generate call text data; An intent analysis module, configured to perform intent recognition processing on the call text data based on an intent recognition model and generate a customer intent result; An emotion recognition module, configured to perform emotion recognition processing on the call text data based on an emotion recognition model to generate an emotion recognition result; a response analysis module, configured to perform service response analysis on the call text data based on a service response analysis model to generate a service response analysis result; The audit generation module is used to generate an audit report including service quality analysis results based on the customer intention results, emotion recognition results and service response analysis results.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a call service quality audit program stored in the memory and capable of running on the processor. When the call service quality audit program is executed by the processor, the steps of the call service quality audit method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a call service quality audit program, which, when executed by the processor, implements the steps of the call service quality audit method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Voice call data analysis system based on cloud computing

    CN120751428A

  • Multi-terminal customer service session quality automatic auditing method

    CN121169199A

  • Sales customer service data classification and arrangement method and system based on AI analysis

    CN121327605A

  • Short text dialogue-oriented multilevel intention recognition agent processing system, method and device, processor and storage medium thereof

    CN121579783A

  • Insurance policy customer allocation method, system and equipment based on relation graph, and medium

    CN122089491A