AI-based Deep Customer Service Dialogue Method and System Based on Generative Speech Dialogue Content Model

CN122575353APending Publication Date: 2026-08-14GUANGZHOU EAPHONETECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

然而,这种架构存在根本性缺陷:ASR模块的语音识别误差会直接传递至后续环节,导致语义理解偏差甚至响应错误;TTS模块独立于语义生成过程,常出现语音韵律与语义情感不匹配的情况,例如用欢快语调播报负面信息,严重影响用户体验

Benefits of technology

[0006]本发明有益效果:通过端到端声义一体建模与反事实轨迹生成技术的融合,本方法在用户语音输入存在口音干扰或环境噪声时,仍能精准捕捉语义核心并生成合规响应,将对话系统的合规响应率提升至99%以上,同时自然度评分达4.6/5.0,显著改善传统级联架构因误差累积导致的对话失效问题。多粒度一致性约束机制通过动态追踪用户意图与情感脉络,确保多轮对话中安抚、解释等策略的连贯性,避免因响应割裂引发的用户信任下降。可信度自校验模块嵌入政策知识图谱后,可实时修正虚构的服务条款或退款政策,从源头消除法律风险,减少企业因信息误导产生的纠纷成本。声学参数与语义情感的协同优化技术,使语音输出的语调、停顿与内容情感精准匹配,彻底消除用机械音播报关怀信息的割裂感,增强用户情感共鸣。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575353A_ABST
    Figure CN122575353A_ABST
Patent Text Reader

Abstract

This invention proposes an AI customer service deep dialogue method and system based on a generative speech dialogue content model. It belongs to the interdisciplinary technical field of generative artificial intelligence, speech synthesis, and cognitive dialogue systems. The method includes: performing multimodal joint acoustic-semantic encoding on user speech input to generate a joint embedding vector that integrates acoustic features and semantic representations; constructing an end-to-end integrated acoustic-semantic generation model based on the joint embedding vector to achieve a direct mapping from speech signals to semantic responses; and, through the fusion of end-to-end integrated acoustic-semantic modeling and counterfactual trajectory generation technology, this method can still accurately capture the semantic core and generate compliant responses even when user speech input contains accent interference or environmental noise, increasing the compliance response rate of the dialogue system to over 99%, while achieving a naturalness score of 4.6 / 5.0, significantly improving the dialogue failure problem caused by error accumulation in traditional cascaded architectures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention proposes an AI customer service deep dialogue method and system based on a generative speech dialogue content model, which belongs to the interdisciplinary field of generative artificial intelligence, speech synthesis and cognitive dialogue systems. Background Technology

[0002] In the field of traditional AI customer service technology, generative dialogue systems often adopt a cascaded architecture, which completes the dialogue process sequentially through automatic speech recognition (ASR), natural language understanding (NLU), natural language generation (NLG), and text-to-speech (TTS) modules. However, this architecture has fundamental flaws: speech recognition errors in the ASR module are directly transmitted to subsequent stages, leading to semantic comprehension deviations or even incorrect responses; the TTS module is independent of the semantic generation process, often resulting in a mismatch between speech rhythm and semantic emotion, such as delivering negative information in a cheerful tone, which seriously affects the user experience.

[0003] More importantly, existing systems lack the ability to model the evolutionary patterns of multi-turn dialogues. Each response is generated independently, making it difficult to maintain consistency between intent and emotion, and easily leading to contradictory behaviors such as reassuring users in the first round and shirking responsibility in the second. Furthermore, large language models are prone to fabricating false information when generating responses, such as fictitious service terms or policies, triggering legal risks and trust crises. Although some research has attempted to improve the system through end-to-end voice dialogue and emotional TTS technologies, core issues such as deep fusion of acoustic features and semantics, consistency constraints in multi-turn dialogues, and reliable verification of factual information remain unresolved. Therefore, there is an urgent need for a new AI customer service deep dialogue method that can break through the limitations of traditional cascaded architectures, achieve integrated sound and semantic modeling, controllable dialogue evolution, and reliable factual information, in order to drive the industry towards a higher-level intelligent paradigm. Summary of the Invention

[0004] This invention provides an AI-powered deep dialogue method and system for customer service based on a generative speech dialogue content model, in order to solve the problems mentioned in the background section above: The present invention proposes an AI customer service deep dialogue method based on a generative speech dialogue content model, the method comprising: S1. Perform multimodal joint acoustic-semantic encoding on the user's voice input to generate a joint embedding vector that integrates acoustic features and semantic representation; construct an end-to-end integrated acoustic-semantic generation model based on the joint embedding vector to complete the direct mapping from speech signal to semantic response; S2. Generate counterfactual trajectories based on joint embedding vectors, simulate potential user inquiry paths by introducing adversarial examples, and generate a multi-dimensional response candidate set covering positive facts and reverse hypotheses; perform semantic consistency verification on the candidate set. S3. Perform factual credibility self-verification on the response candidate set after semantic consistency verification. By embedding policy knowledge graph and business rule engine, cross-verify the authenticity of key information involved in the response; mark and correct fictitious content, and generate a credible response set. S4. Based on the trusted response set, prosodic-semantic co-optimization is performed. By dynamically adjusting the acoustic parameters, the speech output is strongly correlated with semantic emotion, eliminating the phenomenon of emotion mismatch and generating a natural speech response that integrates sound and meaning. S5. Perform multi-granularity quality assessment based on the integrated sound and semantic natural speech response, and generate a response quality score by combining user feedback data and system preset indicators. When the score is lower than the threshold, trigger the model iteration optimization process to update the joint sound and semantic coding parameters and the counterfactual trajectory generation strategy.

[0005] The AI ​​customer service deep dialogue system based on a generative speech dialogue content model proposed in this invention includes: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are made to implement the method described in any one of the above.

[0006] The beneficial effects of this invention are as follows: By integrating end-to-end semantic-verbal modeling with counterfactual trajectory generation technology, this method can accurately capture the semantic core and generate compliant responses even when user voice input is subject to accent interference or environmental noise. This increases the compliance response rate of the dialogue system to over 99%, while achieving a naturalness score of 4.6 / 5.0, significantly improving the dialogue failure problem caused by error accumulation in traditional cascaded architectures. The multi-granularity consistency constraint mechanism dynamically tracks user intent and emotional context, ensuring the coherence of reassurance and explanation strategies in multi-turn dialogues and avoiding a decline in user trust due to fragmented responses. The credibility self-verification module, embedded with a policy knowledge graph, can correct fictitious service terms or refund policies in real time, eliminating legal risks at the source and reducing the cost of disputes caused by misleading information for enterprises. The collaborative optimization technology of acoustic parameters and semantic emotion ensures precise matching of the tone, pauses, and emotional content of the voice output, completely eliminating the disjointed feeling of broadcasting caring information in a mechanical voice and enhancing user emotional resonance. Attached Figure Description

[0007] Figure 1 This is a diagram illustrating the steps of the method described in this invention. Detailed Implementation

[0008] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.

[0009] One embodiment of the present invention, such as Figure 1 As shown, an AI customer service deep dialogue method based on a generative speech dialogue content model includes: S1. Perform multimodal joint acoustic and semantic encoding on the user's voice input to generate a joint embedding vector that integrates acoustic features and semantic representation; construct an end-to-end integrated acoustic and semantic generation model based on the joint embedding vector to complete the direct mapping from speech signal to semantic response and avoid the risk of error accumulation in each stage of the cascaded architecture (ASR→NLU→NLG→TTS). S2. Generate counterfactual trajectories based on joint embedding vectors, simulate potential user inquiry paths by introducing adversarial examples, and generate a multi-dimensional response candidate set covering positive facts and reverse hypotheses; perform semantic consistency verification on the candidate set, eliminate responses that contradict the intention of the historical dialogue, and ensure the consistency of intent and emotion in multi-turn dialogues. S3. Perform factual credibility self-verification on the response candidate set after semantic consistency verification. By embedding a policy knowledge graph and a business rule engine, cross-verify the authenticity of key information involved in the response, such as the policy terms and service terms of the key information; mark and correct fictitious content, and generate a credible response set that meets business specifications and legal requirements. S4. Based on the trusted response set, perform prosodic-semantic co-optimization. By dynamically adjusting acoustic parameters, including speech rate, pitch, and pauses, the speech output is strongly correlated with semantic emotions (such as reassurance, explanation, and confirmation), eliminating emotional mismatch phenomena, such as delivering bad news in a cheerful tone; and generating a natural speech response that integrates sound and meaning. S5. Perform multi-granularity quality assessment based on the integrated sound and semantic natural speech response, and generate a response quality score by combining user feedback data and system preset indicators (such as compliance rate and naturalness score). When the score is lower than the threshold, trigger the model iteration optimization process, update the joint sound and semantic encoding parameters and counterfactual trajectory generation strategy, and continuously improve the credibility and naturalness of the dialogue system.

[0010] The working principle and effects of the above technical solution are as follows: Through multimodal semantic-verbal joint encoding and end-to-end model construction, the accumulation of errors in each stage of the cascaded architecture is reduced, avoiding response deviations caused by error superposition. Optimizing counterfactual trajectory generation and semantic consistency verification enhances the intent and emotional coherence of multi-turn dialogues, reducing responses that contradict historical dialogues. Embedding policy knowledge graphs and business rule engines improves the authenticity and compliance of response information, avoiding business risks caused by fabricated content. Prosodic-semantic collaborative optimization eliminates emotional mismatches, improves the naturalness of voice responses, and enhances the user's dialogue experience. Multi-granularity quality assessment and model iteration continuously improve system credibility, ensuring service compliance while improving response accuracy, reducing the probability of users repeatedly asking questions and dialogue interruptions, and further reducing the cost of human customer service intervention.

[0011] In one embodiment of the present invention, S1 includes: S11. Collect the user's real-time voice input signal, filter out environmental noise and interference ripples in the signal, retain clear and effective voice segments, and ensure the integrity and purity of the input signal. S12. Extract the acoustic low-level features from the effective speech segments, including core features such as Mel frequency cepstral coefficients, fundamental frequency, and speech duration, to form a standardized acoustic feature set; S13. Perform multimodal joint acoustic-semantic encoding on the acoustic feature set and the text semantic information corresponding to the speech, and achieve deep fusion of acoustic features and semantic information through attention mechanism to generate a joint embedding vector that integrates acoustic features and semantic representation. S14. Input the joint embedded vector into the preset model architecture, complete the feature dimension calibration and feature weight allocation, and build an end-to-end integrated semantic generation model, breaking the module segmentation barrier of the traditional cascaded architecture. S15. By using an end-to-end integrated speech-semantic generation model, the mapping and conversion from speech signals to semantic responses can be directly realized, eliminating intermediate translation steps and avoiding the risk of error accumulation in each step of the cascaded architecture.

[0012] The working principle and effects of the above technical solution are as follows: It filters environmental noise and interference in speech input, improving the purity and integrity of the speech signal and avoiding feature extraction deviations caused by noise interference. It extracts standardized acoustic low-level features, enhancing the standardization of the feature set and reducing data chaos in subsequent encoding processes. Multimodal joint acoustic-semantic coding achieves deep fusion of acoustic and semantic features, improving the representation accuracy of joint embedding vectors and making feature transmission more precise. It constructs an end-to-end integrated acoustic-semantic generation model, breaking down the barriers of traditional cascaded architectures, eliminating intermediate translation stages, reducing error accumulation in each stage, and avoiding response distortion caused by error superposition. This not only improves the mapping efficiency from speech to semantic response but also ensures the accuracy of the mapping, further enhancing the basic stability of subsequent dialogue responses and reducing the probability of dialogue anomalies caused by signal or feature problems.

[0013] In one embodiment of the present invention, S13 includes: The generated standardized acoustic feature set is retrieved, and the corresponding text semantic information of the speech input is obtained synchronously to complete the synchronous alignment of the two types of data. Initiate a multimodal acoustic-semantic joint coding process, and perform feature normalization processing on the acoustic feature set and text semantic information respectively to unify the feature dimensions of the two types of data; An attention mechanism is introduced to assign weights to the normalized acoustic features and semantic information, thereby enhancing the representational ability of key features and reducing the interference of redundant features. The weighted acoustic features and semantic information are deeply fused to complete feature interaction and reconstruction, generating a joint embedding vector that integrates acoustic features and semantic representation.

[0014] The working principle and effects of the above technical solution are as follows: By synchronously aligning the acoustic feature set with the text semantic information, the fusion deviation caused by the misalignment of the two types of data is reduced, and encoding errors caused by feature mismatch are avoided. Feature normalization processing is performed on the two types of data to unify the feature dimensions, reduce encoding interference caused by dimensional differences, and make subsequent fusion smoother. An attention mechanism is introduced to allocate feature weights, strengthen the representational ability of key features, weaken the influence of redundant features, improve feature utilization efficiency, and reduce the processing resources occupied by invalid features. Deep fusion and reconstruction are performed on the weighted features to generate accurate joint embedding vectors, improving the comprehensiveness and accuracy of feature representation. This ensures efficient fusion of acoustic and semantic features and improves the reliability of the embedding vectors, laying a solid foundation for subsequent model construction and response generation, and avoiding subsequent dialogue response deviations caused by inadequate feature fusion.

[0015] In one embodiment of the present invention, S2 includes: S21. Extract key information such as user dialogue intent and historical dialogue context from the joint embedding vector, sort out the dialogue logic, and clarify the core needs of the current dialogue. S22. Generate counterfactual trajectories based on the core needs of the dialogue, set counterfactual assumptions in different scenarios, and simulate potential dialogue paths such as possible follow-up questions, challenges, and supplements from users. S23. Introduce an adversarial sample generation mechanism to generate diverse adversarial samples that conform to the user's potential dialogue path, enrich the coverage of dialogue scenarios, and improve the counterfactual trajectory system. S24. Combining positive fact responses and negative hypothesis responses, integrate and generate a multi-dimensional response candidate set covering multiple scenarios and intentions to ensure the comprehensiveness and diversity of the candidate set; S25. Perform semantic consistency verification on the multi-dimensional response candidate set, compare the fit of each candidate response with the historical dialogue intent and contextual logic, eliminate contradictory responses, retain candidate content that conforms to the dialogue logic, and ensure the consistency of intent and emotion in multi-turn dialogues.

[0016] The working principle and effects of the above technical solution are as follows: It extracts key information and core needs from the dialogue, clarifies the logical flow of the dialogue, reduces response deviations caused by biased demand judgments, and avoids irrelevant answers. Based on core needs, it generates counterfactual trajectories, simulates potential user dialogue paths, and combines adversarial examples to enrich scenario coverage, enhancing the comprehensiveness and diversity of the response candidate set and reducing dialogue interruptions caused by scenario omissions. It performs semantic consistency verification on the response candidate set, eliminating content that contradicts historical dialogues, ensuring the consistency of intent and emotion in multi-turn dialogues, and reducing the probability of logical confusion in the dialogue. It can accurately capture users' core needs and anticipate potential follow-up questions, making responses more targeted and forward-looking, further improving dialogue fluency, avoiding user dissatisfaction caused by one-sided or illogical responses, and providing reliable support for the generation of subsequent credible responses.

[0017] In one embodiment of the present invention, S3 includes: S31. Extract key information from the response candidate set after semantic consistency verification, including core content such as policy terms, service terms, business parameters, and processing procedures, to form a key information list; S32. Embed the key information list into the preset policy knowledge graph and business rule engine to establish a mapping relationship between key information and knowledge graph and rule engine, so as to realize rapid matching and retrieval of information; S33. Verify the authenticity and accuracy of key information through two-way cross-validation of policy knowledge graph and business rule engine, and investigate issues such as information deviation and fabrication; S34. Mark any fictitious content or information discrepancies discovered during the verification process, and correct the content in accordance with policy requirements and business specifications to ensure the compliance of the response information; S35. Integrate all the revised candidate responses to generate a set of credible responses that comply with business specifications and legal requirements, ensuring the credibility and compliance of the response content.

[0018] The working principle and effects of the above technical solution are as follows: Key information is extracted from the candidate response set and compiled into a list, reducing omissions and preventing inaccurate responses due to missing information. Key information is linked to policy knowledge graphs and business rule engines to achieve rapid matching and retrieval, improving information verification efficiency and reducing the time cost of manual verification. Through two-way cross-validation, information deviations and fabricated issues are identified, improving the authenticity and accuracy of key information and preventing the transmission of incorrect information to users. Problematic content is marked and corrected to ensure responses comply with policies and business specifications, reducing compliance risks and avoiding disputes arising from non-compliant responses. The corrected candidate responses are integrated to generate a credible response set, enhancing the credibility of the response content. This ensures service compliance, increases user trust in customer service responses, and reduces repeated inquiries and complaints due to unreliable information.

[0019] In one embodiment of the present invention, step S4 includes: S41. Analyze the semantic sentiment of each response in the trusted response set, distinguish different emotional scenarios such as soothing, explaining, confirming, and reminding, and clarify the emotional needs corresponding to each response; S42. Determine the acoustic parameter standards corresponding to different emotional scenarios, clarify the reasonable range of speech rate, pitch, and pauses, and establish the correspondence between emotional scenarios and acoustic parameters; S43. Based on the correspondence, perform prosody and semantic co-optimization, dynamically adjust the acoustic parameters of each response, and achieve matching of speech rate with semantic complexity, tone with emotional tendency, and pause with semantic logic. S44. Investigate and optimize the emotional mismatch problem during the process, eliminate unreasonable situations such as a cheerful tone conveying negative information and a low tone conveying positive information, and ensure that the voice output is strongly related to semantic emotion. S45. Integrate and optimize acoustic parameters and response semantics to generate a natural speech response that integrates sound and meaning, thereby improving the naturalness and emotional relevance of the dialogue.

[0020] The working principle and effects of the above technical solution are as follows: It analyzes the semantic and emotional tendencies of reliable responses, clarifies the emotional needs of different scenarios, reduces emotional judgment bias, and avoids a disconnect between voice output and semantic emotion. It establishes a correspondence between emotional scenarios and acoustic parameters, standardizes the adjustment criteria for speech rate, tone, and pauses, reduces the arbitrariness of parameter adjustments, and makes voice output more standardized. Through prosody and semantic collaborative optimization, it achieves accurate matching of acoustic parameters with semantics and emotion, improving the naturalness of voice response. It identifies and eliminates emotional mismatch problems, avoiding unreasonable situations such as conveying negative information with a cheerful tone, and enhances the correlation between voice output and emotion. It integrates the optimized parameters and semantics to generate a natural voice response that integrates sound and meaning, making the dialogue more in line with human communication habits, improving the user's auditory experience, reducing the decline in user experience caused by emotional mismatch or stiff voice, and further enhancing the friendliness of AI customer service.

[0021] In one embodiment of the present invention, S43 includes: S431. Retrieve the established correspondence between emotional scenes and acoustic parameters, and simultaneously obtain the semantic content and emotional tendency of each response in the trusted response set to complete data association and matching; S432. Initiate the prosody and semantic co-optimization process, perform semantic complexity analysis on the semantic content of each response, and classify it into three levels: simple, medium, and complex. S433. Based on the semantic complexity level and emotional tendency, dynamically adjust the acoustic parameters of the corresponding response to adapt to different semantic and emotional needs. S434. Verify the adjusted acoustic parameters to ensure that the speech rate matches the semantic complexity, the tone matches the emotional tendency, and the pauses match the semantic logic, thus completing the collaborative optimization.

[0022] The working principle and effects of the above technical solution are as follows: It retrieves the correspondence between emotional scenarios and acoustic parameters, synchronously associating the semantic content and emotional tendency of the response, reducing parameter adjustment deviations caused by data misalignment, and avoiding mismatches between acoustic parameters and emotions / semantics. It performs complexity analysis and grading of the response semantics, making parameter adjustments more targeted and reducing adjustment errors caused by ambiguous semantic complexity judgments. It dynamically adjusts acoustic parameters based on semantic level and emotional tendency to adapt to different expression needs and enhance the rationality of voice output. It verifies the adjusted acoustic parameters to ensure that speech rate, pitch, and pauses respectively match semantic complexity, emotional tendency, and semantic logic, achieving collaborative optimization. This not only makes the voice response more aligned with semantic and emotional needs but also improves the fluency and naturalness of voice expression, avoiding user auditory discomfort caused by abrupt adjustments, and further strengthening the integrated sound-meaning expression effect.

[0023] In one embodiment of the present invention, S433 includes: By combining semantic complexity level and sentiment tendency features, the core acoustic parameters that need to be adjusted in the response are selected; the speech rate parameter is adjusted according to the semantic complexity level to improve the fluency of complex semantic expression and adapt to the rhythm of simple semantic transmission. By altering tone parameters in conjunction with emotional inclination, the intensity of expression in negative scenarios is softened, while the effectiveness of conveying positive scenarios is enhanced. Adjust the pause parameters according to semantic logic nodes, add reasonable intervals at semantic transition points, and compress redundant intervals in continuous semantic segments; The adjusted speech rate, pitch, and pause parameters are integrated to form a complete acoustic parameter configuration that fits the current response expression requirements.

[0024] The working principle and effects of the above technical solution are as follows: Core acoustic parameters are selected by combining semantic complexity and emotional tendency, reducing invalid parameter adjustments, lowering adjustment costs, and avoiding parameter chaos caused by unnecessary operations. Speech rate is adjusted according to semantic complexity, making complex semantic expressions smoother and simple semantic transmission more efficient, avoiding excessively fast or slow speech rates that affect information delivery. Pitch is changed according to emotional tendency, softening the intensity of negative scene expressions and strengthening the effect of positive scene expressions, reducing emotional transmission deviations, and avoiding user misunderstandings caused by inappropriate pitch. Pause parameters are corrected according to semantic logic, adding reasonable intervals and compressing redundant intervals, making speech expression more consistent with semantic logic and avoiding semantic breaks caused by chaotic pauses. The adjusted parameters are integrated to form a complete and adapted parameter configuration, ensuring that speech expression conforms to semantic and emotional needs, improving the naturalness and comprehensibility of speech, further optimizing the user's auditory experience, and reducing the decline in dialogue experience caused by inappropriate parameters.

[0025] In one embodiment of the present invention, step S5 includes: S51. Conduct multi-granularity quality testing for natural speech responses that integrate sound and meaning, and set testing standards from four dimensions: compliance, naturalness, emotional fit, and information accuracy. S52. Collect real-time user feedback data, including user satisfaction ratings, conversation interruption rate, number of repeated questions, etc., and extract preset indicators for system operation, covering compliance rate, naturalness score, response latency, etc. S53. Weighted integration of user feedback data and system preset indicators, setting the weight ratio of each indicator, and generating a response quality score through quantitative calculation to achieve objectivity and comprehensiveness in quality assessment. S54. Compare the response quality score with the system's preset threshold to determine whether the current response quality meets the standard. If the score is higher than or equal to the threshold, retain the current model parameters; if the score is lower than the threshold, trigger the model iterative optimization process. S55. Based on the quality assessment results, update the joint semantic and verbal coding parameters and the counterfactual trajectory generation strategy, optimize the model's feature extraction and response generation capabilities, and continuously improve the credibility and naturalness of the dialogue system.

[0026] The working principle and effects of the above technical solution are as follows: It sets multi-dimensional standards for voice response quality detection, making quality assessment more comprehensive and avoiding evaluation bias caused by single-dimensional detection, ensuring accurate identification of various problems in the response. It collects user feedback and system operation indicators to enrich the sources of evaluation data, reduce the subjectivity of the assessment, and make the scoring more relevant to actual usage scenarios. It performs weighted fusion and quantitative calculations on the two types of data to improve the objectivity and accuracy of response quality scoring, avoiding bias from human evaluation. By comparing the score with thresholds, it accurately determines whether the response quality meets the standards, triggering timely model iteration and preventing poor responses from continuously affecting the user experience. Based on the evaluation results, it updates model parameters and strategies, optimizing model performance. This not only continuously improves the credibility and naturalness of the dialogue system but also reduces the probability of users asking repeated questions and dialogue interruptions, further reducing the intervention cost of human customer service and steadily improving the service quality of AI customer service.

[0027] In one embodiment of the present invention, S53 includes: The collected user feedback data and extracted system preset indicators are aggregated and standardized to eliminate calculation bias caused by differences in different data dimensions. Based on the core needs of AI customer service dialogue scenarios, differentiated weightings are set for each detection indicator, strengthening the influence weight of core indicators such as compliance and information accuracy, and reducing the weighting of secondary indicators. The standardized user feedback data and the system's preset indicators are weighted according to their respective weight proportions to obtain the weighted score of each individual indicator. The weighted scores of all individual indicators are summed to complete the quantitative calculation process and generate a comprehensive response quality score. The generated response quality scores are validated for reasonableness, and abnormal calculation results are excluded to ensure that the scores can comprehensively and objectively reflect the response quality level.

[0028] The working principle and effects of the above technical solution are as follows: It aggregates user feedback and system metrics, standardizes them to eliminate differences across data dimensions, reduces calculation bias, and avoids scoring distortion caused by inconsistent data specifications. It sets differentiated weights based on the core needs of AI customer service, strengthening the impact of core indicators such as compliance and information accuracy while reducing interference from secondary indicators, making the scoring more aligned with actual service needs. Weighted calculations are performed according to weights, and scores are aggregated to complete the quantification process, improving the accuracy of response quality scoring and reducing errors from human calculations. The scoring undergoes reasonableness verification to eliminate abnormal results, ensuring that the scoring comprehensively and objectively reflects response quality and preventing abnormal scoring from misleading subsequent model optimization. This approach guarantees both the objectivity and comprehensiveness of the scoring while making it more targeted, providing a reliable basis for subsequent quality judgment and model iteration, and reducing deviations in optimization direction caused by inaccurate scoring.

[0029] In one embodiment of the present invention, S55 includes: Summarize the quality assessment results, identify the core detection dimensions with low scores, and pinpoint the current performance shortcomings of the model; To address the shortcomings in model performance, the acoustic-semantic joint coding parameters were adjusted to optimize the fusion accuracy of acoustic features and semantic information and enhance the extraction effect of key features. Optimize the counterfactual trajectory generation strategy, adjust the adversarial example generation logic, expand the coverage of potential user follow-up questions, and improve the diversity of response candidate sets; Import the updated parameters and strategies into the model and conduct small-scale dialogue tests to verify the optimization effect of the parameters and strategies. Based on the test results, parameters and strategies are fine-tuned to form an iteratively optimized model configuration, continuously improving the credibility and naturalness of the dialogue system.

[0030] The working principle and effects of the above technical solution are as follows: By summarizing and evaluating the results, the performance shortcomings of the model are identified, making optimization more targeted, reducing meaningless parameter adjustments, and avoiding the waste of resources caused by blind optimization. Adjusting the acoustic-semantic joint encoding parameters improves the fusion accuracy of acoustic and semantic features, strengthens the extraction effect of key features, and improves the accuracy of subsequent response generation. Optimizing the counterfactual trajectory generation logic expands the coverage of user follow-up questions, enriches the response candidate set, and reduces the probability of missing dialogue scenarios. Importing the updated configuration for small-scale testing verifies the optimization effect and avoids service anomalies caused by directly deploying untested parameters. Fine-tuning the configuration based on the test results forms an iterative model solution that continuously improves the credibility and naturalness of the dialogue system, gradually reduces response bias and scenario omissions, and ensures a steady improvement in the overall service performance of AI customer service.

[0031] An embodiment of the present invention provides an AI-powered deep dialogue system for customer service based on a generative speech dialogue content model, comprising: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are made to implement the method described in any one of the above.

[0032] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. An AI-powered deep dialogue method for customer service based on a generative speech dialogue content model, characterized in that: The method includes: S1. Perform multimodal joint acoustic and semantic encoding on the user's voice input to generate a joint embedding vector that integrates acoustic features and semantic representation; construct an end-to-end integrated acoustic and semantic generation model based on the joint embedding vector to complete the direct mapping from speech signal to semantic response; S2. Generate counterfactual trajectories based on joint embedding vectors, simulate potential user inquiry paths by introducing adversarial examples, and generate a multi-dimensional response candidate set covering positive facts and reverse hypotheses; perform semantic consistency verification on the candidate set. S3. Perform factual credibility self-verification on the response candidate set after semantic consistency verification. By embedding policy knowledge graph and business rule engine, cross-verify the authenticity of key information involved in the response; mark and correct fictitious content, and generate a credible response set. S4. Based on the trusted response set, prosody-semantic co-optimization is performed. By dynamically adjusting the acoustic parameters, the speech output is strongly correlated with semantic emotion, eliminating the phenomenon of emotion mismatch and generating a natural speech response that integrates sound and meaning. S5. Perform multi-granularity quality assessment based on the integrated sound and semantic natural speech response, and generate a response quality score by combining user feedback data and system preset indicators; when the score is lower than the threshold, trigger the model iterative optimization process to update the joint sound and semantic coding parameters and the counterfactual trajectory generation strategy.

2. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 1, characterized in that, S1 includes: S11. Collect the user's real-time voice input signal, filter out environmental noise and interference ripples in the signal, retain clear and effective voice segments, and ensure the integrity and purity of the input signal. S12. Extract the acoustic low-level features from the effective speech segments to form a standardized acoustic feature set; S13. Perform multimodal joint acoustic-semantic encoding on the acoustic feature set and the text semantic information corresponding to the speech, and achieve deep fusion of acoustic features and semantic information through attention mechanism to generate a joint embedding vector that integrates acoustic features and semantic representation. S14. Input the joint embedding vector into the preset model architecture, complete the feature dimension calibration and feature weight allocation, and build an end-to-end integrated semantic generation model. S15. By utilizing an end-to-end integrated speech-semantic generation model, the mapping and conversion from speech signals to semantic responses can be directly realized, eliminating intermediate translation steps and avoiding the risk of error accumulation in each step of the cascaded architecture.

3. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 1, characterized in that, The S2 includes: S21. Extract key information from the joint embedding vector, sort out the dialogue logic, and clarify the core needs of the current dialogue. S22. Generate counterfactual trajectories based on the core needs of the dialogue, set counterfactual assumptions in different scenarios, and simulate the user's potential dialogue path. S23. Introduce an adversarial sample generation mechanism to generate diverse adversarial samples that conform to the user's potential dialogue path, enrich the coverage of dialogue scenarios, and improve the counterfactual trajectory system. S24. Combining positive fact responses and negative hypothesis responses, integrate and generate a multi-dimensional response candidate set covering multiple scenarios and intentions to ensure the comprehensiveness and diversity of the candidate set; S25. Perform semantic consistency verification on the multi-dimensional response candidate set, compare the fit of each candidate response with the historical dialogue intent and context logic, eliminate contradictory responses, and retain candidate content that conforms to the dialogue logic.

4. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 1, characterized in that, The S3 includes: S31. Extract key information from the response candidate set after semantic consistency verification and form a key information list; S32. Embed the key information list into the preset policy knowledge graph and business rule engine to establish a mapping between key information and knowledge graph and rule engine, so as to realize rapid matching and retrieval of information; S33. Through two-way cross-validation of policy knowledge graph and business rule engine; S34. Mark any fictitious content or information discrepancies discovered during the verification process, and correct the content in accordance with policy requirements and business specifications to ensure the compliance of the response information; S35. Integrate all the revised candidate responses to generate a set of credible responses that comply with business specifications and legal requirements.

5. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 1, characterized in that, The S4 includes: S41. Analyze the semantic sentiment of each response in the trusted response set, distinguish different emotional scenarios, and clarify the emotional needs corresponding to each response; S42. Determine the acoustic parameter standards corresponding to different emotional scenarios, clarify the reasonable range of speech rate, pitch, and pauses, and establish the correspondence between emotional scenarios and acoustic parameters; S43. Based on the correspondence, perform prosody and semantic co-optimization, dynamically adjust the acoustic parameters of each response, and achieve matching of speech rate with semantic complexity, tone with emotional tendency, and pause with semantic logic. S44. Investigate and optimize for emotional mismatch issues, eliminate unreasonable situations, and ensure that the voice output is strongly correlated with semantic emotion. S45. Integrate and optimize acoustic parameters and response semantics to generate a natural speech response that integrates sound and meaning, thereby improving the naturalness and emotional relevance of the dialogue.

6. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 5, characterized in that, S43 includes: S431. Retrieve the established correspondence between emotional scenes and acoustic parameters, and simultaneously obtain the semantic content and emotional tendency of each response in the trusted response set to complete data association and matching; S432. Initiate the prosody and semantic co-optimization process, perform semantic complexity analysis on the semantic content of each response, and classify it into three levels: simple, medium, and complex. S433. Based on the semantic complexity level and emotional tendency, dynamically adjust the acoustic parameters of the corresponding response to adapt to different semantic and emotional needs. S434. Verify the adjusted acoustic parameters to ensure that the speech rate matches the semantic complexity, the tone matches the emotional tendency, and the pauses match the semantic logic, thus completing the collaborative optimization.

7. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 6, characterized in that, S433 includes: By combining semantic complexity level and sentiment tendency features, the core acoustic parameters that need to be adjusted in the response are selected; the speech rate parameter is adjusted according to the semantic complexity level to adapt to the rhythm of simple semantic transmission. By altering tone parameters in conjunction with emotional inclination, the intensity of expression in negative scenarios is softened, while the effectiveness of conveying positive scenarios is enhanced. Adjust the pause parameters according to semantic logic nodes, add reasonable intervals at semantic transition points, and compress redundant intervals in continuous semantic segments; The adjusted speech rate, pitch, and pause parameters are integrated to form a complete acoustic parameter configuration that fits the current response expression requirements.

8. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 1, characterized in that, The S5 includes: S51. Conduct multi-granularity quality testing for natural speech responses that integrate sound and meaning, and set testing standards from four dimensions: compliance, naturalness, emotional fit, and information accuracy. S52. Collect real-time user feedback data and extract preset system operation indicators. S53. Weighted integration of user feedback data and system preset indicators, setting the weight ratio of each indicator, and generating a response quality score through quantitative calculation to achieve objectivity and comprehensiveness in quality assessment. S54. Compare the response quality score with the system's preset threshold to determine whether the current response quality meets the standard. If the score is higher than or equal to the threshold, retain the current model parameters; if the score is lower than the threshold, trigger the model iterative optimization process. S55. Based on the quality assessment results, update the joint semantic and verbal coding parameters and the counterfactual trajectory generation strategy, optimize the model's feature extraction and response generation capabilities, and continuously improve the credibility and naturalness of the dialogue system.

9. The AI ​​customer service deep dialogue method based on a generative speech dialogue content model according to claim 8, characterized in that, The S55 includes: Summarize the quality assessment results, identify the core detection dimensions with low scores, and pinpoint the current performance shortcomings of the model; To address the shortcomings in model performance, the acoustic-semantic joint coding parameters were adjusted to optimize the fusion accuracy of acoustic features and semantic information and enhance the extraction effect of key features. Optimize the counterfactual trajectory generation strategy, adjust the adversarial example generation logic, expand the coverage of potential user follow-up questions, and improve the diversity of response candidate sets; Import the updated parameters and strategies into the model and conduct small-scale dialogue tests to verify the optimization effect of the parameters and strategies. Based on the test results, parameters and strategies are fine-tuned to form an iteratively optimized model configuration, continuously improving the credibility and naturalness of the dialogue system.

10. An AI-powered deep dialogue system for customer service based on a generative speech dialogue content model, including: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method of any one of claims 1 to 9.