Adaptive ai phone answering method and system based on emotion recognition and intelligent decision-making

CN122845718APending Publication Date: 2026-09-29ANHUI JIUGUANG PANORAMIC INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610988930.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0010]鉴于现有技术的不足,本发明实施例提供一种基于情感识别与智能决策的自适应AI电话应答方法及系统,用以解决现有AI电话应答系统在电话窄带信道条件下情感识别精度不足、缺乏个性化适应、非发声段感知缺失、决策响应滞后且策略效果不可控的技术问题

Benefits of technology

[0019]本发明实施例提供的技术方案带来的有益效果至少包括:

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845718A_ABST
    Figure CN122845718A_ABST
Patent Text Reader

Abstract

The application discloses an adaptive AI telephone answering method and system based on emotion recognition and intelligent decision-making, and relates to the technical fields of artificial intelligence and voice signal processing. The method obtains a channel compensation matrix according to a voice coding standard type and extracts emotion invariant features through adversarial training; a frame-level emotion recognition value is converted into a personalized emotion offset by inferring an emotion expression baseline vector from a voiceprint feature; paralanguage features are obtained through typed detection of non-speech events and are fused to generate enhanced emotion representations; an emotion trend vector is constructed from first and second order time derivatives, and when a threshold is met, a weight is selected according to a strategy to trigger a forward-looking intervention strategy; the change in the personalized emotion offset after the strategy is executed is decomposed into a strategy net effect and a natural evolution component to trigger a strategy reinforcement, switching or rollback response. As a result, the emotion recognition accuracy under a narrowband channel is improved, the emotion perception is continuously covered in the speech segment and the non-speech segment, and the forward-looking intervention strategy is modified through a cause-and-effect closed loop.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of artificial intelligence and speech signal processing technology, and in particular to an adaptive AI telephone response method and system based on emotion recognition and intelligent decision-making. Background Technology

[0002] AI-powered telephone response systems are applied in scenarios such as customer service, telemarketing, and follow-up calls, using speech recognition and natural language processing to achieve human-computer voice interaction. Existing systems mostly use deep learning to extract acoustic features from the vocal segment for emotion classification. For personalization, they use a uniform standard for all users or rely on pre-stored profiles in a customer relationship management system. In terms of decision-making, they match preset strategy templates based on the current emotion and update them offline after the call ends.

[0003] However, existing technologies have at least the following shortcomings:

[0004] First, existing models are designed for broadband speech training and do not take into account the nonlinear distortion introduced by telephone codecs such as G.711, G.729, and AMR. This distortion distorts emotion-related acoustic features such as formants, fundamental frequency, and spectral tilt. However, telephone scenarios only have a single speech modality and lack a mechanism to compensate for distortion at the signal level and extract emotion features that are invariant to the coding standard type.

[0005] Secondly, existing technologies use a uniform absolute classification standard, which ignores individual differences in the speaker's emotional expression and is prone to misjudgment. Furthermore, relying on customer relationship management system profiles requires a pre-built database, which cannot handle users who are making their first call and have no history. It also lacks the ability to quickly establish a personalized baseline based solely on the current call voice.

[0006] Third, existing technologies only analyze vocal segments, while silence, pauses, and interruptions also carry emotions. They only use the pause ratio as a single scalar input and do not perform semantic typological modeling of non-vocal events. Emotional output is interrupted during silence, making it impossible to achieve continuous perception throughout the entire call.

[0007] Fourth, existing emotion-decision mapping is reactive and open-loop, failing to model the rate and acceleration of emotion change, thus only able to respond after deterioration and unable to predict and proactively intervene in the early stages. Furthermore, the actual impact is not evaluated after the strategy is implemented, and updates are mostly done offline after the call ends without a closed loop within the call, resulting in uncontrollable effects.

[0008] The fundamental reason is that existing technologies have not coordinated the design of channel compensation, personalized calibration, full-process call awareness, and forward-looking decision-making closed-loop feedback.

[0009] Therefore, there is an urgent need for an adaptive AI telephone response method and system that can achieve telephone channel distortion compensation, speaker-personalized baseline calibration, emotion perception throughout the call, and forward-looking intelligent decision-making and closed-loop feedback adjustment. Summary of the Invention

[0010] In view of the shortcomings of the prior art, the present invention provides an adaptive AI telephone answering method and system based on emotion recognition and intelligent decision-making, in order to solve the technical problems of insufficient emotion recognition accuracy, lack of personalized adaptation, lack of non-voice segment perception, delayed decision response and uncontrollable strategy effect of the existing AI telephone answering system under narrow-band telephone channel conditions.

[0011] In a first aspect, the present invention provides an adaptive AI telephone answering method based on emotion recognition and intelligent decision-making, executed by a telephone answering system, comprising:

[0012] S1: Based on the speech coding standard used in the current telephone call, obtain the corresponding channel compensation matrix, compensate for the nonlinear distortion introduced by the codec in the speech signal of the current call, and obtain the compensated speech signal; input the compensated speech signal into the feature extraction network trained adversarially, extract the emotion invariant features that are invariant to the speech coding standard type, and obtain the frame-level emotion recognition value based on the emotion invariant features; the emotion invariant features refer to the emotion features that make the emotion representation not change with the speech coding standard type after adversarial training; the frame-level emotion recognition value refers to the frame-by-frame emotion evaluation value obtained based on the emotion invariant features in the arousal and valence dimensions;

[0013] S2: Within a preset time window after the call is established, the voiceprint feature vector is extracted from the compensated speech signal. The current speaker's emotional expression baseline vector is inferred through a pre-trained voiceprint-emotion baseline mapping network. The frame-level emotion recognition value is converted into a personalized emotion offset based on the emotional expression baseline vector. The emotional expression baseline vector is a vector that represents the speaker's arousal and valence acoustic parameter baseline values ​​and emotional expression amplitude coefficient in a neutral emotional state. The personalized emotion offset is the standardized difference obtained by subtracting the neutral state baseline of the corresponding dimension in the emotional expression baseline vector from the frame-level emotion recognition value and dividing it by the emotional expression amplitude coefficient of the corresponding dimension.

[0014] S3: Perform typological detection and temporal modeling on non-vocal events during the call to obtain paralinguistic features. Fuse the paralinguistic features with personalized emotional offsets to generate an enhanced emotional representation. Non-vocal events include meaningful silence, hesitation, pauses, and interruptions. The duration classification threshold for meaningful silence is adaptively adjusted based on the emotional expression amplitude coefficient in the emotional expression baseline vector output in step S2. The enhanced emotional representation refers to the full-process emotional representation that integrates personalized emotional offsets for vocal segments and paralinguistic features for non-vocal segments.

[0015] S4: After pre-smoothing the enhanced emotional representation, the first and second time derivatives are calculated to construct an emotional trend vector. When the emotional trend vector meets the preset threshold condition, a strategy type is selected from the candidate strategy type set according to the strategy selection weight, and the corresponding level of prospective intervention strategy is triggered. The emotional trend vector is a vector composed of the first and second time derivatives, representing the rate and acceleration of emotional change. The prospective intervention strategy is a response strategy actively implemented before the valence component of the personalized emotional offset drops to the preset negative emotional threshold.

[0016] S5: After each forward-looking intervention strategy is implemented, the observed changes in personalized emotional shift are decomposed into the net effect of the strategy and the natural evolution component. Based on the direction and magnitude of the net effect of the strategy, a strategy reinforcement, strategy switching, or strategy rollback response is triggered, and the response results are fed back to step S4 to correct the preset threshold conditions and strategy selection weights of subsequent forward-looking intervention strategies. The strategy selection weight refers to the probability weight of each candidate strategy type being selected first. The natural evolution component refers to the inertial extrapolation prediction value of the emotion after the strategy is implemented based on the emotion trend vector in step S4. The net effect of the strategy refers to the difference between the actual observed changes in personalized emotional shift and the natural evolution component.

[0017] Secondly, the present invention provides an adaptive AI telephone answering system based on emotion recognition and intelligent decision-making, comprising: a processor and a memory;

[0018] The memory stores a program or instructions that can run on a processor, which, when executed by the processor, implement the steps of the adaptive AI telephone response method based on emotion recognition and intelligent decision-making as described in the first aspect.

[0019] The beneficial effects of the technical solutions provided in the embodiments of the present invention include at least the following:

[0020] (1) By extracting emotion invariant features that are invariant to the coding standard type through channel compensation and adversarial training, the accuracy of emotion recognition is improved in narrowband telephone channels without the need for additional modalities.

[0021] (2) By mapping the emotional expression baseline vector through voiceprint features and converting the frame-level emotional recognition value into a personalized emotional offset, the recognition bias caused by the speaker's expression difference can be eliminated, and a personalized baseline can also be established for users who are making their first call and have no history.

[0022] (3) By typifying non-vocal events such as silence, pauses, and interruptions and integrating them with personalized emotional offsets to form an enhanced emotional representation, the emotional perception is extended from the vocal segment to continuous coverage throughout the entire call.

[0023] (4) By using the first and second derivatives of the emotional trend vector, a forward-looking intervention strategy is triggered in the early stage of emotional deterioration. The emotional changes after the intervention are decomposed into the net effect of the strategy and the natural evolution component to quantify the causal effect of the strategy. Based on this, the subsequent threshold and strategy selection weight are corrected, forming a closed loop of decision-making, execution, evaluation and feedback within the same call. Attached Figure Description

[0024] Figure 1 This is a flowchart illustrating the adaptive AI telephone response method based on emotion recognition and intelligent decision-making provided in an embodiment of the present invention.

[0025] Figure 2 This is a schematic diagram of the structure of an adaptive AI telephone response system based on emotion recognition and intelligent decision-making, provided in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the present invention will now be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention.

[0027] The method of this invention is executed by a telephone response system, which has the ability to output response voice to the user and obtain the intent tag of the generated response text. In this invention, the user refers to the human party engaging in a conversation with the telephone response system; in the context of voiceprint and speech feature processing, it is also referred to as the speaker; both refer to the same object. The speech signal processed by the method of this invention comes from a narrowband telephone channel with a sampling rate of not less than 8kHz and is compressed and transmitted via a telephone network codec. The phonological and non-phonological segments of the speech are determined by a speech activity detection module, which labels each frame of speech based on a short-time energy threshold or a pre-trained binary classifier; its specific implementation is well-known in the art.

[0028] Reference manual attached Figure 1 This diagram illustrates a flowchart of an adaptive AI telephone answering method based on emotion recognition and intelligent decision-making, provided by an embodiment of the present invention. The method includes steps S1 to S5, each of which is described in detail below.

[0029] Step S1, Channel Compensation and Emotion Feature Extraction: Based on the speech coding standard type used in the current telephone call, obtain the corresponding channel compensation matrix, compensate for the nonlinear distortion introduced by the codec in the speech signal of the current call, and obtain the compensated speech signal; input the compensated speech signal into the feature extraction network trained by adversarial analysis, extract the emotion invariant features that are unchanged by the speech coding standard type, and obtain the frame-level emotion recognition value based on the emotion invariant features.

[0030] Step S101, Encoding Type Acquisition: Acquire the voice encoding standard type used in the current call. The encoding standard type identifier is obtained through the session description protocol media description information carried in the session initiation protocol signaling message. If the signaling negotiation information is unavailable, the voice encoding standard type is automatically inferred through the statistical characteristics of the voice bitstream. In one possible implementation, the set of voice encoding standard types includes, but is not limited to, G.711, G.729, G.726, AMR narrowband, and AMR wideband, indexed by encoding standard type. The sign, among which , The system pre-configures the total number of supported coding standard types. In one possible implementation, the total number of supported coding standard types is set to 5. When signaling is unavailable, automatic inference is performed by extracting at least one statistical measure from the bitstream, including frame length, bit rate, silent frame feature pattern, short-time energy distribution, zero-crossing rate, and inter-frame correlation, and inputting it into a pre-trained classifier. In one possible implementation, the pre-trained classifier uses a supervised classification model such as a support vector machine or a multilayer perceptron, and is trained offline using bitstream samples of known coding standard types.

[0031] Step S102, Channel Distortion Compensation: Based on the speech coding standard type obtained in step S101, select the corresponding channel compensation matrix from the pre-built codec distortion model library, apply the channel compensation matrix to the Mel-spectral feature vector of the speech signal, and compensate for the nonlinear quantization distortion introduced by the coding compression to obtain the compensated speech signal.

[0032] Extract the speech signal frame by frame Vermeer spectral eigenvectors ,in It is a frame index, starting from 0. Let be the number of filters in the Mel filter bank. In one possible implementation, the number of filters is 40, the frame length is 25 milliseconds, the frame shift is 10 milliseconds, and the corresponding frame rate is 100 frames per second. Based on the encoding standard type, the corresponding compensation parameters are selected from the distortion model library, and residual compensation is performed according to equation (1):

[0033]

[0034] in, For the first After frame compensation Vermeer spectral eigenvectors; For encoding standard type corresponding Channel compensation matrix; For encoding standard type corresponding Hidden layer weight matrix; For encoding standard type corresponding dimensional bias vector; The activation function is a nonlinear activation function. For example, the nonlinear activation function can be any one of a modified linear unit, a modified linear unit with leakage, or a Gaussian error linear unit. The channel compensation matrix, hidden layer weight matrix, and bias vector are all learnable parameters obtained through offline training. During pre-construction, for each coding standard type, the Mel spectrum of broadband clean speech is used as the target, and the Mel spectrum after encoding and decoding of the coding standard is used as the input. The parameters are obtained by supervised learning to minimize the mean square error. Each coding standard type corresponds to a set of independent parameters. Equation (1) adopts a residual compensation structure, in which the hidden layer linear transformation term models the distortion coupling relationship between frequency bands. The channel compensation matrix maps the hidden layer output to a frequency band-wise compensation amount, which is added back to the original Mel spectrum feature vector to achieve residual correction. The above residual structure makes the compensation amount only need to learn the distortion increment rather than the complete feature reconstruction, which reduces the model complexity and improves the training stability.

[0035] Step S103, Emotion Invariant Extraction: The compensated Mel-spectrum feature vector obtained in step S102 and the uncompensated Mel-spectrum feature vector are input in parallel into a dual-branch network. The emotion branch extracts emotion representations, and the channel branch predicts the speech coding standard type. This dual-branch network is pre-trained adversarially on the channel branch through a gradient inversion layer, forcing the output of the emotion branch to remain unchanged for the speech coding standard type. During runtime, the emotion branch extracts the emotion invariant features. The uncompensated Mel-spectrum feature vector is the Mel-spectrum feature vector mentioned in step S102.

[0036] In this implementation, the dual-branch network uses the concatenated and uncontracted Mel-spectrum feature vectors as input to the shared encoder, preserving the feature differences before and after compensation. The shared encoder outputs a sentiment representation. The sentiment branch uses this sentiment representation as input to predict the sentiment label, while the channel branch uses the sentiment representation processed by the gradient inversion layer as input to predict the speech coding standard type. In one possible implementation, the shared encoder consists of three cascaded one-dimensional convolutional layers and one long short-term memory network. The three one-dimensional convolutional layers have 64, 128, and 128 output channels respectively, and each has a kernel length of 3. The long short-term memory network has 128 hidden units. Both the sentiment branch and the channel branch are two fully connected networks with a hidden layer dimension of 128. The number of layers and dimensions of the above networks are set according to the actual engineering scenario.

[0037] The training objective function of the dual-branch network is a weighted sum of the emotion recognition loss and the channel classification loss. The channel classification loss is multiplied by an adversarial balance coefficient, which is a positive value used to adjust the weight ratio of the two losses. In one possible implementation, its value ranges from 0.1 to 1.0. The emotion recognition loss is used to constrain the emotion representation output by the emotion branch to approximate the emotion label. The emotion label is a two-dimensional truth vector of manually labeled arousal and valence dimensions in the training data. In one example, the emotion recognition loss uses a consistency correlation coefficient loss, i.e., taking one minus the consistency correlation coefficient between the emotion representation and the emotion label. The larger the coefficient, the higher the consistency and the smaller the loss. The channel classification loss is used to constrain the channel branch to predict the speech coding standard type based on the emotion representation, and uses a cross-entropy loss. The gradient inversion layer performs an identity mapping during forward propagation. During backward propagation, the gradient passing through this layer is multiplied by -1, thus inverting the sign of the gradient component of the emotion representation in relation to the channel classification loss. This causes the emotion branch parameters to be updated simultaneously in the direction of reducing emotion recognition loss and increasing channel classification loss during optimization, forcing the emotion representation to exclude information that can be used to distinguish coding standard types. The channel branch parameters themselves are still updated in the normal direction to maximize channel classification accuracy, forming an adversarial game. After training, the emotion branch parameters and channel compensation parameters are fixed for inference. Based on the emotion invariant features, a regression network obtains frame-by-frame emotion evaluation values, i.e., frame-level emotion recognition values, in both arousal and valence dimensions. In one possible implementation, the regression network is a two-layer fully connected network with a hidden layer dimension of 64 and an output layer containing two nodes corresponding to arousal and valence dimensions, respectively. These correspond to the two dimensions of arousal and valence, respectively. Specifically, the output layer of the regression network is activated using the hyperbolic tangent function, making... Positive values ​​represent positive emotional tendencies, negative values ​​represent negative emotional tendencies, and zero values ​​represent a neutral state.

[0038] Through the cascaded processing of the residual compensation and adversarial training mechanisms described above, step S1 can obtain emotion invariant features that are unaffected by the coding standard type under the single-modal constraint of telephone speech, which is different from the existing technology that relies on multimodal compensation. The obtained frame-level emotion recognition value is used as a unified frame-level input after distortion cancellation, and step S2 uses this input to calibrate the speaker's personalized baseline.

[0039] Step S2, Personalized Emotional Baseline Calibration: Within a preset time window after the call is established, the voiceprint feature vector is extracted from the compensated speech signal. The emotional expression baseline vector of the current speaker is inferred through a pre-trained voiceprint-emotional baseline mapping network. Based on the emotional expression baseline vector, the frame-level emotion recognition value is converted into a personalized emotion offset. Both the compensated speech signal and the frame-level emotion recognition value are obtained from the output of step S1.

[0040] Step S201, Voiceprint Feature Extraction: Within a preset time window after the call is established, a voiceprint feature vector is extracted from the compensated speech signal. The voiceprint feature vector includes the mean basic pitch, pitch dynamic range, average speech rate, average energy, and spectral tilt. The average speech rate is estimated based on a statistical measure of the number of fundamental frequency periods of voiced frames identified by fundamental frequency detection within the vocal segment, as determined by the speech activity detection module, relative to the total duration of the voiced frames. In cases where fundamental frequency periods cannot be extracted from unvoiced frames, the unvoiced frames are not included in the statistical measure. The fundamental frequency detection and speech rate estimation methods are well-known technologies in the field. The starting point of the preset time window is the moment the user first speaks. Specifically, the length of the preset time window is 3 to 5 seconds. In the initial stage of a call, before the preset time window has ended and before the voiceprint feature vector is available, the neutral state baseline is set to the group mean of the neutral state baselines of all speakers in the training dataset, and the emotional expression amplitude coefficient is set to the group mean of the emotional expression amplitude coefficients of all speakers in the training dataset. This reduces the standardized calculation of personalized emotional offset to a form based on group statistical parameters. The group mean is used as a conservative initial estimate of the neutral state baseline and emotional expression amplitude coefficient of the unknown speaker, ensuring that the standardization process in the initial stage maintains moderate sensitivity and avoids making aggressive assumptions about unconfirmed speaker traits. After the preset time window ends, the personalized parameters of the current speaker are switched. If the user does not speak within the preset time window after the call is established, the above group mean is maintained as the default baseline and default emotional expression amplitude coefficient of the current speaker until the user speaks for the first time and is recalculated, or until the call ends. The training dataset refers to the multi-speaker emotional speech dataset used to train the voiceprint-emotion baseline mapping network. The ground truth value of the neutral state baseline for each speaker is the mean of the frame-level emotion recognition value of the speaker in the neutral state in the corresponding dimension, and the ground truth value of the emotion expression amplitude coefficient is the standard deviation of the frame-level emotion recognition value of the speaker in each emotional state in the corresponding dimension.

[0041] Step S202, Emotional Baseline Inference: The voiceprint feature vector extracted in step S201 is input into the voiceprint-emotional baseline mapping network, which outputs an emotional expression baseline vector. The emotional expression baseline vector includes a neutral state arousal baseline. Neutral state valence baseline and the magnitude of emotional expression , For example, the voiceprint-emotion baseline mapping network is a three-layer fully connected network, trained offline on a dataset containing emotional speech from multiple speakers. This dataset includes speech samples from each speaker in neutral and various emotional states. The physical basis of the voiceprint-emotion baseline mapping lies in known statistical laws in the field, namely, there is a statistical correlation between voiceprint features such as a speaker's habitual speaking rate, average energy, and fundamental frequency characteristics and their emotional arousal, valence benchmark value, and emotional expression fluctuation range in a neutral state. Therefore, the emotional expression baseline vector can be inferred from the voiceprint feature vector. The specific network architecture, number of layers, number of speakers, and sample size in the dataset can all be set according to the actual engineering scenario. The output node corresponding to the emotional expression amplitude coefficient uses a softplus activation function to ensure that the output is always positive. The emotional expression amplitude coefficient reflects the speaker's long-term expression habits and does not change significantly within the time scale of a single call; therefore, it remains unchanged as the initial output value of step S202 during a single call.

[0042] Step S203, Emotional Shift Calculation: Subtract the neutral state arousal baseline output in step S202 from the frame-level emotion recognition value output in step S1 in the arousal dimension, and subtract the neutral state valence baseline output in step S202 in the valence dimension. Then divide each of these sub-values ​​by the coefficients of the corresponding dimensions in the emotional expression amplitude coefficients output in step S202 to obtain the personalized emotional shift. The personalized emotional shift is calculated according to formula (2):

[0043]

[0044] in, For the first Frame in dimension Personalized emotional offset; The output of step S1 Frame in dimension Frame-level emotion recognition values; Dimensions in a neutral emotional state The acoustic parameter reference values; For dimension The amplitude coefficient of emotional expression on the surface ; , Represents the wakefulness dimension. The valence dimension is represented. Equation (2) eliminates the systematic bias of individual neutral state by subtracting the baseline and normalizes the individual expression intensity difference by dividing by the amplitude coefficient, making the personalized emotional shift of different speakers comparable.

[0045] Step S204, Baseline Dynamic Correction: After the initial inference of the emotional expression baseline vector is completed in step S202, the neutral state arousal baseline and neutral state valence baseline in the emotional expression baseline vector output in step S202 are dynamically corrected during the call using an exponential moving average method. The dynamic correction is performed according to equation (3):

[0046]

[0047] Equation (3) is applicable to ; For the first Dimensions after frame dynamic correction Baseline value, initial value Output the value for step S202; The smoothing coefficient for the exponential moving average is... ; For the first The personalized emotion offset vector of the frame, where the superscript Indicates vector transpose; for The L2 norm; The preset neutrality threshold is used for determination. Preferably, The value is 0.95. The value is 0.3. Equation (3) only allows the current frame to participate in baseline correction when the magnitude of the personalized emotion offset is lower than the neutrality determination threshold, so as to avoid strong emotion frames polluting the baseline estimation and make the baseline only reflect the neutral state characteristics of the speaker and gradually converge with the call process. When the current frame is determined by the speech activity detection module to be a non-vocal segment and there is no frame-level emotion recognition value update, the baseline correction of Equation (3) is suspended and the current baseline value remains unchanged until the next valid vocal frame appears. During the call, the neutral state baseline in Equation (2) is replaced by the dynamically corrected baseline value after step S204 is started, that is, the personalized emotion offset of subsequent frames is calculated based on the latest dynamic baseline.

[0048] When the user remains in a non-neutral emotional state after step S204 is initiated, the neutral frame participation condition in equation (3) may not be met for an extended period, potentially causing the dynamic baseline to remain at the initial group mean for a prolonged time. Therefore, the system monitors the period from the start of the call (step S204) to the beginning of the call. The number of frames that meet the neutral frame participation condition within a time period of seconds; if this number of frames is lower than the preset minimum number of participating frames. At that time, the period for using the group mean as the neutral baseline is extended to the [number]th day after the start of the call. seconds and relax the neutrality threshold to ,in and These are the preset initial monitoring duration and the extended monitoring duration, respectively, in seconds. As a relaxation factor, If, even after the relaxation, sufficient neutral frames cannot be collected before the end of the extended monitoring period, the group mean baseline is maintained as the speaker's default baseline until the call ends. In one possible implementation, The value is 15 seconds. The value is 30 seconds. The value is 50 frames. The value is 1.5. The above degradation mechanism ensures that the method remains executable even in boundary scenarios where the user is in a non-neutral state at the beginning of the call.

[0049] Therefore, step S2 only requires voice data within a preset time window after the call is established to establish a personalized emotion assessment reference. Compared with existing technologies that use a unified absolute classification standard or rely on customer relationship management system profiles, the personalized calibration in this step is entirely based on the current call voice signal itself, and is equally applicable to users making their first call and having no history. The baseline estimate is continuously and dynamically corrected during the call through the conditional update mechanism of equation (3). The output personalized emotion offset is comparable across speakers and is generated by fusing step S3 with non-vocal segment paralinguistic features to produce an enhanced emotion representation.

[0050] Step S3, Non-voice Event Modeling and Emotion Enhancement: Type detection and temporal modeling are performed on non-voice events during the call to obtain paralinguistic features. These features are then fused with personalized emotion offsets to generate enhanced emotion representations. Both the personalized emotion offset and the emotion expression baseline vector are obtained from the output of step S2. When the user is in a silent period and the voice activity detection module determines it to be a non-voice segment, the personalized emotion offset remains the value calculated from the most recent valid voice frame before the start of the silent period, until the next valid voice frame appears. This ensures that the strategy effectiveness evaluation in step S5 maintains a continuous input data stream across silent periods.

[0051] Non-vocal events include meaningful silence, hesitation, pauses, and interruptions. Vocal segments and non-vocal segments are determined by the aforementioned speech activity detection module. The duration classification threshold for meaningful silence is adaptively adjusted based on the emotional expression amplitude coefficient in the emotional expression baseline vector output in step S2.

[0052] Meaningful silences are categorized into thinking silences, resistant silences, confused silences, and emotionally suppressed silences based on the semantic type of the most recent output from the telephone response system, and are indexed by subtype. The semantic type of the response is directly obtained by the telephone response system based on the intent tag of its most recently generated response text. The category space of the intent tag includes at least tags corresponding to the four semantic types of the response, without the need for additional speech recognition or natural language understanding modules. The intent tag is the internal output metadata attached by the telephone response system when generating the response text, and its category space includes at least categories such as questions, requests or instructions, complex explanations, and rejection or negative responses corresponding to the semantic type of the response. The specific category division and generation method of the intent tag is a technology known in the art. Among them, the semantic type of the response corresponding to thinking silence is a question, the semantic type of the response corresponding to resistant silence is a request or instruction, the semantic type of the response corresponding to confused silence is a complex explanation and is accompanied by a decrease in the user's speech rate before silence, and the semantic type of the response corresponding to emotional suppression silence is a rejection or negative response and is accompanied by a sharp drop in the user's base frequency before silence. The aforementioned decrease in user speech rate is determined by the decrease ratio of the average speech rate in the last sliding window before silence relative to the average speech rate of the previous several windows exceeding a preset decrease ratio threshold; the aforementioned sudden drop in user fundamental frequency is determined by the decrease ratio of the fundamental frequency of the last frame before silence relative to the average fundamental frequency of the previous several consecutive frames exceeding a preset sudden drop ratio threshold; in one possible implementation, the length of the sliding window is 500 milliseconds, the previous several windows are the previous 3 consecutive windows, the previous several consecutive frames are 10 frames, and both the decrease ratio threshold and the sudden drop ratio threshold are 15%. The silence duration threshold is adaptively adjusted according to formula (4). In one possible implementation, the silence duration threshold is mainly related to the speaker's arousal dimension expression amplitude, therefore the arousal dimension emotional expression amplitude coefficient is used. Adaptive adjustments can be made; in other implementations, the affective amplitude coefficient of the valence dimension or a linear combination of the coefficients of the two dimensions mentioned above can also be used for adjustment. The specific adjustment formula for the silence duration threshold is as follows:

[0053]

[0054] in, This indicates taking the larger of the two independent variables within the parentheses; For the adaptively adjusted first The threshold for classifying silence duration, in seconds; For the first The default duration threshold for classifying silence, in seconds. In one possible implementation, the default threshold for thoughtful silence... The default threshold for resistant silence is 2.0 seconds. The default threshold for confused silence is 3.0 seconds. The default threshold for emotionally suppressed silence is 2.5 seconds. It takes 1.5 seconds; This is the threshold sensitivity coefficient. ; The emotional expression amplitude coefficient of the current speaker arousal dimension output in step S2; To train the dataset, we need to find the group mean of the arousal dimension of the affective amplitude coefficient for all speakers. In extreme cases, if the population mean If it is zero, then... Set directly to To avoid division by zero; This is the lower limit of the silence duration threshold. This is used to prevent the adaptive adjustment result from being non-positive or too small. In one possible implementation, The value is 0.5. This was obtained based on statistical analysis of the training data. The value is 0.5 seconds. The mechanism of Equation (4) is as follows: Speakers with a larger amplitude coefficient of emotional expression tend to show stronger intention even in silence, so the classification threshold should be appropriately extended to reduce oversensitivity judgment on such speakers; conversely, speakers with a smaller amplitude coefficient may express obvious emotional changes with shorter silence, so the threshold should be shortened to improve sensitivity; lower limit The classification threshold is ensured to remain physically reasonable even under extreme amplitude coefficients. In actual detection, when the duration of a non-vocal segment determined by the speech activity detection module reaches or exceeds the duration classification threshold corresponding to its type, the non-vocal segment is determined to be meaningful silence of the corresponding type.

[0055] For hesitation pauses, a distinction is made between pauses within a sentence and pauses between sentences. The distinction between pauses within and between sentences is determined jointly based on the prosodic features of the speech segments before and after the pause, and the pause duration: if the pause duration is less than a preset pause-between threshold, and the decrease rate of the fundamental frequency and energy of the last frame before the pause relative to the average of the previous several frames is less than a preset breakage ratio threshold, it is determined to be a pause within a sentence; otherwise, it is determined to be a pause between sentences. In one possible implementation, the preset pause-between threshold is 800 milliseconds, and the preset breakage ratio threshold is 20%. The method extracts the pause duration distribution pattern and acoustic discontinuity of the preceding and following speech segments for hesitation pauses. Acoustic discontinuity is defined as the normalized Euclidean distance between the last frame before the pause and the first frame after the pause, based on three acoustic parameters: fundamental frequency, energy, and spectral centroid. Normalization is performed independently for each dimension based on the standard deviation of each acoustic parameter obtained from training data statistics. When the standard deviation of a certain acoustic parameter in the training data is zero, the normalization factor for that dimension is set to 1 to avoid division by zero. A larger acoustic discontinuity indicates a stronger interruption in the speech flow caused by the pause. Simultaneously, filler words appearing in the call are detected. These filler words are detected by template matching based on a pre-set filler word acoustic template library. This library is constructed from acoustic samples of commonly used filler words in the target language after feature extraction. The size and word composition of the filler word acoustic template library are set according to the actual language and application scenario. The acoustic features of the filler words' fundamental frequency, energy, and duration are extracted.

[0056] For interruption behavior, the timing, frequency, and energy fluctuation patterns before and after the user-initiated interruption are extracted. The user-initiated interruption is determined through dual-talk detection: when the telephone answering system detects a sudden increase in user-side voice energy exceeding a preset energy threshold and lasting for a duration exceeding a preset interruption duration threshold during voice output, it is determined to be a user-initiated interruption. The detection of user-side voice energy is supplemented by echo cancellation and noise suppression preprocessing known in the art to avoid false positives. In one possible implementation, the preset energy threshold is 1.5 times the average energy of the system's output voice, and the preset interruption duration threshold is 300 milliseconds.

[0057] After obtaining the detection results of the above three types of non-voice events, a time-series feature vector based on a sliding window is extracted for each type of non-voice event as a secondary language feature. The time-series feature vector includes the event density of each type of non-voice event within a unit time, the interval distribution of adjacent similar events, and the difference vector between each non-voice event and the personalized emotional offset of the preceding and following vocal segments. The unit time is the duration of the sliding window.

[0058] The temporal feature vector and personalized emotional offset are fused using a multi-head attention mechanism. In one possible implementation, the multi-head attention mechanism has four attention heads, and the hidden layer dimension of each head is consistent with the input dimension. The fusion uses the personalized emotional offset as the query vector, and the temporal feature vector and the personalized emotional offset are linearly projected to a unified dimension to serve as the key vector and value vector, respectively. The output of the multi-head attention is mapped back to the two dimensions of arousal and valence through a linear projection layer. The attention weights are calculated based on the signal-to-noise ratio of the current frame using logistic regression. Adaptive Function Adjustment: The logistic function, also known as the sigmoid function, is an S-shaped monotonically increasing function with a value range between 0 and 1. The steepness parameter multiplied by the difference between the current frame's signal-to-noise ratio (SNR) and a preset SNR threshold serves as the independent variable of this function, and its output is the adaptive SNR weight. The SNR is the logarithm of the energy ratio between the current frame's speech and the adjacent silent frames, measured in decibels (dB). The steepness parameter is a positive value used to control the transition rate of the weights near the threshold. In one possible implementation, the steepness parameter is set to 0.5, and the preset SNR threshold is set to 15 dB. In the multi-head attention mechanism, the attention score corresponding to the personalized emotion offset is multiplied by the adaptive SNR weight, and the attention score corresponding to the temporal feature vector is multiplied by a value minus the adaptive SNR weight, and then normalized. When the signal-to-noise ratio (SNR) is higher than the preset SNR threshold, the adaptive SNR weight approaches 1, with personalized emotional offset as the main weight source; when the SNR is lower than the preset SNR threshold, the adaptive SNR weight approaches 0, increasing the weight of the temporal feature vector. This adaptive mechanism enables the system to prioritize acoustic emotional features under high SNR conditions and shift to relying on noise-insensitive non-auditory event temporal features under low SNR conditions.

[0059] When a user is in the silence period of the non-voice event, the silence period sentiment estimate is output through a two-layer feedforward network based on the current silence type determination result, the enhanced sentiment representation of the last frame of the most recent voice segment before silence, and the historical mean vector of the enhanced sentiment representations output in previous frames. Specifically, the four-dimensional one-hot encoded vector of the current silence type, the two-dimensional enhanced sentiment representation vector of the last frame of the most recent vocal segment before silence, and the two-dimensional historical mean vector of the enhanced sentiment representations output in previous frames are sequentially concatenated into an eight-dimensional input vector. When silence occurs at the beginning of the call and there is no vocal segment before silence, the last frame enhanced sentiment representation vector is replaced by a zero vector. When there is no enhanced sentiment representation output during the call, the historical mean vector is replaced by a zero vector. The eight-dimensional input vector first passes through a first fully connected layer, where the activation function of the first layer is a modified linear unit, i.e., taking the larger value between the input element and zero for each element. The hidden layer dimension is preferably 32. Then, it passes through a second fully connected layer with a hyperbolic tangent activation function, outputting a silence period sentiment estimate containing two dimensions: arousal and valence. The hyperbolic tangent function constrains the output to the range of -1 to 1. The weight matrix and bias vector of the above two feedforward networks are obtained offline through training on an emotional speech dataset labeled according to the correspondence rules between response semantic types and silence subtypes as described in claim 4. The aforementioned feedforward network infers the emotional state during the silence period by integrating three inputs: semantic information of the silence type, acoustic state before silence, and emotional historical trend, thus ensuring continuous output of emotional perception during the silence period.

[0060] By fusing personalized emotional biases from the vocal segment and paralinguistic features from the non-vocal segment through the aforementioned multi-head attention mechanism, the attention fusion output from the vocal segment, activated by the hyperbolic tangent function, is used as an enhanced emotional representation, with its value range constrained within... Within the range, the value range is consistent with the value range of the silent period sentiment estimate; during the silent period, the silent period sentiment estimate is used as the enhanced sentiment representation, thereby achieving continuous and uniform enhanced sentiment representation output throughout the call. , Each frame outputs a two-dimensional vector. .

[0061] Through the typological detection and temporal fusion of the three types of non-vocal events—silence, pauses, and interruptions—step S3 expands emotion perception from covering only vocal segments to the entire call, unlike existing technologies that only treat the proportion of pauses as a single scalar prosodic feature. Equation (4) adaptively adjusts the classification threshold for meaningful silence based on the speaker's personalized baseline, linking the detection threshold for non-vocal events with the speaker's individual characteristics output in step S2. The output enhanced emotion representation remains continuous throughout the call, serving as the time-series input for emotion trend analysis in step S4.

[0062] Step S4, Sentiment Trend Analysis and Prospective Intervention: After pre-smoothing the enhanced sentiment representation, the first and second time derivatives are calculated to construct the sentiment trend vector. When the sentiment trend vector meets the preset threshold conditions, a strategy type is selected from the candidate strategy type set according to the strategy selection weight, and the corresponding level of prospective intervention strategy is triggered. The enhanced sentiment representation is obtained from the output of step S3.

[0063] Step S401, Temporal Derivative Estimation: A linear regression is performed on the frame-level sequence of the enhanced sentiment representation using a sliding window, and the regression slope is used as the estimate of the first-order temporal derivative. This linear regression fitting process itself is the pre-smoothing process described in Step S4, which suppresses frame-by-frame noise by least-squares fitting of multiple frames of data within the window. Let the enhanced sentiment representation be in dimension... The frame-level sequence value on The first-order time derivative is calculated according to equation (5):

[0064]

[0065] in, For the first Frame in dimension The first-order time derivative estimate on the time surface characterizes the rate of change in sentiment. For the first Frame in dimension Enhanced sentiment representation values ​​on; The width of the sliding window. ; This is the frame offset index within the window. ; Indicates the frame offset index within the window From the lower limit Up to the upper limit Summing up each term. The denominator of equation (5) For any All are greater than zero. Equation (5) is essentially based on... The least squares linear regression slope within the symmetrical sliding window centered on the frame is used to suppress frame-by-frame noise fluctuations by linearly fitting the local frame sequence, thus obtaining a robust rate of change estimate.

[0066] Since equation (5) requires the first After the frame The data of the nth frame, in the actual system, is the nth frame. The first time derivative of the frame is at the 1st Calculations can only be performed after a frame arrives, i.e., this introduces... Frame processing latency. Due to this Frame processing latency, the first Equation (5) of the frame is calculated in the first... The process begins after frame acquisition arrives, therefore the calculations required above... Each future frame is already in the buffer and available during computation. During the initial phase of the call... When the required historical frames are unavailable because they do not yet exist, a reduced window is used. That is, the symmetrical alignment calculation is performed using the number of currently available historical frames as half the width. The reduced half width is not less than 1 to ensure that the denominator is always positive. Therefore, equation (5) starts from the second frame. The calculation begins from the first frame. The first time derivative is set to zero.

[0067] The rate of change of the regression slope between adjacent windows is used as an estimate of the second time derivative, which is calculated according to equation (6):

[0068]

[0069] in, For the first Frame in dimension The second-order time derivative estimate on the time surface represents the acceleration of sentiment change; and The front and rear distances are respectively The two first-order time derivatives of the frame. The calculation of equation (6) requires... and Both are applicable, therefore equation (6) is derived from... Start calculating from here, in In the initial stage, the second-order time derivative is set to zero. At this time, the sentiment trend vector is composed only of the first-order time derivative. The preset threshold condition mentioned in step S4 is determined based solely on the first-order time derivative in this initial stage. Equation (6) calculates the rate of change of the slope using the central difference method, which together with equation (5) constitutes the sentiment trend vector. .

[0070] Half width of the sliding window The window size is dynamically adjusted based on the number of dialogue turns in the current call: a smaller window is used in the early stages of the call when there are fewer turns to improve temporal resolution, and the window size is gradually increased as the call progresses to improve smoothness. In one possible implementation, let the current number of dialogue turns be... One dialogue turn is defined as the interaction pair in which the telephone answering system completes one answer output and the user completes one response. ,in This is the floor operator, which takes the largest integer not greater than a given value. This indicates taking the smaller of the two independent variables within the parentheses. and These are the lower and upper limits of the window's half-width, respectively. This represents the number of saturation rounds. In one possible implementation, Take 3 frames. Take 10 frames. Take 5 rounds; the half width of the sliding window At the start of each round of dialogue, the cumulative number of dialogue rounds counted at the end of the previous round is updated once using the above formula. Within the same round of dialogue... The values ​​remain unchanged to ensure scale continuity in the calculation of the first and second time derivatives. At a frame shift of 10 milliseconds... , Under this configuration, the processing delay for the second-order time derivative is 2. A frame of 60 to 200 milliseconds is within an acceptable range for the time scale of strategic decision-making in a telephone answering system.

[0071] Step S402, Emotional State Smoothing: After calculating the first and second time derivatives in step S401, a Hidden Markov Model is constructed, using the emotional state of each frame as the hidden state. The first and second time derivatives obtained in step S401 are combined with the enhanced emotional representation as observations. Wherein, the... The observation vector of the frame is The model comprises six components: an enhanced sentiment representation value with two dimensions, a first-order time derivative, and a second-order time derivative. In one implementation, each component of the observation vector undergoes independent Z-score standardization before being input into the Hidden Markov Model (HMM), ensuring that the mean of each component is zero and the variance is one. This eliminates the impact of the magnitude difference between the enhanced sentiment representation and its time derivative on the emission distribution covariance estimation. The mean and variance parameters used for standardization are obtained offline based on training data during the training phase and are kept fixed and reused during the training of emission distribution parameters and subsequent inference phases to maintain dimensional consistency between the training and inference phases. The Viterbi decoding output of the most likely state sequence is used as the smooth state sequence of the enhanced sentiment representation. When the state transition probability corresponding to the hidden state inference result between adjacent frames is lower than a preset transition probability threshold, the frame is determined to be a non-smooth transition frame. Subsequent mean replacement processing is performed on the non-smooth transition frame. For example, the transition probability threshold is set to 0.05. In real-time processing, Viterbi decoding is performed within a fixed-length sliding window, the window length of which is synchronized with the sliding window in step S401. Each time a new frame is reached, the window slides forward one frame and the observation sequence within the window is re-decoded. In one possible implementation, for the enhanced sentiment representation value of the non-smooth transition frame, the original observation vector is replaced with the mean vector of the six-dimensional multivariate Gaussian emission distribution corresponding to the hidden state of the non-smooth transition frame. Two components of the mean vector corresponding to the enhanced sentiment representation are taken to form a new enhanced sentiment representation sequence. Based on the new enhanced sentiment representation sequence, the first and second time derivatives are recalculated according to equations (5) and (6) to construct the sentiment trend vector. The replacement is only used for the internal sequence when constructing the sentiment trend vector and does not change the enhanced sentiment representation output to downstream steps. In one possible implementation, the hidden state set of the Hidden Markov Model includes three discrete states: positive sentiment, neutral sentiment, and negative sentiment. The specific number and division of the hidden states can be adjusted according to the actual needs of sentiment modeling. The mapping relationship between the hidden states and the continuous values ​​of the enhanced sentiment representation is realized through the six-dimensional multivariate Gaussian emission distribution corresponding to each hidden state. The mean, covariance matrix, and transition probability matrix between hidden states are jointly estimated on the training data using the Baum-Welch algorithm. The training data includes call voice samples with manually labeled sentiment states, and its specific size is determined by those skilled in the art based on the dimension of the observation vector and the complexity of the state space.

[0072] When the sentiment trend vector meets a preset threshold condition, a strategy type is selected from the candidate strategy type set according to the strategy selection weight, and a corresponding level of forward-looking intervention strategy is triggered. The levels include warning level, intervention level, and emergency level. The candidate strategy type set includes four strategy types: speech slowing, empathic language, benefit compensation, and switching to human intervention. When each level is triggered, a strategy type is sampled from the corresponding subset according to the current strategy selection weight: warning level samples from the subset consisting of speech slowing and empathic language; intervention level samples from the subset consisting of empathic language, benefit compensation, and speech slowing; and emergency level always selects switching to human intervention. The forward-looking intervention strategy of the emergency level is executed immediately through a preset polite interruption mechanism. This polite interruption mechanism is triggered when the user's voice energy is detected to be below a preset interruptible energy threshold and the duration exceeds a preset interruptible duration threshold, interrupting the current system output with a preset polite interjection statement. This polite interruption mechanism is a well-known technology in the field of voice dialogue systems.

[0073] When the components of the first-order time derivative in the valence dimension Below the first preset threshold When this occurs, a proactive intervention strategy at the warning level is triggered, and a candidate set of reassurance phrases is preloaded. In one example, the candidate set of reassurance phrases consists of pre-configured empathic and emotionally acknowledging response texts, the specific content of which is set according to the actual application scenario. In one possible implementation, The value is .

[0074] When the components of the first-order time derivative in the valence dimension Below the second preset threshold Furthermore, the components of the second-order time derivative in the valence dimension When the value is less than zero, a proactive intervention strategy at the intervention level is triggered, and a reassurance strategy is actively implemented. Unlike the emergency level, which immediately interrupts the current system output, the intervention level does not interrupt the current system output. The reassurance strategy triggered is executed along with the response in the next round of the telephone answering system. In one possible implementation, The value is .

[0075] When the valence prediction value after a preset number of future frames based on the extrapolation of the sentiment trend vector is lower than the third preset threshold, an emergency-level prospective intervention strategy is triggered, and the plan to switch to manual service is initiated. The extrapolation prediction is calculated according to equation (7):

[0076]

[0077] in, For from the first Frame forward push Post-frame valence prediction; For the first Enhanced sentiment representation value in the frame valence dimension; and The first First and second time derivatives of the frame valence dimension; This is a preset prediction time span, in frames. For example, The value is 20 frames. When When the emergency level is triggered, In one possible implementation, the value is... . This is used only as an extrapolation signal to trigger the determination and does not represent the actual observed value. Its value may exceed the physical range of the measured value. Due to the existence of calculation problems in equations (5) and (6) The frame processing delay, the extrapolation starting point of equation (7) is the latest available time after the delay, therefore the effective future prediction time span relative to the current actual processing time is Frame, when , The effective prediction span is 14 frames. When the value is increased to 10, the effective prediction span is shortened to zero. At this point, the extrapolation calculation of equation (7) is no longer performed, and the valence component of the current enhanced sentiment representation is instead calculated. With the third preset threshold A direct comparison is used to determine whether an emergency level has been triggered. After an emergency level is triggered, the system initiates a transfer request to a human agent. While waiting for a human agent to connect, steps S1 to S5 of this method continue to be executed to maintain emotion perception and response strategy output until a human agent successfully connects, at which point the automatic response process of this method terminates. Equation (7) predicts the short-term evolution of valence based on second-order Taylor expansion. It also utilizes two information sources, the rate of change and the acceleration of change. Compared to reactive judgment based solely on the current emotion label, it can predict the trend of emotion when it is still in the early stages of deterioration.

[0078] When the level of a proactive intervention strategy is switched, the parameters of the preceding and following strategies are weighted using a raised cosine transition weight. The transition weight ranges from 0 to 1, smoothly decreasing from 1 to 0 as the number of frames elapsed since the strategy level switch. Its value is half the sum of one plus a cosine term, where the phase of the cosine term is pi multiplied by the number of elapsed frames divided by the transition duration. The transition duration is a positive integer of frames not less than 1, and in one possible implementation, it is 10 frames. The number of elapsed frames ranges from 0 to the transition duration. During the switch, the strategy parameters for the current frame are calculated by multiplying the transition weight by the strategy parameter vector before the switch, adding one and subtracting the transition weight, and then multiplying this value by the strategy parameter vector after the switch. The strategy parameter vector includes control parameters such as speech rate, tone softness, and reassurance intensity. The strategy parameter vectors corresponding to each level are preset through offline configuration, with the reassurance intensity of the warning level being lower than that of the intervention level, and the intervention level being lower than that of the emergency level. The raised cosine weighted gradual form described above achieves a smooth transition from the previous policy to the next policy, avoiding abrupt changes in response style caused by policy switching.

[0079] Through the processing described in equations (5) to (7) and the raised cosine transition above, step S4 extracts the rate of change and acceleration of emotion from the enhanced emotion representation sequence to construct an emotion trend vector. When the emotion trend vector meets a preset threshold condition, a strategy is selected from the candidate strategy type set according to the strategy selection weight, and a prospective intervention strategy is triggered in a hierarchical manner. The hidden Markov model is used to smooth out and eliminate false trend triggers caused by frame-level noise. This trend-based decision-making mode differs from the reactive mode of matching strategies after identifying the current emotion state in existing technologies. It can predict and actively intervene in the early stage of emotion deterioration. The triggered prospective intervention strategy and the corresponding emotion trend vector at the time are used as inputs for step S5 to evaluate the causal effect of the strategy.

[0080] Step S5, Strategy Effect Evaluation and Closed-Loop Feedback: After each execution of the prospective intervention strategy, the observed changes in personalized emotional shift are decomposed into the net effect of the strategy and the natural evolution component. Based on the direction and magnitude of the net effect, strategy reinforcement, strategy switching, or strategy rollback responses are triggered, and the response results are fed back to Step S4 to correct the preset threshold conditions and strategy selection weights for subsequent prospective intervention strategies. The prospective intervention strategy is triggered by Step S4, the personalized emotional shift is continuously obtained from the output of Step S2, and the emotional trend vector is obtained from the output of Step S4. During the silent period, which is determined by the speech activity detection module to be a non-vocal segment, the personalized emotional shift is maintained according to the maintenance mechanism described in Step S3 at the value calculated from the most recent valid vocal frame before the start of the silent period, for use in calculating the net effect of the strategy in this step.

[0081] Step S501, Observation Window Construction: Define the observation window after the implementation of the prospective intervention strategy, including the length of the observation window. The observation window length is dynamically adjusted based on the currently implemented strategy type. Preferably, the observation window length for intervention-level triggered reassurance strategies, benefit compensation strategies, speech slowing strategies, and empathy-based rhetoric strategies is 15 frames; the observation window length for emergency-level triggered human intervention plans is 25 frames. The above observation window lengths can be adjusted by those skilled in the art based on the prior expectations of the response characteristics of different strategy types. Warning-level strategies only preload the rhetoric candidate set but do not actually implement the strategy and do not enter the observation and evaluation process of this step. When a new round of forward-looking intervention strategies is triggered before the current observation window has ended, the previous round of observation windows is immediately terminated, the net effect of the strategy in this round is calculated based on the collected frame data, and a new round of observation windows is opened; if the number of collected frames is less than the preset minimum effective number of frames... If so, the net effect evaluation of this round of strategy is abandoned, and the strategy selection weights are not updated. Specifically, The value is 5 frames. Enhanced sentiment representations are continuously collected within the observation window, and the sentiment trend vector before the current policy is executed is recorded. The execution start frame is denoted as . The number of completed historical observation windows will be recorded when the call is established. Initialize to zero, and after each observation window completes the calculation of the net effect of the strategy... Add one.

[0082] Step S502, Net Effect Decomposition: The evolution of personalized emotional shifts within the observation window is predicted by inertial extrapolation using the emotional trend vector recorded in Step S501, yielding the natural evolution component. In this step, personalized emotional shifts, rather than enhanced emotional representations, are used as the effect assessment object: the enhanced emotional representations integrate non-vocal segment paralinguistic features, designed to improve the early warning sensitivity of emotional trends to drive the prospective intervention in Step S4; the personalized emotional shifts, after eliminating individual speaker baseline differences, maintain cross-speaker comparability and are more suitable as a standardized indicator to measure the true impact of the strategy intervention. Because the enhanced emotional representations are activated by the hyperbolic tangent function... Within the specified range, and with the derivative of the hyperbolic tangent function near zero being 1, when the amplitude of the personalized emotional shift is small and has not clearly entered the saturation region, the relationship between the enhanced emotional representation and the personalized emotional shift is approximately linear. In this case, the first and second components of the emotional trend vector, divided by the emotional expression amplitude coefficient, can approximately reflect the changing trend of the personalized emotional shift. When the proportion of silent frames within the observation window exceeds the preset silence dominance threshold, since the personalized emotional shift remains at the value of the most recent effective vocal frame during the silence period, the net effect evaluation of the strategy is changed to enhanced emotional representation. according to The personalized emotional shift is obtained by converting the local linear approximation relationship. The measured and counterfactual terms in the subsequent net effect calculation of the strategy are substituted with this converted personalized emotional shift, and the conversion methods for both arousal and valence dimensions are consistent. In one possible implementation, the preset silence dominance threshold is set to 50%. The local linear approximation holds true when the enhanced emotional representation has not entered the hyperbolic tangent function saturation region. Based on the above-mentioned engineering local approximation relationship, the natural evolution component approximates the short-term evolution within the effect observation window through inertial extrapolation. Its accuracy constraint is dynamically compensated by the subsequent prediction confidence assessment mechanism, and the emotional expression amplitude coefficient is also adjusted. Set the preset lower limit of amplitude. When any dimension Less than Instead of the inertial extrapolation, the difference between the mean of the personalized sentiment offset within the observation window and the personalized sentiment offset at the start frame of the policy execution is used as a substitute value for the net effect of the policy. For example... Take 0.05 and calculate according to formula (8):

[0083]

[0084] in, Assuming no policy intervention was implemented, the first Frame in dimension The inertial extrapolation prediction value of personalized sentiment offset; Start frame for strategy execution In Dimensions The personalized emotional offset is used as the starting value for extrapolation; and They are respectively The dimension of the sentiment trend vector output at time step S4 The first and second time derivative components on; The dimension output by step S2 The emotional expression amplitude coefficient is used to convert the rate of change of the enhanced emotional representation space to the personalized emotional offset space. For the frame offset index within the observation window, . The extrapolated predictions used only as counterfactual baselines do not represent actual observations. Equation (8) executes the initial frame according to the strategy. Starting with the personalized emotional offset mentioned above, by dividing by After scaling the sentiment trend vector, a counterfactual baseline is constructed, which is the natural evolution prediction of personalized sentiment bias without policy intervention. When the magnitude of personalized sentiment bias is large enough to cause the hyperbolic tangent function to enter the saturation region, the accuracy of the above linear approximation decreases. At this time, the subsequent prediction confidence mechanism automatically reduces the weight of this round of evaluation to compensate for the approximation error.

[0085] The net effect of the strategy is obtained by subtracting the natural evolution component from the actual observed change in personalized sentiment shift within the observation window. The net effect of the strategy is calculated according to equation (9):

[0086]

[0087] in, This represents the valence component of the net effect of the strategy. For the first observation window The personalized sentiment shift in the valence dimension actually observed in the frame; The corresponding frame natural evolution prediction value calculated by equation (8); The observation window length is given. Equation (9) compares the difference between the actual observation mean and the counterfactual baseline mean to eliminate the interference of natural emotional fluctuations and quantify the causal effect of the strategy intervention itself on the user's emotions.

[0088] In one implementation, after obtaining the net effect of the strategy, the reliability of the net effect is also assessed. The prediction confidence is calculated based on the cumulative deviation between the predicted values ​​of the natural evolutionary component within each historical observation window and the actually observed personalized sentiment shift at the corresponding time. This is done when the number of completed historical observation windows... Insufficient to the preset minimum number of windows At that time, instead of executing equation (10), the prediction confidence level is directly applied. Set to the preset initial confidence level. ;when When the prediction confidence is calculated, the prediction confidence is calculated according to equation (10):

[0089]

[0090] in, To predict confidence levels; This represents the number of historical observation windows that have been completed in the current call. For the first The mean vector of personalized emotional shifts actually observed within a historical observation window in terms of both arousal and valence dimensions; For the first The mean vector of personalized emotional shift predicted by inertial extrapolation within a historical observation window in terms of both arousal and valence dimensions; It is an L2 norm; This represents an exponential function with the natural constant e as its base; It is the numerical stability constant. This is to prevent the denominator from being too small or zero, which could lead to unstable values. In one example, The value is Equation (10) evaluates the predictive ability of the inertial extrapolation model for the current emotional evolution of the call by taking the arithmetic mean of the normalized prediction errors of each historical observation window and then performing a negative exponential mapping. When the prediction bias is large, the prediction confidence approaches zero, indicating that the reliability of the inertial extrapolation model is reduced. At this time, the influence of the response triggered by the net effect of the strategy on the adjustment of the strategy selection weight is reduced; when Below the preset reliability threshold If the inertial extrapolation model is deemed unreliable, the direct change in personalized sentiment shift before and after the observation window is used as the basis for downgrading the net effect of the alternative strategy. In one possible implementation, The value ranges from 0.1 to 0.3, with a typical value of 0.2; The specific value can be adjusted according to the actual engineering scenario. During downgrade evaluation, the difference between the mean of the personalized sentiment shift on the valence dimension within the observation window and the personalized sentiment shift at the start frame of policy execution is used as the substitute value for the net effect of the policy; that is, the substitute value equals... This alternative value does not rely on an inertial extrapolation model, but is based solely on direct observations before and after policy execution. When the number of historical observation windows is less than the preset minimum window number... At that time, the prediction confidence level is set to the preset initial confidence level value. In one example, The value of is 2. The value is 0.5.

[0091] Step S503, Response Trigger: Based on the component values ​​of the net policy effect in the valence dimension obtained in step S502. Trigger one of the following three types of responses:

[0092] when When the strategy is executed to reinforce the response, the strategy selection weight of the current strategy type is increased. Let the set of strategy types contain... Types of strategies, This identifies the current strategy type. The initial value for the strategy selection weight for each strategy type is [value missing]. In one possible implementation, The value is 4, corresponding to four strategy types: slowing down speech, using empathetic language, offering benefits, and transferring to a human agent. The specific types and number of these strategies can be expanded or adjusted according to the actual application scenario. Updated strategy selection weights. Before the update Plus ,in For weighted learning rate, ; Let be the prediction confidence level of equation (10). In one possible implementation, The value is 0.1.

[0093] when When this happens, a policy switching response is executed, reducing the policy selection weight of the current policy type and switching to an alternative policy type. For example, the updated... Equal to before the update Multiply and from the current strategy type The rest Candidate strategy types are selected by sampling from the various strategy types according to the updated weight distribution. To preset the rollback threshold, In one possible implementation, The value is 0.15.

[0094] when When this happens, the execution policy rollback response is initiated, the current policy is terminated, and the updated policy is applied. Set to before update Multiply The process involves restoring the strategy state from the strategy history stack to the most recent strategy state where the net effect of the strategy has a positive valence component. The restored strategy state includes the strategy type and its parameter vector at the corresponding historical moment, and inserts transitional remedial dialogue. This transitional remedial dialogue is a pre-configured transitional script designed to alleviate the discontinuity in the dialogue caused by strategy termination; its specific content is set according to the actual application scenario. The strategy history stack records the strategy types and their corresponding net effects for each round in chronological order. The strategy history stack is initialized to empty when the call is established. In one example, the maximum capacity of the strategy history stack is no more than 100 records; if the maximum capacity is exceeded, the earliest record is discarded based on its recording time. When executing a strategy rollback response, the process traverses backward from the most recent record in the strategy history stack to locate the record where the net effect of the strategy has a positive valence component. If there are no records with positive net effect components in the strategy history stack, the strategy type is resampled and selected according to the current strategy selection weight distribution.

[0095] After updating the strategy selection weights for the three types of responses described above, an iterative truncation-normalization operation is performed to simultaneously satisfy the two constraints: the sum of the weights must be one and the lower bound of the minimum exploration probability must be met. The iterative truncation-normalization operation first normalizes the weights of all strategy types so that their sum is one, and then sets the weights below the preset minimum exploration probability. Strategy type truncation to If the sum of all weights after truncation is still not equal to one, then the weights of the untruncation strategy types are scaled proportionally to restore the sum to one; if new weights appear lower than one after scaling... For the strategy type, repeat the truncation and scaling until all weights are no less than [a certain value]. And the sum is one, or the current result is maintained when the preset maximum number of iterations is reached. In one possible implementation, The value is The maximum number of iterations is 10. This iterative mechanism ensures that the strategy selection weights always satisfy the normalization condition and the minimum exploration probability constraint, preventing strategy selection from degenerating into determinism and losing its exploratory capability. When a strategy reinforcement response increases the weight of the current strategy, subsequent normalization operations correspondingly compress the weight space of other strategies. If the weights of other strategies fall below the lower bound of the minimum exploration probability due to compression, the iterative truncation mechanism will compensate other strategies with a portion of the current strategy's weight increment, achieving a balance between reinforcing effective strategies and maintaining exploratory capabilities.

[0096] Step S504, Feedback Parameter Correction: The response result of step S503 is fed back to step S4 to correct the preset threshold conditions for subsequent prospective intervention strategies. The updated results of the strategy selection weights in step S503 are synchronously used for strategy type selection when triggering the prospective intervention strategy in subsequent step S4. Specifically, the threshold correction rule is: when the net effect of the strategy in two consecutive observation windows with adjacent times within the same call is positive in terms of the valence dimension, the first preset threshold in step S4 is adjusted. Second preset threshold Divide by tightening factor respectively ,because and All are negative values. and Make Therefore Dividing by this factor increases the absolute value of the threshold and makes it more negative, thus making the triggering conditions more difficult to meet and reducing trigger sensitivity, thereby reducing unnecessary intervention; when the net effect of the strategy in two consecutive strategy evaluations is non-positive in the valence dimension, and Multiply by respectively This reduces the absolute value of the threshold, bringing it closer to zero, thus making the triggering conditions easier to meet, increasing trigger sensitivity, and strengthening the intervention. To prevent the threshold from diverging after multiple consecutive adjustments, The correction range is constrained to Within, the superscripts min and max represent the lower and upper limits of the corresponding threshold values, respectively; The correction range is constrained to Within the specified range, values ​​outside the range are truncated to the corresponding boundary value. For example, The value is , The value is , The value is , The value is The third preset threshold The safety threshold for switching to human assistance remains fixed and is not subject to dynamic adjustment.

[0097] Through the processing described in equations (8) to (10), step S5 constructs a counterfactual natural evolution baseline using the inertial extrapolation of the emotional trend vector and separates the causal effect of the strategy itself by the difference between the actual observation and the counterfactual baseline. This differs from the feedback method of offline batch optimization after the end of the call in the existing technology, forming an intra-call closed loop of decision-making, execution, evaluation and feedback in the same call process. Equation (10) further introduces a prediction confidence assessment mechanism, which automatically downgrades to a comparison method based on direct observation when the reliability of the inertial extrapolation model is insufficient, thereby improving the robustness of the causal effect assessment to complex emotional dynamics. Thus, the feedback result of step S5 acts on the threshold conditions and strategy selection weights of step S4, enabling this method to continuously and dynamically correct subsequent forward-looking intervention strategies in the same call process.

[0098] Reference manual attached Figure 2 The diagram shows a schematic representation of the adaptive AI telephone response system based on emotion recognition and intelligent decision-making provided in an embodiment of the present invention.

[0099] This invention also provides an adaptive AI telephone answering system 20 based on emotion recognition and intelligent decision-making, comprising: a processor 201 and a memory 202;

[0100] The memory 202 stores programs or instructions that can run on the processor 201. When the program or instructions are executed by the processor 201, they implement the steps of the adaptive AI telephone response method based on emotion recognition and intelligent decision-making described above, and can achieve the same technical effect. To avoid repetition, the present invention will not elaborate further.

[0101] It should be understood that processor 201 can be a central processing unit, or a digital signal processor, application-specific integrated circuit, field-programmable gate array, or other programmable logic device; memory 202 can be volatile memory or non-volatile memory, such as read-only memory, flash memory, or random access memory.

[0102] The above embodiments can be implemented in whole or in part by software, hardware, firmware or a combination thereof; when implemented in software, they can be implemented in the form of a computer program product, and the computer program can be stored in a computer-readable storage medium such as a magnetic medium, optical medium or semiconductor medium.

[0103] The sequence number of each step does not imply the order of execution. The execution order of each step should be determined by its function and internal logic. The units and algorithm steps in the examples described in this article can be implemented by electronic hardware or a combination of computer software and electronic hardware, depending on the specific application and design constraints of the technical solution.

[0104] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. An adaptive AI telephone answering method based on emotion recognition and intelligent decision-making, executed by a telephone answering system, characterized in that: include: Step S1: Obtain the channel compensation matrix according to the voice coding standard type used in the call, and compensate for the distortion introduced by the codec in the voice signal of the call to obtain the compensated voice signal. The compensated speech signal is input into an adversarially trained feature extraction network to extract emotion invariant features that are invariant to the speech coding standard type. Based on the emotion invariant features, frame-level emotion recognition values ​​for arousal and valence dimensions are obtained. Step S2: Within a preset time window from the moment the user first speaks after the call is established, extract the voiceprint feature vector from the compensated speech signal, and obtain the emotion expression baseline vector including the neutral state baseline and the emotion expression amplitude coefficient through a pre-trained voiceprint-emotion baseline mapping network; subtract the neutral state baseline from the frame-level emotion recognition value according to the dimension and divide it by the emotion expression amplitude coefficient to obtain the personalized emotion offset. Step S3: Perform typological detection and temporal modeling on non-vocal events during the call to obtain paralinguistic features. The non-vocal events include meaningful silence, hesitation, pauses, and interruption behaviors. The paralinguistic features are fused with the personalized emotional offset to obtain an enhanced emotional representation, wherein the classification threshold for the duration of meaningful silence is adjusted according to the emotional expression amplitude coefficient. Step S4: Calculate the first and second time derivatives of the enhanced sentiment representation after pre-smoothing to construct a sentiment trend vector; when the sentiment trend vector meets the preset threshold condition, select from the candidate strategy type set according to the strategy selection weight and trigger a prospective intervention strategy; Step S5: After the prospective intervention strategy is executed, within the preset observation window, the inertial extrapolation prediction value based on the emotional trend vector is used as the natural evolution component, and the difference between the change in the personalized emotional offset and the natural evolution component within the observation window is used as the net effect of the strategy. This triggers a policy enhancement, switching, or rollback response, which is then fed back to step S4 to correct the preset threshold conditions and the policy selection weights.

2. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S1 specifically includes: S101: Obtain the voice coding standard type used in the current call; obtain the coding standard type identifier through the session description protocol media description information carried in the session initiation protocol signaling message; if the signaling negotiation information is unavailable, automatically infer the voice coding standard type through the statistical characteristics of the voice bit stream. S102: Based on the speech coding standard type obtained in step S101, select the corresponding channel compensation matrix from the pre-built codec distortion model library, apply the channel compensation matrix to the Mel-spectral feature vector of the speech signal, compensate for the nonlinear quantization distortion introduced by coding compression, and obtain the compensated Mel-spectral feature vector; the compensated Mel-spectral feature vector constitutes the spectral representation of the compensated speech signal. S103: Input the features of the compensated speech signal obtained in step S102 and the Mel spectrum feature vector before compensation into a dual-branch network in parallel, wherein the sentiment branch extracts sentiment representation and the channel branch predicts the speech coding standard type; the dual-branch network is pre-trained against the channel branch through a gradient inversion layer to force the output of the sentiment branch to remain unchanged for the speech coding standard type; during runtime, the sentiment invariant features are extracted by the sentiment branch.

3. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S2 specifically includes: S201: Within the preset time window after the call is established, extract the voiceprint feature vector from the compensated speech signal; the voiceprint feature vector includes the basic pitch mean, pitch dynamic range, average speech rate, average energy, and spectral tilt. S202: Input the voiceprint feature vector extracted in step S201 into the pre-trained voiceprint-emotion baseline mapping network, and output the emotion expression baseline vector; the emotion expression baseline vector includes a neutral state arousal baseline, a neutral state valence baseline, and an emotion expression amplitude coefficient; the emotion expression amplitude coefficient remains unchanged at the initial output value of step S202 during a single call; S203: Subtract the neutral state arousal baseline output in step S202 from the frame-level emotion recognition value in the arousal dimension and subtract the neutral state valence baseline output in step S202 in the valence dimension, and then divide the value by the coefficient of the corresponding dimension in the emotion expression amplitude coefficient to obtain the personalized emotion offset. S204: During the call, the neutral state arousal baseline and the neutral state valence baseline in the emotion expression baseline vector output in step S202 are dynamically corrected using an exponential moving average method; wherein, the frame-level emotion recognition value of the current frame participates in the dynamic correction only when the magnitude of the personalized emotion offset in the current frame is lower than the preset neutral determination threshold.

4. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S3 involves performing typological detection on the non-voice events, specifically including: The meaningful silence is categorized into thinking silence, resistant silence, confused silence, and emotionally suppressed silence based on the semantic type of the response output most recently generated by the telephone response system before the silence. The semantic type of the response corresponding to the thinking silence is a question; the semantic type of the response corresponding to the resistant silence is a request or instruction; the semantic type of the response corresponding to the confused silence is a complex explanation accompanied by a decrease in the user's speech rate before the silence; and the semantic type of the response corresponding to the emotionally suppressed silence is a rejection or negative response accompanied by a sharp drop in the user's base frequency before the silence. The semantic type of the response is directly obtained by the telephone response system based on the intent tags of its most recently generated response text, and the category space of the intent tags at least contains tags corresponding to the four semantic types of the response. For the hesitation and pause, distinguish between pauses within sentences and pauses between sentences, extract the pause duration distribution pattern and the acoustic discontinuity of the preceding and following speech segments; at the same time, detect filler words appearing in the call, and extract the acoustic features of the fundamental frequency, energy and duration of the filler words; For the interruption behavior, the timing, frequency, and energy mutation pattern before and after the user actively interrupts are extracted; when the telephone answering system detects a sudden increase in the user's voice energy exceeding a preset energy threshold and a duration exceeding a preset interruption duration threshold during the voice output period, it is determined to be an active interruption by the user through dual-talk detection, and the dual-talk detection is supplemented by echo cancellation and noise suppression preprocessing.

5. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S3, which involves fusing the paralinguistic features with the personalized sentiment offset to generate the enhanced sentiment representation, specifically includes: Extract time-series feature vectors based on a sliding window for each type of non-voice event. The time-series feature vectors include the event density of each type of non-voice event per unit time, the interval distribution of adjacent non-voice events of the same type, and the difference vector between each non-voice event and the personalized emotional offset of the preceding and following vocal segments. The unit time is the duration of the sliding window. The temporal feature vector and the personalized emotion offset are fused through a multi-head attention mechanism, wherein the attention weight is automatically adjusted according to the signal-to-noise ratio of the current frame: when the signal-to-noise ratio is higher than a preset signal-to-noise ratio threshold, the personalized emotion offset is used as the main source of weight; when the signal-to-noise ratio is lower than the preset signal-to-noise ratio threshold, the weight of the temporal feature vector is increased. When a user is in the silence period of the non-voice event, based on the current silence type determination result, the enhanced sentiment representation of the last frame of the most recent voice segment before silence, and the historical mean vector of the enhanced sentiment representation output in previous frames, the silence period sentiment estimate is output through the feedforward network as the output of the enhanced sentiment representation during the silence period.

6. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S4, which involves pre-smoothing the enhanced sentiment representation and then calculating the first and second time derivatives, specifically includes: S401: A linear regression is performed on the frame-level sequence of the enhanced sentiment representation using a sliding window. The regression slope is used as an estimate of the first-order time derivative, and the rate of change of the regression slope between adjacent windows is used as an estimate of the second-order time derivative. The length of the sliding window is dynamically adjusted according to the number of dialogue rounds in the current call, and the longer the number of dialogue rounds, the larger the length of the sliding window. S402: Construct a Hidden Markov Model, taking the emotional state of each frame as the hidden state, and using the first-order and second-order time derivatives obtained in step S401 together with the enhanced emotional representation as the observation value; use the Viterbi decoding to output the most likely state sequence as the smooth state sequence of the enhanced emotional representation; when the state transition probability corresponding to the hidden state between adjacent frames is lower than a preset transition probability threshold, determine that the frame is a non-smooth transition frame, replace the original observation vector of the non-smooth transition frame with the emission distribution mean vector corresponding to the hidden state of the non-smooth transition frame, and take the component of the enhanced emotional representation in the mean vector to form a new enhanced emotional representation sequence, and use the new enhanced emotional representation sequence as the input to construct the emotional trend vector.

7. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, In step S4, when the sentiment trend vector meets a preset threshold condition, a corresponding level of forward-looking intervention strategy is triggered. The levels of the forward-looking intervention strategy include warning level, intervention level, and emergency level, specifically including: When the component of the first-order time derivative in the valence dimension is lower than the first preset threshold, the early warning-level forward-looking intervention strategy is triggered, and a set of candidate reassurance phrases is preloaded. When the component of the first-order time derivative in the valence dimension is lower than the second preset threshold and the component of the second-order time derivative in the valence dimension is less than zero, the intervention-level prospective intervention strategy is triggered, and a reassurance strategy is actively executed; the second preset threshold is less than the first preset threshold; the triggering result of the intervention-level prospective intervention strategy is executed when the telephone answering system obtains the next answer turn; for the emergency-level prospective intervention strategy, it is executed immediately through a preset polite interruption mechanism; When the valence prediction value obtained by extrapolating a preset number of frames forward based on the sentiment trend vector is lower than the third preset threshold, the emergency-level forward-looking intervention strategy is triggered, and the plan to switch to manual service is initiated. When the level of the prospective intervention strategy is switched, the parameters of the previous and subsequent strategies are weighted by a smooth transition function.

8. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 1, characterized in that, Step S5 specifically includes: S501: After the prospective intervention strategy is implemented, an observation window is defined, and the length of the observation window is dynamically adjusted according to the type of strategy currently being implemented; within the observation window, the enhanced sentiment representation is continuously collected, and the sentiment trend vector before the current strategy is implemented is recorded; S502: Using the sentiment trend vector recorded in step S501, perform inertial extrapolation prediction on the evolution of the personalized sentiment offset within the observation window to obtain the natural evolution component; subtract the natural evolution component from the actual observed change in the personalized sentiment offset within the observation window to obtain the net effect of the strategy. S503: Trigger one of the following three types of responses based on the component value of the net effect of the strategy in the valence dimension obtained in step S502: When the component value is greater than zero, the strategy reinforcement response is executed, and the strategy selection weight of the current strategy type is increased; When the component value is greater than or equal to a negative preset rollback threshold and less than or equal to zero, a strategy switching response is executed, reducing the strategy selection weight of the current strategy type and switching to an alternative strategy type; the preset rollback threshold is a positive value. When the component value is less than the negative preset rollback threshold, a strategy rollback response is executed, the current strategy is terminated, and the strategy state is restored from the strategy history stack to the most recent strategy state where the component value of the net effect of the strategy in the valence dimension is positive. The restored strategy state includes the strategy type and its strategy parameter vector at the corresponding historical moment of the strategy state, and a transitional remedial statement is inserted. The strategy history stack is used to record the strategy type and its corresponding net effect of each round in chronological order. After updating the strategy selection weights, the strategy reinforcement response, the strategy switching response, and the strategy rollback response all perform an iterative truncation-normalization operation. The iterative truncation-normalization operation includes: normalizing the weights of all strategy types so that their sum is one; truncating the strategy types with weights lower than the preset minimum exploration probability to the minimum exploration probability; scaling the weights of the untruncated strategy types proportionally to restore the sum to one; and repeating the truncation and scaling until all weights are not lower than the minimum exploration probability and their sum is one. S504: Feed back the response result of step S503 to step S4 to correct the preset threshold conditions and strategy selection weights of the subsequent prospective intervention strategy.

9. The adaptive AI telephone answering method based on emotion recognition and intelligent decision-making according to claim 8, characterized in that, Step S502, after obtaining the net effect of the strategy and before executing step S503, further includes evaluating the reliability of the net effect of the strategy: The prediction confidence is calculated based on the cumulative deviation between the predicted mean vector of the natural evolution component in each historical observation window and the mean vector of the personalized sentiment offset actually observed at the corresponding time. The influence of the response triggered by the net effect of the strategy on the correction of the strategy selection weight is adjusted according to the numerical adjustment step S503 based on the predicted confidence level. The lower the predicted confidence level, the smaller the influence. When the predicted confidence level is lower than the preset reliability threshold, the difference between the mean of the personalized sentiment shift in the valence dimension within the observation window and the personalized sentiment shift in the strategy execution start frame is used as the substitute value of the net effect of the strategy. When the number of historical observation windows is less than the preset minimum number of windows, the prediction confidence is set to the preset initial confidence value.

10. An adaptive AI telephone answering system based on emotion recognition and intelligent decision-making, characterized in that, include: Processor and memory; The memory stores programs or instructions that can run on the processor, which, when executed by the processor, implement the steps of the adaptive AI telephone response method based on emotion recognition and intelligent decision-making as described in any one of claims 1 to 9.