A method for intelligent auxiliary agent response service in customer service centers of the financial industry
By comprehensively analyzing the signal-to-noise ratio and speech rate performance coefficients of audio frames and adjusting the gain strategy, the problem of low speech recognition accuracy in existing technologies is solved, and more efficient speech recognition and response services are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-11
- Publication Date
- 2026-03-13
AI Technical Summary
In existing technologies, the call center systems in the financial industry rely on gain control based on amplitude for speech recognition, which results in low speech recognition accuracy in complex background noise and variable speech speed environments, affecting response efficiency.
By comprehensively analyzing the signal-to-noise ratio and speech rate performance coefficients of audio frames, the speech gain requirement is determined, and a gain is applied to each audio frame based on a two-dimensional evaluation model. Combined with speech recognition technology and a financial knowledge base, accurate responses are provided.
It significantly improves the clarity of voice signals and the accuracy of semantic understanding in complex environments, thereby enhancing the overall service efficiency of the financial customer service agent system.
Smart Images

Figure CN121309727B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of telephone communication technology, and more specifically to a method for providing intelligent auxiliary agent response services in customer service centers within the financial industry. Background Technology
[0002] Customer service centers in the financial industry widely adopt intelligent agent systems. These systems use speech recognition technology to understand customer audio content and combine it with a financial knowledge base to provide professional responses, thereby achieving efficient customer consultation and problem-solving. In existing technologies, agent systems typically first process customer audio in frames, and then directly apply corresponding gain control based on the amplitude level of each audio frame—suppressing high amplitude and boosting low amplitude—to enhance the overall recognizability of the speech signal. However, in real-world applications, customer audio often contains complex background noise and varies in speaking speed. This gain control method, which relies solely on amplitude, results in low overall speech recognition accuracy, leading to unsatisfactory response efficiency for agent systems. Summary of the Invention
[0003] To address the issue of low overall accuracy in speech recognition using existing methods that rely solely on amplitude gain modulation, this invention aims to provide an intelligent auxiliary agent response service method for customer service centers in the financial industry. The specific technical solution adopted is as follows:
[0004] In a first aspect, the present invention provides a method for intelligent auxiliary agent response service in a customer service center in the financial industry. The method includes: determining the signal-to-noise ratio (SNR) coefficient of a target audio frame among multiple audio frames; the SNR coefficient characterizes the relative proportion of voice signal to noise signal in the target audio frame; the multiple audio frames are audio frames determined after processing the audio frame to be identified by the customer; the target audio frame is any one of the multiple audio frames; determining the speech rate coefficient of the target audio frame; the speech rate coefficient characterizes the density of phoneme information per unit time in the target audio frame; determining the speech gain requirement of the target audio frame based on the SNR coefficient and the speech rate coefficient; the speech gain requirement characterizes the strength of the gain adjustment required for the target audio frame; determining the speech gain requirement corresponding to each audio frame among the multiple audio frames; applying speech gain to each audio frame based on the speech gain requirement corresponding to each audio frame among the multiple audio frames; performing speech recognition on the multiple audio frames after applying the speech gain, and determining the speech recognition result.
[0005] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: determining the noise tendency of the target audio frame based on its short-time energy and zero-crossing rate; short-time energy characterizes the signal strength of the target audio frame; zero-crossing rate characterizes the number of times the target audio frame crosses zero level per unit time, reflecting the frequency of signal fluctuations in the target audio frame; noise tendency characterizes the degree to which the target audio frame is biased towards noise. Based on the audio periodicity characteristics of the target audio frame, determining the audio periodicity of the target audio frame; audio periodicity characteristics characterize the periodicity of the human voice signal in the audio frame; audio periodicity characterizes the degree to which the audio frame conforms to the periodicity of the human voice signal; based on the noise tendency and audio periodicity, determining the signal-to-noise ratio (SNR) coefficient of the target audio frame; noise tendency is inversely proportional to the SNR coefficient, and audio periodicity is directly proportional to the SNR coefficient.
[0006] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: determining the signal period of the target audio frame based on the autocorrelation function of the audio time-domain waveform of the target audio frame; determining multiple comparison time periods relative to the time period in which the target audio frame is located; the comparison time periods are time periods whose distance from the time period in which the target audio frame is located is an integer multiple of the period; determining the audio periodicity of the target audio frame based on the dynamic time warping (DTW) distance between the target audio frame time period and each comparison time period; the audio periodicity is the mean of the reciprocals of the DTW distances between the target audio frame time period and each comparison time period.
[0007] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: determining a formant euphoria index based on the formant interval of the target audio frame; the formant interval is used to characterize the average distance between formants in the spectrum of the target audio frame; the formant euphoria index is used to characterize the degree of speech rate smoothness of the target audio frame; determining a speech rate acceleration factor based on the fundamental frequency signal amplitude and harmonic quantity of the target audio frame; the fundamental frequency signal amplitude is used to characterize the signal strength within the fundamental frequency band of the target audio frame; the harmonic quantity is the number of harmonic signals generated by the fundamental frequency in the target audio frame; the speech rate acceleration factor is used to characterize the degree of speech rate acceleration of the target audio frame; and determining a speech rate performance coefficient based on the formant euphoria index and the speech rate acceleration factor.
[0008] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: determining the average formant interval of the target audio frame and the average formant interval of the historical audio frames; the historical audio frames are audio frames determined after framing the client's historical audio; and determining the formant optimization index based on the difference between the average formant interval of the target audio frame and the average formant interval of the historical audio frames.
[0009] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: normalizing the signal-to-noise ratio (SNR) performance coefficient and speech rate performance coefficient based on a preset normalization function; placing the normalized SNR performance coefficient and speech rate performance coefficient in a two-dimensional coordinate system; the horizontal axis of the two-dimensional coordinate system is the normalized SNR performance coefficient, and the vertical axis is the normalized speech rate performance coefficient; determining the Euclidean distance between the coordinate point corresponding to the target audio frame and a preset reference point; the preset reference point is the coordinate point in the two-dimensional coordinate system where the normalized SNR performance coefficient reaches its maximum value and the normalized speech rate performance coefficient reaches its minimum value; and determining the speech gain requirement based on the difference between the Euclidean distance and the mean of the Euclidean distances of historical audio frames.
[0010] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: determining the average speech gain requirement of historical audio frames; determining the gain adjustment coefficient of each audio frame based on the difference between the speech gain requirement corresponding to each audio frame and the average speech gain requirement of historical audio frames; and applying speech gain to each audio frame based on the gain adjustment coefficient of each audio frame.
[0011] In conjunction with the first aspect mentioned above, in one possible implementation, the method specifically includes: extracting speech features from multiple audio frames after applying speech gain; identifying the speech features using a preset acoustic model to determine phoneme units; correcting the phoneme units using a preset language model to generate candidate text sequences; and determining the final speech recognition result based on a preset decoding algorithm combined with the phoneme units and candidate text sequences.
[0012] In conjunction with the first aspect mentioned above, in one possible implementation, the method further includes: matching the speech recognition result with a financial business knowledge base to determine the response information corresponding to the customer's question.
[0013] In conjunction with the first aspect mentioned above, in one possible implementation, the method further includes: extracting non-silent segments from the customer's audio to be identified based on a preset voice activity detection algorithm; performing frame segmentation on the non-silent segments to determine multiple audio frames.
[0014] The present invention has the following beneficial effects:
[0015] This invention determines the required speech gain by comprehensively analyzing the signal-to-noise ratio (SNR) and speech rate (PR) performance coefficients of each audio frame, achieving refined processing of the audio signal. This method automatically adjusts the gain strategy based on the noise level and user speech rate characteristics in the actual call environment, significantly improving the clarity of the speech signal in complex environments. Compared to traditional single gain control methods, this solution enables the speech recognition engine to acquire higher-quality audio input, improving the semantic understanding accuracy and overall service efficiency of the financial customer service agent system. It thus solves the technical problem of low overall accuracy in speech recognition caused by existing methods that rely solely on amplitude gain control, resulting in unsatisfactory response efficiency in agent systems. Attached Figure Description
[0016] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart illustrating a method for providing intelligent auxiliary agent response service in a customer service center in the financial industry, as provided in an embodiment of the present invention.
[0018] Figure 2 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0019] Figure 3 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0020] Figure 4 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0021] Figure 5 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0022] Figure 6 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0023] Figure 7 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0024] Figure 8 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0025] Figure 9 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided by an embodiment of the present invention.
[0026] Figure 10 This is a flowchart illustrating another intelligent auxiliary agent response service method for customer service centers in the financial industry, provided as an embodiment of the present invention. Detailed Implementation
[0027] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a smart auxiliary agent response service method for customer service centers in the financial industry proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0029] The following description, in conjunction with the accompanying drawings, details a specific solution for an intelligent auxiliary agent response service method for customer service centers in the financial industry provided by this invention.
[0030] Please see Figure 1 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S101-S106, which will be described in detail below.
[0031] S101. Determine the signal-to-noise ratio performance coefficient of the target audio frame among multiple audio frames.
[0032] Among them, the signal-to-noise ratio performance coefficient is used to characterize the relative proportion of human voice signal and noise signal in the target audio frame; multiple audio frames are audio frames determined after processing the audio frames to be identified by the client; the target audio frame is any one of the multiple audio frames.
[0033] In one possible implementation, the frame energy concentration is first determined based on the short-time energy of the target audio frame and the average short-time energy of historical client audio frames. Simultaneously, the relative value of the zero-crossing rate is determined based on the zero-crossing rate of the target audio frame and the maximum / minimum zero-crossing rate of historical client audio frames. Then, the noise tendency is determined based on the frame energy concentration and the relative value of the zero-crossing rate. Simultaneously, the audio periodicity is determined based on the audio periodicity characteristics of the target audio frame. Specifically, the signal period is determined by calculating the autocorrelation function of the audio time-domain waveform, extracting multiple comparison time periods that are integer multiples of the period from the target audio frame, and determining the audio periodicity based on the dynamic time warping distance between the target audio frame and each comparison time period. Finally, the signal-to-noise ratio (SNR) performance coefficient is determined based on the synergistic analysis of the noise tendency and the audio periodicity.
[0034] Among them, short-time energy is used to characterize the signal strength of the target audio frame; zero-crossing rate is used to characterize the number of times the target audio frame crosses the zero level per unit time, reflecting the frequency of signal fluctuations in the target audio frame; and noise tendency is used to characterize the degree to which the target audio frame is biased towards noise.
[0035] Understandably, noise propensity assessment effectively reflects the sensitivity of audio framing to background noise interference, while audio periodicity assessment reflects the level of preservation of the inherent periodicity of the human voice signal. Noise propensity has a negative correlation with the signal-to-noise ratio (SNR) performance coefficient, while audio periodicity has a positive correlation with the SNR performance coefficient. Together, they constitute a complete audio signal quality assessment framework, ensuring accurate judgment of speech signal quality in complex environments.
[0036] S102. Determine the speech rate performance coefficient of the target audio frame.
[0037] Among them, the speech rate performance coefficient is used to characterize the density of phoneme information per unit time in the target audio frame.
[0038] In one possible implementation, a formant dominance index is first determined based on the formant intervals of the target audio frames, while a speech rate acceleration factor is determined based on the fundamental frequency signal amplitude and harmonic quantity of the target audio frames. The formant dominance index is obtained by comparing the average formant intervals of the target audio frames with the average formant intervals of historical customer audio frames, and is used to characterize the smoothness of vocal tract shape changes. The speech rate acceleration factor is obtained through a comprehensive evaluation of the fundamental frequency signal amplitude and harmonic quantity, and is used to reflect the speed characteristics of vocal cord vibration. Finally, a speech rate performance coefficient is determined based on the synergistic analysis of the formant dominance index and the speech rate acceleration factor.
[0039] Understandably, the formant dominance index effectively reflects the rate of change in vocal tract shape during articulation; a higher index value indicates a slower speech rate. Conversely, the speech rate acceleration factor reflects the fundamental frequency characteristics of vocal cord vibration and its harmonic distribution; a higher factor value indicates a faster speech rate. Together, these two factors constitute a complete speech rate assessment framework from the perspectives of vocal tract movement and sound source characteristics, ensuring accurate capture of the speech rate characteristics of different speakers.
[0040] S103. Based on the signal-to-noise ratio performance coefficient and the speech rate performance coefficient, determine the speech gain requirement for the target audio frame.
[0041] Among them, the speech gain requirement is used to characterize the strength of the gain adjustment that needs to be applied to the target audio frame.
[0042] In one possible implementation, the signal-to-noise ratio (SNR) performance coefficient and speech rate performance coefficient are first normalized. The normalized SNR performance coefficient is used as the first evaluation dimension, and the normalized speech rate performance coefficient is used as the second evaluation dimension, constructing a two-dimensional evaluation model. In this model, the Euclidean distance between the coordinate point corresponding to the target audio frame and a preset reference point is determined. The preset reference point is the coordinate point in the two-dimensional evaluation model where the normalized SNR performance coefficient reaches its maximum value and the normalized speech rate performance coefficient reaches its minimum value. Finally, based on the difference between this Euclidean distance and the average Euclidean distance of historical customer audio frames, the speech gain requirement is determined.
[0043] Understandably, the signal-to-noise ratio (SNR) performance coefficient reflects the relative intensity of the effective speech signal and background noise in audio framing, while the speech rate performance coefficient reflects the density of phoneme information in audio framing. In the two-dimensional evaluation model, the Euclidean distance between the coordinate point and the preset reference point characterizes the deviation of the target audio framing state from the ideal signal quality. The smaller the distance value, the closer the audio framing is to the characteristics of a high-quality speech signal, and the smaller the required gain adjustment; conversely, a larger distance value indicates that a greater gain adjustment is needed to optimize the signal quality. This evaluation method based on a two-dimensional distance metric can accurately reflect the gain requirements of audio framing in practical application environments.
[0044] S104. Determine the speech gain requirement for each audio frame in multiple audio frames.
[0045] One possible implementation involves establishing a frame-by-frame queue processing mechanism to process audio frames sequentially according to time. For each audio frame in the queue, its signal-to-noise ratio (SNR) and speech rate (PR) performance parameters are obtained. Based on a two-dimensional evaluation model, the deviation from the ideal signal state is calculated to determine the speech gain requirement for that frame. During processing, the evaluation of each audio frame is independent, but the historical customer audio frame parameters are kept consistent to ensure the consistency of the evaluation criteria.
[0046] Understandably, employing a frame-by-frame independent processing mechanism can fully consider the individual characteristics of each audio frame, avoiding the problem of ignoring local signal characteristics due to overall evaluation. At the same time, maintaining the consistency of historical reference data ensures the comparability of speech gain requirements across different audio frames, establishing a coordinated and unified foundation for subsequent overall gain adjustment. This processing method guarantees both the accuracy of individual audio frame processing and the consistency of the entire audio stream processing.
[0047] S105. Based on the speech gain requirement corresponding to each audio frame in multiple audio frames, apply speech gain to each audio frame.
[0048] In one possible implementation, the average voice gain requirement of historical customer audio frames is first obtained as a benchmark reference value. For each audio frame to be processed, its voice gain requirement is compared with this historical average, and a gain adjustment coefficient is determined based on the degree of difference between the two. Subsequently, a specific voice gain value is determined according to the gain adjustment coefficient, and gain adjustment operation is performed on the corresponding audio frame.
[0049] Understandably, using historical averages as a benchmark ensures the stability of gain adjustment and avoids inaccuracies caused by fluctuations in single-frame data. The gain adjustment coefficient reflects the degree of deviation of the target audio frame from its historical typical state, and this coefficient is positively correlated with the speech gain requirement. When the speech gain requirement is higher than the historical average, it indicates that the audio frame needs a stronger gain boost; conversely, it requires a smaller gain adjustment or maintaining the original level. This differentiated adjustment strategy based on historical reference can effectively optimize signal quality while maintaining consistency in the processing.
[0050] S106. Perform speech recognition on multiple audio frames after applying speech gain, and determine the speech recognition result.
[0051] In one possible implementation, speech feature parameters, including but not limited to time-frequency features such as Mel-frequency cepstral coefficients and filter bank features, are first extracted from audio frames after applying speech gain. Then, the feature vectors are input into a pre-defined acoustic model for phoneme unit recognition to obtain a basic phoneme unit sequence. Next, a pre-defined language model is used to perform contextual error correction on the recognized phoneme units, generating a candidate text sequence that conforms to linguistic rules. Finally, a pre-defined decoding algorithm fuses the acoustic model output and the language model output to determine the final text recognition result.
[0052] The technical solution provided in the above embodiments can bring at least the following beneficial effects: This embodiment determines the required speech gain by comprehensively analyzing the signal-to-noise ratio performance coefficient and speech rate performance coefficient of each audio frame, thus achieving refined processing of the audio signal. This method can automatically adjust the gain strategy according to the noise level and user speech rate characteristics in the actual call environment, significantly improving the clarity of the speech signal in complex environments. Compared with the traditional single gain control method, this solution enables the speech recognition engine to obtain higher quality audio input, improving the semantic understanding accuracy and overall service efficiency of the financial customer service agent system. This solves the technical problem that existing methods relying solely on amplitude gain control have low overall speech recognition accuracy, leading to unsatisfactory response efficiency in the agent system.
[0053] In one possible implementation, please refer to Figure 2 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S201-S203, which will be described in detail below.
[0054] S201. Based on the short-time energy and zero-crossing rate of the target audio frame, determine the noise tendency of the target audio frame.
[0055] In one possible implementation, the short-time energy of the target audio frame and the average short-time energy of all customer audio frames within a preset historical time period are obtained. The energy concentration of the target audio frame is determined based on the difference between the short-time energy of the target audio frame and the average short-time energy of all customer audio frames within the preset historical time period. Then, the zero-crossing rate of the target audio frame and the maximum value of the zero-crossing rate among all customer audio frames within the preset historical time period are obtained. The noise tendency of the target audio frame is determined based on the energy concentration of the target audio frame, the zero-crossing rate of the target audio frame, and the maximum value of the zero-crossing rate among all customer audio frames within the preset historical time period.
[0056] For example, the frame energy concentration of the target audio frame. Satisfy the following formula 1:
[0057] Formula 1
[0058] in, The short-time energy of the audio frame analysis is physically represented as the sum of squares of the signal amplitude, reflecting the strength of the speech signal. The preset historical time period is the short-term energy average of all customer audio frames within a preset historical time period. The physical meaning is the historical average energy level. It can be understood that the preset historical time period of 2 days is determined based on historical experience and can be optimized and adjusted based on actual test conditions. This invention does not limit this. The function is a normalization function. It uses the maximum and minimum value method to normalize the input values to the range [0,1], making the results dimensionless and easy to compare.
[0059] This indicates the difference between the current frame's short-term energy and the historical average energy; the greater the difference, the more abnormal the current frame's energy level is, which may contain noise or special speech features. The physical meaning of is frame energy concentration, which reflects the degree of deviation of the target audio frame energy level from the historical average level. The larger the value, the more abnormal the energy of the current frame, which may require special attention.
[0060] For example, the noise tendency of the target audio frame. The following formula 2 is satisfied:
[0061] Formula 2
[0062] in, The zero-crossing rate is used to divide the target audio frame. Physically, it is the number of times the signal crosses the zero level per unit time, reflecting the frequency of signal fluctuation. This is the maximum zero-crossing rate of all customer audio frames within a preset historical time period, which physically represents the highest historical zero-crossing rate level. The frame energy concentration for dividing the target audio into frames; The physical meaning of this is the audio noise tendency, which reflects the likelihood that a target audio frame will be contaminated by noise. The larger the value, the more likely the frame segment is to contain noise; For parameter tuning coefficients, if If it is 0, then set it to 0.001. If the value is not 0, then set it to 0. Dimensions and Same; Ψ is the parameter tuning coefficient, if If it is 0, then set it to 0.001. If the value is not 0, then set it to 0. Its dimension is the same as Ψ.
[0063] This represents the ratio of the current zero-crossing rate to the historical maximum zero-crossing rate. The closer the value is to 1, the higher the current zero-crossing rate and the greater the tendency for noise. It is the reciprocal of the degree of energy concentration. The smaller (lower energy concentration), the better. The larger the value, the higher the noise tendency.
[0064] S202. Based on the audio periodicity characteristics of the target audio frame, determine the audio periodicity of the target audio frame.
[0065] Among them, audio periodicity features are used to characterize the periodicity of human voice signals in audio frames; audio periodicity representation degree is used to characterize the degree to which audio frames conform to the periodicity of human voice signals.
[0066] In one possible implementation, the signal period corresponding to the target audio frame is determined by analyzing the autocorrelation function of the audio time-domain waveform signal of the target audio frame. The autocorrelation function is used to characterize the periodic repetition characteristics of the signal itself, and the signal period is the interval corresponding to the peak value of the autocorrelation function. Then, based on the signal period, multiple comparison time periods are extracted from the audio time-domain waveform diagram. The distance between the comparison time period and the target audio frame time period is an integer multiple of the signal period. Finally, the dynamic time warping distance between the target audio frame and each comparison time period is calculated, and the average value is obtained based on the reciprocal of these distance values to obtain the audio periodicity.
[0067] S203. Based on the noise tendency and audio periodicity, determine the signal-to-noise ratio performance coefficient of the target audio frame.
[0068] Among them, the noise tendency is inversely proportional to the signal-to-noise ratio performance coefficient, while the audio periodicity is directly proportional to the signal-to-noise ratio performance coefficient.
[0069] One possible implementation involves establishing a comprehensive evaluation mechanism to determine the signal-to-noise ratio (SNR) performance coefficient through the collaborative analysis of noise tendency and audio periodicity. First, an initial SNR performance coefficient is calculated based on noise tendency and audio periodicity, where noise tendency is negatively correlated with the SNR performance coefficient, and audio periodicity is positively correlated. Subsequently, the initial SNR performance coefficient is corrected based on the broadband characteristics of noise. This is achieved by analyzing the signal distribution characteristics of human voice and non-human voice frequency bands in the audio frame spectrum to obtain a noise analysis ratio parameter, which is then used to adjust the initial SNR performance coefficient.
[0070] Understandably, the two-stage evaluation method combining initial calculation and frequency domain correction can more accurately reflect the actual signal-to-noise ratio level of audio frames. The initial calculation stage effectively distinguishes the basic characteristics of human voice from noise through complementary analysis of temporal features; the frequency domain correction stage further eliminates noise interference that overlaps with the human voice frequency band by analyzing the distribution characteristics of the signal in different frequency bands.
[0071] For example, the initial signal-to-noise ratio performance coefficients of the target audio frame. The following formula 3 is satisfied:
[0072] Formula 3
[0073] in, The noise tendency of the target audio frame; The audio periodicity is physically represented by the average of the inverse of the distance between the target audio frame and the integer multiple period of the dynamic time warping (DTW). The larger the value, the more it conforms to the periodicity of human voice. This is a normalization function that maps input values to the range [0,1] using the maximum and minimum value method. For parameter tuning coefficients, if If it is 0, then set it to 0.001. If the value is not 0, then set it to 0. Dimensions and same.
[0074] It combines noise tendency and periodicity characteristics. If the noise tendency is low and the periodicity is high, the product is large, indicating a good signal-to-noise ratio performance coefficient; Normalizing the product, we get .
[0075] For example, the initial signal-to-noise ratio performance coefficients of the target audio frame. Satisfy the following formula 4:
[0076] Formula 4
[0077] in, This is a noise analysis ratio parameter, which physically represents the ratio of the total amplitude of the audio signal in the non-reference analysis frequency band to the total amplitude of the entire frequency band, reflecting the noise content. It is an exponential function with the natural constant e as its base, and its physical meaning is a decay factor. Its range is (0,1]; for .
[0078] The meaning is if The larger (higher noise content), the better. The smaller the value, the greater the noise impact; To correct the signal-to-noise ratio performance coefficient using an attenuation factor. When the noise content is high, reduce... The value is used to obtain the corrected signal-to-noise ratio performance coefficient. ; The physical meaning of is the corrected signal-to-noise ratio performance coefficient, which takes into account the wideband characteristics of noise.
[0079] The technical solution provided in the above embodiments can bring at least the following beneficial effects: This embodiment constructs a two-dimensional signal-to-noise ratio (SNR) evaluation mechanism by combining noise tendency and audio periodicity. Noise tendency identifies noise from the perspective of signal energy and fluctuation characteristics, while audio periodicity verifies the authenticity of speech from the perspective of the periodicity of human voice. The two work together to effectively distinguish between human voice and noise. This mechanism improves the robustness of SNR evaluation, avoids the limitations of single feature analysis, and provides a more accurate basis for subsequent gain adjustment.
[0080] Please see Figure 3 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S301-S303, which will be described in detail below.
[0081] S301. Determine the signal period of the target audio frame based on the autocorrelation function of the audio time-domain waveform diagram of the target audio frame.
[0082] One possible implementation involves calculating the autocorrelation function of the audio time-domain waveform corresponding to the target audio frame and analyzing the peak distribution characteristics of this function to determine the signal period. Specifically, the autocorrelation function characterizes the similarity between the audio signal and its delayed version; its peak position corresponds to the interval between signal repetitions, i.e., the signal period. For periodic speech signals, the autocorrelation function exhibits a clear peak sequence. By identifying the delay time corresponding to the significant peaks, the signal period of the target audio frame can be determined. This process relies on the quasi-periodic characteristics of speech signals, effectively capturing the periodic patterns caused by vocal cord vibration, and providing fundamental parameters for subsequent calculations of audio periodicity.
[0083] S302. Determine multiple comparison time periods with the time period of the target audio frame.
[0084] In one possible implementation, multiple comparison time periods are selected in the audio time-domain waveform based on a determined signal period. Specifically, the selection of comparison time periods follows the principle of integer multiples of the period, that is, the start time difference between each comparison time period and the time period in which the target audio frame is located is an integer multiple of the signal period.
[0085] For example, let the start time of the target audio frame be... If the signal period is T, then the start time of the kth comparison period is... Where k is an integer and k≠0. The duration of the comparison period is consistent with the frame length of the target audio frame to ensure consistency in the period length. By selecting comparison periods at integer multiples of the period, the periodic repetition characteristics of the speech signal can be effectively captured, establishing a suitable reference benchmark for subsequent periodic similarity comparisons. The advantage of this method is that it fully utilizes the quasi-periodic characteristics of the speech signal, making the selection of comparison periods both conform to acoustic laws and have a clear mathematical basis.
[0086] S303. Determine the audio periodicity of the target audio frame based on the dynamic time warping (DTW) distance between the target audio frame time period and each comparison time period.
[0087] The audio periodicity is the mean of the reciprocals of the DTW distances between the target audio frame time period and each comparison time period.
[0088] In one possible implementation, the audio periodicity is determined by calculating the Dynamic Time Warping (DTW) distance between the target audio frame segment and each comparison segment. Specifically, for each comparison segment, its DTW distance with the target audio frame segment is calculated. This distance measures the similarity between the audio signal sequences within the two segments; a smaller distance indicates a more similar signal pattern. Then, the reciprocal of each DTW distance value is taken, such that a larger reciprocal value corresponds to a higher similarity. Finally, the arithmetic mean of the reciprocal DTW distances for all comparison segments is calculated to obtain the audio periodicity of the target audio frame segment. This parameter reflects the average similarity between the target audio frame and signals at different periodic positions. A higher value indicates a more pronounced periodicity in the signal and a purer human voice component.
[0089] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment analyzes the periodic characteristics of the audio signal and evaluates the similarity between different periodic time segments based on the dynamic time warping distance, thereby determining the audio periodicity of the target audio frame, thus realizing the quantification of the periodicity of human voice in the audio frame, effectively capturing the quasi-periodic pattern of human voice, resisting the interference of non-periodic noise, and further improving the accuracy of signal-to-noise ratio evaluation.
[0090] In one possible implementation, please refer to Figure 4 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S401-S403, which will be described in detail below.
[0091] S401. Determine the formant priority index based on the formant interval of the target audio frame.
[0092] Among them, the formant interval is used to characterize the average distance between formants in the spectrum of the target audio frame; the formant dominance index is used to characterize the degree of speech rate relaxation in the target audio frame.
[0093] In one possible implementation, the mean formant interval of the target audio frame is first obtained. This parameter represents the average distance between formants in the spectrum of the target audio frame, reflecting the vocal tract shape change characteristics caused by speech rate. A larger mean formant interval indicates a slower speech rate. Simultaneously, the mean formant interval of customer audio frames within a historical time period is obtained. This parameter serves as a historical reference benchmark, providing contextual comparison of speech rate performance. The difference between the current mean formant interval and the historical mean formant interval is calculated, and this difference is normalized to obtain the formant optimization index.
[0094] Understandably, by standardizing and comparing the current formant interval with historical data, the impact of individual pronunciation habits and differences in recording environment on speech rate assessment can be effectively eliminated, making the formant optimization index more objective and comparable.
[0095] S402. Determine the speech rate acceleration factor based on the amplitude of the fundamental frequency signal and the number of harmonics in the target audio frame.
[0096] In one possible implementation, the amplitude of the fundamental frequency signal of the target audio frame is obtained. This parameter represents the total signal intensity within the fundamental frequency band, reflecting the basic energy level of vocal cord vibration. Simultaneously, the number of harmonics generated by the fundamental frequency within the target audio frame is obtained. This parameter represents the total number of harmonic signals appearing at integer multiples of the fundamental frequency, reflecting the harmonic richness of vocal cord vibration. Based on the product relationship between the fundamental frequency signal amplitude and the number of harmonics, a speech rate acceleration factor is obtained through normalization. This factor comprehensively reflects the intensity characteristics and spectral richness of vocal cord vibration. When the speech rate increases, the vocal cord vibration frequency rises, the fundamental frequency energy increases, and the harmonic components generated per unit time become more concentrated. Therefore, a larger speech rate acceleration factor γ value indicates a faster speech rate.
[0097] For example, the speech rate acceleration factor The following formula 5 is satisfied:
[0098] Formula 5
[0099] in This refers to the total signal amplitude within the fundamental frequency band corresponding to the current audio frame analysis, which physically represents the energy intensity of the fundamental frequency signal. This represents the number of harmonic signals generated corresponding to the fundamental frequency within the current frame, which physically means the number of harmonics. It is a normalization function that maps the input value to the range [0,1] by normalizing the maximum and minimum values.
[0100] It is the product of the fundamental frequency amplitude and the number of harmonics. If the fundamental frequency amplitude is large and the number of harmonics is large, it indicates that the speech rate may be fast and the sound energy is high. Normalizing the product yields the speech rate acceleration factor. The larger the product, the better. The larger the value; It reflects the speech rate and high-frequency energy of the target audio frames. The higher the value, the faster the speech rate and the more phonemes are accumulated.
[0101] Understandably, although and Different dimensions, but product Dimensions and Same, after normalization Since it is dimensionless, parameters with different dimensions are normalized to become dimensionless, which does not affect the calculation of the speech rate acceleration factor based on the above formula.
[0102] S403. Based on the formant preference index and the speech rate acceleration factor, determine the speech rate performance coefficient.
[0103] In one possible implementation, a speech rate performance coefficient is calculated through synergistic analysis of the formant euphoria index and the speech rate acceleration factor. Specifically, the formant euphoria index characterizes the degree of speech rate slowness in audio framing, with a larger value indicating a slower speech rate; the speech rate acceleration factor characterizes the degree of speech rate acceleration in audio framing, with a larger value indicating a faster speech rate. Based on the inverse relationship between the formant euphoria index and the speech rate performance coefficient, and the direct relationship between the speech rate acceleration factor and the speech rate performance coefficient, the reciprocal of the formant euphoria index is combined with the speech rate acceleration factor, and the speech rate performance coefficient is obtained through normalization. This coefficient comprehensively reflects the speech rate characteristics of audio framing; a larger value indicates a faster speech rate and denser phoneme information.
[0104] For example, speech rate performance coefficient Satisfy the following formula 6:
[0105] Formula 6
[0106] in, The index for prioritizing resonance peaks; As a factor that accelerates speech rate; For parameter tuning coefficients, if If it is 0, then set it to 0.001. If the value is not 0, then set it to 0. Dimensions and Same. If The smaller The larger the value, the faster the speaking speed; Combining formants and speech rate factors. If... large and A larger product indicates a higher speech rate coefficient. The speech rate performance coefficient reflects the speed and high-frequency characteristics of the target audio frames. The larger the value, the faster the speech rate, the more phonemes are accumulated, and the lower the required gain. This is used for collaborative signal-to-noise ratio analysis to determine gain requirements.
[0107] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment achieves a comprehensive evaluation of speech rate performance through the comprehensive calculation of formant dominance index and speech rate acceleration factor. Formant interval reflects the changes in vocal tract caused by speech rate, and fundamental frequency harmonic characteristics reflect the changes in vocal cord vibration frequency. By characterizing speech rate features from different dimensions, the speech rate evaluation results are more consistent with the physical characteristics of actual speech, providing reliable speech rate parameters for gain control.
[0108] Please see Figure 5 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S501-S502, which will be described in detail below.
[0109] S501. Determine the average formant interval of the target audio frame and the average formant interval of the historical audio frames.
[0110] In one possible implementation, the spectrogram corresponding to the target audio frame is obtained, the formant positions are identified, the interval between adjacent formants is calculated, and the mean formant interval of the target audio frame is determined. This parameter reflects the vocal tract shape characteristics of the target audio frame and is directly related to the speech rate. The formant interval data of all customer audio frames within a preset historical time period are extracted from the historical voice data stored in the agent system, and the historical mean formant interval is obtained by averaging these historical data.
[0111] S502. Based on the difference between the average formant interval of the target audio frame and the average formant interval of the historical audio frames, determine the formant optimization index.
[0112] One possible implementation involves calculating the relative difference between the mean formant interval of the target audio frame and the historical mean formant interval, and then standardizing this difference to obtain the formant optimization index. Specifically, firstly, the mean formant interval of the target audio frame and the mean formant interval of the client audio frames within the historical time period are obtained, and the difference between the two is calculated. Subsequently, this difference is normalized to obtain the final formant optimization index. This index quantifies the speech rate performance of the target audio frame relative to the historical average level. The larger the value, the greater the formant interval of the target audio frame relative to the historical level, i.e., the slower the speech rate and the better the formant feature performance.
[0113] For example, the formant dominance index The following formula 7 is satisfied:
[0114] Formula 7
[0115] in, The average interval of each formant in the spectrum corresponding to the frame of the target audio is the physical meaning of the average distance between formants, which reflects the speed of speech. The average interval of the formant peaks in the audio spectrum of customer responses over a two-day historical period of the agent system is used. Physically, it represents the historical average interval of the formant peaks. It is understood that the two-day historical period is determined based on historical experience and can be optimized and adjusted based on actual test results. This invention does not limit this. It is a normalization function that maps the input value to the range [0,1] by normalizing the maximum and minimum values.
[0116] This indicates the degree of deviation between the current resonant interval and the historical average interval. If > This indicates that the formant interval is large, and the speech rate may be slow. Normalizing the differences yields the resonance peak priority index. The greater the difference The larger the value; The physical meaning of this is the formant preference index, which reflects the formant spacing of the target audio frame relative to historical levels. The larger the value, the larger the formant interval, the slower the speech rate, and the better the formant performance.
[0117] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment achieves a relative evaluation of speech rate characteristics by comparing the formant interval of the target audio frame with historical data. This method can adapt to the pronunciation habits of different users, eliminate the influence of individual differences on speech rate evaluation, and make speech rate judgment more universal and comparable.
[0118] Please see Figure 6 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S601-S604, which will be described in detail below.
[0119] S601. Normalize the signal-to-noise ratio performance coefficient and speech rate performance coefficient based on a preset normalization function.
[0120] One possible implementation involves using a pre-defined normalization function to standardize the signal-to-noise ratio (SNR) performance coefficients and speech rate performance coefficients. The purpose of normalization is to convert parameters with different dimensions and numerical ranges to a unified numerical interval, eliminating differences in magnitude between parameters and ensuring the fairness and comparability of subsequent analyses. Specifically, the normalization function is... The function, through a maximum and minimum value normalization method, linearly maps the original parameter values to a closed interval [0,1]. For the signal-to-noise ratio (SNR) performance coefficient, its minimum and maximum value boundaries are determined based on historical data; for the speech rate performance coefficient, its numerical range boundaries are also determined based on historical speech rate data. After normalization, both the SNR and speech rate performance coefficients are converted into dimensionless standardized values, preserving the relative magnitudes of the original parameters while providing a unified dimensional standard.
[0121] S602. Place the normalized signal-to-noise ratio performance coefficient and speech rate performance coefficient in a two-dimensional coordinate system.
[0122] In the two-dimensional coordinate system, the horizontal axis represents the normalized signal-to-noise ratio performance coefficient, and the vertical axis represents the normalized speech rate performance coefficient.
[0123] One possible implementation involves constructing a two-dimensional Cartesian coordinate system with the normalized signal-to-noise ratio (SNR) coefficient as the horizontal axis and the normalized speech rate coefficient as the vertical axis. Each audio frame is mapped to a unique coordinate point in this system based on its normalized values of these two parameters. The horizontal axis represents the SNR level of the audio frame, with higher values indicating better SNR performance; the vertical axis represents the speech rate of the audio frame, with higher values indicating faster speech rate and denser phonemes. Through this mapping method, the acoustic features of each audio frame are transformed into geometric positional relationships in two-dimensional space, providing an intuitive spatial analysis basis for subsequent calculations of speech gain requirements.
[0124] S603. Determine the coordinates of the target audio frame and the Euclidean distance between them and the preset reference point.
[0125] In one possible implementation, a preset reference point with specific acoustic significance is set based on the established two-dimensional coordinate system. The coordinates of this preset reference point are composed of the maximum value of the normalized signal-to-noise ratio (SNR) performance coefficient and the minimum value of the normalized speech rate (PR) performance coefficient, corresponding to the ideal high SNR and low speech rate states. By calculating the Euclidean distance between the coordinate point corresponding to the target audio frame and this preset reference point, a quantitative index characterizing the deviation of the acoustic characteristics of the target audio frame from the ideal state is obtained. The Euclidean distance is calculated based on the formula for the straight-line distance between two points in a two-dimensional plane; the larger the distance value, the farther the acoustic characteristics of the target audio frame deviate from the ideal state.
[0126] Understandably, by using Euclidean distance as a metric, the complex problem of acoustic feature evaluation is transformed into an intuitive problem of spatial distance calculation. This method can simultaneously take into account the characteristics of both signal-to-noise ratio and speech rate, and comprehensively reflect the overall acoustic quality of audio frames through a single numerical value.
[0127] 604. Determine the speech gain requirement based on the difference between the Euclidean distance and the mean Euclidean distance of historical audio frames.
[0128] One possible implementation involves obtaining the average Euclidean distance of customer audio frames over a historical time period. This parameter reflects the average deviation of historical speech data from the ideal acoustic state. The speech gain requirement is obtained by calculating the relative difference between the Euclidean distance of the target audio frame and the historical average Euclidean distance, and then normalizing this difference. Specifically, when the Euclidean distance of the target audio frame is less than the historical average, it indicates that the acoustic characteristics of the target audio frame are closer to the ideal state, requiring relatively low gain control; conversely, when the Euclidean distance is greater than the historical average, it indicates that the target audio frame deviates significantly from the ideal state, requiring relatively high gain control. The speech gain requirement is converted into a standardized numerical representation through normalization, providing a quantitative basis for subsequent gain control.
[0129] For example, voice gain requirement The following formula 8 is satisfied:
[0130] Formula 8
[0131] in, This represents the Euclidean distance between the current target audio frame and the high SNR and low-energy corner point. The high SNR and low-energy corner point corresponds to the point where the normalized value of the SNR performance coefficient is 1 and the normalized value of the speech rate performance coefficient is 0 (because a large speech rate performance coefficient indicates high energy). Physically, this represents the distance between the current frame and the ideal gain point. The distance between customer audio frames and high signal-to-noise ratio, low-energy corner points over the past two days is the average Euclidean distance, which physically represents the historical average distance. It is understood that the two-day historical time period is determined based on historical experience and can be optimized and adjusted based on actual test results. This invention does not limit this. It is a normalization function that maps the input value to the range [0,1] by normalizing the maximum and minimum values.
[0132] This represents the difference between the historical average distance and the current target frame distance. If... < This indicates that the current target frame is closer to a high signal-to-noise ratio, low-energy corner point. If Z is positive, and the larger the difference, the greater the gain requirement of the current frame compared to the historical average level; if Z = Z′: the gain requirement of the current target frame is consistent with the historical average level, the difference is 0, and the gain requirement is at a medium level; if Z > Z′: the current target frame deviates from the optimal gain requirement scenario (such as low signal-to-noise ratio or fast speech speed, requiring low gain), then Z′−Z is negative, and the larger the absolute value of the difference, the weaker the gain requirement.
[0133] and The numerical values are positively correlated. The larger the difference (the closer the current target frame is to the gain scene). The closer it is to 1, the stronger the gain requirement. and It is negatively correlated. The smaller the current frame segment, the better (the closer the current target frame segment is to the gain scene). The closer it is to 1, the stronger the gain requirement. and It is positively correlated. The larger the historical frame size, the further it deviates from the optimal scene. The closer it is to 1, the stronger the gain requirement.
[0134] This indicates that the difference is normalized to obtain the speech gain requirement. The greater the difference (the closer the current frame is to the ideal point), the better. The larger the value; The physical meaning of is the speech gain requirement, which reflects the degree to which the current audio frame needs gain. The larger the value, the higher the gain needs to be applied.
[0135] The technical solution provided in the above embodiments can bring at least the following beneficial effects: This embodiment, by constructing a two-dimensional evaluation model of signal-to-noise ratio and speech rate, transforms the calculation of speech gain requirement into a distance metric from an ideal point in two-dimensional space. This method transforms complex speech feature relationships into intuitive spatial relationships, preserving the synergistic influence mechanism of signal-to-noise ratio and speech rate while simplifying the calculation logic, making the determination of gain requirement more scientific and reasonable.
[0136] Please see Figure 7 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S701-S703, which will be described in detail below.
[0137] S701. Determine the average voice gain requirement of historical audio frames.
[0138] In one possible implementation, voice gain requirement data for all customer audio frames within a preset time period is extracted from historical voice processing records stored in the agent system. These historical voice gain requirement data are then validated and outlier filtered to retain valid data samples that meet preset quality requirements. Subsequently, the arithmetic mean of the filtered valid historical voice gain requirement data is calculated to obtain the historical voice gain requirement mean. This mean parameter reflects the average level of gain requirement during historical voice data processing and serves as a benchmark reference value for evaluating the current audio frame gain requirement.
[0139] It is understood that the preset time period can be any historical time period, and can be selected based on the actual experimental results; this disclosure does not limit this.
[0140] S702. Based on the difference between the corresponding voice gain requirement of each audio frame and the average historical voice gain requirement, determine the gain control coefficient of each audio frame.
[0141] In one possible implementation, a gain adjustment coefficient is obtained by calculating the relative difference between the current audio frame's speech gain requirement and the historical average speech gain requirement, and then standardizing this difference. Specifically, the speech gain requirement of the current audio frame and the historical average speech gain requirement are first obtained, and the difference between them is calculated. Then, this difference is normalized to convert it into a standardized coefficient within a preset range. The gain adjustment coefficient reflects the strength of the gain requirement of the current audio frame relative to the historical average. When the speech gain requirement is higher than the historical average, the gain adjustment coefficient is positive, indicating that the gain needs to be increased; when the speech gain requirement is lower than the historical average, the gain adjustment coefficient is negative, indicating that the gain needs to be reduced.
[0142] For example, the gain control coefficient Satisfy the following formula 9:
[0143] Formula 9
[0144] in, Real-time analysis of voice gain requirements for current customers in voice frame segmentation; The average voice gain requirement of each customer response voice frame in history when providing financial agent response service, which physically means the historical average demand. The rated voice gain is set to 10, which physically means the maximum gain value. It is understood that the value of conventional voice gain is generally selected between 1 and 10. The rated voice gain of 10 in this invention is only an example and can be adjusted based on the actual effect. This invention does not limit this. It is a normalization function that maps the input value to the range [0,1] by normalizing the maximum and minimum values.
[0145] This represents the difference between the current gain demand and the historical average demand. If > This indicates that if the current demand is high, the gain should be large. The difference is normalized to obtain the relative gain adjustment factor, with a value range of [0, 1]. Multiply the adjustment factor by the rated gain to obtain the gain control coefficient. The larger the adjustment factor, The larger the value.
[0146] S703: Apply speech gain to each audio frame based on the gain control coefficient of each audio frame.
[0147] In one possible implementation, the gain control coefficient of each audio frame is obtained, and the gain of the corresponding audio frame is adjusted according to the speech gain control coefficient. When the gain value is greater than the system reference gain, the signal strength is increased; when the gain value is less than the system reference gain, the signal strength is suppressed. The system reference gain is a reference value determined based on historical experience, and the reference value can also be adjusted based on the actual gain situation. This disclosure does not limit this. The gain adjustment process adopts digital signal processing technology and is achieved through time-domain signal amplitude scaling or frequency-domain equalization to ensure that the adjusted audio signal maintains the original spectral characteristics and speech quality.
[0148] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment achieves gain control based on statistical laws by comparing the current gain demand with historical levels. This method can adapt to changes in different audio environments, avoid drastic fluctuations in gain values, ensure the stability and continuity of gain adjustment, and at the same time ensure the rationality of the control range by setting a preset gain rating.
[0149] Please see Figure 8 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S801-S804, which will be described in detail below.
[0150] S801. Extract speech features from multiple audio frames after applying speech gain.
[0151] In one possible implementation, speech features are extracted from the audio frames after speech gain adjustment. Specifically, the extracted speech features include, but are not limited to, Mel-frequency cepstral coefficients and filter bank features. Mel-frequency cepstral coefficients are obtained by simulating human hearing characteristics, converting the spectrum to a Mel-scale, and then performing cepstral analysis to obtain characteristic parameters characterizing the short-time power spectrum of the speech signal. Filter bank features are obtained by setting a set of overlapping triangular bandpass filters on the Mel-scale to compress and extract the spectral features of the speech signal, determining the speech features of multiple audio frames after applying speech gain.
[0152] S802. The speech features are identified by a preset acoustic model to determine the phoneme unit.
[0153] In one possible implementation, the extracted speech feature vector is input into a pre-defined acoustic model. This model models the probability distribution of the input features and outputs the corresponding phoneme unit sequence. The pre-defined acoustic model is constructed based on statistical learning methods, establishing a mapping relationship between speech features and phoneme units through parameter training on a large amount of speech training data. During the recognition process, the acoustic model calculates the state probability of the input speech feature vector. Through joint analysis of state transition probabilities and observation probabilities, it determines the most likely state sequence and then maps this state sequence to the corresponding phoneme units. Phoneme units include, but are not limited to, basic speech units such as vowels and consonants, which constitute the basic building blocks of speech recognition.
[0154] For example, the preset acoustic model adopts a combined architecture of Hidden Markov Model (HMM) and Gaussian Mixture Model (GMM). This HMM-based acoustic modeling method effectively handles the temporal variability of speech signals and the randomness of acoustic features, providing accurate phoneme-level recognition results for subsequent language model processing. The model automatically acquires acoustic knowledge from data through statistical learning, avoiding complex manual rule design and exhibiting good robustness and adaptability.
[0155] S803. The phoneme unit is modified by a preset language model to generate a candidate text sequence.
[0156] In one possible implementation, the pre-defined language model is constructed based on statistical language modeling methods. It learns probabilistic relationships between words by training on a large-scale text corpus. The model receives recognition results composed of phoneme units and performs contextual relevance analysis and grammatical correctness verification on the phoneme unit sequences output by the acoustic model by calculating the joint probability of different word sequences.
[0157] For example, the preset language model is an N-gram language model. This model is trained on professional texts such as financial business documents, customer service dialogue records, and product descriptions to establish statistical characteristics for financial terminology and business expressions. The model calculates the joint probability distribution of word sequences using N-gram statistical methods and uses the Viterbi decoding algorithm to find the optimal word sequence.
[0158] S804. Based on the preset decoding algorithm combined with the phoneme unit and the candidate text sequence, determine the final speech recognition result.
[0159] In one possible implementation, the preset decoding algorithm employs dynamic programming principles. By comprehensively considering the phoneme unit probabilities output by the acoustic model and the word sequence probabilities output by the language model, it seeks the globally optimal text output result. This decoding algorithm establishes a search space, where the acoustic model provides the low-level phoneme recognition confidence and the language model provides the high-level language structure constraints. Through a comprehensive scoring mechanism that balances acoustic evidence and language rules, it selects the sequence with the highest score among all possible text sequences as the final recognition result. Specifically, the decoding process maintains and expands candidate paths, calculates the cumulative score of each path, which is a weighted sum of the acoustic score and the language weight, and finally retains the text corresponding to the optimal path as the speech recognition result.
[0160] For example, the preset decoding algorithm is the Viterbi algorithm. This algorithm uses dynamic programming to find the path with the highest weighted sum of the acoustic model score and the language model score in a search space consisting of phoneme units and candidate text sequences. The text sequence corresponding to this path is then determined as the final speech recognition result. During the decoding process, a pruning strategy is used to control the size of the search space, ensuring a balance between decoding efficiency and accuracy.
[0161] The technical solution provided in the above embodiments can bring at least the following beneficial effects: This embodiment constructs a complete speech recognition process through the collaborative work of an acoustic model and a language model. The acoustic model is responsible for low-level phoneme recognition, and the language model is responsible for high-level semantic correction. The two are optimized and integrated through decoding algorithms to ensure the accuracy and reliability of the conversion from audio features to final text.
[0162] Please see Figure 9 The present invention illustrates a flowchart of a smart auxiliary agent response service method for a customer service center in the financial industry, provided by an embodiment of the present invention. The method includes the following steps S901, which will be described in detail below.
[0163] S901. Match the speech recognition results with the financial business knowledge base to determine the answer information corresponding to the customer's question.
[0164] One possible implementation involves first semantically parsing the text sequence obtained from speech recognition to extract key business elements and user intent. Key business elements include, but are not limited to, structured data such as business type, product name, operation instructions, monetary value, and time information; user intent includes operation types such as query, processing, consultation, and complaint. Subsequently, based on the parsed business elements and user intent, a multi-dimensional matching retrieval is performed in a financial business knowledge base. This knowledge base comprises multiple sub-knowledge bases, including a product information base, a business process base, a frequently asked questions base, and a policy and regulation base. A unified retrieval interface and matching engine are established to enable collaborative querying of disparate knowledge. The matching process employs a semantic similarity-based retrieval algorithm, comprehensively considering multiple factors such as word form matching, semantic association, and contextual relevance, ultimately outputting the most relevant response to the customer's question.
[0165] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment realizes a complete closed loop from voice recognition to business services by intelligently matching the voice recognition results with the financial knowledge base, ensuring the professionalism and accuracy of the information fed back by the agent system.
[0166] Please see Figure 10 The diagram illustrates a flowchart of an intelligent auxiliary agent response service method for a customer service center in the financial industry, according to an embodiment of the present invention. The method includes the following steps S1001-S1002, which will be described in detail below.
[0167] S1001. Extract the non-silent segments from the customer's audio to be identified based on a preset voice activity detection algorithm.
[0168] In one possible implementation, a pre-defined speech activity detection algorithm distinguishes between speech segments and silence segments by analyzing the short-time energy characteristics and zero-crossing rate characteristics of the audio signal. First, the input audio to be identified is segmented into frames. For each audio frame, its short-time energy and zero-crossing rate are calculated. The short-time energy reflects the cumulative sum of the squares of the signal amplitude, and the zero-crossing rate reflects the frequency with which the signal crosses zero levels. Based on a comprehensive assessment of the short-time energy and zero-crossing rate, it is determined whether the frame is a speech frame. Finally, segments consecutively identified as speech frames are merged to form complete non-silence segments, providing valid input for subsequent audio framing processing.
[0169] S1002. Perform frame segmentation on the non-silent segments to determine multiple audio frames.
[0170] In one possible implementation, the non-silent segments extracted by the speech activity detection algorithm are segmented into frames. Continuous non-silent segments are divided into a series of short-time audio frames. The speech signal within each audio frame can be considered a quasi-stationary signal, satisfying the short-time stationarity assumption. During the framing process, a smoothing window function is used to window each frame to reduce spectral leakage. The resulting series of audio frames serves as the basic processing unit for subsequent speech feature analysis and gain adjustment.
[0171] The technical solution provided by the above embodiments can bring at least the following beneficial effects: This embodiment provides a high-quality audio data foundation for subsequent analysis through voice activity detection and frame preprocessing. The preprocessing stage provides necessary guarantees for the reliable operation of the entire system and is an important foundation for achieving high-quality voice processing.
[0172] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. The processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0173] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments.
Claims
1. A method for intelligent assisted agent response service in a financial industry customer service center, characterized in that, The method comprises: determining a signal-to-noise ratio performance coefficient of a target audio frame in a plurality of audio frames; the signal-to-noise ratio performance coefficient is used to represent the relative proportion of human voice signals and noise signals in the target audio frame; the plurality of audio frames are audio frames determined after processing of a client audio frame to be identified; the target audio frame is any one of the plurality of audio frames; determining a speech rate performance coefficient of the target audio frame; the speech rate performance coefficient is used to represent the density of phoneme information in unit time in the target audio frame; normalizing the signal-to-noise ratio performance coefficient and the speech rate performance coefficient based on a preset normalization function; placing the normalized signal-to-noise ratio performance coefficient and the speech rate performance coefficient in a two-dimensional coordinate system; the horizontal axis of the two-dimensional coordinate system is the normalized signal-to-noise ratio performance coefficient, and the vertical axis is the normalized speech rate performance coefficient; determining the Euclidean distance between the coordinate point corresponding to the target audio frame and a preset reference point; the preset reference point is a coordinate point in the two-dimensional coordinate system where the normalized signal-to-noise ratio performance coefficient takes the maximum value and the normalized speech rate performance coefficient takes the minimum value; determining a speech gain requirement degree based on the difference between the Euclidean distance and the average of the Euclidean distances of historical audio frames; the speech gain requirement degree is used to represent the strength of the gain adjustment required by the target audio frame; determining the speech gain requirement degree corresponding to each audio frame in the plurality of audio frames; applying speech gain to each audio frame based on the speech gain requirement degree corresponding to each audio frame in the plurality of audio frames; performing speech recognition on the plurality of audio frames after applying speech gain to determine a speech recognition result.
2. The intelligent assisted agent response service method for financial industry's customer service center as claimed in claim 1, wherein, The method comprises: determining an audio noise tendency degree of the target audio frame based on the short-time energy and the zero-crossing rate of the target audio frame; the short-time energy is used to represent the signal strength of the target audio frame; the zero-crossing rate is used to represent the number of times the target audio frame crosses the zero level in unit time, reflecting the frequency of signal fluctuation of the target audio frame; the audio noise tendency degree is used to represent the degree to which the target audio frame is biased towards noise; determining an audio periodicity presentation degree of the target audio frame based on the audio periodicity feature of the target audio frame; the audio periodicity feature is used to represent the periodicity of human voice signals in the audio frame; the audio periodicity presentation degree is used to represent the degree to which the audio frame conforms to the periodicity of human voice signals; determining the signal-to-noise ratio performance coefficient of the target audio frame based on the audio noise tendency degree and the audio periodicity presentation degree; the audio noise tendency degree is inversely proportional to the signal-to-noise ratio performance coefficient, and the audio periodicity presentation degree is proportional to the signal-to-noise ratio performance coefficient.
3. The intelligent assisted agent response service method for financial industry's customer service center as claimed in claim 2, wherein, The method comprises: determining the signal period of the target audio frame based on the autocorrelation function of the audio time-domain waveform graph of the target audio frame; determine a plurality of comparison time periods with the time period in which the target audio frame is located; the comparison time period is a time period with an integer multiple of the period from the time period in which the target audio frame is located; determine the audio period performance degree of the target audio frame based on the dynamic time warping (DTW) distance between the target audio frame and each comparison time period.
4. The intelligent assisted agent response service method for financial industry's customer service center of claim 1, wherein, The method further comprises: determine a resonance peak optimization indicator based on the formant interval of the target audio frame; the formant interval is used to represent the average distance between the formant in the frequency spectrum of the target audio frame; the resonance peak optimization indicator is used to represent the degree of speed relaxation of the target audio frame; determine a speed-up factor based on the fundamental signal amplitude and the number of harmonics of the target audio frame; the fundamental signal amplitude is used to represent the signal strength in the fundamental frequency band of the target audio frame; the number of harmonics is the number of harmonic signals generated by the fundamental frequency in the target audio frame; the speed-up factor is used to represent the degree of speed-up of the target audio frame; determine the speed performance coefficient based on the resonance peak optimization indicator and the speed-up factor.
5. The intelligent assisted agent response service method for financial industry's customer service center as claimed in claim 4, wherein, The method further comprises: determine the formant interval mean value of the target audio frame and the formant interval mean value of the historical audio frame; the historical audio frame is an audio frame determined by performing frame processing on the historical audio of the customer; determine the resonance peak optimization indicator based on the difference between the formant interval mean value of the target audio frame and the formant interval mean value of the historical audio frame.
6. The intelligent assisted agent response service method for financial industry's customer service center of claim 1, wherein, The method further comprises: determine the historical audio frame gain requirement mean value; determine the gain control coefficient of each audio frame based on the difference between the speech gain requirement degree corresponding to each audio frame and the historical speech gain requirement mean value; apply speech gain to each audio frame based on the gain control coefficient of each audio frame.
7. The intelligent assisted agent response service method in a financial industry's customer service center as claimed in claim 1, wherein, The method further comprises: extract the speech features of the plurality of audio frames to which speech gain is applied; determine the phoneme unit by recognizing the speech features through a preset acoustic model; generate a candidate text sequence by correcting the phoneme unit through a preset language model; determine the final speech recognition result based on a preset decoding algorithm in combination with the phoneme unit and the candidate text sequence.
8. The intelligent assisted agent response service method of a customer service center of a financial industry according to claim 7, characterized by, The method further comprises: match the speech recognition result with the financial business knowledge base to determine the reply information corresponding to the customer's question.
9. The intelligent assisted agent response service method in a financial industry's customer service center of claim 1, wherein, The method further comprises: extract the non-silence segment in the audio to be recognized by the customer based on a preset voice activity detection algorithm; perform frame processing on the non-silence segment to determine the plurality of audio frames.
Citation Information
Patent Citations
Voice interactive chart dynamic generation method and system based on large language model
CN120496560A
Target voice regulation and control method, device and equipment based on intelligent glasses
CN120544552A