An intelligent telephone customer service interaction method for a power enterprise based on artificial intelligence
By dynamically adjusting the quantization step size and sensitivity weight of the A-law quantizer, the bandwidth optimization problem of traditional A-law quantization in high-concurrency scenarios is solved, achieving efficient voice coding and compatibility, and improving the real-time performance and service quality of intelligent telephone customer service for power companies.
Patent Information
- Application Number
- CN202511280070.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-09
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-09-09
AI Technical Summary
Traditional A-law quantification methods are difficult to meet bandwidth optimization requirements in high-concurrency customer service scenarios, leading to a decline in system stability and service quality, especially in environments with limited network resources.
By collecting voice frames during real-time telephone customer service interactions, the masking characteristics of the voice frames are dynamically analyzed using a psychoacoustic model. The quantization step size and sensitivity weight of the A-law thirteen-segment quantizer are adjusted to achieve adaptive encoding. Quantization in high-energy regions is appropriately relaxed to improve compression efficiency, while fine quantization is maintained in sensitive regions to protect speech clarity.
It effectively avoids wasting bit resources and auditory distortion, improves coding efficiency and subjective auditory quality, maintains compatibility with existing PSTN and IP communication systems, and enhances the real-time performance and stability of intelligent telephone customer service voice interaction for power companies.
Smart Images

Figure CN120783769B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of voice signal quantization technology. More specifically, this invention relates to an artificial intelligence-based intelligent telephone customer service interaction method for power companies. Background Technology
[0002] In modern power companies' intelligent telephone customer service systems, voice interaction is the core channel for user service. A large amount of voice call data needs to be transmitted in real time through the Public Switched Telephone Network (PSTN) or IP network, placing high demands on the real-time performance, compatibility, and bandwidth efficiency of voice coding. To achieve effective compression and high-quality restoration of voice signals, A-law or μ-law pulse code modulation (PCM) coding standards are widely adopted. Among them, A-law thirteen-segment quantization, as the core technology of the ITU-T G.711 protocol, is widely used in traditional telephone systems and embedded voice devices due to its simple algorithm, low latency, and strong decoding compatibility.
[0003] The A-law 13-segment quantization method is a non-uniform quantization approach that divides the amplitude range of the input signal into 13 linear segments (8 segments on each of the positive and negative half-axis, with the zero segment shared). This allows for fine quantization of small signals and coarse compression of large signals, thereby expanding the dynamic range and improving overall auditory quality within an 8-bit PCM bitstream. This method has a solid foundation for deployment in the existing customer service communication architecture of power companies, supports seamless integration with traditional PSTN equipment, and is suitable for high-reliability, low-latency voice interaction scenarios.
[0004] However, traditional A-law quantization uses a fixed segmentation structure and quantization step size, lacks perceptual analysis of speech content, and cannot achieve dynamic bit allocation, resulting in a high average bit rate. This makes it difficult to meet the bandwidth optimization requirements of high-concurrency customer service scenarios, especially in remote areas or environments with limited network resources, affecting system stability and service quality. Summary of the Invention
[0005] To address the technical problem that traditional A-law quantization, with its fixed segmentation structure and quantization step size, struggles to meet bandwidth optimization requirements in high-concurrency customer service scenarios, this invention provides an AI-based intelligent telephone customer service interaction method for power companies, comprising:
[0006] The system collects voice frames during real-time telephone customer service interactions and converts them to the frequency domain. It acquires the main masking tone of each voice frame, determines the total masking threshold of each frequency component of the voice frame based on the main masking tone, and determines the overall perceptual masking strength of the voice frame based on the total masking threshold of each frequency component in the frequency domain. For each positive segment in the A-law thirteen-segment line, it determines the sensitivity weight of each positive segment based on the overall perceptual masking strength and the segment number, determines the quantization relaxation coefficient of each positive segment based on the overall perceptual masking strength and the sensitivity weight of each positive segment, adjusts the range size of each positive segment based on the quantization relaxation coefficient, and performs quantization encoding on the voice frame based on the adjusted range size.
[0007] This invention utilizes a psychoacoustic model to dynamically analyze the masking characteristics of speech frames, generating a comprehensive perceived masking strength that reflects the human ear's ability to mask noise. Based on this, the sensitivity weights and quantization relaxation coefficients of each quantization segment are adjusted. This allows for moderate quantization relaxation in high-frequency and low-amplitude segments of high-energy, highly masked speech frames to improve compression efficiency, while maintaining fine quantization in sensitive areas such as voiceless consonants and transitional sounds to protect speech clarity. This effectively avoids the waste of bit resources and auditory distortion caused by traditional fixed quantization, balancing the subjective quality of speech coding with compression efficiency. At the same time, it maintains full compatibility of the decoding end with existing PSTN and IP communication systems, improving the real-time performance and compatibility of intelligent telephone customer service voice interaction in power companies.
[0008] Preferably, determining the overall perceptual masking intensity of the speech frame based on the total masking threshold of each frequency component in the frequency domain includes: dividing the frequency components of the speech frame into several equal-width Bark sub-bands based on the values of each frequency component in the Bark scale in the frequency domain; for any sub-band, obtaining the mean of the total masking thresholds of all frequency components contained in the sub-band as the total masking threshold of that sub-band; dynamically generating the perceptual weights of each sub-band using a Gaussian function; and weighting and summing the total masking thresholds of the sub-bands based on the perceptual weights of each sub-band to obtain the overall perceptual masking intensity of the speech frame.
[0009] This invention obtains the comprehensive sensing masking intensity by weighted summation of sub-band masking thresholds, which not only preserves the integrity of frequency domain sensing information, but also integrates it into scalar parameters that can be used for time domain quantization control. This achieves efficient mapping from multidimensional psychoacoustic analysis to one-dimensional coding adjustment, and provides a scientific and stable sensing basis for subsequent adaptive adjustment of A-law quantization strategies.
[0010] Preferably, the perceptual weights of the sub-bands satisfy the expression: ;in, Indicates the first Perceptual weights of individual bands; Indicates the first The center Bark value of each sub-band This indicates the preset baseline Bark value; is the Gaussian decay factor.
[0011] Preferably, the sensitivity weights of each positive segment satisfy the expression: ;in, Indicates the first The first positive section is in Sensitivity weights for each speech frame; Indicates the sequence number of the main segment; Indicates the first The normalized overall perceptual masking strength of each speech frame.
[0012] This invention utilizes the nonlinear characteristics of power functions to maintain high sensitivity in low-amplitude segments when the masking strength is low, thus protecting voiceless consonants and transitional sounds. When the masking strength is high, it significantly reduces the sensitivity of each segment, especially the low-amplitude segment, allowing the quantization step size to be appropriately increased. This avoids bit waste caused by excessive fine quantization in maskable regions. It retains the A-law protection mechanism for small signals and introduces the adaptive capability of perception-driven processing, significantly improving the balance between coding efficiency and subjective listening quality.
[0013] Preferably, the quantization relaxation coefficient satisfies the expression: ;in, Indicates the first The first positive section is in Quantization relaxation factor per speech frame; This is the preset gain coefficient; Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first The first positive section is in Sensitivity weights for each speech frame.
[0014] This invention uses the normalized integrated perceptual masking strength as the enabling factor for masking capability, ensuring that the relaxation mechanism is activated only when speech is maskable. This allows the adjustment of the range of each positive segment to conform to both human auditory characteristics and engineering controllability. In highly masked speech frames, it effectively increases the quantization step size of high-amplitude segments and some low-amplitude segments, improving compression efficiency. In sensitive speech frames, it maintains quantization accuracy close to standard A-law for each segment, avoiding distortion. This achieves a dynamic balance of coding granularity, taking into account both speech quality and bit utilization.
[0015] Preferably, adjusting the range size of each positive segment according to the quantization relaxation coefficient includes: for each positive segment, calculating the product between the initial range size of each positive segment and the quantization relaxation coefficient, and normalizing the product to obtain the adjusted range size of each positive segment.
[0016] This invention achieves dynamic redistribution of the quantization interval by normalizing the product of the initial range size of each positive segment and the quantization relaxation coefficient. This retains the adaptive adjustment capability based on perceptual intensity while ensuring that the sum of the adjusted ranges of all positive segments remains 1, strictly maintaining the input domain within the range of [0,1]. This avoids signal overflow or decoding mismatch caused by parameter adjustment, ensures full compatibility with the standard A-law decoder, improves the robustness and stability of the algorithm, and supports flexible optimization of coding granularity under different speech content, balancing compression efficiency and speech fidelity.
[0017] Preferably, the step of quantizing and encoding the speech frame according to the adjusted range size includes: determining the adjusted range of each segment and the quantization step size according to the adjusted range size of each segment; and quantizing the first segment... The first in the audio frame The time-domain data were normalized, and the results were used... Indicate; confirm The symbol is searched based on the adjusted range of all main segments. The corresponding positive segment is determined based on its quantization step size. Uniform quantization is performed, the quantization level is calculated, and an 8-bit nonlinear PCM codeword is generated by combining the sign bit to achieve quantization encoding.
[0018] Preferably, determining the adjusted range and quantization step size of each positive segment based on the adjusted range size of each positive segment includes: for the first... Each positive segment has a quantization step size of [number]. ;when At that time, the first The adjusted range of each positive segment is: ;when At that time, the first The adjusted range of each positive segment is: ;when At that time, the first The adjusted range of each positive segment is: ,in Indicates the first The first positive section is in Adjusted range size per speech frame Indicates the first The first positive section is in The adjusted range size for each audio frame.
[0019] Preferably, obtaining the main masking tone of the speech frame includes:
[0020] For each speech frame, the amplitudes corresponding to all frequency components in the speech frame in the spectrum are arranged in ascending order to form an amplitude sequence, and the local maxima in the amplitude sequence are obtained. All local maxima are sorted in descending order, and the frequency components corresponding to the first K local maxima are taken as a main masking tone, where K is the preset number of main masking tones.
[0021] Preferably, the normalized overall perception masking strength satisfies the expression:
[0022] ;in, Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first before normalization The overall perceptual masking strength of each speech frame; Indicates the reference cover level; Indicates the preset maximum masking level; Describes the minimum value function. This represents the maximum value function.
[0023] The beneficial effects of this invention are as follows: This invention can moderately relax quantization when the speech energy is concentrated to improve compression efficiency, and maintain fine quantization in sensitive areas such as voiceless consonants to protect speech clarity. It effectively avoids bit waste and auditory distortion caused by traditional fixed quantization, takes into account both subjective speech quality and coding efficiency, and is intelligent at the encoding end and transparent at the decoding end, which improves the real-time performance and compatibility of intelligent telephone customer service voice interaction for power companies. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating an intelligent telephone customer service interaction method for power companies based on artificial intelligence, as described in this invention.
[0025] Figure 2 This is a flowchart illustrating step 2 of an AI-based intelligent telephone customer service interaction method for power companies according to the present invention. Detailed Implementation
[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0028] This invention discloses an intelligent telephone customer service interaction method for power companies based on artificial intelligence, referring to... Figure 1 This includes steps S1 to S3:
[0029] S1. Real-time acquisition of voice frames during telephone customer service interactions, and conversion of voice frames to the frequency domain.
[0030] Specifically, during the intelligent telephone customer service interaction process of power companies, user voice signals are collected in real time through the telephone interface. These voice signals are generated by sound wave vibrations being converted into continuously varying analog electrical signals via a microphone, and then sampled by an analog-to-digital converter to obtain discrete-time digital voice signals. The sampling rate is set by the implementers according to the actual implementation situation, such as 8000Hz, 16000Hz, 48000Hz, etc.
[0031] It should be noted that speech signals are time-varying and non-stationary; directly performing a Fourier transform on the entire signal will lead to spectral leakage. To accurately analyze local characteristics, this invention divides the speech signal into frames and uses a window function to smooth the frame boundaries, thereby reducing spectral leakage.
[0032] Specifically, the input speech signal is divided into a series of overlapping speech frames. In this embodiment, the duration of each frame (i.e., frame length) is set to 20ms, and the inter-frame shift (i.e., frame shift) is set to 10ms. In other embodiments, implementers can set the frame length and frame shift according to the actual implementation situation.
[0033] Furthermore, windowing is applied to each speech frame separately. The window function is set by the implementer according to the actual implementation situation, such as Hamming window, Hanning window, etc. It should be noted that frame segmentation and windowing are well-known technologies and will not be described in detail here.
[0034] Furthermore, a Short-Time Fourier Transform (STFT) is performed on each windowed speech frame to convert the speech frame from the time domain to the frequency domain, resulting in several frequency components.
[0035] S2. Obtain the main masking tone of the speech frame, determine the total masking threshold of the main masking tone for each frequency component of the speech frame, and determine the comprehensive perceptual masking strength of the speech frame based on the total masking threshold of each frequency component of the speech frame in the frequency domain.
[0036] The flowchart for step S2 is shown below. Figure 2 The process includes steps S201 to S203, specifically as follows:
[0037] S201. Obtain the main masking tone of the speech frame and the power level corresponding to the main masking tone.
[0038] It should be noted that the human auditory system is more sensitive to energy-concentrated regions in the frequency spectrum. These regions often correspond to formants or strong consonant components (such as plosives and fricatives) in speech, possessing strong masking capabilities. For any frequency component of a speech frame in the frequency spectrum, the amplitude is the complex modulus of that frequency component, reflecting the vibration intensity of that frequency component in the speech frame. The larger the amplitude, the stronger the energy of that frequency component, which in speech typically corresponds to energy-concentrated regions such as vowel formants, fundamental harmonics, or strong consonants (such as plosives and fricatives). Therefore, this invention locates the dominant sound components in perception based on the amplitude of each frequency component of the speech frame in the frequency spectrum.
[0039] Specifically, for each speech frame, the amplitudes corresponding to all frequency components in the speech frame's spectrum are arranged in ascending order to form an amplitude sequence, and local maxima are obtained from this sequence. All local maxima are then sorted in descending order, and the frequency components corresponding to the first K local maxima are each used as a primary masking tone. Here, K is a preset number of primary masking tones. In this embodiment, K=6. In other embodiments, implementers can flexibly set the value of K according to the actual application scenario, for example, K=5~10, to adapt to different speech content (e.g., more primary masking tones are needed for dense sections of voiceless consonants).
[0040] It should be noted that the human ear is particularly sensitive to regions of concentrated energy in the frequency spectrum. These regions not only have high loudness themselves, but also exert a strong masking effect on their neighboring frequencies; that is, a loud sound can make weaker sounds nearby inaudible. Therefore, this invention locates the main masking tone by using local maxima in the amplitude sequence.
[0041] It should be further explained that the human ear's perception of sound intensity is non-linear, approximately following a logarithmic law (Weber-Fechner law). Therefore, directly using linear amplitude cannot accurately reflect subjective loudness. Power levels, on the other hand, convert linear amplitude into a logarithmic scale, more closely reflecting the true perceptual characteristics of the human ear. Furthermore, since the masking effect is essentially an energy-level phenomenon—that is, the energy of a loud sound suppresses the energy of a soft sound—this invention converts each main masking tone into a power level.
[0042] Specifically, obtain the power level corresponding to each main masking tone:
[0043] .
[0044] In the formula, Indicates the first The first audio frame The power level corresponding to each main masking tone is expressed in decibels (dB). It means that the first The first audio frame The amplitude corresponding to each main masking tone; This is a smoothing constant used to prevent when... Occasionally This leads to an incorrect formula. In this embodiment, In other embodiments, implementers may set the appropriate parameters according to the actual implementation situation. ,For example ~ It should be noted that the calculation of power stages is a well-known technique, and its specific principles will not be elaborated upon here.
[0045] S202. Determine the total masking threshold for each frequency component of the speech frame by the main masking tone.
[0046] Specifically, the physical frequencies of each main masking tone are converted from the Hz scale to the Bark perception scale.
[0047] It's important to note that the human ear's perception of frequency is non-linear; it has high resolution for low-frequency components but low resolution for high-frequency components. Furthermore, the frequency unit Hz is a linear unit and cannot accurately reflect the true perception of the human ear. Bark is a psychoacoustic frequency scale; 1 bark... A "critical bandwidth" is the smallest frequency band that the human ear can distinguish. Therefore, this invention converts the physical frequencies of each frequency component of a speech frame from the Hz scale to the Bark perceptual scale. The specific conversion process is a well-known technique and will not be described in detail here.
[0048] Furthermore, for any frequency component of a speech frame Based on the asymmetric masking model, the frequency components of each main masking tone are determined. The masking threshold generated at this location:
[0049] .
[0050] In the formula, Indicates the first The first audio frame The main masking tone in the frequency components The masking threshold generated at the location; Indicates the first The first audio frame The power level corresponding to each main masking tone; Indicates the first The first audio frame Individual masking tone and frequency components Distance at the Bark scale , Represents frequency components Values at the Bark scale. Indicates the first The first audio frame The value of the main masking tone at the Bark scale; coefficient 27 is an empirical constant reflecting the masking attenuation rate.
[0051] Furthermore, all the main masking tones of the speech frame are placed in the frequency components. The masking thresholds generated at the location are superimposed to obtain the frequency components. Synthetic masking threshold at:
[0052] .
[0053] in, Indicates the first Each speech frame in frequency components The synthetic masking threshold at the location; Indicates the first The first audio frame The main masking tone in the frequency components The masking threshold generated at the location; Indicates the first The number of frequency components corresponding to each speech frame in the frequency domain. This invention will... The main masking tones of each speech frame in frequency components The masking thresholds generated at each point are converted to linear energy units, summed, and then converted back to a logarithmic scale to obtain the synthetic masking threshold. .
[0054] It should be noted that even in a quiet environment where no other sounds are present, the human ear cannot perceive sounds below a certain minimum intensity. This minimum audible sound intensity varies with frequency and is called the absolute hearing threshold (ATH). For example, the human ear is most sensitive in the 2–5 kHz range, while stronger sound pressure levels are required to detect sounds in the low-frequency (<200 Hz) and high-frequency (>15 kHz) ranges. Therefore, this invention determines the total masking threshold for each frequency component based on the absolute hearing threshold and the synthesis masking threshold of the speech frame at each frequency component.
[0055] Specifically, for any frequency component of a speech frame The frequency components of the speech frame are obtained through the ATH model. The absolute hearing threshold at that point. It should be noted that the ATH model is a well-known technique and will not be elaborated upon here.
[0056] For speech frames in frequency components The synthetic masking threshold at a given point is compared with the absolute hearing threshold, and the maximum value is taken as the frequency component. The total masking threshold at the location.
[0057] It should be noted that the total masking threshold represents the maximum noise intensity that the human ear cannot perceive in the current speech context.
[0058] S203. Determine the overall perceptual masking strength of the speech frame based on the total masking threshold of each frequency component in the frequency domain.
[0059] It should be noted that, in order to achieve perceptual modeling that conforms to the characteristics of human hearing, this invention divides the spectrum of the speech frame into several Bark sub-bands. The frequency range of each sub-band is divided according to the critical bandwidth characteristics of the human ear, so that each sub-band has the same width at the Bark scale, thereby more realistically simulating the filtering response mechanism of the basilar membrane of the human ear.
[0060] Specifically, based on the values of each frequency component in the speech frame at the Bark scale in the frequency domain, the frequency components are divided into several Bark sub-bands of equal width. In this embodiment, the total number of sub-bands is set to 16 to balance the accuracy of perceptual modeling with computational complexity. In other embodiments, implementers can set the total number of sub-bands according to the actual implementation situation.
[0061] Furthermore, for any sub-band, the average of the total masking thresholds at all frequency components contained in the sub-band is obtained as the total masking threshold of that sub-band.
[0062] It should be noted that different frequency bands contribute differently to speech intelligibility across different sub-bands, and cannot be simply weighted equally. The human ear exhibits significant frequency selectivity in its perception of loudness: it is most sensitive in the 1000–1500 Hz frequency range, a critical band for the transition between vowel formants and voiced / unvoiced sounds; while sensitivity gradually decreases towards lower and higher frequencies, forming a bell-shaped response curve. Assigning the same weight to all sub-bands would cause the coding strategy to deviate from the true perceptual patterns of the human ear. Therefore, this invention uses a Gaussian function to dynamically generate the perceptual weights for each sub-band.
[0063] Specifically, the perceptual weights of each sub-band satisfy the expression:
[0064] .
[0065] in, Indicates the first Perceptual weights of individual bands; Indicates the first The center Bark value of each sub-band This indicates the preset baseline Bark value; This is a Gaussian attenuation factor used to control the deviation of the sensing weights from the reference Bark value with frequency. The decay rate.
[0066] For the baseline Bark value When frequency components of 1000–1500Hz belong to the same sub-band, the reference Bark value will be used. Set to the Bark value at the center of the sub-band where the frequency component 1000–1500Hz is located. When the frequency component 1000–1500Hz belongs to a different sub-band, then use the reference Bark value. Set to the mean of the Bark values at the center of the sub-band containing the frequency component 1000–1500Hz. For the Gaussian attenuation factor... The empirical value is 3.0, indicating that when the Bark distance deviates... When the number of units exceeds 3, the perceived weight decays to below approximately 60%, which aligns with the characteristic of a rapid decline in high-frequency / low-frequency sensitivity in the human ear. In other embodiments, implementers can adjust the settings according to the actual application scenario. and The value of is chosen to suit different auditory needs, but it is necessary to ensure that It cannot be 0.
[0067] Furthermore, based on the total masking threshold of each sub-band and the perceptual weights, the overall perceptual masking strength of the speech frame is obtained:
[0068] .
[0069] in, Indicates the first The overall perceptual masking strength of the first speech frame is used to reflect the first... The ability of a single audio frame to mask noise; Indicates the first The first audio frame Total masking threshold for each sub-band; Indicates the number of sub-bands; Indicates the first The perceptual weight of each individual.
[0070] S3. For each positive segment in the A-law thirteen-segment line, determine the sensitivity weight of each positive segment based on the comprehensive perceptual masking strength of the speech frame and the sequence number of the positive segment. Determine the quantization relaxation coefficient of each positive segment based on the comprehensive perceptual masking strength of the speech frame and the sensitivity weight of each positive segment. Adjust the range size of each positive segment based on the quantization relaxation coefficient. Quantize and encode the speech frame based on the adjusted range size.
[0071] It should be noted that the quantization step size structure of the A-law thirteen-segment quantizer is a fixed lookup table mode and does not adjust with changes in speech content. In frames with concentrated speech energy and complex spectra (such as vowels and voiced consonants), using the same quantization precision as for silence or voiceless consonant segments leads to over-quantization in areas imperceptible to the human ear, wasting bit resources. However, this invention integrates the perceived masking strength, which reflects the overall ability of a speech frame to mask noise. Therefore, this invention dynamically adjusts the range of each segment in the A-law thirteen-segment quantizer based on the integrated perceived masking strength, thereby adjusting the quantization step size. This achieves the technical effect of improving speech coding compression efficiency and subjective listening quality while maintaining compatibility with the G.711 standard.
[0072] It should be further explained that when the overall perceptual masking strength is greater, it indicates that the speech energy in the speech frame is more concentrated and the spectrum is more complex, such as being in the vowel segment, voiced segment, or strong resonance region. At this time, the main tone component can effectively mask the quantization noise in its vicinity, allowing for a larger quantization error without causing distortion perceptible to the human ear. Conversely, when the overall perceptual masking strength is smaller, it indicates that the speech is in the sensitive region of voiceless consonants, transitional sounds, or near silence. At this time, the masking effect is weak, and fine quantization must be maintained to protect high-frequency details and speech clarity, and avoid "harsh sounds" or recognition errors.
[0073] Specifically, the intensity of the integrated perception masking is normalized to eliminate the influence of its dimensions. The normalization expression is:
[0074] .
[0075] in, Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first before normalization The overall perceptual masking strength of each speech frame; Indicating the reference masking level, in this invention, dB, this is because when When dB is used, it usually corresponds to voiceless consonants, transitional sounds, or near-silent speech segments. Fine quantization is required to protect speech clarity and intelligibility. This indicates the preset maximum masking level. In this invention, dB is used because this value corresponds to the typical upper limit of high-energy speech (such as vowels, plosives, and strong resonances) in telephone communication systems. Speech frames exceeding this value have extremely strong masking capabilities, allowing for maximum quantization compression. Implementers can also set the value according to the actual implementation situation. as well as However, it must meet the following requirements. ; Describes the minimum value function. Represents the maximum value function. , Used to restrict The range is between [0,1].
[0076] Furthermore, since the positive and negative semi-axes of the A-law are symmetrical about the origin, this invention uses the eight positive segments of the positive semi-axe of the A-law as an example for explanation. For each of the eight positive segments, the initial range size of each segment is denoted as... , … Its value is determined by the standard A-law function.
[0077] Based on the normalized comprehensive perceptual masking strength of the speech frames and the sequence number of the segments, the sensitivity weight of each segment is determined:
[0078] .
[0079] in, Indicates the first The first positive section is in Sensitivity weights for each speech frame; Indicates the sequence number of the main segment; Indicates the first The normalized overall perceptual masking strength of each speech frame.
[0080] It should be noted that, under static conditions without considering the masking effect, the low amplitude range ( Smaller amplitude ranges correspond to small signal regions, to which the human ear is extremely sensitive (e.g., voiceless consonants, transitional sounds), thus exhibiting high perceptual sensitivity; while higher amplitude ranges ( Larger signals (corresponding to large signal regions) have strong energy, but the human ear's ability to distinguish them is relatively low, thus resulting in lower sensitivity. However, in actual speech, the masking effect dynamically changes perceptual sensitivity; when speech energy is concentrated and the spectrum is complex, i.e.,... When the amplitude is large, low-amplitude components may be effectively masked by the main tone, and their sensitivity should be appropriately reduced; while when the speech is sparse or in a silence transition zone, i.e. When the amplitude is small, the low-amplitude range must maintain high sensitivity to prevent quantization noise exposure. Therefore, this invention uses the normalized result of the positive segment number. With base as, The sensitivity weights are dynamically adjusted using a gamma transform, with the result being an exponent. This is achieved when the normalized overall perception masking intensity... The smaller the size, the weaker the concealment capability. The closer it is to 1, the better. The closer the value is to Therefore, sensitivity weight near This approach achieves a higher sensitivity weight in the low-amplitude range, maintaining high sensitivity but requiring fine quantization to preserve speech details, while the higher-amplitude range has a lower sensitivity weight, allowing for coarse quantization. The normalized overall perceptual masking strength... The larger the size, the stronger the concealment capability. The smaller, the better The greater the increase, the more... The value in The larger the base, the greater the sensitivity weight. Reduced. Meanwhile, due to the nonlinear characteristics of the gamma transform, at larger... In the case of, when The smaller the size, the more... The greater the increase on this basis, the more... When it is larger, in The smaller the increase on the basis, the better for the low amplitude range ( (Smaller) sensitivity weight The decrease is more significant, increasing its range and thus allowing for a moderate increase in its quantization step size. This avoids bit waste caused by "overprotection" in a strong masking environment, especially in the high-amplitude range. (Larger) sensitivity weight The decrease is relatively gradual, maintaining a certain degree of coarseness, which aligns with the physiological characteristic that the human ear is inherently less sensitive to high-amplitude ranges. It should be noted that when... When the value is close to 1, it indicates that the speech frame has extremely strong energy and a complex spectrum. The main tone component has a strong masking ability. At this time, even the quantization noise in the low amplitude range is almost completely masked. Therefore, the quantization can be relaxed as a whole to maximize the compression efficiency.
[0081] Furthermore, based on the normalized comprehensive perceptual masking strength of the speech frames and the sensitivity weights of each positive segment, the quantization relaxation coefficient of each positive segment is determined:
[0082] .
[0083] in, Indicates the first The first positive section is in The quantization relaxation factor under each speech frame is used to control the quantization granularity of each segment of the speech frame. When the quantization granularity is greater than 1, it means that the positive segment can be coarsely quantized. When the quantization granularity is equal to 1, it means that the original precision is maintained. The preset gain coefficient is used to limit the step size variation and prevent audible distortion due to excessive adjustment. The value range is [0.01, 0.05] to ensure smooth adjustment of the quantization step size and avoid audible distortion caused by excessive adjustment. In this embodiment, In other embodiments, implementers can set the value of the gain coefficient within the range of the gain coefficient according to the actual implementation situation; Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first The first positive section is in Sensitivity weights for each speech frame.
[0084] In the formula, the sensitivity weight Reflects the first The compressibility of a positive segment, when the sensitivity weight... The lower the value, the better. The more "compressible" a segment is, the better. The larger the value, the greater the range of the positive segment should be, thus increasing the quantization step size of the positive segment. However, this adjustment only takes effect when the masking capability is strong; therefore, a normalized overall perceived masking strength is introduced. As an enabling factor, when The smaller the child, the weaker their regulatory capacity. The larger the size, the stronger the regulatory capacity.
[0085] Furthermore, the adjusted range of each positive segment is determined based on the quantification relaxation coefficient:
[0086] .
[0087] in, Indicates the first The first positive section is in The adjusted range size for each audio frame; Indicates the first The first positive section is in Quantization relaxation factor per speech frame; Indicates the first The initial range size of each positive segment. Used for Normalization is performed to ensure that the sum of the adjustment ranges of all positive segments is 1, that is, the input domain remains [0,1], which is compatible with the standard A-law decoder and avoids sample overflow.
[0088] Furthermore, based on the adjusted range of each positive segment, the adjusted range and quantization step size are determined, including:
[0089] For the A positive section, when At that time, the first The adjusted range of each positive segment is: Quantization step size is ;when At that time, the first The adjusted range of each positive segment is: Quantization step size is ;when At that time, the first The adjusted range of each positive segment is: Quantization step size is .
[0090] Furthermore, the speech frames are quantized and encoded according to the adjusted range of each segment, including:
[0091] For the The first in the audio frame The time-domain data were normalized, and the results were used... Indicate. Confirm. The symbol is searched based on the adjusted range of all main segments. The corresponding positive segment is determined based on its quantization step size. Uniform quantization is performed, the quantization level is calculated, and an 8-bit nonlinear PCM codeword is generated by combining the sign bit. This is achieved by... The speech frame is encoded by quantizing and encoding each time-domain data in each speech frame.
[0092] It should be noted that this invention achieves a technical paradigm of "intelligent adaptation at the encoder and full compatibility at the decoder" by dynamically adjusting the segmentation parameters of the A-law thirteen-segment quantizer. The encoder adaptively optimizes the quantization granularity according to the speech content, while the decoder still uses the standard G.711 inverse quantization table without any modification, and can be directly connected to PSTN or IP networks, significantly improving compression efficiency without sacrificing subjective speech quality.
Claims
1. A method for intelligent telephone customer service interaction in power companies based on artificial intelligence, characterized in that: include: Real-time acquisition of voice frames during telephone customer service interactions, and conversion of voice frames to the frequency domain; Obtain the main masking tone of the speech frame, determine the total masking threshold of the main masking tone for each frequency component of the speech frame, and determine the comprehensive perceptual masking strength of the speech frame based on the total masking threshold of each frequency component of the speech frame in the frequency domain. For each positive segment in the A-law thirteen-segment line, the sensitivity weight of each positive segment is determined based on the comprehensive perceptual masking strength of the speech frame and the sequence number of the positive segment. The quantization relaxation coefficient of each positive segment is determined based on the comprehensive perceptual masking strength of the speech frame and the sensitivity weight of each positive segment. The range size of each positive segment is adjusted based on the quantization relaxation coefficient. The speech frame is then quantized and encoded based on the adjusted range size.
2. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The determination of the comprehensive perceptual masking strength of a speech frame based on the total masking threshold at each frequency component in the frequency domain includes: Based on the values of each frequency component in the frequency domain at the Bark scale, the frequency components are divided into several equal-width Bark sub-bands; for any sub-band, the average of the total masking thresholds of all frequency components contained in the sub-band is obtained as the total masking threshold of that sub-band. The perceptual weights of each sub-band are dynamically generated using a Gaussian function. The total masking threshold of the sub-bands is then weighted and summed based on the perceptual weights of each sub-band to obtain the overall perceptual masking intensity of the speech frame.
3. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 2, characterized in that, The perceptual weights of the sub-bands satisfy the expression: ; in, Indicates the first Perceptual weights of individual bands; Indicates the first The center Bark value of each sub-band This indicates the preset baseline Bark value; is the Gaussian decay factor.
4. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The sensitivity weights of each positive segment satisfy the expression: ; in, Indicates the first The first positive section is in Sensitivity weights for each speech frame; Indicates the sequence number of the main segment; Indicates the first The normalized overall perceptual masking strength of each speech frame.
5. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The quantization relaxation factor satisfies the expression: ; in, Indicates the first The first positive section is in Quantization relaxation factor per speech frame; This is the preset gain coefficient; Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first The first positive section is in Sensitivity weights for each speech frame.
6. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The adjustment of the range size of each positive segment based on the quantization relaxation coefficient includes: For each positive segment, calculate the product between the initial range size of each positive segment and the quantization relaxation coefficient, and normalize the product to obtain the adjusted range size of each positive segment.
7. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The step of quantizing and encoding the speech frame according to the adjusted range size includes: The adjusted range and quantization step size of each positive segment are determined based on the adjusted range size of each positive segment; for the first... The first in the audio frame The time-domain data were normalized, and the results were used... Indicate; confirm The symbol is searched based on the adjusted range of all main segments. The corresponding positive segment is determined based on its quantization step size. Uniform quantization is performed, the quantization level is calculated, and an 8-bit nonlinear PCM codeword is generated by combining the sign bit to achieve quantization encoding.
8. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 7, characterized in that, The process of determining the adjusted range and quantization step size of each positive segment based on the adjusted range size of each positive segment includes: For the Each positive segment has a quantization step size of [number]. ;when At that time, the first The adjusted range of each positive segment is: ;when At that time, the first The adjusted range of each positive segment is: ;when At that time, the first The adjusted range of each positive segment is: ,in Indicates the first The first positive section is in Adjusted range size per speech frame Indicates the first The first positive section is in The adjusted range size for each audio frame.
9. The intelligent telephone customer service interaction method for power companies based on artificial intelligence according to claim 1, characterized in that, The acquisition of the main masking tone of the speech frame includes: For each speech frame, the amplitudes corresponding to all frequency components in the speech frame in the spectrum are arranged in ascending order to form an amplitude sequence, and the local maxima in the amplitude sequence are obtained. All local maxima are sorted in descending order, and the frequency components corresponding to the first K local maxima are taken as a main masking tone, where K is the preset number of main masking tones.
10. A method for intelligent telephone customer service interaction for power companies based on artificial intelligence, as described in claim 4 or 5, characterized in that, The normalized integrated sensing masking strength satisfies the expression: ; in, Indicates the first The normalized overall perceptual masking strength of each speech frame; Indicates the first before normalization The overall perceptual masking strength of each speech frame; Indicates the reference cover level; Indicates the preset maximum masking level; Describes the minimum value function. This represents the maximum value function.
Citation Information
Patent Citations
Method and device for detecting voice endpoint
CN103730110A
Speech recognition method, related device, electronic equipment and storage medium
CN114842833A