User Dialogue Interaction Method and System Based on Deep Learning

By using LMS algorithm and frequency domain analysis technology to remove echoes in voice input in deep learning-based user dialogue interaction methods, the echo interference problem is solved and the accuracy of interaction is improved.

CN119811413BActive Publication Date: 2025-05-27QINGDAO CITY BRAIN INVESTMENT DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510300194.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-05-27
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

In the prior art, deep learning-based user dialogue interaction methods are susceptible to echo interference when processing voice input, affecting the accuracy of model learning.

Method used

The echo signal in the nearest input signal of the speech is removed by the LMS algorithm, the echo probability of each frequencies is analyzed using the frequency domain similarity and frequency domain attenuation degree, and the step size of the LMS algorithm is dynamically adjusted to improve the effect of echo suppression.

Benefits of technology

It effectively improves the accuracy of user dialogue interaction based on deep learning and reduces the impact of echo signals on the interaction process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119811413B_ABST
    Figure CN119811413B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of voice data processing, and particularly to a user dialogue interaction method and system based on deep learning. The method includes the steps of: obtaining the target delay time between the reference signal and the amplitude of the output signal; obtaining the frequency-domain similarity through the amplitude between the reference signal and the frequency of the proximal input signal; obtaining the frequency-domain attenuation degree of the reference signal or the proximal input signal at this frequency through the amount of change in the adjacent amplitudes of the frequency window in the reference signal or the proximal input signal; combining the frequency-domain similarity between the reference signal and the frequency of the proximal input signal and the frequency-domain attenuation degree of the reference signal or the proximal input signal, determining the echo probability of each frequency in the proximal input signal and adjusting the step size, and using the step size of the frequency in the LMS algorithm to eliminate the echo in the proximal input signal, so as to realize user dialogue interaction and effectively improve the accuracy of user dialogue interaction based on deep learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech data processing technology, and in particular to a user dialogue interaction method and system based on deep learning. Background Art

[0002] The deep learning model can obtain the interaction records generated during the user interaction process. By learning the language rules and patterns in the interaction records, it can understand and simulate the complexity of human language, and provide strong support for smart phone conferences, smart homes and other fields. In order to improve the user experience and enhance the intelligence of human-computer interaction, it is necessary to continuously obtain user interaction records to train and optimize the deep learning model.

[0003] There are many studies in the prior art on how to improve the training efficiency of deep learning models. For example, a patent application document with publication number CN116796758A discloses a dialogue interaction method, dialogue interaction device, equipment and storage medium. The application determines the dialogue type of dialogue text information according to the text content of the dialogue text information; fills the dialogue type, target attributes and their corresponding associated attributes into a preset dialogue information template to obtain target dialogue information; inputs the target dialogue information into the dialogue model to obtain dialogue interaction information corresponding to the dialogue text information.

[0004] The above-mentioned prior art realizes interactive learning of the dialogue model by obtaining the text content of the dialogue text information. However, the dialogue text information usually comes from the user's text input and voice input. The text input content can be processed directly, but the voice input content usually has echo interference, which affects the accuracy of model learning.

[0005] Based on this, how to effectively improve the accuracy of user dialogue interaction based on deep learning is an urgent problem to be solved by technical personnel in this field. Summary of the invention

[0006] In order to solve the technical problem of how to effectively improve the accuracy of user dialogue interaction based on deep learning, the present invention provides a user dialogue interaction method and system based on deep learning.

[0007] In a first aspect, the present invention provides a user dialogue interaction method based on deep learning, which adopts the following technical solution:

[0008] The user dialogue interaction method based on deep learning includes the following steps:

[0009] The near-end input signal and output signal of the speech during the interaction are obtained, and the far-end input signal is recorded as the reference signal; the delay time corresponding to the maximum value of the cross-correlation function of the amplitude of the reference signal and the output signal at each delay time is recorded as the target delay time ; Through the reference signal The frequency of the near-end input signal The amplitude difference corresponding to the frequency is used to obtain the frequency domain similarity; the frequency window is obtained, and the frequency domain attenuation degree of the reference signal or the near-end input signal at the frequency is obtained by the average of the adjacent amplitude changes of the frequency window in the reference signal or the near-end input signal; the frequency window in the reference signal is obtained. The frequency of the near-end input signal The frequency domain similarity between The product of The frequency of the near-end input signal The absolute value of the frequency domain attenuation degree difference of the frequencies is normalized by the ratio of the product and the absolute value of the difference to obtain the first frequency in the near-end input signal. The probability of echo at each frequency; ; The near-end input signal The frequency step size, is the preset maximum step size, is a hyperparameter, The near-end input signal The echo probability of a frequency, e is a constant; the frequency step size is used in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction.

[0010] The present invention takes into account that the echo signal generated during the user dialogue interaction will affect the accuracy of the interaction, and therefore removes the echo signal during the user dialogue interaction through the LMS algorithm, effectively improving the accuracy of the user dialogue interaction based on deep learning. In this process, the present invention takes into account that the LMS algorithm uses a fixed step size, and a step size that is too long or too short will affect the convergence efficiency and accuracy of the algorithm. Based on this, the present invention analyzes the possibility that each frequency in the near-end input signal of the voice is an echo based on the echo feature, and adjusts the step size of the frequency based on the possibility that each frequency is an echo, which can effectively improve the accuracy of the LMS algorithm in eliminating echoes, thereby effectively improving the accuracy of the user dialogue interaction based on deep learning.

[0011] According to the user dialogue interaction method based on deep learning provided by the present invention, the near-end input signal and output signal of the speech during the interaction process are obtained, and the far-end input signal is recorded as a reference signal. It also includes: obtaining the amplitude value at each acquisition moment during the dialogue interaction process for preprocessing to obtain the near-end input signal, output signal and far-end input signal.

[0012] The present invention takes into account that the initially collected voice signal is not conducive to data processing, and therefore improves the overall quality of the data through preprocessing to prepare for subsequent data processing.

[0013] According to the user dialogue interaction method based on deep learning provided by the present invention, the method for obtaining the cross-correlation function includes: presetting the neighborhood of the reference signal or the output signal at each acquisition moment; obtaining the mean of the neighborhood amplitude entropy values ​​of the reference signal and the output signal at the same acquisition moment, and using the negative of the mean as the exponent of the exponential function with e as the base to obtain the stable weights of the reference signal and the output signal at the acquisition moment; ; The reference signal and the output signal are delayed by The cross-correlation function value under The reference signal and the output signal are The stable weight of each acquisition moment, The reference signal The amplitude value at each acquisition moment, The output signal The amplitude value at the acquisition moment.

[0014] The present invention provides a precise method for constructing a cross-correlation function of a reference signal and an output signal under a delay time, which can effectively improve the accuracy of constructing the cross-correlation function by weighting based on the stable weights of the reference signal and the output signal at each acquisition moment.

[0015] According to the user dialogue interaction method based on deep learning provided by the present invention, the frequency domain similarity satisfies the relationship: ;

[0016] The reference signal The frequency of the near-end input signal The frequency domain similarity between the frequencies, is the target delay time, , are the reference signals frequency, the near-end input signal The amplitude value corresponding to the frequency is is an exponential function with base e, is an absolute value.

[0017] According to the deep learning-based user dialogue interaction method provided by the present invention, before calculating the step size of the frequency in the near-end input signal, it also includes: removing the frequencies in the near-end input signal whose echo probability is greater than a preset echo threshold.

[0018] The present invention takes into account the large amount of frequency data generated during user interaction. Therefore, before determining the step size based on the frequency echo probability, some frequencies with large echo probabilities can be eliminated by presetting an echo threshold to reduce the data volume and effectively improve the efficiency of algorithm processing.

[0019] According to the user dialogue interaction method based on deep learning provided by the present invention, the frequency step size is used in the LMS algorithm to eliminate the echo in the near-end input signal, including: presetting a weight vector; obtaining a filter output based on the weight vector and the near-end input signal, and iteratively updating the weight vector through the error between the reference signal and the filter output and the frequency step size; using the weight vector that meets the iteration termination condition after the update as the target weight vector, and eliminating the output of the filter corresponding to the target weight vector as the echo, so as to obtain the near-end input signal after the echo is eliminated.

[0020] According to the deep learning-based user dialogue interaction method provided by the present invention, the frequency step size is used in the LMS algorithm to eliminate the echo in the near-end input signal to achieve user dialogue interaction, including: obtaining text information of the near-end input signal after the echo is eliminated, converting the text information into voice information and outputting it to the user.

[0021] The present invention eliminates echoes in near-end voice input signals through the LMS algorithm, which can effectively reduce the impact of echo signals on the user interaction process, thereby effectively improving the accuracy of user dialogue interaction based on deep learning.

[0022] In a second aspect, the present invention provides a user dialogue interaction system based on deep learning, which adopts the following technical solution:

[0023] A user dialogue interaction system based on deep learning includes: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the user dialogue interaction method based on deep learning is implemented.

[0024] By adopting the above technical solution, the above-mentioned deep learning-based user dialogue interaction method is generated into a computer program and stored in a memory so as to be loaded and executed by a processor, thereby making a terminal device based on the memory and the processor for easy use.

[0025] The present invention has the following technical effects:

[0026] Based on the above technical solution, the present invention removes the echo signal in the user dialogue interaction process through the LMS algorithm when implementing the user dialogue interaction based on deep learning, which effectively improves the accuracy of the user dialogue interaction based on deep learning. In this process, the present invention analyzes the possibility of each frequency in the near-end voice input signal being an echo based on the echo feature, and adjusts the frequency step size based on the possibility of each frequency being an echo, which can effectively improve the accuracy of echo elimination, thereby effectively improving the accuracy of user dialogue interaction based on deep learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts.

[0028] Figure 1 A flowchart of a user dialogue interaction method based on deep learning provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0029] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0030] It should be understood that when the terms "first", "second", etc. are used in the claims, descriptions, and drawings of the present invention, they are only used to distinguish different objects, rather than to describe a specific order. The terms "include" and "comprise" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their collections.

[0031] The deep learning model can obtain the interaction records generated during the user interaction process. By learning the language rules and patterns in the interaction records, it can understand and simulate the complexity of human language, and provide strong support for smart phone conferences, smart homes and other fields. In order to improve the user experience and enhance the intelligence of human-computer interaction, it is necessary to continuously obtain user interaction records to train and optimize the deep learning model.

[0032] The interaction records of deep learning models usually include user text input and voice input, but the content of voice input usually has echo interference, which affects the accuracy of model learning.

[0033] The Least Mean Square (LMS) algorithm is an echo cancellation algorithm that estimates an approximate echo path to approximate the real echo path by setting a fixed step size to adjust the filter weight vector, thereby obtaining an estimated echo signal, and removes this signal from the mixed signal of near-end speech and far-end echo to achieve echo cancellation.

[0034] Based on this, an embodiment of the present invention discloses a user dialogue interaction method based on deep learning. The method removes the echo signal in the near-end speech input signal through LMS, and realizes user dialogue interaction based on the near-end speech input signal with the echo signal removed, which can effectively improve the accuracy of user dialogue interaction based on deep learning.

[0035] For details, please refer to Figure 1 As shown, Figure 1 A flowchart of a user dialogue interaction method based on deep learning is provided in an embodiment of the present invention, and the method specifically includes the following steps.

[0036] S1: Obtain the near-end input signal and output signal of the speech during the interaction process, and record the far-end input signal as the reference signal.

[0037] It should be noted that users usually involve three main voice signals in the process of dialogue interaction, including voice input from far-end users, voice input from near-end users, and voice output. When the far-end signal is transmitted to the near-end and played through the speaker, the near-end microphone will pick up the far-end sound played by the speaker and the voice input of the near-end user as the near-end input signal. At this time, the near-end output signal received by the far-end user will contain the delayed echo of the far-end sound emitted by itself, thereby reducing the accuracy of the voice information received by the far-end user and affecting the user experience.

[0038] Based on this, the embodiment of the present invention obtains the far-end user voice input as a reference signal, and screens out the echo in the near-end input signal through the time domain and frequency domain characteristics of the reference signal and the near-end input signal and output signal.

[0039] For example, in an embodiment of the present invention, a near-end input signal and an output signal of speech are obtained during the interaction process, and a far-end input signal is recorded as a reference signal. Previously, the method also includes: obtaining the amplitude value at each acquisition moment during the dialogue interaction process for preprocessing to obtain a near-end input signal, an output signal, and a far-end input signal.

[0040] Among them, the preprocessing can be analog-to-digital conversion of the speech signal, denoising, missing data interpolation, removal of silence by endpoint detection method, etc., which can be specifically set according to actual needs, and the embodiments of the present invention do not impose too many restrictions on this.

[0041] Specifically, the collection time interval can be preset to respectively construct the time domain rectangular coordinate system, frequency domain rectangular coordinate system and time-frequency spectrum of the near-end input signal, output signal and reference signal with the collection time, the frequency and amplitude corresponding to the collection time.

[0042] The collection time interval may be preset to 1 second; it may be specifically set according to actual needs, and the embodiment of the present invention does not impose too many restrictions on this.

[0043] The time-frequency spectrum can show how the signal changes with time and frequency. The colors in the time-frequency spectrum can reflect the amplitude of the signal at a specific time and frequency point.

[0044] After obtaining the near-end input signal, output signal and reference signal based on the above steps, the following steps may be performed.

[0045] It should be further explained that the LMS algorithm is an echo cancellation algorithm, which achieves echo cancellation by setting a fixed step size to adjust the weight vector of the filter. However, if the fixed step size is small, the algorithm converges too slowly, and the echo path cannot be prepared when the delay time between the echo and the reference signal is long; on the contrary, if the fixed step size is large, the filter coefficients will be updated too quickly, making it unstable and difficult to converge to the expected value.

[0046] Based on this, the embodiment of the present invention performs the following steps to analyze the delay characteristics between the reference signal and the output signal, as well as the frequency domain similarity characteristics and frequency domain attenuation characteristics of the reference signal and the near-end input signal at each acquisition moment, thereby obtaining the possibility that each frequency in the near-end input signal belongs to an echo, and obtaining the step size of each frequency according to the distribution of the echo signal.

[0047] S2: The delay time corresponding to the maximum value of the cross-correlation function between the amplitudes of the reference signal and the output signal at each delay time is recorded as the target delay time.

[0048] It should be noted that the characteristics of the reference signal and the near-end input signal containing the echo are similar in some frequency bands, but the echo has a delay characteristic. Therefore, the delay characteristic between the reference signal and the near-end input signal containing the echo can be obtained first. The acquisition equipment of the near-end input signal and the output signal is close. Therefore, the delay characteristic between the reference signal and the output signal can be obtained as the delay time between the reference signal and the output signal. The delay time can be obtained through the phase difference between the reference signal and the output signal in the time domain.

[0049] Based on this, an embodiment of the present invention obtains the similarity between the reference signal and the output signal at different delay times in the time domain. When the similarity between the signals is the largest, it means that the shapes, waveforms and changes between the two are the most similar. At this time, the corresponding delay time is the delay characteristic between the reference signal and the output signal.

[0050] In order to reduce the amount of data processing, the acquisition range of the delay time can be set according to actual needs, and each possible target delay time can be traversed in the acquisition range. delay time.

[0051] It can be understood that the cross-correlation function can be used to evaluate the similarity between two signals at different delay times.

[0052] For example, the method of obtaining the cross-correlation function includes the following two possible implementation methods:

[0053] In a possible implementation, the embodiment of the present invention can directly obtain the amplitude values ​​of the reference signal and the output signal at each acquisition time, and construct a cross-correlation function. For details, see the following relationship:

[0054] ;

[0055] The reference signal and the output signal are delayed by The cross-correlation function value under The reference signal The amplitude value at each acquisition moment, The output signal The amplitude value at each acquisition moment, is positive infinity, is negative infinity, is the differential symbol, is the integral symbol.

[0056] In this way, the embodiment of the present invention directly constructs a cross-correlation function based on the amplitude values ​​of the reference signal and the output signal at each acquisition moment, and can accurately and quickly obtain the similarity between the reference signal and the output signal, thereby determining the target delay time.

[0057] In another possible implementation, an embodiment of the present invention may preset a neighborhood of a reference signal or an output signal at each acquisition moment; obtain the mean of the neighborhood amplitude entropy values ​​of the reference signal and the output signal at the same acquisition moment, and use the negative of the mean as the exponent of an exponential function with e as the base to obtain the stable weights of the reference signal and the output signal at the acquisition moment; and construct a cross-correlation function through the reference signal, the output signal, and the stable weights of the reference signal and the output signal at each acquisition moment.

[0058] The neighborhood of each acquisition moment in the reference signal or output signal can be centered on the position of the current acquisition moment in the time domain rectangular coordinate system, and a number of other acquisition moments are equally obtained on the left and right sides of the current moment as the neighborhood of the current moment. The size of the neighborhood can be 9, and the size of the neighborhood is the number of acquisition moments included in the neighborhood. The size of the neighborhood can be set according to actual needs.

[0059] For example, the stable weights of the reference signal and the output signal at each acquisition moment are determined, and the details can be seen in the following relationship:

[0060] ;

[0061] The reference signal and the output signal are The stable weight of each acquisition moment, The reference signal The neighborhood amplitude entropy value at each acquisition moment, The output signal The neighborhood amplitude entropy value at each acquisition moment, is an exponential function with base e.

[0062] Among them, the specific steps of obtaining the neighborhood amplitude entropy value at the collection time can be obtained through the existing technology, and the embodiment of the present invention will not be described in detail here.

[0063] In the above formula, the reference signal and the output signal are The larger the mean of the neighborhood amplitude entropy value at each acquisition moment, the closer the reference signal and the output signal are at the first acquisition moment. The greater the degree of chaos at each collection moment, the worse the stability and the lower the corresponding stability weight.

[0064] For example, the cross-correlation function value of the reference signal and the output signal at the delay time is determined by the reference signal, the output signal, and the stable weights of the reference signal and the output signal at each acquisition time. For details, see the following relationship:

[0065] ;

[0066] The reference signal and the output signal are delayed by The cross-correlation function value under The reference signal and the output signal are The stable weight of each acquisition moment, The reference signal The amplitude value at each acquisition moment, The output signal The amplitude value at the acquisition moment.

[0067] In this way, the embodiment of the present invention takes into account that the stability of the speech signal at some acquisition moments is poor and the frequency changes of the signals in different time periods are large. Therefore, by setting a weight for each acquisition moment and setting a higher weight for the acquisition moment with better stability, the credibility of the similarity degree obtained based on the cross-correlation function is improved.

[0068] The embodiment of the present invention constructs a cross-correlation function based on the above method, and can obtain the correlation degree between the reference signal and the output signal at each delay time, so as to accurately obtain the delay time corresponding to the maximum value of the cross-correlation function at each delay time, record it as the target delay time, and continue to execute the following steps.

[0069] S3: The frequency domain similarity is obtained by the amplitude difference between the frequencies in the reference signal and the frequencies in the near-end input signal.

[0070] It should be noted that, based on the above steps, the target delay time between the reference signal and the output signal can be obtained, and the acquisition devices of the near-end input signal and the output signal are close, and the delay is very small. Therefore, the target delay time between the reference signal and the output signal can be used as the delay time between the reference signal and the near-end input signal. Based on the delay time between the reference signal and the near-end input signal, the frequency domain similarity between the reference signal and the near-end input signal can be obtained.

[0071] For example, in the embodiment of the present invention, the first The frequency of the near-end input signal The amplitude difference corresponding to the frequency is used to obtain the frequency domain similarity.

[0072] For example, in an embodiment of the present invention, determining the first The frequency of the near-end input signal The frequency domain similarity between the frequencies can be specifically seen in the following relationship:

[0073] ;

[0074] The reference signal The frequency of the near-end input signal The frequency domain similarity between the frequencies, is the target delay time, , are the reference signals frequency, the near-end input signal The amplitude value corresponding to the frequency is is an exponential function with base e, is an absolute value.

[0075] In the above formula, the reference signal The frequency of the near-end input signal The greater the difference in the amplitude values ​​corresponding to the two frequencies, the lower the similarity between them.

[0076] It can be understood that some frequencies in the near-end input signal have no corresponding frequencies in the reference signal, and such frequencies may not be processed.

[0077] After obtaining the frequency domain similarity between each frequency in the near-end input signal and each frequency in the reference signal through the above formula, continue to perform the following steps.

[0078] S4: Obtain a frequency window, and obtain the frequency domain attenuation degree of the reference signal or the near-end input signal at the frequency by averaging the adjacent amplitude changes of the frequency window in the reference signal or the near-end input signal.

[0079] It should be noted that, based on the above steps, by analyzing the target frequency changes of the reference signal and the near-end input signal, the delay characteristics of the echo characterized by domain similarity can be obtained; in addition, the echo also has the characteristic of slow attenuation.

[0080] Based on this, the embodiment of the present invention obtains the corresponding frequency domain attenuation degree by analyzing the amplitude change of the reference signal or the near-end input signal in its window.

[0081] Among them, the size of the window can be set to 10, and the size of the window can be set according to actual needs.

[0082] For example, when obtaining the frequency window, there are two possible implementation methods:

[0083] In a possible implementation, the position of the current frequency in the time-frequency spectrum may be taken as a starting point, and a plurality of frequencies on the right side thereof may be acquired as a window of the current frequency.

[0084] In another possible implementation, the position of the current frequency in the time-frequency spectrum may be taken as a starting point, and a plurality of frequencies on the left side thereof may be acquired as windows of the current frequency.

[0085] For example, if the size of the current frequency window is 10, you can use the current frequency as the starting point and obtain 9 frequencies on the right or left side of the current frequency as the window of the current frequency. The final window contains the current frequency itself.

[0086] It can be understood that since the frequency and amplitude in the frequency window correspond one to one, the adjacent amplitude changes on the time-frequency spectrum are the absolute values ​​of the amplitude differences corresponding to the adjacent frequencies in the window. The larger the mean value of the adjacent amplitude changes in the window of the current frequency, the more drastic the change of the signal amplitude with the frequency, and the greater the frequency domain attenuation of the corresponding current frequency.

[0087] Take the adjacent amplitudes of the frequency window as an example: if the window length is 3, there are two sets of adjacent amplitudes, namely the amplitudes corresponding to the first and second frequencies, and the amplitudes corresponding to the second and third frequencies.

[0088] After obtaining the frequency domain attenuation degree of the reference signal or the near-end input signal at each frequency based on the above steps, continue to perform the following steps.

[0089] S5: Calculate the echo probability and step size of each frequency in the near-end input signal.

[0090] It should be noted that, based on the above steps, by analyzing the delay characteristics and slow attenuation characteristics of the reference signal and the near-end input signal, the frequency domain similarity between the reference signal and the near-end input signal, as well as the frequency domain attenuation degree of the reference signal and the near-end input signal can be obtained. By combining the frequency domain similarity and frequency domain attenuation degree between the reference signal and the near-end input signal, the possibility of each frequency in the near-end input signal being an echo can be accurately obtained.

[0091] For example, in the embodiment of the present invention, when calculating the echo probability of each frequency in the near-end input signal, the first The frequency of the near-end input signal The frequency domain similarity between The product of The frequency of the near-end input signal The absolute value of the frequency domain attenuation degree difference of the frequencies is normalized by the ratio of the product and the absolute value of the difference to obtain the first frequency in the near-end input signal. The probability of an echo at a frequency.

[0092] For example, by combining the frequency domain similarity and frequency domain attenuation degree between the reference signal and the near-end input signal, the possibility that each frequency in the near-end input signal is an echo is obtained, as shown in the following relationship:

[0093] ;

[0094] The near-end input signal The probability of an echo at a frequency, The reference signal The frequency of the near-end input signal The frequency domain similarity between the frequencies, is the target delay time, The reference signal The frequency domain attenuation degree of a frequency, The near-end input signal The frequency domain attenuation degree of a frequency, is the absolute value, is a linear normalization function.

[0095] In the above formula, frequency domain similarity is used to describe the similarity of the change trend between the reference signal and the near-end input signal. The larger the The smaller the value, the greater the near-end input signal. The closer the characteristics of the frequency are to the echo characteristics, the closer the corresponding near-end input signal is to the echo characteristics. The higher the echo probability of a certain frequency.

[0096] After the echo probability of each frequency in the near-end input signal is obtained based on the above formula, the step size in the LMS algorithm can be determined based on the echo probability of each frequency.

[0097] It should be further explained that the amount of data generated during the user interaction process is large. In order to reduce the amount of data processing and improve the efficiency of user interaction based on deep learning, the embodiment of the present invention can also screen out some frequencies with a higher echo probability by presetting an echo threshold before determining its step size based on the echo probability.

[0098] For example, in the embodiment of the present invention, before calculating the step length of the frequency in the near-end input signal, the method further includes: removing the frequencies in the near-end input signal whose echo probability is greater than a preset echo threshold.

[0099] The preset echo threshold may be set to 0.5; the preset echo threshold may be set according to actual needs, and the embodiment of the present invention does not impose too many limitations on this.

[0100] After reducing the amount of data processing based on the above steps, the embodiment of the present invention can determine the step length based on the echo probability of the frequency.

[0101] It is understandable that the embodiment of the present invention takes into account the time-varying and complex characteristics of the echo, so even if some echoes are screened out through the above steps, there are still residual echoes, and their characteristics will also change over time. In addition, a larger step size in the LMS algorithm can speed up the convergence speed, but may cause the algorithm to be unstable; a smaller step size can improve stability, but the convergence speed is slower.

[0102] Based on this, the embodiment of the present invention dynamically adjusts the step size of the LMS algorithm according to the echo probability of the frequency, so that the filter can better adapt to the data change of the echo and improve the effect of echo suppression.

[0103] For example, in the embodiment of the present invention, the step length of each frequency is determined, and the specific details can be referred to the following relationship:

[0104] ;

[0105] The near-end input signal The frequency step size, is the preset maximum step size, is a hyperparameter, The near-end input signal The echo probability of a frequency, e is a constant.

[0106] Among them, the hyperparameter can be set to 0.05, and can be set according to actual needs.

[0107] In the above formula, the greater the echo probability of a frequency, the greater the probability that the frequency is an echo. In order to make the echo signal better adapt to data changes in the delay attenuation region, the step size of the echo signal can be increased, thereby improving the convergence speed of the algorithm.

[0108] After the step length of each frequency in the near-end input signal is obtained based on the above steps, the frequency step length can be used in the LMS algorithm to remove the echo signal in the near-end input signal, that is, continue to perform the following steps.

[0109] S6: The frequency step size is used in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction.

[0110] By way of example, in an embodiment of the present invention, a frequency step is used in an LMS algorithm to eliminate echoes in a near-end input signal, including: presetting a weight vector; obtaining a filter output based on the weight vector and the near-end input signal, and iteratively updating the weight vector through an error between a reference signal and the filter output and a frequency step; using the updated weight vector that meets the iteration termination condition as a target weight vector, and eliminating the output of the filter corresponding to the target weight vector as an echo, to obtain a near-end input signal after the echo is eliminated.

[0111] The specific step of using the frequency step size in the LMS algorithm to remove the echo in the near-end input signal can be implemented by the existing technology, and the embodiment of the present invention will not be described in detail here.

[0112] After processing the echo signal in the near-end input signal based on the above steps, user dialogue interaction based on deep learning can be achieved through the near-end input signal that does not contain the echo signal.

[0113] For example, in an embodiment of the present invention, the frequency step size is used in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction, including: obtaining text information of the near-end input signal after the echo is removed, converting the text information into voice information and outputting it to the user.

[0114] Specifically, after the text information is input into the speech recognition system, the near-end input signal can be converted into acoustic features and mapped to text to obtain its text information; after the text information parsing model recognizes the text, it extracts key information to respond to the user and converts it into voice information to output to the user, thereby realizing user dialogue interaction.

[0115] Among them, when obtaining the text information parsing model, a historical text information dataset can be obtained, and a word segmentation tool can be used to divide the text information in the historical text information dataset into independent vocabulary units, and a statistical method can be used to extract key statistical features in the text information; the key statistical features of each text information in the historical text information dataset are manually labeled to obtain a labeled historical text information dataset; the labeled historical text information dataset is trained through a deep neural network to obtain a text information parsing model.

[0116] Based on the above embodiments, user dialogue interaction based on deep learning can be realized.

[0117] It can be seen that in the embodiment of the present invention, when implementing user dialogue interaction based on deep learning, the near-end input signal and output signal of the speech during the interaction process can be obtained, and the far-end input signal is recorded as the reference signal; the delay time corresponding to the maximum value of the cross-correlation function of the amplitude of the reference signal and the output signal at each delay time is recorded as the target delay time ; Through the reference signal The frequency of the near-end input signal The amplitude difference corresponding to the frequency is used to obtain the frequency domain similarity; the frequency window is obtained, and the frequency domain attenuation degree of the reference signal or the near-end input signal at the frequency is obtained by the average of the adjacent amplitude changes of the frequency window in the reference signal or the near-end input signal; the frequency window in the reference signal is obtained. The frequency of the near-end input signal The frequency domain similarity between The product of The frequency of the near-end input signal The absolute value of the frequency domain attenuation degree difference of the frequencies is normalized by the ratio of the product and the absolute value of the difference to obtain the first frequency in the near-end input signal. The probability of echo at each frequency; ; The near-end input signal The frequency step size, is the preset maximum step size, is a hyperparameter, The near-end input signal The echo probability of a frequency, e is a constant; the frequency step size is used in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction.

[0118] In this way, the embodiment of the present invention removes the echo signal in the user dialogue interaction process through the LMS algorithm, and effectively improves the accuracy of the user dialogue interaction based on deep learning. In this process, the embodiment of the present invention takes into account that the LMS algorithm uses a fixed step size, which will affect the accuracy of the algorithm processing. Based on this, the embodiment of the present invention analyzes the possibility of each frequency in the near-end input signal being an echo based on the echo feature, and adjusts the frequency step size based on the possibility of each frequency being an echo, which can effectively improve the accuracy of echo elimination, thereby effectively improving the accuracy of user dialogue interaction based on deep learning.

[0119] An embodiment of the present invention also discloses a user dialogue interaction system based on deep learning, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a user dialogue interaction method based on deep learning provided by the present invention is implemented.

[0120] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, and their configuration and functions are known in the art, so they will not be described in detail here.

[0121] In the present invention, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory RRAM, a dynamic random access memory DRAM, a static random access memory SRAM, an enhanced dynamic random access memory EDRAM, a high bandwidth memory HBM, a hybrid memory cube HMC, etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium may be part of a device or accessible or connectable to a device.

[0122] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.

[0123] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A user dialogue interaction method based on deep learning, characterized in that: include: Acquire the near-end input signal and output signal of the speech during the interaction, and record the far-end input signal as a reference signal; The delay time corresponding to the maximum value of the cross-correlation function of the reference signal and the output signal amplitude at each delay time is recorded as the target delay time ; Through the reference signal The frequency of the near-end input signal The amplitude difference corresponding to the frequency is used to obtain the frequency domain similarity; Obtain a frequency window, and obtain the frequency domain attenuation degree of the reference signal or the near-end input signal at the frequency by averaging the adjacent amplitude changes of the frequency window in the reference signal or the near-end input signal; Get the reference signal The frequency of the near-end input signal The frequency domain similarity between The product of The frequency of the near-end input signal The absolute value of the frequency domain attenuation degree difference of the frequencies is normalized by the ratio of the product and the absolute value of the difference to obtain the value of the near-end input signal. The probability of echo at each frequency; ; The near-end input signal The frequency step size, is the preset maximum step size, is a hyperparameter, The near-end input signal The echo probability of a frequency, e is a constant; the frequency step size is used in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction.

2. The user dialogue interaction method based on deep learning according to claim 1, characterized in that: The acquiring of the near-end input signal and the output signal of the speech in the interaction process and recording the far-end input signal as a reference signal also includes: During the conversation interaction process, the amplitude value at each acquisition moment is obtained for preprocessing to obtain the near-end input signal, output signal and far-end input signal.

3. The user dialogue interaction method based on deep learning according to claim 1, characterized in that: The method for obtaining the cross-correlation function includes: Preset the neighborhood of the reference signal or output signal at each acquisition moment; Obtain the mean of the neighborhood amplitude entropy values ​​of the reference signal and the output signal at the same acquisition time, use the negative of the mean as the exponent of an exponential function with e as the base, and obtain the stable weights of the reference signal and the output signal at the acquisition time; ; The reference signal and the output signal are delayed by The cross-correlation function value under The reference signal and the output signal are The stable weight of each acquisition moment, The reference signal The amplitude value at each acquisition moment, The output signal The amplitude value at the acquisition moment.

4. The user dialogue interaction method based on deep learning according to claim 1, characterized in that: The frequency domain similarity satisfies the relationship: ; The reference signal The frequency of the near-end input signal The frequency domain similarity between the frequencies, is the target delay time, , are the reference signals frequency, the near-end input signal The amplitude value corresponding to the frequency is is an exponential function with base e, is an absolute value.

5. The user dialogue interaction method based on deep learning according to claim 1, characterized in that: The frequency step size in the near-end input signal is calculated by: The frequencies in the near-end input signal whose echo probability is greater than a preset echo threshold are removed.

6. The user dialogue interaction method based on deep learning according to claim 1, characterized in that: The method of using the frequency step size in the LMS algorithm to remove the echo in the near-end input signal includes: A weight vector is preset; a filter output is obtained based on the weight vector and the near-end input signal, and the weight vector is iteratively updated by the error between the reference signal and the filter output and the frequency step size; the weight vector that meets the iteration termination condition after the update is used as the target weight vector, and the output of the filter corresponding to the target weight vector is removed as the echo to obtain the near-end input signal after the echo is removed.

7. The user dialogue interaction method based on deep learning according to claim 6 is characterized in that: The step size of using the frequency in the LMS algorithm to remove the echo in the near-end input signal to achieve user dialogue interaction includes: The text information of the near-end input signal after the echo is removed is obtained, the text information is converted into voice information, and then output to the user.

8. A user dialogue interaction system based on deep learning, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the user dialogue interaction method based on deep learning according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Dialogue interaction method, dialogue interaction device, equipment and storage medium

    CN116796758A

  • Method and system for processing subband signals using adaptive filters

    CA2437477A1

  • Echo eliminator and echo cancellation method

    CN101043560A