Speech recognition method based on natural speech processing

By segmenting, decomposing and modal feature analysis of the speech signals of AI speech robots, combined with periodic approximation signals, the problem of echo interference is solved, and more efficient and accurate speech recognition is achieved.

CN119068866BActive Publication Date: 2025-05-16SHENZHEN CHENGXUNDE COMM SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411015354.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-26
Publication Date
2025-05-16
Estimated Expiration
2044-07-26

AI Technical Summary

Technical Problem

The speech recognition of AI voice robots in closed rooms is interfered with echo, affecting the accuracy of the recognition.

Method used

Through the method based on natural speech processing, the speech signal is segmented, decomposed, modal feature analysis and periodic approximation signal construction to accurately suppress echo signals and improve the accuracy of speech recognition.

Benefits of technology

It effectively reduces noise and interference in voice signals, improves the accuracy and efficiency of voice recognition, and enhances the real-time communication ability between robots and users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068866B_ABST
    Figure CN119068866B_ABST
Patent Text Reader

Abstract

The present application relates to the field of intelligent speech, and specifically to a speech recognition method based on natural speech processing, which includes: segmenting the call speech data, and then decomposing it to obtain several layers of modal components of each segment of speech signal; detecting the peak difference and periodicity of the speech signal in the modal component through the voice characteristics of the robot, and then identifying the robot speech signal through the similarity of the modal components in multiple segments of speech; segmenting the speech signal through periodicity, constructing a periodic approximate signal through the similar characteristics between the signals, obtaining the degree of echo attenuation, improving the learning rate factor of the filter, and obtaining the recognition result of the robot speech segment. The present application aims to accurately suppress the echo signal in the robot speech signal and improve the accuracy of speech recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent speech, and in particular to a speech recognition method based on natural speech processing. Background Art

[0002] As a product of communication technology and artificial intelligence technology, artificial intelligence (AI) voice robots combine advanced speech recognition and natural language processing technologies and have the ability to understand and respond autonomously. They can automatically perform outbound phone calls and answering tasks, and effectively communicate and interact with users in real time. The development of this technology not only improves the efficiency of corporate customer service, but also provides consumers with a more convenient and personalized service experience.

[0003] When AI voice robots have conversations with users, they usually rely on voice recognition technology to understand the content of the call. However, if the robot is located in a quiet and closed room, this environment may cause echoes from the walls, which interferes with the clarity of the voice signal and affects the accuracy of voice recognition. Summary of the invention

[0004] In view of the above, it is necessary to provide a speech recognition method based on natural speech processing to accurately suppress the echo signal in the robot voice signal and improve the accuracy of speech recognition.

[0005] One embodiment of the present application provides a speech recognition method based on natural speech processing, the method comprising:

[0006] Each call voice data is segmented to obtain several segments of voice signals; each segment of voice signals is decomposed to obtain several layers of modal components of each segment of voice signals;

[0007] Based on the fluctuation amplitude and fluctuation duration of all modal components of each speech signal, determine whether the speech signal is a robot speech segment;

[0008] Based on the difference distribution between modal components in different robot voice segments, the modal components are selected as robot voice signals;

[0009] Segment each robot voice signal to obtain several sub-signals; based on the time domain feature differences and frequency domain feature differences between different sub-signals, filter the periodic approximate signal of each sub-signal;

[0010] The variation range of all sub-signals of all robot voice signals of the robot voice segment and the periodic approximate signal are integrated, and each robot voice signal is filtered, reconstructed, and recognized to obtain the recognition result of the robot voice segment.

[0011] The specific steps of determining whether the voice signal is a robot voice segment include:

[0012] Based on the shape characteristics and numerical distribution characteristics of each layer of modal components of each speech signal, the peak consistency confidence of each layer of modal components is obtained;

[0013] The maximum value of the peak consistency confidence of all modal components of each speech signal is obtained, and the maximum value of all speech signals is threshold segmented to obtain an optimal threshold, and the speech signal with the maximum value greater than the optimal threshold is used as the robot speech segment.

[0014] The peak consistency confidence of each layer modal component is obtained as follows:

[0015] According to the discrete degree of the peak value of each layer of modal components of each speech signal, the peak difference of each layer of modal components is obtained; the time intervals between all adjacent peak values ​​of each layer of modal components are calculated, and the time intervals are classified to obtain the first feature category; according to the discrete degree and the number of elements in the first feature category, the discrete significance value of each layer of modal components is obtained; the inverse proportional mapping result of the fusion result of the peak difference and the discrete significance value is used as the peak consistency confidence of each layer of modal components.

[0016] The screening modal component is used as the robot voice signal, specifically:

[0017] Based on the numerical difference between any two layers of modal components belonging to different robot voice segments, combined with the peak consistency, the modal difference between the any two layers of modal components is obtained;

[0018] Based on the discrete degree of all the modal differences of any modal component of each robot voice segment and in combination with the peak consistency confidence, the synthetic voice confidence of any modal component of each robot voice segment is obtained;

[0019] The modal component with the largest confidence of the synthesized sound in each robot voice segment is used as the robot voice signal.

[0020] The step of obtaining the modal difference between any two layers of modal components includes:

[0021] The modal components of each layer of each speech signal are sampled point by point to form a corresponding speech data sequence;

[0022] For any two layers of modal components belonging to different robot voice segments, the voice data sequences corresponding to the modal components, the peak consistency confidence, and the difference between the peak-to-peak values ​​are fused to obtain the modal difference of the any two layers of modal components.

[0023] The step of screening the periodic approximate signal of each sub-signal is specifically as follows:

[0024] For each robot voice signal, according to the time domain distribution difference between each sub-signal and other sub-signals, other sub-signals are screened to obtain the approximate signal of each sub-signal;

[0025] For each robot voice signal, the difference between all peak values ​​of any sub-signal and any approximate signal thereof is obtained and recorded as the signal frequency difference; all the signal frequency differences of any sub-signal are threshold segmented to obtain a difference threshold, and all approximate signals whose signal frequency difference is less than the difference threshold are taken as the periodic approximate signals of any sub-signal.

[0026] The obtaining of the approximate signal of each sub-signal is specifically:

[0027] The maximum difference value of the corresponding time intervals between any two sub-signals of each robot voice signal is used as the period threshold of each robot voice signal; the sub-signals of all other robot voice signals whose time interval difference with any sub-signal of each robot voice signal is less than the corresponding maximum difference value are used as the approximate signals of any sub-signal of the corresponding robot voice signal.

[0028] The method of obtaining the recognition result of the robot voice segment is as follows:

[0029] Based on the distribution of each sub-signal and all periodic approximate signals in each robot voice signal, the echo attenuation degree of each sub-signal is obtained;

[0030] Based on the echo attenuation degree of each sub-signal, a dynamic learning rate of each sub-signal is obtained;

[0031] Based on the dynamic learning rate, all periodic approximate signals of each robot voice signal are filtered and reconstructed to obtain each pure robot voice signal; natural language processing technology is used for all pure robot voice signals of the robot voice segment to obtain the recognition result of the robot voice segment.

[0032] The step of obtaining the echo attenuation degree of each sub-signal includes:

[0033] The average peak-to-peak value of all sub-signals that are periodic approximate signals in each robot voice signal is obtained as the denominator of the echo attenuation degree of each sub-signal; the peak-to-peak value of each sub-signal is used as the numerator of the echo attenuation degree of the corresponding sub-signal.

[0034] The dynamic learning rate of each sub-signal is obtained as follows:

[0035] When a sub-signal is a periodic approximation signal of other sub-signals, the product of the initial learning rate and the echo attenuation degree is used as the dynamic learning rate of the corresponding sub-signal; otherwise, the initial learning rate is used as the dynamic learning rate of the corresponding sub-signal.

[0036] This application has at least the following beneficial effects:

[0037] The present application first obtains voice call data and segments it, dividing long segments of speech into smaller voice segments, which helps to reduce noise and interference in voice signals; decomposition can more accurately analyze each segment and reduce the possibility of misrecognition; a preliminary evaluation of the voice segments is performed by calculating the change characteristics of the fluctuation amplitude and fluctuation duration in the modal component; considering that the time interval between adjacent response sentences of the robot is large, the periodicity of the voice segments is detected to make the acquired robot voice signal more accurate; since the robot voice signal has high consistency in multiple voice data, the robot voice signal is verified; through the voice features of the robot voice signal and its echo, a periodic approximate signal is constructed to distinguish between normal signals and echo signals, and the echo signal in the robot voice signal is accurately suppressed based on the periodic approximate signal to improve the subsequent voice recognition effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 A flow chart of a speech recognition method based on natural speech processing provided in this application;

[0039] Figure 2 Visualization of call voice data provided for this application;

[0040] Figure 3 Schematic diagram of the voice signal provided for this application;

[0041] Figure 4 Schematic diagram of the modal decomposition results provided for this application. DETAILED DESCRIPTION

[0042] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example" and the like are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "or", "for example" and the like is intended to present related concepts in a concrete manner.

[0043] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the present application. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.

[0044] It should also be noted that the terms "first" and "second" in this application and its drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. The method disclosed in the embodiments of the present application or the method shown in the flow chart includes one or more steps for implementing the method. Without departing from the scope of protection of this application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.

[0045] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0046] The present application embodiment first proposes a speech recognition method based on natural speech processing and is applied to the field of intelligent speech. Figure 1 , the method comprises the following steps:

[0047] S1: Segment each call voice data to obtain several voice signal segments; decompose each voice signal segment to obtain several layers of modal components of each voice signal segment.

[0048] Through the outbound call data collection module in the AI ​​intelligent answering system, the voice call data when using the AI ​​voice robot for outbound calls is collected, and all the collected call voice data is transmitted to the call voice database.

[0049] In some embodiments of the present application, the wth call voice data in the call voice database is taken as an example for analysis, wherein the call voice data visualization diagram is as follows: Figure 2 As shown, the horizontal axis is the call time in seconds; the vertical axis is the normalized amplitude value.

[0050] Furthermore, the wth call voice data is segmented into multiple voice signals by using the VAD (Voice Activity Detection) algorithm, wherein the voice signal schematic diagram is as shown in FIG. Figure 3 As shown, the horizontal axis is time, in seconds; the vertical axis is the normalized amplitude value.

[0051] Taking the rth segment of speech signal after segmentation as an example, the variational mode decomposition algorithm is used to decompose the rth segment of speech signal. During the processing, the regularization parameter is set to 0.1 and the number of decomposition layers is set to 7. The schematic diagram of the mode decomposition result is as follows: Figure 4 As shown, the horizontal axis is time, in seconds; the vertical axis is the normalized amplitude value of each mode.

[0052] It should be noted that the implementer can select the most appropriate speech segmentation algorithm, regularization parameter and decomposition layer number according to the collected call speech data. The VAD algorithm and the variational mode decomposition algorithm are well-known technologies and will not be described here.

[0053] S2: Based on the fluctuation amplitude and fluctuation duration of all modal components of each speech signal, determine whether the speech signal is a robot speech segment.

[0054] Considering that users may interrupt the AI ​​voice robot during a call, it is necessary to further process the voice data to identify the robot voice, so that the VAD algorithm can segment sentences more accurately and improve the accuracy of subsequent semantic recognition and user identity classification. Since speech synthesis technology usually controls the speed, volume and speaking style to maintain a certain degree of stability and naturalness, the volume, speed and speaking style of the AI ​​voice robot remain unchanged, so the peak and peak period of its sound in the time domain diagram are highly consistent.

[0055] According to the discrete degree of the peak value of each layer of modal components of each speech signal, the peak difference of each layer of modal components is obtained; the time intervals between all adjacent peak values ​​of each layer of modal components are calculated, and the time intervals are classified to obtain the first feature category; according to the discrete degree and the number of elements in the first feature category, the discrete significance value of each layer of modal components is obtained; the inverse proportional mapping result of the fusion result of the peak difference and the discrete significance value is used as the peak consistency confidence of each layer of modal components.

[0056] It should be noted that the degree of dispersion can be obtained by methods such as standard deviation, variance, and coefficient of variation; the classification method can be implemented by threshold segmentation or clustering algorithm;

[0057] Specifically, since the processing method for each layer of modal components of each voice signal is the same, some embodiments of the present application take the g-th layer modal component in the r-th voice as an example for analysis, and record the standard deviation of all maximum values ​​in the g-th layer modal component as the peak difference of the g-th layer modal component. The greater the peak difference, the greater the peak difference of the maximum peak, and the less likely the voice signal is a robot signal.

[0058] Furthermore, since the volume and speech speed of the robot change very little, the time intervals between adjacent maxima are almost the same. However, considering that when the robot is speaking, the pause marks such as commas in the sentence will cause a brief interruption in the robot's voice. Since the interruption time is very short, the VAD algorithm may not divide the sentence into two segments. Therefore, the segmented voice signal may contain pauses and the like, and the time intervals between adjacent maxima may also be large. If the fluctuation degree of all time intervals is directly calculated, it will affect the subsequent detection of the robot voice signal. Therefore, the time intervals between all adjacent maximum peaks of the g-th layer modal component are calculated, and all time interval values ​​are used as input. The K-means algorithm is used to cluster them, and the number of clusters is obtained by the elbow rule. Among them, the K-means algorithm and the elbow rule are well-known technologies and will not be repeated here. The cluster that satisfies the minimum standard deviation of all elements in the cluster and the number of elements in the cluster is greater than or equal to 2 is taken as the first feature category of the g-th layer modal component. The ratio of the standard deviation corresponding to the first feature category of the g-th layer modal component to the number of elements is taken as the discrete significant value of the g-th layer modal component.

[0059] It should be understood that the smaller the discrete degree of the element value in the first characteristic category of the g-th layer modal component, the more consistent the time interval distance of each maximum peak; the more elements in the cluster, the more times the time interval appears. At this time, the smaller the discrete significance value of the g-th layer modal component, the more consistent the time intervals between adjacent maximum peaks of the speech signal in the g-th layer modal component, and the more consistent the number of times, the more stable the periodicity of the amplitude, and the greater the possibility of being a robot voice signal. Let Z = 1 / (A×B+τ); where Z is the peak consistency confidence of the g-th layer modal component; A is the peak difference of the g-th layer modal component, B is the discrete significance value of the g-th layer modal component, and τ is a preset parameter adjustment factor greater than zero. In order to avoid the denominator being 0, the value is 0.01. The larger the value of the peak consistency confidence, the greater the possibility that the speech signal in the g-th layer modal component is a robot voice.

[0060] Based on this, all speech signals are screened according to the distribution of peak consistency confidence of all modal components of each speech signal to obtain the robot speech segment:

[0061] The maximum value of the peak consistency confidence of all modal components of each speech signal is obtained, and the maximum value of all speech signals is threshold segmented to obtain an optimal threshold, and the speech signal with the maximum value greater than the optimal threshold is used as the robot speech segment.

[0062] It should be noted that there are many methods for obtaining the optimal threshold, such as the Otsu threshold method, the fixed threshold method, etc. The implementer can select a suitable threshold segmentation method according to the actual situation to obtain the robot voice segment. Some embodiments of the present application use the Otsu threshold method to obtain the optimal threshold.

[0063] S3: Based on the difference distribution between the modal components in different robot voice segments, the modal components are screened as robot voice signals.

[0064] Considering that the robot's voice is a synthesized sound, its volume and speaking speed have obvious consistency characteristics. Therefore, in the robot's speech segment, the speech signals between different modal components are similar.

[0065] The modal components of each layer of each speech signal are sampled point by point to form a corresponding speech data sequence; for any two layers of modal components belonging to different robot speech segments, the modal components corresponding to the speech data sequences, peak consistency confidence, and peak-to-peak value differences are fused to obtain the modal difference of the any two layers of modal components.

[0066] It should be noted that the differences between sequences can be measured by the Euclidean distance, Manhattan distance, DTW distance, etc. between the sequences; the differences between data can be measured by difference, absolute value of difference, ratio, etc.; the fusion between variables can be achieved through multiplication, addition, and mixed operations such as addition and multiplication; implementers can choose at their own discretion, and this application does not impose any restrictions on this.

[0067] Specifically, let C = (D + E) × F. Where C is the modal difference between the modal component of the a-th layer of the ith robot voice segment and the modal component of the b-th layer of the j-th robot voice segment; D is the normalized value of the DTW distance between the voice data sequences of the modal component of the a-th layer of the ith robot voice segment and the modal component of the b-th layer of the j-th robot voice segment; E and F are the normalized values ​​of the absolute value of the difference between the peak-to-peak value and the peak consistency confidence of the modal component of the a-th layer of the ith robot voice segment and the modal component of the b-th layer of the j-th robot voice segment, respectively. The smaller the modal difference value, the smaller the difference in the voice signals in the two layers of modal components. The normalization method selects the maximum and minimum normalization method.

[0068] Based on the discrete degree of all the modal differences of any modal component of each robot voice segment and combined with the peak consistency confidence, the synthetic sound confidence of any modal component of each robot voice segment is obtained.

[0069] The modal component with the largest confidence of the synthesized sound in each robot voice segment is used as the robot voice signal.

[0070] In some embodiments of the present application, the modal component in each robot voice segment with the smallest modal difference from the a-layer modal component of the ith robot voice segment is used as the similar modal component of the a-layer modal component of the ith robot voice segment, and the standard deviation of the modal differences corresponding to all similar modal components of the a-layer modal component of the ith robot voice segment is calculated, and the standard deviation is divided by the peak consistency confidence of the a-layer modal component to obtain the synthetic sound confidence of the a-layer modal component of the ith robot voice segment.

[0071] It should be understood that the smaller the discreteness, the more similar the speech signals between the modal component of the a-th layer of the ith robot voice segment and its similar modal component, that is, the greater the confidence of the synthetic sound, the greater the possibility that the modal component of the a-th layer of the ith robot voice segment is a robot signal. The echo signal is identified and suppressed through the robot language signal.

[0072] S4: Segment each robot voice signal to obtain a number of sub-signals; based on the time domain feature differences and frequency domain feature differences between different sub-signals, filter the periodic approximate signal of each sub-signal.

[0073] When the AI ​​voice robot talks to the user in a closed room, the walls of the room will reflect the robot's voice signal, causing an echo. The robot's voice is similar to the echo, which may also affect the segmentation of the VAD algorithm and the subsequent voice recognition effect. Since the robot does not move and the distance between the robot and the wall remains unchanged, the time delay of the robot's voice signal reflected by the wall is relatively consistent, that is, the delay between the robot's voice signal and the echo signal is relatively consistent. At the same time, since the robot's voice does not produce changing characteristics, the echo signal after energy attenuation after reflection from the wall has a high consistency with the robot's voice signal. Therefore, the echo voice signal can be regarded as a replica of the robot signal that is delayed in time and reduced in amplitude, and the degree of time delay and amplitude change are relatively consistent.

[0074] After variational modal decomposition, the robot's echo signal may not be limited to a single modal component, but distributed in multiple modal components. However, despite the wide distribution, the echo signals in different modal components still have a certain degree of similarity with the robot's original voice signal in time domain characteristics and energy distribution.

[0075] The point corresponding to the peak is used as the segmentation point to segment each robot voice signal to obtain multiple sub-signals; the maximum difference value of the corresponding time intervals between any two sub-signals of each robot voice signal is used as the period threshold of each robot voice signal; the sub-signals of all other robot voice signals whose time interval difference with any sub-signal of each robot voice signal is less than the corresponding maximum difference value are used as approximate signals of any sub-signal of the corresponding robot voice signal.

[0076] It should be understood that a robot voice signal is divided into multiple sub-signals, that is, one modal component corresponds to multiple sub-signals, and the lengths of the sub-signals may not be consistent. The sound amplitude period in the robot voice signal remains consistent in its echo signal, which means that the change characteristics of the sound are preserved in the echo.

[0077] Specifically, some embodiments of the present application calculate the absolute value of the difference in the time lengths of all pairwise combinations of sub-signals in the f-th robot voice signal, and record the maximum value of the absolute value of the difference as the period threshold of the f-th robot voice signal; calculate the absolute value of the difference in time intervals between the u-th sub-signal of the f-th robot voice signal and all sub-signals of other robot voice signals, and record the sub-signal whose absolute value of the difference is less than or equal to the corresponding period threshold as the approximate signal of the u-th sub-signal of the f-th robot voice signal.

[0078] Since the voice of the AI ​​voice robot is a synthetic sound, the frequency of its echo signal is relatively consistent with the frequency of the robot's voice signal.

[0079] For each robot voice signal, according to the difference in frequency domain distribution between any sub-signal and all its approximate signals, the approximate signal of any sub-signal is screened to obtain a periodic approximate signal of any sub-signal.

[0080] Specifically, some embodiments of the present application use Fourier transform to obtain the spectrum of the two sub-signals with the u-th sub-signal and its v-th approximate signal; obtain the significant peaks in the two spectrum graphs through the peak detection algorithm, and use the sequence composed of all peaks as the peak frequency sequence of the two sub-signals in the order of frequency from small to large; record the DTW distance of the two peak frequency sequences as the signal frequency difference, the smaller the signal frequency difference, the more consistent the v-th approximate signal of the u-th sub-signal is with it in frequency distribution, and the greater the possibility that the v-th approximate signal is the echo signal of the u-th sub-signal. The signal frequency difference between the u-th sub-signal and all its approximate signals is used as the input of the Otsu threshold algorithm, and the difference threshold is output; all approximate signals with a signal frequency difference less than the difference threshold are used as the periodic approximate signal of the u-th sub-signal. At this time, since the periodicity and frequency difference between the periodic approximate signal and the sub-signal are both small, the possibility that the periodic approximate signal belongs to the echo signal is greater.

[0081] S5: integrating the variation range of all sub-signals of all robot voice signals of the robot voice segment and the periodic approximate signal, filtering, reconstructing, and recognizing each robot voice signal, and obtaining the recognition result of the robot voice segment.

[0082] By obtaining a dynamic learning rate factor, the robot signal is processed:

[0083] The average level of the peak-to-peak value of all sub-signals that are periodic approximate signals in each robot voice signal is obtained as the denominator of the echo attenuation degree of each sub-signal; the peak-to-peak value of each sub-signal is used as the numerator of the echo attenuation degree of each sub-signal; when the sub-signal is a periodic approximate signal of other sub-signals, the product of the initial learning rate and the echo attenuation degree is used as the dynamic learning rate of the corresponding sub-signal; otherwise, the initial learning rate is used as the dynamic learning rate of the corresponding sub-signal; based on the dynamic learning rate, all periodic approximate signals of each robot voice signal are filtered and reconstructed to obtain each pure robot voice signal; natural language processing technology is used for all pure robot voice signals of the robot voice segment to obtain the recognition result of the robot voice segment.

[0084] In some embodiments of the present application, the ratio of the peak-to-peak value of each sub-signal in the f-th robot voice signal to the average of the peak-to-peak values ​​of all sub-signals that are periodic approximate signals is recorded as the echo attenuation degree of each sub-signal. The greater the echo attenuation degree, the smaller the signal energy of the echo signal. The learning rate of the adaptive filter (Normalized Least Mean Squares, NLMS) is adjusted according to the echo attenuation degree:

[0085]

[0086] Wherein, G is the dynamic learning rate of the sub-signal, α is the initial learning rate, and β is the echo attenuation degree of the sub-signal. After the adjustment, the filter is more flexible in processing echo signals. If the filtered sub-signal is an echo signal and its echo attenuation degree is large and the energy is weak, by increasing the learning rate (i.e., multiplying by the larger echo attenuation degree), the filter can adapt and learn the echo characteristics faster, and then suppress these echoes more quickly and accurately. The periodic approximate signal under each robot voice signal is filtered by the filter, and the filtering results of all robot voice signals are reconstructed to obtain a robot voice signal with echo removed, which is recorded as a pure robot voice signal. It should be noted that signal reconstruction can be achieved by Fourier transform, which is a well-known technology and will not be repeated here.

[0087] The robot voice signal after all echoes are removed from the robot voice segment is used to obtain the recognition text result of the robot voice segment. It should be noted that the Speech-to-Text API is a well-known technology for speech recognition and will not be described in detail here; the implementer may also choose other technologies for speech recognition, and this application does not limit this.

[0088] The embodiment of the present application provides a speech recognition method based on natural speech processing, the method comprising: first acquiring voice call data and segmenting it, dividing long speech segments into smaller speech segments, which helps to reduce noise and interference in the speech signal; performing decomposition, which can more accurately analyze each segment and reduce the possibility of misrecognition; performing a preliminary evaluation of the speech segments by calculating the change characteristics of the fluctuation amplitude and fluctuation duration in the modal component; taking into account the large time interval between adjacent response sentences of the robot, detecting the periodicity of the speech segments, so that the acquired robot speech signal is more accurate; because the robot speech signal has a high consistency in multiple speech data, verifying the robot speech signal; constructing a periodic approximation signal through the speech features of the robot speech signal and its echo to distinguish between normal signals and echo signals, and accurately suppressing the echo signal in the robot speech signal based on the periodic approximation signal to improve the subsequent speech recognition effect.

[0089] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two continuous operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0090] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A speech recognition method based on natural speech processing, characterized in that: The method comprises the following steps: Each call voice data is segmented to obtain several segments of voice signals; each segment of voice signals is decomposed to obtain several layers of modal components of each segment of voice signals; Based on the fluctuation amplitude and fluctuation duration of all modal components of each speech signal, determine whether the speech signal is a robot speech segment; Based on the difference distribution between modal components in different robot voice segments, the modal components are selected as robot voice signals; Segment each robot voice signal to obtain several sub-signals; based on the time domain feature differences and frequency domain feature differences between different sub-signals, filter the periodic approximate signal of each sub-signal; The average peak-to-peak value of all sub-signals that are periodic approximate signals in each robot voice signal is obtained as the denominator of the echo attenuation degree of each sub-signal; the peak-to-peak value of each sub-signal is used as the numerator of the echo attenuation degree of the corresponding sub-signal; When the sub-signal is a periodic approximation signal of other sub-signals, the product of the initial learning rate and the echo attenuation degree is used as the dynamic learning rate of the corresponding sub-signal; otherwise, the initial learning rate is used as the dynamic learning rate of the corresponding sub-signal; Based on the dynamic learning rate, all periodic approximate signals of each robot voice signal are filtered and reconstructed to obtain each pure robot voice signal; natural language processing technology is used for all pure robot voice signals of the robot voice segment to obtain the recognition result of the robot voice segment.

2. The speech recognition method based on natural speech processing according to claim 1, characterized in that: The specific steps of determining whether the voice signal is a robot voice segment include: Based on the shape characteristics and numerical distribution characteristics of each layer of modal components of each speech signal, the peak consistency confidence of each layer of modal components is obtained; The maximum value of the peak consistency confidence of all modal components of each speech signal is obtained, and the maximum value of all speech signals is threshold segmented to obtain an optimal threshold, and the speech signal with the maximum value greater than the optimal threshold is used as the robot speech segment.

3. The speech recognition method based on natural speech processing as claimed in claim 2, characterized in that: The peak consistency confidence of each layer modal component is obtained as follows: According to the discrete degree of the peak value of each layer of modal components of each speech signal, the peak value difference of each layer of modal components is obtained; the time interval between all adjacent peak values ​​of each layer of modal components is calculated, and the time interval is classified to obtain the first feature category; According to the discrete degree and number of elements in the first feature category, the discrete significant value of each layer of modal components is obtained; the inverse proportional mapping result of the fusion result of the peak difference and the discrete significant value is used as the peak consistency confidence of each layer of modal components.

4. The speech recognition method based on natural speech processing as claimed in claim 3, characterized in that: The screening modal component is used as the robot voice signal, specifically: Based on the numerical difference between any two layers of modal components belonging to different robot voice segments, combined with the peak consistency, the modal difference between the any two layers of modal components is obtained; Based on the discrete degree of all the modal differences of any modal component of each robot voice segment and in combination with the peak consistency confidence, the synthetic voice confidence of any modal component of each robot voice segment is obtained; The modal component with the largest confidence of the synthesized sound in each robot voice segment is used as the robot voice signal.

5. The speech recognition method based on natural speech processing as claimed in claim 4, characterized in that: The step of obtaining the modal difference between any two layers of modal components comprises: The modal components of each layer of each speech signal are sampled point by point to form a corresponding speech data sequence; For any two layers of modal components belonging to different robot voice segments, the voice data sequences corresponding to the modal components, the peak consistency confidence, and the difference between the peak-to-peak values ​​are fused to obtain the modal difference of the any two layers of modal components.

6. The method for speech recognition based on natural speech processing according to claim 1, characterized in that: The screening of the periodic approximate signal of each sub-signal is specifically as follows: For each robot voice signal, according to the time domain distribution difference between each sub-signal and other sub-signals, other sub-signals are screened to obtain the approximate signal of each sub-signal; For each robot voice signal, the difference between all peak values ​​of any sub-signal and any approximate signal thereof is obtained and recorded as the signal frequency difference; all the signal frequency differences of any sub-signal are threshold segmented to obtain a difference threshold, and all approximate signals whose signal frequency difference is less than the difference threshold are taken as the periodic approximate signals of any sub-signal.

7. The method for speech recognition based on natural speech processing according to claim 6, characterized in that: The approximate signals of each sub-signal are obtained as follows: The maximum difference value of the corresponding time intervals between any two sub-signals of each robot voice signal is used as the period threshold of each robot voice signal; the sub-signals of all other robot voice signals whose time interval difference with any sub-signal of each robot voice signal is less than the corresponding maximum difference value are used as the approximate signals of any sub-signal of the corresponding robot voice signal.

Citation Information

Patent Citations

  • Signal enhancement processing method of bone conduction earphone

    CN117059120A

  • Intelligent data analysis method for barrier-free Bluetooth controller

    CN117456983A