Voice processing method based on artificial intelligence

By performing detailed division of the speech signal into stable and disturbed segments and analyzing the feature transfer ratio, combined with a multi-channel BP neural network, the problem of low speech recognition accuracy in existing technologies is solved, achieving higher recognition accuracy.

CN120748414AActive Publication Date: 2025-10-03SICHUAN BOCHUANGHUI FRONTIER TECH CO LTD +1
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
CN202511200071.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-03
Estimated Expiration
2045-08-26

AI Technical Summary

Technical Problem

Existing speech recognition technology roughly segments speech signals during the feature extraction and matching stages, fails to finely divide stable and disturbed segments, and ignores the feature migration rules between adjacent speech segments, resulting in low recognition accuracy. It is particularly susceptible to noise interference in complex environments or short speech segments.

Method used

The speech signal is divided into multiple short time frames. The stable segments and disturbance segments are marked by the zero-crossing rate. The fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio are extracted. The migration ratio feature matrix is ​​constructed and a multi-channel BP neural network is used to process these features for speaker identification.

Benefits of technology

By analyzing the feature migration rules of adjacent speech segments, the accuracy of speech recognition is improved, the differentiated information in different speech segment combinations is fully utilized, and the recognition ability in complex environments is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748414A_ABST
    Figure CN120748414A_ABST
Patent Text Reader

Abstract

The invention discloses a voice processing method based on artificial intelligence, and belongs to the technical field of voice recognition. The method comprises the following steps: firstly, dividing a voice signal into a plurality of short-time frames, and marking a stable section and a disturbance section through a zero-crossing rate; then extracting four combinations of a stationary section-stationary section, a stationary section-disturbance section, a disturbance section-stationary section and a disturbance section-disturbance section; for each combination, extracting a fundamental frequency migration ratio, a harmonic gravity center migration ratio, a spectrum bandwidth migration ratio and a power spectrum migration ratio, and constructing a corresponding migration ratio characteristic matrix; calculating similarity coefficients of each column of features and storage features in each combined migration ratio feature matrix, and averaging the similarity coefficients to obtain feature matching vectors; and finally, processing the migration ratio feature matrixes and the feature matching vectors of the four combinations by using a multi-channel-BP neural network to realize speaker identity recognition. According to the invention, through refined feature extraction and combination processing, the speech recognition precision is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of speech recognition technology, and in particular to a speech processing method based on artificial intelligence. Background Art

[0002] In the field of speech processing, speaker recognition technology, as a core support for scenarios such as human-computer interaction, identity authentication, and information security, has always been a research hotspot. Its core goal is to accurately identify the speaker by analyzing the personalized acoustic features contained in the speech signal.

[0003] In existing technologies, the feature extraction process of speaker recognition methods often revolves around the basic acoustic properties of speech signals. For example, fundamental frequency features are used to capture the pitch of speech, Mel-frequency cepstral coefficients (MFCCs) simulate the auditory characteristics of the human ear to extract spectral envelope features, and linear prediction coefficients (LPCs) describe the linear prediction characteristics of speech signals based on a vocal tract model. These features are extracted based on the static properties of consecutive frames. For dynamic feature processing, common methods include performing differential operations on static features to obtain temporal variation information, or modeling the probability distribution of feature sequences using hidden Markov models (HMMs) and Gaussian mixture models (GMMs) to reflect the dynamic evolution of speech.

[0004] However, existing technologies have obvious limitations: First, most methods segment speech signals in a relatively rough manner, and do not make a detailed division based on the essential differences between stable segments (such as vowels) and disturbed segments (such as consonants) that are commonly found in speech. As a result, feature extraction lacks specificity for the characteristics of different speech segments. Second, feature analysis mostly focuses on the acoustic properties within a single speech segment, and pays insufficient attention to the feature migration patterns between adjacent speech segments (such as stable segments and disturbed segments, disturbed segments and stable segments, etc.). This migration pattern is closely related to the speaker's pronunciation habits, vocal cords, and physical characteristics of the vocal tract, and is important personalized distinguishing information. In addition, in the feature matching stage, existing methods often rely on similarity metrics of single-type features, and have weak collaborative analysis capabilities for features in multiple combination scenarios. It is difficult to fully utilize the differentiated information contained in different speech segment combinations. In complex environments or short speech fragments, they are easily affected by noise interference or insufficient feature information, resulting in low recognition accuracy. Summary of the Invention

[0005] In view of the above-mentioned deficiencies in the prior art, the present invention provides an artificial intelligence-based speech processing method that solves the problem of low speech recognition accuracy in the prior art.

[0006] In order to achieve the above-mentioned object of the invention, the technical solution adopted by the present invention is: a speech processing method based on artificial intelligence, comprising the following steps: The speech signal is divided into multiple short time frames, and the short time frames are marked based on the zero crossing rate to obtain the stable segment and the disturbance segment; Extract stationary segment-stationary segment combinations, stationary segment-disturbance segment combinations, disturbance segment-stationary segment combinations, and disturbance segment-disturbance segment combinations; For each combination, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio are extracted to construct the migration ratio feature matrix of the corresponding combination; Calculate the similarity coefficient between each column feature and the stored feature in the migration ratio feature matrix of each combination, and take the average of the similarity coefficients of each column to obtain the feature matching coefficient corresponding to each column to form a feature matching vector; A multi-channel BP neural network is used to process the transfer ratio feature matrix and feature matching vector of the four combinations to obtain the speaker identification result.

[0007] Furthermore, the process of obtaining the stable segment and the disturbed segment includes: Divide the speech signal into multiple short time frames; Count the zero-crossing rate of each short-time frame and take the average of the zero-crossing rates of each short-time frame; The short time frames with zero-crossing rate greater than the mean are marked as disturbance segments; The short time frames with zero crossing rate less than or equal to the mean are marked as stationary segments.

[0008] Furthermore, the process of constructing the migration ratio feature matrix of the corresponding combination includes: Perform Fourier transform on each signal segment in each combination to obtain the spectrum of each signal segment in the combination; In a combination, according to the spectrum of each signal segment, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio of each combination are obtained; The fundamental frequency migration ratios of the same combination at each moment constitute a fundamental frequency migration ratio vector; The harmonic center of gravity migration ratio of the same combination at each moment is used to form a harmonic center of gravity migration ratio vector; The spectrum bandwidth migration ratios of the same combination at each moment constitute a spectrum bandwidth migration ratio vector; The power spectrum migration ratios of the same combination at each moment constitute a power spectrum migration ratio vector; The fundamental frequency migration ratio vector, the harmonic center of gravity migration ratio vector, the spectrum bandwidth migration ratio vector and the power spectrum migration ratio vector are used as row vectors to form a migration ratio characteristic matrix of the combination.

[0009] Furthermore, the process of obtaining the fundamental frequency shift ratio includes: in a combination, taking the frequency corresponding to the maximum amplitude in the spectrum of each segment signal as the fundamental frequency, and taking the ratio of the fundamental frequency of the subsequent segment signal to the fundamental frequency of the previous segment signal as the fundamental frequency shift ratio; The process of obtaining the harmonic gravity center migration ratio includes: extracting each harmonic according to the fundamental frequency in the spectrum of each signal segment in a combination, calculating the harmonic gravity center of each harmonic, and taking the ratio of the harmonic gravity center of the subsequent signal segment to the harmonic gravity center of the previous signal segment as the harmonic gravity center migration ratio; The process of obtaining the spectrum bandwidth migration ratio includes: in a combination, taking the ratio of the spectrum bandwidth of the subsequent segment signal to the spectrum bandwidth of the previous segment signal as the spectrum bandwidth migration ratio; The process of obtaining the power spectrum migration ratio includes: in one combination, taking the ratio of the power spectrum of the subsequent signal segment to the power spectrum of the previous signal segment as the power spectrum migration ratio.

[0010] Furthermore, the process of forming a feature matching vector includes: Obtaining a first similarity coefficient according to a difference between the fundamental frequency migration ratio of each combination and the stored fundamental frequency migration ratio; A second similarity coefficient is obtained according to the difference between the harmonic center of gravity migration ratio of each combination and the stored harmonic center of gravity migration ratio; Obtaining a third similarity coefficient according to a difference between the spectrum bandwidth migration ratio of each combination and the stored spectrum bandwidth migration ratio; Obtaining a fourth similarity coefficient according to a difference between the power spectrum migration ratio of each combination and the stored power spectrum migration ratio; The first similarity coefficient, the second similarity coefficient, the third similarity coefficient and the fourth similarity coefficient corresponding to each column of the migration ratio feature matrix of each combination are averaged to obtain the feature matching coefficient corresponding to each column; The feature matching coefficients of each column form a feature matching vector.

[0011] Furthermore, the process of obtaining the first similarity coefficient includes: subtracting the fundamental frequency migration ratio from the stored fundamental frequency migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the first similarity coefficient; The process of obtaining the second similarity coefficient includes: subtracting the harmonic center of gravity migration ratio from the stored harmonic center of gravity migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the second similarity coefficient; The process of obtaining the third similarity coefficient includes: subtracting the spectrum bandwidth migration ratio from the stored spectrum bandwidth migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the third similarity coefficient; The process of obtaining the fourth similarity coefficient includes: subtracting the power spectrum migration ratio from the stored power spectrum migration ratio, taking the absolute value of the subtraction result, normalizing the value after taking the absolute value to obtain the gap coefficient, and subtracting the gap coefficient from 1 to obtain the fourth similarity coefficient.

[0012] Furthermore, the multi-channel-BP neural network includes: 4 convolutional feature enhancement channels, Concat layer and BP neural network; The first input end of each convolution feature enhancement channel is used to input a combined transfer ratio feature matrix, and the second input end is used to input a feature matching vector of the same combination; The input end of the Concat layer is connected to the output end of the four convolutional feature enhancement channels respectively, and its output end is connected to the input end of the BP neural network; The output end of the BP neural network serves as the output end of the multi-channel-BP neural network.

[0013] Furthermore, the convolutional feature enhancement channels each include: a first convolutional layer, a second convolutional layer, and a multiplier; The input end of the first convolutional layer serves as the first input end of the convolutional feature enhancement channel, and its output end is connected to the input end of the second convolutional layer; The first input end of the multiplier is connected to the output end of the second convolutional layer, the second input end thereof serves as the second input end of the convolutional feature enhancement channel, and the output end thereof serves as the output end of the convolutional feature enhancement channel.

[0014] Furthermore, the convolution kernel size of the first convolution layer is 1×4, and the convolution kernel size of the second convolution layer is 1×1. The first convolution layer is used to perform convolution on each column of the 4×N migration ratio feature matrix to obtain a 1×N convolution compression feature. The second convolution layer is used to extract features from the 1×N convolution compression feature to obtain a 1×N mapping feature. The 1×N mapping feature is element-wise multiplied with the 1×N feature matching vector at the multiplier to obtain a 1×N convolution enhancement feature, where N is the length.

[0015] The beneficial effects of the present invention are: 1. This invention uses four combinations of stationary and disturbed segments for feature analysis, fully exploring the characteristic migration patterns between adjacent speech segments of different types. These migration patterns are directly related to the speaker's pronunciation habits and the physical characteristics of the vocal cords and vocal tract, effectively overcoming the lack of attention paid to this information in existing technologies.

[0016] 2. The fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio, and power spectrum migration ratio extracted by the present invention quantify the migration characteristics of features in different combinations from multiple dimensions. Compared with traditional static features, they can better reflect the unique way of the speaker in the process of continuous pronunciation and improve the discriminative ability of features.

[0017] 3. The present invention calculates the similarity between the migration ratio feature matrix and the stored features of each combination, and obtains the feature matching vector through mean processing. It combines the multi-channel BP neural network to collaboratively process the feature matrices and matching vectors of the four combinations, making full use of the differentiated information contained in different speech segment combinations and improving the accuracy of speech recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flow chart of a speech processing method based on artificial intelligence; Figure 2 Schematic diagram of the plateau-plateau combination; Figure 3 It is a schematic diagram of the combination of stable segment and disturbance segment; Figure 4 Schematic diagram of the disturbance segment-stationary segment combination; Figure 5 Schematic diagram of the disturbance section-disturbance section combination; Figure 6 Schematic diagram of the structure of multi-channel BP neural network. DETAILED DESCRIPTION

[0019] The specific embodiments of the present invention are described below to facilitate understanding of the present invention by those skilled in the art. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, as long as various changes are within the spirit and scope of the present invention as defined and determined by the appended claims, these changes are obvious, and all inventions and creations utilizing the concepts of the present invention are protected.

[0020] like Figure 1 As shown, a speech processing method based on artificial intelligence includes the following steps: The speech signal is divided into multiple short time frames, and the short time frames are marked based on the zero crossing rate to obtain the stable segment and the disturbance segment; Extract stationary segment-stationary segment combinations, stationary segment-disturbance segment combinations, disturbance segment-stationary segment combinations, and disturbance segment-disturbance segment combinations; For each combination, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio are extracted to construct the migration ratio feature matrix of the corresponding combination; Calculate the similarity coefficient between each column feature and the stored feature in the migration ratio feature matrix of each combination, and take the average of the similarity coefficients of each column to obtain the feature matching coefficient corresponding to each column to form a feature matching vector; A multi-channel BP neural network is used to process the transfer ratio feature matrix and feature matching vector of the four combinations to obtain the speaker identification result.

[0021] In this embodiment, the process of obtaining the stable segment and the perturbed segment includes: Divide the speech signal into multiple short-time frames; Statistically calculate the zero-crossing rate of each short-time frame, and take the average value of the zero-crossing rates of each short-time frame; Mark the short-time frames with a zero-crossing rate greater than the average value as the perturbed segment; Mark the short-time frames with a zero-crossing rate less than or equal to the average value as the stable segment.

[0022] In this embodiment, the length range of the short-time frames is: 20 - 30 milliseconds, and there is no overlap between adjacent frames.

[0023] The zero-crossing rate refers to the number of times the waveform of the speech signal crosses the zero level (from positive to negative or from negative to positive) per unit time. Its value is directly related to the proportion of high-frequency components in the signal: for stable segment speech such as vowels, the vocal cords vibrate regularly during pronunciation, the signal is mainly composed of low-frequency harmonic components, the waveform is smooth, and the number of times crossing the zero level is small, so the zero-crossing rate is low; for perturbed segment speech such as consonants (such as the plosives "p", "t", and the fricatives "s", "sh"), during pronunciation, airflows generate turbulence or explosions in the vocal tract, the signal contains a large amount of high-frequency noise components, the waveform fluctuates violently, and the number of times crossing the zero level is large, so the zero-crossing rate is high.

[0024] Two short-time frames in the combination are adjacent.

[0025] Short-time frames with a zero-crossing rate higher than the average value have significant high-frequency perturbation characteristics, which conform to the acoustic attributes of "perturbed segments" such as consonants; short-time frames with a zero-crossing rate lower than or equal to the average value have more prominent low-frequency stable characteristics, which conform to the acoustic attributes of "stable segments" such as vowels.

[0026] Such as Figure 2 shown, both the front and rear segments are stable segments with a low zero-crossing rate (such as the transition from vowel "a" to vowel "i", such as the transition from "a" to "i" in "ài"). Such as Figure 3 shown, the front segment is a stable segment (such as vowel "u"), and the rear segment is a perturbed segment with a high zero-crossing rate (such as voiceless consonant "sh"). Such as Figure 4 shown, the front segment is a perturbed segment (such as voiceless consonant "p"), and the rear segment is a stable segment (such as vowel "a"). Such as Figure 5 shown, both the front and rear segments are perturbed segments with a high zero-crossing rate (such as voiceless consonant "s" followed by voiceless consonant "h").

[0027] In this embodiment, the process of constructing the migration ratio feature matrix corresponding to the combination includes: Perform Fourier transform on each segment of the signal in each combination to obtain the spectrum of each segment of the signal in the combination; In one combination, according to the spectrum of each segment of the signal, obtain the fundamental frequency migration ratio, harmonic centroid migration ratio, spectral bandwidth migration ratio, and power spectrum migration ratio of each combination; The fundamental frequency migration ratios of the same combination at each moment constitute a fundamental frequency migration ratio vector; The harmonic center of gravity migration ratio of the same combination at each moment is used to form a harmonic center of gravity migration ratio vector; The spectrum bandwidth migration ratios of the same combination at each moment constitute a spectrum bandwidth migration ratio vector; The power spectrum migration ratios of the same combination at each moment constitute a power spectrum migration ratio vector; The fundamental frequency migration ratio vector, the harmonic center of gravity migration ratio vector, the spectrum bandwidth migration ratio vector and the power spectrum migration ratio vector are used as row vectors to form a migration ratio characteristic matrix of the combination.

[0028] The process of constructing the migration ratio characteristic matrix of the stationary segment-stationary segment combination is illustrated by taking the stationary segment-stationary segment combination as an example: Fourier transform is performed on each signal segment in the stationary segment-stationary segment combination to obtain the spectrum of each signal segment in the combination; In the stationary segment-stationary segment combination, according to the spectrum of each signal segment, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio of the stationary segment-stationary segment combination are obtained; The fundamental frequency migration ratio of the stationary segment-stationary segment combination at each moment constitutes a fundamental frequency migration ratio vector; The harmonic gravity center migration ratio of the stationary segment-stationary segment combination at each moment is used to form a harmonic gravity center migration ratio vector; The spectrum bandwidth migration ratio of the stationary segment-stationary segment combination at each moment is used to form a spectrum bandwidth migration ratio vector; The power spectrum migration ratio of the stationary segment-stationary segment combination at each moment constitutes a power spectrum migration ratio vector; The fundamental frequency migration ratio vector, the harmonic center of gravity migration ratio vector, the spectrum bandwidth migration ratio vector and the power spectrum migration ratio vector are used as row vectors to form a migration ratio characteristic matrix of the stationary segment-stationary segment combination.

[0029] The process of obtaining the stationary segment-disturbance segment combination, the disturbance segment-stationary segment combination, and the disturbance segment-disturbance segment combination is the same as that of obtaining the migration ratio characteristic matrix of the stationary segment-stationary segment combination.

[0030] In this embodiment, the process of obtaining the fundamental frequency migration ratio includes: in a combination, taking the frequency corresponding to the maximum amplitude in the spectrum of each signal segment as the fundamental frequency, and taking the ratio of the fundamental frequency of the subsequent signal segment to the fundamental frequency of the previous signal segment as the fundamental frequency migration ratio.

[0031] The process of obtaining the fundamental frequency migration ratio is illustrated by taking the stationary segment-stationary segment combination as an example: in the stationary segment-stationary segment combination, the frequency corresponding to the maximum amplitude in the spectrum of each signal segment is taken as the fundamental frequency, and the ratio of the fundamental frequency of the subsequent stationary segment to the fundamental frequency of the previous stationary segment is taken as the fundamental frequency migration ratio of the stationary segment-stationary segment combination.

[0032] The process of obtaining the harmonic center of gravity migration ratio includes: in a combination, according to the fundamental frequency in the spectrum of each signal segment, extracting each harmonic, calculating the harmonic center of gravity of each harmonic, and taking the ratio of the harmonic center of gravity of the subsequent signal segment to the harmonic center of gravity of the previous signal segment as the harmonic center of gravity migration ratio.

[0033] In this embodiment, the harmonic center of gravity is calculated using the existing spectrum frequency center of gravity formula.

[0034] The process of obtaining the harmonic center of gravity migration ratio is illustrated by taking the stationary segment-stationary segment combination as an example: in the stationary segment-stationary segment combination, each harmonic is extracted according to the fundamental frequency in the spectrum of each signal segment, the harmonic center of gravity of each harmonic is calculated, and the ratio of the harmonic center of gravity of the subsequent stationary segment to the harmonic center of gravity of the previous stationary segment is taken as the harmonic center of gravity migration ratio.

[0035] The process of obtaining the spectrum bandwidth migration ratio includes: in one combination, taking the ratio of the spectrum bandwidth of the subsequent segment signal to the spectrum bandwidth of the previous segment signal as the spectrum bandwidth migration ratio.

[0036] The process of obtaining the spectral bandwidth migration ratio is illustrated by taking the stationary segment-stationary segment combination as an example: in the stationary segment-stationary segment combination, the ratio of the spectral bandwidth of the subsequent stationary segment to the spectral bandwidth of the previous stationary segment is taken as the spectral bandwidth migration ratio.

[0037] The process of obtaining the power spectrum migration ratio includes: in one combination, taking the ratio of the power spectrum of the subsequent signal segment to the power spectrum of the previous signal segment as the power spectrum migration ratio.

[0038] The process of obtaining the power spectrum migration ratio is illustrated by taking the stationary segment-stationary segment combination as an example: in the stationary segment-stationary segment combination, the ratio of the power spectrum of the subsequent stationary segment to the power spectrum of the previous stationary segment is taken as the spectrum bandwidth migration ratio.

[0039] The processes of obtaining the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio for the stationary segment-disturbance segment combination, disturbance segment-stationary segment combination, disturbance segment-disturbance segment combination and stationary segment-stationary segment combination are the same.

[0040] The fundamental frequency is directly related to the vibration frequency of the speaker's vocal folds, and its transfer ratio accurately reflects the individual changes in vocal fold vibration during the transition between different speech segments. The harmonic center of gravity transfer ratio focuses on the harmonic components of the fundamental frequency, capturing the migration pattern of harmonic energy distribution through the ratio of the harmonic centers of gravity. Harmonic characteristics are closely related to the resonance characteristics of the speaker's vocal tract and can reflect individual differences in vocal tract structure. The spectral bandwidth transfer ratio focuses on the relative changes in spectral width. The spectral bandwidth reflects the distribution range of the frequency components of the speech signal. Its transfer ratio can characterize the degree of spectral broadening or contraction of the preceding and following signal segments, further enriching the dynamic information of the spectral characteristics. The power spectrum transfer ratio focuses on the relative changes in signal energy distribution. The power spectrum is related to the intensity and articulation force of the speech signal, and its transfer ratio can reflect the individual changes in the energy of the preceding and following signal segments. These four types of transfer ratios extract feature transfer information for different speech segment combinations from four dimensions: vocal fold vibration, vocal tract resonance, spectral distribution range, and energy distribution.

[0041] The transfer ratio in this invention is calculated by calculating the ratio of the signal parameters of the subsequent segment to the signal parameters of the preceding segment. This relative change quantification method effectively captures the transition pattern between the two signal segments in terms of specific acoustic features. Compared to the static parameters of a single signal segment, the transfer ratio can better highlight individual speech patterns across different speech segment combinations (such as steady-state segment to steady-state segment, steady-state segment to disturbed segment, disturbed segment to steady-state segment, and disturbed segment to disturbed segment). These dynamic patterns embody the speaker's unique pronunciation habits and speech rhythm, providing more discriminative features for distinguishing different speakers.

[0042] In this embodiment, the process of forming a feature matching vector includes: Obtaining a first similarity coefficient according to a difference between the fundamental frequency migration ratio of each combination and the stored fundamental frequency migration ratio; A second similarity coefficient is obtained according to the difference between the harmonic center of gravity migration ratio of each combination and the stored harmonic center of gravity migration ratio; Obtaining a third similarity coefficient according to a difference between the spectrum bandwidth migration ratio of each combination and the stored spectrum bandwidth migration ratio; Obtaining a fourth similarity coefficient according to a difference between the power spectrum migration ratio of each combination and the stored power spectrum migration ratio; The first similarity coefficient, the second similarity coefficient, the third similarity coefficient and the fourth similarity coefficient corresponding to each column of the migration ratio feature matrix of each combination are averaged to obtain the feature matching coefficient corresponding to each column; The feature matching coefficients of each column form a feature matching vector.

[0043] The stored fundamental frequency migration ratios include: the stored fundamental frequency migration ratio of the steady segment-steady segment combination, the stored fundamental frequency migration ratio of the steady segment-disturbance segment combination, the stored fundamental frequency migration ratio of the disturbance segment-steady segment combination, and the stored fundamental frequency migration ratio of the disturbance segment-disturbance segment combination. These stored fundamental frequency migration ratios are derived from the fundamental frequency migration ratios of four combinations of pre-recorded user voice signals.

[0044] The stored harmonic center of gravity migration ratio includes: the stored harmonic center of gravity migration ratio of the steady segment-steady segment combination, the stored harmonic center of gravity migration ratio of the steady segment-disturbance segment combination, the stored harmonic center of gravity migration ratio of the disturbance segment-steady segment combination, and the stored harmonic center of gravity migration ratio of the disturbance segment-disturbance segment combination. These stored harmonic center of gravity migration ratios are derived from the harmonic center of gravity migration ratios of four combinations of pre-recorded user voice signals.

[0045] The stored spectrum bandwidth migration ratio includes: the stored spectrum bandwidth migration ratio of the steady segment-steady segment combination, the stored spectrum bandwidth migration ratio of the steady segment-disturbance segment combination, the stored spectrum bandwidth migration ratio of the disturbance segment-steady segment combination, and the stored spectrum bandwidth migration ratio of the disturbance segment-disturbance segment combination. These stored spectrum bandwidth migration ratios are derived from the spectrum bandwidth migration ratios of four combinations of pre-recorded user voice signals.

[0046] The stored power spectrum migration ratio includes: the stored power spectrum migration ratio of the steady segment-steady segment combination, the stored power spectrum migration ratio of the steady segment-disturbance segment combination, the stored power spectrum migration ratio of the disturbance segment-steady segment combination, and the stored power spectrum migration ratio of the disturbance segment-disturbance segment combination. These stored power spectrum migration ratios are derived from the power spectrum migration ratios of four combinations of pre-recorded user voice signals.

[0047] In this embodiment, the process of obtaining the similarity coefficient includes: subtracting the fundamental frequency migration ratio, the harmonic center of gravity migration ratio, the spectrum bandwidth migration ratio, and the power spectrum migration ratio from the stored corresponding migration ratios, taking the absolute value of the difference and normalizing it to obtain the gap coefficient, and then subtracting the gap coefficient from 1 to obtain the first, second, third, and fourth similarity coefficients, respectively.

[0048] In this embodiment, the process of obtaining the first similarity coefficient includes: subtracting the fundamental frequency migration ratio from the stored fundamental frequency migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the first similarity coefficient; The process of obtaining the second similarity coefficient includes: subtracting the harmonic center of gravity migration ratio from the stored harmonic center of gravity migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the second similarity coefficient; The process of obtaining the third similarity coefficient includes: subtracting the spectrum bandwidth migration ratio from the stored spectrum bandwidth migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the third similarity coefficient; The process of obtaining the fourth similarity coefficient includes: subtracting the power spectrum migration ratio from the stored power spectrum migration ratio, taking the absolute value of the subtraction result, normalizing the value after taking the absolute value to obtain the gap coefficient, and subtracting the gap coefficient from 1 to obtain the fourth similarity coefficient.

[0049] In this embodiment, the formula for calculating the similarity coefficient is: , where θ is the similarity coefficient and V is the absolute value. The present invention calculates the similarity coefficient for each combination of the four transfer ratios and takes the average to obtain the feature matching coefficient for that combination. The larger the feature matching coefficient, the closer it is to the stored human voice, thereby improving the recognition accuracy of the multi-channel BP neural network.

[0050] like Figure 6 As shown, the multi-channel-BP neural network includes: 4 convolutional feature enhancement channels, Concat layer and BP neural network; The first input end of each convolution feature enhancement channel is used to input a combined transfer ratio feature matrix, and the second input end is used to input a feature matching vector of the same combination; The input end of the Concat layer is connected to the output end of the four convolutional feature enhancement channels respectively, and its output end is connected to the input end of the BP neural network; The output end of the BP neural network serves as the output end of the multi-channel-BP neural network.

[0051] exist Figure 6 Among them, the first combination is a steady segment-steady segment combination, the second combination is a steady segment-disturbance segment combination, the third combination is a disturbance segment-steady segment combination, and the fourth combination is a disturbance segment-disturbance segment combination.

[0052] The present invention uses each convolution feature enhancement channel to process a combined migration ratio feature matrix and feature matching vector to achieve feature enhancement, then performs feature splicing through the Concat layer, and uses a BP neural network to perform recognition based on the spliced ​​features.

[0053] In this embodiment, the convolutional feature enhancement channel includes: a first convolutional layer, a second convolutional layer and a multiplier; The input end of the first convolutional layer serves as the first input end of the convolutional feature enhancement channel, and its output end is connected to the input end of the second convolutional layer; The first input end of the multiplier is connected to the output end of the second convolutional layer, the second input end thereof serves as the second input end of the convolutional feature enhancement channel, and the output end thereof serves as the output end of the convolutional feature enhancement channel.

[0054] In this embodiment, the convolution kernel size of the first convolution layer is 1×4, and the convolution kernel size of the second convolution layer is 1×1. The first convolution layer is used to perform convolution on each column of the 4×N migration ratio feature matrix to obtain a 1×N convolution compression feature. The second convolution layer is used to extract features from the 1×N convolution compression feature to obtain a 1×N mapping feature. The 1×N mapping feature is element-wise multiplied with the 1×N feature matching vector at the multiplier to obtain a 1×N convolution enhancement feature, where N is the length.

[0055] The present invention uses a first convolutional layer to process a 4×N migration ratio feature matrix, compressing the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectral bandwidth migration ratio, and power spectrum migration ratio within the same column, converting them into 1×N convolutional compressed features. Feature mapping is then performed through a second convolutional layer to obtain mapped features. A multiplier multiplies the mapped features element-wise with the feature matching vector. Positions in the mapped features that have a high degree of match (large matching coefficient) with the stored features are made more prominent, while positions with low matching (small matching coefficient) are suppressed, increasing the gap between different features. This weighting method enables the network to automatically focus on key information consistent with the target speaker's characteristics while attenuating irrelevant or significantly different interference information.

[0056] In this embodiment, it can be set that when the BP neural network output is greater than or equal to a threshold, the speaker's identity is valid; if it is less than the threshold, it is invalid and the identity verification fails.

[0057] This invention uses four combinations of stationary and disturbed segments for feature analysis, fully exploring the characteristic migration patterns between adjacent speech segments of different types. These migration patterns are directly related to the speaker's pronunciation habits and the physical characteristics of the vocal cords and vocal tract, effectively overcoming the lack of attention paid to this information in existing technologies.

[0058] The fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio extracted by the present invention quantify the migration characteristics of features in different combinations from multiple dimensions. Compared with traditional static features, they can better reflect the unique way of the speaker in the process of continuous pronunciation and improve the discrimination ability of features.

[0059] The present invention calculates the similarity between the migration ratio feature matrix and the stored features of each combination, obtains the feature matching vector through mean processing, and combines the multi-channel BP neural network to collaboratively process the feature matrices and matching vectors of the four combinations, making full use of the differentiated information contained in different speech segment combinations and improving the speech recognition accuracy.

[0060] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Those skilled in the art will readily appreciate that various modifications and variations of the present invention are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A speech processing method based on artificial intelligence, characterized in that: The following steps are involved: The speech signal is divided into multiple short time frames, and the short time frames are marked based on the zero crossing rate to obtain the stable segment and the disturbance segment; Extract stationary segment-stationary segment combinations, stationary segment-disturbance segment combinations, disturbance segment-stationary segment combinations, and disturbance segment-disturbance segment combinations; For each combination, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio are extracted to construct the migration ratio feature matrix of the corresponding combination; Calculate the similarity coefficient between each column feature and the stored feature in the migration ratio feature matrix of each combination, and take the average of the similarity coefficients of each column to obtain the feature matching coefficient corresponding to each column to form a feature matching vector; A multi-channel BP neural network is used to process the transfer ratio feature matrix and feature matching vector of the four combinations to obtain the speaker identification result.

2. The artificial intelligence-based speech processing method according to claim 1, characterized in that: The process of obtaining the stable segment and the disturbed segment includes: Divide the speech signal into multiple short time frames; Count the zero-crossing rate of each short-time frame and take the average of the zero-crossing rates of each short-time frame; The short time frames with zero-crossing rate greater than the mean are marked as disturbance segments; The short time frames with zero crossing rate less than or equal to the mean are marked as stationary segments.

3. The speech processing method based on artificial intelligence according to claim 1, characterized in that: The process of constructing the migration ratio feature matrix of the corresponding combination includes: Perform Fourier transform on each signal segment in each combination to obtain the spectrum of each signal segment in the combination; In a combination, according to the spectrum of each signal segment, the fundamental frequency migration ratio, harmonic center of gravity migration ratio, spectrum bandwidth migration ratio and power spectrum migration ratio of each combination are obtained; The fundamental frequency migration ratios of the same combination at each moment constitute a fundamental frequency migration ratio vector; The harmonic center of gravity migration ratio of the same combination at each moment is used to form a harmonic center of gravity migration ratio vector; The spectrum bandwidth migration ratios of the same combination at each moment constitute a spectrum bandwidth migration ratio vector; The power spectrum migration ratios of the same combination at each moment constitute a power spectrum migration ratio vector; The fundamental frequency migration ratio vector, the harmonic center of gravity migration ratio vector, the spectrum bandwidth migration ratio vector and the power spectrum migration ratio vector are used as row vectors to form a migration ratio characteristic matrix of the combination.

4. The artificial intelligence-based speech processing method according to claim 1 or 3, characterized in that: The process of obtaining the fundamental frequency shift ratio includes: in a combination, taking the frequency corresponding to the maximum amplitude in the spectrum of each segment signal as the fundamental frequency, and taking the ratio of the fundamental frequency of the subsequent segment signal to the fundamental frequency of the previous segment signal as the fundamental frequency shift ratio; The process of obtaining the harmonic gravity center migration ratio includes: extracting each harmonic according to the fundamental frequency in the spectrum of each signal segment in a combination, calculating the harmonic gravity center of each harmonic, and taking the ratio of the harmonic gravity center of the subsequent signal segment to the harmonic gravity center of the previous signal segment as the harmonic gravity center migration ratio; The process of obtaining the spectrum bandwidth migration ratio includes: in a combination, taking the ratio of the spectrum bandwidth of the subsequent segment signal to the spectrum bandwidth of the previous segment signal as the spectrum bandwidth migration ratio; The process of obtaining the power spectrum migration ratio includes: in one combination, taking the ratio of the power spectrum of the subsequent signal segment to the power spectrum of the previous signal segment as the power spectrum migration ratio.

5. The artificial intelligence-based speech processing method according to claim 1, characterized in that: The process of forming a feature matching vector includes: Obtaining a first similarity coefficient according to a difference between the fundamental frequency migration ratio of each combination and the stored fundamental frequency migration ratio; A second similarity coefficient is obtained according to the difference between the harmonic center of gravity migration ratio of each combination and the stored harmonic center of gravity migration ratio; Obtaining a third similarity coefficient according to a difference between the spectrum bandwidth migration ratio of each combination and the stored spectrum bandwidth migration ratio; Obtaining a fourth similarity coefficient according to a difference between the power spectrum migration ratio of each combination and the stored power spectrum migration ratio; The first similarity coefficient, the second similarity coefficient, the third similarity coefficient and the fourth similarity coefficient corresponding to each column of the migration ratio feature matrix of each combination are averaged to obtain the feature matching coefficient corresponding to each column; The feature matching coefficients of each column form a feature matching vector.

6. The artificial intelligence-based speech processing method according to claim 1 or 5, characterized in that: The process of obtaining the first similarity coefficient includes: subtracting the fundamental frequency migration ratio from the stored fundamental frequency migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the first similarity coefficient; The process of obtaining the second similarity coefficient includes: subtracting the harmonic center of gravity migration ratio from the stored harmonic center of gravity migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the second similarity coefficient; The process of obtaining the third similarity coefficient includes: subtracting the spectrum bandwidth migration ratio from the stored spectrum bandwidth migration ratio, taking the absolute value of the subtraction result, normalizing the absolute value to obtain a gap coefficient, and subtracting the gap coefficient from 1 to obtain the third similarity coefficient; The process of obtaining the fourth similarity coefficient includes: subtracting the power spectrum migration ratio from the stored power spectrum migration ratio, taking the absolute value of the subtraction result, normalizing the value after taking the absolute value to obtain the gap coefficient, and subtracting the gap coefficient from 1 to obtain the fourth similarity coefficient.

7. The speech processing method based on artificial intelligence according to claim 1, characterized in that: Multi-channel-BP neural network includes: 4 convolutional feature enhancement channels, Concat layer and BP neural network; The first input end of each convolution feature enhancement channel is used to input a combined transfer ratio feature matrix, and the second input end is used to input a feature matching vector of the same combination; The input end of the Concat layer is connected to the output end of the four convolutional feature enhancement channels respectively, and its output end is connected to the input end of the BP neural network; The output end of the BP neural network serves as the output end of the multi-channel-BP neural network.

8. The artificial intelligence-based speech processing method according to claim 7, characterized in that: The convolution feature enhancement channels all include: the first convolution layer, the second convolution layer and the multiplier; The input end of the first convolutional layer serves as the first input end of the convolutional feature enhancement channel, and its output end is connected to the input end of the second convolutional layer; The first input end of the multiplier is connected to the output end of the second convolutional layer, the second input end thereof serves as the second input end of the convolutional feature enhancement channel, and the output end thereof serves as the output end of the convolutional feature enhancement channel.

9. The artificial intelligence-based speech processing method according to claim 8, characterized in that: The convolution kernel size of the first convolution layer is 1×4, and the convolution kernel size of the second convolution layer is 1×1. The first convolution layer is used to convolve each column of the 4×N migration ratio feature matrix to obtain 1×N convolution compression features. The second convolution layer is used to extract features from the 1×N convolution compression features to obtain 1×N mapping features. At the multiplier, the 1×N mapping features are element-wise multiplied with the 1×N feature matching vector to obtain 1×N convolution enhanced features, where N is the length.

Citation Information

Patent Citations

  • Human voice detection method and device

    CN114242074A

  • Intelligent platform for appearance design

    CN115168333A

  • Multi-speaker identification method based on FMCW radar

    CN116203521A

  • Single-channel speech enhancement method for balancing noise reduction amount and speech quality

    CN116913308A

  • Voice decoding recognition method and system for multi-mode throat vibration signal and lip moving point data

    CN119068870A