A multi-feature speech fusion processing method
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HANGZHOU DIANZI UNIVERSTIY INFORMATION ENG SCHOOL
- Filing Date
- 2026-05-18
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]为了弥补以上不足,本发明提供了一种多特征语音融合的处理方法,旨在改善传统的人工耳蜗语音处理大都采用基于频带能量的固定频带选择方式,容易造成关键语义信息对应频带刺激不足的问题
1、本发明中,通过对频带语音分量分别提取频谱包络特征、基频变化特征、瞬态变化特征以及环境噪声特征,并基于基频变化特征以及瞬态变化特征生成动态增强系数,进而基于动态增强系数对频谱包络特征进行加权处理并生成特征贡献值,从而改善了传统的人工耳蜗语音处理大都采用基于频带能量的固定频带选择方式,由于高能量稳态频带长期占用刺激资源,从而造成关键语义信息对应频带刺激不足的问题。
Smart Images

Figure CN122511268A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of cochlear implant speech coding technology, and in particular to a multi-feature speech fusion processing method. Background Technology
[0002] Cochlear implant speech processing technology is an important research direction in auditory signal processing and speech coding. It primarily involves performing frequency band division, feature extraction, and stimulation parameter generation on external speech signals to convert speech information into electrical stimulation sequences suitable for cochlear implant electrode stimulation, thereby assisting hearing-impaired users in regaining their speech perception abilities. Current cochlear implant speech processing solutions typically employ a fixed frequency band division method and select target stimulation channels based on the short-time energy or spectral envelope information of each frequency band, subsequently generating corresponding stimulation parameters based on the target stimulation channel. Some solutions also incorporate frequency variation information or noise suppression strategies to adjust the stimulation parameters, thereby improving speech processing capabilities in complex speech environments.
[0003] Traditional cochlear implant speech processing mostly adopts a fixed frequency band selection method based on frequency band energy. Because high-energy steady-state frequency bands occupy stimulation resources for a long time, this results in insufficient stimulation of the frequency bands corresponding to key semantic information. Summary of the Invention
[0004] To overcome the above shortcomings, this invention provides a multi-feature speech fusion processing method, which aims to improve the problem that traditional cochlear implant speech processing mostly adopts a fixed frequency band selection method based on frequency band energy, which easily leads to insufficient stimulation of the frequency band corresponding to key semantic information.
[0005] This invention provides the following technical solution: a multi-feature speech fusion processing method, comprising the following steps: S1. Acquire the input speech signal, perform frequency band decomposition on the input speech signal, and obtain multiple frequency band speech components; S2. Extract the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component, and perform noise filtering on the frequency band speech components based on the environmental noise features. S3. For the frequency band speech components that have passed noise screening, generate dynamic enhancement coefficients based on fundamental frequency change features and transient change features, and perform weighted processing on the spectral envelope features based on the dynamic enhancement coefficients to generate feature contribution values corresponding to each frequency band speech component. S4. Generate a stimulus channel priority sequence based on the feature contribution value corresponding to each frequency band speech component. Correct the stimulus channel priority sequence according to the historical stimulus state corresponding to each stimulus channel. When the historical stimulus load corresponding to the target stimulus channel exceeds the preset threshold, select the frequency band speech component with lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel to perform priority compensation. S5. Generate stimulation parameters based on the corrected stimulation channel priority sequence.
[0006] The above technical solution extracts spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features from the frequency band speech components. Based on the fundamental frequency variation features and transient variation features, dynamic enhancement coefficients are generated. Then, based on the dynamic enhancement coefficients, the spectral envelope features are weighted and a feature contribution value is generated. This improves the problem that traditional cochlear implant speech processing mostly adopts a fixed frequency band selection method based on frequency band energy. Because high-energy steady-state frequency bands occupy stimulation resources for a long time, the frequency band stimulation corresponding to key semantic information is insufficient.
[0007] Furthermore, in S1, the step of performing frequency band decomposition on the input speech signal includes: The input speech signal is sequentially processed by pre-emphasis, framing, and windowing to obtain continuous speech analysis frames. Frequency domain transformation processing is performed on continuous speech analysis frames to obtain the spectrum signal; Based on the filter bank, frequency division processing is performed on the spectrum signal to obtain multiple frequency band speech components corresponding to different frequency ranges; Establish the correspondence between multiple frequency band speech components and stimulation channels.
[0008] Furthermore, in S2, the steps of extracting the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component include: Spectral envelope features are generated based on the spectral energy distribution corresponding to the speech components in each frequency band. Fundamental frequency variation features are generated based on frequency changes between adjacent speech analysis frames; Transient change features are generated based on short-term energy changes between adjacent speech analysis frames; Environmental noise features are generated based on the proportion of background noise energy in the speech components of each frequency band.
[0009] Further, in S2, the step of performing noise filtering on the frequency band speech components based on environmental noise characteristics includes: Generate the frequency band signal-to-noise ratio corresponding to each frequency band speech component based on environmental noise characteristics; Compare the frequency band signal-to-noise ratio corresponding to each frequency band speech component with the preset filtering conditions; Speech components in frequency bands that meet the preset screening criteria are eligible for subsequent weighted processing. For frequency band speech components that do not meet the preset screening conditions, the priority of subsequent weighted processing is reduced.
[0010] Furthermore, in S3, the step of generating dynamic enhancement coefficients based on fundamental frequency variation characteristics and transient variation characteristics includes: Obtain the fundamental frequency variation characteristics and transient variation characteristics of the speech components in each frequency band; Calculate the frequency offset corresponding to the fundamental frequency variation characteristics and the short-time energy change corresponding to the transient variation characteristics, respectively. Normalization was performed on the magnitude of the changes and the intensity of the mutations; Dynamic enhancement coefficients are generated based on the normalized change magnitude and mutation intensity.
[0011] Furthermore, in S3, the step of weighting the spectral envelope features based on the dynamic enhancement coefficient includes: Obtain the dynamic enhancement coefficients and spectral envelope features corresponding to the speech components in each frequency band; Multiplicative gain processing is performed on the spectral envelope features based on the dynamic enhancement coefficient; Generate enhanced spectral envelope features; The feature contribution values for each frequency band speech component are generated based on the enhanced spectral envelope features.
[0012] Furthermore, in S4, the step of correcting the stimulus channel priority sequence based on the historical stimulus state corresponding to each stimulus channel includes: Obtain the historical stimulus status corresponding to each stimulus channel; The historical stimulus load of each stimulus channel is statistically analyzed based on the historical stimulus state corresponding to each stimulus channel. Compare the historical stimulus load corresponding to each stimulus channel with the preset load threshold; For stimulation channels whose historical stimulus load exceeds a preset load threshold, reduce their current priority.
[0013] Further, in S4, the step of selecting the speech component with a lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel and performing priority compensation includes: Identify target stimulus channels whose historical stimulus load exceeds a preset threshold; Obtain the speech components of the adjacent frequency bands corresponding to the target stimulus channel; Statistical analysis of the historical stimulus load corresponding to speech components in adjacent frequency bands; Prioritize the stimulus channels corresponding to speech components in frequency bands with lower historical stimulus loads.
[0014] The present invention has the following beneficial effects: 1. In this invention, spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features are extracted from the frequency band speech components. Dynamic enhancement coefficients are generated based on the fundamental frequency variation features and transient variation features. Then, the spectral envelope features are weighted based on the dynamic enhancement coefficients to generate feature contribution values. This improves the problem that traditional cochlear implant speech processing mostly adopts a fixed frequency band selection method based on frequency band energy. Since high-energy steady-state frequency bands occupy stimulation resources for a long time, the frequency band stimulation corresponding to key semantic information is insufficient.
[0015] 2. In this invention, noise screening is performed on the frequency band speech components based on environmental noise characteristics, and a stimulation channel priority sequence is generated based on the feature contribution value corresponding to the frequency band speech components. Then, stimulation parameters are generated according to the corrected stimulation channel priority sequence. This improves the traditional cochlear implant speech processing method, which mostly adopts the processing method of directly generating stimulation parameters based on the current frequency band energy. Due to the lack of a competitive screening process for noise frequency bands, the problem of high noise frequency bands participating in the allocation of stimulation resources is caused.
[0016] 3. In this invention, the priority sequence of stimulation channels is corrected according to the historical stimulation state corresponding to each stimulation channel. When the historical stimulation load corresponding to the target stimulation channel exceeds a preset threshold, the speech component with a lower historical stimulation load in the adjacent frequency band speech component corresponding to the target stimulation channel is selected for priority compensation. Then, stimulation parameters are generated according to the corrected stimulation channel priority sequence. This improves the traditional cochlear implant speech processing method, which mostly adopts the method of independently generating stimulation parameters based on the current speech frame. Due to the lack of a constraint process on the historical stimulation load, some stimulation channels are continuously activated for a long time. Attached Figure Description
[0017] Figure 1 This is a flowchart illustrating a multi-feature speech fusion processing method proposed in this invention. Figure 2 This is a schematic diagram of the frequency band decomposition process for multi-feature speech fusion proposed in this invention; Figure 3 This is a schematic diagram of a speech feature extraction and noise screening process for multi-feature speech fusion proposed in this invention. Figure 4 This is a schematic diagram of a dynamic enhancement process for multi-feature speech fusion proposed in this invention; Figure 5 This is a schematic diagram of a stimulation channel priority correction and compensation process based on historical stimulus load for multi-feature speech fusion proposed in this invention. Figure 6 This is a schematic diagram of the architecture of a multi-feature speech fusion processing system proposed in an embodiment of the present invention. Detailed Implementation
[0018] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Example 1: In the first embodiment of the present invention, a multi-feature speech fusion processing method is provided, such as... Figure 1 and Figure 2 As shown, the process includes the following steps: S1, acquiring the input speech signal, performing frequency band decomposition on the input speech signal to obtain multiple frequency band speech components; Furthermore, in S1, the step of performing frequency band decomposition on the input speech signal includes: The input speech signal is sequentially processed by pre-emphasis, framing, and windowing to obtain continuous speech analysis frames. Frequency domain transformation processing is performed on continuous speech analysis frames to obtain the spectrum signal; Based on the filter bank, frequency division processing is performed on the spectrum signal to obtain multiple frequency band speech components corresponding to different frequency ranges; Establish the correspondence between multiple frequency band speech components and stimulation channels.
[0020] Specifically, the input speech signal first undergoes pre-emphasis processing to enhance high-frequency speech components and reduce the impact of speech spectrum tilt on subsequent frequency band division. The pre-emphasized speech signal can be represented as follows: ;in, Indicates the input speech signal at the th The amplitude at each sampling point This indicates the amplitude of the pre-emphasized speech signal. This represents the pre-weighting coefficient, in this scheme, The preferred value is 0.95; subsequently, framing processing is performed on the pre-emphasized speech signal, dividing the continuous speech into multiple continuous speech analysis frames. Each continuous speech analysis frame corresponds to a fixed length of sampling data. In this scheme, the preferred duration of the speech analysis frame is... The preferred overlap duration between adjacent speech analysis frames is... After framing, windowing is applied to each consecutive speech analysis frame. The windowed speech analysis frame can be represented as follows: ;in, This represents the raw sampled values in the speech analysis frame. Represents the window function coefficients. This represents the windowed speech analysis frame. In this scheme, the preferred window function is a Hamming window, which can be expressed as: ;in, This represents the number of sampling points corresponding to a single speech analysis frame. After windowing, frequency domain transformation is performed on consecutive speech analysis frames to obtain the spectral signal. The frequency domain transformation result can be expressed as... ;in, Indicates the first The spectral amplitude corresponding to each frequency point Represents the imaginary unit. The frequency index value is represented; subsequently, frequency division processing is performed on the spectral signal based on the filter bank, dividing the spectral energy within different frequency ranges into different frequency bands. Each frequency band's speech component corresponds to speech information within a different frequency range. In this scheme, the filter bank is preferably a cochlear-like filter bank or... The filter bank preferably has 8 to 22 frequency bands. After the speech components of each frequency band are generated, the correspondence between the speech components of multiple frequency bands and the stimulation channels is further established so that different frequency ranges correspond to different stimulation channels. Among them, the low frequency band speech components correspond to the low-order stimulation channels, and the high frequency band speech components correspond to the high-order stimulation channels. The obtained frequency band speech components are used as input data for subsequent spectral envelope features, fundamental frequency change features, transient change features, and environmental noise features, as well as as input data for feature contribution value generation and stimulation channel priority ranking.
[0021] like Figure 1 and Figure 3 As shown, S2 extracts the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component, and performs noise filtering on the frequency band speech components based on the environmental noise features; Furthermore, in S2, the steps of extracting the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component include: Spectral envelope features are generated based on the spectral energy distribution corresponding to the speech components in each frequency band. Fundamental frequency variation features are generated based on frequency changes between adjacent speech analysis frames; Transient change features are generated based on short-term energy changes between adjacent speech analysis frames; Environmental noise features are generated based on the proportion of background noise energy in the speech components of each frequency band.
[0022] Specifically, the spectral envelope features, fundamental frequency variation features, transient frequency variation features, and environmental noise features are all generated based on the extraction of frequency band speech components. These frequency band speech components originate from the frequency sub-band signals after frequency band decomposition. For the spectral envelope features, the spectral energy distribution corresponding to each frequency band speech component is first statistically analyzed. The frequency band speech component is then... The frequency band energy in a speech analysis frame can be expressed as: ;in, Indicates the first The speech components of each frequency band in the first... Frequency band energy in each speech analysis frame Indicates the first The first frequency band speech component corresponds to the first Amplitude at each frequency point, and They represent the first The starting and ending frequency indices for each frequency band are used. The spectral envelope features are generated based on the frequency band energy sequences corresponding to the speech components of each frequency band, and are used to characterize the steady-state speech structure information corresponding to the speech components of each frequency band. The generated spectral envelope features serve as the basic input data for subsequent weighted processing of feature contribution values. For the fundamental frequency variation features, the change in the fundamental frequency between adjacent speech analysis frames is first statistically analyzed. The fundamental frequency change result corresponding to each speech analysis frame can be expressed as: ;in, Indicates the first The fundamental frequency variation amplitude corresponding to each speech analysis frame. Indicates the first The main frequency value corresponding to each speech analysis frame is extracted based on the spectral peak position corresponding to the speech component in each frequency band. In this scheme, the main frequency extraction range is preferably... to The generated fundamental frequency variation features are used to characterize pitch changes and tone transitions, and serve as input data for subsequent dynamic enhancement coefficient generation. For transient variation features, the short-time energy changes between adjacent speech analysis frames are first statistically analyzed. The transient change results corresponding to each speech analysis frame can be represented as follows: ;in, Indicates the first The transient change intensity corresponding to each speech analysis frame Indicates the first The short-time energy corresponding to each speech analysis frame can be expressed as: ;in, Indicates the first The first speech analysis frame Each sample value, This represents the number of sampling points corresponding to a single speech analysis frame. In this scheme, Preferably 256 or 512; transient change features are used to characterize the changes in consonant boundaries and plosives, and serve as input data in the subsequent dynamic enhancement coefficient generation process; for environmental noise features, the proportion of background noise energy in each frequency band speech component is first statistically analyzed, and the environmental noise features can be expressed as follows: ;in, Indicates the first The speech components of each frequency band in the first... The percentage of background noise energy in each speech analysis frame. Indicates the first The background noise energy corresponding to each frequency band speech component is obtained based on the statistical results of silence segments. Silence segments are determined based on speech analysis frames whose short-term energy is below a preset silence threshold. This represents the total frequency band energy of the corresponding frequency band speech components. The generated environmental noise features are used for subsequent noise screening processing. The degree of noise interference corresponding to each frequency band speech component is determined by the environmental noise features, and it serves as the screening basis for subsequent frequency band speech components to participate in dynamic enhancement processing and stimulation channel priority sorting processing.
[0023] Furthermore, in S2, the step of performing noise filtering on the frequency band speech components based on environmental noise characteristics includes: Generate the frequency band signal-to-noise ratio corresponding to each frequency band speech component based on environmental noise characteristics; Compare the frequency band signal-to-noise ratio corresponding to each frequency band speech component with the preset filtering conditions; Speech components in frequency bands that meet the preset screening criteria are eligible for subsequent weighted processing. For frequency band speech components that do not meet the preset screening conditions, the priority of subsequent weighted processing is reduced.
[0024] Specifically, noise filtering is performed based on environmental noise features to distinguish the noise interference levels corresponding to each frequency band speech component and adjust the priority of subsequent weighted processing based on the noise interference levels. The environmental noise features are derived from the statistical results of the background noise energy proportion in each frequency band speech component, where the background noise energy is obtained through statistical analysis of silent segment speech frames. First, the frequency band signal-to-noise ratio corresponding to each frequency band speech component is generated based on the environmental noise features. The speech components of each frequency band in the first... The frequency band signal-to-noise ratio in each speech analysis frame can be expressed as: ;in, Indicates the first The speech components of each frequency band in the first... Band signal-to-noise ratio in each speech analysis frame Indicates the first The frequency band energy corresponding to each frequency band speech component is obtained statistically based on the sum of the squares of the spectral amplitudes of the corresponding frequency band speech components. Indicates the first The background noise energy corresponding to each frequency band speech component is obtained based on the average noise energy statistics in the speech analysis frames of the silence segment. In this scheme, the silence segment duration is preferably... to After generating the signal-to-noise ratio (SNR) of each frequency band speech component, the SNR is compared with preset screening criteria. In this scheme, the preferred preset screening criterion is that the SNR is not lower than a certain threshold. When the frequency band signal-to-noise ratio meets the preset screening conditions, the speech component of the corresponding frequency band is retained for subsequent weighted processing, allowing it to continue participating in the generation of dynamic enhancement coefficients and feature contribution values. When the frequency band signal-to-noise ratio does not meet the preset screening conditions, the priority of subsequent weighted processing for the speech component of the corresponding frequency band is reduced. The preferred method for this reduction is to decrease the contribution weight of the speech component in the subsequent feature contribution value calculation. The result of the contribution weight adjustment can be expressed as follows: ;in, Indicates the first The original contribution weights corresponding to each frequency band speech component. This indicates the adjusted contribution weight. This represents the weight decay coefficient, in this scheme, The preferred value is 0.2 to 0.5; the adjusted contribution weight is used in the subsequent weighted processing of dynamic enhancement coefficient and spectral envelope features, so that the priority of the frequency band speech component with higher noise interference is reduced in the subsequent stimulus channel priority sorting process.
[0025] like Figure 1 and Figure 4 As shown in Figure S3, for the frequency band speech components that have passed noise screening, dynamic enhancement coefficients are generated based on fundamental frequency change features and transient change features. The spectral envelope features are then weighted based on the dynamic enhancement coefficients to generate feature contribution values corresponding to each frequency band speech component. Furthermore, in S3, the step of generating dynamic enhancement coefficients based on fundamental frequency variation characteristics and transient variation characteristics includes: Obtain the fundamental frequency variation characteristics and transient variation characteristics of the speech components in each frequency band; Calculate the frequency offset corresponding to the fundamental frequency variation characteristics and the short-time energy change corresponding to the transient variation characteristics, respectively. Normalization was performed on the magnitude of the changes and the intensity of the mutations; Dynamic enhancement coefficients are generated based on the normalized change magnitude and mutation intensity.
[0026] Specifically, the dynamic enhancement coefficients are generated based on fundamental frequency variation features and transient variation features, and are used to adjust the weights of different frequency band speech components in the subsequent feature contribution value calculation process. First, the fundamental frequency variation features and transient variation features corresponding to each frequency band speech component are obtained. The fundamental frequency variation features are derived from the main frequency variation results between adjacent speech analysis frames, and the transient variation features are derived from the short-time energy variation results between adjacent speech analysis frames. Then, the frequency offset corresponding to the fundamental frequency variation features and the short-time energy variation corresponding to the transient variation features are calculated respectively. The frequency offset corresponding to each speech analysis frame can be expressed as: ;in, Indicates the first Frequency offset corresponding to each speech analysis frame Indicates the first The main frequency value corresponding to each speech analysis frame is obtained based on the spectral peak position of the speech component in the frequency band. In this scheme, the statistical range of the center frequency is preferably the main energy frequency region within the corresponding frequency band. The short-time energy change corresponding to each speech analysis frame can be expressed as follows: ;in, Indicates the first Short-time energy change corresponding to each speech analysis frame Indicates the first The short-time energy corresponds to each speech analysis frame, and is obtained based on the sum of squares of the sampled values within the speech analysis frame. After calculating the frequency offset and the change in short-time energy, normalization is performed on the frequency offset and the change in short-time energy. The normalized frequency offset result can be expressed as follows: ;in, This represents the normalized frequency shift result. This represents the maximum frequency offset within the current speech processing cycle; the normalized short-time energy change result can be expressed as... ;in, This represents the normalized short-time energy change result. This represents the maximum short-time energy change within the current speech processing cycle; subsequently, dynamic enhancement coefficients are generated based on the normalized frequency offset and the normalized short-time energy change results. The dynamic enhancement coefficients can be expressed as... ;in, Indicates the first The dynamic enhancement coefficients corresponding to each speech analysis frame This represents the frequency offset weighting coefficient. This represents the weighting coefficient for short-term energy changes, in this scheme. Preferably, it is 0.4 to 0.6. Preferably, it is 0.4 to 0.6, and The generated dynamic enhancement coefficients are used for subsequent weighted processing of spectral envelope features. By increasing the enhancement weights of the speech components in the frequency bands with large frequency offsets and large short-term energy changes, the ranking position of the corresponding speech components in the subsequent feature contribution value generation and stimulus channel priority ranking process is improved.
[0027] Furthermore, in S3, the step of weighting the spectral envelope features based on the dynamic enhancement coefficients includes: Obtain the dynamic enhancement coefficients and spectral envelope features corresponding to the speech components in each frequency band; Multiplicative gain processing is performed on the spectral envelope features based on the dynamic enhancement coefficient; Generate enhanced spectral envelope features; The feature contribution values for each frequency band speech component are generated based on the enhanced spectral envelope features.
[0028] Specifically, in the process of allocating speech stimulus resources, the dynamic enhancement coefficients and spectral envelope features corresponding to the speech components of the frequency band after noise filtering are first obtained. The spectral envelope features characterize the spectral energy distribution of the corresponding frequency band speech component in the current speech analysis frame, and the dynamic enhancement coefficients characterize the combined change degree of the fundamental frequency variation features and transient variation features corresponding to the current frequency band speech component. The dynamic enhancement coefficients are generated jointly by the normalized frequency offset and short-time energy variation from the previous step. Subsequently, multiplicative gain processing is performed on the spectral envelope features based on the dynamic enhancement coefficients to obtain the enhanced spectral envelope features. In this embodiment, let the first... The spectral envelope features corresponding to each frequency band speech component are The dynamic enhancement coefficient is The enhanced spectral envelope features are The enhanced spectral envelope features satisfy the following relationship: ;in, By the The spectral energy distribution corresponding to each frequency band speech component is calculated. The dynamic enhancement coefficient is generated by weighting the normalized frequency offset and the normalized short-time energy change. In this scheme, the dynamic enhancement coefficient is... The value is preferably set to a real value between 0.8 and 2.0 to control the enhancement magnitude of the spectral envelope features. After multiplicative gain processing, feature contribution values corresponding to the speech components of each frequency band are generated based on the enhanced spectral envelope features. In this embodiment, let the first... The feature contribution value corresponding to each frequency band speech component is Then the feature contribution values satisfy the following relationship ,in, E' represents the total number of speech components in the frequency band.k Indicates the first The enhanced spectral envelope features corresponding to each frequency band speech component are used to characterize the relative semantic contribution of each frequency band speech component in the current speech analysis frame. After the feature contribution values are generated, the feature contribution values corresponding to each frequency band speech component are input into the subsequent stimulus channel priority sorting process to generate a stimulus channel priority sequence based on the feature contribution values corresponding to each frequency band speech component.
[0029] like Figure 1 and Figure 5 As shown, S4 generates a stimulus channel priority sequence based on the feature contribution value corresponding to each frequency band speech component, corrects the stimulus channel priority sequence according to the historical stimulus state corresponding to each stimulus channel, and performs priority compensation when the historical stimulus load corresponding to the target stimulus channel exceeds the preset threshold, selects the frequency band speech component with lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel. S5. Generate stimulation parameters based on the corrected stimulation channel priority sequence; Furthermore, in S4, the step of correcting the stimulus channel priority sequence based on the historical stimulus states corresponding to each stimulus channel includes: Obtain the historical stimulus status corresponding to each stimulus channel; The historical stimulus load of each stimulus channel is statistically analyzed based on the historical stimulus state corresponding to each stimulus channel. Compare the historical stimulus load corresponding to each stimulus channel with the preset load threshold; For stimulation channels whose historical stimulus load exceeds a preset load threshold, reduce their current priority.
[0030] Specifically, firstly, a stimulus channel priority sequence is generated based on the feature contribution values corresponding to the speech components of each frequency band. The feature contribution values characterize the semantic contribution of the corresponding frequency band speech component in the current speech analysis frame, and the stimulus channel priority sequence characterizes the stimulus allocation order of each stimulus channel in the current speech analysis frame. Then, the historical stimulus state corresponding to each stimulus channel is obtained. The historical stimulus state includes the number of consecutive activations, the most recent activation time, and the cumulative stimulus duration. Based on the historical stimulus state, the historical stimulus load corresponding to each stimulus channel is statistically analyzed. In this embodiment, let the first... The number of consecutive activations corresponding to each stimulation channel is: The most recent activation interval is The cumulative stimulation duration is Historical stimulus load is The historical stimulus load then satisfies the following relationship: ;in, This represents the load weighting coefficient corresponding to the number of consecutive activations. This represents the load weighting coefficient corresponding to the cumulative stimulus duration. This represents the load weighting coefficient corresponding to the most recent activation time interval. In this scheme, The preferred value is 0.4. The preferred value is 0.4. The preferred value is 0.2; number of consecutive activations The cumulative stimulus duration is obtained by counting the number of times the stimulus channel remains active in a continuous speech analysis frame. The most recent activation interval is obtained by statistically analyzing the cumulative activation duration of the stimulus channel within a historical time window. The time difference between the current speech analysis frame and the previous stimulus activation time is used to obtain the data. After completing the historical stimulus load statistics, the historical stimulus load corresponding to each stimulus channel is compared with a preset load threshold. In this scheme, the preset load threshold is preferably 1.2 times the average historical stimulus load. When the historical stimulus load corresponding to the first stimulation channel exceeds the preset load threshold, the first stimulation channel will be affected. The stimulation channel lowers its current priority. In this embodiment, let the first stimulation channel be... The priority of each stimulation channel before correction is The corrected priority is P' i The historical stimulus load modulation coefficient is The corrected priority satisfies the following relationship: Among them, the historical stimulus load modulation coefficient The historical stimulus load modulation coefficient is determined based on the comparison between historical stimulus load and a preset load threshold. When the historical stimulus load exceeds the preset load threshold, the modulation coefficient is adjusted accordingly. Less than 1, in this scheme, the historical stimulus load modulation coefficient The preferred value is a real value between 0.5 and 0.9. After completing the stimulation channel priority correction, a corrected stimulation channel priority sequence is generated, and subsequent stimulation parameter generation processing is performed based on the corrected stimulation channel priority sequence. The stimulation parameters include stimulation current amplitude, pulse width, and pulse firing sequence. The higher the stimulation channel priority, the higher the corresponding stimulation current amplitude and pulse firing priority. The lower the stimulation channel priority, the lower the corresponding stimulation current amplitude and pulse firing frequency.
[0031] Furthermore, in S4, the step of selecting the speech component with a lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel and performing priority compensation includes: Identify target stimulus channels whose historical stimulus load exceeds a preset threshold; Obtain the speech components of the adjacent frequency bands corresponding to the target stimulus channel; Statistical analysis of the historical stimulus load corresponding to speech components in adjacent frequency bands; Prioritize the stimulus channels corresponding to speech components in frequency bands with lower historical stimulus loads.
[0032] Specifically, after the stimulation channel priority correction is completed, priority compensation processing is performed on target stimulation channels whose historical stimulus load exceeds a preset threshold. The target stimulation channel is the stimulation channel whose historical stimulus load exceeds the preset threshold and whose current priority is reduced. First, the target stimulation channel whose historical stimulus load exceeds the preset threshold is determined. In this embodiment, let the first... The historical stimulus load corresponding to each stimulus channel is: The preset load threshold is When the following relationship is satisfied, the first... One stimulation channel was identified as the target stimulation channel: Among them, historical stimulus load The preset load threshold is obtained by statistically analyzing the number of consecutive activations, the cumulative duration of stimulation, and the time interval between the most recent activations. Based on historical average stimulus loads, a preset load threshold is determined in this scheme. Preferably, it is 1.2 times the average historical stimulus load of all stimulus channels; after determining the target stimulus channel, the adjacent frequency band speech components corresponding to the target stimulus channel are obtained. The adjacent frequency band speech components are used to characterize the frequency band speech components adjacent to the frequency range corresponding to the target stimulus channel. In this embodiment, if the frequency band number corresponding to the target stimulus channel is... The frequency band numbers corresponding to the speech components of adjacent frequency bands include as well as Subsequently, the historical stimulus loads corresponding to the speech components of adjacent frequency bands are statistically analyzed. In this embodiment, let the first... The historical stimulus load corresponding to each adjacent frequency band speech component is The candidate frequency band speech components for compensation are determined based on the following relationship: ;in, This represents the minimum historical stimulus load among speech components in adjacent frequency bands. This represents the historical stimulus load corresponding to each adjacent frequency band speech component; after determining the candidate frequency band speech components for compensation, the priority of the stimulus channel corresponding to the frequency band speech component with a lower historical stimulus load is increased. In this embodiment, the priority of the stimulus channel before compensation is set to 1. The priority of the compensated stimulation channels is The priority compensation coefficient is The priority of the compensated stimulation channels then satisfies the following relationship: Among them, the priority compensation coefficient Used to control the magnitude of priority increase, in this scheme, the priority compensation coefficient... The preferred value is a real value between 0.1 and 0.5. After priority compensation is completed, the compensated stimulation channel priority is rewritten into the corrected stimulation channel priority sequence, and subsequent stimulation parameter generation processing is performed based on the corrected stimulation channel priority sequence. The stimulation parameters include stimulation current amplitude, pulse width, and pulse firing sequence.
[0033] Example 2: In the second embodiment of the present invention, the present invention provides a multi-feature speech fusion processing system, such as... Figure 6 As shown, it includes the following modules: The frequency band decomposition module is used to acquire the input speech signal, perform frequency band decomposition on the input speech signal, and obtain multiple frequency band speech components; The feature filtering module is used to extract the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component, and to perform noise filtering on the frequency band speech components based on the environmental noise features. The dynamic enhancement module is used to generate dynamic enhancement coefficients for the frequency band speech components that have passed noise screening, based on the fundamental frequency change characteristics and transient change characteristics, and to perform weighted processing on the spectral envelope features based on the dynamic enhancement coefficients to generate the feature contribution values corresponding to each frequency band speech component. The priority compensation module is used to generate a stimulus channel priority sequence based on the feature contribution value corresponding to each frequency band speech component, correct the stimulus channel priority sequence according to the historical stimulus state corresponding to each stimulus channel, and when the historical stimulus load corresponding to the target stimulus channel exceeds the preset threshold, select the frequency band speech component with the lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel to perform priority compensation. The parameter generation module is used to generate stimulation parameters based on the corrected stimulation channel priority sequence.
[0034] In cochlear implant speech stimulation scenarios, users are often in complex noisy environments such as subway stations, shopping malls, or multiple people conversing. Background noise in these environments can easily interfere with the effective semantic information in the speech audio band. Furthermore, traditional stimulation strategies mostly employ fixed frequency band weight allocation, making it difficult to adjust stimulus resource allocation based on dynamic speech characteristics and the historical load status of the stimulation channels. This can easily lead to insufficient retention of semantic features, indistinct stimulation of key speech information, and prolonged high-load operation of some stimulation channels. To address these issues, this invention employs a multi-feature speech fusion processing system, the structure of which is as follows: Figure 6 As shown. The specific implementation process of this system is as follows: First, the frequency band decomposition module acquires the input speech signal and performs frequency band decomposition processing on the input speech signal to obtain multiple frequency band speech components. This maps the speech information in different frequency ranges to the corresponding stimulus channels, enabling the subsequent speech feature processing to perform analysis and processing for different frequency regions, while improving the ability to distinguish speech frequency information in complex speech environments. Then, the feature filtering module extracts the spectral envelope features, fundamental frequency change features, transient change features, and environmental noise features corresponding to the speech components of each frequency band, and performs noise filtering on the speech components of the frequency band based on the environmental noise features, thereby reducing the impact of high-noise frequency band speech components on the subsequent stimulus resource allocation process, while retaining the main semantic information in the effective speech frequency band and increasing the participation ratio of effective speech features in the subsequent stimulus processing. Subsequently, the dynamic enhancement module generates dynamic enhancement coefficients for the frequency band speech components that have passed noise screening, based on fundamental frequency change features and transient change features. Based on the dynamic enhancement coefficients, the spectral envelope features are weighted to generate feature contribution values corresponding to each frequency band speech component. This increases the stimulation weight of the frequency band speech components corresponding to speech transition regions, consonant boundary regions, and key semantic regions, so that speech features with obvious dynamic changes receive higher priority in the stimulus resource allocation process. Next, the priority compensation module generates a priority sequence of stimulation channels based on the feature contribution values corresponding to the speech components of each frequency band, and corrects the priority sequence of stimulation channels according to the historical stimulation state of each stimulation channel. At the same time, when the historical stimulation load corresponding to the target stimulation channel exceeds the preset threshold, the speech component with the lower historical stimulation load in the adjacent frequency band speech components corresponding to the target stimulation channel is selected for priority compensation. This avoids some stimulation channels from working under high load for a long time, and realizes the dynamic allocation of stimulation resources between adjacent frequency bands, improving the overall load distribution balance of the stimulation channels. Finally, the parameter generation module generates stimulation parameters based on the corrected stimulation channel priority sequence and uses the stimulation parameters in the cochlear implant stimulation output process. This enables the speech stimulation results to simultaneously take into account the semantic emphasis and the stimulation channel load balance, thereby improving the speech clarity and semantic recognition ability in the cochlear implant speech perception process under complex noise environments.
[0035] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A processing method of multi-feature speech fusion, characterized in that, Includes the following steps: S1. Acquire the input speech signal, perform frequency band decomposition on the input speech signal, and obtain multiple frequency band speech components; S2. Extract the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component, and perform noise filtering on the frequency band speech components based on the environmental noise features. S3. For the frequency band speech components that have passed noise screening, generate dynamic enhancement coefficients based on fundamental frequency change features and transient change features, and perform weighted processing on the spectral envelope features based on the dynamic enhancement coefficients to generate feature contribution values corresponding to each frequency band speech component. S4. Generate a stimulus channel priority sequence based on the feature contribution value corresponding to each frequency band speech component. Correct the stimulus channel priority sequence according to the historical stimulus state corresponding to each stimulus channel. When the historical stimulus load corresponding to the target stimulus channel exceeds the preset threshold, select the frequency band speech component with lower historical stimulus load from the adjacent frequency band speech components corresponding to the target stimulus channel to perform priority compensation. S5. Generate stimulation parameters based on the corrected stimulation channel priority sequence.
2. The processing method of claim 1, wherein, In S1, the step of performing frequency band decomposition on the input speech signal includes: The input speech signal is sequentially processed by pre-emphasis, framing, and windowing to obtain continuous speech analysis frames. Frequency domain transformation processing is performed on continuous speech analysis frames to obtain the spectrum signal; Based on the filter bank, frequency division processing is performed on the spectrum signal to obtain multiple frequency band speech components corresponding to different frequency ranges; Establish the correspondence between multiple frequency band speech components and stimulation channels.
3. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S2, the steps of extracting the spectral envelope features, fundamental frequency variation features, transient variation features, and environmental noise features of each frequency band speech component include: Spectral envelope features are generated based on the spectral energy distribution corresponding to the speech components in each frequency band. Fundamental frequency variation features are generated based on frequency changes between adjacent speech analysis frames; Transient change features are generated based on short-term energy changes between adjacent speech analysis frames; Environmental noise features are generated based on the proportion of background noise energy in the speech components of each frequency band.
4. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S2, the step of performing noise filtering on the frequency band speech components based on environmental noise characteristics includes: Generate the frequency band signal-to-noise ratio corresponding to each frequency band speech component based on environmental noise characteristics; Compare the frequency band signal-to-noise ratio corresponding to each frequency band speech component with the preset filtering conditions; Speech components in frequency bands that meet the preset screening criteria are eligible for subsequent weighted processing. For frequency band speech components that do not meet the preset screening conditions, the priority of subsequent weighted processing is reduced.
5. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S3, the step of generating dynamic enhancement coefficients based on fundamental frequency variation characteristics and transient variation characteristics includes: Obtain the fundamental frequency variation characteristics and transient variation characteristics of the speech components in each frequency band; Calculate the frequency offset corresponding to the fundamental frequency variation characteristics and the short-time energy change corresponding to the transient variation characteristics, respectively. Normalization was performed on the magnitude of the changes and the intensity of the mutations; Dynamic enhancement coefficients are generated based on the normalized change magnitude and mutation intensity.
6. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S3, the step of weighting the spectral envelope features based on the dynamic enhancement coefficient includes: Obtain the dynamic enhancement coefficients and spectral envelope features corresponding to the speech components in each frequency band; Multiplicative gain processing is performed on the spectral envelope features based on the dynamic enhancement coefficient; Generate enhanced spectral envelope features; The feature contribution values for each frequency band speech component are generated based on the enhanced spectral envelope features.
7. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S4, the step of correcting the stimulation channel priority sequence based on the historical stimulation state corresponding to each stimulation channel includes: Obtain the historical stimulus status corresponding to each stimulus channel; The historical stimulus load of each stimulus channel is statistically analyzed based on the historical stimulus state corresponding to each stimulus channel. Compare the historical stimulus load corresponding to each stimulus channel with the preset load threshold; For stimulation channels whose historical stimulus load exceeds a preset load threshold, reduce their current priority.
8. The multi-feature speech fusion processing method according to claim 1, characterized in that, In S4, the step of selecting the speech component with a lower historical stimulus load from the adjacent speech components corresponding to the target stimulus channel and performing priority compensation includes: Identify target stimulus channels whose historical stimulus load exceeds a preset threshold; Obtain the speech components of the adjacent frequency bands corresponding to the target stimulus channel; Statistical analysis of the historical stimulus load corresponding to speech components in adjacent frequency bands; Prioritize the stimulus channels corresponding to speech components in frequency bands with lower historical stimulus loads.