Cochlear implant coding method, cochlear implant coding system and cochlear implant equipment
By using two coding modes in the cochlear implant to process the fundamental frequency, formant, fundamental wave and harmonic information of speech and music respectively, the limitations of the existing technology in the performance of Chinese tonal language and music are solved, the recognition rate of tonal language and music melody is improved, and the user's auditory experience is improved.
Patent Information
- Application Number
- CN202510950787.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing cochlear implant coding strategies have significant limitations in the performance of tonal languages and music such as Chinese, especially in noisy environments, where the tone recognition rate is low, the music melody recognition rate is insufficient, and the speech fundamental frequency and music harmonic information cannot be effectively retained.
Two encoding modes are used: the first mode extracts fundamental frequency and formant information for speech data, and the second mode extracts fundamental and harmonic information for music data. Through high-pass filtering, frequency division filtering, full-wave rectification and low-pass filtering, the characteristic envelopes of speech and music are extracted and modulated respectively. Combined with Fourier transform for frequency domain processing and smoothing, the modulated envelope is generated and then down-sampled and mapped to the electrode array.
It significantly improves the tone language recognition rate and music perception experience of cochlear implant recipients in noisy environments, improves the user's auditory experience, especially in terms of speech intelligibility and music perception, and improves the tone language recognition rate and music melody recognition rate.
Smart Images

Figure CN120673772A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of cochlear implant signal processing, and in particular, to a cochlear implant encoding method, system and cochlear implant device. Background Art
[0002] Cochlear implants are an important rehabilitation tool for patients with severe hearing loss, and optimizing their speech coding strategies has long been a research priority. However, traditional audio processing strategies, such as continuous alternating sampling and advanced combinatorial coding, are primarily based on Western language designs and have significant limitations in their performance in tonal languages like Chinese and music. Studies have shown that Chinese speakers' tone recognition rate in noisy environments decreases by up to 40% compared to quiet environments, significantly impacting semantic comprehension.
[0003] Current encoding strategies primarily extract the time-domain envelope of speech, while the tonal characteristics of Chinese rely on fundamental frequency, formants, or harmonics. Researchers and scholars have also explored using fundamental frequency to augment tonal information, but these efforts face challenges balancing computing power and power consumption in embedded devices or difficulties in deploying real-time speech processing. Regarding music perception, existing strategies fail to effectively preserve the harmonic information of various instruments, resulting in melody recognition rates below 30%. Summary of the Invention
[0004] To address the significant limitations of existing cochlear implants in their ability to perform tonal languages like Chinese and music, this application discloses a cochlear implant encoding method, system, and device that effectively preserves speech fundamental frequency and music harmonic information, thereby improving the tonal language recognition rate and musical auditory perception of cochlear implant recipients. Specifically, the technical solutions of this application are as follows:
[0005] In a first aspect, the present application discloses a cochlear implant encoding method, which is applicable to a cochlear implant, wherein the cochlear implant includes a first mode and a second mode; and comprises the following steps:
[0006] receiving target audio data, and performing pre-emphasis processing on the target audio data using a high-pass filter to obtain emphasized audio data after the processing;
[0007] Dividing the emphasized audio data into N audio channels using a frequency division filtering method;
[0008] The N audio channels are subjected to full-wave rectification to obtain absolute values of the data; and a low-pass filter is used to obtain envelopes of each channel;
[0009] In the first mode, the feature data extracted from the weighted audio data includes: a fundamental frequency band and a formant frequency band, and the formant frequency band is smoothed based on the weighted audio data; or, in the second mode, the feature data extracted from the target audio data includes: a fundamental signal and a harmonic signal, and the harmonic signal is smoothed based on the target audio data;
[0010] Using a low-pass filter to obtain a characteristic envelope of the characteristic data; modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope;
[0011] The modulation envelope is downsampled; and an electrode array in the cochlear implant is mapped based on the downsampled data.
[0012] In some embodiments, the cochlear implant encoding method further includes:
[0013] manually or automatically switching the cochlear implant between a first mode and a second mode based on the type of the target audio data currently being received;
[0014] The target audio data includes a first audio and a second audio; and the emphasized audio data includes a first emphasized audio and a second emphasized audio.
[0015] In some implementations, in the first mode, the method of dividing the emphasized audio data into N audio channels using a frequency division filtering method specifically includes:
[0016] Converting the first emphasized audio from the time domain to the frequency domain using a fast Fourier transform, and dividing the spectrum into N frequency bands according to a preset number of channels N and a frequency range;
[0017] The frequency spectra of the N frequency bands are then transformed back to the time domain through inverse Fourier transform to obtain the N audio channels.
[0018] In some embodiments, in the first mode, extracting feature data from the emphasized audio data includes: a fundamental frequency band and a formant frequency band, and performing smoothing on the formant frequency band based on the emphasized audio data; specifically, the step of:
[0019] Converting the first emphasized audio from the time domain to the frequency domain using a fast Fourier transform; dividing the frequency domain signal into a fundamental frequency band and a plurality of formant frequency bands based on a frequency domain decomposition method; and then converting the signal to the time domain using an inverse Fourier transform;
[0020] Based on the first emphasized audio, the formant frequency band is smoothed using the following formula:
[0021] Smoothed formant frequency band = [a*formant frequency band + (1-a)*first emphasized audio]; where a is the weight coefficient;
[0022] A half-wave rectification method is used to remove the negative offset of the smoothed resonance peak frequency band, and then the negative offset is multiplied by a compensation coefficient to compensate for the energy loss of the half-wave rectification.
[0023] In some embodiments, in the first mode, the step of obtaining a characteristic envelope of the characteristic data using a low-pass filter and modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope specifically includes:
[0024] Based on the fundamental frequency band and the formant frequency band after half-wave rectification, using low-pass filters with different cutoff frequencies to extract the fundamental frequency envelope and the formant envelope respectively; normalizing the fundamental frequency envelope and the formant envelope respectively;
[0025] The fundamental frequency envelope is multiplied point by point with all the channel envelopes; and the formant envelope is multiplied point by point with several channel envelopes that match it to obtain the modulation envelope.
[0026] In some other embodiments, in the second mode, the method of dividing the emphasized audio data into N audio channels using a frequency division filtering method includes:
[0027] Divide the second emphasized audio into N frequency band channels using a frequency division filter and perform filtering processing;
[0028] Reversely arranging the order of the channel data of the N frequency band channels after filtering, and performing secondary filtering processing on the reversely arranged channel data using the frequency division filter;
[0029] The channel data after the secondary filtering is reversely arranged twice to restore the original channel order to obtain the N audio channels.
[0030] In some other embodiments, in the second mode, extracting feature data from the target audio data includes: a fundamental signal and a harmonic signal, and smoothing the harmonic signal based on the target audio data; specifically, the step of:
[0031] Converting the second audio from the time domain to the frequency domain using a fast Fourier transform method, and analyzing the power spectrum peak within a specified frequency range to determine the fundamental wave range;
[0032] Based on the fundamental wave range, multiple harmonic ranges are estimated by frequency multiplication calculation; and then the fundamental wave signal and the harmonic signal are obtained by inverse Fourier transformation;
[0033] Based on the second audio frequency, the harmonic signal is smoothed using the following formula:
[0034] Smoothed harmonic signal = [a*harmonic signal + (1-a)*second audio], where a is the weight coefficient;
[0035] A half-wave rectification method is used to remove the negative offset of the smoothed harmonic signal, and then the negative offset is multiplied by a compensation coefficient to compensate for the energy loss of the half-wave rectification.
[0036] In some other embodiments, in the second mode, the step of obtaining a characteristic envelope of the characteristic data using a low-pass filter; and modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope specifically includes:
[0037] Based on the fundamental wave signal and the harmonic signal after half-wave rectification, using low-pass filters with different cutoff frequencies to extract the fundamental wave envelope and the harmonic wave envelope respectively; normalizing the fundamental wave envelope and the harmonic wave envelope respectively;
[0038] The fundamental wave envelope is multiplied point by point with all the channel envelopes; and the harmonic envelope is multiplied point by point with several channel envelopes that match it to obtain the modulation envelope.
[0039] In a second aspect, the present application further discloses a cochlear implant coding system, comprising:
[0040] a mode switching module, configured to receive target audio data and manually or automatically switch the cochlear implant between a first mode and a second mode based on the type of the target audio data currently received;
[0041] A pre-processing module, configured to perform pre-emphasis processing on the target audio data using a high-pass filter to obtain emphasized audio data after processing;
[0042] A frequency division filtering module, configured to divide the weighted audio data into N audio channels using a frequency division filtering method;
[0043] An envelope extraction module is used to obtain the absolute value of the data by using a full-wave rectification method for the N audio channels; and to obtain the envelope of each channel by using a low-pass filter;
[0044] a feature extraction module configured to, in the first mode, extract feature data from the weighted audio data, including a fundamental frequency band and a formant frequency band, and perform smoothing on the formant frequency band based on the weighted audio data; or, in the second mode, extract feature data from the target audio data, including a fundamental signal and a harmonic signal, and perform smoothing on the harmonic signal based on the target audio data;
[0045] The feature extraction module is further configured to obtain a feature envelope of the feature data using a low-pass filter;
[0046] A characteristic modulation module, configured to modulate the channel envelope based on the characteristic envelope to obtain a modulation envelope;
[0047] A data sampling module is used to perform downsampling processing on the modulation envelope; and map the electrode array in the cochlear implant based on the downsampling data.
[0048] In a third aspect, the present application further discloses a cochlear implant device, comprising a cochlear implant coding system as described in any one of the above embodiments.
[0049] Compared with the prior art, this application has at least one of the following beneficial effects:
[0050] 1. The cochlear implant encoding method of this application has two modes, one for processing the characteristics of the first audio and the other for processing the second audio. By extracting the characteristic data of speech (including fundamental frequency and formant information) and the characteristic data of music (including fundamental wave and harmonic wave information) to modulate the output signal, the user's auditory experience can be significantly improved, especially in speech intelligibility, tonal language recognition, and music perception.
[0051] 2. In the first mode, the present application performs channel division by Fourier transform to frequency domain division method, and smoothes the formant frequency band based on the weighted audio data, and modulates the channel envelope by combining the fundamental frequency band and the formant frequency band. This can encode information such as the intonation, rhythm, and stress of the speech, and use this data to dynamically adjust the electrical stimulation pattern, which can significantly improve the recognition rate of tonal language, allowing users to more clearly distinguish different tones, while enhancing the ability to separate speech in noisy environments.
[0052] 3. In the first mode, the present application performs channel division through secondary filtering and secondary reverse arrangement. The harmonic signal is smoothed based on the target audio data, and the channel envelope is modulated in combination with the fundamental signal and the harmonic signal, thereby encoding information such as the rhythm, tempo and timbre of the music to distinguish the music, which can help users perceive the melody contour and the timbre of the instrument, thereby improving the music appreciation experience. The tone and rhythm encoding of the cochlear implant is expected to be closer to natural hearing, providing users with a richer and more natural listening experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] The preferred implementation scheme will be described below in a clear and understandable manner with reference to the accompanying drawings to further illustrate the above-mentioned characteristics, technical features, advantages and implementation methods of the present application.
[0054] Figure 1 This is a flowchart of one embodiment of a cochlear implant encoding method of the present application;
[0055] Figure 2 This is a flowchart of the first mode and the first mode processing steps in one embodiment of a cochlear implant encoding method of the present application;
[0056] Figure 3 This is a schematic structural flow chart of an embodiment of a cochlear implant coding system of the present application;
[0057] Figure 4 This is a structural flow chart of another embodiment of a cochlear implant coding system of the present application. DETAILED DESCRIPTION
[0058] In the following description, specific details such as specific system structures and technologies are provided for illustration rather than limitation to facilitate a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obstructing the description of the present application with unnecessary details.
[0059] It will be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integers, steps, operations, elements and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections.
[0060] To simplify the drawings, only portions relevant to the invention are schematically depicted in each figure; they do not represent the actual structure of the product. Furthermore, to simplify the drawings and facilitate understanding, in some figures, only one component with the same structure or function is schematically depicted or labeled. In this document, "one" not only means "only one" but also "more than one."
[0061] It should be further understood that the term "and / or" used in this specification and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.
[0062] In addition, in the description of the present application, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0063] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the specific implementation methods of the present application will be described below with reference to the accompanying drawings. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings and other implementation methods can be obtained based on these drawings without inventive work.
[0064] Cochlear Implant (CI) is designed for patients with severe to profound hearing loss and is suitable for patients with sensorineural hearing loss who are not well treated by hearing aids. It helps patients regain the ability to perceive sound by directly bypassing the damaged cochlear hair cells through electrical stimulation of the auditory nerve. Modern cochlear implant systems usually consist of two parts: 1. The external part (sound processor): worn behind the ear or on the head, responsible for collecting sound and processing signals. 2. The internal part (implant): surgically implanted in the cochlea to transmit electrical signals to the auditory nerve. Cochlear implants have successfully helped hundreds of thousands of hearing-impaired people around the world restore their hearing, and they play a key role in the language development of children with prelingual deafness. However, there is still room for improvement in the existing technology for the recognition of tonal languages and music such as Chinese.
[0065] The working principles of existing cochlear implants are as follows: 1. Sound collection: The microphone captures environmental sounds. 2. The sound processor converts the sound collected by the microphone into digital electrical signals and extracts key acoustic features, such as frequency, intensity, and time domain envelope. Signal encoding and electrode stimulation, the processor decomposes the sound into different frequency bands and converts them into electrical pulse patterns. 3. The signal is wirelessly transmitted to the internal implant through the remote sensing coil. 4. The implant decodes the digital electrical signal, generates corresponding electrical pulses, and then electrically stimulates the electrode array in the cochlea. The electrode array of the implant stimulates different positions of the cochlea according to the tonal topology. 5. Neural response and auditory perception, the electrical pulses directly activate the auditory nerve fibers, and the signals are transmitted to the brain through the auditory nerve to form hearing.
[0066] Traditional audio processing strategies, such as Continuous Interleaved Sampling (CIS) and Advanced Combination Encoders (ACE), convert sound signals into electrical pulses to stimulate the auditory nerve. These strategies rely on extracting and encoding the sound signal's envelope information, while relatively ignoring or simplifying the encoding of fine structure, such as the fundamental frequency F0. Envelope information represents the slowly changing amplitude profile of the signal over time, corresponding to the syllables, rhythm, and intonation of speech. Fine structure, on the other hand, represents the rapid fluctuations in the signal's high frequencies, corresponding to timbre, pitch, and harmonic details.
[0067] The present invention provides a cochlear implant coding system and coding method with low computational complexity and high real-time performance, which can effectively preserve speech fundamental frequency and music harmonic information, and can improve the tone language recognition rate and music auditory perception of cochlear implant recipients.
[0068] An embodiment of a cochlear implant encoding method of the present application specifically includes the following steps:
[0069] S001, receiving target audio data, and manually or automatically switching the cochlear implant between a first mode and a second mode based on the type of the target audio data currently received.
[0070] In this embodiment, the cochlear implant supports two encoding modes: Mode 1 and Mode 2. Optionally, Mode 1 is a speech mode that processes speech data, while Mode 2 is a music mode that processes music data. These two modes respectively process characteristics such as vocal cord vibration (fundamental frequency) and vocal tract tuning (formants) in speech data, and characteristics such as timbre, rhythm, and melody (fundamental and harmonics) in music data, providing users with a richer and more natural auditory experience.
[0071] Mode switching can be done manually or automatically. Manual switching: The user can manually switch between the first and second modes of the cochlear implant based on their needs. Automatic switching: The system automatically identifies the type of target audio data being received and switches between the first and second modes based on the audio type. During the automatic identification of target audio data, the determination of the audio type relies on an analysis of the differences in the time domain, frequency domain, and statistical characteristics of the received audio.
[0072] In some embodiments, the first audio exhibits short-term non-stationary characteristics in the time domain, and its waveform shows obvious syllables and alternation of voiced and unvoiced segments. The second audio usually has a stronger periodicity or harmonic structure, especially in the steady-state instrument note part. In the frequency domain, the spectral energy of speech is concentrated in a lower frequency range (such as 300Hz-4kHz), and the formant structure is clear. Music may cover a wider frequency band (such as 20Hz-20kHz), with richer harmonic components and a stronger distribution regularity.
[0073] This application is based on two different coding modes to execute different logic processing flows, please refer to the attached manual Figure 1 As shown, an embodiment of a cochlear implant encoding method of the present application further includes the steps of:
[0074] S010: Use a high-pass filter to pre-emphasize the target audio data to obtain emphasized audio data. The target audio data includes a first audio and a second audio. The emphasized audio data includes a first emphasized audio and a second emphasized audio.
[0075] Specifically, the first audio and the second audio are voice data and music data respectively. The first emphasized audio and the second emphasized audio are pre-emphasized voice data and pre-emphasized music data respectively.
[0076] In this embodiment, pre-emphasis processing of the target audio data using a high-pass filter can reduce the impact of lip radiation, improve audio clarity, and optimize the perceived audio quality. Pre-emphasis can attenuate or remove low-frequency noise and interference, such as environmental hum, power line interference, or breathing noise. These noises are typically concentrated in the low-frequency band, and filtering them out can improve the clarity of the speech signal. Furthermore, high-pass filtering can compensate for the high-frequency attenuation characteristics of certain microphones or recording devices, making the target audio data sound more natural and balanced.
[0077] In one embodiment, the first audio is pre-emphasized to obtain the first emphasized audio. In another embodiment, the second audio is pre-emphasized to obtain the second emphasized audio.
[0078] S020: Divide the emphasized audio data into N audio channels using a frequency division filtering method.
[0079] In one implementation of this embodiment, in the first mode, the step S020 specifically includes: S021, using the method of Fourier transform to frequency domain to divide the weighted audio data into N audio channels. Specifically, the first weighted audio is converted from the time domain to the frequency domain using fast Fourier transform, and the spectrum is divided into N frequency bands according to the preset number of channels N and frequency range. Then, the spectrum of the N frequency bands is transformed back to the time domain through inverse Fourier transform to obtain the N audio channels, refer to the attached manual. Figure 2 .
[0080] In another implementation of this embodiment, in the second mode, step S020 specifically includes: S022, using a bidirectional filtering technique that performs two filtering steps, first in a forward direction and then in a reverse direction, to divide the weighted audio data into N audio channels. Specifically, a frequency division filter is used to divide the second weighted audio into N frequency band channels and filter them. The order of the channel data of the N frequency band channels after filtering is reversed, and the reversed channel data is filtered twice using the frequency division filter. The channel data after the secondary filtering is reversed a second time to restore the original channel order, thereby obtaining the N audio channels.
[0081] S030: Full-wave rectify the N audio channels to obtain absolute values of the data, and use a low-pass filter to obtain envelopes of each channel.
[0082] In this embodiment, after obtaining the time-domain signals of N channels, each channel's time-domain signal is first full-wave rectified. This involves taking the absolute value of the data and flipping the negative half-cycle to the positive half-cycle, positioning the signal completely above the time axis to facilitate subsequent envelope extraction. Next, the rectified signal is low-pass filtered to smooth out rapid fluctuations while retaining the envelope that reflects changes in signal energy. The low-frequency component output by each channel is the time-domain envelope of that channel, representing the energy variation of the original signal within that frequency band over time.
[0083] S040, extracting feature data from the audio data or the weighted audio data, and obtaining a feature envelope of the feature data using a low-pass filter. Optionally, in the first mode, extracting feature data from the weighted audio data includes: S041, a fundamental frequency band and a formant frequency band, and smoothing the formant frequency band based on the weighted audio data. Alternatively, in the second mode, extracting feature data from the target audio data includes: a fundamental signal and a harmonic signal, and smoothing the harmonic signal based on the target audio data. Obtaining a feature envelope of the feature data using a low-pass filter.
[0084] In one implementation of the present embodiment, in the first mode, the step S040 specifically includes: S042, using fast Fourier transform to convert the first emphasized audio from the time domain to the frequency domain. Based on the frequency domain decomposition method, the frequency domain signal is divided into a base frequency band and multiple resonance peak frequency bands. Then it is transformed to the time domain through inverse Fourier transform. Based on the first emphasized audio, the resonance peak frequency band is smoothed using the following formula: smoothed resonance peak frequency band = [a* resonance peak frequency band + (1-a)* first emphasized audio]. Wherein a is the weight coefficient. The negative offset of the smoothed resonance peak frequency band is removed by the half-wave rectification method, and then multiplied by the compensation coefficient to compensate for the energy loss of the half-wave rectification. Refer to the appendix of the manual. Figure 2 .
[0085] Based on the fundamental frequency band and the formant frequency band after half-wave rectification, low-pass filters with different cutoff frequencies are used to extract the fundamental frequency envelope and the formant envelope, and the fundamental frequency envelope and the formant envelope are normalized.
[0086] In another implementation of this embodiment, in the second mode, the step S040 specifically includes: converting the second audio from the time domain to the frequency domain by a fast Fourier transform method, analyzing the power spectrum peak within the specified frequency range to determine the fundamental range. Based on the fundamental range, multiple harmonic ranges are estimated by frequency multiplication calculation. The fundamental signal and the harmonic signal are then obtained by inverse Fourier transform. Based on the second audio, the harmonic signal is smoothed using the following formula: smoothed harmonic signal = [a*harmonic signal + (1-a)*second audio], where a is a weight coefficient. The negative offset of the smoothed harmonic signal is removed by a half-wave rectification method, and then multiplied by a compensation coefficient to compensate for the energy loss of the half-wave rectification.
[0087] Based on the fundamental wave signal and the harmonic signal after half-wave rectification, low-pass filters with different cutoff frequencies are used to extract the fundamental wave envelope and the harmonic wave envelope respectively. The fundamental wave envelope and the harmonic wave envelope are normalized respectively.
[0088] S050: Modulate the channel envelope based on the characteristic envelope to obtain a modulated envelope.
[0089] In one implementation of this embodiment, in the first mode, step S050 specifically includes: S051, multiplying the fundamental frequency envelope by all the channel envelopes point by point. Multiplying the formant envelope by several of its matching channel envelopes point by point to obtain the modulation envelope.
[0090] In another implementation of this embodiment, in the second mode, the step S050 specifically includes: S052, multiplying the fundamental envelope with all the channel envelopes point by point. Multiplying the harmonic envelope with several of the channel envelopes that match it point by point to obtain the modulation envelope, refer to the attached manual. Figure 2 .
[0091] S060: Downsample the modulation envelope and map the electrode array in the cochlear implant based on the downsampled data.
[0092] In this embodiment, the modulated data of each channel is downsampled and decimated, that is, a number of data points are extracted from each channel. Uniform sampling is performed according to a set decimation factor M, reducing the data rate to 1 / M of the original value. The downsampled data obtained by the downsampling process, i.e., the sampled modulation envelope data, can significantly reduce the data volume. This can improve system computational efficiency and reduce storage or transmission bandwidth, especially for multi-channel parallel processing systems such as cochlear implants or real-time speech encoders.
[0093] The modulated and decimated downsampled data needs to be further converted into electrical stimulation parameters. Specifically, the downsampled data will be mapped to the physical channels corresponding to the cochlear implant electrode array. The envelope amplitude of each channel will determine the electrical stimulation current intensity or pulse width of the electrode. At the same time, the timing characteristics of the signal are converted into the electrode stimulation rate or time pattern. By retaining the modulated envelope and key spectral features and restoring the perceptual cues of speech as much as possible, the hearing experience of hearing-impaired users can be optimized.
[0094] Based on the same concept, the present application also discloses a cochlear implant coding system. An embodiment of the cochlear implant coding system of the present application includes the following modules:
[0095] The mode switching module is configured to receive target audio data and manually or automatically switch the cochlear implant between a first mode and a second mode based on the type of the target audio data currently received.
[0096] The pre-processing module is used to perform pre-emphasis processing on the target audio data using a high-pass filter to obtain emphasized audio data after processing.
[0097] The frequency division filtering module is used to divide the emphasized audio data into N audio channels using a frequency division filtering method.
[0098] The envelope extraction module is used to obtain the absolute value of the data by using a full-wave rectification method for the N audio channels and to obtain the envelope of each channel by using a low-pass filter.
[0099] The feature extraction module is configured to, in the first mode, extract feature data from the emphasized audio data, including a fundamental frequency band and a formant frequency band, and perform smoothing on the formant frequency band based on the emphasized audio data. Alternatively, in the second mode, extract feature data from the target audio data, including a fundamental signal and a harmonic signal, and perform smoothing on the harmonic signal based on the target audio data.
[0100] Specifically, in some embodiments, the feature extraction module is configured to, in the speech mode, extract feature data from the pre-emphasized audio including a fundamental frequency band and a formant frequency band, and perform smoothing on the formant frequency band based on the pre-emphasized audio. Alternatively, in the music mode, extract feature data from the target audio including a fundamental signal and a harmonic signal, and perform smoothing on the harmonic signal based on the target audio.
[0101] The feature extraction module is further configured to obtain a feature envelope of the feature data using a low-pass filter.
[0102] The characteristic modulation module is used to modulate the channel envelope based on the characteristic envelope to obtain a modulation envelope.
[0103] A data sampling module is configured to perform downsampling processing on the modulation envelope and map the electrode array in the cochlear implant based on the downsampled data.
[0104] The following will discuss the various modules of the system and the steps of the method.
[0105] This application provides another embodiment of a cochlear implant coding system. Based on the above system embodiment, for the processing of the first mode, refer to the attached specification. Figure 3 As shown, the specific steps include:
[0106] S110 , the first audio enters a pre-processing module and undergoes pre-emphasis processing.
[0107] Specifically, a first-order high-pass filter is used to process the first audio signal to increase the high-frequency components of the speech, thereby obtaining a pre-emphasized first audio signal. Pre-emphasis can reduce the impact of lip radiation, improve speech clarity, and optimize the perceived quality of the speech. Optionally, the first audio signal is pre-emphasized speech data.
[0108] S120: The first emphasized audio enters the frequency division filtering module and performs data frequency division filtering.
[0109] Frequency division filtering: Using the fast Fourier transform to frequency domain method, the time domain signal is converted to the frequency domain and then divided into channels. Based on the preset number of channels N and frequency range, the spectrum is divided into N frequency bands, each corresponding to a channel. The division method can be uniform equal-width bands or non-uniform bands, such as those based on the Bark scale. The latter is more consistent with the human auditory characteristics. The spectral components of each channel are extracted using a rectangular window or a frequency domain filter with smooth transitions. The spectrum of each channel is then inverse Fourier transformed back to the time domain, ultimately resulting in N channels.
[0110] S130: Input the N channel data, i.e., the N time-domain frequency-divided signals obtained by frequency-dividing filtering, into an envelope extraction module to extract the channel envelopes. Specifically, a full-wave rectification method is used, and the absolute value of the data is taken. A low-pass filter is further used to obtain the envelopes of each channel.
[0111] S140: The first emphasized audio enters a feature envelope extraction module to extract a feature envelope.
[0112] Specifically, the process includes the following sub-steps: S141, Frequency Domain Decomposition: The first emphasized audio signal is converted to the frequency domain via Fast Fourier Transform (FFT) and divided into N key frequency bands, including the fundamental frequency F0, the first formant F1, the second formant F2, and the third formant F3. Each frequency band is then restored to the time domain via Inverse Fourier Transform (IFT), where the signal is separated into sub-band components containing the fundamental frequency and formant characteristics. High-frequency emphasis compensates for high-frequency attenuation in voice transmission, while frequency domain decomposition precisely isolates the fundamental frequency and formant for subsequent independent processing.
[0113] In some embodiments, the first emphasized audio is fast Fourier transformed into the frequency domain to obtain frequency band data of 80 Hz to 400 Hz, 300 Hz to 900 Hz, 800 Hz to 2500 Hz, and 1700 Hz to 3500 Hz, respectively. The four channel data are then fast Fourier transformed into the time domain. The four signal segments contain information about the fundamental frequency, first formant, second formant, and third formant corresponding to F0, F1, F2, and F3, respectively.
[0114] S142 uses the pre-emphasized first audio to smooth the four frequency bands: smoothed formant frequency band = [a * formant frequency band + (1 - a) * first emphasized audio]. a is a weighting factor. Optionally, a is 0.7. Smoothing the formant frequency band data can highlight the stable structure of the band's formants while increasing the non-formant speech information, thereby smoothing the transition between frequency bands during subsequent envelope extraction.
[0115] S143 uses half-wave rectification to remove the negative offset and then multiplies the signal by a compensation factor to compensate for the energy loss caused by half-wave rectification. Optionally, the compensation factor is 1.7. Compared to full-wave rectification, half-wave rectification can reduce high-frequency harmonic interference and focus the signal's primary energy concentration, making it particularly suitable for preserving the envelope shape of the fundamental frequency and resonance peak.
[0116] S144, applying a low-pass filter to the half-wave rectified data to extract the time domain envelopes of the fundamental frequency and the resonance peak respectively.
[0117] In some embodiments, the half-wave rectified fundamental frequency data is passed through a low-pass filter with a cutoff frequency of 400 Hz to obtain the fundamental frequency envelope, i.e., the F0 envelope. The half-wave rectified first formant data is passed through a low-pass filter with a cutoff frequency of 900 Hz to obtain the first formant envelope, i.e., the F1 envelope. The half-wave rectified second formant data is passed through a low-pass filter with a cutoff frequency of 2500 Hz to obtain the second formant envelope, i.e., the F2 envelope. The half-wave rectified third formant data is passed through a low-pass filter with a cutoff frequency of 3500 Hz to obtain the third formant envelope, i.e., the F3 envelope.
[0118] It helps to filter out high-frequency carrier components and retain the slow-changing features that reflect vocal cord vibration (fundamental frequency) and vocal tract tuning (resonant peaks). The feature envelope is the core parameter of speech perception.
[0119] S145: Normalize the fundamental frequency envelope and the formant envelope. Specifically, the F0, F1, F2, and F3 data are normalized to eliminate the impact of amplitude differences on subsequent modulation and improve feature consistency across different speakers or environments.
[0120] S150: Input the extracted channel envelope and feature envelope into a feature modulation module for feature modulation.
[0121] The fundamental frequency envelope is multiplied point by point with the envelopes of all channels, and the resonance peak envelope is multiplied point by point with the envelopes of several matching channels to generate the modulated channel envelope.
[0122] Specifically, the following steps are included: S151, multiplying the F0 envelope with the envelope of each channel data after frequency division filtering.
[0123] S152: Multiply the F1 envelope by the matched channel data envelope after frequency division filtering. The matching principle is that the F1 bandpass filter range is 300Hz to 900Hz, and any channel data envelope that overlaps with the 300Hz to 900Hz range is considered a matched channel. Assume that the channel data envelope ranges are [100200], [200300], [300500], [500800], [8001200], [12001600], [16002000], [20002500], [25003000], [30003500], etc. The F1 envelope is multiplied by the channel envelopes of [100200], [200300], [300500], [500800], and [8001200].
[0124] In step S153, the F2 envelope is multiplied by the matched channel data envelope after frequency division filtering. The matching principle is that the F2 bandpass filter range is 800Hz to 2500Hz, and the corresponding channel data envelopes that intersect with the 800Hz to 2500Hz range are all matched channels. Based on the assumptions in step S152, the F2 envelope is multiplied by the channel envelopes of [800 1200], [1200 1600], [1600 2000], and [2000 2500].
[0125] In step S154, the F3 envelope is multiplied by the matched channel data envelope after frequency division filtering. The matching principle is that the F3 bandpass filter range is 1700Hz to 3500Hz, and the corresponding channel data envelopes that intersect with the 1700Hz to 3500Hz range are all matched channels. Based on the assumptions in step S152, the F2 envelope is multiplied by the channel envelopes [16002000], [20002500], [25003000], and [30003500].
[0126] S160: The modulation envelope obtained after modulation is input into a data sampling module for downsampling processing.
[0127] The modulated data for each channel is downsampled and extracted. The sampled envelope data for each channel, such as the time-domain energy information after fundamental frequency and formant modulation, is mapped to the corresponding physical channel of the cochlear implant electrode array. The envelope amplitude of each channel determines the electrical stimulation current intensity or pulse width of that electrode.
[0128] This application discloses another embodiment of a cochlear implant coding system. For the processing of the second mode, refer to the attached specification. Figure 4 As shown. Specifically including the steps:
[0129] S210: The second audio enters the pre-processing module and undergoes pre-emphasis processing.
[0130] Specifically, a first-order high-pass filter is used to increase high-frequency components, generating the second pre-emphasized audio. Because cochlear implant recipients typically experience greater high-frequency loss than low-frequency loss, pre-emphasis can reduce the impact of lip radiation, improve speech clarity, and optimize the perceived quality of the music. Optionally, the second audio is music data that has not undergone pre-emphasis.
[0131] S220: The second emphasized audio enters the frequency division filtering module and performs data frequency division filtering.
[0132] The second emphasized audio is divided into N channels using a frequency division filtering method. The N channel data is then reversed and filtered using the same filter. The filter output is then reversed. This preserves the data's phase information, facilitating the expression of musical melody. The second audio is divided into N frequency band channels, each covering a specific frequency range, such as non-uniformly divided according to the Bark scale.
[0133] Subsequently, the data order of each of the N channels is completely reversed, that is, the first sampling point of each channel is swapped with the Nth sampling point, and the second sampling point is swapped with the N-1th sampling point, with N being an optional N = 256. The same analogy is applied to the remaining channels. The same frequency division filter bank is then applied to the reversed channel data for filtering. The first and second filtering processes correct the phase distortion introduced by the initial filtering, as the combination of the reverse operation and the forward filtering is equivalent to zero-phase filtering. Finally, the filter output data is reversed again to restore the original channel order. At this point, the phase characteristics of each channel signal are preserved, and the integrity of the frequency band division is not compromised.
[0134] At step S230, the N channel data, i.e., the N channel time domain signals obtained by frequency division and filtering, are input into an envelope extraction module to extract the channel envelope. First, the time domain signal of each channel is full-wave rectified, i.e., its absolute value is taken. Next, the rectified signal is low-pass filtered. The low-frequency component output by each channel is the channel envelope of that channel.
[0135] S240: The initial second audio enters a feature extraction module to extract a feature envelope.
[0136] Specifically, the process includes the following sub-steps: S241 , first converting the second audio data from the time domain to the frequency domain using a fast Fourier transform, analyzing the power spectrum peak within a specified frequency range to determine the fundamental frequency. This allows for accurate location of the dominant periodic component of the signal, providing a benchmark for subsequent harmonic analysis.
[0137] In some embodiments, the second audio is transformed to the frequency domain through a fast Fourier transform (FFT). The frequency F corresponding to the power peak is found within a frequency range less than 1500 Hz. To encompass the timbres of more instruments, the frequency range between F*0.7 and F*1.3 is used as the prominent fundamental frequency range for each frame of music. If F*1.3 is greater than 1500 Hz, the upper limit of the prominent fundamental frequency range is simply set to 1500. Assuming the peak frequency F = 800 Hz, the fundamental frequency range is [800*0.7800*1.3] = [5601040].
[0138] S242, based on the fundamental wave range, the system estimates the harmonic frequency band through frequency doubling calculation and extracts the spectral components of the corresponding frequency band, and then restores it to the time domain harmonic signal through inverse Fourier transform. This operation separates the harmonic structure of the signal and helps to finely characterize the timbre characteristics.
[0139] In some embodiments, an estimated harmonic frequency band below 8000 Hz for the current frame of data (each frame of data has 256 sampling points) is calculated based on the fundamental wave range, and harmonic signals are extracted based on the frequency band. The specific harmonic calculation method is as follows: Assuming the fundamental wave range is [560 × 1040], its second harmonic range is [560 × 2 × 1040 × 2], and its third harmonic range is [560 × 3 × 1040 × 3]. N harmonic ranges are calculated sequentially until the upper limit of the highest harmonic is greater than 8000 Hz. In this example, the highest harmonic calculated is the 7th harmonic because the upper limit of the 8th harmonic is greater than 8000 Hz.
[0140] S243 , using the second audio frequency, performs time-domain smoothing on all harmonic signals. The smoothed harmonic signal = [a * harmonic signal + (1 - a) * second audio frequency], where a is a weighting factor. The optional value of a is 0.7. Smoothing helps to achieve smoother transitions between harmonics during envelope extraction.
[0141] S244 uses half-wave rectification to remove the negative offset and then multiplies the signal by a compensation factor to compensate for half-wave rectification energy loss. Optionally, the compensation factor is 0.7. Half-wave rectification removes the DC offset and simplifies envelope extraction by truncating the negative portion of the signal while retaining the positive half-cycle.
[0142] S245 calculates the envelope of each harmonic and performs filtering using the upper limit of each harmonic signal's frequency band as the cutoff frequency of a low-pass filter. The envelope of each harmonic signal is obtained by low-pass filtering with the upper limit of the harmonic frequency band as the cutoff frequency. This adaptively preserves the dynamic characteristics of each harmonic, avoiding the loss of detail that would result from a uniform cutoff frequency.
[0143] In some embodiments, the envelope of each harmonic is calculated, and the upper limit of each harmonic signal frequency range is used as the cutoff frequency of a low-pass filter for filtering. For example, if the fundamental frequency range is [560-1040], a low-pass filter with a cutoff frequency of 1040 Hz is used. If the first harmonic frequency range is [1120-2080], a low-pass filter with a cutoff frequency of 2080 Hz is used. The Nth harmonic is processed similarly.
[0144] S246: Normalization processing is performed on the smoothed harmonic signal. Normalization processing eliminates amplitude differences between different harmonics, making the envelope data more suitable for subsequent modulation.
[0145] S250: Input the extracted channel envelope and feature envelope into a feature modulation module to perform feature modulation.
[0146] The fundamental frequency envelope is multiplied point by point with the envelopes of all channels, and the harmonic envelope is multiplied point by point with the envelopes of several channels that match it, to generate a modulated channel envelope.
[0147] In some implementations, the process specifically includes the following step S251: multiplying the fundamental wave envelope F0 with each channel data after frequency division filtering.
[0148] S252: Multiply the first harmonic envelope F1 by the matched channel data. The matching principle is as follows: Assume the first harmonic range is [1120-2080]. Channel data envelopes that intersect with the range 1120 to 2080 are matched channels. Assume the channel data envelope ranges are [100-200], [200-300], [300-500], [500-800], [800-1200], [1200-1600], [1600-2000], [2000-2500], [2500-3000], [3000-3500], etc. The first harmonic envelope F1 is multiplied by the channel envelopes of [800-1200], [1200-1600], [1600-2000], and [2000-2500].
[0149] S253: Process the second harmonic envelope to the Nth harmonic envelope in sequence. The process is the same as step S252 and will not be described in detail here.
[0150] S260: The modulation envelope obtained after modulation is input into a data sampling module for downsampling processing.
[0151] The modulated data for each channel is downsampled and decimated. The sampled envelope data for each channel, such as the time-domain energy information after the fundamental and harmonic waves, is mapped to the corresponding physical channel of the cochlear implant electrode array. The envelope amplitude of each channel determines the electrical stimulation current intensity or pulse width of that electrode.
[0152] Based on the same concept, this application also discloses a cochlear implant device, including a cochlear implant coding system as described in any of the above embodiments. The cochlear implant includes a first mode and a second mode, each processing speech data according to its characteristics. This can significantly improve the user's auditory experience, particularly in speech intelligibility, tonal language recognition, and music perception.
[0153] The cochlear implant encoding method, system and cochlear implant device of the present application have the same technical concept, and the technical details of the embodiments of the two are applicable to each other. To reduce repetition, they will not be described again.
[0154] Those skilled in the art will clearly understand that, for the sake of convenience and brevity of description, only the division of the above-mentioned program modules is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program units or modules to complete all or part of the functions described above. The program modules in the embodiment can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one processing unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software program unit. In addition, the specific names of the program modules are only for the purpose of distinguishing each other and are not used to limit the scope of protection of this application.
[0155] Although the preferred embodiments of the present application have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present application.
Claims
1. A cochlear implant encoding method, characterized in that: The invention is applicable to a cochlear implant, wherein the cochlear implant includes a first mode and a second mode; and comprises the following steps: receiving target audio data, and performing pre-emphasis processing on the target audio data using a high-pass filter to obtain emphasized audio data after the processing; Dividing the emphasized audio data into N audio channels using a frequency division filtering method; The N audio channels are subjected to full-wave rectification to obtain absolute values of the data; And use low-pass filter to obtain the envelope of each channel; In the first mode, the feature data extracted from the weighted audio data includes: a fundamental frequency band and a formant frequency band, and the formant frequency band is smoothed based on the weighted audio data; or, in the second mode, the feature data extracted from the target audio data includes: a fundamental signal and a harmonic signal, and the harmonic signal is smoothed based on the target audio data; Using a low-pass filter to obtain a characteristic envelope of the characteristic data; Modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope; The modulation envelope is downsampled; and an electrode array in the cochlear implant is mapped based on the downsampled data.
2. A cochlear implant encoding method according to claim 1, characterized in that: Also includes: manually or automatically switching the cochlear implant between a first mode and a second mode based on the type of the target audio data currently being received; The target audio data includes a first audio and a second audio; and the emphasized audio data includes a first emphasized audio and a second emphasized audio.
3. A cochlear implant encoding method according to claim 2, characterized in that: In the first mode, the frequency division filtering method is used to divide the emphasized audio data into N audio channels; Specifically include: Converting the first emphasized audio from the time domain to the frequency domain using a fast Fourier transform, and dividing the spectrum into N frequency bands according to a preset number of channels N and a frequency range; The frequency spectra of the N frequency bands are then transformed back to the time domain through inverse Fourier transform to obtain the N audio channels.
4. A cochlear implant encoding method according to claim 3, characterized in that: In the first mode, extracting feature data from the emphasized audio data includes: a fundamental frequency band and a formant frequency band, and performing smoothing processing on the formant frequency band based on the emphasized audio data; specifically, the process includes: Converting the first emphasized audio from the time domain to the frequency domain using a fast Fourier transform; dividing the frequency domain signal into a fundamental frequency band and a plurality of formant frequency bands based on a frequency domain decomposition method; and then converting the signal to the time domain using an inverse Fourier transform; Based on the first emphasized audio, the formant frequency band is smoothed using the following formula: Smoothed formant frequency band = [a*formant frequency band + (1-a)*first emphasized audio]; where a is the weight coefficient; A half-wave rectification method is used to remove the negative offset of the smoothed resonance peak frequency band, and then the negative offset is multiplied by a compensation coefficient to compensate for the energy loss of the half-wave rectification.
5. A cochlear implant encoding method according to claim 4, characterized in that: In the first mode, the method of obtaining a characteristic envelope of the characteristic data by using a low-pass filter and modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope specifically includes: Based on the fundamental frequency band and the formant frequency band after half-wave rectification, using low-pass filters with different cutoff frequencies to extract the fundamental frequency envelope and the formant envelope respectively; normalizing the fundamental frequency envelope and the formant envelope respectively; The fundamental frequency envelope is multiplied point by point with all the channel envelopes; and the formant envelope is multiplied point by point with several channel envelopes that match it to obtain the modulation envelope.
6. A cochlear implant encoding method according to claim 2, characterized in that: In the second mode, the method of dividing the emphasized audio data into N audio channels using a frequency division filtering method includes: Divide the second emphasized audio into N frequency band channels using a frequency division filter and perform filtering processing; Reversely arranging the order of the channel data of the N frequency band channels after filtering, and performing secondary filtering processing on the reversely arranged channel data using the frequency division filter; The channel data after the secondary filtering is reversely arranged twice to restore the original channel order to obtain the N audio channels.
7. A cochlear implant encoding method according to claim 6, characterized in that: In the second mode, extracting feature data from the target audio data includes: fundamental wave signals and harmonic signals, and smoothing the harmonic signals based on the target audio data; specifically, the steps include: Converting the second audio from the time domain to the frequency domain using a fast Fourier transform method, and analyzing the power spectrum peak within a specified frequency range to determine the fundamental wave range; Based on the fundamental wave range, multiple harmonic ranges are estimated by frequency multiplication calculation; and then the fundamental wave signal and the harmonic signal are obtained by inverse Fourier transformation; Based on the second audio frequency, the harmonic signal is smoothed using the following formula: Smoothed harmonic signal = [a*harmonic signal + (1-a)*second audio], where a is the weight coefficient; A half-wave rectification method is used to remove the negative offset of the smoothed harmonic signal, and then the negative offset is multiplied by a compensation coefficient to compensate for the energy loss of the half-wave rectification.
8. A cochlear implant encoding method according to claim 7, characterized in that: In the second mode, the step of obtaining a characteristic envelope of the characteristic data using a low-pass filter and modulating the channel envelope based on the characteristic envelope to obtain a modulated envelope specifically includes: Based on the fundamental wave signal and the harmonic signal after half-wave rectification, using low-pass filters with different cutoff frequencies to extract the fundamental wave envelope and the harmonic wave envelope respectively; normalizing the fundamental wave envelope and the harmonic wave envelope respectively; The fundamental wave envelope is multiplied point by point with all the channel envelopes; and the harmonic envelope is multiplied point by point with several channel envelopes that match it to obtain the modulation envelope.
9. A cochlear implant coding system, characterized in that: include: a mode switching module, configured to receive target audio data and manually or automatically switch the cochlear implant between a first mode and a second mode based on the type of the target audio data currently received; A pre-processing module, configured to perform pre-emphasis processing on the target audio data using a high-pass filter to obtain emphasized audio data after processing; A frequency division filtering module, configured to divide the weighted audio data into N audio channels using a frequency division filtering method; An envelope extraction module, configured to obtain absolute values of the N audio channels using a full-wave rectification method; And use low-pass filter to obtain the envelope of each channel; a feature extraction module configured to, in the first mode, extract feature data from the weighted audio data, including a fundamental frequency band and a formant frequency band, and perform smoothing on the formant frequency band based on the weighted audio data; or, in the second mode, extract feature data from the target audio data, including a fundamental signal and a harmonic signal, and perform smoothing on the harmonic signal based on the target audio data; The feature extraction module is further configured to obtain a feature envelope of the feature data using a low-pass filter; A characteristic modulation module, configured to modulate the channel envelope based on the characteristic envelope to obtain a modulation envelope; A data sampling module is used to perform downsampling processing on the modulation envelope; and map the electrode array in the cochlear implant based on the downsampling data.
10. A cochlear implant device, characterized in that: Including a cochlear implant coding system as described in claim 9.