Cochlear implant sound signal processing method, device and equipment and readable storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
但是,这种方法很难精确表达F0,且电流脉冲之间容易产生干扰,造成低频语音编码效果不佳
[0061]上述人工耳蜗声音信号处理方法、装置、计算机设备、计算机可读存储介质和计算机程序产品,通过将声音通道划分为第一类通道和第二类通道,所述第一类通道覆盖语音基频范围内的频段,所述第二类通道覆盖语音基频范围之外的频段;从而可以将声音信号的接收通道分为低频和中高频,便于后续进行针对性处理,提高第一类通道的基频参数提取准确性。按照第一方式对所述第一类通道的每一帧数据进行数据抽取处理,得到第一音频数据;按照第二方式对所述第二类通道的每一帧数据进行数据抽取处理,得到第二音频数据;从而可以针对不同类的通道采用不同的数据抽取策略,防止低频通道的刺激速率过高而导致基频信息被过采样,继而导致数据模糊或丢失。将第一音频数据和第二音频数据进行非线性压缩后,分别映射为第一电流脉冲序列和第二电流脉冲序列;对所述第一电流脉冲序列进行重新排列,得到重排后的第一电流脉冲序列;从而可以确保低频语音信息(特别是基频F0)的保真度与刺激时序的优化。按照所述重排后的第一电流脉冲序列和所述第二电流脉冲序列进行脉冲发放。从而可以在有限的通道数和刺激速率下,通过对第一类通道实施低于第二类通道的抽取率,实现了刺激脉冲资源在频域上的非均匀分配。使得更多的脉冲资源分配给需要更高时间分辨率的频段,从而舍弃对F0的精确计算,在现有的人工耳蜗算法框架下,一方面把F0固定在前2至3个通道,并采用递增的、分数倍于基准刺激速率的方式,满足人工耳蜗低频锁相理论;另一方面,对于低频通道的脉冲发放方式,在时间和空间两个维度安排刺激脉冲时序,进一步降低低频电极之间的脉冲互扰作用,确保了低频言语编码的效果。
Smart Images

Figure CN122551809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical device technology, and in particular to a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for processing sound signals from a cochlear implant. Background Technology
[0002] A cochlear implant is an electronic device used to compensate for sensorineural hearing loss, and it is widely used for auditory reconstruction in people with severe or higher hearing impairment. Its basic principle is to pick up external sound signals, analyze and process them, and output electrical current to directly act on the auditory nerve, bypassing the damaged inner ear structures, to achieve auditory function replacement.
[0003] Traditional cochlear implant sound coding techniques typically involve delivering current pulses at the same rate from the electrodes corresponding to low and high frequencies, or using fundamental frequency (F0) modulation for low-frequency channel data processing. However, this method struggles to accurately represent F0, and interference between current pulses easily occurs, resulting in poor low-frequency speech coding performance. Summary of the Invention
[0004] Therefore, it is necessary to provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for processing cochlear implant sound signals that can improve the accuracy of F0 parameter extraction and reduce mutual interference of current pulses between low-frequency electrodes, in order to address the above-mentioned technical problems.
[0005] In a first aspect, this application provides a method for processing sound signals from a cochlear implant, the method comprising:
[0006] The audio channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0007] Data extraction processing is performed on each frame of data from the first type of channel according to the first method to obtain the first audio data;
[0008] The second type of channel data is processed by data extraction according to the second method to obtain the second audio data;
[0009] After nonlinear compression of the first audio data and the second audio data, they are mapped into a first current pulse sequence and a second current pulse sequence, respectively.
[0010] The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence;
[0011] Pulse delivery is performed according to the rearranged first current pulse sequence and the second current pulse sequence.
[0012] In one embodiment, the step of extracting data from each frame of the first type of channel according to a first method to obtain first audio data includes:
[0013] Determine the number of sampling points for each of the first-class channels;
[0014] Based on the data extraction ratio constraint, determine the extraction parameters corresponding to each frame of data;
[0015] Based on the extraction parameters and the number of sampling points corresponding to each first type of channel, data extraction processing is performed according to the root mean square algorithm or the mean algorithm to obtain the first audio data.
[0016] Wherein, if the first k channels in each frame of data are the first type of channels, and k is a natural number greater than 1, then the number of sampling points of the first type of channels must satisfy the following constraint:
[0017] The number of sampling points in the first k channels increases incrementally, and the number of sampling points in the kth channel is not greater than the maximum value of the fundamental frequency.
[0018] In one embodiment, the incremental relationship includes:
[0019] From the first channel to the kth channel, the number of sampling points in each channel increases exponentially.
[0020] In one embodiment, the step of extracting data from each frame of the second type of channel according to the second method to obtain the second audio data includes:
[0021] The number of sampling points for each second-class channel is determined based on the channel stimulus rate, data frame length, and audio sampling frequency corresponding to each second-class channel.
[0022] Based on the data extraction ratio constraint, determine the extraction parameters corresponding to each frame of data;
[0023] Based on the extraction parameters and the number of sampling points corresponding to each second type of channel, data extraction processing is performed according to the root mean square algorithm or the mean algorithm to obtain the second audio data.
[0024] In one embodiment, rearranging the first current pulse sequence to obtain the rearranged first current pulse sequence includes:
[0025] The time interval between the first current pulses in the first current pulse sequence is adjusted to satisfy the following first condition:
[0026] Within the same data frame and adjacent data frames, the time distance between each adjacent first current pulse is the largest; and / or, the number of stimulation cycles across different channels by adjacent first current pulses is the largest.
[0027] In one embodiment, rearranging the first current pulse sequence to obtain the rearranged first current pulse sequence includes:
[0028] The time interval between the first current pulses in the first current pulse sequence is adjusted to satisfy the following second condition:
[0029] Within the same data frame and adjacent data frames, the electrode spacing between the first current pulses emitted by each adjacent electrode is the largest.
[0030] In one embodiment, before performing data extraction processing on each frame of data from the first type of channel according to the first method, and before performing data extraction processing on each frame of data from the second type of channel according to the second method, the method further includes:
[0031] The audio signals input to each audio channel are filtered by a bandpass filter to obtain the filtered audio signals.
[0032] The filtered audio signal is enveloped and processed by full-wave rectification and low-pass filtering to obtain multi-frame data to be processed.
[0033] Secondly, this application also provides a cochlear implant sound signal processing device, the device comprising:
[0034] The channel division module is used to divide the audio channel into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0035] The first extraction module is used to perform data extraction processing on each frame of data of the first type of channel in a first manner to obtain the first audio data.
[0036] The second extraction module is used to extract data from each frame of data of the second type of channel according to the second method to obtain the second audio data.
[0037] The mapping module is used to non-linearly compress the first audio data and the second audio data, and then map them into a first current pulse sequence and a second current pulse sequence, respectively.
[0038] The reordering module is used to rearrange the first current pulse sequence to obtain the rearranged first current pulse sequence.
[0039] The pulse delivery module is used to deliver pulses according to the rearranged first current pulse sequence and the second current pulse sequence.
[0040] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0041] The audio channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0042] Data extraction processing is performed on each frame of data from the first type of channel according to the first method to obtain the first audio data;
[0043] The second type of channel data is processed by data extraction according to the second method to obtain the second audio data;
[0044] After nonlinear compression of the first audio data and the second audio data, they are mapped into a first current pulse sequence and a second current pulse sequence, respectively.
[0045] The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence;
[0046] Pulse delivery is performed according to the rearranged first current pulse sequence and the second current pulse sequence.
[0047] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0048] The audio channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0049] Data extraction processing is performed on each frame of data from the first type of channel according to the first method to obtain the first audio data;
[0050] The second type of channel data is processed by data extraction according to the second method to obtain the second audio data;
[0051] After nonlinear compression of the first audio data and the second audio data, they are mapped into a first current pulse sequence and a second current pulse sequence, respectively.
[0052] The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence;
[0053] Pulse delivery is performed according to the rearranged first current pulse sequence and the second current pulse sequence.
[0054] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0055] The audio channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0056] Data extraction processing is performed on each frame of data from the first type of channel according to the first method to obtain the first audio data;
[0057] The second type of channel data is processed by data extraction according to the second method to obtain the second audio data;
[0058] After nonlinear compression of the first audio data and the second audio data, they are mapped into a first current pulse sequence and a second current pulse sequence, respectively.
[0059] The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence;
[0060] Pulse delivery is performed according to the rearranged first current pulse sequence and the second current pulse sequence.
[0061] The aforementioned cochlear implant sound signal processing method, device, computer equipment, computer-readable storage medium, and computer program product divide the sound channel into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the fundamental frequency range of speech, while the second type of channel covers the frequency band outside the fundamental frequency range of speech. This allows the sound signal receiving channel to be divided into low-frequency and mid-to-high-frequency channels, facilitating subsequent targeted processing and improving the accuracy of fundamental frequency parameter extraction for the first type of channel. Data extraction processing is performed on each frame of data from the first type of channel according to a first method to obtain first audio data; data extraction processing is performed on each frame of data from the second type of channel according to a second method to obtain second audio data. This allows different data extraction strategies to be used for different types of channels, preventing excessively high stimulation rates in low-frequency channels from causing oversampling of fundamental frequency information, which could lead to data blurring or loss. After nonlinear compression of the first and second audio data, they are mapped to a first current pulse sequence and a second current pulse sequence, respectively. The first current pulse sequence is rearranged to obtain a rearranged first current pulse sequence. This ensures the fidelity of low-frequency speech information (especially the fundamental frequency F0) and optimizes the stimulation timing. Pulse delivery is performed according to the rearranged first and second current pulse sequences. This allows for a non-uniform distribution of stimulation pulse resources in the frequency domain by implementing a lower decimation rate for the first type of channels compared to the second type, within a limited number of channels and stimulation rate. This allocates more pulse resources to frequency bands requiring higher time resolution, thus abandoning the precise calculation of F0. Within the existing cochlear implant algorithm framework, on the one hand, F0 is fixed in the first 2 to 3 channels, and an incremental stimulation rate, fractionally multiple of the baseline stimulation rate, is used to satisfy the low-frequency phase-locked loop theory of cochlear implants; on the other hand, for the pulse delivery method of low-frequency channels, the stimulation pulse timing is arranged in both time and space dimensions to further reduce pulse interference between low-frequency electrodes, ensuring the effectiveness of low-frequency speech coding. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 This is a flowchart illustrating a method for processing sound signals from an artificial cochlear implant in one embodiment.
[0064] Figure 2 This is a schematic diagram of rearrangement of the first current pulse sequence in one embodiment;
[0065] Figure 3 This is a flowchart illustrating the cochlear implant sound signal processing method in another embodiment;
[0066] Figure 4 A schematic diagram illustrating the principle of a cochlear implant sound signal processing method according to an embodiment of this application;
[0067] Figure 5 This is a schematic diagram illustrating the effect of maximizing the data distance in the first channel in one embodiment;
[0068] Figure 6 This is a schematic diagram illustrating the effect of maximizing the data distance in the second channel in one embodiment;
[0069] Figure 7 This is a schematic diagram illustrating the effect of maximizing the data distance in the third channel in one embodiment;
[0070] Figure 8 This is a schematic diagram of the data arrangement and pulse emission timing of 16 channels in one embodiment;
[0071] Figure 9 This is a schematic diagram illustrating the effect of maximizing the data distance in the first channel in another embodiment;
[0072] Figure 10 This is a schematic diagram illustrating the effect of maximizing the data distance in the second channel in another embodiment;
[0073] Figure 11 This is a schematic diagram illustrating the effect of maximizing the data distance in the third channel in another embodiment;
[0074] Figure 12 This is a schematic diagram of the data arrangement and pulse delivery timing of 16 channels in another embodiment;
[0075] Figure 13 This is a waveform diagram of a speech signal in one embodiment;
[0076] Figure 14 This is a spectrogram of a speech signal in one embodiment;
[0077] Figure 15 This is a schematic diagram of the waveform output after data extraction in one embodiment;
[0078] Figure 16 An electrode diagram from one embodiment;
[0079] Figure 17 This is a schematic diagram of the waveform output after data extraction in another embodiment;
[0080] Figure 18 Electrode diagram in another embodiment;
[0081] Figure 19 This is a schematic diagram of the data arrangement and pulse emission timing of 16 channels in another embodiment;
[0082] Figure 20 This is a structural block diagram of the cochlear implant sound signal device in one embodiment;
[0083] Figure 21 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0084] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0085] It should be noted that the terms "first," "second," etc., used in this application can be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish the first element from the second element. The terms "comprising" and "having," and any variations thereof, used in this application, are intended to cover non-exclusive inclusion. The term "multiple" used in this application refers to two or more. The term "and / or" used in this application refers to one of the embodiments, or any combination of multiple embodiments.
[0086] To facilitate understanding of the technical solutions in the various embodiments of this application, a brief description of the relevant technologies is provided first.
[0087] The main sound coding methods used in cochlear implants include n-of-m (n-of-m) peak selection and Continuous Interleaved Sampling (CIS). The principle of n-of-m is to select only the n channels with the largest spectral amplitudes from multiple analysis channels and send pulses to their corresponding electrodes. The principle of CIS is that all channels (electrodes) are active, and each channel extracts the envelope of its corresponding frequency band, then modulates this envelope using a fixed, high-rate pulse sequence (e.g., 800-2000 pps).
[0088] It should be understood that the common feature of the two sound encoding methods mentioned above is that regardless of whether the sound is low or high, the corresponding electrode output current pulses are emitted at the same rate.
[0089] Another technique used in cochlear implants is to employ fundamental frequency (F0) modulation in low-frequency channel data processing. However, this method also has two problems: the first problem is that F0 extraction is inaccurate, especially in noisy environments, where F0 performance will be greatly reduced; the second problem is that the center frequency and channel stimulation rate of each electrode in traditional cochlear implants are fixed, so even if F0 is calculated very accurately, the working mechanism of existing cochlear implant systems cannot accurately express the value of F0.
[0090] To address the problems existing in related technologies, this application aims to provide a cochlear implant sound signal processing method. This method abandons the precise calculation of F0 and, within the existing cochlear implant algorithm framework, on the one hand, fixes F0 in the first 2 to 3 channels and adopts an incremental method with a stimulation rate that is fractional times greater than the reference stimulation rate to satisfy the low-frequency phase-locked loop theory of cochlear implants; on the other hand, for the pulse delivery method of the low-frequency channels, the stimulation pulse timing is arranged in both time and space dimensions to further reduce the pulse interference between low-frequency electrodes.
[0091] The human ear's recognition of the F0 characteristic signal in speech signals, and further recognition of tonal languages (such as Mandarin), is of great significance. Since the frequency band of the F0 characteristic signal in speech signals is typically no greater than 500Hz, cochlear implants usually use the center frequencies of the first 2 to 3 channels to cover the F0 frequency range. Therefore, the focus can be placed on low-frequency channel data processing, thus abandoning the precise calculation of F0. Within the existing cochlear implant algorithm framework, on the one hand, F0 is fixed in the first 2 to 3 channels, and an incremental stimulation rate fractionally higher than the baseline stimulation rate is used to satisfy the low-frequency phase-locked loop theory of cochlear implants; on the other hand, for the pulse delivery method of the low-frequency channels, the stimulation pulse timing is arranged in both temporal and spatial dimensions to further reduce pulse interference between low-frequency electrodes.
[0092] In one exemplary embodiment, such as Figure 1 As shown, a method for processing sound signals from a cochlear implant is provided, which may include the following steps S101 to S106. Wherein:
[0093] Step S101: Divide the audio channels into a first type of channel and a second type of channel.
[0094] The first type of channel covers frequency bands within the voice baseband range, while the second type of channel covers frequency bands outside the voice baseband range.
[0095] In this embodiment, the sound signals input by the cochlear implant are divided into a first type of channel and a second type of channel according to the sound frequency spectrum range. Among them, the first type of channel can also be understood as the fundamental frequency channel, which is used to cover the fundamental frequency range of speech (usually about 80 - 150 Hz for men and about 150 - 300 Hz for women). This part of the information directly corresponds to the pitch and intonation of speech. The second type of channel can also be understood as the non-fundamental frequency channel, which is used to cover the frequency bands outside the fundamental frequency range and mainly carries the formants, consonant information and high-frequency details of speech, and determines the clarity and intelligibility of speech.
[0096] As an optional example, if the audio sampling frequency of the cochlear implant system is fs, the data frame length is L, and it is divided into m channels, then the low-frequency band within the speech F0 range (usually less than 500 Hz) occupies k (k < m) channels (the first type of channel), and the channels from k + 1 to m are the second type of channels.
[0097] Step S102, perform data extraction processing on each frame of data of the first type of channel according to the first method to obtain the first audio data.
[0098] In this embodiment, for the first type of channel, its main purpose is to accurately capture and retain the periodic timing information of the sound waveform within the frequency band, that is, the fine time structure of F0.
[0099] Exemplarily, determine the number of sampling points of each first type of channel respectively; determine the extraction parameters corresponding to each frame of data according to the data extraction ratio constraint condition; perform data extraction processing according to the extraction parameters and the number of sampling points corresponding to each first type of channel according to the root mean square algorithm or the mean algorithm to obtain the first audio data; where, if the first k channels in each frame of data are the first type of channels and k is a natural number greater than 1, the number of sampling points of the first type of channel needs to meet the following constraint conditions: the number of sampling points of the first k channels is in an increasing relationship, and the number of sampling points of the kth channel is not greater than the maximum value of the fundamental frequency.
[0100] In this embodiment, the extraction parameters can be determined through the data extraction ratio formula, and the data extraction ratio formula is as follows:
[0101]
[0102]
[0103] In the formula: L represents the data frame length, represents the average value (taking a positive integer value), represents the i-th sample value, represents the number of sampling points of the first type of channel, represents the number of sampling points of the second type of channel, k represents the number of channels of the first type of channel, and m represents the total number of channels of the first type of channel and the second type of channel.
[0104] Among them, for frame length of A frame of data, before or The sampling ratio of each sample value is That is, each One data point is extracted from each sample value; the extraction ratio of the last sample value is... That is, the last Each sample value is extracted into one data point.
[0105] In order to make and To make them as close as possible, let the difference between them be D, and then calculate D. min Minimum value:
[0106]
[0107] As an alternative example, the above increasing relationship may include: starting from the 1st channel to the kth channel, the number of sampling points in each channel increases exponentially.
[0108] In this embodiment, if the first k channels of the audio channel are of the first type, then data extraction can be performed on the first k low-frequency channels according to the above constraints. Wherein, the number of sampling points q extracted from each of the first k channels... ch q ch =q1,q2,…,q k In the formula, ch = 1, 2, ..., k; k is a natural number greater than 1, and q k Let q represent the number of sampling points in the k-th channel, and q k The value of is no greater than the number of sampling points in the (k+1)th channel.
[0109] It should be noted that the extraction rate for the k-th channel is q. k In the case of / L, the number of sampling points q extracted for the first k channels of each frame of data. ch The following constraints must be met:
[0110]
[0111]
[0112] In the formula: The number of sampling points for the k-th channel is represented by fs, where fs represents the audio sampling frequency of the cochlear implant system, and L represents the data frame length. This represents the maximum value of the fundamental frequency.
[0113] Optionally, the calculation formula for data extraction processing according to the root mean square algorithm is as follows:
[0114] For the first k channels:
[0115]
[0116] In the formula: This represents the audio data extracted by the root mean square algorithm, where j represents the number of samples. Indicates channel sampling data, This represents the average value (which takes the value of a positive integer). This represents a sample value. This indicates the number of sampling points for the first type of channel.
[0117] Optionally, the calculation formula for data extraction processing according to the mean algorithm is as follows:
[0118] For the first k channels:
[0119]
[0120] In the formula: This represents the audio data extracted using the mean algorithm, where j represents the number of samples. Indicates channel sampling data, This represents the average value (which takes the value of a positive integer). This represents a sample value. This indicates the number of sampling points for the first type of channel.
[0121] Step S103: Perform data extraction processing on each frame of data of the second type channel according to the second method to obtain the second audio data.
[0122] In this embodiment, for the second type of channel, the envelope extraction algorithm is mainly used to focus on the changes in speech signal energy in order to ensure speech clarity.
[0123] For example, the number of sampling points for each second-type channel is determined based on the channel stimulus rate, data frame length, and audio sampling frequency corresponding to each second-type channel; the extraction parameters corresponding to each frame of data are determined based on the data extraction ratio constraint; and the data extraction is performed according to the root mean square algorithm or the mean algorithm based on the extraction parameters and the number of sampling points corresponding to each second-type channel to obtain the second audio data.
[0124] As an example, assuming that the channel stimulation rate (PCSR) of channels k+1 to m is s, in units of pulses per second (pps), the formula for calculating the sampling rate (number of sampling points) of each frame of data for the second type of channel is as follows:
[0125]
[0126] In the formula: The number of sampling points for the second type of channel is represented by fs, the audio sampling frequency of the cochlear implant system is represented by L, the data frame length is represented by s, the channel stimulation rate of the (k+1)th to the mth channel is represented by Z, and Z is represented by a positive integer.
[0127] Optionally, the calculation formula for data extraction processing according to the root mean square algorithm is as follows:
[0128] For channels k+1 to m:
[0129]
[0130] In the formula, This indicates the number of sampling points for the second type of channel.
[0131] Optionally, the calculation formula for data extraction processing according to the mean algorithm is as follows:
[0132] For channels k+1 to m:
[0133]
[0134] In the formula, This indicates the number of sampling points for the second type of channel.
[0135] Step S104: After nonlinear compression of the first audio data and the second audio data, they are mapped to the first current pulse sequence and the second current pulse sequence, respectively.
[0136] In this embodiment, the data processed by the two types of channels are dynamically compressed to adapt to the limited dynamic range of electrical stimulation, and then mapped to the current pulse sequence of the corresponding electrodes.
[0137] For example, the first and second audio data are scaled using channel feature values to bring them to a uniform benchmark. Then, after determining the maximum and minimum values of each channel feature, linear normalization is performed. The nonlinear compression methods for the first audio data include power function compression, logarithmic compression, and piecewise linear-logarithmic compression. The nonlinear compression methods for the second audio data include multi-channel parallel compression, where each frequency band channel is compressed independently.
[0138] Optionally, after obtaining the compressed normalized amplitude values (range 0-1), these amplitude values are then converted into current amplitude values (microamperes). The mapping formula is as follows:
[0139] Current amplitude = I_THR + compression value × (I_MCL - I_THR)
[0140] In the formula: I_THR represents the hearing threshold current of the electrode (μA), I_MCL represents the most comfortable current of the electrode (μA), and the compression value represents the normalized value (0-1) after nonlinear compression.
[0141] It should be understood that timing information is preserved in both the first and second current pulse sequences. This timing information includes: time stamp mapping, pulse timing, and electrode allocation. Specifically, time stamp mapping means that each feature value has its corresponding original time position; pulse timing means that the time stamp of the feature value directly determines the pulse emission time; and electrode allocation means mapping the feature value to the corresponding physical electrode based on its channel.
[0142] Step S105: The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence.
[0143] In this embodiment, the time interval between each pulse in the first current pulse sequence can be adjusted to achieve optimization in both the time and spatial dimensions. Optimization in the time dimension refers to maximizing the time distance between pulses to improve temporal resolution and avoid interference; optimization in the spatial dimension refers to maximizing the electrode spacing to improve spatial resolution and spectral clarity.
[0144] Optionally, optimization in the time and space dimensions can be achieved by constructing an objective function and constraints. For example, the time interval between the first current pulses in the first current pulse sequence can be adjusted to satisfy the following first condition:
[0145] Within the same data frame and adjacent data frames, the time distance between each adjacent first current pulse is the largest; and / or, the number of stimulation cycles across different channels traversed by adjacent first current pulses is the largest. The stimulation cycle is the reciprocal of the stimulation frequency; for example, if the stimulation frequency of the second type of pulse is 900Hz, then the stimulation cycle is approximately 1.111ms.
[0146] For example, the time interval between the first current pulses in the first current pulse sequence is adjusted to satisfy the following second condition:
[0147] Within the same data frame and adjacent data frames, the electrode spacing between the first current pulses emitted by each adjacent electrode is the largest.
[0148] It should be understood that, in addition to the constraints mentioned above, more constraints can be set to improve the uniformity and continuity of the current pulse distribution (e.g., data frame boundary continuity, original timing fidelity) to adapt to different signal types and environmental conditions, improve the accuracy and naturalness of pitch perception, reduce inter-channel interference and masking effects, enhance the intelligibility of speech in noisy environments, and improve the perceived quality of musical melodies.
[0149] Among them, data frame boundary continuity refers to the time relationship maintained between pulses across frames; original timing fidelity refers to the rearranged time being as close as possible to the original time.
[0150] For example, Figure 2 This is a schematic diagram of the rearrangement of the first current pulse sequence in one embodiment, as shown below. Figure 2 As shown, for a single electrode, maximizing distance refers to maximizing in the time dimension. On the one hand, for a single electrode, whether within a data frame or in adjacent data frames, the emitted pulses satisfy the maximization of time distance, thus achieving a uniform or quasi-uniform distribution relationship (e.g., Figure 2 (As shown in t1 and t2). Here, t1 and t2 represent the time interval between adjacent current pulses in the second channel. Optionally, on the other hand, for different channel stimulation cycles, the current pulses of adjacent electrodes should span as many stimulation cycles as possible, for example... Figure 2 The values t3, t4, t5, and t6 are used in the equation. Specifically, t3 and t4 represent the time interval between adjacent current pulses in channel 2 and channel 1; and t3 and t4 represent the time interval between adjacent current pulses in channel 2 and channel 3.
[0151] See also Figure 2 As shown, for multiple electrodes, maximizing distance refers to maximizing pulse delivery in the spatial dimension (s). Specifically, for sequentially delivered adjacent pulses, whether within a data frame or between adjacent data frames, pulses delivered by adjacent electrodes should be as far apart as possible on the electrodes. Here, electrodes include the 1st to the (k+1th)th working electrodes, such as... Figure 2 In this context, 's' represents the firing time corresponding to the electrode.
[0152] Step S106: Pulse delivery is performed according to the rearranged first current pulse sequence and second current pulse sequence.
[0153] In this embodiment, after determining the firing timing of the rearranged first and second current pulse sequences, the current pulses are fired according to the firing timing sequence. This allows for optimization of the current pulse sequence of the first type of channel in both time and / or spatial dimensions. By fusing the two sequences and discarding precise calculations of F0, within the existing cochlear implant algorithm framework, on the one hand, F0 is fixed in the first 2 to 3 channels, and an incremental rate, fractionally multiple of the baseline stimulation rate, is adopted to satisfy the low-frequency phase-locked loop theory of cochlear implants. On the other hand, for the pulse firing method of the low-frequency channel, the stimulation pulse timing is arranged in both time and spatial dimensions, further reducing pulse interference between low-frequency electrodes and significantly improving the low-frequency audio coding effect.
[0154] In the aforementioned cochlear implant sound signal processing method, the sound channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the fundamental frequency range of speech, while the second type of channel covers the frequency band outside the fundamental frequency range of speech. This allows the sound signal receiving channels to be divided into low-frequency and mid-to-high-frequency channels, facilitating subsequent targeted processing and improving the accuracy of fundamental frequency parameter extraction for the first type of channel. Data extraction processing is performed on each frame of data from the first type of channel according to a first method to obtain first audio data; data extraction processing is performed on each frame of data from the second type of channel according to a second method to obtain second audio data. This allows different data extraction strategies to be used for different types of channels, preventing oversampling of fundamental frequency information due to excessively high stimulation rates in low-frequency channels, which could lead to data blurring or loss. After nonlinear compression of the first and second audio data, they are mapped to a first current pulse sequence and a second current pulse sequence, respectively. The first current pulse sequence is rearranged to obtain a rearranged first current pulse sequence, thus ensuring the fidelity of low-frequency speech information (especially the fundamental frequency F0) and optimizing the stimulation timing. Pulses are delivered according to the rearranged first and second current pulse sequences. This allows for a non-uniform distribution of stimulation pulse resources in the frequency domain by implementing a lower decimation rate for the first type of channels compared to the second type, within a limited number of channels and stimulation rates. This enables more pulse resources to be allocated to frequency bands requiring higher time resolution, thus allowing for the abandonment of precise calculation of F0. Within the existing cochlear implant algorithm framework, on the one hand, F0 is fixed in the first 2 to 3 channels, and an incremental stimulation rate, fractionally higher than the baseline stimulation rate, is used to satisfy the low-frequency phase-locked loop theory of cochlear implants. On the other hand, for the pulse delivery method of low-frequency channels, the timing of stimulation pulses is arranged in both time and space dimensions to further reduce pulse interference between low-frequency electrodes, ensuring the effectiveness of low-frequency speech coding.
[0155] In another exemplary embodiment, such as Figure 3 As shown, a method for processing sound signals from a cochlear implant is provided, which may include the following steps S301 to S308. Wherein:
[0156] Step S301: Divide the audio channels into a first type of channel and a second type of channel.
[0157] The first type of channel covers frequency bands within the voice baseband range, while the second type of channel covers frequency bands outside the voice baseband range.
[0158] In this embodiment, the sound signal input from the cochlear implant is divided into a first type of channel and a second type of channel according to the sound spectrum range. The first type of channel can also be understood as the fundamental frequency channel, used to cover the fundamental frequency range of speech (typically approximately 80-150 Hz for males and 150-300 Hz for females). This information directly corresponds to the pitch and intonation of speech. The second type of channel can also be understood as the non-fundamental frequency channel, used to cover frequency bands outside the fundamental frequency range, mainly carrying the formants, consonant information, and high-frequency details of speech, which determine the clarity and intelligibility of speech.
[0159] Step S302: The audio signals input to each audio channel are filtered by a bandpass filter to obtain the filtered audio signals.
[0160] In this embodiment, the original audio signals input to each audio channel are first filtered by a bandpass filter (for example, a second-order Butterworth bandpass filter is used), thereby limiting the audio of each channel to a certain frequency band range to obtain the filtered audio signal (narrowband signal).
[0161] Step S303: The envelope of the filtered audio signal is extracted using full-wave rectification and low-pass filtering to obtain multi-frame data to be processed.
[0162] In this embodiment, the input signal is a narrowband signal after bandpass filtering, and the signal of each channel is an oscillating waveform within its specific frequency range. Full-wave rectification is the process of converting AC audio signals into unidirectional signals. The original audio signal is a waveform that alternates between positive and negative values on the time axis; after rectification, all negative values are flipped to positive values, forming a signal that is always positive or zero.
[0163] In this embodiment, the purpose of low-pass filtering is to eliminate high-frequency oscillations while ensuring that the rectified signal still contains a fast carrier oscillation component. Here, the envelope refers to the contour line of the signal amplitude changing over time. After extracting the envelope, multiple frames of data to be processed can be obtained. Optionally, the amplitude of the envelope corresponds to the stimulation current intensity, the rate of change of the envelope corresponds to the stimulation rate, and multi-channel envelopes correspond to the distribution of multi-electrode currents, etc.
[0164] Step S304: Perform data extraction processing on each frame of data of the first type of channel according to the first method to obtain the first audio data.
[0165] Step S305: Perform data extraction processing on each frame of data of the second type channel according to the second method to obtain the second audio data.
[0166] Step S306: After nonlinear compression of the first audio data and the second audio data, they are mapped to the first current pulse sequence and the second current pulse sequence, respectively.
[0167] Step S307: The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence.
[0168] Step S308: Pulse delivery is performed according to the rearranged first current pulse sequence and second current pulse sequence.
[0169] For the specific implementation process and technical effects of steps S304 to S308 in this embodiment, please refer to [link / reference]. Figure 1 The relevant descriptions of steps S102 to S106 in the method embodiment shown will not be repeated here.
[0170] For example, Figure 4 A schematic diagram of the principle of a cochlear implant sound signal processing method provided in an embodiment of this application is shown below. Figure 4 As shown, the sound signal enters m sound channels (ch1~chm). The input signal of each sound channel is sequentially processed through bandpass filtering, envelope extraction, data extraction, nonlinear compression, and current mapping, and then outputs a corresponding current pulse. Among them, for the first type of channels (the first k channels), the data extraction stage uses fractional multiplication and F0 constraint to perform data extraction processing to obtain the first audio data.
[0171] In this embodiment, by employing constraints such as fractional decimation and fractional amplification to downsample the low-frequency channel, a hierarchical presentation of the fundamental frequency information of sound is achieved. This solves the problems of inaccurate parameter extraction in traditional F0 algorithms and the inability to accurately present F0 modulation information due to limitations in the system's stimulation rate. Furthermore, by reordering the low-frequency electrode stimulation pulses, the mutual interference between stimulation pulses between low-frequency electrodes is further reduced, which is beneficial for improving the recognition and clarity of low-frequency information (such as tone) in speech signals. Regarding low-frequency channel information processing, this embodiment, while adhering to the traditional CIS processing algorithm architecture of cochlear implants, further simplifies the algorithm framework and reduces its complexity; on the other hand, it reduces the number of current stimulation pulses output at the electrode terminals, further reducing the system's power consumption.
[0172] The cochlear implant sound signal processing method provided in this application will be applied and verified in conjunction with specific embodiments below.
[0173] In an optional example, assume the cochlear implant system has an audio sampling frequency of 16kHz, a data frame length L of 128, and is divided into 16 channels. Three channels cover the low-frequency band within the speech F0 range (typically less than 500Hz). The frequency ranges of the first three channels are [100Hz, 200Hz], [200Hz, 350Hz], and [350Hz, 500Hz], corresponding to center frequencies of 175Hz, 275Hz, and 425Hz, respectively. A second-order Butterworth bandpass filter is used. Envelope extraction employs full-wave rectification followed by low-pass filtering (e.g., a second-order Butterworth low-pass filter with a cutoff frequency of 340Hz).
[0174] If the channel stimulation rate for channels 4 to 16 is 1 kpps, then p = 8. According to the fractional sampling rule, q ch The value range is integer [1, 8]. Furthermore, according to the fractional multiplication rule and the F0 constraint rule, we take q1=1, q2=2, and q3=3, which correspond to stimulation rates of 125pps, 250pps, and 375pps, respectively. It can be seen that the stimulation rates of these three channels are all less than 500Hz.
[0175] See Figure 1 The method shown yields the following data extraction ratios for the three channels:
[0176] Avg1=128, Tot1=1, Res1=0;
[0177] Avg1 represents the average value of the first channel, Tot1 represents the number of data blocks in the first channel, and Res1 represents the individual sample value of the first channel.
[0178] Avg2=64, Tot2=2, Res1=64;
[0179] Avg2 represents the average value of the second channel, Tot2 represents the number of data blocks in the second channel, and Res2 represents the individual sample value of the second channel.
[0180] Avg3=42, Tot3=3, Res1=44;
[0181] Avg3 represents the average value of the third channel, Tot3 represents the number of data blocks in the third channel, and Res3 represents the individual sample value of the third channel.
[0182] For channels 4 to 16, Avg=16, Tot=8, Res=16, where Avg represents the average value, Tot represents the number of data blocks in the channel, and Res represents the individual sample value.
[0183] The data extraction method can be either the root mean square algorithm or the mean algorithm.
[0184] Alternatively, logarithmic compression and linear current mapping can be used, as shown in the following formula:
[0185]
[0186] In the formula: x represents the sound envelope amplitude, y represents the electrical stimulation amplitude, and A and B are constants.
[0187] For example, based on the amplitude range of the input sound envelope From the electrical stimulation threshold (T) and the comfort level (C), a constant can be obtained. and Value. Where:
[0188]
[0189]
[0190] or,
[0191]
[0192] In the formula, This represents the maximum value of the input sound envelope amplitude. This represents the minimum value of the input sound envelope amplitude.
[0193] Optionally, the input audio signal here is a normalized WAV audio signal, where , .
[0194] Furthermore, for medium- and high-frequency electrode pulses, current pulses can be delivered in sequence. Based on a single-channel stimulation rate of 1 kpps, the time interval between two consecutive current pulses on the same electrode is 1 ms, and the time interval between two consecutive current pulses on adjacent electrodes is 62.5 μs.
[0195] For example, Figure 5 This is a schematic diagram illustrating the effect of maximizing the data distance in the first channel in one embodiment. Figure 6 This is a schematic diagram illustrating the effect of maximizing the data distance in the second channel in one embodiment. Figure 7 This is a schematic diagram illustrating the effect of maximizing the data distance in the third channel in one embodiment.
[0196] In this embodiment, for low-frequency electrode pulse delivery, current pulses are delivered using a distance-maximizing pulse delivery method. For the same electrode, whether within a single frame or in adjacent frames, the time interval between stimulation pulses delivered by the first electrode (E1) is... The time interval between stimulation pulses emitted by the second electrode (E2) is... The time interval between stimulation pulses emitted by the third electrode (E3) and .
[0197] In this embodiment, for adjacent electrodes, the time interval between the two adjacent pulses emitted by the first electrode (E1) and the second electrode (E2) is 1. The time interval between two adjacent pulses emitted from the second electrode (E2) and the third electrode (E3) is... and The time interval between two adjacent pulses emitted from the third electrode (E3) and the fifth electrode (E5) is... .
[0198] In this embodiment, spatially: within the minimum sequence stimulation pulse time, the first electrode (E1) current pulse is adjacent to the third electrode (E3) current pulse, with an electrode spacing of [missing information]. (Time interval is) The current pulse of the second electrode (E2) is adjacent to the current pulse of the third electrode (E3), and the electrode spacing is [missing information]. (Time interval is) The third electrode (E3) current pulse is adjacent to the fifth electrode (E5) current pulse, with an electrode spacing of [missing information]. (Time interval is) ).
[0199] For example, Figure 8 This is a schematic diagram of the data arrangement and pulse emission timing of 16 channels in one embodiment. Figure 8 As can be seen, the low-frequency electrode maximizes the distance in both time and space dimensions. Whether it is an adjacent pulse of the same electrode or different electrodes of adjacent pulses, it exceeds the two parameter values of 1ms and 62.5us for the pulse spacing of the mid- and high-frequency electrodes, thus achieving the purpose of reducing inter-pulse interference.
[0200] In another alternative example, assume the cochlear implant system has an audio sampling frequency of 16kHz, a data frame length L of 128, and is divided into 16 channels. Three channels cover the low-frequency band within the speech F0 range (typically less than 500Hz). The frequency ranges of the first three channels are [100, 200], [200, 350], and [350, 500], corresponding to center frequencies of 175Hz, 275Hz, and 425Hz, respectively. A second-order Butterworth bandpass filter is used. Envelope extraction employs full-wave rectification followed by low-pass filtering (e.g., a second-order Butterworth low-pass filter with a cutoff frequency of 340Hz).
[0201] If the channel stimulation rate for channels 4 to 16 is 625 pps, then p = 5. According to the fractional decimation rule, q ch The value range is integer [1, 5]. Furthermore, according to the fractional multiplication rule and the F0 constraint rule, we take q1=1, q2=2, and q3=3, which correspond to stimulation rates of 125pps, 250pps, and 375pps, respectively. It can be seen that the stimulation rates of these three channels are all less than 500Hz.
[0202] See Figure 1 The method shown yields the following data extraction ratios for the three channels:
[0203] Avg1=128, Tot1=1, Res1=0;
[0204] Avg1 represents the average value of the first channel, Tot1 represents the number of data blocks in the first channel, and Res1 represents the individual sample value of the first channel.
[0205] Avg2=64, Tot2=2, Res1=64;
[0206] Avg2 represents the average value of the second channel, Tot2 represents the number of data blocks in the second channel, and Res2 represents the individual sample value of the second channel.
[0207] Avg3=42, Tot3=3, Res1=44;
[0208] Avg3 represents the average value of the third channel, Tot3 represents the number of data blocks in the third channel, and Res3 represents the individual sample value of the third channel.
[0209] For channels 4 to 16, Avg=16, Tot=8, Res=16, where Avg represents the average value, Tot represents the number of data blocks in the channel, and Res represents the individual sample value.
[0210] The data extraction method can be either the root mean square algorithm or the mean algorithm.
[0211] Alternatively, logarithmic compression and linear current mapping can be used, as shown in the following formula:
[0212]
[0213] In the formula: x represents the sound envelope amplitude, y represents the electrical stimulation amplitude, and A and B are constants.
[0214] For example, based on the amplitude range of the input sound envelope From the electrical stimulation threshold (T) and the comfort level (C), a constant can be obtained. and Value. Where:
[0215]
[0216]
[0217] or,
[0218]
[0219] In the formula, This represents the maximum value of the input sound envelope amplitude. This represents the minimum value of the input sound envelope amplitude.
[0220] Optionally, the input audio signal here is a normalized WAV audio signal, where , .
[0221] Furthermore, for medium- and high-frequency electrode pulses, current pulses can be delivered in sequence. Based on a single-channel stimulation rate of 625pps, the time interval between two consecutive current pulses on the same electrode is 1.6ms, and the time interval between two consecutive current pulses on adjacent electrodes is 100us.
[0222] For example, Figure 9 This is a schematic diagram illustrating the effect of maximizing the data distance in the first channel in another embodiment. Figure 10 This is a schematic diagram illustrating the effect of maximizing the data distance in the second channel in another embodiment. Figure 11 This is a schematic diagram illustrating the effect of maximizing the data distance in the third channel in another embodiment.
[0223] In this embodiment, for low-frequency electrode pulse delivery, current pulses are delivered using a distance-maximizing pulse delivery method. For the same electrode, whether within a single frame or in adjacent frames, the time interval between stimulation pulses delivered by the first electrode (E1) is... The time interval between stimulation pulses emitted by the second electrode (E2) is... The time interval between stimulation pulses emitted by the third electrode (E3) and .
[0224] In this embodiment, for adjacent electrodes, the time interval between the two adjacent pulses emitted by the first electrode (E1) and the second electrode (E2) is 1. The time interval between two adjacent pulses emitted from the second electrode (E2) and the third electrode (E3) is t. 23 =t 24=1.6ms - 100μs = 1.5ms; the time interval t between two adjacent pulses emitted by the third and fifth electrodes. 33 =t 34 =1.6ms-100μs=1.5ms.
[0225] In this embodiment, spatially: within the minimum sequence stimulation pulse time, the first electrode (E1) current pulse is adjacent to the third electrode (E3) current pulse, with an electrode spacing of [missing information]. (Time interval is) The second electrode (E2) current pulse is adjacent to the fifth electrode (E5) current pulse, with an electrode spacing of [missing information]. (Time interval is) The third electrode (E3) current pulse is adjacent to the fifth electrode (E5) current pulse, with an electrode spacing of one electrode. (Time interval is) ).
[0226] For example, Figure 12 This is a schematic diagram of the data arrangement and pulse delivery timing of 16 channels in another embodiment. Figure 12 As can be seen, the low-frequency electrode maximizes the distance in both time and space dimensions. Whether it is an adjacent pulse of the same electrode or a different electrode with adjacent pulses, it exceeds the two parameter values of 1.6ms and 100us for the pulse spacing of the mid- and high-frequency electrodes, thus achieving the purpose of reducing inter-pulse interference.
[0227] For example, Figure 13 This is a waveform diagram of a speech signal in one embodiment (the horizontal axis represents time in seconds; the vertical axis represents amplitude). Figure 14 This is a spectrogram of a speech signal in one embodiment (the horizontal axis represents time in seconds; the vertical axis represents frequency). Figure 15 This is a schematic diagram of the waveform output after data extraction in one embodiment. Figure 16 This is an electrode diagram from one embodiment (the horizontal axis represents time in seconds; the vertical axis represents the electrode number). For example, Figure 17 This is a schematic diagram of the waveform output after data extraction in another embodiment. Figure 18 This is an electrode diagram from another embodiment (the horizontal axis represents time in seconds; the vertical axis represents the electrode number).
[0228] In another alternative example, for the second type of channel, a process for selecting the maximum value can be added. For example, in the output of the (k+1)th to the mth channel, the largest n values are selected as the output, and the corresponding electrodes output n current pulses.
[0229] In this embodiment, if n is 8, then within 1ms, the 8 channels with the largest amplitude are selected from the 4th to the 16th channels for current mapping, which corresponds to 13 electrodes output current pulses.
[0230] For example, Figure 19 This is a schematic diagram of the data arrangement and pulse emission timing of 16 channels in another embodiment. Figure 19 As can be seen, in addition to maximizing the distance between low-frequency electrodes in both time and space dimensions, the number of pulses in the mid-to-high frequency channels is also reduced, thus lowering the interference between pulses in the high-frequency part.
[0231] It should be understood that although the steps in the flowcharts of the above embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the above embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages in other steps. It is understood that the steps in different embodiments can be freely combined as needed, and all non-contradictory solutions formed by such combinations are within the scope of protection of this application.
[0232] Based on the same inventive concept, this application also provides a cochlear implant sound signal processing device for implementing the cochlear implant sound signal processing method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the cochlear implant sound signal processing device provided below can be found in the limitations of the cochlear implant sound signal processing method described above, and will not be repeated here.
[0233] In one exemplary embodiment, such as Figure 20 As shown, a cochlear implant sound signal processing device is provided, comprising: a channel division module 2001, a first extraction module 2002, a second extraction module 2003, a mapping module 2004, a reordering module 2005, and a pulse delivery module 2006, wherein:
[0234] The channel division module 2001 is used to divide the audio channel into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range.
[0235] The first extraction module 2002 is used to perform data extraction processing on each frame of data of the first type channel according to the first method to obtain the first audio data.
[0236] The second extraction module 2003 is used to perform data extraction processing on each frame of data of the second type channel according to the second method to obtain the second audio data.
[0237] The mapping module 2004 is used to non-linearly compress the first audio data and the second audio data, and then map them into a first current pulse sequence and a second current pulse sequence, respectively.
[0238] The reordering module 2005 is used to rearrange the first current pulse sequence to obtain the rearranged first current pulse sequence.
[0239] The pulse delivery module 2006 is used to deliver pulses according to the rearranged first current pulse sequence and second current pulse sequence.
[0240] For example, the first extraction module 2002 is specifically used to: determine the number of sampling points for each first type of channel; determine the extraction parameters corresponding to each frame of data according to the data extraction ratio constraint; and perform data extraction processing according to the root mean square algorithm or the mean algorithm based on the extraction parameters and the number of sampling points corresponding to each first type of channel to obtain the first audio data; wherein, if the first k channels in each frame of data are first type channels, and k is a natural number greater than 1, then the number of sampling points of the first type of channel must meet the following constraint: the number of sampling points of the first k channels is increasing, and the number of sampling points of the kth channel is not greater than the maximum value of the fundamental frequency.
[0241] For example, the increasing relationship includes: starting from the first channel to the kth channel, the number of sampling points in each channel increases exponentially.
[0242] For example, the second extraction module 2003 is specifically used to: determine the number of sampling points for each second type of channel based on the channel stimulus rate, data frame length, and audio sampling frequency corresponding to each second type of channel; determine the extraction parameters corresponding to each frame of data based on the data extraction ratio constraint; and perform data extraction processing according to the root mean square algorithm or the mean algorithm based on the extraction parameters and the number of sampling points corresponding to each second type of channel to obtain the second audio data.
[0243] For example, the reordering module 2005 is specifically used to: adjust the time distance between first current pulses in the first current pulse sequence to satisfy the following first condition: the time distance between each adjacent first current pulse is the largest in the same data frame and adjacent data frames; and / or, the number of stimulation cycles of different channels traversed by adjacent first current pulses is the largest.
[0244] For example, the reordering module 2005 is specifically used to: adjust the time interval between the first current pulses in the first current pulse sequence to satisfy the following second condition: within the same data frame and adjacent data frames, the electrode interval between the first current pulses emitted by each adjacent electrode is the maximum.
[0245] For example, the above-described apparatus may further include:
[0246] The filtering module 2007 is used to filter the audio signals input from each audio channel through a bandpass filter to obtain the filtered audio signals.
[0247] The envelope extraction module 2008 is used to extract the envelope of the filtered audio signal through full-wave rectification and low-pass filtering to obtain multi-frame data to be processed.
[0248] The modules in the aforementioned cochlear implant sound signal processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the computer device's memory as software, so that the processor can call and execute the corresponding operations of each module.
[0249] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 21 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements a cochlear implant sound signal processing method.
[0250] Those skilled in the art will understand that Figure 21 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0251] In one exemplary embodiment, a computer device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method steps described in the various embodiments above.
[0252] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the method steps of the various embodiments described above.
[0253] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the method steps of the various embodiments described above.
[0254] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0255] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0256] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed. However, they should not be construed as limiting the scope of this application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method of processing a sound signal for a cochlear implant, characterized by, The method includes: The audio channels are divided into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range. Data extraction processing is performed on each frame of data from the first type of channel according to the first method to obtain the first audio data; The second type of channel data is processed by data extraction according to the second method to obtain the second audio data; After nonlinear compression of the first audio data and the second audio data, they are mapped into a first current pulse sequence and a second current pulse sequence, respectively. The first current pulse sequence is rearranged to obtain the rearranged first current pulse sequence; Pulse delivery is performed according to the rearranged first current pulse sequence and the second current pulse sequence.
2. The method according to claim 1, characterized in that, The step of extracting data from each frame of the first type of channel according to the first method to obtain the first audio data includes: Determine the number of sampling points for each of the first-class channels; Based on the data extraction ratio constraint, determine the extraction parameters corresponding to each frame of data; Based on the extraction parameters and the number of sampling points corresponding to each first type of channel, data extraction processing is performed according to the root mean square algorithm or the mean algorithm to obtain the first audio data. Wherein, if the first k channels in each frame of data are the first type of channels, and k is a natural number greater than 1, then the number of sampling points of the first type of channels must satisfy the following constraint: The number of sampling points in the first k channels increases incrementally, and the number of sampling points in the kth channel is not greater than the maximum value of the fundamental frequency.
3. The method according to claim 2, characterized in that, The increasing relationship includes: From the first channel to the kth channel, the number of sampling points in each channel increases exponentially.
4. The method according to claim 1, characterized in that, The step of extracting data from each frame of the second type of channel according to the second method to obtain the second audio data includes: The number of sampling points for each second-class channel is determined based on the channel stimulus rate, data frame length, and audio sampling frequency corresponding to each second-class channel. Based on the data extraction ratio constraint, determine the extraction parameters corresponding to each frame of data; Based on the extraction parameters and the number of sampling points corresponding to each second type of channel, data extraction processing is performed according to the root mean square algorithm or the mean algorithm to obtain the second audio data.
5. The method according to any one of claims 1 to 4, characterized in that, The step of rearranging the first current pulse sequence to obtain the rearranged first current pulse sequence includes: The time interval between the first current pulses in the first current pulse sequence is adjusted to satisfy the following first condition: Within the same data frame and adjacent data frames, the time distance between each adjacent first current pulse is the largest; and / or, the number of stimulation cycles across different channels by adjacent first current pulses is the largest.
6. The method according to any one of claims 1 to 4, characterized in that, The step of rearranging the first current pulse sequence to obtain the rearranged first current pulse sequence includes: The time interval between the first current pulses in the first current pulse sequence is adjusted to satisfy the following second condition: Within the same data frame and adjacent data frames, the electrode spacing between the first current pulses emitted by each adjacent electrode is the largest.
7. The method according to any one of claims 1 to 4, characterized in that, Before performing data extraction processing on each frame of data from the first type of channel according to the first method, and before performing data extraction processing on each frame of data from the second type of channel according to the second method, the method further includes: The audio signals input to each audio channel are filtered by a bandpass filter to obtain the filtered audio signals. The filtered audio signal is enveloped and processed by full-wave rectification and low-pass filtering to obtain multi-frame data to be processed.
8. A cochlear implant sound signal processing device, characterized in that, The device includes: The channel division module is used to divide the audio channel into a first type of channel and a second type of channel. The first type of channel covers the frequency band within the speech fundamental frequency range, and the second type of channel covers the frequency band outside the speech fundamental frequency range. The first extraction module is used to perform data extraction processing on each frame of data of the first type of channel in a first manner to obtain the first audio data. The second extraction module is used to extract data from each frame of data of the second type of channel according to the second method to obtain the second audio data. The mapping module is used to non-linearly compress the first audio data and the second audio data, and then map them into a first current pulse sequence and a second current pulse sequence, respectively. The reordering module is used to rearrange the first current pulse sequence to obtain the rearranged first current pulse sequence. The pulse delivery module is used to deliver pulses according to the rearranged first current pulse sequence and the second current pulse sequence.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.