Auditory signal processing method and device

By employing a multi-path auditory signal processing method, utilizing cross-processing, fusion, and independent processing modes, the problem of existing auditory sensors failing to meet multiple performance indicators under low computational cost is solved, achieving efficient and robust auditory perception.

CN121099243APending Publication Date: 2025-12-09TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410730015.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-06-06
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing auditory sensor hardware struggles to simultaneously meet multiple performance metrics with low computational costs, particularly the trade-off between spectral temporal resolution and spectral frequency resolution, and post-processing algorithms introduce computational redundancy and latency.

Method used

A multi-path auditory signal processing method is adopted to obtain complementary information of primitives through multiple paths, and to perform complementary processing by combining cross-processing, fusion processing and independent processing modes to achieve efficient and robust auditory perception.

Benefits of technology

It achieves excellent auditory perception performance with extremely low computing cost, improves overall energy efficiency, and optimizes multiple performance indicators at the same time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121099243A_ABST
    Figure CN121099243A_ABST
Patent Text Reader

Abstract

The invention provides an auditory signal processing method and device. The auditory signal processing method comprises the steps of obtaining a to-be-processed auditory signal; wherein the auditory signal to be processed corresponds to a plurality of channels, and the properties of at least one primitive among the channels are complementary; the primitives refer to the output signal type, time domain time resolution, frequency domain frequency resolution, spectrum time resolution, spectrum frequency resolution, cepstrum domain channel number, response mode, data precision, sensitivity, signal-to-noise ratio, total harmonic distortion, sound overload point, perceptible frequency range, spatial resolution or spatial dimension of the auditory signal, and whether self-adaption exists or not; one or a combination of more of non-linearity; and on the basis of the primitive, performing complementary processing on the auditory signal to be processed according to a pre-constructed complementary processing mode. According to the invention, primitive complementary information is obtained through multiple channels, and then efficient and robust auditory perception is realized through complementary processing.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of auditory signal processing, and in particular to an auditory signal processing method and device. BACKGROUND

[0002] At present, the mainstream auditory sensor hardware mostly adopts a single channel design. A significant problem of the signal output by this design is the mutual restriction between performance indicators. Taking the spectral output as an example, the spectral time resolution and the spectral frequency resolution cannot be optimal at the same time, and present a relationship of this and that.

[0003] In addition, some researches alleviate the mutual restriction problem of the time-frequency resolution of the spectral output to a certain extent through wavelet transform, frame overlap and other algorithms. However, these post-processing algorithms introduce a large amount of calculation, and are accompanied by a large amount of calculation redundancy. Compared with direct processing at the sensor end, these algorithms not only introduce additional delay, but also are not suitable for systems or tasks with high real-time requirements.

[0004] In summary, the prior art has the problem of being unable to simultaneously meet multiple performance indicators at low computational cost. SUMMARY

[0005] The present application provides an auditory signal processing method and device to solve the defect that multiple performance indicators cannot be simultaneously met at low computational cost in the prior art, and to realize excellent perceptual performance at extremely low cost.

[0006] The present application provides an auditory signal processing method, comprising the following steps: acquiring an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and the channels are complementary in at least one property of a primitive; the primitive is one or more of the following: an output signal type of an auditory signal, a time domain time resolution, a frequency domain frequency resolution, a spectral time resolution, a spectral frequency resolution, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; complementary processing the auditory signal to be processed according to a complementary processing mode constructed in advance based on the primitive; wherein the complementary processing mode comprises one or more of the following: cross processing, fusion processing and independent processing.

[0007] According to the auditory signal processing method provided by the present application, the acquisition of the auditory signal to be processed specifically comprises: acquiring the auditory signal to be processed based on a multi-channel auditory sensor system; Alternatively, the auditory signal to be processed can be generated based on a preset simulation algorithm.

[0008] According to an auditory signal processing method provided by the present invention, the multi-channel auditory sensor system specifically includes: Multiple primitive complementary auditory sensors; And / or, at least one first auditory sensor and at least one second auditory sensor that is complementary to the primitive of the first auditory sensor; And / or, at least one integrated primitive complementary auditory sensor.

[0009] According to the auditory signal processing method provided by the present invention, the output signal type includes time domain signal, frequency domain signal, cepstral domain signal, spectral signal, neuromorphic or brain-like auditory signal; And / or, the response mode includes an intensity signal, a differential signal, and a pulse signal.

[0010] According to an auditory signal processing method provided by the present invention, when the complementary processing mode is cross-processing, the step of performing complementary processing on the auditory signal to be processed based on the primitives and according to a pre-constructed complementary processing mode specifically includes: The auditory signal of the current pathway is supplemented by a target pathway that is complementary to the primitive of the current pathway.

[0011] According to an auditory signal processing method provided by the present invention, when the complementary processing mode is independent processing, the step of performing complementary processing on the auditory signal to be processed based on the primitive and according to a pre-constructed complementary processing mode specifically includes: The auditory signal to be processed is processed independently and simultaneously based on a pre-built independent processing algorithm.

[0012] According to an auditory signal processing method provided by the present invention, when the complementary processing mode is fusion processing, the step of performing complementary processing on the auditory signal to be processed based on the primitives and according to a pre-constructed complementary processing mode specifically includes: The auditory signals to be processed are simultaneously fused based on a pre-built fusion processing algorithm.

[0013] The present invention also provides an auditory signal processing device, comprising: a signal unit configured to acquire a to-be-processed auditory signal, wherein the to-be-processed auditory signal corresponds to a plurality of channels, and at least one property of a primitive is complementary between the channels, the primitive being one or more of a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum domain, a frequency resolution in the spectrum domain, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; a processing unit configured to perform complementary processing on the to-be-processed auditory signal based on the primitive and according to a complementary processing mode constructed in advance, wherein the complementary processing mode comprises one or more of a combination of cross processing, fusion processing, and independent processing.

[0014] The application also provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the auditory signal processing method of any of the above when executing the program.

[0015] The application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program is executable on a processor to implement the auditory signal processing method of any of the above.

[0016] The application also provides a computer program product comprising a computer program, wherein the computer program is executable on a processor to implement the auditory signal processing method of any of the above.

[0017] The auditory signal processing method and device provided by the application acquire a to-be-processed auditory signal, wherein the to-be-processed auditory signal corresponds to a plurality of channels, and at least one property of a primitive is complementary between the channels, the primitive being one or more of a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum domain, a frequency resolution in the spectrum domain, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; the to-be-processed auditory signal is complementary processed based on the primitive and according to a complementary processing mode constructed in advance, wherein the complementary processing mode comprises one or more of a combination of cross processing, fusion processing, and independent processing. The application acquires complementary information of the primitive through a plurality of channels, and then implements efficient and robust auditory perception through complementary processing, thereby achieving excellent perceptual performance at a very low cost. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to make the technical solutions in the present application or prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative work based on these drawings also belong to the protection scope of the present application.

[0019] Figure 1 is a flowchart of the hearing signal processing method provided by the present application.

[0020] Figure 2 is a hearing signal chain in the prior art provided by the present application.

[0021] Figure 3 is a structural schematic diagram of the hearing signal processing device provided by the present application.

[0022] Figure 4 is a structural schematic diagram of the electronic device provided by the present application. DETAILED DESCRIPTION

[0023] In order to make the technical solutions in the present application or prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and all other drawings obtained by those of ordinary skill in the art without creative work based on these drawings also belong to the protection scope of the present application.

[0024] It can be understood that the microphone (MIC) is the most common hearing sensor, and the digital microphone (digital microphone) contains an analog-to-digital converter (Analog-to-digital converter, ADC) inside, and the output is a one-dimensional discrete acoustic signal sequence, called a time-domain signal. The Fourier transform of the time-domain signal obtains the frequency-domain signal. In addition to the time-domain signal and the frequency-domain signal, there are three special signal modes: the cepstrum domain signal, the frequency spectrum signal, and the neural form or brain-like hearing signal. In the like Figure 2In the shown hearing signal chain, generally speaking, the hearing sensor outputs a time domain signal, which can be converted into a frequency domain, a cepstrum domain or a spectrum domain in the conversion and calculation link to extract various features to complete the hearing task. In the conversion and calculation process of the spectrum, since the Fourier transform is performed on the continuous hundreds of signals in the one-dimensional signal and recorded as a time, the time resolution of the spectrum graph quickly decreases, and the frequency resolution and the time resolution are mutually restricted. The cepstrum domain signal can be understood as taking the logarithm of the frequency domain and then performing inverse Fourier transform to return to the time domain, although it returns to the time domain, but the data is reduced in dimension. The neuro-morphic or brain-like hearing signal refers to the efficient perception and representation of the hearing signal by referring to or being inspired by the principle of the human auditory system, which is usually realized by a neuro-morphic or brain-like hearing sensor, or can be obtained by other modal signals through a neuro-morphic or brain-like hearing sensor simulation algorithm. For example, the DAS mentioned below is a kind of neuro-morphic or brain-like hearing sensor.

[0025] Compared with the MIC, the DAS has extremely high spectral time resolution. Since the spectral frequency resolution of the DAS is determined by the number of channels, the limited chip area does not allow too many channels, so the frequency resolution of the DAS is not as good as that of the traditional MIC.

[0026] Compared with the MIC, the DAS has extremely high spectral time resolution. Since the spectral frequency resolution of the DAS is determined by the number of channels, the limited chip area does not allow too many channels, so the frequency resolution of the DAS is not as good as that of the traditional MIC.

[0027] On this basis, the present application provides a kind of hearing signal processing method and device, based on the complementary multi-channel hearing signal processing method of hearing primitive, the information of primitive complement is obtained by multiple paths, then efficient hearing perception is realized by fusion algorithm.

[0028] The hearing signal processing method of the present application will be described below Figures 1-2 The hearing signal processing method of the present application will be described below Figure 1 The hearing signal processing method of the present application will be described below Figure 1 As shown in the flowchart of the hearing signal processing method provided by the present application, the method comprises steps 110 and 120.

[0029] Step 110: obtaining a to-be-processed auditory signal; wherein the to-be-processed auditory signal corresponds to multiple channels, and at least one property of a primitive between the channels is complementary; the primitive is one or more of a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio (SNR), a total harmonic distortion (THD), an acoustic overload point (AOP), a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear.

[0030] It should be noted that, as a whole, the obtained to-be-processed auditory signal essentially belongs to a complementary multi-channel auditory signal, and at least one property of a primitive between the corresponding multiple channels is complementary.

[0031] Further, the definition of the primitive is derived from human hearing. It can be understood that human hearing has a cocktail party effect, and people can select and separate the signal of interest in a noisy environment through attention and other mechanisms, such as in a cocktail party where multiple people are speaking at the same time, only the speech of the person of interest can be focused on and other speeches can be shielded. Or in the case of a large number of signals superimposed, the human auditory system decomposes information into auditory primitives through characteristics such as timbre and sound source, and finally constructs an auditory scene.

[0032] In the embodiments of the present application, the auditory primitive includes but is not limited to: an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, an acoustic overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear.

[0033] In some embodiments, the output signal type includes a time domain signal, a frequency domain signal, a cepstrum domain signal, a spectrum signal, a neuromorphic or brain-like auditory signal.

[0034] It needs to be noted that in some embodiments, the output signal type refers to 5 categories of time domain signals, frequency domain signals, spectral signals, cepstrum domain signals and neuro-morphic or brain-like auditory signals. Among them, the time domain signal focuses on the time resolution in the time domain, and for the traditional MIC, the time resolution in the time domain refers to the sampling rate, that is, the inverse of the difference between two adjacent signal timestamps. The pure frequency domain signal does not have time resolution, and mainly focuses on the frequency resolution in the frequency domain, that is, the minimum frequency difference between two adjacent signals. The spectral signal includes both time resolution and frequency resolution. The cepstrum domain is usually composed of multiple channels, and the main indicator for measuring its resolution is the number of channels. The neuro-morphic or brain-like auditory signal is more complex, and has different resolution indicators according to different subtypes, which are not unique. For example, for DAS, the indicators are spectral frequency resolution and spectral time resolution, because DAS output signal can be understood as a special spectral signal.

[0035] For the perceptible frequency range, the difference between the highest and lowest frequencies that the sensor can respond to is defined.

[0036] In some embodiments, the response modalities include intensity signals, differential signals, and pulse signals.

[0037] It can be understood that the traditional MIC obtains an intensity signal; the differential signal refers to a weighted difference operation between the current time and the current position signal and at least one arbitrary time and arbitrary position signal. The pulse signal is a special case of the intensity signal, that is, in the process of converting the intensity signal into the pulse signal, only when the intensity exceeds a certain threshold, the pulse signal outputs a non-zero value, otherwise it outputs 0. The pulse signal usually has 1-bit precision, and there are also cases of multi-bit precision.

[0038] Data precision refers to the number of bits after signal digitization.

[0039] Sensitivity refers to the relationship between sound pressure signal and output electrical signal, usually in units of dBFS (digital signal) and dBV (analog signal).

[0040] The definition of signal-to-noise ratio in MIC usually refers to the system output signal-to-noise ratio when a 1 Pa standard sinusoidal signal is input.

[0041] Total harmonic distortion THD and acoustic overload point AOP measure the nonlinearity of the system. For MIC, there is a clear definition, which will not be repeated here. For DAS, the relevant indicators need to be redefined.

[0042] A single auditory sensor collects a one-dimensional signal, and through the signals of multiple auditory sensors, the sound source can be located and separated, and the spatial distribution of the sound signal or even the acoustic characteristic distribution in the scene can be analyzed; the spatial resolution / spatial dimension number refers to the number of sampling points.

[0043] Adaptation and nonlinearity are both characteristics of the human auditory system, the human ear controls the sensitivity of the cochlea to a signal at a certain frequency response position through outer hair cells, thereby achieving attention to certain specific frequency signals, in addition, the human auditory system itself has the characteristics of nonlinear response to sound signals, such as the same sound pressure level but different frequency signals, the strength of the human ear may be different, generally the loudness is used to represent the subjective judgment result of the sensor whether it has adaptation and nonlinearity, which can also be complementary, such as considering the nonlinearity of the human ear in path A, and not considering in path B, then the actual effect of the auditory scene and the effect of the human ear can be perceived at the same time.

[0044] Further, in the aspect of obtaining the to-be-processed auditory signal, it can be understood that the auditory signal is first converted into an analog electrical signal through a MEMS diaphragm or other means. In the process of converting the electrical signal, the existing method includes a traditional way of using an ADC for time domain processing, a way of using a DAS or other brain-like auditory sensor for spectral event processing, and other possible ways (such as a differential response mode, directly obtaining a frequency domain signal or a cepstrum domain signal, etc.).

[0045] In actual operation, in some embodiments, the obtaining the to-be-processed auditory signal specifically includes: obtaining the to-be-processed auditory signal based on a multi-path auditory sensor system; Or, generating the to-be-processed auditory signal based on a preset simulation algorithm.

[0046] Specifically, a plurality of auditory sensors can be used to obtain the to-be-processed auditory signal, and the to-be-processed auditory signal is integrated into an auditory sensor system. In terms of the composition of the specific path of the sensor, the acquisition of the complementary multi-path auditory signal can be implemented by using a plurality of primitive complementary ADC-based MICs. A plurality of primitive complementary DAS-based MICs can also be used. Or an ADC and DAS-based MIC can be used. It can also be an integrated primitive complementary auditory sensor. The DAS mentioned here can be replaced by other types of neuro-morphic or brain-like auditory sensors.

[0047] In addition, the data can also be obtained without hardware, but through simulation or generative methods. The present application does not limit the specific generation or simulation method or software algorithm, and a suitable method can be selected in actual operation.

[0048] Based on the above embodiments, in some embodiments, the multi-path auditory sensor system specifically includes: a plurality of primitive complementary auditory sensors; And / or, at least one first hearing sensor is complementary to at least one second hearing sensor with a primitive complementary to the first hearing sensor primitive. And / or, at least one integrated primitive complementary hearing sensor.

[0049] To further illustrate the primitive complementarity, a specific embodiment is given. In this embodiment, it is assumed that the overall system includes a DAS and an ADC simultaneously processing the signal of the MIC output, and the ADC is subsequently connected to a short-time Fourier transform (STFT) module for obtaining the response spectrum, and the two are complementary in adaptation and frequency time and frequency resolution. It can be the same MIC that divides two analog signals, or two MICs can be set up, and if two MICs are set up, there can be primitive complementarity between the MICs in sensitivity and the like. For example, it can also be that multiple ADCs simultaneously process the output signal of the MIC, ADC1 has the characteristics of high time resolution (sampling rate) and low data precision, ADC2 has the characteristics of low time resolution (sampling rate) and high data precision, and the two form primitive complementarity.

[0050] Step 120: based on the primitive, complementary processing is performed on the to-be-processed hearing signal according to a pre-constructed complementary processing mode; wherein the complementary processing mode includes one or more combinations of cross processing, fusion processing, and independent processing.

[0051] A single channel hearing signal usually cannot achieve all performance indicators simultaneously optimal, and multiple channels need to be complementarily processed to achieve the overall optimal effect.

[0052] The embodiment of the application sets multiple complementary processing modes for complementary processing of the to-be-processed hearing signal. It can be understood that, for example, two channels are complementary in time domain time resolution and frequency domain frequency resolution, one channel has high time domain time resolution but low frequency domain frequency resolution, and one channel has low time domain time resolution but low frequency domain frequency resolution, then two channels are expected to achieve high time domain time resolution and high frequency domain frequency resolution after late fusion processing. The cost of achieving two high indicators on one channel may be much higher than achieving one high indicator and one low indicator, so this complementary multi-channel solution can achieve extremely high performance (in some special cases, this high performance may be impossible to achieve in a single channel no matter how much resource cost) at extremely low resource cost (power consumption, bandwidth, volume, computing power, etc.), thereby greatly improving the overall energy efficiency.

[0053] In some embodiments, when the complementary processing mode is cross processing, the complementary processing of the to-be-processed auditory signal based on the primitive and according to the pre-constructed complementary processing mode specifically includes: Supplementing the auditory signal of the current channel with the target channel that is complementary to the primitive of the current channel.

[0054] Specifically, the cross processing (or master-slave processing) includes: while processing the signal of the self channel, introducing the information supplement and improvement of the channel that is complementary to the primitive of the self channel.

[0055] Further, the interrupt function can also be considered in the process of cross processing. In order to further illustrate, a specific embodiment is given: two channels output spectral signals at the same time, the spectral time resolution of channel 1 is high, and the spectral frequency resolution of channel 1 is low, the spectral time resolution of channel 2 is low, and the spectral frequency resolution of channel 2 is high. Then when channel 1 detects that the spectrum changes rapidly in time, channel 2 can be interrupted because the time resolution of channel 2 is insufficient to reflect the time change of the spectrum, and when channel 2 detects that the spectrum changes slightly in frequency, channel 1 can also be interrupted for the same reason. This interruption can be resumed after the abnormal situation ends.

[0056] Among them, the interrupt mechanism can be but not limited to the predictive interrupt mechanism, that is, through the information of a channel, it is judged whether the other channel is likely to have a problem, and if the probability of the other channel having a problem is large, the interrupt is triggered.

[0057] In some embodiments, when the complementary processing mode is independent processing, the complementary processing of the to-be-processed auditory signal based on the primitive and according to the pre-constructed complementary processing mode specifically includes: Simultaneously performing independent processing on the to-be-processed auditory signal based on the pre-constructed independent processing algorithm.

[0058] Specifically, the independent processing includes: processing only the signal of the self channel based on the independent processing algorithm to obtain the processing result of the single channel. The independent processing algorithm can be the signal processing algorithm built in each auditory sensor, or a new independent processing algorithm can be developed according to the need, for example, two channels are first independently processed, and then fused. It can be understood that, since the independent processing process may need to be prepared for subsequent fusion processing, the independent processing process is adjusted based on the content to be prepared. The present application does not make specific limitation on this.

[0059] In some embodiments, when the complementary processing mode is fusion processing, the complementary processing of the to-be-processed auditory signal based on the primitive and according to the pre-constructed complementary processing mode specifically includes: The pre-constructed fusion processing algorithm is used to perform fusion processing on the to-be-processed auditory signal.

[0060] Specifically, the fusion processing includes: inputting information of a plurality of channels of primitive complementarity into the same algorithm (fusion processing algorithm) to perform fusion processing. It needs to be emphasized that the fusion processing algorithm has multiple forms, including pre-fusion algorithm and post-fusion algorithm. For example, a plurality of channels are input through a fusion neural network, and a high-performance signal under a certain signal type is obtained through end-to-end inference. The present application does not make specific limitations thereto.

[0061] On the basis of the above-mentioned embodiments, it needs to be pointed out that different processing modes can be combined, such as first performing independent processing on each channel, and then performing fusion processing on the independent processing results. In detail, for example, each channel is first independently subjected to sound source positioning, and then the positioning results are weighted to obtain the final judgment result.

[0062] Further, the processing method provided by the embodiments of the present application can be deployed in a computing module inside a sensor, or a back-end independent computing and processing module, such as CPU, GPU, brain-like computing chip or other computing chips, etc.

[0063] It needs to be emphasized that the execution hardware of step 110 and step 120 can be completely separated, or can be integrated at the board level, or can be integrated in the same chip, and the present application does not make limitations thereto.

[0064] Specifically, the present application can be used on general-purpose processors, including CPU and GPU. In the CPU, the complementary interrupt mode of the embodiments of the present application is realized by using the interrupt system, and the parallel complementary auditory tasks are realized by using multi-thread and multi-process means; in the GPU, the multi-stream or multi-process service (Multi-process service, MPS) can be used to realize the parallel processing of multiple complementary auditory tasks; the stream and event mechanism of CUDA (Compute Unified Device Architecture, Unified Computing Architecture) can be used to trigger or interrupt another auditory task, so as to realize complementary interruption.

[0065] The present application can also be used on brain-like computing chips. For example, for brain-like many-core chips, a large number of complementary tasks can be deployed in parallel on them by using the characteristics of decentralized distributed computing, so that they can be executed simultaneously. In the brain-like many-core chip, there is also an inter-core instant communication operation, and the complementary interruption can be realized by using this method.

[0066] The application provides an auditory signal processing method, which comprises the following steps: obtaining an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one property of a primitive is complementary between the channels; the primitive is one or more combinations of the following: an output signal type of the auditory signal, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a channel number in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a spatial dimension number, whether adaptive, and whether nonlinear; and the auditory signal to be processed is processed according to a complementary processing mode which is constructed in advance based on the primitive; wherein the complementary processing mode comprises one or more combinations of cross processing, fusion processing and independent processing. The application obtains complementary information of the primitive through multiple channels, and then realizes efficient and robust auditory perception through complementary processing, thereby realizing excellent perception performance at a very low cost.

[0067] The application provides an auditory signal processing device, and the following description of the auditory signal processing device can be referred to in combination with the above description of the auditory signal processing method. Figure 3 The application provides an auditory signal processing device, and the following description of the auditory signal processing device can be referred to in combination with the above description of the auditory signal processing method. Figure 3 As shown in the figure, the device comprises: A signal unit 310 is configured to obtain an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one property of a primitive is complementary between the channels; the primitive is one or more combinations of the following: an output signal type of the auditory signal, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a channel number in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a spatial dimension number, whether adaptive, and whether nonlinear; A processing unit 320 is configured to process the auditory signal to be processed according to a complementary processing mode which is constructed in advance based on the primitive; wherein the complementary processing mode comprises one or more combinations of cross processing, fusion processing and independent processing.

[0068] According to the application, the auditory signal to be processed is obtained, and specifically comprises the following steps: The auditory signal to be processed is obtained based on a multi-channel auditory sensor system; Or, the auditory signal to be processed is generated based on a preset simulation algorithm.

[0069] According to the application, the multi-channel auditory sensor system specifically comprises the following: Multiple auditory sensors with complementary primitives; And / or, at least one first hearing sensor is complementary to at least one second hearing sensor with the first hearing sensor primitive; And / or, at least one integrated primitive complementary hearing sensor.

[0070] According to the application, an auditory signal processing device is provided, and the output signal type includes time domain signal, frequency domain signal, cepstrum domain signal, frequency spectrum signal, neural morphology or brain-like auditory signal; And / or, the response mode includes intensity signal, differential signal and pulse signal.

[0071] According to the application, an auditory signal processing device is provided, and in the case of cross processing of the complementary processing mode, the auditory signal to be processed is processed based on the primitive according to the pre-constructed complementary processing mode, specifically including: The auditory signal of the current channel is supplemented by the target channel complementary to the primitive of the current channel.

[0072] According to the application, an auditory signal processing device is provided, and in the case of independent processing of the complementary processing mode, the auditory signal to be processed is processed based on the primitive according to the pre-constructed complementary processing mode, specifically including: The auditory signal to be processed is simultaneously processed based on the pre-constructed independent processing algorithm.

[0073] According to the application, an auditory signal processing device is provided, and in the case of fusion processing of the complementary processing mode, the auditory signal to be processed is processed based on the primitive according to the pre-constructed complementary processing mode, specifically including: The auditory signal to be processed is simultaneously processed based on the pre-constructed fusion processing algorithm.

[0074] The application provides an auditory signal processing device, which comprises the following steps: obtaining an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one property of a primitive is complementary between the channels; the primitive is one or more of the following: a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; based on the primitive, the auditory signal to be processed is processed according to a complementary processing mode constructed in advance; wherein the complementary processing mode comprises one or more of the following: a combination of cross processing, fusion processing and independent processing. The application obtains complementary information of the primitive through multiple channels, and then realizes efficient and robust auditory perception through complementary processing, so that excellent perception performance is realized at a very low cost.

[0075] Figure 4 An example of a schematic diagram of a physical structure of an electronic device is shown in Figure 4 The electronic device can include a processor 410, a communications interface 420, a memory 430 and a communications bus 440, wherein the processor 410, the communications interface 420 and the memory 430 can communicate with each other through the communications bus 440. The processor 410 can call a logical instruction in the memory 430 to execute an auditory signal processing method, which comprises the following steps: obtaining an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one property of a primitive is complementary between the channels; the primitive is one or more of the following: a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a time resolution in a spectrum, a frequency resolution in a spectrum, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; based on the primitive, the auditory signal to be processed is processed according to a complementary processing mode constructed in advance; wherein the complementary processing mode comprises one or more of the following: a combination of cross processing, fusion processing and independent processing.

[0076] In addition, the logic instructions in the memory 430 described above can be implemented in the form of software function units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that contribute to the prior art or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0077] In another aspect, the present application also provides a computer program product, which comprises a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the hearing signal processing method provided by the above-mentioned methods, the method comprising: obtaining a to-be-processed hearing signal; wherein the to-be-processed hearing signal corresponds to a plurality of channels, and at least one property of a primitive between the channels is complementary; the primitive is one or more combinations of the following: an output signal type of a hearing signal, a time domain time resolution, a frequency domain frequency resolution, a frequency spectrum time resolution, a frequency spectrum frequency resolution, a cepstrum domain channel number, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a spatial dimension number, whether adaptive, whether nonlinear; based on the primitive, the to-be-processed hearing signal is processed according to a pre-constructed complementary processing mode; wherein the complementary processing mode comprises one or more combinations of cross processing, fusion processing and independent processing.

[0078] In yet another aspect, the present application also provides a non-transitory computer readable storage medium having stored thereon a computer program, which, when executed by a processor, implements the hearing signal processing method provided by any of the above methods, the method comprising: obtaining a to-be-processed hearing signal; wherein the to-be-processed hearing signal corresponds to a plurality of channels, and at least one property of a primitive is complementary between the channels; the primitive is one or more of a combination of an output signal type, a time resolution in a time domain, a frequency resolution in a frequency domain, a spectral time resolution, a spectral frequency resolution, a number of channels in a cepstrum domain, a response mode, a data precision, a sensitivity, a signal-to-noise ratio, a total harmonic distortion, a sound overload point, a perceptible frequency range, a spatial resolution or a number of spatial dimensions, whether adaptive, and whether nonlinear; and performing complementary processing on the to-be-processed hearing signal according to a pre-constructed complementary processing mode based on the primitive; wherein the complementary processing mode comprises one or more of a combination of cross processing, fusion processing, and independent processing.

[0079] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.

[0080] From the above description of the embodiments, those skilled in the art can clearly understand that the embodiments can be implemented by means of software plus a necessary general hardware platform, and of course can also be implemented by hardware. Based on such understanding, the above technical solutions can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.

[0081] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An auditory signal processing method, characterized in that, include: Acquire an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one primitive among the channels has complementary properties; the primitive is the output signal type of the auditory signal, time-domain time resolution, frequency-domain frequency resolution, spectral time resolution, spectral frequency resolution, number of cepstral channels, response mode, data precision, sensitivity, signal-to-noise ratio, total harmonic distortion, acoustic overload point, perceptible frequency range, spatial resolution or spatial dimension, whether it is adaptive, whether it is nonlinear, or one or more combinations thereof; Based on the primitives, the auditory signal to be processed is subjected to complementary processing according to a pre-constructed complementary processing mode; wherein, the complementary processing mode includes one or more combinations of cross processing, fusion processing and independent processing.

2. The auditory signal processing method according to claim 1, characterized in that, The acquisition of the auditory signal to be processed specifically includes: The auditory signal to be processed is acquired based on a multi-path auditory sensor system; Alternatively, the auditory signal to be processed can be generated based on a preset simulation algorithm.

3. The auditory signal processing method according to claim 2, characterized in that, The multi-channel auditory sensor system specifically includes: Multiple primitive complementary auditory sensors; And / or, at least one first auditory sensor and at least one second auditory sensor that is complementary to the primitive of the first auditory sensor; And / or, at least one integrated primitive complementary auditory sensor.

4. The auditory signal processing method according to claim 1, characterized in that, The output signal types include time-domain signals, frequency-domain signals, cepstral domain signals, spectral signals, and neuromorphic or brain-like auditory signals. And / or, the response mode includes an intensity signal, a differential signal, and a pulse signal.

5. The auditory signal processing method according to claim 1, characterized in that, When the complementary processing mode is cross-processing, the complementary processing of the auditory signal to be processed based on the primitive and according to the pre-constructed complementary processing mode specifically includes: The auditory signal of the current pathway is supplemented by a target pathway that is complementary to the primitive of the current pathway.

6. The auditory signal processing method according to claim 1, characterized in that, When the complementary processing mode is independent processing, the complementary processing of the auditory signal to be processed based on the primitive and according to the pre-constructed complementary processing mode specifically includes: The auditory signal to be processed is processed independently and simultaneously based on a pre-built independent processing algorithm.

7. The auditory signal processing method according to claim 1, characterized in that, When the complementary processing mode is fusion processing, the complementary processing of the auditory signal to be processed based on the primitive and according to the pre-constructed complementary processing mode specifically includes: The auditory signals to be processed are simultaneously fused based on a pre-built fusion processing algorithm.

8. An auditory signal processing device, characterized in that, include: A signal unit is used to acquire an auditory signal to be processed; wherein the auditory signal to be processed corresponds to multiple channels, and at least one primitive among the channels has a complementary property; the primitive is the output signal type of the auditory signal, time-domain time resolution, frequency-domain frequency resolution, spectral time resolution, spectral frequency resolution, number of cepstral channels, response mode, data accuracy, sensitivity, signal-to-noise ratio, total harmonic distortion, acoustic overload point, perceptible frequency range, spatial resolution or spatial dimension, whether it is adaptive, and whether it is nonlinear, or one or more combinations thereof; The processing unit is configured to perform complementary processing on the auditory signal to be processed according to the primitive and a pre-constructed complementary processing mode; wherein the complementary processing mode includes one or more combinations of cross processing, fusion processing and independent processing.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the auditory signal processing method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the auditory signal processing method as described in any one of claims 1 to 7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the auditory signal processing method as described in any one of claims 1 to 7.