System for processing of sound data
A system processes otoacoustic signals using bilateral audio sensors and neural networks to control actuators, addressing noise interference and enhancing control accuracy by recognizing patterns in otoacoustic emissions.
Patent Information
- Application Number
- PCT/NL2025/050364
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-01
- Filing Date
- 2025-07-24
- Publication Date
- 2026-02-05
AI Technical Summary
Existing systems fail to effectively utilize otoacoustic signals for controlling actuators due to low emission levels and interference from noise, particularly in noisy environments.
A system utilizing bilateral audio sensors and actuators in earpieces, combined with electronic circuits and neural networks, processes otoacoustic signals to control actuators by recognizing patterns in otoacoustic emissions, reducing common mode noise, and converting signals to the frequency domain for enhanced accuracy.
The system enhances the accuracy and reliability of controlling actuators through improved noise cancellation and pattern recognition, allowing for more precise control using otoacoustic signals.
Smart Images

Figure NL2025050364_05022026_PF_FP_ABST
Abstract
Description
[0001] Title: system for processing of sound data
[0002] TECHNICAL FIELD
[0003] The various aspects and examples thereof relate to processing of otoacoustic sound.
[0004] BACKGROUND
[0005] Whereas the large public is aware their ears can be used to process sound and extract data therefrom, few are aware that their ears are also capable of emitting sound. Research on this has been done by for example by S.N. Lovich (Parametric information about eye movements is sent to the ears, Proceedings of the National Academy of Sciences, 2023 Vol. 120 No. 48 e2303562120, https: / / doi.org / 10.1073 / pnas.2303562120).
[0006] SUMMARY
[0007] It is preferred to provide an improved system for controlling matter or actuators by means of otoacoustic signals.
[0008] A first aspect provides a system for processing of audio data. The system comprises a left earpiece comprising a left audio sensor and a left audio actuator, the left audio sensor being arranged to receive an audio signal from a left ear of a subject and a right earpiece comprising a left audio sensor and a right audio actuator, the right audio sensor being arranged to receive an audio signal from a right ear of a subject. The system further comprises an electronic circuit arranged to provide an electronic stimulus signal to at least a first of the left audio actuator and the right audio actuator to produce an audio signal by means of the first audio actuator, receive an electronic signal from a first of the left audio sensor and the right audio sensor, compare data comprised by the received electronic signal to reference data electronically stored in the circuit, and based on the outcome of the comparing, generate an output electronic signal. This system allows for control of an actuator in response to otoacoustic signals received by means of an earpiece. Emitted otoacoustic sounds vary by variation in behaviour of a person. If a person purposely behaves in a particular way - by moving eyes in a particular direction or by focussing attention left, right, up, down, behave in another way of a combination of two or more thereof -, emission of otoacoustic sounds may be actively and consciously controlled by the person - at least to a certain extent. By comparing a particular sound pattern to pre-stored sound patterns coupled to a particular intended action of a user, sound patterns may be recognised. And such known sound patterns may be coupled to particular actions by the system - or a system or single actuator coupled thereto. This allows a user to control an actuator by means of otoacoustic sound emitted by one or more ears of the user.
[0009] In an example, the electronic circuit is arranged to provide a left electronic stimulus signal to the left audio actuator and a right stimulus electronic signal to the right audio actuator to produce a left audio signal by means of the left audio actuator and a right audio signal by means of the right audio actuator, and receive a left electronic signal from a first of the left audio sensor and a right electronic signal from the right audio sensor.
[0010] Use of two ears increases a degree of freedom in the system allowing for more information to be generated, transmitted, recognised and used.
[0011] Another example further comprises a noise cancellation unit arranged to reduce common mode noise received from the left microphone and received from the right microphone. The level of otoacoustic emissions are relatively low and may be lost in noisy environments. By reducing or even cancelling common mode noise simultaneously and bilaterally present in both ear canals, the accuracy of the system is improved.
[0012] A further example further comprises a differential amplifier arranged to produce an amplified signal based on a difference between the left audio signal and the right audio signal, both which comprise both correlated otoacoustic emissions (a bilateral unified differential signal) and (uncorrelated) noise, and wherein the amplified signal is provided to the electronic signal for comparing data comprised by the amplified signal to the reference data. By using a differential amplifier, common mode noise may be reduced or even cancelled and due to the selective amplification of the correlated difference in otoacoustic emissions from the left and right inner ears, the signal of interest is selectively amplified. Such common mode noise may be outside noise, like a siren of an ambulance, as well as inside body noise, like breathing and the beating of the heart.
[0013] In yet another example, the electronic circuit is further arranged to convert the electronic signal received by the electronic circuit to the frequency domain and the electronic circuit is arranged to compare the received data to the reference data in the frequency domain. In the frequency domain, some signal processing and pattern recognition requires less processing. One example is convolution, but this applies to other operations as well. In the frequency domain, not only spectrum data may be available, but also phase data of a signal.
[0014] In again a further example, the electronic circuit is further arranged to generate unwrapped phase data for the converted received electronic signal. With unwrapped phase data available, the system is more accurate than with wrapped phase data. Furthermore, it makes the system less dependent on the initial status of the system. It is noted that the initial position of the outer hair cells and stereocilia is unknown, in the left cochlea as well as the right cochlea. Yet, movement of the outer hair cells and stereocilia is correlated, within the same cochlea as for the right cochlea compared to the left cochlea. By processing unwrapped phase data, a higher accuracy is achieved.
[0015] In again another example, the electronic circuit is arranged to convert the received signal from the analogue domain to the digital domain. The digital domain is less sensitive to external noise. This allows for improved signal processing.
[0016] In yet a further example, the electronic stimulus signal is generated in the digital domain and the electronic circuit further comprises a digital to analogue converter to convert the generated electronic stimulus signal from the digital domain to the analogue domain and the electronic stimulus signal is provided to at least one of the left audio actuator and the right audio actuator in the analogue domain. Whereas processing of data may often be more efficient in the digital domain for various reasons, the external world is analogue and in the end, data has to be presented in analogue way.
[0017] In another example, the electronic circuit is arranged to have operation of the analogue to digital converter and operation of the digital to analogue converter synchronised in time relative to one another. Providing synchronisation, for example by means of a synchronisation unit, accuracy may be increased.
[0018] In a further example, the electronic circuit is arranged to operate a residual neural network arranged to compare data comprised by the received electronic signal to reference data provided to the residual neural network for training the residual neural network. Residual neural networks, deep learning neural networks, neural networks in general and artificial intelligence networks in general, implemented in electronic processing circuits, provide a myriad of option for pattern recognition.
[0019] Whereas they are predominantly known for pattern recognition in images, they may be used for pattern recognition of sounds and in sounds as well. This may be used in the time domain as well as in the frequency domain, optionally with phase data combined.
[0020] In again another example, the electronic circuit is further arranged to generate an electronic instruction signal that causes at least one of the left speaker and the right speaker to generate spoken instructions to a user. This allows a user to be instructed, by means of audible and / or visible instructions - or even tactile instructions.
[0021] In a further example, the electronic circuit is arranged to provide the electronic stimulus signal in conjunction with the spoken instruction. The stimulus signal may trigger generation of otoacoustic sounds by the ear. By providing instructions to the person to whom the ear belongs, the user may be instructed to exhibit certain behaviour. With the behaviour, a particular otoacoustic sound pattern may be generated and emitted. By receiving a particular sound pattern, with know behaviour, a pattern may be received tied to a behaviour and stored for comparison with other sounds patterns with unknown behaviour. This allows the system to be trained.
[0022] In one example, the electronic stimulus signal is provided after the electronic instruction signal has been provided to the at least one of the left speaker and the right speaker. In this case, a user may prepare for exhibiting the instructed behaviour.
[0023] In yet a further example, the electronic circuit is arranged to receive the electronic signal from a first of the left audio sensor and the right audio sensor upon providing the electronic stimulus signal. This example allows instantaneous and full-duplex operation.
[0024] In again another example, data comprised by the received electronic signal is incorporated into the reference data in conjunction with data comprised by the electronic instruction signal. This allows to build a reference model for pattern recognition for received signals from the audio sensors.
[0025] A second aspect provides a control system for an actuator comprising the system according to the first aspect and an interface arranged to receive the electronic output signal; and control the actuator based on the electronic output signal. This provides for a system for control of an actuator by means of otoacoustic emissions. More in detail, it allows for control of the actuator by means of behaviour that triggers particular otoacoustic emissions. Such behaviour may be physically, by means of using particular body muscles, for example connected to an eye, but the behaviour may also be mentally, by actively generating particular thoughts or thinking patterns.
[0026] In an example, the interface is further arranged to generate an electronic input data signal and the electronic circuit is arranged to receive the electronic input signal and generate the electronic stimulus signal in response to the electronic input signal, this allows for direct control of the actuator.
[0027] A third aspect provides a method of processing of audio data. The method comprises providing, via a left audio actuator provided in a left earpiece a left stimulus audio signal to a left ear, receiving, via a left audio sensor provided in a left earpiece, an audio signal from a left ear of a subject and generating, by means of a left audio sensor, a right electronic ear signal, providing, via a right audio actuator provided in a right earpiece a right stimulus audio signal to a right ear, receiving, via a right audio sensor provided in a right earpiece, an audio signal from a right ear of a subject and generating, by means of a left audio sensor, a left electronic ear signal, comparing data comprised by the left electronic ear signal and the right electronic ear signal to reference data electronically stored in the circuit, based on the outcome of the comparing, generate an output electronic signal.
[0028] BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The various aspects and examples thereof will now be discussed in further detail in conjunction with drawings. In the drawings:
[0030] Figure 1: shows a system for processing of otoacoustic data;
[0031] Figure 2: shows a first flowchart depicting a procedure for processing of otoacoustic data; and Figure 3: shows a second flowchart depicting a procedure for training a pattern recognition system for use in the procedure of Figure 2.
[0032] DETAILED DESCRIPTION
[0033] Figure 1 shows a system 100 for processing audio data. The system 100 comprises a headphone comprising a left earpiece 110 and a right earpiece 120. The left earpiece 110 comprises a left speaker 112 as a left audio actuator and a left microphone 114 as a left audio sensor. The right earpiece 120 comprises a right speaker 122 as a right audio actuator and a right microphone 124 as a right audio sensor.
[0034] The left earpiece 110 and the right earpiece 120 may be provided as over-ear, on-ear or in-ear headpieces. The left earpiece 110 is arranged to engage with a left ear 102 and the right earpiece 120 is arranged to engage with a right ear 104. The left speaker 112 and the right speaker 122 are arranged to provide sound to the left ear 102 and the right ear 104, respectively. Sound is in the context of this description to be understood as vibration of air or another gaseous medium, though in certain cases, a solid or liquid medium may be applicable as well.
[0035] The left microphone 114 is arranged to receive sound from the left ear 102 and the right microphone 124 is arranged to receive sound from the right ear 104. As such, the left microphone 114 and the right microphone 124 are arranged on the left earpiece 110 and the right earpiece 120 such that they are enable to capture sound emitted from the left ear 102 and the right ear 104, respectively. In this way, the left microphone 114 and the right microphone 124 are arranged to capture both otoacoustic emissions and noise (correlated and uncorrelated) from the left ear 102 and the right ear 104, respectively.
[0036] The left microphone 114 and the right microphone 124 are connected to a differential amplifier 132 provided in an analogue front end 130 of the system 100. The differential amplifier 132 may be implemented as a high speed low noise instrumental amplifier. This amplifier reduces common mode noise from the bilaterally placed microphones in 2 ears and selectively amplifies a differential otoacoustic emission signal. The differential amplifier 132 may be provided with a high pass filter. In one example, the high pass filter has a cut-off frequency of 10 Hz. In another example, the cut-off frequency is higher, though lower than a general frequency at which otoacoustic signals are provided. In another example, a high-pass filter is provided as well at this stage.
[0037] In practice, the cut-off frequency of such high pass filter may be between 3 Hz and 30 Hz, more in particular between 4 Hz and 25 Hz, even more in particular between 7 and 15 Hz. With such cut-off frequency, DC noise and 1 / / noise is filtered out from the input signals.
[0038] The analogue front end 130 further comprises a left output amplifier 134 and a right output amplifier 136. The left output amplifier 134 is arranged to provide a left output electronic signal to the left speaker 112 to cause the left speaker 112 to produce sound comprising left output audio data carried by the left output electronic signal. The right output amplifier 136 is arranged to provide a right electronic signal to the right speaker 114 to cause the right speaker 122 to produce sound comprising right output audio data carried by the right output electronic output signal.
[0039] The system 100 further comprises a baseband circuit 140 arranged for processing data. The processing of data by the baseband circuit 140 is in this example predominantly handled in a digital domain; in other examples, this may be carried out in an analogue domain. The baseband circuit 140 may be an ASIC (application specific integrated circuit), it may be programmed on an FPGA (field programmable gate array), it may be programmed on a general purpose processing unit (CPU), a neural processing unit or a graphics processing unit (GPU) or may be provided as a combination of two or more thereof. As such, the baseband circuit 140 may be provided as one or more pieces of semiconductor material, like silicon.
[0040] The baseband circuit 140 comprises a baseband processing unit 142 arranged to control the various parts of the baseband circuit 140 as described above and below. The baseband circuit 140 further comprises a baseband memory module 144. The baseband memory module 144 may comprise one or both of volatile and non-volatile memory cells. The baseband memory module 144 is arranged to store data, including audio data received or to be sent. The baseband memory module 144 may further be arranged to store data for programming the baseband processing unit 142 to carry out procedures as discussed above and below.
[0041] The baseband circuit further comprises an analogue to digital converter 152 arranged to convert the output of the differential amplifier 132 from the analogue domain to the digital domain. The analogue to digital converter 152 may operate at very high frequencies, up to 125 MHz, taking 125 million samples per second. In another example, the analogue to digital converter 152 take approximately 44 thousand samples per second at 44 kHz, twice the maximum frequency that may theoretically be heard by a human being. In again another example, the analogue to digital converter 152 operates at 4kHz, twice the frequency of known otoacoustic emissions from human ears. In yet another example, the sampling frequency is one or two orders of magnitude lower, i.e. 400 Hz or 40 Hz - within margins of about 20%; most important for these examples is the order of magnitude, not the exact values.
[0042] The digital signal, carrying input audio data received by the left speaker 114 and the right speaker 124, is provided to a Fast Fourier Transformation - FFT - module 154. The FFT module 154 provides a frequency representation of a set of samples in the time domain and a phase characteristic thereof. In addition to a representation of frequency components of the input audio data, the FFR module 154 is arranged to provide wrapped phase data and unwrapped phase data. This is in Figure 1 shown as a phase-frequency output channel 156.
[0043] The phase-frequency output channel 156 is provided to a residual neural network 158 as a pattern recognition module. In this example, the residual neural network 158 is a Resnetl8, a deep learning neural network having 18 layers. In other examples, other artificial intelhgences and other machine learning techniques may be used. In yet other examples, a conventional pattern recognition module may be used, in which the data provided by the phase-frequency output channel 156 is compared to phasefrequency data that stored by the baseband memory module 144.
[0044] The output of the residual neural network 158 is coupled to a host 170. The host 170 comprises a control interface 176 arranged to control an actuator 190 based on the output of the residual neural network 158. The actuator 190 may for example be an electrical switch for switching a light source on or off, an electrical lock arranged to lock or unlock a door, other, or combination thereof. The host 170 further comprises a trigger module 178 arranged to provide an input signal to the baseband circuit 140.
[0045] The host 170 is in this example controlled by a host processing unit 172. To the host processing unit 172, a host memory module 174 is coupled. The host memory module 174 is arranged to store data, including trigger data to be sent to the baseband circuit 140 and data for controlling the actuator 190. The host memory module 174 may further be arranged to store data for programming the host processing unit 172 to carry out procedures as discussed above and below.
[0046] The input signal provided by the trigger module 178 is provided to a stimulus generator 164 comprised by the baseband circuit 140. The stimulus generator 164 generates a stimulus based on the input data, optionally in conjunction with data stored in the baseband memory module 144. The stimulus comprises output audio data to be provided to the left ear 102 and the right ear 104. In this example, the stimulus comprises audio data in the digital domain.
[0047] The stimulus may comprise data that may be converted to audible and understandable sound, for example spoken instructions. Additionally or alternatively, the stimulus may comprise generation data, arranged to be converted in stimulus sounds that trigger a left cochlea of the left ear 102 and / or a right cochlea of the right ear 104 to provide otoacoustic emissions.
[0048] The stimulus is provided to a digital to a digital to analogue converter 162, which provides a left analogue output signal to the left output amplifier 134 and right analogue output signal to the right output amplifier 136.
[0049] The analogue to digital converter 152 and the digital to analogue converter 162 may be synchronised by means of a synchronising unit 142. The analogue to digital converter 152 and the digital to analogue converter 162 do not necessarily have to operate at the same frequency and do not necessarily need to take the same amount of samples per time unit. Yet, it may be preferred that operations of one are a multiple of operations of the other.
[0050] The left output amplifier 134 and the right output amplifier 136 generate the left output electronic signal and the right output electronic signal, as discussed above.
[0051] The operation of the system 100 will be discussed in further detail in conjunction with a first flowchart 200 (Figure 2) and a second flowchart 300 (Figure 3). The various parts of the first flowchart 200 are briefly summarised below:
[0052] 202 start
[0053] 204 obtain stimuli
[0054] 206 convert stimuli to analogue
[0055] 208 amplify left stimulus 210 amplify right stimulus
[0056] 212 provide left stimulus sound to left ear
[0057] 214 provide right stimulus sound to right ear
[0058] 216 receive sound from left ear
[0059] 218 receive sound from right ear
[0060] 220 subtract and amplify received electronic signals
[0061] 222 convert difference signal to digital
[0062] 224 provide phase and frequency characteristics using FFT
[0063] 226 process phase and frequency characteristics by the Resnet
[0064] 228 provide output for Resnet to host
[0065] 230 control actuator
[0066] 232 end
[0067] The procedure starts in a terminator 202 and continues to step 204 in which the stimulus generator 164 obtains stimulus data. The stimulus data may comprise data enabling the speakers to emit stimulus sounds. Such sound may comprise two tones, which may be provided as sine waves, in the order of one to five kHz. Such stimulus data may be generated by the stimulus generator, may be retrieved from the baseband memory module 144, obtained otherwise or a combination of two or more thereof.
[0068] Next, in step 206, the stimulus data is transformed into a left analogue stimulus signal and a right analogue stimulus signal by means of the digital to analogue transformer 162, which analogue stimulus signals carry the stimulus data. In other examples, only the left analogue stimulus signal or the right analogue stimulus signal is generated.
[0069] In step 208, the left analogue stimulus signal is amplified by means of the left output amplifier 134 and in step 210, the right analogue stimulus signal is amplified by means of the right output amplifier 136. In step 212, the amplified left stimulus signal is provided to the left ear 102 as sound carrying the stimulus data or at least a left part thereof by means of the left speaker 112. In step 214, the amplified right stimulus signal is provided to the right ear 104 as sound carrying the stimulus data or at least a right part thereof by means of the right speaker 122.
[0070] In response to the audio data provided to the ears, the ears provide otoacoustic sound. The otoacoustic sound may carry data. The data may be carried in a single signal provided by a single ear and / or the data may be carried by the set of sounds of the left ear 102 and the right ear 104, for example in a difference between both sound signals. In step 216, the left otoacoustic sound is received from the left ear 102 by means of the left microphone 114 and transformed to a left input electronic signal. In step 218, the right otoacoustic sound is received from the right ear 104 by means of the right microphone 124 and transformed to a right input electronic signal.
[0071] In step 220, the left electronic input signal and the right electronic input signal are subtracted from one another to reduce common mode noise. The left electronic input signal may be subtracted from the right electronic input signal or the other way around. The otoacoustic differential signal is amplified and provided to the analogue to digital converter 152. The subtraction of the two signals allows for cancellation of any common mode noise, i.e. of noise present in both signals. The analogue to digital converter 152 converts the differential signal to the digital domain in step 222.
[0072] The digitised signal is provided, in interval that may or may not overlap, to the FFT module 154. The FFT module 154 provides at least one of a frequency image of the digitised input signal in the interval, a wrapped phase image of the digitised input signal in the interval and an unwrapped phase image of the digitised input signal in the interval in step 224. The time interval may be between 10 milliseconds and 200 milliseconds, more in particular between 100 milliseconds ad 150 milliseconds, for example 125 milliseconds. In step 226, the phase and frequency data thus generated is provided to the residual neural network 158. In this example, the residual neural network 158 has been trained as discussed below in further detail, which allows the residual neural network 158 to recognise patterns in at least one of the frequency and phase data that it receives. More in general, residual neural networks like Resnetl8 are arranged to process image data. Such image data may be provided in square image snippets of 224 by 224 pixels, in three channels comprising data on red, green and blue, for each pixel. For audio data discussed here, 50176 data points may be provided in such data object, using for example the red channel for frequency, the green channel for wrapped phase and the blue channel for unwrapped phase.
[0073] By virtue of the training, the residual neural network is arranged to recognise patterns in the audio data received and arranged to provide an output based on data carried by the differential otoacoustic sounds received by the system 100 in a way that allows to process that data, for example for controlling the actuator 190. The output of the residual neural network 158 is in step 228 provided to the host 170 and the control interface 176 in particular. In step 230, the control interface 176 controls the actuator 190 based on the data received from the residual neural network 158 and the procedure ends in a terminator 232.
[0074] As discussed above, Figure 3 shows a second flowchart 300. The second flowchart 300 depicts a training procedure for training the residual neural network 158. The various parts of the second flowchart 300 are briefly summarised below.
[0075] 302 start procedure
[0076] 304 obtain instruction data
[0077] 306 obtain stimuli
[0078] 308 mix stimuli and instruction data
[0079] 310 converted mixed signal to analogue 312 amplify left sound data
[0080] 314 amplify right sound data
[0081] 316 provide left sound data to left ear
[0082] 318 provide right sound data to right ear
[0083] 320 receive sound from left ear
[0084] 322 receive sound from right ear
[0085] 324 subtract and amplify received electronic signals
[0086] 326 convert difference signal to digital
[0087] 328 provide phase and frequency characteristics using FFT 330 annotate phase and frequency characteristics
[0088] 332 provide phase and frequency characteristics to residual neural network
[0089] 334 end procedure
[0090] The procedure starts in a terminator 302 and proceeds to step 304 in which instruction data is obtained. The instruction data comprises data that enables to generate sound with spoken text. The instruction data may comprise instructions for a user to whom the left ear 102 and the right ear 104 belong to focus on left, right, top, down, other or a combination of two or more thereof. In another example, user instruction data is provided by means of visual indicators.
[0091] The instruction data may be obtained from the baseband memory module 144. Next, in step 206, the stimulus generator 164 obtains stimulus data. Such stimulus data may be generated by the stimulus generator, may be retrieved from the baseband memory module 144, obtained otherwise or a combination of two or more thereof.
[0092] In step 308, the stimulus data and the instruction data are mixed to be provided together in sound carrying both types of data, at the same time. This is an optional step, if indeed the stimulus data and the instruction data are to be provided at the same time to the left ear 102 and the right ear 104. In this example, the instruction data and the stimulus data are mixed in the digital domain. In another example, the instruction data and the stimulus data are mixed in the analogue domain. If the user instructions are provided by means of visual indicators or otherwise not to be provided by means of the left speaker 112 and the right speaker 122, the mixing step may be omitted.
[0093] In step 310, the mixed signal thus provided to converted from the digital domain to the analogue domain by means of the digital to analogue converter 162. In step 312, the left analogue mixed signal is amplified by means of the left output amplifier 134 and in step 314, the right analogue mixed signal is amplified by means of the right output amplifier 136.
[0094] In step 316, the amplified left mixed signal is provided to the left ear 102 as sound carrying the stimulus data and the instruction data or at least a left part thereof by means of the left speaker 112. In step 318, the amplified right mixed signal is provided to the right ear 104 as sound carrying the stimulus data and the instruction data or at least a right part thereof by means of the right speaker 122.
[0095] In response to the sound provided to the ears and the stimulus data in particular, the ears provide otoacoustic sound. The otoacoustic sound may carry data. In step 320, the left otoacoustic sound is received from the left ear 102 by means of the left microphone 114 and transformed to a left input electronic signal. In step 322, the right otoacoustic sound is received from the right ear 104 by means of the right microphone 124 and transformed to a right input electronic signal.
[0096] In step 324, the left electronic input signal and the right electronic input signal are subtracted from one another, as discussed in conjunction with the first flowchart 200 (Figure 2). The left electronic input signal may be subtracted from the right electronic input signal or the other way around. The differential signal is amplified and provided to the analogue to digital converter 152. The analogue to digital converter 152 converts the differential signal to the digital domain in step 326.
[0097] The digitised signal is provided, in interval that may or may not overlap, to the FFT module 154. The FFT module 154 provides at least one of a frequency image of the digitised input signal in the interval, a wrapped phase image of the digitised input signal in the interval and an unwrapped phase image of the digitised input signal in the interval in step 328.
[0098] In step 330, the phase and frequency data thus generated is annotated with the instruction data obtained in step 304. In step 332, the annotated phase and frequency data is provided to the residual neural network for training the residual neural network. The residual neural network - or other artificial intelligence or other pattern recognition system - thus trained may be used for the procedure as depicted by the first flowchart 200 of Figure 2.
[0099] In summary, the disclosure relates to a system that is provided for processing of otoacoustic signals. A pattern recognition module is provided with reference data to recognise known patterns in audio data received. Based on the recognition of particular sound data, the system may control an actuator. By training a user to exhibit particular behaviour that triggers particular otoacoustic sounds with particular patterns, and by building a proper reference database, the user may be able to control the actuator by means of the behaviour, via otoacoustic emissions by the ears of the user. The reference database may be a conventional database, but also a neural network or other artificial intelligence.
Claims
Claims1. A system for processing of audio data, the system comprising: a left earpiece comprising a left audio sensor and a left audio actuator, the left audio sensor being arranged to receive an audio signal from a left ear of a subject; a right earpiece comprising a right audio sensor and a right audio actuator, the right audio sensor being arranged to receive an audio signal from a right ear of a subject; an electronic circuit arranged to: provide an electronic stimulus signal to at least a first of the left audio actuator and the right audio actuator to produce an audio signal by means of the first audio actuator; receive an electronic signal from a first of the left audio sensor and the right audio sensor; compare data comprised by the received electronic signal to reference data electronically stored in the circuit; based on the outcome of the comparing, generate an output electronic signal.
2. The system according to claim 1, wherein the electronic circuit is arranged to: provide a left electronic stimulus signal to the left audio actuator and a right stimulus electronic signal to the right audio actuator to produce a left audio signal by means of the left audio actuator and a right audio signal by means of the right audio actuator; and receive a left electronic signal from a first of the left audio sensor and a right electronic signal from the right audio sensor.
3. The system according to any one of the preceding claims, further comprising a noise cancellation unit arranged to reduce commonmode noise received from the left audio sensor and received from the right audio sensor.
4. The system according to claim 2 or claim 3, further comprising a differential amphfier arranged to produce an amplified signal based on a difference between the left audio signal and the right audio signal, the amplified signal being based on a differential otoacoustic signal and noise, and wherein the amplified signal is provided to the electronic circuit for comparing data comprised by the amplified signal to the reference data.
5. The system according to any of the preceding claims, wherein the electronic circuit is further arranged to convert the electronic signal received by the electronic circuit to the frequency domain and the electronic circuit is arranged to compare the received data to the reference data in the frequency domain.
6. The system according to claim 5, wherein the electronic circuit is further arranged to generate phase data based on the converted received electronic signal.
7. The system according to claim 5 or claim 6, wherein the electronic circuit is further arranged to generate unwrapped phase data based on the converted received electronic signal.
8. The system according to any one of the preceding claims, wherein the electronic circuit is arranged to convert the received signal from the analogue domain to the digital domain.
9. The system according to claim 8 to the extend dependent on any one of claim 5 to claim 7, wherein the electronic circuit comprisesan analogue to digital converter arranged to converted the received signal to the frequency domain in the digital domain.
10. The system according to any one of the preceding claims, wherein the electronic stimulus signal is generated in the digital domain and the electronic circuit further comprises a digital to analogue converter to convert the generated electronic stimulus signal from the digital domain to the analogue domain and the electronic stimulus signal is provided to at least one of the left audio actuator and the right audio actuator in the analogue domain.
11. The system according to claim 10 to the extent dependent on claim 8, wherein the electronic circuit is arranged to have operation of the analogue to digital converter and operation of the digital to analogue converter synchronised in time relative to one another.
12. The system according to any one of the preceding claims, wherein the electronic circuit is arranged to operate a residual neural network arranged to compare data comprised by the received electronic signal to reference data provided to the residual neural network for training the residual neural network.
13. The system according to any one of the preceding claims, wherein the electronic circuit is further arranged to generate an electronic instruction signal that causes at least one of the left speaker and the right speaker to generate spoken instructions to a user.
14. The system according to claim 13, wherein the electronic circuit is arranged to provide the electronic stimulus signal in conjunction with the spoken instruction.
15. The system according to claim 14, wherein the electronic stimulus signal is provided after the electronic instruction signal has been provided to the at least one of the left speaker and the right speaker.
16. The system according to any one of the claims 13 to 15, wherein the electronic circuit is arranged to receive the electronic signal from a first of the left audio sensor and the right audio sensor upon providing the electronic stimulus signal.
17. The system according to claim 16, wherein data comprised by the received electronic signal is incorporated into the reference data in conjunction with data comprised by the electronic instruction signal.
18. A control system for an actuator comprising the system according to any one of the preceding claims and an interface arranged to: receive the electronic output signal; and control the actuator based on the electronic output signal.
19. The control system according to claim 18, wherein the interface is further arranged to generate an electronic input data signal and the electronic circuit is arranged to receive the electronic input signal and generate the electronic stimulus signal in response to the electronic input signal.
20. Method of processing of audio data, the method comprising: providing, via a left audio actuator provided in a left earpiece a left stimulus audio signal to a left ear;receiving, via a left audio sensor provided in a left earpiece, an audio signal from a left ear of a subject and generating, by means of a left audio sensor, a right electronic ear signal; providing, via a right audio actuator provided in a right earpiece, a right stimulus audio signal to a right ear; receiving, via a right audio sensor provided in a right earpiece, an audio signal from a right ear of a subject and generating, by means of a right audio sensor, a right electronic ear signal; comparing data comprised by the left electronic ear signal and comprised by the right electronic ear signal to reference data electronically stored in the circuit; and based on the outcome of the comparing, generate an output electronic signal.
Citation Information
Patent Citations
Information processing device, and information processing method
US20230306095A1
Personalization of auditory stimulus
US9497530B1
Methods and devices for applying hypothermic therapy to a human auditory system
WO2023014342A1