Synthesis of lingsl's sound

By combining the decomposition and whitening synthesis source with shape components, stable and adjustable tone synthetic Lin's tone was generated, solving the recording fluctuations and individual differences of Lin's tests in the prior art, and improving the accuracy and efficiency of Lin's tests.

CN120476296APending Publication Date: 2025-08-12MED-EL ELECTRONIC MEDICAL EQUIP CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380089307.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-29
Filing Date
2023-12-28
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The prior art is difficult to generate stable, natural and customizable synthetic Lin's sounds. It is used in Lin's tests, and there are problems of recording fluctuations and individual differences in the invigilator.

Method used

By decomposing the spectrum of human recordings into shape components and source components, providing a synthetic source and whitening to combine with the shape components, a hybrid Lin's audio spectrum is generated and converted into a time domain signal.

Benefits of technology

A stable, adjustable and fluctuating synthetic Lin's sound is produced, suitable for cochlear implant assembly, reducing individual differences and improving the accuracy and efficiency of the test.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120476296A_ABST
    Figure CN120476296A_ABST
Patent Text Reader

Abstract

A system and method for creating a mixed Lin's sound for a Lin's test using a recording of a human emanating Lin's sound. First, a recording of a human being is decomposed into at least one shape component. The synthesis source is then whitened and multiplied by the shape component of the human recording to create a customizable synthetic Lin's sound with less tone fluctuations.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 435,970, filed on December 29, 2022, the entire contents of which are incorporated herein by reference. Technical Field

[0002] The present invention relates to a system and method for synthesizing a synthetic Lin sound, and more particularly to decomposing the frequency spectrum of a human recording of a Lin sound to provide shape components, and then adding a synthesis source to the shape components to generate a synthetic Lin sound. Background Art

[0003] Audiologists use speech tests to assess a patient's ability to understand speech. This makes speech tests an integral part of the lives of cochlear implant users and others with hearing impairments. Primarily, they are used to determine a hearing-impaired patient's eligibility for a cochlear implant. Subsequently, speech tests measure the success of the implant and can even guide the implant fitting and refitting process.

[0004] A particularly interesting speech test is the Ling test, named after Daniel Ling. The test consists of six sounds chosen to cover the auditory spectrum associated with human (originally North American) speech. The typical six Ling sounds in English are: (1) AH, (2) EE, (3) OO, (4) M, (5) S, and (6) SH. In German, the typical six Ling sounds are: (1) A, (2) I, (3) U, (4) M, (5) S, and (6) SCH.

[0005] The Lin test is used during the cochlear implant fitting process. During fitting, the implant is fitted with a lower and upper limit to ensure that the patient can hear every sound without discomfort. The Lin test is a useful method for confirming a patient's hearing because it breaks down speech into just six tones.

[0006] The Lin test is superior to typical speech-based listening tests for several reasons. Lin sounds are quick and simple, consisting of just a few short items, whereas speech tests are used with a set of words or sentences. Lin sounds are context-independent, so listeners are only measured on what they hear, without any external guesswork. Lin sounds can be used in a variety of languages because they convey no meaning. Speakers of any language will be able to parse the sounds together to determine if their listening is adequate.

[0007] Lin sounds are also specific and informative. Traditional speech tests yield summary metrics, such as the percentage of words understood. Because the spectral characteristics of each Lin sound are known, the Lin test can more narrowly pinpoint hearing deficits. For example, a patient who has difficulty hearing the fifth Lin sound may have problems with high frequencies.

[0008] Typically, the Lin test is administered verbally by an invigilator, such as an audiologist. The invigilator simply says one of the tones and asks the patient if they can hear it, distinguish it from other tones, and repeat it. These questions determine detection, discrimination, and recognition, respectively.

[0009] While the procedure is simple to perform, it can result in significant variation depending on the examiner. Each examiner can have different pronunciations, volume, duration, and interpretive outcomes. Therefore, an alternative approach is to use recorded sounds. Recorded material during the test helps reduce variability, but it also introduces new challenges in administering the Lin test. The duration of the recorded sound depends on the skill of the human recorder. Furthermore, once a sound is recorded, the same sound is repeated over and over. Patients can eventually learn to identify a particular sound from non-speech-related cues, such as fluctuations in the recorded speech.

[0010] Therefore, it is advantageous to develop synthetically created Lin sounds. Synthesized sounds do not have the disadvantages of spoken sounds because they do not have varying fluctuations and can be generated at any time. In addition, the parameters that determine the characteristics of the speech can be adjusted in real time.

[0011] However, synthetic Lin sounds are difficult to produce. Popular speech synthesizers that produce realistic speech are not designed to produce the long vowels required for Lin sounds. Older speech synthesizers (such as the Klatt model) can produce long vowels, but they sound noticeably unnatural. Summary of the Invention

[0012] According to one embodiment of the present invention, a method for generating a mixed Lin sound signal is provided. The method includes decomposing a frequency spectrum of a Lin sound recording of a human speaker to provide a shape component; providing a synthesis source; whitening the synthesis source; combining the whitened synthesis source with the shape component to form a mixed Lin sound spectrum; and converting the mixed Lin sound spectrum into the time domain to form the mixed Lin sound signal.

[0013] According to a related embodiment of the present invention, the method may include adjusting the shape component to zero below a frequency value. Decomposing the recording to provide the shape component may include applying a sliding maximum filter to the magnitude of the spectrum of the recording based on a window size. The window size may be greater than or equal to the fundamental frequency of the synthesized source.

[0014] According to further related embodiments of the present invention, providing the original synthetic source may include providing the original synthetic source having a selected fundamental frequency. Providing the original synthetic source for voiced sounds may include providing a synthetic glottal stream signal. Providing the original synthetic source for unvoiced sounds may include providing Gaussian white noise.

[0015] According to another related embodiment of the present invention, the method may include recording multiple Ring sounds of a human speaker, wherein each recording generates a mixed Ring sound signal of its corresponding Ring sound. Optionally, each mixed Ring sound signal may have the same fundamental frequency, which may be 100 Hz.

[0016] According to another embodiment of the present invention, a non-transitory storage medium is provided, storing instructions that, when executed, establish a computer process or controller that: decomposes a frequency spectrum of a recording of a human speaker into a source component and a shape component; provides a synthesized source; whitens the synthesized source; combines the whitened synthesized source with the shape component to form a mixed Lindsay audio spectrum; and converts the mixed Lindsay audio spectrum into the time domain to form the mixed Lindsay audio signal.

[0017] According to a related embodiment of the invention, the computer process or controller may further include adjusting the shape component to zero below a frequency value. Decomposing the recording to provide the shape component may include applying a sliding maximum filter to the amplitude of the spectrum of the recording based on a window size. The window size may be greater than or equal to the fundamental frequency of the synthesized source.

[0018] According to further related embodiments of the present invention, providing the original synthetic source may include providing the original synthetic source having a selected fundamental frequency. Providing the original synthetic source for voiced sounds may include providing a synthetic glottal stream signal. Providing the original synthetic source for unvoiced sounds may include providing Gaussian white noise.

[0019] According to a further related embodiment of the present invention, the computer process or controller may further include: decomposing a plurality of recordings of a human speaker's Lin sound, wherein each recording generates a corresponding mixed Lin sound signal of its corresponding Lin sound. Each mixed Lin sound signal may have the same fundamental frequency, such as, but not limited to, 100 Hz. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above features of the embodiments will be more easily understood through the following detailed description in conjunction with the accompanying drawings, in which:

[0021] Figure 1 is a flowchart describing the process of a hybrid method for synthesizing Lin's sounds according to an embodiment of the present invention; and

[0022] Figure 2are several graphical depictions of sound spectra and portions of sound spectra according to an embodiment of the present invention. DETAILED DESCRIPTION

[0023] In an illustrative embodiment, a system and method for creating synthetic Lin sounds are provided. For example, the system and method can record a human Lin sound, decompose the recording into shape components, and add a synthetic source to the shape components to create a customizable synthetic Lin sound with minimal pitch fluctuation. Using synthetic Lin sounds is advantageous when fitting a cochlear implant or other hearing aid. Details are provided below.

[0024] Figure 1 This is a flow chart describing a process 100 for a hybrid method for synthesizing Lin sounds according to one aspect of the present invention. The hybrid method combines components of a recorded human speech with a synthetic source. The hybrid method can advantageously allow for the generation of natural-sounding vocalizations with adjustable pitch and duration without unpredictable fluctuations. In embodiments, the various steps of the flow chart can be performed by a computer system or controller. The computer system may include a non-transitory storage medium storing instructions for performing the steps of the process.

[0025] The process 100 first provides a recording of a human speaker, step 101. The recording will include one or more Lin sounds uttered by the human speaker.

[0026] The recorded sound is then converted into a frequency spectrum X(f)=FFT{x(t)} in the frequency domain, step 102 .

[0027] Next, X(f) is decomposed into two components 103: its source component and its shape component. This decomposition of sound into two components is generally based on how humans create speech. The vibrations and noise generated by the larynx are generally the source component. The vocal tract filters out the noise from the larynx to form the sound of speech, which is generally considered the shape component. Because the spectrum of a human speaker typically varies over time, the decomposed spectrum can be either an average spectrum or a momentary spectrum.

[0028] The shape component T(f) can be determined as the spectral envelope of the sound spectrum using the following equation: Where f-Δ / 2<=f' <f+Δ / 2

[0029] This equation determines the shape component by applying a sliding maximum filter to X(f). The source component S(f) can be determined from the equation S(f) = X(f) / T(f). However, it should be noted that the source component is generally not required for synthesizing mixed Lin sounds, and may not even be determined in various embodiments.

[0030] In various embodiments, the shape component is then zeroed in step 104 before being combined with the synthesized source in step 105. Zeroing the shape component helps suppress unwanted low-frequency artifacts that may be present. These low frequencies are generally irrelevant to the intelligibility of the Lin sound. In various embodiments, zeroing the shape component can include setting T(f) to zero for all values of f less than a certain value (e.g., 100 Hz). In other embodiments, the zeroing can be performed at a certain percentage below the fundamental frequency and / or at a certain percentage below the maximum value of T(f) over all values of f, i.e., for some value of f, T(f) is less than a fraction of max(T(f)) over all values of f.

[0031] In step 105, an original synthetic source is provided. The original synthetic source can be obtained from various sources, such as various databases or computers, or synthesized in other ways through an algorithm. The fundamental frequency of the original synthetic source, even if not completely accurate, can still represent the fundamental frequency of the final mixed synthetic Lin sound. Therefore, the original synthetic source can be advantageously customized with a desired fundamental frequency. The original synthetic source can help ensure that the sound is stable in time, that is, it does not fluctuate in pitch or loudness as human speech naturally does. In various embodiments, the original synthetic source for unvoiced sounds can be generated from Gaussian white noise, while the original synthetic source for voiced sounds can be generated from a synthetic glottal flow signal. Samples of synthetic glottal flow signals can be found in G. Fant, J. Liljencrantz, Q. Lin, "A four-parameter model of glottal flow", STL-QPSR, vol. 26, 1985. The entire contents of this document are incorporated herein by reference. In various embodiments, pulse trains can be used for the synthetic source, however, the result may sound more artificial than using glottal flow.

[0032] In step 106, the original synthetic source is whitened. Whitening the source may include adjusting the volume of each frequency to the same or similar value. In some embodiments, the original synthetic source provided may be a signal that has been whitened. Whitening the synthetic source may provide a constant power spectral density of the original synthetic source. In various embodiments, the whitened synthetic source may be calculated by dividing the spectrum of the original synthetic source by the spectrum envelope of the original synthetic source. This will result in uniformity of peaks in the spectrum of the original synthetic source, as each peak will be divided by itself. One advantage of whitening the synthetic source is that it ensures that each frequency component in the synthetic source has a volume, and the synthetic source can be multiplied by the shape component in a non-decreasing manner in step 107.

[0033] The whitened original synthesized source is then combined with the shape component to form a synthesized mixed Lin sound spectrum, step 107. The combination can be accomplished by multiplying the whitened synthesized source with the shape component. The mixed Lin sound spectrum can then be converted back to the time domain, step 108. Step 108 can be implemented by an inverse fast Fourier transform. Since the synthesized mixed Lin sound is generated using the original synthesized source, the synthesized mixed Lin sound typically has the same fundamental frequency as the original synthesized source. Therefore, when generating each different Lin sound, the user can advantageously make each Lin sound have the same fundamental frequency. During the Lin test, the consistent fundamental frequency makes it more difficult for the patient being tested to decipher each Lin sound from another Lin sound using pitch.

[0034] In step 109, a mixed Lin sound is played to perform the Lin test. The Lin test can be used when fitting a cochlear implant or other hearing device to a particular patient. For example, during cochlear implant fitting, the electrical pulses generated by the cochlear implant in response to sound waves are delivered at a level low enough to not harm the patient, yet high enough to ensure that the sounds are heard. Pulses are mapped between these two thresholds so that the patient can appropriately hear the desired range of sounds. Because English and other languages are full of sounds throughout speech, designing a test that breaks down speech into six tones can improve testing efficiency. While these sounds lack context in the form of sentences and words, an audiologist performing the test without the use of mixed Lin sounds can provide natural information about the nature of each sound. For example, the test provider (or recording) may be speech-deaf, speak some sounds in a higher or lower pitch, or have their lips read. The mixed Lin sound test described here avoids this information and provides a more accurate test of the fitting. Furthermore, if necessary, the parameters that determine speech characteristics can be adjusted in real time. In various embodiments, the audiologist may play a mixture of Lin sounds and ask the patient a) if they can hear the tone (detection), b) if they can distinguish the tone from other sounds (discrimination), c) or if they can repeat the sound (recognition).

[0035] Figure 2 1 is a diagram illustrating a plurality of graphs of audio spectra and portions of audio spectra according to an embodiment of the present invention. Each graph depicts an audio spectrum in the frequency domain, allowing for the creation of a mixed Lin sound that is stationary in time. Figure 2 Each plot in represents the value on the x-axis and the frequency on the logarithmic y-axis. Figure 1 As depicted, the initial spectrum 201 can be divided into a source component 202 and a shape component 203. In one embodiment, the initial spectrum 201 is a recording of an expected Lin's sound produced by a human at a specific instant or a recording of an average of expected Lin's sounds produced over a specific time period.

[0036] Source component 202 represents the various frequencies present in initial spectrum 201, typically with a constant power spectral density. Source component 202 contains the overall pitch of the sound and is responsible for temporal fluctuations. Since initial spectrum 201 is based on human speech, there are many temporal fluctuations, graphically represented as the width of the base of the peaks in initial spectrum 201 and source component 202. Furthermore, due to the pitch changes caused by natural tonality, source component 202 is filled with higher frequencies. Not shown in the figure is the time-varying state of human speech, whose frequency and pitch fluctuate. Therefore, it is necessary to provide a synthetic source and discard the initial source.

[0037] An original synthetic source 210 is provided as a source component for synthesizing the Lin sounds. The synthetic source 210 has a customizable fundamental frequency, although in some embodiments, the fundamental frequency is determined from the initial source component. For different sounds, various synthetic sources can be used. In some embodiments, unvoiced sounds such as S and SH are generated by a Gaussian white noise synthetic source, and voiced sounds such as AH, EE, OO, and M are generated using a synthetic glottal stream signal. Further, in some embodiments, it is advantageous to whiten the original synthetic source to create a constant spectral power density, thereby allowing the source to be multiplied non-decreasingly by the shape component. The whitened synthetic source 211 is derived from the synthetic source 210. In one embodiment, the whitened synthetic source 211 is obtained by applying a sliding maximum filter to the amplitude of the original synthetic source spectrum (similar to the process described for obtaining the shape component), with Δ equal to the fundamental frequency to obtain the spectral envelope. The whitened synthetic source is then divided by its own spectral envelope, essentially resulting in each peak being divided by itself.

[0038] The shape component 203 represents the envelope of the source frequencies in the spectrum. In various embodiments, the shape component 203 can be determined by taking the maximum value of X(f) over a certain period of time. For each frequency, the value of the shape component is the maximum value of the initial spectrum between two values. Typically, these two values are expressed as f–Δ / 2 and f+Δ / 2. In various embodiments for voiced sounds (i.e., AH, EE, OO, M), Δ must be at least the fundamental frequency or greater. In various embodiments for unvoiced sounds (i.e., S, SH), Δ is arbitrary, with lower Δ values providing higher quality results. In some embodiments, Δ is selected separately for each Lin sound, but for consistency, Δ can be the same for each Lin sound. For example, the fundamental frequencies of all voiced sounds can be measured in the initial spectrum, so that the highest fundamental frequency can be selected and rounded up to obtain Δ.

[0039] In various embodiments, the shape component 203 undergoes a zeroing process to obtain a zeroed shape component 204. Zeroing the shape component helps suppress unwanted low-frequency artifacts because low frequencies are irrelevant to the intelligibility of the Lin sound. Figure 2In the embodiment shown, the shape component 204 is set to 0 for all frequencies below 100 Hz. In various embodiments, nulling may occur at various frequencies.

[0040] Final spectrum 220 is a combination of the whitened synthesized source 211 and the zeroed shape component 204. In some embodiments, the whitened synthesized source spectrum 211 is multiplied, without limitation, by the shape component 204 to obtain final spectrum 220. Final spectrum 220 is similar to initial spectrum 201 in that the graphical representation is of the same Lindsay. However, final spectrum 220 has several advantageous differences. For example, the fundamental frequency of the first peak shown in final spectrum 220 is customizable. Because the source is synthesized, it can be more easily and accurately altered, for example, on a computer, than a human could alter their speech. Final spectrum 220 also has fewer key changes, as indicated by the cleaner bases of the peaks, and remains constant over time. This prevents listeners from being able to determine which Lindsay is being heard based on auditory cues.

[0041] In some embodiments, in order to play the sound, the final spectrum 220 will be converted back to the time domain.To convert the final spectrum 220 to the time domain, an inverse fast Fourier transform can be performed on the spectrum.

[0042] Various embodiments of the present invention may have the features of the potential claims listed in the paragraphs following this paragraph (and before the actual claims provided at the end of this application). These potential claims form part of the written description of this application. Therefore, in subsequent proceedings involving this application or any application claiming priority based on this application, the subject matter of the following potential claims may be presented as actual claims. The inclusion of such potential claims should not be construed to mean that the actual claims do not cover the subject matter of the potential claims. Therefore, a decision not to present these potential claims in a later proceeding should not be interpreted as donating the subject matter to the public.

[0043] Embodiment can be implemented as computer program product for computer system in whole or in part.This implementation can include a series of computer instructions, which are fixed on tangible media, such as computer readable media (for example, disk, CD-ROM, ROM or fixed disk), or via a modem or other interface device, such as a communication adapter connected to a network by a medium to be transferred to the computer system. The medium can be a tangible medium (for example, an optical communication line or an analog communication line), or it can be a medium implemented using wireless technology (for example, microwave, infrared or other transmission technology). A series of computer instructions embodies all or part of the functions described for the system before this article. It will be understood by those skilled in the art that these computer instructions can be written in many programming languages to be used together with many computer architectures or operating systems. In addition, this instruction can be stored in any memory device, such as a semiconductor, a magnetic memory device, an optical memory device or other memory device, and can be transmitted using any communication technology, such as light, infrared, microwave or other transmission technology. It is contemplated that such a computer program product may be distributed as a removable medium with accompanying printed or electronic documentation (e.g., compressed packaged software), preloaded onto a computer system (e.g., on a system ROM or fixed disk), or distributed from a server or electronic bulletin board over a network (e.g., the Internet or the World Wide Web). Of course, some embodiments of the present invention may be implemented as a combination of software (e.g., a computer program product) and hardware. Yet another embodiment of the present invention may be implemented entirely in hardware or software (e.g., a computer program product).

[0044] The embodiments of the present invention described above are exemplary embodiments only; many variations and modifications will be apparent to those skilled in the art. All such variations and modifications are intended to be within the scope of the present invention as defined in the appended claims.

Claims

1. A method for generating a mixed Lin sound signal, characterized in that: The method comprises: Decomposing the spectrum of a recording of a human speaker's Lin sounds to provide a shape component; Provide synthetic sources; whitening the synthetic source; combining the whitened synthesis source with the shape component to form a mixed Lin audio spectrum; and The mixed Lin's sound spectrum is converted into a time domain to form the mixed Lin's sound signal.

2. The method according to claim 1, characterized in that Also includes: The shape components below the frequency value are adjusted to zero.

3. The method according to claim 1, characterized in that Decomposing the sound recording to provide the shape component includes applying a sliding maximum filter to the magnitude of the frequency spectrum of the sound recording based on a window size.

4. The method according to claim 3, characterized in that The window size is greater than or equal to the fundamental frequency of the synthetic source.

5. The method according to claim 1, wherein Providing the original said synthesized source includes providing the original said synthesized source having a selected fundamental frequency.

6. The method according to claim 5, characterized in that Providing the original said synthetic source for voiced sounds comprises providing a synthetic glottal stream signal.

7. The method according to claim 5, characterized in that Providing the original said synthetic source for the unvoiced speech comprises providing Gaussian white noise.

8. The method according to claim 1, characterized in that Also includes: A plurality of Lin-sound recordings of a human speaker are decomposed, wherein each recording generates a mixed Lin-sound signal of its corresponding Lin-sound.

9. The method according to claim 8, characterized in that Each mixed Lin's tone signal has the same fundamental frequency.

10. The method according to claim 2, characterized in that The frequency value is 100 Hz.

11. A non-transitory storage medium, characterized in that: storing instructions that, when executed, establish a computer process comprising: Decompose the spectrum of a recording of a human speaker into source and shape components; Provide synthetic sources; whitening the synthetic source; combining the whitened synthesis source with the shape component to form a mixed Lin audio spectrum; and The mixed Lin's tone spectrum is converted into a time domain to form a mixed Lin's tone signal.

12. The non-transitory storage medium according to claim 11, wherein: The computer process also includes adjusting the shape components below a frequency value to zero.

13. The non-transitory storage medium according to claim 11, wherein: Decomposing the sound recording to provide the shape component includes applying a sliding maximum filter to the magnitude of the frequency spectrum of the sound recording based on a window size.

14. The non-transitory storage medium according to claim 13, wherein: The window size is greater than or equal to the fundamental frequency of the synthetic source.

15. The non-transitory storage medium according to claim 11, wherein: Providing the original said synthesized source includes providing the original said synthesized source having a selected fundamental frequency.

16. The non-transitory storage medium according to claim 15, wherein: Providing the original said synthetic source for voiced sounds comprises providing a synthetic glottal stream signal.

17. The non-transitory storage medium according to claim 15, wherein: Providing the original said synthetic source for the unvoiced speech comprises providing Gaussian white noise.

18. The non-transitory storage medium according to claim 11, wherein: The computer process also includes decomposing a plurality of human speaker recordings of Lin sounds, wherein each recording generates a corresponding mixed Lin sound signal of its respective Lin sound.

19. The non-transitory storage medium according to claim 18, wherein: Each mixed Lin's tone signal has the same fundamental frequency.

20. The non-transitory storage medium according to claim 12, wherein: The frequency value is 100 Hz.

21. A system for generating mixed Lin sounds, characterized in that The system comprises: Controller for: Decompose the spectrum of a recording of a human speaker into source and shape components; Provide synthetic sources; whitening the synthetic source; combining the whitened synthesis source with the shape component to form a mixed Lin audio spectrum; and The mixed Lin's tone spectrum is converted into a time domain to form a mixed Lin's tone signal.

22. The system according to claim 21, wherein: The controller is further configured to adjust the shape component below a frequency value to zero.

23. The system according to claim 21, wherein: Decomposing the sound recording to provide the shape component includes applying a sliding maximum filter to the magnitude of the frequency spectrum of the sound recording based on a window size.

24. The system according to claim 23, wherein: The window size is greater than or equal to the fundamental frequency of the synthetic source.

25. The system according to claim 21, wherein Providing the original said synthesized source includes providing the original said synthesized source having a selected fundamental frequency.

26. The system according to claim 25, wherein: Providing the original said synthetic source for voiced sounds comprises providing a synthetic glottal stream signal.

27. The system according to claim 25, wherein: Providing the original said synthetic source for the unvoiced speech comprises providing Gaussian white noise.

28. The system according to claim 21, wherein: A recording of a human speaker includes a plurality of Lin sounds, wherein each recording generates a mixed Lin sound signal of its corresponding Lin sound.

29. The system according to claim 28, wherein: Each mixed Lin's tone signal has the same fundamental frequency.

30. The system according to claim 22, wherein: The frequency value is 100 Hz.

31. A system for assembling a cochlear implant into a human body and outputting the mixed sound signal according to claim 21, characterized in that: The controller is further configured to output the mixed sound signal for performing a Lin test.

32. A method for fitting a cochlear implant via a Lin test using the mixed Lin sound signal of claim 1, characterized in that: The method includes outputting the mixed Lin sound to perform a Lin test.

33. The method according to claim 32, characterized in that Also included is requesting the patient to perform a task selected from the group consisting of: detecting the Lin sounds, distinguishing the Lin sounds, and repeating the Lin sounds.