Method for automatic voice tuning and sound system using the method

The method and system address the challenge of personalized voice tuning in karaoke products by using gender recognition and pitch detection to enhance singing performance through real-time adjustments, improving user experience.

WO2026112875A1PCT designated stage Publication Date: 2026-06-04HARMAN INT IND INC +1

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HARMAN INT IND INC
Filing Date
2024-11-28
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing karaoke products fail to provide personalized voice tuning for users, particularly addressing gender-specific pitch differences and pitch-related singing difficulties, leading to off-key and voice crack issues.

Method used

A method and system that utilize gender recognition and pitch detection to automatically adjust microphone input signals in real-time, applying tailored reverberation, echo, and harmonic component adjustments based on gender and pitch analysis.

Benefits of technology

Enhances the singing experience by providing personalized voice tuning, compensating for pitch weaknesses and improving vocal performance for a wider range of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024135194_04062026_PF_FP_ABST
    Figure CN2024135194_04062026_PF_FP_ABST
Patent Text Reader

Abstract

This disclosure provides a method of automatic voice tuning for a sound system and the sound system using the method. The method may comprise: obtaining, via a microphone, an input signal representative of a user's voice; obtaining gender information; obtain a detected pitch by performing a pitch detection on the input signal; and applying a tuning control to the input signal based on the gender information and the detected pitch.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD FOR AUTOMATIC VOICE TUNING AND SOUND SYSTEM USING THE METHODTECHNICAL FIELD

[0001] The present disclosure relates to audio processing, and specifically relates to an optimized method for automatic voice tuning and a sound system using the method.BACKGROUND

[0002] When singing has become an indispensable form of entertainment for people to relieve stress in daily life, portable products with karaoke functions are favored by many people for their unique convenience and friendliness. People can choose to sing at home with friends or take the portable product out to sing outside. There have been many products that support ‘karaoke mode’ , such as JBL Party Box, JBL Party Box encore and so on. However, most of the products on the market mainly focus on the sound performance of the product itself, and rarely pay attention to the user's ability and whether they can sing easily. In daily life, it’s common to see people who are not good at singing may face problems of singing off key or voice crack when singing a difficult song. But for professional singers, how to sing well is much easier than for most ordinary users. Thus, it would be a great feature if karaoke products could compensate for ordinary users’ weakness in singing.

[0003] In the professional field, the sound engineer will make on-site real-time adjustments based on the live sounds they hear. They may also get to know the singer's singing habits through rehearsals in advance to ensure the real-time tuning performance of a live show. Such a method is more professional, but also more inconvenient and not suitable for consumer products. In the field of consumption, it’s a common way to add reverberation or echo to the products to improve the karaoke performance. Such a method can somewhat help improve the user’s singing performance through weakening the impact of defective direct sounds by adding reverberation, making it difficult for singers and audiences to detect some defects. However, these added reverberations are basically fixed, which cannot cover most people. For example, the same reverberation will have different effects on a girl's singing voice and a boy's singing voice, and the evaluation of that will be good or bad.

[0004] Therefore, improved technology is needed to overcome the above defects.SUMMARY

[0005] According to one aspect of the disclosure, a method of automatic voice tuning for a sound system is provided. The method may comprise obtaining, via a microphone, an input signal representative of a user’s voice; obtaining gender information; performing a pitch detection on the input signal to obtain a detected pitch; and applying a tuning control to the input signal based on the gender information and the detected pitch.

[0006] According to another aspect of the present disclosure, a sound system is provided. The sound system may comprise a memory configured to store instructions, and a processor coupled to the memory. The processor may be configured to perform the instructions to obtain, via a microphone, an input signal representative of a user’s voice; obtain gender information; obtain a detected pitch by performing a pitch detection on the input signal; and apply a tuning control to the input signal based on the gender information and the detected pitch.

[0007] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium comprising computer-executable instructions which, when executed by a computer, causes the computer to perform the method disclosed herein.BRIEF DESCRIPTION OF THE DRAWINGS

[0008] FIG. 1 illustrates a schematic diagram of automatic tuning for a sound system according to one or more embodiments of the present disclosure.

[0009] FIG. 2 shows an example of a voice-based gender recognition model using a CNN model.

[0010] FIG. 3 illustrates a method of automatic tuning for a sound system according to one or more embodiments of the present disclosure.

[0011] FIG. 4 illustrates an example that shows the detected pitch in a period of a female’s singing.

[0012] FIG. 5 illustrates the corresponding tuning control of reverberation and echo, e.g. gain control in the period of the female’s singing, based on the detected pitch of FIG. 4.

[0013] FIG. 6 illustrates an example that shows the harmonic components adjustment of one frame of the input voice signal.

[0014] To facilitate understanding, identical reference numerals have been used, where possible, to designate identical elements that are common to the figures. It is contemplated that elements disclosed in one embodiment may be beneficially utilized in other embodiments without specific recitation. The drawings referred to here should not be understood as being drawn to scale unless specifically noted. Also, the drawings are often simplified and details or components are omitted for clarity of presentation and explanation. The drawings and discussion serve to explain the principles discussed below, where similar designations denote similar elements. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Examples will be provided below for illustration. The descriptions of the various examples will be presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those with ordinary skills in the art without departing from the scope and spirit of the described embodiments.

[0016] The disclosure provides a new approach used for the sound system (such as products or portable products with karaoke mode) , which adopts a new method to automatically tune the user’s singing voice based on gender recognition and pitch detection, so that the sound system can adjust the microphone input signal in real-time based on the user's singing situation, and output the adjusted signal to the loudspeaker for playback after being processed with the karaoke algorithm, thereby providing the user with a more pleasant singing experience. This is achieved by using an optimized method of automatic tuning based on a combination of gender recognition and pitch detection. With gender recognition, the method can apply different tuning controls to users with different genders (such as females and males) . With pitch detection, the method can adjust the level / amount of reverberation / echo and compensate for the weakness in pitch. The proposed method and system herein can provide a flexible adjustment for singing in karaoke mode and ensure that this method and system can bring a better singing experience to most people. The proposed method and system will be explained in detail with reference to FIGS. 1-6 as follows.

[0017] FIG. 1 illustrates a schematic diagram of automatic tuning for a sound system according to one or more embodiments of the present disclosure. As shown in FIG. 1, the user's voice is picked up by the microphone and transmitted to the loudspeaker for playback after being processed with the karaoke algorithm. As can be recognized by those skilled in the art, there are additional components / modules to support karaoke algorithms applied to sound systems. In order to concisely and prominently explain the new approach proposed by the present disclosure, only the necessary components / modules for explaining the working principle of the new approach are shown in Process module 100 in FIG. 1.

[0018] Among these components / modules, Gender Recognition module 102 and Pitch Detection module104 are two key components / modules combined to analyze the user’s voice for further tuning. Based on the features of the voice recognized by the Gender Recognition module 102 and detected by the Pitch Detection module 104, Tuning Control module 106 helps control two associated ‘Gain&EQ’ modules generating adjustable gain control factors and equalization filters. Gain&EQ I module 108 is used to adjust the amount of reverberation and echo, and Gain&EQ II module 110 is used to adjust the energy of harmonic components of the overall output which includes direct sound, and adjusted reverberation and echo. Delay module 112 is used to align the delay between the direct sound and the adjusted reverberation and echo.

[0019] Thanks to the development of machine learning techniques, it can help to realize voice-based gender recognition. The key idea is to train a classification model based on the features of voice like Mel-scaled power spectrogram (Mel) , Mel-frequency cepstral coefficients (MFCCs) , power spectrogram chroma (Chroma) , tonal centroid, etc. As can be recognized by those skilled in the art, there are many commonly used classification models, such as a Convolutional Neural Network (CNN) model, a Recurrent Neural Network (RNN) model, a ResNet34 model, etc. Those with ordinary skills in the art can understand that they can choose any of the commonly used classification models to realize the algorithm implemented in the Gender Recognition module 102.

[0020] FIG. 2 illustrates an example of a voice-based gender recognition model using a CNN model. Usually, the convolutional neural network is mainly divided into five layers, i.e., a data input layer, a convolution Layer, a ReLU excitation layer, a pooling layer, and a fully connected layer. These five layers are connected in turn, and each layer accepts the output feature data of the previous layer and provides it to the next layer. The data input layer extracts the features of the input data. In this example, the input data may be audio data. The convolution calculation layer performs convolution mapping on these features. The excitation layer uses a non-linear excitation function to excite neurons to meet the conditions and transmit the feature information to the next neuron. The pool layer is used to compress the amount of data and parameters, thus reducing over-fitting. The fully connected layer is used to connect the feature information of all output layers, and summarize and sort out the information to complete the output. In the example shown in FIG. 2, it can be understood that the class result, i.e., gender information associated with the user’s voice can be obtained using the gender recognition model.

[0021] Users sometimes experience out-of-tune or cracked voices when singing because the pitch of the song is too high or too low. Pitch is a perceptual property of sounds, which is quantified by frequency and measured in Hertz (Hz) . It is a major auditory attribute of musical tones that makes it possible to judge sounds as “higher” and “lower” in the sense associated with musical melodies. In general, a female’s pitch range may be higher, and a male’s may be lower. The proposed method in the present disclosure performs pitch detection on the input voice signal via the Pitch Detection module 104 to realize automatic tuning by analyzing the user's pitch based on the characteristics of the pitch.

[0022] As those skilled in the art can recognize, there have been many studies on pitch detection algorithms. The main idea of pitch detection is to estimate the fundamental frequency of the input signal. This can be done in the time domain, for example, using YIN algorithm, McLeod Pitch Method (MPM) algorithm, etc., in the frequency domain using Harmonic Product Spectrum algorithm, etc., or in both the time domain and the frequency domain using Yet Another Algorithm for Pitch Tracking (YAAPT) algorithm, etc. Alternatively, that can be done with machine learning, such as TensorFlow Self-Supervised Pitch Estimation (TensorFlow SPICE) model. As those skilled in the art can understand, they can choose any of the existing algorithms or models to realize the pitch detection algorithm implemented in the Pitch Detection module 102.

[0023] FIG. 3 illustrates a method of automatic tuning for a sound system according to one or more embodiments of the present disclosure. The sound system may be any kind of system with karaoke functions and may comprise at least one microphone and at least one loudspeaker. The sound system may further comprise a memory and a processor. The memory may be configured to store computer-readable instructions or codes for causing the processor to carry out the aspects of the present disclosure. The processor may be any technically feasible hardware unit configured to process data and execute software applications, including without limitation, a central processing unit (CPU) , a microcontroller unit (MCU) , an application-specific integrated circuit (ASIC) , a digital signal processor (DSP) chip and so forth. It can be recognized that the discussed method in the present disclosure may be realized by a processor included in the sound system.

[0024] At S302, an input signal representative of a user’s voice may be obtained via a microphone when the user starts singing with the karaoke mode of the sound system.

[0025] At S304, gender information may be obtained. In some embodiments, the gender information may be obtained by performing a gender recognition on the input signal, for example, via the Gender Recognition module 102. Thus, the gender information indicative of the gender of the user may be obtained in real-time. Based on the gender information, different tuning controls will be applied to the Gain&EQ I module 108 and Gain&EQ II module 110. In some embodiments, the gender information may be obtained from the use’s input. For example, the user can manually select the pitch range via an input interface of the product. Based on the selected pitch range, the gender information can be obtained.

[0026] At S306, a pitch detection may be performed on the input signal, for example, via the Pitch Detection module 102. Then, a detected pitch can be obtained in real time based on the pitch detection. In some embodiments, the detected pitch may be indicative of a fundamental frequency of the user’s voice. It can be understood that harmonic components related to the fundamental frequency of the user’s voice can also be obtained from the pitch detection, since the harmonic components are multiples of the fundamental frequency, including 2nd harmonic, 3rd harmonic…, etc. The fundamental frequency and its related harmonic components obtained from the pitch detection may be further analyzed in combination with the gender information obtained at S304.

[0027] At S308, a tuning control may be applied to the input signal based on the gender information and the detected pitch. In common, due to differences in physiological structure, males and females have different pitch ranges. Usually, the pitch range of male voices is around 98~350Hz, and the pitch range of female voices is around 220~740Hz. Thus, the operational frequency band on which the turning control is performed will be different for different genders. So, when the gender information is detected, the tuning control module 106 will work on the different operational frequency bands based on the gender information.

[0028] In some embodiments, before the tuning control is applied, it may be determined whether a tuning control should be performed according to the detected pitch in combination with the gender information. This may be done by a comparison between the detected pitch and a pitch threshold selected from preset pitch thresholds. In some embodiments, different pitch thresholds for different genders are preset. For example, for a male, the low-pitch threshold may be preset to 90Hz and the high-pitch threshold may be preset to 250Hz. For a female, the low-pitch threshold may be preset to 220Hz and the high-pitch threshold may be preset to 400Hz. It should be recognized that the values of the above thresholds are presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. These thresholds may be modified according to different system requirements.

[0029] In some embodiments, the determination may comprise selecting the pitch threshold based on the gender information, wherein the pitch threshold may include a low-pitch threshold and a high-pitch threshold. A pitch below the low-pitch threshold may be considered as a low pitch, and a pitch above the high-pitch threshold may be considered as a high pitch. If a high / low pitch is detected, it means that it would be a hard part to sing, and the method / system should automatically tune the input voice signal to compensate for the weakness in pitch and related harmonic components.

[0030] In some embodiments, the determination may further comprise comparing the detected pitch with the selected high-pitch threshold and the selected high-pitch threshold respectively, and determining that the turning control should be performed in response to determining that the detected pitch is below the low-pitch threshold or is above the high-pitch threshold.

[0031] As an example, if the detected gender information indicates the user who is singing is a male, the method will select the first pitch threshold associated with the male gender, wherein the first pitch threshold including the first low-pitch threshold (e.g. 90Hz) and the first high-pitch threshold (e.g. 250Hz) . The tuning control module 106 will work in response to the detected pitch being below 90Hz or above 250Hz.

[0032] As another example, if the detected gender information indicates the user who is singing is a female, the method will select the second pitch threshold associated with the female gender, wherein the second pitch threshold including the second low-pitch threshold (e.g. 220Hz) and the second high-pitch threshold (e.g. 400Hz) . The tuning control module 106 will work in response to the detected pitch being below 220Hz or above 400Hz.

[0033] In some embodiments, the turning control may comprise a first control for the Gain&EQ I module 108 and a second control for the Gain&EQ II module 110. Both the first control and the second control are used to control the two associated modules generating adjustable gain factors and equalization filters. The first control aims to an adjustment of reverberation and echo which will be added to the input signal to help beautify the user’s voice. The second control aims to an adjustment of the fundamental frequency and harmonic components of the overall output which includes direct sound (i.e., the input signal) and the adjusted reverberation and echo to help reduce the difficulties faced by the user when singing.

[0034] In some examples, the first control is configured to adjust the amount / level of reverberation and echo through the adjustable gain factors and equalization filters, and the second control is configured to adjust the energy / magnitude of the fundamental frequency and related harmonic components through the adjustable gain factors and equalization filters.

[0035] For example, if it is determined based on the gender information that a girl is singing, and at the same time, a high / low pitch is detected via the pitch detection, then the method and system will automatically perform the tuning control on the input signal by adding more reverberation and echoes and increasing the energy of the fundamental frequency and harmonic components. The harmonic components are multiples of the fundamental frequency, including 2nd harmonic, 3rd harmonic…, etc. The harmonic components of different orders are independently adjustable, the balances between the adjustment of the harmonic components depend on the gender information, the detected pitch and harmonics.

[0036] FIG. 4 is an example showing the detected pitch in a period of a female’s singing. FIG. 5 shows the corresponding tuning control of reverberation and echo, e.g. gain control in the period of the female’s singing. In this example, the gain of reverberation and echo remains constant when the detected pitch is medium (i.e., the detected pitch is considered not high or not low after a comparison with a high-pitch threshold or a low-pitch threshold) and increases as the pitch gets higher. In some examples, a smoothing process may be performed on the gain, to avoid sudden changes in singing performance.

[0037] FIG. 6 illustrates an example that shows the harmonic components adjustment of one frame of the input voice signal. As shown in FIG. 6, curve 602 indicates the magnitude profile with the adjustment via the Gain&EQ II module 110, and curve 604 indicates the magnitude profile without the adjustment via the Gain&EQ II module 110. For example, the detected pitch for this frame is determined as high pitch, around 474Hz. In response to the determination of high pitch, the method and system may automatically adjust the gain of the fundamental frequency and related harmonics. For example, as the higher-order harmonics become weaker, higher gain will be added for high-order harmonics.

[0038] The descriptions of the various embodiments have been presented for purposes of illustration, but are not intended to be exhaustive or limited to the embodiments disclosed. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

[0039] In the preceding, reference is made to embodiments presented in this disclosure. However, the scope of the present disclosure is not limited to specific described embodiments. Instead, any combination of the preceding features and elements, whether related to different embodiments or not, is contemplated to implement and practice contemplated embodiments. Furthermore, although embodiments disclosed herein may achieve advantages over other possible solutions or over the prior art, whether or not a particular advantage is achieved by a given embodiment is not limiting of the scope of the present disclosure. Thus, the preceding aspects, features, embodiments and advantages are merely illustrative and are not considered elements or limitations of the appended claims except where explicitly recited in a claim (s) .

[0040] Aspects of the present disclosure may take the form of an entire hardware embodiment, an entire software embodiment (including firmware, resident software, micro-code, etc. ) or an embodiment combining software and hardware aspects that may all generally be referred to herein as a “circuit, ” “block” , “module” , “unit” or “system. ”

[0041] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium (or media) having computer-readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.

[0042] The computer-readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer-readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer-readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM) , a read-only memory (ROM) , an erasable programmable read-only memory (EPROM or Flash memory) , a static random access memory (SRAM) , a portable compact disc read-only memory (CD-ROM) , a digital versatile disk (DVD) , a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer-readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber-optic cable) , or electrical signals transmitted through a wire.

[0043] Computer-readable program instructions described herein can be downloaded to respective calculating / processing devices from a computer-readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers.

[0044] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) , and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0045] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0046] The flowchart and block diagrams in the drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function (s) . In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.

[0047] While the foregoing is directed to embodiments of the present disclosure, other and further embodiments of the disclosure may be devised without departing from the basic scope thereof, and the scope thereof is determined by the claims that follow.

[0048] Clause 1. In some embodiments, a method of automatic voice tuning for a sound system, comprising: obtaining, via a microphone, an input signal representative of a user’s voice; obtaining gender information; obtaining a detected pitch by performing a pitch detection on the input signal ; and applying a tuning control to the input signal based on the gender information and the detected pitch.

[0049] Clause 2. The method according to clause 1, wherein the gender information is indicative of a gender of the user, and the detected pitch is indicative of a fundamental frequency of the user’s voice.

[0050] Clause 3. The method according to any one of clauses 1-2, wherein the turning control comprises a first control configured to adjust reverberation and echo which will be added to the input signal, and a second control configured to adjust the fundamental frequency and its related harmonic components.

[0051] Clause 4. The method according to any one of clauses 1-3, wherein the applying a tuning control to the input signal based on the gender information and the detected pitch comprises: selecting one of a first pitch threshold and a second pitch threshold based on the gender information; comparing the detected pitch with the selected pitch threshold; and performing the tuning control based on the comparison.

[0052] Clause 5. The method according to any one of clauses 1-4, wherein the comparing the detected pitch with the selected pitch threshold comprises comparing the detected pitch with the selected pitch threshold to determine whether the turning control should be performed.

[0053] Clause 6. The method according to any one of clauses 1-5, wherein the first pitch threshold includes a first low-pitch threshold and a first high-pitch threshold; and wherein the second pitch threshold includes a second low-pitch threshold and a second high-pitch threshold.

[0054] Clause 7. The method according to any one of clauses 1-6, further comprises: determining the turning control should be performed in response to determining that the detected pitch is below the first low-pitch threshold or is above the first high-pitch threshold; or determining that the turning control should be performed in response to determining that the detected pitch is below the second low-pitch threshold or is above the second high-pitch threshold.

[0055] Clause 8. In some embodiments, a sound system comprising: a memory configured to store instructions; and a processor configured to perform the instructions to: obtain, via a microphone, an input signal representative of a user’s voice; obtain gender information; obtain a detected pitch by performing a pitch detection on the input signal; and apply a tuning control to the input signal based on the gender information and the detected pitch.

[0056] Clause 9. The sound system according to clause 8, wherein the gender information is indicative of the gender of the user, and the detected pitch is indicative of a fundamental frequency of the user’s voice.

[0057] Clause 10. The sound system according to any one of clauses 8-9, wherein the turning control comprises a first control configured to adjust reverberation and echo which will be added to the input signal, and a second control configured to adjust the fundamental frequency and its related harmonic components.

[0058] Clause 11. The sound system according to any one of clauses 8-10, wherein the processor is further configured to select one of a first pitch threshold and a second pitch threshold based on the gender information; compare the detected pitch with the selected pitch threshold; and perform the tuning control based on the comparison.

[0059] Clause 12. The sound system according to any one of clauses 8-11, wherein the processor is further configured to compare the detected pitch with the selected pitch threshold to determine whether the turning control should be performed.

[0060] Clause 13. The sound system according to any one of clauses 8-12, wherein the first pitch threshold includes a first low-pitch threshold and a first high-pitch threshold; and wherein the second pitch threshold includes a second low-pitch threshold and a second high-pitch threshold.

[0061] Clause 14. The sound system according to any one of clauses 8-13, wherein the processor is further configured to: determine that the turning control should be performed in response to determining that the detected pitch is below the first low-pitch threshold or is above the first high-pitch threshold; or determine that the turning control should be performed in response to determining that the detected pitch is below the second low-pitch threshold or is above the second high-pitch threshold.

[0062] Clause 15. In some embodiments, a computer-readable storage medium comprising computer-executable instructions which, when executed by a computer, causes the computer to perform the method according to any one of clauses 1-7.

Claims

1.A method of automatic voice tuning for a sound system, comprising:obtaining, via a microphone, an input signal representative of a user’s voice;obtaining gender information;obtaining a detected pitch by performing a pitch detection on the input signal ; andapplying a tuning control to the input signal based on the gender information and the detected pitch.2.The method according to claim 1, wherein the gender information is indicative of a gender of the user, and the detected pitch is indicative of a fundamental frequency of the user’s voice.3.The method according to claim 2, wherein the turning control comprises a first control configured to adjust reverberation and echo which will be added to the input signal, and a second control configured to adjust the fundamental frequency and its related harmonic components.4.The method according to claim 1, wherein the applying a tuning control to the input signal based on the gender information and the detected pitch comprises:selecting one of a first pitch threshold and a second pitch threshold based on the gender information;comparing the detected pitch with the selected pitch threshold; andperforming the tuning control based on the comparison.5.The method according to claim 4, wherein the comparing the detected pitch with the selected pitch threshold comprises comparing the detected pitch with the selected pitch threshold to determine whether the turning control should be performed.6.The method according to claim 5,wherein the first pitch threshold includes a first low-pitch threshold and a first high-pitch threshold; andwherein the second pitch threshold includes a second low-pitch threshold and a second high-pitch threshold.7.The method according to claim 6, further comprises:determining that the turning control should be performed in response to determining that the detected pitch is below the first low-pitch threshold or is above the first high-pitch threshold; ordetermining that the turning control should be performed in response to determining that the detected pitch is below the second low-pitch threshold or is above the second high-pitch threshold.8.A sound system comprising:a memory configured to store instructions; anda processor configured to perform the instructions to:obtain, via a microphone, an input signal representative of a user’s voice;obtain gender information;obtain a detected pitch by performing a pitch detection on the input signal; andapply a tuning control to the input signal based on the gender information and the detected pitch.9.The sound system according to claim 8, wherein the gender information is indicative of the gender of the user, and the detected pitch is indicative of a fundamental frequency of the user’s voice.10.The sound system according to claim 9, wherein the turning control comprises a first control configured to adjust reverberation and echo which will be added to the input signal, and a second control configured to adjust the fundamental frequency and its related harmonic components.11.The sound system according to claim 8, wherein the processor is further configured to:select one of a first pitch threshold and a second pitch threshold based on the gender information;compare the detected pitch with the selected pitch threshold; andperform the tuning control based on the comparison.12.The sound system according to claim 11, wherein the processor is further configured to compare the detected pitch with the selected pitch threshold to determine whether the turning control should be performed.13.The sound system according to claim 12,wherein the first pitch threshold includes a first low-pitch threshold and a first high-pitch threshold; andwherein the second pitch threshold includes a second low-pitch threshold and a second high-pitch threshold.14.The sound system according to claim 13, wherein the processor is further configured to:determine that the turning control should be performed in response to determining that the detected pitch is below the first low-pitch threshold or is above the first high-pitch threshold; ordetermine that the turning control should be performed in response to determining the detected pitch is below the second low-pitch threshold or is above the second high-pitch threshold.15.A computer-readable storage medium comprising computer-executable instructions which, when executed by a computer, causes the computer to perform the method according to any one of claims 1-7.