Panoramic sound recording karaoke earphone

By adopting multi-microphone array and sound source positioning technology in karaoke headphones, the problem of poor sound pickup effect of existing karaoke headphones is solved, achieving a panoramic surround effect and a higher immersive experience.

CN119967332AInactive Publication Date: 2025-05-09SHENZHEN SEVENSTAR TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510439752.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-05-09
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing karaoke headphones have poor sound pickup effect, insufficient sound field coverage, and errors in sound source judgment and positioning.

Method used

The panoramic sound recording karaoke headphone design is adopted, including a control box, a microphone array, a headphone head and a headphone cable. Multiple time domain signals are collected through the microphone array, and the audio signal rendered by binaural use of sound source positioning technology is used to output the audio signal rendered by binaurals.

Benefits of technology

The panoramic sound surround effect is achieved, reducing the error and noise of single microphone collection, and improving the wearer's immersive experience in the karaoke scene.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967332A_ABST
    Figure CN119967332A_ABST
Patent Text Reader

Abstract

The application discloses a panoramic sound recording karaoke earphone. The earphone comprises a control box, a microphone array, an earphone head and an earphone line. One end of the left earphone line and one end of the right earphone line are both connected with earphone heads, the other end of the left earphone line and the other end of the right earphone line are both connected with one end of the control box, the left microphone is arranged on the left earphone line, and the right microphone is arranged on the right earphone line; a middle microphone is arranged at one end of the control box, the other end of the left earphone line and the other end of the right earphone line are symmetrically arranged on the two sides of the middle microphone, and the length of the earphone line between the left microphone and the middle microphone is equal to the length of the earphone line between the right microphone and the middle microphone; the control box is configured to carry out sound source localization according to a plurality of time domain signals collected by the microphone array, output audio signals rendered by two ears, and transmit the audio signals to the earphone head through the earphone line. According to the application, singing sound of a wearer can be collected to a large extent, time domain signals in all directions in an environment can be collected, and surround of panoramic sound is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of earphone technology, and in particular to an all-around sound recording karaoke earphone. Background Art

[0002] Nowadays, wired-controlled microphone headphones are still the mainstream choice for mobile karaoke. When users are outdoors or traveling, they prefer the lightweight solution of "one machine and one ear". At the same time, wired-controlled microphone headphones have a high hardware penetration rate and strong versatility, and are compatible with almost all karaoke apps without the need for additional debugging and driver.

[0003] However, existing wire-controlled microphone headphones usually collect sound through one microphone, which often has problems such as a single sound pickup direction and insufficient sound field coverage, and there are errors in the judgment and positioning of the sound source. Summary of the invention

[0004] The purpose of this application is to provide a panoramic sound recording karaoke headset to solve the technical problem of poor sound pickup effect of existing karaoke headsets in the prior art. The preferred technical solutions among the many technical solutions provided in this application can produce many technical effects as described below.

[0005] To achieve the above objectives, this application provides the following technical solutions: The present application provides a panoramic sound recording karaoke headset, comprising a control box, a microphone array, an earphone head and an earphone cable; the earphone cable comprises a left earphone cable and a right earphone cable; the microphone array comprises a left microphone, a right microphone and a middle microphone; the earphone head comprises a speaker and an acoustic cavity for encapsulating the speaker; one end of the left earphone cable and one end of the right earphone cable are both connected to the earphone head, the other end of the left earphone cable and the other end of the right earphone cable are both connected to one end of the control box, the left microphone is arranged on the left earphone cable, and the right microphone is arranged on the right earphone cable; the middle microphone is arranged at one end of the control box, the other end of the left earphone cable and the other end of the right earphone cable are symmetrically arranged on both sides of the middle microphone, and the length of the earphone cable between the left microphone and the middle microphone is equal to the length of the earphone cable between the right microphone and the middle microphone; the control box is configured to perform sound source localization according to multiple time domain signals collected by the microphone array, output binaurally rendered audio signals, and transmit them to the earphone head via the earphone cable.

[0006] In some embodiments, the control box includes a shell, and the shell is provided with an acquisition module, an extraction module, a calculation module and a processing module; the acquisition module is configured to acquire the time domain signal collected by the microphone array; the extraction module is configured to extract the spatial acoustic feature information between every two microphones of the microphone array and each microphone itself according to the time domain signal acquired by the acquisition module; the calculation module calculates the azimuth and pitch angle of the sound source according to the spatial acoustic feature information extracted by the extraction module, wherein the spatial acoustic feature information includes time delay information, phase difference information and energy value; the processing module outputs the audio signal according to the azimuth and pitch angle of the sound source acquired by the calculation module.

[0007] In some embodiments, the extraction module includes a delay unit, a phase unit and a volume unit; the delay unit can be used to estimate the delay between every two microphones to obtain the delay information; the phase unit obtains the phase difference information between every two microphones after performing a fast Fourier transform on each time domain signal; the volume unit is used to obtain the energy value of the RMS energy of each time domain signal.

[0008] In some embodiments, the processing module includes a beam forming unit and a HRTF mapping unit; the beam forming unit is used to generate beams pointing in different directions to determine the direction of the sound source; the HRTF mapping unit is used to convolve the beam signal pointing to the direction of the sound source generated by the beam forming unit with the left and right ear HRTFs respectively, simulate the head filtering effect, and output the audio signal.

[0009] In some embodiments, the length of the headphone cable between the left microphone and the corresponding headphone head is greater than the length of the headphone cable between the left microphone and the middle microphone, and the length of the headphone cable between the right microphone and the corresponding headphone head is greater than the length of the headphone cable between the right microphone and the middle microphone.

[0010] In some embodiments, the microphone array includes at least two movable microphones movable along the headphone cable, one movable microphone is arranged on the headphone cable between the left microphone and the corresponding headphone head, and the other movable microphone is arranged on the headphone cable between the right microphone and the corresponding headphone head.

[0011] In some embodiments, a volume adjustment button is disposed on the surface of the control box, and the volume adjustment button is used to adjust the volume of the audio signal output to the earphone head.

[0012] In some embodiments, microphone switch buttons are provided on both side walls of the control box, one of the microphone switch buttons is used to control whether the left microphone collects the time domain signal, and the other microphone switch button is used to control whether the right microphone collects the time domain signal.

[0013] In some embodiments, microphone adjustment buttons are provided on both side walls of the control box, one of the microphone adjustment buttons is used to amplify or reduce the time domain signal collected by the left microphone, and the other microphone adjustment button is used to amplify or reduce the time domain signal collected by the right microphone.

[0014] In some embodiments, the panoramic sound recording karaoke headphones include a cable tie, which is arranged between the control box and the left microphone and the right microphone. The left headphone cable and the right headphone cable both pass through the interior of the cable tie in parallel, and the cable tie bundles the left headphone cable and the right headphone cable.

[0015] Implementing one of the above technical solutions of the present application has the following advantages or beneficial effects: In the present application, by setting a left microphone on the left headphone cable, setting a right microphone on the right headphone cable, and setting a middle microphone at the connection between the control box and the headphone cable, a microphone array is formed to collect multiple time domain signals in different directions, and the sound source is located based on these time domain signals, and finally an audio signal rendered by both ears is output. In this case, the microphone array can collect the singing voice of the wearer to a large extent, as well as the time domain signals in the front, back, left, right, top, bottom and oblique directions in the collection environment, so as to realize the surround of panoramic sound; compared with the method of collecting by using a single microphone, the microphone array of the present application collects time domain signals by each microphone cooperating with each other, and can accurately focus on the sound source in any direction, that is, it is not limited to the specific direction and distance of the sound source, and can reduce the error and noise caused by the single microphone collection, so that the sound direction and distance sense perceived by the wearer are the same as the reality, which improves the wearer's immersive experience in the karaoke scene. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. It is obvious that the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative work. In the drawings: Figure 1 2 is a schematic diagram of the structure of the panoramic sound recording karaoke headset according to an embodiment of the present application; Figure 2It is a structural block diagram of the control box of an embodiment of the present application.

[0017] In the figure: 1. Atmos recording karaoke earphones; 10. Control box; 20. Left earphone cable; 30. Right earphone cable; 40. Left microphone; 50. Right microphone; 60. Middle microphone; 70. Earphone head; 100. Acquisition module; 110. Extraction module; 120. Calculation module; 130. Processing module; 111. Delay unit; 112. Phase unit; 113. Volume unit; 131. Beamforming unit; 132. HRTF mapping unit; 140. Preprocessing module; 141. Frame windowing unit; 142. Filtering unit; 143. Amplification unit; 144. ADC unit; 150. Sound effect processing module; 160. Ear return mixer; 170. Noise reduction enhancement module; 151. Reverberation unit; 152. Sound effect enhancement unit; 171. Adaptive beamforming unit; 172. Howling suppression unit; 173. Embedding unit. DETAILED DESCRIPTION

[0018] In order to make the purpose, technical solutions and advantages of the present application clearer, the various exemplary embodiments to be described below will refer to the corresponding drawings, which constitute a part of the exemplary embodiments, wherein various exemplary embodiments that may be used to implement the present application are described. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation methods described in the following exemplary embodiments do not represent all implementation methods consistent with the present disclosure. It should be understood that they are only examples of processes, methods, and devices that are consistent with some aspects disclosed in the present application as detailed in the attached claims, and other embodiments may also be used, or the embodiments listed herein may be modified in structure and function without departing from the scope and essence of the present application.

[0019] In the description of the present application, it should be understood that the terms "center", "longitudinal", "lateral", etc. indicate the orientation or positional relationship based on the drawings, which is only for the convenience of describing the present application and simplifying the description, rather than indicating or implying that the elements referred to must have a specific orientation, be constructed and operated in a specific orientation. The terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. The term "multiple" means two or more. The terms "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, an integral connection, a mechanical connection, an electrical connection, a communication connection, a direct connection, an indirect connection through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to the specific circumstances.

[0020] In order to illustrate the technical solution described in the present application, a specific embodiment is provided below, and only the parts related to the embodiment of the present application are shown.

[0021] like Figure 1 and Figure 2 As shown, the present application provides a panoramic sound recording karaoke headset 1, including a control box 10, a microphone array, an earphone head 70 and an earphone cable; In some embodiments, the headphone cable may include a left headphone cable 20 and a right headphone cable 30; the microphone array may include a left microphone 40, a right microphone 50 and a center microphone 60; and the headphone head 70 may include a speaker and an acoustic cavity for encapsulating the speaker.

[0022] In some embodiments, one end of the left headphone cable 20 and one end of the right headphone cable 30 are both connected to an earphone head 70, the other end of the left headphone cable 20 and the other end of the right headphone cable 30 are both connected to one end of the control box 10, the left microphone 40 is set on the left headphone cable 20, and the right microphone 50 is set on the right headphone cable 30.

[0023] In some embodiments, a middle microphone 60 may be provided at one end of the control box 10, and the other end of the left headphone cable 20 and the other end of the right headphone cable 30 are symmetrically arranged on both sides of the middle microphone 60, and the length of the headphone cable between the left microphone 40 and the middle microphone 60 is equal to the length of the headphone cable between the right microphone 50 and the middle microphone 60.

[0024] The Cartesian coordinate system is constructed in the normal use posture of the panoramic sound recording karaoke headset 1. The microphone positions of the left microphone 40, the right microphone 50 and the middle microphone 60 of the panoramic sound recording karaoke headset 1 can be expressed by the following formula: L(−d,0,0),R(d,0,0),C(0,0,h), Wherein, L(−d, 0, 0) is the coordinate of the left microphone 40, R(d, 0, 0) is the coordinate of the right microphone 50, C(0, 0, h) is the coordinate of the center microphone 60, the origin position is the center of the wearer's head, d is the left-right distance, and h is the vertical height.

[0025] In some embodiments, the control box 10 may include a housing such as Figure 2 As shown, the housing may be provided with an acquisition module 100, an extraction module 110, a calculation module 120 and a processing module 130. Specifically, a circuit board or a controller carrying the acquisition module 100, the extraction module 110, the calculation module 120 and the processing module 130 may be provided inside the housing.

[0026] In some embodiments, the acquisition module 100 can be configured to acquire a time domain signal collected by a microphone array; the extraction module 110 can be configured to extract the spatial acoustic feature information between every two microphones of the microphone array and each microphone itself based on the time domain signal acquired by the acquisition module 100; the calculation module 120 can calculate the azimuth and pitch angle of the sound source based on the spatial acoustic feature information extracted by the extraction module 110, wherein the spatial acoustic feature information may include time delay information, phase difference information and energy value; the processing module 130 can output an audio signal based on the azimuth and pitch angle of the sound source acquired by the calculation module 120.

[0027] In some embodiments, the acquisition module 100 may be configured to acquire a first time domain signal acquired by the left microphone 40 , a second time domain signal acquired by the right microphone 50 , and a third time domain signal acquired by the center microphone 60 .

[0028] In some embodiments, the acquisition module 100 can synchronously acquire the first time domain signal, the second time domain signal and the third time domain signal, which are three-channel signals in total, thereby avoiding clock deviation of the acquired signals and better maintaining time alignment.

[0029] In some embodiments, the shell may be provided with a preprocessing module 140, and the preprocessing module 140 may include a frame splitting and windowing unit 141 and a filtering unit 142; the frame splitting and windowing unit 141 may be used to perform frame splitting and windowing processing on the first time domain signal, the second time domain signal and the third time domain signal, and the filtering unit 142 may be used to perform denoising processing on the first time domain signal, the second time domain signal and the third time domain signal.

[0030] Specifically, since the time domain signal has a time-varying property, the framing and windowing unit 141 can perform framing processing on the first time domain signal, the second time domain signal, and the third time domain signal according to a preset frame length and a preset frame shift. The preset frame length may range from 15 milliseconds to 40 milliseconds per frame. By setting the preset frame length, the problem of insufficient frequency resolution and decreased time resolution can be better solved. At the same time, by setting the preset frame shift, the loss of frame edge information can be reduced, so that the frames can smoothly transition.

[0031] In some embodiments, the framing windowing unit 141 can perform windowing processing on the time domain signal processed by framing. That is, the framing windowing unit 141 can perform windowing processing on the first time domain signal, the second time domain signal and the third time domain signal processed by framing. Specifically, the framing windowing unit 141 can perform frame-by-frame windowing on each frame of the time domain signal according to a preset window function, and the preset window function can be one of a Hamming window, a Kaiser window or a Blackman window. Thus, the framing windowing unit 141 can significantly reduce spectrum leakage, which is convenient for subsequent processing.

[0032] In some embodiments, the filtering unit 142 may process the time domain signal according to spectral subtraction or adaptive filtering to reduce background noise. Specifically, spectral subtraction may refer to subtracting the noise component from the spectrum of the noisy signal to retain the target signal, and adaptive filtering may refer to minimizing the error and suppressing the noise by dynamically adjusting the filter coefficients.

[0033] In some embodiments, the preprocessing module 140 may further include an amplifying unit 143 and an ADC unit 144. The amplifying unit 143 may amplify the time domain signal, and the ADC unit 144 may perform analog-to-digital conversion on the time domain signal. Thus, the level of the time domain signal can be amplified and increased to be sufficient to drive subsequent circuits. At the same time, after the analog signal is converted into a digital signal, the time domain signal can be subsequently processed in various ways using digital signal processing technology.

[0034] In some embodiments, the extraction module 110 may include a delay unit 111 , a phase unit 112 , and a volume unit 113 .

[0035] In some embodiments, the delay unit 111 can be used to estimate the delay for every two microphones and obtain delay information. That is, the delay unit 111 can obtain the first delay information between the left microphone 40 and the right microphone 50, the second delay information between the left microphone 40 and the middle microphone 60, and the third delay information between the right microphone 50 and the middle microphone 60. Specifically, the delay unit 111 can calculate the cross-correlation function of the three pairs of microphones, such as the generalized cross-correlation, and obtain the first delay information, the second delay information, and the third delay information according to the peak position.

[0036] In some embodiments, the first delay information may be obtained by the following formula: .

[0037] In some embodiments, the phase unit 112 can obtain the phase difference information between each two microphones after performing a fast Fourier transform on each time domain signal. Specifically, the phase unit 112 can obtain the first phase difference information between the left microphone 40 and the right microphone 50, the second phase difference information between the left microphone 40 and the middle microphone 60, and the third phase difference information between the right microphone 50 and the middle microphone 60.

[0038] In some embodiments, the volume unit 113 may be used to obtain the RMS (Root Mean Square) energy of each time domain signal. Specifically, the RMS energy may be used to describe the average power of the time domain signal.

[0039] In some embodiments, assuming a far-field plane wave, the calculation module 120 can construct a set of equations based on the spatial acoustic feature information and the coordinates of each microphone to solve the direction of the sound source. Further, the calculation module 120 can use the least squares method to obtain the azimuth and elevation angle of the sound source. At this time, the direction vector can be u=(sinθcosϕ, sinθsinϕ, cosθ), where θ is the elevation angle and ϕ is the azimuth angle.

[0040] In some embodiments, the processing module 130 may include a beamforming unit 131 and a HRTF (Head-Related Transfer Function) mapping unit 132 .

[0041] In some embodiments, the beamforming unit 131 can be used to generate beams pointing in different directions to determine the direction of the sound source. Specifically, the beamforming unit 131 can adjust the first delay information, the second delay information, and the third delay information according to the azimuth and elevation angle of the sound source so that the sound source direction signals are phase-aligned when superimposed. In this way, the propagation differences of different microphones can be compensated and in-phase superposition can be achieved. After compensating for the delay, the beamforming unit 131 can weight each time domain signal to form a beam signal.

[0042] In some embodiments, the beam forming unit 131 may compare the energy values ​​of each generated beam and determine the direction corresponding to the maximum energy value as the direction of the sound source. The energy value of each beam may be expressed by the following formula: E(θ, ϕ)=∫∣B(θ, ϕ)∣2dtE(θ, ϕ)=∫∣B(θ, ϕ)∣2dt, Among them, B(θ, ϕ) is the beam.

[0043] In some embodiments, the HRTF mapping unit 132 can match the sound source azimuth and pitch angle corresponding to the sound source azimuth with the HRTF database to obtain the left and right ear filter coefficients. The HRTF mapping unit 132 can be used to convolve the beam signal pointing to the sound source azimuth generated by the beam forming unit 131 with the left and right ear HRTFs respectively, simulate the head filtering effect, and output the audio signal.

[0044] In some embodiments, the length of the earphone cable between the left microphone 40 and the corresponding earphone head 70 may be greater than the length of the earphone cable between the left microphone 40 and the middle microphone 60, and the length of the earphone cable between the right microphone 50 and the corresponding earphone head 70 may be greater than the length of the earphone cable between the right microphone 50 and the middle microphone 60. Thus, the left microphone 40 and the right microphone 50 can be brought closer to the wearer's mouth, and a clearer human voice can be collected.

[0045] In some embodiments, the microphone array may include at least two movable microphones movable along the headphone cable, one movable microphone may be disposed on the headphone cable between the left microphone 40 and the corresponding headphone head 70, and another movable microphone may be disposed on the headphone cable between the right microphone 50 and the corresponding headphone head 70. Specifically, the left microphone 40 may be fixedly disposed on the left headphone cable 20, and one movable microphone may adjust its position on the left headphone cable 20 as required, and correspondingly, the right microphone 50 may be fixedly disposed on the right headphone cable 30, and another movable microphone may adjust its position on the right headphone cable 30 as required. In this case, the movable microphone can assist the corresponding microphone in collecting time domain signals, thereby enhancing the sound pickup effect.

[0046] In some embodiments, a volume adjustment button may be provided on the surface of the control box 10 , and the volume adjustment button may be used to adjust the volume of the audio signal output to the earphone head 70 .

[0047] In some embodiments, the panoramic sound recording karaoke headset 1 may include a connector, which can be used to connect to an external device, for example, it can communicate with a mobile phone, tablet and other devices, and the wearer can adjust various parameters of the panoramic sound recording karaoke headset 1 through the external device.

[0048] In some embodiments, the connector can be connected to the control box 10 by wired or wireless communication. For example, the connector can be connected to the other end of the control box 10 by a data cable, or can be connected by wireless transmission methods such as Bluetooth, WiFi, etc.

[0049] In some embodiments, both side walls of the control box 10 may be provided with microphone switch buttons, one of which may be used to control whether the left microphone 40 collects time domain signals, and the other may be used to control whether the right microphone 50 collects time domain signals. Specifically, the microphone switch button near the left earphone line 20 may be used to control the left microphone 40 to turn on or off the sound pickup function, and the microphone switch button near the right earphone line 30 may be used to control the right microphone 50 to turn on or off the sound pickup function. Thus, it is convenient for the wearer to adjust the sound pickup of the device according to the judgment of the on-site sound source.

[0050] In some embodiments, microphone adjustment buttons may be provided on both side walls of the control box 10, one microphone adjustment button is used to amplify or reduce the time domain signal collected by the left microphone 40, and the other microphone adjustment button is used to amplify or reduce the time domain signal collected by the right microphone 50. Specifically, the microphone adjustment button close to the left earphone line 20 can be used to amplify or reduce the first time domain signal, and the microphone adjustment button close to the right earphone line 30 can be used to amplify or reduce the second time domain signal. In this way, the sound pickup effects of the left microphone 40 and the right microphone 50 can be adjusted in real time.

[0051] In some embodiments, the panoramic sound recording karaoke headset 1 may include a cable tie, which can be set between the control box 10 and the left microphone 40 and the right microphone 50. The left headphone cable 20 and the right headphone cable 30 can pass through the inside of the cable tie in parallel, and the cable tie can gather the left headphone cable 20 and the right headphone cable 30.

[0052] In some embodiments, the housing is provided with a sound effect processing module 150 , an ear return mixer 160 and a noise reduction enhancement module 170 .

[0053] In some embodiments, the sound effect processing module 150 may include a reverberation unit 151 and a sound effect enhancement unit 152 .

[0054] In some embodiments, the reverberation unit 151 can use a feedback delay network (FDN) to generate multi-channel reverberation, wherein the reverberation time, pre-delay and high-frequency attenuation factor can be adjusted as required, the reverberation time can be controlled by adjusting the attenuation coefficient, the pre-delay can be used to simulate the space size, and the high-frequency attenuation factor can be used to simulate the air absorption effect.

[0055] The feedback delay network is an advanced digital signal processing technology used to generate artificial reverberation. The feedback delay network simulates the multiple reflections and diffusion processes of sound waves in a closed space through a set of mutually coupled delay lines and feedback matrices. The feedback delay network mainly includes delay lines, feedback matrices, and attenuation filters. Specifically, the reverberation unit 151 can distribute binaural signals to multiple delay lines. After attenuation and filtering, the output of each delay line is redistributed to other delay lines through the feedback matrix. The signals of all delay lines are mixed and output to form a reverberation tail.

[0056] By providing the reverberation unit 151, a reverberation effect with spatial distribution characteristics can be generated, so that the wearer can feel the diffuse reflection of sound from different directions or environments.

[0057] In some embodiments, the sound enhancement unit 152 can perform automatic gain control on the time domain signal and adjust the fundamental frequency using a real-time PSOLA (Pitch Synchronous Addition) algorithm. Specifically, automatic gain control may include signal detection, gain calculation, and smoothing. Signal detection may refer to calculating the amplitude characteristics of the input signal, and gain calculation may refer to dynamically adjusting the gain coefficient according to the difference between the target level and the current level, and preventing gain mutations through attack time and release time. Furthermore, the target level may be the RMS energy of the desired output, the attack time may refer to the response speed of the gain reduction, and the release time may refer to the response speed of the gain increase accordingly.

[0058] Automatic gain control can provide a stable input for the PSOLA algorithm, and the PSOLA algorithm can perform pitch adjustment. Therefore, by setting the sound enhancement unit 152, it can provide dual guarantees of volume stability and pitch flexibility, thereby improving the wearer's experience.

[0059] In some embodiments, the ear return mixer 160 can be used to mix multiple input channels, and the input channels can include microphone arrays, musical instruments, line inputs, and digital signals, etc. The ear return mixer 160 can use an FPGA (field programmable gate array) or DSP (digital signal processor) chip to process multiple signals in real time.

[0060] The ear return signal ratio of the ear return mixer 160 can be controlled by user-adjustable parameters. The ear return mixer 160 can be used for low-latency processing, such as A / D conversion delay and signal transmission delay, so that the full link delay is less than 4 milliseconds.

[0061] In some embodiments, the noise reduction enhancement module 170 may include an adaptive beamforming unit 171 , a howling suppression unit 172 , and an embedding unit 173 .

[0062] In some embodiments, the adaptive beamforming unit 171 may refer to a minimum variance distortionless response (MVDR) beamformer constructed according to a microphone array, and the adaptive beamforming unit 171 may be used to suppress noise in a direction other than human voice.

[0063] In some embodiments, the howling suppression unit 172 may include a notch filter group, which can be used to dynamically track the feedback frequency. Specifically, the notch filter group may include multiple independent notch filters, each of which can be used to suppress signals at specific frequencies. Furthermore, the howling suppression unit 172 can detect and lock the changing interference frequency in real time based on the dynamic tracking feedback frequency technology, automatically adjust the center frequency and bandwidth of the notch filter, and achieve precise suppression. Among them, the interference frequency may include acoustic feedback howling, power line noise, or radio frequency interference.

[0064] In some embodiments, the howling suppression unit 172 can pre-configure the center frequency and bandwidth of each notch filter, and can monitor the howling frequency in real time and suppress it. By setting the howling suppression unit 172, the howling between the microphone array and the speaker can be better eliminated, so that the panoramic sound recording karaoke headset 1 can take into account both real-time and accuracy, and improve the wearer's karaoke experience.

[0065] In some embodiments, the embedding unit 173 may connect the time domain signal and the three-dimensional spatial information where the time domain signal is located in parallel and encode them into the audio stream, wherein the three-dimensional spatial information includes the azimuth angle and the elevation angle of the sound source.

[0066] The panoramic sound recording karaoke earphone 1 of the present application can collect time domain signals in various directions through a microphone array, locate the sound source, and realize the recording of human voice and panoramic sound. The panoramic sound recording karaoke earphone 1 of the present application can switch the usage mode according to different application scenarios, for example, it can include band mode and lead singer mode.

[0067] Specifically, in the scenarios of multi-person chorus and band performances, the microphone array of the present application can synchronously capture the azimuth information of each sound source, completely retain the phase relationship of each sound source in three-dimensional space, preserve the presence of singing and background music accompaniment, improve the accuracy of the lead singer's sound image positioning, and reduce the spatial distribution error rate of the background music accompaniment, thereby presenting an immersive sound field with a clear sense of spatial hierarchy.

[0068] In a solo singing scenario, the microphone array of the present application can distinguish the direction of each sound source, lock the position of the lead singer, retain the singing sound, denoise the ambient sound, and improve the speech clarity index, thereby ensuring a clear vocal recording effect in a noisy environment.

[0069] In the present application, a left microphone 40 is arranged on the left headphone cable 20, a right microphone 50 is arranged on the right headphone cable 30, and a middle microphone 60 is arranged at the connection between the control box 10 and the headphone cable to form a microphone array to collect multiple time domain signals in different directions, and the sound source is located based on these time domain signals, and finally the audio signal rendered by both ears is output. In this case, the microphone array can collect the singing voice of the wearer to a large extent, as well as the time domain signals in the front, back, left, right, top, bottom and oblique directions in the collection environment, so as to realize the surround of panoramic sound; compared with the method of collecting by using a single microphone, the microphone array of the present application collects time domain signals by each microphone cooperating with each other, and can accurately focus on the sound source in any direction, that is, it is not limited to the specific direction and distance of the sound source, and can reduce the error and noise caused by the single microphone collection, so that the sound direction and distance sense perceived by the wearer are the same as the reality, which improves the wearer's immersive experience in the karaoke scene.

[0070] The above is only the preferred embodiment of the present application. Those skilled in the art will appreciate that various changes or equivalent substitutions may be made to these features and embodiments without departing from the spirit and scope of the present application. In addition, under the guidance of the present application, these features and embodiments may be modified to adapt to specific circumstances and materials without departing from the spirit and scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of the present application belong to the protection scope of the present application.

Claims

1. A panoramic sound recording karaoke headset, characterized in that: Includes control box, microphone array, earphone head and earphone cable; The headphone cable includes a left headphone cable and a right headphone cable; the microphone array includes a left microphone, a right microphone and a center microphone; the headphone head includes a speaker and an acoustic cavity for encapsulating the speaker; One end of the left earphone cable and one end of the right earphone cable are both connected to the earphone head, the other end of the left earphone cable and the other end of the right earphone cable are both connected to one end of the control box, the left microphone is arranged on the left earphone cable, and the right microphone is arranged on the right earphone cable; The middle microphone is disposed at one end of the control box, the other end of the left headphone cable and the other end of the right headphone cable are symmetrically disposed on both sides of the middle microphone, and the length of the headphone cable between the left microphone and the middle microphone is equal to the length of the headphone cable between the right microphone and the middle microphone; The control box is configured to perform sound source localization according to a plurality of time domain signals collected by the microphone array, output binaurally rendered audio signals, and transmit the audio signals to the headphone heads via the headphone cable.

2. The panoramic sound recording karaoke headset according to claim 1, characterized in that: The control box includes a shell, and the shell is provided with an acquisition module, an extraction module, a calculation module and a processing module; the acquisition module is configured to acquire the time domain signal collected by the microphone array; the extraction module is configured to extract the spatial acoustic feature information between every two microphones of the microphone array and each microphone itself according to the time domain signal acquired by the acquisition module; The calculation module calculates the azimuth and elevation of the sound source according to the spatial acoustic feature information extracted by the extraction module, wherein the spatial acoustic feature information includes time delay information, phase difference information and energy value; the processing module outputs the audio signal according to the azimuth and elevation of the sound source obtained by the calculation module.

3. The panoramic sound recording karaoke headset according to claim 2, characterized in that: The extraction module includes a delay unit, a phase unit and a volume unit; the delay unit can be used to estimate the delay between every two microphones to obtain the delay information; the phase unit obtains the phase difference information between every two microphones after performing a fast Fourier transform on each time domain signal; the volume unit is used to obtain the energy value of the RMS energy of each time domain signal.

4. The panoramic sound recording karaoke headset according to claim 2, characterized in that: The processing module includes a beam forming unit and a HRTF mapping unit; the beam forming unit is used to generate beams pointing in different directions to determine the direction of the sound source; the HRTF mapping unit is used to convolve the beam signal pointing to the sound source direction generated by the beam forming unit with the left and right ear HRTFs respectively, simulate the head filtering effect, and output the audio signal.

5. The panoramic sound recording karaoke headset according to claim 1, characterized in that: The length of the headphone cable between the left microphone and the corresponding headphone head is greater than the length of the headphone cable between the left microphone and the middle microphone, and the length of the headphone cable between the right microphone and the corresponding headphone head is greater than the length of the headphone cable between the right microphone and the middle microphone.

6. The panoramic sound recording karaoke headset according to claim 1, characterized in that: The microphone array includes at least two movable microphones movable along the headphone cable, one movable microphone is arranged on the headphone cable between the left microphone and the corresponding headphone head, and the other movable microphone is arranged on the headphone cable between the right microphone and the corresponding headphone head.

7. The panoramic sound recording karaoke headset according to claim 1, characterized in that: A volume adjustment button is provided on the surface of the control box, and the volume adjustment button is used to adjust the volume of the audio signal output to the earphone head.

8. The panoramic sound recording karaoke headset according to claim 1, characterized in that: Both side walls of the control box are provided with microphone switch buttons, one of the microphone switch buttons is used to control whether the left microphone collects the time domain signal, and the other microphone switch button is used to control whether the right microphone collects the time domain signal.

9. The panoramic sound recording karaoke headset according to claim 1, characterized in that: Both side walls of the control box are provided with microphone adjustment buttons, one of which is used to amplify or reduce the time domain signal collected by the left microphone, and the other of which is used to amplify or reduce the time domain signal collected by the right microphone.

10. The panoramic sound recording karaoke headset according to claim 1, characterized in that: The panoramic sound recording karaoke earphones include a cable tie, which is arranged between the control box and the left microphone and the right microphone. The left headphone cable and the right headphone cable both pass through the interior of the cable tie in parallel, and the cable tie bundles the left headphone cable and the right headphone cable.

Citation Information

Patent Citations

  • Method and device for realizing karaoke function through earphone, and earphone

    CN106210983A

  • Earphone and audio denoising method thereof

    CN108184182A

  • Wireless panoramic sound mixing earphone

    CN111246331A

  • Audio processing method, device and earphone

    CN119789005A