Ambisonic karaoke headset

CN224555756UActive Publication Date: 2026-07-24SHENZHEN SEVENSTAR TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Utility models(China)
Current Assignee / Owner
SHENZHEN SEVENSTAR TECHNOLOGY CO LTD
Filing Date
2025-07-23
Publication Date
2026-07-24

Smart Images

  • Figure CN224555756U_ABST
    Figure CN224555756U_ABST
Patent Text Reader

Abstract

The application discloses a panoramic sound recording karaoke headset, which comprises a control box, a microphone array, a headset and a headset wire; one end of the left headset wire and one end of the right headset wire are connected with the headset; the other end of the left headset wire and the other end of the right headset wire are connected with one end of the control box; a left microphone is arranged on the left headset wire, and a right microphone is arranged on the right headset wire; one end of the control box is provided with a middle microphone; the other end of the left headset wire and the other end of the right headset wire are symmetrically arranged on the two sides of the middle microphone; the length of the headset wire between the left microphone and the middle microphone is equal to the length of the headset wire between the right microphone and the middle microphone; the control box is configured to output an audio signal rendered by binaural rendering according to multiple time domain signals collected by the microphone array, and the audio signal is delivered to the headset via the headset wire. The application can collect the singing sound of the wearer and the time domain signals in each direction in the environment to the greatest extent, and realize the surround of the panoramic sound.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of headphone technology, and more particularly to a panoramic sound recording karaoke headphone. Background Technology

[0002] Currently, wired microphone headsets remain the mainstream choice for mobile karaoke. When outdoors or traveling, users prefer a lightweight solution of "one device and one headset." At the same time, wired microphone headsets have high hardware penetration and strong versatility, and are compatible with almost all karaoke apps without the need for additional driver adjustments.

[0003] However, existing wired microphone headphones typically use a single microphone to collect sound, which often results in a single sound pickup direction, insufficient sound field coverage, and errors in the identification and localization of sound sources. Utility Model Content

[0004] The purpose of this application is to provide a panoramic sound recording karaoke headset to solve the technical problem of poor sound pickup in existing karaoke headsets. The various technical effects of the preferred technical solutions provided in this application are detailed below.

[0005] To achieve the above objectives, this application provides the following technical solutions: This application provides a panoramic sound recording karaoke headset, including a control box, a microphone array, headset heads, and headset cables; the headset cables include a left headset cable and a right headset cable; the microphone array includes a left microphone, a right microphone, and a center microphone; the headset heads include a speaker and an acoustic cavity for encapsulating the speaker; one end of the left headset cable and one end of the right headset cable are both connected to the headset heads, and the other ends of the left headset cable and the other ends of the right headset cable are both connected to one end of the control box; the left microphone is disposed on the left headset cable, and the right microphone is disposed on the right headset cable; one end of the control box is provided with the center microphone, and the other ends of the left and right headset cables are symmetrically disposed on both sides of the center microphone; the length of the headset cable between the left microphone and the center microphone is equal to the length of the headset cable between the right microphone and the center microphone; the control box is configured to perform sound source localization based on multiple time-domain signals collected by the microphone array, output a binaural-rendered audio signal, and transmit it to the headset heads via the headset cables.

[0006] In some embodiments, the control box includes a housing, which is provided with an acquisition module, an extraction module, a calculation module, and a processing module; the acquisition module is configured to acquire the time-domain signal collected by the microphone array; the extraction module is configured to extract spatial acoustic feature information between every two microphones of the microphone array and for each microphone itself based on the time-domain signal acquired by the acquisition module; the calculation module calculates the azimuth and elevation angles of the sound source based on the spatial acoustic feature information extracted by the extraction module, wherein the spatial acoustic feature information includes time delay information, phase difference information, and energy value; the processing module outputs the audio signal based on the azimuth and elevation angles of the sound source acquired by the calculation module.

[0007] In some embodiments, the extraction module includes a time delay unit, a phase unit, and a volume unit; the time delay unit can be used to estimate the time delay for each pair of microphones and obtain the time delay information; the phase unit performs a fast Fourier transform on each time domain signal and obtains the phase difference information between each pair of microphones; the volume unit is used to obtain the energy value of the RMS energy of each time domain signal.

[0008] In some embodiments, the processing module includes a beamforming unit and an HRTF mapping unit; the beamforming unit is used to generate beams pointing in different directions to determine the location of the sound source; the HRTF mapping unit is used to convolve the beam signals generated by the beamforming unit pointing in the location of the sound source with the HRTFs of the left and right ears respectively to simulate the head filtering effect and output the audio signal.

[0009] In some embodiments, the length of the headphone cable between the left microphone and the corresponding earphone head is greater than the length of the headphone cable between the left microphone and the middle microphone, and the length of the headphone cable between the right microphone and the corresponding earphone head is greater than the length of the headphone cable between the right microphone and the middle microphone.

[0010] In some embodiments, the microphone array includes at least two movable microphones that can move along the headphone cable, one of which is disposed on the headphone cable between the left microphone and the corresponding earpiece, and the other of which is disposed on the headphone cable between the right microphone and the corresponding earpiece.

[0011] In some embodiments, the surface of the control box is provided with a volume adjustment button, which is used to adjust the volume of the audio signal output to the earphone head.

[0012] In some embodiments, microphone switch buttons are provided on both side walls of the control box. One microphone switch button is used to control whether the left microphone collects the time domain signal, and the other microphone switch button is used to control whether the right microphone collects the time domain signal.

[0013] In some embodiments, microphone adjustment buttons are provided on both side walls of the control box. One microphone adjustment button is used to amplify or reduce the time domain signal acquired by the left microphone, and the other microphone adjustment button is used to amplify or reduce the time domain signal acquired by the right microphone.

[0014] In some embodiments, the panoramic sound recording karaoke headphones include a cable bundler disposed between the control box and the left and right microphones. The left and right earphone cables pass parallel through the interior of the cable bundler, which gathers the left and right earphone cables together.

[0015] Implementing one of the technical solutions described above in this application has the following advantages or beneficial effects: In this application, a microphone array is constructed by setting a left microphone on the left earphone cable, a right microphone on the right earphone cable, and a central microphone at the connection between the control box and the earphone cable to collect multiple time-domain signals from different directions. Based on these time-domain signals, sound source localization is performed, and finally, a binaural-rendered audio signal is output. In this case, the microphone array can collect the wearer's singing voice to a maximum extent, as well as time-domain signals from the front, back, left, right, above, below, and diagonal directions in the environment, achieving panoramic surround sound. Compared to using a single microphone for acquisition, the microphone array in this application collects time-domain signals through the cooperation of each microphone, enabling precise focusing on sound sources in any direction, i.e., it is not limited by the specific direction and distance of the sound source. Simultaneously, it reduces the errors and noise caused by single-microphone acquisition, making the wearer's perceived sound direction and distance indistinguishable from reality, thus improving the wearer's immersive experience in karaoke scenarios. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 This is a schematic diagram of the structure of the panoramic sound recording karaoke headphones according to an embodiment of this application; Figure 2This is a structural block diagram of the control box according to an embodiment of this application.

[0017] In the diagram: 1. Immersive sound recording and karaoke headphones; 10. Control box; 20. Left headphone cable; 30. Right headphone cable; 40. Left microphone; 50. Right microphone; 60. Center microphone; 70. Headphone head; 100. Acquisition module; 110. Extraction module; 120. Calculation module; 130. Processing module; 111. Delay unit; 112. Phase unit; 113. Volume unit; 131. Beamforming unit; 132. HRTF mapping unit; 140. Preprocessing module; 141. Frame segmentation and windowing unit; 142. Filtering unit; 143. Amplification unit; 144. ADC unit; 150. Audio effect processing module; 160. Ear monitor mixer; 170. Noise reduction and enhancement module; 151. Reverb unit; 152. Audio effect enhancement unit; 171. Adaptive beamforming unit; 172. Feedback suppression unit; 173. Embedding unit. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, various exemplary embodiments described below will be referenced to the accompanying drawings, which form part of the exemplary embodiments and depict various exemplary embodiments that may be adopted to implement this application. Unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. It should be understood that they are merely examples of processes, methods, and apparatuses consistent with some aspects of this application disclosed as detailed in the appended claims, and other embodiments may be used, or structural and functional modifications may be made to the embodiments listed herein without departing from the scope and spirit of this application.

[0019] In the description of this application, it should be understood that the terms "center," "longitudinal," "lateral," etc., indicate the orientation or positional relationship based on the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the referred element must have a specific orientation, or be constructed and operated in a specific orientation. The terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. The term "multiple" means two or more. The terms "connected" and "linked" should be interpreted broadly, for example, they can be fixed connections, detachable connections, integral connections, mechanical connections, electrical connections, communication connections, direct connections, indirect connections through an intermediate medium, and can be the internal connection of two elements or the interaction relationship between two elements. The term "and / or" includes any and all combinations of one or more of the related listed items. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.

[0020] To illustrate the technical solutions described in this application, specific embodiments are provided below, showing only the parts related to the embodiments of this application.

[0021] like Figure 1 and Figure 2 As shown, this application provides a panoramic sound recording karaoke headset 1, including a control box 10, a microphone array, a headset head 70, and a headset cable; In some embodiments, the headphone cable may include a left headphone cable 20 and a right headphone cable 30; the microphone array may include a left microphone 40, a right microphone 50 and a center microphone 60; and the headphone head 70 may include a speaker and an acoustic cavity for encapsulating the speaker.

[0022] In some embodiments, one end of the left earphone cable 20 and one end of the right earphone cable 30 are both connected to the earphone head 70, and the other end of the left earphone cable 20 and the other end of the right earphone cable 30 are both connected to one end of the control box 10. The left microphone 40 is disposed on the left earphone cable 20, and the right microphone 50 is disposed on the right earphone cable 30.

[0023] In some embodiments, one end of the control box 10 may be provided with a middle microphone 60, and the other end of the left earphone cable 20 and the other end of the right earphone cable 30 are symmetrically arranged on both sides of the middle microphone 60. The length of the earphone cable between the left microphone 40 and the middle microphone 60 is equal to the length of the earphone cable between the right microphone 50 and the middle microphone 60.

[0024] The following uses a Cartesian coordinate system based on the normal usage posture of the panoramic sound recording karaoke headset 1. The positions of the left microphone 40, right microphone 50, and center microphone 60 of the panoramic sound recording karaoke headset 1 can be represented by the following formula: L(−d,0,0), R(d,0,0), C(0,0,h), Where L(−d, 0, 0) is the coordinate of the left microphone 40, R(d, 0, 0) is the coordinate of the right microphone 50, C(0, 0, h) is the coordinate of the middle microphone 60, the origin is the center of the wearer's head, d is the left-right distance, and h is the vertical height.

[0025] In some embodiments, the control box 10 may include a housing, such as Figure 2 As shown, the housing may include an acquisition module 100, an extraction module 110, a calculation module 120, and a processing module 130. Specifically, the housing may contain a circuit board or controller that houses the acquisition module 100, the extraction module 110, the calculation module 120, and the processing module 130.

[0026] In some embodiments, the acquisition module 100 can be configured to acquire the time-domain signal collected by the microphone array; the extraction module 110 can be configured to extract the spatial acoustic feature information between every two microphones of the microphone array and the spatial acoustic feature information of each microphone itself based on the time-domain signal acquired by the acquisition module 100; the calculation module 120 can calculate the azimuth and elevation angles of the sound source based on the spatial acoustic feature information extracted by the extraction module 110, wherein the spatial acoustic feature information may include time delay information, phase difference information and energy value; the processing module 130 can output an audio signal based on the azimuth and elevation angles of the sound source acquired by the calculation module 120.

[0027] In some embodiments, the acquisition module 100 may be configured to acquire a first time-domain signal acquired by the left microphone 40, a second time-domain signal acquired by the right microphone 50, and a third time-domain signal acquired by the middle microphone 60.

[0028] In some embodiments, the acquisition module 100 can simultaneously acquire a first time-domain signal, a second time-domain signal, and a third time-domain signal, for a total of three channels. This avoids clock deviations in the acquired signals and maintains better time alignment.

[0029] In some embodiments, the housing may be provided with a preprocessing module 140, which may include a frame-splitting windowing unit 141 and a filtering unit 142. The frame-splitting windowing unit 141 may be used to perform frame-splitting windowing processing on the first time domain signal, the second time domain signal and the third time domain signal, and the filtering unit 142 may be used to perform noise reduction processing on the first time domain signal, the second time domain signal and the third time domain signal.

[0030] Specifically, since the time-domain signal has time-varying properties, the framing and windowing unit 141 can perform framing processing on the first time-domain signal, the second time-domain signal, and the third time-domain signal according to the preset frame length and the preset frame shift. The preset frame length can range from 15 milliseconds to 40 milliseconds per frame. By setting the preset frame length, the problems of insufficient frequency resolution and decreased time resolution can be effectively solved. At the same time, by setting the preset frame shift, the loss of frame edge information can be reduced, enabling smooth transitions between frames.

[0031] In some embodiments, the framing windowing unit 141 can perform windowing processing on the framed time-domain signal. That is, the framing windowing unit 141 can perform windowing processing on the first time-domain signal, the second time-domain signal, and the third time-domain signal after framed processing. Specifically, the framing windowing unit 141 can window each frame of the time-domain signal frame by frame according to a preset window function, which can be one of a Hamming window, a Kaiser window, or a Blackman window. Thus, the framing windowing unit 141 can significantly reduce spectral leakage, providing convenience for subsequent processing.

[0032] In some embodiments, the filtering unit 142 can process the time-domain signal according to spectral subtraction or adaptive filtering to reduce background noise. Specifically, spectral subtraction can refer to subtracting the noise component from the spectrum of the noisy signal to retain the target signal, while adaptive filtering can refer to suppressing noise by dynamically adjusting the filter coefficients to minimize the error.

[0033] In some embodiments, the preprocessing module 140 may further include an amplification unit 143 and an ADC unit 144. The amplification unit 143 amplifies the time-domain signal, and the ADC unit 144 performs analog-to-digital conversion on the time-domain signal. This amplifies and improves the level of the time-domain signal, making it sufficient to drive subsequent circuits. Furthermore, after converting the analog signal into a digital signal, various processing techniques can be applied to the time-domain signal.

[0034] In some embodiments, the extraction module 110 may include a delay unit 111, a phase unit 112, and a volume unit 113.

[0035] In some embodiments, the delay unit 111 can be used to estimate the delay for every two microphones and obtain delay information. That is, the delay unit 111 can obtain first delay information between the left microphone 40 and the right microphone 50, second delay information between the left microphone 40 and the middle microphone 60, and third delay information between the right microphone 50 and the middle microphone 60. Specifically, the delay unit 111 can calculate the cross-correlation function of the three pairs of microphones, such as generalized cross-correlation, and obtain the first delay information, second delay information, and third delay information based on the peak position.

[0036] In some embodiments, the first time delay information can be obtained by the following formula: .

[0037] In some embodiments, the phase unit 112 can perform a fast Fourier transform on each time-domain signal to obtain the phase difference information between each pair of microphones. Specifically, the phase unit 112 can obtain the first phase difference information between the left microphone 40 and the right microphone 50, the second phase difference information between the left microphone 40 and the middle microphone 60, and the third phase difference information between the right microphone 50 and the middle microphone 60.

[0038] In some embodiments, the volume unit 113 can be used to acquire the RMS (Root Mean Square) energy of each time-domain signal. Specifically, the RMS energy can be used to describe the average power of the time-domain signal.

[0039] In some embodiments, assuming a far-field plane wave, the calculation module 120 can construct a system of equations based on spatial acoustic feature information and the coordinates of each microphone to solve for the direction of the sound source. Further, the calculation module 120 can use the least squares method to obtain the azimuth and elevation angles of the sound source. In this case, the direction vector can be u = (sinθcosϕ, sinθsinϕ, cosθ), where θ is the elevation angle and ϕ is the azimuth angle.

[0040] In some embodiments, the processing module 130 may include a beamforming unit 131 and an HRTF (Head-Related Transfer Function) mapping unit 132.

[0041] In some embodiments, the beamforming unit 131 can be used to generate beams pointing in different directions to determine the location of the sound source. Specifically, the beamforming unit 131 can adjust the first time delay information, the second time delay information, and the third time delay information according to the azimuth and elevation angles of the sound source, so that the sound source direction signals are phase-aligned when superimposed. This compensates for the propagation differences of different microphones, achieving in-phase superposition. After compensating for the time delay, the beamforming unit 131 can weight the various time-domain signals to form a beam signal.

[0042] In some embodiments, the beamforming unit 131 can compare the energy values ​​of the generated beams to determine the direction corresponding to the maximum energy value as the azimuth of the sound source. The energy value of each beam can be expressed by the following formula: E(θ, ϕ)=∫∣B(θ, ϕ)∣2dtE(θ, ϕ)=∫∣B(θ, ϕ)∣2dt, Where B(θ, ϕ) is the beam.

[0043] In some embodiments, the HRTF mapping unit 132 can match the azimuth and elevation angles corresponding to the sound source location with the HRTF database to obtain the left and right ear filter coefficients. The HRTF mapping unit 132 can be used to convolve the beam signal pointing to the sound source location generated by the beamforming unit 131 with the left and right ear HRTFs respectively to simulate the head filtering effect and output the audio signal.

[0044] In some embodiments, the length of the headphone cable between the left microphone 40 and the corresponding earphone head 70 can be greater than the length of the headphone cable between the left microphone 40 and the middle microphone 60, and the length of the headphone cable between the right microphone 50 and the corresponding earphone head 70 can be greater than the length of the headphone cable between the right microphone 50 and the middle microphone 60. This allows the left microphone 40 and the right microphone 50 to be closer to the wearer's mouth, resulting in clearer voice capture.

[0045] In some embodiments, the microphone array may include at least two movable microphones that can move along the headphone cable. One movable microphone may be positioned on the headphone cable between the left microphone 40 and the corresponding earphone head 70, and the other movable microphone may be positioned on the headphone cable between the right microphone 50 and the corresponding earphone head 70. Specifically, the left microphone 40 may be fixedly mounted on the left headphone cable 20, and the position of the movable microphone on the left headphone cable 20 may be adjusted as needed. Correspondingly, the right microphone 50 may be fixedly mounted on the right headphone cable 30, and the position of the other movable microphone on the right headphone cable 30 may be adjusted as needed. In this configuration, the movable microphones can assist the corresponding microphones in acquiring time-domain signals, thereby enhancing the sound pickup effect.

[0046] In some embodiments, the surface of the control box 10 may be provided with a volume adjustment button, which can be used to adjust the volume of the audio signal output to the headphone head 70.

[0047] In some embodiments, the panoramic sound recording karaoke headset 1 may include a connector for connecting to external devices, such as mobile phones, tablets, etc., so that the wearer can adjust various parameters of the panoramic sound recording karaoke headset 1 through the external device.

[0048] In some embodiments, the connector can communicate with the control box 10 via wired or wireless means. For example, the connector can be connected to the other end of the control box 10 via a data cable, or it can communicate via wireless transmission methods such as Bluetooth or WiFi.

[0049] In some embodiments, both side walls of the control box 10 may be provided with microphone switch buttons. One microphone switch button can be used to control whether the left microphone 40 collects time-domain signals, and the other microphone switch button can be used to control whether the right microphone 50 collects time-domain signals. Specifically, the microphone switch button near the left earphone cable 20 can be used to control the left microphone 40 to turn on or off its pickup function, and the microphone switch button near the right earphone cable 30 can be used to control the right microphone 50 to turn on or off its pickup function. This allows the wearer to easily adjust the device's pickup based on the judgment of the ambient sound source.

[0050] In some embodiments, both side walls of the control box 10 may be provided with microphone adjustment buttons. One microphone adjustment button is used to amplify or reduce the time-domain signal acquired by the left microphone 40, and the other microphone adjustment button is used to amplify or reduce the time-domain signal acquired by the right microphone 50. Specifically, the microphone adjustment button near the left earphone cable 20 can be used to amplify or reduce the first time-domain signal, and the microphone adjustment button near the right earphone cable 30 can be used to amplify or reduce the second time-domain signal. Thus, the sound pickup effect of the left microphone 40 and the right microphone 50 can be adjusted in real time.

[0051] In some embodiments, the panoramic sound recording karaoke headset 1 may include a cable bundler, which may be disposed between the control box 10 and the left microphone 40 and the right microphone 50. The left earphone cable 20 and the right earphone cable 30 may both pass parallel to each other through the interior of the cable bundler, and the cable bundler may gather the left earphone cable 20 and the right earphone cable 30 together.

[0052] In some embodiments, the housing includes an audio processing module 150, an in-ear mixer 160, and a noise reduction enhancement module 170.

[0053] In some embodiments, the sound processing module 150 may include a reverberation unit 151 and a sound enhancement unit 152.

[0054] In some embodiments, the reverberation unit 151 may use a feedback delay network (FDN) to generate multi-channel reverberation, wherein the reverberation time, pre-delay and high-frequency attenuation factor can be adjusted as needed. The reverberation time can be controlled by adjusting the attenuation coefficient, the pre-delay can be used to simulate space size, and the high-frequency attenuation factor can be used to simulate air absorption effect.

[0055] Feedback delay networks are an advanced digital signal processing technique for generating artificial reverberation. They simulate the multiple reflections and diffusions of sound waves within a confined space using a set of coupled delay lines and a feedback matrix. A feedback delay network primarily consists of delay lines, a feedback matrix, and an attenuation filter. Specifically, the reverberation unit 151 can distribute binaural signals to multiple delay lines. The output of each delay line, after attenuation and filtering, is redistributed to other delay lines via the feedback matrix. The signals from all delay lines are then mixed and output to form the reverberation tail.

[0056] By setting the reverberation unit 151, a reverberation effect with spatial distribution characteristics can be generated, allowing the wearer to feel the diffusion and reflection of sound from different directions or environments.

[0057] In some embodiments, the audio enhancement unit 152 can perform automatic gain control on the time-domain signal and adjust the fundamental frequency using a real-time PSOLA (Pitch Synchronous Alignment) algorithm. Specifically, automatic gain control can include signal detection, gain calculation, and smoothing processing. Signal detection can refer to calculating the amplitude characteristics of the input signal, and gain calculation can refer to dynamically adjusting the gain coefficient based on the difference between the target level and the current level, preventing sudden gain changes through attack and release times. Further, the target level can be the desired output RMS energy, the attack time can be the response speed of gain reduction, and the corresponding release time can be the response speed of gain increase.

[0058] Automatic gain control provides a stable input for the PSOLA algorithm, which can adjust pitch. Thus, by setting up the sound enhancement unit 152, both volume stability and pitch flexibility can be guaranteed, improving the wearer's experience.

[0059] In some embodiments, the in-ear mixer 160 can be used for mixing multiple input channels, which may include microphone arrays, musical instruments, line inputs, and digital signals. The in-ear mixer 160 can use an FPGA (Field Programmable Gate Array) or DSP (Digital Signal Processor) chip to process multiple signals in real time.

[0060] The ratio of the in-ear signal in the in-ear mixer 160 can be controlled by user-adjustable parameters. The in-ear mixer 160 can be used for low-latency processing, such as A / D conversion delay and signal transmission delay, so that the end-to-end latency is less than 4 milliseconds.

[0061] In some embodiments, the noise reduction enhancement module 170 may include an adaptive beamforming unit 171, a howling suppression unit 172, and an embedding unit 173.

[0062] In some embodiments, the adaptive beamforming unit 171 may refer to a minimum variance distortionless response (MVDR) beamformer constructed according to the microphone array, and the adaptive beamforming unit 171 may be used to suppress noise in non-human voice directions.

[0063] In some embodiments, the howling suppression unit 172 may include a notch filter bank for dynamically tracking the feedback frequency. Specifically, the notch filter bank may include multiple independent notch filters, each of which can be used to suppress signals at a specific frequency. Further, the howling suppression unit 172 may, based on dynamic feedback frequency tracking technology, detect and lock onto changing interference frequencies in real time, automatically adjusting the center frequency and bandwidth of the notch filters to achieve precise suppression. The interference frequencies may include acoustic feedback howling, power line noise, or radio frequency interference, etc.

[0064] In some embodiments, the feedback suppression unit 172 can be pre-configured with the center frequency and bandwidth of each notch filter, and can monitor and suppress feedback frequencies in real time. By setting the feedback suppression unit 172, feedback between the microphone array and the speaker can be effectively eliminated, enabling the panoramic sound recording karaoke headset 1 to balance real-time performance and accuracy, thus improving the wearer's karaoke experience.

[0065] In some embodiments, the embedding unit 173 can concatenate the time-domain signal and the three-dimensional spatial information in which the time-domain signal is located in parallel and encode them into the audio stream, wherein the three-dimensional spatial information includes the azimuth and pitch angles of the sound source.

[0066] The panoramic sound recording karaoke headset 1 of this application can acquire time-domain signals from various directions through a microphone array, locate the sound source, and achieve recording of human voices and panoramic sound. The panoramic sound recording karaoke headset 1 of this application can switch between usage modes according to different application scenarios, such as including a band mode and a lead singer mode.

[0067] Specifically, in scenarios involving multiple singers singing and band performances, the microphone array of this application can simultaneously capture the directional information of each sound source, fully preserve the phase relationship of each sound source in three-dimensional space, and maintain the on-site effect of singing and accompaniment. This can improve the positioning accuracy of the lead singer's sound image and reduce the spatial distribution error rate of the accompaniment, thereby presenting an immersive sound field with a clear sense of spatial hierarchy.

[0068] In solo singing scenarios, the microphone array of this application can distinguish the direction of each sound source, lock the position of the lead singer, preserve the singing sound, reduce ambient noise, and improve the speech intelligibility index, thereby ensuring a clear vocal recording effect even in noisy environments.

[0069] In this application, a microphone array is constructed by setting a left microphone 40 on the left earphone cable 20, a right microphone 50 on the right earphone cable 30, and a center microphone 60 at the connection between the control box 10 and the earphone cable. This array collects multiple time-domain signals from different directions, locates the sound source based on these time-domain signals, and finally outputs an audio signal rendered by both ears. In this case, the microphone array can collect the wearer's singing voice to a maximum extent, as well as time-domain signals from the front, back, left, right, above, below, and diagonal directions in the environment, achieving panoramic surround sound. Compared with the method of collecting signals using a single microphone, the microphone array of this application collects time-domain signals through the cooperation of each microphone, and can accurately focus on sound sources in any direction, that is, it is not limited by the specific direction and distance of the sound source. At the same time, it can reduce the errors and noise caused by single-microphone collection, so that the sound direction and distance perceived by the wearer are no different from reality, improving the wearer's immersive experience in karaoke scenarios.

[0070] The above description is merely a preferred embodiment of this application. Those skilled in the art will understand that various changes or equivalent substitutions can be made to these features and embodiments without departing from the spirit and scope of this application. Furthermore, under the teachings of this application, these features and embodiments can be modified to adapt to specific situations and materials without departing from the spirit and scope of this application. Therefore, this application is not limited to the specific embodiments disclosed herein, and all embodiments falling within the scope of the claims of this application are within the protection scope of this application.

Claims

1. A panoramic sound recording and karaoke headset, characterized in that, Includes control box, microphone array, earphone jack and earphone cable; The headphone cable includes a left headphone cable and a right headphone cable; the microphone array includes a left microphone, a right microphone, and a center microphone; the headphone head includes a speaker and an acoustic cavity for encapsulating the speaker; One end of the left earphone cable and one end of the right earphone cable are both connected to the earphone head, and the other end of the left earphone cable and the other end of the right earphone cable are both connected to one end of the control box. The left microphone is located on the left earphone cable and the right microphone is located on the right earphone cable. One end of the control box is provided with the middle microphone, and the other end of the left earphone cable and the other end of the right earphone cable are symmetrically arranged on both sides of the middle microphone. The length of the earphone cable between the left microphone and the middle microphone is equal to the length of the earphone cable between the right microphone and the middle microphone. The control box is configured to locate the sound source based on multiple time-domain signals collected by the microphone array, output a binaural-rendered audio signal, and transmit it to the earphone head via the earphone cable.

2. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, The length of the headphone cable between the left microphone and the corresponding earphone head is greater than the length of the headphone cable between the left microphone and the middle microphone, and the length of the headphone cable between the right microphone and the corresponding earphone head is greater than the length of the headphone cable between the right microphone and the middle microphone.

3. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, The microphone array includes at least two movable microphones that can move along the headphone cable. One movable microphone is located on the headphone cable between the left microphone and the corresponding earpiece, and the other movable microphone is located on the headphone cable between the right microphone and the corresponding earpiece.

4. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, The surface of the control box is provided with a volume adjustment button, which is used to adjust the volume of the audio signal output to the earphone.

5. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, Both side walls of the control box are equipped with microphone switch buttons. One microphone switch button is used to control whether the left microphone collects the time domain signal, and the other microphone switch button is used to control whether the right microphone collects the time domain signal.

6. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, Both side walls of the control box are equipped with microphone adjustment buttons. One of the microphone adjustment buttons is used to amplify or reduce the time domain signal acquired by the left microphone, and the other microphone adjustment button is used to amplify or reduce the time domain signal acquired by the right microphone.

7. The panoramic sound recording and karaoke headphones according to claim 1, characterized in that, The panoramic sound recording karaoke headphones include a cable bundler, which is disposed between the control box and the left and right microphones. The left and right earphone cables pass parallel through the interior of the cable bundler, which gathers the left and right earphone cables together.