Signal processing method and device, storage medium and electronic equipment
By intelligently dividing the target audio signal and matching the speaker, the audio signals in the karaoke speaker are divided into different frequency bands and allocated to suitable speakers, solving the technical problem that a single speaker cannot meet both karaoke and music playback, and achieving high-quality audio playback and speaker protection.
Patent Information
- Application Number
- CN202510456276.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-29
AI Technical Summary
The existing karaoke speakers cannot meet the large dynamic vocal output during karaoke and the high-frequency details restore during music playback using a single type of tweeter.
By dividing the target audio signal into multiple sub-signals, and inputting sub-signals of different frequency bands into different speakers for output according to signal type and speaker performance information, intelligent matching of speakers is achieved.
It realizes the balanced presentation of dynamic vocals and high-frequency music details, improves the audio quality of karaoke and music playback, protects the speaker from overload damage, and extends the system service life.
Smart Images

Figure CN120390185A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of audio processing, and in particular, to a signal processing method, apparatus, storage medium, and electronic device. Background Art
[0002] In the field of KTV speakers, in traditional designs, a single type of high-frequency speaker, such as a cloth-edge paper-cone high-frequency speaker or a silk-dome high-frequency speaker, is often used to meet the high-frequency sound quality reproduction requirements. However, when faced with two different application scenarios of KTV singing and music playing, the performance limitations of a single type of high-frequency speaker become obvious.
[0003] A cloth-edge paper-cone high-frequency speaker, due to its lightweight paper cone and enhanced cloth-edge design, can provide good anti-overload ability and power reserve, and is very suitable for vocal output in KTV singing scenarios. However, when playing music, its high-frequency upper limit is relatively low, and it cannot fully restore the ultra-high-frequency details in music, such as instrument overtones or sibilants. The tone color is relatively hard, and the detail expression ability is insufficient. On the other hand, a silk-dome high-frequency speaker, with its extremely light diaphragm and excellent transient response, can restore the high-frequency details in music signals with high resolution, and the tone color is soft and natural, making it an ideal choice for music playback. However, in a KTV singing scenario, due to its weak instantaneous high-power bearing capacity, it is easy to cause diaphragm deformation during large-dynamic vocal output, resulting in distortion or even burning out the speaker unit, which limits its use in a KTV environment. The design limitations of the above two types of high-frequency speakers make it difficult for KTV speakers to achieve a balance in the reproduction of vocals and music signals, and it is impossible to achieve the dual experience of warm and thick vocals and delicate and vivid music.
[0004] In view of the above problems, no effective solution has been proposed yet. Summary of the Invention
[0005] The present application provides a signal processing method, apparatus, storage medium, and electronic device, so as to at least solve the technical problem that when a single type of high-frequency speaker is used in an existing KTV speaker, it cannot simultaneously meet the large-dynamic vocal output during KTV singing and the high-frequency detail restoration during music playback.
[0006] According to one aspect of the present application, a signal processing method is provided, including: receiving a target audio signal, where the target audio signal is a music signal and / or a microphone signal; dividing the target audio signal into N sub-signals based on the signal type of the target audio signal, where N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different; when the target audio signal is a microphone signal, inputting S sub-signals among the N sub-signals that are greater than a first frequency division point into S speakers for output, and inputting N - S sub-signals into the same speaker for output, where S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal; when the target audio signal is a music signal, inputting T sub-signals among the N sub-signals that are greater than a second frequency division point into T speakers for output, and inputting N - T sub-signals into the same speaker for output, where T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from the T speakers, the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
[0007] Optionally, dividing the target audio signal into N sub-signals based on the signal type of the target audio signal includes: when it is detected that the target audio signal is a microphone signal, dividing the microphone signal into a first sub-signal and a second sub-signal according to a first frequency division point, where the first frequency division point is the frequency point determined based on the human voice frequency range and the performance information of the speaker for dividing the microphone signal, and the minimum frequency in the frequency band corresponding to the second sub-signal is greater than the first frequency division point.
[0008] Optionally, dividing the target audio signal into N sub-signals based on the signal type of the target audio signal includes: when it is detected that the target audio signal is a music signal, dividing the music signal into a third sub-signal and a fourth sub-signal according to a second frequency division point, where the second frequency division point is the frequency point determined based on the music signal frequency range and the performance information of the speaker for dividing the music signal, and the minimum frequency in the frequency band corresponding to the fourth sub-signal is greater than the second frequency division point.
[0009] Optionally, when the target audio signal is a microphone signal, S sub-signals among the N sub-signals that are greater than the first frequency division point are respectively input into S speakers for output, and the N - S sub-signals are input into the same speaker for output, including: when the target audio signal is a microphone signal, input the first sub-signal in the microphone signal into the first speaker for output, and input the second sub-signal in the microphone signal into the second speaker for output, where the first speaker and the second speaker are independent speakers, the frequency response range and power handling capacity of the first speaker in the frequency band corresponding to the first sub-signal are greater than those of the second speaker, and the anti-overload capacity and the ability to reproduce dynamic signals of the second speaker in the frequency band corresponding to the second sub-signal are greater than those of the first speaker, where the dynamic signal is used to represent the change value of the intensity or amplitude of the signal that is greater than the first preset threshold.
[0010] Optionally, when the target audio signal is a music signal, T sub-signals among the N sub-signals that are greater than the second frequency division point are respectively input into T speakers for output, and the N - T sub-signals are input into the same speaker for output, including: when the target audio signal is a music signal, input the third sub-signal in the music signal into the first speaker for output, and input the fourth sub-signal in the music signal into the third speaker for output, where the first speaker and the third speaker are independent speakers, the frequency response range and power handling capacity of the first speaker in the frequency band corresponding to the third sub-signal are greater than those of the third speaker, and the transient response ability and the target ability of the third speaker in the frequency band corresponding to the fourth sub-signal are greater than those of the first speaker, where the target ability is used to represent the ability to capture and reproduce the sound changes in the music signal.
[0011] Optionally, the method further includes: when the target audio signal is a microphone signal and a music signal, divide the target audio signal according to the first frequency division point and the second frequency division point to obtain the first sub-signal and the second sub-signal of the microphone signal, and the third sub-signal and the fourth sub-signal of the music signal; input the second sub-signal into the second speaker for output, and input the fourth sub-signal in the music signal into the third speaker for output; mix the first sub-signal in the microphone signal and the third sub-signal in the music signal, and input the target signal obtained after the mixing process into the first speaker for output.
[0012] Optionally, after dividing the target audio signal into N sub-signals based on the signal type of the target audio signal, the method further includes: based on the signal type and frequency range of the target audio signal, input the N sub-signals into the corresponding amplifiers of each sub-signal for amplification processing.
[0013] Optionally, before respectively inputting N sub-signals into the amplifiers corresponding to each sub-signal for amplification processing, the method further includes: detecting a target difference value of each sub-signal among the N sub-signals, where the target difference value is used to quantify the difference between the peak level and the average level in the sub-signal; when detecting that the target difference value of the i-th sub-signal is greater than or equal to a second preset threshold, reducing the amplification factor of the amplifier corresponding to the i-th sub-signal, where i is an integer less than or equal to N; when detecting that the target difference value of the i-th sub-signal is less than the second preset threshold, increasing the amplification factor of the amplifier corresponding to the i-th sub-signal.
[0014] Optionally, after receiving the target audio signal, it includes: when the target audio signal is a microphone signal, performing a first preprocessing operation on the microphone signal, where the first preprocessing operation at least includes a first process and a second process, where the first process is used to monitor the signal level change in the microphone signal, determine the instantaneous level based on the signal level change, and compress the signal whose instantaneous level exceeds a third preset threshold, and the second process is used to adjust the signal intensity of different frequency bands in the microphone signal; when the target audio signal is a music signal, performing a second preprocessing operation on the music signal, where the second preprocessing operation at least includes a third process and a fourth process, where the third process is used to align the phases of different frequency signals in the music signal, and the fourth process is used to correct the music signal according to the frequency response curve of the speaker.
[0015] According to another aspect of the present application, there is also provided a signal processing device, including: a receiving unit, configured to receive a target audio signal, where the target audio signal is a music signal and / or a microphone signal; a dividing unit, configured to divide the target audio signal into N sub-signals based on the signal type of the target audio signal, where N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different; a first output unit, configured to, when the target audio signal is a microphone signal, respectively input S sub-signals among the N sub-signals that are greater than a first frequency division point into S speakers for output, and input N - S sub-signals into the same speaker for output, where S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal; a second output unit, configured to, when the target audio signal is a music signal, respectively input T sub-signals among the N sub-signals that are greater than a second frequency division point into T speakers for output, and input N - T sub-signals into the same speaker for output, where T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from the T speakers, and the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
[0016] According to another aspect of the present application, a computer-readable storage medium is provided, which includes a stored executable program, wherein when the executable program runs, the device where the computer-readable storage medium is located is controlled to execute the above-mentioned signal processing method.
[0017] According to another aspect of the present application, an electronic device is also provided, comprising one or more processors and a memory, wherein the memory is used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned signal processing method.
[0018] In the present application, a target audio signal is first received, wherein the target audio signal is a music signal and / or a microphone signal, and then based on the signal type of the target audio signal, the target audio signal is divided into N sub-signals, wherein N is an integer greater than 1, and different sub-signals correspond to different frequency bands, and then when the target audio signal is a microphone signal, S sub-signals greater than the first crossover point in the N sub-signals are respectively input into S speakers for output, and NS sub-signals are input into the same speaker for output, wherein S is an integer greater than or equal to 1 and less than or equal to N, and the speakers corresponding to the NS sub-signals are different from the S speakers, and the first crossover point is the frequency point for dividing the microphone signal, and finally when the target audio signal is a music signal, T sub-signals greater than the second crossover point in the N sub-signals are respectively input into T speakers Output is performed, and NT sub-signals are input into the same loudspeaker for output, wherein T is an integer greater than or equal to 1 and less than or equal to N, the loudspeakers corresponding to the NT sub-signals are different from the T loudspeakers, and the T loudspeakers are different from the S loudspeakers. The second crossover point is the frequency point for dividing the music signal, that is, through intelligent crossover and loudspeaker matching, the high-frequency parts of the music signal and the microphone signal are input into different loudspeakers for output, and the low-frequency parts of the two are input into other loudspeakers for output, thereby achieving the purpose of optimizing the playback quality of signals from different audio sources, thereby achieving the technical effect of balancing the presentation of dynamic human voices and high-frequency music details, and thus solving the technical problem that when existing karaoke speakers use a single type of tweeter, they cannot simultaneously meet the requirements of large dynamic human voice output during karaoke and high-frequency detail restoration during music playback. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 is a flowchart of an optional signal processing method according to an embodiment of the present application;
[0021] Figure 2 It is a schematic diagram of an optional signal processing method according to an embodiment of the present application;
[0022] Figure 3 It is a schematic diagram of an optional signal processing device according to an embodiment of the present application. Detailed implementation manners
[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0025] It should be noted that the information collected in the present application (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for display, data for analysis, etc.) are information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application, etc. of the relevant data all comply with relevant laws, regulations and standards, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse. For example, an interface is set between the present system and relevant users or institutions to provide corresponding operation entrances for users to choose to agree or refuse the automated decision-making results; if the user chooses to refuse, the expert decision-making process will be entered.
[0026] According to an embodiment of the present application, a method embodiment of a signal processing method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0027] It should be noted that an intelligent speaker system can be the execution subject of the signal processing method of the embodiment of the present application. It can be understood that the signal processing method provided by the embodiment of the present application can also be executed by other systems or devices as the execution subject, and the embodiment of the present application does not make specific limitations in this regard.
[0028] Figure 1 is a flowchart of an optional signal processing method according to an embodiment of the present application, as Figure 1 shown, the method includes the following steps:
[0029] Step S101, receive a target audio signal.
[0030] In step S101, the target audio signal is a music signal and / or a microphone signal.
[0031] Optionally, the target audio signal refers to the audio input to be processed by the intelligent speaker system, and it can come from different signal sources, such as music signals emitted by music playback devices or human voice signals obtained by microphones. In the context of an intelligent speaker system, the target audio signal can also include both human voice and accompanying music at the same time.
[0032] Step S102, based on the signal type of the target audio signal, divide the target audio signal into N sub-signals.
[0033] In step S102, N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different.
[0034] Optionally, the signal types include music signals and microphone signals.
[0035] Optionally, the sub-signals are multiple parts divided from the original audio signal according to different frequencies, and each part represents a specific frequency range of the original signal. For example, a music signal can be divided into three sub-signals: bass, midrange, and treble, while a microphone signal may be divided into two sub-signals: midrange and treble; the frequency band refers to a specific set of frequency ranges, and each sub-signal corresponds to one or a group of frequency bands, such as the low-frequency band, the mid-frequency band, and the high-frequency band.
[0036] Optionally, the smart speaker system uses a frequency divider or digital signal processing technology (such as DSP (Digital Signal Processing)) to analyze the received audio signal and divide it into multiple sub-signals, each sub-signal covering a unique frequency band in the audio signal. This is done to enable the system to process audio of different frequencies in a targeted manner, thereby optimizing the audio quality of each frequency band and the usage efficiency of the speakers.
[0037] Step S103, when the target audio signal is a microphone signal, input the S sub-signals among the N sub-signals that are greater than the first frequency division point into S speakers for output respectively, and input the N - S sub-signals into the same speaker for output.
[0038] In step S103, S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal.
[0039] Optionally, the first frequency division point can be used to divide the high-frequency part and the low-frequency part in the microphone signal.
[0040] Optionally, when the smart speaker system detects that the target audio signal is a microphone signal, it inputs the high-frequency part in the microphone signal into the corresponding speaker for output, and inputs the remaining low-frequency part in the sub-signals into other speakers for output.
[0041] Optionally, the speakers corresponding to the high-frequency part and the low-frequency part in the microphone signal are different.
[0042] Step S104, when the target audio signal is a music signal, input the T sub-signals among the N sub-signals that are greater than the second frequency division point into T speakers for output respectively, and input the N - T sub-signals into the same speaker for output.
[0043] In step S104, T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from the T speakers, the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
[0044] Optionally, the second frequency division point can be used to divide the high-frequency part and the low-frequency part in the music signal.
[0045] Optionally, when the smart speaker system detects that the target audio signal is a music signal, it inputs the high-frequency part in the music signal into the corresponding speaker for output, and inputs the remaining low-frequency part in the sub-signals into other speakers for output. Among them, the speakers corresponding to the high-frequency part and the low-frequency part in the music signal are different.
[0046] Optionally, the loudspeaker corresponding to the high frequency part of the music signal and the loudspeaker corresponding to the high frequency part of the microphone signal are different loudspeakers.
[0047] Optionally, the smart speaker system sends the frequency-divided sub-signals to the most appropriate speakers for playback. For example, the high-frequency sub-signal of a music signal is sent to a silk dome tweeter, while the high-frequency sub-signal of a microphone signal is sent to a paper cone fabric tweeter. This allows the speakers to maximize their performance, ensuring the highest quality reproduction of every component of the audio, delivering the best experience in both karaoke and music playback scenarios.
[0048] It can be seen from the contents of steps S101 to S104 that in the present application, a target audio signal is first received, wherein the target audio signal is a music signal and / or a microphone signal, and then based on the signal type of the target audio signal, the target audio signal is divided into N sub-signals, wherein N is an integer greater than 1, and different sub-signals correspond to different frequency bands, and then when the target audio signal is a microphone signal, S sub-signals greater than the first crossover point in the N sub-signals are input into S speakers for output respectively, and NS sub-signals are input into the same speaker for output, wherein S is an integer greater than or equal to 1 and less than or equal to N, and the speakers corresponding to the NS sub-signals are different from the S speakers, and the first crossover point is the frequency point for dividing the microphone signal, and finally when the target audio signal is a music signal, T sub-signals greater than the second crossover point in the N sub-signals are input into the same speaker for output. The NT sub-signals are input into T speakers for output, and NT sub-signals are input into the same speaker for output, wherein T is an integer greater than or equal to 1 and less than or equal to N, the speakers corresponding to the NT sub-signals are different from the T speakers, and the T speakers are different from the S speakers. The second crossover point is the frequency point for dividing the music signal, that is, through intelligent crossover and speaker matching, the high-frequency parts of the music signal and the microphone signal are input into different speakers for output, and the low-frequency parts of the two are input into other speakers for output, thereby achieving the purpose of optimizing the playback quality of signals from different audio sources, thereby realizing the technical effect of giving consideration to both dynamic human voice and high-frequency music details, and thus solving the technical problem that the existing K song speakers cannot simultaneously meet the large dynamic human voice output during K song and the high-frequency detail restoration during music playback when using a single type of tweeter.
[0049] In an optional embodiment, when the smart speaker system detects that the target audio signal is a microphone signal, it divides the microphone signal into a first sub-signal and a second sub-signal according to a first frequency division point, wherein the first frequency division point is a frequency point for dividing the microphone signal determined based on the human voice frequency range and the performance information of the speaker, and the minimum frequency in the frequency band corresponding to the second sub-signal is greater than the first frequency division point.
[0050] Optionally, the first crossover point is determined based on the human voice frequency range and speaker performance information. It is a frequency point used to divide the human voice signal input by the microphone into two frequency bands. The human voice frequency range usually starts from 85Hz to 255Hz and extends to 4kHz to 10kHz, or even higher. The speaker performance information includes key parameters such as frequency response, power carrying capacity, and transient response. By analyzing the above information, the system selects an optimal first crossover point to ensure that each frequency band of the human voice signal can be accurately reproduced by the most suitable speaker unit, while avoiding overload and signal distortion.
[0051] Optionally, the smart speaker system first monitors the audio input interface and identifies the signal source. When it is determined that the signal comes from the microphone, the microphone signal is divided into a first sub-signal and a second sub-signal according to the determined first crossover point. Among them, the first sub-signal corresponds to a lower frequency range, and the second sub-signal covers a higher frequency. The key here is that the minimum frequency in the frequency band corresponding to the second sub-signal (i.e., above the first crossover point) is greater than the maximum frequency in the frequency band corresponding to the first sub-signal (i.e., below the first crossover point). This division method can ensure that the low-frequency and mid-low-frequency vocal parts are processed by one speaker (such as the first speaker, responsible for the playback of low and mid-low frequencies), while the mid-high to high-frequency vocal parts are processed by another speaker (such as the second speaker, responsible for the playback of mid-high to high frequencies). The two do not interfere with each other and each performs at its best.
[0052] From the above content, it can be seen that by implementing the above steps, the smart speaker system can significantly improve the quality of vocal output in the karaoke scene. First, the smart speaker system can accurately identify the microphone signal, avoiding signal confusion and invalid processing. Secondly, by reasonably determining the first crossover point, the effective separation of the signal is ensured, so that the low-frequency part and the high-frequency part can be processed by targeted speakers, avoiding the performance bottlenecks that may be encountered when a single speaker tries to cover the entire frequency band, such as overload risks and insufficient transient response. Ultimately, this signal processing strategy not only improves the clarity and naturalness of the human voice, but also protects the speaker unit, extends the service life of the system, and provides users with a more comfortable and high-quality karaoke experience.
[0053] In an alternative embodiment, when the smart speaker system detects that the target audio signal is a music signal, it divides the music signal into a third sub-signal and a fourth sub-signal according to a second frequency division point, where the second frequency division point is a frequency point determined based on the frequency range of the music signal and the performance information of the speaker for dividing the music signal, and the minimum frequency in the frequency band corresponding to the fourth sub-signal is greater than the second frequency division point.
[0054] Optionally, based on the wide frequency range of the music signal (such as 20 Hz to 20 kHz) and the performance information of the speakers in the system (including but not limited to frequency response, transient response, anti-overload ability, etc.), the smart speaker system calculates and determines an optimal second frequency division point. The setting of the second frequency division point ensures that the music signal can be reasonably distributed among the speakers, allowing the low-frequency part and the high-frequency part to be processed by the speakers most suitable for them respectively.
[0055] Optionally, in the smart speaker system, when the target audio signal is recognized as a music signal, the system will activate a dedicated music signal processing mechanism to ensure that every subtlety of the music can be perfectly presented. The following are the specific implementation steps of this process:
[0056] The signal detection module built into the smart speaker system monitors the audio input. Once it confirms that the signal source is the MusicIN port, it automatically recognizes the target audio signal as a music signal. This recognition process is intelligent, ensuring the accuracy and efficiency of the subsequent processing flow. Then, the music signal is divided into a third sub-signal and a fourth sub-signal according to the second frequency division point. Among them, the third sub-signal covers the audio from the lowest frequency to the second frequency division point, and this frequency band usually includes the low-frequency and mid-low-frequency parts of the music; the fourth sub-signal contains the audio from the second frequency division point to the highest frequency, that is, the mid-high-frequency and high-frequency parts of the music. Such a division ensures that the minimum frequency contained in the fourth sub-signal is higher than the maximum frequency of the third sub-signal, preventing signal overlap or interference between different frequency bands. Subsequently, the third sub-signal and the fourth sub-signal are respectively sent to the corresponding speaker units.
[0057] As can be seen from the above, by implementing the above music signal processing flow, the smart speaker system can significantly improve the quality of music playback. This process not only performs intelligent analysis based on the frequency range of the music signal, but also fully considers the performance limitations of the speakers, ensuring that each piece of music is properly frequency-divided and then distributed to the most suitable speakers for playback. In addition, this method also avoids the problems of overload and distortion that may occur when a single speaker processes the full-frequency music signal, protects the system from damage, and at the same time enhances the user's music appreciation experience.
[0058] In an alternative embodiment, when the target audio signal is a microphone signal, the smart speaker system inputs the first sub-signal in the microphone signal into the first speaker for output, and inputs the second sub-signal in the microphone signal into the second speaker for output. The first speaker and the second speaker are independent speakers. The frequency response range and power handling capacity of the first speaker in the frequency band corresponding to the first sub-signal are greater than those of the second speaker. The anti-overload capacity and the ability to reproduce dynamic signals of the second speaker in the frequency band corresponding to the second sub-signal are greater than those of the first speaker. The dynamic signal is used to characterize that the change value of the intensity or amplitude of the signal is greater than the first preset threshold.
[0059] Optionally, the smart speaker system first determines the source of the audio signal. Once it is confirmed that the signal is input through the microphone, the system switches to the microphone signal processing mode. Next, the system divides the microphone signal into two sub-signals according to a preset first crossover point: the first sub-signal and the second sub-signal. The first sub-signal, that is, the low-frequency to mid-low-frequency part in the microphone signal, is input into the first speaker for output, and the second sub-signal, that is, the mid-high-frequency to high-frequency part in the microphone signal, will be input into the second speaker for output.
[0060] Optionally, the first speaker is a woofer; the second speaker is a paper cone cloth-edge tweeter. Its diaphragm is usually made of pulp material, and the edge is reinforced with cloth and other glue coatings to enhance stability and durability. The frequency range usually covers 800Hz - 8kHz, and some designs can extend to above 10kHz in the high frequency. Because the paper cone is light in mass and the cloth edge has high compliance, it is easy to drive. At the same time, it has a certain anti-overload capacity and power reserve, and is suitable for the output of large dynamic vocals in the KTV scene. However, the high-frequency upper limit of the paper cone cloth-edge tweeter is usually lower than that of the silk diaphragm tweeter. The high-frequency details such as sibilance and instrument overtones are insufficient, and the tone color is relatively hard and not delicate enough. Secondly, the inertia of the paper cone is relatively large, and the followability for fast transient signals such as percussion is slightly inferior to that of the silk diaphragm tweeter, and it is not very suitable for the high-frequency reproduction of music signals. Therefore, the paper cone cloth-edge tweeter is selected as the second speaker in this embodiment.
[0061] Optionally, when using the microphone to speak or sing KTV, due to room effects or placement, howling will inevitably occur in extreme cases. If a conventional silk diaphragm dome tweeter is used to reproduce the microphone signal, because its anti-overload capacity and power reserve are insufficient, it is very easy to burn out. However, the method in this embodiment distributes the high-frequency part of the microphone signal to the paper cone cloth-edge speaker for reproduction. The paper cone cloth-edge speaker can reproduce the human voice clearly and brightly, and because it has a certain anti-overload capacity and power reserve, it can greatly reduce the risk of burning out the unit due to howling.
[0062] As can be seen from the above, by implementing the above specific microphone signal processing strategy, the smart speaker system can significantly improve the output quality of human voices in the KTV scenario. The system can not only intelligently identify microphone signals, but also reasonably allocate them to the most suitable speaker units according to the frequency characteristics of the signals. This design ensures that each frequency band of the human voice signal can be processed efficiently and accurately. The first speaker is responsible for the stable output from low frequency to mid-low frequency, while the second speaker focuses on the delicate and dynamic reproduction from mid-high frequency to high frequency, ensuring the warmth, thickness and clarity of the human voice, and at the same time reducing the risk of speaker overload caused by large dynamic changes (such as howling). In addition, this method reduces the complexity and cost of the overall system by optimizing signal allocation, while improving the comfort of the user's KTV experience and the durability of the speaker system.
[0063] In an optional embodiment, when the target audio signal is a music signal, the smart speaker system inputs the third sub-signal in the music signal into the first speaker for output, and inputs the fourth sub-signal in the music signal into the third speaker for output. Here, the first speaker and the third speaker are independent speakers. The frequency response range and power carrying capacity of the first speaker in the frequency band corresponding to the third sub-signal are greater than those of the third speaker, and the transient response ability and target ability of the third speaker in the frequency band corresponding to the fourth sub-signal are greater than those of the first speaker. Here, the target ability is used to characterize the ability to capture and reproduce the sound changes in the music signal.
[0064] Optionally, when the target audio signal is a music signal, the smart speaker system divides the music signal into a third sub-signal and a fourth sub-signal. Here, the third sub-signal contains the low-frequency and mid-low frequency parts of the music, while the fourth sub-signal covers the mid-high frequency and high frequency parts of the music. These two parts of the signal are then separately sent to the first speaker and the third speaker in the system for processing and reproduction.
[0065] Optionally, the third speaker is a silk dome tweeter. Its diaphragm is usually made of silk or polyester film, and its frequency range generally covers above 2kHz - 20kHz. It has excellent high-frequency extension and rich details. Because the silk film is extremely light in quality, it has a fast transient response and excellent control of splitting vibration, and its tone color is soft and natural, making it suitable for high-resolution music reproduction. However, the silk dome tweeter also has some obvious disadvantages. When applied, it needs to be combined with a mid-bass unit for frequency division. If the design is improper, it is easy to cause connection problems such as mid-high frequency break. In addition, the fibers of the silk diaphragm are relatively thin, and its ability to withstand instantaneous high power and distortion is weak. At extreme volumes, it may cause the diaphragm to deform and cause distortion, or even burn out and be damaged, so it is not suitable for continuous high-power KTV scenarios. Therefore, in this embodiment, the silk dome tweeter is selected as the third speaker.
[0066] As can be seen from the above, by implementing the above optimization distribution strategy for music signals, the smart speaker system can significantly improve the sound quality and auditory experience of music playback. The system cleverly utilizes the performance differences between the first speaker and the third speaker to ensure that each frequency band of the music signal is processed by the speaker unit most suitable for it, thereby achieving the depth and strength of the low frequency, as well as the delicacy and naturalness of the high frequency. This distribution strategy not only maximizes the performance of the speakers, preventing overload and distortion, but also enhances the dynamic range and detail performance of music playback, providing users with a more realistic and immersive music appreciation experience. At the same time, through intelligent signal recognition and professional distribution, the smart speaker system demonstrates its high flexibility and intelligence in audio processing, being able to adapt to different music styles and playback scenarios and providing high-quality audio output.
[0067] In an alternative embodiment, when the target audio signal is a microphone signal and a music signal, the smart speaker system divides the target audio signal according to the first crossover point and the second crossover point to obtain the first sub-signal and the second sub-signal of the microphone signal, as well as the third sub-signal and the fourth sub-signal of the music signal. Then, the second sub-signal is input to the second speaker for output, the fourth sub-signal in the music signal is input to the third speaker for output, and finally, the first sub-signal in the microphone signal and the third sub-signal in the music signal are mixed, and the target signal obtained after the mixing process is input to the first speaker for output.
[0068] Optionally, when the smart speaker system receives both the microphone signal and the music signal simultaneously, the smart speaker system divides the microphone signal and the music signal respectively based on the determined first crossover point and the second crossover point to obtain the first sub-signal (low-frequency to mid-low-frequency human voice) and the second sub-signal (mid-high-frequency to high-frequency human voice) of the microphone signal, as well as the third sub-signal (low-frequency to mid-frequency music) and the fourth sub-signal (high-frequency music) of the music signal. On this basis, the system inputs the second sub-signal of the microphone signal to the second speaker (paper cone edge cloth tweeter) for playback, and inputs the fourth sub-signal in the music signal to the third speaker (silk dome tweeter) for playback. At the same time, the smart speaker system mixes the first sub-signal of the microphone signal and the third sub-signal of the music signal. This process is completed by the DSP, ensuring the seamless integration of the human voice and the background music in the mid-low frequency band and creating a harmonious auditory experience. The target signal after the mixing process is input to the first speaker for output, where the first speaker has the best frequency response for processing signals in this frequency band and sufficient power handling capacity, thus ensuring the accurate playback of the mixed signal.
[0069] As can be seen from the above, the above embodiments significantly improve the audio quality of the smart speaker system when simultaneously processing microphone signals and music signals. Through precise signal detection and intelligent signal division, the system ensures that different types of audio signals are reasonably allocated to the speaker units most suitable for them. The independent playback of the second sub-signal and the fourth sub-signal, as well as the mixing process of the first sub-signal and the third sub-signal, not only protect the speakers from overload damage but also optimize the playback effect of audio signals. Especially in the mid-low frequency parts of human voices and background music, a smooth and natural sound mixture is achieved, avoiding sound breaks and disharmony, and enhancing the overall audio coherence and immersion. In addition, this strategy of multi-speaker collaborative work also effectively utilizes system resources and improves the adaptability and performance of the speaker in multiple audio scenarios.
[0070] In an alternative embodiment, after dividing the target audio signal into N sub-signals based on the signal type of the target audio signal, the smart speaker system inputs the N sub-signals into the amplifiers corresponding to each sub-signal respectively for amplification processing based on the signal source and frequency range of the target audio signal.
[0071] Optionally, when the audio input interface of the smart speaker system receives a target audio signal, the first step of the system is to identify the signal source (type), that is, to determine whether the signal is a music signal or a microphone signal. Subsequently, regardless of the source of the target audio signal, the system will perform signal division according to a preset frequency division point to obtain N sub-signals, and each sub-signal corresponds to a specific frequency band of the audio signal.
[0072] Optionally, the smart speaker system fully considers the characteristics of different speaker units during design, such as the first speaker and the third speaker (responsible for low frequency to mid-low frequency), and the second speaker and the fourth speaker (responsible for mid-high frequency to high frequency). Each speaker unit has its optimal frequency response range and power handling capacity. The system also needs to perform matching analysis on the performance of the amplifier because different amplifiers may be suitable for signals in different frequency ranges to achieve the best amplification effect. After the signal division is completed, the smart speaker system then inputs the N sub-signals into the amplifiers corresponding to each sub-signal for amplification processing according to the signal source and frequency range of the target audio signal. For example, the third sub-signal of the music signal is input into the amplifier connected to the first speaker, while the fourth sub-signal is input into the amplifier connected to the third speaker. Similarly, the first sub-signal and the second sub-signal of the microphone signal are also input into the corresponding amplifiers for amplification respectively, and the corresponding amplifiers can adjust their amplification parameters according to the specific requirements of the sub-signals to achieve the best power output. This process ensures that the audio signals in each frequency band can obtain accurate and sufficient power support, while avoiding problems such as distortion or speaker overload caused by amplifying signals in amplifiers with unsuitable frequency ranges.
[0073] Optionally, Figure 2 is a schematic diagram of an optional signal processing method according to an embodiment of the present application, as Figure 2As shown, Music IN represents the music input signal, which is an audio data stream from an external device or network and mainly contains sound signals of various music types. Mic IN represents the microphone input signal, which usually comes from the audio signal generated when the user sings or speaks and is used to collect human voices in real time. The DSP is used to digitally process the signal, including but not limited to filtering, dynamic range control, phase calibration, etc., to optimize the signal quality and adapt to the characteristics of different speakers. After receiving the audio signal, the audio signal is first subjected to sound effect processing. Among them, the sound effect processing is used to enhance the sound effects of the input music and microphone signals, such as adding a surround feeling, adjusting the balance, etc., to improve the final auditory effect. Then, the processed audio signal is transmitted to the frequency divider for processing. Among them, the distributor is used to divide the audio signal into multiple sub-signals according to the frequency characteristics to ensure that signals in different frequency ranges can be sent to the speaker units most suitable for them. In addition, the frequency divider usually sets multiple crossover points to divide the signal into high, middle, and low frequencies, etc. Then, the sub-signals of different frequency bands after the frequency division processing are respectively transmitted to the corresponding AMP (power amplifier, Amplifier) for amplification processing. Among them, when receiving both music signals and microphone signals at the same time, the mid-bass part divided by the distributor will be mixed, and then the mixed signal will be transmitted to the corresponding amplifier for processing. Finally, the amplified sub-signals are transmitted to the corresponding speakers for output. For example, the high-frequency part of the music signal is output through a silk dome tweeter, the mid-high frequency part of the microphone signal is output through a paper cone edge mid-high speaker, and the mid-low frequency part of the music signal and microphone signal is output through a woofer.
[0074] As can be seen from the above, by implementing the above intelligent amplification processing strategy for the target audio signal source and frequency range, the smart speaker system can significantly improve the playback quality of the audio signal. The system can not only intelligently identify the signal source and perform precise frequency division, but also allocate sub-signals to the most suitable components for processing according to the performance characteristics of the speakers and amplifiers. This strategy ensures that each frequency band of the audio signal can obtain the best power amplification and speaker playback. Whether it is the low-frequency intensity in the music signal or the high-frequency details in the microphone signal, they can be presented delicately and without loss. In addition, by precisely matching the signal with the amplifier, the system also avoids the risks of distortion and speaker overload during the signal processing process, improves the overall sound quality performance and system stability, and provides users with a better audio experience.
[0075] In an alternative embodiment, the smart speaker system detects the target difference value of each of the N sub-signals, where the target difference value is used to quantify the difference between the peak level and the average level in the sub-signal. When the target difference value of the i-th sub-signal is detected to be greater than or equal to the second preset threshold, the amplification factor of the amplifier corresponding to the i-th sub-signal is reduced, where i is an integer less than or equal to N. When the target difference value of the i-th sub-signal is detected to be less than the second preset threshold, the amplification factor of the amplifier corresponding to the i-th sub-signal is increased.
[0076] Optionally, when processing the N sub-signals, the smart speaker system will detect the target difference value (dynamic range) of each sub-signal in real time, so as to intelligently adjust the amplification factor of the amplifier to ensure the smoothness of the sound quality and the safe operation of the speaker.
[0077] Optionally, the smart speaker system continuously monitors the difference between the peak level and the average level of each sub-signal and calculates the target difference value. The target difference value is a key parameter used to quantify the gap between the peak level and the average level in the audio signal, and intuitively reflects the size of the signal dynamic range.
[0078] Optionally, when the target difference value of the i-th sub-signal (where i is an integer less than or equal to N) is detected to be greater than or equal to the second preset threshold set by the system, it indicates that the dynamic range of this sub-signal is too large and there may be a situation of sudden peak values, which may cause the speaker to be overloaded or the sound quality to be distorted. To prevent this from happening, the system will automatically reduce the amplification factor of the amplifier corresponding to the i-th sub-signal to lower the peak level, ensure that the signal is output within the safe range of the speaker, and at the same time maintain the clarity and stability of the sound quality. On the contrary, if the target difference value of the i-th sub-signal is less than the second preset threshold, it means that the dynamic range of the sub-signal is moderate and there are no abnormal peak levels. In this case, the system will increase the amplification factor of the amplifier corresponding to the i-th sub-signal to give full play to the potential of the speaker and improve the fullness and detail performance of the sound quality.
[0079] Optionally, while adjusting the amplification factor of the amplifier, the smart speaker system may also be equipped with a real-time optical feedback mechanism to further fine-tune the working parameters of the amplifier by monitoring the actual output state of the speaker. This mechanism ensures that the amplified signal can not only meet the optimization requirements of the sound quality but also always remain within the safe working range of the speaker.
[0080] From the above content, it can be seen that the above implementation method significantly improves the sound quality performance of the smart speaker system and the service life of the speaker. Through dynamic level optimization and amplification control, the system can automatically adjust the amplification factor of each sub-signal to ensure that even when processing audio signals with a large dynamic range, the speaker overload phenomenon can be avoided while maintaining the purity of the sound quality and richness of details. In addition, this strategy can also optimize the dynamic balance of the signal, so that every note of the music and every ups and downs of the human voice can be properly amplified and replayed, greatly improving the user's listening experience, especially for those music or live broadcast content with drastic dynamic changes. More importantly, this flexible amplification control method not only improves the sound quality, but also extends the service life of the speaker, reduces equipment maintenance costs, and improves the overall reliability and durability of the system.
[0081] In an optional embodiment, after receiving the target audio signal, the smart speaker system performs a first preprocessing operation on the microphone signal when the target audio signal is a microphone signal, wherein the first preprocessing operation includes at least a first processing and a second processing, wherein the first processing is used to monitor the signal level change in the microphone signal, determine the instantaneous level based on the signal level change, and compress the signal whose instantaneous level exceeds a third preset threshold, and the second processing is used to adjust the signal strength of different frequency bands in the microphone signal; when the target audio signal is a music signal, the second preprocessing operation is performed on the music signal, wherein the second preprocessing operation includes at least a third processing and a fourth processing, wherein the third processing is used to align the phases of different frequency signals in the music signal, and the fourth processing is used to correct the music signal according to the frequency response curve of the speaker.
[0082] Optionally, after receiving the microphone signal, the smart speaker system continuously monitors the signal level changes of the microphone signal, wherein the level changes reflect the intensity fluctuations of the signal. Especially in karaoke or speech scenes, the intensity of the human voice may change rapidly, causing speaker overload or sound distortion. Based on the signal level changes, the system can determine the instantaneous level of the microphone signal in real time. When the instantaneous level exceeds the third preset threshold set by the system, the smart speaker system will automatically start the level compression function to prevent the signal intensity from instantaneously overloading, reduce the risk of speaker damage, and ensure the stability and sound quality of the human voice output. After the first processing, the system will adjust the signal strength of different frequency bands in the microphone signal. This adjustment process is usually completed by DSP. Through gain control and other means, it ensures that each frequency band of the human voice can be appropriately enhanced, especially the mid-high frequency part, to achieve a clear and full human voice playback effect, while avoiding potential damage to the speaker by high-frequency signals.
[0083] Optionally, after receiving the music signal, the smart speaker system uses phase correction technology to align the phases of signals with different frequencies in the music signal. This is because the music signal contains multiple frequency components. If the phases of these components are inconsistent, it may lead to phase distortion of the sound, affecting the sound quality. Phase alignment ensures that when the music signal is output by the speaker, each frequency component can be synchronized, avoiding phase distortion and bringing a more natural and harmonious music experience to the user. Then the smart speaker system also corrects the music signal based on the frequency response curve of the speaker. The frequency response curve of the speaker reflects the gain characteristics of the speaker at different frequencies. Through this processing, the system can compensate for possible insufficient gain or excessive gain at certain frequencies of the speaker, ensuring that all frequency components can be accurately and evenly reproduced, improving the resolution of the music signal and the smoothness of the sound.
[0084] As can be seen from the above, through the above embodiments, the audio quality of the smart speaker system in processing microphone signals and music signals has been significantly improved. The first preprocessing operation on the microphone signal, including the monitoring and compression of the instantaneous level, and the adjustment of the signal intensity in the frequency band, ensures the stability and clarity of the human voice output, while protecting the speaker from overload damage. The second preprocessing operation on the music signal, through phase alignment and the correction of the frequency response curve, optimizes the output of the music signal and improves the naturalness and balance of music playback. These preprocessing technologies not only enhance the playback effect of the audio signal, but also further improve the compatibility and adaptability of the entire system, enabling the smart speaker system to work more intelligently and efficiently when processing complex audio signals, and providing a more personalized and immersive audio experience for users.
[0085] The embodiment of the present application also provides a signal processing device. It should be noted that the signal processing device in the embodiment of the present application can be used to execute the signal processing method provided in the embodiment of the present application. The signal processing device provided in the embodiment of the present application is introduced below.
[0086] According to the embodiment of the present application, there is also provided a device for implementing the above signal processing method, Figure 3 which is a schematic diagram of an optional signal processing device according to the embodiment of the present application, as Figure 3As shown in the figure, the device includes: a receiving unit 301, configured to receive a target audio signal, where the target audio signal is a music signal and / or a microphone signal; a dividing unit 302, configured to divide the target audio signal into N sub-signals based on the signal type of the target audio signal, where N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different; a first output unit 303, configured to, when the target audio signal is a microphone signal, input S sub-signals greater than a first frequency division point among the N sub-signals into S speakers for output respectively, and input N - S sub-signals into the same speaker for output, where S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from both the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal; a second output unit 304, configured to, when the target audio signal is a music signal, input T sub-signals greater than a second frequency division point among the N sub-signals into T speakers for output respectively, and input N - T sub-signals into the same speaker for output, where T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from both the T speakers, the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
[0087] Optionally, the dividing unit 302 includes: a first sub-dividing unit, configured to, when detecting that the target audio signal is a microphone signal, divide the microphone signal into a first sub-signal and a second sub-signal according to the first frequency division point, where the first frequency division point is the frequency point for dividing the microphone signal determined based on the vocal frequency range and the performance information of the speaker, and the minimum frequency in the frequency band corresponding to the second sub-signal is greater than the first frequency division point.
[0088] Optionally, the dividing unit 302 includes: a second sub-dividing unit, configured to, when detecting that the target audio signal is a music signal, divide the music signal into a third sub-signal and a fourth sub-signal according to the second frequency division point, where the second frequency division point is the frequency point for dividing the music signal determined based on the music signal frequency range and the performance information of the speaker, and the minimum frequency in the frequency band corresponding to the fourth sub-signal is greater than the second frequency division point.
[0089] Optionally, the first output unit 303 includes: a first output subunit, configured to, when the target audio signal is a microphone signal, input a first sub-signal in the microphone signal to a first speaker for output, and input a second sub-signal in the microphone signal to a second speaker for output, where the first speaker and the second speaker are independent speakers, the frequency response range and power handling capacity of the first speaker in the frequency band corresponding to the first sub-signal are greater than those of the second speaker, and the anti-overload capacity and the ability to reproduce dynamic signals of the second speaker in the frequency band corresponding to the second sub-signal are greater than those of the first speaker, where the dynamic signal is used to characterize that the change value of the intensity or amplitude of the signal is greater than a first preset threshold.
[0090] Optionally, the second output unit 304 includes: a second output subunit, configured to, when the target audio signal is a music signal, input a third sub-signal in the music signal to a first speaker for output, and input a fourth sub-signal in the music signal to a third speaker for output, where the first speaker and the third speaker are independent speakers, the frequency response range and power handling capacity of the first speaker in the frequency band corresponding to the third sub-signal are greater than those of the third speaker, and the transient response ability and the target ability of the third speaker in the frequency band corresponding to the fourth sub-signal are greater than those of the first speaker, where the target ability is used to characterize the ability to capture and reproduce sound changes in the music signal.
[0091] Optionally, the signal processing device further includes: a first division unit, a third output unit, and a fourth output unit. The first division unit is configured to, when the target audio signal is a microphone signal and a music signal, divide the target audio signal according to a first frequency division point and a second frequency division point to obtain a first sub-signal and a second sub-signal of the microphone signal, and a third sub-signal and a fourth sub-signal of the music signal; the third output unit is configured to input the second sub-signal to the second speaker for output, and input the fourth sub-signal in the music signal to the third speaker for output; the fourth output unit is configured to perform a mixing process on the first sub-signal in the microphone signal and the third sub-signal in the music signal, and input the target signal obtained after the mixing process to the first speaker for output.
[0092] Optionally, the signal processing device further includes: a first processing unit, configured to, based on the signal type and frequency range of the target audio signal, input N sub-signals to corresponding amplifiers of each sub-signal for amplification processing.
[0093] Optionally, the signal processing device further includes: a first detection unit, a second processing unit, and a third processing unit. Among them, the first detection unit is configured to detect the target difference value of each of the N sub-signals, where the target difference value is used to quantify the difference between the peak level and the average level in the sub-signal; the second processing unit is configured to reduce the amplification factor of the amplifier corresponding to the i-th sub-signal when it is detected that the target difference value of the i-th sub-signal is greater than or equal to a second preset threshold, where i is an integer less than or equal to N; the third processing unit is configured to increase the amplification factor of the amplifier corresponding to the i-th sub-signal when it is detected that the target difference value of the i-th sub-signal is less than the second preset threshold.
[0094] Optionally, the signal processing device further includes: a fourth processing unit and a fifth processing unit. Among them, the fourth processing unit is configured to perform a first preprocessing operation on the microphone signal when the target audio signal is a microphone signal, where the first preprocessing operation at least includes a first process and a second process. The first process is used to monitor the signal level change in the microphone signal, determine the instantaneous level based on the signal level change, and compress the signal whose instantaneous level exceeds a third preset threshold. The second process is used to adjust the signal strength of different frequency bands in the microphone signal; the fifth processing unit is configured to perform a second preprocessing operation on the music signal when the target audio signal is a music signal, where the second preprocessing operation at least includes a third process and a fourth process. The third process is used to align the phases of different frequency signals in the music signal, and the fourth process is used to correct the music signal according to the frequency response curve of the speaker.
[0095] According to another aspect of the embodiments of the present application, there is also provided a computer-readable storage medium, which includes an executable program stored therein. When the executable program runs, it controls the device where the computer-readable storage medium is located to execute the above-mentioned signal processing method.
[0096] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned signal processing method.
[0097] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages and disadvantages of the embodiments.
[0098] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0099] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings or direct couplings or communication connections shown or discussed with each other can be through some interfaces. The indirect couplings or communication connections of units or modules can be in electrical or other forms.
[0100] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0101] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0102] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0103] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A signal processing method, characterized in that, Including: Receiving a target audio signal, where the target audio signal is a music signal and / or a microphone signal; Based on the signal type of the target audio signal, dividing the target audio signal into N sub-signals, where N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different; When the target audio signal is a microphone signal, inputting S sub-signals among the N sub-signals that are greater than a first frequency division point into S speakers for output, and inputting N - S sub-signals into the same speaker for output, where S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal; When the target audio signal is a music signal, inputting T sub-signals among the N sub-signals that are greater than a second frequency division point into T speakers for output, and inputting N - T sub-signals into the same speaker for output, where T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from the T speakers, and the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
2. The signal processing method according to claim 1, wherein Based on the signal type of the target audio signal, dividing the target audio signal into N sub-signals includes: When it is detected that the target audio signal is the microphone signal, dividing the microphone signal into a first sub-signal and a second sub-signal according to the first frequency division point, where the first frequency division point is the frequency point determined based on the human voice frequency range and the performance information of the speaker for dividing the microphone signal, and the minimum frequency in the frequency band corresponding to the second sub-signal is greater than the first frequency division point.
3. The signal processing method according to claim 2, wherein Based on the signal type of the target audio signal, dividing the target audio signal into N sub-signals includes: When it is detected that the target audio signal is the music signal, dividing the music signal into a third sub-signal and a fourth sub-signal according to the second frequency division point, where the second frequency division point is the frequency point determined based on the music signal frequency range and the performance information of the speaker for dividing the music signal, and the minimum frequency in the frequency band corresponding to the fourth sub-signal is greater than the second frequency division point.
4. The signal processing method according to claim 2, wherein When the target audio signal is a microphone signal, inputting S sub-signals among the N sub-signals that are greater than a first frequency division point into S speakers for output, and inputting N - S sub-signals into the same speaker for output, includes: When the target audio signal is the microphone signal, the first sub-signal in the microphone signal is input into the first speaker for output, and the second sub-signal in the microphone signal is input into the second speaker for output. Herein, the first speaker and the second speaker are independent speakers. The frequency response range and power carrying capacity of the first speaker in the frequency band corresponding to the first sub-signal are greater than those of the second speaker, and the anti-overload capacity and the ability to reproduce dynamic signals of the second speaker in the frequency band corresponding to the second sub-signal are greater than those of the first speaker, where the dynamic signal is used to characterize that the change value of the signal intensity or amplitude is greater than a first preset threshold.
5. The signal processing method according to claim 3, wherein When the target audio signal is a music signal, T sub-signals among the N sub-signals that are greater than a second frequency division point are respectively input into T speakers for output, and N - T sub-signals are input into the same speaker for output, including: When the target audio signal is the music signal, the third sub-signal in the music signal is input into the first speaker for output, and the fourth sub-signal in the music signal is input into the third speaker for output. Herein, the first speaker and the third speaker are independent speakers. The frequency response range and power carrying capacity of the first speaker in the frequency band corresponding to the third sub-signal are greater than those of the third speaker, and the transient response ability and the target ability of the third speaker in the frequency band corresponding to the fourth sub-signal are greater than those of the first speaker, where the target ability is used to characterize the ability to capture and reproduce the sound changes in the music signal.
6. The signal processing method according to claim 3, wherein The method further includes: When the target audio signal is the microphone signal and the music signal, the target audio signal is divided according to the first frequency division point and the second frequency division point to obtain the first sub-signal and the second sub-signal of the microphone signal, and the third sub-signal and the fourth sub-signal of the music signal; The second sub-signal is input into the second speaker for output, and the fourth sub-signal in the music signal is input into the third speaker for output; The first sub-signal in the microphone signal and the third sub-signal in the music signal are subjected to mixing processing, and the target signal obtained after the mixing processing is input into the first speaker for output.
7. The signal processing method according to claim 1, wherein After dividing the target audio signal into N sub-signals based on the signal type of the target audio signal, the method further includes: Based on the signal type and frequency range of the target audio signal, the N sub-signals are respectively input into the amplifiers corresponding to each sub-signal for amplification processing.
8. The signal processing method according to claim 7, characterized in that Before the N sub-signals are respectively input into the amplifiers corresponding to each sub-signal for amplification processing, the method further includes: Detecting the target difference value of each of the N sub-signals, where the target difference value is used to quantify the difference between the peak level and the average level in the sub-signal; When it is detected that the target difference value of the i-th sub-signal is greater than or equal to the second preset threshold, the amplification factor of the amplifier corresponding to the i-th sub-signal is reduced, where i is an integer less than or equal to N; When it is detected that the target difference value of the i-th sub-signal is less than the second preset threshold, the amplification factor of the amplifier corresponding to the i-th sub-signal is increased.
9. The signal processing method according to claim 1, wherein After receiving the target audio signal, it includes: When the target audio signal is the microphone signal, a first preprocessing operation is performed on the microphone signal, where the first preprocessing operation at least includes a first process and a second process, where the first process is used to monitor the signal level change in the microphone signal, determine the instantaneous level based on the signal level change, and compress the signal whose instantaneous level exceeds the third preset threshold, and the second process is used to adjust the signal intensity of different frequency bands in the microphone signal; When the target audio signal is the music signal, a second preprocessing operation is performed on the music signal, where the second preprocessing operation at least includes a third process and a fourth process, where the third process is used to align the phases of different frequency signals in the music signal, and the fourth process is used to correct the music signal according to the frequency response curve of the speaker.
10. A signal processing device, characterized in that, It includes: A receiving unit that receives a target audio signal, where the target audio signal is a music signal and / or a microphone signal; A dividing unit that divides the target audio signal into N sub-signals based on the signal type of the target audio signal, where N is an integer greater than 1, and the frequency bands corresponding to different sub-signals are different; A first output unit that, when the target audio signal is the microphone signal, inputs S sub-signals among the N sub-signals that are greater than the first frequency division point into S speakers for output respectively, and inputs N - S sub-signals into the same speaker for output, where S is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - S sub-signals is different from the S speakers, and the first frequency division point is the frequency point for dividing the microphone signal; A second output unit that, when the target audio signal is the music signal, inputs T sub-signals among the N sub-signals that are greater than the second frequency division point into T speakers for output respectively, and inputs N - T sub-signals into the same speaker for output, where T is an integer greater than or equal to 1 and less than or equal to N, the speaker corresponding to the N - T sub-signals is different from the T speakers, the T speakers are different from the S speakers, and the second frequency division point is the frequency point for dividing the music signal.
11. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, where when the computer program runs, it causes the device where the computer-readable storage medium is located to execute the signal processing method according to any one of claims 1 to 9.
12. An electronic device, characterized in that, Comprising one or more processors and a memory, the memory being used for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the signal processing method according to any one of claims 1 to 9.
Citation Information
Cited By
Collar clip type sound amplification system, squeal removal processing method thereof and related device
CN121001022A