Beating machine control method based on audio signals

By denoising the audio signal and frequency domain expansion, real-time analysis and extraction of beat signals, the problem that existing beat machines cannot match the music rhythm, the match between beat rhythm and music rhythm is achieved, and the user experience and the intelligence level of beat machines are improved.

CN120204567APending Publication Date: 2025-06-27DONGGUAN BEIPAIPAI TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510184170.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When existing slap machines slap according to the music rhythm, they cannot analyze the audio signals in real time, resulting in the slap rhythm that does not match the music rhythm, affecting the user experience.

Method used

By denoising and reconstructing the audio signal, frequency domain expansion and peak detection technology are used to analyze the audio signal in real time, extract the beat signal, and output it to the control unit of the beat machine, to achieve accurate control of the beat intensity and frequency.

Benefits of technology

The slap rhythm and music rhythm are achieved, which significantly improves the user experience effect, and improves the intelligent level of the slap machine and the effect of musical therapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120204567A_ABST
    Figure CN120204567A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and mainly relates to an audio signal-based flapping machine control method, which comprises the following steps of: acquiring an audio signal through a database or external audio sampling, performing de-noising processing to obtain an original signal, and reconstructing the original signal to obtain a base signal; performing frequency domain expansion on the base signal to obtain a short-time energy spectrum of the base signal; afterwards, position signals are extracted from the short-time energy spectrum through peak detection, the energy peak value of each position signal is detected, and a time signal corresponding to each beat is determined according to the energy peak value; and finally, a beat signal is constructed according to the time signal corresponding to each beat, and the beat signal is transmitted to a control unit of the beating machine. The control unit converts the beat signal into a driving signal, and the driving signal further drives the flapping machine to perform flapping action on the user; therefore, the technical defect that an existing beating control method cannot enable the beating machine to be matched with the audio signal is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and particularly relates to a control method for a patting machine based on an audio signal. Background Art

[0002] In recent years, the relationship between music and physical and mental health has received extensive attention. Numerous studies have shown that music can affect people's emotions and physiological responses, and helps to relieve stress, improve mood, and promote physical relaxation and recovery. Based on this, patting machines combined with music rhythms have gradually become popular products in the field of health care and physical therapy. Such devices can provide rhythmic body pats while users enjoy music, enhancing the pleasure of music and improving the physical therapy effect.

[0003] Although traditional patting machines can provide a certain degree of relaxation and comfort, there are still obvious technical defects in the existing devices when patting according to the music rhythm. Existing devices usually cannot analyze audio signals in real time and are difficult to quickly respond to changes in music rhythms. This results in a mismatch between the patting rhythm and the music rhythm, affecting the user experience. Moreover, existing patting machines rely on fixed settings for beat output and lack the ability to dynamically adjust, and cannot adapt to the changing music characteristics, resulting in the inability to accurately control the patting intensity and frequency in actual use.

[0004] Based on this, it is urgent to improve the existing control method for patting machines to solve the technical defects existing in the prior art. Summary of the Invention

[0005] The purpose of the present invention is to provide a control method for a patting machine based on an audio signal to solve the technical defect that the prior art cannot match the audio signal.

[0006] In order to achieve the above technical purpose, the following technical solutions are implemented in this application:

[0007] A control method for a patting machine based on an audio signal includes the following steps:

[0008] S101. Denoise the audio signal obtained through the database or external audio sampling to obtain the original signal, and reconstruct the original signal to obtain the base signal;

[0009] S201. Expand the base signal in the frequency domain through changes in time and frequency, and calculate the spectral energy of the expanded base signal in each time window to obtain the short-time energy spectrum of the base signal;

[0010] S301. Extract the position signal from the short-time energy spectrum through peak detection. Detect the energy peak of each position signal, and compare the energy peak at each position with a preset judgment threshold to obtain the time signal corresponding to each beat.

[0011] S401. Construct a beat signal based on the time signal corresponding to each beat and output the beat signal to the control unit of the flapper.

[0012] The control unit converts the beat signal into a driving signal, and the driving signal drives the flapper to slap the user.

[0013] The above technical solution has the following technical effects:

[0014] The flapper control method based on audio signals of the present invention can quickly respond to changes in music rhythm by real-time analyzing audio signals, ensuring a high degree of matching between the slapping rhythm and the music rhythm, thus significantly improving the user experience. In addition, by dynamically adjusting the beat output, this method adapts to diverse music characteristics and achieves precise control of the slapping intensity and frequency. Compared with traditional flappers, the present invention not only improves the intelligence level of the device but also further enhances the effect of music physiotherapy, bringing a more comfortable and pleasant user experience.

[0015] As a further improvement to the flapper control method based on audio signals of the present invention, in step S101, the audio signal is split layer by layer through discrete wavelet transform, and the approximation coefficient and detail coefficient of each layer of the split audio signal are obtained through layer-by-layer calculation.

[0016] Denoise the detail coefficients in each layer of the audio signal, and combine the detail coefficients in each layer of the audio signal after denoising and the approximation coefficients of each layer of the split audio signal to obtain the original signal.

[0017] As a further improvement to the flapper control method based on audio signals of the present invention, the calculation method of the approximation coefficient in each layer of the audio signal is as follows:

[0018]

[0019] Among them, 2i and 2i + 1 are used to represent the sample pairs processed in the audio signal represented by X[n]; n and i are the sample indices in the audio signal, representing the position of each sample; c j [i] represents the approximation coefficient in the jth layer;

[0020] The calculation method of the detail coefficient is as follows:

[0021]

[0022] Among them, d j [i] represents the detail coefficient at the j-th layer, indicating the high-frequency feature of the audio signal at the current j-th level.

[0023] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the method for denoising the detail coefficient is: retaining the part where the value of the detail coefficient of each layer of the audio signal is greater than or equal to the first threshold, and removing the part where the value of the detail coefficient of each layer of the audio signal is less than the first threshold.

[0024] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the original signal is reconstructed by inverse wavelet transform to obtain a base signal, and the expression form of the inverse wavelet transform is:

[0025]

[0026] Among them, x'[n] is the base signal after denoising processing, L is the number of layers of wavelet transform, c j is the approximation coefficient at the j-th layer, d j is the detail coefficient at the j-th layer.

[0027] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the type of wavelet is the Haar wavelet.

[0028] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, in step S101, the base signal is expanded by the change of the short-time Fourier transform in time and frequency, and the expression form of the short-time Fourier is:

[0029]

[0030] Among them, Z xx [f,t] represents the frequency-domain representation of the base signal at frequency f and time t; x[m] represents the value of the base signal at time m; w[m - n] represents the value of the window function at time m - n; e -j2πfm is used to complete the transformation of the base signal from the time domain to the frequency domain, and f is the frequency.

[0031] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the spectral energy of each time window is calculated based on the spectral energy calculation method;

[0032] The expression of the spectral energy calculation method is:

[0033] E[n] = ∑|Z xx [f,n]| 2

[0034] Among them, E[n] is the spectral energy value of the base signal at the sample indexing node n; Z xx [f, t] represents the frequency-domain representation of the base signal at frequency f and time t.

[0035] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the value range of the judgment threshold in step S301 is 0.5 - 0.7.

[0036] As a further improvement to a method for controlling a flapper based on an audio signal according to the present invention, the expression of the drive signal is:

[0037]

[0038] Among them, P[i] is a discrete sequence and satisfies P[i] ∈ {0, 1}; S(t) is the drive signal output by the control unit, δ(t - t i ) is the Dirac pulse function, corresponding to the time signal in the beat signal, A mode is the back-patting mode information. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:

[0040] Figure 1 is the flowchart of the flapper control method in Embodiment 1 of the present invention;

[0041] Figure 2 is the waveform diagram of collecting the drumbeat signal in Embodiment 2 of the present invention;

[0042] Figure 3 are the high-frequency and low-frequency signals of the 1st - 4th layers in the discrete wavelet transform processing in Embodiment 2 of the present invention;

[0043] Figure 4 is the waveform diagram of the base signal reconstructed by the inverse wavelet transform in Embodiment 2 of the present invention;

[0044] Figure 5 is the wave amplitude spectrogram of the base signal expanded by the short-time Fourier transform in terms of time and frequency in Embodiment 3 of the present invention;

[0045] Figure 6 is the short-time energy spectrogram of the base signal in Embodiment 3 of the present invention;

[0046] Figure 7 is the waveform diagram of the total energy of the base signal and the detected beats in Embodiment 3 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] The present invention will be described in detail below with reference to the accompanying drawings and in combination with the implementation regulations. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0048] The following detailed descriptions are all exemplary descriptions, aiming to provide further detailed explanations for the present invention. Unless otherwise specified, all technical terms adopted by the present invention have the same meanings as those commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the present invention are only for describing specific implementation manners and are not intended to limit the exemplary implementation manners according to the present invention.

[0049] To facilitate the preparation and understanding of the solutions provided in the following embodiments of the present invention, before describing the technical solutions provided by the present invention, the terms related to the present invention are explained as follows:

[0050] A database is a systematic and well-organized collection of data, usually stored electronically in a computer system. It allows users to conveniently store, retrieve, manage, and update data. A database can be operated through a database management system (DBMS), and the DBMS provides a controllable way to access and manage the data in the database.

[0051] Spectral energy refers to the energy distribution contained in each frequency component of a signal in the frequency domain (spectrum). It is usually related to fields such as signal processing and vibration analysis. Spectral energy represents the energy distribution of a signal at different frequencies, usually obtained by performing a Fourier transform (or other transforms) on the signal to obtain its spectrum. In the spectrum, the relationship between the energy of the signal and the frequency can help us analyze the characteristics of the signal, and spectral energy plays an important role in signal analysis, communication, audio processing, and various engineering applications.

[0052] Peak detection is a technique in signal processing used to identify and extract local maxima (peaks) and local minima (valleys) in a signal. It is very important in many application fields, including audio analysis, biosignal processing, image processing, financial data analysis, etc. The following are some key points of peak detection: The goal of peak detection is to find the characteristic points in the signal that are higher or lower than their surrounding values, and these characteristic points represent the extreme values of the signal.

[0053] Beat refers to the way of organizing musical notes and rests in music, usually used to form the rhythm basis of music. It is a measure of time in music, helping musicians and listeners understand the flow and structure of music. At the same time, the beat is the most basic time unit in music, representing the pulsation of music. Usually, it is a uniform accent that can be felt. For example, when performing, people often use their feet to beat the rhythm to keep the tempo.

[0054] A drive signal refers to a signal used to control or activate a system, device, or circuit. Such a signal can be voltage, current, or a pulse, and is typically used to change a physical quantity, such as speed, position, or power. For example, a PWM signal (Pulse Width Modulation): adjusts the output power by changing the width of the pulse, and is often used in motor control and brightness adjustment.

[0055] The Discrete Wavelet Transform (DWT) is an important signal processing technique widely used in fields such as signal analysis, image processing, and data compression. By decomposing a signal in terms of scale and position, DWT decomposes the signal into multiple wavelet coefficients of different frequency components, enabling the analysis of the signal in both the time domain and the frequency domain.

[0056] A wavelet is a function with finite support and rapid decay, having good localization properties in both time and frequency. Different from the Fourier transform, the wavelet transform can provide information about the signal within a specific time-frequency range.

[0057] As is known by common technical knowledge, the present invention can be implemented by other embodiments that do not depart from its spirit or essential features. Therefore, the above-disclosed embodiments are illustrative in all aspects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.

[0058] Embodiment 1

[0059] As Figure 1 shown, to solve the problem that the control method of the slapper in the prior art cannot match the music rhythm, the present application makes the following improvements to the existing slapper control method: Specifically, the slapper control method based on the audio signal of the present application includes the following steps:

[0060] S101. Denoise the audio signal obtained through the database or external audio sampling to obtain the original signal, and reconstruct the original signal to obtain the basic signal;

[0061] S201. Expand the frequency domain of the basic signal through changes in time and frequency, and calculate the spectral energy of the expanded basic signal in each time window to obtain the short-time energy spectrum of the basic signal;

[0062] S301. Extract the position signal from the short-time energy spectrum through peak detection, detect the energy peak of each position signal, and judge and compare the energy peak at each position according to a preset judgment threshold to obtain the time signal corresponding to each beat;

[0063] S401. Construct a beat signal based on the time signal corresponding to each beat and output the beat signal to the control unit of the beater;

[0064] The control unit converts the beat signal into a driving signal, and the driving signal drives the beater to beat the user.

[0065] Furthermore, the working principle of the above technical solution is as follows: The beater control method based on audio signals of the present invention first obtains audio signals through a database or external audio sampling. This step ensures the diversity and real-time nature of the audio signals. After obtaining the audio signals, noise reduction processing is performed to eliminate the influence of background noise on subsequent processing, thereby obtaining a purer original signal. Subsequently, the original signal is reconstructed to obtain a base signal. This process helps to extract the main features of the audio signal and provides a basis for subsequent processing.

[0066] After obtaining the base signal, the present invention adopts a frequency-domain expansion method to change the base signal in terms of time and frequency, and calculates the spectral energy on each time window, thereby obtaining the short-time energy spectrum of the base signal. This step can capture the energy distribution of the audio signal changing with time and provides key information for subsequent beat extraction.

[0067] Next, position signals are extracted from the short-time energy spectrum through peak detection technology. These position signals correspond to the beat positions in the audio signal. Then, the energy peaks of each position signal are detected, and the energy peaks at each position are judged and compared according to a preset judgment threshold, thereby obtaining the time signal corresponding to each beat. This step can accurately identify the beats in the audio signal and provides accurate time information for subsequent construction of the beat signal.

[0068] Finally, a beat signal is constructed based on the time signal corresponding to each beat, and the beat signal is output to the control unit of the beater. The control unit converts the beat signal into a driving signal, and the driving signal further drives the beater to beat the user. This process realizes the synchronization of the audio signal and the beating action, ensures the matching of the beating rhythm and the music rhythm, and thus improves the user experience.

[0069] Furthermore, compared with the prior art, the present application adopts noise reduction processing and peak detection technology in the signal processing process, effectively eliminating the interference of background noise and accurately extracting the beat information in the audio signal. This not only improves the accuracy of beat recognition but also enhances the stability and reliability of the beater's action. Moreover, the present invention realizes precise control of the beating intensity and frequency by dynamically adjusting the beat output. This improvement enables the beater to adaptively adjust according to different music characteristics and provides a more personalized and comfortable user experience.

[0070] In summary, the present invention has achieved significant technological innovations and improvements in aspects such as audio signal processing, beat recognition, and percussion machine control, bringing a better music therapy experience to users.

[0071] Embodiment 2

[0072] As Figures 2 - 4 shown, different from Embodiment 1: in order to further improve the sampling processing accuracy of the audio signal of the present application, further, as a further improvement to a percussion machine control method based on an audio signal of the present invention, in step S101, the audio signal is split layer by layer through discrete wavelet transform and the approximation coefficients and detail coefficients of each layer of the split audio signal are obtained through layer-by-layer calculation;

[0073] Denoising processing is performed on the detail coefficients in each layer of the audio signal, and the original signal is obtained by combining the detail coefficients in each layer of the audio signal after denoising processing and the approximation coefficients of each layer of the split audio signal. Specifically, the calculation method of the approximation coefficients in each layer of the audio signal is:

[0074]

[0075] where 2i and 2i + 1 are used to represent the sample pairs processed in the audio signal represented by X[n]; n and i are the sample indices in the audio signal, expressing the position of each sample; c j [i] represents the approximation coefficient in the jth layer;

[0076] In the specific implementation process,

[0077] The calculation method of the detail coefficients is:

[0078]

[0079] where d j [i] represents the detail coefficient in the jth layer, representing the high-frequency characteristics of the audio signal at the current j level.

[0080] It should be noted that in the above expressions of the detail coefficients and approximation coefficients in the jth layer, there is: for 0 ≤ n < N, i is derived from the position of n. (where n needs to be grouped on the square value) Here, the integer function is used to locate the current sample pair. Example: If n = 3, N = 8, in the first layer (j = 1), then: Corresponding to the calculation of c1[1] and d1[1].

[0081] Further, the denoising method for the detail coefficients is as follows: the part of the detail coefficients of each layer of the audio signal whose value is greater than or equal to the first threshold is retained, and the part of the detail coefficients of each layer of the audio signal whose value is less than the first threshold is removed.

[0082] That is, it satisfies the following discrete mathematics model:

[0083]

[0084] Among them, threshold is the set first threshold, which is used to determine whether to retain the detail coefficients. The part of the detail coefficients that is less than the threshold is considered noise and is therefore removed.

[0085] Further, for multi-layer discrete wavelet transform, in each layer of recursive calculation of the audio signal, the low-frequency part will replace the original signal for further processing. Suppose L layers of processing are performed, and the obtained coefficient set is:

[0086] coeffs L ={(c1,d1),(c2,d2),(c L ,d L )}

[0087] The approximation coefficients and detail coefficients in the above coefficient set only consider the coefficients of each layer processed according to the sample index i, which does not mean that there is a lack of coefficient details here.

[0088] Among them, the length of the audio signal to be processed is n, and the result of the discrete wavelet transform decomposition is as follows:

[0089] The first layer (length n):

[0090]

[0091] After the first layer of processing, the lengths of the approximation coefficient c1 and the detail coefficient d1 are both

[0092] The second layer (length ):

[0093]

[0094] After the second layer of processing, the lengths of the approximation coefficient c2 and the detail coefficient d2 are

[0095] The third layer (length ):

[0096]

[0097] After the third layer of processing, the lengths of the approximation coefficient c3 and the detail coefficient c3 are

[0098] Further, keep recursing until reaching the L-th layer.

[0099] Among them, the last layer c L and d L has a length of When reaching d = log2(N), the final length of the signal is approximately 1 (i.e., the signal cannot be further divided).

[0100] In the specific implementation process, Figure 2 is the time-domain waveform containing noise for the display of drum music in a noisy environment obtained through external audio sampling, where the horizontal axis is time (seconds) and the vertical axis is the signal amplitude, intuitively representing the change of the signal. After being processed by the above-mentioned discrete wavelet transform (only taking the expansion of 4 layers as an example here), the obtained processing result is as Figure 3 shown. It can be observed that the approximation coefficients and detail coefficients after being processed by the discrete wavelet transform, that is, after being processed by the threshold, which helps to compare the noise reduction effect and the result of signal processing. Figure 4 Then it shows the waveform diagram of the audio signal in the time domain obtained after wavelet reconstruction. The position and intensity of the beats can be clearly seen, showing the signal obtained after wavelet reconstruction. This graph allows us to observe the main differences between the denoised signal and the original signal.

[0101] As a further improvement to a method for controlling a percussion machine based on an audio signal of the present invention, the original signal is reconstructed by inverse wavelet transform to obtain a base signal, and the expression form of the inverse wavelet transform is:

[0102]

[0103] Among them, x'[n] is the base signal after being processed by noise reduction, L is the number of layers of wavelet transform, c j is the approximation coefficient of the j-th layer, and d j is the detail coefficient of the j-th layer.

[0104] Further, when specifically referring to the approximation coefficient of the first layer, the simplified expression is:

[0105]

[0106] Specifically, although the approximation coefficients of the first layer are used in the above simplified reconstruction process, in fact, as the level increases, the subsequent low-frequency components can be considered as the "accumulation" of the detailed information. In other words, the approximation coefficients of the first layer provide the "basis" for the subsequent details, making the reconstruction more accurate. In the actual signal reconstruction process, whether it is high-frequency or low-frequency information, they all participate in the complete reconstruction of the signal together. Using the approximation coefficients of the first layer plus all the detailed coefficients is to ensure that the important information in the signal can be captured.

[0107] From the perspective of structural decomposition. In wavelet transform, signal processing is carried out hierarchically step by step. The approximation coefficient c1 of the first layer provides a basic representation of the original signal, representing the overall trend or basic structure of the signal.

[0108] In terms of the hierarchical representation of details, the detailed coefficient d of each layer j captures the changes in the signal at this hierarchical scale. The detailed coefficients of different layers provide different frequency band information of the signal, reflecting the higher frequency components of the signal.

[0109] Specifically, using the approximation coefficient c1 of the first layer to reconstruct the signal is a simplified representation in wavelet decomposition and reconstruction. The purpose is to effectively reconstruct the basic features and changes of the signal by combining the low-frequency information and the high-frequency information of each layer. In this way, when performing multi-level signal processing, the important information in the signal can be retained to the greatest extent. In specific applications, according to the requirements, higher-level approximation coefficients can be selected to promote signal reconstruction, as shown in the original formula.

[0110] Furthermore, the type of wavelet selected in the above discrete wavelet transform is the Haar wavelet. Among them, the Haar wavelet has the advantages of discreteness and fast calculation and is often used in basic signal processing. The high-frequency information and the low-frequency information will still be calculated in a similar way, but different wavelet categories will be different in the calculation process, filter design and effect performance.

[0111] For the rest that is the same as in Embodiment 1, it will not be elaborated in this embodiment.

[0112] Embodiment 3

[0113] As Figures 1 - 7 shown, different from Embodiment 1: In order to further improve the extraction accuracy of the beat signal of the present application, further, in step S101, the base signal is expanded through the changes in time and frequency by short-time Fourier transform, and the expression form of short-time Fourier is:

[0114]

[0115] Among them, Z xx[f,t] represents the frequency-domain representation of the base signal at frequency f and time t; x[m] represents the value of the base signal at time m; w[m-n] represents the value of the window function at time m-n; e -j2πfm is used to complete the transformation of the base signal from the time domain to the frequency domain, and f is the frequency.

[0116] It should be noted that the basic idea of the STFT process is to divide the signal x[n] into overlapping short time periods (windows), and apply a window function (usually a window function such as the Hanning window, Hamming window, etc.) to each segment. Each segment should be regarded as a local signal and can be analyzed independently. In the specific implementation process, the above technical solution defines a window function w[m-n] and applies it to the short-time window of the signal. This window function determines the sample range we consider when performing frequency-domain analysis.

[0117] For each window center position n, calculate all the samples within this window:

[0118]

[0119] Thus, the time signal is transformed into a frequency-domain representation by means of weighted summation. This method can capture the changes of the signal at different times and frequencies.

[0120] Furthermore, calculate the spectral energy of each time window based on the spectral energy calculation method;

[0121] The expression of the spectral energy calculation method is:

[0122] E[n] = ∑|Z xx [f,n]| 2

[0123] where E[n] is the spectral energy value of the base signal at the sample index node n; Z xx [f,n] represents the frequency-domain representation of the base signal at frequency f and time t. It should be noted that the above STFT process Z xx [f,t] represents the frequency-domain representation of the signal at frequency f and time t, and Z xx [f,n] is only the expression form of the function. Since n is an index in time, it does not represent calculating the spectral energy through a new variable, but only a different manifestation form to distinguish whether it is expanded around the window center time.

[0124] Furthermore, in step S301, the value range of the judgment threshold is 0.5 - 0.7. In the specific implementation process, when the judgment threshold reaches 0.5, most of the noisy signals can be eliminated. The specific range of the judgment threshold can be adjusted according to the type of the input audio signal in the specific implementation process.

[0125] Among them,Figure 5 This is the amplitude spectrogram of the short-time Fourier transform (STFT) of the base signal, taking the drumbeat music in Embodiment 2 as an example. By visualizing the amplitude spectrum of the STFT result, it shows the energy distribution of the signal at different frequencies at each time point. The horizontal axis represents time, the vertical axis represents frequency, and the color or brightness represents the energy intensity (in dB). Figure 6 This is the spectrogram of the short-time energy spectrum of the base signal. The horizontal axis represents time, and the vertical axis represents the energy value corresponding to the frequency. By comparing Figure 5 and Figure 6 , it can be observed that in the short-time energy spectrogram, the characteristics of the beat signal are more obvious, and the high-frequency noise is further suppressed. Figure 6 The bright spots in represent the distribution of the beat signal in time and frequency. The energy values corresponding to these bright spots are relatively high, reflecting the distribution of the beat signal.

[0126] Furthermore, Figure 7 This is the waveform diagram of the total energy of the base signal and the detected beats, showing the change of the total energy within the entire signal time range and marking the positions of the detected beats. The horizontal axis represents time, the vertical axis represents energy, and the red dots represent the positions of the detected beats. (At this time, the judgment threshold reaches 0.7).

[0127] Furthermore, the expression of the drive signal is:

[0128]

[0129] where P[i] is a discrete sequence and satisfies P[i] ∈ {0, 1}; S(t) is the drive signal output by the control unit, δ(t - t i ) is the Dirac impulse function, corresponding to the time signal in the beat signal, and A mode is the information of the backbeat mode. In the specific implementation process, A mode is used to adjust the beating pressure of the beater on the user; the discrete sequence formed by P[i] is the most basic expression method of the beat signal.

[0130] Those skilled in the art should understand that the embodiments of the present invention can provide methods, systems, or computer program products. Therefore, the present invention can be implemented in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can be implemented in the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.

[0131] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems) and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device create means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

[0132] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the functions specified in one or more blocks.

Claims

1. A beating machine control method based on audio signal, characterized in that: The following steps are included: S101, performing denoising processing on an audio signal obtained through a database or external audio sampling to obtain an original signal, and reconstructing the original signal to obtain a base signal; S201, expanding the base signal in the frequency domain by changing it in time and frequency and calculating the spectrum energy of the expanded base signal in each time window to obtain a short-time energy spectrum of the base signal; S301, extracting position signals from the short-time energy spectrum through peak detection, detecting the energy peak of each position signal and comparing and judging the energy peak at each position according to a preset judgment threshold to obtain a time signal corresponding to each beat; S401, constructing a beat signal according to the time signal corresponding to each beat and outputting the beat signal to a control unit of a beating machine; The control unit converts the beat signal into a driving signal, and the driving signal drives the tapping machine to tap the user.

2. A method for controlling a beating machine based on an audio signal according to claim 1, characterized in that: In the step S101, the audio signal is split layer by layer by discrete wavelet transform, and the approximate coefficients of the audio signal of each layer after the split and the detail coefficients of the audio signal of each layer are obtained by layer-by-layer calculation; The detail coefficients in each layer of the audio signal are subjected to denoising, and the original signal is obtained by combining the detail coefficients in each layer of the audio signal after denoising and the approximate coefficients of each layer of the audio signal after splitting.

3. A method for controlling a beating machine based on an audio signal according to claim 2, characterized in that: The calculation method of the approximate coefficients in the audio signal of each layer is: Wherein, 2i and 2i+1 are used to represent the sample pairs processed in the audio signal represented by X[n]; n and i are sample indexes in the audio signal, expressing the position of each sample; c j [i] represents the approximation coefficient at the jth level; The calculation method of the detail coefficient is: Among them, d j [i] represents the detail coefficient at the jth layer, indicating the high-frequency characteristics of the audio signal at the current jth level.

4. A method for controlling a beating machine based on an audio signal according to claim 2, characterized in that: The denoising method of the detail coefficient is: retaining the part of the detail coefficient of the audio signal of each layer whose value is greater than or equal to the first threshold, and removing the part of the detail coefficient of the audio signal of each layer whose value is less than the first threshold.

5. The method for controlling a beating machine based on an audio signal according to claim 2, characterized in that: The original signal is reconstructed by inverse wavelet transform to obtain the base signal. The inverse wavelet transform is expressed as: Wherein, x'[n] is the base signal after denoising, L is the number of layers of wavelet transform, c j is the approximation coefficient of the jth layer, d j is the detail coefficient of the j-th layer.

6. The method for controlling a beating machine based on an audio signal according to claim 2, characterized in that: The type of the wavelet is Haar wavelet.

7. The method for controlling a beating machine based on an audio signal according to claim 1, characterized in that: In step S101, the base signal is expanded by the change of short-time Fourier transform in time and frequency. The expression form of the short-time Fourier transform is: Among them, Z xx [f, t] represents the frequency domain representation of the base signal at frequency f and time t; x[m] represents the value of the base signal at time m; w[mn] represents the value of the window function at time mn; e[mn] represents the value of the window function at time mn; -j2πfm It is used to complete the transformation of the base signal from the time domain to the frequency domain, and f is the frequency.

8. The method for controlling a beating machine based on an audio signal according to claim 1, characterized in that: Calculating the spectrum energy of each time window based on a spectrum energy calculation method; The expression of the spectrum energy calculation method is: E[n]=∑|Z xx [f,n]| 2 Wherein, E[n] is the spectrum energy value of the base signal on sample n; Z xx [f,n] ​​represents the frequency domain representation of the base signal at frequency f and time n.

9. The method for controlling a beating machine based on an audio signal according to claim 1, characterized in that: In step S301, the value range of the judgment threshold is 0.5-0.

7.

10. The method for controlling a beating machine based on an audio signal according to claim 1, characterized in that: The expression of the driving signal is: Wherein, P[i] is a discrete sequence and satisfies P[i]∈{0,1}; S(t) is the driving signal output by the control unit, δ(tt i ) is a Dirac pulse function corresponding to the time signal in the beat signal, A mode It is the back pat mode information.