Auxiliary teaching equipment for singing practice

The singing practice equipment, which combines a multi-microphone array and an intelligent evaluation module with a deep learning algorithm, solves the problems of pitch deviation and missed singing details in traditional teaching, realizes precise and personalized singing practice assistance, and improves singing skills and stage performance.

CN120748264APending Publication Date: 2025-10-03BIJIE PRESCHOOL TEACHERS COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510593682.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-09
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

Traditional singing teaching methods cannot achieve precision, personalization and intelligence. Vocal teachers find it difficult to maintain highly concentrated auditory attention for a long time, which leads to missing singing details. There is also a lack of comprehensive recording and in-depth analysis of students' singing process, making it impossible to discover the potential patterns of singing habits and skill development.

Method used

It uses a multi-microphone array, audio processing module, intelligent evaluation module, frequency adjustment module, direction simulation module, feedback output module, interactive operation module and data storage and update module, combined with deep learning algorithms and acoustic measurement technology to achieve accurate collection, analysis and feedback of the practitioner's voice.

Benefits of technology

It provides multi-form and intuitive singing feedback to help practitioners quickly adjust pitch and control sound direction, improve singing skills and stage performance, and realize personalized training plans and systematic singing data recording and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748264A_ABST
    Figure CN120748264A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of music education auxiliary equipment, and discloses auxiliary teaching equipment for singing practice, and the equipment comprises a multi-microphone array which is composed of a specific number of microphones which are arranged according to a preset space layout, and is used for collecting a sound signal of a practicer and transmitting the sound signal to an audio processing module; the audio processing module is connected with the multi-microphone array and has the functions of preprocessing audio signals, performing sound positioning calculation and performing frequency analysis processing; and the intelligent evaluation module is connected with the audio processing module, stores a large amount of singing sample data and obtains an evaluation model through deep learning algorithm training. And accurate adjustment is performed through a sound frequency adjustment unit by adopting a digital frequency synthesis and adaptive filtering hybrid algorithm. According to the algorithm, an appropriate sine wave signal is generated according to the deviation between the standard pitch and the actual singing pitch, and the adjustment process is continuously optimized through the adaptive filter according to the error signal, so that the frequency of the output signal is close to the standard pitch frequency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of music education auxiliary equipment, and in particular to an auxiliary teaching equipment for singing practice. Background Art

[0002] In the field of singing instruction, traditional teaching methods rely primarily on the vocal teacher's auditory judgment and experience. Teachers listen to students singing in class, using their professional expertise to point out problems with pitch, rhythm, and timbre, and provide demonstrations and corrections. However, this approach has many limitations.

[0003] On the one hand, vocal teachers have limited energy and struggle to maintain a high level of auditory focus throughout long teaching sessions, potentially overlooking subtle singing details. Furthermore, teachers' subjective judgments of subtle pitch deviations or rhythmic changes can be subject to error. When determining pitch deviations, teachers can usually only roughly perceive whether they are too high or too low, rather than accurately determining the exact cent value.

[0004] On the other hand, traditional teaching methods lack comprehensive documentation and in-depth analysis of students' singing processes. Teachers can only provide immediate feedback to students during their performances, but are unable to systematically review and study their long-term performance, making it difficult to identify potential patterns and problems in students' singing habits and skill development.

[0005] With the continuous advancement of science and technology, digital audio technology has gradually been applied to music education. Simple audio recording software and basic pitch-checking tools have begun to appear. These tools can, to a certain extent, help students record their performances and conduct preliminary pitch analysis. However, these tools have significant shortcomings in areas such as sound localization simulation, intelligent feedback, and personalized teaching. For example, they are unable to simulate the sound direction effects in different performance environments, nor can they provide personalized training plans and targeted practice repertoire recommendations based on students' singing characteristics. Summary of the Invention

[0006] In response to the shortcomings of the existing technology, the present invention provides an auxiliary teaching device for singing practice, which solves the limitations of traditional singing teaching methods and the problem that it cannot meet the needs of precise, personalized and intelligent teaching.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: an auxiliary teaching device for singing practice, comprising:

[0008] A multi-microphone array, which consists of a specific number of microphones arranged in a predetermined spatial layout, is used to collect the practitioner's voice signal and transmit it to the audio processing module;

[0009] The audio processing module is connected to a multi-microphone array and has the functions of pre-processing audio signals, calculating sound positioning, and performing frequency analysis and processing;

[0010] The intelligent evaluation module is connected to the audio processing module. It stores a large amount of singing sample data and uses a deep learning algorithm to train an evaluation model. This model conducts a multi-dimensional evaluation of the singing based on various characteristics of the audio signal and generates feedback information.

[0011] The frequency adjustment module adjusts the frequency when it detects that the pitch sung by the practitioner deviates from the standard pitch;

[0012] Direction simulation module, based on the vector-based amplitude shift algorithm, changes the direction of sound, simulating the effect of sound coming from different directions according to teaching needs;

[0013] A feedback output module receives feedback information generated by the intelligent evaluation module, displays visual data on a display screen, and informs the practitioner through a voice broadcast device;

[0014] Interactive operation module, where practitioners can select practice tracks, view performance evaluation reports, and adjust equipment parameters through the touch screen of the interactive operation module;

[0015] The data storage and update module is used to store the practitioner's singing history data.

[0016] Preferably, the audio processing module includes:

[0017] An audio acquisition unit, receiving audio signals collected by a multi-microphone array;

[0018] Pre-processing unit, which removes noise and interference and converts the sampling rate;

[0019] Sound localization unit, which determines the direction of sound based on an algorithm;

[0020] The frequency analysis processing unit extracts the fundamental frequency and harmonic frequencies of the audio signal.

[0021] Preferably, the intelligent evaluation module includes:

[0022] The model building unit uses a convolutional neural network and a bidirectional long short-term memory network to build an evaluation model;

[0023] Model training unit, validation set loss function value, training is stopped when the validation set loss function value no longer decreases in consecutive training cycles;

[0024] The model prediction unit extracts and converts the features extracted from the input audio signal to perform a multi-dimensional evaluation of the practitioner's singing.

[0025] Preferably, the frequency adjustment module includes:

[0026] A frequency deviation calculation unit compares the deviation between the standard pitch frequency and the actual frequency sung by the practitioner;

[0027] The frequency adjustment unit uses a hybrid algorithm of digital frequency synthesis and adaptive filtering to adjust the frequency of the audio signal.

[0028] Preferably, the direction simulation module includes:

[0029] A direction calibration unit, which calibrates the direction angle of the sound channel relative to the listener based on acoustic measurement methods;

[0030] The adjustment calculation unit calculates the amplitude adjustment factor of the sound signal of each channel based on the vector basis amplitude shift algorithm.

[0031] Preferably, the feedback output module includes:

[0032] A signal display unit generates a time-fundamental frequency curve in the form of image drawing;

[0033] The information display unit uses different colors and sizes of fonts to display pitch and rhythm errors;

[0034] The voice broadcast unit uses speech synthesis technology to inform the practitioners of feedback information in the form of voice.

[0035] Preferably, the interactive operation module includes:

[0036] Interactive unit, select practice tracks and view singing evaluation reports through touch screen or voice commands.

[0037] Preferably, the data storage and update module includes:

[0038] A data compression unit compresses audio data using an adaptive compression algorithm based on audio signal characteristics;

[0039] The data update unit cleans the new sample data and merges it with the original samples in the repository for storage.

[0040] Preferably, the interactive operation module further comprises: a voice recognition unit, which can listen to the practitioner's voice instructions to adjust the equipment parameters.

[0041] The present invention provides an auxiliary teaching device for singing practice, which has the following beneficial effects:

[0042] 1. When a pitch deviation is detected in the present invention, the sound frequency adjustment unit uses a hybrid algorithm of digital frequency synthesis and adaptive filtering to perform precise adjustments. This algorithm can not only generate a suitable sinusoidal wave signal based on the deviation between the standard pitch and the actual singing pitch, but also continuously optimize the adjustment process according to the error signal through an adaptive filter to ensure that the frequency of the output signal is close to the standard pitch frequency. Practitioners can hear the sound of the correct pitch in real time during the singing process. By comparing it with their actual singing voice, they can perceive the pitch difference more intuitively, thereby quickly adjusting the vocalization method and improving the pitch perception and control ability. Long-term use can effectively correct pitch problems and significantly reduce the incidence of pitch deviation. It plays an extremely critical auxiliary role in pitch training and helps practitioners establish correct pitch concepts and singing habits.

[0043] 2. The present invention uses a sound direction simulation unit to accurately calibrate the direction angle of the sound channel relative to the audience based on the vector basis amplitude translation algorithm, and calculates the amplitude adjustment factor of the sound signal of each channel according to the sound localization result and the preset target direction. It can simulate the effect of sound coming from different directions, and the error is controlled within a very small range. This makes the practitioner feel as if they are in a real singing scene, and better feel the propagation and reflection rules of sound in space, which helps to more accurately control the sense of direction and space of sound during singing. In stage performance training, practitioners can adapt to the changes in sound in different stage layouts and sound environments in advance, improve the authenticity and appeal of the performance, enhance stage expression, and make the singing more vivid and layered.

[0044] 3. The present invention provides practitioners with multi-format, intuitive feedback information through a high-definition touch screen and voice broadcast device. The display screen displays high-resolution visual data of the performance, such as pitch curves, rhythm waveforms, and spectrograms. Real-time rendering technology ensures smooth and timely graphical display. Important feedback information is displayed in different colors, font sizes, and flashing prompts, allowing practitioners to quickly focus on key issues. Voice broadcasts inform practitioners of evaluation results and improvement suggestions in a clear and natural voice. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 A three-dimensional diagram of the auxiliary teaching device of the present invention;

[0046] Figure 2 This is a diagram of the architecture of the auxiliary teaching equipment in the present invention. DETAILED DESCRIPTION

[0047] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0048] Please see the attached Figure 1 -Attached Figure 2 The embodiment of the present invention provides an auxiliary teaching device for singing practice, comprising:

[0049] 1. Multi-microphone array

[0050] In this embodiment, the multi-microphone array adopts a three-dimensional layout, consisting of 8 high-sensitivity microphones arranged in a regular octahedron, with the distance between adjacent microphones ranging from 0.8 to 1.2 meters. This layout can collect the practitioner's voice signals in all directions, avoiding the problem of incomplete or distorted sound collection caused by the single microphone position. The frequency response range of each microphone is 15Hz-22kHz, and the signal-to-noise ratio is not less than 70dB. It can accurately capture various audio details when the practitioner sings, and accurately restore everything from low bass to high treble. When the equipment is installed, each microphone needs to be calibrated using professional acoustic measurement instruments to ensure the consistency and accuracy of its frequency response. For example, a standard audio signal generator is placed in the center of the equipment, and a single-frequency signal from 15Hz to 22kHz is sent to each microphone in turn. The response data of each microphone is recorded, and fine-tuning is performed based on this data to control the response differences of each microphone within a very small range.

[0051] 2. Audio processing module

[0052] In this embodiment, the audio processing module is connected to a multi-microphone array and has the functions of pre-processing the audio signal, calculating sound localization, and performing frequency analysis and processing, as follows:

[0053] 2.1 Audio Acquisition Unit

[0054] It is electrically connected to the multi-microphone array and is responsible for receiving the audio signals collected by the multi-microphone array.

[0055] 2.2 Preprocessing Unit

[0056] For the band-pass filter in the pre-processing, it obtains the actual frequency response curve by performing a full-band frequency sweep test on the microphone, and then accurately calculates the inductance and capacitance parameters based on the least squares method to fit the difference between the ideal frequency response curve and the actual frequency response curve, so that the band-pass filter attenuates the 20Hz-20kHz audio signal by less than 3dB. The specific operation process is to first connect the audio signal generator to the input port of the audio processing unit, set the signal generator output to start from 15Hz and increase the sweep signal to 22kHz in steps of 1Hz. The audio processing unit records the signal amplitude collected by the microphone at each frequency point to obtain the actual frequency response curve. Then, based on the ideal frequency response curve, the least squares method is used to calculate the inductance and capacitance parameters required for the band-pass filter to compensate for the frequency response deviation of the microphone. This can effectively remove low-frequency noise and high-frequency interference in the audio signal and improve the purity of the audio signal.

[0057] In terms of sampling rate conversion, a polyphase filter structure is used according to a specific conversion formula:

[0058]

[0059] Among them, y[n] is the converted sampling point, h k [n] is the polyphase filter coefficient, x k [n] is a subsequence of the original audio signal, and M is the number of phases of the polyphase filter. The audio signal is uniformly converted to a sampling rate of 44.1kHz to meet the requirements of subsequent audio analysis and processing, ensuring the compatibility and accuracy of audio data in different processing links. When implementing the polyphase filter, the number of phases M of the polyphase filter is first determined according to the requirements of the sampling rate conversion. In this embodiment, M=4. Then, based on the selected filter type and design indicators, the coefficient h of the polyphase filter is calculated. k [n]. In the actual conversion process, the input audio signal is decomposed into M subsequences, each subsequence corresponds to a phase, and then multiplied by the corresponding filter coefficients and summed to obtain the converted sampling point y[n].

[0060] 2.3 Sound Localization Unit

[0061] In this embodiment, the sound localization calculation is based on the generalized cross-correlation algorithm combined with the arrival time difference estimation. The cross power spectrum density is weighted using the PHAT weighting function. The weighted generalized cross-correlation function is:

[0062]

[0063] It can enhance the accuracy of sound localization in low signal-to-noise ratio environments. When calculating the sound direction angle, a preliminary direction angle θ0 is first calculated based on the ideal geometry of the microphone array. Then, by placing multiple calibration sound sources at known locations within the device's installation environment, the actual time difference between the sound reaching the microphones is measured and compared with the theoretical value. The angular deviation Δθ caused by installation error is calculated. The final direction angle calculation formula is θ = θ0 + Δθ. In practice, when the calibration sound sources are placed at different locations (e.g., 10 evenly distributed points on a sphere with a radius of 1 meter centered on the device), they are sequentially triggered to emit pulse signals of a specific frequency (e.g., 1kHz). The audio processing unit records the time each microphone receives the signal. Based on the geometry of the microphone array and the known speed of sound (343 m / s), the theoretical time difference between the sound reaching each microphone is calculated. This is compared with the actual measured time difference, and the angular deviation Δθ caused by installation error is calculated based on the geometric relationship. For example, if the theoretical time difference between the sound reaching two microphones is t1 and the actual measured time difference is t2, the magnitude of the angular deviation Δθ can be calculated based on the microphone spacing and the speed of sound. In this way, the direction of the source of the practitioner's voice can be accurately determined, providing accurate basic data for the subsequent sound direction simulation function.

[0064] 2.4 Frequency Analysis Processing Unit

[0065] In this embodiment, the frequency analysis processing unit uses a fast Fourier transform (FFT) to convert the time-domain audio signal into a frequency-domain signal, enabling in-depth analysis of the audio signal's spectral characteristics. Properly setting the frame length and frame shift ensures analysis accuracy while improving processing efficiency. By extracting harmonic frequencies using a search algorithm based on energy thresholds and frequency interval constraints, a comprehensive assessment of the pitch and timbre of a performer's singing can be made. For example, when analyzing a single note sung by a performer, the frequency domain signal is obtained through a fast Fourier transform. Then, based on the set energy threshold and frequency interval constraints, the fundamental frequency and harmonic frequencies of the note are accurately extracted. If the performer's singing pitch is accurate, the extracted fundamental frequency should be close to the standard pitch, and the proportional relationship between the harmonic frequencies should also meet the requirements of music theory. By analyzing these frequency parameters, the performer's pitch, including the richness and distribution of harmonics, as well as the characteristics of the timbre, can be determined.

[0066] 3. Intelligent evaluation module

[0067] In this embodiment, the intelligent evaluation module uses a large amount of stored singing sample data covering different styles and levels and an evaluation model that combines a convolutional neural network with a bidirectional long short-term memory network to comprehensively evaluate the practitioner's singing.

[0068] First, during the model training phase, the different convolutional layers of the convolutional neural network perform multi-level feature extraction on the input singing sample data using convolution kernels of different sizes and strides. For example, smaller convolution kernels (such as 3×3) can extract local detailed features, while larger convolution kernels (such as 7×7) can extract more macroscopic features. The batch normalization layer after each convolution layer normalizes the data, making the data distribution more stable across different layers, which is conducive to model training and convergence. The ReLU activation function layer introduces nonlinear characteristics to enhance the model's expressive power. The bidirectional long short-term memory network can process the temporal information of audio data and perform a comprehensive analysis of the audio features of the preceding and following time series. Through training on a large amount of sample data, the model continuously adjusts the parameters of the convolutional neural network and the bidirectional long short-term memory network to improve the prediction accuracy of singing evaluation indicators (such as pitch deviation, rhythm error, timbre score, and emotional expression intensity).

[0069] When evaluating a performer's singing, the model first extracts and transforms the audio signal features after inputting them. The model then passes them through the various layers of a convolutional neural network (CNN). For example, the audio signal's spectral feature maps are convolved through a convolutional layer to produce a series of new feature maps. These feature maps are then batch normalized and processed with an activation function before being passed to the next convolutional layer. After processing by the CNN, the resulting feature vectors are fed into a bidirectional long-short-term memory (LSTM) network. This network further processes the feature vectors based on the audio's temporal information, analyzing, for example, the duration relationships between notes and trends in pitch variation. Finally, the output feature vectors from the BLSTM network are fed into a fully connected layer, which uses a softmax function to convert the outputs into predicted probability distributions for each category, thereby determining the performer's singing performance in terms of pitch, rhythm, timbre, and emotional expression. For example, if the model predicts a significant deviation in pitch, it could be because the extracted fundamental frequency differs significantly from the standard pitch frequency. A significant rhythm error could be due to the duration ratio between notes not meeting standard rhythm requirements. A low timbre score could be due to an unsatisfactory harmonic structure or noise interference. A lack of emotional expression could be due to subtle variations in the dynamics or speed of the performance. In this way, the intelligent assessment module can comprehensively, objectively, and accurately assess a practitioner's singing proficiency and provide detailed, targeted feedback.

[0070] 4. Frequency adjustment module

[0071] In this embodiment, when it is detected that the pitch sung by the practitioner deviates from the standard pitch, the frequency adjustment is performed as follows:

[0072] When the audio processing unit detects a deviation between the pitch of the vocalist's singing and the standard pitch, the sound frequency adjustment unit activates. First, it calculates the deviation between the standard pitch frequency and the actual fundamental frequency. Then, based on this deviation, a high-precision digital signal generator generates a sine wave signal of the corresponding frequency. This sine wave signal serves as the basis for adjusting the audio signal frequency. For example, if the vocalist's singing pitch is low—that is, the actual fundamental frequency is lower than the standard pitch frequency—the generated sine wave signal frequency will be the difference between the two, with a phase resolution of 0.1°, ensuring high accuracy.

[0073] Next, the original audio signal and the generated sine wave signal are mixed through an adaptive filter. The adaptive filter works by continuously adjusting the filter coefficients based on an error signal, so that the output signal is as close as possible to the desired signal (i.e., an audio signal with a standard pitch). This error signal is calculated using an optimization algorithm based on the minimum mean square error criterion. The error signal is first calculated based on the original error between the desired signal and the current output signal. The weighting coefficients are then dynamically adjusted based on the frequency characteristics of the audio signal to produce the final error signal. Using this error signal, the adaptive filter continuously adjusts its coefficients according to a specific update formula, thereby adjusting the frequency of the original audio signal. For example, if the pitch of the current output signal is still low during the adjustment process, the error signal will prompt the adaptive filter to increase the weighting of the sine wave signal, gradually bringing the output signal's pitch closer to the standard pitch until it reaches or approaches the standard pitch. This allows the practitioner to hear the correct pitch in real time, thereby better perceiving and adjusting their singing pitch.

[0074] 5. Direction simulation module

[0075] In this embodiment, the direction simulation module changes the sound direction based on the vector-based amplitude shift algorithm, simulating the effect of sound coming from different directions according to teaching needs, as follows:

[0076] 5.1 Direction Calibration Unit

[0077] In this embodiment, the direction calibration unit calculates the amplitude adjustment factor of the sound signal of each channel in a 7.1-channel audio system based on the original direction and the preset target direction determined by the sound localization unit using a vector-based amplitude shifting algorithm:

[0078]

[0079] Among them, β k is the direction angle of channel k relative to the listener, θ is the original direction angle, θ t is the target direction angle.

[0080] In calculating the amplitude adjustment factor a of the sound signal of each channelk When the direction angle β of the sound channel relative to the listener is k Calibration is performed using a method based on acoustic measurement. Specifically, when the device is installed, a standard sound source is used to emit sound at no less than 20 evenly distributed points on a spherical surface centered on the device. The actual response of each channel is measured at each sounding point and compared with the theoretical response. The deviation value δ of the directional angle is calculated. k , the actual direction angle in use:

[0081]

[0082] The calculation error of the amplitude adjustment factor is within ±0.05.

[0083] During the actual calibration process, for example, a standard sound source is placed at 25 evenly distributed points on a spherical surface with a radius of 2 meters, and the standard sound source is triggered in turn to emit a stable audio signal of a specific frequency (such as 500Hz). Each channel of the 7.1-channel audio system is connected to a highly sensitive sound sensor to collect the intensity of the audio signal received by the channel. When the standard sound source emits sound at a certain point, the actual signal intensity received by each channel is recorded, and the theoretical direction angle of each channel relative to the sound point is calculated based on the theoretical model of audio signal propagation in space. By comparing the difference between the actual signal strength and the theoretical signal strength, the deviation value δ of the direction angle is calculated using the geometric acoustic principle and the least squares optimization algorithm. k .

[0084] Assume that at a certain sound point, the theoretical direction angle of the left front channel is According to theoretical calculation, the signal strength received at this direction angle should be The actual measured signal strength is The directional angle deviation value v1 of the sound channel is calculated by using the least squares optimization algorithm through the measurement data of a series of sound emission points. Then, the directional angle of the sound channel in actual use is: β1 = 30° + (-2°) = 28°.

[0085] 5.2. Adjusting the calculation unit

[0086] After determining the exact direction angle β k Then, according to the original direction angle θ obtained by the sound positioning unit and the preset target direction angle θ t , calculate the amplitude adjustment factor a of the sound signal of each channel k For example, when the original direction angle θ = 45°, the target direction angle θ t =60°, for a certain channel k, its direction angle β relative to the listener k =35°, then:

[0087]

[0088] By adjusting the sound signal amplitude of each channel according to the corresponding adjustment factor a k Adjustments are made to simulate a change in the sound's direction from the original to the target direction. This precise directional simulation algorithm simulates the effect of sound arriving from different directions, allowing practitioners to better understand the propagation and reflection patterns of sound in space. This helps to better control the sound's sense of direction and space during singing, enhancing the authenticity and appeal of stage performances. For example, in chorus practice, practitioners can more accurately grasp the spatial relationship between their own voices and those of other members, making the chorus more harmonious and three-dimensional.

[0089] 6. Feedback output module

[0090] In this embodiment, the feedback output module receives the feedback information generated by the intelligent evaluation module, displays the visual data on the display screen, and notifies the practitioner through the voice broadcast device, as follows:

[0091] 6.1 Signal Display Unit

[0092] In this embodiment, the signal display unit generates a time-base frequency curve in the form of image drawing, as follows:

[0093] When drawing the pitch curve, first obtain the fundamental frequency data of each frame of audio signal from the audio processing unit, use the time axis as the horizontal axis and the fundamental frequency value as the vertical axis. For example, obtain the fundamental frequency of one frame of audio signal every 10 milliseconds, and connect these fundamental frequency data in sequence over time to form a pitch curve. For the rhythm waveform diagram, according to the beat information of the audio signal, the starting point of the beat is marked as a peak or trough, and the rhythm changes are displayed by drawing a continuous waveform. The spectrum diagram is drawn after performing spectral analysis on the audio signal, with the frequency as the horizontal axis and the amplitude of the frequency component as the vertical axis. Different frequency ranges are represented by different colors or grayscales to intuitively display the frequency distribution characteristics of the audio signal.

[0094] 6.2 Information display unit

[0095] In this embodiment, the information display unit uses different colors and sizes of fonts to display pitch and rhythm errors, as follows:

[0096] When displaying feedback, different colors and font sizes are used to distinguish different types of information. Important information, such as severe pitch deviation and rhythm errors, flashes at a frequency of 2-3Hz, in high-contrast red or orange, and with a font size 1.5-2 times the normal font size. This allows practitioners to focus on key issues at a glance. For example, if the pitch deviation exceeds 50 cents, the display will display "Severe pitch deviation: 50 cents" in large, flashing red text. Simultaneously, the larger deviation portion of the pitch curve will be highlighted with a bold red line segment.

[0097] 6.3. Voice broadcast unit

[0098] In this embodiment, the voice broadcast unit uses speech synthesis technology to inform the practitioner of the feedback information in the form of speech, as follows:

[0099] The voice broadcast device uses clear and natural speech synthesis technology to provide feedback to the practitioner in the form of voice. For example, "During your performance, your pitch deviation was large, with an average deviation of 30 cents. Please refer to the standard pitch for adjustment. The system has adjusted your voice frequency for comparison." The speech synthesis technology first analyzes the feedback text to determine the voice's intonation, speed, and stress parameters. Then, based on a preset voice library, the text is converted into a voice signal for output. This multi-format feedback output method allows practitioners to gain a more intuitive and comprehensive understanding of their performance, allowing them to adjust their singing style in a timely manner and improve learning efficiency.

[0100] 7. Interactive operation module

[0101] In this embodiment, the practitioner can select practice tracks, view performance evaluation reports, and adjust equipment parameters through the touch screen of the interactive operation module;

[0102] 7.1 Interaction Unit

[0103] In this embodiment, the interactive unit features a touchscreen and voice command recognition capabilities. Practitioners can use the touchscreen or voice commands to select practice tracks, view performance evaluation reports, and adjust device parameters. Voice command recognition utilizes a deep learning-based speech recognition model, trained with a dataset containing at least 10,000 voice commands related to singing practice.

[0104] 7.2 Speech Recognition Unit

[0105] In this embodiment, the speech recognition unit utilizes an architecture that combines a convolutional neural network and a recurrent neural network. The convolutional neural network extracts the spectral features of the speech signal, while the recurrent neural network processes the temporal information of the speech signal. During training, the speech command dataset is first preprocessed, including sampling rate unification, noise reduction, and framing. The preprocessed speech signal is then input into the convolutional neural network. The first convolution kernel of the convolutional neural network has a size of 5*5 and a step size of 2, which is used to extract the local spectral features of the speech signal. After passing through several convolutional and pooling layers, a feature map sequence is generated. This feature map sequence is then input into the recurrent neural network, which further processes the feature map sequence based on the temporal characteristics of the speech signal to learn the semantic information of the speech command. The model training uses a contrastive loss function to ensure that the feature vectors of different speech commands have sufficient discrimination in the feature space. The speech command recognition accuracy rate exceeds 95%.

[0106] This convenient interactive operation allows practitioners to use the device more freely and flexibly, eliminating the need for tedious manual operations. This improves the convenience and efficiency of practice and allows practitioners to focus more on the singing itself. For example, during practice, practitioners do not need to stop singing to manually operate the device. Simply say "pause practice" and the device will immediately respond, pausing audio playback and analysis, and awaiting further instructions from the practitioner.

[0107] 8. Data storage and update module

[0108] In this embodiment, the data storage and update module is used to store the practitioner's singing history data. Specifically, the data storage and update unit is used to store the practitioner's singing history data, including audio recordings of each practice session, evaluation results, and practice time information, with a storage capacity of no less than 1TB. The data storage utilizes an efficient database management system to compress and store the audio data. The compression algorithm utilizes an adaptive compression algorithm based on audio signal characteristics, dynamically adjusting the compression ratio based on the frequency content and energy distribution of the audio signal, thereby improving storage efficiency while ensuring audio quality.

[0109] 8.1 Data Compression Unit

[0110] When storing audio data, the audio signal is first analyzed to determine its primary frequency components and energy distribution. For example, for an audio signal primarily composed of mid- and high-frequency components with concentrated energy, a higher compression ratio is used to compress the low-frequency portion, while a relatively lower compression ratio is used for the mid- and high-frequency components to preserve the key audio information. This adaptive compression algorithm can reduce audio data storage space requirements by 30%-50%.

[0111] When updating the singing sample database, the new sample data is first cleaned and preprocessed to remove noise and labeling errors, and then merged and stored with the samples in the original database. For updating the evaluation model, the new sample data is first used to calculate the loss function value of the model on these samples, and then, based on the gradient information of the loss function value, only some parameters in the model related to the new samples are adjusted. For example, if the new samples mainly affect the model performance of the pitch assessment part, then only the parameters of the neural network layer related to the pitch assessment are updated, rather than retraining the entire model. In this way, while ensuring the continuous optimization of the model, the computing resources and time required for the update are greatly reduced, so that the equipment can adapt to new singing styles and techniques in a timely manner, and provide practitioners with more accurate and up-to-date teaching assistance services.

[0112] 8.2 Data Update Unit

[0113] The data storage and update module regularly updates the singing sample database and evaluation model via a network connection, at least once a month. The updated data is sourced from high-quality singing samples annotated by professional vocal teachers and model optimization data. The update process utilizes an incremental learning algorithm, training and adjusting model parameters only on newly added samples, avoiding retraining the entire dataset and improving update efficiency.

[0114] To sum up, this auxiliary teaching equipment for singing practice provides practitioners with a comprehensive, accurate, efficient and interactive singing practice auxiliary environment through the collaborative work of various components, using advanced audio processing technology, intelligent evaluation models and diversified functional modules. It can effectively help practitioners improve their singing skills, pitch, rhythm control ability and expression of emotions in songs. It has great practical value and driving effect for both the training of professional singers and the learning of ordinary singing enthusiasts.

[0115] Working Principle: This auxiliary teaching device for singing practice works collaboratively through a multi-microphone array, an audio processing module, an intelligent evaluation module, a frequency adjustment module, a direction simulation module, a feedback output module, an interactive operation module, and a data storage and update module. The multi-microphone array utilizes eight high-sensitivity microphones arranged in a three-dimensional octahedron configuration. Calibrated to ensure consistent frequency response, it accurately captures the practitioner's voice signals from all directions. The audio acquisition unit of the audio processing module receives the multi-microphone array audio signal. The pre-processing unit uses a bandpass filter to filter out noise interference and convert the sampling rate using frequency sweep testing and least squares parameter calculation. The sound localization unit determines the sound direction angle based on a generalized cross-correlation algorithm combined with time difference of arrival estimation. The frequency analysis processing unit uses fast Fourier transform and related algorithms to extract harmonic frequencies and evaluate singing performance. The intelligent evaluation module utilizes a large amount of singing sample data, a convolutional neural network, and a bidirectional long short-term memory network to train an evaluation model. After training, it comprehensively evaluates the practitioner's performance and provides multi-dimensional feedback. When a pitch deviation is detected, the frequency adjustment module generates a sine wave signal using a high-precision digital signal generator. This signal is then mixed with the original audio signal through an adaptive filter to adjust the frequency. The Direction Simulation Module's Direction Calibration Unit calibrates the vocal channel directional angles using a vector-based amplitude translation algorithm. The Adjustment Calculation Unit then calculates the amplitude adjustment factors for each channel's sound signal to simulate changes in sound direction. The Feedback Output Module's Signal Display Unit plots a time-frequency curve to visualize data. The Information Display Unit highlights singing errors in various ways, and the Voice Announcement Unit uses speech synthesis technology to provide feedback to the practitioner. The Interactive Operation Module's Interaction Unit facilitates user operation through a touchscreen and deep learning-based voice command recognition with over 95% accuracy. The Voice Recognition Unit utilizes a combined convolutional neural network and recurrent neural network architecture to extract speech spectral features and process timing information. The Data Compression Unit in the Data Storage and Update Module utilizes an adaptive compression algorithm to reduce audio data storage space requirements. The Data Update Unit uses an incremental learning algorithm to update the singing sample database and evaluation model at least monthly using data annotated by professional vocal teachers. The Data Update Unit stores and manages historical singing data using an efficient database management system, providing comprehensive, accurate, and convenient singing practice assistance services.

[0116] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. An auxiliary teaching device for singing practice, characterized in that: include: A multi-microphone array, which consists of a specific number of microphones arranged in a predetermined spatial layout, is used to collect the practitioner's voice signal and transmit it to the audio processing module; The audio processing module is connected to a multi-microphone array and has the functions of pre-processing audio signals, calculating sound positioning, and performing frequency analysis and processing; The intelligent evaluation module is connected to the audio processing module. It stores a large amount of singing sample data and uses a deep learning algorithm to train an evaluation model. This model conducts a multi-dimensional evaluation of the singing based on various characteristics of the audio signal and generates feedback information. The frequency adjustment module adjusts the frequency when it detects that the pitch sung by the practitioner deviates from the standard pitch; Direction simulation module, based on the vector-based amplitude shift algorithm, changes the direction of sound, simulating the effect of sound coming from different directions according to teaching needs; A feedback output module receives feedback information generated by the intelligent evaluation module, displays visual data on a display screen, and informs the practitioner through a voice broadcast device; Interactive operation module, where practitioners can select practice tracks, view performance evaluation reports, and adjust equipment parameters through the touch screen of the interactive operation module; The data storage and update module is used to store the practitioner's singing history data.

2. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The audio processing module includes: An audio acquisition unit, receiving audio signals collected by a multi-microphone array; Pre-processing unit, which removes noise and interference and converts the sampling rate; Sound localization unit, which determines the direction of sound based on an algorithm; The frequency analysis processing unit extracts the fundamental frequency and harmonic frequencies of the audio signal.

3. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The intelligent evaluation module includes: The model building unit uses a convolutional neural network and a bidirectional long short-term memory network to build an evaluation model; Model training unit, validation set loss function value, training is stopped when the validation set loss function value no longer decreases in consecutive training cycles; The model prediction unit extracts and converts the features extracted from the input audio signal to perform a multi-dimensional evaluation of the practitioner's singing.

4. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The frequency adjustment module includes: A frequency deviation calculation unit compares the deviation between the standard pitch frequency and the actual frequency sung by the practitioner; The frequency adjustment unit uses a hybrid algorithm of digital frequency synthesis and adaptive filtering to adjust the frequency of the audio signal.

5. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The direction simulation module includes: A direction calibration unit, which calibrates the direction angle of the sound channel relative to the listener based on acoustic measurement methods; The adjustment calculation unit calculates the amplitude adjustment factor of the sound signal of each channel based on the vector basis amplitude shift algorithm.

6. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The feedback output module includes: A signal display unit generates a time-fundamental frequency curve in the form of image drawing; The information display unit uses different colors and sizes of fonts to display pitch and rhythm errors; The voice broadcast unit uses speech synthesis technology to inform the practitioners of feedback information in the form of voice.

7. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The interactive operation module includes: Interactive unit, select practice tracks and view performance evaluation reports through the touch screen.

8. The auxiliary teaching device for singing practice according to claim 1, characterized in that: The data storage and update module includes: A data compression unit compresses audio data using an adaptive compression algorithm based on audio signal characteristics; The data update unit cleans the new sample data and merges it with the original samples in the repository for storage.

9. The auxiliary teaching device for singing practice according to claim 7, characterized in that: The interactive operation module also includes: a voice recognition unit that can listen to the practitioner's voice instructions to adjust equipment parameters.