AI-based singing teaching audio analysis evaluation and intelligent examination system
Through AI-based singing teaching audio analysis and evaluation and intelligent examination system, the topology map is built for similarity analysis, which solves the problem of in-depth capture of audio signal characteristics and providing targeted feedback in the existing technology, and realizes in-depth analysis and scientific evaluation of students' singing audio.
Patent Information
- Application Number
- CN202510426978.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing singing teaching technology relies on teachers' experience and personal judgment, and cannot deeply capture the subtle changes and complex characteristics in the audio signal, and the feedback mechanism is limited and cannot provide targeted guidance.
Using AI-based singing teaching audio analysis and evaluation and intelligent examination system, through data collection, audio preprocessing, feature extraction and feature map construction modules, the pitch, pitch length and rhythm characteristics of students and standard singing audio are extracted, and the topological map is constructed for similarity analysis.
It realizes in-depth analysis of students' singing audio, provides scientific and comprehensive assessment, helps teachers formulate personalized teaching strategies, improve teaching efficiency and students' learning effectiveness.
Smart Images

Figure CN120148552A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of singing teaching, and particularly to an AI-based singing teaching audio analysis and evaluation and intelligent examination system. Background Art
[0002] In the current era of rapid digital development, the field of singing teaching is undergoing a profound transformation. Traditional singing teaching mostly relies on teachers' experience and personal judgment, usually guiding students through face-to-face teaching interactions. However, this method is often restricted by time, location, and the professional level of teachers. With the rapid development of artificial intelligence and audio analysis technology, many educators have started to use digital technology to assist traditional teaching, evaluating and providing feedback on students' singing performances through audio analysis tools, thereby promoting the improvement of their learning. Although the existing technology has improved teaching efficiency to a certain extent, enabling students to obtain more scientific learning support, there are still many deficiencies.
[0003] The existing technology mainly includes simple audio recording and analysis tools, which usually can only provide basic sound feature analysis. Most of these tools rely on basic signal processing algorithms and are unable to deeply capture the subtle changes and complex features in the audio signal. In addition, due to the relatively simple audio analysis function of most systems, the feedback mechanism is also limited within the conventional scoring framework and cannot provide targeted guidance for students. The simplified analysis method often leads to one-sided evaluation and cannot comprehensively reflect the actual performance of students and their potential improvement directions. Summary of the Invention
[0004] The present invention provides an AI-based singing teaching audio analysis and evaluation and intelligent examination system to solve the defects existing in the prior art.
[0005] The present invention provides an AI-based singing teaching audio analysis and evaluation and intelligent examination system, including: A data acquisition module for inputting audio signals, where the audio signals include standard singing audio and student singing audio.
[0006] An audio preprocessing module for denoising and normalizing the audio signal to obtain a normalized audio signal.
[0007] A feature extraction module for extracting the sound features of the normalized audio signal by using digital signal processing technology, where the sound features include student sound features and standard sound features.
[0008] A feature map construction module for quantifying and fusing voice features to obtain a high-dimensional feature vector, performing a non-linear mapping on the high-dimensional feature vector to obtain a feature representation, extracting feature points from the feature representation, calculating the correlation relationship between the feature points, and using the persistent homology method to construct the feature points and the correlation relationship into a topological graph structure to obtain a feature map, where the feature map includes a student feature map and a standard feature map.
[0009] A comparative analysis module for calculating the similarity between the student feature map and the standard feature map.
[0010] According to the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the present invention, the data acquisition module includes a standard singing audio recording unit, and the standard singing audio recording unit is used to capture the standard singing audio, and the process includes: Using a multi-dimensional audio capture technology through an audio capture device to record the standard singing audio.
[0011] Adopting a noise suppression technology based on the quantum acoustics principle to eliminate background noise.
[0012] According to the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the present invention, the data acquisition module further includes a student singing audio recording unit, and the student singing audio recording unit is used to provide a standardized recording environment and capture the student singing audio, and the process includes: Setting up a recording studio in an acoustically optimized environment and recording the student singing audio through an audio capture device.
[0013] Using a dynamic range extension technology to dynamically adjust the recording gain according to the amplitude change of the student singing audio.
[0014] According to the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the present invention, the audio preprocessing unit includes a noise reduction processing unit and a normalization processing unit. The noise reduction processing unit is used to convert the audio signal into the frequency domain through a short-time Fourier transform and identify the noise by applying a noise power spectrum estimation. The normalization processing unit is used to analyze the probability distribution characteristics of the audio signal, perform amplitude normalization processing on the audio signal to obtain a normalized audio signal, and the normalized audio signal includes a normalized standard audio and a normalized student audio.
[0015] According to the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the present invention, the voice features include pitch features, pitch length features and rhythm features, and the process of extracting features from the normalized audio signal includes: Segmenting the normalized audio signal, performing harmonic analysis on each segment of the normalized audio signal to obtain the harmonic amplitude of each segment of the normalized audio signal, multiplying all the harmonic amplitudes to obtain a harmonic product spectrum, and taking the peak value of the harmonic product spectrum as the pitch feature.
[0016] Apply the endpoint detection algorithm to mark the start frame and end frame of each normalized audio signal, calculate the time difference between adjacent start frames and end frames, and obtain the pitch length feature of each normalized audio signal.
[0017] Obtain the time-frequency information of the normalized audio signal. The time-frequency information includes energy peaks and periodic patterns. Use the beat tracking algorithm to obtain the beat positions based on the time-frequency information, and obtain the rhythm feature.
[0018] According to the AI-based singing teaching audio analysis and evaluation and intelligent examination system provided by the present invention, the feature map construction unit includes a feature quantization unit, a feature fusion unit, a feature representation unit, a feature point acquisition unit, a feature point relationship analysis unit, and a topological map construction unit. The feature quantization unit is used to quantize the sound features by using logarithmic compression operations to obtain quantization features. The feature fusion unit is used to perform feature fusion on the quantization features by using kernel principal component analysis to obtain high-dimensional feature vectors. The feature representation unit is used to perform non-linear mapping on the high-dimensional feature vectors through a multi-layer perceptron to obtain feature representations. The feature point acquisition unit is used to perform statistical analysis on the data points in the feature representation, set the key data point threshold by calculating the local density and outlier degree of the data points, and use the data points that reach the key data point threshold as feature points. The feature point relationship analysis unit is used to calculate the Euclidean distance between every two feature points and use the Euclidean distance as the association relationship of the feature points. The topological map construction unit is used to construct a topological map structure from the feature points and the association relationships by using persistent homology. The nodes in the topological map represent feature points, and the edges represent the association relationships between feature points.
[0019] According to the AI-based singing teaching audio analysis and evaluation and intelligent examination system provided by the present invention, the process of constructing a topological map structure from the feature points and the association relationships by using persistent homology includes: Extract feature points from the feature representation and generate a feature point set.
[0020] Calculate the Euclidean distance between the feature points and generate a distance matrix.
[0021] Set the adjacency relationship threshold, construct an adjacency graph, where the vertices in the adjacency graph correspond to feature points, and the edges correspond to the connection relationships where the Euclidean distance between feature points is less than the adjacency relationship threshold.
[0022] Expand the adjacency relationship threshold to obtain multiple scales, and construct a simplicial complex according to the multiple scales.
[0023] Extract the topological features in the simplicial complex and analyze the changes of the topological features at different scales.
[0024] Record the generated topological features for each scale, perform persistence analysis by tracking the appearance and disappearance of the topological features, and obtain persistence features.
[0025] Visualize the persistence features to generate a persistence diagram.
[0026] Set a threshold for the persistence features, and identify and extract the persistent topological features that reach the persistence feature threshold from the persistence diagram.
[0027] Combine the connection relationships of the feature points in the persistent diagram to obtain the topological graph structure.
[0028] According to the AI-based singing teaching audio analysis and evaluation and intelligent examination system provided by the present invention, the comparison analysis module includes a feature map preprocessing unit, a graph matching unit, and a similarity measurement unit. The feature map preprocessing unit is used to standardize the feature map and screen the feature points. The graph matching unit is used to match and align the student feature map with the standard feature map. The similarity measurement unit is used to calculate the similarity between the student feature map and the standard feature map.
[0029] According to the AI-based singing teaching audio analysis and evaluation and intelligent examination system provided by the present invention, the process of calculating the similarity includes: Use the graph edit distance as the similarity measure to calculate the minimum number of edit operations required to transform the student feature map into the standard feature map.
[0030] Generate the similarity based on the minimum number of edit operations.
[0031] According to the AI-based singing teaching audio analysis and evaluation and intelligent examination system provided by the present invention, the intelligent examination system includes: An audio input module for inputting audio data, including standard singing audio and the singing audio of the examinee.
[0032] An audio processing module for denoising and normalizing the audio data to obtain normalized data, including normalized standard data and normalized examinee data.
[0033] A feature acquisition module for extracting the voice features from the normalized data, including standard voice features and the singing features of the examinee.
[0034] A feature graph construction module for constructing a topological graph based on the voice features, including a standard topological graph and an examinee topological graph.
[0035] A scoring module for calculating the similarity between the standard topological graph and the examinee topological graph and generating an examination score based on the similarity.
[0036] The AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the present invention can effectively extract pitch, pitch duration and rhythm features by analyzing students' singing audio in real time. By constructing a topological graph, it provides an intuitive visual representation of the correlation between audio feature points, enabling teachers and students to clearly understand the mutual relationship between different audio features. It can not only display the feature graphs of individual students, but also compare different students with standard voice features to help users identify the advantages and disadvantages of their audio performances. It can capture rich topological features in the audio signal, such as connectivity and cyclic structures. The persistence at different voice scales provides an in-depth understanding of the essence of the audio signal, which helps to analyze the style, delicacy of music performance and the diversity of singing skills. Through the similarity calculation of the structural features in the graph, the similarity of students' singing audio relative to the standard audio can be evaluated more comprehensively. This comparison is not limited to local features, but covers the topological features of the entire audio, making the similarity evaluation more scientific and comprehensive. By analyzing the audio features of students through the topological graph, the respective singing characteristics and possible deficiencies of students can be revealed. Based on these analysis results, teachers can formulate targeted teaching strategies to ensure that each student can receive personalized guidance and help. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 FIG. 1 is a schematic structural diagram of the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by the embodiments of the present invention; Figure 2 FIG. 2 is a schematic structural diagram of the AI-based singing intelligent examination system provided by the present embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] The following will describe the AI-based singing teaching audio analysis, evaluation and intelligent examination system of the present invention in conjunction with Figure 1 - Figure 2 FIG. 1 and FIG. 2.
[0041] Figure 1 It is a schematic structural diagram of an AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by an embodiment of the present invention.
[0042] As Figure 1 shown, the AI-based singing teaching audio analysis, evaluation and intelligent examination system provided by an embodiment of the present invention includes a data acquisition module, an audio preprocessing module, a feature extraction module, a feature map construction module and a comparative analysis module.
[0043] The data acquisition module is used to input audio signals, and the audio signals include standard singing audio and student singing audio.
[0044] The data acquisition module includes a standard singing audio input unit, and the standard singing audio input unit is used to capture standard singing audio. The process includes: Using an audio acquisition device and applying multi-dimensional audio capture technology to record the standard singing audio.
[0045] The audio acquisition device includes a large diaphragm condenser microphone and an audio interface. The process of multi-dimensional audio capture includes: Equipping multiple microphones to achieve spatial acoustic modeling, including adjusting the microphones at different positions and angles to capture the propagation characteristics of sound in space and form a three-dimensional sound field.
[0046] Applying spatial acoustic modeling technology to simulate and analyze the sound propagation, and recording the incident angle and reflection characteristics.
[0047] In this embodiment, three high-quality large diaphragm condenser microphones, namely A, B, and C, are selected and placed at different positions and angles. The microphones are connected to a computer through an audio interface to form a complete audio acquisition system.
[0048] Among them, microphone A is set in the front, 1 meter away, to capture the direct sound source; microphone B is set on the side, 1.5 meters away, to capture the side reflection; microphone C is set obliquely above, 2 meters away, to capture the high-frequency reflection and ambient sound.
[0049] After setting up the microphones, conduct on-site tests to obtain the acoustic wave characteristics at different positions. Optimize the sound field acquisition effect by adjusting the height, tilt angle and distance of the microphones. Mark the positions of each microphone for subsequent experiments, and make adjustments according to the information of each microphone in the final mixing.
[0050] Use the spatial acoustic modeling software EASE to simulate different sound fields. Record the incident angle and reflection characteristics of the sound source during the simulation to determine the optimal microphone layout and settings.
[0051] Record the incident angle of each microphone relative to the sound source, where: the incident angle of microphone A is 0°, the incident angle of microphone B is 30°, and the incident angle of microphone C is 45°.
[0052] Adopt a noise suppression technology based on the principle of quantum acoustics to eliminate background noise, which is expressed by the formula:
[0053] In the formula, represents the filtered signal, represents the original signal, represents the noise power spectral density at frequency f_out, represents the frequency response of the filter.
[0054] The data acquisition module also includes a student singing audio recording unit, which is used to provide a standardized recording environment and capture the student singing audio. The process includes: Set up a recording studio in an acoustically optimized environment and record the student singing audio through an audio acquisition device.
[0055] Use dynamic range extension technology to dynamically adjust the recording gain according to the amplitude change of the student singing audio, which is expressed by the formula:
[0056] In the formula, represents the adjusted gain output, G represents the gain factor, represents the amplitude of the input audio signal.
[0057] In this embodiment, ensure that the recording studio is located in a quiet area, far from noise sources, and use acoustic materials such as sound-absorbing cotton and sound insulation boards for wall, ceiling and floor treatment to reduce sound wave reflection and external noise interference. Use a low-noise air conditioning system and ensure its stable operation to reduce background noise. Ensure good air circulation in the recording studio while maintaining a quiet environment. Configure multiple large diaphragm condenser microphones to improve the capture performance and at the same time achieve a richer audio signal. Use a professional audio interface to connect the microphone to the computer to ensure the conversion of high-quality analog signals to digital signals. Install an acoustic processor and a mixer for real-time monitoring and adjustment of the recording effect. Use headphones for monitoring and adjust the audio effect in time to avoid sound interference caused by delay. Before the formal recording, guide the students to do vocalization exercises, including breathing, vocalization and pitch control. Explain the use of the recording equipment and ensure their understanding of the recording process to reduce the nervousness during recording. Conduct a test recording to evaluate the recording environment and equipment effects and adjust the equipment parameters in time.
[0058] Monitor the audio signal amplitude through software and calculate the amplitude A of the input signal in real time input .
[0059] Set a dynamic gain factor G. When the signal amplitude is below a certain threshold, set the gain factor to 2.0. When the signal amplitude increases, reduce the gain factor to 0.8 to prevent distortion.
[0060]
[0061] During recording, use professional audio processing software to monitor the audio waveform in real time and adjust the value of G to adapt to the changes in the audio signal. With the help of the waveform diagram and input volume indicator provided by the software, perform gain adjustment to ensure that the maximum amplitude of the output signal does not exceed the range allowed by the system. After the initial recording is completed, conduct playback analysis and perform detailed gain adjustment for each segment of the student's audio. Optimize the overall audio quality through equalization and compression of the audio signal.
[0062] In one embodiment, when a student is singing, the amplitude change of the input audio signal is as follows: At a certain moment, A input = 0.3, then according to the formula, the gain adjustment is:
[0063] When the input signal amplitude increases to A input = 1.2, the gain factor changes to 0.8, and the output gain adjustment is:
[0064] Dynamic gain adjustment ensures the naturalness and expressiveness of the sound source and avoids distortion problems caused by excessive volume.
[0065] The audio preprocessing module is used to perform noise reduction and normalization processing on the audio signal to obtain a normalized audio signal.
[0066] The audio preprocessing unit includes a noise reduction processing unit and a normalization processing unit. The noise reduction processing unit is used to convert the audio signal into the frequency domain through short-time Fourier transform and identify the noise by applying noise power spectrum estimation. The normalization processing unit is used to analyze the probability distribution characteristics of the audio signal, perform amplitude normalization processing on the audio signal to obtain a normalized audio signal, and the normalized audio signal includes a normalized standard audio and a normalized student audio.
[0067] In this embodiment, the short-time Fourier transform divides the audio signal into short-time frames, each frame has a length of N, and the sampling frequency is f s = 44.1 kHz, and the number of samples per frame N = 2048, then the duration of each frame is:
[0068] Apply a Hamming window to each frame to reduce spectral leakage.
[0069] Estimate the noise power spectrum based on the statistical characteristics of the real audio signal. At a certain frequency f, the noise power spectrum is:
[0070] In the formula, represents the spectrum obtained by short-time Fourier transform. Through noise estimation, the noise components in the signal are identified.
[0071] Apply Wiener filtering for noise suppression. The filtering coefficient is expressed as:
[0072] In the formula, represents the power spectrum of the signal, represents the frequency response of the filter.
[0073] Conduct a statistical analysis on the amplitude of the audio signal to obtain the mean and standard deviation. The formula is expressed as:
[0074]
[0075] In the formula, represents the mean of the audio signal, represents the standard deviation of the audio signal, represents the amplitude of each sampling point of the audio signal, and M represents the total number of samples.
[0076] Normalize the amplitude of the audio signal so that the signal amplitude is within the range of [-1, 1]. The normalization formula is as follows:
[0077] In the formula, represents the normalized amplitude.
[0078] The feature extraction module is used to extract the voice features of the normalized audio signal by using digital signal processing technology. The voice features include student voice features and standard voice features.
[0079] The voice features include pitch features, duration features, and rhythm features. The process of extracting features from the normalized audio signal includes: Segment the normalized audio signal, conduct harmonic analysis on each segment of the normalized audio signal to obtain the harmonic amplitude of each segment of the normalized audio signal, multiply all the harmonic amplitudes to obtain the harmonic product spectrum, and take the peak value of the harmonic product spectrum as the pitch feature.
[0080] Apply the endpoint detection algorithm to mark the start frame and end frame of each normalized audio signal, calculate the time difference between adjacent start frames and end frames, and obtain the pitch length feature of each normalized audio signal.
[0081] Obtain the time-frequency information of the normalized audio signal. The time-frequency information includes the energy peak and the periodic pattern. Adopt the beat tracking algorithm to obtain the beat position according to the time-frequency information, and get the rhythm feature.
[0082] In this embodiment, the segmented length of the normalized audio signal is set to 1024 sampling points, the sampling frequency is 44.1 kHz, the time corresponding to the length of each audio signal segment is 0.0232 s. At the same time, in order to reduce the influence of segmentation, the overlapping length between each segment is set to 512 sampling points.
[0083] Divide the normalized audio signal into multiple segments. The start point and end point of each segment are expressed as:
[0084] where k represents the index of the segment, is the starting point of the k-th segment.
[0085] Generate multiple overlapping segments through this method to make the data have continuity and feature stability.
[0086] For each normalized audio signal segment, perform a fast Fourier transform to obtain its frequency spectrum.
[0087] Where a certain normalized audio signal segment is x[n], and its discrete Fourier transform is expressed as:
[0088] where X[k] represents the k-th frequency component in the discrete Fourier transform result, represents the complex coefficient of this frequency in the frequency domain. Each X[k] value contains the amplitude and phase information in the frequency domain. x[n] represents the input discrete signal sequence in the time domain. N represents the length of the discrete Fourier transform, that is, the total number of samples of the input signal x[n].
[0089] Calculate the corresponding harmonic amplitude according to the amplitude of the frequency spectrum. The formula is expressed as:
[0090] where represents the amplitude of the frequency component Re(X[k]) represents the real part of the complex number Im(X[k]) represents the imaginary part of the complex number
[0091] Multiply all the harmonic amplitudes to obtain the harmonic product spectrum. The formula is expressed as:
[0092] In the formula, R represents the number of excellent components of the spectrum.
[0093] Extract the peak value of the harmonic product spectrum as the pitch feature.
[0094] Mark the start frame and end frame of each audio signal through an endpoint detection algorithm. Adopt a method based on energy or zero-crossing rate, calculate the energy of each frame of audio, and confirm the start frame and end frame according to the set threshold.
[0095] Record the start frame index and end frame index of each normalized audio signal. Calculate the time difference between adjacent start frames and end frames to obtain the pitch length feature of each segment.
[0096] Extract the time-frequency information from the normalized audio signal, including energy peak and periodic pattern. For each audio signal segment, calculate the short-time energy and find the energy peak. Analyze the signal using the autocorrelation function, extract the periodicity from the autocorrelation function, perform a periodicity determination test on the sampled data, and extract the rhythm information of the audio.
[0097] According to the time-frequency information and periodic pattern, adopt a beat tracking algorithm to obtain the beat position. Adopt a method based on the energy threshold to narrow the analysis range, only consider the strongest energy peak to confirm the beat. Calculate the time interval between adjacent beats to establish the rhythm feature. Analyze the stability and change trend between beats through the interval time obtained by this calculation.
[0098] The feature map construction module is used to quantify and fuse the sound features to obtain a high-dimensional feature vector, perform a non-linear mapping on the high-dimensional feature vector to obtain a feature representation, extract feature points from the feature representation, and calculate the correlation relationship between the feature points. Adopt the persistent homology method to construct the feature points and correlation relationship into a topological graph structure to obtain a feature map, and the feature map includes a student feature map and a standard feature map.
[0099] The feature map construction unit includes a feature quantization unit, a feature fusion unit, a feature representation unit, a feature point acquisition unit, a feature point relationship analysis unit, and a topological map construction unit. The feature quantization unit is used to quantize the sound features by using logarithmic compression operation to obtain quantized features. The feature fusion unit is used to perform feature fusion on the quantized features by using kernel principal component analysis to obtain a high-dimensional feature vector. The feature representation unit is used to perform a non-linear mapping on the high-dimensional feature vector through a multi-layer perceptron to obtain a feature representation. The feature point acquisition unit is used to perform statistical analysis on the data points in the feature representation, set a key data point threshold by calculating the local density and outlier degree of the data points, and use the data points that reach the key data point threshold as feature points. The feature point relationship analysis unit is used to calculate the Euclidean distance between every two feature points and use the Euclidean distance as the association relationship of the feature points. The topological map construction unit is used to construct the feature points and the association relationship into a topological map structure by using persistent homology. The nodes in the topological map represent feature points, and the edges represent the association relationships between feature points.
[0100] In this embodiment, the main task of the feature quantization unit is to quantize the sound features in order to simplify the data representation of the features and reduce the complexity of subsequent processing. Logarithmic compression is used to reduce the eigenvalue range and enhance the influence of small-value features. Given the audio features, the formula for logarithmic compression is expressed as:
[0101] In the formula, represents the i-th compressed eigenvalue, represents the i-th value of the original audio feature.
[0102] The feature fusion unit fuses the quantized features through kernel principal component analysis to generate a high-dimensional feature vector.
[0103] Kernel principal component analysis aims to map the original feature space to a higher-dimensional space through non-linear mapping to extract the most important features. The formula is expressed as:
[0104] In the formula, Z represents the fused high-dimensional feature vector, K(X) represents the kernel mapping matrix of the input feature data, and U represents the feature vector matrix obtained through eigenvalue decomposition.
[0105] The feature representation unit performs a non-linear mapping on the high-dimensional feature vector through a multi-layer perceptron to generate a feature representation.
[0106] The multi-layer perceptron includes an input layer, a hidden layer, and an output layer. The formula for non-linear mapping is expressed as:
[0107] Where O represents the output feature, f represents the sigmoid activation function, Y represents the weight matrix, I represents the input feature, and b represents the bias.
[0108] In the feature point acquisition unit, statistical analysis is performed on the data points in the feature representation to identify key data points as feature points.
[0109] The local density of each point in the feature representation is calculated, and this density measures the distribution of data points in the feature space. The calculation formula for the local density is expressed as:
[0110] Where represents the local density, represents the neighborhood with a certain radius centered on the data point x, represents the kernel function.
[0111] In the feature point relationship analysis unit, the association relationships between the extracted feature points are calculated to determine the distances and relationships between each feature point.
[0112] For each pair of feature points p i and p j , the Euclidean distance calculation formula is expressed as:
[0113] Where represents the Euclidean distance between the feature points and , and respectively represent the first-dimensional eigenvalues of the feature points and , and respectively represent the second-dimensional eigenvalues of the feature points and .
[0114] The process of constructing the feature points and association relationships into a topological graph structure using persistent homology includes: extracting feature points from the feature representation and generating a set of feature points.
[0115] Calculate the Euclidean distance between the feature points and generate a distance matrix.
[0116] Set the adjacency relationship threshold Υ, construct an adjacency graph, where the vertices in the adjacency graph correspond to the feature points and the edges correspond to the connection relationships where the Euclidean distance between the feature points is less than the adjacency relationship threshold.
[0117] Expand the adjacency relationship threshold to obtain multiple scales, and construct a simplicial complex according to the multiple scales.
[0118] At each threshold, for feature points p i , p j ) satisfying (d(p i ) < Υ i and p j , add an edge.
[0119] If d(p i , p j ) < τ and d(p j , p k ) < Υ i , then p i , p j and p k form a triangle, i.e., a 2 - simplex.
[0120] Extract higher - order simplices, such as 3 - simplices, 4 - simplices, etc., to generate a filtered complex.
[0121] Extract topological features in the filtered complex and analyze the changes of topological features at different scales. Topological features include connected components, loops, and holes. Connected components reflect the connectivity between data points. Loops are used to detect loop structures in the filtered complex. Holes represent higher - order features, indicating dimensions in high - dimensional spaces.
[0122] Record the generated topological features for each scale, and perform persistence analysis by tracking the appearance and disappearance of topological features to obtain persistence features.
[0123] Visualize the persistence features to generate a persistence diagram. By visualizing the birth and death records of features as a persistence diagram, the persistence of features can be intuitively observed. Plot the features in a two - dimensional coordinate system, where the X - axis is the birth time of the feature and the Y - axis is the death time of the feature.
[0124] The result obtained by persistent homology is recorded as a persistence diagram, and the persistence diagram represents the birth scale and death scale of each feature.
[0125] Set a persistence feature threshold, and identify and extract persistent topological features that reach the persistence feature threshold from the persistence diagram. Determine a persistence feature threshold by analyzing the persistence diagram to identify and extract persistent topological features with practical significance.
[0126] Combine the connection relationships between feature points and their connections in the persistence diagram to obtain a topological graph structure. Connect feature points with persistent topological features. Establishing connection relationships between feature points is based on the persistent topological features identified in the persistence diagram. Generate the final topological graph by merging feature points with their connection relationships in the persistence diagram.
[0127] The comparison and analysis module is used to calculate the similarity between the student feature map and the standard feature map.
[0128] The comparative analysis module includes a feature map preprocessing unit, a graph matching unit, and a similarity measurement unit. The feature map preprocessing unit is used to normalize the feature map and screen the feature points. The graph matching unit is used to match and align the student feature map with the standard feature map. The similarity measurement unit is used to calculate the similarity between the student feature map and the standard feature map.
[0129] The process of calculating the similarity includes: Using the graph edit distance as the similarity measurement, calculate the minimum number of editing operations required to transform the student feature map into the standard feature map.
[0130] Generate the similarity based on the minimum number of editing operations, and the formula is expressed as:
[0131] In the formula, B represents the minimum number of editing operations calculated by the graph edit distance, and are the numbers of feature points in the student feature map and the standard feature map respectively.
[0132] Through the above process, integrating and comparatively analyzing the features extracted from the audio not only ensures the accurate analysis of the features, but also provides a direction for further improving the audio quality and feature extraction algorithm. In practical applications, it can deeply mine the data features, promote the research and application of audio signal processing, and provide strong support for related model training, audio recognition, and classification tasks.
[0133] Figure 2 is a schematic structural diagram of the AI-based intelligent singing examination system provided in this embodiment.
[0134] As Figure 2 shown, the intelligent examination system includes an audio recording module, an audio processing module, a feature acquisition module, a feature map construction module, and a scoring module.
[0135] The audio recording module is used to record audio data, including standard singing audio and the singing audio of the examinee.
[0136] The audio processing module is used to perform noise reduction and normalization processing on the audio data to obtain normalized data, including normalized standard data and normalized examinee data.
[0137] The feature acquisition module is used to extract the singing features in the normalized data, including standard voice features and the voice features of the examinee.
[0138] The feature map construction module is used to construct a topological map based on the voice features, including a standard topological map and an examinee topological map.
[0139] The scoring module is used to calculate the similarity between the standard topological graph and the examinee's topological graph, and generate an exam score based on the similarity.
[0140] In summary, this embodiment provides an AI-based singing teaching audio analysis and evaluation and intelligent examination system. By analyzing the singing audio of students in real time, it effectively extracts pitch, pitch length, and rhythm features. By constructing a topological graph, it provides an intuitive visual representation of the correlation between audio feature points, enabling teachers and students to clearly understand the mutual relationship between different audio features. It can not only display the feature graphs of individual students, but also compare different students with standard voice features to help users identify the advantages and disadvantages of their audio performances.
[0141] The topological graph constructed by the persistent homology method can capture rich topological features in the audio signal, such as connectivity and cyclic structures. The persistence of these features at different sound scales provides an in-depth understanding of the essence of the audio signal, which helps to analyze the style, fineness of music performance, and the diversity of singing skills.
[0142] By calculating the similarity of the structural features in the graph, it is possible to more comprehensively evaluate the similarity of the student's singing audio relative to the standard audio. This comparison is not limited to local features, but covers the topological features of the entire audio, making the similarity evaluation more scientific and comprehensive.
[0143] By analyzing the audio features of students through the topological graph, the individual singing characteristics and possible deficiencies of students can be revealed. Based on these analysis results, teachers can formulate targeted teaching strategies to ensure that each student can receive personalized guidance and help.
[0144] From the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. AI-based singing teaching audio analysis, evaluation and intelligent examination system, characterized by: include: A data acquisition module, used for recording audio signals, wherein the audio signals include standard singing audio and student singing audio; An audio preprocessing module, used for performing noise reduction and normalization processing on the audio signal to obtain a normalized audio signal; A feature extraction module, used for extracting the sound features of the normalized audio signal by using digital signal processing technology, wherein the sound features include student sound features and standard sound features; A feature map construction module is used to quantify and fuse the sound features to obtain a high-dimensional feature vector, perform nonlinear mapping on the high-dimensional feature vector to obtain a feature representation, extract feature points from the feature representation, calculate the correlation between the feature points, and use a persistent homology method to construct the feature points and the correlation into a topological graph structure to obtain a feature map, wherein the feature map includes a student feature map and a standard feature map; The comparison and analysis module is used to calculate the similarity between the student feature graph and the standard feature graph.
2. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The data acquisition module includes a standard singing audio input unit, and the standard singing audio input unit is used to capture standard singing audio. The process includes: Through the audio acquisition device, multi-dimensional audio capture technology is used to record standard singing audio; Noise suppression technology based on quantum acoustic principles eliminates background noise.
3. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The data collection module also includes a student singing audio recording unit, which is used to provide a standardized recording environment and capture the student singing audio. The process includes: Set up a recording studio in an acoustically optimized environment and record student singing audio using an audio acquisition device; Use dynamic range extension technology to dynamically adjust the recording gain according to the amplitude changes of the student's singing audio.
4. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The audio preprocessing unit includes a noise reduction processing unit and a normalization processing unit; the noise reduction processing unit is used to convert the audio signal into a frequency domain through short-time Fourier transform, and apply noise power spectrum estimation to identify noise; the normalization processing unit is used to analyze the probability distribution characteristics of the audio signal, perform amplitude normalization processing on the audio signal, and obtain a normalized audio signal, and the normalized audio signal includes normalized standard audio and normalized student audio.
5. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The sound features include pitch features, sound length features and rhythm features, and the process of extracting features from the normalized audio signal includes: The normalized audio signal is processed in sections, and harmonic analysis is performed on each section of the normalized audio signal to obtain the harmonic amplitude of each section of the normalized audio signal, and all the harmonic amplitudes are multiplied to obtain a harmonic product spectrum, and the peak value of the harmonic product spectrum is used as the pitch feature; The endpoint detection algorithm is applied to mark the start frame and end frame of each normalized audio signal, and the time difference between adjacent start frames and end frames is calculated to obtain the sound length feature of each normalized audio signal; The time-frequency information of the normalized audio signal is obtained, wherein the time-frequency information includes an energy peak and a periodic law, and a beat tracking algorithm is used to obtain a beat position according to the time-frequency information to obtain a rhythm feature.
6. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The feature map construction unit includes a feature quantization unit, a feature fusion unit, a feature representation unit, a feature point acquisition unit, a feature point relationship analysis unit and a topology map construction unit; the feature quantization unit is used to quantize the sound features by logarithmic compression operation to obtain quantized features; the feature fusion unit is used to perform feature fusion on the quantized features by kernel principal component analysis to obtain a high-dimensional feature vector; the feature representation unit is used to perform nonlinear mapping on the high-dimensional feature vector by a multi-layer perceptron to obtain a feature representation; the feature point acquisition unit is used to perform statistical analysis on the data points in the feature representation, set a key data point threshold by calculating the local density and outlier degree of the data points, and use the data points that reach the key data point threshold as feature points; the feature point relationship analysis unit is used to calculate the Euclidean distance between every two feature points, and use the Euclidean distance as the correlation relationship of the feature points; The topology map construction unit is used to construct the feature points and the association relationships into a topology map structure by using a persistent homology method, wherein nodes in the topology map represent feature points, and edges represent association relationships between feature points.
7. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The process of constructing the feature points and the association relationships into a topological graph structure using the persistent homology method includes: Extracting feature points from the feature representation and generating a feature point set; Calculate the Euclidean distance between feature points and generate a distance matrix; An adjacency threshold is set to construct an adjacency graph, wherein the vertices in the adjacency graph correspond to feature points, and the Euclidean distance between the edge-corresponding feature points is less than the adjacency threshold; Expanding the adjacency threshold to obtain multiple scales, and constructing a simplified complex according to the multiple scales; Extracting topological features from the simplified complex and analyzing changes in the topological features at different scales; Record the topological features generated for each scale, perform persistence analysis by tracking the appearance and disappearance of the topological features, and obtain persistence features; Visualizing the persistence characteristics to generate a persistence graph; Setting a persistence feature threshold, identifying and extracting persistent topological features reaching the persistence feature threshold from the persistence graph; The feature points are combined with the connection relationships of the feature points in the persistence graph to obtain a topological graph structure.
8. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1 is characterized in that: The comparative analysis module includes a feature map preprocessing unit, a map matching unit and a similarity measurement unit; the feature map preprocessing unit is used to standardize the feature map and screen the feature points; the map matching unit is used to match and align the student feature map with the standard feature map; the similarity measurement unit is used to calculate the similarity between the student feature map and the standard feature map.
9. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 8, characterized in that: The process of calculating the similarity includes: Using graph edit distance as a similarity metric, calculating the minimum number of edit operations required to transform the student feature graph into the standard feature graph; A similarity is generated according to the minimum number of editing operations.
10. The AI-based singing teaching audio analysis, evaluation and intelligent examination system according to claim 1, characterized in that: The intelligent examination system comprises: An audio input module is used to input audio data, including standard singing audio and examinee singing audio; An audio processing module, used for performing noise reduction and normalization processing on audio data to obtain normalized data, including normalized standard data and normalized examinee data; A feature acquisition module is used to extract the sound features in the normalized data, including the standard sound features and the examinee's singing features; A feature map building module is used to build a topology map based on sound features, including a standard topology map and an examiner topology map; The scoring module is used to calculate the similarity between the standard topology map and the examinee's topology map, and generate a test score based on the similarity.