A Qinqiang Opera Speech Learning Assistance System and Method Based on Big Data Analysis
By analyzing Qinqiang opera singing audio with big data, extracting pitch segments and generating deviation sequences, setting thresholds to identify problem windows, and calculating intersections, the problem of accurate positioning and quantitative evaluation of pitch coordination deviations in Qinqiang opera singing was solved, achieving precise technical guidance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XIAN UNVERSITY OF ARTS & SCI
- Filing Date
- 2026-03-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing technologies struggle to continuously track pitch changes during specific interval transitions in Qinqiang opera performances, making it difficult to accurately identify and locate pitch coordination deviations, resulting in a lack of targeted and precise analysis and feedback.
Through big data analysis, the fourth and seventh pitch segments in Qinqiang opera singing audio are extracted, compared with preset natural pitches to generate deviation sequences, thresholds are set to identify problem windows, the intersection of windows is calculated, and a quantitative evaluation report is generated.
It has achieved automated, precise time-based positioning and quantitative evaluation of pitch and coordination deviations in Qinqiang opera singing, providing accurate technical guidance and feedback.
Smart Images

Figure CN122135739A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of opera singing analysis technology, and more specifically, this application relates to a Qinqiang opera speech learning assistance system and method based on big data analysis. Background Technology
[0002] In traditional Qinqiang opera singing, the transitions between specific intervals, such as the "joyful tone" (which is a natural pitch) and the "bitter tone" (which is a slight raise of a fourth and a slight lower of a seventh), are key technical aspects that embody its artistic style. Physically, this process manifests as a continuous and rapid dynamic change in the vocal audio.
[0003] Currently, audio analysis technology has been applied to assist vocal music teaching, but existing technologies still have limitations when dealing with continuous and non-linear pitch changes with distinctive operatic characteristics, such as Qinqiang opera's transitions. These systems often focus on the identification and evaluation of stable pitch nodes or statistical analysis of the entire performance segment, making it difficult to synchronously and continuously track the deviation states of the two pitches during the transition.
[0004] Because it is impossible to automatically and accurately segment specific problem areas from continuous audio signals, the analysis feedback and instantaneous technical errors in the performance cannot be precisely matched in the time dimension. The evaluation results lack process descriptions for specific transition stages, resulting in insufficient precision in technical guidance.
[0005] To address the aforementioned issues, there is an urgent need in this field for an audio analysis technology capable of continuously tracking pitch transitions during specific intervals in Qinqiang opera performances and automatically identifying and locating time windows of pitch coordination deviation problems. This technology would solve the technical shortcomings of existing systems, such as insufficient specificity and ambiguous problem location in dynamic pitch process analysis. Summary of the Invention
[0006] To address the aforementioned technical problems, this technical solution provides a Qinqiang opera speech learning assistance system and method based on big data analysis, which solves the problems mentioned in the background section.
[0007] In a first aspect, embodiments of this application provide a Qinqiang opera speech learning assistance system based on big data analysis, comprising: a data acquisition module: used to acquire the first Qinqiang opera singing audio data of the target user in the current period and perform audio extraction to obtain the fourth pitch level and the seventh pitch level, and respectively subtract them from the corresponding preset natural pitches to obtain a first deviation sequence and a second deviation sequence; a first deviation window acquisition module: used to record the first time the first deviation sequence exceeds the preset first deviation threshold as the first starting point if the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, and to record the first time after the first starting point that the first deviation sequence is less than the preset first deviation threshold as the first ending point, based on the first starting point and the first deviation window acquisition module; The first deviation window is obtained from the first cutoff point; the second deviation window acquisition module is used to record the first time the second deviation sequence exceeds the preset second deviation threshold as the second starting point and the first time the second deviation sequence is less than the preset second deviation threshold after the second starting point as the second cutoff point, and obtain the second deviation window based on the second starting point and the second cutoff point; the initial conversion window acquisition module is used to calculate the intersection of the first deviation window and the second deviation window to obtain the initial conversion window; the Qinqiang learning auxiliary report output module is used to generate and output the first Qinqiang learning auxiliary report based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
[0008] Secondly, embodiments of this application provide a Qinqiang opera speech learning assistance method based on big data analysis, including the following steps: acquiring the first Qinqiang opera singing audio data of the target user in the current period and extracting the audio to obtain the fourth pitch level and the seventh pitch level, and subtracting them from the corresponding preset natural pitches to obtain the first deviation sequence and the second deviation sequence; if the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, the first time the first deviation sequence exceeds the threshold is recorded as the first starting point, and the moment when the first deviation sequence first falls below the preset first deviation threshold after the first starting point is recorded as the first segment. The first deviation window is obtained based on the first starting point and the first ending point. If the duration of the second deviation sequence exceeding the preset second deviation threshold exceeds the preset second duration threshold, the moment when the second deviation sequence first exceeds the threshold is recorded as the second starting point, and the moment when the second deviation sequence first falls below the preset second deviation threshold after the second starting point is recorded as the second ending point. The second deviation window is obtained based on the second starting point and the second ending point. The intersection of the first deviation window and the second deviation window is calculated to obtain the initial conversion window. The first Qinqiang opera learning assistance report is generated and output based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
[0009] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned Qinqiang opera speech learning assistance system based on big data analysis.
[0010] One or more technical solutions provided in the embodiments of this application have at least the following technical effects or advantages:
[0011] 1. This application generates a first deviation sequence and a second deviation sequence throughout the singing process by extracting the fourth and seventh pitch segments and continuously comparing them with the corresponding preset natural pitches. This transforms subjective auditory evaluation into continuous pitch deviation data of two key pitches changing over time, achieving objective quantitative monitoring of the core elements of the transition process and solving the problem of evaluation relying on subjective experience and lacking data support.
[0012] 2. This application intelligently identifies the periods of pitch problems that exceed the tolerance and occur continuously from the first and second deviation sequences by setting preset first deviation thresholds, preset second deviation thresholds, and corresponding preset first and second duration thresholds; these are known as the first deviation window and the second deviation window. An initial conversion window is obtained by calculating the intersection of these two windows. This window accurately represents the core time period during the conversion from the fourth to the seventh tone when pitch problems occur simultaneously in both tone classes. This solves the problem of unclear localization to a specific time interval.
[0013] 3. This application generates a first Qinqiang opera learning assistance report by calculating the overlap between the window and the preset conversion window, providing a quantitative score or indicator that directly and objectively reflects the gap between the user's vocal performance and the standard or expectation. It realizes the leap from fuzzy evaluation to quantitative feedback based on a precise time window, providing clear goals and benchmarks for targeted practice. Attached Figure Description
[0014] Figure 1 A schematic diagram of the structure of the Qinqiang Opera speech learning assistance system based on big data analysis provided in the embodiments of this application;
[0015] Figure 2 A schematic diagram of the logical flow of the Qinqiang Opera speech learning assistance system based on big data analysis provided in the embodiments of this application;
[0016] Figure 3 A schematic diagram illustrating the steps of the Qinqiang Opera speech learning assistance method based on big data analysis provided in the embodiments of this application. Detailed Implementation
[0017] This application provides a Qinqiang opera speech learning assistance system and method based on big data analysis, which solves the technical problems in the prior art where the analysis of the dynamic pitch process in Qinqiang opera singing is not targeted enough and the problem positioning is vague.
[0018] To address the problem that existing audio analysis techniques struggle to accurately locate and quantify pitch coordination deviations during specific interval transitions in Qinqiang opera singing, the underlying logic of this solution stems from a clear observation: effective technical diagnosis must first deconstruct the continuous, mixed singing audio stream into independently analyzable signal components corresponding to specific technical actions.
[0019] In the context of Qinqiang opera's vocal transitions, the core action involves the dynamic connection of two key pitches. Therefore, the starting point is to perform targeted signal separation on the singing audio, extracting the core data streams representing the performance of these two pitches—their respective pitch segment information. Simply obtaining the pitch trajectory is insufficient to diagnose "deviation"; an objective reference system must be established. Thus, each extracted pitch segment data is continuously compared to its corresponding ideal pitch value, pre-established using standard data, calculating the difference between the actual pitch and the ideal pitch at each moment. This operation transforms the abstract concept of "pitch accuracy" into two continuously changing, quantified data sequences over time, representing the independent and synchronous pitch deviations of the two target pitches during the singing process.
[0020] The next step is to intelligently identify the truly meaningful "problem periods" that require attention from these two consecutive deviation data sequences. A simple, momentary exceedance might be an occasional fluctuation, not a stable technical error. Therefore, the key to logical refinement lies in introducing "persistence" as a criterion. By setting reasonable deviation tolerance thresholds and minimum duration thresholds for each pitch deviation sequence, the system can automatically screen out audio segments that both exceed the allowable range and persist for a sufficient duration. This establishes two independent "problem windows" on the timeline: one window identifies the period when the fourth pitch is consistently sung incorrectly, and the other identifies the period when the seventh pitch is consistently sung incorrectly.
[0021] However, simply identifying problem windows for each of the two pitches is insufficient to accurately characterize the technical flaws in the coordinated action of "vocal transition." The core manifestation of the vocal transition problem in physical time lies precisely in the overlapping timeframe during which the problems of these two pitches occur simultaneously. Therefore, the key logical leap lies in correlating and analyzing the two independently identified problem windows, calculating their intersection on the timeline. This overlapping time interval is defined as the "initial transition window," which precisely characterizes the exact start and end times of the coordinated technical problem of simultaneous instability in pitch control of both pitches during the singer's transition from the fourth to the seventh pitch.
[0022] Ultimately, to achieve an operational and quantitative assessment, this diagnosed problem window needs to be compared with an ideal, standard, or desired technical action timing template—a preset transition window. By calculating the overlap ratio between the actual diagnosed problem window and the standard window, the originally vague evaluation of "inaccurate pitch transitions" can be transformed into an objective quantitative score based on precise timing anchors. This score directly reflects the gap between the singer's actual performance and the target technical standard in specific core aspects.
[0023] In summary, the core innovative logic of this solution lies in its ability to achieve, for the first time, automated and precise temporal localization and quantitative assessment of the specific technical problem of "pitch coordination deviation" in Qinqiang opera singing through a progressive analysis chain of "signal separation – independent quantization – continuous deviation judgment – window intersection." The most unique technical problem it solves is that existing general-purpose audio analysis systems cannot automatically identify and accurately define the specific time period of technical error caused by the simultaneous loss of control of two related pitches during transition in a continuous singing stream. This provides a previously unattainable data benchmark for precise and process-oriented guidance in opera singing.
[0024] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0025] Figure 1This is a schematic diagram of the structure of the Qinqiang Opera speech learning assistance system based on big data analysis provided in this application embodiment. The Qinqiang Opera speech learning assistance system based on big data analysis includes: a data acquisition module: used to acquire the first Qinqiang Opera singing audio data of the target user in the current period and perform audio extraction to obtain the fourth pitch level and the seventh pitch level, and respectively subtract them from the corresponding preset natural pitch to obtain the first deviation sequence and the second deviation sequence; a first deviation window acquisition module: used to record the first time the first deviation sequence exceeds the preset first deviation threshold as the first starting point if the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, and to record the first time after the first starting point the first time the first deviation sequence is less than the preset first deviation threshold as the first cutoff point. The system comprises the following modules: a first deviation window is obtained based on a first starting point and a first ending point; a second deviation window acquisition module is used to record the first time the second deviation sequence exceeds a preset second deviation threshold as the second starting point and the first time the second deviation sequence falls below the preset second deviation threshold after the second starting point as the second ending point, and to obtain the second deviation window based on the second starting point and the second ending point; an initial conversion window acquisition module is used to calculate the intersection of the first deviation window and the second deviation window to obtain the initial conversion window; and a Qinqiang opera learning assistance report output module is used to generate and output a first Qinqiang opera learning assistance report based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
[0026] Figure 2 A schematic diagram of the logical flow of the Qinqiang Opera speech learning assistance system based on big data analysis provided in the embodiments of this application.
[0027] The preset natural pitch specifically refers to the theoretical standard pitch value in a particular Qinqiang opera aria. It is not obtained from a single frequency value, but rather from a standard sample library composed of pitch standards determined according to the mode and singing style. A high-precision fundamental frequency extraction algorithm is used to batch process the sample library, extracting the pitch values of the "fourth pitch" and "seventh pitch" in stable sustained notes from all samples. Statistical analysis is performed on all the extracted pitch values for each pitch, such as calculating the median or robust mean, and fine-tuning is made according to musical acoustics theory to finally determine the preset natural pitch used for comparison in that aria.
[0028] The preset transition window defines the ideal time interval for the transition from stable fourth-tone vocalization to stable seventh-tone vocalization in this Qinqiang opera excerpt. Its construction is based on the analysis of a standard sample library. Using automatic note segmentation and fundamental frequency tracking technology, the moments when the "stable fourth-tone segment ends / begins to change" and the moments when the "stable seventh-tone pitch is reached" are precisely marked in each standard singing sample. The intervals of all samples are statistically analyzed, and the central trends of their start times and durations are calculated to ultimately determine a fixed time interval as the preset transition window.
[0029] For the input first Qinqiang opera singing audio data, a robust fundamental frequency estimation algorithm is used to perform continuous fundamental frequency estimation throughout the entire duration, resulting in an original pitch curve that varies over time. Simultaneously, the system uses score-following technology to obtain structural information about the "fourth pitch" and "seventh pitch" in the singing segment from MIDI scores or pre-annotated notes. Guided by these positions, relatively stable pitch segments corresponding to these two pitches are located on the original pitch curve. This requires precise segmentation using note initiation detection, ultimately yielding two pitch-time sequences representing the user's performance of the fourth and seventh pitches, respectively: the fourth pitch segment and the seventh pitch segment.
[0030] This solution achieves full-process, dual-channel, and collaborative pitch deviation diagnosis for the key technical aspect of specific interval transitions in Qinqiang opera singing. Specifically, the solution deconstructs continuous singing audio into independent pitch trajectories for two target pitch levels and generates two parallel pitch deviation monitoring sequences through continuous comparison with preset standard pitches. A "continuous over-threshold" criterion is introduced to identify the effective problem periods for each pitch level, and finally, by calculating the intersection of the two problem periods, the collaborative problem time window where pitch instability occurs simultaneously during the transition between the two pitch levels is accurately and automatically located. This transforms the description of the complex problem of inaccurate pitch transitions from a vague overall evaluation to a quantitative "problem interval" location based on precise time coordinates. This lays an irreplaceable and objective data foundation for subsequent accurate feedback, correction, and in-depth analysis, fundamentally solving the pain points of existing general technologies in this specific scenario, which suffer from vague problem location and lack of specificity.
[0031] Furthermore, after generating the first Qinqiang opera learning assistance report, the first correction module is executed: The sequences of the first and second deviation sequences that are within the initial conversion window are extracted to obtain the first and second pitch transition sequences, and first-order difference operations are performed on them respectively to obtain the first and second rate of change sequences; if the first rate of change sequence exceeds the first rate of change threshold, the moment when the first rate of change sequence first exceeds the first rate of change threshold is taken as the starting point for the fourth pitch refinement; if the second rate of change sequence exceeds the second rate of change threshold, the moment when the second rate of change sequence first falls below the second rate of change threshold is taken as the starting point for the seventh pitch refinement; The starting time of the first Qinqiang opera performance audio data acquisition is recorded as the starting time. The larger of the differences between the starting point of the fourth-level refinement and the starting point of the seventh-level refinement and the starting time is obtained, and the corresponding time is determined as the correction starting time. The starting time and ending time of the initial conversion window are obtained. If the absolute difference between the correction starting time and the starting time of the initial conversion window exceeds a preset difference threshold, the starting time of the initial conversion window is replaced with the correction starting time to obtain the first correction conversion window. The second Qinqiang opera learning assistance report is generated and output based on the ratio of the union of the first correction conversion window and the preset conversion window to the preset conversion window.
[0032] In this embodiment, the first and second rate of change thresholds are set based on the fact that, during a standard and smooth Qinqiang opera transition, when the pitch transitions from one stable pitch to another, the rate of change will experience a phase of rapid increase relative to the starting pitch or rapid decrease relative to the target pitch. By analyzing the rate of change distribution of the first and second transition sequences in a large number of standard singing samples near a preset transition window, the high percentile value of this distribution, such as the 85th percentile, is taken as the threshold to identify significant starting points of dynamic pitch changes. Typical values range from 100 to 500 cents per second. For example, the first rate of change threshold can be set to +300 cents / sec to indicate that the fourth pitch begins to decline significantly, and the second rate of change threshold can be set to -280 cents / sec to indicate that the seventh pitch begins to stabilize significantly.
[0033] The preset difference threshold is used to determine whether the difference between the starting point of automatic refinement and the starting point of the initial judgment is significant and needs to be corrected. Its setting takes into account the slight delays that may exist in the singer's physiological reactions and breath transitions, ensuring that correction is only triggered when the time deviation is large enough to indicate that the initial threshold judgment may be affected by pre-existing background noise or slow shifts. A typical value range is 50 milliseconds to 200 milliseconds; for example, it can be set to 100 milliseconds.
[0034] First-order difference operations are a common technique in the field of signal processing, used to calculate the difference between adjacent data points in a discrete-time series, thereby obtaining the rate of change of the series.
[0035] This embodiment does not simply rely on static pitch deviation amplitude to mark the boundary of the problem window, but delves into the problem window to analyze the dynamic characteristics of pitch changes. By calculating the rate of change and setting an appropriate dynamic threshold, it can more accurately capture the physical moment when "the pitch begins to undergo a substantial change." Taking the larger of the two values as the correction starting point logically ensures that the corrected starting point is at least no earlier than the moment when either of the two pitches begins to undergo significant dynamic change. This is more consistent with the actual starting point of the "co-conversion" problem and filters out problem windows that are prematurely triggered by the slow drift of a single pitch.
[0036] This embodiment performs dynamic analysis on the pitch change process within the initially located problem window, and refines the starting time of the problem using a pitch change rate threshold. When the refined time differs significantly from the original time, the starting point of the problem window is corrected. This advances the localization of the starting time of the "inaccurate pitch transition" problem from judging the "deviation state" based on static amplitude to recognizing the "change action" based on the dynamic process. Thus, in the fast-paced dynamic scenario of Qinqiang opera pitch transitions, a more accurate problem time anchor point that is closer to the actual physiological reality of singing and has stronger anti-interference ability is obtained.
[0037] Furthermore, the cutoff time of the first correction conversion window is obtained, and within the subsequent preset time interval, the third starting point acquisition module is executed: the ratio of high-frequency energy to low-frequency energy of each time frame in the second Qinqiang singing audio data within the preset time interval is obtained and calculated to generate a high-frequency energy ratio sequence; when the duration of the high-frequency energy ratio sequence continuously exceeding the preset roaring energy threshold reaches the preset stable duration, the corresponding time is recorded as the third starting point.
[0038] The Qinqiang opera's bitter-tone singing style is often accompanied by explosive "roaring" techniques. Identifying the starting point of this roaring is crucial for assessing the transition between vocal shifts. After switching the problem window, the starting point detection for the roaring is performed.
[0039] In this embodiment, high-frequency energy generally refers to the energy components above 4kHz in the spectrum, which are related to the brightness, penetration, and overtones unique to shouting; low-frequency energy generally refers to the energy components below 500Hz, which are mainly related to the fundamental frequency and chest resonance. High-frequency energy and low-frequency energy can be obtained through statistical analysis of historical data.
[0040] A preset energy threshold for shouting is used to distinguish between shouting and regular singing. During shouting, the vocal cords vibrate differently, and the high-frequency overtone energy is significantly enhanced. By extracting a large number of Qinqiang opera audio samples labeled with the start time of shouting, the high-frequency energy ratio sequence before and after the start of shouting is calculated, and the threshold is determined through statistical classification. The typical value range is 2.0 to 8.0, for example, set to 4.5.
[0041] The preset stabilization duration is designed to prevent false triggering by brief bursts of sound or noise. It requires the high-frequency energy ratio sequence to continuously exceed a threshold for a certain duration before a stable start to the shouting is confirmed. This duration is set based on the energy build-up time of the shouting's onset. A typical value range is 30 milliseconds to 150 milliseconds, for example, set to 80 milliseconds.
[0042] The acquisition and calculation of the ratio of high-frequency energy to low-frequency energy in each time frame of the second Qinqiang opera singing audio data within a preset time interval includes spectral energy calculation: the audio signal of each time frame is subjected to short-time Fourier transform to obtain the spectrum, and then the spectrum amplitude is integrated or summed within a specified frequency range to obtain the energy value of that frequency band.
[0043] This embodiment transforms the important stylistic feature of Qinqiang opera, "roaring," into quantifiable audio signal features for automated detection. This not only achieves objective identification of distinctive techniques but, more importantly, provides a crucial time benchmark—the third starting point—for subsequent analysis of the "state transition" between vocal transitions and roaring. By performing detection within a preset time interval, for example, 0.5 to 2 seconds immediately following the cutoff time of the first correction transition window, it can intelligently focus on the reasonable time range within which roaring may occur.
[0044] This embodiment calculates the ratio of high-frequency to low-frequency energy in audio frames to form a feature sequence, and automatically detects the starting moment of iconic shouting techniques in Qinqiang opera singing based on energy ratio thresholds and stable duration conditions. This achieves objective identification of characteristic techniques in opera singing, laying a crucial data foundation for expanding from single pitch analysis to a comprehensive evaluation of the connection quality between "pitch transitions" and "style techniques".
[0045] Furthermore, if a third starting point is obtained, the second correction module is executed: using the third starting point as a reference point, backtracking for a preset backtracking time is performed to obtain a backtracking cutoff candidate point; if the difference between the cutoff time and the starting time of the first correction conversion window is greater than the difference between the backtracking cutoff candidate point and the starting time, the cutoff time of the first correction conversion window is updated to the backtracking cutoff candidate point, and the second correction conversion window is obtained.
[0046] In this embodiment, the preset backtracking duration represents a reasonable time interval tracing back from the starting point of the shouting performance, assuming the singer has completed the pitch transition and is ready to shout. This setting is based on statistical analysis of the breath adjustment and resonance transition time required by professional singers between transitioning from a stable pitch to the start of the shouting performance. A typical value range is 100 milliseconds to 400 milliseconds, for example, set to 250 milliseconds.
[0047] This embodiment uses a third starting point, which is a subsequent, clearly defined singing event, to deduce and constrain the reasonable end boundary of the preceding problem window. If the pitch problem window initially determined by the system ends too late, or even overlaps with the preparation stage for singing, this is logically unreasonable.
[0048] Therefore, by taking the third starting point as a benchmark, backtracking by one preparation time yields a candidate point for the backtracking cutoff. The original window cutoff time is updated to this candidate point only when the original cutoff time is later. In essence, this is a "posterior" calibration of the cutoff boundary of the problem window based on the objective laws of singing. This ensures that the analyzed "pitch problem period" will not unreasonably encroach on the clear subsequent singing stage, making the problem location more consistent with the temporal structure of the actual singing.
[0049] This embodiment, based on the detected starting point of the shouting performance, traces back a preset preparation time to reasonably calibrate the cutoff time of the pitch conversion problem window. This introduces a posterior correction mechanism based on the logical relationship of the singing event, making the located "problem window" not only based on the signal itself in the time dimension, but also consistent with the sequential logic of the singing technique execution, thus improving the authenticity and interpretability of the problem description.
[0050] Furthermore, after obtaining the second correction conversion window, the multi-dimensional quantitative analysis module is executed: obtaining the cutoff time of the second correction conversion window; extracting the spectral features of the audio frames of the second Qinqiang opera singing audio data corresponding to the cutoff time of the second correction conversion window to obtain the first spectral feature vector; extracting the spectral features of the audio frames of the second Qinqiang opera singing audio data corresponding to the third starting point to obtain the second spectral feature vector; calculating the cosine distance between the first spectral feature vector and the second spectral feature vector as an audio state coherence index; calculating the average slope of the linear fitting of the high-frequency energy ratio sequence as the steepness feature of the roaring start; calculating the difference between the audio state coherence index and the preset coherence threshold to obtain the state connection score; calculating the difference between the steepness feature of the roaring start and the preset steepness threshold to obtain the start intensity score.
[0051] In this embodiment, the spectral feature vector can be obtained by extracting the Mel-frequency cepstral coefficients of the audio frame. For example, 13-dimensional MFCC coefficients can be calculated to form a 13-dimensional feature vector.
[0052] The calculation process of Mel frequency cepstral coefficients mainly includes: pre-emphasizing the audio frame, windowing the frame, obtaining the spectrum through fast Fourier transform, passing the spectrum through a Mel-scale filter bank, taking the logarithm of the output of each filter, and finally performing discrete cosine transform to obtain the MFCC coefficients.
[0053] The audio state coherence metric uses cosine distance to measure the difference between two spectral feature vectors. The smaller the cosine distance, the more similar the spectra and the more coherent the state.
[0054] A linear fit is performed on a segment of the high-frequency energy ratio sequence near the third starting point. The slope of the resulting fitted line is the steepness characteristic of the onset of the shouting, reflecting the speed at which the shouting energy is built up.
[0055] Preset coherence threshold and preset kurtosis threshold: By calculating the above indicators on a large number of labeled singing samples and correlating them with expert scores, the thresholds that distinguish between the two are statistically derived. For example, the coherence threshold may be 0.35 unit cosine distance, and the kurtosis threshold may be 25 energy ratios / second.
[0056] This embodiment, when time intervals permit, quantitatively evaluates the crucial transition segment from "end of pitch transition" to "beginning of shouting" from two dimensions: spectral similarity and energy build-up speed, generating a state transition score and an onset intensity score. This is the first time that the analysis of Qinqiang opera's pitch transition problem has been expanded from the single dimension of "whether it is on time and in pitch" to a multi-dimensional quality assessment of "whether the state after the transition smoothly transitions to the next technique," achieving a refined depiction of the continuity of the performance and the quality of the explosive start.
[0057] Furthermore, if the state transition score and the attack intensity score are obtained, the third correction module is executed: the state transition score and the attack intensity score are normalized and weighted to obtain the overall transition quality score; when the overall transition quality score is lower than the preset quality threshold, the difference between the preset quality threshold and the overall transition quality score is recorded as the quality difference, and the window length of the preset transition window is shortened according to the ratio of the quality difference to the overall transition quality score and used for the next cycle.
[0058] In this embodiment, a preset quality threshold is used to determine whether the transition quality is "qualified". This threshold is obtained by calibrating the overall transition quality score of a large number of singing samples with the "passing grade" of expert subjective evaluation.
[0059] This embodiment not only outputs user evaluations but also uses the evaluation results to optimize the system's own analysis parameters. If a user's transition quality is poor, it indicates that they cannot complete a good transition under the time pressure defined by the current preset transition window. Based on this, the system dynamically shortens the preset transition window length for the next cycle, effectively raising the expected standard for the user's transition speed. This is a personalized, progressive training difficulty adjustment mechanism that guides users towards more compact and efficient standard transition techniques, making the auxiliary system intelligent in its teaching strategy.
[0060] This embodiment integrates multi-dimensional scores into an overall quality score, and adaptively adjusts the standard time window length for the next training cycle based on this score. This achieves dynamic matching between system evaluation parameters and the user's actual learning progress, upgrading the auxiliary system from a static "measuring instrument" to an "intelligent coach" with personalized teaching strategy adjustment capabilities, thus promoting the targeted and progressive nature of training.
[0061] Furthermore, when the overall connection quality score is lower than the preset quality threshold, the process also includes: using the overall connection quality score as the query condition, performing a graph traversal search in the pre-constructed Qinqiang opera singing teaching knowledge graph to locate related error pattern nodes, correct demonstration nodes, and training method nodes; integrating the content contained in the error pattern nodes, correct demonstration nodes, and training method nodes, generating a third Qinqiang opera learning assistance report according to preset rules, and outputting it.
[0062] In this embodiment, the construction process of the pre-constructed Qinqiang opera singing teaching knowledge graph is as follows: First, common technical errors, correct technical points, and targeted training methods related to "vocal transitions" and "roaring onset" are extracted from Qinqiang opera textbooks, academic literature, and teaching video texts of renowned teachers to form entities (nodes). Then, the relationships between entities are defined. Finally, the graph database is used for storage. During querying, the overall transition quality score is mapped to specific error pattern keywords or score ranges, and the associated nodes and paths are queried in the graph.
[0063] This embodiment deeply integrates quantitative analysis results with structured domain knowledge. It doesn't simply provide a score, but uses the score as a key to unlock a pre-organized, rich knowledge base. By locating error pattern nodes, it tells the user "You may have made XX mistake"; by associating with correct demonstration nodes, it shows "The correct approach should be YY"; and by recommending training method nodes, it provides "You can improve by practicing ZZ". This makes the output third-party Qinqiang opera learning assistance report highly targeted and actionable, directly translating data analysis conclusions into teaching actions, greatly enhancing the practical value and teaching effectiveness of the assistance system.
[0064] This embodiment, based on quantitative scoring, automatically retrieves relevant error causes, correct examples, and training plans from the Qinqiang opera teaching knowledge graph and integrates them to generate an in-depth tutoring report. This achieves intelligent tutoring across the entire chain, from "identifying problems" to "explaining the root causes" and then to "providing solutions," significantly improving the amount of information and guidance value of learning feedback, and mimicking the complete tutoring process of expert teachers—"diagnosis-attribution-prescription."
[0065] Furthermore, it also includes a tutoring visualization interface module: displaying the fitting curves of the first deviation sequence and the second deviation sequence, and overlaying the first starting point, the first ending point, the second starting point, and the second ending point; displaying the curve of the high-frequency energy ratio sequence, and marking the third starting point; displaying the overall connection quality score; and displaying all Qinqiang learning assistance reports.
[0066] In this embodiment, through the above technical solution, the core data sequences, key event points, analysis windows, and evaluation results generated throughout the analysis process are visualized in a graphical and integrated manner. This provides users and mentors with an intuitive and comprehensive analysis view, making the complex signal processing and logical judgment results in the background transparent, greatly optimizing the human-computer interaction experience, and helping users understand their own problems and track their progress.
[0067] Figure 3 This is a schematic diagram illustrating the steps of the Qinqiang opera speech learning assistance method based on big data analysis provided in this application embodiment. The Qinqiang opera speech learning assistance method based on big data analysis includes the following steps: acquiring the first Qinqiang opera singing audio data of the target user in the current period and performing audio extraction to obtain the fourth pitch level and the seventh pitch level, and subtracting them from the corresponding preset natural pitches to obtain the first deviation sequence and the second deviation sequence; if the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, then the first time the first deviation sequence exceeds the threshold is recorded as the first starting point, and the first time the first deviation sequence after the first starting point is less than the preset first deviation threshold is recorded as the first time the first deviation sequence exceeds the threshold. The moment when the second deviation sequence deviates from the threshold is recorded as the first cutoff point. A first deviation window is obtained based on the first start point and the first cutoff point. If the duration of the second deviation sequence exceeding the preset second deviation threshold exceeds the preset second duration threshold, the moment when the second deviation sequence first exceeds the threshold is recorded as the second start point, and the moment when the second deviation sequence first falls below the preset second deviation threshold after the second start point is recorded as the second cutoff point. A second deviation window is obtained based on the second start point and the second cutoff point. The intersection of the first deviation window and the second deviation window is calculated to obtain the initial conversion window. A first Qinqiang opera learning assistance report is generated and output based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
[0068] This application also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements a Qinqiang opera speech learning assistance system based on big data analysis.
[0069] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0070] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0071] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0072] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0073] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0074] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A Qinqiang Opera pronunciation learning assistance system based on big data analysis, characterized in that, include: Data acquisition module: used to acquire the first Qinqiang opera singing audio data of the target user in the current period and perform audio extraction to obtain the fourth pitch level and the seventh pitch level, and then calculate the difference between them and the corresponding preset natural pitch to obtain the first deviation sequence and the second deviation sequence. First Deviation Window Acquisition Module: If the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, the first time the first deviation sequence exceeds the threshold is recorded as the first starting point, and the first time the first deviation sequence after the first starting point is less than the preset first deviation threshold is recorded as the first ending point, and the first deviation window is obtained based on the first starting point and the first ending point. The second deviation window acquisition module is used to record the first time the second deviation sequence exceeds the preset second deviation threshold as the second starting point and the first time the second deviation sequence is less than the preset second deviation threshold after the second starting point as the second ending point, and to obtain the second deviation window based on the second starting point and the second ending point. Initial conversion window acquisition module: used to calculate the intersection of the first deviation window and the second deviation window to obtain the initial conversion window; The Qinqiang Opera Learning Assistance Report Output Module is used to generate and output the first Qinqiang Opera learning assistance report based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
2. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 1, characterized in that, After generating the first Qinqiang opera learning assistance report, execute the first correction module: Extract the first deviation sequence and the second deviation sequence that are within the initial transformation window to obtain the first tone-turning sequence and the second tone-turning sequence, and perform first-order difference operation on them respectively to obtain the first rate of change sequence and the second rate of change sequence. If the first rate of change sequence exceeds the first rate of change threshold, then the moment when the first rate of change sequence first exceeds the first rate of change threshold is taken as the starting point for the fourth tone refinement. If the second rate of change sequence exceeds the second rate of change threshold, then the moment when the second rate of change sequence first falls below the second rate of change threshold is taken as the starting point for the seventh tone refinement. The starting time for acquiring the first Qinqiang opera performance audio data is denoted as the starting time. Obtain the larger of the differences between the starting point of the fourth tone refinement and the starting point of the seventh tone refinement and the starting time, and determine the time corresponding to the larger of the two as the starting time of the correction. Obtain the start and end times of the initial conversion window; If the absolute difference between the correction start time and the start time of the initial conversion window exceeds a preset difference threshold, then the start time of the initial conversion window is replaced with the correction start time to obtain the first correction conversion window. The second Qinqiang opera learning assistance report is generated and output based on the ratio of the union of the first modified conversion window and the preset conversion window to the preset conversion window.
3. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 2, characterized in that, Obtain the cutoff time of the first correction conversion window and, within the subsequent preset time interval, execute the third starting point acquisition module: The ratio of high-frequency energy to low-frequency energy in each time frame of the second Qinqiang opera performance audio data within a preset time interval is obtained and calculated to generate a high-frequency energy ratio sequence. When the duration for which the high-frequency energy ratio sequence continuously exceeds the preset shouting energy threshold reaches the preset stable duration, the corresponding moment is recorded as the third starting point.
4. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 3, characterized in that, If a third starting point is obtained, then execute the second correction module: Using the third starting point as a reference point, backtrack forward for a preset backtracking time to obtain candidate backtracking cutoff points; If the difference between the cutoff time and the start time of the first corrected conversion window is greater than the difference between the backtracking cutoff candidate point and the start time, then the cutoff time of the first corrected conversion window is updated to the backtracking cutoff candidate point, and a second corrected conversion window is obtained.
5. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 4, characterized in that, After obtaining the second correction and transformation window, execute the multi-dimensional quantitative analysis module: Obtain the cutoff time of the second correction conversion window; Extract the spectral features of the audio frames of the second Qinqiang opera performance audio data corresponding to the cutoff time of the second correction conversion window to obtain the first spectral feature vector; Extract the spectral features of the audio frame of the second Qinqiang opera performance corresponding to the third starting point to obtain the second spectral feature vector; Calculate the cosine distance between the first spectral eigenvector and the second spectral eigenvector as an indicator of audio state coherence. The average slope of the linear fit of the high-frequency energy ratio sequence is calculated as a feature of the steepness of the onset of the shouting; The difference between the audio state coherence index and the preset coherence threshold is calculated to obtain the state continuity score. The difference between the steepness feature of the shouting onset and the preset steepness threshold is calculated to obtain the onset intensity score.
6. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 5, characterized in that, If the state transition score and the onset intensity score are obtained, then the third correction module is executed: The overall transition quality score is obtained by normalizing and weighting the transition score and the attack intensity score; When the overall connection quality score is lower than the preset quality threshold, the difference between the preset quality threshold and the overall connection quality score is recorded as the quality difference. The window length of the preset conversion window is shortened according to the ratio of the quality difference to the overall connection quality score and used for the next cycle.
7. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 6, characterized in that, When the overall connection quality score is lower than the preset quality threshold, it also includes: Using the overall connection quality score as the query condition, a graph traversal search is performed in the pre-constructed Qinqiang singing teaching knowledge graph to locate related error pattern nodes, correct demonstration nodes, and training method nodes; The content contained in the error mode node, correct demonstration node, and training method node is integrated, and a third Qinqiang opera learning assistance report is generated and output according to preset rules.
8. The Qinqiang Opera Speech Learning Assistance System based on Big Data Analysis according to claim 6, characterized in that, It also includes a tutoring visualization interface module: Display the fitted curves of the first deviation sequence and the second deviation sequence, and overlay the first start point, the first stop point, the second start point, and the second stop point; Display a graph of the high-frequency energy ratio sequence, and mark the third starting point; Displays an overall connection quality score; Display all Qinqiang opera learning assistance reports.
9. A method for assisting Qinqiang opera pronunciation learning based on big data analysis, characterized in that, Includes the following steps: The first Qinqiang opera singing audio data of the target user in the current period is obtained and audio is extracted to obtain the fourth pitch level and the seventh pitch level. Based on this, the difference is calculated with the corresponding preset natural pitch to obtain the first deviation sequence and the second deviation sequence. If the duration of the first deviation sequence exceeding the preset first deviation threshold exceeds the preset first duration threshold, then the first time the first deviation sequence exceeds the threshold is recorded as the first starting point, and the first time the first deviation sequence after the first starting point is less than the preset first deviation threshold is recorded as the first ending point. The first deviation window is obtained based on the first starting point and the first ending point. If the duration of the second deviation sequence exceeding the preset second deviation threshold exceeds the preset second duration threshold, then the moment when the second deviation sequence first exceeds the threshold is recorded as the second starting point, and the moment when the second deviation sequence first falls below the preset second deviation threshold after the second starting point is recorded as the second ending point. The second deviation window is obtained based on the second starting point and the second ending point. Calculate the intersection of the first deviation window and the second deviation window to obtain the initial conversion window; The first Qinqiang opera learning assistance report is generated and output based on the ratio of the union of the initial conversion window and the preset conversion window to the preset conversion window.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the system as described in any one of claims 1-8.