Artificial intelligence-based vocal music training method and system

By collecting and analyzing the body posture and pronunciation characteristic parameters of vocal trainees, an abnormal training set is generated and visualized, which solves the problem that existing technologies cannot intuitively display posture and pronunciation abnormalities, and achieves scientific and precise vocal training results.

CN120636465BActive Publication Date: 2026-02-17SHANGLUO UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510954824.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-11
Publication Date
2026-02-17
Estimated Expiration
2045-07-11

AI Technical Summary

Technical Problem

Existing vocal training analysis techniques cannot intuitively display trainees' body posture and vocal abnormalities, making it difficult for trainees to self-correct.

Method used

By collecting body posture parameters and vocal feature parameters of the target vocal trainees, an abnormal training set is generated using artificial intelligence, and then visualized to provide a correction set.

Benefits of technology

It provides a visual representation of posture and pronunciation abnormalities, offering in-depth insights to help trainees make scientific and precise adjustments and corrections, thus improving the scientific nature and accuracy of vocal training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636465B_ABST
    Figure CN120636465B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of voice analysis, and particularly discloses a vocal music training method and system based on artificial intelligence, which collects and analyzes body posture parameters of a target vocal music training person, detects posture abnormalities after preprocessing, forms a first abnormal training set, simultaneously captures and analyzes pronunciation feature parameters, identifies pronunciation abnormalities, forms a second abnormal training set, comprehensively generates an abnormal parameter training total set, and matches a correction set, and the abnormal degree and causes of body posture and the abnormal degree and causes of pronunciation are visually displayed, so that the target vocal music training person can directly see his / her own posture and pronunciation problems, and is facilitated to make self-adjustment and correction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech analysis technology, specifically to a vocal training method and system based on artificial intelligence. Background Technology

[0002] With the rapid development of science and technology, artificial intelligence technology has permeated all areas of society. The introduction of artificial intelligence technology has brought innovative development to the traditional vocal music teaching model, providing a more personalized, efficient and scientific vocal music training experience, and bringing new opportunities for the innovation and development of vocal music education.

[0003] For example, the invention patent with announcement number CN106448701B announces a comprehensive vocal training system, including a sound acquisition module, a laryngeal position detection module, a breathing frequency acquisition module, a feature signal acquisition module, a data processing module, a pitch and rhythm extraction module, an overtone feature extraction module, a note pitch and note duration model establishment module, a score generation module, a primary singing skill assessment module, a comprehensive singing skill assessment module, a comprehensive training program output module, a mathematical model establishment module, a virtual parameter actuation module, a virtual sensor, and a simulation analysis module.

[0004] For example, the invention patent with publication number CN119229895A discloses a vocal pronunciation training system, which relates to the field of vocal training technology. The vocal pronunciation training system includes a sound acquisition module, a sound preprocessing module, a feature extraction module, a model construction module, and a standard pronunciation database module. The system uses standard vocal pronunciation data to produce pronunciations through the trainee's personalized voice model. This can intuitively show the difference between the trainee's current vocal pronunciation and the standard vocal pronunciation. The system also analyzes the difference coefficient between the trainee's voice and the standard pronunciation through a comparative analysis module, thereby clearly understanding the size of the knowledge gap and developing different types of vocal pronunciation training programs based on the size of the difference.

[0005] However, in the process of implementing the embodiments of this application, it was found that the above-mentioned technology has at least the following technical problems: In the existing vocal training analysis technology, although the existing analysis technology can capture data such as the trainee's movements, when presenting this data, it usually only provides the original records or simple analysis results, which often leads to trainees having to rely on self-observation or guidance generated by the analysis results to make adjustments, rather than intuitively understanding the abnormalities in their own movements. Summary of the Invention

[0006] In view of the shortcomings of the prior art, the present invention provides a vocal training method and system based on artificial intelligence, which can effectively solve the problems involved in the above-mentioned background art.

[0007] To achieve the above objectives, the present invention provides the following technical solution: The first aspect of the present invention provides a vocal training method based on artificial intelligence, comprising: collecting a set of body posture parameters of a target vocal trainee and labeling it as a target first dataset; using a first artificial intelligence scheme to determine whether to preprocess the target first dataset, wherein the set of body posture parameters is a sequence of video frames from the vocal training process of the target vocal trainee; based on the target first dataset, using a second artificial intelligence scheme to determine whether to generate a first abnormal training set for the target vocal trainee, wherein the first abnormal training set is a set used to record abnormal body postures of the target vocal trainee during vocal training; obtaining the pronunciation feature parameters of the target vocal trainee and using a third artificial intelligence scheme to determine whether... Whether to generate a second abnormal training set for the target vocal trainee, the second abnormal training set being a collection used to record pronunciation abnormalities of the target vocal trainee during vocal training; combining the first and second abnormal training sets of the target vocal trainee to generate a total abnormal parameter training set for the target vocal trainee; using a fourth artificial intelligence scheme to match a correction set for the target vocal trainee; and visually displaying the correction set and the total abnormal parameter training set of the target vocal trainee. The total abnormal parameter training set is a set composed of the first and second abnormal training sets, and the correction set is a set of methods used to correct abnormalities that occur in the target vocal trainee during vocal training.

[0008] As a further method, the determination of whether to generate a first abnormal training set for the target vocal trainee using the second artificial intelligence scheme specifically involves: obtaining the second attribute parameters of the target first dataset; evaluating the posture deviation assessment index of the target vocal trainee based on the second attribute parameters of the target first dataset; matching the posture deviation assessment threshold of the target vocal trainee from the vocal database based on the body posture attribute parameter set of the target vocal trainee; comparing the posture deviation assessment index of the target vocal trainee with the posture deviation assessment threshold of the target vocal trainee; if the posture deviation assessment index of the target vocal trainee is less than or equal to the posture deviation assessment threshold of the target vocal trainee, then no first abnormal training set for the target vocal trainee is generated; if the posture deviation assessment index of the target vocal trainee is greater than the posture deviation assessment threshold of the target vocal trainee, then a first abnormal training set for the target vocal trainee is generated; the specific process for generating the first abnormal training set of the target vocal trainee is as follows: from the posture deviation of the target vocal trainee... The evaluation index extracts the posture deviation evaluation factors of the target vocal trainee at each moment during the training period. The posture deviation evaluation factors and the posture deviation evaluation threshold of the target vocal trainee at each moment during the training period are then processed by difference and absolute value analysis to obtain the posture deviation evaluation difference of the target vocal trainee at each moment during the training period. Contour curves of the target vocal trainee at each moment during the training period are extracted from the second attribute parameters of the first target dataset. Contour curves of the target vocal trainee at historical adjacent moments during the training period are obtained. These historical adjacent moment contour curves are then integrated with the contour curves of the target vocal trainee at each moment during the training period to obtain the contour curve of the target vocal trainee at each moment during the training period minus the historical adjacent contour curve. Finally, the posture deviation evaluation difference of the target vocal trainee at each moment during the training period and the contour curve of the target vocal trainee at each moment during the training period minus the historical adjacent contour curve are combined to generate the first abnormal training set of the target vocal trainee.

[0009] As a further method, the process of using a third artificial intelligence scheme to determine whether to generate a second abnormal training set for the target vocal trainee is as follows: Based on the vocal characteristic parameters of the target vocal trainee, analyze the vocal audio stability factor of the target vocal trainee; extract the vocal audio stability threshold from the vocal database; compare the vocal audio stability factor of the target vocal trainee with the vocal audio stability threshold; if the vocal audio stability factor of the target vocal trainee is greater than the vocal audio stability threshold, then it is determined not to generate a second abnormal training set for the target vocal trainee; if the vocal audio stability factor of the target vocal trainee is less than or equal to the vocal audio stability threshold, then it is determined to generate a second abnormal training set for the target vocal trainee; the generation of the second abnormal training set for the target vocal trainee... The second abnormal training set is generated as follows: The pronunciation audio stability factor of the target vocal trainee is matched with the pronunciation abnormality level corresponding to each pronunciation audio stability factor interval stored in the vocal database to obtain the pronunciation abnormality level of the target vocal trainee; the fundamental frequency of the target vocal trainee at each moment during the training period is extracted from the pronunciation feature parameters of the target vocal trainee and averaged to obtain the average fundamental frequency of the target vocal trainee during the training period; a reference average fundamental frequency of the target vocal trainee during the training period is obtained and compared with the average fundamental frequency of the target vocal trainee during the training period to obtain the average fundamental frequency comparison result; and the second abnormal training set of the target vocal trainee is generated based on the pronunciation abnormality level and the average fundamental frequency comparison result.

[0010] A second aspect of the present invention provides an artificial intelligence-based vocal training system, comprising: a data preprocessing module, configured to collect a set of body posture parameters of a target vocal trainee and label it as a target first dataset, and use a first artificial intelligence scheme to determine whether to preprocess the target first dataset, wherein the set of body posture parameters is a sequence of video frames from the vocal training process of the target vocal trainee; a first abnormal training set generation module, configured to, based on the target first dataset, use a second artificial intelligence scheme to determine whether to generate a first abnormal training set for the target vocal trainee, wherein the first abnormal training set is a set used to record abnormal body postures of the target vocal trainee during vocal training; and a second abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and use a third artificial intelligence scheme to determine whether to preprocess the target vocal trainee's body posture parameters ... third abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and use a third artificial intelligence scheme to determine whether to preprocess the target vocal trainee's body posture parameters; and a fourth abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and use a third artificial intelligence scheme to determine whether to preprocess the target vocal trainee's body posture parameters; and a fifth abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and use a third artificial intelligence scheme to determine whether to preprocess the target vocal trainee's body posture parameters; and a sixth abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and use a third artificial intelligence scheme to determine whether to preprocess the target vocal trainee's body posture parameters; and a seventh abnormal training set generation module, configured to acquire the pronunciation feature parameters of the target vocal trainee and acquire the pronunciation feature parameters of the target vocal trainee. The system determines whether to generate a second abnormal training set for the target vocal trainee. The second abnormal training set is a collection used to record pronunciation abnormalities of the target vocal trainee during vocal training. The visualization module is used to generate a total abnormal parameter training set for the target vocal trainee by combining the first and second abnormal training sets. The fourth artificial intelligence scheme is used to match the correction set for the target vocal trainee. The correction set and the total abnormal parameter training set of the target vocal trainee are visualized. The total abnormal parameter training set is a set composed of the first and second abnormal training sets. The correction set is a set of methods used to correct abnormalities that occur in the target vocal trainee during vocal training.

[0011] Compared with the prior art, the embodiments of the present invention have at least the following advantages or beneficial effects:

[0012] (1) This invention provides a vocal training method and system based on artificial intelligence. By collecting and analyzing the body posture parameters of the target vocal trainee, and after preprocessing, abnormal posture is detected to form a first abnormal training set. At the same time, the pronunciation feature parameters are captured and analyzed to identify pronunciation abnormalities and form a second abnormal training set. By combining the two, a total abnormal parameter training set is generated and a correction set is matched. The degree and cause of abnormal body posture and the degree and cause of abnormal pronunciation are displayed through visualization. The target vocal trainee can intuitively see his / her posture and pronunciation problems, which is convenient for self-adjustment and correction.

[0013] (2) This invention analyzes the posture deviation assessment index of the target vocal training personnel through data analysis. It not only reveals the existence of deviation, but also accurately locates the specific degree of deviation through quantitative analysis, providing trainees with unprecedented in-depth insight. This insight is incomparable to visual observation and makes vocal training more scientific and precise.

[0014] (3) This invention combines the posture deviation assessment index of the target vocal trainee with the pronunciation audio stability factor, which not only accurately quantifies the stability of the pronunciation audio, but also reveals the direct impact of posture deviation on pronunciation quality. This allows us to start from the underlying causes affecting pronunciation, accurately locate and correct the deficiencies in vocal trainees' vocal performance, clearly understand how posture deviation subtly affects the pronunciation process, and then take targeted improvement measures to fundamentally improve the stability and accuracy of pronunciation. Attached Figure Description

[0015] The present invention will be further described with reference to the accompanying drawings, but the embodiments in the drawings do not constitute any limitation on the present invention. For those skilled in the art, other drawings can be obtained based on the following drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the method steps of the present invention.

[0017] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0018] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0019] Reference Figure 1 As shown, the first aspect of the present invention provides a vocal training method based on artificial intelligence, comprising: collecting a set of body posture parameters of a target vocal trainee and labeling it as a target first dataset; and using a first artificial intelligence scheme to determine whether to preprocess the target first dataset, wherein the set of body posture parameters is a sequence of video frames of the vocal training process of the target vocal trainee.

[0020] Specifically, the process of determining whether to preprocess the target first dataset using the first artificial intelligence solution is as follows: First attribute parameters of the target first dataset are obtained; based on these parameters, a data collection quality factor is analyzed; a data collection quality threshold is extracted from the vocal database; the data collection quality factor of the target first dataset is compared with the data collection quality threshold; if the data collection quality factor is greater than the threshold, preprocessing is not required; if the data collection quality factor is less than or equal to the threshold, preprocessing is required; the aforementioned data collection quality threshold represents the minimum value within a reasonable range for the data collection quality factor.

[0021] Specifically, the data collection quality factor of the target first dataset is analyzed as follows: the first attribute parameter of the target first dataset includes the motion coordinates of each target point in the target first dataset; the target first dataset contains a sequence of video frames recording the vocal training process of the target vocal trainees through a camera; each target point refers to the key position points of the torso of the target vocal trainees identified and located in the target first dataset using image processing and positioning algorithms (such as human key point recognition algorithms), and marked as each target point. These target points usually cover multiple important landmarks on the torso, such as shoulders, waist, and hips, aiming to comprehensively and accurately reflect the body posture characteristics of the trainees; the motion coordinates refer to the coordinate position points of the target points in the video frames, which can be obtained through image recognition technology (such as corner detection technology).

[0022] The motion coordinates of all target points belonging to the first dataset of the target are combined and marked as the motion coordinate set of each target point belonging to the first dataset of the target. Specifically, all motion coordinates of each target point are summarized to form an independent set for each target point, and this independent set is marked as the motion coordinate set.

[0023] The process involves obtaining the body posture attribute parameter set of the target vocal trainee and matching the motion coordinate reference set of each target point from the vocal database. The motion coordinate set of each target point in the first target dataset is compared with the motion coordinate reference set. This comparison analysis identifies abnormal motion coordinate points for each target point in the first target dataset. The total number of abnormal motion coordinate points in each target point in the first target dataset is counted and marked as the total number of abnormal motion coordinate points in the first target dataset. The ratio of the total number of abnormal motion coordinate points in the first target dataset to the total number of motion coordinate points in the first target dataset is then calculated to obtain the motion abnormality rate for each target point in the first target dataset. The aforementioned body posture attribute parameter set of the target vocal trainee refers to the collection of body posture-related information input by the target vocal trainee, including but not limited to weight, waist circumference, height, and... Shoulder width; The motion coordinate reference sets of each target point matched above are specifically matched as follows: The vocal database stores the motion coordinate reference sets of each target point corresponding to each body posture attribute parameter set. A similarity algorithm (such as cosine similarity) is used to calculate the similarity between the body posture attribute parameter set of the target vocal trainee and the body posture attribute parameter sets in the vocal database. The body posture attribute parameter set in the vocal database with the highest similarity to the target trainee's parameter set is selected, and the motion coordinate reference set of each target point corresponding to the body posture attribute parameter set in the vocal database is used as the motion coordinate reference set of each matched target point; The specific comparison process for the above abnormal motion coordinate points is as follows: The motion coordinate sets of each target point in the first target dataset are compared with the motion coordinate reference sets of each target point matched from the vocal database. If a motion coordinate point does not appear in the corresponding motion coordinate reference set, it is determined to be an abnormal motion coordinate point; The motion anomaly rate of each target point in the first target dataset represents the proportion of abnormal motion coordinate points in each target point in the first target dataset.

[0024] The total number of motion coordinates of each target point in the first target dataset is processed by standard deviation, and the result is marked as the distribution discrete value of the target points in the first target dataset. This is used to quantify the statistical dispersion of the total number of target points in the target dataset.

[0025] The process involves obtaining the timestamps of each motion coordinate point of each target point belonging to the first target dataset, analyzing the data to determine the average data update duration of each target point in the first target dataset, obtaining the reference data update duration of each target point in the first target dataset from the vocal database, and performing difference and absolute value processing on the reference and reference data update durations. This yields the data update deviation duration of each target point in the first target dataset, which is then summed to obtain the total data update deviation duration. The timestamps mentioned above are the recording timestamps, which can be extracted from the timestamp log. The average data update duration of each target point is determined through the following data analysis process: for each target point, the duration between the timestamps of its adjacent motion coordinate points is calculated, i.e., the data update duration of each target point, and then the average is calculated to obtain the average data update duration of each target point. The reference data update duration represents the reference value of the average data update duration. The data update deviation duration represents the absolute value of the difference between the average data update duration and the reference data update duration. The total data update deviation duration represents the cumulative result of the data update deviation duration.

[0026] By combining the motion anomaly rate of each target point in the first target dataset, the distribution dispersion value of the target points in the first target dataset, and the total data update deviation time of the target points in the first target dataset, a data collection quality factor for the first target dataset is derived. The data collection quality factor for the first target dataset represents the quantitative data on the degree of influence of the motion anomaly rate, distribution dispersion value, and total data update deviation time on the data quality collected by the first target dataset, and is used to quantify the data quality collected by the first target dataset.

[0027] The specific analysis method for the data collection quality factor of the target first dataset is as follows:

[0028]

[0029] In the formula, PT_DQ is the data collection quality factor of the first target dataset, d is the number of each target point, d={1,2,3,...,w}, w is the total number of target points, and MA_DQ d Let BP_DQ be the motion anomaly rate of the d-th target point in the first dataset, CQ_DQ be the discrete distribution value of the target point in the first dataset, B_F1 be the normalized influence factor of motion anomaly rate preset in the vocal database, B_F2 be the normalized influence factor of discrete distribution value preset in the vocal database, and B_F3 be the normalized influence factor of total data update deviation time preset in the vocal database.

[0030] The aforementioned motion anomaly rate normalization influence factor is used to normalize the motion anomaly rate, quantifying the influence of a unit value of the motion anomaly rate on the data collection quality factor; the aforementioned distribution discrete value normalization influence factor is used to normalize the distribution discrete value, quantifying the influence of a unit value of the distribution discrete value on the data collection quality factor; the aforementioned data update total deviation duration normalization influence factor is used to normalize the data update total deviation duration, quantifying the influence of a unit value of the data update total deviation duration on the data collection quality factor; the vocal database stores the correspondence between the motion anomaly rate, distribution discrete value, and data update total deviation duration and their corresponding normalization influence factors. For example, by inputting the stored motion anomaly rate, distribution discrete value, and data update total deviation duration into the vocal database, the vocal database can match the motion anomaly rate normalization influence factor, distribution discrete value normalization influence factor, and data update total deviation duration normalization influence factor, all with values ​​ranging from 0 to 1.

[0031] It needs to be explained that during the collection of the target's first dataset, the collection equipment (such as cameras) may be constrained by multiple external conditions, including environmental network factors, leading to issues such as missing frames in the data collection. By quantifying the data collection quality, targeted preprocessing schemes can be proposed, thereby avoiding the inaccuracies that may arise from a uniform preprocessing strategy. In the process of quantifying data collection quality, there is a close and complex interrelationship and influence among the target point's motion anomaly rate, distribution dispersion, and total data update deviation time. A high motion anomaly rate often indicates that there are more discontinuities or abnormal changes during the data collection process. This may lead to a more discrete distribution of the data in the time series, i.e., an increase in the distribution dispersion. At the same time, an increase in the motion anomaly rate may also be accompanied by... The extended duration of the total data update deviation is due to the fact that abnormal motion data often requires more time to correct or compensate, thus affecting the normal update frequency. Conversely, the increased duration of the total data update deviation may further exacerbate the motion anomaly rate, as untimely data updates may cause abnormal states to be continuously recorded without timely correction. The increased dispersion of the distribution also increases the difficulty of synchronizing data at each target point, which may lead to more inconsistencies or redundant information in the target's primary dataset, which is regarded as "noise" in the dataset. This not only increases the complexity of subsequent data processing but may also interfere with the accurate judgment of the motion state of the target points, thereby reducing the overall data collection quality of the target's primary dataset. This noise may manifest as frame rate fluctuations, discontinuities between frames, or abnormal changes in image content.

[0032] Furthermore, the preprocessing of the target first dataset specifically involves the following steps: The vocal database stores a first preprocessing scheme corresponding to data collection quality factor interval one, a second preprocessing scheme corresponding to data collection quality factor interval two, and a third preprocessing scheme corresponding to data collection quality factor interval three; the data collection quality factors of the target first dataset are matched with the data collection quality factor intervals one, two, and three stored in the vocal database; if the data collection quality factors of the target first dataset belong to data collection quality factor interval one, then the first preprocessing scheme for the target first dataset is obtained, and the target first dataset is preprocessed according to the first preprocessing scheme ... If the data collection quality factor belongs to the second data collection quality factor interval, then a second preprocessing scheme for the target first dataset is matched, and the target first dataset is preprocessed according to the second preprocessing scheme. If the data collection quality factor of the target first dataset belongs to the third data collection quality factor interval, then a third preprocessing scheme for the target first dataset is matched, and the target first dataset is preprocessed according to the third preprocessing scheme. In an example embodiment, assuming the data collection quality factor of the target first dataset is G, which belongs to the second data collection quality factor interval [G-20%, G+30%], the matched second preprocessing scheme for the target first dataset includes: frame interpolation and restoration: the interpolation algorithm type is spline interpolation, and the interpolation window size is set to... The interpolation threshold is H (used to determine whether interpolation is needed). Noise filtering and smoothing: the filter type is a low-pass filter, the filter strength is dynamically adjusted according to the noise level in the G value (the strength needs to be increased by G+10% when the noise is high), and the smoothing window size is... Outlier detection and handling: The outlier detection threshold is K (used to determine whether replacement is needed; if the outlier is greater than the outlier detection threshold, replacement is required). The replacement strategy is to use the average value of adjacent frames for replacement.

[0033] Based on the first target dataset, the second artificial intelligence scheme is used to determine whether to generate a first abnormal training set for the target vocal trainee. The first abnormal training set is a collection used to record abnormal body postures of the target vocal trainee during vocal training.

[0034] In one specific embodiment, the present invention analyzes data to determine the posture deviation assessment index of the target vocal trainee. This not only reveals the existence of the deviation but also, through quantitative analysis, precisely pinpoints the specific degree of the deviation, providing trainees with unprecedented depth of insight. This insight is incomparable to visual observation and makes vocal training more scientific and precise.

[0035] Furthermore, the specific process for determining whether to generate a first abnormal training set for the target vocal trainee using the second artificial intelligence scheme is as follows: First, obtain the second attribute parameters of the target first dataset; second, evaluate the posture deviation assessment index of the target vocal trainee based on the second attribute parameters of the target first dataset; third, match the posture deviation assessment threshold of the target vocal trainee from the vocal database based on the target vocal trainee's body posture attribute parameter set; fourth, compare the posture deviation assessment index of the target vocal trainee with the posture deviation assessment threshold of the target vocal trainee; if the posture deviation assessment index of the target vocal trainee is less than or equal to the posture deviation assessment threshold of the target vocal trainee, then the target vocal trainee is not generated. The first abnormal training set of the trainees; the posture deviation assessment threshold of the target vocal trainees, which represents the maximum value of the reasonable range of the posture deviation assessment index of the target vocal trainees. The specific matching process is as follows: the posture deviation assessment thresholds corresponding to each body posture attribute parameter set are stored in the vocal database. The similarity between the body posture attribute parameter set of the target vocal trainees and each body posture attribute parameter set in the vocal database is calculated using a similarity algorithm (such as cosine similarity). The body posture attribute parameter set in the vocal database with the highest similarity to the parameter set of the target trainees is selected, and the posture deviation assessment threshold corresponding to the body posture attribute parameter set in the vocal database is used as the posture deviation assessment threshold of the target vocal trainees.

[0036] If the posture deviation assessment index of the target vocal trainee is greater than the posture deviation assessment threshold of the target vocal trainee, then the first abnormal training set of the target vocal trainee is generated. The specific generation process is as follows: extract the posture deviation assessment factor of the target vocal trainee at each moment in the training period from the posture deviation assessment index of the target vocal trainee, and perform difference and absolute value processing on the posture deviation assessment factor of the target vocal trainee at each moment in the training period and the posture deviation assessment threshold of the target vocal trainee in turn to obtain the posture deviation assessment difference of the target vocal trainee at each moment in the training period; the above posture deviation assessment difference is used to quantify the degree of posture deviation of the target vocal trainee at each moment in the training period.

[0037] The contour curves of the target vocal trainee at each moment within the training period are extracted from the second attribute parameters of the first target dataset. Contour curves of the target vocal trainee at historically adjacent moments within the training period are also obtained. These historically adjacent contour curves are then integrated with the contour curves of the target vocal trainee at each moment within the training period to obtain the contour curve of the target vocal trainee at each moment within the training period minus the historically adjacent contour curve. The aforementioned contour curve minus the historically adjacent contour curve refers to presenting the contour curve at each moment together with the contour curve of the corresponding historically adjacent moment. The contour curve at each moment is represented by a solid line, and the corresponding historically adjacent contour curve is represented by a solid line. The contour curves at different times are represented by dashed lines, providing contextual information about the changes in the target vocal trainee's body posture over time. For example, within a training cycle, T1 and T2 are adjacent times, and T1 is a historical adjacent time of T2. The contour curve of the target vocal trainee at time T2 is drawn with a solid line, and the contour curve of the target vocal trainee at time T1 is drawn with a dashed line using the same coordinate system. In this way, the contour curve of each specific time forms a clear visual contrast with the contour curves of its historical adjacent times. The contour curves of the target vocal trainee at the historical adjacent times within the training cycle can be obtained by directly querying the corresponding contour curves of the historical adjacent times.

[0038] The first abnormal training set of the target vocal trainee is generated by combining the posture deviation assessment difference of the target vocal trainee at each moment during the training period and the contour curve-historical adjacent contour curve of the target vocal trainee at each moment during the training period; wherein, in the first abnormal training set of the target vocal trainee, the posture deviation assessment difference at each moment is marked on the contour curve-historical adjacent contour curve of the corresponding moment.

[0039] Specifically, the evaluation process for the postural deviation assessment index of the target vocal trainee is as follows: the second attribute parameters of the target first dataset include the contour curve of the target vocal trainee at each moment during the training cycle, the key point coordinates of the target vocal trainee at each moment during the training cycle, and the chest movement frequency of the target vocal trainee at each moment during the training cycle; the aforementioned training cycle represents the time period during which the target vocal trainee conducts vocal training. Generally, each training cycle focuses on a specific type of vocal training. For example, in one training cycle, the target vocal trainee will focus on practicing... The pronunciation technique for the vowel 'ah'; the aforementioned contour curve, a curve describing the edge of the target vocal trainee, can be obtained through image recognition technology (such as nonlinear dimensionality reduction image recognition technology); the aforementioned key point refers to a point randomly selected from each target point using a random algorithm (such as a random number algorithm); the aforementioned key point coordinates refer to the coordinate position of the key point in the video frame of the first dataset of the target, which can be obtained through image recognition technology (such as corner detection technology); the aforementioned chest cavity movement frequency, i.e., the number of chest cavity movements per moment, can be obtained by the target vocal trainee wearing devices such as accelerometers.

[0040] It should be noted that the target dataset may have been preprocessed, therefore the keypoint coordinates need to be re-derived.

[0041] By comparing the contour curves of the target vocal trainee at each moment during the training period with the corresponding historical adjacent contour curves, the contour deviation area of ​​the target vocal trainee at each moment during the training period is obtained. The specific comparison process is as follows: the comparison is achieved through image registration and other techniques. Based on the comparison, image processing software (such as Matrix Lab) is used to calculate the difference area between the region contained in the contour curve at each moment and the region contained in the historical adjacent contour curve. This difference area is the contour deviation area of ​​the target vocal trainee at each moment during the training period, which is used to quantify the degree of contour change of the target vocal trainee at each moment during the training period.

[0042] By performing distance analysis between the key point coordinates of the target vocal trainee at each moment during the training period and the corresponding historical adjacent key point coordinates, the displacement of the key points of the target vocal trainee at each moment during the training period is obtained, which represents the distance between the key point coordinates of the target vocal trainee at each moment during the training period and the corresponding historical adjacent key point coordinates.

[0043] The vocal frequency of the target vocal trainee at each moment during the training period is extracted from the vocal characteristic parameters of the target vocal trainee. The vocal frequency and chest movement frequency of the target vocal trainee at each moment during the training period are normalized to obtain the vocal frequency-chest movement frequency factor of the target vocal trainee at each moment during the training period, which represents the relationship coefficient between the vocal frequency and chest movement frequency of the target vocal trainee at each moment during the training period.

[0044] Based on the data collection quality factors of the target first dataset, the posture deviation correction values ​​of the target vocal trainees are matched from the vocal database. The posture deviation correction value is used to quantify the degree of correction to the posture deviation assessment index of the target vocal trainees. The specific matching process is as follows: the vocal database stores the posture deviation correction values ​​corresponding to each data collection quality factor interval. The data collection quality factor interval of the target first dataset is queried from the vocal database. The posture deviation correction value corresponding to the data collection quality factor interval stored in the vocal database is the posture deviation correction value of the target vocal trainees.

[0045] Based on the body posture attribute parameter set of the target vocal trainee, the bounding contour deviation area, bounding key point displacement, and reference vocal frequency-chest movement frequency factor are matched from the vocal database. The specific matching process is as follows: The vocal database stores the bounding contour deviation area, bounding key point displacement, and reference vocal frequency-chest movement frequency factor corresponding to each body posture attribute parameter set. The similarity between the target vocal trainee's body posture attribute parameter set and each body posture attribute parameter set in the vocal database is calculated using a similarity algorithm (such as cosine similarity). The body posture attribute parameter set in the vocal database with the highest similarity to the target trainee's parameter set is selected, and the bounding contour deviation area, bounding key point displacement, and reference vocal frequency-chest movement frequency factor corresponding to the body posture attribute parameter set in the vocal database are used as the matched bounding contour deviation area, bounding key point displacement, and reference vocal frequency-chest movement frequency factor.

[0046] By comprehensively analyzing the contour deviation area, key point displacement, vocal frequency-chest movement frequency factor, posture deviation correction value, defined contour deviation area, defined key point displacement, and reference vocal frequency-chest movement frequency factor at various times during the training cycle of the target vocal trainee, a posture deviation assessment index is derived. This index represents the quantitative data on the influence of contour deviation area, key point displacement, vocal frequency-chest movement frequency factor, posture deviation correction value, defined contour deviation area, defined key point displacement, and reference vocal frequency-chest movement frequency factor on the posture deviation of the target vocal trainee, and is used to comprehensively quantify the degree of posture deviation of the target vocal trainee.

[0047] The posture deviation assessment index of the target vocal trainee is specifically obtained by correcting the first, second, and third results using posture deviation correction values. The first result represents the influence of the deviation of the contour deviation area component on the posture deviation assessment index, specifically obtained through a comprehensive analysis of the contour deviation area, the defined contour deviation area, and the normalized influence factor of the contour deviation area. The second result represents the influence of the deviation of the key point displacement component on the posture deviation assessment index, specifically obtained through a comprehensive analysis of the key point displacement, the defined key point displacement, and the normalized influence factor of the key point displacement. The third result represents the influence of the deviation of the vocal frequency-chest movement frequency factor component on the posture deviation assessment index, specifically obtained through a comprehensive analysis of the vocal frequency-chest movement frequency factor, the reference vocal frequency-chest movement frequency factor, and the normalized influence factor of the vocal frequency-chest movement frequency factor. The specific assessment method for the posture deviation assessment index of the target vocal trainee is as follows:

[0048]

[0049]

[0050] In the formula, AD_SEC is the posture deviation assessment index of the target vocal trainee, and AD_ST a CDA is an assessment factor for postural deviation of the target vocal trainee at time a within the training cycle. a VKP represents the area of ​​contour deviation of the target vocal trainee at time a within the training period. a VFT represents the key point displacement of the target vocal trainee at time 'a' within the training cycle. aLet ΔCDAJ be the vocal frequency-chest movement frequency factor of the target vocal trainee at time a within the training cycle, ΔVKPJ be the defined contour deviation area, ΔVFT be the defined key point displacement, ΔVFT be the reference vocal frequency-chest movement frequency factor, fr1 be the preset contour deviation area normalization influence factor in the vocal database, fr2 be the preset key point displacement normalization influence factor in the vocal database, fr3 be the preset vocal frequency-chest movement frequency factor normalization influence factor in the vocal database, T be the posture deviation correction value of the target vocal trainee, a represent each time within the training cycle, a∈[a1, a2], a1 is the start time of the training cycle, a2 is the end time of the training cycle, and ADC be the posture deviation assessment threshold of the target vocal trainee.

[0051] The above definition of contour deviation area represents the maximum allowable area of ​​contour deviation; the above definition of key point displacement represents the maximum allowable displacement of key points; the above reference phonation frequency-chest movement frequency factor represents the reference value of the phonation frequency-chest movement frequency factor.

[0052] The aforementioned contour deviation area normalization influence factor is used to normalize the contour deviation area and quantify the influence of the unit value of the contour deviation area on the posture deviation assessment index. The aforementioned key point displacement normalization influence factor is used to normalize the key point displacement and quantify the influence of the unit value of the key point displacement on the posture deviation assessment index. The aforementioned vocal frequency-chest movement frequency factor normalization influence factor is used to normalize the vocal frequency-chest movement frequency factor and quantify the influence of the vocal frequency-chest movement frequency factor on the posture deviation assessment index. The vocal database stores the correspondence between the contour deviation area, key point displacement, and vocal frequency-chest movement frequency factor and their corresponding normalization influence factors. For example, by inputting the contour deviation area, key point displacement, and vocal frequency-chest movement frequency factor into the vocal database, the vocal database can match the contour deviation area normalization influence factor, key point displacement normalization influence factor, and vocal frequency-chest movement frequency factor normalization influence factor, all of which have values ​​between 0 and 1.

[0053] It needs to be explained that during targeted vocal training, incorrect vocal posture has a significant negative impact on training effectiveness. Incorrect vocal posture often leads to an increase in the area of ​​contour deviation, reflecting instability and lack of control in the trainee's body posture during vocalization. When the area of ​​contour deviation increases, it indicates significant positional changes in body parts related to vocalization, which may include twisting, tilting, or swaying of the body. These changes all lead to increased displacement of key points, indicating problems with the coordination and stability of the trainee's body posture during vocalization. Furthermore, unstable body posture also adversely affects the coordination between vocal frequency and chest cavity movement frequency. In the process of correct vocalization, the vocalization... The vocal frequency and chest movement frequency should maintain a certain degree of coordination and stability. However, when the trainee's torso sways abnormally, their chest movement frequency may become abnormally intense, causing the ratio between the vocal frequency and chest movement frequency to deviate significantly from the corresponding reference value. More importantly, these effects are not isolated but interconnected. For example, incorrect vocal posture may simultaneously lead to an increase in the area of ​​contour deviation and an increase in the displacement of key points; while unstable body posture may simultaneously affect the coordination between the vocal frequency and chest movement frequency and the change in the area of ​​contour deviation. This allows for the accurate capture and identification of abnormal torso swaying that is difficult to detect with the naked eye, ensuring the accuracy and effectiveness of vocal training.

[0054] The pronunciation feature parameters of the target vocal trainee are obtained, and a third artificial intelligence scheme is used to determine whether to generate a second abnormal training set of the target vocal trainee. The second abnormal training set is a set used to record the pronunciation abnormalities of the target vocal trainee during vocal training.

[0055] In one specific embodiment, this invention combines the postural deviation assessment index of the target vocal trainee with the vocal audio stability factor. This not only precisely quantifies the stability of the vocal audio but also reveals the direct impact of postural deviation on vocal quality. This allows for precise identification and correction of vocal trainees' shortcomings in vocal performance by addressing the underlying causes affecting vocal quality. It provides a clear understanding of how postural deviation subtly influences the vocal process, enabling targeted improvement measures to fundamentally enhance vocal stability and accuracy. Given the differences in physical conditions among trainees, this invention personalizes and matches the required parameters for vocal training based on the target vocal trainee's set of body posture attribute parameters. This improves the relevance and efficiency of vocal training, effectively avoiding the generalized guidance problem prevalent in traditional training—that is, avoiding the shortcomings of using uniform standards while ignoring individual differences.

[0056] Furthermore, the specific process for determining whether to generate a second abnormal training set for the target vocal trainee using the third artificial intelligence solution is as follows: Based on the vocal characteristic parameters of the target vocal trainee, analyze the vocal audio stability factor of the target vocal trainee; extract the vocal audio stability threshold from the vocal database; compare the vocal audio stability factor of the target vocal trainee with the vocal audio stability threshold; if the vocal audio stability factor of the target vocal trainee is greater than the vocal audio stability threshold, then it is determined not to generate a second abnormal training set for the target vocal trainee; if the vocal audio stability factor of the target vocal trainee is less than or equal to the vocal audio stability threshold, then it is determined to generate a second abnormal training set for the target vocal trainee; the aforementioned vocal audio stability threshold represents the minimum value within the reasonable range of the vocal audio stability factor of the target vocal trainee.

[0057] The specific process for generating the second abnormal training set of the target vocal trainee is as follows: matching the pronunciation audio stability factor of the target vocal trainee with the pronunciation abnormality level corresponding to each pronunciation audio stability factor interval stored in the vocal database to obtain the pronunciation abnormality level of the target vocal trainee, which is used to quantify the degree of pronunciation abnormality of the target vocal trainee. The specific matching process is as follows: querying the pronunciation audio stability factor interval stored in the vocal database to which the pronunciation audio stability factor of the target vocal trainee belongs, and the pronunciation abnormality level corresponding to the pronunciation audio stability factor interval stored in the vocal database is the pronunciation abnormality level of the target vocal trainee.

[0058] The fundamental frequency of the target vocal trainee at each moment during the training period is extracted from the vocal feature parameters of the target vocal trainee, and averaged to obtain the average fundamental frequency of the target vocal trainee during the training period. A reference average fundamental frequency of the target vocal trainee during the training period is obtained and compared with the average fundamental frequency of the target vocal trainee during the training period to obtain the average fundamental frequency comparison result. A second abnormal training set for the target vocal trainee is generated based on the vocal abnormality level and the average fundamental frequency comparison result. The aforementioned average fundamental frequency represents the average value of the fundamental frequency; the aforementioned reference average fundamental frequency represents the reference value of the average fundamental frequency; average... The fundamental frequency comparison results are specifically: Comparison Result 1 (average fundamental frequency is greater than the reference average fundamental frequency), Comparison Result 2 (average fundamental frequency is equal to the reference average fundamental frequency), and Comparison Result 3 (average fundamental frequency is less than the reference average fundamental frequency). In an example embodiment, assuming that the pronunciation abnormality level of the target vocal trainee is level three and the average fundamental frequency comparison result is Comparison Result 1, then the second abnormal training set of the target vocal trainee includes: Abnormality type: pitch too high, Abnormality level: level three (indicating that the problem is relatively serious and requires special attention and correction), and specific manifestation: the average fundamental frequency is higher than the reference average fundamental frequency, reflecting that the trainee may be using excessive force when vocalizing.

[0059] Specifically, the analysis process for the vocal audio stability factor of the target vocal trainee is as follows: the vocal characteristic parameters of the target vocal trainee include the vocal frequency of the target vocal trainee at each moment during the training period, the fundamental frequency of the target vocal trainee at each moment during the training period, and the pitch curve of the target vocal trainee during the training period; the vocal frequency represents the number of sounds emitted per second by the target vocal trainee during the training process, which can be obtained by analyzing the collected audio of the target vocal trainee through audio analysis software (such as audio editing software); the fundamental frequency, also known as the fundamental tone frequency, or simply fundamental frequency F0, is the lowest frequency periodic vibration frequency in the sound waveform, and is the most basic frequency component of the sound signal. In vocal training, the fundamental frequency determines the pitch of the sound. The fundamental frequency (FFM) values ​​collected from the target vocal trainees can be analyzed using audio analysis software (such as audio editing software). During the training period, when the FFM values ​​are found to be the same or very close at certain times, these FFM values ​​are uniformly marked to simplify analysis and discussion. Specifically, these identical FFM values ​​are labeled as the FFM values ​​at the corresponding times, so that these values ​​can be quickly identified and referenced in subsequent analysis. This marking method helps to better understand the changing trend of the FFM during the training process and can also improve the accuracy and efficiency of the analysis. The pitch curve mentioned above represents a graph showing the change of pitch of the target vocal trainee's voice over time during the training period, which can be obtained by analyzing the collected audio of the target vocal trainees using audio analysis software (such as audio editing software).

[0060] The fundamental frequency of the target vocal trainee at each moment during the training period is processed by standard deviation to obtain the fundamental frequency dispersion value of the target vocal trainee during the training period. This dispersion value is then compared with the average fundamental frequency of the target vocal trainee during the training period to obtain the fundamental frequency jitter variation coefficient of the target vocal trainee during the training period. The aforementioned fundamental frequency dispersion value is used to quantify the degree of dispersion of the fundamental frequency of the target vocal trainee during the training period. The aforementioned fundamental frequency jitter variation coefficient is used to quantify the degree of fundamental frequency jitter of the target vocal trainee during vocalization.

[0061] Obtain the reference pitch curve of the target vocal trainee during the training period. Compare the target vocal trainee's pitch curve during the training period with the reference pitch curve during the training period to obtain the overlap rate of the pitch curves during the training period. The reference pitch curve refers to the reference curve of the pitch curve, which can be extracted from professional teaching materials in the field of vocal music. The overlap rate of the pitch curves refers to the proportion of overlap between the pitch curve and the reference pitch curve. It can be obtained by inputting the pitch curve and the reference pitch curve into image processing software (such as Matrix Lab) and analyzing it through the curve overlap comparison function of the image processing software.

[0062] Several data points are located on the pitch curve of the target vocal trainee during the training period. The pitch values ​​of these data points are obtained and processed by standard deviation. The processing result is marked as the pitch dispersion value of the target vocal trainee during the training period. The aforementioned location of several data points refers to several data points located by a random algorithm (such as simple random sampling). The aforementioned pitch dispersion value is used to quantify the degree of dispersion of the pitch values ​​of several data points.

[0063] Based on the data collection quality factor of the target first dataset, the pronunciation audio stability correction value of the target vocal trainee is matched from the vocal database. This is used to quantify the degree of correction to the pronunciation audio stability factor of the target vocal trainee. The specific matching process is as follows: The vocal database stores the pronunciation audio stability correction value corresponding to each data collection quality factor interval. The vocal database stores the data collection quality factor interval to which the data collection quality factor of the target first dataset belongs. The pronunciation audio stability correction value corresponding to the data collection quality factor interval stored in the vocal database is the pronunciation audio stability correction value of the target vocal trainee.

[0064] It should be explained that, given the synchronization requirements of video and audio in multimedia content, any quality fluctuations during video collection or network conditions during uploading may affect the accompanying audio data. To ensure the accuracy and completeness of audio data analysis, the impact of the data collection quality factor of the target primary dataset must be considered and evaluated when conducting audio analysis.

[0065] By comprehensively considering the posture deviation assessment index, the pronunciation audio stability correction value, the fundamental frequency jitter coefficient of variation, the pitch curve overlap rate, and the pitch dispersion value of the target vocal trainee during the training period, a pronunciation audio stability factor is derived for the target vocal trainee. This pronunciation audio stability factor represents the quantitative data on the influence of the posture deviation assessment index, pronunciation audio stability correction value, fundamental frequency jitter coefficient of variation, pitch curve overlap rate, and pitch dispersion value on the pronunciation audio stability of the target vocal trainee, and is used to comprehensively quantify the degree of pronunciation audio stability of the target vocal trainee.

[0066] The specific analysis method for the vocal audio stability factor of the target vocal trainee is as follows:

[0067]

[0068] In the formula, ASR is the vocal audio stability factor of the target vocal trainee, AD_SEC is the posture deviation assessment index of the target vocal trainee, K is the normalized weight factor of the posture deviation assessment index preset in the vocal database, b1 is the normalized influence factor of the fundamental frequency jitter variation coefficient preset in the vocal database, b2 is the normalized influence factor of the pitch curve overlap rate preset in the vocal database, b3 is the normalized influence factor of the pitch dispersion value preset in the vocal database, CV is the fundamental frequency jitter variation coefficient of the target vocal trainee during the training period, DT is the pitch curve overlap rate of the target vocal trainee during the training period, SV is the pitch dispersion value of the target vocal trainee during the training period, CVJ is the preset defined fundamental frequency jitter variation coefficient in the vocal database, DTJ is the preset defined pitch curve overlap rate in the vocal database, SVJ is the preset defined pitch dispersion value in the vocal database, and GY is the vocal audio stability correction value of the target vocal trainee.

[0069] The above definition of fundamental frequency jitter variation coefficient represents the maximum allowable value of fundamental frequency jitter variation coefficient; the above definition of pitch curve coincidence rate represents the minimum allowable value of pitch curve coincidence rate; the above definition of pitch dispersion value represents the maximum allowable value of pitch dispersion value.

[0070] The aforementioned normalized weighting factor for the posture deviation assessment index is used to normalize the posture deviation assessment index, representing the proportion of the posture deviation assessment index in the audio stability factor. The aforementioned normalized influence factor for the fundamental frequency jitter coefficient of variation is used to normalize the fundamental frequency jitter coefficient of variation, quantifying the influence of the unit value of the fundamental frequency jitter coefficient of variation on the audio stability factor. The aforementioned normalized influence factor for the pitch curve overlap rate is used to normalize the pitch curve overlap rate, quantifying the influence of the unit value of the pitch curve overlap rate on the audio stability factor. The aforementioned normalized influence factor for pitch dispersion is used to normalize the pitch dispersion value, quantifying the influence of the unit value of the pitch dispersion value on the audio stability factor. The vocal database stores the correspondence between posture deviation evaluation index and its corresponding normalized weighting factor, as well as the correspondence between fundamental frequency jitter variation coefficient, pitch curve overlap rate, and pitch dispersion value and their corresponding normalized influence factors. For example, by inputting the posture deviation evaluation index, fundamental frequency jitter variation coefficient, pitch curve overlap rate, and pitch dispersion value into the vocal database, the vocal database can match the normalized weighting factor of the posture deviation evaluation index, the normalized influence factor of the fundamental frequency jitter variation coefficient, the normalized influence factor of the pitch curve overlap rate, and the normalized influence factor of the pitch dispersion value. The values ​​are all between 0 and 1. In this embodiment, normalization refers to de-normalization.

[0071] It should be explained that the posture deviation assessment index directly reflects the vocal trainee's posture during training. Poor posture often affects breathing control and vocal cord vibration, leading to an increase in the fundamental frequency jitter coefficient. Furthermore, this fundamental frequency instability significantly reduces the overlap between the pitch curve and the reference curve, meaning that the trainee's actual singing pitch deviates significantly from the expected target. At the same time, the pitch dispersion value, as an indicator of the degree of pitch distribution dispersion, is also affected by posture deviation and fundamental frequency jitter. An increase in the pitch dispersion value means a decrease in pitch control ability. Ultimately, the combined effect of these parameters profoundly affects the vocal audio stability factor, thereby affecting the overall effect of vocal training and sound quality.

[0072] The first and second abnormal training sets of the target vocal trainees are combined to generate a total abnormal parameter training set for the target vocal trainees. A fourth artificial intelligence scheme is used to match and generate a correction set for the target vocal trainees. The correction set and the total abnormal parameter training set of the target vocal trainees are then visualized. The total abnormal parameter training set is a set composed of the first and second abnormal training sets. The correction set is a set of methods used to correct abnormalities that occur in the target vocal trainees during vocal training.

[0073] Specifically, the process of matching the correction set of the target vocal trainee using the fourth artificial intelligence scheme involves the following steps: First, the overall training set of abnormal parameters of the target vocal trainee is compared with the overall training sets of abnormal parameters stored in the vocal database to obtain the similarity between the overall training set of abnormal parameters of the target vocal trainee and each of the other training sets. Then, the similarity scores of the overall training set of abnormal parameters of the target vocal trainee are sorted in descending order, and the correction set corresponding to the highest similarity score is extracted and marked as the correction set of the target vocal trainee. This similarity comparison can be performed using a cosine similarity algorithm. In one example embodiment, the correction set of the target vocal trainee includes: during the high-note challenge phase of the training cycle, the trainee needs to constantly... Maintain a high level of body awareness, paying particular attention to keeping the torso upright. During this stage, it is crucial to avoid any form of excessive forward or backward leaning. These improper postures not only interfere with the smoothness of breathing but may also disrupt the stability of the vocal mechanism. When practicing complex techniques involving rapid scale transitions during the training cycle, trainees must strictly control the amplitude of body swaying. Excessive body swaying not only distracts attention but may also cause instability in breathing and vocalization, leading to pitch deviations and inconsistent timbre during scale transitions. It is recommended to adopt a more stable posture and maintain relative stillness during rapid scale transitions to ensure vocal accuracy and smoothness of scale transitions. If trainees experience pitch deviations when attempting high notes, it is recommended to adopt a strategy of gradually lowering the pitch. Practice by first lowering the pitch to a more comfortable range and then gradually approaching the original high note target.

[0074] In one specific embodiment, the present invention provides a vocal training method and system based on artificial intelligence. By collecting and analyzing the body posture parameters of the target vocal trainee, and after preprocessing, detecting posture abnormalities to form a first abnormal training set, the present invention also captures and analyzes pronunciation feature parameters to identify pronunciation abnormalities to form a second abnormal training set. By combining the two, a total abnormal parameter training set is generated, and a correction set is matched. The degree and cause of abnormal body posture and pronunciation are visualized, allowing the target vocal trainee to intuitively see their posture and pronunciation problems, facilitating self-adjustment and correction.

[0075] Reference Figure 2 As shown, the second aspect of the present invention provides an artificial intelligence-based vocal training system, including: a data preprocessing module, a first abnormal training set generation module, a second abnormal training set generation module, a visualization module, and a vocal database.

[0076] The vocal database stores data collection quality thresholds, a first preprocessing scheme corresponding to data collection quality factor interval one, a second preprocessing scheme corresponding to data collection quality factor interval two, a third preprocessing scheme corresponding to data collection quality factor interval three, motion coordinate reference sets for each target point, data update reference duration for each target point in the first target dataset, motion anomaly rate normalization influence factor, distribution discrete value normalization influence factor, total data update deviation duration normalization influence factor, posture deviation assessment threshold for the target vocal trainee, posture deviation correction value for the target vocal trainee, defined contour deviation area, and defined key points. The following parameters are considered: reference vocal frequency-chest movement frequency factor, contour deviation area normalization influence factor, key point displacement normalization influence factor, vocal frequency-chest movement frequency factor normalization influence factor, vocal audio stability threshold, vocal abnormality level corresponding to each vocal audio stability factor interval, vocal audio stability correction value for the target vocal trainee, posture deviation assessment index normalization weight factor, fundamental frequency jitter coefficient of variation normalization influence factor, pitch curve overlap rate normalization influence factor, pitch dispersion value normalization influence factor, defining fundamental frequency jitter coefficient of variation, defining pitch curve overlap rate, defining pitch dispersion value, and the total training set of each abnormal parameter.

[0077] The data preprocessing module is connected to the first abnormal training set generation module, the first abnormal training set generation module is connected to the second abnormal training set generation module, the second abnormal training set generation module is connected to the visualization module, and the data preprocessing module, the first abnormal training set generation module, the second abnormal training set generation module, and the visualization module are all connected to the vocal music database.

[0078] The data preprocessing module is used to collect the body posture parameter set of the target vocal trainee and label it as the target first dataset. A first artificial intelligence scheme is used to determine whether to preprocess the target first dataset. The body posture parameter set is a sequence of video frames from the vocal training process of the target vocal trainee. The first abnormal training set generation module is used to determine, based on the target first dataset, whether to generate a first abnormal training set for the target vocal trainee using a second artificial intelligence scheme. The first abnormal training set is a collection used to record abnormal body postures of the target vocal trainee during vocal training. The second abnormal training set generation module is used to obtain the target... The vocal training personnel's pronunciation feature parameters are used to determine whether to generate a second abnormal training set for the target vocal training personnel using a third artificial intelligence scheme. The second abnormal training set is a collection used to record the pronunciation abnormalities of the target vocal training personnel during vocal training. The visualization module is used to combine the first abnormal training set and the second abnormal training set of the target vocal training personnel to generate a total abnormal parameter training set of the target vocal training personnel. A fourth artificial intelligence scheme is used to match the correction set of the target vocal training personnel, and the correction set and the total abnormal parameter training set of the target vocal training personnel are visualized.

[0079] The above content is merely an example and illustration of the structure of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the structure of the invention or exceed the scope defined by the present invention, they should all fall within the protection scope of the present invention. The abnormal parameter training set is a set composed of the first abnormal training set and the second abnormal training set. The correction set is a set of methods for correcting abnormalities that occur in the target personnel during vocal training.

Claims

1. A vocal training method based on artificial intelligence, characterized in that, include: Collect the body posture parameter set of the target vocal training personnel and mark it as the target first dataset. Use the first artificial intelligence scheme to determine whether to preprocess the target first dataset. The body posture parameter set is a video frame sequence of the vocal training process of the target vocal training personnel. Based on the first target dataset, the second artificial intelligence scheme is used to determine whether to generate a first abnormal training set for the target vocal trainee. The first abnormal training set is a set used to record abnormal body postures of the target vocal trainee during vocal training. The pronunciation feature parameters of the target vocal trainee are obtained, and a third artificial intelligence solution is used to determine whether to generate a second abnormal training set of the target vocal trainee. The second abnormal training set is a set used to record the pronunciation abnormalities of the target vocal trainee during vocal training. The first and second abnormal training sets of the target vocal trainees are combined to generate a total abnormal parameter training set for the target vocal trainees. A fourth artificial intelligence scheme is used to match the correction set for the target vocal trainees. The correction set and the total abnormal parameter training set of the target vocal trainees are visualized. The total abnormal parameter training set is a set composed of the first and second abnormal training sets. The correction set is a set of methods used to correct the abnormalities that occur in the vocal training process of the target trainees. The specific process for determining whether a second abnormal training set of the target vocal training personnel has been generated using a third artificial intelligence solution is as follows: Based on the pronunciation characteristic parameters of the target vocal trainees, the pronunciation audio stability factor of the target vocal trainees is analyzed. Extract the pronunciation audio stability threshold from the vocal database, compare the pronunciation audio stability factor of the target vocal trainee with the pronunciation audio stability threshold, and if the pronunciation audio stability factor of the target vocal trainee is greater than the pronunciation audio stability threshold, then it is determined that a second abnormal training set for the target vocal trainee will not be generated. If the pronunciation audio stability factor of the target vocal trainee is less than or equal to the pronunciation audio stability threshold, then a second abnormal training set of the target vocal trainee is generated. The specific process for generating the second abnormal training set of the target vocal trainee is as follows: matching the pronunciation audio stability factor of the target vocal trainee with the pronunciation abnormality level corresponding to each pronunciation audio stability factor interval stored in the vocal database to obtain the pronunciation abnormality level of the target vocal trainee. The fundamental frequency of the target vocal trainee at each moment during the training period is extracted from the pronunciation feature parameters of the target vocal trainee, and the mean is processed to obtain the average fundamental frequency of the target vocal trainee during the training period. The reference average fundamental frequency of the target vocal trainee during the training period is obtained and compared with the average fundamental frequency of the target vocal trainee during the training period to obtain the average fundamental frequency comparison result. Based on the pronunciation abnormality level of the target vocal trainee and the average fundamental frequency comparison result, a second abnormal training set of the target vocal trainee is generated. The specific analysis process for the vocal audio stability factor of the target vocal trainee is as follows: The vocal characteristic parameters of the target vocal trainee include the vocal frequency of the target vocal trainee at each moment during the training period, the fundamental frequency of the target vocal trainee at each moment during the training period, and the pitch curve of the target vocal trainee during the training period. The fundamental frequency of the target vocal trainee at each moment during the training period is processed by standard deviation to obtain the discrete value of the fundamental frequency of the target vocal trainee during the training period. This discrete value is then compared with the average fundamental frequency of the target vocal trainee during the training period to obtain the fundamental frequency jitter variation coefficient of the target vocal trainee during the training period. Obtain the reference pitch curve of the target vocal trainee during the training period, compare the pitch curve of the target vocal trainee during the training period with the reference pitch curve of the target vocal trainee during the training period, and analyze the overlap rate of the pitch curve of the target vocal trainee during the training period. Several data points are located on the pitch curve of the target vocal trainee during the training period. The pitch values ​​of the data points are obtained and the standard deviation is processed. The processing result is marked as the pitch discrete value of the target vocal trainee during the training period. Based on the data collection quality factor of the target first dataset, the pronunciation audio stability correction value of the target vocal trainees is matched from the vocal database; By comprehensively considering the posture deviation assessment index, the pronunciation audio stability correction value, the fundamental frequency jitter coefficient of variation, the pitch curve overlap rate, and the pitch dispersion value of the target vocal trainee during the training period, a pronunciation audio stability factor is derived for the target vocal trainee. This pronunciation audio stability factor represents the quantitative data on the influence of the posture deviation assessment index, pronunciation audio stability correction value, fundamental frequency jitter coefficient of variation, pitch curve overlap rate, and pitch dispersion value on the pronunciation audio stability of the target vocal trainee, and is used to comprehensively quantify the degree of pronunciation audio stability of the target vocal trainee.

2. The vocal training method based on artificial intelligence according to claim 1, characterized in that: The process of determining whether to preprocess the target first dataset using the first artificial intelligence solution is as follows: Obtain the first attribute parameter of the target first dataset, and analyze the data collection quality factor of the target first dataset based on the first attribute parameter of the target first dataset; Extract the data collection quality threshold from the vocal database, compare the data collection quality factor of the first target dataset with the data collection quality threshold, and if the data collection quality factor of the first target dataset is greater than the data collection quality threshold, it is determined that no preprocessing will be performed on the first target dataset. If the data collection quality factor of the target first dataset is less than or equal to the data collection quality threshold, then it is determined that the target first dataset should be preprocessed.

3. The vocal training method based on artificial intelligence according to claim 2, characterized in that: The preprocessing of the first target dataset is as follows: The vocal database stores the first preprocessing scheme corresponding to the first data collection quality factor interval, the second preprocessing scheme corresponding to the second data collection quality factor interval, and the third preprocessing scheme corresponding to the third data collection quality factor interval. Match the data collection quality factors of the first target dataset with the data collection quality factor intervals 1, 2, and 3 stored in the vocal database; If the data collection quality factor of the target first dataset belongs to the data collection quality factor interval one, then the first preprocessing scheme of the target first dataset is matched and the target first dataset is preprocessed according to the first preprocessing scheme of the target first dataset. If the data collection quality factor of the target first dataset belongs to the second interval of the data collection quality factor, then the second preprocessing scheme of the target first dataset is obtained, and the target first dataset is preprocessed according to the second preprocessing scheme of the target first dataset. If the data collection quality factor of the target first dataset belongs to the data collection quality factor interval three, then the third preprocessing scheme of the target first dataset is matched and the target first dataset is preprocessed according to the third preprocessing scheme.

4. The vocal training method based on artificial intelligence according to claim 3, characterized in that: The specific analysis process for the data collection quality factor of the target first dataset is as follows: The first attribute parameter of the target first dataset includes the motion coordinate points of each target point to which the target first dataset belongs; The motion coordinates of all target points in the first target dataset are combined and labeled as the motion coordinate set of each target point in the first target dataset. The body posture attribute parameter set of the target vocal training personnel is obtained, and the motion coordinate reference set of each target point is matched from the vocal database. The motion coordinate set of each target point in the first target dataset is compared with the motion coordinate reference set of each target point. The abnormal motion coordinates of each target point in the first target dataset are identified through comparison and analysis. The total number of abnormal motion coordinates of each target point in the first target dataset is counted and labeled as the total number of abnormal motion coordinates of each target point in the first target dataset. The ratio of the total number of abnormal motion coordinates of each target point in the first target dataset to the total number of motion coordinates of each target point in the first target dataset is calculated to obtain the motion abnormality rate of each target point in the first target dataset. The total number of motion coordinate points of each target point in the first dataset is processed by standard deviation, and the processing result is marked as the discrete distribution value of the target points in the first dataset. Obtain the timestamps of each motion coordinate point of each target point in the first target dataset, analyze the data to obtain the average data update time of each target point in the first target dataset, obtain the data update reference time of each target point in the first target dataset from the vocal database, and perform difference and absolute value processing on the corresponding average data update time of each target point in the first target dataset to obtain the data update deviation time of each target point in the first target dataset, and sum them to obtain the total data update deviation time of the target points in the first target dataset; By combining the motion anomaly rate of each target point in the first target dataset, the distribution dispersion value of the target points in the first target dataset, and the total data update deviation time of the target points in the first target dataset, a data collection quality factor for the first target dataset is derived. The data collection quality factor for the first target dataset represents the quantitative data on the degree of influence of the motion anomaly rate, distribution dispersion value, and total data update deviation time on the data quality collected by the first target dataset, and is used to quantify the data quality collected by the first target dataset.

5. The vocal training method based on artificial intelligence according to claim 1, characterized in that: The specific process for determining whether to generate a first abnormal training set for the target vocal training personnel using the second artificial intelligence scheme is as follows: Obtain the second attribute parameters of the target first dataset, and evaluate the posture deviation assessment index of the target vocal trainee based on the second attribute parameters of the target first dataset; Based on the body posture attribute parameter set of the target vocal trainee, the posture deviation assessment threshold of the target vocal trainee is matched from the vocal database. The posture deviation assessment index of the target vocal trainee is compared with the posture deviation assessment threshold of the target vocal trainee. If the posture deviation assessment index of the target vocal trainee is less than or equal to the posture deviation assessment threshold of the target vocal trainee, then the first abnormal training set of the target vocal trainee is not generated. If the posture deviation assessment index of the target vocal trainee is greater than the posture deviation assessment threshold of the target vocal trainee, then the first abnormal training set of the target vocal trainee is generated. The specific process for generating the first abnormal training set of the target vocal trainee is as follows: extract the posture deviation assessment factor of the target vocal trainee at each moment in the training cycle from the posture deviation assessment index of the target vocal trainee, and process the difference and absolute value of the posture deviation assessment factor of the target vocal trainee at each moment in the training cycle and the posture deviation assessment threshold of the target vocal trainee in turn to obtain the posture deviation assessment difference of the target vocal trainee at each moment in the training cycle. Extract the contour curves of the target vocal trainer at each moment during the training period from the second attribute parameters of the first target dataset, obtain the contour curves of the target vocal trainer at historical adjacent moments during the training period, and integrate the contour curves of the target vocal trainer at historical adjacent moments during the training period with the contour curves of the target vocal trainer at each moment during the training period to obtain the contour curves of the target vocal trainer at each moment during the training period - historical adjacent contour curves. The first abnormal training set of the target vocal trainee is generated by combining the difference in posture deviation assessment at each moment during the training period and the contour curve-historical adjacent contour curve of the target vocal trainee at each moment during the training period.

6. The vocal training method based on artificial intelligence according to claim 5, characterized in that: The assessment process for the postural deviation evaluation index of the target vocal training personnel is as follows: The second attribute parameters of the target first dataset include the contour curve of the target vocal trainee at each moment during the training period, the key point coordinates of the target vocal trainee at each moment during the training period, and the chest movement frequency of the target vocal trainee at each moment during the training period. By comparing the contour curves of the target vocal trainee at each moment during the training period with the corresponding historical adjacent contour curves, the contour deviation area of ​​the target vocal trainee at each moment during the training period can be obtained. Distance analysis is performed between the key point coordinates of the target vocal trainee at each moment during the training cycle and the corresponding historical adjacent key point coordinates to obtain the key point displacement of the target vocal trainee at each moment during the training cycle. The vocal frequency of the target vocal trainee at each moment during the training period is extracted from the vocal characteristic parameters of the target vocal trainee. The vocal frequency of the target vocal trainee at each moment during the training period is normalized to the ratio of the chest movement frequency of the target vocal trainee at each moment during the training period, and the vocal frequency-chest movement frequency factor of the target vocal trainee at each moment during the training period is obtained. Based on the data collection quality factor of the first target dataset, the posture deviation correction values ​​of the target vocal trainees are matched from the vocal database; Based on the body posture attribute parameter set of the target vocal training personnel, the area of ​​the defined contour deviation, the displacement of the defined key points, and the reference vocal frequency-chest movement frequency factor are matched from the vocal database. By comprehensively analyzing the contour deviation area, key point displacement, vocal frequency-chest movement frequency factor, posture deviation correction value, defined contour deviation area, defined key point displacement, and reference vocal frequency-chest movement frequency factor at various times during the training cycle of the target vocal trainee, a posture deviation assessment index is derived. This index represents the quantitative data on the influence of contour deviation area, key point displacement, vocal frequency-chest movement frequency factor, posture deviation correction value, defined contour deviation area, defined key point displacement, and reference vocal frequency-chest movement frequency factor on the posture deviation of the target vocal trainee, and is used to comprehensively quantify the degree of posture deviation of the target vocal trainee.

7. The vocal training method based on artificial intelligence according to claim 1, characterized in that: The specific analysis process for matching the correction set of the target vocal training personnel using the fourth artificial intelligence scheme is as follows: The similarity between the total set of abnormal parameters of the target vocal trainee and the total set of abnormal parameters stored in the vocal database is compared to obtain the similarity between the total set of abnormal parameters of the target vocal trainee and the total set of abnormal parameters. The similarity between the total training set of abnormal parameters of the target vocal trainee and each training set of abnormal parameters is sorted in descending order. The correction set of the training set of abnormal parameters corresponding to the first-ranked similarity is extracted and marked as the correction set of the target vocal trainee.

8. A system applying the artificial intelligence-based vocal training method as described in any one of claims 1-7, characterized in that: include: The data preprocessing module is used to collect the body posture parameter set of the target vocal trainee and mark it as the target first dataset. The first artificial intelligence scheme is used to determine whether to preprocess the target first dataset. The body posture parameter set is a video frame sequence of the vocal training process of the target vocal trainee. The first abnormal training set generation module is used to determine whether to generate a first abnormal training set for the target vocal trainee based on the target first dataset and using the second artificial intelligence scheme. The first abnormal training set is a set used to record abnormal body postures of the target vocal trainee during vocal training. The second abnormal training set generation module is used to obtain the pronunciation feature parameters of the target vocal trainee and use the third artificial intelligence scheme to determine whether to generate the second abnormal training set of the target vocal trainee. The second abnormal training set is a collection used to record the pronunciation abnormalities of the target vocal trainee during vocal training. The visualization module is used to generate a total training set of abnormal parameters for the target vocal trainee by combining the first abnormal training set and the second abnormal training set of the target vocal trainee. It then uses a fourth artificial intelligence scheme to match the correction set of the target vocal trainee and visualizes the correction set and the total training set of abnormal parameters of the target vocal trainee. The total training set of abnormal parameters is a set composed of the first and second abnormal training sets. The correction set is a set of methods used to correct abnormalities that occur in the vocal training process of the target trainee.

Citation Information

Patent Citations

  • A comprehensive vocal training system

    CN106448701B

  • Vocal music pronunciation training system

    CN119229895A

  • Teaching method and device, electronic equipment and storage medium

    CN111695777A