Autonomous language learning system based on big data speech recognition

By adopting big data speech recognition technology and deep learning models in the language learning system, combined with clustering algorithms, we can identify and distinguish learners' pronunciation bias and dialect differences, the problem of misjudgment in the existing system when dealing with learners with specific dialect accents is solved, and more accurate pronunciation evaluation and personalized feedback are achieved.

CN120236589AInactive Publication Date: 2025-07-01XIAMEN DUOXIANG ANIMATION CO LTD

Patent Information

Application Number
CN202510712410.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

When dealing with learners with specific dialect accents, existing language learning systems cannot correctly distinguish the source of pronunciation offset, resulting in continuous misjudgment of learners' learning situation, affecting their learning confidence and practical application ability.

Method used

A language autonomous learning system based on big data speech recognition is adopted. Through the acquisition module, processing module, optimization module and analysis module, the speech signal is cleaned, framed, feature extraction and clustered, and combined with deep learning models and clustering algorithms, we can identify and distinguish learners' pronunciation bias and dialect differences.

Benefits of technology

The system can accurately identify learners' pronunciation bias and dialect differences, provide personalized improvement suggestions, enhance learners' learning confidence and continuous learning motivation, and improve their practical language application ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236589A_ABST
    Figure CN120236589A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of speech recognition, and particularly discloses a language autonomous learning system based on big data speech recognition, and the system comprises an acquisition module which carries out the grouping of users according to regions, marks the users in the same group as target users, and obtains a speech signal when the target users carry out the language learning; the processing module is used for performing framing processing on the voice signals to obtain voice frames, extracting voice features of the voice frames based on a voice recognition technology, and determining sub-features based on a deep learning model; the optimization module is used for generating coordinate points, clustering the coordinate points to obtain clusters, performing descending sorting on the clusters according to the density, and determining a target cluster based on the sorting; and the analysis module is used for acquiring the target character, acquiring an analysis result of the target character based on the target cluster, and not sending an error prompt when the analysis result is dialect abnormity. According to the invention, frequent sending of error prompt information caused by dialects of language learners can be avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech recognition, and particularly to a language autonomous learning system based on big data speech recognition. Background Art

[0002] Big data speech recognition is a method that combines a large amount of speech data with advanced machine learning technologies to provide machines with powerful "listening" capabilities. In the field of language learning, big data speech recognition technology provides accurate, efficient, and personalized assistance to learners. By analyzing a large number of pronunciation samples of native speakers and learners, the system can accurately capture subtle speech differences and phoneme mispronunciations, and immediately provide pronunciation scores and improvement suggestions.

[0003] However, in current language learning methods, it mainly relies on the distance metric from the standard pronunciation template or acoustic features, ignoring the diversity of individual speech characteristics. For learners with specific dialect accents, the system often fails to correctly distinguish the source of their pronunciation deviation and simply marks it as "mispronunciation" or "low score", which easily leads to continuous misjudgment of the learning situation of learners and affects their learning confidence and practical application ability. Summary of the Invention

[0004] The purpose of the present invention is to provide a language autonomous learning system based on big data speech recognition to solve the above technical problems.

[0005] The purpose of the present invention can be achieved through the following technical solutions: A language autonomous learning system based on big data speech recognition includes a collection module, a processing module, an optimization module, and an analysis module. Specifically: Collection module: Group users according to regions, mark the users in the same group as target users, and obtain the speech signals of the target users during language learning. Processing module: Perform frame splitting on the speech signal to obtain speech frames, extract the speech features of the speech frames based on speech recognition technology, and determine the speech features corresponding to a single character based on a deep learning model, denoted as sub-features. Optimization module: For the same character, generate coordinate points based on the sub-features, perform clustering on the coordinate points to obtain clustering clusters, sort the clustering clusters in descending order according to density, and use the first n clustering clusters in the sorting as the target clusters of the current character, where n is a preset quantity. Analysis module: Obtain all the characters read by the current user during language learning, denoted as target characters, and obtain the analysis results of the target characters based on the target clusters. The analysis results include dialect anomalies and learning anomalies. When the analysis result is a dialect anomaly, no error prompt is sent.

[0006] As a further solution of the present invention: Before obtaining the speech frames, it further includes: Clean and denoise the collected speech signals. For missing data, fill it with the mean, median, or predicted values based on machine learning algorithms.

[0007] As a further solution of the present invention: Generating coordinate points includes: Normalize the sub-features; Establish a coordinate system of m dimensions, where m is the total number of types of sub-features. Generate coordinate points (A1, A2,..., Am) in the coordinate system, and Am represents the m-th type of sub-feature after normalization.

[0008] As a further solution of the present invention: Obtaining the analysis result of the target text includes: Obtain the sub-features of the target text i corresponding to the current user during the language learning process, denoted as comparison features; Obtain the target cluster corresponding to the target text i, denoted as the standard cluster, and extract the sub-features corresponding to the theoretical center of gravity of the standard cluster, denoted as the standard features; Obtain the first difference between the comparison features and the corresponding standard features, and perform weighted summation on the first difference to obtain the first difference degree; Obtain the second difference between the comparison features and the preset reference features, and perform weighted summation on the second difference to obtain the second difference degree; When the second difference degree is less than or equal to the preset difference degree threshold, do not output the analysis result; When the second difference degree is greater than the preset difference degree threshold, perform the following steps: When the first difference degree is greater than the preset difference degree threshold, determine that the analysis result is an abnormal learning; When the first difference degree is less than or equal to the preset difference degree threshold, determine that the analysis result is a dialect abnormality.

[0009] As a further solution of the present invention: After determining that the analysis result is a dialect abnormality, it further includes: Denote the target text with the analysis result of dialect abnormality as the improved text, generate pronunciation improvement suggestions for the improved text, and push them to the user.

[0010] As a further solution of the present invention: During the process of sorting the clustering clusters, the clustering clusters with the same density occupy the same sorting position in the sorting.

[0011] The beneficial effects of the present invention: Compared with the prior art: Through the cleaning, framing, feature extraction, and normalization of the user's voice signal, the present invention accurately identifies the multi-dimensional sub-features of individual characters in combination with a deep learning model, and uses a clustering algorithm to automatically divide high-density target clusters in the feature space, effectively capturing diverse speech patterns with similar pronunciation features. On this basis, the system can dynamically compare the difference between the pronunciation features of each learner and the theoretical center of gravity of the target cluster where they are located, accurately identify the sources of learning deviations and dialect differences, and give personalized improvement suggestions for dialect anomalies, while indicating the correction direction for real learning deviations. In this way, the present invention not only overcomes the misjudgment problem of the traditional method's one-size-fits-all approach to dialect accents, but also respects and retains the individual speech feature differences of learners, significantly improving the fairness and accuracy of pronunciation evaluation; at the same time, through targeted feedback and guidance, it enhances the learning confidence and continuous learning motivation of learners, and ultimately promotes the steady improvement of their practical language application ability. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The present invention will be further described below with reference to the accompanying drawings.

[0013] Figure 1 It is a schematic diagram of the modules of a language autonomous learning system based on big data speech recognition according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0015] Please refer to Figure 1 As shown, the present invention is a language autonomous learning system based on big data speech recognition, including a collection module, a processing module, an optimization module, and an analysis module. Specifically: Collection module: In actual operation, firstly, learners living in Beijing, Tianjin, and Hebei are classified as northern dialect areas according to the geographical location of the users, and learners living in Guangdong, Guangxi, and Fujian are classified as southern dialect areas (other divisions are also possible, such as districts, cities, and counties). Then, all learners in each group are uniformly labeled with regional tags and marked as target users in the area. Then, when a target user starts to practice reading aloud online, the front-end recording module automatically collects the continuous voice stream during the reading process, and uploads the recorded original audio file to the back-end server in real time. On the server side, the corresponding interface will store the uploaded audio together with the audio samples of other users in the same group in the corresponding regional voice database according to the user's regional tag, so as to facilitate subsequent unified management and analysis. It is understandable that grouping users by region and centrally acquiring the target users' voice signals is based on the fact that dialect differences in different regions will show obvious group commonalities in acoustic features. Regional grouping can aggregate voice data with similar accents to avoid misjudging pronunciation deviations with the same dialect background as learning errors. This can provide a more accurate distribution of training samples for subsequent feature clustering and evaluation based on deep learning, thereby helping the system to more accurately distinguish dialect habits from actual pronunciation deviations, and achieve the ultimate goal of providing personalized feedback to learners. Processing module: Divide the uploaded continuous voice stream into a series of short-time windows of appropriate width, and keep a certain overlap between adjacent short-time windows. Then apply the windowing function to each short-time window to reduce the distortion caused by signal truncation, so as to obtain several continuous voice frames; then, for each voice frame, the system calls the pre-trained speech recognition module, and extracts the feature vector representing the voice amplitude and spectrum shape based on acoustic algorithms such as filter banks and logarithmic energy, just like drawing acoustic contours in different frequency bands; then, the system inputs these frame-level features into the deep learning model, and the model aligns the correspondence between phonemes and characters through a multi-layer neural network, finds the frame sequence located in a certain text pronunciation segment, and synthesizes the features of each frame in this segment to generate a sub-feature vector representing the pronunciation features of the text, which is recorded as the sub-feature of the text. For example, when the user reads the word "Beijing", the model can automatically intercept the continuous frame features corresponding to the word "North" and output a multi-dimensional vector for subsequent analysis; It should be noted that by first performing frame segmentation and extracting sub-features corresponding to each character based on speech recognition technology and deep learning models, it is possible to finely divide the speech signal in the time domain and frequency domain, and map complex and continuous acoustic information to discrete text units. This not only allows for the capture of subtle changes during the pronunciation process but also enables subsequent clustering algorithms to operate in a feature space with a higher granularity, avoiding treating the naturally flowing acoustic signal rigidly as a whole. This improves the ability to identify the sources of pronunciation differences and lays a foundation for personalized evaluation and precise feedback; In a preferred embodiment of the present invention, before obtaining the speech frames, it further includes: Preprocess the collected original audio. When background noise or wind noise and other interferences are detected, an adaptive filtering algorithm is called to denoise the audio signal. For example, by estimating the noise spectrum and suppressing the energy in the corresponding frequency band in the frequency domain to restore the main components of the speech; if audio sample loss is found due to network jitter or microphone distortion during waveform analysis, the system will first use the average waveform of adjacent frames to preliminarily fill in the missing time segments. When necessary, the median smoothing method is combined to remove the influence of extreme values, or a trained time series prediction model is used to interpolate and complete the missing area; after cleaning and denoising, the system converts the processed audio to a unified sampling rate and quantization accuracy, and outputs a high-quality speech signal after denoising and data filling for subsequent frame segmentation processing.

[0016] It should be noted that cleaning and denoising the original speech before frame segmentation and filling in the missing data with the mean, median, or machine learning algorithm-based methods are to eliminate the noise and data gaps caused by environmental interference and device jitter, enabling subsequent feature extraction to be based on real and continuous speech information. This avoids mistakenly treating interference signals or missing segments as pronunciation deviations, ensuring that the extracted speech features are more stable and reliable. This not only improves the system's recognition accuracy of the learner's true pronunciation pattern but also provides a clear and continuous data basis for subsequent feature clustering and pronunciation evaluation, thus better supporting personalized feedback and refined analysis.

[0017] Optimization module: All sub-feature vectors of a certain character are regarded as coordinate points in multidimensional space. For example, when a user reads the character "北" multiple times, the system will normalize the sub-features extracted each time to generate several coordinate points. Then, a suitable clustering algorithm, such as K-means or DBSCAN, is selected to assign these coordinate points to different clusters. DBSCAN will automatically identify core points and boundary points based on the neighborhood density of the points. GMM estimates the shape of the cluster by fitting multiple Gaussian distributions. After clustering, the system calculates the point density inside each cluster. A cluster with high density means that the pronunciation features represented by the cluster are more concentrated and typical in multiple readings. Then, all clusters are arranged in order from high to low density, and clusters with the same density are given the same sorting position. Finally, the system selects the first several clusters from the sorting results as the target clusters of the character, so that the user's real-time pronunciation can be compared with the center of gravity of these high-density clusters in subsequent analysis. It is understandable that the purpose of using sub-features to generate coordinate points and screening target clusters based on cluster density is to identify the most representative pronunciation pattern among diverse pronunciation samples and avoid mistaking occasional abnormal pronunciations or noise interference as standards. Doing so not only preserves the commonalities in dialects or individual differences, but also helps to accurately distinguish learners' pronunciation deviations from real group pronunciation patterns in the future, thereby providing a stable and representative reference cluster for the evaluation module, enabling the system to give more accurate and targeted feedback for each reading, ultimately achieving the goal of personalized and adaptive language learning assessment and guidance.

[0018] In another preferred embodiment of the present invention, generating coordinate points includes: Collect sub-feature vectors corresponding to a certain word, which may include various types such as spectral envelope, resonance peak frequency, and duration features; then for each sub-feature, calculate its minimum and maximum values ​​in the entire sub-feature set, and map each sub-feature value to the interval from zero to one by subtracting the minimum value and then dividing by the range. For example, when processing the resonance peak feature of the word "北", its original frequency value will be linearly converted to a normalized value; after normalization, determine the coordinate system dimension according to the total number of sub-feature types. For example, if five sub-features are extracted, establish a five-dimensional coordinate system, in which each dimensional axis corresponds to a normalized sub-feature; finally, for each sub-feature vector of the word "北", map it one by one to the coordinate system to form a coordinate point, so as to obtain the characteristic position of the reading in the high-dimensional space; It is worth noting that the purpose of normalizing the sub-features and generating coordinate points in the multidimensional space is to eliminate the differences in different acoustic dimensions and compare various features at the same scale, so that the comprehensive performance of multiple pronunciation attributes can be intuitively displayed in the same coordinate system. This can not only avoid the numerical amplitude of a certain dimension being too large and dominating the overall distance calculation, but also enable the subsequent clustering algorithm to more accurately identify the concentrated areas of pronunciation patterns, thereby laying a solid foundation for the subsequent screening of target clusters and the accurate assessment of learners' pronunciation deviations, and realizing personalized and adaptive language learning assessment.

[0019] Analysis module: obtain all the texts read aloud by the current user during the language learning process, record them as target texts, and obtain analysis results of the target texts based on the target clusters. The analysis results include dialect anomalies and learning anomalies. When the analysis result is dialect anomalies, no error prompt is sent; In another preferred embodiment of the present invention, obtaining the analysis result of the target text includes: For the target text i read by the user, its corresponding sub-feature vector is extracted as the comparison feature, and the standard pronunciation template feature vector of the text is obtained from the pre-established reference feature library as the reference feature. The reference feature is equivalent to an expert-level pronunciation demonstration and is used to determine whether the user's pronunciation meets the standard. Then, the standard cluster of the text is located from the target cluster obtained by clustering multiple reading samples and the sub-feature at its theoretical center of gravity is extracted as the standard feature to represent the typical pronunciation in the same dialect background. Then, the comparison feature is subtracted from the reference feature dimension by dimension and multiplied by the corresponding weight and then summed. The weight can be set between 0.2 and 0.5. The selection is made based on experience, thereby obtaining a second difference degree, which is used to determine whether the user has deviated from the standard pronunciation template as a whole; when the second difference degree is less than or equal to the preset difference degree threshold, it is considered that the user's pronunciation is consistent with the template, and no analysis result is output; and when the second difference degree is greater than the threshold, the comparison feature is further subtracted from the standard feature and weighted summed to obtain the first difference degree. When the first difference degree is greater than the threshold, it is determined to be a learning anomaly, and when the first difference degree is less than or equal to the threshold, it is determined to be a dialect anomaly, and specific pronunciation correction suggestions are provided to the user in the case of learning anomalies, and no error prompt is sent in the case of dialect anomalies to respect their accent characteristics; It should be noted that the above process can first effectively screen whether there are basic pronunciation errors through the reference feature template, and then subdivide the source of deviation in combination with the theoretical center of gravity of the standard cluster. This can not only avoid misjudging natural accent deviations as learning errors due to dialect differences, but also ensure that real pronunciation deviations are corrected in a timely manner; the introduction of reference features enables the system to compare user pronunciation with standard demonstrations at multiple levels, enhance the accuracy of error detection, and cooperate with deviation analysis based on dialect clusters to provide more comprehensive and reliable data support for personalized feedback and subsequent learning path planning, thereby ultimately promoting the steady improvement of learners' pronunciation quality and actual language application ability.

[0020] Another preferred embodiment of the present invention, after determining that the analysis result is a dialect abnormality, further includes: After determining that a target text belongs to a dialect anomaly, the text will first be marked as an improved text and added to the list to be fed back. Then the system calls the pre-trained pronunciation improvement generation model, which combines deep learning and rule base methods. First, the sub-feature difference distribution of the improved text is extracted, and then the mouth shape and tone deviation are located through phoneme-level alignment. Then, specific improvement suggestions are generated in combination with polishing rules. For example, for the problem of the high tone of the word "北", the model will output "try to move the center of gravity of the tone downward, extend the ending time of the closed mouth sound, and imitate the rhythm of the standard audio example", and attach the corresponding standard voice example clip; the generated suggestions will be organized according to the natural language template and packaged with the sample audio link, and then sent to users in real time through the message push interface of the online learning platform or the mobile notification module, so that learners can immediately obtain and apply them in subsequent practice; It should be noted that through the above method, under the premise of respecting the characteristics of learners' dialects, an operational, audible and learnable pronunciation optimization path is provided for them. Learners will not feel denied due to dialect differences, and the key points of standard pronunciation can be cleverly integrated into dialect habits, thereby maintaining learners' learning enthusiasm and enhancing their motivation to practice. At the same time, through example guidance and targeted suggestions, learners can be helped to gradually correct pronunciation deviations in daily situations, and ultimately achieve personalized and adaptive language learning goals.

[0021] The above formulas are all dimensionless and numerical calculations. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters and thresholds in the formula are set by technicians in this field according to actual conditions.

[0022] The above is a detailed description of an embodiment of the present invention, but the content is only a preferred embodiment of the present invention and cannot be considered to limit the scope of implementation of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the patent coverage of the present invention.

Claims

1. A language autonomous learning system based on big data speech recognition, characterized in that, It includes a collection module, a processing module, an optimization module and an analysis module. Specifically: Collection module: Group users by region, mark the users in the same group as target users, and obtain the speech signals of the target users during language learning; Processing module: Perform frame processing on the speech signal to obtain speech frames, extract the speech features of the speech frames based on speech recognition technology, and determine the speech features corresponding to a single character based on a deep learning model, denoted as sub-features; Optimization module: For the same character, generate coordinate points based on the sub-features, cluster the coordinate points to obtain clustering clusters, sort the clustering clusters in descending order according to density, and take the first n clustering clusters in the sorting as the target clusters of the current character, where n is a preset quantity; Analysis module: Obtain all the characters read by the current user during language learning, denoted as target characters, and obtain the analysis results of the target characters based on the target clusters. The analysis results include dialect anomalies and learning anomalies. When the analysis result is a dialect anomaly, no error prompt is sent.

2. The language autonomous learning system based on big data speech recognition according to claim 1, characterized in that, Before obtaining the speech frames, it also includes: Clean and denoise the collected speech signal. For missing data, fill it with the mean, median or predicted value based on a machine learning algorithm.

3. A language autonomous learning system based on big data speech recognition according to claim 1, characterized in that, Generating coordinate points includes: Normalize the sub-features; Establish an m-dimensional coordinate system, where m is the total number of types of sub-features. Generate coordinate points (A1, A2,..., Am) in the coordinate system, and Am represents the m-th type of sub-feature after normalization.

4. A language autonomous learning system based on big data speech recognition according to claim 1, characterized in that, Obtaining the analysis results of the target characters includes: Obtain the sub-features of the target character i corresponding to the current user during language learning, denoted as comparison features; Obtain the target cluster corresponding to the target character i, denoted as the standard cluster, and extract the sub-features corresponding to the theoretical center of gravity of the standard cluster, denoted as the standard features; Obtain the first difference between the comparison features and the corresponding standard features, and perform weighted summation on the first difference to obtain the first difference degree; obtain the second difference between the comparison features and the preset reference features, and perform weighted summation on the second difference to obtain the second difference degree; When the second difference degree is less than or equal to the preset difference degree threshold, no analysis result is output; When the second difference degree is greater than the preset difference degree threshold, perform the following steps: When the first difference degree is greater than the preset difference degree threshold, determine that the analysis result is a learning anomaly; When the first difference degree is less than or equal to the preset difference degree threshold, determine that the analysis result is a dialect anomaly.

5. A language autonomous learning system based on big data speech recognition according to claim 4, characterized in that, After determining that the analysis result is a dialect anomaly, it also includes: Denote the target characters with dialect anomalies in the analysis results as improved characters, generate pronunciation improvement suggestions for the improved characters, and push them to the users.

6. A language autonomous learning system based on big data speech recognition according to claim 1, characterized in that, During the process of sorting the clustering clusters, the clustering clusters with the same density occupy the same sorting position in the sorting.

Citation Information

Patent Citations

  • Chinese dialect identification method based on active learning

    CN114005432A

  • Speech recognition method and device for screening cognitive impairment

    CN116153298A

  • Chinese intelligent teaching method

    CN116682421A

  • Artificial intelligence interaction method and artificial intelligence interaction system

    CN117690416A

  • Efficient self-adaptive speech recognition engine-oriented hot word error correction method and system

    CN118471201A

Cited By

  • Safety box unlocking method based on facial recognition

    CN120977038A