Multimodal artificial intelligence agent driven mandarin pronunciation posture modeling method

By constructing a dynamic pose manifold space for phonemes and a cross-modal mapping function, the problems of difficulty in quantifying tongue position and aligning continuous speech flow in Mandarin pronunciation assessment are solved, achieving accurate pronunciation correction and a visualized training path.

CN121962306BActive Publication Date: 2026-06-23四川吉利学院
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
四川吉利学院
Filing Date
2026-04-03
Publication Date
2026-06-23

AI Technical Summary

Technical Problem

Existing Mandarin pronunciation assessment technologies cannot accurately depict the dynamic posture of key articulatory organs such as tongue position, making it difficult to perform precise time alignment and quantitative correction under continuous speech conditions, and lacking a continuous and traceable training path.

Method used

By collecting tongue position coordinates, lip shape key points, and mandibular opening and closing angle data from standard speakers and learners, a dynamic pose manifold space for phonemes is constructed. A cross-modal mapping function is established for data mapping, and time alignment is performed through a time deformation function. The structured pose deviation vector field and a two-dimensional convergent trajectory model are calculated for correction.

Benefits of technology

It achieves computable modeling and structured alignment of tongue position and posture, provides a continuous and traceable pronunciation correction path, and improves the accuracy of pronunciation assessment and the visualization optimization of the training process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962306B_ABST
    Figure CN121962306B_ABST
Patent Text Reader

Abstract

The application discloses a mandarin pronunciation posture modeling method driven by a multi-modal artificial intelligence intelligent agent, and relates to the field of pronunciation posture estimation. The method unifies the tongue position map coordinates, lip key points and jaw opening angle for modeling by constructing a phoneme dynamic posture manifold space, so that the standard pronunciation forms a continuous geometric trajectory expression. Through a cross-modal mapping function, the learner's observable data is mapped to a posture space consistent with the standard, realizing numerical estimation of the tongue position state. Through a time deformation function, the continuous speech flow is time-aligned for structure preservation, so that the posture comparison is established under a consistent time framework. A minimum opposite pair phoneme structure distance difference is introduced to construct a structured posture deviation vector field containing a discrimination constraint, so that the pronunciation correction direction satisfies the dual constraints of approaching the target and moving away from the confusion area. Through a two-dimensional convergence trajectory model, the posture deviation energy and the phoneme distinguishability change are quantified, and continuous tracking and path planning of the training process are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pronunciation pose estimation, specifically a method for modeling Mandarin pronunciation pose driven by a multimodal artificial intelligence agent. Background Technology

[0002] Existing Mandarin and classical recitation pronunciation training techniques primarily focus on acoustic parameter matching. They extract features from speech signals, such as fundamental frequency, formants, energy envelope, or Mel-frequency cepstral coefficients, and then perform dynamic time warping or similarity calculations with standard speech templates to provide a pronunciation accuracy score. Some techniques combine video capture with two-dimensional keypoint tracking of lip shapes, determining whether the mouth shape conforms to the demonstration standard by calculating the distance or angle changes between keypoints. Other methods employ simple frame-by-frame image comparison, calculating the difference between the learner's current image and a standard mouth shape image. These methods mainly rely on acoustic features or surface-level visible mouth shape features for calculation, typically comparing at the discrete syllable level and processing the temporal dimension using frame-level alignment or overall speech segment scoring.

[0003] The aforementioned existing technologies have the following shortcomings: First, there is no one-to-one correspondence between acoustic parameters and tongue movements inside the oral cavity. When a learner's pronunciation is close to the target acoustic features but the tongue posture deviates significantly, it is difficult to accurately identify the source of the error. Second, pronunciation under continuous speech conditions is mostly evaluated syllable by syllable, without unified modeling of the transitional movements between phonemes. This makes it difficult to handle the time misalignment caused by differences in speech rate and liaison in classic recitation scenarios. Third, existing comparison methods based on lip shape key points can only reflect externally visible movements and cannot quantitatively infer the tongue position, a key articulation point. In structurally easily confused scenarios such as retroflex consonants and aspirated consonants, there is a lack of directional constraints, making it difficult to determine whether the learner is biased towards the target phoneme or the opposing phoneme region. The training process relies heavily on subjective prompts and lacks a continuous and traceable quantitative path. Summary of the Invention

[0004] This invention proposes a Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent. It aims to solve the problems of existing pronunciation evaluation technologies that rely solely on acoustic features or external lip shape information and cannot accurately depict the dynamic posture of key articulatory organs such as tongue position. At the same time, it overcomes the technical defects of continuous speech flow conditions, such as difficulty in precise time alignment due to differences in speech rate and connected speech, difficulty in quantifying the boundary of the minimum opposition phoneme structure, lack of discriminative constraints on the direction of pronunciation correction, and lack of continuous visual optimization path in the training process. Thus, it achieves computable modeling, structured alignment, and quantitative correction of pronunciation pose.

[0005] The method for modeling Mandarin pronunciation pose driven by a multimodal AI agent includes the following steps:

[0006] S1. Collect tongue position coordinates, lip shape key points and mandibular opening and closing angle data of standard speakers under monosyllabic and continuous speech conditions. Through phoneme annotation, time segmentation and normalization processing, perform dynamic function modeling on each phoneme and construct phoneme dynamic pose manifold space.

[0007] Specifically, existing Mandarin pronunciation training systems mostly rely on acoustic parameters such as fundamental frequency, formants, or simple mouth shape diagrams as references, lacking unified modeling of tongue position, lip shape, and mandibular coordinated movements. This is especially problematic under continuous speech conditions, where it's difficult to form calculable standard pronunciation trajectories. This step collects tongue position diagram coordinates, key lip points, and mandibular opening and closing angle data from standard speakers under both monosyllabic and continuous speech conditions. Combined with phoneme annotation and time normalization, dynamic function modeling of the complete pronunciation cycle for each phoneme is performed, transforming discrete pronunciation posture data into continuously expressible pose functions. Furthermore, low-dimensional embedding is used to construct a dynamic pose manifold space for phonemes, enabling different phonemes to have measurable structural relationships within a unified geometric space. This overcomes the shortcomings of existing technologies that only focus on the acoustic level, achieving a shift from sound evaluation to pose modeling.

[0008] S2. Collect data on key points around the mouth and jaw movement during learners’ Mandarin and classical recitation training. Establish a cross-modal mapping function from the learner’s coordinate system to the standard tongue position map coordinate system based on the phoneme dynamic pose manifold space. Map the learner’s real-time pose data to the phoneme dynamic pose manifold space to obtain the learner’s mapped pose sequence with the same structure as the standard space.

[0009] Specifically, existing pronunciation correction techniques typically rely on acoustic comparisons or simple video demonstrations, making it difficult to structurally correlate learners' perioral movement data with standard tongue position diagrams. This step establishes a cross-modal mapping function from the learner's coordinate system to the standard tongue position diagram coordinate system, transforming the perioral key points and mandibular movement data acquired by the camera into a geometric representation space consistent with the standard model. This allows the learner's real-time pose data to be embedded into the dynamic pose manifold space of phonemes. This processing method solves the problem of data incomparability caused by differences in facial scale between different acquisition devices and individuals, enabling comparison between learner poses and standard poses within the same structural space.

[0010] S3. Collect word-level timestamp information of classic recitation texts, generate standard dynamic pose time series based on the timestamps of classic recitation texts, and take the learner mapped pose series as input. Solve the time deformation function according to the principle of minimizing the overall pose error, and align the learner mapped pose series with the corresponding standard dynamic pose time series to obtain the aligned standard dynamic pose series.

[0011] Specifically, in classical recitation scenarios, pronunciation exhibits continuous speech characteristics, with co-construction and transition effects between phonemes. Existing technologies mostly employ static comparison methods on a syllable-by-syllable basis, which struggle to achieve stable correction in continuous contexts. This step generates a standard dynamic pose time series based on the timestamps of the classical recitation text. Using the learner's mapped pose sequence as input, the time deformation function is solved through the principle of minimizing overall pose error, achieving dynamic alignment between the learner's trajectory and the standard trajectory in the time dimension. Through time alignment, the influence of differences in speech rate and rhythm can be eliminated, ensuring point-to-point correspondence between the two sequences under a unified time index. This provides a foundation for subsequent continuous deviation analysis and overcomes the incomparability caused by different speech rates in existing technologies.

[0012] S4. Calculate the pose deviation vector over continuous time based on the learner's mapped pose sequence and the standard dynamic pose sequence after time alignment; calculate the structural distance difference between the preset minimum set of opposing phonemes and the learner's pose; construct a discrimination constraint term based on the structural distance difference; couple the discrimination constraint term with the pose deviation energy to obtain the structured pose deviation vector field.

[0013] Specifically, traditional pronunciation assessments often rely on a single error index, failing to distinguish the structural differences between near-target phonemes and near-incorrect phonemes. This step, based on calculating the continuous-time deviation vector between the learner's mapped pose sequence and the standard dynamic pose sequence, introduces a minimum set of opposing phonemes. By calculating the structural distance difference between the learner's pose and the target phoneme and its opposing phonemes in manifold space, a discriminative constraint term is constructed and coupled with the pose deviation energy to form a structured pose deviation vector field. This approach considers not only the magnitude of the deviation but also its direction, enhancing the directionality and discriminative power of pronunciation correction.

[0014] S5. Based on the structured pose deviation vector field, the pose deviation energy is statistically analyzed in conjunction with the training cycle, and the phoneme discrimination improvement index is calculated according to the structural distance difference corresponding to the minimum opposing phoneme set; a two-dimensional convergent trajectory model is constructed based on the pose deviation energy and the phoneme discrimination improvement index to generate the articulation pose correction path.

[0015] Specifically, existing pronunciation training systems typically lack long-term quantitative tracking mechanisms, relying mainly on periodic scoring or subjective evaluation. This step, based on a structured pose deviation vector field, statistically analyzes the pose deviation energy within the training period and calculates the phoneme discrimination improvement index by combining the changes in structural distance differences corresponding to the minimum opposing phoneme set. By constructing a two-dimensional convergence trajectory model, the convergence degree of pose deviation and the improvement degree of phoneme discrimination ability are jointly expressed, thereby forming a repeatable and quantifiable pronunciation pose correction path.

[0016] The beneficial effects of the invention are:

[0017] This scheme constructs a dynamic pose manifold space for phonemes, unifying the modeling of tongue position map coordinates, lip shape key points, and mandibular opening and closing angles to create a continuous geometric trajectory representation of standard pronunciation. Through a cross-modal mapping function, learner-observable data is mapped to a pose space consistent with the standard, enabling numerical estimation of tongue position states. A time deformation function is used to perform temporal alignment of the continuous speech flow, ensuring pose comparisons are established within a consistent time frame. The scheme introduces the minimum opposition pair phoneme structural distance difference to construct a structured pose deviation vector field with discriminative constraints, ensuring that the pronunciation correction direction simultaneously satisfies the dual constraints of moving closer to the target and moving away from easily confused regions. Finally, a two-dimensional convergent trajectory model quantifies the changes in pose deviation energy and phoneme discrimination, enabling continuous tracking and path planning during the training process. Attached Figure Description

[0018] Figure 1 This is a flowchart of the method for modeling Mandarin pronunciation pose driven by a multimodal artificial intelligence agent proposed in this embodiment of the invention. Detailed Implementation

[0019] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings, but the scope of protection of the present invention is not limited to the following description.

[0020] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention; that is, the described embodiments are only a part of the embodiments of the invention, and not all of them. The components of the embodiments of the invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0021] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention. It should be noted that relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations.

[0022] Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or machine that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or machine. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or machine that includes said element.

[0023] The features and performance of the present invention will be further described in detail below with reference to embodiments.

[0024] Example 1

[0025] Among them, such as Figure 1 A method for modeling Mandarin pronunciation pose driven by a multimodal AI agent includes the following steps:

[0026] S1. Collect tongue position coordinates, lip shape key points and mandibular opening and closing angle data of standard speakers under monosyllabic and continuous speech conditions. Through phoneme annotation, time segmentation and normalization processing, perform dynamic function modeling on each phoneme and construct phoneme dynamic pose manifold space.

[0027] S2. Collect data on key points around the mouth and jaw movement during learners’ Mandarin and classical recitation training. Establish a cross-modal mapping function from the learner’s coordinate system to the standard tongue position map coordinate system based on the phoneme dynamic pose manifold space. Map the learner’s real-time pose data to the phoneme dynamic pose manifold space to obtain the learner’s mapped pose sequence with the same structure as the standard space.

[0028] S3. Collect word-level timestamp information of classic recitation texts, generate standard dynamic pose time series based on the timestamps of classic recitation texts, and take the learner mapped pose series as input. Solve the time deformation function according to the principle of minimizing the overall pose error, and align the learner mapped pose series with the corresponding standard dynamic pose time series to obtain the aligned standard dynamic pose series.

[0029] S4. Calculate the pose deviation vector over continuous time based on the learner's mapped pose sequence and the standard dynamic pose sequence after time alignment; calculate the structural distance difference between the preset minimum set of opposing phonemes and the learner's pose; construct a discrimination constraint term based on the structural distance difference; couple the discrimination constraint term with the pose deviation energy to obtain the structured pose deviation vector field.

[0030] S5. Based on the structured pose deviation vector field, the pose deviation energy is statistically analyzed in conjunction with the training cycle, and the phoneme discrimination improvement index is calculated according to the structural distance difference corresponding to the minimum opposing phoneme set; a two-dimensional convergent trajectory model is constructed based on the pose deviation energy and the phoneme discrimination improvement index to generate the articulation pose correction path.

[0031] Specifically, the above implementation method collects multidimensional data on tongue position, lip shape, and mandible of standard speakers in multiple contexts. Using time normalization and dynamic function modeling techniques, discrete sampling points are transformed into continuous and smooth motion trajectories, thereby constructing a phoneme dynamic pose manifold space containing the topological structure of pronunciation and establishing a benchmark for standard pronunciation. On this basis, to address the difficulty of learners' invisible tongue position, external observation data such as key points around the mouth and mandibular movements are collected. Using the prior distribution characteristics of the standard manifold space, a cross-modal mapping function is trained, nonlinearly projecting the learner's only external features into the standard coordinate system, reconstructing a full pose sequence containing inferred tongue position information, so that learner data has the basis for direct comparison with standard data in the same geometric space. Furthermore, to eliminate the comparison error caused by differences in speech rate, a reference trajectory is generated based on the timestamps of classic recitation texts. With the learner's reconstructed sequence as the goal, the standard sequence is nonlinearly time-warped by solving the time deformation function that minimizes the overall pose error, achieving strict alignment between the two in the time domain. In addition, not only is the instantaneous pose deviation vector after alignment calculated, i.e., the imitation error, but the concept of minimum opposite pairs in linguistics is also introduced to calculate the structural distance difference between the learner's pose and the center of easily confused phonemes. Based on this, a discriminative constraint term pointing away from the erroneous phonemes is constructed. The traction force towards the standard and the repulsive force away from confusion are weighted and coupled in the spatiotemporal domain to generate a high-dimensional structured pose deviation vector field, which can accurately describe the correction trend of pronunciation movement.

[0032] Furthermore, based on the vector field described in the above embodiments, the deviation energy within the training period is integrated in the time domain to quantify the overall error level. The improvement index of phoneme discrimination is calculated by combining the structural distance change of the minimum opposing pair. The key evaluation dimensions are mapped to the two-dimensional convergence trajectory model to evaluate the gap between the learner's current pronunciation state and the ideal convergence point. By calculating the gradient direction pointing to the optimal state in the model and mapping it back to the physical pronunciation space, a visual pronunciation posture correction path is generated. This path can intuitively guide the learner to adjust the tongue position and lip shape, so that their pronunciation action approaches the standard pronunciation along the trajectory of minimizing energy and maximizing discrimination, thereby achieving accurate correction of pronunciation posture.

[0033] Further, it should be noted that in Chinese, contrastive phonemes refer to two or more phonemes that, if they appear in the same phonetic environment, such as in the same positions of initial and final consonants, and can change the meaning of a word after substitution, then they form a contrastive relationship and belong to different phonemes respectively. Exemplarily, [p] (the initial consonant of "爸" in Mandarin) and [pʰ] (the initial consonant of "怕" in Mandarin) are typical contrastive phonemes. Because through the difference between aspirated and unaspirated sounds, the two words "爸" and "怕" with completely different meanings are distinguished. This contrastive relationship is the basis for Chinese pinyin to record phonetic sounds with different letters and distinguish words.

[0034] Further, the step S1 specifically includes the following sub-steps:

[0035] S101. Collect the original data of a standard speaker under single-syllable and continuous speech flow conditions. The original data includes a sequence of tongue map coordinates, a sequence of lip key points, and a sequence of mandibular opening and closing angles;

[0036] S102. Perform phoneme annotation and time segmentation on the original data, intercept the valid data segments corresponding to each phoneme, and uniformly map the time axis of the valid data segments to a normalized interval;

[0037] S103. Perform parametric fitting on each phoneme's valid data segment after normalization processing, establish a dynamic function model for each phoneme, and obtain a set of dynamic functions describing the standard pronunciation movement trajectory;

[0038] S104. Take the set of dynamic functions as sample points, maintain the local neighborhood relationship between sample points through manifold dimensionality reduction, and construct a phoneme dynamic pose manifold space containing the topological structure of standard pronunciation.

[0039] Specifically, the above implementation method converts the multi-modal motion data discretely collected during the standard pronunciation process into a phoneme motion model with a continuous geometric structure. By uniformly collecting the sequence of tongue map coordinates, the sequence of lip key points, and the sequence of mandibular opening and closing angles under single-syllable and continuous speech flow conditions, a time-series motion trajectory covering the entire pronunciation cycle is obtained; after performing phoneme annotation and time segmentation on the original data, the valid data segment corresponding to each phoneme is intercepted and mapped to a unified normalized time interval, making the pronunciation processes under different speech rates and different context conditions comparable on the time scale; further, by performing parametric fitting on the normalized data segments, the discrete trajectory is converted into a continuously differentiable dynamic function model, enabling the pronunciation process of each phoneme to be expressed in the form of a function, thereby achieving a stable description of the standard pronunciation movement law.

[0040] Furthermore, by embedding the dynamic functions of each phoneme as high-dimensional samples into the same feature space and maintaining the local neighborhood relationship between samples through manifold dimensionality reduction, the topological features between phonemes can be preserved while reducing dimensionality, so that the differences in articulation motion can be manifested in the form of geometric distance. The resulting phoneme dynamic pose manifold space not only expresses the dynamic evolution law within each phoneme, but also characterizes the structural relationship between different phonemes.

[0041] Furthermore, in step S103, for the first... The modeling process of the dynamic function model of phonemes is expressed as follows:

[0042] ;

[0043] In step S104, the constructed dynamic pose manifold space of the phonemes Represented as:

[0044] ;

[0045] The Indicates the first Each phoneme at the normalization time The standard dynamic pose vector, the Indices representing phonemes The value range is 1 to The This indicates the total number of preset standard phonemes. This represents the normalized time variable, with a value range of 1. The Represents the normalization time. The corresponding standard tongue position diagram horizontal coordinate, the Represents the normalization time. The corresponding standard tongue position diagram longitudinal coordinates and lip shape key point coordinates, the Represents the normalization time. The corresponding standard mandibular opening angle value, the angle mark The data attribute representing the standard speaker, wherein the superscript indicates the transpose operation of a vector or matrix, and the... The dynamic pose manifold space of phonemes, the This represents a manifold construction operator used to map high-dimensional dynamic functions to low-dimensional manifold structures. This represents set operations.

[0046] Additionally, it should be noted that in morphological modeling and computer vision, the aforementioned It is not a single numerical value, but a high-dimensional column vector. This is necessary for the same dynamic function... The text describes the spatial morphology of "key elements of tongue position contour and lip shape" simultaneously, using these discrete physical points. The axis coordinates are cascaded. Specifically, in physical space, using... Use scattered points to depict the outline of the tongue, using Using feature points to depict the lips, for example, the upper and lower lip peaks, and the corners of the mouth, then at time... :

[0047] ;

[0048] Similarly, It is also a dimension The column vector contains the x-axis coordinates of all these points. Therefore, In reality, it is a high-dimensional vector containing dozens or even hundreds of dimensions, at any given point in time. It can instantly recreate the complete 2D geometric topology of the entire oral cavity, including the tongue and the outside, including the lips and jaw.

[0049] Specifically, in step S103, the effective data segments of each phoneme after time normalization are expressed continuously. The original discrete sampling sequence is transformed into a dynamic function model describing the complete articulation cycle. By uniformly parameterizing and fitting the trajectory of tongue position map coordinate changes, the spatial distribution changes of lip shape key points, and the curve of mandibular opening and closing angle changes, each phoneme forms a continuous pose evolution function within the normalized time interval. The articulation process is no longer represented as a discrete set of points, but as a motion trajectory function that changes smoothly with time. The differences between different phonemes are reflected in the function shape, change amplitude, and trajectory path structure. Further, in step S104, the dynamic functions of all phonemes are input as a whole sample set into the manifold construction process. By maintaining the local neighborhood relationship between samples, dimensionality reduction mapping is performed, so that the phoneme motion patterns originally in the high-dimensional function space form a topologically structured distribution in the low-dimensional space. Phonemes with similar motion patterns are distributed in close proximity in this space, while phonemes with significant differences form relatively separated regions. The entire phoneme set constitutes a continuous and structurally stable dynamic pose manifold in the low-dimensional space.

[0050] Furthermore, step S2 specifically includes the following sub-steps:

[0051] S201. Real-time acquisition of learner's perioral key point data and mandibular movement data to form learner observation vector sequence;

[0052] S202. Based on the standard data distribution in the dynamic pose manifold space of phonemes, establish a cross-modal mapping function from learner observation vectors to the standard tongue position map coordinate system;

[0053] S203. Input the learner's real-time observation vector sequence into the cross-modal mapping function to calculate the learner's estimated coordinates in the standard coordinate system;

[0054] S204. Arrange the calculated estimated coordinates in chronological order to generate a learner mapping pose sequence consistent with the standard spatial structure.

[0055] Specifically, the above implementation method collects real-time data on key points around the learner's mouth and mandibular movement to form a time-varying sequence of observation vectors, enabling the learner's pronunciation process to be expressed in a unified vector form. Furthermore, using the distribution structure of standard pronunciation data in the phoneme dynamic pose manifold space as a reference, a cross-modal mapping relationship is established from the learner's observation vector space to the standard tongue position diagram coordinate system, ensuring that data from different modal sources have a geometrically corresponding basis. The learner's real-time observation vector sequence is input into the cross-modal mapping relationship, and coordinate transformation and scale matching are performed on the observation data at each moment to obtain the learner's estimated pose expression in the standard tongue position diagram coordinate system. In addition, the estimation results at each moment are arranged in chronological order to form a continuous learner mapped pose sequence, ensuring that this sequence remains consistent with the standard phoneme dynamic pose space in both structural and temporal dimensions.

[0056] Furthermore, step S202 specifically includes the following sub-steps:

[0057] S2021. Extract paired training sample sets from standard speaker data, each sample including perioral key points and mandibular data and corresponding phoneme dynamic pose manifold coordinates;

[0058] S2022. Construct a mapping network structure that includes a nonlinear feature mapping layer and a linear regression layer, and define a kernel function to map the source domain data to a high-dimensional regenerating kernel Hilbert space;

[0059] S2023. Establish a mapping error loss function and introduce a Laplacian regularization term. Train the mapping network using the gradient descent algorithm to generate a cross-modal mapping function.

[0060] It should be noted that its core optimization objective can be expressed as:

[0061] ;

[0062] ;

[0063] Wherein, it represents the first The observation vector corresponding to each standard pronunciation sample contains data on key points around the mouth and mandibular opening and closing angles. Indicates and The corresponding coordinates of the standard dynamic pose vector in the manifold space; Represents a cross-modal mapping network; Indicates the total number of training samples; This represents the nonlinear eigenmap induced by the kernel function; This represents the set of weight parameters for a cross-modal mapping network. This represents the weighting coefficient of the Laplace regularization term; This represents the graph Laplacian matrix constructed based on the adjacency relationships of the dynamic pose manifold space of phonemes; This represents the matrix trace operation.

[0064] The above implementation method specifically describes a process that unfolds the nonlinear distribution in the observation space to a high-dimensional feature space through kernel mapping. A linear regression relationship is established in this space, creating a stable mapping between the perioral and mandibular data and the pose of the standard tongue position map. Simultaneously, the mapping results are constrained by a Laplace regularization term to maintain the neighborhood topology in the manifold space, ensuring that adjacent phoneme samples maintain their geometric proximity after mapping. This, in turn, guarantees both mapping accuracy and phoneme structural consistency.

[0065] Furthermore, in step S204, the process of generating the learner's mapped pose sequence is specifically represented as follows:

[0066] ;

[0067] ;

[0068] Among them, the Indicates the time of real-time data acquisition The calculated full pose vector of the learner in standard space includes the inferred tongue coordinates and the visible lip and mandible coordinates. This represents the discrete-time frame index during the real-time acquisition process, where the time represents... The observation vector is composed of the coordinates of key perioral points and the vertical displacement data of the mandible collected from the learner. This represents the cross-modal mapping function constructed by partial least squares regression. The set of trained weight parameters representing the cross-modal mapping function is obtained through training based on the constructed manifold space data. This represents the set of mapped pose sequences of a learner within a complete pronunciation segment. This represents the total number of sampling frames for the current pronunciation segment. Specifically, the external motion information acquired by the learner at the observable level is transformed into an internal pose representation isomorphic to the standard pronunciation space. During continuous pronunciation, the learner forms an observation vector at a discrete-time frame index t. This vector, composed of the coordinates of key points around the mouth and the vertical displacement data of the mandible, reflects the external motion state that can be directly obtained at the current moment; it is obtained through a pre-trained cross-modal mapping function. Under the constraints of the weight parameter set, regression calculation is performed on the observation vector at each time step to obtain the estimated pose vector in the standard space. This vector includes not only visible lip and jaw information but also tongue position coordinates inferred through statistical correlation, thus providing a numerical representation of the tongue position state, which cannot be directly observed. Furthermore, the values ​​corresponding to each discrete time step t are... Arranged chronologically, this forms a learner-mapped pose sequence within a complete pronunciation segment. This sequence maintains consistency with the standard phoneme dynamic pose manifold space in the structural dimension and continuity in the temporal dimension, so that the learner's pronunciation process is represented as a motion trajectory in the standard space. In this way, the data that was originally limited to the visible movements is transformed into a holistic pose path that includes the coordination of tongue position, lip shape and mandible, so that the pronunciation movement is regarded as a unified geometric movement process, rather than a collection of separate local movements.

[0069] Furthermore, step S3 specifically includes the following sub-steps:

[0070] S301. Analyze the word-level timestamp information of classic recitation texts, extract the corresponding standard dynamic functions from the phoneme dynamic pose manifold space, and generate a standard dynamic pose time series corresponding to the text;

[0071] S302. Using the learner's mapped pose sequence as the target and the standard dynamic pose time series as a reference, establish a target functional that includes a pose distance term and a time smoothing term;

[0072] S303. Based on the principle of minimizing the overall pose error, the target functional is optimized to obtain the optimal time deformation function;

[0073] S304. The time axis of the standard dynamic pose time sequence is resampled using the optimal time deformation function to obtain a standard dynamic pose sequence that is strictly aligned with the learner sequence in time.

[0074] Specifically, by analyzing word-level timestamp information in classic recitation texts, the start and end intervals of each phoneme in the speech flow can be determined. Corresponding standard dynamic functions are extracted from the phoneme dynamic pose manifold space and expanded according to the text's temporal order to form a standard dynamic pose time sequence corresponding to the text's rhythm. This sequence not only reflects the internal movement patterns of individual phonemes but also includes the natural connections between phonemes, giving standard pronunciation a continuous structural expression in the temporal dimension. Furthermore, using the learner's mapped pose sequence as the target trajectory and the standard dynamic pose time sequence as the reference trajectory, a target functional containing pose distance and temporal smoothing terms is constructed. This ensures that the temporal deformation process maintains the continuity and monotonicity of the temporal mapping while reducing pose differences. By optimizing this target functional, the optimal temporal deformation function is obtained. This function is then used to resample the standard dynamic pose time sequence along the time axis, ensuring that the resampled standard sequence corresponds point-by-point with the learner's sequence in the time index. Through this method, speech rate and rhythm differences are transformed into computable temporal deformation problems, allowing pronunciation comparison to be established on a structurally consistent temporal framework, rather than a simple frame-to-frame correspondence.

[0075] Furthermore, in step S303, the objective functional optimization process of the optimal time deformation function is specifically expressed as follows:

[0076] ;

[0077] Among them, the This represents the optimal time deformation function that minimizes the overall pose error. This represents the operator for minimizing the parameters of a function, the... Derived from the learner mapping pose sequence generated in step S2, the This indicates that, based on the timestamps of the classic recitation text, the standard reference trajectory continuous function, constructed by concatenating the dynamic function model in step S1, is invoked. This represents the time deformation function to be solved, used to map the learner's time axis to the time axis of a standard reference trajectory. The Euclidean norm of a vector is denoted as . The weighting coefficients of the time warp regularization term are used to limit the severity of time distortion. This represents the derivative of the time-varying function, i.e., the local time scaling factor.

[0078] Specifically, the goal of the above implementation method is to minimize the overall difference between the learner's mapped pose trajectory and the standard reference trajectory while maintaining the smoothness of the temporal mapping. This is achieved by constructing an objective functional that includes a pose distance term and a time variation regularization term, using the time deformation function as the variable to be optimized, and continuously solving for it throughout the entire pronunciation segment interval. The pose distance term measures the cumulative Euclidean distance between the learner's pose vector and the standard reference trajectory under the current temporal mapping relationship; the time regularization term constrains the first derivative of the time deformation function, keeping the time scaling rate close to a unit value and avoiding excessive local time compression or stretching.

[0079] Within this framework, the optimal time deformation function that minimizes the overall error is obtained through variational solving or discrete optimization iteration of the objective functional. This function describes the continuous mapping relationship between the learner's time axis and the standard time axis, and its derivative reflects the degree of scaling of the local speech rate. Furthermore, it should be noted that the optimized time deformation function, while satisfying the monotonically increasing constraint, achieves overall alignment between pose trajectories. This ensures that temporal differences are absorbed into the mapping relationship in a functional form, rather than being directly added to the pose error, thus guaranteeing that pose comparisons are based on consistent temporal correspondence.

[0080] Furthermore, in step S304, the elements in the aligned standard dynamic pose sequence are represented as follows:

[0081] ;

[0082] Among them, the This indicates that after time alignment, at time... The corresponding standard dynamic pose vector.

[0083] Specifically, the processing logic of the standard dynamic pose sequence is based on the premise that the optimal time deformation function has been determined. The time parameters of the standard reference trajectory are remapped onto the learner's time axis. For each discrete moment in the learner sequence... Through the optimal time deformation function Calculate its corresponding time position in the standard reference trajectory, and then substitute that time position into the standard reference trajectory function. The standard dynamic pose vector after resampling is obtained. This process is not a simple truncation or interpolation, but rather a functional-level repositioning of the standard trajectory based on a continuous temporal mapping relationship, ensuring that the temporal structure of the standard trajectory aligns with the learner's temporal structure. After this mapping, the standard dynamic pose sequence corresponds point-by-point with the learner's mapped pose sequence in terms of temporal index. At the same time t, the two vectors represent the standard pronunciation state and the learner's pronunciation state in a unified temporal context, respectively. The temporal differences are fully absorbed into the temporal deformation function, and the pose vector itself no longer carries rhythmic difference factors.

[0084] Furthermore, step S4 specifically includes the following sub-steps:

[0085] S401. Based on the learner's mapped pose sequence and the aligned standard dynamic pose sequence, calculate the pose deviation vector frame by frame, and extract the magnitude of the vector as the pose deviation energy.

[0086] S402. Obtain the preset set of minimum contrasting phonemes, calculate the Euclidean distance between the learner's current pose and the center of each phoneme in the minimum contrasting pair, and determine the structural distance difference;

[0087] S403. Construct a discrimination constraint term based on the structural distance difference. When the structural distance difference is less than the safety threshold, generate a repulsive force vector pointing away from the opposing pitch position.

[0088] S404. Weighted coupling of the repulsion vector corresponding to the discrimination constraint term and the pose deviation vector is performed to generate a structured pose deviation vector field for describing the correction trend.

[0089] Specifically, the above implementation calculates the difference vector between the learner's mapped pose sequence and the time-aligned standard dynamic pose sequence frame by frame under a unified time index, clearly representing the spatial offset at each moment. The magnitude of this difference vector is calculated to obtain the pose deviation energy at the corresponding moment, which characterizes the spatial distance between the learner's current pronunciation state and the standard pronunciation state. This process transforms the pronunciation error into a measurable vector form, rather than a single scalar score. Furthermore, combining a preset minimum set of opposing phonemes, the Euclidean distance between the learner's current pose and the center of the opposing phoneme is calculated in the dynamic pose manifold space. By comparing the distance difference between the target phoneme and the opposing phoneme, a structural distance difference index is obtained. When the structural distance difference is below a safety threshold, it indicates that the current pose is approaching the opposing phoneme region in the manifold space. At this time, a repulsive vector pointing away from the center of the opposing phoneme is generated based on the distance gradient direction, giving the offset direction a distinguishing significance. In addition, the repulsive force vector is coupled with the original pose deviation vector according to the weighting coefficient to obtain a structured pose deviation vector field that comprehensively considers the magnitude and direction of the error. This makes the correction trend reflect both the need to move closer to the standard pose and the structural constraints of moving away from the easily confused pitch region.

[0090] It should be noted that the time-aligned standard dynamic pose sequence corresponds point-by-point with the learner's mapped pose sequence in terms of time index.

[0091] Furthermore, in step S404, the structured pose deviation vector field is specifically represented as follows:

[0092] ;

[0093] Among them, the Represents the pose deviation vector, the This represents the discriminant constraint term, the... The coupling weight coefficients representing the pose deviation vector, the The coupling weight coefficients of the discrimination constraint vector are represented by the following: Indicates time The synthesized structured pose deviation vector field is used to indicate the optimal direction for pronunciation correction.

[0094] Specifically, the above implementation method uses vector superposition to uniformly represent the deviation information from two different sources. At each discrete time t, the pose deviation vector... Reflects the direct spatial difference between the learner's current pose and the standard pose, with its direction pointing towards the standard trajectory position; discrimination constraint terms. This originates from the structural distance analysis of the minimum opposing phonemes. When the learner's pose approaches the opposing phoneme region in the manifold space, this vector is generated along a direction away from the center of the opposing phoneme. Geometrically, these two represent the tendency to approach the target and the tendency to avoid entering the opposing region, respectively. By introducing coupling weight coefficients and... A linear combination of the two types of vectors is performed to form a composite vector. The weighting coefficients are used to adjust the proportion of influence of the two types of constraints at different training stages, so that the spatial correction maintains both the convergence trend towards the standard pose and the stability of the phoneme distinction boundary; thus, the obtained... It is no longer a simple error vector pointing to the standard position, but a corrective direction expression that comprehensively considers the proximity of the target and the structural distinguishability, so that the pronunciation adjustment evolves along a discriminative trajectory in the manifold space.

[0095] Furthermore, the specific calculation process for the aforementioned pose deviation vector and discrimination constraint term is as follows:

[0096] In step S401, the time is calculated. pose deviation vector Represented as:

[0097] ;

[0098] In step S403, the constructed discriminant constraint vector for the minimum set of contrasting phonemes Represented as:

[0099] ;

[0100] The Derived from the aligned standard dynamic pose sequence generated in step S3, the This represents the preset minimum set of opposite phoneme indices corresponding to the current target phoneme. Represents a set The index of the contrasting phonemes in the text, the This represents the Heaviside step function, which outputs 1 when the input is greater than 0, and 0 otherwise. This indicates the preset safe distance threshold for phoneme confusion. Indicates contrasting phonemes The cluster center coordinate vectors in the constructed manifold space, the This represents the numerical stability constant to prevent the denominator from being zero.

[0101] Furthermore, step S5 specifically includes the following sub-steps:

[0102] S501. Based on the structured pose deviation vector field, the pose deviation energy is integrated over time within a single training cycle to obtain the total energy value of the training cycle.

[0103] S502. Based on the structural distance difference corresponding to the minimum set of opposite phonemes, calculate the average interval between the learner's pose and the boundary of the opposite phonemes, and calculate the phoneme discrimination improvement index.

[0104] S503. Construct a two-dimensional convergence trajectory model with the total energy value on the horizontal axis and the phoneme discrimination improvement index on the vertical axis, and map the state of the current training cycle to this model;

[0105] S504. In the two-dimensional convergent trajectory model, plan the optimal gradient direction from the current state to the target state, and map the gradient direction back to the tongue position map coordinate system to generate a visualized pronunciation pose correction path.

[0106] Specifically, within a single training cycle, the changes in the structured pose deviation vector field over time are accumulated. The pose deviation energy at each moment is integrated over time to obtain the total energy value for that cycle. This reflects the learner's overall deviation from the standard pose trajectory during that stage. This value reflects the comprehensive intensity of spatial errors during continuous pronunciation, rather than a single instantaneous error. Furthermore, based on the structural distribution of each phoneme in the minimum set of contrasting phonemes, the average interval between the learner's pose trajectory and the boundary of the contrasting phonemes in the manifold space is statistically analyzed. By comparing the changes in this interval within different training cycles, a phoneme discrimination improvement index is formed to characterize the learner's stability in moving away from easily confused regions in the structural space. A two-dimensional convergent trajectory model is constructed using the total energy value as the horizontal axis and the phoneme discrimination improvement index as the vertical axis. Each training cycle corresponds to a state point in the model, and the trajectory of this state point in the plane reflects the synchronous changes in pronunciation quality and discrimination ability. In addition, the optimal descent direction is determined based on the gradient distribution of the target convergence region within this plane, and this direction is transformed back to the tongue bitmap coordinate space through a geometric mapping relationship, so that the training trend is presented as a continuous pose adjustment trajectory in the form of a spatial path.

[0107] Furthermore, in step S504, the pronunciation posture correction path is specifically represented as follows:

[0108] ;

[0109] ;

[0110] Among them, the This indicates the articulation posture correction path, the Indicates the first The state coordinate vector in the two-dimensional convergent trajectory model for each training cycle, with its horizontal axis representing the total energy of pose deviation and its vertical axis representing the phoneme discrimination improvement index, where represents the number of elements in the minimum set of contrasting phonemes. This represents the preset minimum set of opposite phoneme indices corresponding to the current target phoneme, the set being represented. The index of the contrasting phonemes in the text, the Indicates contrasting phonemes The coordinate vectors of the cluster centers in the constructed manifold space.

[0111] Specifically, the above implementation transforms the state change direction in the two-dimensional convergent trajectory model into a specific spatial correction path. The state vector of the next training cycle It consists of two components, the horizontal axis of which is the energy integral value of the structured pose deviation vector field over the entire speech segment, i.e., for The cumulative results over the time interval reflect the overall spatial deviation intensity of the current cycle; the vertical axis represents the learner's pose relative to the centers of each phoneme in the minimum set of phonemes. The average minimum distance represents the degree of separation between the state vector and easily confused regions in the manifold space. The gradient of this state vector in the two-dimensional plane... It reflects the changing trend of the current training state relative to the target convergence region; further, it maps the gradient direction in the two-dimensional plane through a mapping operator. This is converted into a directional expression consistent with the pose space, establishing a correspondence between changes in the plane and motion directions in the manifold space; the pose deviation vector is then structured. Provides real-time revised trends at various times, and compares them with... By combining these steps, a vocalization pose correction path can be obtained. This path remains continuous in the time dimension and comprehensively considers the overall energy decline trend and the differentiation improvement trend in the spatial dimension, so that the pronunciation adjustment manifests as a trajectory of gradual evolution along a specific geometric direction, rather than an isolated instantaneous correction.

[0112] Example 2

[0113] Furthermore, as a preferred embodiment of the above embodiments, a Mandarin pronunciation pose modeling system driven by a multimodal artificial intelligence agent is proposed. This system is implemented based on the Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in any one of the embodiments above, and specifically includes:

[0114] The system comprises a standard pronunciation manifold construction module, a learner cross-modal mapping module, a spatiotemporal alignment and deviation analysis module, and a correction path generation module. These functional modules achieve a complete closed loop from pronunciation mechanism modeling to visualization feedback through tight data flow coupling. The system is equipped with a high-standard standard pronunciation manifold construction module, which integrates a multimodal physiological signal acquisition unit to capture tongue position coordinates, lip shape key points, and mandibular opening and closing angle data of standard pronunciation speakers under different pronunciation conditions, such as monosyllabic and continuous speech. The module's internal processing unit performs fine phoneme annotation and time segmentation on the massive amount of raw data collected, mapping all valid data segments to a unified normalized time interval. Furthermore, it utilizes parametric fitting techniques to dynamically model the motion trajectory of each phoneme, generating a continuous and smooth set of reference trajectories. Based on this, the module further applies a manifold learning algorithm to mine local neighborhood relationships between sample points, constructing a dynamic pose manifold space that can characterize the topological structure of standard pronunciation, thus providing a geometrically constrained physical reference system for the pronunciation evaluation of the entire system.

[0115] To address the physical limitations of directly observing the tongue position inside the learner's mouth, the system also includes a learner cross-modal mapping module. This module is responsible for reconstructing the learner's full vocal poses without intrusion. During Mandarin and classical recitation training, the module collects real-time data on key points around the learner's mouth and mandibular movement data through external visual sensors, and retrieves a pre-constructed dynamic pose manifold space as a prior knowledge base. Through a trained cross-modal mapping function, the module nonlinearly projects the learner's only external dimension observation data into a high-dimensional coordinate system of the standard manifold space, calculating the full pose vector including the inferred tongue position coordinates. A mathematical mapping is established from the learner coordinate system to the standard tongue position map coordinate system, generating a learner mapped pose sequence with the same geometric structure as the standard space, enabling the learner's vocal movements to be quantitatively compared with standard pronunciation in the same mathematical dimension.

[0116] To eliminate the impact of individual speech rate differences and rhythmic fluctuations on assessment accuracy, and to deeply analyze the essence of pronunciation defects, the system integrates a spatiotemporal alignment and deviation analysis module. This module first extracts corresponding standard trajectories from the manifold space based on word-level timestamp information of classic recitation texts, generating a theoretically standard dynamic pose time series. Further, the module takes the learner's mapped pose sequence as input and, based on the principle of minimizing overall pose error, solves for the optimal time deformation function through an optimization algorithm. This performs nonlinear scaling and normalization on the time axis of the standard sequence, resulting in a standard dynamic pose sequence strictly aligned with the learner in the time dimension. After completing time alignment, the module not only calculates the instantaneous pose deviation vector over continuous time but also introduces the concept of minimum opposite pairs from linguistics to calculate the structural distance difference between the learner's current pose and the center of easily confused phonemes. Based on this distance difference, the module constructs a discriminative constraint term pointing away from erroneous phonemes and weightedly couples this constraint term with the pose deviation energy to generate a structured pose deviation vector field that contains both a pulling force towards standard pronunciation and a repulsive force away from confused phonemes.

[0117] Furthermore, the system transforms complex vector fields into intuitive training guidance through a correction path generation module. This module, based on a structured pose deviation vector field, performs temporal integration statistics on pose deviation energy throughout the entire training cycle to quantify the overall degree of pronunciation errors. Simultaneously, it calculates the phoneme discrimination improvement index based on the structural distance difference change of the minimum set of opposite pairs. The module maps these two key performance indicators to a two-dimensional convergence trajectory model to evaluate the distance and orientation of the learner's current pronunciation state relative to the ideal convergence target. Based on this evaluation result, the module plans the optimal gradient direction pointing to the low-energy, high-discrimination state and maps this gradient back to the physical pronunciation space, generating a dynamic pronunciation pose correction path. This path can be presented visually on the terminal interface, precisely guiding the learner to adjust their tongue trajectory and lip opening and closing, so that their pronunciation movements approach standard pronunciation along the optimal path, thereby achieving efficient and accurate pronunciation correction training.

[0118] Example 3

[0119] Furthermore, as a preferred embodiment of the above embodiments, an application scenario for a method for mapping the pronunciation postures of Mandarin and classical recitation based on a visual tongue position map is proposed. This method can be applied to reading training in primary and secondary school Chinese classrooms, a training platform for broadcasting and hosting majors, a Mandarin proficiency test PSC preparation system, and an online classical recitation training platform for adult learners.

[0120] Taking a middle school's Mandarin reading improvement project as an example, the system collected training data from 30 students over 12 weeks, with 3 training sessions per week, each lasting approximately 8 minutes, for a total sampling time of approximately 8640 minutes. During the training, each student's facial video data (1920×1080 resolution, 60 fps sampling rate) was acquired via a high-definition camera. Combined with a key point detection algorithm, 68 perioral key points and mandibular opening and closing angle data were extracted, forming approximately 60 sets of observation vectors per second. Simultaneously, the system incorporates a phoneme dynamic pose manifold space constructed by a standard speaker, containing standard dynamic trajectory data for approximately 80 phoneme units, including 21 initials, 39 finals, and common variations such as neutral tone and retroflexion in Mandarin. Each phoneme unit contains an average of 3000–5000 standard pose samples.

[0121] In practical applications, the system utilizes a cross-modal mapping function to map approximately 60 frames per second of perioral and mandibular data from the learner to the tongue position map manifold space, generating a corresponding learner-mapped pose sequence. Subsequently, the standard dynamic pose time sequence is resampled using a time deformation function to achieve strict time alignment with the learner sequence. Based on frame-by-frame calculation of pose deviation energy, the system combines a minimum set of contrasting phonemes (e.g., approximately 15 sets of high-frequency easily confused phoneme pairs such as b-p, d-t, g-k, etc.) to calculate the structural distance difference and construct a structured pose deviation vector field. Within a single training cycle (approximately 480 seconds), the pose deviation energy is integrated over time, resulting in a decrease in the total energy value from the initial E1=1850 to E2=620. Simultaneously, the phoneme discrimination improvement index increases from 0.42 to 0.81, and a visualized convergence path is formed in the two-dimensional convergence trajectory model.

[0122] After 12 weeks of training, experimental data showed that: the misjudgment rate of retroflex consonants decreased from the initial 28.6% to 6.4%; the confusion rate of front and back nasal consonants decreased from 22.1% to 5.9%; the average score of students in the simulated Mandarin proficiency test increased by 9.3 points (out of 100); the total energy of positional deviation decreased by an average of about 65%; and the phoneme discrimination index increased by an average of about 78%.

[0123] In the practical training scenario of broadcasting and hosting majors, this method can generate pronunciation posture correction paths in real time and display the optimization direction in the form of dynamic trajectory in the tongue position map coordinate system on the display screen. This allows learners to observe the adjustment trend of the deviation vector field indication within a single sentence training session (about 10 seconds), realizing the transformation from auditory error correction to visual structural error correction.

[0124] Therefore, this implementation method is not suitable for scenarios such as Mandarin teaching, classical poetry recitation training, language rehabilitation auxiliary training, and professional broadcasting voice training that require high-precision pronunciation modeling, continuous speech flow structure alignment, and visual correction path feedback. It is especially suitable for deployment in smart terminals or teaching all-in-one machine environments with video acquisition and real-time computing capabilities.

[0125] The above description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the forms disclosed herein and should not be construed as excluding other embodiments. It can be used in various other combinations, modifications, and environments, and can be altered within the scope of the concept described herein through the above teachings or related technologies or knowledge. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention should be within the protection scope of the appended claims.

Claims

1. A method for modeling Mandarin pronunciation pose driven by a multimodal artificial intelligence agent, characterized in that, Includes the following steps: S1. Collect tongue position coordinates, lip shape key points and mandibular opening and closing angle data of standard speakers under monosyllabic and continuous speech conditions. Through phoneme annotation, time segmentation and normalization processing, perform dynamic function modeling on each phoneme and construct phoneme dynamic pose manifold space. S2. Collect data on key points around the mouth and jaw movement during learners’ Mandarin and classical recitation training. Establish a cross-modal mapping function from the learner’s coordinate system to the standard tongue position map coordinate system based on the phoneme dynamic pose manifold space. Map the learner’s real-time pose data to the phoneme dynamic pose manifold space to obtain the learner’s mapped pose sequence with the same structure as the standard space. S3. Collect word-level timestamp information of classic recitation texts, generate standard dynamic pose time series based on the timestamps of classic recitation texts, and take the learner mapped pose series as input. Solve the time deformation function according to the principle of minimizing the overall pose error, and align the learner mapped pose series with the corresponding standard dynamic pose time series to obtain the aligned standard dynamic pose series. S4. Calculate the pose deviation vector over continuous time based on the learner's mapped pose sequence and the standard dynamic pose sequence after time alignment; calculate the structural distance difference between the preset minimum set of opposing phonemes and the learner's pose; construct a discrimination constraint term based on the structural distance difference; couple the discrimination constraint term with the pose deviation energy to obtain the structured pose deviation vector field. S5. Based on the structured pose deviation vector field, the pose deviation energy is statistically analyzed in conjunction with the training cycle, and the phoneme discrimination improvement index is calculated according to the structural distance difference corresponding to the minimum opposing phoneme set; a two-dimensional convergent trajectory model is constructed based on the pose deviation energy and the phoneme discrimination improvement index to generate the articulation pose correction path. Step S5 specifically includes the following sub-steps: S501. Based on the structured pose deviation vector field, the pose deviation energy is integrated over time within a single training cycle to obtain the total energy value of the training cycle. S502. Based on the structural distance difference corresponding to the minimum set of opposite phonemes, calculate the average interval between the learner's pose and the boundary of the opposite phonemes, and calculate the phoneme discrimination improvement index. S503. Construct a two-dimensional convergence trajectory model with the total energy value on the horizontal axis and the phoneme discrimination improvement index on the vertical axis, and map the state of the current training cycle to this model; S504. In the two-dimensional convergent trajectory model, plan the optimal gradient direction from the current state to the target state, and map the gradient direction back to the tongue position map coordinate system to generate a visualized articulation pose correction path. In step S504, the pronunciation posture correction path is specifically represented as follows: ; ; Among them, the This indicates the articulation posture correction path, the Indicates the first The state coordinate vector in the two-dimensional convergent trajectory model for each training cycle, with the horizontal axis representing the total energy of pose deviation and the vertical axis representing the phoneme discrimination improvement index. The number of elements in the minimal set of opposite phonemes, the This represents the preset minimum set of opposite phoneme indices corresponding to the current target phoneme. Represents a set The index of the contrasting phonemes in the text, the Show contrasting phonemes The cluster center coordinate vectors in the constructed manifold space, the Indicates time The synthesized structured pose deviation vector field is used to indicate the optimal direction for pronunciation correction. This indicates the total number of sampled frames for the current pronunciation segment. Indicates the time of real-time data acquisition The calculated full pose vector of the learner in standard space includes the inferred tongue coordinates and the visible lip and mandible coordinates. This represents the discrete time frame index during the real-time acquisition process.

2. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 1, characterized in that, Step S1 specifically includes the following sub-steps: S101. Collect raw data from standard speakers under monosyllabic and continuous speech conditions, the raw data including tongue position map coordinate sequence, lip shape key point sequence and mandibular opening and closing angle sequence; S102. The original data is phoneme-labeled and time-segmented, the effective data segments corresponding to each phoneme are extracted, and the time axis of the effective data segments is uniformly mapped to the normalized interval; S103. Perform parameterized fitting on the effective data segments of each phoneme after normalization, establish a dynamic function model for each phoneme, and obtain a set of dynamic functions describing the motion trajectory of standard pronunciation; S104. Using the set of dynamic functions as sample points, and maintaining the local neighborhood relationships between sample points through manifold dimensionality reduction, a dynamic pose manifold space of phonemes containing the standard pronunciation topology is constructed.

3. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 2, characterized in that, In step S103, for the first The modeling process of the dynamic function model of phonemes is expressed as follows: ; In step S104, the constructed dynamic pose manifold space of the phonemes Represented as: ; The Indicates the first Each phoneme at the normalization time The standard dynamic pose vector, the Indices representing phonemes The value range is 1 to The This indicates the total number of preset standard phonemes. This represents the normalized time variable, with a value range of 1. The Represents the normalization time. The corresponding standard tongue position diagram horizontal coordinate, the Represents the normalization time. The corresponding standard tongue position diagram longitudinal coordinates and lip shape key point coordinates, the Represents the normalization time. The corresponding standard mandibular opening angle value, the angle mark The superscript represents the data attribute of the standard speaker. The transpose operation represents a vector or matrix. The dynamic pose manifold space of phonemes, the This represents a manifold construction operator used to map high-dimensional dynamic functions to low-dimensional manifold structures. This represents set operations.

4. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 1, characterized in that, Step S2 specifically includes the following sub-steps: S201. Real-time acquisition of learner's perioral key point data and mandibular movement data to form learner observation vector sequence; S202. Based on the standard data distribution in the dynamic pose manifold space of phonemes, establish a cross-modal mapping function from learner observation vectors to the standard tongue position map coordinate system; S203. Input the learner's real-time observation vector sequence into the cross-modal mapping function to calculate the learner's estimated coordinates in the standard coordinate system; S204. Arrange the calculated estimated coordinates in chronological order to generate a learner mapping pose sequence consistent with the standard spatial structure.

5. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 4, characterized in that, Step S202 specifically includes the following sub-steps: S2021. Extract paired training sample sets from standard speaker data, each sample including perioral key points and mandibular data and corresponding phoneme dynamic pose manifold coordinates; S2022. Construct a mapping network structure that includes a nonlinear feature mapping layer and a linear regression layer, and define a kernel function to map the source domain data to a high-dimensional regenerating kernel Hilbert space; S2023. Establish a mapping error loss function and introduce a Laplacian regularization term. Train the mapping network using the gradient descent algorithm to generate a cross-modal mapping function.

6. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 4, characterized in that, In step S204, the process of generating the learner's mapped pose sequence is specifically represented as follows: ; ; Among them, the Indicates time The observation vector is composed of the coordinates of key perioral points and the vertical displacement data of the mandible collected from the learner. This represents the cross-modal mapping function constructed by partial least squares regression. The set of trained weight parameters representing the cross-modal mapping function is obtained through training based on the constructed manifold space data. This represents the set of mapped pose sequences of a learner within a complete pronunciation segment.

7. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 1, characterized in that, Step S3 specifically includes the following sub-steps: S301. Analyze the word-level timestamp information of classic recitation texts, extract the corresponding standard dynamic functions from the phoneme dynamic pose manifold space, and generate a standard dynamic pose time series corresponding to the text; S302. Using the learner's mapped pose sequence as the target and the standard dynamic pose time series as a reference, establish a target functional that includes a pose distance term and a time smoothing term; S303. Based on the principle of minimizing the overall pose error, the target functional is optimized to obtain the optimal time deformation function; S304. The time axis of the standard dynamic pose time sequence is resampled using the optimal time deformation function to obtain a standard dynamic pose sequence that is strictly aligned with the learner sequence in time.

8. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 7, characterized in that, In step S303, the objective functional optimization process of the optimal time deformation function is specifically expressed as follows: ; Among them, the This represents the optimal time deformation function that minimizes the overall pose error. Represents about the function The operator for minimizing parameters, the Derived from the learner mapping pose sequence generated in step S2, the This indicates that, based on the timestamps of the classic recitation text, the standard reference trajectory continuous function, constructed by concatenating the dynamic function model in step S1, is invoked. This represents the time deformation function to be solved, used to map the learner's time axis to the time axis of a standard reference trajectory. The Euclidean norm of a vector is denoted as . The weighting coefficients of the time warp regularization term are used to limit the severity of time distortion. This represents the derivative of the time-varying function, i.e., the local time scaling factor.

9. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 8, characterized in that, In step S304, the elements in the aligned standard dynamic pose sequence are represented as follows: ; Among them, the This indicates that after time alignment, at time... The corresponding standard dynamic pose vector.

10. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 1, characterized in that, Step S4 specifically includes the following sub-steps: S401. Based on the learner's mapped pose sequence and the aligned standard dynamic pose sequence, calculate the pose deviation vector frame by frame, and extract the magnitude of the vector as the pose deviation energy. S402. Obtain the preset set of minimum contrasting phonemes, calculate the Euclidean distance between the learner's current pose and the center of each phoneme in the minimum contrasting pair, and determine the structural distance difference; S403. Construct a discrimination constraint term based on the structural distance difference. When the structural distance difference is less than the safety threshold, generate a repulsive force vector pointing away from the opposing pitch position. S404. Weighted coupling of the repulsion vector corresponding to the discrimination constraint term and the pose deviation vector is performed to generate a structured pose deviation vector field for describing the correction trend.

11. The Mandarin pronunciation pose modeling method driven by a multimodal artificial intelligence agent as described in claim 10, characterized in that, In step S404, the structured pose deviation vector field is specifically represented as follows: ; Among them, the Represents the pose deviation vector, the The term represents the discrimination constraint, the... The coupling weight coefficients representing the pose deviation vector, the This represents the coupling weight coefficient of the discrimination constraint vector.

Citation Information

Patent Citations

  • Pronunciation correction method based on vocal organ form and behavior deviation visualization

    CN108877319A

  • AI-driven personalized voice training and pronunciation correction system

    CN120015053A