Multi-modal fusion oral English fluency real-time evaluation method and system
By using a multimodal fusion assessment method, voice, visual, and interactive data are collected and processed simultaneously, which solves the problems of the single assessment method and insufficient data processing in the existing technology. It realizes efficient fusion of multimodal information and real-time feedback, and improves the accuracy of English oral fluency assessment and learning assistance effect.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HENAN POLICE ACAD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-04-17
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing methods for assessing English speaking fluency rely on a single speech modality, which cannot effectively quantify visual stuttering and interactive lag. Furthermore, multimodal fusion methods suffer from issues such as timestamp misalignment, insufficient robustness of feature extraction, and high computational resource requirements at the data processing level.
A multimodal fusion evaluation method is adopted. By simultaneously collecting speech, visual and interactive modal data, extracting key facial geometric feature points and performing spatial correction, the time synchronization of visual and speech data is achieved. Feature fusion is performed using a cross-modal attention mechanism to generate a fluency evaluation vector and output an online adaptive guidance signal.
It achieves temporal consistency of multimodal data and high-quality feature extraction, outputs structured fluency assessment results, provides instant feedback and personalized learning guidance, and improves the accuracy of assessment and the timeliness of learning assistance.
Smart Images

Figure CN121883221A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent education technology, and in particular to a multimodal fusion method and system for real-time assessment of English oral fluency. Background Technology
[0002] English speaking fluency is a key indicator for measuring language communication competence, and its assessment requires comprehensive consideration of multiple dimensions, including speech continuity, prosody, and nonverbal behavior. Current mainstream assessment methods mainly rely on acoustic analysis of a single speech modality, which cannot effectively quantify visual pauses and interaction delays caused by nervousness, thinking, etc., and thus have a limited assessment dimension.
[0003] More advanced multimodal fusion methods attempt to integrate visual and interactive information, but they still face the following shortcomings at the data processing level: First, there are millisecond to hundreds of millisecond-level timestamp misalignments between multi-source data streams (audio, visual, and interactive), stemming from device, network, and processing delays, which cause cross-modal correlation analysis to fail and may make it difficult to accurately determine causal relationships; Second, complex application environments severely affect the robustness of feature extraction, such as changes in lighting and occlusion causing facial geometric feature detection drift, and non-stationary background noise interfering with the effective separation of acoustic features (especially voiceless consonants and weak pronunciations); Third, the model's ability to generalize to individual physical characteristics and cultural differences is insufficient. For example, the small facial movement range of middle-aged and elderly people and the special regional accent rhythms may cause feature distribution shifts, making the fusion weight allocation based on attention mechanisms inaccurate; In addition, the high requirements for computing resources for real-time processing on the edge and the privacy and security risks involving facial biometric information may also restrict the large-scale deployment of the technology. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a multimodal fusion method and system for real-time assessment of English speaking fluency, which realizes a closed loop from multimodal synchronous acquisition, fusion analysis to real-time feedback.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows: Firstly, a multimodal fusion method for real-time assessment of English spoken fluency, the method comprising: Collect learners' multimodal synchronous acquisition streams, which include speech modality data, visual modality data, and interaction modality data; Key facial geometric feature points are extracted from the visual modality data, and spatial correction parameters are calculated based on these points. The visual modality data is then geometrically corrected using these spatial correction parameters and synchronized temporally with the speech modality data to obtain visual features. Acoustic features are extracted from the speech modality data to obtain speech features. Interaction features are extracted based on the temporal information in the interaction modality data. Multimodal feature data is obtained based on the visual features, speech features, and interaction features. The multimodal feature data is input into a multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector. The multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features, and interaction features to form a fusion feature embedding. Fluency analysis is then performed based on the fusion feature embedding to obtain a fluency evaluation vector that includes a fluency value and sub-dimensional evaluation values. Based on the fluency assessment vector, an online adaptive guidance signal is generated and output. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their spoken expression strategies; or trigger a personalized learning task.
[0006] Secondly, a multimodal fusion real-time English speaking fluency assessment system includes: The acquisition synchronization module is used to acquire learners' multimodal synchronous acquisition streams, which include speech modality data, visual modality data, and interaction modality data; The feature alignment module is used to extract key facial geometric feature points from the visual modality data and calculate spatial correction parameters based on the key facial geometric feature points; perform geometric correction on the visual modality data using the spatial correction parameters and synchronize it with the speech modality data in time to obtain visual features; extract acoustic features from the speech modality data to obtain speech features; extract interaction features based on the temporal information in the interaction modality data; and obtain multimodal feature data based on the visual features, speech features, and interaction features. The fusion evaluation module is used to input the multimodal feature data into the multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector. The multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features, and interaction features to form a fusion feature embedding. Fluency analysis is then performed based on the fusion feature embedding to obtain a fluency evaluation vector that includes a fluency value and sub-dimensional evaluation values. The feedback module is used to generate and output an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their oral expression strategies; or trigger a personalized learning task.
[0007] Thirdly, a computing device includes: One or more processors; A storage device for storing one or more programs that, when executed by one or more processors, cause the one or more processors to implement the method.
[0008] Fourthly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0009] The above-described solution of the present invention has at least the following beneficial effects: Simultaneous acquisition of multimodal data including speech, vision, and interaction provides comprehensive multi-source information reflecting spoken expression. This lays a complete data foundation for subsequent multimodal fusion analysis. The synchronous acquisition feature ensures the temporal consistency of multi-source data, establishing a basis for cross-modal data processing. Extracting key facial geometric feature points accurately captures core visual information. Spatial correction parameter calculation and geometric correction improve the quality of visual modality data, eliminating the impact of spatial distortion. Temporal synchronization of visual and speech data ensures temporal alignment of cross-modal data, avoiding feature misalignment. Targeted extraction of features from each modality transforms raw data into effective features. The integrated multimodal feature data provides standardized, high-quality data for subsequent fusion evaluation. The system provides input; a cross-modal attention mechanism dynamically mines the relationships between visual, speech, and interactive features, enabling adaptive weighted fusion of multimodal information; fusion feature embedding integrates core complementary information from multiple modalities, enhancing the expressive power of features; the output includes an evaluation vector containing fluency values and sub-dimensional evaluation values, achieving a structured presentation of evaluation information and comprehensively carrying fluency-related evaluation dimension information; online adaptive guidance signals generated based on the evaluation vectors can accurately match learners' fluency performance; the guidance signals combine expression strategy adjustment guidance with personalized task triggering functions, covering different learning needs; real-time output characteristics enable instant feedback during the learning process, improving the timeliness and adaptability of learning assistance. Attached Figure Description
[0010] Figure 1 This is a flowchart illustrating a multimodal fusion method for real-time assessment of spoken English fluency provided by an embodiment of the present invention.
[0011] Figure 2 This is a schematic diagram of a multimodal fusion real-time English speaking fluency assessment system provided by an embodiment of the present invention. Detailed Implementation
[0012] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0013] like Figure 1 As shown, embodiments of the present invention propose a multimodal fusion method for real-time assessment of English speaking fluency, the method comprising the following steps: Step 100: Collect the learner's multimodal synchronous acquisition stream, which includes speech modality data, visual modality data, and interaction modality data; Step 200: Extract key facial geometric feature points from the visual modality data, and calculate spatial correction parameters based on the key facial geometric feature points; perform geometric correction on the visual modality data using the spatial correction parameters, and synchronize it with the speech modality data in time to obtain visual features; extract acoustic features from the speech modality data to obtain speech features; extract interaction features based on the temporal information in the interaction modality data; and obtain multimodal feature data based on the visual features, speech features, and interaction features. Step 300: The multimodal feature data is input into the multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector; wherein, the multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features and interaction features to form a fusion feature embedding, and performs fluency analysis based on the fusion feature embedding to obtain a fluency evaluation vector containing a fluency value and sub-dimensional evaluation values; Step 400: Generate and output an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their spoken expression strategies; trigger a personalized learning task.
[0014] In this embodiment of the invention, parallel acquisition of multi-dimensional data of voice and visual interaction is achieved, ensuring the original correlation of data from different modalities and providing complete data support for the effective extraction of subsequent multi-modal features; spatial correction parameters are calculated and visual data is corrected by facial key geometric feature points, improving the effectiveness of visual modal data; temporal synchronization of visual and voice modal data is achieved, ensuring the temporal consistency of cross-modal data; acoustic features and interaction features based on temporal information are extracted separately, fully preserving the core information of each modality; multi-modal feature data is integrated to achieve orderly fusion of information from different modalities; visual and voice interaction features are dynamically correlated and weighted through a cross-modal attention mechanism to achieve adaptive integration of features from each modality; fused feature embedding is formed and fluency analysis is performed to fully leverage the synergistic effect of multi-modal information; an evaluation vector containing fluency value and sub-dimensional evaluation values is output to achieve structured output of multi-dimensional evaluation information; online adaptive guidance signals are generated based on the fluency evaluation vector to achieve instant connection between evaluation results and learning guidance; learners are guided to adjust their oral expression strategies to help optimize the learning process in real time; personalized learning tasks are triggered to accurately match learning guidance with multi-modal evaluation results.
[0015] In a preferred embodiment of the present invention, step 100 above involves collecting the learner's multimodal synchronous acquisition stream, which includes speech modality data, visual modality data, and interaction modality data, including: Step 101 involves acquiring the learner's original audio stream using a terminal microphone array and performing high-frequency pre-enhancement and overlapping window frame division on the original audio stream to obtain speech modal data. Specifically, this includes: First, acquiring the original audio stream using the microphone array deployed on the terminal, utilizing the spatial pickup characteristics of the microphone array to capture the audio signal in the direction of the learner's voice, while suppressing noise signals from irrelevant directions in the surrounding environment; Next, performing high-frequency pre-enhancement processing on the acquired original audio stream, by selectively adjusting the gain of high-frequency bands related to voiceless consonants and weak readings in the audio signal to enhance this type of audio information that is easily masked by noise; Based on this, overlapping window frame division is performed on the high-frequency pre-enhancement audio signal. Specifically, a window function of a preset length can be selected to segment the audio stream segment by segment, and a preset proportion of overlapping area is set between adjacent window frames to avoid loss of audio information during the window frame division process. Through the above processing, speech modal data suitable for subsequent acoustic feature extraction is finally obtained.
[0016] Step 102: Acquire the original video stream containing the learner's face and upper body area through the terminal camera, and perform video stream decoding and contrast-limited adaptive histogram equalization on the original video stream to obtain visual modality data. Specifically, this includes: First, controlling the terminal camera to continuously acquire the original video stream of the learner's face and upper body area. This area is selected because it contains key facial geometric feature points and upper body limb movement information, which can meet the needs of subsequent visual feature extraction. Further, perform video stream decoding processing on the acquired original video stream, specifically decompressing the encoded and compressed original video stream into an uncompressed video frame sequence, completing the conversion of the original video data from storage format to a processable format. Subsequently, perform facial region geometric cropping algorithm processing on the decoded video frame sequence. Specifically, first, detect key facial feature points in each video frame to determine the pixel coordinates of six core feature points: the inner corners of both eyes, the outer corners of both eyes, the tip of the nose, and the left and right corners of the mouth, and then calculate these... The bounding rectangle of the polygon formed by six feature points is used to determine the cropping range, with the upper left pixel coordinates of the bounding rectangle as the starting point and the lower right pixel coordinates as the ending point. The cropping range expands outward by a preset pixel value in the horizontal direction to cover the facial edge area, expands upward by a preset pixel value in the vertical direction to cover the forehead area, and expands downward by a preset pixel value to cover the jaw area. Based on the determined cropping range, pixel regions are cropped for each video frame to remove redundant background information outside the facial area. Then, contrast-limited adaptive histogram equalization is performed on the cropped video frame sequence. Specifically, the video frame is divided into multiple sub-regions, and histogram equalization is performed on each sub-region to improve local contrast. At the same time, the slope of the cumulative distribution function of the histogram is limited to avoid noise amplification caused by over-enhancement. Through the above processing, the detail clarity of the facial area in the video frame under different lighting conditions can be effectively improved, and the impact of lighting changes and background redundancy on subsequent facial geometric feature detection can be reduced, ultimately obtaining high-quality visual modality data.
[0017] Step 103: Capture learners' response behaviors in dialogue practice in real time through platform interaction logs, and extract discrete time-stamped sequences including question-and-answer turn switching times and topic continuity markers to obtain interaction modality data. Specifically, this includes: First, based on the interaction log system of the online oral practice platform, monitor and capture various response behaviors of learners in the dialogue practice process in real time, including but not limited to learners initiating voice input, completing voice input, and providing text responses, and synchronously record the original timestamps corresponding to each behavior. These original timestamps use the same system benchmark as the acquisition timestamps of the voice modality data and visual modality data. Then, introduce geometric interpolation of multimodal timestamps. The calibration algorithm performs time synchronization calibration. The specific process is as follows: First, standard timestamps of key time nodes of the voice modality and key time nodes of the visual modality collected at the same time are extracted. Among them, key time nodes of the voice modality are such as the start and stop times of the voice signal, and key time nodes of the visual modality are such as the trigger times of facial movements. Using these standard timestamps as reference nodes, a time calibration coordinate system is constructed. The deviation value between the original timestamp of the interaction behavior and the timestamp of the reference node is calculated. Then, a linear time mapping relationship is established between two adjacent reference nodes through geometric interpolation. Substituting the deviation value, the calibrated interaction timestamp is calculated to ensure the temporal consistency between the interaction modality timestamp and the voice and visual modality timestamps.
[0018] Furthermore, a temporal analysis is performed on the calibrated response behavior to extract key time nodes and feature information closely related to the interaction. Specifically, the start and stop times of the voice signals of both parties in the dialogue are detected to determine the question-and-answer turn switching time, and the topic continuity marker is determined by analyzing the correlation between the learner's response content and the current dialogue topic. The extracted information is then organized into a discrete time marker sequence in chronological order. Through the above processing, the core temporal information in the interaction process can be filtered out, and irrelevant and redundant interaction data can be removed. Finally, interaction modal data that can reflect the core features of dialogue interaction is obtained, providing effective support for subsequent time synchronization and correlation analysis of multimodal data.
[0019] Step 104 involves timestamping the speech modal data, visual modal data, and interaction modal data to form a multimodal synchronous acquisition stream. Specifically, this includes adding corresponding timestamp information to each window frame of the speech modal data, each video frame of the visual modal data, and each discrete time marker of the interaction modal data, using the system clock of the terminal acquisition device as a unified time reference. The timestamp information is accurate to the millisecond level. Further, based on the added timestamp information, preliminary temporal alignment of the three types of modal data is performed. By adjusting the time axis of each modal data, speech, visual, and interaction data corresponding to the same timestamp can be accurately matched. This integrates the speech modal data, visual modal data, and interaction modal data into a temporally consistent multimodal synchronous acquisition stream. This setup effectively compensates for the timetamp misalignment defect in multi-source data streams, ensuring the effectiveness of subsequent cross-modal correlation analysis and laying a reliable data foundation for multimodal fusion evaluation.
[0020] In this embodiment of the invention, a microphone array is used to acquire the original audio stream, ensuring the directionality and integrity of the audio acquisition; high-frequency pre-enhancement improves the identifiability of high-frequency information in the audio; overlapping window frame segmentation ensures the continuity and integrity of the audio data, providing a high-quality speech modality data foundation for subsequent acoustic feature extraction; a camera is used to acquire the original video stream containing the learner's face and upper body, ensuring full coverage of key visual areas; video stream decoding enables the processable transformation of the original video data; contrast-limited adaptive histogram equalization improves the detail clarity of video frames under different lighting conditions, enhancing the face and upper body. The extractability of features forms high-quality visual modal data; response behavior is captured in real time through platform interaction logs to ensure the timeliness and authenticity of interaction data; discrete time marker sequences of question-and-answer turn switching times and topic continuity markers are extracted to filter key temporal information in the interaction process, clearly retain the core features of dialogue interaction, and generate effective interaction modal data; timestamp annotation: the temporal alignment of voice modal data, visual modal data, and interaction modal data is achieved to ensure the spatiotemporal correlation of multimodal data, and a well-structured multimodal synchronous acquisition stream is constructed to provide a reliable data foundation for subsequent cross-modal data fusion processing.
[0021] In a preferred embodiment of the present invention, step 200 involves extracting key facial geometric feature points from the visual modality data and calculating spatial correction parameters based on the key facial geometric feature points; performing geometric correction on the visual modality data using the spatial correction parameters and synchronizing it temporally with the speech modality data to obtain visual features; extracting acoustic features from the speech modality data to obtain speech features; extracting interaction features based on the temporal information in the interaction modality data; and obtaining multimodal feature data based on the visual features, speech features, and interaction features, including: Step 201 involves extracting facial region images from the visual modality data and performing image preprocessing to obtain standardized image data. Specifically, this includes: first, extracting regions containing the learner's face from the visual modality data to remove redundant background information unrelated to facial features and reduce interference from irrelevant regions on subsequent feature detection; further, performing image preprocessing operations on the extracted facial region images, specifically including converting color images to grayscale images to reduce data processing dimensionality, performing Gaussian smoothing on the grayscale images to suppress image noise, and scaling the images to a preset size to achieve image specification uniformity. Through the above series of preprocessing operations, standardized image data is finally obtained, providing a consistent and high-quality image foundation for subsequent detection of key facial geometric feature points.
[0022] Step 202: Based on the standardized image data, candidate facial landmarks are located and selected through multi-scale corner detection and feature descriptor matching. Specifically, this includes: First, a multi-scale corner detection method is used to traverse and detect the standardized image data. By setting different scale parameters, different sizes of local facial regions in the image are adapted to ensure that corner features at different scales of the face can be fully captured, thereby improving the range adaptability of facial landmark detection. Further, feature descriptors are extracted from the detected corner features. By calculating the feature descriptors of each corner and matching them with a preset facial landmark feature template, candidate points that conform to the feature rules of facial landmarks are selected based on the matching similarity. On this basis, abnormal points with matching similarity below a preset threshold are removed, and finally, candidate facial landmarks are obtained. The combination of multi-scale detection and feature matching effectively improves the accuracy and reliability of candidate facial landmark location.
[0023] Step 203 involves applying geometric constraints to the candidate facial landmarks and performing iterative optimization to obtain a set of key facial geometric feature points. Specifically, this includes: First, based on the inherent characteristics of human facial structure, applying geometric constraints to the candidate facial landmarks, including constraints on the distance between eye landmarks, the relative positional relationship between eye and nose landmarks, and the spatial distribution of mouth and nose landmarks. These geometric constraints are used to eliminate abnormal candidate points that do not conform to human facial structure characteristics. Further, for the candidate facial landmarks filtered by geometric constraints, an iterative nearest-point optimization algorithm is introduced for iterative optimization. Specifically, the filtered set of candidate facial landmarks is first used as the set to be optimized, while a preset facial structure feature template point set is called. This template point set is based on key feature points from a large number of normal facial samples, statistically obtained, and conforms to the standard facial structure feature ratio. Based on this, the nearest landmark point in the template point set is calculated for each landmark point in the set to be optimized, establishing an initial point-to-point correspondence. Subsequently, based on this initial point-to-point correspondence... Based on the given relationship, the rigid transformation parameters that minimize the average distance between the point set to be optimized and the template point set are calculated. These parameters include rotation angle and translation amount. The calculated rigid transformation parameters are applied to all marker points in the point set to be optimized to obtain an updated candidate facial marker point set. Then, the point pair correspondence between the updated point set and the template point set is recalculated, and the rigid transformation parameters are solved again to update the point set. The process of establishing point pair correspondence, solving rigid transformation parameters, and updating the point set is repeated. After each iteration, the overall average distance between the point set to be optimized and the template point set is calculated as the geometric error. The iteration process continues until the overall geometric error is less than a preset threshold or the number of iterations reaches a preset upper limit. Finally, a set of key facial geometric feature points is obtained through the synergistic optimization of geometric constraints and the iterative nearest point optimization algorithm. Through the synergistic effect of dual optimization, the localization accuracy of key facial geometric feature points is further improved, while the fit between the feature point set and the real facial structure features is strengthened, ensuring the reliability of subsequent spatial correction parameter calculations.
[0024] Step 204: Pair each feature point in the set of key facial geometric feature points with feature points in the predefined standard frontal face model that have the same semantic annotation to establish a semantic alignment relationship for feature points. Specifically, this includes: First, pre-constructing a standard frontal face model. The pre-definition, construction, training, and implementation process of this model is as follows: First, determine that the core purpose of the model is to provide a unified semantic benchmark and geometric reference for facial feature points. Therefore, the predefined model needs to cover general facial structural features of different age groups and skin colors, and simultaneously determine the semantic categories and number of feature points in the model. Then, conduct basic data collection for model construction, collecting a large number of frontal face image samples of healthy adults of different genders, ages, and skin colors to ensure the samples have broad representativeness to improve the model's generalization ability. Preprocess the collected frontal face image samples, including grayscale conversion, Gaussian smoothing noise reduction, and size standardization, to eliminate the impact of image quality differences on model construction. Next, perform manual semantic annotation, where professional annotators annotate the inner corners of both eyes and both eyes according to unified annotation specifications on each preprocessed frontal face image. Feature points of key areas such as the outer corners of the eyes, the tip of the nose, and the left and right corners of the mouth are assigned unique semantic labels. Model training is conducted based on labeled sample data, using the coordinates of the labeled feature points as the training target. Statistical learning methods are used to perform cluster analysis and mean calculation on the feature point coordinates of all samples to obtain the average coordinates of feature points for each semantic category. Furthermore, the relative distances, angles, and other geometric relationships between feature points are calculated to form geometric constraint rules that conform to the general facial structure characteristics. The average coordinates and geometric constraint rules are integrated into an initial standard frontal facial model. The initial model is then validated using a validation set. Frontal facial image samples that were not used in training are selected, and the standard feature points output by the model are compared with the manually labeled feature points of the samples to calculate the coordinate error. If the error exceeds a preset threshold, the average coordinates of the model's feature points and the geometric constraint parameters are adjusted. The training and validation process is repeated until the model error is less than the preset threshold, completing the training and construction of the standard frontal facial model. This model contains multiple semantically labeled facial feature points, covering key areas such as the inner corners of the eyes, the outer corners of the eyes, the tip of the nose, and the left and right corners of the mouth.
[0025] Furthermore, each feature point in the set of key geometric feature points of the face is traversed, and each feature point is paired with the feature points with the same semantic annotation in the standard frontal face model according to the positional and functional attributes of each feature point on the face. This pairing method establishes the semantic alignment relationship of the feature points. This relationship can determine the correspondence between the facial feature points and the standard feature points under different poses, avoid semantic confusion of feature points, and provide a feature correspondence basis for the accurate calculation of subsequent spatial correction parameters.
[0026] Step 205: Based on all feature point pairs associated with the semantic alignment relationship of the feature points, a similarity transformation matrix is solved using the least squares fitting method to minimize the coordinate error between the set of key facial geometric feature points and the standard frontal facial model. Specifically, this includes: First, based on the semantic alignment relationship of the feature points, all paired feature point pairs are extracted, and each feature point pair contains two sets of coordinate information, namely the image pixel coordinates of the actual key facial geometric feature points and the standard coordinates of the corresponding semantically labeled feature points in the standard frontal facial model. Further, the coordinate data of all feature point pairs are preprocessed. First, the centroid coordinates of the actual set of key facial geometric feature points and the centroid coordinates of the feature point set of the standard frontal facial model are calculated respectively. Then, the actual feature point coordinates and the standard feature point coordinates in each feature point pair are subtracted from their respective centroid coordinates to complete the coordinate centering process, thereby eliminating the interference of the difference in the position of the coordinate origin on the subsequent fitting calculation.
[0027] Subsequently, the least squares fitting method was used to solve for the specific transformation matrix parameters. First, based on the coordinates of the centered feature points, the covariance matrix between the actual feature points and the standard feature points was calculated. By performing statistical analysis on the covariance matrix, relevant statistics for determining the rotation parameters were obtained. Then, based on these statistics, the rotation angle parameters that align the actual feature points with the standard feature points were determined. On this basis, the mean modulus of all centered actual feature points and the mean modulus of the standard feature points were calculated. The ratio of the two was used as the scaling parameter for the similarity transformation to ensure that the actual facial feature points and the standard model feature points are matched in scale.
[0028] Next, combining the centroid coordinates of the actual feature points, the centroid coordinates of the standard feature points, and the determined rotation and scaling parameters, the translation parameters are calculated through inverse coordinate transformation. These translation parameters can compensate for the spatial offset between the actual face and the standard model. Finally, the obtained rotation, scaling, and translation parameters are integrated to construct a similarity transformation matrix, minimizing the overall coordinate error between the set of key geometric feature points of the actual face and the set of feature points of the standard frontal face model after transformation by this matrix. The entire process solves the components of the transformation matrix step by step, fully utilizing the coordinate information of all feature point pairs for statistical fitting. This ensures that the obtained similarity transformation matrix can accurately represent the pose difference between the actual face and the standard model, guaranteeing the accuracy of the transformation matrix and providing reliable parameter support for subsequent spatial correction.
[0029] Step 206 involves resolving spatial correction parameters representing rotation, translation, and scaling transformations from the similarity transformation matrix. Specifically, this includes: First, performing a structured analysis operation on the similarity transformation matrix to determine that it consists of a rotation-scaling composite component and a translation component. The rotation-scaling composite component represents the rotation and scale changes of the facial posture, while the translation component represents the offset of the facial position. Further, the rotation-scaling composite component is extracted from the similarity transformation matrix, and its parameters are decomposed to separate the rotation and scaling parameters. First, by calculating the proportional relationship of relevant elements in the rotation-scaling composite component, a rotation parameter representing the degree of facial deflection around the image plane coordinate axis is derived. This parameter accurately reflects the deflection angle of the actual face relative to the standard frontal facial model. Then, based on the element modulus of the rotation-scaling composite component, a scaling parameter representing the proportional relationship between the actual facial region and the standard model size is obtained. This parameter can be directly used to adjust the overall size of the facial region to maintain consistency with the scale of the standard model.
[0030] Subsequently, the translation component in the similarity transformation matrix is extracted. This component directly corresponds to the positional offset information of the facial region in the image plane. Through numerical analysis of the translation component, horizontal translation parameters representing the horizontal positional deviation of the face and vertical translation parameters representing the vertical positional deviation of the face are obtained. These two types of translation parameters together constitute a complete set of translation parameters used to calibrate the spatial position of the facial region in the image. Finally, the obtained rotation, scaling, and translation parameters are integrated to form a complete set of spatial correction parameters. The entire analysis process, through the structured decomposition of the components of the similarity transformation matrix, transforms the abstract matrix operation results into intuitive and operable specific parameters. Each parameter corresponds to a clear geometric transformation meaning, providing a targeted adjustment basis for the subsequent geometric correction of visual modality data, ensuring that facial pose correction work can accurately match the actual facial pose differences.
[0031] Step 207: Construct an affine transformation matrix based on the spatial correction parameters, and use the affine transformation matrix to perform geometric transformation on each frame of the visual modality data, mapping the facial region in the image to a predefined normalized coordinate system to obtain a pose-normalized facial region image sequence. Specifically, this includes: First, constructing an affine transformation matrix based on the spatial correction parameters. Specifically, first, determining the geometric transformation meaning of the rotation, translation, and scaling parameters included in the spatial correction parameters. The rotation parameter represents the deflection angle of the face around the image plane coordinate axis; the scaling parameter represents the size ratio between the facial region and a standard frontal facial model; and the translation parameter represents the position of the facial region within the image plane. Position offset; furthermore, the three types of parameters are integrated according to the composite logic of geometric transformation. First, a scaling transformation component is constructed based on the scaling parameter to adjust the overall size of the facial region to match the scale of the standard model. Then, a rotation transformation component is constructed by combining the rotation parameter to correct the deflection posture of the face to align with the orientation of the standard model. Finally, a translation transformation component is constructed by incorporating the translation parameter to calibrate the spatial position of the facial region to fit the coordinates of the standard model. On this basis, the scaling, rotation, and translation transformation components are integrated into a complete affine transformation matrix according to the operation rules of geometric transformation. This ensures that the matrix can perform comprehensive geometric transformation on the image pixel coordinates and accurately match the posture difference between the actual face and the standard model.
[0032] Furthermore, the constructed affine transformation matrix is applied to each frame of the visual modality data. The specific steps are as follows: First, the construction process of a predefined standardized coordinate system is determined. This coordinate system is pre-defined based on a standard frontal facial model, with the core purpose of providing a unified spatial reference for facial regions in different poses. The origin of the coordinate system is first set at the centroid of all feature points in the standard frontal facial model, ensuring that the center of the coordinate system coincides with the geometric center of the standard face. Then, the horizontal direction is set as the x-axis, and the vertical direction as the y-axis, with the x-axis and y-axis perpendicular to each other and conforming to the conventional coordinate logic of image pixels. Subsequently, based on the feature point distribution range of the standard frontal facial model, a fixed pixel size range is determined to ensure that the... All geometrically transformed facial regions can be fully adapted to this coordinate range, and the mapped facial regions of different learners maintain consistent size. This standardized coordinate system is pre-stored in the system for use during geometric transformations. Based on this, the pixel boundary range of the facial region in each frame is determined according to key facial geometric feature points, avoiding invalid processing of non-facial regions in the image to improve computational efficiency. Subsequently, for each pixel coordinate in the facial region, the scaling, rotation, and translation transformation rules in the affine transformation matrix are applied sequentially to map the original pixel coordinates to the predefined standardized coordinate system, so that facial regions of different poses can be uniformly aligned to standard spatial positions and size specifications.
[0033] Through the above frame-by-frame and pixel-by-pixel geometric transformation processing, a facial region image sequence with pose normalization is obtained. This sequence completely eliminates the differences in facial poses between different learners, as well as the feature differences caused by the pose changes of the same learner at different times, improving the consistency and comparability of visual features and providing a standardized visual data foundation for subsequent multimodal feature fusion.
[0034] Step 208 involves matching and aligning the timestamps of the pose-normalized facial region image sequence with the timestamps of the speech modality data to obtain spatiotemporally aligned visual features. Specifically, this includes: first, extracting the timestamp information of each frame in the pose-normalized facial region image sequence, where the timestamp is a real-time time record of the terminal system clock when the image frame acquisition is completed; simultaneously extracting the timestamp information of each audio frame in the speech modality data obtained in step 101, where the timestamp is a real-time time record of the terminal system clock when the audio frame is divided into window frames; further, standardizing the format of the two types of timestamp information to unify the time unit and data format of the timestamps, ensuring the consistency of subsequent time difference calculations.
[0035] Based on this, using the system clock of the terminal acquisition device as a unified time reference, a timestamp matching and alignment operation is performed: First, a time error tolerance threshold is preset. This threshold is set comprehensively based on factors such as the hardware latency of terminal data acquisition, network transmission fluctuations, and data processing time to ensure coverage of normal time deviations that may occur in actual application scenarios. Then, all image frames of the pose-normalized facial region image sequence are traversed. For the timestamp of each image frame, a matching item is searched in the audio frame timestamp set of the speech modality data. The absolute value of the difference between the image frame timestamp and each audio frame timestamp is calculated. Audio frames with an absolute difference value less than or equal to the preset error tolerance threshold are determined as matching candidate frames. If there are multiple candidate frames, the audio frame with the smallest absolute difference value is selected as the final matching frame. If there is no candidate frame that meets the error threshold requirements, the image frame is marked as an unmatched frame, and subsequent processing can be performed through interpolation or time axis stretching.
[0036] After matching and pairing all image frames and audio frames, each successfully matched image frame and audio frame is associated and bound together, retaining their respective feature information and timestamp correspondence. Through the above timestamp matching and alignment process, spatiotemporally aligned visual features are obtained. These features establish a precise temporal correlation with the speech modal data, ensuring that visual and speech information in the same time dimension can accurately correspond, providing a temporally consistent data foundation for the effective fusion of subsequent multimodal features, and ensuring the effectiveness of cross-modal association analysis.
[0037] Step 209 involves performing a short-time Fourier transform on the speech modal data to calculate the Mel frequency cepstral coefficients, fundamental frequency profile, and short-time energy of each audio frame, thereby obtaining speech features characterizing articulation, prosody, and rhythm. Specifically, this includes: first, performing short-time Fourier transform preprocessing on the speech modal data; second, dividing the speech modal data into frames according to a preset frame length and frame shift; third, adding a Hanning window to each frame of audio signal to suppress spectral leakage; and fourth, performing a Fourier transform on each windowed frame signal to convert the audio signal in the time domain into a frequency domain signal, clearly presenting the amplitude distribution characteristics of each frequency component, thus laying the foundation for subsequent acoustic parameter extraction.
[0038] Furthermore, based on the transformed frequency domain signal, Mel frequency cepstral coefficients are calculated. The specific process is as follows: First, a preset number of Mel filter banks are constructed. The frequency response of these filter banks conforms to the characteristics of human hearing and covers the effective frequency range of the speech signal. The frequency domain signal is filtered through the Mel filter banks, and the logarithmic energy of each filter output is calculated. Discrete cosine transform is performed on the logarithmic energy sequence of all filters, and the first preset order of the transform result is extracted as the Mel frequency cepstral coefficient. This coefficient can effectively characterize the spectral envelope features of speech and reflect the detailed information of pronunciation. At the same time, the fundamental frequency profile of each audio frame is calculated. Specifically, the framed speech signal is preprocessed by removing low-frequency noise through high-pass filtering, and then half-wave rectification and low-pass filtering are performed to extract the harmonic envelope of the signal. Based on the harmonic envelope, candidate fundamental frequencies are searched within a preset frequency range, and the optimal fundamental frequency is selected by calculating the harmonic matching degree corresponding to the candidate values. The optimal fundamental frequency of each frame is smoothed to eliminate isolated outliers and form a continuous fundamental frequency profile, thereby characterizing the prosodic variation features of speech.
[0039] In addition, the short-time energy of each audio frame is calculated as follows: for each frame of speech signal after framing, the sum of squares of all sampling points within the frame is calculated, and then the sum of squares is normalized to obtain the short-time energy value of that frame; the temporal variation of the short-time energy is used to characterize the intensity fluctuations and rhythmic features of the speech; finally, the Mel frequency cepstral coefficients, fundamental frequency profile, and short-time energy are aligned and integrated frame by frame to form speech features that can comprehensively characterize pronunciation, prosody, and rhythm, effectively strengthening the core information of the speech modality and reducing the interference of non-stationary background noise.
[0040] Step 210: For the discrete-time marker sequence in the interactive modal data, statistically analyze the question-and-answer response delay distribution, topic keyword repetition rate, and turn-taking frequency within a preset time window to obtain interactive features characterizing dialogue coherence and interactivity. Specifically, this includes: First, based on the regular rhythm and response duration patterns of spoken dialogue practice, after setting the preset time window, a sliding time window is set. The specific preset method for the time window is as follows: Considering the average duration of typical question-and-answer turns in spoken dialogue, the reasonable time consumption of normal learner responses, and the timeliness requirements for feature extraction, the basic preset duration is first set to a range of 3 to 5 seconds, while the window sliding step size is set to 1 to 2 seconds to ensure reasonable overlap between adjacent windows to avoid missing feature information; then... The system makes targeted adjustments based on the different dialogue scenarios. If the dialogue scenario is a fast-paced casual conversation, the preset duration is reduced to 3 seconds to improve the timeliness of features. If the dialogue scenario is a scenario that requires in-depth thinking, such as academic discussions or business negotiations, the preset duration is increased to 5 seconds to fully cover the entire question-and-answer response process. In addition, the preset duration can be dynamically adjusted according to the actual dialogue rhythm. When it is detected that the duration of multiple consecutive question-and-answer rounds exceeds the current window duration, the window duration is automatically extended by 1 second. When it is detected that the duration of multiple consecutive question-and-answer rounds is less than half of the current window duration, the window duration is automatically shortened by 0.5 seconds, ensuring that the window can always focus on the key time segments in the dialogue practice, taking into account both the timeliness and completeness of features.
[0041] Furthermore, based on the discrete time marker sequence in the interaction modal data obtained in step 103, multi-dimensional statistical analysis is conducted within the aforementioned preset sliding time window: First, the distribution of question-and-answer response delays is statistically analyzed. The time difference between the question marker time and the learner's response marker time in each question-and-answer round within the window is extracted, and the distribution intervals of these differences and the proportion of each interval are statistically analyzed to reflect the learner's response efficiency. Second, the repetition rate of topic keywords is statistically analyzed. First, the preset keyword set of the current dialogue topic is extracted, and the learner's response content markers within the window are traversed. The number of occurrences and repetition frequencies of keywords are statistically analyzed, and the ratio of repetition frequency to total occurrence frequency is calculated as the topic keyword repetition rate to characterize the learner's mastery of the topic and the coherence of expression. Third, the turn-taking frequency is statistically analyzed. The total number of turn-taking between the two parties within the window is statistically analyzed, and the total number is divided by the window duration to obtain the turn-taking frequency per unit time to reflect the interactive rhythm of the dialogue. Through the above multi-dimensional statistical analysis, the statistical results are standardized and integrated to finally obtain interactive features that can characterize the coherence and interactivity of the dialogue.
[0042] Step 211 involves temporally synchronizing the visual features, speech features, and interaction features, and fusing them through channel splicing to form multimodal feature data with a unified dimension. Specifically, this includes: First, performing temporal synchronization processing on the visual features, speech features, and interaction features. First, a unified temporal granularity is determined. Using the frame time interval of the speech features as a benchmark, the visual features are resampled according to the frame time interval. For the interaction features, the missing feature values are filled in by linear interpolation to ensure that the sampling times of the three types of features are completely matched in the temporal dimension. Then, the temporal alignment accuracy of each modality feature is verified. The timestamp difference of different modal features at the same time after alignment is calculated. If the difference exceeds a preset threshold, the resampling and interpolation parameters are readjusted until the temporal consistency of all features meets the requirements, further enhancing the temporal consistency of the multimodal features.
[0043] Furthermore, a channel-based concatenation method is used to fuse the three types of features after time synchronization: first, the dimensional information of visual features, speech features, and interaction features is confirmed to ensure that the number of samples for each modality feature is completely consistent with the time step; then, the three types of features are concatenated in parallel according to the channel dimension, that is, while keeping the time step and the number of samples unchanged, the feature dimensions of each modality are superimposed and integrated; the concatenated feature data is subjected to dimensional normalization processing, and the feature values are mapped to a preset range through normalization operations to eliminate the influence of differences in the numerical range of different modal features, forming multimodal feature data with a unified dimension.
[0044] Through the above processing, the resulting multimodal feature data can fully carry the core information of speech, vision and interaction, and the unified dimensions facilitate the processing of subsequent multimodal fusion evaluation models. It fully leverages the synergistic advantages of multimodal data and provides comprehensive feature support for accurate English speaking fluency assessment.
[0045] In this embodiment of the invention, extracting facial region images removes redundant background information, and image preprocessing achieves image standardization and noise suppression, resulting in standardized image data that provides a consistent and high-quality image foundation for subsequent facial feature point detection. Multi-scale corner detection adapts to facial regions of different sizes, improving the range adaptability of feature point detection. Feature descriptor matching enhances the localization accuracy of candidate facial landmarks, and a reliable candidate point set is obtained after screening, providing a high-quality prerequisite for determining key feature points. Geometric constraints eliminate abnormal candidate points that do not conform to facial structural features, and iterative optimization gradually corrects feature point localization deviations, resulting in a set of key facial geometric feature points, ensuring the reliability of subsequent spatial correction. Establishing semantic alignment relationships for feature points enables semantic matching between facial feature points of different poses and standard models, avoiding semantic confusion of feature points and providing a feature correspondence basis for the accurate calculation of subsequent spatial correction parameters. The least squares fitting method fully utilizes the information of all feature point pairs, minimizing coordinate errors in the solved similarity transformation matrix and ensuring the accuracy of the transformation matrix. The analytically obtained rotation, translation, and scaling transformation parameters can accurately quantify the facial features. Posture deviation provides an operational basis for the geometric correction of subsequent visual modality data; affine transformation realizes the mapping of facial regions to a standardized coordinate system, completes posture normalization processing, eliminates feature differences caused by different facial postures, and obtains a standardized facial region image sequence, improving the consistency of visual features; timestamp matching and alignment ensures the temporal synchronization of visual and speech features, establishes the spatiotemporal correlation of cross-modal data, and provides a temporally consistent data foundation for subsequent multimodal feature fusion; short-time Fourier transform can effectively extract frequency domain information of audio frames, and multi-dimensional acoustic parameters comprehensively characterize pronunciation details, prosodic changes and rhythmic features, fully preserving the core information of the speech modality and forming effective speech features; preset time window statistics focus on interactive information within key time ranges, and multi-dimensional statistical indicators comprehensively characterize dialogue response efficiency, topic coherence and interaction rhythm, accurately extracting the core features of the interaction modality; temporal synchronization further strengthens the temporal consistency of multimodal features, and channel splicing realizes the orderly fusion of features, forming multimodal feature data of a unified dimension, which is convenient for subsequent model processing and fully leverages the synergistic effect of multimodal data.
[0046] In a preferred embodiment of the present invention, step 300 involves inputting the multimodal feature data into a multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector. The multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically correlate and weightedly fuse the visual features, speech features, and interaction features to form a fused feature embedding. Fluency analysis is then performed based on the fused feature embedding to obtain a fluency evaluation vector containing a fluency value and sub-dimensional evaluation values, including: Step 301 involves inputting the corresponding visual features, speech features, and interaction features from the multimodal feature data into independent trainable linear projection layers, projecting the feature vectors of each modality onto a unified feature subspace to obtain a set of query vectors, key vectors, and value vectors corresponding to each modality. Specifically, this includes: first, determining that visual features, speech features, and interaction features in the multimodal feature data each have different dimensional specifications and data distribution characteristics; direct fusion could easily lead to modal information imbalance. Therefore, independent trainable linear projection layers are configured for each of the three types of features. These linear projection layers are specifically implemented for each modality. Each modality has a corresponding independent fully connected neural network layer, and the projection layer structure is exactly the same. Its weight parameters have trainable properties and are initialized using Xavier initialization. The core function of this projection layer is to uniformly map the original feature vectors of each modality with different dimensions to a preset common feature subspace. The dimension D of this common feature subspace is a configurable hyperparameter, which can be set to 256 or 512 dimensions according to the model training requirements. In addition to following the Xavier initialization rules, the initial values of the parameters of each projection layer are also fine-tuned in combination with the statistical distribution characteristics of the corresponding modality features to ensure that the initial projection direction fits the essence of the features.
[0047] Furthermore, visual features, speech features, and interaction features are input into their respective trainable linear projection layers. Through linear transformation operations within the projection layers, the feature vectors of each modality, which originally had different dimensions, are uniformly mapped to the aforementioned common feature subspace of dimension D. During this process, each projection layer can dynamically adjust its internal weight parameters through subsequent model training, making the mapped features more suitable for the needs of cross-modal association mining. At the same time, the feature vector output by each modality through the projection layer is subjected to linear transformation through three different weight matrices within the projection layer, splitting it into three sets of vectors of dimension D, which serve as the query vector, key vector, and value vector corresponding to that modality, respectively. These three sets of vectors together constitute the vector set of the corresponding modality. Through the above processing, not only is the spatial unification of features from different modalities achieved, but a standardized input suitable for subsequent attention calculations is also generated, laying the foundation for effective interaction of cross-modal features.
[0048] Step 302: Based on the set of query vectors, key vectors, and value vectors, calculate the cross-attention weight matrix between any two modal features through a cross-modal multi-head attention mechanism. Specifically, this includes: First, building a multi-head attention structure based on the set of query vectors, key vectors, and value vectors for each modality. The number of attention heads H in this structure is a configurable hyperparameter, for example, H=8. Each attention head has an independent feature processing channel, and the weight parameters of each attention head are independent of each other and can be trained to ensure that different attention heads can capture the correlation features between different dimensions of modality.
[0049] Furthermore, for each attention head, cross-modal association calculations are performed: taking any two modalities as a group, the query vector of the first modality and the key vector of the second modality are multiplied by a dot product. The result is then divided by a scaling factor, which is the square root of the key vector dimension, to avoid the problem of excessively large values caused by high vector dimensions. The scaled result is then normalized using the Softmax function to obtain an initial weight matrix representing the association strength between the two modal features. The above calculation process is performed in parallel on all eight attention heads, with each attention head outputting a set of initial weight matrices of different dimensions. After all attention heads have completed their calculations, the initial weight matrices output by all attention heads are concatenated and integrated. The concatenated matrix is then input into a linear projection layer for dimension normalization, ultimately obtaining the cross-attention weight matrix between any two modal features. Through multi-head parallel calculation and integration, the multi-dimensional and multi-level association relationships between different modalities can be comprehensively captured, avoiding the omission of association information by a single attention head, and providing a basis for the association strength for subsequent dynamic weighted fusion.
[0050] Step 303: Based on the cross-attention weight matrix, the value vectors from different modalities are weighted and summed to obtain a dynamically weighted modal fusion vector. Specifically, this includes: First, extracting all cross-attention weight matrices and determining the positions and correlation strengths of the two modal features corresponding to each element in the matrix, ensuring precise matching between the weights and the value vector positions; Further, executing the fusion process according to the logic of a single modality as the core and other modalities as supplements, specifically, for each modal value vector, weighting and integrating the value vectors of other modalities based on its cross-attention weight matrix with the other two modalities; taking the visual modality as an example, first calling the cross-attention weight matrices corresponding to speech and interaction, and using these two matrices as weights, performing element-wise multiplication and weighting on the corresponding speech value vector and interaction value vector to obtain two weighted single-modal value vectors; then... The two weighted vectors are superimposed and integrated with the value vector of the visual modality itself to obtain a fusion vector with vision as the core that dynamically reflects the influence of speech and interaction modal information. Similarly, with speech and interaction modalities as the core, the corresponding cross-attention weight matrices are called to weight the value vectors of the other two modalities, and then superimposed with their own value vectors to obtain fusion vectors with speech and interaction as the core. The above process is repeated to complete the cross-weighted integration between all modalities, resulting in three sets of fusion vectors with different modalities as the core. Finally, a linear fusion operation is used to integrate the three sets of vectors into a unified dynamic weighted modal fusion vector. Through this dynamic weighting method based on the correlation strength, the fusion process can respond to the correlation differences of different modal features, while adapting to the feature distribution shifts caused by individual physical characteristics and cultural differences, ensuring that closely related modal information is fully highlighted.
[0051] Step 304: The dynamically weighted modal fusion vector is input into a multilayer perceptron. Through nonlinear transformation and feature dimensionality reduction, a unified fusion feature embedding is formed. Specifically, this includes: First, the dynamically weighted modal fusion vector is input into a preset multilayer perceptron, which consists of an input layer, at least two fully connected hidden layers, and an output layer. ReLU or GELU nonlinear activation functions are used between layers to achieve nonlinear transformation of features. Through the action of these nonlinear activation functions, the linear correlation limitation of shallow features in the modal fusion vector can be broken, effectively mining the deep semantic features after the fusion of different modal information, and improving the richness of feature expression. Furthermore, the multilayer perceptron design includes a bottle... The neck structure, where the dimension of the middle fully connected hidden layer is smaller than that of the input layer, achieves feature compression and dimensionality reduction. This bottleneck structure is the embedded feature dimensionality reduction module. Through this design of decreasing dimensionality, the feature dimension is gradually reduced, redundant information in the fusion vector is eliminated, and key features that can represent the core association of spoken fluency are selected based on the importance weight of the features, ensuring that the core fusion information is not lost while simplifying data complexity. Through the collaborative processing of the fully connected structure of the multilayer perceptron, nonlinear transformation and bottleneck dimensionality reduction structure, a fixed-dimensional, highly abstract fusion feature embedding is finally output. This embedding not only integrates the core information of the three modalities of speech, vision and interaction, but also has a streamlined dimensional specification.
[0052] Step 305: The fused features are embedded into the fluency analysis network. The fluency analysis network, through fully connected layers and regression layers, obtains a fluency evaluation vector containing the comprehensive fluency value and the evaluation values of each sub-dimension. Specifically, this includes: First, the fused features are embedded into the fluency analysis network, which is composed of a fully connected layer module and a regression output layer. The structure and implementation logic of each module are as follows: The fully connected layer module is a deep integration structure, specifically composed of 2 to 3 fully connected layers connected in series. The number of neurons in the first fully connected layer is consistent with the dimension of the fused feature embedding to ensure complete feature input. The number of neurons in subsequent layers gradually decreases according to a preset ratio. Finally, the number of neurons in the output layer strictly matches the total number of dimensions to be evaluated (including the comprehensive value and all sub-dimensions). ReLU activation function is used between layers to introduce non-linear expression, which enhances the ability to extract high-level semantic features. Through the progressive mapping of the weight matrices of each layer, the fused features are gradually transformed into a feature space directly related to the evaluation target, accurately strengthening the correspondence between features and fluency evaluation indicators.
[0053] Furthermore, the output of the fully connected layer module is directly input into the regression output layer, which adopts a dual parallel branch structure. The neuron configuration and computational logic of each branch are precisely adapted to the evaluation requirements. The first branch is the comprehensive value output branch, which contains only one neuron. By performing weighted summation and linear mapping regression calculation on the high-level semantic features output by the fully connected layer, it directly outputs a quantitative value that can comprehensively represent the learner's oral fluency level. The range of this quantitative value is preset to a continuous interval of 0 to 100 points according to the evaluation standard. The second branch is the sub-dimensional output branch, which contains the same number of neurons as the preset number of sub-evaluation dimensions. Each neuron corresponds to a sub-dimensional dimension (such as speech continuity, prosody, and interactive coherence). By configuring a targeted weight matrix for each neuron, dimension-specific regression calculation is performed on the high-level semantic features, and the quantitative score of each sub-dimensional is output. Each score also follows a unified value standard of 0 to 100 points.
[0054] To ensure the accuracy of the evaluation results, this fluency analysis network is trained using an end-to-end supervised learning approach. The training dataset is a multimodal spoken fluency dataset with precise annotations, containing multimodal spoken samples from learners of different learning levels, regional accents, and age groups. The annotations cover the overall fluency level and specific performance scores for each sub-dimension. During training, the mean squared error loss function is used to calculate the deviation between the model's predicted values and the labeled true values. The Adam optimizer iteratively updates the weight parameters of each layer of the network. At the same time, a learning rate decay strategy and an early stopping mechanism are introduced to avoid model overfitting and ensure that the network has stable evaluation performance in different scenarios. Through the parallel output of the two branches, a structured fluency evaluation vector is finally obtained. The vector arranges the overall fluency value and the evaluation values of each sub-dimension in a fixed order. This vector not only realizes the overall fluency evaluation but also refines the evaluation results of each core influencing dimension, effectively making up for the shortcomings of traditional single-modal evaluation dimensions and making the evaluation more in line with the needs of real language communication scenarios.
[0055] In this embodiment of the invention, an independently trainable linear projection layer adapts to the characteristics of various modal features, achieving accurate mapping of feature vectors from different modalities to a unified feature subspace; the generated query vector key vector value vector set provides standardized input for cross-modal attention calculation, ensuring the feasibility of subsequent cross-modal interactions; the cross-modal multi-head attention mechanism can simultaneously capture multi-dimensional correlations between multiple modalities, improving the comprehensiveness of inter-modal correlation mining; the calculation of the cross-attention weight matrix quantifies the correlation strength between features of different modalities, providing accurate weight basis for dynamic fusion; and the dynamic adaptation and fusion of value vectors is achieved based on the weighted summation of the weight matrix, enabling the fusion process to respond to the correlation strength of features of different modalities. The system effectively integrates complementary information from different modalities, enhancing the fusion vector's capacity to support multimodal collaborative features. The nonlinear transformation of the multilayer perceptron can uncover deep feature relationships within the fusion vector, enriching feature representation. Feature dimensionality reduction simplifies data complexity while preserving core fusion information, improving subsequent processing efficiency. The resulting unified fusion feature embedding provides a structured feature input for fluency analysis. The fully connected layer of the fluency analysis network fully integrates deep information from the fusion feature embedding, strengthening the connection between features and evaluation targets. The regression layer achieves ordered output of the overall fluency value and sub-dimensional evaluation values, forming a structured fluency evaluation vector that comprehensively carries evaluation information.
[0056] In a preferred embodiment of the present invention, step 400 above involves generating and outputting an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guiding learners to adjust their spoken expression strategies; triggering personalized learning tasks, including: Step 401 involves parsing the fluency assessment vector to extract the comprehensive fluency value and sub-dimensional evaluation values. Specifically, this includes: First, receiving the fluency assessment vector and determining that it is a structured multi-dimensional data set containing a comprehensive fluency value and multiple sub-dimensional evaluation values. These sub-dimensional values correspond to core assessment dimensions directly related to fluency, such as speech continuity, expressive prosody, and interactive responsiveness. Further, performing a structured parsing operation on the fluency assessment vector, according to a preset data format specification, splits the data elements in the vector by semantic category, separating the comprehensive fluency value representing the overall fluency level, and the evaluation values corresponding to each sub-dimensional. During this process, the validity of each parsed value is verified, eliminating abnormal values or missing data to ensure the accuracy and completeness of the extracted comprehensive fluency value and sub-dimensional evaluation values. Through the above parsing process, the integrated assessment vector is decomposed into independently analyzable sub-indicators, providing accurate and standardized data input for subsequent threshold comparison and weakness identification, laying the foundation for personalized guidance decision-making.
[0057] Step 402: Based on the overall fluency score and the evaluation scores of each dimension, compare them with the preset overall threshold and the thresholds of each dimension to determine whether each dimension is lower than its corresponding threshold. Specifically, this includes: First, retrieving a pre-stored set of threshold parameters, which contains the overall fluency threshold and the corresponding thresholds for each dimension. The specific preset method for the overall threshold and the thresholds of each dimension is as follows: First, conduct large-scale oral sample data collection, collecting a large number of oral samples from learners of different levels at different learning stages and in different target scenarios, covering multiple fluency levels such as beginner, intermediate, and advanced, to ensure that the samples have broad representativeness and coverage; then, annotate and analyze the collected oral samples, which are then evaluated by professional assessors. Based on unified assessment standards, each sample's overall fluency level and performance in sub-dimensions such as speech continuity, prosody, and interactive responsiveness are graded. Statistical analysis is conducted based on the graded sample data to calculate the distribution range and critical values of overall performance and performance in each sub-dimension under each fluency level. The lowest critical value corresponding to the qualified level is used as the initial comprehensive threshold and threshold for each dimension. Furthermore, the initial thresholds are calibrated and adjusted in conjunction with teaching objectives and learning patterns to ensure that the thresholds not only meet the general fluency assessment standards but also adapt to the ability requirements of different learning stages. After all thresholds are determined, they are stored in a threshold parameter set and can be dynamically adjusted according to the differences in target scenarios at different learning stages to ensure the adaptability of the thresholds.
[0058] Furthermore, the overall fluency score is compared with a preset overall threshold to determine whether the overall fluency meets the standard. At the same time, the evaluation value of each sub-dimension is compared with the corresponding dimension threshold one by one, and the comparison result of each dimension is recorded. The comparison results are summarized and organized to identify and mark the dimensions whose evaluation values are lower than the corresponding thresholds, i.e., the fluency weakness dimensions. Through a standardized threshold comparison process, the overall and sub-dimension compliance status can be accurately determined, the core directions that need improvement can be clearly identified, blind guidance can be avoided, and the pertinence and efficiency of subsequent processing can be improved.
[0059] Step 403: For each dimension judged to be below the threshold, calculate the standardized difference between the sub-dimensional evaluation value and the threshold as the deviation. Specifically, this includes: First, for each sub-dimensional dimension judged to be below the threshold, extract the evaluation value of that dimension and its corresponding preset threshold. Use these two values as two feature points in the preset one-dimensional feature space to establish the correspondence between the value and the spatial position, so that the dimensional difference can be intuitively represented by the spatial distance.
[0060] Furthermore, based on the aforementioned feature space mapping results, geometric distance calculation is performed, specifically calculating the Euclidean distance between two feature points. This distance value is the standardized difference. The reason for using Euclidean distance as the geometric distance calculation method is that it can reflect the straight-line distance between two points in one-dimensional space, and can directly and equivalently characterize the standardized gap between the evaluation value and the threshold, effectively eliminating the influence of threshold differences in different dimensions, and making the deviation of different dimensions have a unified and comparable standard. Through this standardization process achieved by geometric distance calculation, the qualitative judgment of the shortcoming dimension is transformed into a quantitative description of the degree of deviation. The obtained deviation can more accurately quantify the gap between each shortcoming dimension and the standard requirements, providing a reliable quantitative basis for subsequent strategy matching and task generation, and ensuring the objectivity of personalized guidance.
[0061] Step 404 integrates all dimension labels below the threshold and their corresponding deviations to obtain a decision feature vector containing the short-board dimension labels and their corresponding deviations. Specifically, this includes: First, collecting all labeled short-board dimension labels and the corresponding deviations for each short-board dimension to ensure a one-to-one correspondence between dimension labels and deviations, avoiding information misalignment; Second, according to a preset vector structure specification, combining the short-board dimension labels and their corresponding deviations in a fixed order to form a structured decision feature vector; Each element in the vector contains the association information between the dimension label and the deviation, while retaining the semantic attributes of each dimension; Through integration processing, the scattered short-board information is transformed into unified structured data, which facilitates rapid matching of feedback strategies and learning tasks, improves data processing efficiency, and provides complete decision support data for personalized guidance, ensuring that subsequent processing can comprehensively cover all short-board dimensions.
[0062] Step 405: Based on the decision feature vector, match and generate targeted oral expression adjustment guidance information from the preset multimodal feedback strategy library. The oral expression adjustment guidance information includes at least one of the following specific improvement suggestions: for pronunciation continuity, for expressive rhythm, and for interactive responsiveness. Specifically, it includes: First, retrieving the preset multimodal feedback strategy library. The specific preset method of this multimodal feedback strategy library is as follows: it is pre-constructed by combining common shortcomings in oral language learning, teaching practice experience, and language acquisition expert knowledge; first, systematically sorting out the core assessment dimensions of oral fluency, determining key shortcomings such as pronunciation continuity, expressive rhythm, and interactive responsiveness, and then dividing different deviation levels for each shortcoming dimension, and designing differentiated improvement guidance based on the learning difficulties corresponding to different deviation levels; then collecting and organizing high-quality oral language teaching cases, expert guidance plans, and solutions to common learner problems, and transforming these contents into specific and actionable improvement suggestions, such as a steady speech rate suggestion for low deviation in pronunciation continuity, and a suggestion for high deviation... The training program provides guidance on pause rhythm for speech improvement, including exercises on intonation and stress placement for rhythmic expression, and timing and transition techniques for interactive responsiveness. It then categorizes and archives all guidance content according to a logical structure of weakness dimensions, deviation levels, and improvement suggestions. Each strategy entry is labeled with its corresponding dimension and deviation range, forming a structured strategy library framework. Finally, the effectiveness of the strategy library is validated through a teaching pilot program. Learners of different levels are selected to try it out, and feedback is collected to optimize the accuracy and practicality of the improvement suggestions. Redundant content is removed, and guidance for missing scenarios is added, completing the pre-constructed multimodal feedback strategy library. This strategy library pre-stores oral expression improvement guidance for different weakness dimensions and deviation levels, covering specific improvement suggestions for multiple dimensions such as pronunciation continuity, rhythmic expression, and interactive responsiveness. For example, it provides suggestions on adjusting speech rate and pausing control for pronunciation continuity, intonation and stress placement for rhythmic expression, and timing and transition techniques for interactive responsiveness.
[0063] Furthermore, the decision feature vector is matched with the strategy entries in the strategy library. The corresponding strategy category is identified based on the weakness dimension, and then specific improvement suggestions are selected based on the deviation degree. The multiple improvement suggestions are integrated and optimized, duplicate content is removed, and they are sorted by importance to form a clear and organized guide to adjusting spoken expression. Through precise strategy matching, it is ensured that the generated guidance information is highly adapted to the learner's weakness dimension and deviation degree.
[0064] Step 406: Based on the weakness dimension identifier and deviation degree in the decision feature vector, locate the target region in the preset multi-dimensional task feature space; calculate the center vector of the target region, and select the basic template from the pre-stored task templates through vector similarity matching. Specifically, this includes: First, retrieving the preset multi-dimensional task feature space. The specific preset method of this multi-dimensional task feature space is as follows: it is pre-constructed by combining personalized oral training needs, weakness improvement logic, and teaching practice experience; First, sort out the core weakness dimensions and deviation degree levels of oral fluency, and determine the weakness dimension identifier and deviation degree as the core feature axes of the space. The weakness dimension identifier axis covers the identifiers corresponding to all evaluation dimensions such as pronunciation continuity, expression rhythm, and interactive responsiveness. The deviation degree axis is divided into three intervals: low, medium, and high according to the degree of difference; Then, set the corresponding task difficulty gradient and training focus for different intervals of each feature axis. For example, the weakness dimension identifier... When pronunciation continuity is recognized and the deviation is in the low range, basic difficulty speech rate control training is provided; when the deviation is in the high range, advanced difficulty pause rhythm reinforcement training is provided. Subsequently, various pre-set learning task templates are associated with different regions of the multi-dimensional feature space according to their training objectives, difficulty levels, and applicable scenarios, ensuring that the task templates in each region accurately match the training needs of the corresponding feature axis interval. Finally, the rationality of the space construction is verified through teaching pilots. Learners with different types of weaknesses and different degrees of deviation are selected to try it out, and feedback on task suitability is collected. The relationship between the feature axis interval division and the task template association is adjusted, and the space structure is optimized to improve matching accuracy, completing the pre-set construction of the multi-dimensional task feature space. This space uses weakness dimension identification and deviation as the core feature axis, with different intervals on each axis corresponding to different task difficulties and training focuses. Each region within the space is associated with a specific type of learning task template.
[0065] Furthermore, based on the weakness dimension identifier and deviation degree in the decision feature vector, the corresponding target region is located in the multi-dimensional task feature space. This region accurately matches the learner's current weakness type and gap degree. The center vector of the target region is calculated and used as the core reference for task matching. The similarity of the center vector with the feature vectors of all pre-stored task templates is compared, and the template with the highest similarity is selected as the base template. Through spatial localization and similarity matching, it is ensured that the selected base template is highly consistent with the learner's current level and weakness needs, providing a high-quality basic framework for subsequent personalized task generation and reducing the randomness of task generation.
[0066] Step 407: The information entropy density and cognitive load coefficient of the basic template are calibrated based on the deviation to obtain adaptation parameters. A structured personalized learning task is generated based on these adaptation parameters, specifically including: First, determining that the basic template contains two core parameters: information entropy density and cognitive load coefficient. Information entropy density represents the amount of information and complexity of the task, while cognitive load coefficient represents the degree of cognitive ability required of the learner. Further, the two core parameters are calibrated based on the deviation: If the deviation is large, it indicates a significant gap between the learner and the standard requirements, thus reducing the information entropy density and cognitive load coefficient of the basic template. The cognitive load coefficient simplifies task content and reduces learning difficulty; if the deviation is small, two parameters are appropriately increased to enhance the task's challenge; through the linkage calibration of deviation and parameters, a parameter combination suitable for the learner's current level is obtained; based on the calibrated suitable parameters, the content structure, training intensity, and completion time of the basic template are adjusted to generate a structured, personalized learning task; through dynamic calibration and adjustment, the learning task is precisely matched to the learner's weaknesses, avoiding both excessively difficult tasks that lead to learner frustration and tasks that are too easy to achieve the desired improvement, thus effectively enhancing the adaptability and effectiveness of the learning task.
[0067] Step 408 involves encapsulating the spoken language adjustment guidance information and the personalized learning task, and outputting an online adaptive guidance signal to the learner through the platform's multi-channel human-computer interaction interface. Specifically, this includes: first, integrating and encapsulating the spoken language adjustment guidance information and the personalized learning task to form a structured online adaptive guidance signal, ensuring a clear logical connection between the two types of information during the encapsulation process for easy understanding and execution by the learner; further, calling the platform's multi-channel human-computer interaction interface, which supports multiple output methods such as text, voice, pop-ups, and message pushes, adapting to different learning terminals and usage scenarios; automatically selecting the optimal output channel based on the learner's preset preferences and the current terminal environment, and pushing the encapsulated online adaptive guidance signal to the learner; through multi-channel output, ensuring that the guidance signal can reach the learner efficiently and promptly, while adapting to different usage scenarios and terminal devices, improving the convenience and coverage of online adaptive guidance, and meeting the needs of real-time processing on the terminal side.
[0068] In this embodiment of the invention, the fluency assessment vector is structured and parsed to extract the overall fluency value and the evaluation values of each sub-dimension, achieving hierarchical decomposition of assessment information. This ensures the integrity and accuracy of the assessment data, providing standardized data input for subsequent threshold comparison, deviation calculation, and other processing, laying the foundation for accurate decision-making. By comparing the overall fluency value and the evaluation values of each sub-dimension with preset thresholds, the compliance status of each dimension can be standardized. The weaker dimensions in the fluency assessment are clearly identified, determining the core direction for subsequent guidance and improvement, making targeted processing goal-oriented and improving decision-making efficiency. The standardized difference of dimensions below the threshold is calculated as the deviation, which can quantitatively characterize the gap between each weak dimension and the standard requirements. The qualitative dimension compliance judgment is transformed into a quantitative description of the degree of deviation, providing a quantitative basis for subsequent strategy matching and task generation. Integrating the weak dimension identifiers and their deviations to form a decision feature vector can structurally integrate scattered weak information. This achieves the orderly organization of multi-dimensional weak information, facilitating rapid matching of feedback strategies and learning tasks, improving the efficiency of data processing, and providing complete decision support data for personalized guidance. Based on the decision feature... The system generates guidance information by matching vectors from a pre-defined feedback strategy library, enabling precise matching of improvement prompts with the dimensions of weaknesses. It provides targeted improvement directions for weaknesses in different dimensions such as pronunciation continuity, expressive rhythm, and interactive responsiveness, ensuring the practicality and adaptability of the guidance information and improving the effectiveness of learners' expression strategies. By locating target regions in a multi-dimensional task feature space, it can match basic task templates that are compatible with the dimensions of weaknesses and deviations. Selecting basic templates based on vector similarity ensures the task templates fit the learner's current level, providing a high-quality foundation for subsequent personalized task generation. Calibrating the information entropy density and cognitive load coefficient of the basic templates based on deviations allows for dynamic adjustment of task difficulty and information capacity. This ensures that the generated personalized learning tasks accurately match the learner's degree of weakness, avoiding tasks that are too difficult or too easy, thus improving the adaptability and effectiveness of the learning tasks. Encapsulating and integrating oral expression adjustment guidance information with personalized learning tasks enables structured output of guidance content. Outputting through multi-channel human-computer interaction interfaces adapts to different interaction scenarios and learning terminals, ensuring that guidance signals can efficiently reach learners and improving the convenience and coverage of online adaptive guidance.
[0069] like Figure 2 As shown, embodiments of the present invention also provide a multimodal fusion real-time English speaking fluency assessment system, comprising: The acquisition synchronization module is used to acquire learners' multimodal synchronous acquisition streams, which include speech modality data, visual modality data, and interaction modality data; The feature alignment module is used to extract key facial geometric feature points from the visual modality data and calculate spatial correction parameters based on the key facial geometric feature points; perform geometric correction on the visual modality data using the spatial correction parameters and synchronize it with the speech modality data in time to obtain visual features; extract acoustic features from the speech modality data to obtain speech features; extract interaction features based on the temporal information in the interaction modality data; and obtain multimodal feature data based on the visual features, speech features, and interaction features. The fusion evaluation module is used to input the multimodal feature data into the multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector. The multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features, and interaction features to form a fusion feature embedding. Fluency analysis is then performed based on the fusion feature embedding to obtain a fluency evaluation vector that includes a fluency value and sub-dimensional evaluation values. The feedback module is used to generate and output an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their oral expression strategies; or trigger a personalized learning task.
[0070] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0071] Embodiments of the present invention also provide a computing device, including: a processor and a memory storing a computer program, wherein the computer program, when executed by the processor, performs the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effect. The above descriptions are preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A multimodal fusion-based real-time assessment method for oral English fluency, characterized in that, The method includes: Step 100: Collect the learner's multimodal synchronous acquisition stream, which includes speech modality data, visual modality data, and interaction modality data; Step 200: Extract key facial geometric feature points from the visual modality data, and calculate spatial correction parameters based on the key facial geometric feature points; perform geometric correction on the visual modality data using the spatial correction parameters, and synchronize it with the speech modality data in time to obtain visual features; extract acoustic features from the speech modality data to obtain speech features; extract interaction features based on the temporal information in the interaction modality data; and obtain multimodal feature data based on the visual features, speech features, and interaction features. Step 300: The multimodal feature data is input into the multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector; wherein, the multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features and interaction features to form a fusion feature embedding, and performs fluency analysis based on the fusion feature embedding to obtain a fluency evaluation vector containing a fluency value and sub-dimensional evaluation values; Step 400: Generate and output an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their spoken expression strategies; trigger a personalized learning task.
2. The method for real-time assessment of oral English fluency in multi-modal fusion according to claim 1, wherein, Step 100 includes: The learner's original audio stream is acquired through a terminal microphone array, and the original audio stream is subjected to high-frequency pre-enhancement and overlapping window frame division to obtain speech modal data. The original video stream containing the learner's face and upper body area is acquired by the terminal camera, and the original video stream is decoded and contrast-limited adaptive histogram equalization is performed to obtain visual modality data. By capturing learners' response behavior in dialogue practice in real time through platform interaction logs, discrete time marker sequences including question-and-answer turn switching times and topic continuity markers are extracted to obtain interaction modality data; The voice modal data, visual modal data, and interaction modal data are timestamped to form a multimodal synchronous acquisition stream.
3. The multimodal fusion method for real-time assessment of English spoken fluency according to claim 2, characterized in that, Step 200 includes: Facial region images are extracted from the visual modality data and preprocessed to obtain standardized image data; Based on the standardized image data, candidate facial landmarks are located and screened through multi-scale corner detection and feature descriptor matching. Geometric constraints are applied to the candidate facial landmarks and iterative optimization is performed to obtain a set of key facial geometric feature points; Each feature point in the set of key facial geometric feature points is paired with a feature point in a predefined standard frontal face model that has the same semantic annotation to establish a semantic alignment relationship between the feature points. Based on all feature point pairs associated with the semantic alignment relationship of the feature points, a similarity transformation matrix is solved by the least squares fitting method to minimize the coordinate error between the set of key facial geometric feature points and the standard frontal facial model. Spatial correction parameters characterizing rotation, translation, and scaling transformations are extracted from the similarity transformation matrix.
4. The multimodal fusion method for real-time assessment of English spoken fluency according to claim 3, characterized in that, Step 200 further includes: Based on the spatial correction parameters, an affine transformation matrix is constructed, and the affine transformation matrix is used to perform geometric transformation on each frame of the visual modality data, mapping the facial region in the image to a predefined normalized coordinate system to obtain a pose-normalized facial region image sequence. The timestamps of the pose-normalized facial region image sequence are matched and aligned with the timestamps of the speech modality data to obtain spatiotemporally aligned visual features. The speech modal data is subjected to short-time Fourier transform to calculate the Mel frequency cepstral coefficients, fundamental frequency profile, and short-time energy of each audio frame, thereby obtaining speech features that characterize pronunciation, prosody, and rhythm. For the discrete time-stamped sequence in the interactive modal data, the question-and-answer response delay distribution, topic keyword repetition rate, and turn-taking frequency within a preset time window are statistically analyzed to obtain interactive features characterizing dialogue coherence and interactivity. The visual features, voice features, and interaction features are synchronized in time and fused through channel splicing to form multimodal feature data with a unified dimension.
5. The multimodal fusion method for real-time assessment of English spoken fluency according to claim 4, characterized in that, Step 300 includes: The visual features, speech features, and interaction features corresponding to the multimodal feature data are respectively input into independent trainable linear projection layers, and the feature vector of each modality is projected onto a unified feature subspace to obtain a set of query vectors, key vectors, and value vectors corresponding to each modality. Based on the set of query vector, key vector, and value vector, a cross-modal multi-head attention mechanism is used to calculate the cross-attention weight matrix between any two modal features. Based on the cross-attention weight matrix, the value vectors from different modalities are weighted and summed to obtain a dynamically weighted modality fusion vector; The dynamically weighted modal fusion vector is input into a multilayer perceptron, and a unified fusion feature embedding is formed through nonlinear transformation and feature dimensionality reduction. The fusion features are embedded into the fluency analysis network, which, through a fully connected layer and a regression layer, yields a fluency evaluation vector containing a comprehensive fluency value and evaluation values for each of the sub-dimensions.
6. The multimodal fusion method for real-time assessment of English spoken fluency according to claim 5, characterized in that, Step 400 includes: The fluency evaluation vector is parsed to extract the overall fluency value and the evaluation values of each sub-dimension. Based on the overall fluency value and the evaluation values of each dimension, the values are compared with the preset overall threshold and the threshold of each dimension to determine whether each dimension is lower than its corresponding threshold. For each dimension that is judged to be below the threshold, the standardized difference between the sub-dimensional evaluation value and the threshold is calculated as the deviation. Integrate all dimension identifiers below the threshold and their corresponding deviations to obtain a decision feature vector containing the shortest dimension identifier and its corresponding deviation. Based on the decision feature vector, targeted spoken expression adjustment guidance information is matched and generated from a preset multimodal feedback strategy library. The spoken expression adjustment guidance information includes at least one of the following specific improvement suggestions: for pronunciation continuity, for expression rhythm, and for interactive responsiveness.
7. The multimodal fusion method for real-time assessment of English spoken fluency according to claim 6, characterized in that, Step 400 further includes: Based on the shortest dimension identifier and deviation in the decision feature vector, the target region is located in the preset multi-dimensional task feature space; the center vector of the target region is calculated, and a basic template is selected from the pre-stored task templates by vector similarity matching; The information entropy density and cognitive load coefficient of the basic template are calibrated based on the deviation to obtain the adaptation parameters; a structured personalized learning task is generated based on the adaptation parameters. The spoken language adjustment guidance information is encapsulated with the personalized learning task, and online adaptive guidance signals are output to learners through the platform's multi-channel human-computer interaction interface.
8. A multimodal fusion real-time English speaking fluency assessment system, the system implementing the method as described in any one of claims 1 to 7, characterized in that, include: The acquisition synchronization module is used to acquire learners' multimodal synchronous acquisition streams, which include speech modality data, visual modality data, and interaction modality data; The feature alignment module is used to extract key facial geometric feature points from the visual modality data, and calculate spatial correction parameters based on the key facial geometric feature points; perform geometric correction on the visual modality data using the spatial correction parameters, and synchronize with the speech modality data in time to obtain visual features; Acoustic features are extracted from the speech modal data to obtain speech features; Based on the temporal information in the interaction modal data, interaction features are extracted; based on the visual features, voice features, and interaction features, multimodal feature data is obtained. The fusion evaluation module is used to input the multimodal feature data into the multimodal fusion evaluation model for real-time fusion and evaluation to obtain a fluency evaluation vector. The multimodal fusion evaluation model uses a cross-modal attention mechanism to dynamically associate and weight the visual features, speech features, and interaction features to form a fusion feature embedding. Fluency analysis is then performed based on the fusion feature embedding to obtain a fluency evaluation vector that includes a fluency value and sub-dimensional evaluation values. The feedback module is used to generate and output an online adaptive guidance signal based on the fluency assessment vector. The online adaptive guidance signal is used to perform at least one of the following: guide learners to adjust their oral expression strategies; or trigger a personalized learning task.
9. A computing device, characterized in that, include: One or more processors; A storage device for storing one or more programs, which, when executed by one or more processors, cause the one or more processors to perform the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 7.