English auxiliary teaching software multi-modal optimization method based on embedded system

By extracting and classifying multimodal content feature parameters, and combining teaching objectives with learners' cognitive patterns, the content resource set of English teaching aids is constructed and optimized. This solves the problems of static content organization, logical breaks, and cognitive mismatch in existing systems, and achieves efficient, scientific, and device-friendly multimodal teaching.

CN121456799BActive Publication Date: 2026-05-19SHENYANG NORMAL UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENYANG NORMAL UNIV
Filing Date
2025-10-29
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Existing embedded English teaching software lacks the ability to systematically extract and intelligently schedule key parameters such as text density, speech rhythm and image resolution in multimodal content organization, resulting in chaotic content presentation logic, unbalanced cognitive load, poor device adaptability, and difficulty in dynamically generating scientific, coherent and efficient learning sequences based on learners' cognitive stages and teaching objectives.

Method used

By extracting feature parameters from multimodal content, classifying it using data grouping methods, and combining teaching objectives with learners' cognitive patterns, a set of content resources is constructed. The content presentation structure is then adjusted through sequence generation, rearrangement, and optimization models to ensure logical coherence and device compatibility, dynamically matching learning needs.

Benefits of technology

It realizes intelligent grouping and dynamic optimization of English teaching content in an embedded environment, which improves the hierarchy, coherence and personalization of teaching content, reduces learners' cognitive load, enhances the synergistic effect of multimodal information, and ensures the real-time operation and stability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121456799B_ABST
    Figure CN121456799B_ABST
Patent Text Reader

Abstract

The application discloses an English auxiliary teaching software multi-modal optimization method based on an embedded system, aiming at solving the problem that the existing embedded English auxiliary teaching software lacks systematic extraction and intelligent scheduling capability for key parameters such as text density, speech rhythm and image resolution, resulting in content logic confusion, cognitive load imbalance and poor device adaptability. The method comprises the following steps: extracting the text density, speech rhythm and image resolution of multi-modal content and grouping; combining teaching target matching and supplementary resources to determine the optimization set; constructing an initial presentation structure based on the cognitive level of learners; rearranging the difficulty gradient to ensure logical coherence; dynamically replacing resources according to individual needs; and adjusting the loading order for the low-power characteristic of the embedded system. By adopting the above technical scheme, hierarchical, coherent and personalized presentation of teaching content can be realized, cognitive load is reduced, collaborative effect is improved, and efficient and stable operation on resource-limited devices is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of computer software and educational technology integration, specifically relating to a multimodal optimization method for English teaching aid software based on an embedded system. Background Technology

[0002] With the deepening of educational informatization, English teaching software, as a core carrier of the digital transformation of language education, is gradually being integrated into various smart terminals and embedded devices, providing learners with a flexible and convenient language training environment. Embedded systems, with their advantages of low power consumption, high integration, and portability, have become an important platform for mobile English teaching and are widely used in students' daily learning scenarios. However, limited by hardware resources and algorithm efficiency, existing embedded English teaching software has significant shortcomings in content organization logic and multimodal presentation strategies, making it difficult to simultaneously meet the dual needs of scientific teaching and user experience.

[0003] Among these, the dynamic optimization of multimodal teaching content has become a key direction for improving the effectiveness of supplementary teaching. English teaching materials typically integrate multiple modalities such as text, audio, and images, and their effective synergy directly affects learners' depth of understanding and cognitive load. An ideal teaching system should be able to automatically construct a logically coherent sequence of content based on learners' cognitive stages and teaching objectives, and achieve reasonable matching between the attributes of each modality. However, current systems generally adopt a static, pre-set content arrangement method, lacking the ability to quantitatively analyze and intelligently schedule the inherent characteristics of the materials.

[0004] Existing technologies struggle to systematically extract and group key parameters of multimodal resources, such as text density, speech rhythm, and image resolution, resulting in a lack of hierarchy and adaptability in content organization. Furthermore, the teaching content is disconnected from the knowledge point system, failing to dynamically adjust the presentation order based on difficulty gradients and time distribution, easily leading to cognitive overload for beginners or loss of interest for advanced learners. In addition, in embedded environments, systems must balance memory usage and processing efficiency, but existing solutions often neglect the collaborative optimization of multimodal attributes (such as visual complexity and speech rate), resulting in significant deficiencies in the logical coherence, cognitive adaptability, and device compatibility of the content presentation structure. Therefore, there is an urgent need for a method that can intelligently organize and dynamically optimize multimodal English teaching content in resource-constrained embedded systems. Summary of the Invention

[0005] The purpose of this invention is to provide a multimodal optimization method for English teaching software based on embedded systems, which can effectively solve the problems mentioned in the background. Existing embedded English teaching software lacks the ability to systematically extract and intelligently schedule key parameters such as text density, speech rhythm, and image resolution in multimodal content organization, resulting in logically chaotic content presentation, cognitive load imbalance, poor device adaptability, and difficulty in dynamically generating scientific, coherent, and efficient learning sequences based on learners' cognitive stages and teaching objectives.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A multimodal optimization method for English teaching software based on an embedded system includes the following specific steps:

[0008] Step 1: Extract feature parameters of multimodal content resources, including text density, speech rhythm and image resolution. Use data grouping method to group the feature parameters to obtain a classified set of content resources.

[0009] Step 2: Match the categorized content resource set with the preset teaching objectives, which include knowledge point coverage and interactivity requirements. If the matching degree is higher than the preset threshold, the set is retained; otherwise, supplementary resources are obtained from the content resource library. The supplementary resources meet the requirements of topic relevance and diversity, thus determining the optimized content resource set.

[0010] Step 3: Obtain the difficulty level and duration distribution attributes of each resource in the optimized content resource set. The attributes are evaluated based on the learner's cognitive level and the content duration. The initial content presentation structure is constructed by using a sequence generation method combined with the principle of progressive learning to obtain the preliminary presentation structure.

[0011] Step 4: For the initial presentation structure, the distribution of difficulty levels reflects the gradient of resources from simple to complex. If the distribution is uneven, the adjacent resource positions are adjusted and reordered. The logical coherence of the presentation structure is judged to ensure smooth connection of knowledge points.

[0012] Step 5: Analyze the visual complexity and speech speed attributes from the reordered presentation structure. These attributes involve the density of image elements and the rate of speech unit. Use an optimization model to adjust the combination of these attributes to obtain the enhanced content presentation structure.

[0013] Step 6: Based on the relevance between the enhanced content presentation structure and learning needs, including personalized progress and interest preferences, if the needs change, the resources in the presentation structure are dynamically replaced, and the replacement maintains overall balance to determine the final content presentation structure.

[0014] Step 7: Obtain the operating parameters of the final content presentation structure in the embedded system. The parameters include memory usage and processing speed. Adjust the loading order of the presentation structure for low power consumption characteristics to obtain a content organization structure adapted to mobile devices.

[0015] Preferably, in step 1, text density is calculated by weighting the number of effective words per unit area with syntactic complexity, speech rhythm is quantified based on syllable interval time and stress distribution entropy, and image resolution is comprehensively evaluated by combining pixel density and edge gradient information. After normalization of the three, the data are input into the K-means clustering algorithm, with the number of clusters set to 5, which is used to divide the content resource sets into five categories: primary, intermediate, advanced, enhanced, and extended.

[0016] Preferably, in step 2, the knowledge point coverage is calculated by comparing the node matching rate between the resource tags and the knowledge graph of the curriculum standard. The interactivity requirement is determined based on the number of embedded interactive elements in the resource and the response latency threshold. The matching threshold is set to 85%. When it is lower than this threshold, the system retrieves supplementary resources from the cloud content resource library with a topic relevance score greater than 0.9 and a diversity index greater than 0.7 and injects them.

[0017] Preferably, in step 3, the difficulty level is determined by the frequency of words, the average sentence length, and the depth of the syntax tree. The duration distribution attribute limits the duration of a single resource to no more than 120 seconds. The sequence generation method adopts a path planning algorithm based on Markov decision process. The state space is the knowledge point node, the action space is the resource selection, and the reward function integrates the smoothness of the difficulty gradient and the balance of duration to ensure that the initial presentation structure conforms to the cognitive law of gradual deepening.

[0018] Preferably, in step 4, the uniformity of the difficulty level distribution is determined by calculating the standard deviation of the difficulty difference between adjacent resources. If the standard deviation is greater than 0.3, a local rearrangement mechanism is initiated, and resource positions are exchanged within a window length of 3 using a sliding window. Logical coherence is verified through a knowledge point dependency graph. The resource position is only allowed to be fixed when all the preceding dependency nodes of the subsequent knowledge point have been covered.

[0019] Preferably, in step 5, visual complexity is calculated by combining the number of objects in the image, color saturation variance, and texture energy. Speech speed is quantified in terms of syllables per minute. The optimization model is a multi-objective constraint satisfaction problem solver. The objective function minimizes the covariance between visual complexity and speech speed. The constraints are that the speech speed is not less than 120 syllables per minute and not more than 200 syllables per minute, and the number of image objects does not exceed 8.

[0020] Preferably, in step 6, learning needs are dynamically modeled using user historical interaction logs and real-time feedback signals; personalized progress is inferred from the combined proportion of completed knowledge points and error rate trends; interest preferences are calculated based on a weighted average of resource click frequency and dwell time; and the dynamic replacement mechanism adopts a sliding buffer strategy, replacing low-relevance resources at a rate not exceeding 20% ​​while maintaining the overall difficulty curve and modal balance.

[0021] Preferably, in step 7, memory usage is obtained by statically analyzing the size of resource files and dynamically monitoring runtime stack peaks. Processing speed is based on the measurement of resource decoding and rendering time. Loading order adjustment adopts a layered preloading strategy, placing high-priority resources (such as the first screen content and key knowledge points) in the memory resident area, and sorting the remaining resources according to their usage probability. The delayed loading threshold is set to complete loading within 500 milliseconds after the user operation.

[0022] Preferably, the method deploys a lightweight feature extraction module in the embedded system, the text density calculation module occupies less than 2 megabytes of memory, the speech rhythm analysis module uses fixed-point arithmetic, and the processing time for a single audio segment is less than 800 milliseconds, the image resolution evaluation module is based on edge detection and region segmentation, and the single-frame processing latency is less than 30 milliseconds. The entire system can run stably on an ARM Cortex-A7 processor with a main frequency of 800 MHz.

[0023] Preferably, the content resource library contains more than 50,000 labeled resources. Each resource is associated with knowledge point tags, difficulty levels, modal attributes, and teaching objective mapping relationships. The system supports offline caching of commonly used resource sets, and the cache capacity can be configured from 512 megabytes to 2 gigabytes to ensure that the multimodal optimization process can still be completed in network-limited environments.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] This invention systematically extracts feature parameters such as text density, speech rhythm, and image resolution from multimodal content, and combines these with teaching objectives and learners' cognitive patterns to achieve intelligent grouping, dynamic matching, and sequence optimization of English supplementary teaching content in embedded environments. This method effectively solves the problems of static content organization, logical breaks, and cognitive mismatch in existing systems, significantly improving the hierarchy, coherence, and personalization of teaching content. Simultaneously, through the synergistic optimization of visual complexity and speech speed, it reduces learners' cognitive load and enhances the synergistic effect of multimodal information. Furthermore, considering the low power consumption and memory constraints of embedded systems, this invention designs a lightweight feature extraction and hierarchical loading mechanism, ensuring both optimization effectiveness and the real-time performance and stability of the system, providing efficient, scientific, and device-friendly technical support for mobile English teaching. Attached Figure Description

[0026] Figure 1 This is a schematic diagram of the overall technical solution architecture of the multimodal optimization method for English teaching software based on embedded systems proposed in this invention;

[0027] Figure 2 This is a schematic diagram of the core principle framework of multimodal feature parameter extraction and intelligent grouping in this invention;

[0028] Figure 3 This is a flowchart illustrating the content selection logic based on teaching objective matching and resource optimization in this invention.

[0029] Figure 4 This is a flowchart illustrating the logical process of generating and rearranging the initial content sequence in this invention, which combines cognitive patterns and difficulty gradients.

[0030] Figure 5 This is a schematic diagram of the multi-level interaction relationship and data flow of visual complexity and voice speed collaborative optimization and embedded adaptation in this invention. Detailed Implementation

[0031] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0032] Currently, with the deepening of educational informatization, English tutoring software, as a core carrier of digital transformation in language education, is gradually integrating into various smart terminals and embedded devices, providing learners with a flexible and convenient language training environment. Embedded systems, with their advantages of low power consumption, high integration, and portability, have become an important platform for mobile English teaching and are widely used in students' daily learning scenarios. However, limited by hardware resources and algorithm efficiency, existing embedded English tutoring software has significant shortcomings in content organization logic and multimodal presentation strategies, making it difficult to simultaneously meet the dual needs of scientific teaching and user experience. To address these technical problems, this invention proposes a method for intelligent grouping, dynamic matching, and sequence optimization of English tutoring content in an embedded environment by systematically extracting feature parameters such as text density, speech rhythm, and image resolution of multimodal content, and combining this with teaching objectives and learners' cognitive patterns. This effectively solves the problems of static content organization, logical breaks, and cognitive mismatch in existing systems and is applied to a multimodal optimization method for English tutoring software based on embedded systems.

[0033] refer to Figure 1The overall technical architecture of the multimodal optimization method for English teaching software based on embedded systems proposed in this invention includes a multimodal feature extraction module, a teaching objective matching module, a sequence generation and rearrangement module, a multimodal collaborative optimization module, a dynamic replacement module, and an embedded adaptation module. These modules work collaboratively to form a closed-loop optimization process, ensuring efficient, scientific, and device-friendly content organization on resource-constrained embedded devices.

[0034] In the above method, step 1 involves extracting feature parameters from multimodal content resources. These feature parameters include text density, speech rhythm, and image resolution. A data grouping method is then used to group these feature parameters to obtain a categorized set of content resources. Specifically, in step 1, text density is calculated by weighting the number of effective words per unit area with syntactic complexity. The number of effective words refers to the number of content words after removing stop words. Syntactic complexity is comprehensively evaluated based on the average branching factor and nesting depth of the dependency parsing tree. Speech rhythm is quantified based on syllable interval time and stress distribution entropy. Syllable interval time is obtained by using a forced alignment algorithm to obtain the time difference between the center points of adjacent syllables. Stress distribution entropy is calculated based on acoustic features (such as fundamental frequency and energy envelope) to determine the uncertainty of stress position. Image resolution is comprehensively evaluated by combining pixel density and edge gradient information. Pixel density refers to the number of effective pixels per unit area. Edge gradient information is calculated using the Sobel operator to calculate the mean and variance of the image gradient magnitude. The three types of feature parameters mentioned above, after normalization, are input into the K-means clustering algorithm. The number of clusters is set to 5 to divide the content resource sets into five categories: primary, intermediate, advanced, enhanced, and extended. Normalization uses the min-max scaling method to map each parameter to the interval between 0 and 1, ensuring the comparability of features with different dimensions. The initial centroids of K-means clustering are determined using the K-means++ algorithm, and the iteration termination condition is that the centroid change is less than 1e-4 or the maximum number of iterations reaches 100. This process is executed by a lightweight feature extraction module in the embedded system. The text density calculation module occupies less than 2 megabytes of memory. The speech rhythm analysis module uses fixed-point arithmetic, processing a single audio segment in less than 800 milliseconds. The image resolution evaluation module is based on edge detection and region segmentation, with a single frame processing latency of less than 30 milliseconds. The entire system can run stably on an 800 MHz ARM Cortex-A7 processor. (Reference) Figure 2 The figure illustrates the core principle framework of multimodal feature parameter extraction and intelligent grouping, including three feature extraction channels: text, speech, and image, and their fusion and clustering process.

[0035] In the above method, step 2 involves matching the categorized content resource set with preset teaching objectives, including knowledge point coverage and interactivity requirements. If the matching degree is higher than a preset threshold, the set is retained; otherwise, supplementary resources are retrieved from the content resource library. These supplementary resources meet the requirements of topic relevance and diversity, thus determining an optimized content resource set. Specifically, in step 2, knowledge point coverage is calculated by comparing the node matching rate between resource tags and the curriculum standard knowledge graph. The curriculum standard knowledge graph is stored in the form of a directed acyclic graph, where nodes represent knowledge points and edges represent dependencies. The matching rate is equal to the ratio of the number of knowledge points covered by the resource tags to the total number of knowledge points required by the teaching objectives. Interactivity requirements are determined based on the number of embedded interactive elements in the resources and a response latency threshold. Interactive elements include multiple-choice questions, drag-and-drop questions, and voice follow-up feedback. The response latency threshold is set to 300 milliseconds, meaning the system feedback time after a user operation must not exceed this threshold. The matching degree threshold is set to 85%. When it falls below this threshold, the system retrieves supplementary resources from the cloud content resource library that have a topic relevance score greater than 0.9 and a diversity index greater than 0.7 for injection. The topic relevance score is obtained by calculating the cosine similarity between the resource text and the keywords of the teaching topic. The diversity index is calculated based on the Shannon entropy of the resource modality type, difficulty level, and source. The content resource library contains over 50,000 annotated resources, each associated with knowledge point tags, difficulty level, modal attributes, and a mapping relationship with teaching objectives. The system supports offline caching of frequently used resource sets, with a configurable cache capacity of 512 megabytes to 2 gigabytes, ensuring that the multimodal optimization process can still be completed in network-constrained environments. (Reference) Figure 3 The diagram illustrates the logical process framework for content selection based on teaching objective matching and resource optimization, including the complete process of matching degree calculation, threshold judgment, supplementary resource retrieval and injection.

[0036] In the above method, step 3 involves obtaining the difficulty level and duration distribution attributes of each resource in the optimized content resource set. These attributes are assessed based on the learner's cognitive level and content duration. An initial content presentation structure is constructed using a sequence generation method combined with progressive learning principles, resulting in a preliminary presentation structure. Specifically, in step 3, the difficulty level is determined by vocabulary frequency, average sentence length, and syntax tree depth. Vocabulary frequency is mapped to a difficulty value of 0 to 1 based on the COCA corpus word frequency ranking. The average sentence length refers to the average number of words in a sentence. The syntax tree depth is obtained through dependency parsing to determine the maximum nesting level of the syntactic structure. The duration distribution attribute limits the duration of a single resource to no more than 120 seconds to avoid distracting the learner. The sequence generation method employs a path planning algorithm based on Markov decision processes. The state space represents knowledge point nodes, and the action space represents resource selection. The reward function integrates the smoothness of the difficulty gradient with the balance of duration, ensuring that the preliminary presentation structure conforms to the cognitive pattern of gradual progression from simple to complex. The state transition probability of the Markov decision process is constructed based on a knowledge point dependency graph, allowing only transitions from preceding knowledge points to subsequent knowledge points. The reward function is defined as:

[0037]

[0038] Where di represents the difficulty level of the i-th resource, ti represents the duration of the i-th resource, and α and β are weight coefficients, set to 0.6 and 0.4 respectively, to ensure a comprehensive optimization of difficulty gradient smoothness and duration balance. The algorithm is solved using Q-learning, with a learning rate of 0.1 and a discount factor of 0.9. The exploration strategy employs... -greedy, The initial value is 0.3 and decreases with the number of iterations.

[0039] In the above method, step 4 addresses the difficulty level distribution in the initial presentation structure. This distribution reflects the gradient of resources from simple to complex. If the distribution is uneven, adjacent resource positions are adjusted and reordered to determine the logical coherence of the presentation structure and ensure smooth connection of knowledge points. Specifically, in step 4, the uniformity of the difficulty level distribution is determined by calculating the standard deviation of the difficulty difference between adjacent resources. If the standard deviation is greater than 0.3, a local reordering mechanism is initiated. Resource positions are exchanged within a window of length 3 using a sliding window, and logical coherence is verified through a knowledge point dependency graph. A resource position is only allowed to be fixed when all preceding dependency nodes of a subsequent knowledge point have been covered. The sliding window starts from the beginning of the sequence and moves one resource position at a time. Resources within the window are arranged in ascending order of difficulty level, but must satisfy knowledge point dependency constraints. The knowledge point dependency graph is stored in the form of an adjacency list, and the time complexity for querying preceding dependency nodes is O(1). The reordering process is executed iteratively until the standard deviation drops below 0.3 or a complete scan is completed without improvement. This mechanism ensures that the content sequence satisfies both the smoothness of the difficulty gradient and the logical dependence of knowledge points, avoiding cognitive leaps or knowledge gaps.

[0040] In the above method, step 5 analyzes the visual complexity and speech velocity attributes from the reordered presentation structure. These attributes involve image element density and speech unit rate. An optimization model is used to adjust the combination of these attributes to obtain an enhanced content presentation structure. Specifically, in step 5, visual complexity is calculated by comprehensively considering the number of objects in the image, color saturation variance, and texture energy. The number of objects is obtained using a lightweight object detection model (such as MobileNet-SSD), color saturation variance is calculated based on the HSV color space, and texture energy is obtained by weighting the contrast and correlation indices of the gray-level co-occurrence matrix. Speech velocity is quantified in terms of syllables per minute, and the number of syllables per unit time is counted using a syllable segmentation algorithm. The optimization model is a multi-objective constraint satisfaction problem solver. The objective function minimizes the covariance between visual complexity and speech velocity. The constraints are that the speech velocity is not less than 120 syllables per minute and not more than 200 syllables per minute, and the number of image objects does not exceed 8. The covariance calculation formula is:

[0041]

[0042] Where vi is the visual complexity of the i-th resource, and si is the speech speed of the i-th resource. and These are the means. The solver uses a genetic algorithm with a population size of 50, a crossover probability of 0.8, a mutation probability of 0.1, and a maximum of 200 generations. This optimization ensures a balanced load between the visual and auditory modalities, avoiding cognitive fatigue caused by overload of a single modality. (Reference) Figure 5The diagram illustrates the multi-level interaction and data flow between visual complexity and voice speed co-optimization and embedded adaptation, including attribute analysis, optimization model solving, and result feedback.

[0043] In the above method, step 6 involves dynamically replacing resources in the presentation structure based on the relevance of the enhanced content presentation structure to learning needs, which include personalized progress and interest preferences. If needs change, resources in the presentation structure are dynamically replaced while maintaining overall balance, thus determining the final content presentation structure. Specifically, in step 6, learning needs are dynamically modeled using user historical interaction logs and real-time feedback signals. Personalized progress is inferred jointly from the proportion of completed knowledge points and the error rate trend. The proportion of completed knowledge points is calculated based on the number of knowledge points where the user's answer accuracy exceeds 80%, and the error rate trend is determined using the slope of a sliding window linear regression. Interest preferences are calculated based on a weighted average of resource click frequency and dwell time. Click frequency refers to the number of times a resource is accessed per unit time, and dwell time refers to the average time a user spends on a resource page, with weighting coefficients of 0.4 and 0.6, respectively. The dynamic replacement mechanism employs a sliding buffer strategy, replacing low-relevance resources at a rate not exceeding 20% ​​while maintaining overall difficulty curve and modal balance. Low-relevance resources are defined as resources whose interest preference scores are 0.5 standard deviations below the current sequence mean. During replacement, resources with similar difficulty levels (difference not exceeding 0.2), the same modality type, and the highest interest preference score are selected from the cache resource library for replacement to ensure the overall stability of the sequence structure.

[0044] In the above method, step 7 involves obtaining the operating parameters of the final content presentation structure in the embedded system. These parameters include memory usage and processing speed. The loading order of the presentation structure is adjusted to suit low-power characteristics, resulting in a content organization structure adapted to mobile devices. Specifically, in step 7, memory usage is obtained through static analysis of resource file size and dynamic monitoring of runtime stack peaks. Static analysis is performed when resources are loaded into the database, while dynamic monitoring is conducted in real-time through the embedded operating system's memory management unit. Processing speed is measured based on resource decoding and rendering time. Decoding time refers to the time from reading from the storage medium to decompression in memory, and rendering time refers to the processing time from memory data to screen display. The loading order adjustment employs a layered preloading strategy, placing high-priority resources (such as the first screen content and key knowledge points) in the resident memory area, and sorting the remaining resources according to their usage probability. The delayed loading threshold is set to complete loading within 500 milliseconds after a user operation. The usage probability is calculated based on a Markov chain prediction model, with the state being the current knowledge point and the transition probability being the knowledge point transition frequency in the historical sequence. This strategy significantly reduces peak memory usage, improves system response speed, and ensures smooth operation on low-power embedded devices.

[0045] To verify the effectiveness of this invention, a specific application scenario is given as an example: A second-year junior high school student uses a tablet computer equipped with the method of this invention to learn English grammar, with the teaching objective being to master the present perfect tense. The system first loads 500 multimodal resources related to the present perfect tense from the local cache library, performs step 1 for feature extraction and grouping, resulting in five sets: beginner (120 items), intermediate (180 items), advanced (100 items), reinforcement (60 items), and extension (40 items). Step 2 calculates the knowledge point coverage to be 82%, which is below the 85% threshold. The system then retrieves 30 supplementary resources from the cloud with a topic relevance score greater than 0.9 and a diversity index greater than 0.7, increasing the coverage to 88%. Step 3 assesses the student's cognitive level as lower-intermediate based on their historical error rate (45% error rate for present perfect tense related questions), generating a preliminary presentation structure containing 20 resources with difficulty levels smoothly increasing from 0.4 to 0.7, and each resource lasting less than 90 seconds. Step 4: The standard deviation of the difficulty was 0.35. A reordering mechanism was initiated, and after adjusting the positions of three resources, the standard deviation decreased to 0.28, while all knowledge point dependencies were satisfied. Step 5: After optimization, the covariance between visual complexity and speech speed decreased from 0.15 to 0.05, the speech speed stabilized at 150 syllables per minute, and the number of image objects did not exceed 6. Step 6: Based on students' recent high click frequency (3 times per day) and long dwell time on animation resources (average 120 seconds), two low-relevance text resources were replaced with animation resources. Step 7: The three key resources on the first screen (basic structure of the present perfect tense, common verb conjugations, and typical example sentences) were placed in the resident memory area, and the remaining resources were sorted according to usage probability to ensure loading within 500 milliseconds when scrolling through pages. The entire process took 4.2 seconds on an 800 MHz ARM Cortex-A7 processor, with a peak memory usage of 180 megabytes, significantly better than existing static arrangement schemes.

[0046] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal optimization method for English teaching software based on embedded systems, characterized in that: The specific steps include the following: Step 1: Extract feature parameters of multimodal content resources, including text density, speech rhythm and image resolution. Use data grouping method to group the feature parameters to obtain a classified set of content resources. Step 2: Match the categorized content resource set with the preset teaching objectives, which include knowledge point coverage and interactivity requirements. If the matching degree is higher than the preset threshold, the set is retained; otherwise, supplementary resources are obtained from the content resource library. The supplementary resources meet the requirements of topic relevance and diversity, thus determining the optimized content resource set. Step 3: Obtain the difficulty level and duration distribution attributes of each resource in the optimized content resource set. The attributes are evaluated based on the learner's cognitive level and the content duration. The initial content presentation structure is constructed by using a sequence generation method combined with the principle of progressive learning. Step 4: For the initial presentation structure, the distribution of difficulty levels reflects the gradient of resources from simple to complex. If the distribution is uneven, the adjacent resource positions are adjusted and reordered. The logical coherence of the presentation structure is judged to ensure smooth connection of knowledge points. Step 5: Analyze the visual complexity and speech speed attributes from the reordered presentation structure. These attributes involve the density of image elements and the rate of speech unit. Use an optimization model to adjust the combination of these attributes to obtain the enhanced content presentation structure. Step 6: Based on the relevance between the enhanced content presentation structure and learning needs, including personalized progress and interest preferences, if the needs change, the resources in the presentation structure are dynamically replaced, and the replacement maintains overall balance to determine the final content presentation structure. Step 7: Obtain the operating parameters of the final content presentation structure in the embedded system. The parameters include memory usage and processing speed. Adjust the loading order of the presentation structure for low power consumption characteristics to obtain a content organization structure adapted to mobile devices.

2. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 1, text density is calculated by weighting the number of effective words per unit area with syntactic complexity, speech rhythm is quantified based on syllable interval time and stress distribution entropy, and image resolution is comprehensively evaluated by combining pixel density and edge gradient information. After normalization, the three are input into the K-means clustering algorithm, with the number of clusters set to 5, which is used to divide the content resource sets into five categories: primary, intermediate, advanced, enhanced, and extended.

3. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 2, the knowledge point coverage is calculated by comparing the node matching rate between the resource tags and the knowledge graph of the curriculum standard. The interactivity requirement is determined based on the number of embedded interactive elements in the resource and the response latency threshold. The matching threshold is set to 85%. When it is lower than this threshold, the system retrieves supplementary resources from the cloud content resource library with a topic relevance score greater than 0.9 and a diversity index greater than 0.7 and injects them.

4. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 3, the difficulty level is determined by the frequency of words, the average sentence length, and the depth of the syntax tree. The duration distribution attribute limits the duration of a single resource to no more than 120 seconds. The sequence generation method adopts a path planning algorithm based on Markov decision process. The state space is the knowledge point node, the action space is the resource selection, and the reward function integrates the smoothness of the difficulty gradient and the balance of duration to ensure that the initial presentation structure conforms to the cognitive law of gradual deepening.

5. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 4, the uniformity of the difficulty level distribution is determined by calculating the standard deviation of the difficulty difference between adjacent resources. If the standard deviation is greater than 0.3, a local rearrangement mechanism is initiated, and resource positions are exchanged within a window length of 3 using a sliding window. Logical coherence is verified through a knowledge point dependency graph. The resource position is only allowed to be fixed when all the preceding dependency nodes of the subsequent knowledge point have been covered.

6. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 5, visual complexity is calculated by combining the number of objects in the image, color saturation variance, and texture energy. Speech speed is quantified in terms of syllables per minute. The optimization model is a multi-objective constraint satisfaction problem solver. The objective function minimizes the covariance between visual complexity and speech speed. The constraints are that the speech speed is not less than 120 syllables per minute and not more than 200 syllables per minute, and the number of image objects does not exceed 8.

7. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 6, learning needs are dynamically modeled using user historical interaction logs and real-time feedback signals. Personalized progress is inferred from the combined proportion of completed knowledge points and error rate trends. Interest preferences are calculated based on a weighted average of resource click frequency and dwell time. The dynamic replacement mechanism adopts a sliding buffer strategy, replacing low-relevance resources at a rate not exceeding 20% ​​while maintaining the overall difficulty curve and modal balance.

8. The multimodal optimization method for English teaching software based on an embedded system according to claim 1, characterized in that: In step 7, memory usage is obtained by statically analyzing the size of resource files and dynamically monitoring runtime stack peaks. Processing speed is based on the measurement of resource decoding and rendering time. Loading order adjustment adopts a layered preloading strategy, placing high-priority resources in the resident memory area and sorting the remaining resources according to their usage probability. The delayed loading threshold is set to complete loading within 500 milliseconds after the user operation.

9. The multimodal optimization method for English teaching software based on an embedded system according to claim 2, characterized in that: The text density calculation module occupies less than 2 megabytes of memory, the speech rhythm analysis module uses fixed-point arithmetic, and the processing time for a single audio segment is less than 800 milliseconds. The image resolution evaluation module is based on edge detection and region segmentation, and the single-frame processing latency is less than 30 milliseconds. The entire system can run stably on an ARM Cortex-A7 processor with a main frequency of 800 MHz.

10. The multimodal optimization method for English teaching software based on an embedded system according to claim 3, characterized in that: The content resource library contains more than 50,000 labeled resources. Each resource is associated with knowledge point tags, difficulty level, modal attributes, and teaching objective mapping relationships. The system supports offline caching of commonly used resource sets, and the cache capacity can be configured from 512 megabytes to 2 gigabytes.