Multi-modal interactive education content generation method and system

By constructing a time-scale model and optimizing algorithms to adjust resource allocation, the problem of uneven resource distribution in the generation of multimodal educational content was solved, achieving an efficient and stable generation process and improving the quality and fluency of educational content.

CN121366064APending Publication Date: 2026-01-20ZHONGKE HAOBO INTERNATIONAL EDUCATION TECHNOLOGY (BEIJING) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511502766.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-01-20

AI Technical Summary

Technical Problem

Existing multimodal educational content generation methods suffer from low generation efficiency when resources are unevenly distributed, making it difficult to meet the real-time needs of large-scale and diverse educational scenarios. Furthermore, the generation process is not flexible enough, affecting the fluency and consistency of content output.

Method used

Time series forecasting methods are used to process multimodal generated data, a time-scale model is constructed, resource allocation is adjusted through probabilistic simulation and optimization algorithms, a backup resource pool is activated, and bottlenecks are identified by combining feedback loops and cluster analysis to optimize the generation process.

Benefits of technology

It improves the efficiency and stability of resource allocation for the generation of multimodal educational content, ensures the efficiency and consistency of the generation process, and reduces delays and resource waste.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121366064A_ABST
    Figure CN121366064A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent education, and discloses a multi-modal interactive education content generation method and system. The method comprises the steps of obtaining multi-modal generation data, performing time sequence prediction processing and constructing a time scale model, and determining a modal dynamic value; simulating a resource allocation scene by adopting a probability simulation method based on the modal dynamic value to obtain a cost increasing rule, and adjusting a resource allocation scheme through an optimization algorithm; the time scale model is updated through a feedback cycle mechanism, and the generation time estimation precision is improved; acquiring education content input in real time, judging a resource demand, and if the resource demand exceeds a threshold value, activating the standby resource pool to perform cost control; according to the method, bottleneck points are identified through a clustering analysis method, efficiency improvement indexes are extracted, the multi-modal fusion stability is judged in combination with a sequence modeling method, a resource optimization path is generated, and the resource allocation efficiency and stability of multi-modal education content generation are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent education, in particular to a multi-modal interactive education content generation method and system. BACKGROUND

[0002] Multi-modal education content generation is an important field of artificial intelligence and education technology integration. By integrating various forms of data such as text, images, and audio, it provides learners with a rich and interactive learning experience. This field is crucial for improving the attractiveness and personalized learning effect of education content, especially in online education, virtual classrooms, and other scenarios, and has irreplaceable value. However, existing methods often face significant technical challenges in actual application, especially in terms of resource allocation and generation efficiency, making it difficult to meet the real-time needs of large-scale and diverse education scenarios.

[0003] Current multi-modal content generation methods often face low efficiency due to uneven resource allocation when processing multiple data sources. For example, when generating a teaching video containing text explanation, images, and background music, the system may delay audio generation due to excessive time-consuming image processing, or cause overall process lag due to excessive computing resources occupied by text generation. This imbalance not only increases the generation cost, but also may affect the smoothness and consistency of content output. Many existing solutions attempt to address this issue by pre-setting resource allocation rules, but these rules often cannot adapt to the dynamic changes in resource demand during the generation process of different modal data, resulting in inflexible generation processes.

[0004] The core technical difficulty lies in accurately capturing and balancing the dynamic changes in resource consumption during the generation process of multi-modal data. The resource requirements of each modality's data (such as semantic processing of text, rendering of images, and synthesis of audio) fluctuate over time and task complexity during generation. For example, when generating a mathematical teaching animation, the system may require a large amount of video memory when rendering complex geometric figures, while requiring more processor resources when generating explanatory speech. This dynamic nature of resource requirements makes it difficult for the system to predict and allocate appropriate resources in advance, resulting in slow generation or quality degradation of some modalities, affecting the overall presentation effect of the final content.

[0005] Therefore, how to dynamically adjust the allocation strategy according to the real-time resource requirements of different modalities during multi-modal content generation to ensure the balance between generation efficiency and content quality has become a key problem in the field of education content generation. SUMMARY

[0006] The present application provides a multi-modal interactive education content generation method and system to improve the resource allocation efficiency and stability of multi-modal education content generation.

[0007] In a first aspect, the present application provides a multi-modal interactive educational content generation method, which comprises: S1, acquiring multi-modal generation data, processing the data to obtain resource path representation by using a time series prediction method, constructing a time scale model according to the resource path representation, determining a modal dynamic value based on the output of the time scale model; S2, simulating a resource allocation scenario by a probability simulation method according to the modal dynamic value to obtain a cost increment law, and if the cost increment law shows that the resource allocation is uneven, adjusting a cross-modal coordination mechanism by an optimization algorithm to obtain an optimized resource allocation scheme; S3, extracting key path data from the resource allocation scheme, updating the time scale model by using a feedback loop mechanism, and determining the generation time estimation accuracy; S4, acquiring real-time educational content generation input, and judging whether the resource demand characteristics exceed a preset threshold, and if so, activating a standby resource pool to obtain a cost control output; S5, integrating monitoring logs according to the cost control output, identifying bottleneck points by using a clustering analysis method, and obtaining efficiency improvement indicators according to the bottleneck points; S6, extracting associated features from the efficiency improvement indicators, iteratively processing the associated features by using a recurrent application sequence modeling method, judging the multi-modal fusion stability according to the processed associated features, and obtaining an optimized resource path.

[0008] In a second aspect, the present application provides a multi-modal interactive educational content generation system, which comprises: A data acquisition and modeling module is configured to acquire multi-modal generation data, process the data to obtain resource path representation by using a time series prediction method, construct a time scale model according to the resource path representation, and determine a modal dynamic value based on the output of the time scale model; A simulation and optimization module is configured to simulate a resource allocation scenario by a probability simulation method according to the modal dynamic value to obtain a cost increment law, and if the cost increment law shows that the resource allocation is uneven, adjust a cross-modal coordination mechanism by an optimization algorithm to obtain an optimized resource allocation scheme; A path extraction and updating module is configured to extract key path data from the resource allocation scheme, update the time scale model by using a feedback loop mechanism, and determine the generation time estimation accuracy; A demand judgment and scheduling module is configured to acquire real-time educational content generation input, judge whether the resource demand characteristics exceed a preset threshold, and if so, activate a standby resource pool to obtain a cost control output; A bottleneck identification and analysis module is configured to identify bottleneck points by using a clustering analysis method according to the cost control output integrated monitoring log, and obtain an efficiency improvement index according to the bottleneck points. A feature iteration and optimization module is configured to extract associated features from the efficiency improvement index, iteratively process the associated features by using a cyclic application sequence modeling method, and determine a resource optimization path according to the processed associated features based on multi-modal fusion stability.

[0009] In the technical solution provided in the present application, an intelligent optimization method is proposed for the resource allocation problem in the multi-modal education content generation process, which improves the generation efficiency and content quality. The time series prediction method is used to process multi-modal generation data, and the resource demand changes of different modes are captured in real time. The time scale model is constructed to dynamically adjust the resource allocation, ensuring the resource collaborative optimization between modes. Secondly, based on the probability simulation method, different resource allocation scenarios are simulated, the cost increment law is identified, and the cross-modal collaborative mechanism is adjusted through the optimization algorithm to ensure the balance of resource allocation and reduce the generation cost. The feedback loop mechanism further optimizes the generation time estimation accuracy, improves the predictability of the generation process, and reduces the generation delay. The system judges the resource demand characteristics in real time, and if it exceeds the preset threshold, the standby resource pool is activated to ensure the stability and continuity of the generation process. In addition, the clustering analysis method is used to integrate the monitoring log, identify the bottleneck points and calculate the efficiency improvement index, further optimizing the system performance. The sequence modeling method is used to determine the stability of multi-modal fusion, ensuring the balance of resource allocation, and generating stable and optimized education content. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.

[0011] Figure 1 A flowchart of a multi-modal interactive education content generation method of the present application; Figure 2 A flowchart of a multi-modal resource balanced allocation method based on Monte Carlo simulation and genetic algorithm of the present application; Figure 3 An education video generation resource usage time sequence monitoring of the present application; Figure 4 A structural schematic diagram of a multi-modal interactive education content generation system of the present application. DETAILED DESCRIPTION

[0012] The embodiments of the present application provide a multi-modal interactive education content generation method and system. The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" or "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0013] For ease of understanding, the specific process of the embodiments of the present application is described below. Please refer to Figure 1 One embodiment of the multi-modal interactive education content generation method in the embodiments of the present application includes: Step S1, acquiring multi-modal generation data, processing the data to obtain resource path representation by using a time series prediction method, and constructing a time scale model according to the resource path representation, and determining a modal dynamic value based on the output of the time scale model.

[0014] Specifically, in the process of acquiring multi-modal generation data, historical generation data including text, image and audio and other modalities need to be collected, which covers detailed records of resource consumption. In order to ensure the accuracy of data processing, time series prediction method is used to analyze these multi-modal data to obtain resource path representation. The resource path representation reflects the resource consumption mode and flow trajectory of each modality in the generation process. Through time series prediction, the trend and periodic changes of resource consumption of different modalities can be captured, especially the change law of resource consumption value of each modality with time in the generation process. Based on these prediction results, a time scale model is further constructed, which uses technologies such as recurrent neural network (RNN) to capture the time dependence in the resource path representation and generate dynamic changes of each modality at different time points.

[0015] Specifically, the output of the model can accurately reflect the fluctuation amplitude and frequency of resource consumption of each modality through nonlinear feature extraction, and eliminate noise interference through smoothing processing to obtain stable modal dynamic values. These modal dynamic values provide necessary data basis for resource allocation optimization, which can effectively support real-time decision and cost control, so as to optimize the use of resources in multi-modal generation task and ensure the efficiency and controllability of the generation process.

[0016] Step S2, simulate the cost increment law by a probability simulation method according to the modal dynamic value, and if the cost increment law shows that the resource allocation is uneven, adjust the cross-modal coordination mechanism by an optimization algorithm to obtain an optimized resource allocation scheme.

[0017] Specifically, by the probability simulation method based on the modal dynamic value, the generated scenarios under different resource allocation strategies can be simulated to reveal the cost increment law. This process simulates multiple resource allocation scenarios by means of Monte Carlo method and other probability simulation methods, and records the resource consumption behavior of each modal in each simulation process. According to these simulation results, the peak fluctuation of resource allocation can be extracted to calculate the growth mode of resource consumption, i.e. the cost increment law. This law reflects the growth rate and inflection point of resource allocation, and the stability and volatility indicators of resource allocation are obtained through statistical analysis. If the simulation results show that there is significant resource allocation unevenness in the cost increment law, i.e. the resource consumption of some modal is significantly higher than that of other modal, the optimization algorithm needs to be adjusted.

[0018] Specifically, the optimization algorithm adjusts the cross-modal coordination mechanism, usually uses intelligent optimization techniques such as genetic algorithm, and adjusts the resource allocation proportion of each modal through iteration to ensure that the resource allocation among modal is more balanced. This process continuously corrects the resource allocation weight and priority of each modal by the optimization model to obtain an optimized resource allocation scheme, reduces the fluctuation of resource consumption, improves the efficiency and stability of the overall generation process, and thus ensures the rational use of resources and avoids waste of resources.

[0019] Step S3, extract key path data from the resource allocation scheme, and update the time scale model using a feedback loop mechanism to determine the generation time estimation accuracy.

[0020] Specifically, the key path data is extracted from the optimized resource allocation scheme, including the core time nodes and consumption of each modal resource allocation. These data reveal the key time points in the generation process and the resource usage of each stage. The extracted key path data is input into the time scale model as the input of the feedback loop mechanism for iterative update. The feedback loop mechanism is based on optimization methods such as gradient descent method, which continuously iterates and adjusts the model parameters to make the model's prediction of the actual generation process more accurate. In each feedback iteration, the deviation between the estimated value of the generation time and the actual time consumption is calculated to optimize the prediction ability of the model.

[0021] Specifically, a predicted value of generation time is calculated and compared with the actual generation time consumption to obtain a deviation value. By analyzing the deviation value, the generation time estimation accuracy is quantified, and the model parameters are further adjusted to improve the accuracy of time estimation. The key to this process is to continuously optimize the time scale model to more accurately predict the time required for the generation task, effectively control the time and resource consumption in the generation process, and ensure that the generation task can be completed efficiently within the specified time.

[0022] Step S4, obtain real-time education content generation input, and determine whether the resource demand characteristics exceed the preset threshold. If they exceed, activate the standby resource pool to obtain the cost control output.

[0023] Specifically, when obtaining real-time education content generation input, it is necessary to analyze the content structure of user queries and teaching materials, including text, images, audio and other multi-modal data. The resource demand characteristics of these inputs mainly include computational complexity and storage requirements. Computational complexity can be quantified by analyzing the number of keywords in user queries and the data volume in materials, and storage requirements can be determined by estimating the cache size of generated content. For example, when processing inputs containing a large number of image or audio materials, the computational complexity and storage requirements are high. Next, determine whether these resource demand characteristics exceed the preset threshold. The preset threshold is usually set based on historical generation data, such as a computational complexity threshold of 1000 operation cycles and a storage requirement threshold of 500MB. If the computational complexity or storage requirement exceeds these thresholds, the activation mechanism of the standby resource pool is triggered.

[0024] The standby resource pool contains pre-allocated computing and storage resources, usually formed by dividing part of the server cluster capacity, including idle computing nodes and storage units. After activating the standby resource pool, the system will dynamically adjust resource allocation according to real-time conditions. For example, if the computational complexity exceeds the threshold, additional computing resources will be called from the standby pool, and if the storage requirement exceeds, the standby storage unit will be called. This adjustment ensures the continuity of the generation process and prevents generation delays or interruptions due to insufficient resources. Finally, after adjustment, the system outputs the cost control results, which include the adjusted resource allocation ratio and generation priority, to effectively control the cost of the generation process, ensure efficient use of resources, and avoid unnecessary resource waste.

[0025] Step S5, integrate monitoring logs according to the cost control output and use clustering analysis method to identify bottleneck points, and obtain efficiency improvement indicators according to the bottleneck points.

[0026] Specifically, the monitoring logs during the generation process are outputted according to the cost control, which record the usage of each modal resource and the timestamp during the generation process, including the computational resource consumption, storage usage and generation delay of each task. The integrated monitoring logs provide detailed running data for subsequent analysis, which is used to identify possible bottleneck points. The monitoring logs are processed using clustering analysis methods, and the commonly used method is K-means clustering algorithm. In this step, the parameters of K-means algorithm are first initialized, and the appropriate number of cluster centers K is selected, for example, K can be set to 3, corresponding to low, medium and high resource usage levels. Feature vectors are extracted from the monitoring logs, including resource consumption data and generation delay time of each timestamp. K-means algorithm assigns data points to different clusters according to Euclidean distance, and iteratively updates the cluster center until the difference between data points in the cluster is minimized. In this way, the clustering algorithm can identify different resource consumption patterns in the logs and help locate parts that may have performance bottlenecks. According to the results of clustering analysis, the bottleneck points in the generation process can be identified, which may be manifested as peak periods of resource usage or nodes with long generation delay. Further, by calculating the efficiency improvement indicators, the impact of these bottleneck points on the overall generation process can be quantified. Efficiency improvement indicators usually include resource utilization and generation delay quantification values. Resource utilization is the ratio of total actual resource usage to total available resource, while generation delay is the deviation of actual generation time from expected time. Through the analysis of these indicators, the efficiency of the generation process can be evaluated, and it can be identified which parts need to be optimized to improve the balance of resource allocation and generation efficiency, so as to achieve the goals of cost control and performance improvement.

[0027] Step S6, extract the associated features from the efficiency improvement indicators, and iteratively process the associated features using the recurrent application sequence modeling method. According to the processed associated features, the resource optimization path is determined according to the multi-modal fusion stability.

[0028] Specifically, the associated features are extracted from the efficiency improvement indicators, which may include the frequency of bottleneck points, the peak of resource consumption, and the generation delay, etc. After extracting these features, a recurrent application sequence modeling method, especially the long short-term memory network (LSTM), is used for processing. The LSTM network can effectively capture the long-term dependencies in the time series through its forget gate, input gate, and output gate mechanisms, thereby extracting the dynamic patterns of resource consumption fluctuations. These patterns reveal the mutual influence of resource consumption between different modalities. In each iteration, the LSTM continuously updates and adjusts its state to ensure that the extracted features accurately reflect the relationship between resource allocation and bottleneck points among modalities. By analyzing these processed associated features, the stability of multi-modal fusion can be determined next. The stability evaluation is achieved by calculating the balance degree of resource allocation among modalities and the consistency of generation results. If both the balance degree and the consistency exceed the preset threshold, it means that the fusion process is stable. According to the stability analysis result, a resource optimization path is generated. The path includes the optimized resource allocation strategy and the generation priority, which is used to guide the continuous improvement of resource management, ensuring more balanced resource allocation, effectively improving generation efficiency, and reducing resource waste.

[0029] It can be understood that the execution subject of the present application can be a multi-modal interactive education content generation system, and can also be a terminal or a server, which is not limited here. The server is taken as an example for illustration in the embodiments of the present application.

[0030] In a specific embodiment, the process of executing step S1 can specifically include the following steps: Collect historical multi-modal generation data including text, images, and audio, and extract resource consumption records for each modality of the data; Analyze the resource consumption records using a time series prediction method to determine the resource consumption sequence of each modality in the generation process; Generate resource path representation based on the resource consumption sequence, which reflects the flow trajectory and time distribution characteristics of each modality resource in the generation process; Standardize the resource path representation to obtain a normalized resource path representation, wherein the standardization process includes uniformly formatting the timestamps and consumption amounts of the resource consumption sequence; Obtain the resource path representation, extract non-linear features from the resource path representation using a sequence modeling method, and construct a time scale model, which includes a multi-layer recurrent neural network for capturing time-dependent relationships in the resource path representation; According to the output of the time scale model, the change intensity of each modality at different time points is calculated, and the modality dynamic value is determined, including the fluctuation amplitude and frequency of resource consumption of each modality. The modality dynamic value is smoothed to eliminate noise interference, and the modality dynamic value is obtained.

[0031] Specifically, historical multi-modal generation data is collected, covering three modalities of text, image and audio. These data record the resource consumption of each modality during the generation process, including the usage of computing resources and storage resources. For the data of each modality, resource consumption records are extracted respectively to ensure that the resource usage characteristics of each modality can be reflected in detail. Time series prediction methods such as ARIMA model are used to analyze these resource consumption records, thereby determining the resource consumption sequence of each modality during the generation process. By analyzing these sequences, the resource consumption pattern of each modality on the time axis can be obtained.

[0032] According to these resource consumption sequences, corresponding resource path representations are generated. The resource path represents the trajectory and time distribution characteristics of resource flow of each modality during the generation process, reflecting the change trend of resource consumption of different modalities. In order to ensure the comparability of data between different modalities, standardization processing is performed on these resource path representations, mainly including unified timestamp format and consumption amount unit, to ensure the consistency and standardization of data.

[0033] After obtaining the standardized resource path representation, a sequence modeling method is used for nonlinear feature extraction to construct a time scale model. The model is based on a multi-layer recurrent neural network (RNN) and is used to capture the time-dependent relationship in the resource path representation, thereby revealing the dynamic change law of resource consumption of each modality. According to the output of the time scale model, the change intensity of each modality at different time points is calculated to obtain the modality dynamic value, including the fluctuation amplitude and frequency of resource consumption. The modality dynamic value is smoothed to eliminate noise interference, and stable modality dynamic value is obtained, providing a reliable basis for resource optimization and decision-making.

[0034] In a specific embodiment, the process of step S2 can specifically include the following steps: Obtain the modality dynamic value, and for the multi-modal fusion scene, construct a probability simulation model, which simulates the consumption behavior of each modality under different resource allocation strategies through the Monte Carlo method to determine the peak fluctuation of resource allocation; According to the peak fluctuation, the cost increment law of resource usage is calculated, including the growth rate and inflection point of resource allocation; Statistical analysis is performed on the cost increment law to extract mean and variance features, and stability and volatility indicators of resource allocation are obtained to determine whether the resource allocation is uniform; An increasing cost law is obtained, and it is determined whether the increasing cost law exceeds a preset resource allocation balance threshold; If it exceeds, an optimization model of a cross-modal collaborative mechanism is constructed, the optimization model is based on a genetic algorithm, iteratively adjusts the resource allocation proportion of each modality, determines an optimal resource allocation scheme, and the resource allocation scheme includes the resource allocation weight and priority of each modality; The resource allocation scheme is verified to ensure the collaboration balance between modalities, the verified resource allocation scheme is evaluated and adjusted through a simulation experiment, and an optimized resource allocation scheme is obtained, which reduces the fluctuation amplitude of resource allocation.

[0035] Specifically, dynamic values of each modality are obtained, which represent the resource consumption fluctuation amplitude and frequency of different modalities in the generation process, and reflect the change trend of resource consumption. Based on these dynamic values, a probability simulation model is constructed for a multi-modal fusion scene, which simulates the consumption behavior of each modality under different resource allocation strategies through the Monte Carlo method. In the simulation process, the Monte Carlo method generates resource allocation scenes through multiple random sampling, simulates resource consumption under different strategies, and then determines the peak fluctuation of resource allocation. According to the peak fluctuation, the cost increasing law of resource use can be calculated, which usually includes the growth rate and inflection point. The growth rate can be expressed by a linear regression formula as follows:

[0036] wherein, is the growth rate, and are the resource consumption amounts at time and time , respectively, is the time interval. The inflection point is identified by calculating the second derivative of the change in resource consumption, that is, when the sign of the second derivative changes, it indicates that the growth rate of resource consumption has changed.

[0037] Statistical analysis is performed on the cost increasing law, and mean and variance features are extracted. For example, in one simulation, the resource consumption data is [10, 12, 15, 18, 22] units, the calculated mean is 15.4 units, and the variance is 17.04 units. According to these statistical characteristics, the stability and volatility indicators of resource allocation can be calculated. If the variance is large, it indicates that there is a large fluctuation in resource allocation, indicating that further optimization is needed. On this basis, if the cost increasing law exceeds the preset balance threshold, for example, the threshold is set to 0.1 units / time, an optimization model based on a genetic algorithm will be constructed.

[0038] For example, assume that the current resource allocation ratio is text 50%, image 30%, and audio 20%, and the calculation result shows that the resource consumption of image and audio fluctuates greatly, with variances of 20 and 18, respectively, exceeding the preset fluctuation threshold. In this case, the genetic algorithm will adjust these ratios through iteration, and assume that after 10 generations of iteration, the optimization scheme adjusts the text resource ratio to 40%, the image resource ratio to 40%, and the audio resource ratio to 20%. After simulation verification, the optimization scheme is further adjusted, and the simulation result shows that the optimized resource allocation scheme reduces the resource fluctuation in the generation process by 15%, ensuring the stability and efficiency of the generation task. Figure 2 The figure shows a multi-modal resource equalization allocation process based on Monte Carlo simulation and genetic algorithm.

[0039] In a specific embodiment, the process of performing step S3 can specifically include the following steps: Obtain the optimized resource allocation scheme, extract the critical path data, and the critical path data includes the core time nodes and consumption of each modal resource allocation; Iteratively update the time scale model using a feedback loop mechanism, which adjusts the model parameters based on the gradient descent method; For the updated time scale model, calculate the generation time estimation accuracy, which is determined by comparing the deviation of the predicted time consumption and the actual time consumption; Analyze the deviation to obtain the error distribution of the generation time estimation, which is used for cyclic optimization of the generation process.

[0040] Specifically, first, obtain the optimized resource allocation scheme and extract the critical path data, which includes the core time nodes and resource consumption of each modal in the generation process, providing a basis for generation time estimation. For example, assume that in the optimization scheme, the core time node of the text modal is the 3rd time point, and the resource consumption is 30 units, and the core time node of the image modal is the 4th time point, and the consumption is 40 units. Next, iteratively update the time scale model using a feedback loop mechanism, which adjusts the model parameters based on the gradient descent method. In each iteration, the gradient descent method calculates the gradient of the loss function to gradually adjust the model parameters to reduce the prediction error.

[0041] Assume that the loss function is defined as the difference between the predicted time consumption and the actual time consumption, and the formula is:

[0042] wherein, is the loss function, is the number of samples, is the predicted time consumption of the th sample, The actual time consumption for the first The actual time consumption for the first

[0043] wherein, is the actual generation time, is the predicted generation time. By analyzing the deviation, the error distribution of the generation time estimation can be obtained, which helps to optimize the generation process. If the error is large, resource allocation or optimization strategies may need to be adjusted.

[0044] Specifically, the feedback loop mechanism is an iterative process that optimizes the time scale model by repeatedly inputting critical path data. Gradient descent, as an optimization algorithm, adjusts the model's parameters step by-step by calculating the gradient of the loss function, thereby improving the accuracy of the prediction. In this process, the time scale model handles the nonlinear characteristics of the modal dynamic values, and the feedback loop mechanism iteratively optimizes multiple times based on the extracted critical path data. Initially, the learning rate of the model is set to 0.01, and the loss between the current prediction value and the actual data is calculated, and the model parameters are updated according to the gradient descent method. The update process is completed by subtracting the product of the learning rate and the gradient from the current parameter value, with the goal of reducing the prediction error. As the iteration progresses, the model is continuously optimized until the loss converges, for example, after 100 iterations, the loss drops below 0.001, and the model's prediction accuracy is significantly improved.

[0045] In the multi-modal fusion scenario, the first iteration may mainly adjust the resource allocation parameters of the text and image modalities to reduce the dynamic value deviation of the audio modality, which helps to improve the balance of resource allocation. The second iteration focuses on optimizing the peak fluctuation in resource allocation, further refining the model parameters, thereby optimizing the incremental cost law, making the generation process more efficient and stable. In order to adapt to changes in the real-time generation process, such as input exceeding the preset threshold, the learning rate may be adjusted to 0.005, thereby slowing down the iteration speed but improving the model accuracy. Through multiple iterations, the model can better reflect the resource path representation, further enhance the cross-modal collaboration mechanism, reduce the bottleneck points in the generation process, and improve the generation efficiency.

[0046] In a specific embodiment, the process of performing step S4 can specifically include the following steps: Obtain real-time education content generation input, the input includes user query and teaching materials, extract resource demand features of the input, the features include computational complexity and storage requirements; determining whether the resource requirement features exceed preset resource thresholds; if so, activating a backup resource pool, the backup resource pool including pre-allocated computing and storage resources; adjusting resource allocation of the real-time generation process according to the activation state of the backup resource pool, to obtain a cost control output, the cost control output including an adjusted resource allocation ratio and a generation priority.

[0047] Specifically, when obtaining the real-time education content generation input, including user queries and teaching materials, the system first extracts the resource requirement features of the input, including the computing amount and storage requirements. For example, when the user query is "historical event explanation" and relevant text and picture materials are provided, the system analyzes the complexity of the text and the size of the picture. Assuming that the text content computing amount is 500 computing periods, the storage requirement is 346.4MB, and the picture material computing amount is 200 computing periods, the storage requirement is 492MB. Through comprehensive evaluation, the total computing amount is calculated as 700 periods, and the storage requirement is 752.43MB.

[0048] Next, the system determines whether these resource requirement features exceed the preset resource thresholds. Assuming that the preset computing amount threshold is 600 computing periods and the storage requirement threshold is 700MB. In this example, the total computing amount is 700 computing periods, which exceeds the preset threshold, and the storage requirement is 752.43MB, which also exceeds the storage threshold. Therefore, the system decides to activate the backup resource pool.

[0049] The backup resource pool includes pre-allocated computing and storage resources, and there are 200 computing periods and 300MB of storage resources available for scheduling in the pool. The system adjusts the resource allocation of the real-time generation process according to the activation state of the backup resource pool. Specifically, the system will allocate computing resources from the main resource pool and the backup resource pool in proportion to ensure that the computing task can be completed in time. For example, assuming that the adjustment ratio of computing resources is 70% from the main resource pool and 30% from the backup resource pool, and the adjustment ratio of storage resources is 60% from the main resource pool and 40% from the backup resource pool. After adjustment, the generation priority and resource allocation ratio are as follows: the priority of the computing task is high, ensuring that the user query and material generation are processed in priority, and the storage resource is allocated to the picture material in priority to ensure its efficient storage. The system outputs a cost control scheme according to the adjusted resource allocation ratio and generation priority, ensuring that the resource utilization in the generation process is maximized while avoiding exceeding the budget, and improving the overall generation efficiency.

[0050] In a specific embodiment, the process of performing step S5 can specifically include the following steps: obtaining a cost control output, integrating a generation process monitoring log, the monitoring log including resource usage records and generation timestamps of each modality; The monitoring logs are processed using a clustering analysis method, which is based on the K-means algorithm, to identify bottleneck points in the generation process, including resource usage peaks and generation delay nodes; According to the bottleneck points, efficiency improvement indicators are calculated, including resource utilization and quantitative values of generation delay; The efficiency improvement indicators are sorted to determine the generation link that needs to be improved first.

[0051] Specifically, the cost control output is obtained, and the monitoring logs of the generation process are integrated, which record the resource usage of each modality and the generation timestamp, helping to track the real-time status of each generation link. For example, the monitoring logs may include the CPU usage of text generation, the memory consumption of image processing, and the storage usage of audio generation. These data reflect the resource consumption and generation progress of each modality at different time points. Then, the monitoring logs are processed using a clustering analysis method, usually using the K-means algorithm for clustering analysis. By grouping the resource usage records and generation delay nodes at different time points, the K-means algorithm can identify possible bottleneck points in the generation process. Bottleneck points include peak periods of resource usage and nodes with long generation delays, such as CPU usage exceeding 80% in a certain time period, or a generation link with a delay exceeding a predetermined threshold.

[0052] According to the identified bottleneck points, the system further calculates efficiency improvement indicators, including resource utilization and quantitative values of generation delay. For example, resource utilization can be measured by calculating the ratio of the total amount of resources actually used to the total amount of available resources, while generation delay is quantified by calculating the deviation between the actual generation time and the expected time. By sorting these efficiency improvement indicators, the system can determine the generation link that needs to be improved first. If it is found that the resource usage peak and generation delay of image processing are serious, the system will take it as the priority optimization target to reduce the bottleneck in the generation process, optimize the overall performance, and ensure that the generation process is more stable and efficient.

[0053] Taking the educational video generation process as an example, the monitoring log records the resource usage of each modality and the generation timestamp. For example, during the generation process, the text modality uses 50% of the CPU resources, the image processing stage consumes 6 GB of memory, and the audio processing stage occupies 1.2 GB of storage. The timestamp record shows that text generation started at minute 1, image processing reached its maximum resource consumption at minute 4, and audio processing reached its peak generation delay at minute 7. By clustering these monitoring log data using the K-means algorithm, it is found that the resource consumption peak of the image generation stage occurs at minute 4, with a memory usage rate of 95%, while the delay node of audio generation occurs at minute 7, with an actual generation time that is 3 seconds longer than the expected delay. According to these bottleneck points, the system calculates efficiency improvement indicators, including resource utilization, which is the ratio of the actual memory usage of the image generation stage, 6 GB, to the total available memory, 8 GB, resulting in a resource utilization of 75%. The generation delay is the difference between the actual time of audio generation, 12 seconds, and the expected time, 9 seconds, resulting in a delay of 3 seconds. According to the ranking of these efficiency improvement indicators, the system identifies that image generation and audio processing are the processes that need to be improved first. To optimize the generation process, the system increases the memory quota of the image generation stage to 8 GB and optimizes the audio processing algorithm to reduce the generation delay. Reference Figure 3 The figure shows the monitoring of the timing of resource usage for the generation of educational videos.

[0054] In an embodiment, the process of performing step S6 can specifically include the following steps: Obtain efficiency improvement indicators and extract associated features, including the causal relationship between bottleneck points and resource allocation; Use a sequence modeling method to process the associated features in a loop, and the sequence modeling method is based on a long short-term memory network; Determine the stability of multi-modal fusion, which is determined by the balance of resource allocation between modalities and the consistency of generation results; According to the stability analysis, generate a resource optimization path, which includes an adjusted resource allocation strategy and a generation priority, to guide the continuous improvement of resource management.

[0055] Specifically, efficiency improvement indicators are obtained and relevant features are extracted, including the causal relationship between bottleneck points and resource allocation. For example, when generating an educational video, the bottleneck occurs in the image processing stage, and the memory usage reaches 90% at the 4th minute, with resource consumption of 8GB in the image processing stage, while the audio generation stage has a delay at the 7th minute, with an actual generation time of 15 seconds, which is 4 seconds longer than expected. The relevant features include the memory consumption of image processing and the delay of audio generation, indicating that high memory usage may cause delay in subsequent audio generation. Through these features, the system can identify resource consumption bottlenecks and their impact on the overall generation process. A sequence modeling method based on long short-term memory network (LSTM) is used to process the relevant features in a loop. LSTM can capture long-term dependencies in time series, such as when sharing memory resources between image and audio modalities, delays may accumulate over time. After multiple iterations, the LSTM network can more accurately predict the trend of resource consumption in image and audio modalities.

[0056] Stability of multi-modal fusion is determined by evaluating the balance of resource allocation between modalities and the consistency of generation results. For example, the image modality resource occupancy ratio is 57%, the audio modality is 33%, and the text modality is 10%. The system calculates the deviation of resource allocation between image and audio modalities as 30%, which is a large value, indicating that the resource allocation is not balanced. At the same time, by calculating the generation result consistency, assuming that the similarity of image and audio generation results is 0.85, which is lower than the threshold value of 0.9, indicating that the fusion stability is poor. According to the stability analysis, the system generates a resource optimization path. The adjustment scheme includes adjusting the resource allocation of image modality to 45%, audio modality to 45%, and text modality to maintain 10% resource allocation, and optimizing the generation priority, adjusting the generation priority of image and audio to high to ensure a smoother generation process.

[0057] The above describes a multi-modal interactive educational content generation method in an embodiment of the present application, and the following describes a multi-modal interactive educational content generation system in an embodiment of the present application. Please refer to Figure 4 An embodiment of a multi-modal interactive educational content generation system in an embodiment of the present application includes: A data acquisition and modeling module 201 is configured to acquire multi-modal generation data, process the data using a time series prediction method to obtain a resource path representation, and construct a time scale model based on the resource path representation. The output of the time scale model is used to determine the modal dynamic value; A simulation and optimization module 202 is configured to simulate resource allocation scenarios using a probabilistic simulation method based on the modal dynamic value to obtain a cost increment rule. If the cost increment rule shows that the resource allocation is not balanced, an optimization algorithm is used to adjust the cross-modal collaboration mechanism to obtain an optimized resource allocation scheme; The path extraction and update module 203 is configured to extract critical path data from the resource allocation scheme, update the time scale model using a feedback loop mechanism, and determine the accuracy of the generation time estimate. The demand judgment and scheduling module 204 is configured to obtain real-time education content generation input, judge whether the resource demand characteristics exceed the preset threshold, and if so, activate the standby resource pool to obtain the cost control output. The bottleneck identification and analysis module 205 is configured to integrate monitoring logs according to the cost control output, identify bottleneck points using clustering analysis methods, and obtain efficiency improvement indicators according to the bottleneck points. The feature iteration and optimization module 206 is configured to extract associated features from the efficiency improvement indicators, iteratively process the associated features using a cyclic application sequence modeling method, judge the multi-modal fusion stability according to the processed associated features, and obtain the optimized resource path.

[0058] Through the cooperation of the above components, the system can efficiently process and optimize the resource allocation problem in the multi-modal education content generation process. The data acquisition and modeling module 201 uses time series prediction methods to obtain and process multi-modal generation data, constructs resource path representation, and determines modal dynamic values through a time scale model, providing accurate basic data for subsequent resource allocation and optimization. The simulation and optimization module 202 simulates resource allocation scenarios based on modal dynamic values, identifies cost increment rules, and adjusts cross-modal coordination mechanisms according to optimization algorithms to ensure the balance of resource allocation and reduce generation costs. The path extraction and update module 203 extracts critical path data and updates the time scale model through a feedback loop mechanism to improve the accuracy of generation time estimates. The demand judgment and scheduling module 204 judges whether the resource demand characteristics exceed the preset threshold based on real-time input, and if so, activates the standby resource pool to ensure the flexibility and real-time nature of resource scheduling. The bottleneck identification and analysis module 205 integrates monitoring logs using clustering analysis methods to identify bottleneck points in the generation process and calculate efficiency improvement indicators to guide subsequent optimization. The feature iteration and optimization module 206 extracts associated features from the efficiency improvement indicators, uses sequence modeling methods to judge the stability of multi-modal fusion, and generates optimized resource paths to guide continuous improvement of resource management. Overall, the system can effectively improve the efficiency, stability, and resource utilization of the multi-modal education content generation process through the close cooperation of the modules.

[0059] The above-described embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method and system for multi-modal interactive educational content generation, characterized in that, The method comprises: S1, acquiring multi-modal generation data, processing the data by using a time series prediction method to obtain resource path representation, and constructing a time scale model according to the resource path representation, and determining a modal dynamic value based on the output of the time scale model; S2, according to the modal dynamic value, a resource allocation scene is simulated by a probability simulation method to obtain a cost increment rule, if the cost increment rule shows that the resource allocation is uneven, an optimization algorithm is used to adjust the cross-modal coordination mechanism to obtain an optimized resource allocation scheme; S3, extracting key path data from the resource allocation scheme, updating the time scale model by using a feedback loop mechanism, and determining the generation time estimation accuracy; S4, acquiring real-time education content generation input, and judging whether the resource demand characteristics exceed the preset threshold, if yes, activating the standby resource pool to obtain the cost control output; S5, according to the cost control output, integrating the monitoring log, using clustering analysis method to identify the bottleneck point, and according to the bottleneck point, obtaining the efficiency improvement index; S6, extracting the correlation characteristics from the efficiency improvement index, using a cyclic application sequence modeling method to iteratively process the correlation characteristics, and according to the processed correlation characteristics, judging the multi-modal fusion stability to obtain the resource optimization path. 2.The method of claim 1, wherein, In S1, the multi-modal generation data is acquired, and the data is processed by using a time series prediction method to obtain resource path representation, comprising: Collecting historical multi-modal generation data including text, image and audio, extracting resource consumption records for each modality of the data; using a time series prediction method to analyze the resource consumption records, determining the resource consumption sequence of each modality in the generation process; generating resource path representation according to the resource consumption sequence, the resource path representation reflects the flow trajectory and time distribution characteristics of each modality resource in the generation process; standardizing the resource path representation to obtain normalized resource path representation, wherein the standardization processing includes uniformly formatting the timestamp and consumption amount of the resource consumption sequence. 3.The method of claim 2, wherein, In S1, the time scale model is constructed according to the resource path representation to determine the modal dynamic value, comprising: Obtaining the resource path representation, using a sequence modeling method to extract nonlinear features of the resource path representation, constructing a time scale model, the time scale model includes a multi-layer recurrent neural network for capturing the time dependence in the resource path representation; according to the output of the time scale model, the change intensity of each modality at different time points is calculated, and the modal dynamic value is determined, the modal dynamic value includes the resource consumption fluctuation amplitude and frequency of each modality, the modal dynamic value is smoothed to eliminate noise interference, and the modal dynamic value is obtained. 4.The method of claim 3, wherein, In S2, the probability simulation method is used to simulate the resource allocation scene according to the modal dynamic value to obtain the cost increment rule, comprising: The modal dynamic value is acquired, a probability simulation model is constructed for a multi-modal fusion scene, the probability simulation model simulates consumption behaviors of each modality under different resource allocation strategies through a Monte Carlo method, and peak fluctuation conditions of resource allocation are determined; according to the peak fluctuation conditions, a cost increment law of resource use is calculated, the cost increment law includes a growth rate and an inflection point of resource allocation; statistical analysis is performed on the cost increment law, mean and variance characteristics are extracted, and stability and volatility indexes of resource allocation are obtained, which are used to determine whether resource allocation is uniform. 5.The method of claim 4, wherein, In S2, if the cost increment law shows that resource allocation is not uniform, an optimized resource allocation scheme is obtained through an optimization algorithm adjusting a cross-modal collaborative mechanism, including: The cost increment law is acquired, and it is judged whether the cost increment law exceeds a preset resource allocation balance threshold; if yes, an optimization model of the cross-modal collaborative mechanism is constructed, the optimization model is based on a genetic algorithm, iteratively adjusts resource allocation proportions of each modality, and determines an optimal resource allocation scheme, the resource allocation scheme includes resource allocation weights and priorities of each modality; the resource allocation scheme is verified to ensure collaboration balance among the modalities, the verified resource allocation scheme is adjusted through simulation experiments, and an optimized resource allocation scheme is obtained, which reduces fluctuation amplitude of resource allocation. 6.The method of claim 3, wherein, S3 includes: The optimized resource allocation scheme is acquired, and key path data is extracted, the key path data includes core time nodes and consumption amounts of resource allocation of each modality; a feedback loop mechanism is used to iteratively update the time scale model, the feedback loop mechanism adjusts model parameters based on a gradient descent method; for the updated time scale model, time estimation accuracy is calculated, the time estimation accuracy is determined by comparing deviations of predicted time consumption and actual time consumption; the deviations are analyzed to obtain error distribution of time estimation, which is used for cyclic optimization of the generation process. 7.The method of claim 1, wherein, S4 includes: Real-time education content generation input is acquired, the input includes user queries and teaching materials, resource demand characteristics of the input are extracted, the characteristics include calculation amount and storage demand; it is judged whether the resource demand characteristics exceed a preset resource threshold; if yes, a standby resource pool is activated, the standby resource pool includes pre-allocated calculation and storage resources; according to the activation state of the standby resource pool, resource allocation of the real-time generation process is adjusted, and a cost control output is obtained, the cost control output includes adjusted resource allocation proportions and generation priorities. 8.The method of claim 1, wherein, S5 includes: The cost control output is acquired, a monitoring log of the generation process is integrated, the monitoring log includes resource usage records and generation timestamps of each modality, a clustering analysis method is used to process the monitoring log, the clustering analysis method is based on a K-means algorithm, bottleneck points in the generation process are identified, the bottleneck points include resource usage peaks and generation delay nodes, efficiency improvement indicators are calculated according to the bottleneck points, the indicators include quantitative values of resource utilization and generation delay, and the efficiency improvement indicators are sorted to determine a generation link to be improved preferentially. 9.The method of claim 1, wherein, The S6 includes: The efficiency improvement indicators are acquired, and associated features are extracted, the associated features include a causal relationship between the bottleneck points and resource allocation, a sequence modeling method is used to cyclically process the associated features, the sequence modeling method is based on a long short-term memory network, stability of multi-modal fusion is judged, the stability is determined by balance of resource allocation between modalities and consistency of generation results, a resource optimization path is generated according to the stability analysis, the resource optimization path includes an adjusted resource allocation strategy and a generation priority, and is used to guide continuous improvement of resource management.

10. A multi-modal interactive educational content generation system for implementing a multi-modal interactive educational content generation method according to any one of claims 1 to 9, characterized in that, The multi-modal interactive education content generation system includes: A data acquisition and modeling module is configured to acquire multi-modal generation data, process the data to obtain resource path representation by using a time series prediction method, construct a time scale model according to the resource path representation, and determine a modality dynamic value based on an output of the time scale model; A simulation and optimization module is configured to simulate a resource allocation scenario by using a probability simulation method to obtain a cost increment law according to the modality dynamic value, adjust a cross-modality coordination mechanism to obtain an optimized resource allocation scheme by using an optimization algorithm if the cost increment law shows that resource allocation is uneven, and extract key path data from the resource allocation scheme; A path extraction and update module is configured to update the time scale model by using a feedback loop mechanism and determine generation time estimation accuracy based on the key path data; A demand judgment and scheduling module is configured to acquire real-time education content generation input, judge whether resource demand characteristics exceed a preset threshold, and activate a backup resource pool to obtain a cost control output if the resource demand characteristics exceed the preset threshold; A bottleneck identification and analysis module is configured to integrate monitoring logs according to the cost control output, identify bottleneck points by using a clustering analysis method, and obtain efficiency improvement indicators based on the bottleneck points; A feature iteration and optimization module is configured to extract associated features from the efficiency improvement indicators, iteratively process the associated features by using a cyclic sequence modeling method, judge stability of multi-modal fusion based on the processed associated features, and obtain a resource optimization path.

Citation Information

Patent Citations

  • Online resource management system and method based on AI large model

    CN120256143A

  • Multi-modal content generation method and system and storage medium

    CN120316725A

  • Teaching resource dynamic allocation method and system based on education software platform data

    CN120562830A