A multi-modal teaching resource intelligent fusion and distribution method for in-service education

CN122367078BActive Publication Date: 2026-09-18SHANXI YIHEXUE EDUCATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610822698.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-09-18
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

[0005]为了解决单纯依靠进度百分比记录无法跨越异构模态实现语义对齐、且缺乏对隐性工况预判能力导致跨模态切换存在记忆断点的现有技术问题,本申请提供一种面向在职教育的多模态教学资源智能融合与分发方法

Benefits of technology

[0019] Preferably, the method provided in this application further includes: after the learning session ends, capturing the actual drift degree and review behavior feedback data of this cross-modal switching; transmitting the feedback data to the server through a silent upload channel after lightweight encryption; updating the drift sample set of each student on the server in an incremental manner, and recalculating the set threshold for system self-correction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122367078B_ABST
    Figure CN122367078B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of in-service education resource processing, in particular to a multi-modal teaching resource intelligent fusion and distribution method for in-service education, which comprises the following steps: cutting and mounting heterogeneous teaching resources under the knowledge point nodes of a teaching outline, extracting semantic features and mapping to a latent semantic space to construct a semantic skeleton; collecting hardware sensing signals, combining behavior logs to construct a three-dimensional tensor and generating a time period modal environment preference portrait; extracting modal anchor triplets, calculating an alignment score and a cross-modal drift degree when cross-modal switching occurs, and triggering the generation of a bridging segment when the cross-modal drift degree exceeds a set threshold; predicting the expected occurrence time and the target modal type based on the time period modal environment preference portrait, pre-generating the bridging segment and pushing it to an edge node cache. The application effectively eliminates the heterogeneity barrier of different modal resources and significantly improves the memory fault phenomenon of cross-modal continuing learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of in-service education resource processing technology, specifically to a method for intelligent fusion and distribution of multimodal teaching resources for in-service education. Background Technology

[0002] In the context of in-service education, learners' learning processes are often severely fragmented by fragmented time such as commuting, lunch breaks, or short rests. Learning activities require frequent switching between heterogeneous multimodal teaching resources such as videos, audio, and slides. For example, audio may be used while walking, while switching to video in a quiet environment or at night. This dual switching of modality and environment can easily cause severe memory drift and semantic gaps at the learner's cognitive level, greatly disrupting the continuity of learning.

[0003] Existing technologies generally employ simple physical timestamps or coarse progress percentage markers when handling the seamless transition between different modal resources. However, these resources exhibit inherent heterogeneity in information density and distribution, making precise semantic alignment impossible with mere progress recording. Furthermore, existing technologies fail to effectively identify implicit learning preferences among employees during fragmented time periods and lack preparedness for sudden resource switching requests, resulting in common issues such as initial packet delays and abrupt transitions when switching between devices or modalities.

[0004] Therefore, how to achieve precise semantic alignment at the inherent heterogeneity of different modal resources in terms of information density and distribution, and how to solve the problems of first packet delay and awkward connection that are commonly faced when switching between devices or modalities are urgent technical problems to be solved. Summary of the Invention

[0005] To address the existing technical problems that relying solely on progress percentage records cannot achieve semantic alignment across heterogeneous modalities and that the lack of implicit condition prediction capabilities leads to memory breakpoints during cross-modal switching, this application provides a method for intelligent fusion and distribution of multimodal teaching resources for in-service education.

[0006] In the first aspect, this application provides a method for intelligent fusion and distribution of multimodal teaching resources for in-service education, comprising: segmenting heterogeneous teaching resources and attaching them to knowledge point nodes of the teaching syllabus; extracting semantic features from the segments and mapping them to a latent semantic space to construct a semantic skeleton of multimodal knowledge atoms; collecting hardware sensor signals from the student's device and constructing a three-dimensional tensor including time period buckets, modality types, and physical environments in conjunction with learning behavior logs; generating a time period modality environment preference profile based on the three-dimensional tensor; extracting modality anchor triples of the current learning session; calculating the alignment score and cross-modality drift degree between the anchor left by the previous modality and the candidate anchor of the next modality when a cross-modality switch occurs; triggering the generation of bridging segments when the cross-modality drift degree exceeds a set threshold; predicting the expected time of the next switch and the target modality type based on the time period modality environment preference profile; pre-generating the bridging segments and pushing them to edge nodes for caching before the expected time of switch arrives.

[0007] By uniformly segmenting and mapping heterogeneous teaching resources to a latent semantic space, a purely objective semantic positioning foundation is provided for multimodal resources. Combined with the hardware sensing signals of the device, the three-dimensional tensor of learning behavior is reconstructed, which truly captures the stable preference profile of in-service learners under fragmented time periods. A drift compensation mechanism is introduced during cross-modal switching. By bridging segments to align contextual gaps, and by using the prediction of switching time to push the generated segments to edge nodes in advance, learners can obtain semantically continuous opening content in the first packet of cross-modal switching, which significantly improves the memory gap phenomenon of cross-modal learning continuation.

[0008] Preferably, the step of segmenting and attaching heterogeneous teaching resources to the knowledge point nodes of the teaching syllabus involves segmenting, extracting semantic features, and mapping them to the latent semantic space. This includes: segmenting video streams into video segments and extracting salient region vectors and lecturer posture description vectors; segmenting audio resources into audio segments based on silence intervals and extracting speech textification results and spectral peak envelope vectors; segmenting documents and slides according to paragraph boundaries and extracting title, key sentence, and layout structure vectors; and mapping the extracted features to the latent semantic space through a shared embedding network to form a cross-modal feature skeleton.

[0009] Preferably, the hardware sensor signals collected from the student's device are combined with the learning behavior log to construct a three-dimensional tensor including time period buckets, modal types, and physical environment. This includes: distinguishing physical states based on the steady-state of the synthesized modal length from the device's gyroscope and accelerometer; determining the acoustic scene based on the root mean square of the ambient sound power collected by the microphone; determining the illumination state based on the screen brightness sensor and system time; and reconstructing the learning behavior into a three-dimensional tensor, wherein the boundaries of the time period buckets are naturally generated by the trough positions of the probability density curve of the learning start timestamp, and the physical environment is determined by the joint sample clustering of sensor signals.

[0010] Preferably, the step of generating a time-segment modal environment preference profile based on the three-dimensional tensor includes: statistically analyzing the effective learning time and completion rate in each cell of the three-dimensional tensor under the current time period, current modality, and current physical environment; calculating the completion rate based on the ratio of the actual effective playback time to the segment's own duration; and creating a profile of the feature distribution on the three-dimensional tensor to obtain the time-segment modal environment preference profile.

[0011] Preferably, the step of extracting modal anchor triples of the current learning session, and calculating the alignment score and cross-modal drift between the anchors left by the previous modality and the candidate anchors of the next modality when a cross-modal switch occurs, includes: extracting modal anchor triples containing visual anchors, auditory anchors, and semantic anchors; calculating the alignment score using a cross-attention formula that introduces a modal system offset compensation term, wherein the modal system offset compensation term is obtained by projecting the difference between the feature mean vectors of the source modality and the target modality on historical samples along the direction of the candidate anchors; and defining the cross-modal drift as the inverse sign of the cosine similarity between the anchors of the previous modality and the anchors of the next modality in the latent semantic space after alignment.

[0012] By introducing a modal system offset compensation term composed of the difference in the mean of historical sample features into the alignment calculation, the global center distribution differences between the visual, auditory, and text domains are effectively removed, so that the alignment score between cross-modal anchors can accurately represent the degree of pure semantic difference correlation at the atomic level, and avoid the attention results being biased by heterogeneous modalities.

[0013] Preferably, the step of triggering the generation of a bridging segment when the cross-modal drift exceeds a set threshold includes: when the number of individual review samples of a new learner is insufficient, using the upper quartile of the drift sample pool of the same image group as the group prior threshold; when the confidence interval width of the individual drift mean is not greater than the confidence interval width of the group drift mean, using the upper quartile of the individual drift sample as the individual posterior threshold; using linear interpolation to smoothly transition between the group prior threshold and the individual posterior threshold to obtain the set threshold; when the cross-modal drift exceeds the set threshold, calculating the cosine similarity sequence between the feature vector of each sub-segment and the feature vector of the anchor point of the previous modality within the same skeleton segment of the target modality, and cropping outwards from the peak of the cosine similarity sequence to generate the bridging segment.

[0014] The two-stage processing method, which uses linear interpolation to transition from group prior to individual posterior, smoothly navigates the cold start period when new user features are sparse. Furthermore, it relies on the rule of expanding the trough of the objective similarity curve to extract segment boundaries, without depending on any manually preset experience duration, so that the length of the compensated generation naturally converges to the semantically coherent interval.

[0015] Preferably, the step of predicting the expected occurrence time and target modality type of the next switch based on the time-segment modality environment preference profile includes: locating the corresponding cell in the three-dimensional tensor using the index of the current time-segment bucket and the index of the current environment bucket as joint conditions; reading the offset distribution of the occurrence time of historical switch events relative to the start time of the time-segment bucket under the cell; taking the median of the offset distribution as the expected switch offset; adding the expected switch offset to the start time of the current time-segment bucket to obtain the expected occurrence time; and taking the mode of the target modality in the historical switch event records as the target modality type.

[0016] Preferably, the step of pre-generating the bridging segment and pushing it to the edge node for caching before the expected occurrence time arrives includes: calculating the historical average round-trip latency of the target device handshaking with each edge node in historical sessions; calculating the historical average generation time of the bridging segment from triggering to completion; adding the historical average round-trip latency to the average generation time to obtain an advance; and triggering the pre-generation of the bridging segment at the moment when the expected occurrence time is subtracted from the advance.

[0017] By objectively adding the statistical average of historical network round-trip latency to the historical average of the time taken to generate the bridging content in the background, the lead time is obtained, ensuring that the pre-generated tasks always have a sufficient safety redundancy window, while avoiding the waste of invalid cache on edge nodes caused by premature generation.

[0018] Preferably, the step of pre-generating the bridging segment and pushing it to the edge nodes for caching before the expected occurrence time arrives further includes: selecting a number of edge nodes with the highest historical access frequency of the target device to form a candidate set; selecting the node with the fewest routing hops under the current network topology in the candidate set as the cache target node to push the bridging segment; when a switch actually occurs and the bridging segment has not yet been pre-generated, retrieving the most recently accessed segment of the same skeleton in the target modality from the edge node cache as a temporary start, and continuing to complete the generation of the bridging segment in the background.

[0019] Preferably, the method provided in this application further includes: after the learning session ends, capturing the actual drift degree and review behavior feedback data of this cross-modal switching; transmitting the feedback data to the server through a silent upload channel after lightweight encryption; updating the drift sample set of each student on the server in an incremental manner, and recalculating the set threshold for system self-correction.

[0020] This application utilizes low-cost sensor signals from on-the-job trainees' own devices to reconstruct multidimensional spatiotemporal learning states, uncovering statistically significant fragmented work environment profiles. By implementing cross-attention drift compensation to eliminate distribution differences during cross-modal switching, it accurately locates semantic breakpoints in the context and objectively produces semantically smooth bridging segments that conform to the trainees' cognitive coherence.

[0021] This application combines the pre-compensation method with the sinking edge distribution network based on dynamic calculation of the expected switching time. It achieves accurate pre-positioning of target resources at the optimal routing node. Even in the face of sudden changes in requests caused by user interruption at any time, it can ensure the continuity of knowledge acquisition by using degraded scheduling and delayed insertion mechanisms, effectively improving the end-to-end service stability of the multimodal system. Attached Figure Description

[0022] The above and other objects, features, and advantages of exemplary embodiments of the present invention will become readily apparent from the following detailed description taken in conjunction with the accompanying drawings. In the drawings, several embodiments of the invention are illustrated by way of example and not limitation, and like or corresponding reference numerals denote like or corresponding parts, wherein: Figure 1 This is a flowchart illustrating an intelligent fusion and distribution method for multimodal teaching resources for in-service education according to the present invention.

[0023] Figure 2 This is a schematic diagram illustrating the architecture flow of a method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to the present invention.

[0024] Figure 3 This diagram schematically illustrates a performance comparison of a multimodal teaching resource intelligent fusion and distribution method and system for in-service education according to the present invention. Detailed Implementation

[0025] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0027] This invention discloses an intelligent fusion and distribution method for multimodal teaching resources in continuing education, referring to... Figure 1 This includes steps S1-S4: S1. Semantic skeleton construction of multimodal knowledge atoms.

[0028] In an optional embodiment, knowledge point nodes in the syllabus are used as primary keys to uniformly segment heterogeneous teaching resources such as videos, audios, PPTs, and lecture notes into the smallest independently learnable knowledge atomic units. Video resources are segmented using shot boundary detection. Specifically, the pixel histogram difference between adjacent frames is calculated. When the difference exceeds a preset inter-frame difference threshold, it is determined as a shot boundary, thereby segmenting the continuous video stream into several semantically complete video segments. The inter-frame difference threshold can be determined using the Otsu thresholding method. Audio resources are segmented based on silence intervals. Short-time energy detection identifies silent intervals where the energy is consistently lower than the average background noise, and the midpoint of the silent interval is used as the segmentation boundary. PPTs are segmented by page, and documents are segmented based on paragraph heading levels and semantic paragraph boundaries. Priority is given to semantic paragraphs defined by first- and second-level headings as segmentation units. When a paragraph is too long, semantically complete natural paragraphs are further used as the smallest segmentation unit.

[0029] Each segment is matched with the knowledge point nodes in the outline based on semantic similarity and attached to the same knowledge atom, forming an index structure of one atom, multiple modalities, and multiple segments. This index structure uses the knowledge atom ID as the unique primary key, with modality type as the secondary index and segment number as the tertiary index, forming a clear hierarchical tree-like attachment relationship.

[0030] Specifically, three types of semantic features are extracted for each modal segment of each knowledge atom. For video segments, the salient region vector and lecturer posture description vector are extracted from the salient map obtained by existing visual saliency detection of keyframes. Visual saliency detection can use the Itti saliency model family or the spectral residual frequency domain saliency method. The salient region vector is composed of the centroid coordinates and area ratio of connected regions on the saliency map whose response value exceeds the mean plus one standard deviation of the entire map. The lecturer posture description vector is obtained by normalizing the upper body joint coordinate sequence output by the existing human keypoint detection model. For audio segments, the speech textification result and spectral peak envelope vector are extracted. After the speech textification result is transcribed by the existing speech recognition model, noun phrases and verb phrases are further extracted as a keyword set. The spectral peak envelope vector is formed by concatenating the frame-by-frame mean of the Mel frequency cepstral coefficient sequence. The document and PPT are segmented to extract titles, key sentences, and layout structure vectors. Key sentences are obtained by sorting the paragraphs internally using the existing TextRank algorithm and taking the first few sentences. The layout structure vectors encode layout attributes such as title level, list indentation depth, and text-image ratio.

[0031] Furthermore, the three types of features are mapped to a latent semantic space of the same dimension through a standard shared embedding network, forming the cross-modal feature skeleton of the knowledge atom. The shared embedding network is pre-trained using an existing multimodal contrastive learning framework, using feature pairs of segments from different modalities under the same knowledge atom as positive sample pairs and feature pairs of segments from different knowledge atoms as negative sample pairs. Contrastive loss drives semantically related features from different modalities to converge in the latent space. The dimension of the latent semantic space is determined by the shared embedding network during the pre-training stage based on the convergence of the contrastive loss, ensuring cross-modal semantic discriminability while also considering the efficiency of subsequent cosine similarity calculation. After the skeleton is constructed, any learning action in any modality can locate a unique knowledge atom ID on the skeleton, and locate the current segment feature under that ID, providing a unified semantic coordinate system for subsequent steps such as characterizing memory anchors, measuring drift, and generating bridging segments.

[0032] In this way, by uniformly segmenting and mounting heterogeneous teaching resources onto the knowledge atomic skeleton, a semantic positioning base independent of modality is provided for multimodal resources, enabling subsequent cross-modal learning to have a unified coordinate system, effectively eliminating heterogeneous barriers in progress positioning of different modal resources.

[0033] S2, Sensor-coordinated reconstruction of the spatiotemporal tensor of student behavior.

[0034] In an optional embodiment, low-cost hardware sensor signals already present on the device of the trainee are introduced as objective working condition criteria, forming the basis for behavior description together with the software-side behavior logs. The steady-state composite modulus of the mobile phone's gyroscope and accelerometer can distinguish three physical states: stationary, walking, and vehicle-mounted. The composite modulus is obtained by averaging the Euclidean norm of the triaxial accelerometer readings within a sliding window. The mean value for the stationary state is close to the gravitational acceleration and has a very small variance. The mean value for the walking state is slightly higher than the gravitational acceleration and exhibits periodic fluctuations. The mean value for the vehicle-mounted state fluctuates between walking and stationary states, but its frequency characteristics are significantly different from those for walking.

[0035] The root mean square (RMS) power of ambient sound collected by the microphone in non-recording mode can distinguish between three acoustic scenarios: quiet, moderately noisy, and very noisy. The RMS power is obtained by taking the square root of the square mean of the audio samples within the sampling window. The discrimination boundary between the three scenarios is determined by the natural distribution of the student's historical ambient sound samples. The screen brightness sensor, in conjunction with the system time, can distinguish between daytime and nighttime lighting. When the ambient light level is lower than the nighttime switching threshold for automatic screen brightness adjustment and the system time is within the student's historical nighttime learning period, it is determined to be a nighttime lighting condition.

[0036] Specifically, the learning behavior of students over a period of time is reconstructed from a traditional one-dimensional time-series log into a three-dimensional tensor of time-segment buckets, modality types, and physical environment. The number and boundaries of the time-segment buckets are not manually defined, but rather derived naturally from the probability density curve obtained by estimating the probability density of all the student's historical learning start timestamps on a 24-hour axis using existing kernel density estimation. The local troughs of this curve are used as bucket boundaries, and both the number and width of buckets objectively converge with the student's schedule. The bandwidth parameter of the kernel density estimation is automatically determined using the existing Silverman empirical formula based on the sample standard deviation and sample size. Modality types are enumerated into four categories: audio, video, PPT, and document. The physical environment is determined objectively from the joint samples of the three sensor signals using the existing K-Means elbow rule, with no preset number of scenes. The elbow rule calculates the sum of squares within each cluster under different cluster numbers, and the number of clusters corresponding to the point where the absolute value of the slope of the broken line first significantly narrows is taken as the final number of scenes.

[0037] Next, each cell in the tensor calculates the student's effective learning time and completion rate within that time period, modality, and environment. The completion rate is defined as the ratio of the actual effective playback time of that segment to the segment's own total duration. .

[0038] Among them, effective playback time The data is objectively collected from the player's existing callbacks, after deducting the cumulative duration of paused and background states. The total duration of this segment is given. By creating a feature distribution profile for each student on this three-dimensional tensor, a stable profile of their time-segment modal environmental preferences can be obtained. This profile is naturally converged from objective sensor data with sufficient samples. For example, suppose a student's historical learning start timestamps show three distinct density peaks on the 24-hour axis: morning, noon, and night. The kernel density troughs naturally divide the 24-hour axis into three time-segment buckets. The student's profile shows a strong preference for audio modality in the intersection of the morning bucket and the noisy in-vehicle environment, and a strong preference for video modality in the intersection of the night bucket and the quiet environment, providing objective statistical basis for subsequent switching predictions.

[0039] In this way, by introducing existing sensor signals from devices such as gyroscopes, ambient sound, and screen brightness, the one-dimensional behavior log is reconstructed into a three-dimensional spatiotemporal tensor without increasing hardware costs. This objectively captures the stable preference profiles of working students during fragmented time periods such as commuting, lunch breaks, and nighttime, providing a reliable statistical foundation for subsequent switching prediction and compensation triggering.

[0040] S3, Alignment and Drift Compensation of Cross-Modal Memory Anchors.

[0041] In an optional embodiment, a modal anchor triple is extracted in real time for each learning session, comprising three components: a visual anchor, an auditory anchor, and a semantic anchor. The visual anchor is taken from the salient region features corresponding to the peak position of the saliency map obtained from existing visual saliency detection of the keyframe of the current video segment. Specifically, it is represented by the feature vector of the connected region with the highest response value on the saliency map, which integrates the spatial location, texture gradient, and color contrast information of the region. The auditory anchor is taken from the keywords and spectral segments at energy mutation points and speech rate mutation points in the current audio segment. Energy mutations are detected by the first-order difference of the short-time energy sequence using the 3-sigma outlier criterion; that is, when the first-order difference value exceeds the mean of the sequence plus three times the standard deviation, it is determined to be an energy mutation point. Speech rate mutations are detected by the first-order difference of the phoneme count per unit time after existing VAD endpoint detection using the 3-sigma outlier criterion, with the determination logic consistent with that of energy mutations. Semantic anchors are selected from key sentences in the current document segment that are underlined, paused, or highlighted by the learner. Sentences whose pause duration exceeds the average reading speed of that segment are included in the semantic anchor candidate set. Finally, the sentences with the highest TextRank scores in the candidate set are selected as semantic anchors. The triple is written to the learner's anchor file under that knowledge atom at the end of the session for subsequent cross-modal switching.

[0042] Specifically, when trainees switch between modalities, it is necessary to measure the alignment between the anchor points left by the previous modality and the candidate anchor points of the next modality in a unified semantic space. Directly applying standard cross-attention will be affected by the systematic impact of the non-overlapping distribution centers of different modalities. The existing cross-attention formula is:

[0043] in This is the set of query vectors for the previous modality anchor point. , This is the set of key-value vectors for candidate anchor points of the next modality. The feature dimension; the implicit assumption of this formula is that... and Although they are in the same distribution, the centers of the visual, auditory, and text domains objectively exhibit a systematic shift in the latent semantic space in cross-modal scenarios, causing the attention scores to be skewed overall. To address this, a modal system shift compensation term is introduced, and the improved relationship is as follows:

[0044] in , Let be the feature mean vector of the source modality over the student's historical samples. The target modality is the feature mean vector on the trainee's historical samples. Both are directly obtained from historical data. The compensation term projects the global central difference between the source modality and the target modality along the candidate anchor point direction to compensate, thereby removing the modality domain difference from the attention score and retaining only the atomic semantic difference, so that the alignment score truly reflects the semantic proximity of the two anchor points.

[0045] Furthermore, define cross-modal drift degree To reverse the sign of the cosine similarity between the previous and next anchor points in the unified semantic space after alignment, i.e., the higher the drift, the farther apart the two anchor points are in the semantic space, the threshold for triggering compensation based on the drift is determined by the upper quartile of the review sample set. The time window for judging review behavior does not use any fixed number of minutes, but is objectively determined by the median of the time interval distribution of the first review action relative to the switch time after all the student's historical switches, as the upper bound of the time window.

[0046] To address the issue of an empty sample set for individual review during the cold start period for new learners, a two-stage degradation approach combining group prior and individual posterior can be adopted. The determination of insufficient sample size uses a single objective criterion: when the confidence interval width of the individual drift mean is still greater than the confidence interval width of the drift mean for the same profile group, it is determined that the individual sample is insufficient, triggering the group prior; when the individual confidence interval width narrows to no greater than the group confidence interval width, it is determined that the sample is sufficient, switching to the individual posterior. During the insufficient sample stage, the trigger threshold is taken as the upper quartile of the drift sample pool for the same profile group (all historical learners in the same time period and environment bucket). As a group prior; during the period of sufficient sample size, the trigger threshold is smoothly switched to the upper quartile of the individual drift sample. The smooth transition between the two stages uses linear interpolation, with the ratio of the confidence interval widths as the interpolation weights, to avoid the threshold from changing abruptly at the switching time.

[0047] Next, once the current If the threshold for the upper quartile is exceeded, the system triggers the bridging fragment generation process. It retrieves the same skeleton segment of the knowledge atom containing the anchor point of the previous modality in the target modality, and calculates the cosine similarity sequence between the feature vector of each sub-segment and the feature vector of the anchor point of the previous modality within that segment. ,position Peak position, and expand outwards to both sides along the peak. The first time the value falls below the mean of the sequence, the interval enclosed by this interval is used as the start and end boundary of the bridging segment. The segment division unit is aligned with the smallest semantic unit within the segment. Video segments are divided into fixed-duration sub-windows within the shot, audio segments are divided into single-sentence speech segments detected by VAD, and document segments are divided into natural sentences. After the start and end boundaries of the bridging segment are determined, the system encapsulates the target modality content within the interval into an independent bridging resource package for use in the next edge distribution process. The selection of this interval is entirely driven by the objective similarity curve, and the bridging length naturally converges with the semantic density of knowledge atoms and the similarity distribution of anchor points.

[0048] Thus, by introducing a modal system offset compensation term obtained directly from historical data into the existing cross-attention, the distribution center difference between the visual, auditory, and text domains is removed, so that the alignment score of cross-modal anchor points truly reflects the semantic difference rather than the modal domain difference, effectively suppressing memory anchor point drift. Furthermore, the objective statistical quantile threshold and the two-stage degradation approach ensure the stable operation of the compensation trigger throughout the entire lifecycle.

[0049] S4. Switch to predictive-driven bridging resource edge distribution.

[0050] In one optional embodiment, modality switching events of trainees within a future period are predicted based on the constructed spatiotemporal tensor profile. The prediction does not rely on complex prediction models; it only requires a conditional lookup table on the profile tensor for the current time period, current environment, and current learning duration to output the expected occurrence time and target modality type of the next switch. The specific logic of the conditional lookup table is as follows: using the current time period bucket index and the current environment bucket index as joint conditions, the corresponding cell in the profile tensor is located, and the offset distribution of the occurrence time of historical switching events relative to the start time of the time period bucket under that cell is read. The median of this distribution is taken as the expected switching offset, and added to the start time of the current time period bucket to obtain the expected switching time. The target modality type is taken as the mode of the target modalities in the switching event records. For working trainees, the stability of this profile is sufficient to support minute-level accuracy in predicting switching times, especially in the switching scenario between commuting hours and nighttime intensive learning hours, where the distribution concentration of historical switching times is usually high, and the prediction accuracy further converges with the increase of the historical sample size.

[0051] Specifically, a lead time is set before the predicted switching time. This lead time is the historical round-trip delay average from the edge node to the target device's first packet handshake. Historical statistical average of bridging fragment generation time The sum is objectively given as follows:

[0052] The average value is obtained by taking the round-trip delay samples of the target device's handshakes with each edge node in historical sessions. The average time taken from triggering to completion of all historical bridging segment generation tasks is obtained; both values ​​are directly derived from historical system measurement data. This addresses cold start scenarios such as initial device connection or insufficient historical samples. The median of the historical round-trip delay samples of all devices under the same edge node is taken as the population replacement value. The median of the time taken to generate samples through global bridging on the server is taken as the group substitute value. Both group substitute values ​​are gradually replaced by the individual mean as individual samples accumulate. The transition method is consistent with the linear interpolation logic from group prior to individual posterior.

[0053] Before the system switches to the expected time This involves invoking the bridging fragment generation process. Using the modal anchor triplet already formed in the current session as input, a bridging fragment for the target modality is pre-generated, and this fragment is pushed to the nearest edge node on the network path of the student's target device for caching. The selection of edge nodes is directly and objectively derived based on the historical distribution of the target device's frequently used access points. Specifically, several edge nodes with the highest access frequency of the target device in historical sessions are selected as a candidate set. From this candidate set, the node with the fewest routing hops in the current network topology is selected as the caching target node, avoiding a return to the origin server and thus preventing the first packet delay caused by cross-domain streaming during the switch.

[0054] Furthermore, when a student actually switches, the target device can retrieve the bridging segment from the nearest edge node during the same handshake when it first requests continued learning resources. This segment serves as the opening playback for the new modal session, followed seamlessly by the normal segmentation of the original knowledge atom in the target modality. If there is a discrepancy between the student's actual switching time and the predicted time, and this discrepancy results in the bridging segment not yet being pre-generated, the system initiates a degradation process: it prioritizes retrieving the most recently accessed segment of the same skeleton by another student in the target modality from the edge node cache as a temporary opening, while continuing to generate the bridging segment in the background. Once generated, supplementary bridging content is pushed to the student at the next natural pause point of the current session via an inserted prompt, ensuring that bridging compensation can still complete the closed loop even with delays.

[0055] Next, after the session ends, the system writes back feedback data such as the actual drift degree of this switch, whether a replay occurred, and whether the bridging segment was played completely to the student's personal historical drift sample set and the sample pool of the same image group, for use in the next round of self-updating of the upper quartile threshold. The write-back is triggered by the target device session ending event or the student actively exiting the session; both types of events are captured by the player's existing callbacks. The write-back data is transmitted to the server through a silent background upload channel after lightweight local encryption, without consuming the student's foreground operation bandwidth.

[0056] After receiving the write-back data, the server updates the student's drift sample set and the sample pool of the same image group in an incremental manner, and recalculates the upper quartile threshold, forming a continuous correction path driven solely by objective statistics, so that the system's compensation triggering accuracy continues to converge as the student's historical samples accumulate.

[0057] Figure 2 and Figure 3 These are schematic diagrams illustrating the architecture flow and performance comparison of a multimodal teaching resource intelligent fusion and distribution method for in-service education according to the present invention. As can be seen, Figure 2 The process of heterogeneous teaching resource segmentation and mapping, multi-channel sensor signal acquisition, spatiotemporal fingerprint generation, and data flow based on prediction-triggered bridging segment push is described in detail. Figure 3 The intuitive comparison shows the first packet response latency data curves of this solution and existing conventional methods when making cross-device and cross-modal switching requests, demonstrating that it significantly eliminates the memory gaps in the learning process compared to existing technologies.

[0058] Thus, by switching prediction-driven bridging segment edge pre-distribution, learners can obtain an opening segment semantically continuous with the previous modality anchor point in the first packet of cross-device switching. The memory gap in cross-modal learning is significantly reduced, and the system's end-to-end response stability and the smoothness of learners' learning curves are significantly improved.

Claims

1. A method for intelligent fusion and distribution of multimodal teaching resources for in-service education, characterized in that, include: Heterogeneous teaching resources are segmented and attached to the knowledge point nodes of the teaching syllabus, and semantic features are extracted from the segments and mapped to the latent semantic space to construct the semantic skeleton of multimodal knowledge atoms. The hardware sensor signals from the student's device are collected, and a three-dimensional tensor including time period buckets, modality types, and physical environment is constructed in combination with the learning behavior log. Based on the three-dimensional tensor, a time period modality environment preference profile is generated. Extract modal anchor triples from the current learning session. When a cross-modal switch occurs, calculate the alignment score and cross-modal drift between the anchor left by the previous modality and the candidate anchor of the next modality. This includes: extracting modal anchor triples containing visual, auditory, and semantic anchors; calculating the alignment score using a cross-attention formula that introduces a modal system offset compensation term, wherein the modal system offset compensation term is obtained by projecting the difference between the feature mean vectors of the source modality and the target modality on historical samples along the direction of the candidate anchor; and defining the cross-modal drift as the inverse sign of the cosine similarity between the anchor of the previous modality and the anchor of the next modality in the latent semantic space after alignment. When the cross-modal drift exceeds a set threshold, a bridging segment is generated, including: when the number of individual review samples of a new learner is insufficient, the upper quartile of the drift sample pool of the same image group is used as the group prior threshold; when the confidence interval width of the individual drift mean is not greater than the confidence interval width of the group drift mean, the upper quartile of the individual drift sample is used as the individual posterior threshold; the set threshold is obtained by smoothing the transition between the group prior threshold and the individual posterior threshold using linear interpolation; when the cross-modal drift exceeds the set threshold, the cosine similarity sequence between the feature vector of each sub-segment and the feature vector of the anchor point of the previous modality is calculated within the same skeleton segment of the target modality, and the bridging segment is generated by cropping outwards from the peak of the cosine similarity sequence to both sides. Based on the time-segment modal environment preference profile, the expected time of the next switch and the target modal type are predicted. Before the expected time of switch arrives, the bridging segment is pre-generated and pushed to the edge node for caching.

2. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1, wherein segmenting heterogeneous teaching resources and attaching them to knowledge point nodes of the teaching syllabus, and extracting semantic features from the segments and mapping them to the latent semantic space, includes: The video stream is segmented into video segments, and salient region vectors and lecturer posture description vectors are extracted. The audio resources are segmented into audio segments based on silence intervals, and the speech textification results and spectral peak envelope vectors are extracted. The document and slides are segmented according to paragraph boundaries, and the headings, key sentences, and layout structure vectors are extracted. The extracted features are mapped to the latent semantic space through a shared embedding network to form a cross-modal feature skeleton.

3. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1, wherein the step of collecting hardware sensor signals from student devices and constructing a three-dimensional tensor including time period buckets, modality types, and physical environments in conjunction with learning behavior logs includes: Physical states can be distinguished based on the steady-state composition of the combined modulus of the device-side gyroscope and accelerometer. Acoustic scene discrimination is based on the root mean square power of ambient sound collected by microphone; Lighting status is determined by combining screen brightness sensor and system time; The learning behavior is reconstructed into a three-dimensional tensor, where the boundaries of the time buckets are naturally derived from the troughs of the probability density curves of the learning start timestamps, and the physical environment is determined by joint sample clustering of sensor signals.

4. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 3, wherein generating a time-specific modal environment preference profile based on the three-dimensional tensor includes: In each cell of the three-dimensional tensor, the effective learning duration and completion rate under the current time period, current modality, and current physical environment are statistically analyzed. The completion rate is calculated based on the ratio of the actual effective playback time to the duration of the segment itself. A profile of the time-segment modal environment preference is obtained by mapping the feature distribution on the three-dimensional tensor.

5. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1, wherein predicting the expected time of the next switch and the target modality type based on the time-segment modal environment preference profile includes: The corresponding cell is located in the three-dimensional tensor by using the index of the current time bucket and the index of the current environment bucket as joint conditions; Read the offset distribution of the occurrence time of historical switching events under this cell relative to the start time of the time period bucket; The median of the offset distribution is taken as the expected switching offset. The expected switching offset is added to the start time of the current time period bucket to obtain the expected occurrence time. The mode of the target modality in the historical switching event record is taken as the target modality type.

6. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1, wherein the step of pre-generating the bridging segment and pushing it to the edge node for caching before the expected occurrence time arrives includes: The average historical round-trip latency of the target device in handshakes with each edge node during historical sessions; The historical average generation time of bridging segments from triggering to completion was statistically analyzed. The lead time is obtained by adding the average historical round-trip delay to the average generation time. The pre-generation of the bridging segment is triggered at the moment when the expected occurrence time is subtracted from the lead time.

7. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1, wherein the step of pre-generating the bridging segment and pushing it to the edge node for caching before the expected occurrence time arrives, further includes: A candidate set is formed by selecting several edge nodes with the highest historical access frequency of the target device. The node with the fewest routing hops under the current network topology in the candidate set is selected as the cache target node to push the bridging segment. When a switch actually occurs and the bridging segment has not yet been pre-generated, the most recently accessed segment of the same skeleton in the target modality of the current knowledge atom is retrieved from the edge node cache as a temporary start, and the generation of the bridging segment continues in the background.

8. The method for intelligent fusion and distribution of multimodal teaching resources for in-service education according to claim 1 further includes: After the learning session ends, capture the actual drift degree and review behavior feedback data of this cross-modal switch; The feedback data is lightly encrypted and then transmitted to the server via a silent upload channel; The server updates the drift sample set of each student incrementally and recalculates the set threshold for system self-correction.

Citation Information

Patent Citations

  • Deep learning-based smart tunnel multi-modal data collaborative management method and device

    CN120744470A

  • Multi-mode self-adaptive heterogeneous data intelligent weaving system

    CN122112958A