Intelligent home learning track guiding method fused with multi-modal behavior perception

By constructing a cross-modal weakly coupled difference function and a semantic transfer calibration mechanism, the problem of multimodal signal noise processing in the home learning environment was solved, achieving stable and consistent signal enhancement and personalized learning trajectory prediction, thereby improving the accuracy and interpretability of learning intervention.

CN121836983AInactive Publication Date: 2026-04-10SHENZHEN YUESHANG EDUCATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing family learning analysis methods cannot effectively handle noise in multimodal signals, resulting in distorted representations of learning states and making it difficult to perform high-quality trend predictions and interpretable learning trajectory guidance.

Method used

By constructing a cross-modal weakly coupled difference function, semantic transfer calibration, and behavioral residual completion mechanism, we enhance the processing of noisy modalities and combine semantic-behavioral dual-path representation and multi-timescale learning trajectory prediction to generate personalized learning guidance plans.

Benefits of technology

Maintaining stable consistency of multimodal signals under low-quality conditions improves the accuracy and comprehensibility of learning interventions, generating personalized learning guidance programs with clear behavioral basis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121836983A_ABST
    Figure CN121836983A_ABST
Patent Text Reader

Abstract

The invention relates to the field of intelligent education, in particular to a family learning track intelligent guiding method fusing multi-mode behavior perception. Comprising the steps of hierarchical acquisition of multi-modal behavior signals, modal collaborative enhancement preprocessing, learning semantic-behavior dual-channel representation construction, key behavior driving factor identification, trend prediction of a multi-time scale learning track and intelligent guidance generation of strategy interpretability. By constructing a cross-modal weak coupling difference function, semantic migration calibration and a behavior residual error complementation mechanism, a noise mode interfered by a family environment is enhanced, and a mode confidence map is generated, so that multi-modal signals can still keep stable consistency under the condition of low quality, and the reliability of basic data is remarkably improved; through semantic-behavior double-channel representation construction, key behavior driving factor identification and multi-time scale learning track prediction, unified representation of a learning state can be obtained, and the accuracy, the understandability and the practicability of learning intervention are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent education, specifically to an intelligent guidance method for family learning trajectories that integrates multimodal behavioral perception. Background Technology

[0002] With the increasing prevalence of digital devices in home learning settings, the collection of learning behaviors has gradually expanded from simple operation records to multimodal signals such as visual, audio, and application logs. However, existing home learning analysis methods still have the following problems:

[0003] (1) Multimodal signals are easily affected by the natural environment of the home, and the noise mode cannot be accurately identified and processed:

[0004] Affected by factors such as changes in illumination, equipment vibration, occlusion, environmental voice, and operation delay, the quality of signals in different modes is uneven. Traditional single-mode filtering methods cannot achieve cross-modal noise localization and consistency enhancement, resulting in distorted characterization results.

[0005] (2) Existing methods struggle to form stable and accurate representations of learning states, and lack high-quality trend predictions of learning trajectories and interpretable policy generation.

[0006] Existing learning behavior assessment methods are mostly based on single behaviors or short-term data, which cannot integrate semantic information, operational behaviors and learning pace into a unified model. They are difficult to identify key behavioral drivers that affect learning performance and also difficult to generate explainable learning guidance programs for family scenarios, resulting in guidance strategies that lack a basis and are not traceable. Summary of the Invention

[0007] To address the above issues and overcome the shortcomings of existing technologies, this invention provides an intelligent guidance method for family learning trajectories that integrates multimodal behavioral perception. Addressing the problem of inaccurate identification and processing of noisy modalities, this invention enhances noisy modalities affected by family environmental interference by constructing a cross-modal weakly coupled difference function, semantic transfer calibration, and behavioral residual completion mechanism, generating a modal confidence map. This ensures stable consistency among multimodal signals even under low-quality conditions, significantly improving the reliability of basic data. Addressing the lack of high-quality trend prediction and interpretable strategy generation for learning trajectories, this invention obtains a unified representation of the learning state through semantic-behavioral dual-pathway representation construction, key behavioral driving factor identification, and multi-timescale learning trajectory prediction. Furthermore, by combining a strategy rule base and feature contribution analysis methods, it generates personalized learning guidance plans with clear behavioral and trend bases, thereby improving the accuracy, understandability, and practicality of learning interventions.

[0008] The technical solution adopted by this invention is as follows: The intelligent guidance method for family learning trajectories that integrates multimodal behavior perception provided by this invention includes the following steps:

[0009] Step S1: Hierarchical acquisition of multimodal behavioral signals. Collect multimodal behavioral signals in the home environment, including visual, audio, operational behavior, learning application logs and physiological micro-expressions. Use a modal quality assessment model to identify noise modalities caused by lighting, occlusion, device jitter, background voice and operation delay. Construct a three-level data structure of basic layer - association layer - higher-order mode layer.

[0010] Step S2: Modal collaborative enhancement preprocessing. Based on the noise mode and the corresponding multimodal behavioral signal, a cross-modal consistency enhancement mechanism is proposed. Semantic transfer calibration and behavioral residual completion are used to process the noise mode to obtain the enhanced noise mode. At the same time, a modal confidence map is generated.

[0011] Step S3: Learn the semantic-behavioral dual-pathway representation construction. A dual-pathway representation network is constructed using semantic encoding and behavioral feature extraction methods, including a semantic path and a behavioral path. The semantic path encodes the task content, learning corpus, and screen text using knowledge graphs, while the behavioral path encodes attention, operating habits, and learning rhythm. The two paths are then fused through adaptive alignment and mutual attention to generate a unified learning state representation.

[0012] Step S4: Identification of key behavioral drivers. Based on the unified learning state representation and combined with behavioral analysis methods, identify key behavioral drivers that affect learning performance, including procrastination, increased cognitive load, and changes in focus.

[0013] Step S5: Trend prediction of learning trajectory at multiple time scales. Construct a time prediction framework at the hourly, daily, and weekly levels. Utilize a segmented variational prediction structure to simultaneously predict learning performance trends, risk events, and rhythm fluctuation ranges to obtain prediction results.

[0014] Step S6: Generate intelligently interpretable guidance based on key behavioral drivers and prediction results. Set up a strategy rule base and decision logic, match learning behaviors with strategies, generate guidance schemes, and demonstrate the interpretability of the strategy's basis based on the aforementioned unified learning state representation and key behavioral drivers. Display the behavioral indicators, trend characteristics, and related conditions that trigger the strategy through conventional feature contribution calculation methods, and finally output a personalized learning guidance plan.

[0015] Furthermore, step S2 specifically includes the following steps:

[0016] Step S21: Cross-modal weak coupling localization of noise-sensitive segments. Extract time segments of noise modes on different multimodal behavioral signals, and construct a weak coupling difference function, as follows:

[0017] ;

[0018] ;

[0019] in, Indicates the difference in cross-modal weakly coupled noise. Represents noise modes At any moment Multimodal behavioral signals This represents a local smoothing operator based on the statistical stability of adjacent segments. Indicates based on modality Weights that are dynamically generated based on stability. Represent any positive integer, Represents the set of noise modes;

[0020] By using threshold determination, a set of noise-sensitive segments is obtained based on the peak interval of the weakly coupled difference function, as shown below:

[0021] ;

[0022] in, Represents the set of noise-sensitive segments. Indicates the noise difference threshold;

[0023] Step S22: Semantic transfer calibration. Using the multimodal behavioral signals corresponding to the set of noise-sensitive time segments as input, extract the corresponding noise modality semantic representation for each noise modality, as shown below:

[0024] ;

[0025] in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain noise mode Semantic projection transformation;

[0026] Step S23: Behavioral residual completion mechanism. For the semantic representation of the noise mode, the changing trend of the noise mode in the time dimension is modeled, and cross-modal evolution residuals are constructed, as follows:

[0027] ;

[0028] in, Represents noise modes The time evolution residuals between the noise modes and other noise modes This represents the cross-modal evolution difference operator. noise mode The rate of evolution, noise mode The rate of evolution;

[0029] Step S24: Dynamic weighted calculation. Based on the semantic representation of the noise mode, assign semantic consistency weights to each mode residual, as shown below:

[0030] ;

[0031] in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain This represents a semantic consistency measure function. Representing modes For modes The contribution weight of residual completion;

[0032] Step S25: Enhanced mode generation. Using cross-modal residuals and semantic consistency weights, the noise mode is reconstructed and completed to obtain the enhanced noise mode. The formula used is as follows:

[0033] ;

[0034] in This represents the enhanced noise mode. This represents the signal after initial smoothing of the noise modes. This represents the residual completion amount based on semantic weights.

[0035] Step S26: Modal confidence map generation. Based on the enhanced noise modes, construct a modal confidence map.

[0036] Furthermore, step S3 specifically includes the following steps:

[0037] Step S31: Semantic path encoding. The task content, learning corpus, and screen text are serialized to obtain the semantic input sequence, as shown below:

[0038] ;

[0039] in, Indicates the first Each text unit Indicates the length of the text sequence;

[0040] The semantic pathway representation is obtained through the semantic encoder, as follows:

[0041] ;

[0042] in, For semantic encoding functions;

[0043] Step S32: Behavioral pathway feature extraction. Behavioral features are extracted from attention, operational behavior, and learning rhythm signals to obtain a behavioral feature sequence. The formula used is as follows:

[0044] ;

[0045] in, Indicates the first A behavioral characteristic, Represents behavioral characteristic dimensions;

[0046] The behavioral pathway representation obtained through the behavior encoder is as follows:

[0047] ;

[0048] in Encoding function for behavioral features;

[0049] Step S33: Semantic-behavior alignment fusion, based on semantic representation Behavioral representations are computed across pathways, and adaptively weighted and fused to obtain a unified learning state representation:

[0050] ;

[0051] in, This represents the learning state after fusion. α∈[0,1] For adaptive fusion weights.

[0052] The beneficial effects achieved by the present invention using the above solution are as follows:

[0053] (1) To address the problem of inaccurate identification and processing of noise modes, this invention enhances noise modes affected by home environment interference by constructing cross-modal weak coupling difference function, semantic transfer calibration and behavioral residual completion mechanism, and generates modal confidence spectrum, so that multimodal signals can still maintain stable consistency under low quality conditions, significantly improving the reliability of basic data;

[0054] (2) In response to the lack of high-quality trend prediction and interpretable strategy generation for learning trajectories, this invention constructs a semantic-behavioral dual-pathway representation, identifies key behavioral driving factors, and predicts learning trajectories at multiple time scales. This enables the generation of a unified representation of the learning state. Furthermore, by combining a strategy rule base and feature contribution analysis method, a personalized learning guidance plan with clear behavioral and trend basis is generated, thereby improving the accuracy, comprehensibility, and practicality of learning intervention. Attached Figure Description

[0055] Figure 1 This is a flowchart illustrating the intelligent guidance method for family learning trajectories that integrates multimodal behavior perception provided by the present invention.

[0056] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used together with the embodiments of the invention to explain the invention and do not constitute a limitation thereof. Detailed Implementation

[0057] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0058] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", "top", "bottom", "inner", "outer", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this invention.

[0059] Example 1, see Figure 1 The present invention provides an intelligent guidance method for family learning trajectories that integrates multimodal behavior perception. The method includes the following steps:

[0060] Step S1: Hierarchical acquisition of multimodal behavioral signals;

[0061] Step S2: Modal co-enhancement preprocessing;

[0062] Step S3: Learn the construction of semantic-behavioral dual-pathway representations;

[0063] Step S4: Identification of key behavioral drivers;

[0064] Step S5: Trend prediction of multi-timescale learning trajectories;

[0065] Step S6: Generate strategy-interpretable intelligent guidance.

[0066] In this embodiment, the actual application process of the method of the present invention is illustrated by taking the scenario of a fifth-grade student completing math homework at home as an example.

[0067] A student uses a tablet to complete online math homework at home every day, and the learning process is monitored in real time by a home smart learning terminal. First, the student's multimodal behavioral signals are collected in the learning scenario, including upper body posture and facial expressions captured by the camera in front of the learning table, real-time voice environment collected by the microphone, touch operation records of the tablet, application learning logs, and micro-expression electromyography changes detected by the smart watch strap.

[0068] During a learning session, due to dim room lighting and the student leaning forward, the camera footage showed localized shadows and slight blurring. At the same time, background noise from the younger brother playing caused some audio signal quality to degrade. Based on the built-in modal quality assessment model, the system identified these disturbed visual and audio segments as noise modalities and automatically marked them as "abnormal inputs in the base layer" in the three-level data structure.

[0069] Subsequently, taking the visual modality as an example, although facial details are lost due to shadows, the system uses synchronously extracted micro-expression electromyography signals and touch operation rhythm to perform semantic transfer calibration on the visual signals, completing key behavioral features. This allows the visual modality to recover the student's attention direction and posture changes. For audio affected by background noise, the system combines screen operation logs to infer the student's problem-solving rhythm and recovers effective speech segments related to behavior in the audio modality through a residual completion mechanism. The completed modality will obtain a corresponding confidence score, forming a modality confidence map.

[0070] After modality enhancement, the system simultaneously models "semantic content" and "behavioral performance" through a dual-pathway representation network. The semantic path encodes the question text and the student's past error types into a knowledge graph based on the student's current "fraction addition and subtraction" task. The behavioral path extracts behavioral features from indicators such as posture stability, touch speed, dwell time, and micro-expression tension. When the two are fused through an adaptive alignment module, the system generates a unified learning state representation. For example, this representation shows that when students encounter questions involving operations with unlike denominators, their dwell time significantly increases, and they exhibit signs of increased workload such as slight frowning and forward head tilting.

[0071] Based on the unified representation described above, the system automatically identifies key behavioral drivers. In this example, the student exhibited a significant increase in cognitive load, a slower pace of operation, and a slight decline in attention when encountering cognitive difficulties. These behaviors were identified by the system as key factors affecting learning performance.

[0072] The system uses hourly, daily, and weekly prediction frameworks to predict the student's learning trajectory. After analyzing the student's learning pace and error distribution over the past week, the model predicts that if no adjustments are made, the student's error rate on fraction calculation questions may further increase in the next two days, and the student may procrastinate in the later stages of assignments.

[0073] Finally, based on the prediction results and key behavioral drivers, the system matches a suitable guidance scheme from the policy rule base. For example, the system generates the following interpretable learning guidance strategy:

[0074] (1) After completing the current set of questions, arrange a 3-minute micro-relaxation period to relieve cognitive load;

[0075] (2) Provide automated "example animations for converting denominators" for error-prone fraction problems to reduce mental blocks;

[0076] (3) Add "small goal reminders for fraction calculation" to the learning plan to guide students to maintain a consistent pace;

[0077] The system also displays the basis for the strategy trigger, such as: cognitive load index increased by 18%, operation interval increased by 27%, facial tension increased, etc., so that parents can clearly understand the generation logic of the strategy.

[0078] Example 2, based on the above example, specifically includes the following steps in step S2:

[0079] Step S21: Cross-modal weak coupling localization of noise-sensitive segments. Extract time segments of noise modes on different multimodal behavioral signals, and construct a weak coupling difference function, as follows:

[0080] ;

[0081] ;

[0082] in, Indicates the difference in cross-modal weakly coupled noise. Represents noise modes At any moment Multimodal behavioral signals This represents a local smoothing operator based on the statistical stability of adjacent segments. Indicates based on modality Weights that are dynamically generated based on stability. Represent any positive integer, Represents the set of noise modes;

[0083] By using threshold determination, a set of noise-sensitive segments is obtained based on the peak interval of the weakly coupled difference function, as shown below:

[0084] ;

[0085] in, Represents the set of noise-sensitive segments. Indicates the noise difference threshold;

[0086] Step S22: Semantic transfer calibration. Using the multimodal behavioral signals corresponding to the set of noise-sensitive time segments as input, extract the corresponding noise modality semantic representation for each noise modality, as shown below:

[0087] ;

[0088] in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain noise mode Semantic projection transformation;

[0089] Step S23: Behavioral residual completion mechanism. For the semantic representation of the noise mode, the changing trend of the noise mode in the time dimension is modeled, and cross-modal evolution residuals are constructed, as follows:

[0090] ;

[0091] in, Represents noise modes The time evolution residuals between the noise modes and other noise modes This represents the cross-modal evolution difference operator. noise mode The rate of evolution, noise mode The rate of evolution;

[0092] Step S24: Dynamic weighted calculation. Based on the semantic representation of the noise mode, assign semantic consistency weights to each mode residual, as shown below:

[0093] ;

[0094] in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain This represents a semantic consistency measure function. Representing modes For modes The contribution weight of residual completion;

[0095] Step S25: Enhanced mode generation. Using cross-modal residuals and semantic consistency weights, the noise mode is reconstructed and completed to obtain the enhanced noise mode. The formula used is as follows:

[0096] ;

[0097] in This represents the enhanced noise mode. This represents the signal after initial smoothing of the noise modes. This represents the residual completion amount based on semantic weights.

[0098] Step S26: Modal confidence map generation. Based on the enhanced noise modes, construct a modal confidence map.

[0099] In this embodiment, the code used is as follows:

[0100] import numpy as np

[0101] from scipy.ndimage import gaussian_filter1d

[0102] from sklearn.decomposition import PCA

[0103] import matplotlib.pyplot as plt

[0104] # ============================================

[0105] # Example of simulating three modalities: visual, audio, and operational behavior

[0106] # ============================================

[0107] t = np.linspace(0, 10, 500)

[0108] visual = np.sin(t) + np.random.normal(0, 0.05, len(t))

[0109] audio = np.cos(t) + np.random.normal(0, 0.05, len(t))

[0110] action = 0.5 * np.sin(2 * t) + np.random.normal(0, 0.05, len(t))

[0111] # Inject noise

[0112] noise_range = (150, 230)

[0113] visual[noise_range[0]:noise_range[1]] += np.random.normal(0, 0.8,noise_range[1] - noise_range[0])

[0114] audio[noise_range[0]:noise_range[1]] += np.random.normal(0, 0.6,noise_range[1] - noise_range[0])

[0115] action[noise_range[0]:noise_range[1]] += np.random.normal(0, 0.7,noise_range[1] - noise_range[0])

[0116] modalities = {

[0117] "visual": visual,

[0118] "audio": audio,

[0119] "action": action

[0120] }

[0121] # ============================================

[0122] # S21: Weak coupling difference function D(t)

[0123] # ============================================

[0124] def local_smooth(x):

[0125] return gaussian_filter1d(x, sigma=3)

[0126] def weak_coupling_D(mods):

[0127] D = np.zeros_like(list(mods.values())[0])

[0128] for m_name, Xm in mods.items():

[0129] Xm_s = local_smooth(Xm)

[0130] w = 1 / (np.var(Xm) + 1e-6)

[0131] D += w * np.abs(Xm - Xm_s)

[0132] return D

[0133] D = weak_coupling_D(modalities)

[0134] tau = np.percentile(D, 85)

[0135] T_noise = np.where(D >= tau)[0]

[0136] # ============================================

[0137] # S22: Semantic transfer calibration

[0138] # ============================================

[0139] def semantic_projection(mods, T_idx):

[0140] X = np.vstack([mods[m][T_idx] for m in mods]).T

[0141] pca = PCA(n_components=2)

[0142] Z = pca.fit_transform(X)

[0143] # Assign modal semantic vectors

[0144] z = {}

[0145] for i, m in enumerate(mods):

[0146] z[m] = Z[:, :2].mean(axis=0) + np.random.normal(0, 0.01, 2)

[0147] return z

[0148] z_semantic = semantic_projection(modalities, T_noise)

[0149] # ============================================

[0150] # S23: Cross-modal residual modeling

[0151] # ============================================

[0152] def time_derivative(x):

[0153] return np.gradient(x)

[0154] derivative = {m: time_derivative(modalities[m]) for m in modalities}

[0155] def residual(m, mods):

[0156] others = [k for k in mods if k != m]

[0157] R = np.zeros_like(mods[m])

[0158] for k in others:

[0159] R += np.abs(derivative[m] - derivative[k])

[0160] return R / len(others)

[0161] R = {m: residual(m, modalities) for m in modalities}

[0162] # ============================================

[0163] # S24: Dynamic weighting w(m,k)

[0164] # ============================================

[0165] def sim(a, b):

[0166] return np.exp(-np.linalg.norm(a - b))

[0167] weights = {}

[0168] for m in modalities:

[0169] weights[m] = {}

[0170] others = [k for k in modalities if k != m]

[0171] sims = np.array([sim(z_semantic[m], z_semantic[k]) for k in others])

[0172] sims_exp = np.exp(sims)

[0173] sims_exp / = sims_exp.sum()

[0174] for i, k in enumerate(others):

[0175] weights[m][k] = sims_exp[i]

[0176] # ============================================

[0177] # S25: Enhanced Modality Generation

[0178] # ============================================

[0179] def enhance_modality(m):

[0180] x = modalities[m]

[0181] x_s = local_smooth(x)

[0182] comp = sum(weights[m][k] * R[m] for k in weights[m])

[0183] return x_s + comp

[0184] enhanced = {m: enhance_modality(m) for m in modalities}

[0185] # ============================================

[0186] # S26: Modal confidence plot

[0187] # ============================================

[0188] def modality_confidence(orig, enh):

[0189] # The closer the residuals are to the original state before and after smoothing, the smaller they are → the higher the confidence level.

[0190] return 1 / (np.mean(np.abs(orig - enh)) + 1e-6)

[0191] confidence = {m: modality_confidence(modalities[m], enhanced[m]) form in modalities}

[0192] # ============================================

[0193] # Output example results

[0194] # ============================================

[0195] print("Index to noise-sensitive segment:", T_noise[:20])

[0196] print("\nModal semantic vector z_m:", z_semantic)

[0197] print("\nModal confidence graph:", confidence)

[0198] # ============================================

[0199] # Visualization

[0200] # ============================================

[0201] plt.figure(figsize=(12, 6))

[0202] plt.plot(D, label="D(t) Weakly Coupled Difference Function")

[0203] plt.axhline(tau, color='r', linestyle='--', label="Threshold τ")

[0204] plt.title

[0205] plt.legend()

[0206] plt.show()

[0207] for m in modalities:

[0208] plt.figure(figsize=(12, 5))

[0209] plt.plot(modalities[m], label=f"{m} primitive")

[0210] plt.plot(enhanced[m], label=f"{m} Enhanced")

[0211] plt.title(f"S25: Modal Enhancement Effects ({m})")

[0212] plt.legend()

[0213] plt.show()

[0214] plt.figure(figsize=(6, 4))

[0215] plt.bar(confidence.keys(), confidence.values())

[0216] plt.title("S26: Modal Confidence Mapping")

[0217] plt.show().

[0218] Example 3, based on the above examples, specifically includes the following steps in step S3:

[0219] Step S31: Semantic path encoding. The task content, learning corpus, and screen text are serialized to obtain the semantic input sequence, as shown below:

[0220] ;

[0221] in, Indicates the first Each text unit Indicates the length of the text sequence;

[0222] The semantic pathway representation is obtained through the semantic encoder, as follows:

[0223] ;

[0224] in, For semantic encoding functions;

[0225] Step S32: Behavioral pathway feature extraction. Behavioral features are extracted from attention, operational behavior, and learning rhythm signals to obtain a behavioral feature sequence. The formula used is as follows:

[0226] ;

[0227] in, Indicates the first A behavioral characteristic, Represents behavioral characteristic dimensions;

[0228] The behavioral pathway representation obtained through the behavior encoder is as follows:

[0229] ;

[0230] in Encoding function for behavioral features;

[0231] Step S33: Semantic-behavior alignment fusion, based on semantic representation Behavioral representations are computed across pathways, and adaptively weighted and fused to obtain a unified learning state representation:

[0232] ;

[0233] in, This represents the learning state after fusion. α∈[0,1] For adaptive fusion weights.

[0234] In this embodiment, taking the process of a sixth-grade elementary school student learning math word problems at home as an example, the system first collects the text of a word problem about unit price and quantity calculation that the student is currently solving from the learning terminal, such as "Xiaoming bought 3 pencils, each costing 2.5 yuan, find the total price". After text serialization processing, the system obtains a semantic input sequence composed of "'Xiaoming', 'bought', '3 pencils', 'each', 'price', '2.5 yuan', 'find the total price'", and inputs it into a semantic encoder for encoding to generate a 128-dimensional semantic representation vector Es. This vector is used to represent the semantic topic and knowledge structure of the current task. At the same time, the system extracts multi-source behavioral signals of the student from the camera, touch screen, and learning log, including attention stability and gaze deviation from head posture and gaze trajectory, and collects operational behaviors such as screen click interval and input speed. Features such as the degree of focus and the duration of time spent on the current question collectively constitute the behavioral feature sequence B. After processing by the behavioral encoder, a 128-dimensional behavioral representation vector Eb is obtained, reflecting the student's current focus, operational fluency, and changes in learning rhythm. Subsequently, the system automatically adjusts the fusion ratio of the semantic and behavioral pathways based on fluctuations in the student's behavioral state. In this example, based on the observed slight decrease in attention and uneven behavioral rhythm, the system calculates a fusion weight α of 0.42 and uses this to weight and fuse the vectors of the two pathways, thereby obtaining a unified learning state representation H = 0.42Es + 0.58Eb. The fused 128-dimensional feature vector comprehensively characterizes the semantic attributes of the learning task and the student's real-time behavioral state, providing a reliable underlying state input for subsequent learning trend prediction, behavioral driving factor identification, and personalized guidance strategy generation.

[0235] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0236] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

[0237] The present invention and its embodiments have been described above. This description is not restrictive, and the accompanying drawings are only one embodiment of the present invention; the actual structure is not limited thereto. In conclusion, if those skilled in the art are inspired by this description and design similar structures and embodiments without departing from the spirit of the invention, such designs should fall within the protection scope of the present invention.

Claims

1. A family learning trajectory intelligent guidance method integrating multimodal behavior perception, characterized by: The present invention provides an intelligent guidance method for family learning trajectories that integrates multimodal behavior perception. The method includes the following steps: Step S1: Hierarchical acquisition of multimodal behavioral signals. Collect multimodal behavioral signals in the home environment, use a modal quality assessment model to identify noise modes caused by lighting, occlusion, device jitter, background voice and operation delay, and construct a three-level data structure of basic layer - correlation layer - higher-order mode layer. Step S2: Modal collaborative enhancement preprocessing. Based on the noise mode and the corresponding multimodal behavioral signal, a cross-modal consistency enhancement mechanism is proposed. Semantic transfer calibration and behavioral residual completion are used to process the noise mode to obtain the enhanced noise mode. At the same time, a modal confidence map is generated. Step S3: Learn the semantic-behavioral dual-pathway representation construction. Use semantic encoding and behavioral feature extraction methods to construct a dual-pathway representation network and generate a unified learning state representation. Step S4: Identification of key behavioral drivers: Based on the unified learning state representation and combined with behavioral analysis methods, identify key behavioral drivers that affect learning performance. Step S5: Trend prediction of learning trajectory at multiple time scales. Construct a time prediction framework at the hourly, daily, and weekly levels. Utilize a segmented variational prediction structure to simultaneously predict learning performance trends, risk events, and rhythm fluctuation ranges to obtain prediction results. Step S6: Generate intelligently interpretable guidance based on key behavioral drivers and prediction results. Set up a strategy rule base and decision logic, match learning behaviors with strategies, generate guidance schemes, and demonstrate the interpretability of the strategy's basis based on the aforementioned unified learning state representation and key behavioral drivers. Display the behavioral indicators, trend characteristics, and related conditions that trigger the strategy through conventional feature contribution calculation methods, and finally output a personalized learning guidance plan.

2. The intelligent guidance method for family learning trajectories integrating multimodal behavior perception as described in claim 1, characterized in that: Step S2 specifically includes the following steps: Step S21: Cross-modal weak coupling localization of noise-sensitive segments. Extract time segments of noise modes on different multimodal behavioral signals, and construct a weak coupling difference function, as follows: ; ; in, Indicates the difference in cross-modal weakly coupled noise. Represents noise modes At any moment Multimodal behavioral signals This represents a local smoothing operator based on the statistical stability of adjacent segments. Indicates based on modality Weights that are dynamically generated based on stability. Represent any positive integer, Represents the set of noise modes; By using threshold determination, a set of noise-sensitive segments is obtained based on the peak interval of the weakly coupled difference function, as shown below: ; in, Represents the set of noise-sensitive segments. Indicates the noise difference threshold; Step S22: Semantic transfer calibration. Using the multimodal behavioral signals corresponding to the set of noise-sensitive time segments as input, extract the corresponding noise modality semantic representation for each noise modality, as shown below: ; in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain noise mode Semantic projection transformation; Step S23: Behavioral residual completion mechanism. For the semantic representation of the noise mode, the changing trend of the noise mode in the time dimension is modeled, and cross-modal evolution residuals are constructed, as follows: ; in, Represents noise modes The time evolution residuals between the noise modes and other noise modes This represents the cross-modal evolution difference operator. noise mode The rate of evolution, noise mode The rate of evolution; Step S24: Dynamic weighted calculation. Based on the semantic representation of the noise mode, assign semantic consistency weights to each mode residual, as shown below: ; in, Represents noise modes Noisy modal semantic representation after projection onto the shared semantic domain This represents a semantic consistency measure function. Representing modes For modes The contribution weight of residual completion; Step S25: Enhanced mode generation. Using cross-modal residuals and semantic consistency weights, the noise mode is reconstructed and completed to obtain the enhanced noise mode. The formula used is as follows: ; in This represents the enhanced noise mode. This represents the signal after initial smoothing of the noise modes. This represents the residual completion amount based on semantic weights. Step S26: Modal confidence map generation. Based on the enhanced noise modes, construct a modal confidence map.

3. The intelligent guidance method for family learning trajectories integrating multimodal behavior perception as described in claim 1, characterized in that: Step S3 specifically includes the following steps: Step S31: Semantic path encoding. The task content, learning corpus, and screen text are serialized to obtain the semantic input sequence, as shown below: ; in, Indicates the first Each text unit Indicates the length of the text sequence; The semantic pathway representation is obtained through the semantic encoder, as follows: ; in, For semantic encoding functions; Step S32: Behavioral pathway feature extraction. Behavioral features are extracted from attention, operational behavior, and learning rhythm signals to obtain a behavioral feature sequence. The formula used is as follows: ; in, Indicates the first A behavioral characteristic, Represents the dimensions of behavioral characteristics; The behavioral pathway representation obtained through the behavior encoder is as follows: ; in Encoding function for behavioral features; Step S33: Semantic-behavior alignment fusion, based on semantic representation Behavioral representations are computed across pathways, and adaptively weighted and fused to obtain a unified learning state representation: ; in, This represents the learning state after fusion. For adaptive fusion weights.