Digital skill intelligent learning platform based on multi-modal interaction
Through the multi-modal input receiving unit, a cross-modal intention fusion engine and an adaptive feedback system, the problems of single interaction, lag in feedback and rigid learning paths in the digital skills learning platform are solved, and intelligent fusion of multi-modal inputs and personalized learning path planning are realized, which improves learning efficiency and experience.
Patent Information
- Application Number
- CN202510555203.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing digital skills learning platforms lack multimodal perception, cross-modal fusion, adaptive feedback and knowledge structured representation, resulting in a single interaction method, lagging feedback mechanism, and rigid learning paths, and being unable to accurately identify user operation intentions, increasing learning difficulty and reducing efficiency.
The multi-modal input receiving unit, a cross-modal intention fusion engine, a multi-dimensional adaptive feedback system and a digital skill knowledge graph are adopted to realize intelligent fusion of multi-modal input, real-time adaptive feedback and personalized learning path planning. Through dynamic weight allocation, conflict detection and feedback modal combination, it provides accurate operation intention recognition and personalized learning experience.
Improve the efficiency and experience of digital skills learning, and through intelligent fusion of multimodal input and real-time adaptive feedback, ensure the accuracy of operational intent recognition, and enhance the immersion and learning efficiency of the learning experience.
Smart Images

Figure CN120491924A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital skills learning technology, and in particular to a digital skills intelligent learning platform based on multimodal interaction. Background Art
[0002] Although technologies such as artificial intelligence, multimodal interaction and knowledge graphs have made significant progress, their comprehensive application in the field of digital skills learning is still in its early stages; existing technical solutions mostly focus on the application of a single technology, lacking the systematic integration of multimodal perception, cross-modal fusion, adaptive feedback and structured knowledge representation, and are unable to comprehensively solve the complex problems in digital skills learning.
[0003] Traditional learning platforms lack intelligent intent understanding mechanisms and are unable to accurately identify user operational intent. This is especially true in complex task scenarios, often requiring users to perform tedious steps, increasing learning difficulty and frustration. When users simultaneously utilize multiple input methods, the system struggles to effectively resolve conflicts between modalities, leading to increased misjudgment rates. Existing platforms often utilize pre-set, fixed learning paths, ignoring differences in user skill sets and learning characteristics. These platforms are unable to intelligently adjust learning content and difficulty based on real-time user performance, resulting in low learning efficiency.
[0004] Therefore, there is an urgent need for a digital skills intelligent learning platform that can comprehensively utilize multimodal interaction, cross-modal fusion, adaptive feedback and knowledge graph technology to provide a personalized, efficient and socially recognized digital skills learning experience. Summary of the Invention
[0005] The present invention provides a digital skills intelligent learning platform based on multimodal interaction, which is used to solve the problems of single interaction mode, lagging feedback mechanism, rigid learning path and insufficient modal conflict recognition ability in existing digital skills learning platforms, and realizes intelligent fusion of multimodal input, real-time adaptive feedback and data-driven personalized learning path planning, thereby improving the efficiency, experience and effect of digital skills learning.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] A digital skills intelligent learning platform based on multimodal interaction, the platform comprising: a multimodal input receiving unit, a cross-modal intention fusion engine, a multi-dimensional adaptive feedback system and a digital skills knowledge graph.
[0008] The multimodal input receiving unit is used to collect and pre-process at least two modal signals among the user's voice command signal, gesture operation signal, eye tracking signal and tactile input signal.
[0009] The cross-modal intention fusion engine includes: a weight distribution unit and an operation intention recognition unit.
[0010] The weight allocation unit allocates dynamic weights to each input modality under the current learning task based on the user's historical interaction data, where the sum of the weights of each modality is 1 and all are non-negative values; the weights are calculated by comprehensively considering the historical accuracy of each modality, task complexity parameters and modality priority coefficient.
[0011] The operation intention recognition unit calculates a composite operation instruction according to each modal signal and its corresponding weight.
[0012] The multi-dimensional adaptive feedback system includes a visual feedback unit, a tactile feedback unit and a voice feedback unit, and adjusts the feedback mode combination according to the real-time error rate.
[0013] The digital skill knowledge graph includes a set of skill points, a set of associated edges between skill points, and an association strength weight matrix, which is used to generate personalized learning trajectories in combination with user interaction behavior data.
[0014] Furthermore, the cross-modal intent fusion engine includes: a modal conflict detection unit, which is used to identify conflicts between input instructions of different modalities within the same time window, and to determine whether the distance between expected operation vectors of different modalities exceeds a preset conflict threshold by calculating the distance between them.
[0015] When there is a conflict between different modalities, the corresponding instructions are executed based on the priority of the modalities; the priority of the modalities is sorted from large to small according to the user's historical usage preference for the modalities.
[0016] Furthermore, by calculating the distance D between the expected operation vectors of different modes c (S i ,S j ) The specific method for determining whether the preset conflict threshold is exceeded is: through the function: D c (S i ,S j )=||O(S i )-O(S j )||, quantifies the distance between the expected operation vectors of different modes within the same time window Δt, O(S i ) and O(S j ) represent the expected operation vectors of different modes i and j respectively. When D c 9S i ,S j 0>θ c , it is determined that there is a conflict between different modes.
[0017] θ c is the conflict threshold parameter:
[0018] in, represents the historical conflict root mean square value between modality i and modality j in the past n interactions; w i and w j are the current weights of mode i and mode j respectively; C task is the task complexity factor, which is used to increase the threshold tolerance in complex tasks; C task is the task complexity factor: E task Indicates the estimated number of steps to complete the current learning task, ranging from 1 to 10, B task Indicates the task branching factor, that is, the average number of operation choices per step, ranging from 1 to 5; C task Taking into account the length and width complexity of the task, the range is [0.1, 5].
[0019] α is the historical fluctuation coefficient, β is the weight influence coefficient, and γ is the task complexity coefficient;
[0020]
[0021] Among them, λ α =0.05, is the learning rate parameter; N user It represents the number of interactions of the user in the system. It increases with the user's experience in using the system and has a value range of [0.5, 0.8]. The initial value for new users is 0.5. As the user's experience grows, the system gradually increases its dependence on the user's historical behavior pattern. user The value gradually increases; ACC avg represents the average recognition accuracy of all modalities, with a value range of [0.2, 0.4]; D task Indicates the depth level of the current task in the digital skills knowledge graph. This coefficient increases with the complexity of the task and ranges from [0, 0.3].
[0022] Furthermore, the weight allocation unit dynamically allocates the weights of each input modality in the following manner:
[0023] For the input modality set M=m1,m2,...,m k , the corresponding weight vector is W=w1,w2,...,w k , satisfying the constraints And w s ≥0, k is the total number of input modalities; the weight calculation model is defined as: Among them, φ(m s ) is the mode m s The comprehensive scoring function.
[0024] Among them, ACC s Represents mode ms The historical accuracy of P is in the range of [0,1]; s Represents mode m s The priority coefficient is determined by user preference and modality applicability, and its value range is [0,1]. s Represents mode m s The error rate is obtained by counting the last 100 input modes; the system also sets the weight adjustment frequency: Indicates that each execution of f a Update the weight distribution once after each operation.
[0025] Furthermore, the operation intention recognition unit calculates the composite operation instruction in the following manner:
[0026] D c (S i ,S j )≤θ c When it is determined that there is no conflict between different modes, the signal set corresponding to the input mode set M is: S=S1,S2,...,S k .
[0027] For each modal signal S s , through the feature extraction function: Extract n s dimensional feature vector;
[0028] Through the mode-specific mapping function: Φ s :F s →O s , the mode m s The eigenvector F s Mapping to the operation vector space O s , where Φ s (F s )=σ(W s ·F s +b s ), W s and b s are the mapping weight matrix and bias vector respectively, and σ is the activation function.
[0029] Unified operation vector space conversion function T s :O s →O u , convert the operation vectors of different modes into a unified operation vector space to ensure dimensional consistency;
[0030] Compound operation instruction O co The calculation formula is: By the decision function D(O co )→A, the compound operation instruction O coMapped to the final operation instruction A, where the decision function D( co ) is expressed as: Among them, V A is the standard operation vector corresponding to operation A, <O co ,V A > is the inner product operation of two vectors; A collection of all operation instructions supported by the platform.
[0031] Furthermore, the visual feedback unit displays dynamic prompt information and error operation marks on the screen, providing graphical guidance when the user's operation deviates from the expected path, including color change prompts, trajectory comparison display and step completion visualization;
[0032] The tactile feedback unit transmits operation correctness information through vibration and force feedback devices, adjusts the vibration frequency and intensity according to the operation accuracy, provides tactile confirmation at key operation points, and simulates the correct operation feel through force feedback.
[0033] The voice feedback unit provides real-time guidance and error correction suggestions through voice commands, uses natural language to analyze user confusion points, provides targeted voice prompts based on error types, and provides encouraging voice feedback when completing staged tasks;
[0034] The system continuously monitors the user's operation error rate, and automatically enhances the feedback intensity when the error rate exceeds the preset threshold. It dynamically adjusts the feedback ratio of each modality according to different task types and operation stages to obtain an adjusted feedback modality combination.
[0035] Furthermore, the adjustment model of the feedback modality combination in the multi-dimensional adaptive feedback system is as follows: assuming the current error rate is E(t), the system calculates the weight vector of each modality feedback W(t) = [wv(t), wt(t), ws(t)], where wv(t), wt(t), and ws(t) are the weight values of visual, tactile, and voice feedback, respectively; when E(t)>θE, the weight vector W(t+1) = W(t) + h·(E(t)-θE)·ΔW is updated, where θE is the error rate threshold, h = 0.05, and ΔW is the gradient vector.
[0036] Furthermore, the method of combining the digital skill knowledge graph with user interaction behavior data to generate personalized learning trajectories is as follows: the system records the user's interaction behavior data during the skill point practice process, including completion time, number of errors, operation fluency and modal usage preference; analyzes the user's mastery of different skill points and identifies weak links in skills; calculates the optimal learning path based on the correlation strength weight matrix between skill points to ensure that the learning order of related skill points is reasonable; dynamically adjusts the difficulty of learning content and recommended practice methods based on the user's skill mastery characteristics; combines the user's historical learning trajectory to predict skill development trends and actively recommends the next stage of learning goals.
[0037] Furthermore, the model for generating the personalized learning trajectory is:
[0038] Define the user skill state vector H(u) = [h1,h2,...], where h i Indicates the user's mastery of skill point i, with a value range of [0,1]. Based on the user interaction behavior feature set F(u) = v1, v2, v3, v4, the skill state H is predicted through a deep neural network function.
[0039] The user interaction behavior feature set F(u)=v1, v2, v3, v4; v1 is the operation speed index, v2 is the operation accuracy, v3 is the number of repeated attempts, and v4 is the frequency of seeking help.
[0040] The set of optional learning tasks is represented as: B = b1, b2, ..., the transition function is P(H′|H, b), which represents the probability that the user will transition from skill state H to state H′ after performing action b; the reward function R(H, b) = ΔH, where ΔH is the skill improvement. The iterative optimization strategy is: Among them, H′ is the new state after executing action b, is the maximum Q value in the new state.
[0041] Generate personalized learning trajectories based on the knowledge graph structure and user interaction behavior data:
[0042] Among them, π * (H) is the optimal personalized learning trajectory strategy, that is, the action to be selected in state H; E b is the skill point edge set involved in action b; w_{ij} is the association strength from skill point i to skill point j; μ is the knowledge graph structure influencing factor.
[0043] Furthermore, the platform sets up a skill certification and social incentive mechanism, and issues skill NFT certificates after users complete a specific skill learning path.
[0044] The method for generating skill NFT certificates is as follows: the platform uses the tamper-proof skill certificate generator to generate the skill NFT certificate through the function N(u,s,a)=Ha(E(I u ||T s ||D a ||r)) Generate skill NFT certificate, where I u is the user identifier, I u is the skill evaluation result tensor, D a is the hash value of the operation behavior data, r is a random number, and E is the asymmetric encryption function E(m)=m e modn, Ha is the SHA-256 hash function, and || represents a concatenation operation.
[0045] Compared with the prior art, the present invention has the following beneficial effects:
[0046] The present invention can dynamically adjust the weights of different modalities based on the user's historical interaction data, and resolve the contradictions that may arise from multimodal input through precise conflict detection and processing mechanisms, thereby ensuring the accuracy of operation intention recognition; intelligently adjust the feedback modality combination according to the user's real-time error rate, provide instant and accurate operation guidance, and enhance the immersive learning experience; realize the intelligent fusion of multimodal input, real-time adaptive feedback and data-driven personalized learning path planning, thereby improving the efficiency of digital skills learning, and has great promotion value. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 This is a schematic diagram of the composition of the digital skills intelligent learning platform based on multimodal interaction of the present invention. DETAILED DESCRIPTION
[0048] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of the present invention are described clearly and completely below. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of the present invention.
[0049] It should be noted that the multimodal interaction paradigm of this invention is not only applicable to digital skills learning, but can also be expanded to multiple fields such as professional skills training, distance education, and assisted learning for special groups. The technical solution of this invention relies on modern artificial intelligence, knowledge graphs, reinforcement learning, and multimodal interaction technologies to achieve an intelligent, personalized, and immersive learning experience.
[0050] Hereinafter, the various components of the present invention and their working principles will be described in detail with reference to the accompanying drawings.
[0051] like Figure 1As shown, the digital skills intelligent learning platform based on multimodal interaction of the present invention includes: a multimodal input receiving unit, a cross-modal intention fusion engine, a multi-dimensional adaptive feedback system and a digital skills knowledge graph.
[0052] The multimodal input receiving unit is used to collect and pre-process at least two modal signals among the user's voice command signal, gesture operation signal, eye tracking signal and tactile input signal.
[0053] The multimodal input receiving unit utilizes a multi-channel parallel processing architecture to collect and preprocess user input signals in different modalities. Voice command signals are collected by a microphone array and then processed for noise reduction, signal enhancement, and voice activity detection. Gesture operation signals are collected by a depth camera and then subjected to skeleton point extraction and action recognition. Eye tracking signals are captured by an infrared eye tracker to capture user gaze points and pupil changes. Tactile input signals are detected by pressure sensors and touch screens to detect physical contact. Each modal signal utilizes a separate preprocessing pipeline to ensure signal quality and real-time processing. For mobile device scenarios, the system automatically tailors and adjusts the received modality based on available sensors.
[0054] The cross-modal intent fusion engine includes: a weight allocation unit and an operation intention recognition unit; the system maintains a user modality usage history database, recording the usage frequency, accuracy and operation efficiency of different modalities in various learning tasks; weight allocation adopts a sliding window method, taking into account both long-term historical performance and recent interaction trends, and ensuring the smoothness and dynamic adaptability of weight adjustment through weighted averaging; for new users, the system adopts a cold start strategy and initializes based on the average modality preference of demographically similar users.
[0055] The weight allocation unit allocates dynamic weights to each input modality under the current learning task based on the user's historical interaction data, where the sum of the weights of each modality is 1 and all are non-negative values; the weights are calculated by comprehensively considering the historical accuracy of each modality, task complexity parameters and modality priority coefficient.
[0056] The weight allocation unit dynamically allocates the weights of each input modality in the following manner:
[0057] For the input modality set M=m1,m2,...,m k , the corresponding weight vector is W=w1,w2,...,w k , satisfying the constraints And w s ≥0, k is the total number of input modalities; the weight calculation model is defined as: Among them, φ(m s ) is the mode m sComprehensive scoring function of
[0058] Among them, ACC s Represents mode m s The historical accuracy of P is in the range of [0,1]; s Represents mode m s The priority coefficient is determined by user preference and modality applicability, and its value range is [0,1]; E s Represents mode m s The error rate is obtained by counting the last 100 input modes; the system also sets the weight adjustment frequency: Indicates that each execution of f a Update the weight distribution once after each operation.
[0059] The operation intention recognition unit calculates a composite operation instruction according to each modal signal and its corresponding weight.
[0060] The operation intention recognition unit calculates the composite operation instruction in the following manner:
[0061] D c (S i ,S j )≤θ c When it is determined that there is no conflict between different modes, the signal set corresponding to the input mode set M is: S=S1,S2,...,S k ;
[0062] For each modal signal S s , through the feature extraction function: Extract n s dimensional feature vector;
[0063] Through the mode-specific mapping function: Φ s :F s →O s , the mode m s The eigenvector F s Mapping to the operation vector space O s , where Φ s (F s )=σ(W s ·F s +b s ), W s and b s are the mapping weight matrix and bias vector respectively, and σ is the activation function;
[0064] Unified operation vector space conversion function T s :O s →O u, convert the operation vectors of different modes into a unified operation vector space to ensure dimensional consistency;
[0065] Compound operation instruction O co The calculation formula is: By the decision function D(O co )→A, the compound operation instruction O co Mapped to the final operation instruction A, where the decision function D( co ) is expressed as: Among them, V A is the standard operation vector corresponding to operation A, <O co ,V A > is the inner product operation of two vectors; A collection of all operation instructions supported by the platform.
[0066] The cross-modal intention fusion engine includes: a modal conflict detection unit, which is used to identify conflicts between input instructions of different modalities within the same time window, and to determine whether the distance between expected operation vectors of different modalities exceeds a preset conflict threshold by calculating the distance between them.
[0067] The modal conflict detection unit adopts a two-level detection architecture.
[0068] The first level is timing detection, which identifies concurrent modal instructions through a configurable time window; the second level is semantic detection, which analyzes the operational meaning of different modal instructions. Figure 1 Consistency.
[0069] The conflict threshold parameters are dynamically adjusted according to the user's proficiency. Expert users set lower thresholds to accurately capture subtle conflicts, while beginners set higher thresholds to improve fault tolerance. When a modal conflict is detected, the system not only executes the dominant modal instructions based on priority, but also explains the conflict situation to the user through a multi-dimensional feedback system to help the user understand and adjust their operating behavior.
[0070] When there is a conflict between different modalities, the corresponding instructions are executed based on the priority of the modalities; the priority of the modalities is sorted from large to small according to the user's historical usage preference for the modalities.
[0071] By calculating the distance D between the expected operation vectors of different modes c (S i ,S j ) The specific method for determining whether the preset conflict threshold is exceeded is: through the function: D c (S i ,S j )=||O(S i )-O(S j )||, quantifies the distance between the expected operation vectors of different modes within the same time window Δt, O(Si ) and O(S j ) represent the expected operation vectors of different modes i and j respectively. When D c (S i ,S j )>θ c When , it is determined that there is a conflict between different modes;
[0072] θ c is the conflict threshold parameter:
[0073] in, represents the historical conflict root mean square value between modality i and modality j in the past n interactions; w i and w j are the current weights of mode i and mode j respectively; C task is the task complexity factor, which is used to increase the threshold tolerance in complex tasks; C task is the task complexity factor: E task Indicates the estimated number of steps to complete the current learning task, ranging from 1 to 10, B task Indicates the task branching factor, that is, the average number of operation choices per step, ranging from 1 to 5; C task Taking into account the length and width complexity of the task, the range is [0.1, 5];
[0074] α is the historical fluctuation coefficient, β is the weight influence coefficient, and γ is the task complexity coefficient;
[0075]
[0076] Among them, λ α =0.05, is the learning rate parameter; N user It represents the number of interactions of the user in the system. It increases with the user's experience in using the system and has a value range of [0.5, 0.8]. The initial value for new users is 0.5. As the user's experience grows, the system gradually increases its dependence on the user's historical behavior pattern. user The value gradually increases; ACC avg represents the average recognition accuracy of all modalities, with a value range of [0.2, 0.4]; D task Indicates the depth level of the current task in the digital skills knowledge graph. This coefficient increases with the complexity of the task and ranges from [0, 0.3].
[0077] The multi-dimensional adaptive feedback system includes a visual feedback unit, a tactile feedback unit, and a voice feedback unit, and adjusts the feedback mode combination according to the real-time error rate. The visual feedback unit displays dynamic prompt information and error operation marks on the screen, and provides graphical guidance when the user operation deviates from the expected path, including color change prompts, trajectory comparison display, and step completion visualization. The tactile feedback unit transmits operation correctness information through vibration and force feedback devices, adjusts the vibration frequency and intensity according to the operation accuracy, provides tactile confirmation at key operation points, and simulates the correct operation feel through force feedback. The voice feedback unit provides real-time guidance and error correction suggestions through voice commands, uses natural language to analyze user confusion points, provides targeted voice prompts based on the error type, and provides motivational voice feedback when completing stage tasks.
[0078] The speech modality uses a dual-path processing approach, combining acoustic and semantic features; the gesture modality combines static posture features and dynamic trajectory features; the eye movement modality analyzes gaze distribution and shift patterns; and the tactile modality extracts pressure, direction, and duration features. Feature vectors are converted to a unified operational vector space using a pre-trained modality-specific mapping network, achieving cross-modal semantic alignment.
[0079] The system continuously monitors the user's operation error rate, and automatically enhances the feedback intensity when the error rate exceeds the preset threshold. It dynamically adjusts the feedback ratio of each modality according to different task types and operation stages to obtain an adjusted feedback modality combination.
[0080] The adjustment model of the feedback modality combination in the multidimensional adaptive feedback system is as follows: assuming the current error rate is E(t), the system calculates the weight vector of each modality feedback W(t) = [wv(t), wt(t), ws(t)], where wv(t), wt(t), and ws(t) are the weight values of visual, tactile, and voice feedback, respectively; when E(t)>θE, the weight vector W(t+1) = W(t) + h·(E(t)-θE)·ΔW is updated, where θE is the error rate threshold, h = 0.05, and ΔW is the gradient vector.
[0081] The digital skill knowledge graph includes a set of skill points, a set of associated edges between skill points, and an association strength weight matrix, which is used to generate personalized learning trajectories in combination with user interaction behavior data.
[0082] The method of combining the digital skill knowledge graph with user interaction behavior data to generate a personalized learning trajectory is as follows: the system records the user's interaction behavior data during skill point practice, including completion time, number of errors, operation fluency and modal usage preference; analyzes the user's mastery of different skill points and identifies weak links in skills; calculates the optimal learning path based on the correlation strength weight matrix between skill points to ensure that the learning order of related skill points is reasonable; dynamically adjusts the difficulty of learning content and recommends practice methods based on the user's skill mastery characteristics; combines the user's historical learning trajectory to predict skill development trends and actively recommends learning goals for the next stage.
[0083] The model for generating the personalized learning trajectory is:
[0084] Define the user skill state vector H(u) = [h1,h2,...], where h i Indicates the user's mastery of skill point i, with a value range of [0,1]. Based on the user interaction behavior feature set F(u) = v1, v2, v3, v4, the skill state H is predicted through a deep neural network function.
[0085] The user interaction behavior feature set F(u) = v1, v2, v3, v4; v1 is the operation speed index, v2 is the operation accuracy, v3 is the number of repeated attempts, and v4 is the frequency of seeking help;
[0086] The set of optional learning tasks is represented as: B = b1, b2, ..., the transition function is P(H′|H, b), which represents the probability that the user will transition from skill state H to state H′ after performing action b; the reward function R(H, b) = ΔH, where ΔH is the skill improvement. The iterative optimization strategy is: Among them, H′ is the new state after executing action b, is the maximum Q value in the new state;
[0087] Generate personalized learning trajectories based on the knowledge graph structure and user interaction behavior data:
[0088] Among them, π * (H) is the optimal personalized learning trajectory strategy, that is, the action to be selected in state H; E b is the skill point edge set involved in action b; w_{ij} is the association strength from skill point i to skill point j; μ is the knowledge graph structure influencing factor.
[0089] The platform sets up a skill certification and social incentive mechanism, and issues skill NFT certificates after users complete a specific skill learning path;
[0090] The method for generating skill NFT certificates is as follows: the platform uses the tamper-proof skill certificate generator to generate the skill NFT certificate through the function N(u,s,a)=Ha(E(I u ||T s ||D a ||r)) Generate skill NFT certificate, where I u is the user identifier, I u is the skill evaluation result tensor, D a is the hash value of the operation behavior data, r is a random number, and E is the asymmetric encryption function E(m)=m e modn, Ha is the SHA-256 hash function, and || represents a concatenation operation.
[0091] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A digital skills intelligent learning platform based on multimodal interaction, characterized by: The platform includes: a multimodal input receiving unit, a cross-modal intent fusion engine, a multi-dimensional adaptive feedback system, and a digital skills knowledge graph; The multimodal input receiving unit is used to collect and pre-process at least two modal signals among the user's voice command signal, gesture operation signal, eye tracking signal and tactile input signal; The cross-modal intention fusion engine includes: a weight distribution unit and an operation intention recognition unit; The weight assignment unit assigns dynamic weights to each input modality under the current learning task based on the user's historical interaction data, where the sum of the weights of each modality is 1 and all are non-negative values; the weights are calculated by comprehensively considering the historical accuracy of each modality, task complexity parameters, and modality priority coefficients; The operation intention recognition unit calculates a composite operation instruction based on each modal signal and its corresponding weight; The multi-dimensional adaptive feedback system includes a visual feedback unit, a tactile feedback unit and a voice feedback unit, and adjusts the feedback mode combination according to the real-time error rate; The digital skill knowledge graph includes a set of skill points, a set of associated edges between skill points, and an association strength weight matrix, which is used to generate personalized learning trajectories in combination with user interaction behavior data.
2. The digital skills intelligent learning platform based on multimodal interaction according to claim 1 is characterized in that: The cross-modal intention fusion engine includes: a modal conflict detection unit for identifying conflicts between input instructions of different modalities within the same time window, and determining whether the distance between expected operation vectors of different modalities exceeds a preset conflict threshold by calculating the distance between them; When there is a conflict between different modalities, the corresponding instructions are executed based on the priority of the modalities; the priority of the modalities is sorted from large to small according to the user's historical usage preference for the modalities.
3. The digital skills intelligent learning platform based on multimodal interaction according to claim 2 is characterized in that: By calculating the distance D between the expected operation vectors of different modes c (S i ,S j ) The specific method for determining whether the preset conflict threshold is exceeded is: through the function: D c (S i ,S j )=||O(S i )-O(S j )||, quantifies the distance between the expected operation vectors of different modes within the same time window Δt, O(S i ) and O(S j ) represent the expected operation vectors of different modes i and j respectively. When D c (S i ,S j )>θ c When , it is determined that there is a conflict between different modes; θ c is the conflict threshold parameter: in, represents the historical conflict root mean square value between modality i and modality j in the past n interactions; w i and w j are the current weights of mode i and mode j respectively; C task is the task complexity factor, which is used to increase the threshold tolerance in complex tasks; C task is the task complexity factor: E task Indicates the estimated number of steps to complete the current learning task, ranging from 1 to 10, B task Indicates the task branching factor, that is, the average number of operation choices per step, ranging from 1 to 5; C task Taking into account the length and width complexity of the task, the range is [0.1, 5]; α is the historical fluctuation coefficient, β is the weight influence coefficient, and γ is the task complexity coefficient; Among them, λ α =0.05, is the learning rate parameter; N user It represents the number of interactions of the user in the system. It increases with the user's experience in using the system and has a value range of [0.5, 0.8]. The initial value for new users is 0.
5. As the user's experience grows, the system gradually increases its dependence on the user's historical behavior pattern. user The value gradually increases; ACC avg represents the average recognition accuracy of all modalities, with a value range of [0.2, 0.4]; D task Indicates the depth level of the current task in the digital skills knowledge graph. This coefficient increases with the complexity of the task and ranges from [0, 0.3].
4. The digital skills intelligent learning platform based on multimodal interaction according to claim 2 is characterized in that: The weight allocation unit dynamically allocates the weights of each input modality in the following manner: For the input modality set M=m1,m2,...,m k , the corresponding weight vector is W=w1,w2,...,w k , satisfying the constraints and k is the total number of input modalities; the weight calculation model is defined as: Among them, φ(m s ) is the mode m s Comprehensive scoring function of Among them, ACC s Represents mode m s The historical accuracy of P is in the range of [0,1]; s Represents mode m s The priority coefficient is determined by user preference and modality applicability, and its value range is [0,1]; E s Represents mode m s The error rate is obtained by counting the last 100 input modes; the system also sets the weight adjustment frequency: Indicates that each execution of f a Update the weight distribution once after each operation.
5. The digital skills intelligent learning platform based on multimodal interaction according to claim 4 is characterized in that: The operation intention recognition unit calculates the composite operation instruction in the following manner: D c (S i ,S j )≤θ c When it is determined that there is no conflict between different modes, the signal set corresponding to the input mode set M is: S=S1,S2,...,S k ; For each modal signal S s , through the feature extraction function: Extract n s dimensional feature vector; Through the mode-specific mapping function: Φ s :F s →O s , the mode m s The eigenvector F s Mapping to the operation vector space O s , where Φ s (F s )=σ(W s ·F s +b s ), W s and b s are the mapping weight matrix and bias vector respectively, and σ is the activation function; Unified operation vector space conversion function T s :O s →O u , convert the operation vectors of different modes into a unified operation vector space to ensure dimensional consistency; Compound operation instruction O co The calculation formula is: By the decision function D(O co )→A, the compound operation instruction O co Mapped to the final operation instruction A, where the decision function D( cO ) is expressed as: Among them, V A is the standard operation vector corresponding to operation A, <O co ,V A > is the inner product operation of two vectors; A collection of all operation instructions supported by the platform.
6. The digital skills intelligent learning platform based on multimodal interaction according to claim 5 is characterized in that: The visual feedback unit displays dynamic prompts and error operation marks on the screen, providing graphical guidance when the user's operation deviates from the expected path, including color change prompts, trajectory comparison display, and step completion visualization; The tactile feedback unit transmits operation correctness information through vibration and force feedback devices, adjusts the vibration frequency and intensity according to the operation accuracy, provides tactile confirmation at key operation points, and simulates the feel of correct operation through force feedback; The voice feedback unit provides real-time guidance and error correction suggestions through voice commands, uses natural language to analyze user confusion points, provides targeted voice prompts based on error types, and provides encouraging voice feedback when completing staged tasks; The system continuously monitors the user's operation error rate, and automatically enhances the feedback intensity when the error rate exceeds the preset threshold. It dynamically adjusts the feedback ratio of each modality according to different task types and operation stages to obtain an adjusted feedback modality combination.
7. The digital skills intelligent learning platform based on multimodal interaction according to claim 6 is characterized in that: The adjustment model of the feedback modality combination in the multidimensional adaptive feedback system is as follows: assuming the current error rate is E(t), the system calculates the weight vector of each modality feedback W(t) = [wv(t), wt(t), ws(t)], where wv(t), wt(t), and ws(t) are the weight values of visual, tactile, and voice feedback, respectively; when E(t)>θE, the weight vector W(t+1) = W(t) + h·(E(t)-θE)·ΔW is updated, where θE is the error rate threshold, h = 0.05, and ΔW is the gradient vector.
8. The digital skills intelligent learning platform based on multimodal interaction according to claim 1 is characterized in that: The method for generating personalized learning trajectories by combining the digital skill knowledge graph with user interaction behavior data is as follows: the system records the user's interaction behavior data during skill point practice, including completion time, number of errors, operation fluency, and modality usage preference; Analyze the user's mastery of different skills and identify weak links in skills; Calculate the optimal learning path based on the correlation strength weight matrix between skill points to ensure the learning order of related skill points is reasonable; Based on the user's skill mastery characteristics, the difficulty of learning content and recommended practice methods are dynamically adjusted; combined with the user's historical learning trajectory, skill development trends are predicted, and the next stage of learning goals are actively recommended.
9. The digital skills intelligent learning platform based on multimodal interaction according to claim 8 is characterized in that: The model for generating the personalized learning trajectory is: Define the user skill state vector H(u) = [h1,h2,...], where h i Indicates the user's mastery of skill point i, with a value range of [0,1]. Based on the user interaction behavior feature set F(u) = v1, v2, v3, v4, the skill state H is predicted through a deep neural network function. The user interaction behavior feature set F(u) = v1, v2, v3, v4; v1 is the operation speed index, v2 is the operation accuracy, v3 is the number of repeated attempts, and v4 is the frequency of seeking help; The set of optional learning tasks is expressed as: B = b1, b2, ..., and the transfer function is P(G ′ |G,b), represents the probability of the user transferring from skill state G to state H′ after performing action b; reward function R(G,b) = ΔH; ΔH is the skill improvement; through iterative optimization strategy: Among them, H′ is the new state after executing action b, is the maximum Q value in the new state; Generate personalized learning trajectories based on the knowledge graph structure and user interaction behavior data: Among them, π * (H) is the optimal personalized learning trajectory strategy, that is, the action to be selected in state H; E b is the skill point edge set involved in action b; w_{ij} is the association strength from skill point i to skill point j; μ is the knowledge graph structure influencing factor.
10. The digital skills intelligent learning platform based on multimodal interaction according to claim 1 is characterized in that: The platform sets up a skill certification and social incentive mechanism, and issues skill NFT certificates after users complete a specific skill learning path; The method for generating skill NFT certificates is as follows: the platform uses the tamper-proof skill certificate generator to generate the skill NFT certificate through the function N(u,s,a)=Ha(E(I u ||T s ||d a ||r)) Generate skill NFT certificate, where I u is the user identifier, I u is the skill evaluation result tensor, D a is the hash value of the operation behavior data, r is a random number, and E is the asymmetric encryption function E(m)=m e modn, Ha is the SHA-256 hash function, and || represents a concatenation operation.
Citation Information
Patent Citations
Navigation type experiment interaction device with cognitive function
CN110286763A
Multi-mode interactive intelligent control system
CN118226967A
Visual interaction system based on multiple modes
CN118535023A
Knowledge association learning method and system based on knowledge graph and virtual reality
CN119166830A
Intelligent service dialogue method and system fusing generative large model and knowledge graph
CN119226469A
Cited By
Old-age care robot interaction interface design method based on multiple modes
CN121349585A
Multi-modal interactive teaching simulation training system
CN121415659A
A multi-modal interactive teaching simulation training system
CN121415659B