Multi-mode fusion intelligent video course creation system

Through multimodal data fusion and meta-learning algorithms, a cross-modal knowledge graph is constructed, which solves the problems of insufficient practical experience and high scenario adaptation costs in traditional industrial training, and realizes an efficient and immersive industrial training solution.

CN120509997APending Publication Date: 2025-08-19GUANGZHOU YOUMI TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510583601.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Traditional industrial training courses lack practical experience and difficult to quantify learning effects. The cost of cross-scene courses is high, and they cannot quickly respond to the training needs of new equipment.

Method used

The multimodal fusion intelligent video course creation system is adopted to obtain video, text, three-dimensional models, tactile feedback and sarcotic reaction data through the multimodal data input module, and a cross-modal knowledge graph is built, combining meta-learning algorithms to quickly adapt to new industrial scenarios, output time-aligned video and text and space-registered three-dimensional models, and integrate tactile feedback effects.

Benefits of technology

It realizes an immersive multi-dimensional training experience, significantly improves learners' muscle memory and cognitive feedback on complex operations, reduces the cost of course development in new scenarios, improves training efficiency and accuracy, and adapts to the rapid response of intelligent manufacturing equipment updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509997A_ABST
    Figure CN120509997A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal fusion intelligent video course creation system, which relates to the technical field of virtual reality digital technology system playing and comprises a multi-modal data input module, a multi-modal data processing module and a multi-modal course output module. A cross-modal knowledge graph is constructed by collecting data such as an industrial equipment operation video, a maintenance manual, a three-dimensional model, tactile feedback and galvanic skin response, multi-modal data semantic equivalent conversion is realized, and a new industrial scene is quickly adapted based on a meta-learning algorithm. The output course comprises videos and texts which are aligned in time sequence and a three-dimensional model which is subjected to spatial registration, and a tactile feedback effect is fused and is optimally presented according to a galvanic skin reaction. According to the system, the problems of insufficient practical operation experience and high scene adaptation cost of traditional training are solved, the training accuracy and efficiency are improved, courses can be applied to video monitoring, virtual reality equipment, intelligent consumption equipment and the like, and the system is suitable for industrial training, medical teaching and other scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality digital technology production and playback technology, and specifically to a multi-modal fusion intelligent video course creation system. Background Art

[0002] Traditional industrial training courses mainly rely on two-dimensional videos and text manuals, which have the problems of lack of practical experience and difficult to quantify learning effects. For example, the timing correlation accuracy of equipment operation processes is insufficient, the spatial positioning of complex parts is vague, and it is impossible to capture learners' cognitive difficulties in real time. Although existing technologies attempt to improve training effects through three-dimensional modeling or single modality conversion, they lack the fusion processing of multi-dimensional data such as tactile feedback and physiological reactions, making it difficult to meet the demand for high-precision and immersive training in the era of intelligent manufacturing. In addition, cross-scenario course adaptation requires reliance on a large amount of labeled data, which has high development costs and long cycles, and cannot quickly respond to the training needs of new equipment.

[0003] In view of this, this application is hereby filed. Summary of the Invention

[0004] The purpose of the present invention is to provide a multimodal fusion intelligent video course creation system to solve the problems raised in the above background technology.

[0005] To solve the above technical problems, the present invention provides a multimodal fusion intelligent video course creation system, which includes a multimodal data input module, a multimodal processing module, and a multimodal course output module:

[0006] Multimodal data input module: used to obtain multimodal original course data containing at least video, text, and three-dimensional models. The multimodal original course data includes industrial equipment operation videos, maintenance manual text, and three-dimensional models of parts; it is also used to obtain tactile feedback data generated by force feedback gloves during simulated experimental operations, and physiological sensors to monitor the skin electrodermal response data of students when watching courses.

[0007] Multimodal processing module: connected to the multimodal data input module, used to perform semantic-perceptual consistency processing on the multimodal original course data, including:

[0008] Semantic-perceptual joint representation unit: used to build a unified cross-modal representation space to enhance the semantic association of multimodal data and transform perceptual channels, including:

[0009] Cross-modal knowledge processing subunit: used to construct a cross-modal knowledge graph based on knowledge in the fields of industrial equipment operation procedures, parts assembly relationships, etc., and generate scene semantic constraint vectors. The cross-modal knowledge graph contains temporal association rules between video frames and text paragraphs, and spatial mapping rules between three-dimensional models and video images; at the same time, it combines the tactile feedback data of force feedback gloves and the galvanic skin response data of physiological sensors to further optimize the perceptual association rules in the knowledge graph.

[0010] Perception channel conversion subunit: used to perform semantic equivalent conversion of the multimodal original course data between visual, textual and device operation action channels through a multi-task learning pre-training model. The conversion includes converting device operation video frames into textual operation steps, or mapping three-dimensional model size parameters to video annotations; and can convert the tactile data of the force feedback glove simulation operation into the corresponding video teaching scene annotations, and adjust the weight strategy of modal conversion according to the skin electrode response data.

[0011] Dynamic Scenario Adaptation Unit: This unit is used to dynamically adapt the processing of the semantic-perceptual joint representation unit based on the configuration parameters of target application scenarios such as industrial training and equipment maintenance, including:

[0012] Scenario parameter configuration subunit: used to define pluggable scenario configuration files, which contain semantic constraint parameters specific to the industrial field (such as the timing accuracy threshold of operation steps and the spatial positioning error threshold of components) and perception conversion rules (such as the terminology mapping table of equipment action videos and maintenance manuals); at the same time, analysis rules and application thresholds for tactile feedback data and galvanic skin response data can be set.

[0013] Rapid adaptation operator method subunit: used to optimize the initial parameters of the semantic-perceptual joint representation unit based on the meta-learning algorithm, so that it can complete adaptation in new industrial scenarios (such as different types of equipment, different maintenance processes) through a small number of samples; and can quickly adjust the model parameters according to the tactile and physiological sensor data characteristics in the new scenario to adapt to the immersive course creation needs in the new scenario.

[0014] Multimodal course output module: Connected to the multimodal processing module, it is used to output adapted multimodal course content, including industrial equipment operation training courses and parts maintenance guides. The course content includes a time-aligned video and text combination, and a spatially registered 3D model and video combination. It can also integrate the tactile feedback effects of force feedback gloves simulated operation into the course presentation, and optimize the presentation of course content based on the results of galvanic skin response analysis.

[0015] A comprehensive and immersive data collection system has been built, breaking through the limitations of traditional industrial training in terms of practical experience and quantitative evaluation of learning outcomes. By incorporating tactile and physiological feedback data, it provides a more authentic and diverse information foundation for course creation, helping to create course content that is more relevant to actual operations and precisely adapted to the learner's state.

[0016] Furthermore, the cross-modal knowledge processing subunit is specifically used to:

[0017] The text of the industrial maintenance manual is semantically parsed to extract equipment component entities (such as "bearing" and "bolt"), operation action entities (such as "disassembly" and "installation"), and parameter entities (such as "torque 15N·m" and "accuracy level"), and generate knowledge graph nodes containing entities, relationships, and attributes. At the same time, the actions corresponding to the force feedback glove simulation operation are associated with the equipment component entities, and the course nodes corresponding to abnormal skin conduction response data are marked.

[0018] An industrial equipment maintenance knowledge graph is constructed based on the knowledge graph nodes. The knowledge graph contains timing constraints of equipment operations (such as "inspection" must be after "disassembly") and spatial constraints (such as "the distance between the bolt installation position and the bearing center is 50mm±2mm"); and adds perceptual constraint relationships constructed based on tactile feedback data and galvanic skin response data.

[0019] The industrial equipment maintenance knowledge graph is encoded into a scene semantic constraint vector and input into the Transformer encoding layer to adjust the attention weight when extracting features from the industrial equipment operation video frames and the maintenance manual text. Simultaneously, the semantic constraint vector is modified based on tactile and physiological sensor data, expanding multimodal data association from a purely semantic level to a deep fusion of "semantic-perceptual-physiological" dimensions. This significantly improves the accuracy of the industrial operation process modeling, especially in complex operations. It can accurately locate learners' knowledge weaknesses and operational difficulties based on tactile feedback and physiological response data, making the course design more targeted and focused.

[0020] Furthermore, the sensing channel conversion subunit is specifically configured to:

[0021] For video frames of industrial equipment operation, a multi-task learning model is used to generate text-formatted operating step instructions or a three-dimensional model motion trajectory of the equipment action; and the tactile feedback data generated by the force feedback glove simulation operation can be converted into corresponding text descriptions or three-dimensional model motion features.

[0022] By combining contrastive learning with domain knowledge constraints, the semantic equivalence of the converted step text and the original video actions is ensured. This semantic equivalence is quantitatively assessed through the temporal accuracy of the operational process (e.g., step sequence accuracy ≥ 95%). The matching degree between tactile feedback data and galvanic skin response data is also incorporated into the semantic equivalence evaluation metric. This approach breaks the limitation of traditional industrial video training, which relies solely on two-dimensional visual presentation, and cleverly transforms tactile feedback data into visual or interactive course elements, such as intuitive 3D model vibration effects and vivid textual tactile descriptions. This allows learners to enhance their memory and understanding of the operational process through multiple sensory channels. Furthermore, by incorporating physiological data as an evaluation basis, the conversion quality between different modalities can be dynamically calibrated, effectively ensuring the accuracy and effectiveness of course information delivery.

[0023] Furthermore, the industrial scene configuration file defined by the scene parameter configuration subunit includes:

[0024] Semantic constraint parameters are used to limit the semantic association accuracy between industrial equipment operation videos and maintenance manuals, including:

[0025] Timing correlation accuracy threshold (e.g., the timestamp deviation between the video operation step and the text step is ≤ 500ms);

[0026] Spatial positioning accuracy threshold (e.g., the position error between the 3D model component and the video image is ≤ 3mm);

[0027] Tactile feedback and operation action matching accuracy threshold (for example, the matching error between the strength and position of the tactile feedback and the actual operation action is ≤ 10%);

[0028] The correlation threshold between galvanic skin response data and course content (such as the correlation accuracy between the galvanic skin response fluctuation amplitude and the key knowledge points of the course ≥ 80%).

[0029] Perception conversion rules are used to define the conversion mapping relationship of multimodal data in the industrial field, including:

[0030] A mapping table of video frames of equipment operation actions and maintenance manual terminology (e.g., "rotate the wrench clockwise" corresponds to the text "torque calibration");

[0031] Automatic association rules between component 3D model size parameters and video annotations (e.g., automatic annotation of bearing inner diameter parameters on video disassembly images);

[0032] Mapping rules between force feedback glove tactile feedback data and video operation scenarios (e.g., tactile feedback of a specific force corresponds to the tightening or loosening operation scenario of a device component);

[0033] Mapping rules between galvanic skin response data characteristics and course content adjustment strategies (such as increased galvanic skin response corresponds to increased course explanations or repetitions of key steps); establishing a multimodal fusion standardized evaluation system suitable for industrial scenarios. This system accurately quantifies the matching degree of tactile feedback and the correlation between physiological data and course content. It can flexibly and accurately adjust various parameters according to the operating characteristics of different industrial equipment (such as the high-precision requirements of precision instrument operation and the strong force perception needs of heavy machinery operation), so that the course strictly follows industrial technical specifications and is highly consistent with learners' cognitive laws, significantly improving the practicality and applicability of the course.

[0034] Furthermore, the fast adaptation operator method subunit is based on the Model-Agnostic Meta-Learning (MAML) algorithm, specifically used to:

[0035] For maintenance training scenarios of new industrial equipment (such as robots of different models and CNC machine tools), the initial parameters of the semantic-perception joint representation unit are gradient updated by inputting 20-100 equipment operation video clips and brief maintenance instructions; at the same time, the tactile feedback samples of the force feedback gloves corresponding to the operation of the new equipment and the galvanic skin response samples of students when watching the new equipment course are input to jointly optimize the model parameters.

[0036] Generate a dedicated processing model adapted to new industrial scenarios, so that the accuracy of the timing correlation of the operation steps of the dedicated processing model in the new scenario is increased to more than 92%, and the spatial positioning error of the parts is reduced to within 5mm; and the matching accuracy of tactile feedback and operating actions reaches more than 90%, and the accuracy of the correlation analysis between skin electrodermal response data and course content reaches more than 85%; with the help of the meta-learning mechanism that integrates tactile and physiological data, when facing the development of new industrial equipment courses, the required sample size can be greatly reduced (compared to the traditional method, it is reduced by more than 60%), significantly reducing development costs. At the same time, ensure the continuity and stability of the immersive interactive experience in the new scenario, so that the system can quickly adapt to the characteristics of frequent equipment updates in the era of intelligent manufacturing, effectively improve the efficiency of enterprise training, and provide strong support for enterprises to quickly cultivate talents who can adapt to new technologies and equipment.

[0037] Furthermore, the multimodal original course data also includes device sensor data (such as torque sensor values, displacement sensor coordinates) and audio explanation data. The sensor data is used to assist in verifying the compliance of device operations (such as whether the torque value meets the manual standards). At the same time, force feedback gloves and physiological sensor data are used to assess students' immersion in and understanding of the course. By building a complete closed loop of "operation process-result verification", device sensor data can verify in real time whether the operation meets technical standards, ensuring the accuracy and reliability of the course content at the technical level. The data collected by force feedback gloves and physiological sensors provides an objective and quantitative learning effect evaluation dimension from the perspective of learners' operating experience and psychological reactions, laying a solid data foundation for the subsequent implementation of personalized course recommendations and adaptive learning models.

[0038] Furthermore, the industrial field course content output by the multimodal course output module includes:

[0039] A temporally aligned device operation video and maintenance manual combination, where the temporal correlation accuracy between video frames and text paragraphs is ≥92%;

[0040] A combination of a spatially registered 3D model of a component and an operation video, where the component position error between the 3D model and the video image is ≤3mm;

[0041] Linking equipment sensor data with operating steps, such as displaying in real time on the video screen whether the torque sensor value meets the parameter range in the maintenance manual;

[0042] Demonstrate the tactile feedback effect of force feedback gloves during simulated operations, such as displaying the corresponding tactile feedback intensity and pattern during key operation steps in the course;

[0043] The course content is dynamically adjusted based on the results of galvanic skin response analysis, such as adding prompts or repeatedly playing relevant operation clips in areas where students experience abnormal galvanic skin response. This achieves a comprehensive and in-depth integration of "visual-text-tactile-physiological feedback." Displaying tactile feedback effects during key course operations, such as simulating the resistance felt when tightening a bolt, can effectively promote the formation of muscle memory in learners and deepen their grasp of key operational points. Real-time dynamic adjustment of course content presentation based on physiological data can keenly capture changes in learners' cognitive states, providing timely and targeted assistance, significantly improving the learning efficiency and mastery of complex industrial skills. The accuracy of memory for maintenance procedures can be increased by approximately 35%.

[0044] Furthermore, the force feedback glove features at least five force feedback points, capable of simulating at least five different intensities of force feedback, with a force feedback response time of ≤50ms and a force feedback intensity error of ≤5%. The physiological sensor monitors galvanic skin response data in real time, with a sampling frequency of ≥10Hz and a data transmission delay of ≤100ms. The high-precision force feedback glove ensures a high degree of realism in the simulation, while the fast response time and minimal intensity error prevent learners from misleading due to distorted tactile feedback. The high-sampling-frequency physiological sensor accurately captures learners' momentary cognitive fluctuations, providing real-time, reliable data support for course optimization. This hardware-level guarantee ensures the system's efficiency and stability in both immersive teaching experiences and learning outcome monitoring, enhancing the overall system's practicality and application value.

[0045] The creation method of the multimodal fusion intelligent video course creation system includes the following steps:

[0046] S1 Data Acquisition: This module uses a multimodal data input module to acquire industrial equipment operation videos, maintenance manual texts, component 3D models, force feedback glove tactile feedback data, and galvanic skin response data from physiological sensors.

[0047] S2 Graph Construction: Utilize the cross-modal knowledge processing subunit in the multimodal processing module to construct a cross-modal knowledge graph, and optimize the association rules in the graph based on tactile feedback data and galvanic skin response data;

[0048] S3 Modal Conversion: Through the perceptual channel conversion subunit, the multimodal original course data is semantically equivalently converted between different perceptual channels, and tactile and physiological sensor data are integrated into the conversion process;

[0049] S4 Scenario Adaptation: With the help of the dynamic scenario adaptation unit, based on the parameters and rules in the industrial scenario configuration file, the fast adaptation operator method subunit optimizes the parameters of the semantic-perception joint representation unit to adapt to the new industrial scenario;

[0050] S5 course output: The multimodal course output module outputs multimodal course content that integrates tactile feedback effect display and optimization based on skin electrical response; forming a complete closed loop of "data collection-knowledge modeling-modal conversion-scenario adaptation-intelligent output", upgrading industrial course creation to a data-driven model, improving development standardization and intelligence, and suitable for high-precision practical scenarios.

[0051] The multimodal courses produced by the multimodal fusion intelligent video course creation system are applied to at least one of video surveillance processing equipment, interactive televisions, virtual reality / digital technology playback equipment, and smart consumer devices. This multimodal course is adaptable to a wide range of application scenarios and can be used for course presentation and learning on a variety of common devices. On video surveillance processing equipment, it can be used for real-time monitoring and training review of industrial operating specifications; interactive televisions facilitate group learning and interactive communication in corporate training rooms, classrooms, and other places; virtual reality / digital technology playback equipment can create a highly immersive learning environment and enhance learners' practical experience; and smart consumer devices facilitate learners to conduct fragmented learning anytime, anywhere, greatly expanding the scope of course dissemination and audience, and improving the popularity and accessibility of industrial training education.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] 1. Building an immersive multimodal experience: Force feedback gloves are used to simulate tactile operation (such as bolt resistance and bearing engagement), and physiological sensors are used to capture galvanic skin responses (such as cognitive fluctuations corresponding to learning difficulties). These are deeply integrated with industrial equipment operation videos and three-dimensional models to build a full-dimensional "visual-text-tactile-physiological feedback" curriculum system. This breaks through the experience limitations of traditional two-dimensional training and strengthens learners' muscle memory of complex operations and real-time cognitive feedback.

[0054] 2. Precise adaptation and efficient development for industrial scenarios: Based on the meta-learning algorithm (MAML), only 20-100 new equipment operation samples (including tactile / electrodermal data) are required to complete model optimization, improving the timing correlation accuracy to over 92% and the tactile matching error to ≤10%. Compared with traditional solutions, this reduces sample requirements by 60%, significantly lowering the cost of developing new scenario courses and quickly responding to training needs for equipment updates in intelligent manufacturing.

[0055] 3. Multi-dimensional data-driven course optimization: Cross-modal knowledge graphs integrate the temporal constraints of industrial processes (such as "inspection must be performed after disassembly"), spatial constraints (such as three-dimensional model positioning error ≤ 3mm), and perceptual constraints (such as the association between tactile feedback and operating status). Combined with galvanic skin response, the course content is dynamically adjusted (such as automatic replay of difficult points), realizing a closed loop of "data collection-modeling-optimization". This has increased the accuracy of maintenance step memory by 35% and the pass rate of operation assessments from 65% to 92%.

[0056] 4. Full-scenario terminal coverage and interactive upgrade: The generated multimodal courses can be adapted to video surveillance equipment (real-time review of operating specifications), interactive TVs (group teaching), virtual reality equipment (immersive training) and smart consumer devices (fragmented learning). Through 3D model linkage (error ≤ 2mm) and real-time annotation of sensor data (such as dynamic verification of torque values), they can meet the diversified interactive needs of scenarios such as industrial training and medical surgery teaching, and expand the scope of course dissemination and practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 The principle block diagram of the multimodal fusion intelligent video course creation system;

[0058] Figure 2 A flowchart of the creation method of the multimodal fusion intelligent video course creation system;

[0059] Figure 3 This is a schematic diagram of the structure of a configuration file example in the first embodiment of the present invention;

[0060] Figure 4 This is a structural diagram of an example configuration file in the second embodiment of the present invention. DETAILED DESCRIPTION

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0062] See also Figure 1-4 The present invention provides a technical solution: a multimodal fusion intelligent video course creation system, comprising:

[0063] Example 1:

[0064] In the industrial scenario: Creation of training courses on industrial robot bearing replacement.

[0065] 1. Implementation of Multimodal Data Input Module

[0066] Scenario: Developing a training course for bearing replacement operations for a KUKA KR1000 industrial robot.

[0067] Data collection steps:

[0068] Video data:

[0069] A video of the entire robotic arm bearing replacement process (1080p, 30fps) was recorded, covering the four key steps of "disassembling the robotic arm housing → removing the old bearing → installing the new bearing → torque calibration". The timestamps are 00:01:00-00:03:30, 00:03:30-00:05:00, 00:05:00-00:07:00, and 00:07:00-00:08:30 respectively.

[0070] The tactile data of the force feedback gloves simulated operations were collected: resistance feedback when removing bolts (vibration frequency corresponding to 20 N·m torque) and positioning feedback when installing bearings (5 levels of force).

[0071] Text data:

[0072] Import the maintenance manual PDF and extract text sections such as "Bearing model 6205-2RS", "Bolt torque 15N·m±2N·m", and "Safety operating specifications".

[0073] Synchronously collect students' skin electrical response data while watching the course (through the Empatica E4 bracelet, sampling frequency 10Hz), and mark the skin electrical fluctuation points corresponding to the operational difficulties (such as the skin electrical increase area in the torque calibration link).

[0074] 3D model data:

[0075] Import the STL model of the robotic arm and bearing, and mark the bearing installation position coordinates (X=150mm, Y=80mm, Z=60mm) and the three-dimensional dimensions of the bolt hole.

[0076] Sensor data:

[0077] The real-time data of the torque sensor (sampling frequency 100 Hz) is connected to record the torque value change curve during the operation.

[0078] 2. Implementation of Semantic-Perceptual Joint Representation Unit

[0079] 2.1 Cross-modal Knowledge Processing Subunit

[0080] step:

[0081] 1. Text semantic analysis:

[0082] Extract entities using the spaCy industrial extension library:

[0083] Component entities: "Robotic arm housing", "6205-2RS bearing", "M8×30 bolt";

[0084] Action entities: "disassemble", "remove", "install", "calibrate";

[0085] Parametric entities: "Torque 15 N·m" and "Bearing inner diameter 25 mm".

[0086] Associating tactile data with entities: Mapping the "bolt removal" action to the 20 N·m resistance feedback of the force feedback glove. Marking the "torque calibration" step corresponding to the electrodermal fluctuation points is a learning difficulty.

[0087] 2. Knowledge graph construction:

[0088] Build timing constraints: "Remove old bearing" must be after "Disassemble housing" (interval ≤ 2 minutes); "Calibrate torque" must be after "Install bearing" (interval ≤ 1 minute).

[0089] Build spatial constraints: The center distance tolerance between the bearing installation position and the bolt hole position is ±0.5mm, and the spatial error between the 3D model and the video image is ≤3mm.

[0090] Increase perception constraints: Associate the tactile feedback intensity (such as 5-level locking tactile sensation) with the "bearing installed in place" state, and associate the area of increased skin electricity with the "torque calibration method" knowledge point.

[0091] 3. Constraint vector generation:

[0092] The knowledge graph is encoded into a 2048-dimensional vector and input into the Transformer encoding layer, so that the model prioritizes images related to "torque calibration" (such as the torque wrench beeping prompt frame) when processing video frames.

[0093] 2.2 Perception Channel Conversion Subunit

[0094] step:

[0095] 1. Video-to-text / model conversion:

[0096] The YOLOv8 model is used to identify the “disassembly housing” action in the video and generate text steps: “Use an 8mm socket to rotate the housing bolt counterclockwise until it is completely removed (tactile feedback: 20N·m resistance)”.

[0097] The motion trajectory of the three-dimensional model of the bearing installation (XYZ coordinate changes) is superimposed on the video screen, and the bearing inner diameter parameter "φ25mm" is marked.

[0098] 2. Tactile-visual mapping:

[0099] The 5-level snap-in tactile sensation of the force feedback glove is converted into highlighted annotations on the video screen (for example, a green aperture flashes on the edge of the model when the bearing is snapped in).

[0100] 3. Semantic equivalence assessment:

[0101] Through comparative learning, the timing accuracy of text steps and video actions is ensured to reach 95%, and the matching error between tactile feedback and operation actions is verified to be ≤10% (such as the deviation between actual torque value and feedback intensity).

[0102] 3. Implementation of Dynamic Scene Adaptation Unit

[0103] 3.1 Scene Parameter Configuration Subunit

[0104] Reference Figure 3 The following is a sample configuration file (JSON).

[0105] 3.2 Fast Adaptation Algorithm Unit

[0106] step:

[0107] Meta-learning initialization: Pre-training the semantic-perceptual joint representation unit based on 200 video samples of 10 historical bearing replacement cases.

[0108] Adapting to a new scenario: 20 Kuka KR1000 bearing replacement operation clips (including tactile and electroskin data) were input, and the model parameters were updated using the MAML algorithm to achieve:

[0109] The accuracy of time series correlation increased from 85% to 93%;

[0110] Haptic feedback matching accuracy increased from 80% to 92%;

[0111] The accuracy of skin electrodermal response correlation analysis increased from 75% to 88%.

[0112] 4. Implementation of Multimodal Course Output Module

[0113] Course content and interaction design:

[0114] 1. Time alignment of video and text:

[0115] When the video reaches the "Bearing Installation" step (00:05:00), the text on the right automatically locates to "Section 4.2 Bearing Installation Specifications" and highlights the paragraph "Torque calibration requires the use of a special wrench."

[0116] 2. 3D model and video linkage:

[0117] A three-dimensional model of the robotic arm is embedded in the lower right corner of the video screen, and the bearing installation path is displayed in real time (the red arrow marks the motion trajectory). The position error between the model and the video components is ≤2mm.

[0118] 3. Haptic feedback effect display:

[0119] When the video plays to the "bolt removal" segment, the force feedback glove synchronously triggers a 20N·m resistance vibration, the duration of which is consistent with the bolt rotation time in the video (error ≤ 50ms).

[0120] 4. Adjustment of skin electric drive content:

[0121] When the system detects an increase in the student's galvanic skin response (such as in the torque calibration phase), it automatically marks the yellow area on the video progress bar and pushes a "Torque Calibration Special Exercise" clip after class.

[0122] 5. Sensor data annotation:

[0123] The video screen is superimposed with the torque sensor curve. When the measured torque value falls within the range of 13-17N·m, a green prompt box is displayed, and when it exceeds the range, a red warning is marked.

[0124] 5. Effect Verification and Optimization

[0125] Timing accuracy: A random sample of 50 video-text pairs showed that the timestamp deviation of 48 pairs was ≤500ms, achieving a compliance rate of 96%.

[0126] Spatial accuracy: Measured by PolyWorks, the average positioning error of the 3D model is 1.8mm, meeting the requirement of ≤3mm.

[0127] Tactile matching: In students' subjective feedback, 90% believed that the tactile feedback was consistent with the actual resistance feeling during operation.

[0128] Learning efficiency: Tests at an automobile factory showed that the pass rate for trainees in bearing replacement operations increased from 65% to 92%, and the average operation time was shortened by 40%.

[0129] Summary of the embodiments:

[0130] This example integrates industrial equipment operation videos, maintenance manuals, 3D models, tactile feedback, and physiological data to build a complete process of "data acquisition - knowledge modeling - modal conversion - scenario adaptation - intelligent output". The core innovations include:

[0131] 1. Multi-dimensional data fusion: Incorporating tactile feedback (force feedback gloves) and physiological responses (electrodermal data) into course creation can solve the problems of "lack of operational experience" and "ambiguous identification of learning difficulties" in traditional training.

[0132] 2. Dynamic scene adaptation: Through meta-learning algorithms, new devices (such as different models of robots) can be quickly adapted, reducing sample size requirements by 60% and significantly reducing course development costs.

[0133] 3. Immersive interactive design: Through 3D model linkage, tactile feedback synchronization, and electrodermal-driven content adjustment, students' muscle memory and cognitive efficiency for complex operations are improved.

[0134] The system can be directly reused in industrial scenarios such as CNC machine tool maintenance and automated production line operations. By adjusting the scenario configuration files and knowledge graphs, customized training courses can be quickly generated.

[0135] Example 2:

[0136] In the field of medical teaching videos: Creation of training courses for liver tumor resection surgery.

[0137] 1. Implementation of Multimodal Data Input Module

[0138] Scenario: Develop a liver tumor resection surgical training course for hepatobiliary surgeons, focusing on the core process of "tumor localization → liver pedicle management → tumor resection → resection margin assessment".

[0139] Data collection steps:

[0140] Video data:

[0141] Recording of 4K live video of laparoscopic surgery (60fps), including key steps such as "ultrasound positioning of tumor", "blocking the hepatic artery", "tumor resection", and "pathological section analysis", with time stamps of 00:10:00-00:15:00, 00:15:00-00:25:00, 00:25:00-00:40:00, and 00:40:00-00:45:00 respectively.

[0142] Collect tactile data of the force feedback surgery simulator: tissue resistance feedback when the ultrasound probe contacts the liver (5-10N pressure touch) and force feedback when vascular clamping (3-level positioning feeling).

[0143] Text data:

[0144] Import the PDF of "Guidelines for Liver Tumor Resection Surgery" and extract surgical standard texts such as "tumor safe resection margin ≥ 1 cm" and "hepatic artery occlusion time ≤ 15 minutes".

[0145] The electrodermal response data of medical students while watching the course (sampling frequency 10 Hz) were collected synchronously, and the electrodermal fluctuation points of complex steps such as "liver pedicle dissection" were marked.

[0146] Medical imaging data:

[0147] The patient's preoperative CT / MRI images (1 mm slice thickness) were imported, and the liver and tumor models were reconstructed in three dimensions, with the tumor coordinates (X = 120 mm, Y = 75 mm, Z = 50 mm) and vascular distribution marked.

[0148] Sensor data:

[0149] Access the displacement sensor data of the surgical simulator to record the movement trajectory of the ultrasound probe and the tumor positioning error.

[0150] 2. Implementation of Semantic-Perceptual Joint Representation Unit

[0151] 2.1 Cross-modal Knowledge Processing Subunit

[0152] step:

[0153] 1. Semantic analysis of medical texts:

[0154] Extract entities using the MedSpaCy tool:

[0155] Anatomical entities: "right lobe of the liver tumor", "hepatic artery", "portal vein";

[0156] Action entities: "ultrasound positioning", "blocking", "resection", "suturing";

[0157] Pathological parameters: "Resection margin 1.5 cm from tumor" and "Blocking time 12 minutes".

[0158] Associating tactile data with entities: Mapping the "ultrasound positioning tumor" action with the 5N resistance tactile sensation of force feedback, and marking the "vascular anatomical variation identification" corresponding to the area of increased skin electrical activity as a teaching difficulty.

[0159] 2. Construction of medical knowledge graph:

[0160] Construct timing constraints: "Tumor resection" must be performed after "liver pedicle occlusion" (interval ≤ 5 minutes); "resection margin assessment" must be performed after "tumor resection" (interval ≤ 3 minutes).

[0161] Construct spatial constraints: ultrasound probe positioning error ≤ 2 mm, and spatial error between tumor resection boundary and CT model ≤ 1 mm.

[0162] Increase perceptual constraints: associate the tactile feedback intensity (such as level 3 locking sensation) with the "vascular clamp in place" state, and associate the galvanic skin response with the "anatomical structure identification" knowledge point.

[0163] 3. Constraint vector generation:

[0164] The knowledge graph is encoded into a 1024-dimensional vector and input into the Transformer encoding layer, so that the model prioritizes images related to "vascular variation identification" (such as ultrasound images and surgical field of view superimposed frames) when processing surgical videos.

[0165] 2.2 Perception Channel Conversion Subunit

[0166] step:

[0167] 1. Video-to-text / model conversion:

[0168] The SlowFast network was used to identify the "hepatic pedicle occlusion" action in the surgical video and generate text steps: "Use a vascular clamp to occlude the hepatic artery, and the occlusion time was recorded as 12 minutes (tactile feedback: level 3 clamping feeling)."

[0169] The tumor resection path (translucent red area) of the 3D liver model is superimposed on the video screen, and the safe resection margin range is marked.

[0170] 2. Tactile-visual mapping:

[0171] The tissue resistance tactile sensation (5N pressure) of the force feedback simulator is converted into a highlight of the ultrasound probe contact area in the video screen (blue semi-transparent mask).

[0172] 3. Semantic equivalence assessment:

[0173] Through comparative learning, the timing accuracy of the surgical step text and video actions is ensured to reach 98%, and the matching error between tactile feedback and operation actions is verified to be ≤8% (such as the deviation between clamping force and feedback intensity).

[0174] 3. Implementation of Dynamic Scene Adaptation Unit

[0175] 3.1 Scene Parameter Configuration Subunit

[0176] Reference Figure 4 , configuration file example (JSON).

[0177] 3.2 Fast Adaptation Algorithm Unit

[0178] step:

[0179] Meta-learning initialization: Pre-training the semantic-perceptual joint representation unit based on 300 video samples of 20 historical liver surgeries.

[0180] Adapting to new scenarios: We input 30 operation clips (including tactile and electrodermal data) of a new laparoscopic surgery in a hospital and updated the model parameters using the MAML algorithm to achieve:

[0181] The accuracy of temporal correlation increased from 88% to 96%;

[0182] Haptic feedback matching accuracy increased from 85% to 94%;

[0183] The accuracy of skin electrodermal response correlation analysis increased from 80% to 90%.

[0184] 4. Implementation of Multimodal Course Output Module

[0185] Course content and interaction design:

[0186] 1. Timing alignment of video and guide text:

[0187] When the video plays to the "Tumor Resection" step (00:25:00), the guide on the right automatically locates to "Section 3.3 Principles of Tumor Resection" and highlights the paragraph "Resection along 1 cm outside the tumor capsule."

[0188] 2. Linkage between medical imaging and surgical videos:

[0189] A real-time CT / MRI fusion image is embedded on the left side of the video screen, showing the three-dimensional spatial relationship between the ultrasound probe position and the tumor (error ≤ 1.5 mm), and the resection path is marked with a red arrow.

[0190] 3. Haptic feedback effect display:

[0191] When the video plays to the "ultrasound localization of tumors" segment, the force feedback glove synchronously simulates 5N tissue resistance vibration, and the duration is consistent with the probe movement time in the video (error ≤30ms).

[0192] 4. Adjustment of skin electric drive content:

[0193] The system detected that the medical students' skin electricity increased during the "liver pedicle anatomy" session, automatically inserted a 3D vascular anatomy animation (showing the variant course of the hepatic artery), and pushed virtual anatomy exercises for this session after class.

[0194] 5. Pathology data annotation:

[0195] The video is superimposed with real-time pathology section analysis results (such as assessment of resection margin cell atypia) and displayed synchronously with the surgical step timestamps.

[0196] 5. Effect Verification and Optimization

[0197] Timing accuracy: A random sampling of 40 surgical video-guideline text pairs revealed that 39 of them had timestamp deviations ≤ 300ms, achieving a compliance rate of 97.5%.

[0198] Spatial accuracy: Verified by 3DSlicer software, the average tumor positioning error was 0.8 mm, meeting the requirement of ≤1 mm.

[0199] Tactile matching: In the subjective feedback of medical students, 88% believed that tactile feedback truly reflected the resistance of surgical operation.

[0200] Learning efficiency: A test at a medical university showed that the pass rate of students in liver tumor resection operation assessment increased from 70% to 95%, and the time to master complex steps (such as vascular anatomy) was shortened by 50%.

[0201] Summary of the embodiments:

[0202] This embodiment targets medical teaching scenarios and builds a high-precision surgical training course creation process by integrating surgical videos, medical images, tactile feedback, and physiological data. The core innovations include:

[0203] 1. Deep fusion of medical images: Through submillimeter registration of CT / MRI and surgical videos, the spatial cognitive bias between anatomical structures and actual operations in traditional teaching can be resolved.

[0204] 2. Pathology-operation closed loop: Synchronize pathology slide data and surgical steps in real time, strengthen the "operation-evaluation" connection, and improve medical students' ability to predict surgical efficacy.

[0205] 3. Intelligent response to difficulties: Dynamically insert anatomical animations and virtual exercises based on galvanic skin response to achieve a real-time closed loop of "cognitive difficulties-teaching intervention", significantly improving the efficiency of learning complex skills.

[0206] The system can be reused in medical training scenarios such as orthopedic surgery and endoscopic diagnosis and treatment. By adjusting the medical knowledge graph and tactile feedback parameters, it can quickly generate immersive teaching content that meets clinical standards.

[0207] In summary: The present invention integrates industrial equipment operation videos, maintenance manuals, three-dimensional models, tactile feedback, and physiological data through a multimodal data input module, uses a semantic-perceptual joint representation unit to construct a cross-modal knowledge graph and implement modal conversion, and combines the parameter configuration of the dynamic scene adaptation unit with the meta-learning algorithm to quickly generate immersive training courses adapted to different industrial scenarios. Through the closed loop of "data acquisition-knowledge modeling-modal conversion-scene adaptation-intelligent output", the system has achieved the upgrade of industrial training courses from experience-driven to data-driven, significantly improving course development efficiency and learning effects. It is suitable for high-precision scenarios such as robot maintenance and surgical training, and can be widely used through multiple terminal devices.

Claims

1. A multimodal fusion intelligent video course creation system, including a multimodal data input module, a multimodal processing module, and a multimodal course output module, is characterized by: Multimodal data input module: acquires multimodal original course data including video, text, and 3D models, as well as tactile feedback data from force feedback gloves and galvanic skin response data from physiological sensors; Multimodal processing module: performs semantic-perceptual consistency processing on multimodal data, including: Semantic-perception joint representation unit: Constructs a unified cross-modal representation space, including: Cross-modal knowledge processing sub-unit: Constructs a cross-modal knowledge graph based on industry domain knowledge, generates scene semantic constraint vectors, and optimizes perceptual association rules by combining tactile and electrodermal data; Perception channel conversion sub-unit: A multi-task learning model is used to achieve semantic equivalent conversion between multimodal data channels, converting tactile data into video annotations, and adjusting modal conversion weights based on electrodermal data; Dynamic scene adaptation unit: This unit configures parameter adaptation processing based on industrial scene configuration, including: scene parameter configuration sub-unit: defines scene configuration files containing semantic constraint parameters and perception conversion rules, and sets analysis rules for tactile / electrodermal data; fast adaptation operator method sub-unit: based on meta-learning algorithms, optimizes model parameters through a small number of samples to adapt to the needs of immersive course creation in new industrial scenes; Multimodal course output module: Outputs multimodal courses containing time-aligned videos and texts, and spatially registered 3D models and videos, incorporating tactile feedback effects and optimizing course presentation based on galvanic skin response.

2. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: The cross-modal knowledge processing subunit is used to parse the text of the industrial maintenance manual, extract equipment components, operating actions, parameter entities and generate knowledge graph nodes, associate tactile actions with equipment components, and mark skin electrodermal abnormality course nodes; construct an industrial equipment maintenance knowledge graph containing temporal / spatial constraints and tactile / skin electrodermal perception constraints; encode the knowledge graph into a semantic constraint vector and modify it based on physiological data.

3. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: The perception channel conversion subunit is used to convert industrial equipment operation video frames into at least one of text steps and three-dimensional model motion trajectories, and convert tactile data into text / 3D model features; use contrastive learning and domain knowledge constraints to ensure semantic equivalence, and incorporate tactile / electrodermal data matching into evaluation indicators.

4. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: In the definition of the scene parameter configuration subunit, the industrial scene configuration file includes semantic constraint parameters and perception conversion rules.

5. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: In the fast adaptation operator method subunit, based on the MAML algorithm, by inputting 20-100 new device operation videos and tactile / skin electrical samples, the model parameters are updated to generate a special model adapted to the new scenario.

6. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: The multimodal original course data also includes device sensor data and audio lecture data. Sensor data assists in verifying operational compliance, and tactile / electrical skin data evaluates learning immersion.

7. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: In the multimodal course output module, the output course content includes time-aligned videos and manuals, spatially registered three-dimensional models and videos, sensor data linkage annotation, tactile feedback effect display, and dynamic course adjustment driven by galvanic skin response.

8. The multimodal fusion intelligent video course creation system according to claim 1, characterized in that: The force feedback gloves have ≥5 force feedback points, simulate ≥5 force feedback effects of different intensities, with a response time of ≤50ms and an intensity error of ≤5%; the physiological sensor monitors skin electrical data in real time, with a sampling frequency of ≥10Hz and a transmission delay of ≤100ms.

9. The creation method of the multimodal fusion intelligent video course creation system according to any one of claims 1 to 8, characterized in that: The following steps are involved: S1 Data Acquisition: Acquire industrial equipment operation videos, maintenance manual texts, 3D models, tactile feedback data, and galvanic skin response data; S2 Graph Construction: Construct a cross-modal knowledge graph and optimize association rules based on tactile / electrodermal data; S3 Modal Conversion: Achieve semantic equivalent conversion between multimodal data channels and integrate tactile / physiological data; S4 Scenario Adaptation: Optimize model parameters based on industrial scenario configuration parameters to adapt to new scenarios; S5 course output: Output multimodal courses that integrate tactile feedback and skin electrical optimization.

10. The multimodal course output by the multimodal fusion intelligent video course creation system according to any one of claims 1 to 8, characterized in that: Applicable to at least one of video surveillance processing equipment, interactive television, virtual reality / digital technology playback equipment, and smart consumer equipment.

Citation Information

Patent Citations

  • Puncture surgery teaching system, implementation method, teaching terminal and teaching device

    CN110807968A

  • Virtual practical training method for automobile vocational education

    CN119477615A

  • Fan fault simulation and diagnosis training system and method

    CN119723977A

  • Quality-oriented education creative course visualization method and device based on artificial intelligence

    CN119809883A

  • Dynamic data pipeline construction method based on artificial intelligence and multi-modal data processing

    CN119830200A