Intelligent analysis method, device and medium for user teaching based on landscape large model

By using spatiotemporal alignment of multimodal data and similarity calculation of scene topology, the shortcomings of user operation logic and accuracy diagnosis in intelligent teaching systems are solved, and high-precision teaching feedback and skills training are achieved.

CN122365356APending Publication Date: 2026-07-10GUANGDONG VCOM EDUCATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUANGDONG VCOM EDUCATION TECH
Filing Date
2026-04-11
Publication Date
2026-07-10

Smart Images

  • Figure CN122365356A_ABST
    Figure CN122365356A_ABST
Patent Text Reader

Abstract

This application relates to the fields of artificial intelligence and education technology, and in particular to a user-based intelligent teaching analysis method, device, and medium based on a large-scale user action graph model. The method includes: responding to the execution of a teaching task; acquiring and determining semantic anchors based on multimodal interaction data generated by the user; obtaining an aligned multimodal data stream based on the semantic anchors; acquiring a first feature set from the multimodal data stream and analyzing it to obtain a second feature set; determining dynamic weight coefficients and processing the second feature set to obtain a third feature set; constructing a user action graph topology based on the first, second, and third feature sets and calculating similarity; determining branch diagnosis results based on the similarity and the third feature set and generating action improvement suggestions; and associating the action improvement suggestions with the user action graph topology to generate a user analysis report. This application has the effect of improving the accuracy of user action diagnosis and enhancing the precision of teaching feedback during skills training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of artificial intelligence and education technology, and in particular to a user teaching intelligent analysis method, device and medium based on a large-scale picture model. Background Technology

[0002] In vocational skills training and advanced skills competitions, the effectiveness of teaching operational skills such as information network cabling and fiber optic splicing highly depends on the subjective observation of teachers or coaches. Traditional intelligent teaching systems typically only record task completion time or final score, lacking in-depth understanding and quantitative analysis of the entire operational process.

[0003] Existing intelligent teaching analysis solutions typically combine video monitoring with system log entry, using a unified timestamp to reproduce and evaluate the operation process. However, due to inherent response delays between different acquisition hardware, it is difficult to achieve high-precision spatiotemporal alignment between image data, operation metadata, and device sensor data, making it impossible for the analysis system to accurately reconstruct the true state at the moment of the operation. Furthermore, existing solutions often rely on threshold judgments of single-dimensional result data, lacking in-depth analysis of the kinematic characteristics and logical topology of user operation behavior. This prevents the system from automatically identifying and tracing the causes of logical and accuracy deviations exhibited by the user during operation.

[0004] The existing technical solutions mentioned above have the following drawbacks: they cannot achieve deep intelligent diagnosis of user operation logic and accuracy based on high-precision aligned multimodal data, thus there is room for improvement. Summary of the Invention

[0005] To improve the accuracy of user operation diagnosis and enhance the precision of teaching feedback during skills training, this application provides a user teaching intelligent analysis method, device, and medium based on a large-scale picture model.

[0006] The above-mentioned objective of this application is achieved through the following technical solution:

[0007] A user-led intelligent teaching analysis method based on a large-scale picture model, the user-led intelligent teaching analysis method comprising:

[0008] In response to the execution of teaching tasks, multimodal interaction data generated by users is acquired, semantic anchors are determined based on the multimodal interaction data, and the multimodal interaction data is spatiotemporally aligned based on the semantic anchors to obtain an aligned multimodal data stream.

[0009] A first feature set is obtained from the aligned multimodal data stream, and the first feature set is analyzed to obtain a second feature set;

[0010] Determine the dynamic weight coefficients associated with the teaching task, and use the dynamic weight coefficients to perform weighted fusion processing on the second feature set to obtain the third feature set;

[0011] Based on the first feature set, the second feature set, and the third feature set, a user operation scene topology is constructed, and the similarity between the user operation scene topology and the preset standard topology is calculated.

[0012] Based on the similarity and the third feature set, the branch diagnosis result is determined, and operation improvement suggestions are generated based on the branch diagnosis result;

[0013] The operation improvement suggestions are associated with the spatiotemporal nodes in the user operation scene topology to generate a user analysis report.

[0014] By adopting the above technical solution, and by determining semantic anchors and performing spatiotemporal alignment based on multimodal interaction data, the asynchronous problem caused by hardware latency at heterogeneous data acquisition terminals can be eliminated. By analyzing the first feature set, a second feature set is obtained, enabling preliminary quantification of user operation behavior. By utilizing dynamic weight coefficients to obtain a third feature set, physical operation features can be sublimated into higher-order features representing psychological state and proficiency, thereby achieving a leap from physical to cognitive level assessment. By calculating the similarity between the user operation scene topology and the standard topology and determining the branch diagnosis results, defects can be classified and judged from two dimensions: global logic and local precision. This generates highly targeted operation improvement suggestions and significantly improves the accuracy of teaching feedback.

[0015] In a preferred embodiment, this application can be further configured such that: the multimodal interaction data includes at least image data, operational metadata, and device sensor data; the step of determining semantic anchors based on the multimodal interaction data, and performing spatiotemporal alignment of the multimodal interaction data based on the semantic anchors to obtain an aligned multimodal data stream, specifically includes:

[0016] Identify pixel feature mutation points in the image data within the multimodal interactive data, and determine the time when the pixel feature mutation points occur as the semantic anchor points;

[0017] Obtain the original timestamps of the semantic anchor points corresponding to the operation metadata and the device sensor data, and calculate the time offset of each original timestamp relative to the time of the semantic anchor point.

[0018] Based on the time offset, the operation metadata and the device sensor data are time-referenced to generate the aligned multimodal data stream.

[0019] By adopting the above technical solution, semantic anchor points are determined by identifying pixel feature mutation points in image data and calculating time offsets. This enables the establishment of a precise mapping relationship between unstructured video streams and structured logs and sensor data, thereby achieving resampling alignment of multi-source data streams and effectively solving the problem of misjudgment in the analysis of multimodal data caused by false synchronization.

[0020] In a preferred embodiment, this application can be further configured as follows: the first feature set includes at least time information, spatial coordinates, and tool identification information; the second feature set includes at least an operation rhythm stability index and a motion smoothness index; the step of obtaining the first feature set from the aligned multimodal data stream and analyzing the first feature set to obtain the second feature set specifically includes:

[0021] Based on the time information and the spatial coordinates, the instantaneous velocity characteristics and instantaneous acceleration characteristics at each corresponding moment are obtained;

[0022] Calculate the standard deviation of the instantaneous velocity characteristic within a preset sliding window, and determine the standard deviation as the operation rhythm stability index;

[0023] Calculate the rate of change of the local slope of the instantaneous acceleration feature within the preset sliding window, and determine the motion smoothness index based on the positive and negative transformation frequency of the local slope within the preset threshold range.

[0024] By adopting the above technical solution, and by performing multi-order derivative processing on the displacement vector and calculating the standard deviation and local slope change rate within the sliding window, it is possible to capture micro-motion jitters and rhythm fluctuations that are difficult to detect with the naked eye. This transforms subjective operational feel into traceable physical parameters, thereby improving the objectivity of the evaluation of operational stability and smoothness.

[0025] In a preferred embodiment, this application can be further configured such that: the third feature set includes at least a proficiency index and a cognitive load index; the determination of dynamic weight coefficients associated with the teaching task, and the weighted fusion processing of the second feature set using the dynamic weight coefficients to obtain the third feature set, specifically includes:

[0026] Based on the difficulty coefficient of the teaching task and the user's historical operation performance data, dynamic weight coefficients are assigned to the operation rhythm stability index and the action smoothness index.

[0027] The proficiency index is obtained by weighting and summing the second feature set using the dynamic weight coefficients.

[0028] The cognitive load index is obtained by obtaining the time interval between two consecutive tool identification information switching in the first feature set and performing correlation analysis between the time interval and the local slope change rate in the second feature set.

[0029] By adopting the above technical solution, and by assigning dynamic weight coefficients based on task difficulty and historical performance, and combining them with tool switching intervals for correlation analysis, it is possible to perform flexible assessment of proficiency in different teaching contexts, and accurately extract the cognitive load state of users during operation, thereby identifying whether users have cognitive obstacles such as logical retrieval lag or decision hesitation.

[0030] In a preferred embodiment, this application can be further configured as follows: The step of constructing a user operation scene topology based on the first feature set, the second feature set, and the third feature set, and calculating the similarity between the user operation scene topology and a preset standard topology, specifically includes:

[0031] Based on the tool identification information, spatial coordinates and action execution order in the first feature set, multiple operation units are identified, and an initial operation graph is constructed with each operation unit as a node and the temporal sequence and usage order between each operation unit as directed edges.

[0032] The operation rhythm stability index and the action smoothness index in the second feature set, as well as the proficiency index and the cognitive load index in the third feature set, are used as multi-dimensional attribute feature vectors and mapped to the corresponding nodes. The directed edges are weighted according to the switching time between each operation unit to generate the user operation scene topology.

[0033] Using a graph matching algorithm, the user operation scene topology is compared with the preset standard topology, and the topology similarity in the node arrangement order dimension and the attribute distribution similarity in the multi-dimensional attribute feature vector dimension corresponding to each node are calculated respectively.

[0034] By adopting the above technical solution, by identifying the operation unit to construct the topology and mapping the multi-dimensional attribute feature vectors to the nodes, the linear operation sequence can be transformed into a structured model with business logic meaning. In this way, the graph matching algorithm can be used to realize the synchronous comparison of the operation process logic and the quality of the operation nodes, which significantly improves the ability to identify systemic operation pattern defects.

[0035] In a preferred embodiment, this application can be further configured as follows: determining the branch diagnosis result based on the similarity and the third feature set, and generating operation improvement suggestions based on the branch diagnosis result, specifically includes:

[0036] Based on the topological similarity, the attribute distribution similarity, and the proficiency, the branch diagnosis result is determined, wherein the branch diagnosis result includes logical defect results and precision defect results;

[0037] If the branch diagnosis result is the logical defect result, compare the node differences and edge connection differences between the user operation scene topology and the standard topology to identify missing operation units and operation units with abnormal order.

[0038] Retrieve process reshaping suggestions that match the missing operation unit and the sequence abnormal operation unit from the preset teaching strategy library, and use the process reshaping suggestions as the operation improvement suggestions;

[0039] If the branch diagnosis result is the accuracy defect result, locate the abnormal node in the second feature set whose local slope change rate exceeds the preset residual range, and obtain the motion smoothness index corresponding to the abnormal node;

[0040] The deviation of the motion smoothness index from the standard smoothness constant is calculated. Based on the deviation, a corresponding training guide is matched from the teaching strategy library, and the training guide is used as a suggestion for improving the operation.

[0041] By adopting the above technical solution, and by executing recursive branch diagnostic logic and retrieving matching reshaping suggestions or training guidelines from the teaching strategy library, it is possible to automatically distinguish and locate whether the user's defect type is a logical error at the understanding level or an insufficient precision at the execution level, thereby providing personalized correction solutions and greatly enhancing the guiding value of feedback information.

[0042] In a preferred embodiment, this application can be further configured such that: associating the operation improvement suggestions with spatiotemporal nodes in the user operation scene topology to generate a user analysis report specifically includes:

[0043] Obtain the identification information of the missing operation unit, the sequentially abnormal operation unit, or the abnormal node corresponding to the operation improvement suggestion, and obtain the evidence data corresponding to the identification information in the aligned multimodal data stream, wherein the evidence data includes the original image slices and / or sensor feature curves;

[0044] The operation improvement suggestions are fused and encapsulated with the corresponding evidence data, and arranged according to the temporal logic of the user operation scenario topology to generate the user analysis report.

[0045] By adopting the above technical solution, and by acquiring the image slices or sensor feature curves corresponding to the identification information and fusing and encapsulating them, intuitive physical facts can be provided to support operational improvement suggestions. This makes it easier for users to quickly recall abnormal moments during the review process, thereby enhancing the credibility of the analysis report and the depth of teaching traceability.

[0046] In a preferred embodiment, this application can be further configured such that the user-teaching intelligent analysis method also includes:

[0047] Based on the user analysis report, a corresponding specialized training task is matched and pushed from a preset task library, and the specialized training module is pushed to the user terminal to guide the user to train according to the specialized training task.

[0048] By adopting the above technical solutions and establishing an automatic matching and push mechanism between analysis reports and specialized training tasks, precise reinforcement interventions can be implemented based on the diagnosed weaknesses, thereby forming a data-driven teaching analysis closed loop and accelerating the process of cultivating users' skill proficiency.

[0049] The second objective of this invention is achieved through the following technical solution:

[0050] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described user teaching intelligent analysis method based on a large picture model.

[0051] The above-mentioned objective three of this application is achieved through the following technical solution:

[0052] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described intelligent analysis method for user instruction based on a large-scale picture model.

[0053] In summary, this application includes at least one of the following beneficial technical effects:

[0054] 1. By adopting the above technical solution, by determining semantic anchors and performing spatiotemporal alignment based on multimodal interaction data, the asynchronous problem caused by hardware latency at heterogeneous data acquisition terminals can be eliminated. By analyzing the first feature set, the second feature set is obtained, achieving preliminary quantification of user operation behavior. By using dynamic weight coefficients to obtain the third feature set, physical operation features can be sublimated into higher-order features representing psychological state and proficiency, thereby achieving an evaluation leap from the physical level to the cognitive level. By calculating the similarity between the user operation scene topology and the standard topology and determining the branch diagnosis results, defects can be classified and judged from two dimensions: global logic and local precision, thereby generating highly targeted operation improvement suggestions and significantly improving the accuracy of teaching feedback.

[0055] 2. By establishing an automatic matching and push mechanism between analysis reports and specialized training tasks, precise reinforcement interventions can be implemented based on the diagnosed weaknesses, thereby forming a data-driven teaching analysis closed loop and accelerating the process of cultivating users' skill proficiency. Attached Figure Description

[0056] Figure 1 This is a flowchart illustrating the implementation of a user teaching intelligent analysis method based on a large scene model in one embodiment of this application;

[0057] Figure 2 This is a schematic diagram of the internal structure of a computer device according to an embodiment of this application. Detailed Implementation

[0058] The following embodiments will help those skilled in the art to further understand the function of this application, but do not limit this application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of this application. These all fall within the protection scope of this application.

[0059] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0060] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0061] The present application will be further described in detail below with reference to the accompanying drawings.

[0062] In one embodiment, such as Figure 1 As shown, this application discloses a user-based intelligent teaching analysis method based on a large-scale picture model, which specifically includes the following steps:

[0063] S10. In response to the execution of the teaching task, acquire the multimodal interaction data generated by the user, determine the semantic anchor point based on the multimodal interaction data, and perform spatiotemporal alignment of the multimodal interaction data based on the semantic anchor point to obtain the aligned multimodal data stream.

[0064] Specifically, when a user initiates a teaching task, multiple heterogeneous sensors are simultaneously triggered to collect data. The multimodal interactive data includes image data captured by a camera, operation metadata recorded by the operating software, and equipment sensor data uploaded by the training station sensors. By performing real-time pixel-level scanning of the image data, feature mutation points that can represent key turning points in physical actions are identified, and these mutation moments are defined as semantic anchors. Subsequently, the original timestamps corresponding to the anchor event are extracted from the operation metadata and equipment sensor data streams. The time offset of these timestamps relative to the semantic anchors is calculated, and based on this offset, time axis resampling and phase compensation are performed on all non-image data to generate an aligned multimodal data stream aligned under a unified logical time base.

[0065] S20. Obtain the first feature set from the aligned multimodal data stream, and analyze the first feature set to obtain the second feature set.

[0066] Specifically, basic data containing task execution time information, three-dimensional spatial coordinates, and current operation tool identification information are extracted in real time from the aligned data stream to form a first feature set. Mathematical operations are performed on the first feature set using preset kinematic analysis logic to extract instantaneous velocity features and instantaneous acceleration features. Then, the standard deviation of instantaneous velocity is calculated within a preset sliding window to quantify the operation rhythm stability index. At the same time, the local slope change rate of instantaneous acceleration is calculated to determine the motion smoothness index, thereby transforming the original spatiotemporal coordinates into a second feature set that characterizes the physical execution quality.

[0067] S30. Determine the dynamic weight coefficients associated with the teaching task, and use the dynamic weight coefficients to perform weighted fusion processing on the second feature set to obtain the third feature set.

[0068] Specifically, based on the difficulty coefficient of the current task and the user's historical operation performance data, the influence ratio of the operation rhythm stability index and the action smoothness index obtained from the previous steps is dynamically allocated, i.e., dynamic weight coefficients. The proficiency index is obtained by performing a weighted summation calculation on the second feature set. Furthermore, the time interval of tool identifier switching in the first feature set is extracted, and the correlation analysis is performed between the interval duration and the local slope change rate in the second feature set to derive the cognitive load index. The proficiency index and the cognitive load index together constitute the third feature set.

[0069] S40. Based on the first feature set, the second feature set and the third feature set, construct the user operation scene topology and calculate the similarity between the user operation scene topology and the preset standard topology.

[0070] Specifically, continuous operations are divided into multiple independent operation units using tool identifiers and spatial coordinates in the first feature set. An initial operation graph is constructed with each operation unit as a node and the temporal sequence as directed edges. Then, the stability and smoothness indices in the second feature set and the proficiency and cognitive load indices in the third feature set are encapsulated into multi-dimensional attribute feature vectors and mapped to the corresponding nodes. At the same time, the directed edges are weighted according to the switching time between nodes to generate a user operation graph topology. Finally, the graph matching algorithm is used to compare the structure with the expert-level standard topology to calculate the topology similarity and attribute distribution similarity.

[0071] S50. Based on similarity and the third feature set, determine the branch diagnosis results and generate operation improvement suggestions based on the branch diagnosis results.

[0072] Specifically, the similarity of the topological structure is first compared with the consistency threshold. If the similarity is not met, it is judged as a logical defect. By analyzing the differences in nodes and edges in the topological graph, missing operation units or operation units with abnormal order are identified, and suggestions for process reshaping are matched from the teaching strategy library. If the topological structure meets the standard, the similarity of attribute distribution, proficiency index and accuracy threshold are further compared. If the similarity is not met, it is judged as an accuracy defect. By locating nodes with abnormal slope in the second feature set and calculating their deviation from the standard, training guidance is matched from the strategy library, thereby generating differentiated operation improvement suggestions for different defect types.

[0073] S60. Associate the operation improvement suggestions with the spatiotemporal nodes in the user operation scene topology to generate a user analysis report.

[0074] Specifically, based on the defective nodes or anomaly identification information pointed to by the improvement suggestions, the corresponding original image slices and sensor feature curves are back-indexed from the aligned multimodal data stream as evidence data. The improvement suggestions and evidence data are encapsulated and arranged and displayed according to the time logic axis of the user operation scene topology, thereby generating a visualized user analysis report containing diagnostic conclusions, evidence support, and training guidance.

[0075] In one embodiment, in step S10, the multimodal interaction data includes at least image data, operational metadata, and device sensor data; semantic anchors are determined based on the multimodal interaction data, and the multimodal interaction data is spatiotemporally aligned based on the semantic anchors to obtain an aligned multimodal data stream, specifically including:

[0076] S11. Identify pixel feature mutation points in image data in multimodal interactive data, and determine the time when the pixel feature mutation points occur as semantic anchor points.

[0077] Specifically, computer vision processing algorithms are used to analyze the acquired continuous image frames to capture drastic changes in pixel color and brightness caused by tools contacting workpieces, indicator lights illuminating, or target object displacement. Such visual moments with clear physical meaning are defined as semantic anchor points. For example, in fiber optic fusion splicing tasks, the moment when the fiber end face contacts the fusion splicer slot, causing a high-light reflection in the image, or the moment when the fusion splicer status indicator light changes from off to on, causing a jump in the brightness value of a local pixel, is used as a semantic anchor point.

[0078] S12. Obtain the original timestamps of the semantic anchor points corresponding to the operation metadata and device sensor data, and calculate the time offset of each original timestamp relative to the time of the semantic anchor point.

[0079] Specifically, the system retrieves the corresponding instruction issuance logs recorded by the task management platform when the physical event represented by the semantic anchor occurs, as well as the physical quantity change records produced by the hardware sensors on the training device. It extracts the original timestamps generated by their respective hardware clocks carried in these structured data streams and calculates the absolute difference between them and the semantic anchor time in the image stream as the time offset. For example, if the time when the indicator light in the image turns on is T=10.050s, and the timestamp of the "start fusion" instruction recorded in the operation log is T=10.100s, then the time offset of the operation log is +50ms.

[0080] S13. Based on the time offset, align the operation metadata and device sensor data with the time reference to generate an aligned multimodal data stream.

[0081] Specifically, the time offset determined in the preceding steps is used to perform time axis translation compensation or resampling processing on the operation metadata and device sensor data. That is, the log commands and sensor values ​​that were originally misaligned on the time scale are forwarded or backward mapped according to the offset, so that they completely coincide with the corresponding actions in the image data on the logical time axis. For example, the operation log data with the aforementioned offset of +50ms is shifted forward by 50ms, thereby ensuring that the "start welding" command is at the same time point as the visual image of the indicator light illuminating in the aligned data stream.

[0082] In one embodiment, in step S20, the first feature set includes at least time information, spatial coordinates, and tool identification information, and the second feature set includes at least an operation rhythm stability index and a motion smoothness index. The first feature set is obtained from the aligned multimodal data stream, and the second feature set is obtained by analyzing the first feature set, specifically including:

[0083] S21. Based on the time information and spatial coordinates, obtain the instantaneous velocity characteristics and instantaneous acceleration characteristics at each corresponding moment.

[0084] Specifically, a numerical differential algorithm is used to calculate the rate of change of time between the time series and the three-dimensional spatial coordinates carried in the first feature set. The instantaneous velocity vector is solved by calculating the ratio of the displacement increment to the time increment between adjacent sampling points. The obtained velocity vector is then differentiated again to determine the instantaneous acceleration characteristics. For example, in a stripping action, the velocity curve of the user's hand can be obtained by taking the first derivative, and the change of the applied force can be obtained by taking the second derivative.

[0085] S22. Calculate the standard deviation of the instantaneous velocity characteristics within a preset sliding window, and determine the standard deviation as the operation rhythm stability index.

[0086] Specifically, a fixed step-size analysis interval is established along the time axis. The discreteness of all instantaneous speed values ​​falling within the preset sliding window is statistically calculated. The standard deviation is used to measure whether the user's speed is uniform when performing a specific action. This value is then mapped to an operation rhythm stability index. The standard deviation reflects the fluctuation range of the user's operation speed within a specific period. For example, when a skilled worker performs the action of continuously crimping RJ45 connectors, their speed standard deviation will be very small, indicating that they have formed a stable operation rhythm. On the other hand, a novice may experience large speed fluctuations due to hesitation, resulting in a larger standard deviation.

[0087] S23. Calculate the rate of change of the local slope of the instantaneous acceleration feature within the preset sliding window, and determine the motion smoothness index based on the positive and negative transformation frequency of the local slope within the preset threshold range.

[0088] Specifically, local linear regression or difference operations are performed on the instantaneous acceleration characteristics to obtain the trend of the acceleration curve on a microscopic time scale, i.e., the local slope change rate. The number of times the slope switches polarity within a preset positive and negative value range is monitored in real time. The frequency of positive and negative changes represents the frequency of tremors or corrections at the microscopic level. If the frequency is too high and the rate of change exceeds the preset residual range, it is determined that the user has obvious hesitation, probing, or hand tremors at the moment of operation. The motion smoothness index determined in this way can capture uncontrolled movements that are difficult to detect with the naked eye. For example, when aligning precision components, the slight tremor of the user's hand will cause the acceleration slope to frequently switch between positive and negative, thus obtaining a low smoothness index.

[0089] In one embodiment, in step S30, the third feature set includes at least a proficiency index and a cognitive load index. Dynamic weighting coefficients associated with the teaching task are determined, and the second feature set is weighted and fused using these dynamic weighting coefficients to obtain the third feature set. Specifically, this includes:

[0090] S31. Based on the difficulty coefficient of the teaching task and the user's historical operation performance data, assign dynamic weight coefficients to the operation rhythm stability index and the action smoothness index.

[0091] Specifically, a coefficient adjustment mechanism based on task scenario characteristics and individual ability benchmarks is established. The difficulty coefficient refers to the level of precision or speed requirements for different operational steps preset by teaching experts, while the user's historical operational performance data reflects the user's ability growth curve in past training cycles. By comprehensively judging the focus of the current task, differentiated influence ratios are assigned to the kinematic indicators produced by the preceding steps. For example, when assessing the "timed obstacle removal" task, the operation rhythm stability index is assigned a weight of 0.7, and the action smoothness index is assigned a weight of 0.3; while when assessing the "precision welding" task, the smoothness index is assigned a weight of 0.8.

[0092] S32. The proficiency index is obtained by weighting and summing the second feature set using dynamic weight coefficients.

[0093] Specifically, by using the weighting factors allocated in the preceding steps, a linear superposition operation is performed on the quantitative indicators in the second feature set to aggregate the discrete kinematic parameters into a comprehensive value that can represent the user's mastery of the skill, namely the proficiency index. This index achieves a quantitative evaluation of the user's familiarity with the operation by integrating the physical performance of rhythm and smoothness. For example, if the stability index is 90 points and the smoothness index is 80 points, with weights of 0.7 and 0.3 respectively, the calculated proficiency index is (90×0.7+80×0.3)=87 points.

[0094] S33. Obtain the time interval between two adjacent tool identification information switching in the first feature set, and perform correlation analysis between the time interval and the local slope change rate in the second feature set to obtain the cognitive load index.

[0095] Specifically, by monitoring the time points when tool identifiers change within the first feature set, the non-execution blank period between two consecutive operation actions, i.e., the time interval, is extracted. The correlation calculation is then performed between this time interval and the acceleration slope fluctuation before and after this period to obtain the cognitive load index. For example, if a user pauses for 3 seconds between putting down the wire stripper and picking up the crimping pliers, and the initial action acceleration slope after picking up the crimping pliers fluctuates violently, it indicates that the pause is a period of hesitation with high cognitive load, rather than a normal physical adjustment.

[0096] In one embodiment, step S40, namely, constructing a user operation scene topology based on the first feature set, the second feature set, and the third feature set, and calculating the similarity between the user operation scene topology and a preset standard topology, specifically includes:

[0097] S41. Based on the tool identification information, spatial coordinates and action execution order in the first feature set, identify multiple operation units, and construct an initial operation graph with each operation unit as a node and the temporal sequence and usage order between each operation unit as directed edges.

[0098] Specifically, by utilizing the tool identifier switching points and the clustering characteristics of spatial coordinate displacement recorded in the first feature set, each tool use is defined as an independent operation unit node, and each operation unit is defined as a node in the graph structure. At the same time, based on the temporal logic of action execution and the physical order of tool use, directional connecting lines are established between adjacent nodes as directed edges, thereby generating an initial operation graph that reflects the skeleton of the operation logic. For example, in the task of making network cables, steps such as "using wire strippers", "arranging wire sequence", and "using crimping pliers" are identified as sequential nodes in the graph.

[0099] S42. The operation rhythm stability index and action smoothness index in the second feature set, as well as the proficiency index and cognitive load index in the third feature set, are used as multi-dimensional attribute feature vectors and mapped to the corresponding nodes. The directed edges are weighted according to the switching time between each operation unit to generate the user operation scene topology.

[0100] Specifically, the stability and smoothness indices in the second feature set, and the proficiency and cognitive load indices in the third feature set, are vectorized and encapsulated to form a multi-dimensional attribute feature vector that can comprehensively characterize the operation quality of the node. This vector is then used as the internal attributes of the node. At the same time, the switching time data between each operation unit is extracted and mapped to the weight values ​​of the topological edges. For example, the "Organize Line Sequence" node will be attached with its corresponding smoothness index and cognitive load index, and the directed edge from "Organize Line Sequence" to "Use Wire Crimping Pliers" will be assigned a weight value, such as 1.5 seconds, thereby constructing the user operation scene topology.

[0101] S43. Using a graph matching algorithm, the topological structure of the user's operation scene is compared with the preset standard topological structure, and the similarity of the topological structure in the node arrangement order dimension and the similarity of the attribute distribution in the multi-dimensional attribute feature vector dimension corresponding to each node are calculated respectively.

[0102] Specifically, the graph matching algorithm or subgraph isomorphism comparison logic is invoked to perform multi-dimensional benchmarking operations on the generated user operation scene topology and the expert-level standard topology pre-stored in the cloud. By calculating the degree of overlap between the two graph structures in the node connection order and logical path, the topology similarity reflecting the correctness of the process is obtained. At the same time, by calculating the Euclidean distance or cosine exponent of the multi-dimensional attribute feature vectors between corresponding nodes, the attribute distribution similarity reflecting the standardization of the actions is obtained. For example, if the user's operation order is completely correct, the topology similarity is 100%, but if the proficiency index of the "organizing the line sequence" node is much lower than that of the standard model, the attribute distribution similarity will be lower.

[0103] In one embodiment, step S50, namely, determining the branch diagnosis result based on similarity and the third feature set, and generating operation improvement suggestions based on the branch diagnosis result, specifically includes:

[0104] S51. Based on topological similarity, attribute distribution similarity, and proficiency, determine the branch diagnosis results, which include logical defect results and precision defect results.

[0105] Specifically, the branch diagnosis result is determined primarily based on topological similarity, supplemented by attribute distribution similarity and proficiency. For example, the threshold for topological similarity is set at 95%. If the calculated topological similarity is lower than this threshold, the branch diagnosis result is primarily judged as a logical defect, indicating that the user's core problem lies in an error in the operational process. If the topological similarity is higher than or equal to this threshold, the attribute distribution similarity and proficiency are further examined to see if they meet the preset accuracy benchmark. If any indicator is lower than its corresponding accuracy threshold, such as proficiency being lower than 85 points, the branch diagnosis result is judged as an accuracy defect, indicating that the user's operational technique is insufficient.

[0106] S52. When the branch diagnosis result is a logical defect result, compare the node differences and edge connection differences between the user operation scene topology and the standard topology to identify missing operation units and operation units with abnormal order.

[0107] Specifically, by calculating the difference between the user-generated topology graph and the standard topology graph, it automatically identifies specific nodes that exist in the standard path but are missing in the user's path, or identifies directed edges whose connection order between nodes does not conform to the standard logic. Missing operation units and sequence abnormal operation units accurately locate the blind spots of the user at the knowledge memory level. For example, it identifies the missing operation unit that the user did not perform "put on the crystal head protective cover".

[0108] S53. Retrieve process reshaping suggestions that match missing operation units and sequentially abnormal operation units from the preset teaching strategy library, and use the process reshaping suggestions as operation improvement suggestions.

[0109] Specifically, based on the identified type of logical error, the system automatically retrieves corresponding process demonstration animations, standard operating procedure lists, or logical mnemonic devices from the teaching resource library. The process reshaping suggestions aim to correct users' operational cognition through targeted knowledge reinforcement. Compared to generalized guidance, these suggestions directly inform users which steps have been forgotten and what the correct execution order is, effectively assisting users in reconstructing a standard operational model in their minds. For example, for the missing "protective cover" step, a process reshaping suggestion is generated: "Operation Reminder: Please install the protective cover on the crystal head before the final test."

[0110] S56. If the branch diagnosis result is a precision defect result, locate the abnormal node in the second feature set whose local slope change rate exceeds the preset residual range, and obtain the motion smoothness index corresponding to the abnormal node.

[0111] Specifically, by tracing back to specific operation nodes with low attribute similarity in the topology, and by retrieving the acceleration slope data stored inside the node, the micro-time intervals where acceleration fluctuations are violent and do not conform to physical laws are accurately located, i.e. abnormal nodes. For example, it is found that the user's hand trembles at a high frequency in the last 0.2 seconds of the "pressing" action, resulting in an extremely low smoothness index during that period.

[0112] S57. Calculate the deviation between the motion smoothness index and the standard smoothness constant. Based on the deviation, match the corresponding training guide from the teaching strategy library and use the training guide as a suggestion for operation improvement.

[0113] Specifically, the mathematical difference between the user's action performance at abnormal nodes and the expert-level standard smoothness is quantified, and the corresponding physical skill enhancement scheme is retrieved from the strategy library according to the severity of the deviation. The training guide includes specific action correction parameters or special practice suggestions. For example, if the smoothness deviation value exceeds 50%, the training guide "It is recommended to carry out special training on hand grip stability, focusing on practicing uniform force" is generated.

[0114] In one embodiment, step S60 involves associating operation improvement suggestions with spatiotemporal nodes in the user operation scenario topology to generate a user analysis report, specifically including:

[0115] S61. Obtain the identification information of missing operation units, sequentially abnormal operation units, or abnormal nodes corresponding to the operation improvement suggestions, and obtain the evidence data corresponding to the identification information in the aligned multimodal data stream, wherein the evidence data includes the original image slices and / or sensor feature curves.

[0116] Specifically, by using the unique identification number of the logical offset point or physical performance anomaly point locked in the preceding diagnostic process, the original resource pool in a high-precision synchronization state is retrieved in reverse. By locating the start and end coordinates of these identification information on the global time axis, the original video clip containing the moment of the action is executed is automatically extracted, or the sensor numerical curve reflecting the physical and mechanical fluctuations during that period is extracted. Thus, the diagnostic conclusion is transformed into evidence data with intuitive and perceptible characteristics. For example, for the diagnosis of "shaky hands during pressing", a close-up video of the hand at the moment of pressing is automatically extracted as evidence data.

[0117] S62. Integrate and encapsulate the operation improvement suggestions with the corresponding evidence data, and arrange them according to the temporal logic of the user operation scenario topology to generate a user analysis report.

[0118] Specifically, textual guidance and multimedia physical evidence are logically integrated and distributed across a visualized timeline based on the chronological order of business processes represented by the user's operational landscape. This creates a user analysis report that includes overall scoring, defect location distribution, segmented evidence playback, and expert improvement plans. This chronologically arranged report format allows users to access corresponding professional guidance and concurrent evidence at each key time point, following the evolution of their actual operations. This significantly enhances the intuitiveness and relevance of the teaching review, providing users with a digital learning archive with deep interactive capabilities. For example, on the report's timeline, "press-in" nodes are marked in red and accompanied by clickable evidence videos and specific improvement suggestions.

[0119] In one embodiment, the user-led intelligent teaching analysis method further includes:

[0120] S70. Based on the user analysis report, match and push the corresponding special training task from the preset task library, and push the special training module to the user terminal to guide the user to train according to the special training task.

[0121] Specifically, an automatic mapping index logic is established between analysis conclusions and training resources. Based on the defect categories and severity identified in the user analysis report, reinforcement exercises aimed at solving specific problems are automatically retrieved from a preset task library. For example, process simulation exercises are automatically pushed for logic defect results, while high-frequency repetitive action feel enhancement modules are automatically pushed for accuracy defect results. By accurately distributing the matched specialized training modules to the user interface, users are guided to conduct targeted secondary correction exercises. For example, "standard process simulation" tasks are automatically pushed for logic defect results, while "pressing feel enhancement" specialized training modules are automatically pushed for accuracy defect results, thus realizing a complete teaching closed loop from data collection, intelligent diagnosis to precise intervention.

[0122] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0123] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 2As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database stores data information such as multimodal interaction data, semantic anchors, and multimodal data streams. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a user-centered intelligent analysis method for teaching based on a large-scale graphical model.

[0124] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0125] In response to the execution of teaching tasks, the system acquires multimodal interaction data generated by users, determines semantic anchors based on the multimodal interaction data, and performs spatiotemporal alignment of the multimodal interaction data based on the semantic anchors to obtain an aligned multimodal data stream.

[0126] The first feature set is obtained from the aligned multimodal data stream, and the second feature set is obtained by analyzing the first feature set.

[0127] Determine the dynamic weight coefficients associated with the teaching tasks, and use the dynamic weight coefficients to perform weighted fusion processing on the second feature set to obtain the third feature set;

[0128] Based on the first feature set, the second feature set, and the third feature set, a user operation scene topology is constructed, and the similarity between the user operation scene topology and the preset standard topology is calculated.

[0129] Based on similarity and the third feature set, the branch diagnosis results are determined, and operation improvement suggestions are generated based on the branch diagnosis results;

[0130] The operation improvement suggestions are associated with the spatiotemporal nodes in the user operation scenario topology to generate a user analysis report.

[0131] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:

[0132] In response to the execution of teaching tasks, the system acquires multimodal interaction data generated by users, determines semantic anchors based on the multimodal interaction data, and performs spatiotemporal alignment of the multimodal interaction data based on the semantic anchors to obtain an aligned multimodal data stream.

[0133] The first feature set is obtained from the aligned multimodal data stream, and the second feature set is obtained by analyzing the first feature set.

[0134] Determine the dynamic weight coefficients associated with the teaching tasks, and use the dynamic weight coefficients to perform weighted fusion processing on the second feature set to obtain the third feature set;

[0135] Based on the first feature set, the second feature set, and the third feature set, a user operation scene topology is constructed, and the similarity between the user operation scene topology and the preset standard topology is calculated.

[0136] Based on similarity and the third feature set, the branch diagnosis results are determined, and operation improvement suggestions are generated based on the branch diagnosis results;

[0137] The operation improvement suggestions are associated with the spatiotemporal nodes in the user operation scenario topology to generate a user analysis report.

[0138] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0139] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0140] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A user-centered intelligent teaching analysis method based on a large-scale picture model, characterized in that, The user-led intelligent analysis method includes: In response to the execution of teaching tasks, multimodal interaction data generated by users is acquired, semantic anchors are determined based on the multimodal interaction data, and the multimodal interaction data is spatiotemporally aligned based on the semantic anchors to obtain an aligned multimodal data stream. A first feature set is obtained from the aligned multimodal data stream, and the first feature set is analyzed to obtain a second feature set; Determine the dynamic weight coefficients associated with the teaching task, and use the dynamic weight coefficients to perform weighted fusion processing on the second feature set to obtain the third feature set; Based on the first feature set, the second feature set, and the third feature set, a user operation scene topology is constructed, and the similarity between the user operation scene topology and the preset standard topology is calculated. Based on the similarity and the third feature set, the branch diagnosis result is determined, and operation improvement suggestions are generated based on the branch diagnosis result; The operation improvement suggestions are associated with the spatiotemporal nodes in the user operation scene topology to generate a user analysis report.

2. The user-led intelligent teaching analysis method according to claim 1, characterized in that, The multimodal interaction data includes at least image data, operational metadata, and device sensor data; the step of determining semantic anchors based on the multimodal interaction data and performing spatiotemporal alignment of the multimodal interaction data based on the semantic anchors to obtain an aligned multimodal data stream specifically includes: Identify pixel feature mutation points in the image data within the multimodal interactive data, and determine the time when the pixel feature mutation points occur as the semantic anchor points; Obtain the original timestamps of the semantic anchor points corresponding to the operation metadata and the device sensor data, and calculate the time offset of each original timestamp relative to the time of the semantic anchor point. Based on the time offset, the operation metadata and the device sensor data are time-referenced to generate the aligned multimodal data stream.

3. The user-led intelligent teaching analysis method according to claim 1, characterized in that, The first feature set includes at least time information, spatial coordinates, and tool identification information; the second feature set includes at least an operation rhythm stability index and a motion smoothness index. The process of obtaining the first feature set from the aligned multimodal data stream and analyzing the first feature set to obtain the second feature set specifically includes: Based on the time information and the spatial coordinates, the instantaneous velocity characteristics and instantaneous acceleration characteristics at each corresponding moment are obtained; Calculate the standard deviation of the instantaneous velocity characteristic within a preset sliding window, and determine the standard deviation as the operation rhythm stability index; Calculate the rate of change of the local slope of the instantaneous acceleration feature within the preset sliding window, and determine the motion smoothness index based on the positive and negative transformation frequency of the local slope within the preset threshold range.

4. The user-centered intelligent teaching analysis method according to claim 3, characterized in that, The third feature set includes at least a proficiency index and a cognitive load index. The determination of dynamic weight coefficients associated with the teaching task, and the weighted fusion processing of the second feature set using these dynamic weight coefficients to obtain the third feature set, specifically includes: Based on the difficulty coefficient of the teaching task and the user's historical operation performance data, dynamic weight coefficients are assigned to the operation rhythm stability index and the action smoothness index. The proficiency index is obtained by weighting and summing the second feature set using the dynamic weight coefficients. The cognitive load index is obtained by obtaining the time interval between two consecutive tool identification information switching in the first feature set and performing correlation analysis between the time interval and the local slope change rate in the second feature set.

5. The user-led intelligent teaching analysis method according to claim 4, characterized in that, The step of constructing a user operation scene topology based on the first feature set, the second feature set, and the third feature set, and calculating the similarity between the user operation scene topology and a preset standard topology, specifically includes: Based on the tool identification information, spatial coordinates and action execution order in the first feature set, multiple operation units are identified, and an initial operation graph is constructed with each operation unit as a node and the temporal sequence and usage order between each operation unit as directed edges. The operation rhythm stability index and the action smoothness index in the second feature set, as well as the proficiency index and the cognitive load index in the third feature set, are used as multi-dimensional attribute feature vectors and mapped to the corresponding nodes. The directed edges are weighted according to the switching time between each operation unit to generate the user operation scene topology. Using a graph matching algorithm, the user operation scene topology is compared with the preset standard topology, and the topology similarity in the node arrangement order dimension and the attribute distribution similarity in the multi-dimensional attribute feature vector dimension corresponding to each node are calculated respectively.

6. The user-led intelligent teaching analysis method according to claim 5, characterized in that, The process of determining the branch diagnosis result based on the similarity and the third feature set, and generating operation improvement suggestions based on the branch diagnosis result, specifically includes: Based on the topological similarity, the attribute distribution similarity, and the proficiency, the branch diagnosis result is determined, wherein the branch diagnosis result includes logical defect results and precision defect results; If the branch diagnosis result is the logical defect result, compare the node differences and edge connection differences between the user operation scene topology and the standard topology to identify missing operation units and operation units with abnormal order. Retrieve process reshaping suggestions that match the missing operation unit and the sequence abnormal operation unit from the preset teaching strategy library, and use the process reshaping suggestions as the operation improvement suggestions; If the branch diagnosis result is the accuracy defect result, locate the abnormal node in the second feature set whose local slope change rate exceeds the preset residual range, and obtain the motion smoothness index corresponding to the abnormal node; The deviation of the motion smoothness index from the standard smoothness constant is calculated. Based on the deviation, a corresponding training guide is matched from the teaching strategy library, and the training guide is used as a suggestion for improving the operation.

7. The user-led intelligent teaching analysis method according to claim 6, characterized in that, The step of associating the operation improvement suggestions with spatiotemporal nodes in the user operation scenario topology to generate a user analysis report specifically includes: Obtain the identification information of the missing operation unit, the sequentially abnormal operation unit, or the abnormal node corresponding to the operation improvement suggestion, and obtain the evidence data corresponding to the identification information in the aligned multimodal data stream, wherein the evidence data includes the original image slices and / or sensor feature curves; The operation improvement suggestions are fused and encapsulated with the corresponding evidence data, and arranged according to the temporal logic of the user operation scenario topology to generate the user analysis report.

8. The user-led intelligent teaching analysis method according to claim 1, characterized in that, The user-led intelligent analysis method also includes: Based on the user analysis report, a corresponding specialized training task is matched and pushed from a preset task library, and the specialized training module is pushed to the user terminal to guide the user to train according to the specialized training task.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the user teaching intelligent analysis method based on a large picture model as described in any one of claims 1 to 8.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the user teaching intelligent analysis method based on the large picture model as described in any one of claims 1 to 8.