Mechanical engineering practical training scoring system, method and device based on intelligent machine

By adopting embodied intelligence and multimodal perception technology in the mechanical engineering training scoring system, the existing system's slow detection speed and inability to meet real-time requirements are solved, and efficient and personalized training scoring and feedback are achieved.

CN120183050AActive Publication Date: 2025-06-20先进计算与关键软件(信创)海河实验室
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510662367.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-06-20
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

The existing mechanical engineering training scoring system has problems such as slow detection speed, inability to meet real-time requirements, difficulty in dealing with the alignment and computational redundancy of multimodal features, and inability to achieve high accuracy, multi-dimensional evaluation and personalized evaluation.

Method used

The mechanical engineering training and scoring system based on embodied intelligence is adopted, combined with edge computing and multimodal perception technology, real-time motion detection and scoring is realized through image data acquisition and preprocessing, human target position detection, human joint recognition and posture estimation, human movement recognition and report generation modules.

Benefits of technology

It achieves fast detection speed, meets the real-time requirements of practical training scenarios, improves the effectiveness of standardized and personalized ability cultivation of students' practical training process, and reduces teachers' repetitive teaching tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120183050A_ABST
    Figure CN120183050A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and provides a mechanical engineering practical training scoring system, method and device based on intelligence. Comprising an image data acquisition and preprocessing module, a human body target position detection module, a human body joint recognition and posture estimation module, a human body action recognition module and a report generation module, wherein the image data acquisition and preprocessing module is used for carrying out standardization processing on original images and video data and outputting action sequence data in a unified format; the human body target position detection module is used for marking target frames of human bodies and components; the human body joint recognition and posture estimation module is used for recognizing key points of the trunk and the hands of a human body; the human body action recognition module is used for classifying actions; and the report generation module is used for generating a score table of mechanical engineering practical training. According to the invention, by designing a modular stacking structure, the effects of supporting flexible adjustment of model depth and multi-task extension and considering edge device deployment and high-performance computing requirements are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and provides a mechanical engineering training scoring system, method and device based on embodied intelligence. Background Art

[0002] In the current field of vocational education, practical teaching has always been an important way to cultivate students' operational skills, but the traditional teaching model generally has problems such as excessive reliance on teachers, delayed feedback information, and decentralized data processing, which to a certain extent restricts the teaching efficiency and the improvement of students' skills. In the traditional model, since the teaching process mainly relies on the on-site guidance of teachers, it is difficult to timely and comprehensively grasp the status of students in the process of operating complex equipment, so that teaching feedback often cannot meet the requirements of real-time and comprehensiveness, thus affecting the teaching effect and learning quality.

[0003] At the same time, existing teaching equipment has problems such as insufficient intelligence. Most systems are still at the level of basic information recording or one-way demonstration, lacking the ability to perceive and make dynamic decisions on multi-dimensional operational behaviors. For example, although some training platforms are equipped with video monitoring functions, they are limited by the lag of the data processing architecture, so they can only be replayed and reviewed afterwards, and cannot provide immediate feedback during the operation process; in addition, there is a lack of unified data integration and coordination mechanisms between the dispersed training equipment, which makes it difficult for teachers to control the real-time training status of all students at the same time, and the management of practical training teaching still relies on fragmented manual inspection methods.

[0004] In recent years, with the continuous development of intelligent technology and information technology, edge computing, multimodal perception and large models have provided new ideas for the intelligent upgrade of teaching equipment. However, the existing technology still has the following shortcomings: (1) Existing target detection models have large computational complexity and slow detection speed, and cannot meet the real-time requirements of practical training scenarios.

[0005] (2) Traditional action recognition methods have difficulty dealing with the alignment and computational redundancy problems of multimodal features.

[0006] (3) Automated training report generation technology still cannot meet the requirements of high accuracy, multi-dimensional evaluation and personalized assessment. Therefore, how to efficiently conduct comprehensive and personalized evaluation and feedback on each student's training operation process is still a technical difficulty that needs to be overcome. Summary of the invention

[0007] The present invention aims to solve at least one of the technical problems existing in the related art. To this end, the present invention provides a mechanical engineering training scoring system, method and device based on embodied intelligence, which achieves fast detection speed and meets the real-time requirements of the training scene.

[0008] The present invention provides a mechanical engineering training scoring system based on embodied intelligence, including the following modules: an image data acquisition and preprocessing module; a human target position detection module; a human joint recognition and pose estimation module; a human action recognition module; and a report generation module: The image data acquisition and preprocessing module is used to perform standardization processing on the original image and video data, and output action sequence data with a unified format; The human target position detection module is used to mark the target boxes of the human body and components according to the action sequence data to obtain a human body block diagram; The human joint recognition and pose estimation module is used to perform key point recognition on the torso and hands of the human body block diagram to obtain a human body key point diagram; The human action recognition module is used to classify the actions of the human body key point diagram to obtain action classification data; The report generation module is used to generate a scoring form for mechanical engineering training according to the action classification data and the action sequence data.

[0009] According to a mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the image data acquisition and preprocessing module includes: an image data acquisition module and an image preprocessing module: The image data acquisition module is used to acquire images and videos of mechanical engineering training; The image preprocessing module is used to synchronize the images and videos of the mechanical engineering training using the NTP protocol, extract key frames, and obtain action sequence data.

[0010] According to a mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the human target position detection module uses the CGYOLO human target detection network to detect the action sequence data to obtain a human body block diagram.

[0011] According to a mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the CGYOLO human target detection network includes a backbone network, a neck network, and a detection head: The backbone network includes a C3k2 module and a C2PSA module. The input feature is split into two branches. One branch is processed by grouped convolution and the C2PSA attention mechanism, and the other branch retains the original feature. Finally, cross-stage information reuse is achieved through splicing; The neck network uses the CARAFE lightweight upsampling operator to replace the traditional bilinear interpolation, and at the same time uses the WIoU loss function to replace the original EIoU loss function; The detection head integrates a CSPRGB (CSPNet Red Green Blue) module for feature fusion.

[0012] A mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the working process of the human joint recognition and pose estimation module is as follows: S31: Input the human block diagram into the SRTMPose model, and classify the video actions according to the spatio-temporal change characteristics of the action poses to obtain classified video data; S32: Calculate the classified video data, generate an attention mask for the classified video data according to the SimAM attention mechanism, and suppress the attention mask area to obtain a human key point map.

[0013] A mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the SRTMPose model adds a multi-scale feature fusion Inception module and a parameter-free attention optimization SimAM module to the backbone network.

[0014] A mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the working process of the human action recognition module is as follows: S41: Input the human key point map into the STAR-Transformer model, and classify the video actions according to the spatio-temporal change characteristics of the action poses to obtain classified action video data; S42: Input the classified action video data into the encoder to obtain encoded data, and the encoder includes a full attention FAttn module and a zigzag attention ZAttn module; S43: Input the encoded data into the decoder to obtain action classification data, and the decoder includes a full attention FAttn module and a binary attention BAttn module.

[0015] A mechanical engineering training scoring system based on embodied intelligence provided by the present invention, the report generation module includes a text affine sub-module, an image prompt cross-affine sub-module and a fusion module, and the working process of the report generation module is as follows: S51: Input the action classification data and action sequence data into the image prompt cross-affine sub-module to generate image fusion data, and the image prompt cross-affine sub-module includes a self-attention module, a cross-attention module, a feed-forward neural network module and a residual connection module; S52: Input the text prompt into the text affine sub-module to obtain text fusion data, and the text affine sub-module includes a self-attention module, a normalization module, a feed-forward neural network module and a residual connection module; S53: Input the action classification data, image fusion data and text fusion data into the fusion module to generate an experiment scoring table.

[0016] The present invention also provides a method for scoring mechanical engineering training based on embodied intelligence, including: S1: The mechanical engineering trainee conducts mechanical engineering training in the image data acquisition and preprocessing module, and the image data acquisition and preprocessing module performs image acquisition; S2: The image data acquisition and preprocessing module performs standardization processing on the original image and video data, and outputs action sequence data with a unified format; the human target position detection module marks the target frames of the human body and components according to the action sequence data to obtain a human body frame diagram; the human joint recognition and pose estimation module performs key point recognition on the torso and hands of the human body frame diagram to obtain a human key point diagram; the human action recognition module classifies the actions of the human key point diagram to obtain action classification data; the report generation module is used to generate a scoring form for mechanical engineering training according to the action classification data and the action sequence data.

[0017] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for scoring mechanical engineering training based on embodied intelligence as described in any one of the above are implemented.

[0018] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: The scoring system, method, and device for mechanical engineering training based on embodied intelligence provided by the present invention construct a full-process automated guidance and evaluation system for training operations through edge computing-driven multimodal perception (image recognition, target detection) and large model decision fusion technology, realizing the efficient release of teachers from repetitive teaching tasks and simultaneously improving the standardization and personalized ability training efficiency of students' training processes.

[0019] The additional aspects and advantages of the present invention will be partially given in the following description, partially become apparent from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic structural diagram of the device for scoring mechanical engineering training based on embodied intelligence provided by the present invention.

[0022] Figure 2It is a schematic flowchart of the mechanical engineering training scoring method based on embodied intelligence provided by the present invention.

[0023] Figure 3 It is a schematic diagram of system device deployment.

[0024] Figure 4 It is a flowchart of preprocessing by the image preprocessing module.

[0025] Figure 5 It is the structure of the human joint recognition and pose estimation module.

[0026] Figure 6 It is a schematic diagram of the structure of the electronic device provided by the present invention.

[0027] Reference numerals: 101, Image data acquisition and preprocessing module; 102, Human target position detection module; 103, Human joint recognition and pose estimation module; 104, Human action recognition module; 105, Report generation module; 810, Processor; 820, Communication interface; 830, Memory; 840, Communication bus. Detailed implementation manners

[0028] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0029] In the description of this specification, the description referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0030] The following combines Figures 1 to 6 to describe the present invention.

[0031] Embodiment As Figure 1 shown, Figure 1Schematic diagram of the structure of the mechanical engineering training scoring device based on embodied intelligence provided by the embodiments of the present invention, including the following modules: Image data acquisition and preprocessing module 101, Human target position detection module 102, Human joint recognition and pose estimation module 103, Human action recognition module 104, and Report generation module 105: The image data acquisition and preprocessing module 101 is used to perform standardization processing on the original image and video data, and output action sequence data with a unified format; The human target position detection module 102 is used to mark the target frames of the human body and components; The human joint recognition and pose estimation module 103 is used to identify the key points of the torso and hands of the human body; The human action recognition module 104 is used to classify actions; The report generation module 105 is used to generate a scoring form for mechanical engineering training.

[0032] Specifically, the image data acquisition and preprocessing module 101 includes: an image data acquisition module and an image preprocessing module: The image data acquisition module is used to obtain the images and videos of mechanical engineering training; The image preprocessing module is used to synchronize the images and videos of the mechanical engineering training using the NTP protocol, extract key frames, and obtain action sequence data.

[0033] The goal of this module is to collect and preprocess the action image data of students' training operations, construct a multi-perspective visual data acquisition system for mechanical engineering training scenarios, and achieve high-fidelity recording of the operation process and standardized data conversion through heterogeneous camera deployment and edge-side preprocessing. As Figure 3 shown, the system circularly deploys 4 high-resolution three-dimensional motion cameras around each training bench to cover the detailed areas of the operator's hand movements and machine tool components. Each camera synchronously acquires RGB images with a resolution of 640x640 or higher at a frame rate of 30fps or more, focusing on capturing two types of key data: one is the dynamic action sequence during the disassembly and assembly operation, such as moving and aligning components, etc.; the other is the static comparison images before and after the installation of machine tool components, such as the alignment and tightening of screws. At the same time, 6 wide-angle surveillance cameras are installed on the top of the training room. These cameras can record 1080p panoramic videos at least at 25fps, which are used to record the movement trajectory of the student's operation position and the global spatial distribution state of tools and components.

[0034] The collected image and video data are transmitted to the edge server cluster in real time through 5G or gigabit high-speed Ethernet, and multi-level preprocessing is performed to adapt to subsequent analysis requirements. As Figure 4As shown in the figure, first, the NTP (Network Time Protocol) network transmission protocol is adopted to achieve millisecond-level time synchronization of all cameras, eliminating the time offset of cross-device data streams. Subsequently, according to the Normalized Cross-Correlation (NCC algorithm) key frame dynamic extraction module, motion saliency analysis is performed on the video stream: full-frame rate data is retained during the stage of drastic changes in disassembly and assembly actions, while key frames are extracted at 1fps during static stages such as waiting for guidance, greatly reducing the data volume without losing effective information to obtain the video after key frame extraction. The picture quality optimization module performs adaptive correction for fluctuations in the lighting conditions of the training room, such as ceiling lamp glare and metal component reflection: the local contrast enhancement algorithm based on the Retinex theory is used to suppress high-light overexposure, and at the same time, the non-local mean denoising algorithm is adopted to eliminate image grain noise in low-light environments, obtaining a multi-view action sequence with time sequence synchronization and a comparison atlas of component disassembly and assembly states. The above two constitute the action sequence data.

[0035] Specifically, the human target position detection module 102 uses the CGYOLO human target detection network to detect the action sequence data to obtain a human body frame diagram. Among them, the CGYOLO human target detection network includes a backbone network, a neck network, and a detection head: The backbone network includes a C3k2 module and a C2PSA module. The input features are split into two branches. One branch is processed by grouped convolution and the C2PSA attention mechanism, and the other branch retains the original features. Finally, cross-stage information reuse is achieved through splicing; The neck network uses the CARAFE lightweight upsampling operator to replace the traditional bilinear interpolation, and at the same time uses the WIoU loss function to replace the original EIoU loss function; by introducing a quality-aware weight factor, WIoU imposes a higher penalty on low-quality prediction boxes (such as occluded or motion-blurred human targets), thereby reducing the missed detection rate of dense crowds. Its calculation expression is as follows: Among them, is the WIoU loss result, is the dynamic weight factor, is the intersection over union, is the center distance, which represents the Euclidean distance between the center points of the prediction box and the ground truth box, is the diagonal length, defined as the diagonal length of the smallest bounding box of the prediction box and the ground truth box.

[0036] The detection head integrates the CSPRGB module for feature fusion.

[0037] The CSPRGB Block is an efficient visual feature extraction module. Through the combination of the Cross Stage Partial Network (CSPNet) structure and the fine-grained optimization of the RGB channels, it first divides the input features into two paths. The main path convolution extracts local features, and the bypass path retains global information. Then, the two are fused and complemented. At the same time, channel attention, channel shuffle, or grouped convolution is introduced to strengthen the RGB color interaction, and color space conversion is supported to decouple the luminance and chrominance information. Finally, it can improve the model's ability to capture details such as color and texture while maintaining a low computational cost, especially suitable for visual tasks with high requirements for real-time performance and accuracy.

[0038] After the above series of improvements to the YOLOv11 network, this module can achieve fast and accurate annotation of the positions of training students and components. For example, whether a component is correctly taken out can be judged by the change in the position of the component.

[0039] Specifically, the working process of the human joint recognition and pose estimation module 103 is as follows: S31: Input the human body block diagram into the SRTMPose model, and classify the video actions according to the spatio-temporal change characteristics of the action poses to obtain classified video data; S32: Calculate the classified video data, generate an attention mask for the classified video data according to the SimAM attention mechanism, and suppress the attention mask area to obtain a human body key point diagram.

[0040] Among them, the SRTMPose model is an improvement of the RTMPose human pose estimation model. A multi-scale feature fusion Inception module and a parameter-free attention optimization SimAM module are added to the backbone network of RTMPose.

[0041] First, the human body block diagram is input into the backbone network of the SRTMPose model. The backbone network of SRTMPose is an improved CSPNeXt structure. In this structure, the input image first undergoes initial convolution and downsampling operations through the backbone layer, and then features are extracted from the stacked layers in the CSPNeXt block. In this layer, multiple Inception modules are introduced into the model to capture features at different scales (such as large body joints and small finger joints), alleviate the scale differences between the body and the hand, and multiple SimAM attention modules are used to dynamically enhance key regions (such as the hand contour) through a parameter-free energy function and suppress background interference (such as clothing occlusion or complex lighting). Specifically, the data of each CSPNeXt module first passes through the SimAM attention module, then through the Inception module, and finally the results of all Inception modules are added element-wise to obtain an addition result. The addition result is added element-wise to the feature extraction result of the CSPNeXt block stacked layer to obtain the finally extracted features.

[0042] Finally, the extracted features are dimensionally reduced and passed to the neck network. The neck network is only responsible for simple feature splicing. The spliced features are output through a dual-task by the detection head, and the output includes 17 body key points and 42 hand key points.

[0043] Among them, the core idea of the Inception module is parallel multi-scale convolution. Features are extracted simultaneously through convolution kernels of different sizes, enhancing the network's ability to capture multi-scale targets. It can capture local details such as finger joints and global structures simultaneously, improving the model's adaptability to targets of different sizes. At the same time, through convolution dimensionality reduction, the computational amount of large convolution kernels is reduced.

[0044] SimAM is a parameter-free attention mechanism that dynamically assigns feature weights through an energy function. The energy function can calculate the importance score of each position in the feature map: Among them, is the value of the energy function, is the total number of feature values, is the th feature value, is the ordinal number of the feature value, , is the average value of the feature values, is the standard deviation of the feature values, is the smoothing term.

[0045] Calculating according to the above formula, the positions with lower energy are less important. An attention mask is generated by normalizing the energy values to suppress the low-energy regions. By introducing the SimAM attention mechanism, the memory occupancy can be reduced, while enhancing the response of the key regions and improving the anti-interference ability of the model.

[0046] In the present invention, the Inception module solves the scale difference between the body and the hand through parallel processing of multi-scale convolutions, and the SimAM attention enhances the anti-interference ability through parameter-free dynamic weighting. The collaborative design of the two in the SRTMPose backbone network significantly improves the accuracy and robustness of hand keypoint detection, while maintaining real-time performance, providing an efficient solution for refined pose estimation in complex scenarios.

[0047] Specifically, the working process of the human action recognition module 104 is as follows: S41: Input the human keypoint map into the STAR-Transformer model, and classify the video action according to the spatio-temporal change characteristics of the action pose to obtain classified video data; S42: Input the classified video data into the encoder to obtain encoded data. The encoder includes a full attention FAttn module and a zigzag attention ZAttn module; S43: Input the encoded data into the decoder to obtain action classification data. The decoder includes a full attention FAttn module and a binary attention BAttn module.

[0048] The human action recognition module 104 inputs the video frame and the key point sequence of the video action into the STAR-Transformer model, and classifies the video action according to the spatio-temporal change characteristics of the action pose.

[0049] STAR-Transformer is a deep neural network model for video action classification tasks. Its architecture efficiently processes the input of video frame sequences and human skeleton sequences through multi-modal feature fusion and spatio-temporal attention mechanisms. As Figure 5As shown, the human action recognition module 104 as a whole adopts an encoder-decoder structure. The overall working process of the STAR-Transformer model is on the left side of the dotted line in the figure, while the right side of the dotted line represents the detailed composition of the encoder and decoder of the model, including the full spatio-temporal attention mechanism, residual connection, layer normalization, token decoupling, feed-forward network, etc. The arrow indicates the data flow direction. In the input and feature preprocessing stage, the model inputs two types of multi-modal data, the video frame sequence and the human skeleton sequence, and both are subjected to preliminary feature extraction through the token decoupling module. Then, it is processed by the encoder. The core components of the encoder include the full attention FAttn and the zigzag attention ZAttn. The full attention FAttn refers to the standard self-attention mechanism in the Transformer.

[0050] Among them, the full attention FAttn refers to the standard self-attention mechanism in the Transformer, that is, the formula for each token to pay attention to all other position tokens is: Among them, is the attention function, is the activation function, is the query matrix, is the key matrix, is the value matrix, is the transpose of the key matrix, is the key vector dimension, is the video spatio-temporal feature tensor, is the query matrix weight matrix, represents matrix concatenation, is the human skeleton sequence feature tensor, is the key matrix weight matrix, is the value matrix weight matrix.

[0051] The zigzag attention ZAttn means that the attention range is restricted through a specific "zigzag" path. In the image Transformer, adjacent regions are alternately covered according to the spatial coordinates to reduce the computational complexity. Its calculation formula is as follows: Among them, is the sparse mask, is the position at the time step where it is located, is the position at the time step where it is located, is the position The two-dimensional spatial coordinates of is the position The two-dimensional spatial coordinates of is the domain radius.

[0052] Furthermore, the processing result of the encoder is passed into the decoder, and the decoder generates a classification result through cross-modal interaction. The decoder also contains a full attention module FAttn. The difference is that it replaces the ZAttn module with a binary attention module BAttn. The BAttn module binarizes the attention weights (0 or 1), only retaining the most critical positional relationships, which can significantly reduce memory occupancy. The formula is: Among them, is the binarized value, is the binarization function, is the set threshold.

[0053] After the above processing, the model is activated by Softmax, and the obtained action category probability results are output. In summary, STAR-Transformer realizes the efficient fusion of video and skeleton sequences through the collaborative work of multiple groups of parameters.

[0054] Specifically, the report generation module 105 includes a text affine sub-module, an image prompt cross-affine sub-module, and a fusion module. The working process of the report generation module 105 is as follows: S51: Input the action classification data and action sequence data into the image prompt cross-affine sub-module to generate image fusion data. The image prompt cross-affine sub-module includes a self-attention module, a cross-attention module, a feed-forward neural network module, and a residual connection module; S52: Input the text prompt into the text affine sub-module to obtain text fusion data. The text affine sub-module includes a self-attention module, a normalization module, a feed-forward neural network module, and a residual connection module; S53: Input the action classification data, image fusion data, and text fusion data into the fusion module to generate an experiment score table.

[0055] The main purpose of this module is to use large model technology to judge the action classification results obtained by the human action recognition module, score the standardization of each step of each student's operation during the training according to the preset scoring criteria, and finally output the training score table of each student. The format of the training score table is shown in Table 1: Table 1 Schematic diagram of training score The total score of the experiment is 100 points and it is divided into 8 sub-items, including the complete process from taking out the components for installation to disassembling and putting back the components at the end of the experiment. Each sub-item is assigned different scores according to its importance in the experiment.

[0056] Specifically, at the beginning stage of the training, student A enters the training room for preparation, and the multi-view video surveillance cameras will capture the real-time positions of each student in real time. During the training process, student A conducts the experiment according to the steps described in the experiment scoring form.

[0057] Exemplarily, taking the installation of the Y-axis module by student A as an example. First, student A wears gloves and takes out the Y-axis module from the drawer. At this time, the video surveillance camera will record the comparison pictures before and after the module is taken out, and automatically judge whether the Y-axis module is taken out correctly.

[0058] Further, student A places the taken-out Y-axis module at an appropriate position on the disassembly and assembly workbench, and performs operations such as installation alignment and screwing. At this time, the 3D motion cameras deployed around the workbench will capture the hand movements of student A in real time and upload them to the edge server for analysis to judge which step student A is performing and whether the movements comply with the specifications.

[0059] Further, after detecting that the module installation is completed, the system will automatically identify whether the installed module meets the standards, whether there are missing screws and how many screws are missing, and finally automatically score this step according to the scoring form.

[0060] Further, after student A completes a sub-step, without pausing, directly proceeds to the next step operation. The system will automatically identify which step student A is completing and conduct corresponding evaluation and scoring. After the system identifies that student A has completed all the training operations, it will automatically summarize the scores of all sub-steps and generate a final experiment scoring form for student A. The scoring form is shown in Table 2 for example.

[0061] Specifically, Table 2 details the final scores and reasons for deductions of student A in each step during the training process. For example, during the operation of step 2, student A installed the X-axis module in the reverse direction, and according to the scoring standard, 5 points need to be deducted, and the score for this step is 5 points. Finally, the total score of all steps is summarized to 64 points.

[0062] Table 2 Experiment Scoring Form for Student A Furthermore, the automatic report generation based on the multimodal large model proposed by this module includes three sub-modules. This model accepts three inputs, namely the classification result obtained from the previous module, the image representation, i.e., various video frame data collected during the training process, and the text representation, i.e., the specific content and prompt words of each item in the scoring table. The image representation is input into the image-prompt cross-affine sub-module, and the text representation is input into the text affine sub-module. Finally, the two are fused with the classification result for output to obtain the final scoring table. Both the text affine and the image-prompt cross-affine each contain 12 blocks. The text affine includes a self-attention module, a normalization module, a feed-forward neural network module, and a residual connection module; the image-prompt cross-affine includes a self-attention module, a cross-attention module, a feed-forward neural network module, and a residual connection module. Among them, the self-attention module shares weights with the self-attention module in the text affine transformer, and the cross-attention module in each block also requires an input from an external prompt word.

[0063] Through the above innovative multimodal interaction architecture design, text, images, and classification results are deeply fused. The self-attention module with shared weights is used to achieve cross-modal semantic alignment, effectively improving the consistency of feature expression. At the same time, a learnable dynamic prompt word mechanism is introduced in the image sub-module, combined with cross-attention to focus on key visual information, enhancing the semantic parsing ability for complex training scenarios, and significantly optimizing the computational efficiency. The hierarchical residual design ensures training stability and still maintains reliable output when there is input noise or partial data loss, while the modular stacking structure supports flexible adjustment of the model depth and multi-task expansion, taking into account the deployment on edge devices and high-performance computing requirements.

[0064] As Figure 2 shown, the present invention also provides a method for scoring mechanical engineering training based on embodied intelligence. The steps include: S1: The mechanical engineering trainer conducts mechanical engineering training in the image data acquisition and preprocessing module, and the image data acquisition and preprocessing module performs image acquisition; S2: Score the image acquisition data in the mechanical engineering training scoring system based on embodied intelligence to generate an experimental scoring table.

[0065] Figure 6 Illustrates a schematic diagram of the physical structure of an electronic device. As Figure 6 shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830, and a communication bus 840. Among them, the processor 810, the communication interface 820, and the memory 830 complete mutual communication through the communication bus 840. The processor 810 can call the logical instructions in the memory 830 to execute a method for scoring mechanical engineering training based on embodied intelligence. The method includes: S1: The mechanical engineering trainee conducts mechanical engineering training within the image data acquisition and preprocessing module, and the image data acquisition and preprocessing module performs image acquisition; S2: Score the image acquisition data in the mechanical engineering training scoring system based on embodied intelligence to generate an experiment scoring form.

[0066] In addition, when the logical instructions in the above-mentioned memory 830 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0067] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the technical solution, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical discs, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments.

[0069] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

[0070] It should be noted that the embodiments of the present disclosure can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic: the software part can be stored in a memory and executed by an appropriate instruction execution system such as a microprocessor or dedicated designed hardware. Those skilled in the art can understand that the above devices and methods can be implemented using computer-executable instructions and / or included in processor control code, for example, such code is provided on a programmable memory or a data carrier such as an optical or electronic signal carrier.

[0071] In addition, although the operations of the method of the present disclosure are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution. It should also be noted that the features and functions of two or more devices according to the present disclosure can be embodied in one device. Conversely, the features and functions of one device described above can be further divided and embodied by multiple devices.

[0072] Although the present disclosure has been described with reference to several specific embodiments, it should be understood that the present disclosure is not limited to the specific embodiments disclosed. The present disclosure aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims.

Claims

1. A mechanical engineering training scoring system based on embodied intelligence, characterized in that, It includes the following modules: an image data acquisition and preprocessing module, a human target position detection module, a human joint recognition and pose estimation module, a human action recognition module, and a report generation module; The image data acquisition and preprocessing module is used to perform standardization processing on the original image and video data, and output action sequence data with a unified format; The human target position detection module is used to mark the target boxes of the human body and components according to the action sequence data to obtain a human body block diagram; The human joint recognition and pose estimation module is used to identify the key points of the torso and hands in the human body block diagram to obtain a human body key point diagram; The human action recognition module is used to classify the actions in the human body key point diagram to obtain action classification data; the working process of the human action recognition module is as follows: S41: Input the human body key point diagram into the STAR-Transformer model, and classify the video actions according to the spatio-temporal change characteristics of the action posture to obtain classified action video data; S42: Input the classified action video data into the encoder to obtain encoded data. The encoder includes a full attention FAttn module and a zigzag attention ZAttn module; S43: Input the encoded data into the decoder to obtain action classification data. The decoder includes a full attention FAttn module and a binary attention BAttn module; The report generation module is used to generate a scoring table for the mechanical engineering training according to the action classification data and the action sequence data.

2. The mechanical engineering training scoring system based on embodied intelligence according to claim 1, characterized in that, The image data acquisition and preprocessing module includes: an image data acquisition module and an image preprocessing module: The image data acquisition module is used to acquire images and videos of the mechanical engineering training; The image preprocessing module is used to synchronize the images and videos of the mechanical engineering training using the NTP protocol, extract key frames, and obtain action sequence data.

3. The mechanical engineering training scoring system based on embodied intelligence according to claim 2, characterized in that, The human target position detection module uses the CGYOLO human target detection network to detect the action sequence data to obtain a human body block diagram.

4. The mechanical engineering training scoring system based on embodied intelligence according to claim 3, characterized in that, The CGYOLO human target detection network includes a backbone network, a neck network, and a detection head: The backbone network includes a C3k2 module and a C2PSA module. The input features are split into two branches. One branch is processed by grouped convolution and the C2PSA attention mechanism, and the other branch retains the original features. Finally, cross-stage information reuse is achieved through splicing; The neck network uses the CARAFE lightweight upsampling operator to replace the traditional bilinear interpolation, and at the same time uses the WIoU loss function to replace the original EIoU loss function; The detection head integrates the CSPRGB module for feature fusion.

5. The mechanical engineering training scoring system based on embodied intelligence according to claim 3, characterized in that, The working process of the human joint recognition and pose estimation module is as follows: S31: Input the human body block diagram into the SRTMPose model, and classify the video actions according to the spatio-temporal change characteristics of the action posture to obtain classified video data; S32: Calculate the classified video data, generate an attention mask for the classified video data according to the SimAM attention mechanism, and suppress the attention mask area to obtain a human body key point diagram.

6. The mechanical engineering training scoring system based on embodied intelligence according to claim 5, characterized in that, The SRTMPose model adds a multi-scale feature fusion Inception module and a parameter-free attention optimization SimAM module to the backbone network.

7. The mechanical engineering training scoring system based on embodied intelligence according to claim 6, characterized in that, The report generation module includes a text affine sub-module, an image prompt cross-affine sub-module, and a fusion module. The working process of the report generation module is as follows: S51: Input the action classification data and action sequence data into the image prompt cross-affine sub-module to generate image fusion data. The image prompt cross-affine sub-module includes a self-attention module, a cross-attention module, a feed-forward neural network module, and a residual connection module; S52: Input the text prompt into the text affine sub-module to obtain text fusion data. The text affine sub-module includes a self-attention module, a normalization module, a feed-forward neural network module, and a residual connection module; S53: Input the action classification data, image fusion data, and text fusion data into the fusion module to generate an experimental score table.

8. A mechanical engineering training scoring method based on embodied intelligence for implementing the mechanical engineering training scoring system based on embodied intelligence according to any one of claims 1 to 7, characterized in that, Including: S1: The mechanical engineering trainee conducts mechanical engineering training in the image data acquisition and preprocessing module, and the image data acquisition and preprocessing module performs image acquisition; S2: The image data acquisition and preprocessing module standardizes the original image and video data and outputs action sequence data with a unified format; the human target position detection module marks the target boxes of the human body and components according to the action sequence data to obtain a human body box diagram; the human joint recognition and pose estimation module recognizes the key points of the torso and hands on the human body box diagram to obtain a human key point diagram; the human action recognition module classifies the actions of the human key point diagram to obtain action classification data; the report generation module is used to generate a score table for mechanical engineering training according to the action classification data and action sequence data.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method for scoring mechanical engineering training based on embodied intelligence according to claim 8.

Citation Information

Patent Citations

  • Crane operator practical operation assessment system based on artificial intelligence

    CN117237982A

  • Medical quality control report intelligent generation method and system based on retrieval enhancement

    CN118571402A

  • Multi-modal scientific experiment evaluation auxiliary method and system based on model coupling

    CN118968382A

  • Image processing apparatus and motion estimation method thereof

    US20240412382A1