An intelligent compliance detection system and detection method for assembly work

CN122657785APending Publication Date: 2026-08-28ZHEJIANG TSING-JET TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610713442.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-22
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

首先,视觉模型的构建依赖AI编程和深度学习专业知识,技术门槛高,中小企业难以配备相应技术人员

Benefits of technology

[0042] The lightweight model file is then distributed to the edge-side real-time detection and decision-making terminal. After training, channel pruning and parameter quantization are performed on the multi-task compliance judgment model, compressing the model size and reducing inference computation. This model is then converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal and serialized and packaged to generate a lightweight model file, which is then distributed to the edge-side real-time detection and decision-making terminal. The multi-task compliance judgment model, which originally required high computing power, can now run in real-time on embedded edge computing devices near the production line, reducing hardware costs and simplifying model deployment and upgrade processes, while ensuring the real-time performance of detection and the maintainability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122657785A_ABST
    Figure CN122657785A_ABST
Patent Text Reader

Abstract

The application discloses an intelligent compliance detection system for assembly work, characterized by comprising a camera, a no-code multi-task model construction platform and an edge-side real-time detection and decision terminal, the no-code multi-task model construction platform comprising a graphical collaborative labeling unit and a joint training engine, the graphical collaborative labeling unit receiving labeling information of a standard work video through graphical interaction of a user and outputting structured labeling data, and the joint training engine automatically training a multi-task compliance judgment model based on the structured labeling data; the edge-side real-time detection and decision terminal being deployed on a production line, loading and running the model to analyze real-time assembly video streams and output detailed compliance judgment and alarm signals containing part errors, position errors, action errors and sequence errors. The application has the advantages of greatly reducing technical threshold and deployment cost, being capable of accurately identifying various error types in the assembly process in real time, and effectively improving assembly quality and production efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a manufacturing production line inspection system and method, and more particularly to an intelligent compliance inspection system and method for assembly operations. Background Technology

[0002] On assembly lines in small and medium-sized manufacturing enterprises, assembly operations are generally performed manually, characterized by labor intensity, high repetition, and susceptibility to fatigue and errors. Frequent product changes necessitate workers rotating between different workstations and processes, increasing the risk of operational errors. Line supervisors rely on visual monitoring of multiple workstations, making it difficult to achieve simultaneous and continuous supervision of the operations at each station. Quality inspection is typically located at the back end of the production line, leading to delayed detection of assembly errors, resulting in rework, wasted resources, and even impacting delivery cycles.

[0003] Existing machine vision-based assembly inspection solutions have significant shortcomings when applied to the aforementioned scenarios. First, building visual models relies on AI programming and deep learning expertise, presenting a high technical barrier that small and medium-sized enterprises (SMEs) struggle to equip themselves with the necessary technical personnel for. Second, model training and inference depend on high-performance graphics processors, resulting in high hardware costs and hindering widespread deployment across production line workstations. Furthermore, most existing solutions can only detect single types of assembly errors, failing to provide integrated compliance assessments across multiple dimensions, including part position deviations, operational compliance, and process execution sequence. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide an intelligent compliance detection system and detection method for assembly operations that effectively improves assembly quality and production efficiency, realizing a complete closed loop from data annotation, model training, model deployment to online real-time compliance detection.

[0005] The technical solution adopted by the present invention to solve the above-mentioned technical problems is: an intelligent compliance detection system for assembly operations, comprising: a no-code multi-task model building platform, an edge-side real-time detection and decision-making terminal, and a camera;

[0006] The no-code multi-task model building platform includes a graphical collaborative annotation unit and a joint training engine;

[0007] The graphical collaborative annotation unit is used to provide a user interface to receive annotation information from users on imported standard operation videos through graphical interaction, and output structured annotation data; the annotation information includes parts, installation positions and assembly postures, assembly actions and process sequences.

[0008] The aforementioned joint training engine is used to automatically train a multi-task compliance judgment model based on the aforementioned structured labeled data.

[0009] The camera is used to collect a real-time assembly video stream consisting of real-time images of parts and real-time action images of workers on the production line, and transmit it to the edge-side real-time detection and decision-making terminal.

[0010] The edge-side real-time detection and decision-making terminal is an embedded edge computing device deployed next to the production line. It is used to load and run the multi-task compliance judgment model, analyze the real-time assembly video stream, and output detailed compliance judgments and alarm signals including part errors, position errors, action errors, and sequence errors.

[0011] Compared with existing technologies, the advantages of this invention are as follows: The graphical collaborative annotation unit in the no-code multi-task model building platform receives annotation operations performed by users on standard operation videos via graphical interaction, outputs structured annotation data, and the joint training engine automatically trains a multi-task compliance judgment model based on the structured annotation data; the no-code multi-task model building platform can also input CAD drawings and SOP text files as selectable auxiliary templates, improving the accuracy and scenario adaptability of model training; and it captures a real-time assembly video stream composed of real-time images of parts and real-time action images of workers through a camera, and transmits it to the edge for real-time detection and decision-making. The edge-side real-time detection and decision-making terminal is deployed on the production line, loading and running a multi-task compliance judgment model. It analyzes the real-time assembly video stream and outputs detailed compliance judgments and alarm signals, including part errors, position errors, action errors, and sequence errors. This detection system realizes a complete closed loop from data annotation, model training, model deployment to online real-time compliance detection. The construction and deployment of the visual model can be completed without the need for subsequent actual users to write code, which greatly reduces the technical threshold and deployment cost of intelligent compliance detection in assembly operations. At the same time, it can identify various error types in the assembly process in real time and accurately, effectively improving assembly quality and production efficiency.

[0012] Preferably, the graphical collaborative annotation unit includes:

[0013] The display module is used to display the standard operation video imported by the user and the annotation results generated after the user performs annotation operations on the standard operation video. The annotation results include part category name, installation position and assembly posture, assembly action and process sequence.

[0014] The part labeling sub-unit is used by the user to select the target part on the keyframe of the standard operation video by box selection, and to edit the part category name for the selected target part.

[0015] The positional relationship annotation subunit is used by the user to annotate the correct installation position and assembly posture of the target part relative to the main body component;

[0016] The action definition and annotation subunit is used to allow users to select target action segments on the video timeline of the standard operation video, annotate the corresponding assembly actions for the target action segments, and form a standard assembly action library.

[0017] The process sequence definition sub-unit allows users to edit the sequence of all labeled assembly actions in chronological order to form a standard process sequence, and automatically generates an adjustable standard process template;

[0018] The standard operation video is pre-recorded by senior technicians according to the standard operating procedure, and the standard process sequence corresponds one-to-one with the assembly steps defined in the standard operating procedure.

[0019] The annotation results generated by the display module, the part annotation subunit, the positional relationship annotation subunit, the action definition and annotation subunit, and the process sequence definition subunit together constitute the structured annotation data. Specifically, the part category name constitutes the part label; the positional relationship annotation subunit generates the true relative position offset and true relative angle based on the user-annotated installation position and assembly posture; the assembly action constitutes the action label; and the standard process sequence constitutes the process sequence label. By decomposing the complex annotation task into intuitive graphical operations, non-algorithm personnel can easily annotate part labels, true relative position offsets and true relative angles, action labels, and process sequence labels. Standard operation videos are pre-recorded by senior technicians according to standard operating procedures, ensuring the authority of the supervision signal. The standard process sequence corresponds one-to-one with the assembly steps defined in the standard operating procedures, ensuring strict consistency between the training data and process standards. This refined collaborative annotation method provides the multi-task compliance judgment model with complete structured annotation data covering parts, positions, postures, actions, and assembly sequences, effectively improving the accuracy and reliability of subsequent model training.

[0020] Preferably, the joint training engine adopts a convolutional neural network architecture based on multi-task learning, and uses the structured labeled data to train the multi-task compliance judgment model. The joint training engine specifically includes:

[0021] A shared spatiotemporal feature extraction backbone network is constructed with a MobileNetV3 network integrating a temporal displacement module as its core. This network is used to receive the standard operation video, extract spatial features frame by frame, and perform feature displacement in the channel dimension through the temporal displacement module to fuse temporal context information and output a shared spatiotemporal feature map sequence.

[0022] The part detection head, connected to the shared spatiotemporal feature extraction backbone network, is used to predefine a set of anchor boxes at each scale of the shared spatiotemporal feature map sequence, predict the bounding box coordinate offset, object presence confidence, and part category probability of each anchor box, and calculate the candidate bounding box prediction value of each target part based on the bounding box coordinate offset and the corresponding anchor box coordinate. The dimension of the part category probability is pre-configured by the joint training engine according to the total number of part categories annotated by the user in the structured annotation data. The part detection head multiplies the object presence confidence of each object with the corresponding part category probability to obtain the candidate comprehensive confidence of each anchor box for each part category, and outputs the target bounding box prediction value and the corresponding part presence confidence prediction value of each target part after non-maximum suppression processing based on the candidate bounding box prediction value and the candidate comprehensive confidence value of each target part.

[0023] The spatial relationship regression head is connected to the shared spatiotemporal feature extraction backbone network. Guided by the bounding box prediction value output by the part detection head, the spatial relationship regression head extracts the local feature map of each target part from the shared spatiotemporal feature map sequence using the region of interest alignment operation. The local feature map is then fed into the lightweight convolutional network inside the spatial relationship regression head. The lightweight convolutional network performs regression processing on the local feature map to obtain the predicted value of the relative position offset and the relative angle of the target part relative to the main component.

[0024] The temporal action classification head is connected to the shared spatiotemporal feature extraction backbone network. The temporal action classification head performs global average pooling on the shared spatiotemporal feature map sequence in the time and space dimensions to obtain temporal feature vectors. The temporal feature vectors are then input into the Softmax classification layer set inside the temporal action classification head, and the predicted action matching degree of the standard operation video corresponding to each assembly action in the standard assembly action library is output.

[0025] The process diagram state verification module takes as input a comprehensive output sequence from the part detection head, the spatial relationship regression head, and the temporal action classification head over a preset time period. The module processes this comprehensive output sequence using a pre-constructed directed acyclic graph (DAG) graph convolutional network, which contains at least one graph convolutional layer. The module temporally aggregates and concatenates the predicted part presence confidence, relative position offset, relative angle, and action matching degree from the comprehensive output sequence to form initial feature vectors for each assembly step. After neighborhood information aggregation by the graph convolutional layer, the module outputs the predicted probability distribution of each assembly step in the standard process sequence corresponding to the comprehensive output sequence. The maximum value in the predicted probability distribution is the sequence progress confidence of the comprehensive output sequence.

[0026] During training, the joint training engine compares the predicted values ​​of part presence confidence, relative position offset, and relative angle, as well as the predicted values ​​of action matching and the predicted probability distribution, with the structured labeled data and updates the network weights of the shared spatiotemporal feature extraction backbone network, the part detection head, the spatial relationship regression head, the temporal action classification head, and the process diagram state verification module until the preset maximum number of training iterations is reached, thus completing the training process and obtaining a multi-task compliance judgment model. The no-code multi-task model building platform distributes the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal. By uniformly processing standard operation videos through the shared spatiotemporal feature extraction backbone network, the computational complexity caused by building independent models for parts, positions, actions, and sequences is avoided. The part detection head, spatial relationship regression head, temporal action classification head, and process diagram state verification module share underlying features and each outputs the predicted values ​​of part presence confidence, relative position offset, and relative angle, the predicted values ​​of action matching and the predicted probability distribution, achieving efficient parallel extraction of multi-dimensional compliance information. By jointly updating network weights during training, the various tasks are optimized collaboratively, improving the overall performance and compactness of the model, which facilitates effective improvement of detection efficiency on edge-side real-time detection and decision-making terminals.

[0027] Preferably, the training process in the joint training engine is as follows:

[0028] The joint training engine compares the predicted confidence value of the part with the corresponding part label in the structured annotation data to obtain the part detection loss;

[0029] The joint training engine compares the predicted relative position offset with the actual relative position offset and the predicted relative angle with the actual relative angle to obtain the position regression loss.

[0030] The joint training engine compares the predicted action matching degree with the action label to obtain the action classification loss;

[0031] The joint training engine compares the predicted probability distribution with the process sequence labels to obtain the sequence verification loss;

[0032] The joint training engine updates the network weights of the shared spatiotemporal feature extraction backbone network, the part detection head, the spatial relationship regression head, the temporal action classification head, and the process diagram state verification module based on the calculated part detection loss, position regression loss, action classification loss, and sequence verification loss, until the preset maximum number of training iterations is reached, thus completing the training process and obtaining a multi-task compliance judgment model. The no-code multi-task model building platform then distributes the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal. This multi-loss collaborative update strategy enables supervised joint optimization of the four tasks, balancing the learning progress of each task and preventing any one task from dominating the training. This makes the multi-task compliance judgment model more comprehensive in terms of part detection, position determination, action recognition, and sequence verification.

[0033] Preferably, the directed acyclic graph structured graph convolutional network is constructed in the following manner:

[0034] Each node in the directed acyclic graph corresponds to an assembly step defined in the standard operating procedure, and the node number corresponds one-to-one with the step number of the assembly step.

[0035] The edges in the directed acyclic graph are established according to the sequence of steps and logical dependencies defined in the standard operating procedure, including directed edges connecting consecutive steps and directed edges connecting non-consecutive steps with logical dependencies.

[0036] The process diagram state verification module performs temporal pooling on the part existence confidence prediction value, relative position offset prediction value, relative angle prediction value, and action matching degree prediction value output by the part detection head, the spatial relationship regression head, and the temporal action classification head in the corresponding video segments within a preset time period to obtain the corresponding feature components. The module then concatenates the feature components to form the initial feature vector of each step node.

[0037] The first layer of the graph convolutional network performs a linear transformation on the initial feature vector and aggregates neighborhood features based on the normalized adjacency matrix, and outputs the first layer node features after processing by the ReLU activation function; the second layer of the graph convolutional network performs a linear transformation on the first layer node features and aggregates neighborhood features based on the normalized adjacency matrix, and outputs the second layer node features after processing by the ReLU activation function.

[0038] The process graph state verification module performs global average pooling on the second-layer node features to obtain a global graph representation vector. This global graph representation vector is then input into the Softmax classification layer within the process graph state verification module. The module outputs the probability distribution of each assembly step corresponding to a video segment within a preset time period, and takes the maximum value of this probability distribution as the sequence progress confidence score for that video segment. Assembly steps are explicitly modeled as graph nodes, and directed edges represent the sequence order and logical dependencies between steps defined in the standard operating procedure, enabling the model to capture complex constraints between discontinuous steps. The final output sequence progress confidence score accurately reflects the actual progress of the current video segment in the standard process sequence, significantly improving the ability to identify process skips and sequence reversals.

[0039] Preferably, the joint training engine further includes a lightweight model synthesis module, which is used to perform the following operations after training:

[0040] The trained multi-task compliance judgment model is subjected to channel pruning and parameter quantization to obtain an optimized multi-task compliance judgment model.

[0041] The optimized multi-task compliance judgment model is converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal, and then serialized and encapsulated to generate a lightweight model file.

[0042] The lightweight model file is then distributed to the edge-side real-time detection and decision-making terminal. After training, channel pruning and parameter quantization are performed on the multi-task compliance judgment model, compressing the model size and reducing inference computation. This model is then converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal and serialized and packaged to generate a lightweight model file, which is then distributed to the edge-side real-time detection and decision-making terminal. The multi-task compliance judgment model, which originally required high computing power, can now run in real-time on embedded edge computing devices near the production line, reducing hardware costs and simplifying model deployment and upgrade processes, while ensuring the real-time performance of detection and the maintainability of the system.

[0043] Preferably, the edge-side real-time detection and decision-making terminal includes:

[0044] A multi-dimensional confidence fusion and decision-making unit is used to receive the real-time output of the multi-task compliance judgment model to the real-time assembly video stream. The real-time output includes: the confidence of part existence output by the part detection head, the position offset output by the spatial relationship regression head based on relative position offset and relative angle, the action matching degree output by the temporal action classification head, and the confidence of sequence progress output by the process diagram status verification module. The multi-dimensional confidence fusion and decision-making unit pre-stores decision logic rules. The decision logic rules preset a first threshold corresponding to the confidence of part existence, a second threshold corresponding to the position offset, a third threshold corresponding to the action matching degree, and a fourth threshold corresponding to the confidence of sequence progress. The multi-dimensional confidence fusion and decision-making unit performs a comprehensive calculation by comparing the confidence of part existence, the position offset, the action matching degree, and the confidence of sequence progress with the corresponding first threshold, second threshold, third threshold, and fourth threshold, respectively, and generates a detailed compliance judgment containing four defect dimensions: part error, position error, action error, and sequence error based on the comprehensive calculation result.

[0045] The error tracing and alarm generator, connected to the multi-dimensional confidence fusion and decision-making unit, determines at least one major defect dimension among the four defect dimensions that leads to non-compliance when the detailed compliance judgment is unqualified. It then maps each major defect dimension to a specific error type, triggering corresponding visual and voice alarm prompts. This effectively reduces the risk of false alarms based on a single dimension. When a judgment is unqualified, the error tracing and alarm generator automatically identifies the major defect dimension causing the non-compliance and maps each major defect dimension to a specific error type, triggering visual and voice alarm prompts to quickly guide the operator to correct the problem.

[0046] A detection method based on the intelligent compliance detection system for assembly operations described above includes the following steps:

[0047] Step S1: The camera is used to collect real-time assembly video streams from the assembly station and transmits the real-time assembly video streams to the edge-side real-time detection and decision-making terminal.

[0048] Step S2: Generate a multi-task compliance judgment model through the no-code multi-task model building platform. The generation step includes: receiving structured annotation data generated by the user annotating the standard operation video through the graphical collaborative annotation unit, and automatically training the multi-task compliance judgment model based on the structured annotation data by the joint training engine.

[0049] Step S3: Deploy the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal;

[0050] Step S4: The edge-side real-time detection and decision-making terminal loads and runs the multi-task compliance judgment model to analyze the real-time assembly video stream to identify whether there are part errors, position errors, action errors and sequence errors in the current assembly operation.

[0051] Step S5: The edge-side real-time detection and decision-making terminal outputs detailed compliance judgments based on the analysis results and triggers corresponding alarm signals. A complete end-to-end standardized workflow is defined. This detection method organically links graphical collaborative annotation, automated training, edge-side real-time detection, and decision-making, eliminating the need for subsequent code writing by actual users. This significantly lowers the technical implementation threshold, enabling rapid replication and deployment across different production lines and ensuring the consistency and reliability of assembly compliance detection.

[0052] Preferably, the specific steps of step S4 are as follows:

[0053] Step S4.1: Receive the confidence level, position offset, action matching degree, and sequence progress confidence level of the parts output by the multi-task compliance judgment model for the real-time assembly video stream.

[0054] Step S4.2: The multi-dimensional confidence fusion and decision-maker performs a comprehensive calculation on the confidence of the existence of the part, the position offset, the action matching degree, and the sequence progress confidence based on the pre-stored decision logic rules, and generates a preliminary qualified or unqualified decision.

[0055] Step S4.3: When the judgment is unqualified, the error tracing and alarm generator determines at least one major defect dimension causing the unqualification based on the comprehensive calculation result, and maps each major defect dimension to a specific error type. The multi-dimensional confidence fusion and judgment unit, based on pre-stored judgment logic rules, performs comprehensive calculations on the received part existence confidence, position offset, action matching degree, and sequence progress confidence to generate a preliminary qualified or unqualified judgment, realizing the logical fusion of multi-dimensional information. When the judgment is unqualified, the error tracing and alarm generator determines at least one major defect dimension based on the comprehensive calculation result and directly maps each major defect dimension to a specific error type, making the alarm information accurate to the specific defect dimension and error form, improving the relevance and traceability of the alarm, and significantly shortening the operator's response and error correction time. Attached Figure Description

[0056] Figure 1 This is a block diagram of the composition structure of the detection system of the present invention;

[0057] Figure 2 This is a flowchart of the detection method of the present invention;

[0058] Figure 3 This is the user interface of the no-code multitasking model building platform in the configuration state in this embodiment.

[0059] Figure 4 This is the user interface in this embodiment when annotation is not completed on the no-code multi-task model building platform;

[0060] Figure 5 This is the user interface for action annotation and process sequence annotation on the no-code multi-task model building platform in this embodiment;

[0061] Figure 6 This is the user interface for completing a step-related annotation task on a no-code multi-task model building platform in this embodiment;

[0062] Figure 7 This embodiment shows the real-time detection results of the second assembly operation at the actual assembly station on a display screen connected to the edge-side real-time detection and decision-making terminal.

[0063] Figure 8 This embodiment shows the real-time detection results of the fifth assembly operation at the actual assembly station on a display screen connected to the edge-side real-time detection and decision-making terminal.

[0064] Figure 9 This is a physical image of the edge-side real-time detection and decision-making terminal used in this embodiment. Detailed Implementation

[0065] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.

[0066] Example: Figures 1-9 As shown in the figure, this embodiment provides an intelligent compliance detection system for assembly operations, the architecture of which is as follows: Figure 1 As shown, it mainly consists of two parts: a no-code multi-task model building platform deployed in the cloud or locally, and an edge-side real-time detection and decision-making terminal deployed on the production line.

[0067] I. No-code multi-task model building platform.

[0068] The no-code multi-task model building platform includes a graphical collaborative annotation unit 1 and a joint training engine 2.

[0069] The graphical collaborative annotation unit 1 provides a user interface that supports users in importing standard operating procedure (SOP) videos, CAD drawings, and SOP text files. The SOP videos are pre-recorded by senior technicians according to standard operating procedures. Users annotate parts, installation positions and assembly postures, assembly actions, and process sequences in the videos through graphical interaction within the graphical collaborative annotation unit 1. The graphical collaborative annotation unit 1 outputs structured annotation data based on the user's annotation operations. Annotation information includes parts, installation positions and assembly postures, assembly actions, and process sequences.

[0070] The graphical collaborative annotation unit 1 further includes a display module 11, a part annotation subunit 12, a positional relationship annotation subunit 13, an action definition and annotation subunit 14, and a process sequence definition subunit 15. The display module 11 displays the standard operation video imported by the user and the annotation results generated after the user performs annotation operations on the standard operation video. The part annotation subunit 12 allows the user to select target parts on keyframes of the standard operation video using a box selection operation and edit the part category name for the selected target parts. The positional relationship annotation subunit 13 allows the user to annotate the correct installation position and assembly posture of the target parts relative to the main component. The action definition and annotation subunit 14 allows the user to select target action segments on the video timeline of the standard operation video, annotate the corresponding assembly actions for the target action segments, and form a standard assembly action library. The process sequence definition subunit 15 allows the user to edit the sequence of all annotated assembly actions in chronological order to form a standard process sequence and automatically generates an adjustable standard process template. The annotation results generated by the above sub-units together constitute structured annotation data. Among them, the part category name constitutes the part label, the positional relationship annotation sub-unit 13 generates the actual relative position offset and actual relative angle based on the installation position and assembly posture annotated by the user, the assembly action constitutes the action label, and the standard process sequence constitutes the process sequence label. The standard process sequence corresponds one-to-one with the assembly steps defined in the standard operating procedure.

[0071] The joint training engine 2 is connected to the graphical collaborative annotation unit 1 to automatically train a multi-task compliance judgment model based on structured labeled data. The joint training engine 2 employs a convolutional neural network architecture based on multi-task learning, using structured labeled data to train the multi-task compliance judgment model. During training, the joint training engine 2 automatically loads a pre-trained model centered on a MobileNetV3 network integrating a temporal shift module as a shared spatiotemporal feature extraction backbone network 21. This shared spatiotemporal feature extraction backbone network 21 receives standard operation videos, extracts spatial features frame by frame, and performs feature shifting in the channel dimension through the temporal shift module to fuse temporal context information, outputting a sequence of shared spatiotemporal feature maps.

[0072] On top of the shared spatiotemporal feature extraction backbone network 21, a joint training engine 2 is constructed to build four task heads for collaborative training:

[0073] The part detection head 22 is connected to the shared spatiotemporal feature extraction backbone network 21. At each scale of the shared spatiotemporal feature map sequence, a set of anchor boxes is predefined. The bounding box coordinate offset, object presence confidence, and part category probability of each anchor box are predicted. The candidate bounding box prediction value of each target part is calculated based on the bounding box coordinate offset and the corresponding anchor box coordinate. The candidate comprehensive confidence value is obtained by multiplying the object presence confidence value by the corresponding part category probability. After non-maximum suppression processing based on the candidate bounding box prediction value and the candidate comprehensive confidence value of each target part, the target bounding box prediction value and the corresponding part presence confidence value of each target part are output.

[0074] The spatial relationship regression head 23 is connected to the shared spatiotemporal feature extraction backbone network 21. Guided by the target bounding box prediction value output by the part detection head 22, it extracts the local feature map of each target part from the shared spatiotemporal feature map sequence using a region of interest alignment operation. This region of interest alignment operation preserves the spatial position information of the part by uniformly setting sampling points within the region of interest and calculating the precise feature value of each sampling point using bilinear interpolation. The spatial relationship regression head 23 feeds the local feature map into its internal lightweight convolutional network, which performs regression processing on the local feature map to obtain the predicted values ​​of the relative position offset and relative angle of the target part relative to the main component, used to detect position errors. During the training phase, the lightweight convolutional network is trained based on the true relative position offset and true relative angle in the structured labeled data as supervision signals. The network weights are gradually adjusted through backpropagation until the position and angle information can be accurately regressed from the local feature map.

[0075] The temporal action classification head 24 is connected to the shared spatiotemporal feature extraction backbone network 21. It performs global average pooling on the shared spatiotemporal feature map sequence in the time and space dimensions to obtain the temporal feature vector. The temporal feature vector is then input into the Softmax classification layer set inside it, and outputs the action matching degree prediction value of each assembly action in the standard assembly action library corresponding to the standard operation video, which is used to detect action errors.

[0076] The input to the process diagram state verification module 25 is the comprehensive output sequence of the part detection head 22, the spatial relationship regression head 23, and the temporal action classification head 24 within a preset time period. The process diagram state verification module 25 processes the comprehensive output sequence based on a pre-constructed directed acyclic graph (DAG) structure graph convolutional network. It performs temporal aggregation and concatenation on the predicted values ​​of part existence confidence, relative position offset, relative angle, and action matching degree in the comprehensive output sequence to obtain corresponding feature components. These feature components are then concatenated to form the initial feature vectors for each assembly step. After neighborhood information aggregation through the graph convolutional layer, the output comprehensive output sequence corresponds to the predicted probability distribution of each assembly step in the standard process sequence. The maximum value of this distribution is the sequence progress confidence of the comprehensive output sequence, used to detect sequence errors.

[0077] The graph convolutional network described above is constructed as follows: Each node in the directed acyclic graph (DAG) corresponds to an assembly step defined in the standard operating procedure (SOP), with a one-to-one correspondence between node numbers and step numbers. Edges in the DAG are established based on the sequence and logical dependencies between steps as defined in the SOP, including directed edges connecting consecutive steps and directed edges connecting non-consecutive steps with logical dependencies. The graph convolutional network contains at least one graph convolutional layer. The first layer performs a linear transformation on the initial feature vector and performs neighborhood feature aggregation based on a normalized adjacency matrix, then processes it using a ReLU activation function to output the first-layer node features. The second layer performs a linear transformation on the first-layer node features and performs neighborhood feature aggregation based on a normalized adjacency matrix, then processes it using a ReLU activation function to output the second-layer node features. After two layers of graph convolution, the feature vector of each step node incorporates information from its two-hop neighbors, achieving global process context awareness. Finally, global average pooling is performed on the features of the second layer nodes to obtain the global graph representation vector. The global graph representation vector is then input into the Softmax classification layer inside the process graph state verification module 25 to output the probability distribution corresponding to each assembly step.

[0078] During the training process, for each batch of standard operation video data, the part detection head 22 outputs the predicted value of the existence confidence of each target part, the spatial relationship regression head 23 outputs the predicted value of the relative position offset and the relative angle, the temporal action classification head 24 outputs the predicted value of the action matching degree, and the process diagram state verification module 25 outputs the predicted probability distribution.

[0079] The joint training engine 2 compares each of the predicted values ​​with the corresponding part labels, true relative position offsets and true relative angles, action labels, and process sequence labels in the structured labeled data, and calculates the part detection loss L_obj, position regression loss L_pos, action classification loss L_act, and sequence validation loss L_seq. Specifically, the part detection loss uses the focus loss function, the position regression loss uses the smoothing L1 loss function, and both the action classification loss and sequence validation loss use the cross-entropy loss function.

[0080] The joint training engine 2 calculates the joint loss function L_total: L_total = 1.0 × L_obj + 1.5 × L_pos + 1.0 × L_act + 0.5 × L_seq. The initial weights are set based on the loss ratio of the first round of training, and are fine-tuned every 5 training cycles based on the rate of loss decrease for each task. The joint training engine 2 updates the network weights of the shared spatiotemporal feature extraction backbone network 21, the part detection head 22, the spatial relationship regression head 23, the temporal action classification head 24, and the process diagram state verification module 25 based on the calculated joint loss until convergence, resulting in the trained multi-task compliance judgment model.

[0081] After training, the lightweight model synthesis module in the joint training engine 2 performs compression and optimization operations on the trained multi-task compliance judgment model: Based on the impact of each network layer on model accuracy, the pruning rate and quantization accuracy are configured differently. Network layers with a greater impact on accuracy are pruned with a lower pruning rate and higher quantization accuracy, while those with a smaller impact are pruned with a higher pruning rate and lower quantization accuracy. In this embodiment, the pruning rate ranges from 10% to 45%, and the quantization accuracy is INT8. The optimized model is then converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal 4, and serialized and packaged to generate a lightweight model file. In this embodiment, the edge computing device used is the RK3588 embedded AI computing device, and the lightweight model file format is the RKNN model file. Finally, the lightweight model file is distributed to the edge-side real-time detection and decision-making terminal 4.

[0082] II. Edge-side real-time detection and decision-making terminal 4 and camera 3.

[0083] The edge-side real-time detection and decision-making terminal 4 is an embedded edge computing device deployed next to the production line, including a multi-dimensional confidence fusion and decision-making unit 41 and an error tracing and alarm generator 42.

[0084] Camera 3 is used to collect real-time assembly video streams consisting of real-time images of parts and real-time action images of workers at the assembly station, and transmit them to the internal processing unit of the edge-side real-time detection and decision-making terminal 4.

[0085] The edge-side real-time detection and decision-making terminal 4 loads and runs a lightweight multi-task compliance judgment model to analyze the real-time assembly video stream. The multi-task compliance judgment model outputs real-time results in four dimensions to the real-time assembly video stream: the confidence level of part existence output by the part detection head 22, the position offset output by the spatial relationship regression head 23 based on the relative position offset and relative angle, the action matching degree output by the temporal action classification head 24, and the confidence level of sequence progress output by the process diagram status verification module 25.

[0086] The multi-dimensional confidence fusion and decision-making unit 41 has pre-stored decision logic rules, including a first threshold corresponding to the confidence of the part's existence, a second threshold corresponding to the positional offset, a third threshold corresponding to the action matching degree, and a fourth threshold corresponding to the confidence of the sequence progress. The multi-dimensional confidence fusion and decision-making unit 41 performs a comprehensive calculation by comparing the confidence of the above four dimensions with their corresponding thresholds, and generates a detailed compliance decision based on the comprehensive calculation result, which includes four defect dimensions: part error, position error, action error, and sequence error.

[0087] The error tracing and alarm generator 42 is connected to the multi-dimensional confidence fusion and decision-making unit 41. When the detailed compliance decision is unqualified, the error tracing and alarm generator 42 determines at least one major defect dimension among the four defect dimensions that leads to the unqualification based on the comprehensive calculation results, and maps each major defect dimension to a specific error type, triggering the corresponding visual and voice alarm prompts. Human-computer interaction can be achieved through at least one of the following: indicator lights, display screen, voice speaker, or Andon system.

[0088] III. Overall System Workflow.

[0089] like Figure 2 As shown, the system's workflow is divided into a training phase and an inference phase. In the training phase, users annotate standard operation videos and output structured annotation data through the graphical collaborative annotation unit 1. The joint training engine 2 trains a multi-task compliance judgment model based on the structured annotation data. The lightweight model synthesis module then performs lightweight processing on the model and distributes the lightweight model file to the edge-side real-time detection and decision-making terminal 4. In the inference phase, the edge-side real-time detection and decision-making terminal 4 loads and runs the lightweight model file, performing online detection on the real-time assembly video stream captured by the camera 3. The multi-dimensional confidence fusion and decision-maker 41 generates a compliance judgment, and the error tracing and alarm generator 42 performs error tracing and alarm prompts when non-compliance occurs, forming a closed loop of perception, decision-making, and feedback. The edge-side real-time detection and decision-making terminal 4 is typically connected to a display screen, enabling real-time display of the detection results.

[0090] IV. Case Study 1: Intelligent Detection of Gear Installation Station in Paper Shredder.

[0091] The following section uses the shredder gearbox assembly station of an office equipment manufacturing plant as an example to provide a detailed description of the specific implementation method of this embodiment.

[0092] The assembly task at this workstation requires the operator to correctly install the fixing bolts, locating shims, driven pinion, and driving gear. The standard operating procedure defines the assembly sequence as follows: first, place the fixing bolts; second, install the locating shims; then, press in the driven pinion; tighten the driven pinion with a wrench; press in the driven pinion again; next, press in the driving gear; and finally, clamp with pliers to ensure gear engagement.

[0093] (a) Data collection and construction of standard sample library.

[0094] Two industrial cameras (3) are deployed at the shredder gearbox assembly station, one providing a top-down view and the other a side view, both aimed at the operating area. A senior technician performs the complete assembly operation according to standard operating procedures, and the system records the standard operation video via cameras (3). Simultaneously, users can input AutoCAD drawings and standard operating procedure text files into the no-code multi-task model building platform as supplementary information. The recorded video and input files constitute the foundational data for subsequent annotation and training.

[0095] (ii) Definition and training of no-code models.

[0096] The process engineer performs the following annotation operations on the graphical collaborative annotation unit 1 interface of the no-code multi-task model building platform:

[0097] Part selection and labeling: In the video keyframe, select the main components and then edit the part category names of the fixing bolt, positioning shim, driven pinion, driving pinion, wrench, and pliers in sequence as: "Bolt", "GasketL", "GearS", "GearL", "Wrench", and "Pliers".

[0098] Positional relationship annotation: In another keyframe, the alignment keyway position of "Bolt" with the shaft after full insertion is annotated as the correct installation position, as well as the standard meshing position and attitude of "GearS" and "GearL". Positional relationship annotation sub-unit 13 generates the true relative position offset and true relative angle based on the installation position and assembly attitude annotated by the user.

[0099] Action definition and annotation: On the video timeline, select the time period from "hand places Bolt" to "Bolt is inserted into place" and select the "Place" label from the standard assembly action library; select the time period from "wrench contacts GearS" to "wrench stops" and select the "Tighten" label from the standard assembly action library.

[0100] Process sequence definition: The graphical collaborative annotation unit 1 automatically generates the sequence template: Step 1: Place Bolt; Step 2: Place GasketL; Step 3: Place GearS; Step 4: Tighten Wrench; Step 5: Place GearS; Step 6: Place GearL; Step 7: Clamp Pliers. This standard process sequence corresponds one-to-one with the assembly steps defined in the standard operating procedure. The process engineer confirms that this order cannot be reversed.

[0101] The joint training engine 2 automatically executes the training process based on the structured annotation data generated by the above annotations. Among them, the part category names in the structured annotation data constitute part labels, the values ​​generated by the positional relationship annotation subunit 13 constitute the true relative position offset and the true relative angle, the assembly actions constitute action labels, and the standard process sequence constitutes process sequence labels.

[0102] During training, the shared spatiotemporal feature extraction backbone network 21 takes the MobileNetV3 network with integrated temporal shift module as its core, receives standard operation videos, extracts spatial features frame by frame, and performs feature shifting in the channel dimension through the temporal shift module to fuse temporal context information and output a shared spatiotemporal feature map sequence.

[0103] The part detection head 22 predefines a set of anchor boxes at each scale of the shared spatiotemporal feature map sequence, predicting the bounding box coordinate offset, object presence confidence, and part category probability for each anchor box. The dimension of the part category probability is pre-configured by the joint training engine 2 based on the total number of user-annotated part categories in the structured labeled data; in this case, the total number of part categories is 7. The part detection head 22 calculates the candidate bounding box prediction value based on the bounding box coordinate offset and the corresponding anchor box coordinates. It multiplies the object presence confidence by the corresponding part category probability to obtain the candidate comprehensive confidence, and outputs the target bounding box prediction value and the corresponding part presence confidence prediction value for each target part after non-maximum suppression processing.

[0104] The spatial relationship regression head 23, guided by the target bounding box prediction value output by the part detection head 22, extracts the local feature map of each target part from the shared spatiotemporal feature map sequence using the region of interest alignment operation. This local feature map is then fed into an internal lightweight convolutional network for regression processing to obtain the predicted values ​​of relative position offset and relative angle. During the training phase, the lightweight convolutional network performs supervised learning based on the true relative position offset and true relative angle. By calculating the position regression loss and backpropagating the gradient, the network weights are gradually adjusted until the position and angle information can be accurately regressed from the local feature map.

[0105] The temporal action classification head performs global average pooling on 24 pairs of shared spatiotemporal feature map sequences in both time and space dimensions to obtain temporal feature vectors, which are then output as action matching degree predictions by the internal Softmax classification layer.

[0106] The process diagram state verification module 25 takes the comprehensive output sequence of the part inspection head 22, spatial relationship regression head 23, and temporal action classification head 24 over a preset time period as input, and processes it based on a pre-constructed directed acyclic graph (DAG) structured graph convolutional network. The DAG contains 7 nodes, each corresponding to one of the 7 assembly steps defined in the standard operating procedure. The node feature vector is obtained by temporally pooling the predicted values ​​of part existence confidence, relative position offset, relative angle, and action matching degree in the comprehensive output sequence and then concatenating them. After two layers of graph convolution to aggregate neighborhood information, the output comprehensive output sequence corresponds to the predicted probability distribution of each assembly step in the standard process sequence.

[0107] The joint training engine 2 compares the above predicted values ​​with the corresponding part labels, true relative position offsets and true relative angles, action labels and process sequence labels in the structured labeled data, respectively, and calculates the part detection loss, position regression loss, action classification loss and sequence verification loss accordingly. The network weights are updated based on each loss until convergence, and the trained multi-task compliance judgment model is obtained.

[0108] (III) Lightweight model synthesis and distribution.

[0109] After training, the lightweight model synthesis module in the joint training engine 2 performs compression and optimization operations on the trained multi-task compliance judgment model: the pruning rate and quantization accuracy are configured differently based on the impact of each network layer on the model's accuracy. For layers with a significant impact on accuracy, such as the network layer responsible for target classification in the part detection head 22, a lower pruning rate and higher quantization accuracy are used; for layers with a smaller impact on accuracy, such as the feature pruning layer in the spatial relationship regression head 23, a higher pruning rate and lower quantization accuracy are used. In this embodiment, the pruning rate ranges from 10% to 45%, and the quantization accuracy is INT8. The optimized model is then converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal 4 and serialized and packaged. In this embodiment, the edge computing device is the RK3588 embedded AI computing device, and the lightweight model file format is the RKNN model file. Finally, the lightweight model file is distributed to the edge-side real-time detection and decision-making terminal 4.

[0110] (iv) Online real-time detection and judgment.

[0111] The edge-side real-time detection and decision-making terminal 4 loads and runs a lightweight model file to analyze the 30FPS real-time assembly video stream acquired by the camera 3.

[0112] The multi-task compliance judgment model outputs real-time results in four dimensions for each video segment: the part detection head 22 outputs the confidence of the existence of each part, the spatial relationship regression head 23 calculates and outputs the position offset based on the relative position offset and relative angle, the temporal action classification head 24 outputs the action matching degree, and the process diagram status verification module 25 outputs the confidence of the sequence progress.

[0113] The multi-dimensional confidence fusion and decision maker 41 has pre-stored decision logic rules. In this practical application case 1, the preset first threshold is 0.9 (corresponding to the confidence of part existence), the second threshold is 0.1 (corresponding to the position offset), the third threshold is 0.85 (corresponding to the action matching degree), and the fourth threshold is 0.9 (corresponding to the confidence of sequence progress). When the confidence of part existence of all parts is greater than 0.9, the position offset is less than 0.1, the action matching degree is greater than 0.85, the confidence of sequence progress is greater than 0.9, and the sequence order is correct, a qualified decision is generated; otherwise, an unqualified decision is generated.

[0114] When generating a non-conformance judgment, the error tracing and alarm generator 42 determines at least one major defect dimension among the four defect dimensions that leads to non-conformance based on the comprehensive calculation results, and maps each major defect dimension to a specific error type. For example: if an incorrect part is detected, the confidence level of the corresponding part is too low, triggering a part error and prompting an alarm "Part Error: Mistakenly picked up the drive gear"; if the gear is not fully engaged, the position offset exceeds the standard, triggering a position error and prompting an alarm "Position Error: Gear not fully engaged"; if the bolt is not tightened, the motion matching is too low, triggering a motion error and prompting an alarm "Motion Error: Bolt not tightened"; if the operation sequence is reversed (e.g., placing the gear before placing the shim), the sequence progress confidence level is too low or the sequence is incorrect, triggering a sequence error and prompting an alarm "Sequence Error: Process sequence reversed". The error tracing and alarm generator 42 illuminates the red light at that workstation and plays the specific error prompt voice through the loudspeaker.

[0115] (v) As a further optimization method, this system can also support active learning closed loop.

[0116] After the system had been running for a period of time, the no-code multi-task model building platform detected that the action matching degree of the multi-task compliance judgment model for the action of "clamping with pliers" was consistently low in newly added operator videos. The edge-side real-time detection and decision-making terminal 4 automatically transmitted the aforementioned low-confidence video segments and their associated prediction data back to the no-code multi-task model building platform. The no-code multi-task model building platform automatically filtered out difficult samples according to the preset confidence threshold and pushed the difficult samples to the annotation interface of the graphical collaborative annotation unit 1. After verification, the process engineer found that the newly purchased pliers on the production line were of a different model and had different appearances than those used during training. Therefore, several new pliers clamping action video segments were added to the original annotations. The joint training engine 2 added the newly annotated samples to the training set and incrementally fine-tuned the multi-task compliance judgment model. The updated lightweight model file was redeployed to the edge-side real-time detection and decision-making terminal 4, resolving the false alarm problem.

[0117] V. Case Study 2: Intelligent Inspection at the Assembly and Packaging Station of Data Acquisition Gateway

[0118] The following example, using the final assembly and packaging station of a data acquisition gateway from a certain equipment manufacturer, illustrates another application scenario of this embodiment.

[0119] The assembly and packaging task at this workstation requires the operator to correctly place the gateway host, antenna, power adapter, and instruction manual into the packaging box and seal it. The standard operating procedure defines the following sequence of steps: Step 1, remove the empty packaging box and open the lid; Step 2, place the gateway host inside; Step 3, install and tighten the antenna; Step 4, place the power adapter inside; Step 5, place the instruction manual and warranty card in the correct orientation; Step 6, close the lid.

[0120] (a) Data collection and construction of standard sample library.

[0121] A high-definition industrial camera is deployed directly above the packaging station of the data acquisition gateway, ensuring its field of view fully covers the workbench and packaging box area. Packers perform multiple complete assembly and packaging operations according to standard operating procedures, and the system records these standard operation videos via camera 3. Simultaneously, a corresponding standard operation template is created in the no-code multi-task model building platform. The recorded videos and standard operating procedures constitute the foundational data for subsequent annotation and training.

[0122] (ii) Definition and training of no-code models.

[0123] The process engineer performs the following annotation operations on the graphical collaborative annotation unit 1 interface of the no-code multi-task model building platform:

[0124] Part selection and labeling: In the video keyframes, select and label the target parts in sequence, including the packaging box, gateway host, antenna, power adapter, instruction manual, and warranty card, and edit the corresponding part category name for each target part.

[0125] Positional relationship annotation: In the keyframes after the items are placed inside the packaging box, the correct installation position and assembly posture of each part are annotated. For example, it is annotated that the instruction manual should be placed in the designated area inside the box with the cover facing up, and that the power adapter should be inserted into the corresponding groove inside the box. The positional relationship annotation sub-unit 13 generates the actual relative position offset and actual relative angle of each part based on the user annotation.

[0126] Action definition and annotation: Select action segments on the video timeline and associate them with action tags. Select the time segment from "hand picks up the instruction manual" to "instruction manual is placed in the box and released" and associate it with the action tag "place"; select the time segment from "hand picks up the antenna" to "antenna is fixed and hand is released" and associate it with the action tag "install"; similarly, annotate other operation steps to form a standard assembly action library.

[0127] Process sequence definition: The graphical collaborative annotation unit 1 automatically generates an initial process sequence template based on the time sequence of the action annotations. The process engineer compares it with the standard operating procedure and fine-tunes the sequence to ensure that the standard process sequence corresponds one-to-one with the six assembly steps defined in the standard operating procedure.

[0128] The joint training engine 2 trains a multi-task compliance judgment model based on the structured labeled data generated by the above annotations. The model architecture is the same as that in Case 1. Among them, the part detection head 22 learns to identify all target parts such as packaging boxes, gateway hosts, antennas, power adapters, instruction manuals, and warranty cards; the spatial relationship regression head 23 learns to judge the angular offset of the instruction manual, whether the power adapter is placed in the specified groove, and other positional accuracy; the temporal action classification head 24 learns to distinguish different operation actions such as placement, installation, and packaging; and the process diagram state verification module 25 learns the compliance of the assembly sequence based on a pre-built graph convolutional network with a 6-node directed acyclic graph structure.

[0129] (III) Examples of online real-time detection and erroneous judgments.

[0130] After loading the lightweight model file, the edge-side real-time detection and decision-making terminal 4 performs online analysis on the real-time assembly video stream captured by camera 3. Two typical error scenarios are listed below:

[0131] Scenario 1: Warranty card missing and instruction manual placed in the wrong orientation. The multi-task compliance judgment model outputs: Warranty card presence confidence is 0.05 (lower than the first threshold of 0.9 set in this practical application case 2), and instruction manual position offset is 0.8 (higher than the second threshold of 0.2 set in this practical application case 2). The multi-dimensional confidence fusion and decision-making unit 41 generates a non-compliance judgment after comprehensive calculation. The error tracing and alarm generator 42 determines "part presence confidence" and "position offset" as the two main defect dimensions leading to non-compliance based on the comprehensive calculation results, mapping them to "part error: warranty card missing" and "position error: instruction manual placed in the wrong orientation," respectively, triggering corresponding visual and voice alarm prompts.

[0132] Scenario 2: Improper application of sealing tape and skipped steps. The multi-task compliance judgment model outputs: the current action matching degree is 0.6 (lower than the third threshold of 0.85 set in this practical application case 2), and the sequence progress confidence shows that step 3 (antenna installation) should be executed, but step 4 (placing the power adapter) has already been detected. Multi-dimensional confidence fusion and decision generator 41 generate a non-compliance judgment. The error tracing and alarm generator 42 identifies "action matching degree" and "sequence progress confidence" as the main defect dimensions, mapping them to "Action error: Improper antenna installation" and "Sequence error: Placing the power adapter without installing the antenna," respectively, triggering an alarm.

[0133] (iv) As a further optimization method, this system can also support active learning closed loop.

[0134] In the initial stage of system operation, the multi-task compliance judgment model frequently issued false alarms for the operation of "correctly placing the instruction manual" (which requires both the part to be present and its correct position) in videos of new employees performing the operation, misjudging some skewed but acceptable instruction manuals as missing. The edge-side real-time detection and decision terminal 4 automatically transmitted these low-confidence video clips back to the no-code multi-task model building platform. The no-code multi-task model building platform filtered them as difficult samples and pushed them to the graphical collaborative annotation unit 1. Process engineers supplemented the annotation with a batch of instruction manual samples with different skew angles but considered compliant. After incremental fine-tuning by the joint training engine 2, the model's judgment of the instruction manual placement posture is more in line with the actual requirements of the production line, and the false alarm rate has decreased significantly.

Claims

1. An intelligent compliance inspection system for assembly operations, characterized in that... include: No-code multi-task model building platform, edge-side real-time detection and decision-making terminal and camera; The no-code multi-task model building platform includes a graphical collaborative annotation unit and a joint training engine; The graphical collaborative annotation unit is used to provide a user interface to receive annotation information from users on imported standard operation videos through graphical interaction, and output structured annotation data; the annotation information includes parts, installation positions and assembly postures, assembly actions and process sequences. The aforementioned joint training engine is used to automatically train a multi-task compliance judgment model based on the aforementioned structured labeled data. The camera is used to collect a real-time assembly video stream consisting of real-time images of parts and real-time action images of workers on the production line, and transmit it to the edge-side real-time detection and decision-making terminal. The edge-side real-time detection and decision-making terminal is an embedded edge computing device deployed next to the production line. It is used to load and run the multi-task compliance judgment model, analyze the real-time assembly video stream, and output detailed compliance judgments and alarm signals including part errors, position errors, action errors, and sequence errors.

2. The intelligent compliance detection system for assembly operations according to claim 1, characterized in that, The graphical collaborative annotation unit includes: The display module is used to display the standard operation video imported by the user and the annotation results generated after the user performs annotation operations on the standard operation video. The annotation results include part category name, installation position and assembly posture, assembly action and process sequence. The part labeling sub-unit is used by the user to select the target part on the keyframe of the standard operation video by box selection, and to edit the part category name for the selected target part. The positional relationship annotation subunit is used by the user to annotate the correct installation position and assembly posture of the target part relative to the main body component; The action definition and annotation subunit is used to allow users to select target action segments on the video timeline of the standard operation video, annotate the corresponding assembly actions for the target action segments, and form a standard assembly action library. The process sequence definition sub-unit allows users to edit the sequence of all labeled assembly actions in chronological order to form a standard process sequence, and automatically generates an adjustable standard process template; The standard operation video is pre-recorded by senior technicians according to the standard operating procedure, and the standard process sequence corresponds one-to-one with the assembly steps defined in the standard operating procedure. The annotation results generated by the display module, the part annotation subunit, the positional relationship annotation subunit, the action definition and annotation subunit, and the process sequence definition subunit together constitute the structured annotation data. Among them, the part category name constitutes the part label, the positional relationship annotation subunit generates the true relative position offset and the true relative angle based on the installation position and assembly posture annotated by the user, the assembly action constitutes the action label, and the standard process sequence constitutes the process sequence label.

3. The intelligent compliance detection system for assembly operations according to claim 2, characterized in that, The joint training engine employs a convolutional neural network architecture based on multi-task learning, and uses the structured labeled data to train the multi-task compliance judgment model. The joint training engine specifically includes: A shared spatiotemporal feature extraction backbone network is constructed with a MobileNetV3 network integrating a temporal displacement module as its core. This network is used to receive the standard operation video, extract spatial features frame by frame, and perform feature displacement in the channel dimension through the temporal displacement module to fuse temporal context information and output a shared spatiotemporal feature map sequence. The part detection head, connected to the shared spatiotemporal feature extraction backbone network, is used to predefine a set of anchor boxes at each scale of the shared spatiotemporal feature map sequence, predict the bounding box coordinate offset, object presence confidence, and part category probability of each anchor box, and calculate the candidate bounding box prediction value of each target part based on the bounding box coordinate offset and the corresponding anchor box coordinate. The dimension of the part category probability is pre-configured by the joint training engine according to the total number of part categories annotated by the user in the structured annotation data. The part detection head multiplies the object presence confidence of each object with the corresponding part category probability to obtain the candidate comprehensive confidence of each anchor box for each part category, and outputs the target bounding box prediction value and the corresponding part presence confidence prediction value of each target part after non-maximum suppression processing based on the candidate bounding box prediction value and the candidate comprehensive confidence value of each target part. The spatial relationship regression head is connected to the shared spatiotemporal feature extraction backbone network. Guided by the bounding box prediction value output by the part detection head, the spatial relationship regression head extracts the local feature map of each target part from the shared spatiotemporal feature map sequence using the region of interest alignment operation. The local feature map is then fed into the lightweight convolutional network inside the spatial relationship regression head. The lightweight convolutional network performs regression processing on the local feature map to obtain the predicted value of the relative position offset and the relative angle of the target part relative to the main component. The temporal action classification head is connected to the shared spatiotemporal feature extraction backbone network. The temporal action classification head performs global average pooling on the shared spatiotemporal feature map sequence in the time and space dimensions to obtain temporal feature vectors. The temporal feature vectors are then input into the Softmax classification layer set inside the temporal action classification head, and the predicted action matching degree of the standard operation video corresponding to each assembly action in the standard assembly action library is output. The process diagram state verification module takes as input a comprehensive output sequence from the part detection head, the spatial relationship regression head, and the temporal action classification head over a preset time period. The module processes this comprehensive output sequence using a pre-constructed directed acyclic graph (DAG) graph convolutional network, which contains at least one graph convolutional layer. The module temporally aggregates and concatenates the predicted part presence confidence, relative position offset, relative angle, and action matching degree from the comprehensive output sequence to form initial feature vectors for each assembly step. After neighborhood information aggregation by the graph convolutional layer, the module outputs the predicted probability distribution of each assembly step in the standard process sequence corresponding to the comprehensive output sequence. The maximum value in the predicted probability distribution is the sequence progress confidence of the comprehensive output sequence. During training, the joint training engine compares the predicted confidence values ​​of the parts, the predicted values ​​of relative position offset and relative angle, the predicted values ​​of action matching degree, and the predicted probability distribution with the structured labeled data, and updates the network weights of the shared spatiotemporal feature extraction backbone network, the parts detection head, the spatial relationship regression head, the temporal action classification head, and the process diagram state verification module until the preset maximum number of training iterations is reached, thus completing the training process and obtaining the multi-task compliance judgment model. The no-code multi-task model building platform then distributes the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal.

4. The intelligent compliance detection system for assembly operations according to claim 3, characterized in that, The training process in the aforementioned joint training engine is as follows: The joint training engine compares the predicted confidence value of the part with the corresponding part label in the structured annotation data to obtain the part detection loss; The joint training engine compares the predicted relative position offset with the actual relative position offset and the predicted relative angle with the actual relative angle to obtain the position regression loss. The joint training engine compares the predicted action matching degree with the action label to obtain the action classification loss; The joint training engine compares the predicted probability distribution with the process sequence labels to obtain the sequence verification loss; The joint training engine updates the network weights of the shared spatiotemporal feature extraction backbone network, the part detection head, the spatial relationship regression head, the temporal action classification head, and the process diagram state verification module based on the calculated part detection loss, position regression loss, action classification loss, and sequence verification loss, until the preset maximum number of training iterations is reached, thus completing the training process and obtaining a multi-task compliance judgment model. The no-code multi-task model building platform then distributes the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal.

5. The intelligent compliance detection system for assembly operations according to claim 3, characterized in that, The graph convolutional network with the directed acyclic graph structure described above is constructed as follows: Each node in the directed acyclic graph corresponds to an assembly step defined in the standard operating procedure, and the node number corresponds one-to-one with the step number of the assembly step. The edges in the directed acyclic graph are established according to the sequence of steps and logical dependencies defined in the standard operating procedure, including directed edges connecting consecutive steps and directed edges connecting non-consecutive steps with logical dependencies. The process diagram state verification module performs temporal pooling on the part existence confidence prediction value, relative position offset prediction value, relative angle prediction value, and action matching degree prediction value output by the part detection head, the spatial relationship regression head, and the temporal action classification head in the corresponding video segments within a preset time period to obtain the corresponding feature components. The module then concatenates the feature components to form the initial feature vector of each step node. The first layer of the graph convolutional network performs a linear transformation on the initial feature vector and aggregates neighborhood features based on the normalized adjacency matrix, and outputs the first layer node features after processing by the ReLU activation function; the second layer of the graph convolutional network performs a linear transformation on the first layer node features and aggregates neighborhood features based on the normalized adjacency matrix, and outputs the second layer node features after processing by the ReLU activation function. The process graph state verification module performs global average pooling on the second layer node features to obtain a global graph representation vector. The global graph representation vector is then input into the Softmax classification layer inside the process graph state verification module. The module outputs the probability distribution of each assembly step corresponding to the video segment within a preset time period, and takes the maximum value in the probability distribution as the sequence progress confidence of the video segment.

6. The intelligent compliance detection system for assembly operations according to claim 3, characterized in that, The joint training engine also includes a lightweight model synthesis module, which performs the following operations after training is complete: The trained multi-task compliance judgment model is subjected to channel pruning and parameter quantization to obtain an optimized multi-task compliance judgment model. The optimized multi-task compliance judgment model is converted into an inference engine format suitable for the edge-side real-time detection and decision-making terminal, and then serialized and encapsulated to generate a lightweight model file. The lightweight model file is then sent to the edge-side real-time detection and decision-making terminal.

7. The intelligent compliance detection system for assembly operations according to claim 3, characterized in that, The edge-side real-time detection and decision-making terminal includes: A multi-dimensional confidence fusion and decision-making unit is used to receive the real-time output of the multi-task compliance judgment model to the real-time assembly video stream. The real-time output includes: the confidence of part existence output by the part detection head, the position offset output by the spatial relationship regression head based on relative position offset and relative angle, the action matching degree output by the temporal action classification head, and the confidence of sequence progress output by the process diagram status verification module. The multi-dimensional confidence fusion and decision-making unit pre-stores decision logic rules. The decision logic rules preset a first threshold corresponding to the confidence of part existence, a second threshold corresponding to the position offset, a third threshold corresponding to the action matching degree, and a fourth threshold corresponding to the confidence of sequence progress. The multi-dimensional confidence fusion and decision-making unit performs a comprehensive calculation by comparing the confidence of part existence, the position offset, the action matching degree, and the confidence of sequence progress with the corresponding first threshold, second threshold, third threshold, and fourth threshold, respectively, and generates a detailed compliance judgment containing four defect dimensions: part error, position error, action error, and sequence error based on the comprehensive calculation result. The error tracing and alarm generator is connected to the multi-dimensional confidence fusion and decision-making unit. When the detailed compliance decision is unqualified, the error tracing and alarm generator determines at least one major defect dimension among the four defect dimensions that leads to unqualification based on the comprehensive calculation result of the multi-dimensional confidence fusion and decision-making unit, and maps each of the major defect dimensions to a specific error type, triggering the corresponding visual and voice alarm prompts.

8. A detection method based on any one of claims 3 to 7 for an intelligent compliance detection system for assembly operations, characterized in that... Includes the following steps: Step S1: The camera is used to collect real-time assembly video streams from the assembly station and transmits the real-time assembly video streams to the edge-side real-time detection and decision-making terminal. Step S2: Generate a multi-task compliance judgment model through the no-code multi-task model building platform. The generation step includes: receiving structured annotation data generated by the user annotating the standard operation video through the graphical collaborative annotation unit, and automatically training the multi-task compliance judgment model based on the structured annotation data by the joint training engine. Step S3: Deploy the multi-task compliance judgment model to the edge-side real-time detection and decision-making terminal; Step S4: The edge-side real-time detection and decision-making terminal loads and runs the multi-task compliance judgment model to analyze the real-time assembly video stream to identify whether there are part errors, position errors, action errors and sequence errors in the current assembly operation. Step S5: The edge-side real-time detection and decision-making terminal outputs detailed compliance judgments based on the analysis results and triggers corresponding alarm signals.

9. The detection method according to claim 8, characterized in that, The specific steps of step S4 are as follows: Step S4.1: Receive the confidence level, position offset, action matching degree, and sequence progress confidence level of the parts output by the multi-task compliance judgment model for the real-time assembly video stream. Step S4.2: The multi-dimensional confidence fusion and decision-maker performs a comprehensive calculation on the confidence of the existence of the part, the position offset, the action matching degree, and the sequence progress confidence based on the pre-stored decision logic rules, and generates a preliminary qualified or unqualified decision. Step S4.3: When the judgment is unqualified, the error tracing and alarm generator determines at least one major defect dimension that caused the unqualification based on the comprehensive calculation result, and maps each major defect dimension to a specific error type.