An industrial standard operation step generation method and system based on human behavior sequence recognition

By synchronizing and aligning industrial operation videos in time and space, the contact relationship between tools and workpieces is identified, and standard operating procedures are generated. This solves the problem of inaccurate identification of precision assembly actions in existing technologies and improves the completeness and standardization of operating procedures.

CN122637486APending Publication Date: 2026-08-25SHANGHAI MOPAN TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202611131001.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify subtle assembly movements during precision assembly, distinguish between effective and ineffective actions, and generate work steps that are prone to omissions, repetitions, or sequence errors, making them unsuitable for direct use in standard operating procedures.

Method used

By acquiring continuous video of the industrial operation area, time synchronization and spatial alignment are performed. The local image area is determined by the wrist position, the contact relationship between the tool working end and the workpiece operation position is determined, the action state is identified based on the contact duration and relative position changes, and standard operation steps are generated by verifying the continuity of action segments and matching the process sequence.

Benefits of technology

It improves the ability to distinguish subtle assembly actions and the accuracy of determining the boundaries of continuous actions, and enhances the completeness, standardization, and verifiability of the generated work steps, reducing omissions, repetitions, and sequence errors, and ensuring that the work steps accurately reflect the actual process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122637486A_ABST
    Figure CN122637486A_ABST
Patent Text Reader

Abstract

The present application relates to the technical fields of computer vision and human behavior recognition, and discloses a kind of industrial standard operation step generation method and system based on human behavior sequence recognition;The present application, first, the continuous video of operation area is acquired, and video frame is time-synchronized and spatially aligned;Determine local image area with wrist position, identify the contact relationship and position change between hand, tool and workpiece;According to the continuous change of action state, action segment is divided, and invalid action is filtered out and continuous repeated action is combined according to workpiece state and process sequence before and after action;According to the time sequence of reserved action segment, tool category, action type and workpiece position are determined, and industrial standard operation step is generated;Thereby, subtle action misrecognition, step omission, repetition and sequence error are reduced, and the accuracy of standard operation step generation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and human behavior recognition technology, specifically to a method and system for generating industrial standard operating procedures based on human behavior sequence recognition. Background Technology

[0002] As the level of digitalization and intelligence in the manufacturing industry continues to improve, using industrial operation video to identify operation processes and generate standard operating procedures is gradually becoming an important means of standardizing production operations, accumulating process experience, and assisting personnel training. For example, authorized Chinese invention patent application CN111062364B discloses a method and device for monitoring assembly operations based on deep learning, which judges the type and number of times of assembly actions by identifying the motion information of assembly tools and human joints. Another example is authorized Chinese invention patent application CN114155610B, which discloses a key action recognition method for panel assembly based on upper body posture estimation, which determines the type of action in the panel assembly process through target detection, extraction of key points of the upper body, and continuous action recognition. The aforementioned existing technologies can identify, count, or monitor pre-set assembly actions, but they still have certain shortcomings when generating standard operating procedures from long-term, continuous industrial videos that have not been manually edited. In precision assembly processes, the range of motion for actions such as fastening, insertion, removal, and alignment is usually small, and the human torso posture is quite similar between different actions. Relying solely on the overall posture of the human body or large-scale movement characteristics makes it difficult to accurately identify the contact relationship and positional changes between hands, tools, and workpieces. This can easily lead to different operations being identified as the same action, or the same operation being identified as different actions. Furthermore, adjacent assembly processes are usually performed continuously, with little pause between processes, making it difficult to accurately identify the completion process of the previous action and its transition from one action to another. The starting processes of subsequent actions may overlap, making it difficult to accurately determine the action switching position. When the recognition results lack continuity verification, the action category is prone to changing repeatedly between adjacent video frames, thus incorrectly dividing a complete operation into multiple short action segments. As a result, existing technologies struggle to accurately distinguish between valid operations, repetitive operations, trial-and-error operations, and irrelevant actions. The generated work steps are prone to omissions, repetitions, incorrect sequences, or unreasonable step divisions, making them unsuitable for direct use in standard work guidance. Therefore, improving the ability to distinguish subtle assembly actions, accurately determining the switching position between consecutive actions, and reducing the impact of invalid actions on the step generation results have become technical problems that need to be solved. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method and system for generating industrial standard operating procedures based on human behavior sequence recognition, thus solving the problems mentioned in the background section.

[0004] To achieve the above objectives, the present invention provides the following technical solution: A method for generating industrial standard operating procedures based on human behavior sequence recognition includes: S1. Acquire continuous video of the industrial operation area, and perform time synchronization and spatial alignment on each video frame in the continuous video; S2. Determine the local image area based on the wrist position, determine the working end of the tool and the operation position of the workpiece, and judge the contact relationship based on the distance between them, relative movement, contact duration, multi-view height difference and workpiece state changes. S3. Determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and divide the action segment. S4. Determine the workpiece state based on the stable video frames before and after the action segment. Match the action type, tool type, workpiece position, previous process, and state before and after the action with the process sequence record. Retain valid action segments and necessary action segments. When multiple operations correspond to the same target state, only retain the action segment that forms the state, and merge action segments with the same tool, position, and target state and no other state changes between adjacent segments. S5. Select representative video frames from the start, tool application, and end stages of the action for verification, generate industry standard operating procedures in chronological order, and establish a correlation with the corresponding original video range.

[0005] Preferably, S1 includes: The acquisition frame rate of the visual acquisition device is determined based on the shortest effective motion phase, and the exposure time is determined based on the maximum moving speed and the allowable motion blur width. The main view video frame is used as the time reference to match the auxiliary view video frame. The transformation relationship between image coordinates and workbench coordinates is established by using fixed feature points, and the transformation relationship is verified by check points. The data quality status is determined based on the duration of consecutive missing frames, the effective state of the conversion relationship, and the visibility of the target object, thus forming a synchronized video frame group.

[0006] Preferably, S2 includes: The wrist position is verified based on the shoulder and elbow positions and the wrist movement trajectory in adjacent video frames, and a local image area is set along the direction from the elbow to the wrist, adjusted according to the exposed length of the tool. When the local image regions of the two hands overlap or the working end of the tool is occluded, the correspondence between the hand, the tool and the workpiece is determined based on the movement trajectory before and after the overlap or occlusion, the tool holding state and the workpiece position.

[0007] Preferably, determining the working end of the tool and the operating position of the workpiece includes: The length, width, and main direction of the tool outline within the local image area are matched with the tool information record, and the tool that moves synchronously with the hand is identified as the current tool; The working end of a tool is determined based on its shape. When the two ends of a tool are similar in shape, the end furthest from the hand gripping area is determined as the working end of the tool. Read the workpiece position record corresponding to the product model, convert the workpiece operation position to the worktable coordinate system, and update the workpiece operation position according to the position and direction of the workpiece reference point when the workpiece moves or rotates.

[0008] Preferably, S3 includes: The contact relationship and relative position changes are matched according to preset action rules or standard action segments, or the contact relationship and relative position changes are input into the action state recognition model to determine candidate action states; When the action state alternates repeatedly or is in a state to be verified, the action state is determined according to the tool position, workpiece state and movement relationship before and after the corresponding time period, and the action segment is re-divided at the position where the tool category or workpiece position changes.

[0009] Preferably, S4 includes: Read the workpiece position number corresponding to adjacent action segments. When the workpiece position numbers are the same, it is determined that the adjacent action segments act on the same workpiece position. When no workpiece position number is set, the workpiece operation position corresponding to adjacent action segments is converted to the worktable coordinate system, and the distance between the position coordinates is calculated. When the distance does not exceed the position allowable deviation, it is determined that adjacent action segments act on the same workpiece position. The allowable positional deviation is determined based on product assembly tolerances, spatial mapping errors, and image position recognition errors.

[0010] Preferably, S5 includes: When the specific specifications of a tool cannot be determined, record the corresponding tool function category; When the action type is inconsistent with the workpiece state before and after the action, the action boundary and workpiece state are re-verified. If it still cannot be determined, the corresponding action segment is marked as to be verified. Industrial standard operating procedures are generated by combining the time sequence of retained action segments with the process sequence. Record each overlapping action that results in an independent change in the state of the workpiece separately, and merge auxiliary actions that do not result in an independent change in the state of the workpiece into the corresponding main action.

[0011] Preferably, the industrial standard operating procedure is generated according to the time sequence of the retained action segments and in combination with the process sequence, including: Generate step numbers according to the time sequence of each retained action segment; Record the results when the same tool is used repeatedly at different workpiece positions; When there are multiple attempts, the video range corresponding to the action segment that causes the change in the state of the target workpiece will be used as the original video range for that step. Merge duplicate steps that meet the merging criteria; If the tool category, action type, workpiece position, or action time cannot be confirmed, the corresponding record will be marked as pending verification and associated with the range of original videos that can be confirmed.

[0012] On the other hand, the present invention provides an industrial standard operating procedure generation system based on human behavior sequence recognition, comprising: Video processing module: used to acquire continuous video of the industrial operation area, and to perform time synchronization and spatial alignment of each video frame in the continuous video; Local interaction recognition module: used to determine the local image area based on the wrist position, determine the operation position of the tool working end and the workpiece, and determine the contact relationship based on the distance between the two, relative movement, contact duration, multi-view height difference and workpiece state changes; Action segment segmentation module: used to determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and segment the action segments; Motion Segment Processing Module: This module determines the workpiece state based on stable video frames before and after the motion segment. It matches the motion type, tool type, workpiece position, preceding process, and state before and after the motion with the process sequence record, retaining valid motion segments and necessary motion segments. When multiple operations correspond to the same target state, it only retains the motion segment that forms that state and merges motion segments with the same tool, position, and target state and no other state changes between adjacent segments. The work step generation module is used to select representative video frames from the start, tool action, and end stages of an action for verification, generate industry standard work steps in chronological order, and establish a correlation with the corresponding original video range.

[0013] Compared with existing technologies, this invention provides a method and system for generating industrial standard operating procedures based on human behavior sequence recognition, which has the following beneficial effects: 1. This invention, by performing time synchronization and spatial alignment on continuous operation videos, extracts local image regions based on the wrist position, and identifies subtle assembly actions by combining the contact relationship between the tool's working end and the workpiece's operating position, relative position changes, and workpiece state changes; by verifying the continuity of action states in adjacent video frames and determining the action switching time based on the state duration and the proportion of effective video frames, it reduces the impact of short-term occlusion and recognition fluctuations on action segmentation; by comparing the workpiece state before and after the execution of action segments and matching them with the process sequence, it filters out invalid and trial-and-error actions, and merges continuous repetitive actions corresponding to the same workpiece state change, thereby improving the ability to distinguish subtle assembly actions and the accuracy of determining continuous action boundaries, and reducing omissions, repetitions, sequence errors, and unreasonable divisions in industrial standard operation steps.

[0014] 2. This invention, by selecting representative video frames from the start, tool action, and end stages of each retained action segment, verifies the tool type, action type, workpiece position, and workpiece state before and after the action. It establishes a correspondence between industrial standard operating procedures and action segments, representative video frames, and the original video range according to the actual work sequence. Based on whether overlapping actions result in independent workpiece state changes, they are recorded separately or auxiliary actions are incorporated into the corresponding main actions. Actions using the same tool but acting on different workpiece positions generate separate steps. Missing information or workpiece state uncertainties are marked for verification. The invention avoids inferring work content based on missing information, thereby improving the completeness, standardization, and verifiability of industrial standard operating procedures, ensuring that the generated steps accurately reflect the actual work process. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the process for generating industrial standard operating procedures based on human behavior sequence recognition according to the present invention. Figure 2 This is a schematic diagram illustrating the multi-view video frame time synchronization and spatial alignment of the present invention; Figure 3 This is a schematic diagram illustrating the local interaction relationship recognition based on wrist position according to the present invention; Figure 4 This is a schematic diagram illustrating the continuity verification of action states and the segmentation of action segments in this invention; Figure 5 This is a flowchart of the action segment validity judgment and merging process of the present invention; Figure 6 This is a schematic diagram illustrating the generation and data association of industrial standard operating procedures according to the present invention; Figure 7 This is a schematic diagram of the structure of an industrial standard operating procedure generation system based on human behavior sequence recognition according to the present invention. Detailed Implementation

[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0017] Example 1: Figures 1-6 A method for generating industrial standard operating procedures based on human behavior sequence recognition is presented, including: S1. Acquire continuous video of the industrial operation area, and perform time synchronization and spatial alignment on each video frame in the continuous video; S2. Determine the local image area based on the wrist position, determine the working end of the tool and the operation position of the workpiece, and judge the contact relationship based on the distance between them, relative movement, contact duration, multi-view height difference and workpiece state changes. S3. Determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and divide the action segment. S4. Determine the workpiece state based on the stable video frames before and after the action segment. Match the action type, tool type, workpiece position, previous process, and state before and after the action with the process sequence record. Retain valid action segments and necessary action segments. When multiple operations correspond to the same target state, only retain the action segment that forms the state, and merge action segments with the same tool, position, and target state and no other state changes between adjacent segments. S5. Select representative video frames from the start, tool application, and end stages of the action for verification, generate industry standard operating procedures in chronological order, and establish a correlation with the corresponding original video range.

[0018] This application provides a method for generating industrial standard operating procedures based on human behavior sequence recognition. It is applicable to industrial production scenarios such as electronic component assembly, wire harness splicing, fastener installation, precision device mounting, medical device assembly, and other industrial production scenarios where the status changes of operators, tools, and workpieces can be reflected through operation videos. This method can be executed by an industrial computer, industrial controller, edge computing device, image processing server, or other electronic devices with video processing and data storage capabilities. The processing objects of this method include operators, tools used by operators, workpieces being processed or assembled, and vision acquisition devices used to collect data on the operation process. The input data includes at least unedited continuous operation videos, vision acquisition device calibration information, product model, workstation identification, tool information, workpiece operation position information, and process sequence information. The output is an industrial standard operation procedure arranged according to the actual operation sequence and associated with the corresponding original video clips. The frame rate, exposure time, time synchronization error, spatial mapping tolerance, local image area range, contact distance, state holding time, action observation window, and action switching ratio are determined based on the performance of the visual acquisition equipment, the performance of the synchronization device, product drawings, tooling configuration, tool specifications, product assembly tolerances, and qualified operation videos of the current workstation. Among these, the frame rate, exposure time, and time synchronization error are determined based on the performance of the visual acquisition equipment and the synchronization device; the size, position, and orientation parameters are determined based on product drawings, tooling configuration, and tool specifications; and the duration and ratio parameters are calibrated based on qualified operation videos of the same product and the same workstation. In one implementation, parameter calibration uses no fewer than 30 qualified operation videos, and each qualified operation video covers at least 3 operators. For each action to be identified, the start time of the action, the tool contact time, the workpiece state change time, and the end time of the action are recorded. The duration and positional changes of effective actions, non-contact passages, short-term occlusions, and repeated trial and error actions are statistically analyzed. The parameter calibration results are associated with the product model, workstation identifier, tool version, and vision acquisition device version. When the product model, tool specifications, tooling position, or vision acquisition device position changes, the affected parameters are recalibrated. The data quality status of video frames and their processing results is divided into normal, missing, occluded, and pending verification. Pending verification indicates that the current data is insufficient to support a deterministic judgment. The corresponding data retains the original video range and the cause of the occurrence, but is not directly used to generate industrial standard operating procedures with the meaning of action completion. The specific values ​​in the following text are used to illustrate an implementable parameter configuration. In actual application, it can be adjusted according to the product, workstation, and calibration results.

[0019] Specifically, such as Figure 2 As shown: When the electronic device receives the work start signal, it acquires continuous video of the industrial work area, performs time synchronization and spatial alignment on the video frames output by each vision acquisition device, and forms a video frame group with a unified time base and unified spatial coordinates. The start signal for a job can be input by the operator through the workstation terminal, generated by the workstation sensor when it detects a workpiece entering the predetermined work area, or sent by the production control equipment when the current process starts. The electronic equipment establishes a video task record for the current job. The video task record includes at least the following fields: video task identifier, product number, workstation number, job start time, and vision acquisition device number. Different products or different job batches use different video task identifiers to prevent the mixing of video data from different job tasks. In one embodiment, the industrial work area is captured by a main vision acquisition device located above the workbench and an auxiliary vision acquisition device located to the side or diagonally above the workbench; the main vision acquisition device is used to obtain the positional relationship of the hand, tool and workpiece in the plane of the workbench, and the auxiliary vision acquisition device is used to obtain the height change of the tool relative to the workpiece surface and the insertion depth of the part, and to provide supplementary images when the main view is obstructed. When the actions to be identified are mainly picking, moving, or placing, and the changes in the workpiece state mainly occur within the worktable plane, only the main vision acquisition device can be used; when the actions to be identified include insertion, pressing, pulling out, or other operations with significant height changes, both the main vision acquisition device and the auxiliary vision acquisition device can be used simultaneously; when the operator needs to operate with both hands simultaneously, or when the workpiece structure easily obscures the operation position, an auxiliary vision acquisition device can be added on the other side of the worktable. The acquisition frame rate of the visual acquisition device is determined based on the shortest effective phase in the action to be identified; the electronic device determines the duration of phases such as the tool approaching the workpiece, the tool contacting the workpiece, the workpiece state changing, and the tool leaving the workpiece from the qualified operation video, and obtains at least 8 consecutive images within the shortest effective phase. For example, in wire harness splicing or miniature fastener installation stations, the time required for the tool's working end to contact the workpiece and for the workpiece to stabilize can be 150 to 300 milliseconds. To obtain at least 8 frames of images within this time range, the acquisition frame rate can be selected from 60 to 120 frames per second. Under the acquisition condition of 60 frames per second, approximately 9 to 18 frames of images can be obtained within 150 to 300 milliseconds. Under the acquisition condition of 120 frames per second, approximately 18 to 36 frames of images can be obtained, thereby continuously reflecting the process of tool approaching, contacting, being subjected to force, and leaving. When the shortest effective phase in the current workstation lasts longer than 300 milliseconds, the acquisition frame rate can be reduced; when the shortest effective phase lasts less than 150 milliseconds, the acquisition frame rate can be increased, and the exposure time, time synchronization error, and allowable displacement of adjacent video frames can be re-determined based on the adjusted acquisition frame rate. To reduce image blur caused by rapid hand or tool movement, the single-frame exposure time is determined based on the maximum movement speed of the current workstation and the allowable motion blur width. The single-frame exposure time can be calculated by dividing the allowable motion blur width by the maximum movement speed. For example, if the maximum movement speed of the operator's wrist or the working end of the tool is 2 meters per second and the allowable motion blur width is 1 millimeter, the single-frame exposure time can be selected to be no more than 0.5 milliseconds. If the ambient lighting cannot meet this exposure time, workstation lighting can be increased, or an image position error margin can be added based on the actual exposure time. The image resolution and field of view of the vision acquisition device should be able to distinguish the smallest effective operating position in the current workstation; the actual size corresponding to a unit pixel is determined according to the image resolution and the field of view of the worktable; for example, when the smallest effective operating position in the current workstation is 4 mm, the actual size corresponding to a unit pixel can be selected to be no more than 0.5 mm, so that the operating position with a size of 4 mm covers at least 8 pixels in the image, so as to determine the position of holes, slots and tool working ends; Each visual acquisition device outputs video frames that include at least the acquisition time, device number, and frame sequence number. When multiple visual acquisition devices are used, the same hardware synchronization device can send exposure trigger signals to each visual acquisition device so that images are acquired from different perspectives according to the same time reference. The allowable exposure time difference between different visual acquisition devices is determined based on the maximum moving speed of the wrist or the working end of the tool and the allowable dynamic position error. The allowable exposure time difference can be calculated by dividing the allowable dynamic position error by the maximum moving speed. For example, when the maximum moving speed is 2 meters per second and the allowable dynamic position error is 1 millimeter, the exposure time difference between different visual acquisition devices can be selected to be no more than 0.5 milliseconds. When the field equipment does not support hardware exposure synchronization, a unified clock source can be used to record the acquisition time of video frames, and the video frame time can be corrected according to the time deviation of each visual acquisition device relative to the unified clock source. If the dynamic position error caused by the corrected time deviation exceeds the contact judgment tolerance used for the current action, the corresponding video frame will not participate in the multi-view spatial position calculation, but will only be used to supplement the image information of the occluded area. The electronic device uses the video frames output by the main vision acquisition device as the time reference; for each main view video frame, it searches for video frames in the auxiliary view where the acquisition time difference does not exceed the allowable time error, and forms a synchronous video frame group with corresponding video frames; when there are no video frames in the auxiliary view that meet the time error requirements, the data quality status of the auxiliary view at the current moment is set to missing, and video frames with time differences exceeding the allowable range are not used as substitutes. Electronic devices handle video anomalies based on the duration of consecutive missing frames. The duration of consecutive missing frames is determined based on the acquisition time of the valid video frames before and after the missing period. A fixed number of missing frames is not used as a judgment condition to avoid the same number of missing frames corresponding to different actual times at different acquisition frame rates. When the duration of consecutive missing frames does not exceed 30 milliseconds, the electronic device forms transition position data based on the wrist and tool positions in the video frames before and after the missing period. 30 milliseconds is less than one-fifth of the shortest effective contact phase of 150 milliseconds, and usually does not cover the complete contact process. The transition position data is only used to maintain the continuity of the movement trajectory and is not used to determine a new contact state or the moment of action switching. When the duration of consecutive missing frames exceeds 30 milliseconds but is less than 150 milliseconds, the data quality status of the corresponding time period is set to missing, and the positional change within that time period is determined in combination with other normal perspectives; when the duration of missing frames reaches or exceeds 150 milliseconds, the missing time period may cover the complete critical contact process, so the specific contact status is not inferred based on a single movement trajectory before and after the missing frame. If a visual acquisition device fails to output a valid video frame for 300 consecutive milliseconds, its use in determining the spatial position of the current action will be suspended, and processing will continue using only other normal viewpoints. 300 milliseconds is not less than the duration of most short-term contact phases, and continuing to use this viewpoint may result in the loss of critical action information. If a visual acquisition device fails to output a valid video frame or device status information for 1 consecutive second, the visual acquisition device will be identified as a faulty device, and device inspection information will be generated. The 1 second is used to distinguish between short-term data interruption and continuous failure of the visual acquisition device or communication link. After time synchronization is completed, the electronic device spatially aligns video frames from different perspectives according to the calibration parameters of the visual acquisition device. The calibration parameters include at least focal length, principal point position, lens distortion parameters, and the position and orientation of the visual acquisition device relative to the worktable. The electronic device first performs distortion correction on the video frames according to the focal length, principal point position, and lens distortion parameters to reduce the positional deviation of the image edge areas. For scenarios where the main action occurs within the workbench plane, at least four non-collinear fixed feature points can be selected on the workbench or tooling fixture. The pixel position of each fixed feature point in the image and its actual position in the workbench coordinate system can be recorded, and the transformation relationship between the image plane and the workbench plane can be established based on the corresponding positions. For scenarios requiring determination of tool height or component insertion depth, calibration points distributed on at least two different height planes can be set within the workspace to cover the height range that the tool may traverse; electronic devices acquire the image positions of the same tool working end or the same workpiece feature point in at least two viewpoints, and determine the spatial position of the corresponding object in the worktable coordinate system based on the positional relationship between each vision acquisition device; The worktable coordinate system can use a fixed positioning point on the tooling fixture as the coordinate origin; the fixed positioning point remains unchanged in different work batches, so that the position data of different batches adopt a unified reference; for workpieces that can move or rotate during operation, the electronic equipment determines the position and direction of the workpiece reference point and represents the workpiece operation position as the position relative to the workpiece reference point. After spatial calibration is completed, the electronic equipment uses inspection points that have not participated in the establishment of spatial transformation relationships to verify the spatial mapping results; the inspection points are distributed in the main activity areas of the tool and the workpiece; the electronic equipment determines the actual position and the mapped position of the inspection points respectively, and calculates the positional deviation of each inspection point; For example, in a workstation where the smallest effective operating position is 4 mm and the actual size corresponding to a unit pixel is no more than 0.5 mm, the average position deviation of each inspection point can be selected to be no more than 0.8 mm, and the position deviation of any inspection point can be selected to be no more than 1.2 mm; this error range is less than half of the smallest effective operating size, and reserves space for time synchronization error, tool working end size and image position recognition error; When the positional deviation of the inspection point exceeds the allowable range, the electronic equipment checks whether the fixed feature point is obstructed, whether the position of the visual acquisition device has changed, and whether the lens focal length has been adjusted. For planar mapping, when there are fewer than 4 visible valid fixed feature points or when the fixed feature points are collinear, the current spatial transformation relationship is stopped. For spatial mapping, when the valid calibration points of different height planes cannot cover the current tool activity range, the tool height and part insertion depth are stopped from being output. When the spatial mapping of a single auxiliary visual acquisition device fails, while the main visual acquisition device is normal, it can be downgraded to the main viewpoint processing to continue determining the planar position and motion status, but the height or depth determined by multiple viewpoints will not be output; when neither the main visual acquisition device nor the auxiliary visual acquisition device can complete the spatial mapping, the automatic generation of job content will be paused, and the original video and abnormal information will be retained. The electronic device determines the data quality status of the synchronized video frame group based on the object visibility, image sharpness, time synchronization results, and spatial mapping results; the visibility ratio can be calculated according to the ratio of the visible outline length of the object in the current image to the complete outline length in the corresponding clear reference image; the clear reference image is a video frame obtained under the same tool, the same workpiece position, and the same viewpoint, without obstruction and without obvious motion blur; For example, when the visibility ratio of both the tool working end and the workpiece operation position is not less than 75%, the edge width of the tool working end does not exceed 1.5 times the corresponding edge width in the clear reference image, and time synchronization and spatial mapping are effective, the data quality status can be set to normal. The 75% visibility ratio and 1.5 times the edge width can be determined by comparing the position recognition errors in clear operation videos, occluded videos, and motion-blurred videos, so that the position recognition error under normal conditions does not exceed the contact judgment tolerance used subsequently. When the visibility of the tool's working end or the workpiece's operating position is less than 75% but not less than 25%, the data quality status is set to pending verification; when the visibility of the tool's working end is less than 25%, the data quality status is set to occlusion; the 25% visibility ratio can be determined based on the position recognition error of the tool's working end in the occluded sample. When the position recognition error reaches or exceeds the contact judgment tolerance, the tool contact position is no longer determined solely based on the current viewing angle. The synchronized video frame group includes at least fields such as unified acquisition time, main view video frame, auxiliary view video frame, spatial coordinate transformation parameters, and data quality status. The electronic device writes the synchronized video frame group into the video frame buffer according to the acquisition time, which is used for subsequent wrist positioning, local image region determination, and recognition of the relative relationship between the hand, tool, and workpiece.

[0020] Specifically, such as Figure 3 As shown: When a continuous group of synchronized video frames has been formed in the video frame buffer, and the data quality status of at least one viewpoint is normal, the electronic device determines the operator's wrist position, establishes a local image region based on the wrist position, and identifies the contact relationship and relative position changes between the hand, tool and workpiece within the local image region. The electronic device first determines the area where the operator is located in the main view video frame and identifies the positions of the shoulder, elbow and wrist; the shoulder and elbow positions are used to determine the direction of arm extension, and the wrist position is used to determine the position of the local image area; When there are multiple operators in the screen, the electronic equipment determines the operator corresponding to the current work station based on the pre-set work station work area; when multiple operators enter the work station work area at the same time, they can be distinguished based on personnel identification, the order in which they enter the work area and the distance between them and the current workpiece, so as to avoid identifying the hands of personnel in adjacent work stations as the current work object. Wrist position can be determined using a human key position recognition model, upper limb contour analysis, or visible wrist markers. When using a human key position recognition model, the model input is a video frame containing the operator's upper limb region, and the model output is the position of the left and right wrists in the corresponding video frame. Training samples can include upper limb images of different operators, different arm postures, different work equipment positions, and different lighting conditions, and the wrist positions in the samples are manually marked. The model output is used after checking the positional continuity of adjacent video frames. When using upper limb contour analysis, the electronic device first determines the upper limb contour, then searches for areas where the contour width gradually decreases along the direction from the elbow to the hand, and uses the corresponding positions as candidate wrist positions; when the candidate wrist position is continuous with the wrist movement direction in the preceding and following video frames, it is determined as the wrist position; The allowable wrist displacement between adjacent video frames is determined based on the maximum wrist movement speed of the current workstation, the video frame acquisition cycle, the wrist position recognition error, and the coordinate transformation margin. For example, when the acquisition frame rate is 100 frames per second, the time interval between adjacent video frames is 10 milliseconds; when the maximum wrist movement speed is 2 meters per second, the theoretical movement distance between adjacent video frames does not exceed 20 millimeters; when the wrist position recognition error is 2 millimeters and the coordinate transformation and viewing angle change margin is 3 millimeters, the allowable wrist displacement between adjacent video frames can be selected as 25 millimeters. When the wrist position in the current video frame is displaced by more than 25 mm relative to the previous video frame, the electronic device makes a judgment based on the next video frame. If the wrist position in the next video frame returns to the original direction of movement, the data quality status of the current video frame is set to pending verification. If the wrist position changes along the new direction of movement for more than 3 consecutive frames, it is determined that the operator has moved rapidly, and the wrist movement trajectory is updated. The 3 consecutive frames are used to eliminate single-frame positioning errors while retaining the actual rapid movement process. After determining the wrist position, the electronic device continues to extend towards the palm along the direction from the elbow to the wrist, and uses the extended position as the center of the local image area; the extension distance is determined according to the distance between the wrist and the main gripping area of ​​the palm; for example, when the distance between the wrist and the main gripping area of ​​the palm is 3 cm to 5 cm, the extension distance can be selected as 3 cm to 5 cm, so that the center of the local image area is close to the palm and the tool gripping position; The extent of the local image region is determined based on the maximum distance between the region center and the tool working end, the wrist position recognition error, and the spatial mapping error; the radius of the local image region is not less than the sum of the maximum distance between the region center and the tool working end and the positioning margin, and is not greater than half the width of the effective working area of ​​the current station, so as to reduce the interference of adjacent tools, other workpieces, and the workbench background on the recognition results; For example, in workstations using miniature tweezers, small screwdrivers, and splicing tools, the maximum distance from the center of the local image region to the fingertip is approximately 8 cm, and the tool extends outward from the fingertip by no more than 5 cm. The positioning margin corresponding to the wrist position recognition error and spatial mapping error is 2 cm. Therefore, the radius of the local image region can be selected as 15 cm. For tools with shorter exposed lengths, the radius of the local image region can be selected as 10 cm. When the tool type is not yet determined, the initial region can be determined based on the longest tool allowed to be used at the current workstation. After the tool is identified, the local image region is then adjusted according to the exposed length of the tool. When operating with both hands, the electronic device establishes local image regions based on the left and right wrists respectively, and records the tool and workpiece operation positions corresponding to each hand; when two local image regions overlap, they continue to retain their respective wrist markers, tool holding status, and movement trajectory before entering the overlapping area. The electronic device identifies the hand contour, tool body, tool working end, and candidate operation position of the workpiece within a local image area, and calls the tool information record to determine the tool category and tool working end; the tool information record includes at least the following fields: tool category, overall length range, body width range, working end shape, position of the working end relative to the tool body, and allowed action type; The electronic device records tool information based on the length, width, main direction, and hand grip position of the tool outline; when the tool body has a clear working end structure, the working end of the tool is determined according to the shape of the working end; when the two ends of the tool are similar in shape, the end furthest from the hand grip area is determined as the working end of the tool; when there are multiple tools in the image, the tool that moves synchronously with the hand and is located in a local image area is determined as the current tool. The candidate operation position of the workpiece is determined based on the workpiece position record; the workpiece position record includes at least the following fields: product model, workpiece reference point, operation position number, position of operation position relative to workpiece reference point, allowed tool category, and allowed workpiece state changes. For workpieces with fixed positions, the electronic equipment converts each operating position to the worktable coordinate system before the operation begins; for workpieces that can move or rotate during the operation, the electronic equipment redetermines the position and orientation of the workpiece reference point based on the synchronous video frame group, and updates each operating position accordingly. When the working end of the tool is partially obstructed, the electronic device determines the candidate position of the working end of the tool based on the main direction of the tool body, the tool length in the tool information record, and the hand holding position; the data quality status of the candidate position is set to pending verification, and it is only used for contact status judgment when the position can be verified from other perspectives or in subsequent video frames; The electronic device determines the contact state based on the spatial distance, contour relationship, relative movement direction, and local changes of the workpiece between the tool working end and the workpiece operation position; the contact distance threshold is jointly determined by the maximum deviation of spatial mapping, the maximum displacement error generated by time synchronization, the radius of the tool working end, and the image position recognition margin; after converting all errors to the same worktable coordinate system, they are superimposed according to the direction that may cause the contact position deviation. For example, when the maximum deviation of spatial mapping is 1.2 mm, the maximum displacement error caused by time synchronization is 1 mm, the radius of the tool working end is 2 mm, and the image position recognition margin is 0.8 mm, the contact distance threshold can be selected as 5 mm; this value covers the comprehensive position deviation caused by spatial mapping, time synchronization, tool size, and image recognition. When the distance between the working end of the tool and the operation position of the workpiece is not greater than the contact distance threshold, and this state continues until the contact holding time is reached, the corresponding state is determined as a candidate contact state; the contact holding time is determined by comparing the duration of the effective contact action and the duration of the non-contact passing action. When calibrating the contact holding time, the duration during which the working end of the tool is within the contact distance range during a qualified contact action, and the overlap duration when the tool passes over the workpiece without making contact, are counted separately. The contact holding time is set between the longest non-contact overlap time and the shortest effective contact time to exclude short-term position overlap caused by rapid passage. For example, when the longest non-contact overlap time is 180 milliseconds and the shortest effective contact time is 350 milliseconds, the contact holding time can be selected as 300 milliseconds; for tapping, instantaneous triggering or other operations with short duration, the contact holding time can be re-determined based on qualified samples of the corresponding action and confirmed in conjunction with the workpiece status response; When the working end of the tool and the operating position of the workpiece overlap in the main view, but the height difference between the two is still greater than the contact distance threshold as shown in the auxiliary view, it is not considered as contact; when both the planar distance and the height distance meet the contact conditions, a candidate contact state can be determined; when only a single view is used, it should also be confirmed by combining the changes in tool movement speed and the local state response of the workpiece. The change in tool movement speed can be determined by comparing the average movement speed of the tool before and after the working end enters the contact distance range; for example, if the average movement speed after entering the contact distance range is not higher than 50% of the average movement speed during the approach phase, tool deceleration can be used as an auxiliary condition for contact judgment; the 50% ratio can be set by comparing the speed changes in effective contact actions and non-contact passing actions. The electronic device records the changes in distance between the tool's working end and the workpiece's operating position, the changes in tool direction, the synchronous movement relationship between the hand and the tool, and the changes in the local state of the workpiece according to the acquisition time. For fastening operations, a candidate state for fastening operations can be determined when the distance between the working end of the tool and the fastening position is no greater than the contact distance threshold, the position change of the working end of the tool does not exceed the contact distance threshold, and the tool body undergoes continuous directional changes around its axis; the continuous directional changes can be judged based on the edge of the tool body, surface texture, or the rotation state output by the tool. For insertion operations, when the angle between the moving direction of the tool or the part to be inserted and the interface axis is not greater than the allowable skew angle specified in the product assembly specification, the amount of movement along the interface axis reaches the insertion depth specified in the product drawing, and the offset perpendicular to the interface axis does not exceed the product assembly tolerance, the candidate state of the insertion operation can be determined; for example, when the product assembly specification specifies that the allowable skew angle does not exceed 10 degrees, the direction judgment threshold can be selected as 10 degrees. For pick-up and place operations, a candidate state for the pick-up and place operation can be determined when the hand first approaches the target material, then the material moves synchronously with the hand, and finally the material stops at the target position while the hand leaves. Synchronous movement means that in consecutive video frames, the difference in the movement direction of the hand and the material does not exceed the directional tolerance, and the change in the distance between them does not exceed the positional tolerance. The directional tolerance and positional tolerance are determined based on the material size, the acquisition frame rate, and the spatial mapping error. When the local image regions corresponding to the left and right hands overlap, the electronic device retains the tool holding state, movement direction and corresponding workpiece position of both hands before entering the overlapping area; if one hand keeps the workpiece stable before the overlap, and the other hand continues to move towards the workpiece operation position, the supporting hand is maintained during the overlap, and the movement change consistent with the tool working end is associated with the hand performing the operation. After the overlap is removed, the electronic device re-determines the position of the left and right wrists and the tool holding status, and compares the tool type, movement trajectory and workpiece position after the overlap is removed with the record before the overlap. When the left and right hand identifiers cannot correspond continuously, the association between the hand and the tool is re-established according to the tool type, movement trajectory and position of action, and the left and right position of the hand in the image is not used as the only judgment condition. When the operation position of the hand, tool, or workpiece is obstructed, the electronic device calculates the visible ratio of the tool's working end; the visible ratio is the ratio between the current visible outline length of the tool's working end and the expected complete outline length; the expected complete outline length is scaled based on the tool information record and the tool's position in the worktable coordinate system to reduce image size changes caused by the tool approaching or moving away from the vision acquisition device; When the visible proportion of the tool's working end is less than 25%, the data quality status of the corresponding local image area is set to occlusion; the 25% proportion is determined based on the relationship between the position recognition error of the tool's working end in the occluded sample and the contact distance threshold; when the visible proportion is less than 25% and the position recognition error reaches or exceeds the contact distance threshold, the contact position is no longer determined solely based on the current viewpoint. When the duration of occlusion does not exceed 500 milliseconds, the electronic device maintains the last normal state before occlusion and prioritizes using other perspectives to determine the positional changes of the tool and workpiece; 500 milliseconds can be set according to the duration of the short alignment or contact phase so that short-term occlusion does not cause the action state to be interrupted immediately. When the occlusion duration exceeds 500 milliseconds but does not exceed 2 seconds, the local image area is expanded to 1.5 to 2 times the original range, so that the expanded area covers the wrist, forearm and elbow, but does not exceed the effective working area of ​​the current workstation, and the candidate state is determined according to the arm movement direction and other perspectives; 2 seconds can be set according to the main operation cycle of a single insertion or fastening action at the current workstation. When the tool working end and workpiece operation position cannot be observed from any viewpoint for more than 2 consecutive seconds, the data quality status of the corresponding time period is set to pending verification, and the specific contact status is uncertain. After completing the above processing, the electronic device generates local interactive data for each group of synchronized video frames. The local interactive data includes at least the following fields: wrist position, local image area, hand identifier, tool type, tool working end position, workpiece operation position, contact state, relative distance, movement direction, and data quality status, which are used for subsequent determination of action status and division of action boundaries.

[0021] Specifically, such as Figure 4 As shown: The electronic device determines the action state corresponding to each video frame based on local interactive data, performs continuity verification on adjacent action states, determines the action switching time, and divides the action segments according to the adjacent action switching time. The action status is set according to the operations allowed to be performed at the current workstation, which may include pick up, move, align, contact, insert, pull out, tighten, loosen, place and release; the action status name adopts the common operation name in the corresponding production process and is consistent with the final standard operation record; In one implementation, the electronic device pre-establishes a correspondence between local interactive data and action states; when the distance between the hand and the tool continuously decreases, and the tool subsequently moves synchronously with the hand, the corresponding video frame is determined as the pickup state; when the working end of the tool moves towards the workpiece operation position, and the distance between them continues to decrease but the contact condition is not yet met, it is determined as the alignment state; when the working end of the tool and the workpiece operation position meet the contact condition and continue to move along the operation direction specified by the product, it is determined as the insertion state or the pressing state; when the working end of the tool remains in contact with the fastening position and the tool body continues to rotate, it is determined as the fastening state; when the working end of the tool leaves the workpiece and the workpiece remains in the state after the operation, it is determined as the release state. In another implementation, standard action segments confirmed by process engineers can be saved in advance. The standard action segments include at least fields such as action type, tool category, workpiece position, contact state change sequence, movement direction change sequence, workpiece state change, and allowable duration. When the tool category, workpiece position, and change sequence of each stage in the current local interaction data match the standard action segment, and the duration is within the allowable range of the corresponding action, the corresponding action is determined as a candidate action state. In another implementation, a motion state recognition model can be used to output candidate motion states. The input to the motion state recognition model is local interactive data arranged according to the acquisition time. The input fields include at least the distance between the tool working end and the workpiece operation position, the tool movement direction, the tool direction change, the distance between the hand and the tool, the contact state, and the data quality state. Each input sample corresponds to a state observation window, and the training samples are manually labeled with the motion state corresponding to the center video frame of the window. The action state recognition model outputs the recognition value corresponding to each candidate action state; the electronic device takes the action state with the highest recognition value that meets the output requirements determined in the model verification stage as the candidate action state; the model output requirements can be determined according to the distribution of recognition values ​​of correct and incorrect recognition results in the verification samples, so that the action states that meet the output requirements have stable category discrimination ability; the candidate action states obtained by rule matching, standard action segment matching or action state recognition model are all determined as valid action states after subsequent continuous verification; Electronic devices set a status observation window in the action state sequence; the length of the status observation window is determined according to the duration of independent actions in the qualified work sample of the same work station; for example, when the main operation phases of actions such as picking, placing, aligning, inserting and fastening are concentrated in 1.2 seconds to 1.5 seconds, the length of the status observation window can be selected as 1.5 seconds; The state observation window can include a historical interval of 0.5 seconds before the evaluation time and a subsequent observation interval of 1 second after the evaluation time; the historical interval is used to determine the action state that has been stable before the evaluation time, and the subsequent observation interval is used to determine whether the new action state continues to appear; 0.5 seconds and 1 second are set according to the duration of the stable state before the action starts and the duration of the main change after the action starts, so that the observation window can cover the original action state and the new action state at the same time. The electronic device calculates the duration from the first video frame in which the new action state appears, and counts the number of valid video frames corresponding to the new action state in the subsequent observation interval; the action state ratio is the ratio of the number of valid video frames corresponding to the new action state to the total number of valid video frames in the subsequent observation interval; video frames with missing data quality status are not included in the total number of valid video frames. When the duration of the effective video frames in the subsequent observation interval is less than 70% of the total length of the observation interval, automatic action switching is not performed, and the data quality status of the corresponding interval is set to pending verification. For example, when the length of the subsequent observation interval is 1 second, a 70% effective data ratio can retain at least 0.7 seconds of effective video, covering the shortest stable duration of actions such as picking, inserting, and clamping. The 70% ratio can be determined based on the stability of the action boundary recognition results under different frame loss ratios. When a new action state that differs from the historical action state accounts for 80% of the subsequent observation interval, and the duration of the new action state reaches the minimum duration of the corresponding action, the first frame in which the new action state appears consecutively is determined as the candidate action switching moment. The 80% action state ratio can be determined by comparing qualified action boundary samples with abnormal samples that include occlusion, reflection, and short-term error recognition. For example, within the candidate ratio range of 70% to 90%, 80% can allow a small number of video frames to have erroneous states, while ensuring that the new action state maintains a significant dominant ratio in the subsequent observation interval. During the duration of a new action state, other states with a continuous duration of no more than 200 milliseconds are allowed, and the video frames corresponding to other states are not counted in the new action state frame count. When the continuous duration of other states exceeds 200 milliseconds, the continuous start time of the new action state is re-determined. 200 milliseconds is lower than the minimum stable duration of each valid action and can be used to exclude state interruptions caused by single-frame misjudgment, short-term reflection, or short-term occlusion. The minimum duration of different actions is set according to the corresponding qualified work samples; the electronic equipment can statistically analyze the duration distribution of each type of effective action and determine the corresponding value based on the time difference between the effective action and the short-term error identification. For example, the shortest stable duration for picking and releasing actions can be selected as 300 milliseconds, the shortest stable duration for inserting and pressing actions can be selected as 400 milliseconds, and the shortest duration for fastening actions from stable tool contact to completion of the main rotation can be selected as 600 milliseconds; the above values ​​can be obtained from the lower stable range of the corresponding action duration in qualified work samples, with a margin of no more than 10% for recognition error. When historical action states and new action states alternate repeatedly within the subsequent observation interval, and neither of the new action states meets the requirements for action state ratio and minimum duration, the electronic device maintains the historical action state and does not form a new action segment at the alternation position; if the new action state subsequently meets the corresponding conditions, the video frame in which it first consecutively meets the minimum duration is taken as the candidate action switching moment. When the data quality status of local interactive data is pending verification, the electronic device makes a judgment based on the action status, tool position and workpiece status changes before and after the period. If the alignment status is before the obstruction and the tool has entered the workpiece after the obstruction, and the workpiece status changes in a way consistent with the insertion operation, then the obstruction period is determined as a candidate insertion status and confirmed during the workpiece status verification process. If the tool type, workpiece position, and action status remain consistent before and after the period to be verified, the original action status is maintained; if changes that do not conform to the continuous movement relationship occur before and after the period to be verified, the state to be verified is retained, and no definite action switching time is formed. When the action state repeatedly changes between two or more states within 3 consecutive seconds and fails to form a dominant state that meets the action switching conditions, the electronic device expands its observation range to check whether the state change is caused by repeated alignment, repeated trial and error, continuous occlusion, spatial mapping failure, or visual acquisition device jitter. 3 seconds corresponds to two consecutive state observation windows of 1.5 seconds each, indicating that the state instability has exceeded the observation range of a normal action boundary. The electronic device can also call up standard action segments corresponding to the same product and the same workstation for comparison. If a dominant action consistent with the current process stage can be determined, the action state of the corresponding video frame is corrected. If it still cannot be determined, the corresponding time period is set to be verified and is not divided into multiple determined action segments. Electronic devices divide action segments based on the switching time of adjacent candidate actions; an action segment includes at least the following fields: start time, end time, action type, tool category, workpiece position, main contact state, original video range, and data quality status; When the tool category changes within the same action segment, the electronic device checks whether the tool has been changed; when the workpiece position changes, it checks whether the object being operated on has been moved to another operation position; if the tool or workpiece position is confirmed to have changed, the action switching moment is added at the corresponding changed position, and the action segment is redefined; if it cannot be confirmed, the corresponding position is set to be verified. Each action segment is arranged according to its acquisition time to form a candidate action segment sequence; each candidate action segment retains the start time, end time, and video frame index of the corresponding original video, so that subsequent action filtering, action merging, and standard operation record generation can all be traced back to the original video.

[0022] Specifically, such as Figure 5 As shown: The electronic device judges the validity of candidate action segments based on the workpiece state before and after each candidate action segment is executed and the process sequence information corresponding to the current product, and merges the action segments that meet the merging conditions. Process sequence information can come from product process documents, product drawings, bill of materials, tooling configuration files, or standard operating procedure records confirmed by process engineers. Each process sequence record should include at least the following fields: operation number, workpiece status before the operation, allowed operation type, allowed tool category, workpiece operation position, workpiece status after the operation, preceding operation number, whether repeated execution is allowed, and whether it is a necessary operation. Necessary actions refer to operations that must be performed in the production process but do not necessarily result in a permanent visible change in the state of the workpiece, such as inspection, measurement, cleaning, and temporary clamping. By marking necessary actions in the process sequence record, it is possible to avoid erroneously deleting corresponding action segments based solely on changes in the appearance of the workpiece. The workpiece status is set according to the workpiece type and assembly requirements. For connector assembly, the workpiece status can include whether the terminal is in the interface, the terminal insertion depth, the latch closure status, and the wire harness direction. For fastener assembly, the workpiece status can include whether the screw is in the fastening hole, the screw head height, and the fastening position. For electronic component mounting, the workpiece status can include whether the component is present at the target position, the component orientation, and the deviation of the component center from the target position. For each candidate action segment, the electronic device determines the workpiece state before the action and the workpiece state after the action. The workpiece state before the action is determined first from the last video frame before the start of the action segment that clearly shows the workpiece operation position and the workpiece state remains stable. The workpiece state after the action is determined first from the first video frame after the tool and hand leave the operation position that clearly shows the workpiece state and the workpiece state remains stable. When a suitable video frame cannot be obtained directly, a video frame before the action segment can be searched within the range of 100 to 300 milliseconds before the action segment begins, and a video frame after the action segment can be searched within the range of 200 to 500 milliseconds after the action segment ends. The search range before the action segment is determined based on the stabilization time before the tool approaches the workpiece, and the search range after the action segment is determined based on the time required for the hand to leave the operating position and for the workpiece to complete its rebound or reset. The search range after the action segment does not cross the start time of the next action segment to avoid incorrectly classifying state changes caused by subsequent operations into the current action. The electronic device compares the component status, contour position, orientation, depth, and state stability of the workpiece before and after the action to determine whether the candidate action segment causes a change in the workpiece state that conforms to the production process. For component placement operations, if there is no component to be installed at the target position before the action, and a contour matching the component appears at the target position after the action, and the contour remains stable in position and orientation for 500 consecutive milliseconds, it is determined that a component placement state change has occurred; the 500 milliseconds can be set according to the springback time after the workpiece is placed, the time the hand leaves the workpiece, and the state stabilization time in the qualified operation video. For mating operations, when the terminal before the action is outside the interface and the terminal after the action enters the interface and reaches the insertion depth specified in the product drawing or assembly specification, a change in mating status is determined. The insertion depth and allowable deviation are directly adopted from the design parameters of the corresponding product. For example, if the effective insertion depth of a connector is 4 mm and the allowable deviation is ±0.5 mm, the effective insertion range can be selected from 3.5 mm to 4.5 mm. For fastening operations, a change in fastening status is determined when the screw head is higher than the target mounting surface before the operation, and the screw head height enters the range specified in the assembly specification after the operation and remains stable after the tool is removed; the screw head height range is determined based on the screw specifications, mounting hole structure, target mounting surface position, and product assembly tolerances. For planar placement or mounting operations, the electronic equipment determines the workpiece state based on the distance between the component center and the target position and the component orientation; the allowable deviation for position judgment is jointly determined by the product assembly allowable deviation, spatial mapping error and image position recognition error. After converting each error to the same coordinate system, they are superimposed according to the direction that may cause position deviation. For example, if the allowable deviation for the installation position specified in the product drawings is 1 mm, the spatial mapping error is 0.5 mm, and the image position recognition margin is 0.5 mm, the allowable deviation for position judgment can be selected as 2 mm; if the distance between the center of the component and the target position does not exceed 2 mm, and the component orientation meets the product assembly requirements, it can be determined that the component is in the target installation position. When on-site tools or fixtures can output status signals, these status signals can be used as auxiliary information for judging the status of the workpiece. Status signals may include torque arrival signals output by fastening tools, displacement arrival signals output by pressing equipment, and component arrival signals output by tooling fixtures, etc. Electronic devices match the generation time of status signals with the time when the workpiece status changes in the video according to a unified time reference. When the time difference between the two does not exceed the allowable synchronization error between the device signal and the video acquisition, the status signal can be used as an auxiliary judgment condition for the validity of the action. When the status signal and the video recognition result are inconsistent, the time synchronization status, the tool's position, and the workpiece's occlusion are checked. The completion of the action is not determined solely based on the status signal. The electronic device matches the workpiece state before the action, action type, tool category, workpiece operation position, and workpiece state after the action with the process sequence record; when the above content is consistent with the same process sequence record, and the corresponding preceding process has been completed, it is determined that the candidate action segment conforms to the process sequence. Based on the workpiece state changes and process sequence matching results, candidate action segments can be divided into valid action segments, necessary action segments, invalid action segments, and action segments to be verified. An effective motion segment refers to a motion segment in which a change in the workpiece state occurs before and after the motion, in accordance with the process sequence record, and in which the motion type, tool category, and workpiece operation position all match the change in state. Necessary action segments refer to action segments that do not result in a permanently visible change in the workpiece state, but are marked as mandatory in the process sequence record, and whose tool type, workpiece operation position, duration, and relationship with preceding and following processes all meet the requirements. Invalid motion segments refer to motion segments that do not result in a workpiece state change that meets the requirements and are not necessary actions. Examples include taking the wrong material and putting it back, moving tools in the air, hand movements unrelated to the current task, and repeatedly probing near the workpiece without completing the assembly. The electronic equipment deletes invalid motion segments from the sequence used to generate standard work records and records the reason for deletion. Action segments to be verified refer to action segments whose validity cannot be determined due to occlusion, missing video, abnormal position recognition, or unstable workpiece status; the corresponding original video and the cause of the abnormality are retained for the action segments to be verified, and standard operation records with the meaning of action completion are not directly generated; For repeated alignment or multiple insertion attempts, the electronic device checks the workpiece state after each action segment in sequence; if the previous actions have not made the workpiece reach the target state, but the last action makes the workpiece reach the target state, the previous actions are determined to be invalid trial and error actions, and only the last valid action that leads to the change of the target state is retained. When multiple consecutive action segments together form a single workpiece state change, the corresponding action segments can be merged. The conditions for merging action segments include at least using the same tool, acting on the same workpiece position, corresponding to the same target state change, and there are no other workpiece state changes between adjacent action segments. When two action segments correspond to the same workpiece position number, it can be determined that they act on the same workpiece position; when the workpiece position does not have a fixed number, the coordinates of the two positions in the worktable coordinate system can be compared; when the distance between the coordinates does not exceed the position judgment allowable deviation of the current operation position, it can be determined that they act on the same workpiece position. For example, when an operator tightens the same screw by making multiple small rotations in succession, the multiple rotational motion segments can be combined into one tightening motion segment; when the motion type and tool category are the same, but they are applied to different workpiece positions, they are not combined. For actions that are repeatedly performed at the same workpiece position, if the first action has already brought the workpiece to the target state and subsequent actions have not resulted in further state changes, determine whether the subsequent actions are necessary re-inspection or re-tightening based on the process sequence record; if the process sequence record does not specify re-inspection or re-tightening, the subsequent actions are identified as repetitive actions and deleted. When it is detected that the state of the subsequent workpiece has appeared, but the corresponding state of the preceding workpiece or the preceding action segment does not exist, the electronic device does not supplement the action that did not actually appear in the video, but instead checks back the original video and candidate action segments before the current action. The review time range is determined based on the maximum normal interval between adjacent processes of the current product; for example, if the maximum normal interval between adjacent processes of a certain workstation is 90 seconds, the review time can be selected as 120 seconds, of which the additional 30 seconds is used to cover the time fluctuations caused by operators picking up materials, changing tools, or short-term pauses. During the first review, the visibility ratio requirement of the tool's working end can be appropriately reduced to find actions missed due to partial occlusion. For example, if the visibility ratio used for normal recognition is 25%, the visibility ratio used for review can be selected as 20%. The 20% ratio is determined based on the minimum visibility level that can still keep the tool position error within the contact judgment tolerance in the occluded sample. If no preceding action is found during the first review, a second review can be performed, and the minimum duration of the corresponding action can be appropriately reduced. The minimum duration can be adjusted by 10% of the original calibration value, and the adjusted value should not be lower than 70% of the original calibration value. The 10% adjustment range is determined based on the difference in action speed between different operators, and the lower limit of 70% is used to prevent short-term noise, reflection, or instantaneous position overlap from being identified as valid actions. The automatic review of the same action segment shall not exceed 2 times in total; the first review shall review the local interactive data and action boundaries, and the second review shall review the workpiece status and process sequence. Historical video review shall be counted in the number of automatic reviews; if the preceding action that matches the process sequence record cannot be found after 2 automatic reviews, the corresponding action segment shall be set as pending verification and the automatic generation of subsequent work records that depend on the preceding status shall be stopped. When a workpiece is affected by glare, obstruction, or elastic rebound, and its stable state cannot be determined based on a single video frame, the electronic device compares the workpiece state in consecutive video frames. If the workpiece state remains consistent for 500 milliseconds after the tool leaves the workpiece, the workpiece state can be determined to be stable. The 500 milliseconds are determined based on the workpiece's elastic rebound time, the time required for the hand to leave the operating position, and the state maintenance time in the qualified operation video. If the stable state still cannot be determined before the next action begins, the corresponding action segment is set to be verified. After filtering and merging, the electronic device forms a sequence of retained action segments. The retained action segments inherit fields such as start time, end time, action type, tool category, workpiece position, original video range, and data quality status from the candidate action segments, and add fields such as workpiece status before action, workpiece status after action, and corresponding process number, for use in the generation of subsequent standard operation records.

[0023] Specifically, such as Figure 6 As shown: When a sequence of retained action segments is formed and there are no process sequence anomalies that cannot be automatically handled, the electronic equipment generates industry standard operating procedures according to the execution order of each retained action segment; For each retained motion segment, the electronic device selects video frames that can represent the start of the motion, the tool action, and the end of the motion. The video frames at the start of the motion are used to represent the workpiece state before the motion is executed and the tool used by the operator. The video frames at the tool action stage are used to represent the tool action position and the motion type. The video frames at the end of the motion stage are used to represent the workpiece state after the motion is executed. The tool action phase can be determined based on the moment when the distance between the tool working end and the workpiece operation position is the smallest, the moment when the workpiece state begins to change, or the period when the tool remains in the contact position for the longest time. When there are multiple video frames that meet the conditions, the video frame with clear image, low degree of occlusion, and ability to display both the tool working end and the workpiece operation position should be selected first. The tool category is determined by the results of the preceding identification process; when the tool category is in the verification state, the electronic device verifies it by combining video frames before and after the start and end of the action segment; when the tool function can be determined but the specific specifications cannot be determined, a general tool category such as screw fastening tool, clamping tool, plugging tool or measuring tool is used, and the specific model is not inferred based on blurry or occluded images. The action type is the result after action continuity verification and workpiece status verification. When the action type is inconsistent with the change in workpiece status before and after the action, the electronic equipment will re-check the action boundary or workpiece status based on the remaining automatic verification count. When the same action segment still cannot form a consistent result after accumulating 2 automatic verifications, the action segment will be set to be verified and no operation steps with action completion meaning will be generated. The workpiece position is determined during the preceding identification process and can be recorded according to the workpiece part name, process position number, or position relative to the workpiece reference point. When there are many similar operation positions, the position number and position coordinates can be recorded simultaneously to distinguish different operation objects. The electronic device generates step content based on tool type, action type, workpiece object, workpiece position, and workpiece state after the action; for example, it can generate step content for using a plug-in tool to insert a wire harness terminal into a connector interface, or using a screw fastening tool to tighten a screw at the first mounting hole of a circuit board; the action completion status is based on the workpiece state changes that can be confirmed in the video, without supplementing information that does not appear or cannot be confirmed in the original video. Each industry standard operating procedure is associated with a corresponding reserved action segment; the industry standard operating procedure includes at least the following fields: step number, action start time, action end time, tool category, action type, workpiece object, workpiece position, workpiece state before action, workpiece state after action, representative video frame, original video range, and data quality status; the above records can be converted into standard operating procedure descriptions in text form or into operation cards containing representative images, and different output formats can call the same set of step data; The step number is generated according to the execution order of the retained action segments; when two actions overlap in time, the electronic device determines the recording method based on the relationship between each action and the change in the workpiece state. For example, if one hand holds the workpiece while the other hand performs the tightening, and the workpiece holding does not result in an independent change in workpiece state, the workpiece holding can be recorded as an auxiliary action of the tightening operation, and no separate step is generated; if the left and right hands perform operations on different workpiece positions respectively, and each operation results in an independently verifiable change in workpiece state, corresponding steps are generated respectively, and the action time and original video range of each action are recorded. When multiple consecutive motion clips have been merged into a single retained motion clip, a step is generated based on the start and end times of the merged clip. When the same tool is used to operate on the same type of part, but applied to different workpiece positions, corresponding steps are generated separately. In multiple alignments or trial operations, if only the last operation results in a change in the state of the target workpiece, the video range corresponding to the step is determined from the last valid alignment, and previous invalid attempts are not written into the industry standard operation steps. After generating preliminary industrial standard operating procedures, check whether the electronic equipment inspection tool category is clear, whether the action type is consistent with the workpiece state change, whether the workpiece position can be uniquely determined, whether adjacent steps are repeated, and whether the step sequence conforms to the process sequence record. When the tool category, action type, workpiece position, and workpiece state change are all the same in two adjacent steps, the electronic device checks whether the corresponding action segment meets the merging conditions; if the merging conditions are met, the two steps are merged; if the merging conditions are not met, they are retained and the reason for not merging is recorded. When the required prerequisite workpiece state for a certain step does not exist, or the workpiece position and tool type cannot be determined, the corresponding record is set to be verified, and the step content is not automatically filled in based on the missing information; when the image quality is low but the workpiece state change can be confirmed, the action type, workpiece position and tool function type can be recorded; when the workpiece state change cannot be confirmed, no industrial standard operation steps with completion meaning are generated. When the original video is missing and the start or end time of an action cannot be fully determined, record the video time range that can be confirmed, the missing period, and the data quality status to avoid using the estimated time to replace the actual acquisition time. The correction results generated by manual review can be used to update tool information, workpiece position, or action correspondence; for the same product, the same workstation, and the same type of identification problem, when the number of manually confirmed records reaches a set number, candidate update content is generated; the set number is determined based on the frequency of similar errors in historical operation videos and the sample size required for parameter verification; For example, to reduce the impact of occasional corrections on recognition parameters, the number of manual verification records can be no less than 10; candidate update content should be verified using qualified operation videos that did not participate in the formation of the candidate update content; the verification video should cover the corresponding tool, workpiece position and action type, and be independent of the video used to form the candidate update content; When the candidate update content can reduce the corresponding identification error and does not reduce the identification consistency of the original normal action, a new parameter version is generated after confirmation by the process personnel; if the verification fails, the original parameter version continues to be used; the parameter version record shall include at least the fields of update time, applicable product, applicable workstation, modified content, verification result and previous version identifier, so as to restore to the previous version when an anomaly occurs after the update; The final industrial standard operating procedure is arranged according to the actual operation sequence and retains the correlation between the procedure and the action segments, representative video frames and the original video range, so as to facilitate the viewing, review and traceability of the procedure content.

[0024] Example 2: Figure 7 This invention discloses an industrial standard operating procedure generation system based on human behavior sequence recognition, comprising: Video processing module: used to acquire continuous video of the industrial operation area, and to perform time synchronization and spatial alignment of each video frame in the continuous video; Local interaction recognition module: used to determine the local image area based on the wrist position, determine the operation position of the tool working end and the workpiece, and determine the contact relationship based on the distance between the two, relative movement, contact duration, multi-view height difference and workpiece state changes; Action segment segmentation module: used to determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and segment the action segments; Motion Segment Processing Module: This module determines the workpiece state based on stable video frames before and after the motion segment. It matches the motion type, tool type, workpiece position, preceding process, and state before and after the motion with the process sequence record, retaining valid motion segments and necessary motion segments. When multiple operations correspond to the same target state, it only retains the motion segment that forms that state and merges motion segments with the same tool, position, and target state and no other state changes between adjacent segments. The work step generation module is used to select representative video frames from the start, tool action, and end stages of an action for verification, generate industry standard work steps in chronological order, and establish a correlation with the corresponding original video range.

[0025] Example 3: Based on Examples 1 and 2, the specific application process of an industrial standard operating procedure generation method and system based on human behavior sequence recognition is further explained: This embodiment takes the continuous assembly operation of wire harness terminal plugging and controller housing screw tightening as an example to illustrate the industrial standard operation step generation method based on human behavior sequence recognition and the coordinated operation process of the system. The system is set up at the controller assembly station. The video processing module, local interaction recognition module, motion segmentation module, motion segment processing module, and work step generation module can be deployed on the same industrial computer, or they can be deployed separately on the workstation edge computing device and the production server. Each module associates the video frames, tool positions, motion states, motion segments, workpiece states, and industrial standard work steps generated by the current job through video task identifiers, so that the data of different processing stages correspond to the same product and the same work batch. The controller assembly station is equipped with a main vision acquisition device and an auxiliary vision acquisition device. The main vision acquisition device is located above the workbench and is used to acquire the positional changes of the operator's hands, wire harness terminals, connector interfaces, screws, and screw fastening tools in the plane of the workbench. The auxiliary vision acquisition device is located on the side of the workbench and is used to acquire the height changes of the wire harness terminals, the insertion depth, and the contact state between the screw fastening tools and the screws. The current product's process sequence record specifies that the wire harness terminals should first be moved to the connector interface and plugged in, and then the screws at the first and second mounting holes of the controller housing should be tightened in sequence; the target insertion depth of the connector terminals is 4 mm, the allowable deviation is ±0.5 mm, and the effective insertion range can be selected from 3.5 mm to 4.5 mm; the operation of the operator using the other hand to fix the controller housing during the tightening process is marked as an auxiliary operation and is not generated as a separate industry standard operating procedure. After the workstation sensor detects that the controller to be installed has entered the tooling fixture, it sends a job start signal to the video processing module. The video processing module establishes a video task identifier and associates the video task identifier with the product number, workstation number, job start time, and vision acquisition device number. The main visual acquisition device and the auxiliary visual acquisition device output continuous video at a frame rate of 100 frames per second, with a single frame exposure time of 0.5 milliseconds. The time interval between adjacent video frames corresponding to 100 frames per second is 10 milliseconds, which can acquire 30 frames of images during a continuous contact process of 300 milliseconds. The single frame exposure time is calculated based on the maximum moving speed of the tool's working end of 2 meters per second and the allowable motion blur width of 1 millimeter, in order to reduce image trailing caused by rapid movement. The video processing module reads the acquisition time, device number, and frame number of each video frame. Using the video frame output by the main visual acquisition device as the time reference, it searches for video frames with an acquisition time difference of no more than 0.5 milliseconds in the auxiliary viewpoint and groups the corresponding video frames into a synchronous video frame group. The 0.5 milliseconds is determined based on the maximum moving speed of 2 meters per second and the allowable dynamic position error of 1 millimeter, so that the position deviation caused by different acquisition times from different viewpoints does not exceed 1 millimeter. After time synchronization is completed, the video processing module transforms the image positions from different perspectives to a unified workbench coordinate system based on the pre-completed spatial calibration. The workbench coordinate system uses the fixed positioning point on the tooling fixture as the coordinate origin, and the connector interface, the first mounting hole, and the second mounting hole are respectively set with position numbers and position coordinates. The video processing module writes the synchronized video frame group into the video frame buffer according to the acquisition time and transmits the video task identifier to the local interaction recognition module. The local interaction recognition module reads the synchronous video frame group, determines the position of the operator's shoulder, elbow, and wrist, and establishes local image regions based on the left and right wrists respectively. The center of the local image region can be determined by extending 4 cm from the wrist position along the direction from the elbow to the wrist towards the palm, and the region radius can be selected as 15 cm. The 15 cm is obtained by adding the maximum distance of 8 cm from the region center to the fingertip, the exposed tool length of 5 cm, and the wrist recognition and spatial mapping margin of 2 cm. When the operator's right hand approaches the wire harness terminal, the local interaction recognition module detects that the distance between the right hand and the wire harness terminal is continuously decreasing. Subsequently, the wire harness terminal and the right hand move synchronously, thus determining that the right hand has picked up the wire harness terminal. The generated local interaction data includes at least the following fields: wrist position, hand identification, wire harness terminal position, connector interface position, direction of movement, relative distance, and data quality status. When the wire harness terminal moves toward the connector interface, the local interaction recognition module calculates the spatial distance between the front end of the wire harness terminal and the connector interface; the contact distance threshold can be selected as 5 mm, which is determined by superimposing the maximum spatial mapping deviation of 1.2 mm, the maximum displacement error generated by time synchronization of 1 mm, the front end size of the wire harness terminal of 2 mm, and the image position recognition margin of 0.8 mm; When the operator first attempts to plug in the connector, the front end of the wire harness terminal enters within 5 mm of the connector interface and remains there for approximately 220 milliseconds before exiting the interface. Since this contact holding time is less than 300 milliseconds and the auxiliary viewing angle does not detect an effective insertion depth, the local interaction recognition module identifies this process as a candidate contact and does not confirm that the terminal has been successfully plugged in. After the operator adjusts the direction of the wiring harness, a second insertion is performed. The front end of the wiring harness terminal enters the contact distance range and remains there for approximately 620 milliseconds. The auxiliary viewing angle detects that the angle between the terminal's movement direction and the interface axis is 6 degrees, which is less than the product's specified allowable 10-degree skew angle. The terminal's movement along the interface axis is 4.1 millimeters, which is within the effective insertion range of 3.5 millimeters to 4.5 millimeters. Based on this, the local interaction recognition module generates contact status, movement direction, and insertion position data, and transmits the continuous local interaction data to the action segmentation module. The action segmentation module determines the action states such as picking, moving, aligning, contacting, inserting, and releasing according to the acquisition time; the state observation window can be selected as 1.5 seconds, of which 0.5 seconds before the evaluation time is used to determine the original action state, and 1 second after the evaluation time is used to determine whether the new action state persists; When the proportion of valid video frames in the moving state within the subsequent observation interval reaches 80% and the duration exceeds 300 milliseconds, the action segmentation module determines the first frame in which the moving state appears consecutively as the action switching moment; when the proportion of valid video frames in the inserted state within the subsequent observation interval reaches 80% and the duration exceeds 400 milliseconds, the corresponding video frame is determined as the start position of the inserted action. The contact state in the first insertion attempt lasted only 220 milliseconds, which did not reach the minimum stable duration of the insertion action, so it was classified as an alignment and short-time contact action segment; the contact state and axial movement in the second insertion were continuous, so it was classified as an insertion action segment; after the terminal entered the interface, the operator's right hand was removed from the terminal, and the terminal position remained unchanged, thus determining the release action segment; During the insertion process, the operator's hand briefly obstructs the connector interface for about 180 milliseconds. Since the obstruction time does not exceed the allowable interruption time of 200 milliseconds, the action segmentation module combines the terminal position before and after the obstruction with the insertion depth detected by the auxiliary view to maintain the continuity of the insertion action and not form a new action boundary at the obstruction position. After the wiring harness is connected, the operator uses their left hand to hold the controller housing in place and their right hand to pick up the screw tightening tool and tighten the screws at the first and second mounting holes in sequence. The local interactive recognition module establishes local image areas corresponding to the left and right hands respectively. The controller housing in the left hand area remains stable, while the screw tightening tool in the right hand area moves towards the mounting hole, then contacts the screw and rotates continuously. The motion segmentation module determines the tool picking state based on the synchronous movement of the right hand and the screw fastening tool, the alignment state based on the continuous decrease in the distance between the tool working end and the mounting hole, and the fastening state based on the tool working end maintaining contact with the screw and continuously rotating around the tool axis. The operator uses three consecutive small rotations to tighten the screw at the first mounting hole. The pause time between adjacent rotations is about 150 milliseconds. Since 150 milliseconds is less than the allowable interruption time of 200 milliseconds, the action segmentation module keeps the tightened state continuous and does not form independent action boundaries between rotations. The action segment processing module receives candidate action segments and judges their validity based on the workpiece status and process sequence records before and after the action execution. After the first insertion attempt, the wire harness terminal is still outside the connector interface and has not reached the target insertion depth. Therefore, the action is determined to be an invalid trial action. Before the second insertion, the wire harness terminal was located outside the connector interface; after the second insertion, the terminal entered the interface by 4.1 mm and remained in a stable position for 500 milliseconds after the right hand was removed; this state change conformed to the process sequence record, so the action segment corresponding to the second insertion was determined as a valid action segment; The left hand fixing the controller housing does not form an independent permanent workpiece state change, but the process sequence record marks this operation as an auxiliary operation in the fastening process. Therefore, this action is associated with the corresponding fastening action and does not generate an industrial standard operating procedure separately. The three consecutive rotations at the first mounting hole use the same screw fastening tool, act on the same mounting hole, and together change the same screw from an unfastened state to a fastened state. Since there are no other workpiece state changes between adjacent rotations, the motion segment processing module merges the corresponding segments into a single fastening motion segment. Although the fastening operation at the second mounting hole uses the same tool and the same action type, the corresponding workpiece position number is different. Therefore, the fastening operation at the first mounting hole is retained separately. When the screw tightening tool can output a torque signal, the motion segment processing module matches the torque signal with the moment when the tool stops its main rotation in the video, according to a unified time reference. When the time difference does not exceed the allowable synchronization error and the tool's position is consistent with the corresponding mounting hole, the torque signal is used as auxiliary information to determine that tightening is complete. When the torque signal is inconsistent with the tool's position in the video, tightening is not determined solely based on this signal. After filtering and merging, the action segment processing module forms a sequence of retained action segments. The retained action segments correspond in sequence to the picking and moving of the wire harness terminal, the effective insertion of the wire harness terminal, the tightening of the screw in the first mounting hole, and the tightening of the screw in the second mounting hole. The trial action that failed to complete the insertion on the first attempt is deleted, and the action of fixing the controller housing with the left hand is retained as auxiliary information for the tightening operation. The work step generation module generates industry standard work steps according to the time sequence of the retained action segments, and selects video frames that can reflect the start of the action, the tool's action, and the end of the action for each retained action segment; For the wire harness insertion action, the video frame at the beginning of the action shows the wire harness terminal outside the connector interface, the video frame during the tool action stage shows the wire harness terminal aligned with the interface and entering the interface, and the video frame at the end of the action shows the wire harness terminal remaining inside the interface; the operation step generation module generates the industry standard operation steps for inserting the wire harness terminal into the controller connector interface and reaching the specified insertion position accordingly. For the first mounting hole tightening action, the video frame at the beginning of the action shows that the screw has not been tightened yet, the video frame during the tool action stage shows that the screw tightening tool is applied to the screw at the first mounting hole, and the video frame at the end of the action shows that the tool is removed and the screw remains tightened. The operation step generation module generates the industry standard operation steps for tightening the screw at the first mounting hole of the controller housing with the screw tightening tool, and records the left hand fixing the controller housing as an auxiliary operation. The tightening operation at the second mounting hole is formed as an independent industrial standard operating procedure in the same way; since the first and second mounting holes have different position numbers, the two tightening operations are recorded separately. The final standard operating procedure can be as follows: move the wire harness terminal to the controller connector interface; insert the wire harness terminal into the connector interface and reach the specified insertion position; fix the controller housing and tighten the screws at the first mounting hole; fix the controller housing and tighten the screws at the second mounting hole; Each standard operating procedure includes at least the following fields: step number, action start time, action end time, tool category, action type, workpiece object, workpiece position, workpiece status before action, workpiece status after action, representative video frame, original video range, and data quality status, and is associated with the corresponding time range in the original video. When a change in workpiece status cannot be confirmed due to occlusion or missing video, the local interaction recognition module sets the corresponding data to be verified, the action segment division module retains the corresponding video range, the action segment processing module suspends the confirmation of subsequent actions that depend on the workpiece status, and the work step generation module does not generate industrial standard work steps that have the meaning of action completion. After the workpiece status is manually verified, the action segment screening and industrial standard work step generation can be carried out again. Through the above coordinated operation, the video processing module provides video frame groups using a unified time and space reference, the local interaction recognition module forms the contact relationship and positional changes between the hand, tool and workpiece, the motion segment segmentation module forms candidate motion segments with time boundaries, the motion segment processing module forms a sequence of retained motion segments based on the workpiece state and process sequence, and the work step generation module generates industry standard work steps that can be traced back to the original video based on the retained motion segments.

[0026] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0027] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for generating industrial standard operating procedures based on human behavior sequence recognition, characterized in that, include: S1. Acquire continuous video of the industrial operation area, and perform time synchronization and spatial alignment on each video frame in the continuous video; S2. Determine the local image area based on the wrist position, determine the working end of the tool and the operation position of the workpiece, and judge the contact relationship based on the distance between them, relative movement, contact duration, multi-view height difference and workpiece state changes. S3. Determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and divide the action segment. S4. Determine the workpiece state based on the stable video frames before and after the action segment. Match the action type, tool type, workpiece position, previous process, and state before and after the action with the process sequence record. Retain valid action segments and necessary action segments. When multiple operations correspond to the same target state, only retain the action segment that forms the state, and merge action segments with the same tool, position, and target state and no other state changes between adjacent segments. S5. Select representative video frames from the start, tool application, and end stages of the action for verification, generate industry standard operating procedures in chronological order, and establish a correlation with the corresponding original video range.

2. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, S1 includes: The acquisition frame rate of the visual acquisition device is determined based on the shortest effective motion phase, and the exposure time is determined based on the maximum moving speed and the allowable motion blur width. The main view video frame is used as the time reference to match the auxiliary view video frame. The transformation relationship between image coordinates and workbench coordinates is established by using fixed feature points, and the transformation relationship is verified by check points. The data quality status is determined based on the duration of consecutive missing frames, the effective state of the conversion relationship, and the visibility of the target object, thus forming a synchronized video frame group.

3. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, S2 includes: The wrist position is verified based on the shoulder and elbow positions and the wrist movement trajectory in adjacent video frames, and a local image area is set along the direction from the elbow to the wrist, adjusted according to the exposed length of the tool. When the local image regions of the two hands overlap or the working end of the tool is occluded, the correspondence between the hand, the tool and the workpiece is determined based on the movement trajectory before and after the overlap or occlusion, the tool holding state and the workpiece position.

4. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, Determine the working end of the tool and the operating position of the workpiece, including: The length, width, and main direction of the tool outline within the local image area are matched with the tool information record, and the tool that moves synchronously with the hand is identified as the current tool; The working end of a tool is determined based on its shape. When the two ends of a tool are similar in shape, the end furthest from the hand gripping area is determined as the working end of the tool. Read the workpiece position record corresponding to the product model, convert the workpiece operation position to the worktable coordinate system, and update the workpiece operation position according to the position and direction of the workpiece reference point when the workpiece moves or rotates.

5. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, S3 includes: The contact relationship and relative position changes are matched according to preset action rules or standard action segments, or the contact relationship and relative position changes are input into the action state recognition model to determine candidate action states; When the action state alternates repeatedly or is in a state to be verified, the action state is determined according to the tool position, workpiece state and movement relationship before and after the corresponding time period, and the action segment is re-divided at the position where the tool category or workpiece position changes.

6. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, S4 includes: Read the workpiece position number corresponding to adjacent action segments. When the workpiece position numbers are the same, it is determined that the adjacent action segments act on the same workpiece position. When no workpiece position number is set, the workpiece operation position corresponding to adjacent action segments is converted to the worktable coordinate system, and the distance between the position coordinates is calculated. When the distance does not exceed the position allowable deviation, it is determined that adjacent action segments act on the same workpiece position. The allowable positional deviation is determined based on product assembly tolerances, spatial mapping errors, and image position recognition errors.

7. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 1, characterized in that, S5 includes: When the specific specifications of a tool cannot be determined, record the corresponding tool function category; When the action type is inconsistent with the workpiece state before and after the action, the action boundary and workpiece state are re-verified. If it still cannot be determined, the corresponding action segment is marked as to be verified. Industrial standard operating procedures are generated by combining the time sequence of retained action segments with the process sequence. Record each overlapping action that results in an independent change in the state of the workpiece separately, and merge auxiliary actions that do not result in an independent change in the state of the workpiece into the corresponding main action.

8. The method for generating industrial standard operating procedures based on human behavior sequence recognition according to claim 7, characterized in that, Industrial standard operating procedures are generated by combining the time sequence of retained action segments with the process sequence, including: Generate step numbers according to the time sequence of each retained action segment; Record the results when the same tool is used repeatedly at different workpiece positions; When there are multiple attempts, the video range corresponding to the action segment that causes the change in the state of the target workpiece will be used as the original video range for that step. Merge duplicate steps that meet the merging criteria; If the tool category, action type, workpiece position, or action time cannot be confirmed, the corresponding record will be marked as pending verification and associated with the range of original videos that can be confirmed.

9. An industrial standard operating procedure generation system based on human behavior sequence recognition, used to implement the industrial standard operating procedure generation method based on human behavior sequence recognition as described in any one of claims 1-8, characterized in that, include: Video processing module: used to acquire continuous video of the industrial operation area, and to perform time synchronization and spatial alignment of each video frame in the continuous video; Local interaction recognition module: used to determine the local image area based on the wrist position, determine the operation position of the tool working end and the workpiece, and determine the contact relationship based on the distance between the two, relative movement, contact duration, multi-view height difference and workpiece state changes; Action segment segmentation module: used to determine the candidate action state based on the contact relationship, and verify its continuity based on the duration, effective video frame ratio and allowed interruption duration, determine the action switching time and segment the action segments; Motion Segment Processing Module: This module determines the workpiece state based on stable video frames before and after the motion segment. It matches the motion type, tool type, workpiece position, preceding process, and state before and after the motion with the process sequence record, retaining valid motion segments and necessary motion segments. When multiple operations correspond to the same target state, only the motion segments that form that state are retained, and motion segments with the same tool, position, and target state and no other state changes between adjacent segments are merged. The work step generation module is used to select representative video frames from the start, tool action, and end stages of an action for verification, generate industry standard work steps in chronological order, and establish a correlation with the corresponding original video range.

Citation Information

Patent Citations

  • A method and apparatus for monitoring assembly operations based on deep learning

    CN111062364B

  • A panel assembly key action recognition method based on upper body posture estimation

    CN114155610B