A method for analyzing the efficiency of action combinations based on computer vision recognition
Video data is collected through the camera and preprocessed and hierarchical progressive strategies to identify workers' actions, generate orderly arrangement and combinations, and compare them using computer vision models, solving the problem of low productivity caused by unstandard workers' operations, and realizing standardized analysis and efficiency improvement of workers' actions.
Patent Information
- Application Number
- CN202111388679.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2041-11-22
AI Technical Summary
In industrial production, workers' operational movements are not standard or unskilled, resulting in low production efficiency, and it is difficult for the prior art to effectively standardize and improve the worker's operation efficiency on the assembly line.
The camera is used to collect video data, and then decompose it into multiple continuous images of translational change frames. The hierarchical progression strategy is used to generate an orderly arrangement and combination of basic action units, and it is recognized and compared through a computer vision model to standardize action behavior.
The standardized analysis of workers' actions has been realized, the efficiency of assembly line has been improved, and the reasons for inefficiency have been discovered and management and planning have been carried out.
Smart Images

Figure CN114140875B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically, it relates to a method for analyzing the efficiency of action combinations based on computer vision recognition. Background Art
[0002] Computer vision is a science that studies how to enable machines to "see". More specifically, it refers to machine vision that uses cameras and computers to replace human eyes to identify, track, and measure targets, etc., and further performs graphic processing to make the computer-processed images more suitable for human eye observation or transmission to instrument detection.
[0003] In industrial production, most assembly lines require manual operation. The biggest drawback of manual operation is the low production efficiency. There are many reasons for the low production efficiency. One of the main reasons is that the actions of workers are not standard or not proficient during the production process. Therefore, standardizing the operation actions of workers has become an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for analyzing the efficiency of action combinations based on computer vision recognition. This efficiency analysis method uses a camera to evaluate the action efficiency of workers on the factory assembly line, so as to quickly discover the reasons for the low work efficiency of the assembly line and assist in improving production efficiency.
[0005] The above technical purpose of the present invention is achieved through the following technical solutions: A method for analyzing the efficiency of action combinations based on computer vision recognition, including the following steps:
[0006] 1) Use a camera to collect real-time video data and preprocess it;
[0007] 2) Decompose the continuous actions in the preprocessed video data in step 1) into multiple consecutive images of translational change frames;
[0008] 3) Decompose the actions related to the time context into image classifications, and generate an ordered permutation and combination A i of the basic action units C i = {C1, C2, C3... C i , C i+1 ,...} through a hierarchical progressive strategy for sequence combination of image categories;
[0009] 4) Train the ordered permutation and combination A i = {C1, C2, C3... C i , C i+1 ,...} using a model to obtain a computer vision model;
[0010] 5) Identify the basic units, interpret and process the recognition result sequence of the basic units, compare it with the standard database, and obtain the matching rate between the sequence and the database, that is, achieve the analysis of the specified action behavior.
[0011] Further, in step 3), the generation of the action basic unit C i The way of the ordered permutation and combination is to regard the process of completing a work task at each job site as a complete action process A i , and the action process A i is decomposed into several action steps S that make up the action process i , and then the action step S i is decomposed into multiple action basic units C i , and multiple basic units C i form an ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…}.
[0012] Further, the specific steps of the hierarchical progressive strategy in step 3) are as follows:
[0013] A) Detect whether the camera is turned on. If the camera is on, go to step B). If the camera is off, stop the subsequent operations;
[0014] B) Continuously read the camera video data and identify the basic units. If the identification is successful, input each frame into the target detection model and output the basic unit recognition result list, and add it to the recognition sequence set detectSet, and then go to step C); if no result is identified, mark it as an "error" basic unit;
[0015] C) Compare the size of the sequence set detectSet with the window length L. If the size of detectSet is greater than the window L, call the window smoothing strategy, and at this time, smoothing and increasing the recognition sequence set are alternated; when the smoothing times reach the threshold or when the currently recognized basic unit ends, convert the smoothed result resultSet into a mapping relationship baseSet, and perform a time comparison to determine whether it times out. If it times out, add a mark and go to step D); if the size of detectSet is less than the window L, directly go to step D);
[0016] D) Perform the shortest match between baseSet and the steps in the database. If it matches correctly with a certain step information in the database, it is defined as a "valid" step, discard baseSet, and add the matched step to the stepSet matching sequence, and go to step E); if the match fails, mark the first basic unit of baseSet as "invalid";
[0017] E) Then, match the "valid" step sequence stepSet with the action process information in the database. If it matches the step sequence in a certain action information in the database correctly, it is a "valid" action process; if stepSet does not match the step sequence in any action information in the database, it is an "invalid" action process.
[0018] Further, the specific method of the preprocessing described in step 1) includes the following steps:
[0019] (1) Obtain a video sequence set V (V ∈ Rnx×ny) from the native video data set R, and ensure that each video sequence contains only valid information and the features of a single action;
[0020] (2) Cut the action content in each video sequence according to the start frame and end frame of the hand movement change of the actor, and sequentially obtain video sequence subsets: Vclip (Vclip ∈ Rkx×ky, 0 < kx < nx, 0 < ky < ny), and ensure that the action of each sub-sample in each sampled Vclip set can be within the kx×ky video frame box;
[0021] (3) Assume that the frame video sequence at the i-th time Ti of the video sequence V is V(:,:,Ti). After sequential sampling operations, the resulting sequence Vresult(:,:,Ti) of the action after Δt is obtained, Vresult(:,:,Ti) = |V(:,:,Ti) - V(:,:,Ti+Δt)|; where the time sequence T = (T1, T2, T3,......Tt), and Δt represents the calculation distance difference between two frames.
[0022] In summary, the present invention has the following beneficial effects:
[0023] 1. By identifying the basic units and comparing them with the standard database, the matching rate between the sequence and the database is obtained, so as to achieve the effect of standardizing action behavior analysis;
[0024] 2. By adopting a hierarchical comparison strategy, the error rate and efficiency of the actor's work can be clearly observed, which is helpful for pipeline operation management and planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 is the network topology diagram of the factory deployment in the embodiment of the present invention;
[0026] Figure 2 is the schematic diagram of the action decomposition strategy in the embodiment of the present invention;
[0027] Figure 3 is the schematic diagram of the working process of the hierarchical comparison strategy in the embodiment of the present invention. Detailed implementation mode
[0028] The following will further elaborate on the present invention in conjunction with the attached Figures 1-3 drawings.
[0029] Example: A method for analyzing the efficiency of action combinations based on computer vision recognition, as Figures 1 to 3 shown, includes the following steps:
[0030] 1) Use a camera to collect real-time video data and preprocess it;
[0031] 2) Decompose the continuous actions in the preprocessed video data in step 1) into multiple consecutive images of translational change frames;
[0032] 3) Decompose the actions related to the time context into image classifications, and generate an ordered permutation and combination A of the action basic units C i by means of a hierarchical progressive strategy for sequence combination of the image categories, where A i ={C1, C2, C3…C i , C i+1 ,…};
[0033] 4) Train the ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…} using a model to obtain a computer vision model; this computer vision model is the training model:
[0034] a. According to the result data label set in step 1), put it into the ResNet model;
[0035] b. The output result is the recognition model identifier corresponding to the Ai set sequence.
[0036] 5) Identify the basic units, interpret and process the recognition result sequence of the basic units, and compare it with the standard database to obtain the matching rate between the sequence and the database, that is, achieve the analysis of the standard action behavior.
[0037] In this embodiment, due to the need for video detection, two processing methods of video frame-by-frame detection and frame skipping detection are proposed. Frame-by-frame detection means that each frame of the video stream is recognized by the model, and frame skipping detection is to perform detection every few frames. The detailed process is as follows:
[0038] Input: Video recognition result sequence C, window length L
[0039] Process: Function WindowSmooth(C, L)
[0040]
[0041]
[0042] Output: The smoothed recognition result set resultSet.
[0043] The ordered permutation and combination of generating the action basic unit C in step 3) i The way of ordered permutation and combination is to regard the process of completing a work task at each job site as a complete action process A i , and regard the action process A i Decompose it into several action steps S that make up the action process i , and then decompose the action steps S i Into multiple action basic units C i , multiple basic units C i Form an ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…}.
[0044] In this embodiment, the sequence combination has an order relationship. For example, Figure 2 In, the label on the connection line represents the combination order. C1→C3 forms S1, but C3→C1 is another basic unit combination S different from S1 i , which represents the sequence of occurrence of the basic units. The same applies to the step sequence combination, corresponding to the order of the hand movement events in the job site. The basic unit and the step sequence combination are based on the real job site and follow the coding rules of prefix coding, stipulating that no sequence is the prefix of other coding sequences at the same level. On the one hand, it prevents multiple matching results from appearing in the subsequent sequence matching work. On the other hand, it ensures that the sequence is the shortest combination at the current layer, which is beneficial to recombination and upward extraction. In this article, the process of upward extraction of basic units is defaulted to use the form of the shortest match.
[0045] The specific steps of the hierarchical progressive strategy in step 3) are as follows:
[0046] A) Detect whether the camera is turned on. If the camera is turned on, go to step B). If the camera is turned off, stop the subsequent operations;
[0047] B) Continuously read the camera video data and identify the basic unit. If the identification is successful, after inputting each frame into the target detection model, output the basic unit recognition result list and add it to the recognition sequence set detectSet, and then go to step C); if no result is recognized, mark it as an "error" basic unit;
[0048] C) Compare the size of the sequence set detectSet with the window length L. If the size of detectSet is greater than window L, call the window smoothing strategy, and at this time, smoothing and adding the recognition sequence set are alternated; if the number of smoothing times reaches the threshold or when the current recognized basic unit ends, convert the smoothed result resultSet into a mapping relationship baseSet, and perform time comparison to determine whether it times out. If it times out, add a mark and proceed to step D); if the size of detectSet is less than window L, directly proceed to step D);
[0049] D) Perform the shortest match between baseSet and the steps in the database. If it matches correctly with a certain step information in the database, it is defined as a "valid" step. Discard baseSet, and at the same time add the matched step to the stepSet matching sequence, and proceed to step E); if the match fails, mark the first basic unit of baseSet as "invalid";
[0050] E) Match the "valid" step sequence stepSet with the action process information in the database again. If it matches correctly with the step sequence in a certain action information in the database, it is a "valid" action process; if stepSet does not match correctly with the step sequence in any action information in the database, it is an "invalid" action process.
[0051] In this embodiment, during the hierarchical comparison stage, it is determined whether the basic unit times out, and the counted basic unit sequence proceeds to step S i Compare with the action process A i Compare. Upward extraction is shown as Figure 2 The main idea of its hierarchical comparison is shown in the following pseudocode:
[0052] Input: Real-time camera stream camera
[0053] Process: Function Compare(camera)
[0054]
[0055]
[0056] The specific method of the preprocessing in step 1) includes the following steps:
[0057] (1) Obtain the video sequence set V (V ∈ Rnx×ny) from the native video data set R, and ensure that each video sequence only contains valid information and the features of a single action;
[0058] (2) Cut the action content in each video sequence according to the starting frame and ending frame of the hand movement change of the actor, and sequentially obtain video sequence subsets: Vclip (Vclip ∈ Rkx×ky, 0 < kx < nx, 0 < ky < ny), and ensure that the action of each subsample in each Vclip set sampled can be within the video frame box of kx×ky;
[0059] (3) Assume that the frame video sequence at the i-th time Ti of the video sequence V is V(:,:,Ti). After sequential sampling operations, the resulting sequence Vresult(:,:,Ti) of the action after Δt is obtained, Vresult(:,:,Ti) = |V(:,:,Ti) - V(:,:,Ti+Δt)|; where the time sequence T = (T1, T2, T3,......Tt), and Δt represents the calculation distance difference between two frames.
[0060] In this embodiment, there may be identical or invalid information-containing videos in the video data captured on the same production line. Therefore, in this solution, some videos with incorrect actions or only containing invalid targets, and even other excessive interference items will be detected and removed. If there is a large amount of redundant feature information in the data, this solution will handle it by elimination. Finally, after a series of operations such as detection, merging, and elimination, all video action data sets are obtained in this solution.
[0061] Working principle: By identifying the basic unit and comparing it with the standard database, the matching rate of the sequence and the database is obtained, so as to achieve the effect of standardizing action behavior analysis; by adopting a hierarchical comparison strategy, the error rate and efficiency of the actor's work can be clearly observed, which is helpful for the management and planning of pipeline operations.
[0062] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications without creative contributions to this embodiment according to needs after reading this specification, but as long as it is within the scope of the claims of the present invention, it is protected by the patent law.
Claims
1. A method for analyzing the efficiency of action combinations based on computer vision recognition, characterized in that: Specifically, it includes the following steps: 1) Use a camera to collect real-time video data and preprocess it; 2) Decompose the continuous actions in the preprocessed video data in step 1) into multiple consecutive images of translational change frames; 3) Decompose the action related to the time context into image classification, and generate the basic action unit C through the hierarchical progressive strategy to perform sequential combination on the image categories i to form an ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…}; The specific steps of the hierarchical progressive strategy are as follows: A) Detect whether the camera is turned on. If the camera is on, proceed to the next step. If the camera is off, stop the subsequent operations; B) Continuously read the camera video data and identify the basic units. If the identification is successful, input each frame into the target detection model and output a list of basic unit identification results, and add it to the identification sequence set detectSet, and then proceed to the next step; If no result is identified, mark it as an "error" basic unit; C) Compare the size of the sequence set detectSet with the window length L. If the size of detectSet is greater than the window L, call the window smoothing strategy. At this time, smoothing and increasing the identification sequence set are alternated. When the smoothing times reach the threshold or when the currently identified basic unit ends, convert the smoothed result resultSet into a mapping relationship baseSet, and perform a time comparison to determine whether it times out. If it times out, add a mark and proceed to the next step. If the size of detectSet is less than the window L, directly proceed to the next step; D) Perform the shortest match between baseSet and the steps in the database. If it matches correctly with a certain step information in the database, define it as a "valid" step, discard baseSet, and at the same time add the matched step to the stepSet matching sequence, and then proceed to the next step. If the match fails, mark the first basic unit of baseSet as "invalid"; E) Match the "valid" step sequence stepSet with the action process information in the database again. If it matches correctly with the step sequence in a certain action information in the database, it is a "valid" action process. If stepSet does not match correctly with the step sequence in any action information in the database, it is an "invalid" action process; 4) Train the ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…} using a model to obtain a computer vision model; 5) Identify the basic units, interpret and process the identification result sequence of the basic units, and compare it with the standard database to obtain the matching rate between the sequence and the database, that is, to achieve the analysis of the standard action behavior.
2. The method for analyzing the efficiency of action combinations based on computer vision recognition according to claim 1, wherein: The generation of the action basic unit C in step 3) i The way of the ordered permutation and combination is to regard the process of completing a work task at each job site as a complete action process A i , and regard the action process A i as decomposed into several action steps S that make up the action process i , and then decompose the action step S i into multiple action basic units C i . Multiple basic units C i form an ordered permutation and combination A i ={C1, C2, C3…C i , C i+1 ,…}.
3. A method for analyzing the efficiency of action combinations based on computer vision recognition according to claim 1, characterized in that: The specific method of the preprocessing described in step 1) includes the following steps: (1) Obtain a video sequence set V (V ∈ Rnx×ny) from the original video data set R, and ensure that each video sequence only contains valid information and the characteristics of a single action; (2) Perform a clipping operation on the action content in each video sequence according to the starting frame and ending frame of the hand movement change of the actor, and sequentially obtain video sequence subsets: Vclip (Vclip ∈ Rkx×ky, 0 < kx < nx, 0 < ky < ny), and ensure that the action of each sub-sample in each sampled Vclip set can be within the kx×ky video frame box; (3) Suppose the frame video sequence at the $i$-th time $T_i$ of the video sequence $V$ is $V(:,:,T_i)$. After successive sampling operations, the result sequence $V_{result}(:,:,T_i)$ of the action after $\Delta t$ is obtained, where $V_{result}(:,:,T_i) = |V(:,:,T_i) - V(:,:,T_i+\Delta t)|$. Here, the time sequence $T=(T_1,T_2,T_3,\cdots,T_t)$, and $\Delta t$ represents the computational distance difference between two frames.
Citation Information
Patent Citations
Bridge-type crane main-beam local-damage locating method
CN110596242A
Building worker strain early warning analysis method and system based on computer vision
CN113469063A