Working hour identification model training method and process working hour optimization method, system and equipment

Through the working time identification model training method, the timing attention mechanism and loss function are used to solve the problems of inefficiency and difficulty in standardization caused by relying on manual measurement in working time analysis, and efficient and standardized working time measurement and analysis are achieved.

CN119942642APending Publication Date: 2025-05-06广域铭岛数字科技有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510010097.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The existing working hours analysis relies on manual measurement, which is inefficient and is easily affected by subjective factors of the measuring staff, resulting in the inability to standardize the measurement process, difficult to meet actual guidance requirements, and poor optimization results.

Method used

A working-time recognition model training method is adopted. By obtaining the action image sequence, using a sequence encoder and feature decoder, combining the timing attention mechanism, action loss function and time regression loss function, model training is carried out to obtain the process working-time recognition model.

Benefits of technology

It improves the efficiency of working hours measurement, realizes the standardization of working hours measurement process, and enhances the guiding effect of working hours analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942642A_ABST
    Figure CN119942642A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of production working hours, and discloses a working hour identification model training method and a process working hour optimization method, system and device. According to the application, a to-be-trained model is obtained through a sequence encoder for encoding a received action image sequence by using a time sequence attention mechanism and a feature decoder for performing action recognition on the action image sequence according to action time sequence features; the attention loss function is utilized to enable the model to capture action information of different image frames through a time sequence attention mechanism, the action loss function is utilized to enable the model to learn an action change rule of the image frames, and the time regression loss function is utilized to enable the model to learn the positioning capability of a time boundary; according to the method, the process man-hour identification model is obtained through model training, so that the process man-hour identification model obtained through training realizes process man-hour measurement, the man-hour measurement efficiency is improved, the standardization of the man-hour measurement process is realized, and the guiding effect of man-hour analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of production working hours, and in particular to a working hour recognition model training method and a process working hour optimization method, system and equipment. Background Art

[0002] As a key link in the automobile manufacturing process, automobile final assembly is responsible for accurately assembling all parts into a complete automobile. Its complexity and the numerous processes involved are self-evident. Among them, work time analysis plays a vital role. It is directly related to the improvement of production efficiency, the reduction of costs and the optimization of production processes, and involves multiple levels from the formulation of standard work hours to continuous improvement.

[0003] At present, man-hour analysis builds a scientific standard man-hour system by accurately measuring and analyzing actual operation time and work intensity, so as to accurately identify the "bottleneck" process with the longest time or the lowest efficiency on the production line and make targeted improvements. At the same time, based on accurate man-hour data, enterprises can prepare more realistic production plans, such as material requirement plans and master production plans, to ensure that resources are used efficiently while meeting market demand. In addition, man-hour analysis can also help enterprises reasonably control labor costs and provide a strong basis for formulating pricing strategies.

[0004] However, the measurement of working hours mainly relies on manual measurement, which is not only inefficient but also easily affected by subjective factors of the measuring personnel, resulting in the inability to standardize the working hour measurement process. This series of problems makes it difficult for working hour analysis to meet actual guidance requirements, and the optimization effect is greatly reduced. Summary of the invention

[0005] In order to provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. The summary is not an extensive review, nor is it intended to identify key / critical components or delineate the scope of protection of these embodiments, but rather serves as a prelude to the detailed description that follows.

[0006] In view of the shortcomings of the prior art mentioned above, the present application discloses a work time recognition model training method and a process work time optimization method, system, and equipment to improve the work time measurement efficiency and standardize the measurement process, thereby improving the guidance effect of work time analysis.

[0007] The present application discloses a method for training a work time recognition model, comprising: obtaining a model to be trained, the model to be trained comprising a sequence encoder and a feature decoder, wherein the sequence encoder is used to encode a received action image sequence using a temporal attention mechanism to obtain action temporal features, and the feature decoder is used to perform action recognition on the action image sequence according to the action temporal features to obtain the current work time corresponding to the process node in the action image sequence; constructing a model loss function corresponding to the model to be trained, the model loss function is composed of at least one of an attention loss function, an action loss function and a time regression loss function, wherein the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism, the action loss function is used to characterize the probability distribution difference loss of the current work time between a predicted interval and an actual interval, and the time regression loss function is used to characterize the time boundary difference loss of the current work time between a predicted interval and an actual interval; and using the model loss function to perform model training on the model to be trained to obtain a process work time recognition model.

[0008] In one embodiment of the present application, the sequence encoder includes in sequence: a temporal position encoding unit, which is used to embed the corresponding temporal position code into the action image sequence to obtain a temporal embedding feature; a spatiotemporal feature extraction unit, which is composed of multiple layers of Transformer encoders arranged in sequence, wherein the Transformer encoder is used to encode the data input into the Transformer encoder, and use the temporal attention mechanism to decompose it into static attention and dynamic attention, fuse the static attention and the dynamic attention, and output the fused attention feature.

[0009] In one embodiment of the present application, the feature decoder includes: a classification head, which is used to predict the action probabilities corresponding to different moments in the action image sequence according to the action timing features, and obtain multiple candidate working time frames; a regression head, which is used to predict the time boundaries corresponding to different process nodes in the action image sequence according to the action timing features, and obtain multiple candidate boundary boxes; a detection unit, which is used to smooth each of the candidate working time boxes and each of the candidate boundary boxes using a soft non-maximum suppression algorithm, and determine the current working time corresponding to the process node based on the smoothed candidate boundary boxes and the candidate boundary boxes.

[0010] In one embodiment of the present application, the attention loss function is expressed by the following formula: Where, L TAU_dis is the attention difference loss, A is the number of image segments in the action image sequence, H a is the result of the a-th image segment after the temporal attention weight is applied, λ is a preset scalar value, I is a diagonal matrix, ‖‖F is the Frobenius norm.

[0011] In one embodiment of the present application, the action loss function is expressed by the following formula: In the formula, is the KL divergence loss between the predicted interval and the actual interval of the current working hours, is the feature map corresponding to the prediction interval, M represents the total number of frames corresponding to the prediction interval, C is the vector dimension, W is the image width of the action image sequence, H is the image height of the action image sequence, y∈R (N×C)×W×H is the feature map corresponding to the actual interval, N represents the total number of frames corresponding to the actual interval, n represents the frame number corresponding to the actual interval, is the feature map difference set of the prediction interval, Δy is the feature map difference set of the actual interval, is the probability distribution corresponding to the predicted interval, and σ(Δy) is the probability distribution corresponding to the actual interval.

[0012] In one embodiment of the present application, the time regression loss function is expressed by the following formula: Where, L reg is the time boundary difference loss between the predicted interval and the actual interval of the current working hours, tIoU is the time boundary intersection-over-union ratio between the predicted interval and the actual interval of the current working hours, ρ is the Euclidean distance between the predicted interval and the actual interval, offsets is the time boundary of the predicted interval, offsets gt is the time boundary of the actual interval, and c is a hyperparameter.

[0013] The present application discloses a method for optimizing process working hours. A process working hour recognition model is obtained based on the working hour recognition model training method described in any one of claims 1 to 6, and the process working hour recognition model is adopted. The method for optimizing process working hours includes: obtaining a process flow, wherein the process flow includes one or more process nodes and benchmark working hours corresponding to each of the process nodes; performing image acquisition when a worker executes the process flow to obtain an action image sequence; annotating a human body area and a tool area in the action image sequence to obtain a mask image sequence, and inputting the mask image sequence into the process working hour recognition model to obtain a current working hour corresponding to the process node; generating a value analysis result corresponding to the process node according to a comparison result between the current working hour and the benchmark working hour, and / or optimizing the process balance rate corresponding to the process node according to the current working hour and the benchmark working hour.

[0014] In one embodiment of the present application, the process balance rate is optimized by the following formula: In the formula, LOB is the optimized process balance rate, n is the number of process nodes, c i is the process cycle corresponding to the i-th process node, c max is the process tact corresponding to the process node with the longest process tact, and γ is the ratio between the current working time and the benchmark working time.

[0015] The present application discloses a work time recognition model training system, comprising: an acquisition module, used to acquire a model to be trained, the model to be trained comprising a sequence encoder and a feature decoder, wherein the sequence encoder is used to encode a received action image sequence using a temporal attention mechanism to obtain action temporal features, and the feature decoder is used to perform action recognition on the action image sequence according to the action temporal features to obtain the current work time corresponding to the process node in the action image sequence; a construction module, used to construct a model loss function corresponding to the model to be trained, the model loss function is composed of at least one of an attention loss function, an action loss function and a time regression loss function, wherein the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism, the action loss function is used to characterize the probability distribution difference loss of the current work time between the predicted interval and the actual interval, and the time regression loss function is used to characterize the time boundary difference loss of the current work time between the predicted interval and the actual interval; a training module, used to perform model training on the model to be trained using the model loss function to obtain a process work time recognition model.

[0016] The present application discloses an electronic device, comprising: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above method.

[0017] Beneficial effects of this application:

[0018] The model to be trained is obtained by using a sequence encoder that encodes the received action image sequence using a temporal attention mechanism and a feature decoder that recognizes the action image sequence according to the action temporal features, and constructing an attention loss function for the temporal attention mechanism in the sequence encoder, an action loss function for the probability distribution difference of the prediction interval in the feature decoder, and a time regression loss function for the time boundary difference in the feature decoder, and then training the model to obtain a process time recognition model. In this way, the attention loss function is used to enable the model to capture the action information of different image frames through the temporal attention mechanism, the action loss function is used to enable the model to learn the action change law of the image frame, and the time regression loss function is used to enable the model to learn the positioning ability of the time boundary, so that the trained process time recognition model can realize the process time measurement, which not only improves the efficiency of time measurement, but also realizes the standardization of the time measurement process, thereby improving the guidance effect of time analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 It is a flowchart of a working time recognition model training method in an embodiment of the present application;

[0020] Figure 2 is a schematic diagram of the structure of a sequence encoder in an embodiment of the present application;

[0021] Figure 3 It is a flow chart of a temporal attention mechanism in an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of a mask image sequence in an embodiment of the present application;

[0023] Figure 5 It is a schematic diagram of a process flow of a method for optimizing working hours in an embodiment of the present application;

[0024] Figure 6 It is a structural diagram of a working time recognition model training system in an embodiment of the present invention;

[0025] Figure 7 It is a schematic diagram of the structure of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The following describes the embodiments of the present invention through specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and sub-samples in the embodiments can be combined with each other without conflict.

[0027] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0028] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0029] The terms "first", "second", etc. in the specification and claims of the embodiments of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged where appropriate, so that the embodiments of the embodiments of the present disclosure described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions.

[0030] Unless otherwise stated, the term "plurality" means two or more.

[0031] In the embodiment of the present disclosure, the character " / " indicates that the preceding and following objects are in an "or" relationship. For example, A / B indicates: A or B.

[0032] The term "and / or" is a description of the association relationship between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or, A and B.

[0033] Combination Figure 1 As shown, the embodiment of the present disclosure provides a method for training a work-hour recognition model, including:

[0034] Step S101, obtaining a model to be trained, where the model to be trained includes a sequence encoder and a feature decoder;

[0035] Among them, the sequence encoder is used to encode the received action image sequence using the temporal attention mechanism to obtain the action temporal features;

[0036] Among them, the feature decoder is used to perform action recognition on the action image sequence according to the action timing characteristics, and obtain the current working hours corresponding to the process nodes in the action image sequence;

[0037] Step S102, constructing a model loss function corresponding to the model to be trained, where the model loss function is composed of at least one of an attention loss function, an action loss function, and a time regression loss function;

[0038] Among them, the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism;

[0039] Among them, the action loss function is used to characterize the probability distribution difference loss of the current working hours between the predicted interval and the actual interval;

[0040] Among them, the time regression loss function is used to characterize the time boundary difference loss between the predicted interval and the actual interval of the current working hours;

[0041] Step S103, using the model loss function to perform model training on the model to be trained to obtain a process time recognition model.

[0042] The method for training a work time recognition model provided by the embodiment of the present disclosure is adopted. The model to be trained is obtained by using a sequence encoder that encodes the received action image sequence using a temporal attention mechanism and a feature decoder that recognizes the action image sequence according to the action temporal features, and constructing an attention loss function for the temporal attention mechanism in the sequence encoder, an action loss function for the probability distribution difference of the prediction interval in the feature decoder, and a time regression loss function for the time boundary difference in the feature decoder, and then performing model training to obtain a process work time recognition model. In this way, the attention loss function is used to enable the model to capture the action information of different image frames through the temporal attention mechanism, the action loss function is used to enable the model to learn the action change law of the image frame, and the time regression loss function is used to enable the model to learn the positioning ability of the time boundary, so that the trained process work time recognition model can realize the process work time measurement, which not only improves the work time measurement efficiency, but also realizes the standardization of the work time measurement process, thereby improving the guidance effect of work time analysis.

[0043] Optionally, the sequence encoder includes in sequence: a temporal position encoding unit, used to embed the corresponding temporal position code into the action image sequence to obtain a temporal embedding feature; a spatiotemporal feature extraction unit, composed of multiple layers of Transformer encoders arranged in sequence, wherein the Transformer encoder is used to encode the data input into the Transformer encoder, and use the temporal attention mechanism to decompose it into static attention and dynamic attention, fuse the static attention and dynamic attention, and output the fused attention feature.

[0044] Combination Figure 2 As shown, an embodiment of the present disclosure provides a sequence encoder, including a temporal position encoding unit 201 and a spatiotemporal feature extraction unit 202.

[0045] The temporal position encoding unit 201 is used to embed the corresponding temporal position code into the action image sequence to obtain a temporal embedding feature.

[0046] In some embodiments, temporal position encoding is a technique used when processing time series data, which is mainly used to add unique position information to each position of the input sequence to help the model understand the sequential relationship between the data. The temporal position encoding is composed of a set of learnable vectors, the total number of vectors is the time step T, and the dimension of each vector is C. Then the set of vectors P∈R T×C , where R is a real number set. When processing image sequence features, the embedding vector is determined from the vector set according to the time position t in the time series.

[0047] In some embodiments, the embedding vector is determined from the set of vectors by the following formula:

[0048] E position (t,H,W,C)=P(t,C) Formula (1)

[0049] In formula (1), E position (t,H,W,C) is the embedding vector, t is the time position, H is the image height of the action image sequence, W is the image width of the action image sequence, and C is the vector dimension.

[0050] In some embodiments, temporal position coding is used as an explicit method to embed information of different temporal positions in a feature set, effectively improving the ability to distinguish similar features. Compared with the implicit capture of temporal convolution, this encoding method is more direct and efficient, especially in long-distance temporal modeling, where temporal position embedding shows greater advantages and avoids the additional space-time overhead of temporal convolution due to stacking multiple convolution blocks. Therefore, temporal position coding can improve the performance of feature decoders in temporal modeling.

[0051] In some embodiments, the spatiotemporal feature extraction unit 202 is constructed based on the Transformer encoder, which sequentially uses temporal attention mechanism (Temporal Attention), multi-head attention mechanism (Mult-Head Attention), residual connection (Add) and normalization (Norm), feedforward neural network (Feed Forward), residual connection and normalization, and downsampling (Downsample) to complete encoding.

[0052] Combination Figure 3As shown in the figure, the temporal attention mechanism is different from the recurrent neural network. It processes time changes in parallel. The temporal attention mechanism decomposes the spatiotemporal attention into two parts: intra-frame static attention and inter-frame dynamic attention. The intra-frame static attention uses small-core deep convolution, dilated convolution and 1×1 convolution to simulate large-core convolution, and obtains a larger receptive field within the frame to capture the long-distance dependencies within the frame. The inter-frame dynamic attention uses the inter-channel attention method to learn the channel weights between different frames, thereby capturing the change trend between frames. The fused attention value is equal to the product of the intra-frame dynamic attention and the intra-frame static attention.

[0053] In some embodiments, the temporal attention mechanism is expressed by the following formula:

[0054]

[0055] In formula (2), h is the input feature, SA is the static attention within the frame, Conv(1×1) is the 1×1 convolution layer, DW-DConv is the depthwise convolution, DWConv is the depthwise convolution, DA is the dynamic attention within the frame, FC is the fully connected layer, AvgPool is the average pooling, is the Hadamard product, ⊙ is the Kronecker product, where the depthwise dilated convolution combines the characteristics of depthwise convolution and dilated convolution. In depthwise convolution, each input channel is convolved independently, the number of output channels is the same as the number of input channels, and each output channel is generated only by the corresponding input channel. The dilated convolution inserts holes in the convolution kernel to increase the receptive field of the convolution kernel while keeping the number of parameters and the amount of computation unchanged.

[0056] In some embodiments, the sequence encoder ultimately outputs action temporal features in a feature pyramid structure.

[0057] Optionally, the feature decoder includes: a classification head, used to predict the action probabilities corresponding to different moments in the action image sequence according to the action timing features, and obtain multiple candidate working time boxes; a regression head, used to predict the time boundaries corresponding to different process nodes in the action image sequence according to the action timing features, and obtain multiple candidate boundary boxes; a detection unit, used to smooth each candidate working time box and each candidate boundary box using a soft non-maximum suppression algorithm, and determine the current working time corresponding to the process node based on the smoothed candidate boundary boxes and the candidate boundary boxes.

[0058] In some embodiments, the feature decoder is a lightweight convolutional network using a classification head and a regression head.

[0059] In some embodiments, the classification head predicts the action probability at each moment t by checking each moment t of the L layers on the feature pyramid. The classification network consists of three layers of one-dimensional convolution with a convolution kernel size of 3. The first two layers use layer normalization and ReLU (Rectified Linear Unit) activation function. A Sigmoid function is attached to each output dimension to predict the probability of C action categories. The classification network is attached to each pyramid layer, and its parameters are shared among all layers.

[0060] In some embodiments, the regression head is similar to the classification head, checking each time t of the N layers on the pyramid, but predicting the distance to the start and end of the action only when the current time step t is in the action. The regression head is also implemented using a one-dimensional convolutional network with a ReLU activation function attached at the end.

[0061] In some embodiments, in the target detection task, the traditional NMS (Non-Maximum Suppression) algorithm is used to eliminate overlapping bounding boxes to reduce redundancy and improve detection accuracy. When the target objects are dense or overlapping, the traditional NMS algorithm may mistakenly remove some correct detection boxes, resulting in target loss. The core idea of ​​the soft non-maximum suppression algorithm (SoftNMS) is to consider the influence of all other detection boxes when calculating the score of each detection box, and retain more candidate boxes by smoothing the scores. The Soft-NMS algorithm applies a score decay function (such as an exponential or Gaussian function) to each bounding box, instead of directly setting the low-scoring boxes that overlap with other high-scoring bounding boxes to zero. This smoothing process ensures that even if some bounding boxes overlap with high-scoring boxes to a certain extent, they will not be completely ignored, but their scores will be reduced, allowing these boxes to participate in subsequent detections to a certain extent.

[0062] Optionally, the attention loss function is expressed by the following formula:

[0063]

[0064] In formula (3), L TAU_dis is the attention difference loss, A is the number of image segments in the action image sequence, and H a is the result of the a-th image segment after the temporal attention weight is applied, λ is a preset scalar value, I is a diagonal matrix, ‖‖ F is the Frobenius norm.

[0065] In some embodiments, the attention loss function forces the temporal attention mechanism to capture motion information of different video frames in different frame images. When the attention weights in the temporal attention mechanism tend to a similar distribution in different steps, the penalty term will be calculated to obtain a larger value, thereby guiding the differences between the attention weights to capture different motion features in the same video clip.

[0066] Optionally, the action loss function is expressed by the following formula:

[0067]

[0068] In formula (4), is the KL divergence loss between the predicted interval and the actual interval of the current working hours, is the feature map corresponding to the prediction interval, M represents the total number of frames corresponding to the prediction interval, m represents the frame number corresponding to the prediction interval, C is the vector dimension, W is the image width of the action image sequence, H is the image height of the action image sequence, y∈R (N×C)×W×H is the feature map corresponding to the actual interval, N represents the total number of frames corresponding to the actual interval, and n represents the frame number corresponding to the actual interval. is the feature map difference set of the prediction interval, Δy is the feature map difference set of the actual interval, is the probability distribution corresponding to the prediction interval, and σ(Δy) is the probability distribution corresponding to the actual interval.

[0069] In some embodiments, the prediction interval and the feature map difference set of the prediction interval are expressed by the following formula:

[0070]

[0071] In some embodiments, the action loss function is established based on differential KL divergence regularization, taking into account both intra-frame errors and inter-frame changes, and is used to optimize the loss function of the temporal regression task, wherein the model is forced to learn the changing pattern of the video frames by converting the difference between the predicted frame interval and the true frame interval into a probability distribution and calculating the KL divergence between them.

[0072] Optionally, the temporal regression loss function is expressed by the following formula:

[0073]

[0074] In formula (6), L reg is the time boundary difference loss between the predicted interval and the actual interval of the current working hours, tIoU is the time boundary intersection-over-union ratio between the predicted interval and the actual interval of the current working hours, ρ is the Euclidean distance between the predicted interval and the actual interval, offsets is the time boundary of the predicted interval, and offsetsgt is the time boundary of the actual interval, c is a hyperparameter, and the hyperparameter c is used to prevent the loss function value from being too large and to speed up the convergence.

[0075] In some embodiments, a model loss function is established based on the attention loss function, the action loss function, and the time regression loss function, and the model loss function is expressed by the following formula:

[0076]

[0077] In formula (7), α and β are both hyperparameters, and the hyperparameters α and β are used to adjust the relative importance of different loss terms.

[0078] Optionally, after the model loss function is used to train the model to be trained and the process working time recognition model is obtained, the method also includes: obtaining the process flow, wherein the process flow includes one or more process nodes and the benchmark working time corresponding to each process node; performing image acquisition when the staff executes the process flow to obtain an action image sequence; annotating the human body area and the tool area in the action image sequence to obtain a mask image sequence, and inputting the mask image sequence into the process working time recognition model to obtain the current working time corresponding to the process node; generating a value analysis result corresponding to the process node according to the comparison result between the current working time and the benchmark working time, and / or optimizing the process balance rate corresponding to the process node according to the current working time and the benchmark working time.

[0079] In some embodiments, the action image sequence is passed through a preset segmentation model to obtain a mask image sequence such as Figure 4 As shown, the mask image sequence can clearly identify the position and outline of the human body and the tools used in the video, which is convenient for subsequent analysis and processing of the human actions in the original video.

[0080] Combination Figure 5 As shown, the embodiment of the present disclosure provides a method for optimizing process time, including:

[0081] Step S501, obtaining a process flow, the process flow including one or more process nodes and the benchmark working hours corresponding to each process node;

[0082] Among them, the operating procedures of the assembly workshop personnel are disassembled in detail, down to the most basic action units, so as to formulate standardized operating specifications;

[0083] Among them, based on the MTM (Method-Time-Measurement) method of the predetermined time system, a standard combined action time library is customized, and the standard actions after detailed decomposition are coded and summarized. The standard actions cover a variety of types, including but not limited to body movements, parts installation, threaded fastening, and buckle fixing.

[0084] Among them, the standard actions are summarized into the combined action working time library, which contains the standard actions and the scheduled time parameters corresponding to the standard actions. The scheduled time in the combined action working time library is used as the basis for optimizing the efficiency of the action improvement;

[0085] Step S502, when the worker performs the process flow, image acquisition is performed to obtain an action image sequence, and the human body area and tool area in the action image sequence are marked to obtain a mask image sequence;

[0086] Among them, the designed time-series motion detection method is used to extract features from video data, calculate various motion behaviors, and realize real-time motion tracking and positioning of personnel;

[0087] Step S503, inputting the mask image sequence into the process time recognition model to obtain the current time corresponding to the process node;

[0088] Among them, by capturing and locating actions, determining the action types and the order in which they occur, each action has a corresponding actual working time;

[0089] Step S504, generating a value analysis result corresponding to the process node according to the comparison result between the current working hours and the benchmark working hours;

[0090] Among them, the actual action working hours after analysis are compared with the standard action working hours in the combined action working hours library, so as to analyze each subsequent action according to value-added working hours, necessary non-value-added working hours and auxiliary working hours;

[0091] Step S505, optimizing the process balance rate corresponding to the process node according to the current working hours and the benchmark working hours;

[0092] Among them, by analyzing the waste points identified and organizing them into an optimization list, an operation instruction manual for the process is compiled, which can be used for training and guidance of employees.

[0093] Optionally, the process balance rate is optimized by the following formula:

[0094]

[0095] In formula (8), LOB is the optimized process balance rate, n is the number of process nodes, c iis the process cycle corresponding to the i-th process node, c max is the process tact corresponding to the process node with the longest process tact, and γ is the ratio between the current working time and the benchmark working time.

[0096] Combination Figure 6 As shown, the embodiment of the present disclosure provides a work time recognition model training system, including:

[0097] The acquisition module 601 is used to acquire the model to be trained, and the model to be trained includes a sequence encoder and a feature decoder, wherein the sequence encoder is used to encode the received action image sequence using the temporal attention mechanism to obtain the action temporal features, and the feature decoder is used to perform action recognition on the action image sequence according to the action temporal features to obtain the current working hours corresponding to the process nodes in the action image sequence;

[0098] The construction module 602 is used to construct a model loss function corresponding to the model to be trained, and the model loss function is composed of at least one of an attention loss function, an action loss function, and a time regression loss function, wherein the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism, the action loss function is used to characterize the probability distribution difference loss of the current working hours between the predicted interval and the actual interval, and the time regression loss function is used to characterize the time boundary difference loss of the current working hours between the predicted interval and the actual interval;

[0099] The training module 603 is used to perform model training on the model to be trained using the model loss function to obtain a process time recognition model.

[0100] The work time recognition model training system provided by the embodiment of the present disclosure is adopted, and the model to be trained is obtained by using a sequence encoder that encodes the received action image sequence according to the temporal attention mechanism, and a feature decoder that recognizes the action image sequence according to the action temporal features, and constructing an attention loss function for the temporal attention mechanism in the sequence encoder, an action loss function for the probability distribution difference of the prediction interval in the feature decoder, and a time regression loss function for the time boundary difference in the feature decoder, and then performing model training to obtain a process work time recognition model. In this way, the attention loss function is used to enable the model to capture the action information of different image frames through the temporal attention mechanism, the action loss function is used to enable the model to learn the action change law of the image frame, and the time regression loss function is used to enable the model to learn the positioning ability of the time boundary, so that the trained process work time recognition model can realize the process work time measurement, which not only improves the work time measurement efficiency, but also realizes the standardization of the work time measurement process, thereby improving the guidance effect of work time analysis.

[0101] An embodiment of the present disclosure further provides an electronic device, including: a processor and a memory; the memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the electronic device executes the above method.

[0102] Figure 7 The structure diagram of the computer system suitable for implementing the electronic device of the embodiment of the present application is shown. It should be noted that: Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0103] like Figure 7 As shown, the computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 702 or the program loaded from the storage part 708 to the random access memory (RAM) 703, such as executing the method in the above embodiment. In the random access memory 703, various programs and data required for system operation are also stored. The central processing unit 701, the read-only memory 702 and the random access memory 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0104] The following components are connected to the input / output interface 705: an input section 706 including a keyboard, a mouse, etc.; an output section 707 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 708 including a hard disk, etc.; and a communication section 709 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the input / output interface 705 as needed. A removable medium 711, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 710 as needed so that a computer program read therefrom is installed into the storage section 708 as needed.

[0105] In particular, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section 709, and / or installed from a removable medium 711. When the computer program is executed by a central processing unit (CPU) 701, various functions defined in the system of the present application are executed.

[0106] The above description and the accompanying drawings fully illustrate the embodiments of the present disclosure so that those skilled in the art can practice them. Other embodiments may include structural, logical, electrical, process and other changes. The embodiments represent possible changes only. Unless explicitly required, separate components and functions are optional, and the order of operation can be changed. Parts and sub-samples of some embodiments may be included in or replace parts and sub-samples of other embodiments. Moreover, the words used in this application are only used to describe the embodiments and are not used to limit the claims. As used in the description of the embodiments and the claims, unless the context clearly indicates, the singular forms of "a", "an" and "the" are intended to include plural forms as well. Similarly, the term "and / or" as used in this application refers to any and all possible combinations of listings containing one or more associated ones. In addition, when used in the present application, the term "comprise" and its variants "comprises" and / or comprising refer to the presence of a stated sub-sample, whole, step, operation, element, and / or component, but do not exclude the presence or addition of one or more other sub-samples, wholes, steps, operations, elements, components and / or groups of these. In the absence of further restrictions, the elements defined by the sentence "comprising a ..." do not exclude the presence of other identical elements in the process, method or device comprising the elements. In this article, each embodiment may focus on the differences from other embodiments, and the same and similar parts between the various embodiments may refer to each other. For the methods, products, etc. disclosed in the embodiments, if they correspond to the method part disclosed in the embodiments, then the relevant parts can refer to the description of the method part.

[0107] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software may depend on the specific application and design constraints of the technical solution. Technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of the present disclosure. Technicians can clearly understand that for the convenience and simplicity of description, the specific working process of the above-described system, device and unit can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0108] In the embodiments disclosed herein, the disclosed methods and products (including but not limited to devices, equipment, etc.) can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of units can be only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some sub-samples can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between each other shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to implement this embodiment. In addition, each functional unit in the embodiment of the present disclosure may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit.

[0109] The flowchart and block diagram in the accompanying drawings show the possible architecture, functions and operations of the system, method and computer program product according to the embodiment of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and a part of the module, program segment or code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two consecutive operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

Claims

1. A method for training a work-hour recognition model, characterized in that: include: Acquire a model to be trained, the model to be trained comprising a sequence encoder and a feature decoder, wherein the sequence encoder is used to encode the received action image sequence using a temporal attention mechanism to obtain action temporal features, and the feature decoder is used to perform action recognition on the action image sequence according to the action temporal features to obtain the current working hours corresponding to the process nodes in the action image sequence; Constructing a model loss function corresponding to the model to be trained, wherein the model loss function is composed of at least one of an attention loss function, an action loss function, and a time regression loss function, wherein the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism, the action loss function is used to characterize the probability distribution difference loss of the current working hours between the predicted interval and the actual interval, and the time regression loss function is used to characterize the time boundary difference loss of the current working hours between the predicted interval and the actual interval; The model loss function is used to perform model training on the model to be trained to obtain a process time recognition model.

2. The method for training a man-hour recognition model according to claim 1, characterized in that: The sequence encoder comprises in sequence: A temporal position encoding unit, used for embedding corresponding temporal position codes into the action image sequence to obtain temporal embedding features; The spatiotemporal feature extraction unit is composed of multiple layers of Transformer encoders arranged in sequence, wherein the Transformer encoder is used to encode the data input into the Transformer encoder, and decompose it into static attention and dynamic attention using a temporal attention mechanism, fuse the static attention and the dynamic attention, and output the fused attention feature.

3. The method for training a man-hour recognition model according to claim 1, characterized in that: The feature decoder comprises: A classification head, used for predicting the action probabilities corresponding to different moments in the action image sequence according to the action timing features, to obtain a plurality of candidate working time frames; A regression head, used for predicting the time boundaries corresponding to different process nodes in the action image sequence according to the action timing characteristics, to obtain a plurality of candidate boundary boxes; The detection unit is used to smooth each of the candidate working time boxes and each of the candidate bounding boxes by using a soft non-maximum suppression algorithm, and determine the current working time corresponding to the process node according to the smoothed candidate bounding boxes and the candidate bounding boxes.

4. The method for training a man-hour recognition model according to claim 1, characterized in that: The attention loss function is expressed by the following formula: Where, L TAU_dis is the attention difference loss, A is the number of image segments in the action image sequence, H a is the result of the a-th image segment after the temporal attention weight is applied, λ is a preset scalar value, I is a diagonal matrix, ‖‖ F is the Frobenius norm.

5. The method for training a man-hour recognition model according to claim 1, characterized in that: The action loss function is expressed by the following formula: In the formula, is the KL divergence loss between the predicted interval and the actual interval of the current working hours, is the feature map corresponding to the prediction interval, M represents the total number of frames corresponding to the prediction interval, C is the vector dimension, W is the image width of the action image sequence, H is the image height of the action image sequence, y∈R (N×C)×W×H is the feature map corresponding to the actual interval, N represents the total number of frames corresponding to the actual interval, n represents the frame number corresponding to the actual interval, is the feature map difference set of the prediction interval, Δy is the feature map difference set of the actual interval, is the probability distribution corresponding to the prediction interval, and σ(Δy) is the probability distribution corresponding to the actual interval.

6. The method for training a man-hour recognition model according to claim 1, characterized in that: The time regression loss function is expressed by the following formula: Where, L reg is the time boundary difference loss between the predicted interval and the actual interval of the current working hours, tIoU is the time boundary intersection-over-union ratio between the predicted interval and the actual interval of the current working hours, ρ is the Euclidean distance between the predicted interval and the actual interval, offsets is the time boundary of the predicted interval, offsets gt is the time boundary of the actual interval, and c is a hyperparameter.

7. A method for optimizing process time, characterized in that: Based on the working time recognition model training method according to any one of claims 1 to 6, a process working time recognition model is obtained, and the process working time recognition model is adopted, and the process working time optimization method includes: Obtaining a process flow, wherein the process flow includes one or more process nodes and the benchmark working hours corresponding to each of the process nodes; When the worker performs the process flow, image acquisition is performed to obtain an action image sequence; Annotating the human body area and the tool area in the action image sequence to obtain a mask image sequence, and inputting the mask image sequence into the process time recognition model to obtain the current time corresponding to the process node; A value analysis result corresponding to the process node is generated based on a comparison result between the current working hours and the benchmark working hours, and / or a process balance rate corresponding to the process node is optimized based on the current working hours and the benchmark working hours.

8. The process time optimization method according to claim 7, characterized in that: The process balance rate is optimized by the following formula: In the formula, LOB is the optimized process balance rate, n is the number of process nodes, c i is the process cycle corresponding to the i-th process node, c max is the process tact corresponding to the process node with the longest process tact, and γ is the ratio between the current working time and the benchmark working time.

9. A work time recognition model training system, characterized in that: include: An acquisition module is used to acquire a model to be trained, wherein the model to be trained includes a sequence encoder and a feature decoder, wherein the sequence encoder is used to encode the received action image sequence using a temporal attention mechanism to obtain action temporal features, and the feature decoder is used to perform action recognition on the action image sequence according to the action temporal features to obtain the current working hours corresponding to the process nodes in the action image sequence; A construction module, used to construct a model loss function corresponding to the model to be trained, wherein the model loss function is composed of at least one of an attention loss function, an action loss function and a time regression loss function, wherein the attention loss function is used to characterize the attention difference loss corresponding to the temporal attention mechanism, the action loss function is used to characterize the probability distribution difference loss of the current working hours between the predicted interval and the actual interval, and the time regression loss function is used to characterize the time boundary difference loss of the current working hours between the predicted interval and the actual interval; The training module is used to perform model training on the model to be trained using the model loss function to obtain a process time recognition model.

10. An electronic device, characterized in that: include: Processor and memory; The memory is used to store computer programs, and the processor is used to execute the computer programs stored in the memory, so that the electronic device executes the working time recognition model training method as described in any one of claims 1 to 6, or the process working time optimization method as described in claim 7 or 8.