Repetitive behavior identification method for industrial production line scene

The modified C3D network with SimAM attention and LeakyReLU activation, combined with a dual-stage optimization, addresses low recognition accuracy in industrial repetitive behavior by enhancing feature extraction and training efficiency, ensuring precise and consistent assembly processes.

CN120318196APending Publication Date: 2025-07-15SUZHOU HANGSHENGJIA INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510470630.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The existing convolutional neural network model has low repetitive behavior recognition accuracy in industrial scenarios, making it difficult to meet the real-time monitoring and feedback requirements of industrial production lines.

Method used

Using a multi-view video dataset and an improved baseline C3D network, combined with the SimAM attention mechanism module and the LeakyReLU activation function, the fine-grained feature extraction and recognition of repetitive behavior is enhanced through dynamic differentiated learning rate scheduling strategy.

Benefits of technology

It realizes efficient and accurate repetitive behavior recognition in industrial scenarios, ensures consistency and accuracy of assembly quality, provides real-time monitoring and feedback mechanisms, and improves assembly quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318196A_ABST
    Figure CN120318196A_ABST
Patent Text Reader

Abstract

The invention discloses a repetitive behavior recognition method for an industrial production line scene, and relates to the technical field of computer visual behavior recognition, and the method comprises the steps: constructing an improved network based on a baseline C3D network, fusing a SimAM attention mechanism module in front of a pooling layer after an activation function of each layer, carrying out the feature extraction operation of an input video stream, and carrying out the recognition of the scene. Extracting the focusing capability of each repetitive behavior fine-grained feature improvement model on the key spatio-temporal features; changing a ReLU activation function of the baseline model; a training optimizer and a learning rate scheduling strategy are improved, the model training efficiency is improved, and a technical approach is provided for model real-time updating training. According to the generation method, the attention mechanism is fused with the three-dimensional convolutional neural network to capture repetitive behavior fine-grained features in the assembly line, the model recognition precision is improved by improving an activation function and a model training iteration mode, and an implementation path and a framework are provided for application of a behavior recognition technology in an industrial scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention mainly relates to the technical field of computer vision behavior recognition, and specifically to a method for recognizing repetitive behaviors in an industrial production line scenario. Background Art

[0002] Due to the labor-intensive nature of the assembly production line, manual assembly operations are easily affected by adverse factors such as fatigue and inattention, resulting in unstable assembly quality of products. At the same time, there will also be situations where operators do not operate in a standardized manner or do not follow the process flow, making it difficult to ensure the assembly accuracy and consistency of products. Currently, for manual assembly production lines, it mainly relies on manual inspection and post-event traceability mechanisms, and these methods have inherent defects such as lagging response and strong subjectivity. Therefore, there is an urgent need to establish a real-time monitoring and feedback mechanism through intelligent and digital monitoring and recognition means to ensure the accuracy and consistency of each assembly link, thereby improving the assembly quality and providing a new paradigm for the high-quality development of the manufacturing industry. This need has become the key to optimizing the manual assembly production line. With the extensive development and application of computer technology, behavior recognition technology based on computer vision has received extensive attention. By processing the video data collected by the camera to obtain meaningful visual cues, the computer can simulate humans to achieve information perception and understanding and has the ability to adapt to the environment autonomously. In the field of computer vision, human behavior recognition is a very important research branch. Using computer vision technology to automatically extract visual information related to behaviors from video sequences and analyze and interpret this information to obtain specific human behavior categories, the research on human behavior recognition has profound theoretical significance. Therefore, behavior recognition technology based on computer vision is becoming one of the core technologies to promote the intelligent upgrade of industry.

[0003] At present, for behavior recognition technology, deep learning methods based on convolutional neural networks are mainly adopted, and public datasets are used for model training. However, traditional convolutional neural network models can achieve accurate recognition of daily behaviors in public datasets similar to UCF101 and HMD851. However, in specific application scenarios such as industrial production lines, the repetitive behaviors of workers have fine-grained characteristics due to small behavior differences, and the recognition accuracy of traditional behavior recognition models is relatively low. According to a method and system for human abnormal behavior recognition based on a hybrid attention mechanism provided by the publication number: CN113516028A, feature extraction is performed on the original image to obtain low-level detail features F; the low-level detail features F are screened to obtain main significant features F" and input into a convolutional feature extraction module to obtain high-level semantic features, which are fused with the low-level detail features to obtain fused features; the loss between the predicted value and the actual value of the training sample is calculated to obtain a loss value, and the training parameters are optimized according to the loss value; the neural network model is trained based on the optimized training parameters and the fused features to obtain a trained abnormal behavior recognition model to realize the recognition of human abnormal behaviors. And a method for human behavior recognition based on multi-modal data and deep learning provided by the publication number: CN119600689A, which relates to the field of computer vision technology, mainly preprocesses two types of data, skeleton and RGB, extracts features from the RGB video stream using the Resnet3D18 network, and inputs the proposed RGB features and skeleton features into a feature fusion network to obtain fused features; different categories of actions are divided through a softmax classifier; the interference of noise and background on the classification performance is reduced through preprocessing, and the Resnet3D18 network and the adaptive graph convolutional network dual-stream feature extraction network effectively integrate and extract feature information, and the fusion network dynamically weights the two heterogeneous features with confidence to enable the model to dynamically adjust the importance of each modality according to the actual situation.

[0004] In summary, a method for recognizing repetitive behaviors in industrial assembly production lines with high precision and high efficiency is a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a method for recognizing repetitive behaviors in industrial production line scenarios to solve the technical problems proposed in the above background technology.

[0006] To achieve the above purpose, the present invention provides the following technical solutions:

[0007] A method for recognizing repetitive behaviors in industrial production line scenarios includes the following steps:

[0008] Step 1: Shoot the overall process of the assembly operation behavior. Use an industrial camera array with orthogonal distribution to synchronously shoot the assembly line, and construct a multi-view video dataset by aligning timestamps. The industrial camera array includes top-view, side-view, and front-view perspectives;

[0009] Step 2: Construct an improved network based on the baseline C3D network. Incorporate the SimAM attention mechanism module after the activation function and before the pooling layer in each layer, perform feature extraction operations on the input video stream, and extract fine-grained features of each repetitive behavior to enhance the model's focusing ability on key spatio-temporal features;

[0010] Step 3: According to the feature map extracted by the SimAM attention mechanism, to overcome the limitations of the ReLU activation function, change the ReLU activation function of the baseline model and replace it with the LeakyReLU activation function;

[0011] Step 4: According to the training defects of the baseline model, improve the training optimizer and learning rate scheduling strategy, which speeds up the model training efficiency and provides a technical approach for the model to be updated in real time. The optimization method of the improved training optimizer and learning rate scheduling strategy is specifically a two-stage adaptive optimization framework.

[0012] Preferably, the multi-view video dataset is a self-defined repetitive behavior dataset established by the industrial camera array for shooting repetitive behaviors in the industrial production line scenario. The self-defined repetitive behavior dataset needs to be strengthened by adding Gaussian noise, Gaussian blur, and contrast enhancement operations.

[0013] Preferably, the SimAM attention mechanism in the SimAM attention mechanism module specifically enhances the response of significant regions and suppresses redundant information by dynamically evaluating the importance of each position in the feature map, so as to accurately capture the subtle differences of repetitive behaviors, directly quantifies the feature differences through an energy function, and avoids increasing the model complexity;

[0014] The SimAM attention mechanism module evaluates the importance of the spatio-temporal positions of each pixel point in the feature map through a parameter-free energy function and dynamically adjusts the feature response. For the neurons at any spatio-temporal position considered in the extracted feature map, its energy value is defined as the mean square of the difference between the neuron and other neurons in the local domain.

[0015] Preferably, the feature map is extracted as shown in the following formula:

[0016] F∈R T′×H′×W′×C′ ;

[0017] Among them, T′, H′, and W′ represent the spatio-temporal dimensions of the feature map after downsampling; T′ represents the number of time frames of the feature map after sampling; H represents the height of the feature map after sampling, and W represents the width of the feature map after sampling; C represents the number of sampling channels.

[0018] Preferably, the formula for calculating the mean square difference is as follows:

[0019]

[0020] Among them,

[0021] |Ω| represents the total number of neurons in the domain;

[0022] e t,h,w represents the neuron energy value at any position, x t,h,w represents the neuron at t, h, w at any position; x t′,h′,w′ represents the neuron at the feature extraction location.

[0023] Preferably, in order to convert the energy value into a differentiable attention weight later, a scaling factor is introduced, and a non-linear mapping is performed according to the attention weight calculation formula to obtain the attention weight. Finally, the generated attention weight matrix is applied to each channel through Equation (4).

[0024] Preferably, the formula for calculating the attention weight is as follows:

[0025]

[0026] Among them,

[0027] w t,h,w represents the attention weight;

[0028] λ represents the scaling factor, set to 10 -4 ;

[0029] The attention weight matrix is as follows:

[0030] W ∈ R T′×H′×W′ .

[0031] Preferably, the ReLU activation function is as follows:

[0032] f(x) = max(0, x);

[0033] The LeakyReLU activation function is as follows:

[0034]

[0035] Among them, a represents the gradient, a > 0.

[0036] Preferably, the two-stage adaptive optimization framework specifically realizes the efficient training of the model through a dynamic differential learning rate strategy;

[0037] In the first stage, an improved adaptive moment estimation optimizer is adopted, and a decoupled weight decay mechanism is introduced. First, the network parameters are initialized, and the first-order moment estimation (gradient mean) and the second-order moment estimation (gradient variance) are calculated respectively by the exponential moving average method; during the time step iteration, the gradient of the current batch of data is calculated, and then the first-order moment estimation is updated, the second-order moment estimation is updated, and a bias correction term is used to avoid the moment estimation bias in the initial stage of training. Finally, the parameters are updated;

[0038] In the second stage, the cosine annealing strategy is integrated, and the learning rate is adjusted periodically so that it first decreases, then increases, and then decreases in the form of a cosine function, adjusts the parameters and regulates the convergence speed and model stability, and then performs the learning rate scheduling.

[0039] In summary, the present invention mainly has the following beneficial effects:

[0040] The present invention is directed to the repetitive behaviors in the industrial scenario production line, and accurately and efficiently realizes the recognition of repetitive behaviors in the industrial scenario in the assembly line, establishes a real-time monitoring and feedback mechanism for digital monitoring and recognition means, ensures the accuracy and consistency of each assembly link, thereby improving the assembly quality, and provides a new paradigm for the high-quality development of the manufacturing industry;

[0041] The present invention constructs a self-made data set for repetitive behaviors. Since the application scenarios and application objects are relatively unique, there is no publicly available data set for repetitive behaviors in the industrial scenario at present, so it is particularly important to construct a self-made data set. By using multiple cameras to record the repetitive behavior data set, it provides data support for the training of the behavior recognition model, and at the same time reflects the feasibility of the data augmentation method;

[0042] The present invention is a convolutional network model that can be continuously improved and iteratively optimized. It can continuously expand the data set according to the recognition situation of the actual scenario, apply the updated data set to model training, continuously optimize the model recognition accuracy, thereby continuously improving the robustness and adaptability of the algorithm model, and finally forming an intelligent recognition model with environmental perception ability, solving the problem of model decline of traditional convolutional networks in open scenarios, and enabling the system to still maintain a high recognition accuracy in the real environment where the data distribution changes continuously. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 is the overall flowchart of the present invention;

[0044] Figure 2 is a schematic diagram of the three-dimensional convolutional neural network attention addition module of the present invention. Detailed implementation manners

[0045] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0046] Embodiment

[0047] As Figure 1 and Figure 2 shown, the multi-view video dataset is a self-defined repetitive behavior dataset established by an industrial camera array for shooting repetitive behaviors in the industrial production line scenario. The self-defined repetitive behavior dataset needs to be enhanced by adding Gaussian noise, Gaussian blur, and contrast enhancement operations.

[0048] In the SimAM attention mechanism module, the SimAM attention mechanism specifically enhances the response of the significant region and suppresses redundant information by dynamically evaluating the importance of each position in the feature map, so as to accurately capture the subtle differences of repetitive behaviors. The feature difference is directly quantified through an energy function, avoiding increasing the complexity of the model;

[0049] The SimAM attention mechanism module evaluates the importance of the spatio-temporal positions of each pixel point in the feature map through a parameter-free energy function and dynamically adjusts the feature response, considering neurons at any spatio-temporal position in the extracted feature map; its energy value is defined as the mean square difference between the neuron and other neurons in the local domain.

[0050] The feature map is extracted as shown in the following formula:

[0051] F∈R T′×H′×W′×C′ ;

[0052] where T′, H′, and W′ represent the spatio-temporal dimensions of the feature map after downsampling; T′ represents the number of time frames of the feature map after sampling; H represents the height of the feature map after sampling, and W represents the width of the feature map after sampling; C represents the number of sampling channels.

[0053] The formula for calculating the mean square difference is as follows:

[0054]

[0055] where

[0056] |Ω| represents the total number of neurons in the domain;

[0057] e t,h,w represents the neuron energy value at any position, x t,h,w represents the neuron at any position of t, h, w; x t′,h′,w′ represents the neuron at the feature extraction position.

[0058] After that, in order to convert the energy value into a differentiable attention weight, a scaling factor is introduced, and a non-linear mapping is performed according to the attention weight calculation formula to obtain the attention weight. Finally, the generated attention weight matrix applies this attention weight to each channel through Equation (4).

[0059] The attention weight calculation formula is as follows:

[0060]

[0061] Among them,

[0062] w t,h,w : represents the attention weight;

[0063] λ: represents the scaling factor, set to 10 -4 ;

[0064] The attention weight matrix is as follows:

[0065] W∈R T′×H′×W′ 。

[0066] The ReLU activation function is as follows:

[0067] f(x) = max(0, x);

[0068] The LeakyReLU activation function is as follows:

[0069]

[0070] Among them, a: represents the gradient, a > 0.

[0071] The two-stage adaptive optimization framework specifically realizes the efficient training of the model through the dynamic differential learning rate strategy;

[0072] In the first stage, an improved adaptive moment estimation optimizer is adopted, and a decoupled weight decay mechanism is introduced. First, the network parameters are initialized, and the first-order moment estimation (gradient mean) and the second-order moment estimation (gradient variance) are calculated respectively by the exponential moving average method; at the time step iteration, the gradient of the current batch of data is calculated, and then the first-order moment estimation is updated, the second-order moment estimation is updated, and in order to use the bias correction term to avoid the moment estimation bias in the initial stage of training, finally the parameters are updated;

[0073] In the second stage, the cosine annealing strategy is fused. By periodically adjusting the learning rate, it first decreases and then increases, and then decreases in the form of a cosine function, adjusts the parameters and regulates the convergence speed and model stability, and then performs the learning rate scheduling.

[0074] A method for identifying repetitive behaviors in industrial production line scenarios, whose baseline model is the C3D model, uses 3D convolution to extract spatio-temporal features of videos. The 3D convolution calculation formula is shown in Equation (1):

[0075] F l = σ(W l * F l-1 + b l ) (1)

[0076] Wherein,

[0077] F l : represents the feature map of the l-th layer; W l : represents the weight of the 3D convolution kernel; *: represents the 3D convolution operation; b l : represents the bias term; σ: represents the activation function.

[0078] For the input video, it is expressed as: X ∈ R T×H×W×C ; wherein, T: represents the number of time frames; H: represents the height, W: represents the width; C: represents the number of channels.

[0079] The process of C3D feature extraction can be written as: F = Conv3D(X, W) + b; the final output feature map is expressed as: F ∈ R T ′×H′×W′×C′ , where T′, H′, W′ represent the spatio-temporal dimensions after downsampling.

[0080] However, when identifying repetitive behaviors in industrial production line scenarios, this type of fine-grained features cannot be well extracted. To enhance the focusing ability of C3D on key spatio-temporal features, a SimAM module is inserted after each layer of 3D convolution.

[0081] The SimAM module evaluates the importance of the spatio-temporal positions of each pixel point in the feature map through a parameter-free energy function and dynamically adjusts the feature response. For the extracted feature map: F ∈ R T′×H′×W′×C′ Considering the neuron x t,h,w at any spatio-temporal position (t, h, w); its energy value e t,h,w is defined as the mean square of the difference between this neuron and other neurons in the local domain, as shown in Equation (2).

[0082]

[0083] Wherein,

[0084] Ω: a k t × k h × k w three-dimensional domain centered on (t, h, w), not including the center point itself; |Ω|: represents the total number of neurons in the domain, and the energy value e t,h,wThe smaller value indicates that the neuron is more similar to the surrounding spatio-temporal features and its importance is lower, so it should be suppressed. On the contrary, the larger value indicates that it is a locally significant feature and needs to be enhanced. Then, in order to convert the energy value e t,h,w into a differentiable attention weight, a scaling factor is introduced and a non-linear mapping is performed according to Equation (3) to obtain the attention weight w t,h,w .

[0085]

[0086] where λ is the scaling factor, set to 10 -4 , which is used to adjust the smoothness of the weight distribution. When λ is small, the weight distribution is smoother, which can avoid over-suppressing non-significant regions. According to Equation (3), when e t,h,w → +∞, w t,h,w → -1, and the feature at this position will be completely retained; when e t,h,w → 0, w t,h,w → 0.5, and the feature is partially suppressed. When e t,h,w → -∞, w t,h,w → 0, and the feature is completely suppressed.

[0087] Finally, the generated attention weight matrix W ∈ R T′×H′×W′ applies this attention weight to each channel through Equation (4) to enhance significant features and suppress redundant information, achieving spatio-temporal feature calibration across channels.

[0088] F′ = F ⊙ W (4)

[0089] The schematic diagram of adding the specific SimAM attention mechanism module is as shown in Figure 2 .

[0090] At the same time, it can only be applied to the ReLU activation function used by the baseline model. During the recognition process of repetitive behaviors in industrial scenarios, obvious limitations are exposed, and its formula is shown in Equation (5).

[0091] f(x) = max(0, x) (5)

[0092] When x ≥ 0, the output is x; when x < 0, the output is 0. Although ReLU can accelerate model convergence in most scenarios through the unilateral inhibition characteristic of identity mapping in the positive interval and hard saturation in the negative interval, the zero-gradient phenomenon in the negative gradient domain will cause serious "neuron death" problems. Analyzing from the mechanism level, when the input of the network layer is continuously in the negative interval, the gradient will completely disappear during the backpropagation process, resulting in the weights of the corresponding neurons being unable to be updated, and finally making the node permanently ineffective, leading to training stagnation.

[0093] Therefore, when identifying repetitive behaviors in industrial scenarios, this defect will significantly weaken the robustness of the model against interferences such as dim lighting and partial occlusion. Moreover, due to the attenuation of the effective number of parameters during the training process, the generalization ability of the model shows a systematic degradation. For this reason, we adopted the LeakyReLU activation function to transform the model, and its formula is shown in Equation (6).

[0094]

[0095] When x ≥ 0, the output is the same as the ReLU activation function. When x < 0, the output is ax and the gradient is a, where a is a very small positive number used to ensure that there is still a small gradient when x < 0.

[0096] Compared with the traditional ReLU, LeakyReLU introduces a linear leakage mechanism in the negative interval, maintains the identity mapping property in the forward propagation stage, ensures the integrity of feature expression, and in the negative compensation mechanism, increases the gradient value in this interval from 0 to the order of a, effectively avoiding the gradient chain break during backpropagation and maintaining the power of parameter update. At the same time, the suppression intensity of negative features can be controlled by adjusting a, achieving a balance between activation sparsity and feature retention, thereby improving the robustness and generalization ability of the model.

[0097] Finally, in the training process of the baseline C3D network, a fixed learning rate or a linear decay strategy is often adopted, and this coarse-grained learning rate regulation method has significant defects. In the initial stage of training, due to the lack of dynamic adaptability of the learning rate, the model is prone to falling into local optima or having insufficient convergence speed. In the later stage of training, too small a learning rate will lead to limited parameter update step sizes, causing frequent oscillations in the gradient direction, that is, the gradient oscillation problem.

[0098] This problem is particularly prominent in the task of identifying repetitive behaviors in industrial scenarios, specifically manifested as two contradictions: First, the C3D network has as many as 79 million parameters, and the fixed learning rate strategy results in low utilization of computing resources and is difficult to meet the real-time requirements. Second, the video data of repetitive behaviors has both spatial and temporal complexities, and there is a magnitude difference in the sensitivity of its low-level motion features (such as joint trajectories and temporal displacements) and high-level semantic features (such as action intentions and assembly stage divisions) to the learning rate, and a single learning rate is difficult to achieve the collaborative optimization of multi-level features.

[0099] Therefore, a two-stage adaptive optimization framework is proposed to achieve efficient training of the model through a dynamic differential learning rate strategy. In the first stage, an improved adaptive moment estimation (AdamW) optimizer is adopted, introducing a decoupled weight decay mechanism to effectively alleviate the overfitting tendency of the traditional Adam optimizer in gradient-sparse scenarios.

[0100] Specifically, first, initialize the network parameters θ, and calculate the first - moment estimate m0 (gradient mean) and the second - moment estimate v0 (gradient variance) respectively by the exponential moving average method. At the t - th iteration of the time step, calculate the gradient g of the current batch of data according to Equation (7). t ;

[0101]

[0102] Subsequently, update the first - moment estimate through Equation (8) and update the second - moment estimate through Equation (9), where β1 and β2 control the decay rate of historical gradient information respectively. To avoid the moment estimate bias in the initial stage of training, calculate the bias correction terms using Equations (10) and (11).

[0103] m t =β1·β t-1 +(1 - β1)·g t (8)

[0104]

[0105] Finally, update the parameter θ according to Equation (12). t , where ε is the numerical stability constant and λ decay is the decoupled weight decay coefficient.

[0106]

[0107] In the second stage, fuse the cosine annealing strategy. By periodically adjusting the learning rate, making it first decrease, then increase, and then decrease in the form of a cosine function, which helps the model to more finely adjust the parameters and regulate the convergence speed and model stability. The learning rate scheduling is carried out according to Equation (13).

[0108]

[0109] where η t is the learning rate at the current time step (or iteration number) t, η min : is the minimum value of the learning rate, usually close to 0, η max is the initial maximum value of the learning rate, T cur : is the number of steps that have been completed within a single period currently, T max is the total number of steps in a period;

[0110] When T cur =0 (at the start of the period), the learning rate is η max ; when T cur =T max (at the end of the period), the learning rate drops to η min, the learning rate in the intermediate process decays smoothly according to the cosine function. Through the above operations, it is possible to effectively avoid training oscillations caused by sudden changes in the learning rate and provide the possibility for fine-tuning the learning rate precisely in the later stage. The training strategy of the model is optimized through the cooperation of the Adam optimizer and cosine annealing.

[0111] The present invention proposes a method for identifying repetitive behaviors in industrial scenarios. First, a multi-camera perspective repetitive behavior dataset is defined. In terms of model innovation, the SimAM attention mechanism module is innovatively integrated into the baseline C3D model, the original ReLU activation function is improved, the model training optimizer and learning rate scheduling strategy are optimized, effectively strengthening the model's ability to extract fine-grained features, enhancing the stability and generalization of the model, and finally accelerating the model convergence speed and model recognition accuracy. It provides an efficient and accurate behavior recognition model for the repetitive behavior preparation of the industrial scenario production line, realizes real-time supervision of operators by using industrial cameras, and effectively expands the application boundary of behavior recognition technology in industrial scenarios.

[0112] The above embodiments are only used to illustrate the technical idea of the present invention, and the protection scope of the present invention cannot be limited thereby. Any changes made on the basis of the technical solution according to the technical idea proposed by the present invention shall fall within the protection scope of the present invention.

Claims

1. A method for identifying repetitive behaviors in an industrial production line scenario, characterized in that, It includes the following steps: Step 1: Shoot the overall process of the assembly operation behavior. Use an industrial camera array with orthogonal distribution to synchronously shoot the assembly line, and construct a multi-view video dataset by aligning timestamps. The industrial camera array includes top view, side view, and front view perspectives; Step 2: Construct an improved network based on the baseline C3D network. Incorporate the SimAM attention mechanism module after the activation function and before the pooling layer in each layer, perform feature extraction operations on the input video stream, and extract fine-grained features of each repetitive behavior to enhance the model's focusing ability on key spatio-temporal features; Step 3: According to the feature map extracted by the SimAM attention mechanism, to overcome the limitations of the ReLU activation function, change the ReLU activation function of the baseline model and replace it with the LeakyReLU activation function; Step 4: According to the training defects of the baseline model, improve the training optimizer and learning rate scheduling strategy, which speeds up the model training efficiency and provides a technical approach for real-time model update training. The optimization method of the improved training optimizer and learning rate scheduling strategy is specifically a two-stage adaptive optimization framework.

2. The repetitive behavior recognition method for industrial production line scenarios according to claim 1, characterized in that The multi-view video dataset is a self-defined repetitive behavior dataset established by the industrial camera array for shooting repetitive behaviors in the industrial production line scenario. Gaussian noise, Gaussian blur, and contrast enhancement operations need to be added to the self-defined repetitive behavior dataset for enhancement.

3. The repetitive behavior recognition method for industrial production line scenarios according to claim 1, wherein In the SimAM attention mechanism module, the SimAM attention mechanism specifically captures the subtle differences of repetitive behaviors accurately by dynamically evaluating the importance of each position in the feature map, enhancing the response of significant regions and suppressing redundant information, and directly quantifying the feature difference through the energy function, while avoiding increasing the model complexity; The SimAM attention mechanism module evaluates the importance of the spatio-temporal positions of each pixel point in the feature map through a parameter-free energy function and dynamically adjusts the feature response. For the neurons at any spatio-temporal position considered in the extracted feature map, its energy value is defined as the mean square of the difference between the neuron and other neurons in the local neighborhood.

4. The method for identifying repetitive behaviors in an industrial production line scenario according to claim 3, wherein The feature map is extracted as shown in the following formula: F ∈ R T′×H′×W′×C′ ; where, T′, H′, W′ represent the spatio-temporal dimensions of the feature map after downsampling; T′ represents the number of time frames of the feature map after sampling; H represents the height of the feature map after sampling, and W represents the width of the feature map after sampling; C represents the number of sampling channels.

5. The method for identifying repetitive behaviors in an industrial production line scenario according to claim 3, characterized in that, The formula for calculating the mean square of the difference is as follows: where, |Ω| represents the total number of neurons in the neighborhood; e t,h,w : represents the neuron energy value at any position, x t,h,w : represents the neurons of t, h, w at any position; x t′,h′,w′ : represents the neurons at feature extraction.

6. The repetitive behavior recognition method for industrial production line scenarios according to claim 3, wherein After that, in order to convert the energy value into a differentiable attention weight, a scaling factor is introduced, and non-linear mapping is performed according to the attention weight calculation formula to obtain the attention weight. Finally, the generated attention weight matrix applies this attention weight to each channel through Equation (4).

7. The method for identifying repetitive behaviors in an industrial production line scenario according to claim 6, wherein The attention weight calculation formula is as follows: where, w t,h,w : represents the attention weight; λ: Represents the scaling factor, set to 10 -4 ; The attention weight matrix is as follows: W ∈ R T′×H′×W′ .

8. A method for identifying repetitive behaviors in an industrial production line scenario according to claim 1, characterized in that, The ReLU activation function is as follows: f(x) = max(0, x); The LeakyReLU activation function is as follows: where, a represents the gradient, a > 0.

9. A method for identifying repetitive behaviors in an industrial production line scenario according to claim 1, characterized in that, The specific implementation of the two-stage adaptive optimization framework for efficient training of the model is achieved through a dynamic differential learning rate strategy; In the first stage, an improved adaptive moment estimation optimizer is adopted, and a decoupled weight decay mechanism is introduced. First, the network parameters are initialized, and the first-order moment estimation (gradient mean) and the second-order moment estimation (gradient variance) are calculated respectively by the exponential moving average method; during the time step iteration, the gradient of the current batch of data is calculated, and then the first-order moment estimation is updated, the second-order moment estimation is updated, and in order to use the bias correction term to avoid the moment estimation bias in the initial stage of training, finally the parameters are updated; In the second stage, the cosine annealing strategy is integrated. By periodically adjusting the learning rate, it first decreases and then increases, and then decreases again in the form of a cosine function, adjusts the parameters and regulates the convergence speed and model stability, and then performs the learning rate scheduling.

Citation Information

Patent Citations

  • Human body abnormal behavior recognition method and system based on mixed attention mechanism

    CN113516028A

  • Human body behavior recognition method based on multi-modal data and deep learning

    CN119600689A