Workshop worker operation safety detection method, device and equipment based on optical flow estimation and storage medium

By applying safety detection methods based on optical flow estimation in the workshop, real-time identification of workers' behavior and evaluation of safety, the problem of difficulty in identifying non-obvious safety hazards in the prior art is solved, and the workshop safety and worker safety are improved.

CN120220042APending Publication Date: 2025-06-27DONGFANG ELECTRIC CHENGDU INTELLIGENT TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510139197.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing workshop safety monitoring relies on passive monitoring and manual inspection, making it difficult to identify indistinct safety hazards and abnormal behaviors in a timely manner, and lacks intelligence and initiative.

Method used

The workshop workers' operation safety detection method based on optical flow estimation is adopted, and image behavior data is obtained through the workshop monitoring system, preprocessing and feature extraction is performed, and a fast behavior recognition model based on optical flow is established, and workers' behavior is identified in real time and hazard probability values ​​are output. If it is higher than the set threshold, emergency stop and alarm are performed.

Benefits of technology

It realizes rapid identification and safety assessment of workshop workers' operations, early warning of potential dangerous operations, reduces accidental injuries caused by workers in the workshop, and improves workshop safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220042A_ABST
    Figure CN120220042A_ABST
Patent Text Reader

Abstract

The invention discloses a workshop worker operation safety detection method and device based on optical flow estimation, and relates to the field of smart factory safety protection. Image behavior data of operation of each post personnel in a workshop is acquired and preprocessed; taking the preprocessed image behavior data as a training set, carrying out model training through deep learning and an optical flow estimation algorithm, and obtaining a trained rapid behavior recognition model; inputting image behavior data of each post person obtained in real time into the rapid behavior recognition model, then outputting a danger probability value of an operation behavior of each post person, judging whether the currently obtained behavior of each post person in the workshop has potential safety hazards and violation or not according to the danger probability value, and feeding back to the control center for processing; according to the invention, dangerous operation of workers can be warned in advance, and the equipment can be stopped at the first time when an accident occurs, so that accidental injury to the workers during workshop operation is avoided and reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of safety protection of intelligent factories, and particularly relates to a method and device for detecting the operation safety of workshop workers based on optical flow estimation. Background Art

[0002] In a production workshop, incorrect actions of workers may cause personal injuries to workers and also lead to abnormal shutdowns of the workshop. Therefore, protecting the lives and health of employees and preventing accidents and unexpected events are inevitable requirements for the continuous development of industrial production and the improvement of the automation level. The role of video surveillance technology in workshop safety management is becoming increasingly prominent. As an important tool for workshop safety, the video surveillance system can greatly improve workshop safety through technical means such as active target detection and behavior recognition, which has very practical significance.

[0003] Currently, workshop safety monitoring mainly relies on a combination of passive surveillance videos and manual inspections. This method can handle obvious safety problems and accidents, but it is difficult to immediately identify some hidden safety hazards or abnormal behaviors, which are easily overlooked. Especially, these systems only provide protection when a danger occurs. The premise of protecting the personal safety of workers is that workers strictly follow the standardized operation procedures for production operations. This method lacks intelligence and initiative. Summary of the Invention

[0004] To solve the deficiencies of the prior art, the purpose of the present invention is to provide a method, device, equipment and storage medium for detecting the operation safety of workshop workers based on optical flow estimation. By using the technologies of optical flow calculation and image recognition for behavior recognition of operators, the monitoring of dangerous operations of personnel in each position can be realized, potential safety risks can be detected in a timely manner, alarms can be given for dangerous behaviors and abnormal behaviors, the safety of the workshop can be improved, and the safety of employees in each position can be ensured.

[0005] The present invention is realized through the following technical solutions:

[0006] The first aspect of the present invention provides a method for detecting the operation safety of workshop workers based on optical flow estimation, including the following steps:

[0007] Step 1: Obtain the image behavior data of operators in each position and the time series information of the image behavior data within a period of time through the workshop monitoring system, and preprocess the image behavior data, where the image behavior data includes the RGB information of the operator optical flow information depth information where i represents the i-th type of position in the production workshop;

[0008] Step 2: After extracting the features of the preprocessed image behavior data in Step 1, use the image behavior data after feature extraction as the training set for model training to establish a fast behavior recognition model based on optical flow for identifying the standard action behaviors of workers in each position.

[0009] Step 3: Real-time obtain the image behavior data of workers in each position in the workshop and the time series information of the image behavior data through the workshop monitoring system. After preprocessing the real-time obtained image behavior data according to Step 1, input the corresponding feature information extracted according to Step 2 and the real-time obtained time series information into the fast behavior recognition model in Step 2. The fast behavior recognition model outputs a risk probability value and feeds the risk probability value back to the control center.

[0010] Step 4: If the risk probability value is higher than the set risk threshold V, the control center performs an emergency stop and issues an alarm, notifying the management personnel for emergency handling.

[0011] Further, the preprocessing of the image behavior data in Step 1 includes: First, encode the RGB information optical flow information depth information in chronological order, and then perform noise reduction processing, image stitching, and three-dimensional reconstruction processing on these image behavior data to obtain a three-dimensional model of the data after noise reduction.

[0012] Further, the specific steps for establishing the fast behavior recognition model based on optical flow in Step 2 include:

[0013] Step 2.1: Use a deep neural network composed of a convolutional neural network and an activation function to extract RGB image features from the RGB information in the image behavior data and establish an RGB recognition model; use convolutional layers and deconvolutional layers to extract optical flow features from the optical flow information in the image behavior data and establish an optical flow recognition model.

[0014] Step 2.2: After establishing the RGB recognition model and the optical flow recognition model in Step 2.1, match the two models through the time series information collected by the workshop monitoring system and the depth information in the scene to form the final fast behavior recognition model based on optical flow.

[0015] Further, in Step 2.1, the specific steps for extracting features from the RGB information include: Take the RGB information as the input of the convolutional layer, perform convolution using an odd number of different convolutional kernels, and then perform a max-pooling operation after obtaining the corresponding convolution feature data g(x, y), as shown in Equation 2:

[0016] g(x,y)=∑ k,lf(x - k, y - l)h(k, l), Equation 2;

[0017] In Equation 2: f(x, y) represents the corresponding RGB information, x and y respectively represent the quantity in the x - direction and the quantity in the y - direction of this point in the image, h(k, l) is the convolution kernel, k represents the width of the convolution kernel, and l represents the height of the convolution kernel.

[0018] Furthermore, in step 2.1, when extracting features from the optical flow information, by imposing constraints in the color, gradient, and velocity spaces, a constraint equation, Equation 3, is constructed:

[0019]

[0020] where γ and α are learnable parameters, is expressed as

[0021] shown in Equations 4 - 6:

[0022]

[0023] where represents the gradient direction, G1 represents the first frame, G2 represents the second frame, represents the quantity in the x - direction of this point in the image, represents the mapping of the optical flow vector from the image space to the real - number space, represents the displacement of this optical flow vector in the x - direction, respectively represent the displacement of this optical flow vector in the y - direction;

[0024] Add the HOG constraint between images, as shown in Equation 7:

[0025]

[0026] is the discriminant function, is the optical flow vector of the first frame, is the feature of the second - frame image, is the feature of the first - frame image;

[0027] Then, use the gradient - descent method to solve the system of equations of Equation 3, Equation 4, Equation 5, Equation 6, and Equation 7 to obtain the optical - flow feature map, and then use the same method of convolution and pooling operations as for the RGB information to extract the relevant optical - flow features.

[0028] Before performing the integral calculation on the above - mentioned constraint equation, introduce the empirical formula Actually, introduce a constant 10 -6 for empirical processing.

[0029] Further, the RGB information and optical flow information after feature extraction are used as the training set and trained through a deep learning algorithm. After the training is completed, the RGB feature model parameters obtained using the RGB information are represented by Equation 15, that is, the RGB recognition model:

[0030]

[0031] The optical flow feature model parameters obtained using the optical flow information are represented by Equation 16, that is, the optical flow recognition model:

[0032]

[0033] Where: n and m are both the number of image sequences, t is the time series information, expressed as the t-th moment, represents the n RGB feature model information of the i-th job at time t; represents the m optical flow feature model information of the i-th job at time t.

[0034] Further, in step 2.2, the specific steps are as follows:

[0035] The obtained RGB feature model parameters, optical flow feature model parameters and depth information with the same timestamp are fused in the channel dimension to obtain the final feature data as shown in Equation 17, and then a fully connected layer is performed, as shown in Equation 18:

[0036]

[0037] Where: η is a learnable parameter used to learn the weight of the depth information when extracting features; represents the image depth information with time series information; represents the original image feature information of the first job from 1 to t in the time series, represents the fused feature information of the first job from 1 to n in the time series after learning; weight parameters W 11 ~W nt and the bias parameter are all parameters to be learned; `

[0038] After passing through the fully connected layer, it is input into the softmax function to obtain the corresponding probability distribution, as shown in Equation 19:

[0039]

[0040] Softmax(z i ) represents the risk probability value, a i represents the fused feature information of the i-th job after learning, C represents C action categories, c = 1 represents a safe action, Denote Z c as an exponential distribution.

[0041] Furthermore, it also includes updating the image behavior data obtained in step 3 to the historical dataset and improving the fast behavior recognition model.

[0042] The second aspect of the present invention provides a safety detection device for workshop workers' operations based on optical flow estimation, including:

[0043] The first module is used to obtain the image behavior data of the operators at each position and the time series information of the image behavior data within a period of time through the workshop monitoring system, and preprocess the image behavior data, where the image behavior data includes the RGB information of the operators optical flow information depth information where i represents the i-th type of position in the production workshop;

[0044] The second module is used to extract features from the preprocessed image behavior data in the first module, and use the image behavior data after feature extraction as a training set for model training to establish a fast behavior recognition model based on optical flow for identifying the standard action behaviors of workers at each position;

[0045] The third module is used to obtain the image behavior data of the workers at each position in the workshop and the time series information of the image behavior data in real time through the workshop monitoring system, preprocess the real-time obtained image behavior data and then extract features, and then input the corresponding feature information extracted by the feature extraction and the real-time obtained time series information into the fast behavior recognition model in the second module. The fast behavior recognition model outputs a risk probability value and feeds the risk probability value back to the control center;

[0046] The fourth module is that if the risk probability value is higher than the set risk threshold V, the control center will perform an emergency stop, alarm, and notify the management personnel for emergency handling.

[0047] The third aspect of the present invention provides a computer device, including a processor, an input device, an output device, and a memory. The processor, input device, output device, and memory are interconnected. Among them, the memory is used to store a computer program, and the computer program includes program instructions. The processor is configured to call the program instructions to execute some or all of the steps described in the first aspect of the present invention.

[0048] The fourth aspect of the present invention provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute some or all of the steps described in the first aspect of the present invention.

[0049] The beneficial effects of the present invention are as follows:

[0050] The present invention uses optical flow estimation to quickly identify the operation behaviors of personnel at each post in the workshop, and judge the standardization and safety of the operations of the post personnel. On the one hand, it can give early warnings of dangerous operations of the post personnel. On the other hand, when an accident occurs, the equipment can be stopped immediately, thereby avoiding and reducing the accidental injuries suffered by on-the-job personnel during workshop operations, and ensuring the safety of the workshop and the on-the-job personnel. Description of the Drawings

[0051] Figure 1 It is a diagram showing the camera settings in the workshop. The workshop is photographed at all angles by multiple cameras to obtain the action behaviors of the workers.

[0052] Figure 2 It is a diagram showing the camera settings in a certain workshop. The workshop is photographed without dead angles in all directions by multiple cameras to obtain the action behaviors of the workers.

[0053] Figure 3 It is a diagram showing the behavior judgment method based on optical flow estimation according to an exemplary embodiment of the present invention.

[0054] Figure 4 It is a flowchart showing the use of the method and device for detecting the operation safety of workshop workers based on optical flow estimation according to an exemplary embodiment of the present invention. Detailed Embodiments

[0055] The technical solution of the present invention will be further elaborated in detail below in combination with specific embodiments and the accompanying drawings of the specification. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work fall within the scope of protection of the present invention.

[0056] Embodiment 1

[0057] As Figure 2 shown, the first aspect of the present invention provides a method for detecting the operation safety of workshop workers based on optical flow estimation, including the following steps:

[0058] Step 1: Obtain the image behavior data and the time series information of the image behavior data of the operators at each post within a period of time through the workshop monitoring system, and preprocess the image behavior data, where the image behavior data includes the RGB information of the operator optical flow information depth information where i is the number of the job type;

[0059] Step 2: After extracting the features from the preprocessed image behavior data in Step 1, use the image behavior data after feature extraction as the training set for model training to establish a fast behavior recognition model based on optical flow for identifying the standard action behaviors of workers at each position.

[0060] Step 3: Real-time obtain the image behavior data of workers at each position in the workshop through the workshop monitoring system, as well as the time series information of the image behavior data. After preprocessing the real-time obtained image behavior data according to Step 1, input the corresponding feature information extracted according to Step 2 and the real-time obtained time series information into the fast behavior recognition model in Step 2. The fast behavior recognition model outputs a risk probability value and feeds the risk probability value back to the control center.

[0061] Step 4: If the risk probability value is higher than the set risk threshold V, the control center performs an emergency stop and issues an alarm, notifying the management personnel for emergency handling.

[0062] Embodiment 2

[0063] This embodiment further elaborates and supplements the implementation manner of the present invention on the basis of Embodiment 1.

[0064] The preprocessing of the image behavior data in Step 1 includes: First, encode the RGB information optical flow information depth information in chronological order, and then perform noise reduction processing, image stitching, and three-dimensional reconstruction processing on these image behavior data to obtain a three-dimensional model of the data after noise reduction.

[0065] Embodiment 3

[0066] This embodiment further elaborates and supplements the implementation manner of the present invention on the basis of Embodiment 1 or Embodiment 2.

[0067] The specific steps for establishing the fast behavior recognition model based on optical flow in Step 2 include:

[0068] Step 2.1: Use a deep neural network composed of a convolutional neural network and an activation function to extract RGB image features from the RGB information in the image behavior data and establish an RGB recognition model; use a convolutional layer and a deconvolutional layer to extract optical flow features from the optical flow information in the image behavior data and establish an optical flow recognition model.

[0069] Step 2.2: After establishing the RGB recognition model and the optical flow recognition model in Step 2.1, match the two models through the time series information collected by the workshop monitoring system and the depth information in the scene to form the final fast behavior recognition model based on optical flow.

[0070] In step 2.1, the specific steps for feature extraction of RGB information include: taking the RGB information as the input of the convolutional layer, performing convolution using an odd number of different convolutional kernels, and then performing a max pooling operation after obtaining the corresponding convolution feature data g(x,y), as shown in Equation 2:

[0071] g(x,y) = ∑ k,l f(x - k, y - l)h(k,l) Equation 2;

[0072] In Equation 2: f(x,y) is the corresponding RGB information, x and y respectively represent the quantity in the x - direction and the quantity in the y - direction of this point in the image, h(k,l) is the convolutional kernel, k represents the width of the convolutional kernel, and l represents the height of the convolutional kernel.

[0073] Furthermore, in step 2.1, when performing feature extraction on the optical flow information, by imposing constraints in the color, gradient, and velocity spaces, a constraint equation 3 is constructed:

[0074]

[0075] where γ and α are learnable parameters, is expressed as

[0076] shown in Equations 4 - 6:

[0077]

[0078] where represents the gradient direction, G1 represents the first frame, G2 represents the second frame, represents the quantity in the x - direction of this point in the image, represents the mapping of the optical flow vector from the image space to the real - number space, represents the displacement of this optical flow vector in the x - direction, respectively represent the displacement of this optical flow vector in the y - direction;

[0079] Add the HOG constraint between images, as shown in Equation 7:

[0080]

[0081] is the discriminant function, is the optical flow vector of the first frame, is the feature of the second - frame image, is the feature of the first - frame image;

[0082] Then, the gradient descent method is used to solve the system of equations in Equations (3), (4), (5), (6), and (7) to obtain the optical flow feature map, and then the relevant optical flow features are extracted using the same convolution and pooling operations as the RGB information.

[0083] Before integrating and calculating the above constraint equations, an empirical formula is introduced Actually, a constant 10 is introduced -6 For empirical processing.

[0084] In this embodiment, the general gradient descent method is used as the optimization method to calculate the minimization of the loss function. For example, other methods applicable to deep learning optimization and minimization of the loss function can also be used, such as batch gradient descent method, stochastic gradient descent method, Newton method, quasi-Newton method, conjugate gradient method, and so on.

[0085] Furthermore, the RGB information and optical flow information after feature extraction are used as the training set and trained through a deep learning algorithm. After the training is completed, the RGB feature model parameters obtained using the RGB information are represented by Equation (15), that is, the RGB recognition model:

[0086]

[0087] The optical flow feature model parameters obtained using the optical flow information are represented by Equation (16), that is, the optical flow recognition model:

[0088]

[0089] Where: n and m are both the number of image sequences, t is the time series information, representing the t-th moment, represents the n RGB feature model information of the i-th position at time t; represents the m optical flow feature model information of the i-th position at time t.

[0090] Furthermore, in Step 2.2, the specific steps are as follows:

[0091] The obtained RGB feature model parameters, optical flow feature model parameters, and depth information with the same timestamp are fused in the channel dimension to obtain the final feature data as shown in Equation (17), and then a fully connected layer is performed, as shown in Equation (18):

[0092]

[0093]

[0094] Where: η is a learnable parameter used to learn the weight size of the depth information when extracting features; represents the image depth information with time series information; It represents the original image feature information of the first type of post from 1 to t in the time series. It represents the fused feature information of the first type of post with the time series from 1 to n after learning; the weight parameter W 11 ~W nt and the bias parameter are all parameters to be learned; `

[0095] After passing through the fully connected layer, it is input into the softmax function to obtain the corresponding probability distribution, as shown in Equation 19:

[0096]

[0097] Softmax(z i ) represents the risk probability value, a i represents the fused feature information of the i-th type of post after learning, C represents C action categories, that is, the total number of categories, which are divided into safe actions, dangerous actions, and actions that cannot be judged. c = 1 represents a safe action. represents Z c 's exponential distribution.

[0098] Furthermore, it also includes updating the image behavior data obtained in step 3 to the historical dataset and improving the fast behavior recognition model.

[0099] Example 4

[0100] This example further elaborates and supplements the implementation manner of the present invention on the basis of Example 1, Example 2, or Example 3.

[0101] The monitoring system in the workshop includes cameras, which are usually placed at the entrance and exit of the workshop to observe the entry and exit of personnel, and are also centrally placed around the production line, processing area, assembly line, and other main working areas to monitor the production process, especially the operation behavior of personnel. As Figure 1 shown, depth cameras required are installed at multiple angles above the workshop, and multiple cameras are used jointly to achieve on-line monitoring of multiple angles and all time periods in the workshop.

[0102] The frame rate of general monitoring videos is 25 frames per second - 30 frames per second. However, when performing image processing, too high a frame rate will increase the burden on the computer and instead result in errors such as slow recognition and recognition jams. Therefore, before obtaining the image behavior data, the video data is first preprocessed using the video frame extraction method, such as

[0103] shown in Equation 1:

[0104]

[0105] Where represents the image data at the j-th frame. T is an adjustable parameter. When T = 1, the original frame rate is maintained. The larger T is, the greater the frame extraction and the smaller the required computational amount. is the output data after key frame processing. Among them, k is the timestamp used to mark the time of the image, and this data will be further processed. The obtained image information is decomposed into three types of key information. The RGB information is represented by , the optical flow information , the depth information , and then preprocessing is performed; for the RGB information optical flow information depth information are encoded in chronological order, and then these image behavior data are subjected to noise reduction processing, image stitching, and three-dimensional reconstruction processing to obtain a three-dimensional model of the noise-reduced data. Among them, i represents the i-th position in the production workshop, and k represents the time series information.

[0106] The preprocessed RGB information and the optical flow information are further processed. Since image preprocessing has been performed before, the RGB information is adjusted to 960*960 as the input of the convolutional layer and processed using three different convolutional kernels of 9*9, 7*7, and 3*3 respectively. As shown in Equation 2, f(x, y) is the corresponding RGB information, x and y respectively represent the quantities in the x direction and y direction of this point in the image, h(x, l) is the convolutional kernel, k represents the width of the convolutional kernel, and l represents the height of the convolutional kernel.

[0107] g(x, y) = ∑ k,l f(x - k, y - l)h(k, l) Equation 2;

[0108] That is, the input data is convolved using the convolution h(k, l), and then the corresponding convolved feature data g(x, y) is obtained, and then max pooling operation is performed after convolution. For example, the operations taken respectively after input are to convolve using a 9*9 convolutional kernel, pool the convolved result using a max pooling layer, then convolve using a 7*7 convolutional kernel, and then pool using max pooling again. After two convolutions and poolings, two 3*3 convolutional kernels are respectively used for convolution, and finally, after passing through a pooling layer and a fully connected layer, the final RGB image features are obtained.

[0109] When extracting features from the optical flow information, constraints are imposed in the color, gradient, and velocity spaces to construct the constraint equation 3:

[0110]

[0111] where γ and α are learnable parameters, which is expressed as

[0112] shown in Equation 4 - Equation 6:

[0113]

[0114] where represents the gradient direction, G1 represents the first frame, and G2 represents the second frame. represents the quantity in the x - direction of this point in the image, represents the mapping of the optical flow vector from the image space to the real - number space, represents the displacement of this optical flow vector in the x - direction, respectively represent the displacement of this optical flow vector in the y - direction;

[0115] Add the HOG constraint between images, as shown in Equation 7:

[0116]

[0117] is the discriminant function, is the optical flow vector of the first frame, is the feature of the second - frame image, is the feature of the first - frame image, representing the constraint between frames G1 images.

[0118] Then, use the gradient - descent method to solve the system of equations of Equation 3, Equation 4, Equation 5, Equation 6, and Equation 7 to obtain the optical - flow feature map, and then use the same convolution and pooling operation methods as the RGB information to extract the relevant optical - flow features.

[0119] Before performing the integral calculation on the above - mentioned constraint equation, introduce the empirical formula which is actually introducing a constant 10 -6 for empirical processing.

[0120] In this embodiment, the functions of Equation 3 - Equation 7 are to predict the data of the next frame from the data of one frame.

[0121] Take the RGB information and optical - flow information after feature extraction as the training set, and train through a deep - learning algorithm. During training, first randomly initialize the neuron weights W and the bias parameter b, then calculate the partial derivative of the cost function with respect to each neuron weight W from back to front, and then update each weight by subtracting the product of a learning - rate operator and the partial derivative from the preset neuron weight.

[0122] For example, after the input set m of the training set is learned by the model determined by the parameter set θ, the output set is obtained as shown in Equation 8:

[0123] h θ (m) = σm + b Equation 8; σ is a learnable parameter that is continuously and automatically adjusted by the model.

[0124] According to the least squares rule, the sum of squared differences between the output set and the true value set is used as the cost function, as

[0125] shown in Equation 9:

[0126]

[0127] where m (p) is the p-th vector in the input set m, l (p) is the p-th vector in the true value set, and q is the dimension of the set vector;

[0128] Use the learning method shown in Equation 10 to learn J(θ) in Equation 9, where α is the set learning rate;

[0129]

[0130] Take the partial derivative of Equation 9 and expand it, and substitute Equation (10) into it, as shown in Equation 11:

[0131]

[0132] Obtain the final training update strategy.

[0133] θ j is a parameter that is automatically adjusted, and its purpose is to make closest to 0.

[0134] Use the RGB information and optical flow information after feature extraction as the training set, train through a deep learning algorithm, update using the training update strategy shown in Equation 11, and according to the above formula, obtain the final cost function formed by the neuron weights W and the bias parameter b as shown in Equation 12, and solve it:

[0135]

[0136] Solve Equation 12 to obtain Equations 13 and 14:

[0137]

[0138] where, in Equation 12, z represents the number of vectors in the input set m, λ is a learnable parameter, and n l-1 represents The number of vectors in the output set, i.e., the superscript in the upper right corner: f, denotes starting from to add them up; s l denotes the number of vertical vectors in denotes starting from to add them up; s s+1 denotes the number of horizontal vectors in denotes starting from to add them up; J(W, b) is the cost function of the neuron weights W and the bias parameter b, denotes the obtained neuron weights, denotes the obtained bias parameter.

[0139] In the above formula, the roles of i and j are used for counting. Assuming the size of a two-dimensional matrix is x * y (i.e., the number of horizontal vectors is x, and the number of vertical vectors is y, then i and j respectively represent the j-th number of the horizontal vector and the i-th number of the vertical vector. Through i and j, the number at the corresponding position in this two-dimensional matrix can be uniquely determined).

[0140] After completing the training, the RGB feature model parameters obtained using RGB information are represented by Equation 15, i.e., the RGB recognition model:

[0141]

[0142] The optical flow feature model parameters obtained using optical flow information are represented by Equation 16, i.e., the optical flow recognition model:

[0143]

[0144] Among them: n and m are both the number of image sequences, t is the time series information, representing the t-th moment, represents the n RGB feature model information of the i-th position at time t; represents the m optical flow feature model information of the i-th position at time t.

[0145] The obtained RGB feature model parameters, optical flow feature model parameters and depth information with the same timestamp are fused in the channel dimension to obtain the final feature data as shown in Equation 17, and then a fully connected layer is performed, as shown in Equation 18:

[0146]

[0147] Where: η is a learnable parameter used to learn the weight of depth information when extracting features; represents the depth information of the image with time series information; represents the original image feature information of the time series from 1 to t on the first type of post, represents the fused feature information of the time series from 1 to n on the first type of post after learning; weight parameter W 11 ~W nt and the bias parameter are all parameters to be learned; `

[0148] After passing through the fully connected layer, it is input into the softmax function to obtain the corresponding probability distribution, as shown in Equation 19:

[0149]

[0150] Softmax(z i ) represents the hazard probability value, a i represents the fused feature information of the i-th type of post after learning, C represents C action categories, c = 1 represents a safe action, represents Z c 's exponential distribution.

[0151] As Figure 4 shown, the specific process of using the optical flow estimation algorithm for operation safety detection in the workshop. This process description is only an illustrative example, and its deployment method is also applicable to other workshops. As Figure 4 shown, after the detection device is started, the camera is first checked. If the camera is not started, the camera is started and ensured to be in the recording state. If it cannot be started normally, the camera abnormality is immediately reported; after startup, it can be observed whether there are workers in the workshop. If there are no workers in the workshop, the recording state can be maintained. If there are workers in the workshop, the video monitoring is immediately started, that is, Figure 3 the process shown, to monitor the workshop in real time. If the video monitoring starts abnormally, an abnormality is reported and the technician is called to handle it; after the video monitoring is started, it is detected whether the warning module is turned on while performing the detection. If the warning module is normally turned on, an abnormal situation is also reported; after the warning module is confirmed to be turned on, finally, the emergency situation automatic processing module of the control center is started for inspection. When all the above startup steps are completed, the software and hardware system for operation safety detection using the optical flow estimation algorithm can start to work normally and continuously monitor the workshop.

[0152] Embodiment 5

[0153] This embodiment further elaborates and supplements the implementation manner of the present invention on the basis of Embodiment 1, Embodiment 2, Embodiment 3 or Embodiment 4.

[0154] In a second aspect of the present invention, a safety detection device for workshop workers' operations based on optical flow estimation is provided, including:

[0155] The first module is used to obtain the image behavior data and the time series information of the image behavior data of workers at each post within a period of time through the workshop monitoring system, and preprocess the image behavior data, where the image behavior data includes the RGB information, optical flow information, and depth information of the operator;

[0156] The second module is used to extract features from the preprocessed image behavior data in the first module, and then use the image behavior data after feature extraction as a training set to train a model to establish a fast behavior recognition model based on optical flow for identifying the standard action behaviors of workers at each post;

[0157] The third module is used to obtain the image behavior data of workers at each post in the workshop and the time series information of the image behavior data in real time through the workshop monitoring system, preprocess the real-time obtained image behavior data and then extract features, and then input the corresponding feature information extracted by the feature extraction and the real-time obtained time series information into the fast behavior recognition model in the second module. The fast behavior recognition model outputs a risk probability value, and feeds back the risk probability value to the control center;

[0158] The fourth module is that if the risk probability value is higher than the set risk threshold V, the control center will perform an emergency stop and alarm, and notify the management personnel for emergency handling.

[0159] Embodiment 6

[0160] This embodiment further elaborates and supplements the implementation manner of the present invention on the basis of Embodiment 1, Embodiment 2, Embodiment 3, Embodiment 4 or Embodiment 5.

[0161] In a third aspect of the present invention, a computer device is provided, including a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. Among them, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to call the program instructions to execute some or all of the steps described in the first aspect of the present invention.

[0162] In this embodiment, the processor may be a Central Processing Unit (CPU). The processor may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc., or a combination of the above types of chips.

[0163] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs, non-transitory computer-executable programs, and units, such as the corresponding program units in the above method embodiments of the present invention. By running the non-transitory software programs, instructions, and modules stored in the memory, the processor can execute various functional applications and work data processing of the processor, that is, implement the methods in the above method embodiments.

[0164] The memory may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created by the processor, etc. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.

[0165] Embodiment 7

[0166] This embodiment further elaborates and supplements the implementation manner of the present invention on the basis of Embodiment 1, Embodiment 2, Embodiment 3, Embodiment 4, Embodiment 5, or Embodiment 6.

[0167] The fourth aspect of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute some or all of the steps described in the first aspect of the present invention.

Claims

1. A workshop worker operation safety detection method based on optical flow estimation, characterized by: The following steps are involved: Step 1: Obtain the image behavior data and time series information of operators at various positions over a period of time through the workshop monitoring system, and pre-process the image behavior data, where the image behavior data includes the RGB information of the operators. Optical flow information Depth Information Where i is the i-th position in the production workshop; Step 2: After extracting features from the image behavior data preprocessed in step 1, the image behavior data after feature extraction is used as a training set for model training, and a fast behavior recognition model based on optical flow is established to identify the standardized action behaviors of workers in various positions; Step 3: The image behavior data of workers at various positions in the workshop and the time series information of the image behavior data are obtained in real time through the workshop monitoring system. After preprocessing the image behavior data obtained in real time according to step 1, the corresponding feature information is extracted according to step 2, and the real-time time series information is input into the rapid behavior recognition model in step 2. The rapid behavior recognition model outputs the danger probability value, and the danger probability value is fed back to the control center; Step 4: If the danger probability value is higher than the set danger threshold V, the control center will perform an emergency stop, alarm, and notify the management personnel to take emergency measures.

2. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 1, characterized in that: The preprocessing of the image behavior data in step 1 includes: first, the RGB information Optical flow information Depth Information The images are encoded in chronological order, and then subjected to denoising, image stitching and three-dimensional reconstruction to obtain a three-dimensional model of the denoised data.

3. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 1, characterized in that: The specific steps of establishing a fast action recognition model based on optical flow in step 2 include: Step 2.1, using a deep neural network composed of a convolutional neural network and an activation function to extract RGB image features from the RGB information in the image behavior data, and establishing an RGB recognition model; using a convolutional layer and a deconvolution layer to extract optical flow features from the optical flow information in the image behavior data, and establishing an optical flow recognition model; After the RGB recognition model and the optical flow recognition model are established in step 2.2 and step 2.1, the two models are matched through the time series information collected by the workshop monitoring system and the depth information in the scene to form the final optical flow-based fast behavior recognition model.

4. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 3, characterized in that: In step 2.1, the specific steps of extracting features from RGB information include: using the RGB information as the input of the convolution layer, performing convolution using an odd number of different convolution kernels, and then performing a maximum pooling operation after obtaining the corresponding post-convolution feature data g(x, y), as shown in Formula 2: g(x,y) = ∑ k,l f(x - k, y - l)h(k, l) Equation 2; In Formula 2, f(x, y) is the corresponding RGB information, x and y represent the amount in the x direction and y direction of the point in the image, respectively, h(k, l) is the convolution kernel, k represents the width of the convolution kernel, and l represents the height of the convolution kernel.

5. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 4, characterized in that: In step 2.1, when extracting features from the optical flow information, constraint equation 3 is constructed by constraining the color, gradient, and velocity spaces: Among them, γ and α are learnable parameters. The expression is shown in Formula 4-6: in represents the gradient direction, G1 represents the first frame, G2 represents the second frame, Represents the amount in the x direction of the point in the image, represents the mapping of optical flow vector from image space to real number space, express The displacement of this optical flow vector in the x direction, Respectively The displacement of this optical flow vector in the y direction; Add HOG constraints between images, as shown in Equation 7: is the discriminant function, is the optical flow vector of the first frame, is the feature of the second frame image, is the feature of the first frame image; Then, the gradient descent method is used to solve the equations of Equation 3, Equation 4, Equation 5, Equation 6, and Equation 7 to obtain the optical flow feature map, and then the same convolution and pooling operations as the RGB information are used to extract the relevant optical flow features.

6. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 3, characterized in that: The RGB information and optical flow information after feature extraction are used as training sets and trained through a deep learning algorithm. After the training is completed, the RGB information is used to obtain the RGB feature model parameters expressed by formula 15, that is, the RGB recognition model: The optical flow feature model parameters obtained using the optical flow information are expressed as formula 16, that is, the optical flow recognition model: Where: n and m are the number of image sequences, t is the time series information, expressed as time t, Represents the n RGB feature model information of the i-th job at time t; Represents the m optical flow feature model information of the i-th position at time t.

7. The method for detecting workshop worker operation safety based on optical flow estimation according to claim 6, characterized in that: In step 2.2, the specific steps are: The obtained RGB feature model parameters, optical flow feature model parameters and depth information of the same timestamp are fused in the channel dimension to obtain the final feature data as shown in Formula 17, and then the full connection layer is performed as shown in Formula 18: Where: η is a learnable parameter used to learn the weight of depth information when extracting features; Represents image depth information with time series information; Represents the original image feature information from 1 to t in the time series of the first position, It is represented as the fusion feature information of the time series from 1 to n in the first position after learning; the weight parameter W 11 ~W nt and bias parameters These are all parameters that need to be learned;` After passing through the fully connected layer, it is input into the softmax function to obtain the corresponding probability distribution, as shown in Formula 19: Softmax(z i ) represents the risk probability value, a i It is represented as the fusion feature information of the i-th position after learning, C represents C action categories, c = 1 represents a safe action, Represents Z c The exponential distribution of .

8. A workshop worker operation safety detection device based on optical flow estimation, characterized in that: include: The first module is used to obtain the image behavior data and time series information of operators at various positions over a period of time through the workshop monitoring system, and pre-process the image behavior data, where the image behavior data includes the RGB information of the operators. Optical flow information Depth Information Where i represents the i-th position in the production workshop; The second module is used to extract features from the image behavior data preprocessed in the first module, and use the extracted image behavior data as a training set for model training, so as to establish a fast behavior recognition model based on optical flow, which is used to identify the standardized action behaviors of workers in various positions; The third module is used to obtain the image behavior data of workers at various positions in the workshop and the time series information of the image behavior data in real time through the workshop monitoring system, perform feature extraction after preprocessing the real-time image behavior data, and then input the corresponding feature information extracted from the features and the real-time time series information into the fast behavior recognition model in the second module, and the fast behavior recognition model outputs the danger probability value, and feeds the danger probability value back to the control center; In the fourth module, if the danger probability value is higher than the set danger threshold V, the control center will control the automatic processing module to perform emergency stop and the early warning module to alarm, and notify the management personnel to carry out emergency processing.

9. A computer device, characterized in that: The method comprises a processor, an input device, an output device and a memory, wherein the processor, the input device, the output device and the memory are interconnected, wherein the memory is used to store a computer program, the computer program comprises program instructions, and the processor is configured to call the program instructions to execute the method according to any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to any one of claims 1 to 8.