A method and system for early detection of power violation jobs through intent recognition
By detecting target sources and key nodes of the human skeleton in power operations, and using a spatiotemporal convolutional network of the skeleton to predict the interaction between the operator and the objects, the problem of identifying and predicting violations in power operations is solved, and safety early warning for power operations is achieved.
Patent Information
- Application Number
- CN202211317081.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-26
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2042-10-26
AI Technical Summary
Existing technologies struggle to identify and predict whether operator-object interactions during power operations are in violation of regulations, especially when determining whether power operations are in violation of regulations with a high degree of semantic understanding.
By detecting points of interest features in surveillance image data, suspected target sources are identified. Combining key nodes of the human skeleton and key nodes of the hand skeleton, a spatiotemporal convolutional network of the skeleton is used to predict the interaction relationship between the operator and the object, and a prediction model is built to determine and predict actions.
It enables early warning of violations of power operation rules several seconds in advance, ensuring the safety of power operations.
Smart Images

Figure CN115830703B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a method and system for judging power illegal operation in advance through intention recognition. BACKGROUND
[0002] The current human behavior detection is mostly used in engineering to detect simple warning events, such as standing, lying, sitting and falling, etc. These detections can simply realize the identification of the current behavior state of workers, but cannot solve the tasks under higher semantic understanding, especially the tasks that need to interact with objects, such as judging whether the power operation is illegal. In this case, the interaction between the operator and the operating object needs to be understood, and the action of the operator needs to be understood and predicted in advance. Therefore, on the basis of only analyzing the human, the analysis of the interaction with the object and the prediction of the human behavior need to be added. SUMMARY
[0003] The present application provides a method and system for judging power illegal operation in advance through intention recognition, which defines, detects and predicts the whole power operation process sequence. Compared with the simple detection system, the method and system are more perfect, practical and meet the task requirements.
[0004] The present application is realized by the following technical solutions:
[0005] A method for judging power illegal operation in advance through intention recognition, comprising the following steps:
[0006] According to the work task, the first interest point feature group in the monitoring image data is detected to determine the range of the carrier carrying at least one suspected target source, so as to preliminarily exclude non-target sources. Through this step, the determination range of the target source can be reduced, and the determination result when the interactive object is a non-target source outside the carrier range can be quickly obtained, so as to avoid too many objects and complicate the determination process.
[0007] The position of the suspected target source and the second interest point feature group are detected, and the second interest point feature groups of a plurality of suspected target sources are compared to distinguish a plurality of suspected target sources with similar or different shapes, and the distinguished suspected target sources are marked. In this way, the target sources in the carrier range can be better distinguished, and the marking error of objects with similar shapes can be avoided.
[0008] The human body skeleton key node and the hand skeleton key node in the monitoring image data are detected, and the hand data confirmed according to the hand skeleton key node is obtained. The position of the worker is obtained according to the human body skeleton key node, and the position relationship data set of the hand skeleton key node and the suspected target source is obtained according to the time dimension based on the skeleton space-time convolution network. The hand information and the suspected target source are one-to-one corresponding, so that the subsequent prediction is more accurate.
[0009] The position relationship data set and the corresponding time dimension unit time are taken as the input training parameters of the neural network model to train the prediction model.
[0010] The current action is determined, and the next action is predicted and determined based on the prediction model.
[0011] As an optimization, the hand skeleton key node includes a wrist skeleton key node composed of wrist joints, a palm skeleton key node composed of a palm center, and a finger skeleton key node composed of several finger joints and finger tips. The hand data includes wrist data, palm data, and finger data. The wrist data includes first positioning data of the wrist skeleton key node. The palm data includes the opening degree of the palm confirmed by the finger skeleton key node, and the angle difference between the maximum angle and the minimum angle between the palm center and the upper surface of the outline of a certain suspected target source. The finger data includes the distance change value of the index finger tip to the thumb tip of the same hand, and the positioning data of each finger joint. Each finger joint includes a metacarpophalangeal joint and a finger interphalangeal joint.
[0012] Since the worker needs to bend down to take the corresponding tool placed on the upper surface of the carrier for work in the power operation, the palm height of the worker must be higher than the upper surface of the tool (suspected target source). When the palm is close to a certain suspected target source, the palm will open, and the palm is towards the target source. Therefore, the angle difference between the maximum angle and the minimum angle between the palm center and the upper surface of the outline of the suspected target source must be smaller than the angle difference between the maximum angle and the minimum angle between the upper surfaces of the outlines of other suspected target sources. The distance change value of the index finger tip to the thumb tip of the same hand will increase in the unit time when the palm opens.
[0013] As an optimization, the specific way to determine the position relationship data set is:
[0014] The closest distance of the wrist skeleton key node to a certain suspected target source in each unit time is determined according to the time dimension and the first positioning data, so as to detect the motion trajectory of the wrist.
[0015] The palm opening degree value and the angle difference in each unit time are determined according to the time dimension.
[0016] determining a distance change value of the fingertip of the index finger to the suspected target source in each unit time according to the time dimension;
[0017] combining the closest distance, the palm opening value, the angle difference, the distance change value in each unit time and the corresponding meaning target source into a set of position relationship data sets respectively, the position relationship data set is represented as: {A i,j , B i,j , C i,j , D i,j}, wherein i represents the target source, j represents the unit time in the time dimension, A represents the closest distance, B represents the palm opening value, C represents the angle difference, and D represents the distance change value.
[0018] As an optimization, the palm opening value and the closest distance satisfy: when the closest distance is less than a first threshold value, the palm opening value satisfies a first condition, the angle difference and the closest distance satisfy: when the closest distance is less than a second threshold value, the angle difference satisfies a second condition, the distance change value and the closest distance satisfy: when the closest distance is less than a third threshold value, the distance change value satisfies a third condition, and when the position relationship data set satisfies the conditions at the same time, the suspected target source is determined as a target source.
[0019] As an optimization, the specific way of determining the position relationship data set is:
[0020] determining the closest distance of the wrist skeleton key node to a suspected target source in each unit time according to the time dimension, so as to detect the motion trajectory of the wrist;
[0021] determining the palm opening value in each unit time according to the time dimension;
[0022] determining the connection trajectory of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger in each unit time according to the time dimension;
[0023] determining the relative position of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger and the wrist joint node according to the time dimension;
[0024] combining the closest distance, the connection trajectory of each finger joint node, the relative position of each finger joint node and the wrist joint node in each unit time and the corresponding suspected target source into a set of position relationship data sets respectively, the position relationship data set is represented as: {A i,j , X i,j , Y i,j,k , Z i,j,k,m,n} wherein i represents a suspected target source, j represents a unit time in the time dimension, A represents the closest distance, k represents the finger type, m represents the wrist joint node, n represents the finger joint node, X represents the palm opening value, Y represents the connection trajectory, and Z represents the relative position of each finger joint node to the wrist joint node.
[0025] As an optimization, the palm opening value, the connection trajectory and the closest distance satisfy the following conditions: when the closest distance is less than a fourth threshold value, the palm opening value satisfies a first condition, the connection trajectory of the thumb finger joint node is curved with the connection trajectory of the index finger or middle finger joint node, the angle of the interphalangeal joint of the thumb is between 120° and 170°, the angle at the proximal interphalangeal joint of the index finger or middle finger is between 90° and 160°, the angle at the interphalangeal joint of the index finger or middle finger is greater than 130°, and the curved radii of the connection trajectory of the thumb finger joint node and the connection trajectory of the index finger or middle finger joint node are both directed towards a suspected target source, it is determined that the suspected target source is a target source.
[0026] As an optimization, the skeleton spatio-temporal convolutional network is specifically:
[0027]
[0028] wherein f out (v tp ) represents the motion prediction output of the skeleton key node p at time t; B(v tp ) represents the set of all associated nodes of the key node p at time t, Z tp (v tq ) represents the normalized parameter for the key node p and its associated node q at time t, f in (v tq ) represents the input feature of the key node q at time t, w(v tp , v tq ) represents the path weight of the key node p to its associated node q at time t.
[0029] As an optimization, according to the task planning design, a predetermined task flow label is formulated, which includes setting a current work task and setting a next step work task, wherein the current work task includes a current interaction target, and the next step work task includes a next step interaction target.
[0030] As an optimization, E1: determine whether the current interaction target is a suspected target source within the carrier range, if yes, jump to E2, otherwise, alarm;
[0031] E2: determine whether the current interaction target and the current target source are the same, if yes, continue working, jump to E3, otherwise, alarm.
[0032] E3: predicting the next target source according to the prediction model, and determining whether the next interaction target and the next target source are the same, if yes, returning to the step, until ending, otherwise, performing a warning prompt.
[0033] The application further discloses a system for judging power violation operation in advance through intention recognition, comprising:
[0034] The carrier screening module is configured to detect a first interest point feature group in the monitoring image data according to the work task, determine a carrier carrying at least one suspected target source, and preliminarily exclude non-target sources.
[0035] The marking module is configured to detect the positions of the suspected target sources and a second interest point feature group, compare the second interest point feature groups of the suspected target sources, distinguish the suspected target sources with similar or different morphologies, and mark the distinguished suspected target sources.
[0036] The data establishing module is configured to detect human skeleton key nodes, hand skeleton key nodes and hand data confirmed according to the hand skeleton key nodes in the monitoring image data, obtain the positions of the workers according to the human skeleton key nodes, and obtain a position relationship data set of the hand skeleton key nodes and the suspected target sources according to a time dimension based on a skeleton space-time convolution network.
[0037] The model generating module is configured to train the position relationship data set and a unit time corresponding to the time dimension as input training parameters of a neural network model, to obtain a prediction model.
[0038] The determination module is configured to determine a current action and predict and determine a next action based on the prediction model.
[0039] Compared with the prior art, the application has the following advantages and beneficial effects:
[0040] The application uses an artificial intelligence method to understand and predict the current operation by using human hand movement and human and article interaction, to determine whether the power operation is in violation, and to perform a warning for the violation operation in advance for several seconds, so as to ensure safety. BRIEF DESCRIPTION OF DRAWINGS
[0041] In order to more clearly illustrate the technical solutions of the example embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be considered as a limitation to the scope. For those skilled in the art, other related drawings can also be obtained without creative labor. In the drawings:
[0042] Figure 1 An example diagram for predicting interaction with a correct tool (target source);
[0043] Figure 2 An example diagram for predicting interaction with an incorrect tool (target source);
[0044] Figure 3 Another example diagram for predicting interaction with an incorrect tool (target source);
[0045] Figure 4 An example diagram of an action target suspected area obtained by the present application;
[0046] Figure 5 Another example diagram of an action target suspected area obtained by the present application. DETAILED DESCRIPTION
[0047] In order to make the objects, technical solutions and advantages of the present application clearer, further detailed description will be given to the present application in combination with embodiments and drawings, and the illustrative embodiments of the present application and the description thereof are only used to explain the present application, and do not limit the present application.
[0048] EMBODIMENT
[0049] A method for judging power violation operation in advance through intention recognition, comprising the following steps:
[0050] The input video stream is extracted through a deep convolutional neural network to obtain key parameters of desktop objects, including two steps of S1 and S2.
[0051] S1, according to the work task, a first interest point feature group in the monitoring image data is detected to determine the range of a carrier carrying at least one suspected target source, so as to preliminarily exclude non-target sources; specifically, in the embodiment, according to the work task, a carrier carrying a target source is determined, wherein the carrier refers to an object carrying a suspected target source, for example, as shown in FIG. 1, the carrier is a workbench, and the suspected target source is each object placed on the workbench, such as pliers, paper, and wire coils. Figure 1 For example, the work task this time is to take the wire coil on the table in FIG. 1, and then the determined carrier is the workbench in FIG. 1. Figure 1 The objects outside the workbench are non-target sources, and the objects on the workbench are all suspected target sources. Figure 1 The target source is included in the suspected target source, and the target source and the suspected target source both refer to the objects on the workbench. Before the object to be taken is determined, it is a suspected target source, and after the work staff determines that the suspected target source is to be taken, the suspected target source becomes a target source.
[0052] Through this step, the determination range of the target source can be reduced, and the determination result when the interactive object is outside the carrier range can be quickly obtained, for example, when the object close to the hand information is outside the carrier range, direct alarm or early warning is performed to avoid too many objects and complicate the determination process.
[0053] According to the first interest point feature group contained in the carrier, the range of the carrier containing at least one suspected target source including the target source is determined, and the objects outside the carrier range are discharged as non-target sources. The feature subset of the first interest point feature group includes the contour of the carrier, the color of the carrier, the texture of the carrier surface, etc.
[0054] S2, detecting the position of the suspected target source and the second interest point feature group, and comparing the second interest point feature groups of several suspected target sources to distinguish several suspected target sources with similar or different shapes, and marking the distinguished suspected target sources; in this way, the target sources in the carrier range can be better distinguished, and the marking error of objects with similar shapes can be avoided, for example, if there are multiple pliers with similar shapes in the carrier range, the pliers need to be distinguished by comparing the second interest point feature groups of the multiple pliers.
[0055] In this embodiment, the interest point features in the second interest point feature groups of several suspected target sources are compared according to a preset rule. The interest point features include contour shape, size, color, texture (if there is no texture, comparison is not performed) and label (if there is no label, comparison is not performed).
[0056] The human body skeleton key points and hand skeleton key points (hand skeleton key nodes) in the input video are extracted. The key points are predicted by a specially designed skeleton space-time convolution network to predict the next stage action and action target suspected area, and then combined with the recognized desktop object parameters, the maximum probability interactive object can be predicted. Finally, whether to alarm is judged through warning logic, including S3-S5.
[0057] S3, detecting the human body skeleton key node, the hand skeleton key node and the hand data confirmed according to the hand skeleton key node in the monitoring image data, obtaining the position of the staff according to the human body skeleton key node, and obtaining the position relationship data set of the hand skeleton key node and the suspected target source according to the time dimension based on the skeleton space-time convolution network; the hand information and the suspected target source are one-to-one corresponding, so that the subsequent prediction is more accurate.
[0058] Specifically, in the embodiment, the hand skeleton key nodes include a wrist skeleton key node composed of wrist joints, a palm skeleton key node composed of a palm center, and finger skeleton key nodes composed of finger joints and finger tips, the hand skeleton key nodes include wrist data, palm data, and finger data, the wrist data includes first positioning data of the wrist skeleton key node, the palm data includes an opening degree of the palm confirmed by the finger skeleton key nodes and an angle difference between the maximum angle and the minimum angle between the palm center and the upper surface of a profile of a certain suspected target source, and the finger data includes a distance change value of a fingertip of an index finger to a fingertip of a thumb of the same hand, and positioning data of finger joints, each of the finger joints includes a metacarpophalangeal joint of each finger and a proximal interphalangeal joint of each finger. The palm opening degree refers to whether the palm is open, if the palm is open, the palm opening degree value is 1, if the palm is not open, the palm opening degree value is 0, in a natural state, the fingers are bent, which is considered as the palm not being open, when the fingers are stretched, it is considered as the palm being open, and the nearest distance here refers to the nearest straight line distance from the wrist joint to the surface of the suspected target source.
[0059] In the embodiment, the specific manner of determining the position relationship data set is as follows:
[0060] According to the time dimension and the first positioning data, the nearest distance of the wrist skeleton key node to a certain suspected target source in each unit time is determined to detect the motion trajectory of the wrist, and the motion trajectory of the wrist can be understood as the motion trajectory of the whole hand.
[0061] According to the time dimension, the palm opening degree value and the angle difference in each unit time are determined.
[0062] According to the time dimension, the distance change value of the fingertip of the index finger to the suspected target source in each unit time is determined.
[0063] The nearest distance, the palm opening degree value, the angle difference, the distance change value, and the corresponding meaning target source in each unit time are combined into a set of position relationship data sets, and the position relationship data set is represented as: {A i,j , B i,j , C i,j , D i,j}, wherein i represents the target source, j represents the unit time on the time dimension, the unit time can be one second or several seconds, for example, when the unit time is one second, j takes 1, which is the first second, j takes 2, which is the second second, if the unit time is 2 seconds, j takes 1, which represents 1-2 seconds, j takes 2, which represents 3-4 seconds, A represents the nearest distance, B represents the palm opening degree value, C represents the angle difference, and D represents the distance change value.
[0064] Since the worker needs to bend down to take the corresponding tool on the surface of the carrier to work in the power operation, the palm height of the worker must be higher than the upper surface of the tool (suspected target source), when the palm is close to a suspected target source, the palm will open, and the palm is towards the target source, therefore, the angle difference between the maximum angle and the minimum angle between the palm center and the profile upper surface of the suspected target source must be smaller than the angle difference between the maximum angle and the minimum angle between the profile upper surfaces of other suspected target sources, and the distance change value between the fingertip of the index finger and the fingertip of the thumb of the same hand will increase in the unit time when the palm opens.
[0065] Specifically, in the embodiment, the palm opening value and the closest distance satisfy: when the closest distance is less than the first threshold value, the palm opening value satisfies the first condition, for example, when the closest distance from the wrist joint to the suspected target source is less than 20 cm, the palm opens, the palm opening value is 1, and the first condition is that the palm opening value is 1; the angle difference and the closest distance satisfy: when the closest distance is less than the second threshold value, the angle difference satisfies the second condition, for example, when the closest distance from the wrist joint to the suspected target source is less than 15 cm, the angle difference is less than 90°, and the second condition here is that the angle difference is less than 90°, since the worker's hand takes the object from top to bottom, when the palm is close to the object, the palm is located obliquely above the object, at this time, the angle difference between the maximum angle and the minimum angle between the palm center and the upper surface of the object (this angle difference will be less than 90°) will be smaller than the angle difference between the maximum angle and the minimum angle between the palm and the upper surface of other objects (this angle difference will be greater than 90°), the distance change value and the closest distance satisfy: when the closest distance is less than the third threshold value, the distance change value satisfies the third condition, for example, when the closest distance from the wrist joint to the suspected target source is between 13-18 cm, the distance change value in one of the corresponding unit times will be greater than 5 cm / s, and the third condition here is that the distance change value in the unit time is greater than 5 cm / s, that is, it indicates that the fingers are open when the hand is in this position, indicating that the worker is ready to take the object, when the position relationship data set satisfies the conditions at the same time, it is determined that the suspected target source is the target source.
[0066] Here, another case is provided to determine the target source.
[0067] Since the wrist joint includes multiple joints, here any one joint is selected as the wrist joint, and the subsequent wrist joint is also based on this joint.
[0068] In the embodiment, the specific way to determine the position relationship data set is:
[0069] determine the closest distance of the wrist joint of the radial styloid process to a certain suspected target source in each unit time according to the time dimension, so as to detect the motion trajectory of the wrist;
[0070] determine the palm opening value in each unit time according to the time dimension;
[0071] determine the connection trajectory of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger in each unit time according to the time dimension;
[0072] determine the relative position of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger to the wrist joint node according to the time dimension;
[0073] combine the closest distance, the connection trajectory of each finger joint node, the relative position of each finger joint node to the wrist joint node and the corresponding suspected target source in each unit time into a set of position relationship data sets, which is represented as: {A i,j , X i,j , Y i,j,k , Z i,j,k,m,n}, wherein i represents the suspected target source, j represents the unit time in the time dimension, A represents the closest distance, k represents the finger type, m represents the wrist joint node, n represents the finger joint node, X represents the palm opening value, Y represents the connection trajectory, and Z represents the relative position of each finger joint node to the wrist joint node.
[0074] In the embodiment, the palm opening value, the connection trajectory and the closest distance satisfy the following conditions: when the closest distance is less than a fourth threshold value, the palm opening value satisfies a first condition, i.e., the palm opening value is 1, the palm is in an open state, the connection trajectory of the thumb finger joint node is curved with the connection trajectory of the index finger or middle finger joint node, the angle of the interphalangeal joint of the thumb is between 120° and 170°, the angle of the proximal interphalangeal joint of the index finger or middle finger is between 90° and 160°, the angle of the interphalangeal joint of the index finger or middle finger is greater than 130°, and the curved radii of the connection trajectories of the thumb finger joint node and the index finger or middle finger joint node are both directed to a certain suspected target source, it is determined that the suspected target source is a certain target source.
[0075] When the nearest distance is less than the fourth threshold, it indicates that the hand is approaching a suspected target source. At this time, if the trajectory of the thumb's finger joint node is curved with the trajectory of the index or middle finger's finger joint node, and the trajectory of the thumb's finger joint node is curved with the trajectory of the index or middle finger's finger joint node, and the angle of the thumb's interphalangeal joint is between 120° and 170°, the angle of the proximal interphalangeal joint of the index or middle finger is between 90° and 160°, the angle of the index or middle finger's interphalangeal joint is greater than 130°, and the curvature of the trajectory of the thumb's finger joint node and the trajectory of the index or middle finger's finger joint node are all pointing towards a suspected target source, it means that the suspected target source is about to be grasped. In this way, the suspected target source can be determined as the target source to be grasped.
[0076] In this embodiment, the spatiotemporal convolutional network of the skeleton is specifically:
[0077]
[0078] Among them, f out (v tp B(v) represents the motion prediction output for the key node p of the skeleton at time t; tp Z represents the set of all associated nodes of the critical node p at time t. tp (v tq Z represents the normalized parameter for the key node p and its associated node q at time t, and its meaning is: tp (v tq The value of ) is the sum of the values of all neighboring nodes v of the critical node p at time t. tq Quantity; f in (v tq ) represents the input feature of the key node q at time t, and its meaning is: v tq →R c An N-dimensional vector mapping, R c Let w(v) be a c-dimensional vector on the set of real numbers, and let its elements include the node's sequence number, position, velocity vector, and motion vector relative to other nodes. ti v tj ) represents the path weight from key node i to its associated node j at time t. Specifically, w(v ti v tj It is obtained through training a neural network. Its calculation method is as follows:
[0079] w(v tp v tq )=w′(d(v tp v tq )),
[0080] Wherein d(vtp , v tq ) is the distance between the bone nodes (key nodes) p and q at time t, distance 0 means that p and q are the same, distance 1 means that p and q are adjacent nodes, and distance is large, w'(d(v tp , v tq )) is the derivative function, the weight function about the distance between the key nodes, and the derivative of time. v tp represents the key node p at time t, v tq represents the key node q at time t.
[0081] The position relationship data set is trained as an input training parameter of a neural network model to obtain a prediction model; the training process is a prior process, and the neural network model is also prior, and various thresholds herein can be set according to limited experiments, and the relationship between the relative positions of the finger joint nodes and the carpal joint node and the closest distance can also be obtained by bringing into an LSTM neural network model (long short-term memory neural network model) to train the related relationship data. The training process is prior art, which will not be described here. As shown in FIG. 8, when the hand is far away from the target source, a plurality of suspected target sources of the target source are obtained through the LSTM neural network model, that is, Figures 4-5 Figures 4-5 The suspected target source in the suspected area can be the target source, and after the hand is closer and closer to a certain suspected target source, the suspected target source can be determined as the target source. After the target source is confirmed, the path between the tip of the index finger or the middle finger and the target source can be marked, that is, Figures 1-3 the ray between the tip of the index finger or the middle finger and the target source in FIG. 9.
[0082] The positioning and identification of each joint and the like described above can be realized through a detector and a tracker, and the identification of the object can be realized through an object detector, which will not be described here.
[0083] The current action is determined based on the prediction model, and the next action is predicted and determined.
[0084] In the embodiment, an established task flow label is formulated according to task planning design, and the established task flow label includes a set current work task and a set next work task, wherein the current work task includes a current interactive target, and the next work task includes a next interactive target and a carrier carrying the interactive target. According to the task planning, the carrier and the interactive target detected each time can be known, and subsequently it is only necessary to judge whether the detected target source and the interactive target are the same, and the specific process is as follows:
[0085] In this embodiment, E1: determine whether the current interaction target belongs to the suspected target source within the carrier range, if yes, jump to E2, otherwise, alarm;
[0086] E2: determine whether the current interaction target and the current target source are the same, if yes, continue to work, jump to E3, otherwise, alarm;
[0087] E3: predict the next target source according to the prediction model, and determine whether the next interaction target and the next target source are the same, if yes, return to this step until the end, otherwise, perform early warning prompt.
[0088] Embodiment 2
[0089] The application further discloses a system for judging power violation operation in advance through intention recognition, comprising:
[0090] A carrier screening module is configured to detect a first interest point feature group in the monitoring image data according to a work task, determine a carrier carrying at least one suspected target source, and preliminarily exclude non-target sources.
[0091] A marking module is configured to detect the position of the suspected target source and a second interest point feature group, compare the second interest point feature groups of a plurality of suspected target sources, distinguish a plurality of suspected target sources with similar or different morphologies, and mark the distinguished suspected target sources.
[0092] A data establishing module is configured to detect human skeleton key nodes, hand skeleton key nodes and hand data confirmed according to the hand skeleton key nodes in the monitoring image data, obtain the position of a worker according to the human skeleton key nodes, and obtain a position relationship data set of the hand skeleton key nodes and the suspected target source according to a time dimension based on a skeleton space-time convolution network.
[0093] A model generating module is configured to train the position relationship data set and a unit time corresponding to the time dimension as input training parameters of a neural network model to obtain a prediction model.
[0094] A determination module is configured to determine a current action, and predict and determine a next action based on the prediction model.
[0095] As shown in Figures 1-3 , the white bold box indicates the tool that needs to be correctly selected next, the black bold box indicates the tool that is predicted to be interacted with incorrectly, the white thin box indicates other tools on the table that are not involved in human-tool interaction, and the hand joint position is indicated by a circle, and the ray between the finger tip and the tool indicates the interaction relationship between the target source predicted by the neural network.
[0096] As shown in Figure 1As shown, the predicted next interaction target is the right bold white frame clipper, and the predicted worker is about to take the right bold white frame clipper, at this time, it is predicted that the interaction target and the target source are the same, and no alarm is reported.
[0097] As shown, the predicted interaction target and the target source are not the same, and an alarm is given. Figures 2-3
[0098] Through the method, prediction can be made in advance, and then warning can be given.
[0099] The present application uses human-object interaction prediction technology: first, a human motion detector is used to extract the key nodes of the human skeleton, then a human palm detector and tracker are used to position and joint detect the human palm in more detail, and at the same time, an object detector is used to position and classify the objects in the picture in space and semantics.
[0100] After obtaining the above information as input information, the spatial-temporal two-dimensional information is used to model whether the detection operation is wrong. In space, the neural network is used to model the human skeleton, the trend of human motion, the trend of palm motion, the trend of approaching an object, and the spatial distance between the detection objects. In time sequence, a long short-term memory time network cycle model is used to model different actions and objects contacted before and after in time sequence, for enhancing the classification accuracy of the whole model for identifying the events before and after.
[0101] Experimental process: real-time system test on multiple operation processes, verify the system alarm success rate and the average early warning time under different operation speed conditions
[0102] Experimental input data: about 60 minutes of video stream in total for multiple operations, and the number of interactions with objects is more than 100 times;
[0103] Experimental variable parameters: average moving speed of operator's hand: low (<10 cm / s), medium (10-30 cm / s), high (>30 cm / s)
[0104] Experimental results:
[0105]
[0106] The results show that for medium and low operation speed, the warning success rate is satisfactory, and the warning advance time is about 2 seconds. For high speed operation, due to the distance between objects, occlusion, system response time and other factors, the warning effect is slightly worse than that of medium and low operation speed.
[0107] During the neural network model training process, improving the prediction sensitivity of the suspected area will increase the early warning time, but will cause a greater false alarm probability. Therefore, the warning time, system sensitivity and other parameters need to be adjusted according to the actual working environment.
[0108] The application not only detects simple actions, but also utilizes semantic information of each object in space and correlation information in time sequence, so that the whole model has stronger universality. As long as the detector and tracker can obtain information, the application is applicable in various application scenarios, and the application also provides a labeled data set, which is beneficial to make the whole model more accurate and stable.
[0109] The above specific embodiments further illustrate the purpose, technical solutions and beneficial effects of the application. It should be understood that the above description is only a specific embodiment of the application and is not used to limit the protection scope of the application. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the application should be included in the protection scope of the application.
Claims
1. A method of early detection of power violation operations by intent recognition, characterized by, The method comprises the following steps: According to the work task, the first interest point feature group in the monitoring image data is detected to determine the range of the carrier carrying at least one suspected target source, so as to preliminarily exclude non-target sources; The position of the suspected target source and the second interest point feature group are detected, and the second interest point feature groups of a plurality of suspected target sources are compared to distinguish a plurality of suspected target sources with similar or different morphologies, and the distinguished suspected target sources are marked; The human skeleton key node, the hand skeleton key node in the monitoring image data, and the hand data confirmed according to the hand skeleton key node are detected, the position of the staff is obtained according to the human skeleton key node, and the position relationship data set of the hand skeleton key node and the suspected target source is obtained according to the time dimension based on the skeleton space-time convolution network; The position relationship data set and the corresponding time dimension unit time are taken as the input training parameters of the neural network model to train the prediction model; The current action is determined, and the next action is predicted and determined based on the prediction model.
2. The method of claim 1, wherein the method further comprises: The hand skeleton key node comprises a wrist skeleton key node composed of wrist joints, a palm skeleton key node composed of a palm center, and a finger skeleton key node composed of a plurality of finger joints and finger tips, the hand data comprises wrist data, palm data, and finger data, the wrist data comprises first positioning data of the wrist skeleton key node, the palm data comprises an opening degree of the palm confirmed by the finger skeleton key node and an angle difference between the maximum angle and the minimum angle between the palm center and the upper surface of the outline of a certain suspected target source, and the finger data comprises a distance change value of the tip of the index finger to the tip of the thumb of the same hand, and positioning data of each finger joint, each finger joint comprising a metacarpophalangeal joint of each finger and a finger interphalangeal joint of each finger.
3. The method of claim 2, wherein the method further comprises: The specific manner for determining the position relationship data set is: According to the time dimension and the first positioning data, the closest distance of the wrist skeleton key node to a certain suspected target source in each unit time is determined to detect the motion trajectory of the wrist; According to the time dimension, the palm opening degree value and the angle difference in each unit time are determined; According to the time dimension, the distance change value of the tip of the index finger to the suspected target source in each unit time is determined; The closest distance, the palm opening value, the angle difference, the distance change value in each unit time and the corresponding suspected target source are combined into a set of position relationship data sets, and the position relationship data sets are represented as: {A i,j , B i,j , C i,j , D i,j}, wherein i represents the suspected target source, j represents the unit time in the time dimension, A represents the closest distance, B represents the palm opening value, C represents the angle difference, and D represents the distance change value.
4. The method of claim 3, wherein the method further comprises: The palm opening degree value and the closest distance satisfy the first condition when the closest distance is less than a first threshold value, the angle difference and the closest distance satisfy the second condition when the closest distance is less than a second threshold value, and the distance change value and the closest distance satisfy the third condition when the closest distance is less than a third threshold value, and when the position relationship data set satisfies the conditions at the same time, it is determined that the suspected target source is a target source.
5. The method of claim 2, wherein the method further comprises: The specific manner for determining the position relationship data set can also be: According to the time dimension, the closest distance of the wrist skeleton key node to a certain suspected target source in each unit time is determined to detect the motion trajectory of the wrist; According to the time dimension, the palm opening degree value in each unit time is determined; Determine the connection track of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger in each unit time according to the time dimension; Determine the relative position of each finger joint node of the thumb, index finger, middle finger, ring finger and little finger and the carpal joint node according to the time dimension; The closest distance in each unit time, the connection track of each finger joint node, the relative position of each finger joint node and the wrist bone joint node, and the corresponding suspected target source are respectively combined into a set of position relationship data sets, and the position relationship data sets are represented as: {A i,j , X i,j , Y i,j,k , Z i,j,k,m,n}, wherein i represents the suspected target source, j represents the unit time in the time dimension, A represents the closest distance, k represents the finger type, m represents the wrist bone joint node, n represents the finger joint node, X represents the palm opening value, Y represents the connection track, and Z represents the relative position of each finger joint node and the wrist bone joint node.
6. A method of identifying power violation activities in advance through intent recognition according to claim 5, characterized in that, The palm opening value, the connection track and the nearest distance satisfy: when the nearest distance is less than a fourth threshold value, the palm opening value satisfies a first condition, the connection track of the thumb finger joint node is curved with the connection track of the index finger or middle finger joint node, the angle of the interphalangeal joint of the thumb is between 120° and 170°, the angle of the proximal interphalangeal joint of the index finger or middle finger is between 90° and 160°, the angle of the interphalangeal joint of the index finger or middle finger is greater than 130°, and the curved radii of the connection track of the thumb finger joint node and the connection track of the index finger or middle finger joint node are both towards a suspected target source, and the suspected target source is determined as a fixed target source.
7. The method of claim 5, wherein the method further comprises: The skeleton space-time convolution network specifically is: where f out (v tp ) denotes the motion prediction output of the skeleton key node p at time t; B(v tp ) denotes the set of all associated nodes of the key node p at time t, Z tp (v tq ) denotes the normalized parameters for the key node p and its associated node q at time t, f in (v tq ) denotes the input features of the key node q at time t, w(v tp , v tq ) denotes the path weight from the key node p to its associated node q at time t.
8. The method for early detection of power violation activities through intent recognition as claimed in claim 1, wherein, According to the task planning design, an established task flow label is formulated, and the established task flow label includes a set current work task and a set next work task, wherein the current work task includes a current interactive target, and the next work task includes a next interactive target.
9. The method of claim 8, wherein, E1: determining whether the current interactive target is a suspected target source within the carrier range, if yes, jumping to E2, otherwise, alarming; E2: determining whether the current interactive target and the current fixed target source are the same, if yes, continuing to work, jumping to E3, otherwise, alarming; E3: predicting the next fixed target source according to the prediction model, and determining whether the next interactive target and the next fixed target source are the same, if yes, returning to the step until the end, otherwise, warning.
10. A system for early detection of power violation operations through intent recognition, characterized by, Comprise: A carrier screening module is configured to detect a first interest point feature group in the monitoring image data according to a work task, determine a carrier carrying at least one suspected target source, and preliminarily exclude non-target sources; A marking module is configured to detect the position of the suspected target source and a second interest point feature group, and compare the second interest point feature groups of a plurality of suspected target sources to distinguish a plurality of suspected target sources with similar or different morphologies, and mark the distinguished suspected target sources; A data establishing module is configured to detect human skeleton key nodes, hand skeleton key nodes and hand data confirmed according to the hand skeleton key nodes in the monitoring image data, obtain the position of the worker according to the human skeleton key nodes, and obtain the position relationship data set of the hand skeleton key nodes and the suspected target source according to the time dimension based on a skeleton space-time convolution network; A model generating module is configured to train the position relationship data set and the unit time of the corresponding time dimension as input training parameters of a neural network model to obtain a prediction model. A determination module is configured to make a current action determination and a next action prediction and determination based on the prediction model.
Citation Information
Patent Citations
Method for acquiring hand region of interest and handprint recognition method
CN110728232A
Hand skeleton learning, lifting, and denoising from 2d images
US20190180473A1