Method and system for determining personnel to whom a tool for electric power operation belongs
By improving the Yolov5 network, the problem of determining the affiliation relationship of tools in power operations is solved, and accurate positioning and normative judgment of the personnel to which tools are affiliated is achieved.
Patent Information
- Application Number
- CN202211067731.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-01
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-09-01
AI Technical Summary
The existing deep learning object detection methods cannot accurately determine the affiliation between power operators and tools, especially when multiple operators appear in the image, resulting in the inability to determine whether the personnel use tools correctly.
Improve the Yolov5 network, add prediction branches to predict the position of the personnel to which the tool belongs, and determine the affiliation relationship of the tool by calculating the offset distance between the center point offset and the predicted offset.
It realizes accurate positioning of the affiliation relationship of tools, and can automatically identify whether the operator uses tools correctly, providing a basis for identifying irregular behavior.
Smart Images

Figure CN115424299B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of tool detection, and particularly to a method and system for determining the personnel to whom power operation personnel tools belong. Background Art<s
[0002] Using computer vision means to identify the use of tools by power operation personnel can effectively detect non-standard behaviors in the use of tools by operation personnel (such as safety helmets, insulating gloves, operating rods, etc.). Existing deep learning object detection methods can relatively accurately find the operation personnel and tools that appear in the image, but cannot determine the subordinate relationship of the tools. Especially when there are multiple operation personnel in the image and there is overlap between the personnel, it is more difficult to determine the subordinate relationship of the tools, resulting in the inability to judge whether the personnel correctly use the tools based on the personnel as the main body. Summary of the Invention
[0003] The purpose of the present invention is to provide a method and system for determining the personnel to whom power operation personnel tools belong, which can determine the subordinate relationship of the tools, and further, based on the personnel as the main body, determine whether the personnel correctly use the tools.
[0004] To achieve the above purpose, the present invention provides the following solutions:
[0005] A method for determining the personnel to whom power operation personnel tools belong, comprising:
[0006] Obtain an image to be recognized; the image to be recognized includes multiple personnel and multiple tools;
[0007] Input the image to be recognized into a trained improved Yolo v5 network to obtain a human target set of the image to be recognized and a tool target set of the image to be recognized; the human target set of the image to be recognized includes the central point position coordinates, detection frame width, and detection frame height of each personnel in the image to be recognized, and the tool target set of the image to be recognized includes the central point position coordinates, category, detection frame width, detection frame height, and predicted offset of each tool in the image to be recognized; the improved Yolo v5 network includes a backbone network and a feature fusion network connected in sequence, and three prediction branches respectively connected to the output end of the feature fusion network;
[0008] For any one tool in the image to be recognized, according to the central point position coordinates of the tool, the predicted offset of the tool, and the central point position coordinates of each personnel in the image to be recognized, calculate the offset distance between the central point offset of the tool and each personnel and the predicted offset, and determine that the tool belongs to the personnel corresponding to the smallest offset distance.
[0009] Optionally, for any tool in the image to be recognized, calculate the offset distance between the center point offset of the tool and each person and the predicted offset according to the center point position coordinates of the tool, the predicted offset of the tool, and the center point position coordinates of each person in the image to be recognized. Specifically, it includes:
[0010] According to the formula Calculate the offset distance between the center point offset of tool i and person j and the predicted offset, where px j Represents the abscissa of the center point of person j, tx i Represents the abscissa of the center point of tool i, py i Represents the ordinate of the center point of person j, ty i Represents the ordinate of the center point of tool i, tu i Represents the abscissa of the predicted offset, tv i Represents the ordinate of the predicted offset, d ij Represents the offset distance between the center point offset of tool i and person j and the predicted offset.
[0011] Optionally, the determination process of the trained improved Yolo v5 network specifically includes:
[0012] Obtain training samples; the training samples include multiple training pictures, and each training picture includes multiple people and multiple tools;
[0013] Label each training picture in the training samples to obtain the human target set of each training picture and the tool target subset of each training picture; the human target set of the training picture includes the center point position coordinates, detection box width, and detection box height of each person in the training picture, and the tool target subset of the training picture includes the center point position coordinates, category, detection box width, and detection box height of each tool in the training picture;
[0014] For any tool in any training picture in the training samples, determine the true owner of the tool according to the distance between the center point position coordinates of the area where the tool is preset to belong to the person and the center point position coordinates of each person in the training picture;
[0015] Obtain the predicted offset of the tool according to the center point position coordinates of the true owner of the tool and the center point position coordinates of the area where the tool is preset to belong to the person;
[0016] Determine the predicted offset, center point position coordinates, category, detection box width, and detection box height of each tool in the training picture as the tool target set of the training picture;
[0017] Using the set of tool objects and the set of human objects in each training image of the training samples as the training set, train the improved Yolo v5 network to obtain a trained improved Yolo v5 network.
[0018] Optionally, determining the actual owner of the tool based on the distance between the center point position coordinates of the area where the pre-set owner of the tool is located and the center point position coordinates of each person in the training image specifically includes:
[0019] Calculating the distance between the area where the pre-set owner of the tool is located and each person in the training image according to the center point position coordinates of the area where the pre-set owner of the tool is located and the center point position coordinates of each person in the training image;
[0020] Determining the person corresponding to the smallest such distance as the actual owner of the tool.
[0021] A system for determining the owner of tools for electric power operation personnel includes:
[0022] An image acquisition module, configured to acquire an image to be recognized; the image to be recognized includes multiple persons and multiple tools;
[0023] A set determination module, configured to input the image to be recognized into the trained improved Yolo v5 network to obtain the set of human objects and the set of tool objects of the image to be recognized; the set of human objects of the image to be recognized includes the center point position coordinates, detection frame width, and detection frame height of each person in the image to be recognized, and the set of tool objects of the image to be recognized includes the center point position coordinates, category, detection frame width, detection frame height, and prediction offset of each tool in the image to be recognized; the improved Yolo v5 network includes a backbone network and a feature fusion network connected in sequence and three prediction branches respectively connected to the output end of the feature fusion network;
[0024] A result determination module, configured to, for any tool in the image to be recognized, calculate the offset distance between the center point offset of the tool and each person and the prediction offset according to the center point position coordinates of the tool, the prediction offset of the tool, and the center point position coordinates of each person in the image to be recognized, and determine that the tool belongs to the person corresponding to the smallest such offset distance.
[0025] Optionally, the result determination module specifically includes:
[0026] An offset distance determination unit, configured to according to the formula Calculate the distance between the center point offset of the tool i and the predicted offset of the person j, where px j represents the abscissa of the center point of the person j, tx i represents the abscissa of the center point of the tool i, py i represents the ordinate of the center point of the person j, ty i represents the ordinate of the center point of the tool i, tu i represents the abscissa of the predicted offset, tv i represents the ordinate of the predicted offset, d ij represents the offset distance between the center point offset of the tool i and the person j and the predicted offset.
[0027] Optionally, the system for determining the person to whom the electric power operation tool belongs further includes:
[0028] A sample acquisition module, configured to acquire training samples; the training samples include multiple training pictures, and each training picture includes multiple persons and multiple tools;
[0029] A set annotation module, configured to annotate each training picture in the training samples to obtain a human target set of each training picture and a tool target subset of each training picture; the human target set of the training picture includes the center point position coordinates, detection frame width, and detection frame height of each person in the training picture, and the tool target subset of the training picture includes the center point position coordinates, category, detection frame width, and detection frame height of each tool in the training picture;
[0030] A true result annotation module, configured to, for any tool in any training picture in the training samples, determine the person to whom the tool truly belongs according to the distance between the center point position coordinates of the area where the person to whom the tool is preset to belong and the center point position coordinates of each person in the training picture;
[0031] A predicted offset calculation module, configured to obtain the predicted offset of the tool according to the center point position coordinates of the person to whom the tool truly belongs and the center point position coordinates of the area where the person to whom the tool is preset to belong;
[0032] A tool target subset determination unit, configured to determine the predicted offset, center point position coordinates, category, detection frame width, and detection frame height of each tool in the training picture as the tool target set of the training picture;
[0033] A training module, which uses the set of tool objects and the set of human objects in each training image in the training samples as a training set to train the improved Yolo v5 network to obtain a trained improved Yolov5 network.
[0034] Optionally, the true result annotation module specifically includes:
[0035] A distance calculation unit, which is used to calculate the distance between the center point position coordinates of the area where the tool's preset affiliated person is located and the center point position coordinates of each person in the training image;
[0036] A true result annotation unit, which is used to determine that the person corresponding to the smallest distance is the true affiliated person of the tool.
[0037] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention: The present invention improves the Yolo v5 network and adds a prediction branch to predict the subject vector, so that while the network completes the detection task, it can predict the position of the person to whom the tool belongs, and then match it with the detected operators, so as to determine the operator to whom the tool belongs, complete the identification of the operator's tool, and provide a basis for further automatically identifying the non-standard behavior of the operator using the tool. Description of the Drawings
[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0039] Figure 1 It is a flowchart of the method for determining the person to whom the tool of the electric power operator belongs provided by the embodiment of the present invention;
[0040] Figure 2 It is a structural block diagram of the improved Yolo v5 network described in the present invention;
[0041] Figure 3 It is a structural schematic diagram of the improved Yolo v5 network described in the present invention. Detailed Embodiments
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0044] As Figure 1 shown, a method for determining the personnel to whom the tools of electric power operation personnel provided by an embodiment of the present invention specifically includes:
[0045] Obtain an image to be recognized; the image to be recognized includes multiple personnel and multiple tools.
[0046] Input the image to be recognized into a trained improved Yolo v5 network to obtain the human target set of the image to be recognized and the tool target set of the image to be recognized The human target set of the image to be recognized includes the center point position coordinates, detection frame width, and detection frame height of each person in the image to be recognized. The tool target set of the image to be recognized includes the center point position coordinates, category, detection frame width, detection frame height, and predicted offset of each tool in the image to be recognized. The improved Yolov5 network includes a backbone network and a feature fusion network connected in sequence, and three prediction branches respectively connected to the output end of the feature fusion network. Among them, px j , py j respectively represent the abscissa and ordinate of the center point position of person j in the image to be recognized, pw j and ph j respectively represent the detection frame width and detection frame height of person j in the image to be recognized, Np′ represents the total number of people in the image to be recognized, Nt′ represents the total number of tools in the image to be recognized, tx i , ty i , respectively represent the abscissa and ordinate of the center point position of tool i in the image to be recognized, tc i represents the category of tool i in the image to be recognized, tw i , and th i respectively represent the detection frame width and detection frame height of tool i in the image to be recognized, (tu i , tv i ) represents the predicted offset of tool i in the image to be recognized.
[0047] For any tool in the image to be recognized, calculate the central point offset v between the tool and the central point of each person according to the central point position coordinates of the tool, the predicted offset of the tool, and the central point position coordinates of each person in the image to be recognized. ij and the predicted offset (tu i , tv i ) The offset distance d between ij And determine that the tool belongs to the person corresponding to the smallest offset distance.
[0048] In practical applications, for any tool in the image to be recognized, according to the central point position coordinates of the tool, the predicted offset of the tool, and the central point position coordinates of each person in the image to be recognized, calculate the offset distance between the central point offset of the tool and each person and the predicted offset. Specifically, it includes:
[0049] According to the formula Calculate the distance between the central point offset of tool i and person j and the predicted offset, where px j represents the abscissa of the central point of person j, tx i represents the abscissa of the central point of tool i, py j represents the ordinate of the central point of person j, ty i represents the ordinate of the central point of tool i, tu i represents the abscissa of the predicted offset, tv i represents the ordinate of the predicted offset, d ij represents the offset distance between the central point offset of tool i and person j and the predicted offset.
[0050] In practical applications, the central point offset is: for example, for each element t in T' i , traverse all elements in the set P', and calculate the central point offset vector from element t i to the elements in P'. Let the element in P' be p j , then the central point offset vector is: vij = (px j -tx i , py j -ty i ).
[0051] In practical applications, the determination process of the trained improved Yolov5 network specifically includes:
[0052] Obtain training samples; the training samples include multiple training pictures, and each training picture includes multiple people and multiple tools.
[0053] The training samples are labeled with subject vectors to obtain the labeled training samples, so that each tool has a subject vector pointing to its user. For each image I, two sets are finally obtained: the human target set and the tool target set. The specific labeling process is as follows:
[0054] Label each training image in the training samples to obtain the human target set of each training image and the tool target subset The human target set of the training image includes the center point position coordinates, detection box width, and detection box height of each person in the training image. The tool target subset of the training image includes the center point position coordinates, category (safety helmet, insulating gloves, operating rod, etc.), detection box width, and detection box height of each tool in the training image, px j ′, py j ′ respectively represent the abscissa and ordinate of the center point position of person j in the training image, pw j ′ and ph j ′ respectively represent the detection box width and detection box height of person j in the training image. Np represents the total number of people in the training image, and Nt represents the total number of tools in the training image. tx i ′, ty i ′ respectively represent the abscissa and ordinate of the center point position of tool i in the training image, tc i ′ represents the category of tool i in the training image, twi′ and th i ′ respectively represent the detection box width and detection box height of tool i in the training image, (tu i ′, tv i ′) represents the predicted offset of tool i in the training image.
[0055] For any tool in any training image of the training samples, determine the actual user of the tool according to the distance between the center point position coordinates of the preset area where the user of the tool is located and the center point position coordinates of each person in the training image.
[0056] Obtain the predicted offset of the tool (the vector from the center point of the tool to the center point of the user) according to the center point position coordinates of the actual user of the tool and the center point position coordinates of the preset area where the user of the tool is located.
[0057] Determine the predicted offset, center point position coordinates, category, detection box width, and detection box height of each tool in the training image as the tool target set (labeled training samples) of the training image.
[0058] Using the set of tool objects and the set of human objects in each training image of the training samples as the training set, train the improved Yolov5 network to obtain a trained improved Yolov5 network.
[0059] In practical applications, such as Figure 2 and Figure 3 As shown, the improved Yolov5 network includes: a backbone network, a feature fusion network, and a prediction network connected in sequence. The prediction network includes three prediction branches, and each prediction branch is a Conv. The backbone network includes a first CBS, a second CBS, a first C3, a third CBS, a second C3, a third C3, a fourth CBS, a fourth C3, a fifth C3, a sixth C3, a fifth CBS, a seventh C3, an SPPF, and a sixth CBS connected in sequence. The feature fusion network includes: a first Up, a first Concat, an eighth C3, a seventh CBS, a second Up, a second Concat, a ninth C3, an eighth CBS, a third Concat, a tenth C3, a ninth CBS, a fourth Concat, and an eleventh C3 connected in sequence; the output end of the third C3 is also connected to the input end of the second Concat, the output end of the sixth C3 is connected to the input end of the first Concat, the output end of the sixth CBS is respectively connected to the input end of the first Up and the input end of the fourth Concat, the output end of the seventh CBS is also connected to the input end of the third Concat, the output end of the ninth C3 is connected to the first branch of the prediction network, the output end of the tenth C3 is also connected to the second branch of the prediction network, and the output end of the eleventh C3 is connected to the third branch of the prediction network. Among them, CBS includes Conv, BN, and SiLU, C3 includes Bottlenck, Concat, and three Convs, SPPF includes three Max Pools and one Concat, Bottlenck includes two Convs and one Add. A main vector regression prediction loss is added to the original Yolo v5 network. The calculation formula of this main vector regression prediction loss is as follows: Where and are the predicted main vectors, and tu i ′ and tv i ″ are the true main vectors obtained by annotation, that is, the predicted offset of tool i in the set of tool objects in the training image. The final loss function is the sum of the original loss function (object loss, classification loss, and bounding box regression loss) and the main vector regression prediction loss lcv. During training, the improved Yolo v5 network is trained with the goal of minimizing the final loss function to obtain a trained improved Yolo v5 network. Among them, the calculation formula of the object loss is as follows: lobj = ∑1 noobj [clog(c)+(1 - c)log(1 - c)]+∑1obj [clog(c)+(1 - c)log(1 - c)], where 1 noobj and 1 obj takes values of 0, 1. If the current box is the target, 1 noobj = 0, 1 obj = 1. If the current box is not the target, 1 noobj = 1, 1 obj = 0. c is the predicted object existence prediction value, and the classification loss calculation formula is: lcls = λ cls ∑1 obj BCE, where λ cls is the classification loss constant 1 obj , indicating whether the current prediction box contains the target, BCE represents the binary cross - entropy loss, and the bounding box regression loss is: lbox = CIoU. The calculation formula of CIoU is as follows where
[0060] b, b gt are the ground truth box and the prediction box respectively, ρ 2 calculates the distance between the center points of the two rectangular boxes, w gt , h gt are the width and height of the ground truth box respectively, w, h are the width and height of the prediction box respectively. IoU represents the intersection of the areas of the ground truth box and the prediction box divided by the union of the areas, α represents the aspect ratio coefficient, υ represents the aspect ratio difference, and c is a constant.
[0061] In practical applications, determining the true owner of the tool according to the distance between the center point position coordinates of the area where the owner of the tool is preset and the center point position coordinates of each person in the training picture specifically includes:
[0062] Calculating the distance between the area where the owner of the tool is preset and each person in the training picture according to the center point position coordinates of the area where the owner of the tool is preset and the center point position coordinates of each person in the training picture.
[0063] Determining the person corresponding to the minimum of the distances as the true owner of the tool.
[0064] In practical applications, the process of determining the center point position coordinates of the area where the owner of the tool is preset is as follows: For each element t in the set T of tool targets in each training image I i , manually click on the target human area to which t i belongs, and record the coordinates of this position as (Mx i , My i)Obtain the central point position coordinates of the area where the preset owner of the tool is located. Based on the original detection box annotation, the present invention only needs to click the mouse once on each image to complete the annotation, which is very simple.
[0065] In practical applications, determine the person corresponding to the smallest of the said distances as the real owner of the tool. Specifically: calculate the distances from the position clicked by the mouse to each person in the human target set P and find the person p with the smallest distance jm as the real owner. For example, when calculating the distance from the position clicked by the mouse to the j-th person in the human target set P, the formula:
[0066] Obtain the predicted offset of the tool according to the central point position coordinates of the real owner of the tool and the central point position coordinates of the area where the preset owner of the tool is located. Specifically, according to the formula tu i ′ = px jm - Mx i , tv i ′ = py jm - My i Obtain the predicted offset of the tool, where px jm and py jm are the abscissa and ordinate of the central point position of person p jm respectively.
[0067] For the above method, an embodiment of the present invention also provides a system for determining the owner of a tool for electric power operation personnel, including:
[0068] An image acquisition module, configured to acquire an image to be recognized; the image to be recognized includes multiple persons and multiple tools.
[0069] A set determination module, configured to input the image to be recognized into a trained improved Yolov5 network to obtain the human target set of the image to be recognized and the tool target set of the image to be recognized; the human target set of the image to be recognized includes the central point position coordinates, detection box width, and detection box height of each person in the image to be recognized, and the tool target set of the image to be recognized includes the central point position coordinates, category, detection box width, detection box height, and predicted offset of each tool in the image to be recognized; the improved Yolov5 network includes a backbone network and a feature fusion network connected in sequence and three prediction branches respectively connected to the output end of the feature fusion network.
[0070] A result determination module, configured to, for any tool in the image to be recognized, calculate the offset distance between the center point offset of the tool and each person in the image to be recognized and the predicted offset, and determine that the tool belongs to the person corresponding to the smallest offset distance according to the center point position coordinates of the tool, the predicted offset of the tool, and the center point position coordinates of each person in the image to be recognized.
[0071] In practical applications, the result determination module specifically includes:
[0072] An offset distance determination unit, configured to calculate the distance between the center point offset of tool i and person j and the predicted offset according to the formula where px j represents the abscissa of the center point of person j, tx i represents the abscissa of the center point of tool i, py i represents the ordinate of the center point of person j, ty i represents the ordinate of the center point of tool i, tu i represents the abscissa of the predicted offset, tv i represents the ordinate of the predicted offset, and d ij represents the offset distance between the center point offset of tool i and person j and the predicted offset.
[0073] In practical applications, the system for determining the person to whom a power operation tool belongs further includes:
[0074] A sample acquisition module, configured to acquire training samples; the training samples include multiple training pictures, and each training picture includes multiple persons and multiple tools.
[0075] An aggregation annotation module, configured to annotate each training picture in the training samples to obtain the human target set of each training picture and the tool target subset of each training picture; the human target set of the training picture includes the center point position coordinates, detection frame width, and detection frame height of each person in the training picture, and the tool target subset of the training picture includes the center point position coordinates, category, detection frame width, and detection frame height of each tool in the training picture.
[0076] A true result annotation module, configured to, for any tool in any training picture in the training samples, determine the true person to whom the tool belongs according to the distance between the center point position coordinates of the area where the preset person to whom the tool belongs is located and the center point position coordinates of each person in the training picture.
[0077] A prediction offset calculation module, configured to obtain a predicted offset of the tool according to the center point position coordinates of the actual owner of the tool and the center point position coordinates of the area where the preset owner of the tool is located.
[0078] A tool target subset determination unit, configured to determine the predicted offset, center point position coordinates, category, detection frame width, and detection frame height of each tool in the training image as the tool target set of the training image.
[0079] A training module, configured to use the tool target set of each training image in the training sample and the human target set of each training image as a training set to train the improved Yolo v5 network to obtain a trained improved Yolov5 network.
[0080] In practical applications, the real result annotation module specifically includes:
[0081] A distance calculation unit, configured to calculate the distance between the area where the preset owner of the tool is located and each person in the training image according to the center point position coordinates of the area where the preset owner of the tool is located and the center point position coordinates of each person in the training image.
[0082] A real result annotation unit, configured to determine the person corresponding to the smallest distance as the actual owner of the tool.
[0083] In this specification, each embodiment is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0084] In this article, specific examples are used to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A method for determining the personnel to whom the tools of electric power operation personnel belong, characterized in that, Including: Obtain the image to be recognized; the image to be recognized includes multiple personnel and multiple tools; Input the image to be recognized into the trained improved Yolo v5 network to obtain the human target set of the image to be recognized and the tool target set of the image to be recognized; the human target set of the image to be recognized includes the center point position coordinates, detection box width, and detection box height of each person in the image to be recognized, and the tool target set of the image to be recognized includes the center point position coordinates, category, detection box width, detection box height, and prediction offset of each tool in the image to be recognized; the improved Yolo v5 network includes a backbone network and a feature fusion network connected in sequence, and three prediction branches respectively connected to the output end of the feature fusion network; each prediction branch is a Conv, and the backbone network includes a first CBS, a second CBS, a first C3, a third CBS, a second C3, a third C3, a fourth CBS, a fourth C3, a fifth C3, a sixth C3, a fifth CBS, a seventh C3, an SPPF, and a sixth CBS connected in sequence, and the feature fusion network includes: a first Up, a first Concat, an eighth C3, a seventh CBS, a second Up, a second Concat, a ninth C3, an eighth CBS, a third Concat, a tenth C3, a ninth CBS, a fourth Concat, and an eleventh C3 connected in sequence; the output end of the third C3 is also connected to the input end of the second Concat, the output end of the sixth C3 is connected to the input end of the first Concat, the output end of the sixth CBS is respectively connected to the input end of the first Up and the input end of the fourth Concat, the output end of the seventh CBS is also connected to the input end of the third Concat, the output end of the ninth C3 is connected to the first branch of the prediction network, the output end of the tenth C3 is also connected to the second branch of the prediction network, and the output end of the eleventh C3 is connected to the third branch of the prediction network; For any tool in the image to be recognized, calculate the offset distance between the center point offset of the tool and each person and the prediction offset, and determine that the tool belongs to the person corresponding to the smallest offset distance according to the center point position coordinates of the tool, the prediction offset of the tool, and the center point position coordinates of each person in the image to be recognized.
2. The method for determining the personnel to whom the tools for power operation personnel according to claim 1 belong, characterized in that, For any tool in the image to be recognized, calculate the offset distance between the center point offset of the tool and each person and the prediction offset according to the center point position coordinates of the tool, the prediction offset of the tool, and the center point position coordinates of each person in the image to be recognized, specifically including: According to the formula calculate the offset distance between the center point offset of the tool i and the person j and the predicted offset, where px j represents the abscissa of the center point of the person j, tx i represents the abscissa of the center point of the tool i, py i represents the ordinate of the center point of the person j, ty i represents the ordinate of the center point of the tool i, tu i represents the abscissa of the predicted offset, tv i represents the ordinate of the predicted offset, d ij represents the offset distance between the center point offset of the tool i and the person j and the predicted offset.
3. The method for determining the personnel to whom the tools for power operation personnel according to claim 1 belong, characterized in that, The determination process of the trained improved Yolo v5 network specifically includes: Obtain training samples; the training samples include multiple training pictures, and each training picture includes multiple personnel and multiple tools; Annotate each training image in the training sample to obtain the human target set of each training image and the tool target subset of each training image; the human target set of the training image includes the central point position coordinates, detection box width, and detection box height of each person in the training image, and the tool target subset of the training image includes the central point position coordinates, category, detection box width, and detection box height of each tool in the training image; For any tool in any training image in the training sample, determine the actual owner of the tool according to the distance between the central point position coordinate of the area where the tool is preset to belong to the person and the central point position coordinates of each person in the training image; Obtain the predicted offset of the tool according to the central point position coordinate of the actual owner of the tool and the central point position coordinate of the area where the tool is preset to belong to the person; Determine the predicted offset, central point position coordinate, category, detection box width, and detection box height of each tool in the training image as the tool target set of the training image; Use the tool target set of each training image and the human target set of each training image in the training sample as the training set to train the improved Yolo v5 network to obtain a trained improved Yolo v5 network.
4. A method for determining the personnel to whom the tools of power operation personnel described in claim 3 belong, characterized in that, The step of determining the actual owner of the tool according to the distance between the central point position coordinate of the area where the tool is preset to belong to the person and the central point position coordinates of each person in the training image specifically includes: Calculate the distance between the area where the tool is preset to belong to the person and each person in the training image according to the central point position coordinate of the area where the tool is preset to belong to the person and the central point position coordinates of each person in the training image; Determine the person corresponding to the smallest distance as the actual owner of the tool.
5. A system for determining the personnel to whom the tools of electric power operation personnel belong, characterized in that, Including: An image acquisition module for acquiring an image to be recognized; the image to be recognized includes multiple people and multiple tools; A set determination module, configured to input the image to be recognized into a trained improved Yolo v5 network to obtain a human target set of the image to be recognized and a tool target set of the image to be recognized; the human target set of the image to be recognized includes the central point position coordinates, detection box width, and detection box height of each person in the image to be recognized, and the tool target set of the image to be recognized includes the central point position coordinates, category, detection box width, detection box height, and predicted offset of each tool in the image to be recognized; the improved Yolo v5 network includes a backbone network and a feature fusion network connected in sequence, and three prediction branches respectively connected to the output end of the feature fusion network; each prediction branch is a Conv, and the backbone network includes a first CBS, a second CBS, a first C3, a third CBS, a second C3, a third C3, a fourth CBS, a fourth C3, a fifth C3, a sixth C3, a fifth CBS, a seventh C3, an SPPF, and a sixth CBS connected in sequence, and the feature fusion network includes: a first Up, a first Concat, an eighth C3, a seventh CBS, a second Up, a second Concat, a ninth C3, an eighth CBS, a third Concat, a tenth C3, a ninth CBS, a fourth Concat, and an eleventh C3 connected in sequence; the output end of the third C3 is also connected to the input end of the second Concat, the output end of the sixth C3 is connected to the input end of the first Concat, the output end of the sixth CBS is respectively connected to the input end of the first Up and the input end of the fourth Concat, the output end of the seventh CBS is also connected to the input end of the third Concat, the output end of the ninth C3 is connected to the first branch of the prediction network, the output end of the tenth C3 is also connected to the second branch of the prediction network, and the output end of the eleventh C3 is connected to the third branch of the prediction network; A result determination module, configured to, for any tool in the image to be recognized, calculate the offset distance between the central point offset of the tool and each person and the predicted offset, and determine that the tool belongs to the person corresponding to the smallest offset distance according to the central point position coordinates of the tool, the predicted offset of the tool, and the central point position coordinates of each person in the image to be recognized.
6. The system for determining the personnel to whom the tools for power operation personnel described in claim 5 belong, characterized in that, The result determination module specifically includes: An offset distance determination unit for calculating the offset distance between the center point offset of the tool i and the person j and the predicted offset according to the formula where px j represents the abscissa of the center point of the person j, tx i represents the abscissa of the center point of the tool i, py i represents the ordinate of the center point of the person j, ty i represents the ordinate of the center point of the tool i, tu i represents the abscissa of the predicted offset, tv i represents the ordinate of the predicted offset, and d ij represents the offset distance between the center point offset of the tool i and the person j and the predicted offset.
7. The system for determining the personnel to whom the tools for electric power operation personnel described in claim 5 belong, characterized in that, It further includes: A sample acquisition module, configured to acquire training samples; the training samples include multiple training pictures, and each training picture includes multiple people and multiple tools; A set annotation module, configured to annotate each training picture in the training samples to obtain a human target set of each training picture and a tool target subset of each training picture; the human target set of the training picture includes the central point position coordinates, detection box width, and detection box height of each person in the training picture, and the tool target subset of the training picture includes the central point position coordinates, category, detection box width, and detection box height of each tool in the training picture; The true result annotation module is used to determine the true owner of any tool in any training image of the training samples according to the distance between the central point position coordinates of the area where the tool's preset owner is located and the central point position coordinates of each person in the training image; The predicted offset calculation module is used to obtain the predicted offset of the tool according to the central point position coordinates of the true owner of the tool and the central point position coordinates of the area where the tool's preset owner is located; The tool target subset determination unit is used to determine the predicted offset, central point position coordinates, category, detection frame width, and detection frame height of each tool in the training image as the tool target set of the training image; The training module is used to train the improved Yolo v5 network with the tool target set of each training image in the training samples and the human target set of each training image as the training set to obtain a trained improved Yolo v5 network.
8. The system for determining the personnel to whom the tools for electric power operation personnel according to claim 7 belong, characterized in that, The true result annotation module specifically includes: The distance calculation unit is used to calculate the distance between the area where the tool's preset owner is located and each person in the training image according to the central point position coordinates of the area where the tool's preset owner is located and the central point position coordinates of each person in the training image; The true result annotation unit is used to determine the person corresponding to the smallest distance as the true owner of the tool.
Citation Information
Patent Citations
Behavior identification method and device for multi-category engineering vehicle
CN112800934A
Mask wearing state identification method, device and equipment and readable storage medium
CN112818953A