A method for image recognition of transmission lines
By combining Mask R-CNN, human-like concept learning and Bayesian semantic network methods, the image recognition results of the patrol robot on the transmission line are corrected and decomposed, solving the problem of low image recognition accuracy of the patrol robot and insufficient recognition effect under extreme conditions, achieving higher recognition effect and semantic logic.
Patent Information
- Application Number
- CN202011348272.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-26
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2040-11-26
AI Technical Summary
The image recognition accuracy of the inspection robot is not high, and intelligence is still in perceptual intelligence rather than cognitive intelligence, and the recognition effect is insufficient in extreme weather and extreme situations.
Images of transmission line equipment are collected, rough identification is performed using the Mask R-CNN method, and a correction mechanism is constructed based on human-like concept learning and the correlation between transmission line equipment, relationship correction is made to correct the recognition results, and Bayesian semantic network is used to decompose the final recognition results.
It improves the semantic logic of the recognition results, improves the recognition effect, and enhances the recognition ability under extreme conditions.
Smart Images

Figure CN112733592B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power equipment image recognition, and particularly to a method for image recognition of transmission lines. Background Art
[0002] Transmission lines play a crucial role in the power system. The operation status of power equipment on transmission lines directly affects the safe and stable operation of the power grid. Therefore, great importance must be attached to the daily operation and maintenance work of transmission lines. The inspection work of transmission lines is highly dangerous and repetitive. Moreover, with the rapid economic development, the scale of the power system is constantly expanding, the voltage level is continuously increasing, and the requirements for the stability of the power system are also continuously rising. Traditional manual inspection is time-consuming and laborious, and has certain potential safety hazards and problems of scattered inspection quality, making it increasingly difficult to meet the requirements. The intelligent inspection robots currently under research and testing, combined with a variety of high-tech technologies, are expected to comprehensively replace traditional manual inspection and become an inevitable trend in future development.
[0003] At the same time, currently, due to the prominent problem of "structural shortage of personnel" in power grid enterprises, the shortage of high-quality operation and maintenance staff is serious, and the application of intelligent inspection robots is of great significance for solving this problem. Worldwide, intelligent inspection robots have broad application prospects and markets. Intelligent inspection robots usually include a navigation system, a vision system, a detection system, a communication system, a control system, and a motion system, etc. This article mainly focuses on the research of the vision system of inspection robots, that is, the target detection and image recognition of power equipment on transmission lines. Summary of the Invention
[0004] The purpose of this part is to outline some aspects of the embodiments of the present invention and briefly introduce some preferred embodiments. Some simplifications or omissions may be made in this part, as well as in the abstract and title of the present application, to avoid obscuring the purpose of this part, the abstract, and the title. However, such simplifications or omissions cannot be used to limit the scope of the present invention.
[0005] In view of the above existing problem of low image recognition accuracy, the present invention is proposed.
[0006] Therefore, the technical problems solved by the present invention are: the image recognition accuracy of inspection robots still needs to be improved, the intelligence of inspection robots is still in the stage of perceptual intelligence rather than cognitive intelligence, and inspection robots still have deficiencies when facing some extreme weather and extreme situations.
[0007] To solve the above technical problems, the present invention provides the following technical solutions: collect images of transmission line equipment; roughly identify the collected transmission line images by using the Mask R-CNN method; use human-like concept learning and combine the relevant relationships between transmission line equipment to construct a correction mechanism to correct the relationship of the recognition results; use the Bayesian semantic network for decomposition to obtain the final recognition result.
[0008] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: the using the Mask R-CNN method to recognize the collected transmission line images includes that the Mask R-CNN method first uses the convolutional layer of the object detection algorithm to extract the overall features of the transmission line equipment images, generates regions of interest through the region proposal network, applies ROI Align to ensure the alignment and consistency between the candidate boxes and the extracted features, and uses the classification and regression layers to obtain the object categories and bounding box regressions of the regions of interest, and finally realizes the classification and segmentation of each pixel through the Mask branch; the rough recognition result of the Mask R-CNN includes the classification result of the object, the bounding box, and its corresponding confidence level.
[0009] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: the human-like concept learning includes, by learning the idea of human-like concept learning, using the relevant relationships existing between transmission line electrical equipment to correct the recognition results of the Mask R-CNN; applying more information and relationships to make the recognition results more reasonable and reliable, so that some instance objects with weak intrinsics due to environmental factors have strong semantic support. For the modeling of relationships, corresponding correction mechanisms are constructed from two perspectives: strong correlation relationships and weak correlation relationships.
[0010] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: the relevant relationships existing between the transmission line electrical equipment include current transformers and support insulators, capacitors and support insulators, GIS and bushings, switches and support insulators, and capacitors and capacitors.
[0011] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: constructing the correction mechanism from strong correlation relationships includes that 5 aspects are involved in the process of constructing the correction mechanism from strong correlation relationships, including the Mask R-CNN recognition results, neighborhood, position normalization, relative size, and position relationship; wherein the Mask R-CNN recognition results: it is assumed that there are N recognition results as d (1) , d (2) ,... d (N) , and each recognition result is called an instance, and each instance has three attributes, the classification result Bounding box and confidence Bounding box attribute b b =(x, y, w, h), where (x, y), w, and h represent the center coordinates, width, and height of the bounding box respectively, and d s ∈(0, 1); the neighborhood is defined as:
[0012]
[0013] where: α w , α h > 1 is a parameter for controlling the size of the instance neighborhood; the position normalization method includes setting the normalized bounding box of the target to be corrected as: And at this time for b (j) ∈ε(d (i) ) satisfies the following formula:
[0014]
[0015] The relative size includes a calculation method that is the ratio of the areas of the bounding boxes of two instances. In a strong correlation relationship, it is the area ratio of the determinant instance to the determined instance; in a weak correlation relationship, it is the target instance to other instances within the neighborhood; the position relationship includes: defining the position relationships of "above", "below", "left and right" based on the normalized spatial relationship between instances in the image, which together with the relative size are used as criteria for determining whether there is a strong correlation relationship between instances.
[0016] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: the construction of the correction mechanism from the strong correlation relationship further includes that for two instances with a strong correlation relationship, if one party depends on the correct recognition of the other party to be corrected based on the strong correlation relationship; some of the rough recognition results are correct and have a relatively high confidence, some of the targets are relatively large themselves, and their neighborhoods are also relatively large. A certain correction target may be within several neighborhoods at the same time, and simply relying on its presence or absence will cause ambiguity in the correction.
[0017] As a preferred embodiment of the transmission line image recognition method of the present invention, wherein: the relationship correction of the recognition result includes, for the confidence of the recognition result of the Mask R-CNN network, correcting the recognition result with a confidence lower than 0.90 through a strong correlation relationship; for instances that meet the position conditions and relative size conditions, it is necessary to consider whether there is a strong correlation relationship with other instances within the neighborhood and whether the relationship is "strong" enough. Its formula is expressed as follows:
[0018]
[0019] where: s is the confidence level, which represents the unreliability degree of the rough recognition result, and are both one-dimensional vectors, is a sparse vector that stores in which strong correlation relationships the target to be corrected is included, is the weight occupied by each strong correlation relationship, and the weights are all greater than 1. y is the necessity of strong relationship correction. When y is less than 1, it means that there is no strong correlation relationship. When y is higher than a certain threshold, it means that the recognition result itself is not reliable enough and there is a certain strong correlation relationship, and it can be corrected according to the strong correlation relationship.
[0020] As a preferred solution of the transmission line image recognition method described in the present invention, where: the Bayesian semantic network includes that the Bayesian semantic network is composed of nodes and edges representing objects and relationships between objects respectively. In strong correlation relationships, the neighborhood is as small as possible and is not necessarily a standard square or circle, while in weak correlation relationships, the neighborhood is generally taken as a circle or rectangle and has a relatively larger area compared to the strong correlation relationships; define d (i) the instance set within the neighborhood of Since the classification result of d (i) may be incorrect, especially when the confidence level is low, the correct classification corresponding to the instance d should be (i)
[0021] As a preferred solution of the transmission line image recognition method described in the present invention, where: the construction of the correction mechanism from weak correlation relationships includes that, based on the human-like concept learning, the detection result is corrected from the perspective of object relationships. It is difficult to infer the correct categories of all instances at once. Therefore, the correction of instances in the weak correlation relationship correction method is completed step by step; there is set an instance whose correct classification is and the neighborhood set is From the perspective of object relationships, the instance and its neighborhood are in the highest possibility when the correct class is given, and it can be written as:
[0022]
[0023] Advantages of the present invention: A method for identifying transmission line images is proposed, with the image recognition of power equipment as the application scenario. By constructing a relationship model between power equipment on the transmission line, a strong correlation relationship correction method based on a strong correlation relationship library and a weak relationship correction method based on Bayesian inference are established. Using the spatial, positional, logical, etc. relationships existing between power equipment on the transmission line, the recognition results of Mask R-CNN are corrected, improving the semantic logic of the recognition results and enhancing the recognition effect. Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings. Among them:
[0025] Figure 1 It is a schematic flowchart of the transmission line image recognition method described in the first embodiment of the present invention;
[0026] Figure 2 It is an overall framework diagram of the transmission line image recognition method described in the first embodiment of the present invention;
[0027] Figure 3 It is a recognition result diagram of the Mask R-CNN of the transmission line equipment image for the transmission line image recognition method described in the first embodiment of the present invention;
[0028] Figure 4 It is a transmission line scene diagram of the transmission line image recognition method described in the first embodiment of the present invention;
[0029] Figure 5 It is an example boundary box and neighborhood graph of the transmission line image recognition method described in the first embodiment of the present invention (taking as an example);
[0030] Figure 6 It is a Bayesian semantic network graph constructed according to the recognition results for the transmission line image recognition method described in the first embodiment of the present invention;
[0031] Figure 7 It is a strong correlation relationship correction process diagram of the transmission line image recognition method described in the second embodiment of the present invention;
[0032] Figure 8 It is a Bayesian inference process diagram of the transmission line image recognition method described in the second embodiment of the present invention;
[0033] Figure 9The Bayesian correction process of the transmission line image recognition method described in the second embodiment of the present invention;
[0034] Figure 10 The Mask R-CNN recognition result diagram of the transmission line image recognition method described in the third embodiment of the present invention;
[0035] Figure 11 The strong correlation relationship correction diagram of the transmission line image recognition method described in the third embodiment of the present invention;
[0036] Figure 12 The Bayesian inference correction of the transmission line image recognition method described in the third embodiment of the present invention Figure 1 ;
[0037] Figure 13 The Bayesian inference correction of the transmission line image recognition method described in the third embodiment of the present invention Figure 2 ;
[0038] Figure 14 The relationship diagram between the recognition effect and the sample size of the transmission line image recognition method described in the third embodiment of the present invention. Detailed implementation manners
[0039] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following will describe the detailed implementation manners of the present invention with reference to the accompanying drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0040] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Those skilled in the art can make similar extensions without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below.
[0041] Secondly, the so-called "one embodiment" or "embodiment" herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation manner of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it an individual or alternative embodiment that is mutually exclusive with other embodiments.
[0042] The present invention will be described in detail with reference to the schematic diagrams. When describing the embodiments of the present invention, for the convenience of explanation, the cross-sectional views showing the device structure will be enlarged locally not in accordance with the general scale, and the schematic diagrams are only examples and should not limit the scope of protection of the present invention herein. In addition, in actual production, three-dimensional spatial dimensions including length, width, and depth should be included.
[0043] Meanwhile, in the description of the present invention, it should be noted that the orientation or positional relationship indicated by terms such as "upper, lower, inner, and outer" is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be construed as a limitation to the present invention. In addition, the terms "first, second, or third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0044] Unless otherwise clearly defined and limited in the present invention, the terms "mounted, connected, and coupled" shall be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can also be a mechanical connection, an electrical connection, or a direct connection, and can also be indirectly connected through an intermediate medium, or can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0045] Embodiment 1
[0046] Referring to Figures 1 - 6 , which is the first embodiment of the present invention. This embodiment provides a method for identifying transmission line images, including:
[0047] S1: Collect images of transmission line equipment. It should be noted that
[0048] Use an inspection robot to take pictures of transmission line equipment.
[0049] S2: Coarsely identify the collected transmission line images using the Mask R-CNN method. It should be noted that
[0050] Referring to Figure 3, using the Mask R-CNN method to identify the collected transmission line images includes that the Mask R-CNN method first uses the convolutional layer of the object detection algorithm (Faster RCNN) to extract the overall features of the transmission line equipment images, generates regions of interest (ROIs) through the region proposal network (RPN), applies ROI Align to ensure the alignment and consistency between the candidate boxes and the extracted features, and uses the classification and regression layers to obtain the object categories and bounding box regression of the ROIs. Finally, the classification and segmentation of each pixel are achieved through the Mask branch; the rough recognition result of Mask R-CNN includes the classification result of the object, the bounding box, and its corresponding confidence level.
[0051] Furthermore, Mask RCNN generally follows the idea of Faster RCNN as a whole. The feature extraction adopts the ResNet-FPN architecture and adds a fully convolutional network (FCN) branch for the segmentation task of the candidate regions, that is, Mask.
[0052] Among them, Faster RCNN is a two-stage object detection algorithm, including the Region Proposal Network in the first stage and the bounding box regression and classification in the second stage. Faster RCNN uses the CNN architecture to extract image features, then uses the RPN (Region Proposal Network) to extract the ROIs (regions of interest), and then uses ROI pooling to make all these ROIs into a fixed size, and then passes them to the fully connected layer for bounding box regression and classification prediction. Simply put, in the first stage, the bounding boxes of the candidate objects are proposed, and in the second stage, features are extracted from each candidate box and classified and the bounding box is regressed. Its feature is that it can share the data used in the two stages for faster inference.
[0053] FPN (Feature Pyramid Network) is a multi-scale object detection method. The FPN structure includes three parts: bottom-up, top-down, and lateral connections. This structure can fuse the features of each level, making it have both strong semantic information and strong spatial information at the same time. FPN is a general architecture, and ResNet is a backbone network based on FPN.
[0054] MaskR-CNN and Faster RCNN adopt the same two steps: First, find the RPN, and then classify, localize each RoI found by the RPN, and find the binary mask. This is different from other networks that first find the mask and then classify. In addition to adding a Mask (essentially a convolutional layer) after the pooling layer for semantic segmentation, an important improvement is to use ROIAlign instead of ROI pooling. The problem with FasterR-CNN is that the feature map is mis-aligned with the original image, so it will affect the detection accuracy. MaskR-CNN proposes the method of ROIAlign to replace ROI pooling, and ROIAlign can retain the approximate spatial position.
[0055] S3: Use human-like concept learning and combine the relevant relationships between transmission line equipment to construct a correction mechanism to correct the relationship of the recognition results. It should be noted that
[0056] Human-like concept learning includes learning the idea of human-like concept learning and using the relevant relationships existing between transmission line electrical equipment to correct the recognition results of MaskR-CNN. Its core idea is to use the relationship between the objects in the target neighborhood and the target to improve the semantic logic of the MaskR-CNN recognition results, and infer and correct the recognition results below a certain confidence level according to probability; applying more information and relationships makes the recognition results more reasonable and reliable, so that some instance targets with weak intrinsics due to environmental factors have strong semantic support. For the modeling of relationships, corresponding correction mechanisms are constructed from two perspectives of strong correlation relationships and weak correlation relationships.
[0057] Furthermore, the relevant relationships existing between transmission line electrical equipment include constructing a correction mechanism by combining the relationships between 5 types of transmission line equipment and components and the confidence of the rough recognition results of Mask R-CNN. The relationships between the 5 types of transmission line equipment and components include: CT (current transformer) and support insulator, capacitor and support insulator, GIS and bushing, switch and support insulator, and capacitor and capacitor.
[0058] Even further, referring to Figures 4 - 5 , constructing the correction mechanism from strong correlation relationships includes 5 aspects involved in the process of constructing the correction mechanism from strong correlation relationships, including Mask R-CNN recognition results, neighborhood, position normalization, relative size, and position relationship; the bounding box is a rectangular box containing the instance segmentation result of the target, expressed as coordinates (x 1 , y 1 ; x 2 , y 2), the upper left corner coordinates of the rectangular box are (x 1 , y 1 ), and the lower right corner coordinates of the rectangular box are (x 2 , y 2 ). Thus, the spatial position and size of the rectangular box can be determined. The confidence level is directly given by Mask R-CNN, indicating the reliability of the recognition result. Obviously, the recognition result with a higher confidence level is more likely to be correct. In the process of constructing the correction model, the present invention makes the following agreements: It is assumed that there are N recognition results d (1) , d (2) , … d (N) . Each recognition result is called an instance, and each instance has three attributes: classification result bounding box and confidence level The bounding box attribute b b = (x, y, w, h), where (x, y), w, and h respectively represent the center coordinates, width, and height of the bounding box, and d s ∈ (0, 1); Neighborhood is defined as:
[0059]
[0060] where: α w , α h > 1 is a parameter for controlling the size of the instance neighborhood. When the (j) of another instance d intersects with , it is considered that the instance d (j) is within the neighborhood of the instance d (i) . Whether there is an intersection can be calculated from the diagonal coordinates of the bounding box (provided by Mask R-CNN); The normalization method of the position includes setting the bounding box after normalizing the target to be corrected as: And at this time, for b (j) ∈ ε(d (i) ) satisfies the following formula:
[0061]
[0062] The relative size includes the calculation method as the ratio of the areas of the bounding boxes of two instances. In a strong correlation relationship, it is the area ratio of the determinant to the determined; In a weak correlation relationship, it is the target instance to other instances within the neighborhood; The position relationship includes: Defining the position relationships of "above", "below", "left and right" according to the normalized spatial relationship between instances in the image, which together with the relative size are used as criteria for determining whether there is a strong correlation relationship between instances;
[0063] Factors to be considered when constructing a correction mechanism include: for two entities with a strong correlation, if one party depends on the correct recognition of the other party to be corrected based on the strong correlation; some rough recognition results are correct and have a relatively high confidence level, such as above 0.95, but in a photo, due to reasons such as the shooting angle, it may generate an incorrect strong relationship. At this time, if correction is performed, it may lead to incorrect correction; some targets are relatively large themselves, and their neighborhoods are also relatively large. A correction target may be within several neighborhoods at the same time. Simply relying on its presence or absence will make the correction result in an ambiguous recognition. Relationship correction includes, for the confidence level of the recognition result of the Mask R-CNN network, for recognition results with a confidence level lower than 0.90, perform strong correlation relationship correction; for instances that meet the position conditions and relative size conditions, consider whether there is a strong correlation relationship with other instances in the neighborhood, and whether the relationship is strong enough. Its formula is as follows:
[0064]
[0065] where: s is the confidence level, represents the unreliability degree of the rough recognition result, and are both one-dimensional vectors, is a sparse vector, storing which strong correlation relationships the target to be corrected is included in, is the weight occupied by each strong correlation relationship, and the weights are all greater than 1. y is the necessity of strong relationship correction. When y is less than 1, it means there is no strong correlation relationship. When y is higher than a certain threshold, it means that the recognition result itself is not reliable enough and there is a certain strong correlation relationship, and it can be corrected based on the strong correlation relationship;
[0066] Constructing a correction mechanism from weak correlation relationships includes, based on human-like concept learning, correcting the recognition result from the perspective of object relationships. It is relatively difficult to infer the correct categories of all instances at once. Therefore, in the weak correlation relationship correction method, the correction of instances is completed step by step; assume there is an instance Its corresponding correct classification is The neighborhood set is From the perspective of object relationships, the instance and its neighborhood are in the highest possibility when the correct class is given, and it can be written as:
[0067]
[0068] where, represents the correct classification of the instance, and g c represents the classification set.
[0069] S4: Decompose using the Bayesian semantic network to obtain the final recognition result. It should be noted that
[0070] Referring to Figure 6 , the Bayesian semantic network includes that before object relationship modeling, it is necessary to clearly and efficiently represent a scene composed of many objects and their relationships. As a representation form, a graph contains rich relationship information between elements. Therefore, a scene graph constructed based on the Mask R-CNN detection result is proposed, which is called the Bayesian semantic network. The Bayesian semantic network consists of nodes and edges that respectively represent objects and the relationships between objects; it is easy to understand that a single object is represented by a node, and regarding the relationships between objects, some explanations are needed. The relationship (that is, the edge in the graph) is defined by the two nodes it connects. From a mathematical perspective, there should be an edge between every two nodes. However, this strict rule makes the network quite complex, and the processes of learning and prediction become very time-consuming; in addition, for a single object, not all relationships are equally important. The closer two objects are, the stronger their connection is. When learning a small number of labeled samples, it is necessary to focus on the stronger and more useful relationships rather than all relationships. Therefore, nodes are only connected to their adjacent nodes in the Bayesian semantic network;
[0071] After defining the Bayesian semantic network, assume that there are N recognition results in a graph, d (1) , d (2) , … d (N) , each instance has three attributes, the classification result bounding box and confidence In strong correlation relationships, we hope that the neighborhood of the target contains instances with strong correlation relationships that are closely related to it. Therefore, the neighborhood is as small as possible and not necessarily a standard square or circle. In weak correlation relationships, we hope that the neighborhood of the target contains as many related and reliable recognition results as possible. Therefore, the neighborhood is generally taken as a circle or rectangle, and the relative area is larger than that in strong correlation relationships; define the set of instances within the neighborhood of d (i) Since the classification result of d Since d (i) the classification result of may be incorrect, especially when the confidence is low, the correct classification corresponding to the instance d (i) should be
[0072] Embodiment 2
[0073] Referring to Figures 7 - 9 , this is the second embodiment of the present invention, and a simple schematic diagram is used to represent the process of correcting using strong correlation relationships and the process of correcting using weak correlation relationships;
[0074] Reference Figure 7 is a process corrected using strong correlation relationships. The correction consists of 4 steps, the order of which is given by Roman numerals. In all steps, white nodes represent the original results of Mask R-CNN, blue nodes represent reliable or already corrected results, and red nodes represent results that need to be corrected. Note: Needing correction does not mean misidentification. It just meets the correction conditions set by the present invention. It can be considered that the identification results represented by red nodes are unreliable, but not necessarily incorrect; from Figure 7 it can be seen that by using strong correlation relationships, some of the instances to be corrected in the Mask R-CNN identification results have been corrected and become reliable identification results. The remaining results that cannot be corrected relying on strong correlation relationships rely on weak correlation relationships for Bayesian inference; among them, in Figure (II), there is a unidirectional strong correlation relationship between d 3 and d 4 as shown in the figure. However, since d 3 is an unreliable instance, and according to the directionality, the strong correlation relationship between d 3 and d 4 cannot be used to correct d 3 ; although there is a strong correlation relationship between d 8 and d 9 and it meets the directionality condition, but since d 9 itself has a relatively high confidence level (but less than 0.9), is small, and the calculated y value does not meet the strong correlation relationship threshold, so no strong correlation relationship correction is performed.
[0075] Reference Figures 8 - 9 is a process corrected using weak correlation relationships. First, discuss the correction order in a single picture. When inferring the correct category of a node, it is obvious that if there are more reliable instances in its neighborhood, the inference is more likely to be accurate. That is to say, if is more reliable, then the (i) inferred from the maximum likelihood probability P(d (i) )|g c ) is more likely to be correct. Reliable instances come from two ways: it has a high detection confidence level, has been corrected by the model; during the correction process, for unreliable instances, even if they are in the neighborhood of the target to be corrected, their relationship with the target to be corrected is not considered. Figures Figure 8 and Figure 9 respectively illustrate the inference process and correction process of human concept learning. Weak correlation relationship correction uses the correction results of strong correlation relationships as input;
[0076] For N instances d (1) ,d (2) ,…d(N) , with 0.90 as the confidence threshold t, for Consider the instance (i) is reliable, or for instances that have been corrected for strong correlation, it is considered reliable. Instances of (j) , and correct them in turn; assuming that there is a reliable instance in a graph and the unreliable instances that will be corrected in turn The correction process is as follows:
[0077] For j = 1, 2, ... n,
[0078]
[0079] Where: S j is a reliable instance set consisting of the initial reliable instance and the corrected instances The reasoning about the target only relies on nearby instances, which means that the closer the two objects are, the stronger the connection between them.
[0080] Examples to be corrected The correction order of is determined by the number of reliable instances in its neighborhood. For j = 1, 2, ... n,
[0081]
[0082] t∈({1,2,…N}-{r 1 ,r 2 ,…,r m}-{c 1 ,c 2 ,…,c j-1})
[0083] Where: S j is a reliable instance set, #(S j ∩ε(d (t) )) is d (t) The number of reliable instances nearby, d (t) The rest are all instances that need to be corrected; it is not feasible to directly deal with the equation now, mainly for two reasons: it is impossible to assume the probability distribution, and the probability distribution is in a very high-dimensional space; this is very unfavorable for small sample learning problems, so here the equation is decomposed and simplified, for clarity, with Figure 7 Take step II in the example to illustrate. For this step, write the formula as:
[0084]
[0085] P(d 5 ,d6 , d 8 , d 9 | g c ) = P(d 5 , d 6 , d 8 , d 9 , g c ) / P(g c )
[0086] Then the joint probability can be further written as:
[0087] P(d 5 , d 6 , d 8 , d 9 , g c ) = P(d 9 , g c ) P(d 5 | d 9 , g c ) P(d 6 | d 5 , d 9 , g c ) P(d 8 | d 5 , d 6 , d 9 , g c )
[0088] Note that there is no connection between d 6 and d 8 , so P(d 8 | d 5 , d 6 , d 9 , g c ) can be simplified to P(d 8 | d 5 , d 9 , g c ). In fact, since not all nodes are connected, all conditional probabilities in the correction process can be simplified in a similar way; the normalized bounding box d' b = (x', y', w', h'), to make the relationship model consistent with human cognition, here d' b is decomposed into the position d' p = (x', y'), and the size d' size = (w', h'), so for an instance, d c , d' p and d' size are considered in the correction process;
[0089] However, using the above three attributes simultaneously can also cause problems: the probability distribution is difficult to assume, so it is not friendly to small-sample learning. Therefore, to solve this problem based on human understanding of the relationships between electrical equipment, assume there is a picture of a busbar, tubular busbars that provide support, and column insulators. Their attributes can be described hierarchically as follows: Class: Tubular busbars and column insulators tend to be together; Location: If there are two instances of a tubular busbar and a column insulator, the column insulator is very likely to be below the tubular busbar; Size: If there are two devices, namely a tubular busbar and a column insulator, their sizes are very likely to be similar. The equation for this process is given here. For one instance, the probability that another instance appears nearby can be written as:
[0090]
[0091] Example 3
[0092] Refer to Figures 10 - 14 , which is the third embodiment of the present invention. To better verify and illustrate the technical effects adopted in the method of the present invention, in this embodiment, only the Mask R-CNN network is selected to be tested on a small-sample dataset, and the experimental results are compared by means of scientific demonstration to verify the actual effects of this method;
[0093] When conducting the experiment, the parameters of the software and hardware used are shown in Table 1 below:
[0094] Table 1: Data of software and hardware used for the experiment.
[0095]
[0096]
[0097] To verify whether the correction method based on a strongly correlated relationship library and the correction method based on Bayesian inference proposed by the present invention have improved recognition accuracy compared to only using the Mask R-CNN network on a small-sample dataset; a total of 750 pictures are selected from the dataset. The pictures are all taken from the one-week inspection task of a transmission line inspection robot. Among them, 150 pictures are used as the test set, 600 pictures are used as the training set and the 4-fold cross-validation set. To eliminate the adverse effects caused by accidental factors, the K-fold cross-validation method is used to learn and train these parameters;
[0098] The K-fold cross-validation method evenly divides the entire dataset into K non-overlapping subsets. Each time, one subset is selected as the test set from the divided subsets, and the other K - 1 subsets are used as the training set. The models or parameters learned during the K training processes are averaged and then adopted. The total capacity of the dataset is 750, and the parameters required for training using the 4-fold cross-validation method are selected; during the experiment, we found that for the recognition process of the Mask R-CNN network, since it includes two parts: object detection and classification, there are not only problems such as whether the object is correctly classified, but also problems such as the object not being detected, detecting the wrong object, and inaccurate bounding box framing. Only measuring the precision of the transmission line image recognition result cannot accurately describe the quality of the recognition effect. Therefore, the F1 score is used to describe the rough recognition result of Mask R-CNN and the final recognition effect after strong relationship correction and Bayesian correction; among them, the F1 score is an index used in statistics to measure the precision of a binary classification model. It takes into account both the precision and recall of the classification model. The F1 score can be regarded as a harmonic mean of the model's precision and recall. The F1 Score is defined as follows:
[0099]
[0100] where: precision is the accuracy rate, that is, the ratio of the correctly recognized results to all detected objects, and recall is the recall rate, which represents the ratio between the correctly recognized objects and the real objects existing in the sample.
[0101] Refer to Figure 10 , taking a transmission line image as an example, the correction process of the transmission line image is given. Among them, blue represents reliable recognition results or already corrected instances, and red represents instances to be corrected. Here, the confidence threshold T = 0.90 is used as the confidence threshold. When the confidence of the MaskR-CNN recognition result is greater than 0.9, it is considered a reliable recognition result and no correction is performed; in Figure 10 , there are 4 types of recognition results: line insulators, CTs, support insulators, and busbars. Among the rough recognition results obtained by Mask R-CNN, the support insulators of 2 busbars and 2 CTs do not meet the confidence conditions and are regarded as targets to be corrected. There is a strong correlation between the CT and the support insulator to be corrected, and at this time the CT is reliably recognized. The CT and the instance below it meet the relative size condition and position condition, and the instance to be corrected below the CT can be corrected by the strong correlation relationship according to the strong correlation relationship; after the strong correlation relationship correction, refer to Figure 11, the instance classification labels below the 2 CTs are corrected to "insulator" and become reliable instances as the corrected instances. Since there is no matching strong correlation relationship in the strong correlation relationship library for the 2 busbars, the recognition results are retained and they remain as instances to be corrected. The results after the strong correlation relationship correction are used as the input to the Bayesian inference correction network (weak correlation relationship correction). In the strong correlation relationship correction mechanism, only the detection and judgment of the strong correlation relationship are performed on all the recognition results in the image, and there is no order in the correction process. In the Bayesian inference correction process based on the Bayesian semantic network and human-like concept learning, the correction is carried out one by one according to the number of reliable instances in the neighborhood of the target to be corrected. Refer to Figure 12 , there are more reliable instances in the neighborhood of the busbar in the lower right corner of the image. Semantically, it has more reliable relationships. Therefore, it is corrected by inference first, and after the correction, it also participates in the correction as a reliable instance. It should be noted here that in human cognition, although it is common sense for busbars to appear side by side in a transmission line, since the 2 instances with the true value of busbar in the figure are both considered unreliable instances, neither of them is used regardless of whether their classification labels are "busbar"; From Figure 13 it can be seen that finally, the last instance to be corrected in the figure is corrected by inference. It should be noted that at this time, since the busbar in the lower right corner has been corrected and becomes a reliable instance, the relationship between the two busbars is used in the correction process of the busbar in the upper right corner, that is, the conditional probability between them is calculated. Figure 13 is the final correction result.
[0102] The final recognition results for the test set are as follows. Table 2 shows the comparison of the F1 scores of only using Mask R-CNN and using the correction schemes of strong correlation relationship and weak correlation relationship (collectively referred to as the method of human-like concept learning) here;
[0103] Table 2: Comparison of F1 scores of recognition results.
[0104]
[0105]
[0106] As can be seen from Table 2, the correction method based on a strongly correlated relationship library and the correction method based on Bayesian inference proposed in the present invention have improved the overall and F1 scores of each device compared to only using the recognition results of Mask R-CNN; from the results, the image recognition method based on humanoid concept learning proposed in the present invention, when applied to the daily operation and maintenance inspection images of real transmission line inspection robots, corrects some results that cannot be correctly recognized by only using methods such as neural network training. Compared to Mask R-CNN, the overall F1 score has increased from 0.6951 to 0.7921, taking a step closer to reaching the engineering application level; in addition, as can be seen from Table 2, for different devices, the F1 scores are different and vary greatly, and there are the following reasons:
[0107] In the inspection images of transmission line stations, different device types appear with different frequencies. Devices such as insulators appear extremely frequently, almost in every image and often there are many in one image, while devices such as GIS, switches, and lightning arresters appear less frequently. This actually has a greater impact on image recognition research, especially for recognition methods such as deep learning neural networks that require large sample data for learning. The transmission line image recognition method proposed in the present invention performs semantic enhancement on the recognition results based on the relationships existing between devices, and also has a relatively high F1 score for the recognition results of some devices with low appearance frequencies, and has a relatively large improvement compared to Mask R-CNN. For example, for capacitor devices in the table, the F1 score of humanoid concept learning has increased from 0.4989 to 0.7783 compared to Mask R-CNN;
[0108] In the inspection images of transmission lines, affected by environmental and scene factors, as well as the shooting position and pose, the devices often cannot be fully presented in the image, and there are interferences such as occlusion, stacking, and poor lighting. This brings great difficulties to the learning of neural networks in small samples. For humanoid concept learning, we often do not need the recognition object to present sufficient eigenfeatures. As long as we use the relationship between its reliable recognition results around it, we can predict and infer what it should be;
[0109] Different from traditional image recognition research, in the image recognition of transmission line devices, there are also distinctions between "easy to recognize" and "difficult to recognize" for different recognition objects. Some devices are relatively difficult to be confused with other devices, that is, they have strong eigenfeatures; while some devices such as certain support insulators and bushings, support insulators and columnar PTs, have high similarities in shape, color, size, etc. Under the condition of small samples, it is difficult for machine learning to distinguish them, and errors are easily generated during classification. The recognition method of humanoid concept learning has strong advantages in dealing with such problems because people can accurately determine the category of the target based on the relationship between the target and its surroundings rather than some other things similar to it.
[0110] Referring to Figure 14 , the correction method based on the strong correlation relationship library and the correction method based on Bayesian inference proposed by the present invention have improved the recall rate and precision rate of the recognition results of transmission line electrical equipment compared with only using Mask R-CNN, and have obvious advantages in the recognition of small-sample electrical equipment images. It can be clearly seen from the figure that the recognition effect of Mask R-CNN is greatly affected by the sample size. When the sample size is only 400, the F1 score, precision rate and recall rate of Mask R-CNN are much lower than those of HLCL. Although with the increase of the sample size, these three indicators of Mask R-CNN increase more rapidly, they are still lower than those of HLCL. In the inspection task of transmission line inspection robots, a large number of images are not taken for learning and training each time, and the sample size is only dozens or even less in most cases. At this time, it is undoubtedly very difficult to use the image recognition method of deep learning. The image recognition method based on human-like concept learning proposed by the present invention performs outstandingly in small samples and is more suitable for solving the problem of transmission line inspection image recognition.
[0111] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for identifying images of transmission lines, characterized in that: It includes, collecting images of transmission line equipment; coarsely identifying the collected transmission line images using the MaskR-CNN method; using human-like concept learning and combining the relevant relationships between transmission line equipment to construct a correction mechanism to correct the relationships in the recognition results; using a Bayesian semantic network for decomposition to obtain the final recognition result; The human-like concept learning includes: By learning the idea of human-like concept learning, using the relevant relationships existing between transmission line electrical equipment to correct the recognition results of the MaskR-CNN; for the modeling of relationships, corresponding correction mechanisms are constructed from two perspectives: strong relevant relationships and weak relevant relationships; The relevant relationships existing between the transmission line electrical equipment include: current transformers and support insulators, capacitors and support insulators, GIS and bushings, switches and support insulators, and capacitors and capacitors; The construction of the correction mechanism from strong relevant relationships includes: The process of constructing a correction mechanism from a strong correlation relationship includes the MaskR-CNN recognition result, neighborhood, position normalization, relative size, and position relationship; among them, the MaskR-CNN recognition result: it is assumed that there are N recognition results as d (1) , d (2) , … d (N) , each recognition result is called an instance, and each instance has three attributes, classification result bounding box and confidence Bounding box attribute b b = (x, y, w, h), where (x, y), w, and h represent the center coordinates, width, and height of the bounding box, respectively, and The neighborhood is defined as: where: α w , α h > 1 is a parameter for controlling the size of the instance neighborhood; The normalization method for the position includes setting the normalized bounding box of the target to be corrected as: For b (j) ∈ ε(d (i) ) satisfies the following formula: The calculation method of the relative size is the ratio of the areas of the bounding boxes of two instances. In strong relevant relationships, it is the area ratio of the determinant to the determined; in weak relevant relationships, it is the target instance compared with other instances in the neighborhood; the positional relationship includes: defining the positional relationships of "above", "below", "left and right" based on the normalized spatial relationships between instances in the image, which together with the relative size are used as criteria for determining whether there are strong relevant relationships between instances; The construction of the correction mechanism from weak relevant relationships includes: Based on the human-like concept learning, correcting the recognition results from the perspective of object relationships, and the correction of instances in the weak relevant relationship correction method is completed step by step; There is an instance set Its correct classification is The neighborhood set is From the perspective of object relationships, the instance and its neighborhood are most likely when the correct class is given, and can be written as:
2. The method for identifying images of transmission lines according to claim 1, characterized in that: The use of the Mask R-CNN method to identify the collected transmission line images includes: The MaskR-CNN method first uses the convolutional layer of the object detection algorithm to extract the overall features of the transmission line equipment image, generates regions of interest through the region proposal network, applies ROIAlign to ensure the alignment and consistency between the candidate boxes and the extracted features, and uses the classification and regression layers to obtain the object category and bounding box regression of the regions of interest. Finally, the Mask branch is used to classify and segment each pixel; the coarse recognition result of the MaskR-CNN includes the classification result, bounding box, and corresponding confidence of the target.
3. The method for identifying images of transmission lines according to claim 1 or 2, characterized in that: The construction of the correction mechanism from strong relevant relationships further includes: For two with strong relevant relationships, if one party depends on the correct recognition of the other party, it can be corrected relying on the strong relevant relationship.
4. The method for identifying images of transmission lines according to claim 3, characterized in that: The relationship correction of the recognition results includes: Regarding the confidence of the recognition results of the MaskR-CNN network, for the recognition results with a confidence lower than 0.90, strong correlation relationship correction is performed; for the instances that meet the position condition and the relative size condition, consider whether there is a strong correlation relationship with other instances in the neighborhood and whether the relationship is strong enough. Its formula is expressed as follows: where: s is the confidence level, which represents the unreliability of the rough recognition result, and are both one-dimensional vectors, is a sparse vector that stores in which strong correlation relationships the target to be corrected is included, is the weight occupied by each strong correlation relationship, and the weights are all greater than 1. y is the necessity of strong relationship correction. When y is less than 1, it means that there is no strong correlation relationship. When y is higher than a certain threshold, it means that the recognition result itself is not reliable enough and there is a certain strong correlation relationship, and it can be corrected according to the strong correlation relationship.
5. The transmission line image recognition method according to claim 4, characterized in that: The Bayesian semantic network includes: The Bayesian semantic network consists of nodes and edges that represent objects and the relationships between objects respectively. In a strongly correlated relationship, the neighborhood is as small as possible and not necessarily a standard square or circle. In a weakly correlated relationship, the neighborhood is taken as a circle or rectangle and has a relatively larger area than in the strongly correlated relationship; define the instance set within the neighborhood of d (i) When the confidence level is low, for instance d the correct classification corresponding to it is (i) denotes the correct classification of the instance, and g c denotes the classification set.
Citation Information
Patent Citations
Power equipment target recognition method
CN111209864A