Substation equipment point cloud segmentation method, system, equipment and medium
Through the point cloud segmentation method of substation equipment, the visual segmentation model and point cloud projection technology are used to solve the problem of inaccurate segmentation of similar geometric feature devices in the existing methods, and efficient and accurate point cloud segmentation effect is achieved.
Patent Information
- Application Number
- CN202510514951.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-07-18
AI Technical Summary
The existing substation point cloud segmentation method requires relying on a large amount of training data, making it difficult to accurately distinguish equipment with similar geometric features in complex environments, and is poorly robust and inefficient in work.
By obtaining the whole-domain three-dimensional point cloud model and the whole-domain video frame image of the substation, the positive and negative point sets are segmented using the visual segmentation model, a multi-value mask is generated, and the point cloud is projected to the image plane for object label assignment, and finally the outlier point label is eliminated to achieve accurate distinction of the equipment.
No large amount of training data is required to improve the robustness and work efficiency of point cloud segmentation, and accurately distinguish similar geometric feature devices in complex environments.
Smart Images

Figure CN120339398A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of substations, and in particular, to a method, system, device and medium for point cloud segmentation of substation equipment. Background Art
[0002] With the popularization of smart grids and the increasing complexity of substation equipment, the digital and intelligent management of substations has become increasingly important. As an important part of the power system, substations are responsible for the transformation, distribution and protection of electricity, and the operating status of their equipment directly affects the stability and security of the power grid. In the field of substation inspection, accurately obtaining parameters such as the geometric dimensions, relative positions, and postures of equipment is crucial for fault diagnosis and defect identification, and is related to equipment maintenance and asset management.
[0003] Currently, the monitoring and measurement of substation equipment mainly rely on manual inspections and regular checks. This method is not only time-consuming and laborious, but also has problems such as manual measurement errors, low efficiency, and safety hazards, and cannot meet the requirements of efficient and precise management. With the application of advanced sensor technologies such as lidar and RGB cameras, more and more substations have begun to attempt to digitally model and measure equipment through automated means. However, due to the variety and complexity of substation equipment, as well as possible occlusions between equipment and the uneven distribution of point cloud data, existing automated point cloud segmentation and measurement methods still face many challenges in practical applications.
[0004] Point cloud segmentation aims to divide the point cloud into multiple subsets with physical meanings according to the overall and local characteristics of the point cloud distribution, so as to provide basic support for 3D modeling, object classification and tracking.
[0005] Existing substation point cloud segmentation methods need to rely on a large amount of training data and are difficult to accurately distinguish equipment with similar geometric features in complex environments, with poor robustness and low working efficiency. Summary of the Invention
[0006] In view of this, the present invention provides a method, system, device and medium for point cloud segmentation of substation equipment, which solves the technical problems that existing substation point cloud segmentation methods need to rely on a large amount of training data and are difficult to accurately distinguish equipment with similar geometric features in complex environments, with poor robustness and low working efficiency.
[0007] The first aspect of the present invention provides a method for point cloud segmentation of substation equipment, including:
[0008] Obtaining a global 3D point cloud model and global video frame images of a substation;
[0009] Perform visual segmentation on the global video frame image according to the positive and negative point sets of multiple objects to be segmented, and obtain the masks of each object to be segmented in each frame of the global video frame image;
[0010] For each frame, assign values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtain the multi-value masks of each object to be segmented in the frame;
[0011] Convert the global three-dimensional point cloud model into a point cloud in the camera coordinate system through coordinate transformation, and project the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane;
[0012] Assign object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame, and obtain the object labels of each point on the point cloud plane;
[0013] Remove the outlier labels of the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after removing the outlier labels.
[0014] Preferably, the performing visual segmentation on the global video frame image according to the positive and negative point sets of multiple objects to be segmented, and obtaining the masks of each object to be segmented in each frame of the global video frame image includes:
[0015] Determine the positive point sets and negative point sets of each object to be segmented according to the positions of each object to be segmented in the global video frame image, and integrate the positive point sets and the negative point sets into positive and negative point sets;
[0016] Initialize the global video frame image through a visual segmentation model;
[0017] Convert the data format of the positive and negative point sets into the data format required for inputting into the visual segmentation model;
[0018] Input the positive and negative point sets with the converted data format into the visual segmentation model, and use the visual segmentation model to determine the masks of each object to be segmented in each frame of the global video frame image by using the input positive and negative point sets and the image features embedded generated from the initialized global video frame image.
[0019] Preferably, the assigning values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtaining the multi-value masks of each object to be segmented in the frame includes:
[0020] Assign masks for the objects to be segmented at the same pixel position in the same frame according to the overlapping situation of multiple objects to be segmented at the same pixel position in the same frame; wherein, when multiple objects to be segmented overlap at the same pixel position in the same frame, assign the mask for the objects to be segmented at the same pixel position as 0; when multiple objects to be segmented do not overlap at the same pixel position in the same frame, assign the mask for the objects to be segmented at the same pixel position as 1;
[0021] Determine the multi - value masks for each object to be segmented in the frame according to the masks of all pixel positions of each object to be segmented in the frame.
[0022] Preferably, the step of converting the global three - dimensional point cloud model into a point cloud in the camera coordinate system through coordinate transformation and projecting the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane includes:
[0023] Convert each point in the global three - dimensional point cloud model into a point cloud in the lidar coordinate system according to the pose relationship of the lidar in the scene point cloud;
[0024] Convert the point cloud in the lidar coordinate system into a point cloud in the camera coordinate system according to the extrinsic parameter relationship between the camera and the lidar;
[0025] Project the point cloud in the camera coordinate system onto the image plane using the camera intrinsic matrix to obtain a point cloud plane.
[0026] Preferably, the step of assigning object labels to each point on the point cloud plane according to the multi - value masks of each object to be segmented in each frame to obtain the object labels of each point on the point cloud plane includes:
[0027] Determine the multi - value masks of each frame that match the pixel coordinates of each pixel point on the point cloud plane according to the pixel coordinates of each pixel point on the point cloud plane;
[0028] Determine the object labels of each pixel point on the point cloud plane in each frame according to the pixel coordinates of each pixel point on the point cloud plane and the multi - value masks of each frame;
[0029] Determine the object labels of the last frame as the object labels of each pixel point on the point cloud plane according to the object labels of each pixel point on the point cloud plane in each frame.
[0030] Preferably, the step of removing outlier labels from the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after removing outlier labels includes:
[0031] Cluster the object labels of each point on the point cloud plane through an outlier detection algorithm to determine the neighborhood of each point and the density corresponding to the point;
[0032] Determine multiple outliers on the point cloud plane according to the neighborhood of each point and the density corresponding to the point;
[0033] Remove multiple outliers and the outlier labels to obtain the point cloud of each object to be segmented after removing the outlier labels.
[0034] Preferably, the method further includes:
[0035] Determine the shortest point cloud coordinate distance between the first object and the second object as the measurement distance between the first object and the second object according to the point cloud coordinate distances of the first object and the second object in each frame.
[0036] In a second aspect, the present invention also provides a substation equipment point cloud segmentation system, including:
[0037] An information acquisition module for acquiring the global three-dimensional point cloud model and global video frame images of the substation;
[0038] A visual segmentation module for visually segmenting the global video frame images according to the positive and negative point sets of multiple objects to be segmented to obtain the masks of each object to be segmented in each frame of the global video frame images;
[0039] A multi-value mask determination module for assigning values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in each frame to obtain the multi-value masks of each object to be segmented in each frame;
[0040] A point cloud projection module for converting the global three-dimensional point cloud model into a point cloud in the camera coordinate system and projecting the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane;
[0041] A label assignment module for assigning object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame to obtain the object labels of each point on the point cloud plane;
[0042] An outlier removal module for removing outlier labels from the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after removing the outlier labels.
[0043] In a third aspect, the present invention further provides an electronic device, which includes a memory and a processor. A computer program is stored in the memory. When the computer program is executed by the processor, the processor is caused to execute the steps of the substation equipment point cloud segmentation method as described in the first aspect.
[0044] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the substation equipment point cloud segmentation method as described in the first aspect are implemented.
[0045] As can be seen from the above technical solutions, the present invention performs visual segmentation on the global video frame image through the positive and negative point sets of multiple objects to be segmented, and obtains the mask of each object to be segmented in each frame of the global video frame image. For each frame, according to the overlapping situation of each object to be segmented within the frame, the mask of each object to be segmented is assigned a value to form a multi-value mask of each object to be segmented in this frame. The global three-dimensional point cloud model is projected onto the image plane, and according to the multi-value mask of each object to be segmented in each frame, object labels are assigned to the points on the point cloud plane. Subsequently, outlier labels are removed from the object labels of the points on the point cloud plane, so as to accurately distinguish devices with similar geometric features in a complex environment. This method does not require relying on a large amount of training data, and effectively improves the robustness and working efficiency of point cloud segmentation. Description of the Drawings
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is an application environment of a substation equipment point cloud segmentation method provided by an embodiment of the present invention;
[0048] Figure 2 It is a flowchart of a substation equipment point cloud segmentation method provided by an embodiment of the present invention;
[0049] Figure 3 It is a schematic structural diagram of a substation equipment point cloud segmentation system provided by an embodiment of the present invention;
[0050] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0051] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0052] The substation equipment point cloud segmentation method provided by the embodiments of the present application can be applied to an application environment as Figure 1 shown. Among them, the terminal 101 communicates with the server 102 through a network. The data storage system can store the data that the server 102 needs to process. The data storage system can be integrated on the server 102, or placed in the cloud or other network servers. The terminal 101 or the server 102 acquires the global three-dimensional point cloud model and the global video frame images of the substation; performs visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, and obtains the masks of each object to be segmented in each frame of the global video frame images; for each frame, assigns values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtains the multi-value masks of each object to be segmented in the frame; converts the global three-dimensional point cloud model into the point cloud in the camera coordinate system, and projects the point cloud in the camera coordinate system onto the image plane to obtain the point cloud plane; assigns object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame, and obtains the object labels of each point on the point cloud plane; removes the outlier labels of the object labels of each point on the point cloud plane, and obtains the point cloud of each object to be segmented after removing the outlier labels.
[0053] The terminal 101 can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, etc.
[0054] The server 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0055] As Figure 2 shown, the embodiments of the present application provide a substation equipment point cloud segmentation method. Taking the method applied to the Figure 1 terminal 101 or the server 102 in it as an example, it includes the following steps S1 to S6. Among them:
[0056] Step S1, acquire the global three-dimensional point cloud model and the global video frame images of the substation.
[0057] Among them, a global three-dimensional point cloud model of the substation can be obtained by the inspection robot using lidar three-dimensional reconstruction, and RGB images of consecutive video frames are obtained through the camera as global video frame images.
[0058] Step S2: Perform visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, and obtain the masks of each object to be segmented in each frame of the global video frame images.
[0059] Among them, the positive and negative point sets of the object to be segmented include a positive point set and a negative point set. The positive point set is
[0060] the set of actual position points of the object to be segmented in the global video frame images, and the negative point set is the set of position points where the object to be segmented does not exist in the global video frame images. The determination of the positive and negative point sets can be obtained through manual annotation or learning using existing data.
[0061] By performing visual segmentation on the global video frame images, the mask of each object to be segmented in each frame can be obtained. The mask represents the position and shape of the object to be segmented in the image.
[0062] Step S3: For each frame, assign values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtain the multi-value masks of each object to be segmented in the frame.
[0063] Among them, according to the mask overlapping situation of multiple objects to be segmented in the same frame, special processing is performed on the overlapping mask regions. When the masks of multiple objects to be segmented overlap at the same pixel position in the same frame, the mask values at this pixel position are used to distinguish different objects to be segmented.
[0064] By performing such multi-value assignment processing on the masks of each frame, the multi-value masks of each object to be segmented in each frame can be obtained. The multi-value masks not only contain the position and shape information of the object to be segmented in the image, but also contain the overlapping relationship information between different objects to be segmented.
[0065] Step S4: Convert the global three-dimensional point cloud model into a point cloud in the camera coordinate system, and project the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane.
[0066] Among them, according to the relative position relationship between the lidar and the camera, the coordinate system conversion is first performed. Each point in the global three-dimensional point cloud model obtained by the lidar has three-dimensional coordinates in the original coordinate system. In order to convert these points to the camera coordinate system, the external parameter relationship between the lidar and the camera, including the rotation matrix and the translation vector, needs to be utilized. By applying these conversion parameters, each point in the point cloud model can be converted from the lidar coordinate system to the camera coordinate system. Then, using the internal parameter matrix of the camera, the point cloud in the camera coordinate system is projected onto the image plane.
[0067] Step S5: Assign object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame, and obtain the object labels of each point on the point cloud plane.
[0068] Among them, according to the information in the multi-value mask, the object to be segmented corresponding to each pixel position is determined. Since the multi-value mask contains the overlapping relationship information between different objects to be segmented, the points in the overlapping area can be accurately distinguished as belonging to which object to be segmented. Then, the object label of each pixel position is assigned to the corresponding point on the point cloud plane, and the object labels of each point on the point cloud plane are obtained. In this way, each point on the point cloud plane is assigned a corresponding object label, indicating the object to be segmented to which it belongs.
[0069] Step S6: Remove the outlier labels of the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after removing the outlier labels.
[0070] Among them, by statistically analyzing the object labels of each point on the point cloud plane, the main label and the corresponding number of points of each object to be segmented are determined. Points with object labels inconsistent with the main label and fewer surrounding points are regarded as outliers. The labels of these outliers are removed, and only the main label is retained, so as to obtain the point cloud of each object to be segmented after removing the outlier labels. This step effectively reduces the influence of noise and misclassified points, and improves the accuracy and robustness of point cloud segmentation.
[0071] It should be noted that in the embodiment of the present application, the global video frame image is visually segmented by the positive and negative point sets of multiple objects to be segmented, and the mask of each object to be segmented in each frame of the global video frame image is obtained. For each frame, according to the overlapping situation of each object to be segmented in the frame, the mask of each object to be segmented is assigned to form the multi-value mask of each object to be segmented in this frame. The global three-dimensional point cloud model is projected onto the image plane, and according to the multi-value masks of each object to be segmented in each frame, object labels are assigned to each point on the point cloud plane. Subsequently, the outlier labels of the object labels of each point on the point cloud plane are removed, so as to accurately distinguish devices with similar geometric features in a complex environment. This method does not need to rely on a large amount of training data, and effectively improves the robustness and working efficiency of point cloud segmentation.
[0072] In some embodiments, visual segmentation of the global video frame image is performed according to the positive and negative point sets of multiple objects to be segmented, and masks of each object to be segmented in each frame of the global video frame image are obtained, including:
[0073] Step S201: Determine the positive point set and negative point set of each object to be segmented according to the position of each object to be segmented in the global video frame image, and integrate the positive point set and negative point set into a positive and negative point set.
[0074] Among them, for each object to be segmented in the video, according to the position and shape of the object in the video frame, the positive point set inside the object and the negative point set outside the object are determined manually or through a specific algorithm. The positive point set represents the pixel positions inside the object, and the negative point set represents the pixel positions in the background area outside the object. The coordinates of the point set are determined based on the image pixel coordinate system.
[0075] Exemplarily, by interactively annotating the positive and negative point sets of each object to be segmented, there are X RGB video frames, and each frame image is represented by an index There are Y objects to be measured, and each object is represented by an index For each object i, in a certain frame ti containing the object, a plurality of positive and negative points are selected through interactive annotation (k and l represent the numbers of positive and negative points respectively).
[0076] Using And These two sets are used to represent these points. Therefore, the annotation of object i in frame ti can be expressed as:
[0077] (positive points)
[0078] (negative points)
[0079] For each object i, the user only needs to perform annotation (select positive and negative points) in one frame ti containing the object. Therefore, the annotation of each object can be expressed as a tuple , and the annotation information of all objects is organized into a set S, which contains the annotation tuples of each object. That is , where each represents the annotation information of object i in frame ti.
[0080] The process of interactive annotation is as follows: (1) The user selects a frame. (2) The user selects an object name. (3) The user selects a positive or negative point for annotation. (4) Annotate one or more positive or negative points of the selected object. (5) Return the annotation tuple of the object. (6) And so on, annotate all objects and return a set S containing the annotation tuples of each object.
[0081] Step S202: Initialize the global video frame image through a visual segmentation model.
[0082] Among them, the visual segmentation model can adopt the SAM2 model. The Segment Anything Model 2 (SAM2), as a basic model in the field of image and video visual segmentation, has powerful segmentation capabilities and can generate high-quality object masks based on given prompts. In the substation equipment point cloud segmentation method, the SAM2 model is used to perform visual segmentation on the global video frame image, and by combining the positive and negative point set information of the object, generate the masks of each object to be segmented in the image.
[0083] Among them, convert the global video frame image into the PyTorch tensor format and add a dimension to match the batch dimension requirements of the model input. Input this tensor data into the SAM2 model. The model extracts features from the current frame image through its internal image encoder and generates a feature embedding representation of the frame image. This feature embedding will be used as the basic data for subsequent processing and combined with the positive and negative point set information of the object to generate the object mask.
[0084] Step S203: Convert the data format of the positive and negative point sets into the data format required for input into the visual segmentation model.
[0085] Among them, convert the determined positive and negative point sets into a data format suitable for input into the SAM2 model.
[0086] Specifically, convert the point set coordinates into the PyTorch tensor format and assign corresponding labels to each point according to the nature of the positive and negative points. The positive point label is set to 1, and the negative point label is set to 0.
[0087] Step S204: Input the positive and negative point sets with the converted data format into the visual segmentation model, and through the visual segmentation model, use the input positive and negative point sets and the image feature embedding generated by the initialized global video frame image to determine the masks of each frame of each object to be segmented in the global video frame image.
[0088] Among them, the formatted positive and negative point set tensors and their corresponding label tensors are sequentially input into the SAM2 model that has been initialized for the current video frame. The model processes these point set information and the previously generated image feature embeddings through its internal mask decoder module. During the processing, the mask decoder combines the point set prompts and the image features to generate mask prediction results corresponding to the objects.
[0089] After receiving the positive and negative point set information and the image feature embeddings, the mask decoder of the SAM2 model generates mask predictions for each object through a series of calculation processes (including but not limited to convolution operations, attention mechanism operations, etc.). The mask prediction results are represented in the form of tensors, where each element corresponds to a pixel position in the image, and the element value represents the probability or confidence that the pixel belongs to the object.
[0090] To obtain a clear object segmentation mask, the generated mask prediction tensor is binarized. Usually, the threshold method is adopted, and a suitable threshold (such as 0.5) is set. The elements in the mask prediction tensor that are greater than the threshold are set to 1, indicating that the pixel belongs to the object; the elements less than the threshold are set to 0, indicating that the pixel belongs to the background. After binarization, the obtained tensor is the clear object mask, which can accurately identify the contour and area of the object in the video frame.
[0091] Exemplarily, first input all continuously captured RGB frames of the substation into the inference state of SAM2, and then sequentially take out the annotation tuples of each object from the set S obtained in S2 above. For an object i, input the frame number ti, the positive and negative point sets Pi and Ni in Si, and the inference state into the predictor.add_new_points_or_box() function of SAM2 to obtain the mask of object i in frame ti, and then obtain the masks of object i in all tracked frames through the predictor.propagate_in_video() function. And so on, by sequentially inputting the above information of different objects, the masks of all objects in all frames can be obtained.
[0092] For convenience of representation, the letter M represents the mask, and the set of masks of object i in all frames is where, is the set of frames in which object i is annotated, representing the sequence numbers of all frames in which object i appears; is the mask of object i in the -th frame; the total masks of all objects can be expressed as
[0093]
[0094] In the formula, is the total mask of all objects, and Y is the number of objects.
[0095] In some embodiments, the masks of the objects to be segmented are assigned according to the overlapping situation of the objects to be segmented in the frame, and multi-valued masks of the objects to be segmented in the frame are obtained, including:
[0096] Step S301: Assign a value to the mask of the object to be segmented at the same pixel position in the same frame according to the overlapping situation of multiple objects to be segmented at the same pixel position in the same frame; wherein, when multiple objects to be segmented overlap at the same pixel position in the same frame, the mask of the object to be segmented at the same pixel position is assigned a value of 0; when multiple objects to be segmented do not overlap at the same pixel position in the same frame, the mask of the object to be segmented at the same pixel position is assigned a value of 1;
[0097] Step S302: Determine the multi-valued masks of the objects to be segmented in the frame according to the masks of all pixel positions of the objects to be segmented in the frame.
[0098] Exemplarily, let the mask of object i in the t-th frame be denoted as , which is a binary matrix. A value of 1 represents the pixel area of object i, and a value of 0 represents the background area.
[0099] To represent multiple objects using only one frame of mask for each frame of the image, when the masks of multiple objects overlap at the same pixel position in the same frame, in order to avoid conflicts and ensure that each pixel belongs to only one object, we set the label value of this pixel to 0, indicating that this pixel is an invalid area or background.
[0100] For non-conflicting areas, the label value of object i is assigned , according to the previous assumption, there are Y objects, and in the t-th frame, the final multi-valued mask is:
[0101]
[0102] For X RGB frames, there are X multi-valued masks, which can be expressed as:
[0103]
[0104] In the formula, is the multi-valued mask of X RGB frames.
[0105] In some embodiments, the global three-dimensional point cloud model is transformed into a point cloud in the camera coordinate system through coordinate transformation, and the point cloud in the camera coordinate system is projected onto the image plane to obtain a point cloud plane, including:
[0106] Step S401: According to the pose relationship of the lidar in the scene point cloud, convert each point in the global three-dimensional point cloud model into a point cloud in the lidar coordinate system;
[0107] Step S402: According to the extrinsic parameter relationship between the camera and the lidar, convert the point cloud in the lidar coordinate system into a point cloud in the camera coordinate system;
[0108] Step S403: Project the point cloud in the camera coordinate system onto the image plane using the camera intrinsic matrix to obtain the point cloud plane.
[0109] Exemplarily, after the lidar performs three-dimensional reconstruction on the entire substation, the pose correspondence relationship between the point cloud obtained by each frame (corresponding to the camera image) of the lidar and the entire scene point cloud can be obtained. Among them, the three-dimensionally reconstructed scene point cloud is defined as , the point cloud collected by each frame of the lidar is defined as , the pose transformation of each frame of the lidar in the scene point cloud is , the extrinsic parameter relationship between the camera and the lidar is , the camera intrinsic matrix is K, and the multi-value mask of the t-th frame remains .
[0110] For each point in the scene point cloud, first convert it into the coordinate system of the lidar in the t-th frame, that is:
[0111] ,
[0112] Then further convert the above point cloud into the camera coordinate system, that is:
[0113] ,
[0114] Obtain the point cloud of the scene point cloud in the camera coordinate system , and then use the camera intrinsic matrix K to project the point cloud onto the image plane:
[0115] .
[0116] In some embodiments, according to the multi-value mask of each object to be segmented in each frame, assign object labels to each point on the point cloud plane to obtain the object labels of each point on the point cloud plane, including:
[0117] Step S501: According to the pixel coordinates of each pixel point on the point cloud plane, determine the multi-value masks of each frame that match the pixel coordinates of each pixel point;
[0118] Step S502: According to the pixel coordinates of each pixel point on the point cloud plane and the multi-value masks of each frame, determine the object labels of each pixel point on the point cloud plane in each frame;
[0119] Step S503: Determine the object label of the last frame as the object label of each pixel point on the point cloud plane according to the object labels of each pixel point on the point cloud plane in each frame.
[0120] Among them, if a certain point is assigned labels in multiple frames and these labels conflict (that is, the labels in different frames are different), then the final label is selected according to a preset rule. A common rule is to adopt the principle of "the later one takes precedence", that is, select the label of this point in the last frame that is labeled as the final label. This is because in practical applications, as the video plays, the position and state of the object in the scene may change, so the later-labeled frame often better reflects the current state and position of the object.
[0121] After determining the object labels of each point on the point cloud plane, the point cloud can be segmented according to these labels to obtain point cloud sets of different objects. These segmented point cloud sets can be used for subsequent applications such as substation equipment recognition and status monitoring, providing strong support for the intelligent operation and maintenance of substations.
[0122] Exemplarily, according to the pixel coordinates obtained by projection , query the multi-value mask of the t-th frame , and through obtain the points of the scene point cloud the object label at the t-th frame. If a certain point has labels in multiple frames, the label is updated according to the time sequence, and the label of the later frame is preferentially used, that is:
[0123]
[0124] Finally, each point of the scene point cloud is assigned an object label :
[0125] .
[0126] In some embodiments, outlier label elimination is performed on the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after eliminating outlier labels, including:
[0127] Step S601: Cluster the object labels of each point on the point cloud plane through an outlier detection algorithm to determine the neighborhood of each point and the density corresponding to the point;
[0128] Step S602: Determine multiple outliers on the point cloud plane according to the neighborhood of each point and the density corresponding to the point;
[0129] Step S603: Eliminate multiple outliers and outlier labels to obtain the point cloud of each object to be segmented after eliminating outlier labels.
[0130] Exemplarily, here \(i\) represents the serial number of the points in the point cloud, and \(k\) is used to represent the object label value. For outlier removal, it will be carried out in the order of different objects. First, initialize \(k = 1\), that is, the first object, and filter the points in the scene point cloud with the label value of \(k\), that is , Denote the outlier detection algorithm as OLRM, then the outliers of object \(k\) are . For these points, set their label values to zero, that is, set the outliers as the background. And so on, loop until , and obtain the scene point cloud after removing the outlier labels .
[0131] Among them, for the specific outlier detection algorithm, DBSCAN (Density-Based Spatial Clustering of Applications with Noise) is used here. It is a commonly used density clustering algorithm that can effectively detect outliers. It is based on the density distribution of points, regards the density-dense regions as clusters, and regards the points in the sparse regions as outliers. There are two key parameter values. One is the neighborhood radius, which defines the density range around a certain point, and the other is the minimum number of neighborhood points, which represents the minimum number of neighbors required for a point to become a core point. The specific selection of parameter values needs to be based on the actual situation.
[0132] In some embodiments, it is necessary to measure the distance to be measured (the distance between the closest points of the point clouds of two objects) and parameters such as the length of the object to be measured (the length, width, and height of the bounding box of the point cloud) of different objects to be measured. Specifically, this method further includes:
[0133] According to the point cloud coordinate distances of the first object and the second object in each frame, determine the shortest point cloud coordinate distance between the first object and the second object as the measurement distance between the first object and the second object.
[0134] Exemplarily, when it is necessary to measure the distance between object \(a\) and object \(b\), it is necessary to obtain the points representing object \(a\) and the points representing object \(b\) respectively, and traverse and all point pairs in , calculate the distances between all point pairs and take the minimum value as the distance between the two objects:
[0135]
[0136] In the formula, represents the distance between object \(a\) and object \(b\).
[0137] In other embodiments, if it is necessary to calculate the length of some special objects in a substation, such as rod-shaped objects, the point cloud of the object can be obtained through tags first, and then the main axis direction of the point cloud of the object can be obtained through principal component analysis. Next, the point cloud is projected onto the main axis, and the maximum and minimum values after projection are calculated to obtain the length.
[0138] Based on the same inventive concept, an embodiment of the present application further provides a substation equipment point cloud segmentation system for implementing the substation equipment point cloud segmentation method involved above.
[0139] The implementation solution provided by this system to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the substation equipment point cloud segmentation system provided below can refer to the limitations on the substation equipment point cloud segmentation method in the above text, and will not be repeated here.
[0140] As Figure 3 shown, an embodiment of the present application provides a substation equipment point cloud segmentation system, including:
[0141] An information acquisition module 100, configured to acquire a global three-dimensional point cloud model and global video frame images of a substation;
[0142] A visual segmentation module 200, configured to perform visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, and obtain masks of each object to be segmented in each frame of the global video frame images;
[0143] A multi-value mask determination module 300, configured to assign values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in each frame, and obtain multi-value masks of each object to be segmented in the frame;
[0144] A point cloud projection module 400, configured to convert the global three-dimensional point cloud model into a point cloud in the camera coordinate system and project the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane;
[0145] A label assignment module 500, configured to assign object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame, and obtain object labels of each point on the point cloud plane;
[0146] An outlier rejection module 600, configured to reject outlier labels of object labels of each point on the point cloud plane, and obtain point clouds of each object to be segmented after rejecting outlier labels.
[0147] In some embodiments, the visual segmentation module 200 is configured to:
[0148] Determine the positive point set and negative point set of each object to be segmented according to the positions of the objects to be segmented in the global video frame image, and integrate the positive point set and negative point set into the positive and negative point set;
[0149] Initialize the global video frame image through a visual segmentation model;
[0150] Convert the data format of the positive and negative point set into the data format required for input into the visual segmentation model;
[0151] Input the positive and negative point set with the converted data format into the visual segmentation model, and use the image feature embedding generated by the visual segmentation model from the input positive and negative point set and the initialized global video frame image to determine the mask of each object to be segmented in each frame of the global video frame image.
[0152] In some embodiments, the multi-value mask determination module 300 is used for:
[0153] Assign a value to the mask of the object to be segmented at the same pixel position in the same frame according to the overlapping situation of multiple objects to be segmented at the same pixel position in the same frame; wherein, when multiple objects to be segmented overlap at the same pixel position in the same frame, assign a value of 0 to the mask of the object to be segmented at the same pixel position; when multiple objects to be segmented do not overlap at the same pixel position in the same frame, assign a value of 1 to the mask of the object to be segmented at the same pixel position;
[0154] Determine the multi-value mask of each object to be segmented in the frame according to the masks of all pixel positions of each object to be segmented in the frame.
[0155] In some embodiments, the point cloud projection module 400 is used for:
[0156] According to the pose relationship of the lidar in the scene point cloud, convert each point in the global three-dimensional point cloud model into a point cloud in the lidar coordinate system;
[0157] According to the extrinsic parameter relationship between the camera and the lidar, convert the point cloud in the lidar coordinate system into a point cloud in the camera coordinate system;
[0158] Project the point cloud in the camera coordinate system onto the image plane using the camera intrinsic matrix to obtain the point cloud plane.
[0159] In some embodiments, the label assignment module 500 is used for:
[0160] Determine the multi-value mask of each frame that matches the pixel coordinates of each pixel point of the point cloud plane according to the pixel coordinates of each pixel point of the point cloud plane;
[0161] Determine the object label of each pixel point of the point cloud plane in each frame according to the pixel coordinates of each pixel point of the point cloud plane and the multi-value mask of each frame;
[0162] Determine the object label of the last frame as the object label of each pixel point on the point cloud plane according to the object labels of each pixel point on the point cloud plane in each frame.
[0163] In some embodiments, the outlier rejection module 600 is configured to:
[0164] Cluster the object labels of each point on the point cloud plane through an outlier detection algorithm to determine the neighborhood of each point and the density corresponding to the point;
[0165] Determine multiple outliers on the point cloud plane according to the neighborhood of each point and the density corresponding to the point;
[0166] Reject the multiple outliers and the outlier labels to obtain the point cloud of each object to be segmented after removing the outlier labels.
[0167] In some embodiments, the system further includes: a distance measurement module, configured to determine the shortest point cloud coordinate distance between the first object and the second object as the measurement distance between the first object and the second object according to the point cloud coordinate distances of the first object and the second object in each frame.
[0168] As Figure 4 shown, an embodiment of the present application provides an electronic device. The electronic device 10 includes a memory 20 and a processor 30. When the computer program stored in the memory 20 is executed by the processor 30, the processor 30 is caused to execute the steps of the substation equipment point cloud segmentation method in the above embodiments.
[0169] An embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed, the steps of the substation equipment point cloud segmentation method in the above embodiments are implemented.
[0170] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described system, electronic device, and computer storage medium can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0171] It should be noted that the terms "first", "second", etc. in the description, claims and the above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0172] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps do not necessarily have to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, at least some of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily have to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily have to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps or stages in other steps.
[0173] In several embodiments provided by the present invention, it should be understood that the disclosed systems, electronic devices, computer storage media and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0174] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0175] In addition, each functional unit in various embodiments of the present invention may be integrated into a processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0176] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the method described in various embodiments of the present invention by a computer device (which may be a personal computer, a server, or a network device, etc.). The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks, or optical discs.
[0177] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A method for point cloud segmentation of substation equipment, characterized in that, Including: Obtain the global three-dimensional point cloud model and global video frame images of the substation; Perform visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, and obtain the masks of each object to be segmented in each frame of the global video frame images; For each frame, assign values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtain the multi-value masks of each object to be segmented in the frame; Convert the global three-dimensional point cloud model into a point cloud in the camera coordinate system through coordinate transformation, and project the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane; Assign object labels to each point on the point cloud plane according to the multi-value masks of each object to be segmented in each frame, and obtain the object labels of each point on the point cloud plane; Remove the outlier labels of the object labels of each point on the point cloud plane to obtain the point cloud of each object to be segmented after removing the outlier labels.
2. The substation equipment point cloud segmentation method according to claim 1, wherein The step of performing visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, and obtaining the masks of each object to be segmented in each frame of the global video frame images includes: Determine the positive point set and negative point set of each object to be segmented according to the position of each object to be segmented in the global video frame images, and integrate the positive point set and the negative point set into a positive and negative point set; Initialize the global video frame images through a visual segmentation model; Convert the data format of the positive and negative point sets into the data format required for input into the visual segmentation model; Input the positive and negative point sets with the converted data format into the visual segmentation model, and use the visual segmentation model to determine the masks of each object to be segmented in each frame of the global video frame images by using the input positive and negative point sets and the image features embedded generated from the initialized global video frame images.
3. The substation equipment point cloud segmentation method according to claim 1, characterized in that The step of assigning values to the masks of each object to be segmented according to the overlapping situation of each object to be segmented in the frame, and obtaining the multi-value masks of each object to be segmented in the frame includes: Assign values to the masks of the objects to be segmented at the same pixel position in the same frame according to the overlapping situation of multiple objects to be segmented at the same pixel position in the same frame; wherein, when multiple objects to be segmented overlap at the same pixel position in the same frame, assign the value of 0 to the masks of the objects to be segmented at the same pixel position; when multiple objects to be segmented do not overlap at the same pixel position in the same frame, assign the value of 1 to the masks of the objects to be segmented at the same pixel position; Determine the multi-value masks of each object to be segmented in the frame according to the masks of all pixel positions of each object to be segmented within the frame.
4. The substation equipment point cloud segmentation method according to claim 1, wherein The step of converting the global three-dimensional point cloud model into a point cloud in the camera coordinate system through coordinate transformation, and projecting the point cloud in the camera coordinate system onto the image plane to obtain a point cloud plane includes: According to the pose relationship of the lidar in the scene point cloud, convert each point in the global three-dimensional point cloud model into a point cloud in the lidar coordinate system; According to the extrinsic parameter relationship between the camera and the lidar, the point cloud in the lidar coordinate system is converted into the point cloud in the camera coordinate system; The point cloud in the camera coordinate system is projected onto the image plane by using the camera intrinsic matrix to obtain the point cloud plane.
5. The substation equipment point cloud segmentation method according to claim 1, characterized in that The assigning object labels to each point on the point cloud plane according to the multi-value masks of each of the objects to be segmented in each of the frames to obtain the object labels of each point on the point cloud plane includes: According to the pixel coordinates of each pixel point on the point cloud plane, determining the multi-value masks of each of the frames that match the pixel coordinates of each pixel point; According to the pixel coordinates of each pixel point on the point cloud plane and the multi-value masks of each of the frames, determining the object labels of each pixel point on the point cloud plane in each of the frames; According to the object labels of each pixel point on the point cloud plane in each of the frames, determining the object label of the last frame as the object label of each pixel point on the point cloud plane.
6. The method for substation equipment point cloud segmentation according to claim 1, wherein The removing the outlier labels from the object labels of each point on the point cloud plane to obtain the point cloud of each of the objects to be segmented after removing the outlier labels includes: Clustering the object labels of each point on the point cloud plane through an outlier detection algorithm to determine the neighborhood of each point and the density corresponding to the point; According to the neighborhood of each point and the density corresponding to the point, determining multiple outlier points on the point cloud plane; Removing the multiple outlier points and the outlier labels to obtain the point cloud of each of the objects to be segmented after removing the outlier labels.
7. The method for segmenting point clouds of substation equipment according to claim 1, wherein Further includes: According to the point cloud coordinate distances between the first object and the second object in each of the frames, determining the shortest point cloud coordinate distance between the first object and the second object as the measurement distance between the first object and the second object.
8. A substation equipment point cloud segmentation system, characterized in that, Includes: An information acquisition module, configured to acquire the global three-dimensional point cloud model and the global video frame images of the substation; A visual segmentation module, configured to perform visual segmentation on the global video frame images according to the positive and negative point sets of multiple objects to be segmented, to obtain the masks of each of the objects to be segmented in each frame in the global video frame images; A multi-value mask determination module, configured to, for each frame, assign values to the masks of each of the objects to be segmented according to the overlapping situations of each of the objects to be segmented in the frame, to obtain the multi-value masks of each of the objects to be segmented in the frame; A point cloud projection module, configured to convert the global three-dimensional point cloud model into the point cloud in the camera coordinate system and project the point cloud in the camera coordinate system onto the image plane to obtain the point cloud plane; A label assignment module, configured to assign object labels to each point on the point cloud plane according to the multi-value masks of each of the objects to be segmented in each of the frames, to obtain the object labels of each point on the point cloud plane; An outlier removal module, configured to remove the outlier labels from the object labels of each point on the point cloud plane to obtain the point cloud of each of the objects to be segmented after removing the outlier labels.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, and a computer program is stored in the memory. When the computer program is executed by the processor, the processor executes the steps of the substation equipment point cloud segmentation method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the substation equipment point cloud segmentation method according to any one of claims 1-7.