Power grid target body identification method and device
By using an aircraft to carry multiple image acquisition units in the power grid monitoring system to acquire images from different angles, generate body depth maps and adjust positions, the problem of difficulty in identifying the target body with a single angle acquisition image is solved, and high accuracy recognition of the target body is achieved.
Patent Information
- Application Number
- CN202510391603.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-08-12
AI Technical Summary
In the power grid monitoring system, images collected at a single angle are affected by factors such as object occlusion and personnel activities, making it difficult to accurately identify the target object.
The aircraft is equipped with multiple image acquisition units to acquire images from different angles, perform stereoscopic recovery processing to generate a body depth map, use the target recognition model to identify it, and adjust the aircraft position to re-acquire the image when it is not recognized until the target body is recognized.
It improves the recognition accuracy and reliability of the target body, and can still obtain high-quality images for recognition in the case of environmental changes or occlusion.
Smart Images

Figure CN120472336A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of target detection technology, and in particular to a method and device for identifying power grid targets. Background Art
[0002] In power grid monitoring systems, image acquisition units are typically installed at specific locations within the work area. These units capture images from a single angle within the monitoring area. The system then uses these images to identify objects within the monitoring area, such as equipment, personnel, and tools. However, due to factors such as obstructions and human activity, images captured from a single angle are ineffective, making it difficult to accurately identify objects. Summary of the Invention
[0003] In view of this, an object of the embodiments of the present application is to provide a method and apparatus for identifying a target object in a power grid, so as to solve the problem of low target object identification accuracy.
[0004] Based on the above objectives, an embodiment of the present application provides a method for identifying a target object in a power grid, including:
[0005] Acquire multiple images captured by multiple image acquisition units at different shooting angles; wherein each image acquisition unit is mounted at a different position of the aircraft;
[0006] Perform stereo restoration processing on multiple images to generate a stereo depth map;
[0007] Using a preset target recognition model to identify the stereo depth map, to obtain a target recognition result;
[0008] When the target recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body.
[0009] Optionally, performing stereo restoration processing on the multiple images to generate a stereo depth map includes:
[0010] Perform imaging restoration processing on multiple images to obtain a restored stereo depth map;
[0011] The restored stereo depth map is denoised to obtain a denoised stereo depth map.
[0012] Optionally, the imaging restoration process is performed on the multiple images to obtain a restored stereo depth map, and the method is:
[0013]
[0014] Among them, (X ′ ,Y ′) is the index of the pixel affected by noise and including the offset; G(X,Y) is the cumulative overlap number at the (X,Y) pixel point; M and N are the total number of images collected by the image acquisition unit in the preset row and column directions respectively: H (m,n) is the image of the mth row and nth column; Z is the restored depth; ∈ (m,n) (X′, Y′) is the zero-mean additive Gaussian noise at the pixel point (X′, Y′) in the mth row and nth column of the image.
[0015] Optionally, after generating the stereo depth map, the method further includes:
[0016] The palmprint recognition model is used to identify the three-dimensional depth map to obtain a palmprint recognition result.
[0017] Optionally, using a preset palmprint recognition model to recognize the three-dimensional depth map to obtain a palmprint recognition result includes:
[0018] Extracting palmprint encoding features from the stereo depth map;
[0019] Calculating a differentiation index for each feature bit in the palmprint encoding feature;
[0020] Construct identity information based on the differentiation indicators of each feature bit;
[0021] According to the identity identification information, the target object corresponding to the identity identification information and the identity information of the target object are identified.
[0022] Optionally, the method for calculating the differentiation index is:
[0023]
[0024] Among them, R i is the differentiation index of the i-th feature bit, D i is the difference between categories of the i-th feature bit, W i is the intra-category difference of the i-th feature bit.
[0025] Optionally, after obtaining the target recognition result, the following steps are included:
[0026] Determining a final recognition result based on the palmprint recognition result and the target recognition result;
[0027] When the target recognition result is unrecognized, adjusting the shooting angle of each image acquisition unit by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body, including:
[0028] When the final recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the final recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angles includes the target object.
[0029] Optionally, the method for determining a final recognition result based on the palmprint recognition result and the target recognition result is:
[0030] F fusion =α·F image +(1-α)·F palm (13)
[0031] Among them, F image is the feature vector generated by the target recognition model, F palm is the feature vector generated by the palmprint recognition model, α is the weight value corresponding to the target recognition result; F fusion The fused feature vector is classified by a classifier to determine the final recognition result.
[0032] Optionally, a final recognition result is determined based on the palmprint recognition result and the target recognition result, in the following method:
[0033] P final =α·P image +(1-α)·P palm (14)
[0034] Among them, P omage is the predicted probability value corresponding to the target recognition result, P palm is the predicted probability value corresponding to the palmprint recognition result.
[0035] The present application also provides a power grid target object identification device, including:
[0036] An acquisition module, configured to acquire multiple images captured by multiple image acquisition units at different shooting angles; wherein each image acquisition unit is mounted at a different position of the aircraft;
[0037] An image processing module is used to perform stereo restoration processing on multiple images to generate a stereo depth map;
[0038] A recognition module, configured to recognize the stereo depth map using a preset target recognition model to obtain a target recognition result;
[0039] The adjustment module is used to adjust the shooting angle of each image acquisition unit by controlling the position of the aircraft when the target recognition result is unrecognized, until the target recognition result obtained by recognizing multiple images captured based on the adjusted shooting angle includes the target body.
[0040] From the above description, it can be seen that the power grid target object identification method and device provided in the embodiment of the present application obtain multiple images captured by multiple image acquisition units at different shooting angles, perform stereo restoration processing on the multiple images, generate a stereo depth map, and use the target recognition model to identify the stereo depth map to obtain a target recognition result. When the target recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the target recognition result obtained based on the multiple images captured at the adjusted shooting angle includes the target object; the present application can improve the accuracy of identifying the target object. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0042] Figure 1 This is a schematic diagram of the method flow of an embodiment of the present application;
[0043] Figure 2 This is a block diagram of the device structure of an embodiment of the present application;
[0044] Figure 3 This is a schematic diagram of the electronic device structure frame of an embodiment of the present application. DETAILED DESCRIPTION
[0045] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0046] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present application should have the usual meanings understood by people with ordinary skills in the field to which this disclosure belongs. The "first", "second" and similar words used in the embodiments of the present application do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative positional relationships. When the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0047] like Figure 1 As shown, an embodiment of the present application provides a method for identifying a target object in a power grid, comprising:
[0048] S101: Acquire multiple images captured by multiple image acquisition units at different shooting angles; wherein each image acquisition unit is mounted at a different position of the aircraft;
[0049] In this embodiment, an aircraft is used to assist in the monitoring function of a power grid monitoring system. Image acquisition units are mounted at different locations on the aircraft, each with a different shooting angle. For a work area to be monitored, multiple image acquisition units on the aircraft capture images within their respective monitoring ranges at different shooting angles.
[0050] S102: Performing stereo restoration processing on the multiple images to generate a stereo depth map;
[0051] In this embodiment, for multiple images captured by multiple image acquisition units at different shooting angles, a stereo depth map is constructed through stereo restoration processing, which can effectively restore the structure, form and spatial relationship between different target objects, etc., which is conducive to accurate identification of the target object.
[0052] In some embodiments, stereo restoration is performed based on multiple images captured by each image acquisition unit at different angles and positions. The imaging restoration process can be expressed as:
[0053]
[0054] Where F(X, Y; Z) is the pixel value restored by imaging, (x, Y) is the pixel index, and Z is the restored depth; G(X, Y) is the cumulative overlap number at the pixel point (X, Y) (i.e., the number of images covering the same pixel point (X, Y) in multiple images); M and N are the total number of images in the row and column directions, respectively. Multiple images are sorted according to the spatial positions of the image acquisition units, and the arrangement order of the image acquisition units in the horizontal direction (e.g., horizontal direction) is regarded as the row, and the arrangement order in the vertical direction (e.g., vertical direction) is regarded as the column. The number of images M and N in the row and column directions are counted respectively, or the number of images obtained by counting the row and column order set in the preset coordinate system is: H (m,n) is the image of the mth row and nth column; L1 and L2 are images H (m,n) The total number of pixels in the mth row and nth column; S is the magnification of the image acquisition unit, which is equal to Z / G; G is the focal length, p1 is the distance between the center points of two adjacent image acquisition units in the row direction, and p2 is the distance between the center points of two adjacent image acquisition units in the column direction; c1 and c2 are the width and height of the image acquisition unit respectively; m(L1p1) and n(Y2p2) are the offsets of the image in the mth row and nth column respectively, which are used to adjust the position of pixels and the position of the image in space during the imaging restoration process.
[0055] In some embodiments, the image captured by the image acquisition unit can be represented as:
[0056] H(X,Y)=F(X,Y)·R(X,Y) (2)
[0057] Among them, F(X,Y)>0 is the illumination factor, R(X,Y) is the reflection coefficient, and its value range is 0-1; H(X,Y) is the pixel value of the image at the pixel point (X,Y), that is, the light intensity or brightness received by the image acquisition unit, which is the comprehensive result of the lighting conditions and the reflection characteristics of the object surface.
[0058] Since the image acquisition unit has read noise, which conforms to the zero-mean Gaussian distribution, the noise influences the imaging recovery formula, which can be expressed as:
[0059]
[0060] Among them, ∈ (m,n) (X′, Y′) is the zero-mean additive Gaussian noise at the pixel point (X′, Y′) in the image at the mth row and nth column; (X′, Y′) is the pixel affected by the noise and includes the offset effect (i.e. )’s pixel index.
[0061] Assuming that the read noise is wide-sense stationary, for a fixed recovery depth Z, the variance of the noise component is:
[0062]
[0063] where σ 2 is the variance of the read noise.
[0064] As the number of acquired images increases, the variance of the read noise decreases, and the final signal-to-noise ratio (SNR) can be expressed as:
[0065]
[0066] Among them, G m is the average energy of the target area where the target object is located in the acquired image, M 2 The M in the equation is the number of photons on the pixel, M 2 The expected value of is defined as:
[0067]
[0068] Where Ψ1 is the photon flux density in the target area where the target object is located, Ψ b is the photon flux density of the background area excluding the target object, in photons / pixel / second; ΔΔ d is the dark current, in electrons / pixel / second; η q is the photoelectric conversion efficiency, in electrons / photons; τ is the exposure time, in seconds; σ r is the read noise in RMS electrons / pixel / second.
[0069] The number of photons per pixel N ph The estimate can be expressed as:
[0070]
[0071] Among them, N ph To estimate the number of photons in a pixel, we can use the estimated number of photons N ph , by adjusting the relevant parameters of the image acquisition unit, the number of captured photons can be increased, thereby improving the image quality.
[0072] The stereo depth map restored according to formula (3) contains noise components. The stereo depth map is denoised using a preset denoising algorithm to obtain a denoised stereo depth map, thereby improving image quality and the accuracy of target recognition. In some methods, a pre-built image denoising model can be used to denoise the stereo depth map. For example, a convolutional neural network (CNN) or a generative adversarial network (GAN) is trained based on a number of stereo depth map samples containing noise components. After training, an image denoising model that can remove noise components in the stereo depth map is obtained.
[0073] The imaging restoration process shown in formula (3) is to simulate the light in the image fragment and transmit it back to the preset stereo depth plane of the virtual imaging sensor, thereby obtaining images at different depth positions. That is, the light information of the three-dimensional object or scene is projected or mapped onto a two-dimensional plane in some way, so as to reconstruct the image of the scene on the two-dimensional plane. In technical implementation, two-dimensional images of the three-dimensional scene can be captured from multiple perspectives, and the data of the two-dimensional images from multiple perspectives can be mapped to the preset depth plane using back projection technology. Through the above process, the three-dimensional scene can be reconstructed in different depth layers to obtain an image with depth information.
[0074] S103: Using a preset target recognition model to recognize the stereo depth map, and obtaining a target recognition result;
[0075] S104: When the target recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body.
[0076] In this embodiment, after a stereo depth map is reconstructed based on multiple images acquired from multiple shooting angles, the stereo depth map is input into a pre-built target recognition model, and the target recognition model is used to perform recognition processing on the stereo depth map to output a target recognition result.
[0077] If the quality of the multiple images is high and the reconstructed stereo depth image can effectively reflect the target object, the target recognition result includes the identified target object and its related information. In some embodiments, in the application scenario of the power grid monitoring system, the target objects to be identified include equipment, personnel, tools, etc., and the related information of the target object includes the name, type, status, trajectory, etc. of the target object. For example, the target recognition model determines the name, department, movement trajectory, etc. of the person through facial feature recognition, and determines the name, type, current status, etc. of the tool through object feature recognition.
[0078] If the quality of multiple images is poor due to factors such as environmental conditions and shooting angles, and the reconstructed stereo depth image cannot effectively represent the target, the target recognition result will be an unidentified or unidentified suspected target. For example, in urban areas, the presence of obstructions between the image acquisition unit and the target may cause the captured image to include a partially or completely obscured target, making it impossible to effectively identify the target. Alternatively, due to lighting conditions, the captured image may not include a clear target, resulting in low image quality and making it difficult to accurately identify the target.
[0079] When the target recognition result is an unrecognized or unidentified suspected target, the aircraft is controlled to move to another position, and multiple images are re-collected at the other position. A stereo depth map is reconstructed based on the multiple images, and the stereo depth map is recognized using the target recognition model to obtain the target recognition result at that position. If the target is still not recognized, the process is repeated until the aircraft moves to a specific position, a better image can be collected at the specific position, and recognition is performed based on the reconstructed stereo depth, and the target recognition result obtained includes the determined target.
[0080] In some embodiments, the target recognition model includes a feature extraction module, a classifier, etc. The target recognition model extracts features from the input stereo depth map, calculates the probability value of belonging to one or more target objects based on the extracted features, and compares the probability value with a preset classification threshold. If it is greater than the classification threshold, the target object corresponding to the feature is determined, that is, the category of the target object is identified. If it is less than the classification threshold, no target object of any category is identified. Alternatively, a first classification threshold and a second classification threshold can be set, and the first classification threshold is greater than the second classification threshold. When the probability value is greater than the first classification threshold, the identified target object can be determined. When the probability value is greater than the second classification threshold and less than the first classification threshold, it can be classified as a suspected target object. The suspected target object has some features of the target object but is incomplete due to occlusion or other reasons, and therefore cannot be accurately determined to belong to a certain target object. Optionally, the score threshold, the first score threshold, and the second score threshold can be set fixed values, or can be dynamically adjusted according to image quality to adapt to different external conditions.
[0081] In some embodiments, the flight trajectory of an aircraft can be controlled according to a preset path planning algorithm. In some embodiments, within the aircraft's monitoring work area, the aircraft's flight trajectory is planned based on the aircraft's distance, height, and angle relative to the target, as well as the angles and shooting parameters of each image acquisition unit relative to the target. In other embodiments, the aircraft's flight trajectory can be dynamically planned based on the characteristics and quality of the images captured by the image acquisition units. For example, based on the lighting characteristics in the image, the aircraft's position and angle relative to the target can be adjusted to an optimal position, and / or the angles and shooting parameters of each image acquisition unit can be adjusted so that each image acquisition unit can capture an image with appropriate lighting at the optimal position, with the target clearly visible in the image, thereby improving the target recognition success rate. Alternatively, when image recognition determines the presence of a possible obstruction, the aircraft's position and angle relative to the target can be adjusted to an optimal position so that the images captured by each image acquisition unit at the optimal position are clear of obstructions and can clearly and completely represent the target, thereby improving target recognition accuracy. Furthermore, high-quality images can support less complex target recognition models, reducing resource usage and performance requirements, and improving the overall performance of the monitoring system.
[0082] The power grid target object recognition method provided by this embodiment uses multiple image acquisition units carried by an aircraft to capture multiple images of the working area, reconstructs a stereo depth map based on the multiple images, and uses a target recognition model to recognize the stereo depth map to obtain a target recognition result. If the target recognition result is not recognized, that is, the aircraft does not recognize the target object at the current position, the aircraft is controlled to another position, and multiple image acquisition units are used to re-capture multiple images at multiple shooting angles at other positions, and the target object is re-identified based on the multiple images. This process is repeated until when the aircraft is at a specific position, each image acquisition unit captures multiple images with better results, so that the target object can be recognized based on the multiple images. Using the method of the present application, even if factors such as environmental conditions change or are not conducive to shooting, high-quality images can still be obtained by adjusting the shooting position, shooting angle, etc., and a stereo depth map is obtained by stereo restoration of multiple images. Identification based on the stereo depth map can improve the recognition accuracy of the target object.
[0083] In some embodiments, after generating the stereo depth map, the method further includes:
[0084] The preset palmprint recognition model is used to identify the three-dimensional depth map to obtain the palmprint recognition result.
[0085] In this embodiment, when the target object includes a person, one or more of the multiple images captured by the multiple image acquisition units include the hand texture information of the person, and the three-dimensional depth map reconstructed based on the multiple images includes the hand texture information of the person. The three-dimensional depth map is identified using the palmprint recognition model to obtain the person and his / her palmprint information, and then the identity of the person can be determined based on the palmprint information, thereby improving the accuracy and reliability of target recognition.
[0086] In some implementations, using a preset palmprint recognition model to recognize the 3D depth map to obtain a palmprint recognition result includes:
[0087] Extract palmprint encoding features from stereo depth map;
[0088] Calculate the differentiation index of each feature bit in the palmprint encoding feature;
[0089] Construct identity information based on the differentiation indicators of each feature bit;
[0090] According to the identity identification information, the target object corresponding to the identity identification information and the identity information of the target object are identified.
[0091] In this embodiment, a palmprint recognition model is used to extract palmprint encoding features from a 3D depth map. These features consist of multiple binary feature bits. For each feature bit, a corresponding differentiation index is calculated. Based on the differentiation index for each feature bit, identity information is constructed. If the 3D depth map includes clearly visible palmprint information, valid palmprint encoding features can be extracted from the 3D depth map, and identity information can be constructed. Based on this identity information, the corresponding person and their related identity information can be identified.
[0092] In some methods, the binary palmprint encoding feature is represented as: V k =[v k,1 ,v k,2 ,v k,3 ,…,v k,L ],k∈[1,2,…,T], where V k is the kth palmprint encoding feature, T is the number of palmprint encoding features, v k,L is the Lth feature bit, L is the length of the palmprint encoding feature, that is, the number of bits contained in the palmprint encoding feature, and the value set of each feature bit is {0,1}.
[0093] Define the differentiation index of the feature bit, which is used to describe the contribution of the feature bit to the distinguishing ability of the palmprint encoding feature. The larger the value of the differentiation index, the greater the contribution of the feature bit to the distinguishing ability. The method for calculating the differentiation index of each feature bit in the palmprint encoding feature is:
[0094]
[0095] Among them, R i is the differentiation index of the i-th feature bit, D i is the inter-category difference of the i-th feature bit, which is used to measure the distribution difference between different categories in this feature bit; W i is the intra-category difference of the i-th feature bit, which is used to measure the distribution difference of the palmprint encoding features in the same category at this feature bit.
[0096] Assuming there are P categories (e.g., identity categories), each category includes Q palmprint encoding features, the calculation method for the inter-category difference of the feature bits is:
[0097]
[0098] Among them, z b is the mean value of the b-th category at the i-th feature position, indicating the average value of the palmprint encoding feature of this category at this feature position; z is the mean value of all palmprint encoding features at the i-th feature position, and the calculation method is:
[0099]
[0100] Where PQ is the total number of all palmprint encoding features, v k,i is the i-th feature bit in the k-th palmprint encoding feature.
[0101] The calculation method of intra-category difference is:
[0102]
[0103] In some approaches, after calculating the differentiation index for each feature bit, efficient and accurate identification information is generated through sorting and optimization based on the differentiation index of each feature bit, thereby improving the accuracy of identifying the identity of the target object. Optionally, the feature bits can be re-sorted in descending order of differentiation index, and the sorted code is used as the optimized identification information.
[0104] In some embodiments, after obtaining a target recognition result by using a preset target recognition model to recognize a stereo depth map, the following steps may be performed:
[0105] Determine the final recognition result based on the palmprint recognition result and the target recognition result;
[0106] When the target recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body, including:
[0107] When the final recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the final recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angles includes the target object.
[0108] In this embodiment, after determining the stereo depth map, the target recognition model and the palmprint recognition model can be used simultaneously to identify the stereo depth map, and the final recognition result can be determined based on the combined recognition results of the two. If neither the target recognition result nor the palmprint recognition result identifies the target object, the aircraft is controlled to move to another position, and multiple images are recaptured using each image acquisition unit. A stereo depth map is constructed based on the multiple images. The target recognition model and the palmprint recognition model are re-used to recognize the stereo depth map, and the final recognition result is determined based on the recognition results of the two. This process is repeated until the aircraft moves to the optimal position, the multiple images captured by each image acquisition unit are of high quality, and the constructed stereo depth map can more clearly display the target object, thereby enabling the target recognition model and the palmprint recognition model to identify the target object and identity information based on the stereo depth map. By fusing the recognition results of the target recognition model and the palmprint recognition model, the recognition accuracy of the target object and its related information can be improved.
[0109] In some embodiments, determining a final recognition result based on the palmprint recognition result and the target recognition result includes:
[0110] According to the image quality, the weight values of the palmprint recognition result and the target recognition result are determined respectively;
[0111] The final recognition result is determined based on the palmprint recognition result and the corresponding weight value, the target recognition result and the corresponding weight value.
[0112] In this embodiment, the palmprint recognition result and the target recognition result are weighted and summed to determine the final recognition result, thereby achieving the fusion of the recognition results. The weight values of the palmprint recognition result and the target recognition result can be determined according to the image quality. For example, if the image quality is high, the weight value of the target recognition result can be set to 0.7, and the weight value of the palmprint recognition result can be set to 0.3; if the image is not clear, the weight value of the target recognition result can be set to 0.3, and the weight value of the palmprint recognition result can be set to 0.7. In other ways, the specific weight value can also be determined through experience in combination with the specific application scenario. The above is only an exemplary description, and this application does not limit the specific setting method of the weight value.
[0113] In some embodiments, a method for determining the final recognition result based on the palmprint recognition result and the corresponding weight value, the target recognition result and the corresponding weight value may be feature-level fusion, which is expressed as:
[0114] F fusion=α·F image +(1-α)·F palm (13)
[0115] Among them, F image is the feature vector generated by the target recognition model, F palm is the feature vector generated by the palmprint recognition model, α is the weight value corresponding to the target recognition result; F fusion The fused feature vector is then classified by the classifier to determine the final recognition result.
[0116] In some embodiments, the method for determining the final recognition result based on the palmprint recognition result and the corresponding weight value, the target recognition result and the corresponding weight value can be a decision-level fusion, which is expressed as:
[0117] P final =α·P image +(1-α)·P palm (14)
[0118] Among them, P image is the predicted probability value corresponding to the target recognition result, P palm is the predicted probability value corresponding to the palmprint recognition result.
[0119] In some embodiments, the training method of the target recognition model includes:
[0120] Get image samples;
[0121] Performing stereo restoration processing on the image samples to obtain stereo depth map samples;
[0122] Performing data enhancement processing on the stereo depth map samples to obtain enhanced stereo depth map samples;
[0123] The neural network model is trained based on the enhanced stereo depth map samples, and an object recognition model is obtained after training.
[0124] This embodiment provides a method for training a target recognition model. A plurality of images of a target object captured by multiple image acquisition units in a working area under different environmental conditions are obtained, and an image sample set is constructed based on the captured images. Multiple images of the working area can also be generated based on adversarial learning, and the captured images and generated images are mixed in a predetermined ratio to construct an image sample set. Stereo restoration processing is performed on the images in the image sample set to obtain a plurality of stereo depth maps. The plurality of stereo depth maps are divided into a training set and a test set as training sample sets. Some stereo depth maps with a low signal-to-noise ratio can be included in the test set to enhance the generalization ability of the model.
[0125] In some methods, for the processed stereo depth maps, data samples are expanded through data augmentation to improve the generalization ability of the model. Among them, the methods for data augmentation of stereo depth maps include one or more of the following:
[0126] Multiple independent areas of the image are cropped according to regular geometric shapes (such as squares, hexagons, etc.) to obtain cropped stereo depth maps; multiple independent areas of the image are cropped according to dynamically generated shapes to obtain cropped stereo depth maps; wherein, the independent areas can be areas that do not include the target object or key areas that include the target object. Such samples can increase the model's attention to the key areas where the target object is located and reduce interference information.
[0127] The stereo depth map is cropped by randomly selecting cropping regions to obtain a cropped stereo depth map. The cropped regions may overlap to a certain extent. The stereo depth map is cropped uniformly to obtain a uniformly cropped stereo depth map. The stereo depth map is cropped to varying degrees based on the key areas where the target is located. For example, the key areas may be cropped multiple times according to different shapes. The cropped regions may overlap, and non-key areas may be cropped slightly. This type of sample can simulate different states of the target (e.g., the working state of a person, the use state of a tool, the interactive state of a person using a tool, etc.), enhancing the model's ability to recognize targets in different states.
[0128] Key areas of the image are randomly cropped according to different shapes, with overlapping areas. This type of sample allows the model to learn the various forms of people or tools at different positions and angles, enhancing the model's adaptability to the target and improving recognition stability.
[0129] Stereo depth maps are processed using color space conversion, contrast adjustment, and / or random noise addition to generate processed stereo depth maps. This allows the model to focus more on structural information rather than color during learning. For example, stereo depth maps can be converted to grayscale images and the color histogram can be perturbed under specific conditions. Generating diverse data samples through different data augmentation methods enables the model to learn the characteristics of the target under different visual conditions, improving its adaptability and robustness and enabling it to more effectively adapt to diverse environments.
[0130] In some approaches, data samples can be adjusted at different stages of model training. For example, during the initial training phase, random cropping can be used to augment data samples, allowing for some overlap to enhance the model's ability to recognize local features. As training progresses, a strict cropping strategy can be employed to reduce overlap and improve the model's ability to understand the overall structure.
[0131] In some embodiments, during the training process, the target recognition model performs local area perturbations on the input stereo depth map samples, that is, introduces slight changes or perturbations in the local area of the stereo depth map samples, such as adding noise, geometric deformation, depth offset, changing color, etc., to enhance the robustness of the model in identifying the target body. The model extracts different types of features based on the local area, and splices the extracted features into high-dimensional feature vectors, so that features of the same category can be close in the feature space, while features of different categories have greater discrimination, thereby improving the model's ability to extract more discriminative feature representations. According to the above method, different transformation methods are used to transform the local area, and the corresponding high-dimensional feature vectors are obtained through feature extraction and splicing. The high-dimensional feature vectors corresponding to different transformation methods are weighted and calculated to obtain fused features, thereby improving the model's ability to comprehensively identify the target body and improving the stability of the model.
[0132] In some embodiments, during the training process, a preset optimization framework is used to establish a metric relationship between feature distributions between different types of stereo depth map samples. A feature embedding mechanism of a neural network is employed, and a specific metric method is used to optimize the model so that samples of the same category are closer in the feature space, and samples of different categories are further apart. This ensures that targets of the same category (e.g., the same tool, a person in the same posture) are clustered in the feature space, and targets of different categories have greater distinction. For example, cosine similarity is used to measure the directional similarity of two feature vectors. By maximizing the cosine similarity of samples of the same category and minimizing the similarity of samples of different categories, the model is helped to identify the directionality of features rather than simply numerical differences, thereby improving the model's ability to distinguish different targets.
[0133] In some embodiments, a local transformation operation is applied to the labeled data samples, and a specific feature mapping method is used to map the transformed data samples to a feature representation of a preset dimension. The transformed feature representation is used to train the model, and the stability is further improved by optimizing the framework in the subsequent training stage. Optionally, the locally transformed data samples are processed by a convolution layer, a pooling layer, etc. to generate a high-dimensional feature vector, and the high-dimensional feature vector is mapped to a low-dimensional feature space to obtain a feature representation of the required dimension. By converting the feature dimension, different features of different target objects (such as the shape of a tool, the movement of a person, etc.) can be mapped to corresponding feature representations, thereby improving the model's ability to recognize different target objects.
[0134] In some embodiments, the target recognition model can perform multi-level feature matching based on features from multiple perspectives, and determine the target recognition result based on the multi-level feature matching results. For example, a multi-layer convolutional neural network is used to extract features of different scales from a stereo depth map. In each layer, convolution kernels are used to process image areas of different sizes, and features of different levels are extracted from image areas of different sizes. Afterwards, a preset matching algorithm (such as a scale-invariant feature transformation algorithm) is used to perform preliminary matching at a lower level using fewer features. Based on the results of the preliminary matching, a preset matching algorithm (such as a VGG model, Visual Geometry Group Network) is used to match again at a high level to obtain corresponding matching results. Combining the multi-level feature matching results to determine the target recognition result can reduce the amount of calculation and improve the matching accuracy. Optionally, a voting mechanism is used to merge multiple feature matching results to improve the accuracy of the recognition results.
[0135] In some embodiments, a neural network model is trained based on enhanced stereo depth map samples. The neural network model includes an input layer, a first convolutional layer, a first pooling layer, a second convolutional layer, a second pooling layer, a fully connected layer, an output layer, etc. Among them, the first convolutional layer includes 32 13×13 convolution kernels, uses a sliding window with a step size of 1 for feature extraction, and performs nonlinear transformation through the ReLU activation function. The first pooling layer includes a 2×2 maximum pooling layer with a step size of 2, which is used to reduce the dimension of the feature map while retaining key features. The second convolutional layer includes 64 11×11 convolution kernels for extracting deep features and is processed using the ReLU activation function. The second pooling layer includes a 2×2 maximum pooling layer with a step size of 2 for compressing the feature map and improving computational efficiency. The fully connected layer includes 128 neurons and is processed through the ReLU activation function. The output layer is composed of a SoftMax layer with 2 output units for outputting target bodies of different categories to achieve the final classification decision. During the training process, the stochastic gradient descent (SGD) optimization algorithm was used, the learning rate was set to 0.0001, and the cross entropy loss function was used for model evaluation.
[0136] The power grid target object recognition method provided in the embodiment of the present application utilizes an aircraft equipped with multiple image acquisition units to capture multiple images from different angles, performs stereo restoration based on the multiple images to construct a stereo depth map, and uses a target recognition model to recognize the stereo depth map. If the target object is not recognized, the position of the aircraft is controlled, the position and angle of each image acquisition unit are adjusted, multiple images are recaptured, and recognition is re-performed based on the multiple images. This process is repeated until the aircraft can recognize the target object at a specific position. By continuously adjusting the shooting position and angle, a clearly visible image of the target object can be obtained, thereby improving the accuracy of target object recognition. At the same time, palm print recognition can also be performed on the stereo depth map, and the target recognition result and the palm print recognition result are merged to obtain the final recognition result, thereby improving the accuracy and reliability of target object recognition.
[0137] It should be noted that the method of the embodiment of the present application can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario and performed by multiple devices working together. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present application, and the multiple devices will interact with each other to complete the method.
[0138] It should be noted that the foregoing description of this specification is based on specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0139] like Figure 2 As shown, an embodiment of the present application provides a power grid target object identification device, comprising:
[0140] An acquisition module, configured to acquire multiple images captured by multiple image acquisition units at different shooting angles; wherein each image acquisition unit is mounted at a different position of the aircraft;
[0141] An image processing module is used to perform stereo restoration processing on multiple images to generate a stereo depth map;
[0142] The recognition module is used to recognize the stereo depth map using a preset target recognition model to obtain a target recognition result;
[0143] The adjustment module is used to adjust the shooting angle of each image acquisition unit by controlling the position of the aircraft when the target recognition result is unrecognized, until the target recognition result obtained by recognizing multiple images captured based on the adjusted shooting angle includes the target body.
[0144] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0145] The apparatus of the above embodiment is used to implement the corresponding method in the above embodiment and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0146] Figure 3 10 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0147] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0148] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0149] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0150] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (such as USB, network cable, etc.) or a wireless method (such as mobile network, WiFi, Bluetooth, etc.).
[0151] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0152] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0153] The electronic devices of the above embodiments are used to implement the corresponding methods in the above embodiments and have the beneficial effects of the corresponding method embodiments, which will not be described in detail here.
[0154] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0155] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the above embodiments or technical features in different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present application as described above, which are not provided in detail for the sake of simplicity.
[0156] In addition, to simplify the description and discussion, and in order not to make the embodiment of the present application difficult to understand, the known power supply / ground connection with the integrated circuit (IC) chip and other components may or may not be shown in the accompanying drawings provided. In addition, the device can be shown in the form of a block diagram to avoid making the embodiment of the present application difficult to understand, and this also takes into account the following fact, that is, the details of the implementation method of these block diagram devices are highly dependent on the platform to be implemented in the embodiment of the present application (that is, these details should be fully within the scope of understanding of those skilled in the art). When specific details (for example, circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiment of the present application can be implemented without these specific details or when these specific details are changed. Therefore, these descriptions should be considered to be illustrative rather than restrictive.
[0157] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0158] The embodiments of the present application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present application should be included in the scope of protection of this disclosure.
Claims
1. A method for identifying a target object in a power grid, characterized in that: include: Acquire multiple images captured by multiple image acquisition units at different shooting angles; wherein each image acquisition unit is mounted at a different position of the aircraft; Perform stereo restoration processing on multiple images to generate a stereo depth map; Using a preset target recognition model to identify the stereo depth map, to obtain a target recognition result; When the target recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body.
2. The method according to claim 1, characterized in that The performing stereo restoration processing on the multiple images to generate a stereo depth map includes: Perform imaging restoration processing on multiple images to obtain a restored stereo depth map; The restored stereo depth map is denoised to obtain a denoised stereo depth map.
3. The method according to claim 2, characterized in that The imaging restoration process is performed on the multiple images to obtain the restored stereo depth map, and the method is: Where (X′, Y′) is the index of the pixel affected by noise and including the offset; G(X, Y) is the cumulative overlap number at the (X, Y) pixel point; M and N are the total number of images collected by the image acquisition unit in the preset row and column directions respectively: H (m,n) is the image of the mth row and nth column; Z is the restored depth; ∈ (m,n) (X′, Y′) is the zero-mean additive Gaussian noise at the pixel point (X′, Y′) in the mth row and nth column of the image.
4. The method according to claim 1, wherein After generating the stereo depth map, the method further includes: The palmprint recognition model is used to identify the three-dimensional depth map to obtain a palmprint recognition result.
5. The method according to claim 4, characterized in that The three-dimensional depth map is recognized using a preset palmprint recognition model to obtain a palmprint recognition result, including: Extracting palmprint encoding features from the stereo depth map; Calculating a differentiation index for each feature bit in the palmprint encoding feature; Construct identity information based on the differentiation indicators of each feature bit; According to the identity identification information, the target object corresponding to the identity identification information and the identity information of the target object are identified.
6. The method according to claim 5, characterized in that The method for calculating the differentiation index is: Among them, R i is the differentiation index of the i-th feature bit, D i is the difference between categories of the i-th feature bit, W i is the intra-category difference of the i-th feature bit.
7. The method according to claim 4, characterized in that After obtaining the target recognition result, the following steps are included: Determining a final recognition result based on the palmprint recognition result and the target recognition result; When the target recognition result is unrecognized, adjusting the shooting angle of each image acquisition unit by controlling the position of the aircraft until the target recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angle includes the target body, including: When the final recognition result is unrecognized, the shooting angle of each image acquisition unit is adjusted by controlling the position of the aircraft until the final recognition result obtained by recognizing multiple images acquired based on the adjusted shooting angles includes the target object.
8. The method according to claim 7, characterized in that The method for determining the final recognition result based on the palmprint recognition result and the target recognition result is as follows: F fusion =α·F image +(1-α)·F palm (13) Among them, F image is the feature vector generated by the target recognition model, F palm is the feature vector generated by the palmprint recognition model, α is the weight value corresponding to the target recognition result; F fusion The fused feature vector is classified by a classifier to determine the final recognition result.
9. The method according to claim 7, characterized in that The final recognition result is determined based on the palmprint recognition result and the target recognition result, in the following method: P final =α·P image +(1-α)·P palm (14) Among them, P image is the predicted probability value corresponding to the target recognition result, P palm is the predicted probability value corresponding to the palmprint recognition result.
10. A power grid target object identification device, characterized in that: include: An acquisition module, configured to acquire multiple images captured by multiple image acquisition units at different shooting angles; Among them, each image acquisition unit is mounted at a different position of the aircraft; An image processing module is used to perform stereo restoration processing on multiple images to generate a stereo depth map; A recognition module, configured to recognize the stereo depth map using a preset target recognition model to obtain a target recognition result; The adjustment module is used to adjust the shooting angle of each image acquisition unit by controlling the position of the aircraft when the target recognition result is unrecognized, until the target recognition result obtained by recognizing multiple images captured based on the adjusted shooting angle includes the target body.