Detection method, detection device, electronic equipment and computer storage medium

Through the object detection model trained by Transformer model, the problem of insufficient accuracy of traditional power equipment detection methods in complex scenarios is solved, and efficient and stable detection of equipment attitudes and states is achieved, adapting to environmental changes, reducing operation and maintenance costs, and improving power system safety.

CN120339567APending Publication Date: 2025-07-18BEIJING ESWIN COMPUTING TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510338760.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional power equipment detection methods are difficult to accurately detect equipment status and attitude in complex scenarios and irregular changes, especially in the face of environmental changes such as light and weather. Moreover, traditional visual change detection methods are poorly generalized and difficult to adapt to changeable operating scenarios.

Method used

The object detection model is trained using the Transformer model, and the detection is carried out by obtaining the benchmark image and the image to be tested, and the visual characteristics of the equipment's posture and state are automatically learned, the changing areas are identified, and the impact of environmental changes is resisted in complex scenarios.

Benefits of technology

It improves the accuracy and stability of equipment attitude and status detection, reduces operation and maintenance costs, improves the safety and reliability of the power system, and can output equipment change information in a timely and accurate manner.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339567A_ABST
    Figure CN120339567A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a detection method, a detection device, electronic equipment and a computer storage medium, and the detection method comprises the steps: obtaining a picture pair which comprises a reference picture and a to-be-detected picture; the picture pair is obtained by shooting at least one piece of to-be-tested equipment at the same position and at different moments; detecting the picture pair by using a target detection model so as to determine a changed target area of the to-be-detected equipment; the target detection model is obtained by training an initial detection model by using at least one group of sample picture pairs; the change of the to-be-tested equipment comprises the change of the posture of the to-be-tested equipment and / or the change of the state of the to-be-tested equipment; the initial detection model is a Transform model. According to the embodiment of the invention, the accuracy of posture and / or state change recognition of the to-be-tested equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of substation change detection, and in particular, to a detection method, a detection device, an electronic device, and a computer storage medium. Background Art

[0002] In the scenario of a power substation, accurately detecting the state and attitude changes of equipment has important safety and operation and maintenance significance. Traditional power equipment detection methods have some deficiencies in dealing with complex scenarios and irregular changes. For example, accurately detecting the state and attitude of equipment often requires a large number of manually defined rules and is difficult to adapt to changing operation scenarios.

[0003] In recent years, the rapid development of deep learning technology has provided new ideas for solving this problem. However, due to poor generalization, the change detection method based on traditional vision is difficult to accurately perceive the state changes of equipment, resulting in a decrease in accuracy. Summary of the Invention

[0004] Embodiments of the present disclosure provide a detection method, a detection device, an electronic device, and a computer storage medium for improving the detection accuracy in the scenario of detecting the attitude and / or state of a device to be measured.

[0005] In a first aspect, embodiments of the present disclosure provide a detection method, the detection method including:

[0006] Obtaining a pair of pictures, the pair of pictures including a reference picture and a picture to be measured; the pair of pictures is obtained by photographing at least one device to be measured at the same position and different times;

[0007] Using a target detection model to detect the pair of pictures to determine a target area where the device to be measured has changed; the target detection model is obtained by training an initial detection model using at least one set of sample picture pairs; the change of the device to be measured includes: the attitude of the device to be measured has changed, and / or, the state of the device to be measured has changed; the initial detection model is a Transformer model.

[0008] In some embodiments, the method further includes:

[0009] Judging whether the target area meets a preset condition;

[0010] If the target area meets the preset condition, an alarm is issued;

[0011] If the target area does not meet the preset condition, no alarm is issued.

[0012] In some embodiments, the preset condition is that the area of the target region is greater than a preset threshold, and / or the target device to be measured among the at least one device to be measured changes.

[0013] In some embodiments, the method further includes:

[0014] Obtain a sample data set; the sample data set includes the at least one set of sample picture pairs, and the sample picture pair includes a reference sample picture and a sample picture to be measured;

[0015] Label at least one sample device in each of the sample picture pairs to obtain labeling information; the labeling information is used to indicate that the sample device has changed;

[0016] Use the at least one set of sample picture pairs to train the initial detection model to obtain the target detection model.

[0017] In some embodiments, the initial detection model includes a first feature extraction module, a second feature extraction module, a difference module, an upsampling module, and a decoding module; both the first feature extraction module and the second feature extraction module are Transformer blocks;

[0018] The first feature extraction module is used to extract features from the reference sample picture in the sample picture pair to obtain first feature information;

[0019] The second feature extraction module is used to extract features from the sample picture to be measured in the sample picture pair to obtain second feature information;

[0020] The difference module is used to perform difference processing on the first feature information and the second feature information to obtain difference information;

[0021] The upsampling module is used to upsample the difference information to obtain upsampled information;

[0022] The decoding module is used to decode the upsampled information to obtain the changed area where the sample device has changed.

[0023] In some embodiments, the changed area is represented by a target box or a segmentation mask;

[0024] If the changed area is represented by the target box, a first type of loss function is used in the process of training the initial detection model;

[0025] If the changed area is represented by the segmentation mask, a second type of loss function is used in the process of training the initial detection model;

[0026] Among them, the first type of loss function and the second type of loss function are different.

[0027] In some embodiments, after obtaining the sample data set, the method further includes:

[0028] Performing data augmentation on the sample data set to increase the data volume of the sample data set.

[0029] In some embodiments, in the pair of pictures, the reference picture and the picture to be tested are separated by a preset number of frames, and the reference picture is a historical picture and the picture to be tested is a current picture.

[0030] In a second aspect, an embodiment of the present disclosure provides a detection device, and the detection device includes:

[0031] An acquisition unit, configured to acquire a pair of pictures, the pair of pictures including a reference picture and a picture to be tested; the pair of pictures is obtained by photographing at least one device to be tested at the same position and different times;

[0032] A detection unit, configured to use a target detection model to detect the pair of pictures to determine a target area where the device to be tested has changed; the target detection model is obtained by training an initial detection model using at least one set of sample picture pairs; the change of the device to be tested includes: the posture of the device to be tested changes, and / or, the state of the device to be tested changes; the initial detection model is a Transformer model.

[0033] In a third aspect, an embodiment of the present disclosure provides an electronic device, and the electronic device includes a memory and a processor;

[0034] The memory is used to store a computer program that can run on the processor;

[0035] The processor is configured to execute the steps of the method according to any one of the first aspects when running the computer program.

[0036] In a fourth aspect, an embodiment of the present disclosure provides a computer storage medium, and the computer storage medium stores a computer program, and the computer program realizes the steps of the method according to any one of the first aspects when executed by at least one processor.

[0037] Embodiments of the present disclosure provide a detection method, a detection device, an electronic device, and a computer storage medium. The detection method includes: obtaining a pair of pictures, where the pair of pictures includes a reference picture and a picture to be measured; the pair of pictures is obtained by photographing at least one device to be measured at the same position and different times; using a target detection model to detect the pair of pictures to determine a target area where the device to be measured has changed; the target detection model is obtained by training an initial detection model using at least one set of sample picture pairs; a change in the device to be measured includes: a change in the posture of the device to be measured and / or a change in the state of the device to be measured; the initial detection model is a Transformer model. In this way, the reference picture and the picture to be measured obtained by photographing at least one device to be measured are input into the target detection model to obtain the target area where the posture and / or state of the device to be measured has changed, so that the advantages of the Transformer model can be fully utilized to achieve efficient detection of the posture and / or state change of the device to be measured; at the same time, the visual features of the device posture and / or state can be automatically learned in complex scenarios to improve the accuracy and adaptability of detection; and it can resist the influence of environmental changes such as light and weather, improving the stability and accuracy of the posture and / or state detection of the device to be measured; in addition, the change information of the device to be measured can be output in a timely and accurate manner, reducing the operation and maintenance cost and improving the safety and reliability of the power system. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a schematic flowchart of a detection method provided by an embodiment of the present disclosure Figure 1 ;

[0039] Figure 2 is a schematic flowchart of a detection method provided by an embodiment of the present disclosure Figure 2 ;

[0040] Figure 3 is a schematic flowchart of a model training method provided by an embodiment of the present disclosure Figure 1 ;

[0041] Figure 4 is a schematic structural diagram of an initial detection model provided by an embodiment of the present disclosure;

[0042] Figure 5 is a schematic flowchart of a model training method provided by an embodiment of the present disclosure Figure 2 ;

[0043] Figure 6 is a schematic structural diagram of a detection device provided by an embodiment of the present disclosure;

[0044] Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure Figure 1 ;

[0045] Figure 8 Structural schematic of an electronic device provided by an embodiment of the present disclosure Figure 2 。 Specific embodiments

[0046] Next, the technical solutions in the embodiments of the present disclosure will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present disclosure. It can be understood that the specific embodiments described herein are only used to explain the relevant disclosure, rather than limiting the disclosure. In addition, it should be noted that for the convenience of description, only the parts related to the relevant disclosure are shown in the drawings.

[0047] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this disclosure belongs. The terms used herein are only for the purpose of describing the embodiments of the present disclosure and are not intended to limit the present disclosure.

[0048] In the following description, reference is made to "some embodiments", which describe a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.

[0049] It should be noted that the terms "first / second / third" involved in the embodiments of the present disclosure are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second / third" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein.

[0050] Currently, with the development of social economy, the requirements for power supply in various industries are getting higher and higher. The stable operation of substations plays a crucial role in providing stable and reliable power supply, so it is necessary to accurately detect the state and attitude changes of secondary equipment, fire protection equipment and other equipment. However, traditional equipment state detection methods mainly rely on data collected by sensors and rule definitions, and these methods often fail to achieve satisfactory results when faced with complex scenarios and changing operating states.

[0051] In recent years, the Transformer model has made many breakthroughs in the field of computer vision. Among them, the Transformer model, with its powerful modeling ability and parallel computing advantages, performs excellently in capturing long-distance dependencies and processing large-scale data, providing a more generalizable algorithm solution for change detection in scenarios such as substations.

[0052] However, due to poor generalization, the change detection methods based on traditional vision are difficult to accurately perceive the changes in the device state. For example, the interference of environmental changes on traditional vision change detection methods. The detection of device state and pose changes based on traditional vision algorithms is often interfered by environmental changes such as light and weather, resulting in a decrease in accuracy. And the problem of device change detection in complex scenarios. In complex scenarios such as substations, the state and pose changes of devices are usually affected by multiple factors, and traditional change detection methods are difficult to adapt to the changing operation scenarios. In the related art, the object detection network designed based on Convolutional Neural Networks (CNN) can only identify devices within the training data set and cannot detect changes in all devices.

[0053] Based on this, the embodiments of the present disclosure provide a detection method, which includes: obtaining a pair of pictures, the pair of pictures includes a reference picture and a picture to be tested; the pair of pictures is obtained by photographing at least one device to be tested at the same position and different times; using an object detection model to detect the pair of pictures to determine the target area where the device to be tested has changed; the object detection model is obtained by training an initial detection model with at least one set of sample picture pairs; the change of the device to be tested includes: the pose of the device to be tested changes, and / or, the state of the device to be tested changes; the initial detection model is a Transformer model. In this way, the reference picture and the picture to be tested obtained by photographing at least one device to be tested are input into the object detection model to obtain the target area where the pose and / or state of the device to be tested has changed, so that the advantages of the Transformer model can be fully utilized to achieve efficient detection of the pose and / or state changes of the device to be tested; at the same time, the visual features of the device pose and / or state can be automatically learned in complex scenarios to improve the accuracy and adaptability of detection; and it can resist the influence of environmental changes such as light and weather, improve the stability and accuracy of the detection of the pose and / or state of the device to be tested; in addition, the change information of the device to be tested can be output in a timely and accurate manner, reducing the operation and maintenance costs, and improving the safety and reliability of the power system.

[0054] The following will describe each embodiment of the present disclosure in detail with reference to the drawings.

[0055] In one embodiment of the present disclosure, refer to Figure 1 which shows a schematic flow chart of a detection method provided by an embodiment of the present disclosure. Figure 1 As Figure 1 shown, the detection method may include:

[0056] S101: Obtain a pair of pictures, the pair of pictures includes a reference picture and a picture to be tested; the pair of pictures is obtained by photographing at least one device to be tested at the same position and different times.

[0057] It should be noted that the detection method provided by the embodiments of the present disclosure may specifically be a method for detecting the pose and state changes of equipment in a substation scenario based on deep learning, which can be used in, but not limited to, the detection of the pose and / or state of equipment in a substation scenario or other similar industrial fields. This method can be implemented by a detection device or an electronic device integrated with this detection device. The electronic device can be, for example: a server, a smart phone, a computer host, etc., and no specific limitation is made thereto.

[0058] When performing the detection of the pose and / or state of equipment, first, a pair of pictures including a reference picture and a picture to be measured are obtained. Here, a pair of pictures can be photos taken by the same on-site inspection camera at different times, and this inspection camera can be a fixed-pose monitoring camera deployed in a substation. After obtaining the pair of pictures, it is sent to the detection device in a certain way, or can also be sent to the detection device by another electronic device, and no specific limitation is made thereto.

[0059] It should also be noted that the equipment to be measured can be secondary equipment or fire-fighting equipment, or other types of equipment, and no specific limitation is made thereto.

[0060] In some embodiments, in the pair of pictures, the reference picture and the picture to be measured are separated by a preset number of frames, and the reference picture is a historical picture and the picture to be measured is the current picture.

[0061] It should be noted that for a certain camera, at least one piece of equipment to be measured is detected at intervals of a preset number of frames m (m is an integer greater than 0). Every m frames, two frames of pictures taken by the camera are saved. These two frames of pictures are separated by m frames and represent the historical picture and the current picture respectively. The historical picture is used as the reference picture, and the current picture is used as the picture to be measured. It can be understood that for two adjacent frames of pictures, the current picture is detected with the historical picture as the reference.

[0062] It should also be noted that the preset number of frames m can be manually set according to the sensitivity required by the deployment scenario. For example, it can be detected once or several times per minute, and no specific limitation is made thereto.

[0063] In some embodiments, after obtaining the pair of pictures, the pair of pictures can be preprocessed.

[0064] It should be noted that the preprocessing can include CV (Computer Vision) model preprocessing methods such as resizing and normalization.

[0065] In the embodiments of the present disclosure, the pose and / or state of each piece of equipment to be measured in the pair of pictures may change, and it is necessary to use an object detection model to detect the pair of pictures.

[0066] S102: Detect the picture pair using the target detection model to determine the target area where the device under test has changed; the target detection model is obtained by training an initial detection model using at least one set of sample picture pairs; the change in the device under test includes: a change in the posture of the device under test, and / or, a change in the state of the device under test; the initial detection model is a Transformer model.

[0067] It should be noted that, based on the basic structure of the Transformer model, the model structure design is carried out in the embodiments of the present disclosure to obtain the target detection model. The Transformer model is a deep learning model based on the self-attention mechanism, mainly used for natural language processing and other sequence-to-sequence tasks. The core innovation of the Transformer model lies in its self-attention mechanism, which enables the model to consider all positions in the input sequence simultaneously and perform parallel calculations.

[0068] It should also be noted that the posture of the device under test may include the position and angle of the device. The state of the device under test may include a normal state or a fault state, and may also include transitions between different normal states, such as the pointer of a pointer meter pointing to different areas of the dial, the opening and closing of a box door, the color change of an indicator light, etc. During the operation of the device, different postures and states will be presented, and these postures and states reflect the operation situation and performance of the device.

[0069] It should also be noted that the target detection model uses the reference picture as a reference and outputs the coordinates of the part on the picture to be tested that is different from the reference picture, that is, the coordinates of the target area where the change occurs on the device. The target area can be represented in the form of a segmentation mask or in the form of a target box.

[0070] Among them, the segmentation mask is a technique in computer vision used to precisely separate objects in an image from the background. The segmentation mask refers to a pixel-level mask used in image processing to identify each target object in the image. It can precisely represent the contour and boundary of each target and is usually used in instance segmentation tasks to ensure that each target is independently identified and segmented.

[0071] The target box is a rectangular box or polygon box used to describe the position and size of the target object in the image. The rectangular box is represented by the coordinates (x1, y1) of the upper left corner and the coordinates (x2, y2) of the lower right corner of the rectangle, or by the center coordinates (xc, yc) and the width and height (w, h) of the rectangle. The polygon box is used for target objects with irregular shapes and defines the boundary through multiple points.

[0072] In some embodiments, the detection method may further include:

[0073] Determine whether the target area meets the preset conditions;

[0074] If the target area meets the preset conditions, issue an alarm;

[0075] If the target area does not meet the preset conditions, do not issue an alarm.

[0076] In the embodiments of the present disclosure, after training the initial detection model, the trained target detection model is deployed and used in scenarios such as substations to achieve change detection and real-time alarm for some key devices, such as meters, fire-fighting equipment, etc.

[0077] It should be noted that for the target area output by the target detection model, the alarm rule (i.e., the preset condition) can be set manually. The change detection system can select whether to alarm according to whether the alarm rule is met. If an alarm is issued, the maintenance engineer will conduct a manual inspection and eliminate the alarm on the system.

[0078] In some embodiments, the preset condition is that the area of the target area is greater than a preset threshold, and / or, a target to-be-detected device among at least one to-be-detected device has changed.

[0079] Here, an alarm is issued when the area of the target area is greater than the preset threshold, and / or, an alarm is issued when a change occurs within a certain device area (i.e., the target to-be-detected device). That is to say, an alarm is issued when the target to-be-detected device of concern has changed, or when the change amplitude of the target area is relatively large.

[0080] It should also be noted that the preset threshold can be set according to actual needs, and no specific limitation is made thereto.

[0081] See Figure 2 , which shows the flowchart of a detection method provided by the embodiments of the present disclosure Figure 2 . As Figure 2 shown, the target detection model is deployed to an edge device or a server-side inference device. This device can access multiple patrol cameras. For a certain camera, detection is performed at a certain frame interval m, and the nth frame (n is an integer greater than 0) picture and the (n + m)th frame picture are captured. This pair of pictures (i.e., the picture pair) is input into the Transformer change detection model (i.e., the target detection model). This model will output a target area, which can be represented by a segmentation mask or by a target box. For the output target area, an alarm rule determination is performed. If the alarm rule is met, an alarm is issued.

[0082] Embodiments of the present disclosure provide a detection method that fully utilizes the advantages of the Transformer model to achieve efficient detection of the posture and / or state changes of secondary equipment, fire-fighting equipment, and other equipment. This method can automatically learn the visual features of the posture and / or state of the equipment in complex scenarios to improve the accuracy and adaptability of detection, while overcoming some limitations in traditional methods. Additionally, when the posture and / or state of the equipment changes, this method can output the changed position (i.e., the target area), providing strong support for the safe operation of the power substation.

[0083] Furthermore, embodiments of the present disclosure also provide a model training method for training an initial detection model to obtain a target detection model using at least one set of sample picture pairs. As Figure 3 shown, it specifically includes the following steps:

[0084] S201: Obtain a sample data set; the sample data set includes at least one set of sample picture pairs, and each sample picture pair includes a reference sample picture and a sample picture to be tested.

[0085] It should be noted that the number of sample picture pairs usually needs to be rich enough to achieve a good training effect on the initial detection model. The sample data set needs to include multiple sets of picture pairs taken of various postures and / or states of multiple different types of equipment to be tested.

[0086] It also should be noted that to train a visual Transformer model for change detection (i.e., the initial detection model), corresponding training samples (i.e., the sample data set) need to be prepared. The sample data set includes at least one set of sample picture pairs. In the substation scenario, a large number of pictures of substation equipment are collected using multiple cameras, including pictures of equipment in various states and / or postures such as normal operation state and fault state. Multiple pictures taken by the camera at the same position and posture are used as a set of training and validation data, and any two pictures within the set form a set of sample picture pairs, serving as the reference sample picture and the sample picture to be tested respectively. Multiple sets of sample picture pairs are randomly generated from the collected data as the sample data set to train the initial detection model. Here, using a set of training and validation data taken by the same and fixed camera can correctly identify changes in the posture and / or state of the equipment when the camera has not changed.

[0087] In some embodiments, after obtaining the sample data set, the detection method may further include:

[0088] Performing data augmentation on the sample data set to increase the data volume of the sample data set.

[0089] It should be noted that enhancing the sample data set can include specific methods such as random cropping, rotation, and brightness adjustment to increase data diversity. Additionally, the intensity of data augmentation should be controlled at a relatively low level to ensure that while improving the generalization of the detection model, it does not interfere with the learning of change detection.

[0090] S202: Label at least one sample device in each pair of sample images to obtain labeling information; the labeling information is used to indicate that a change has occurred to the sample device.

[0091] It should be noted that manually label the positions where the posture and / or state of the sample device in each pair of sample images change to form labeled data, ensuring the accuracy and comprehensiveness of the labeling. Additionally, object box labeling or segmentation mask labeling can be used, and no specific limitation is made here.

[0092] It should also be noted that pairs of sample images where the state and posture of the sample device have not changed are used as negative samples and do not need to be labeled.

[0093] In the embodiments of the present disclosure, steps S201 and S202 correspond to the steps of preparing the data set before model training. Taking the substation scenario as an example, collect image data in the substation scenario and perform detailed labeling to ensure that each pair of sample images accurately marks the positions where the posture and / or state of the sample device change.

[0094] S203: Train an initial detection model using at least one pair of sample images to obtain a target detection model.

[0095] It should be noted that before inputting the pair of sample images into the initial detection model, the pair of sample images can be preprocessed.

[0096] It should also be noted that input each pair of sample images into the initial detection model, and use the initial detection model to detect the pair of sample images to determine the changed area where the sample device has changed. The labeled sample data set can be used to perform supervised training on the initial detection model to optimize the network weights and improve the model performance.

[0097] In some embodiments, refer to Figure 4 , which shows a schematic structural diagram of an initial detection model provided by the embodiments of the present disclosure. As Figure 4 shown, the initial detection model can include a first feature extraction module (transformer block 1), a second feature extraction module (transformer block 2), a difference module, an upsampling module, and a decoder; both the first feature extraction module and the second feature extraction module are Transformer blocks;

[0098] The first feature extraction module is used to extract features from the reference sample image in the sample image pair to obtain first feature information;

[0099] The second feature extraction module is used to extract features from the sample image to be tested in the sample image pair to obtain second feature information;

[0100] The difference module is used to perform difference processing on the first feature information and the second feature information to obtain difference information;

[0101] The upsampling module is used to upsample the difference information to obtain upsampled information;

[0102] The decoding module is used to decode the upsampled information to obtain the changed area where the sample device has changed.

[0103] It should be noted that both the first feature extraction module and the second feature extraction module are transformer blocks in the Transformer model, mainly composed of input encoding, positional encoding, multi-head attention mechanism, feed-forward neural network, layer normalization, residual connection, etc. The upsampling module can specifically be a Multilayer Perceptron (MLP) module.

[0104] It also should be noted that the first feature information can be image features, and the second feature information can also be image features. After obtaining the two image features, perform image difference on the two image features. Specifically, subtract the corresponding pixel values of the two images to weaken the similar parts of the images and highlight the changed parts of the images.

[0105] In the embodiments of the present disclosure, a set of sample image pairs is input into the initial detection model. For these two input images (the reference sample image and the sample image to be tested), use transformer block 1 and transformer block2 to extract features respectively, and use the difference module to perform difference on features of different sizes to extract the feature differences between the two images. Then, use MLP to upsample the difference features, and use the decoder head classifier to parse them into target boxes or segmentation masks.

[0106] In some embodiments, the changed area can be represented by a target box or a segmentation mask;

[0107] If the changed area is represented by a target box, then use the first type of loss function during the process of training the initial detection model;

[0108] If the change region is represented by a segmentation mask, a second type of loss function is used during the process of training the initial detection model;

[0109] Among them, the first type of loss function is different from the second type of loss function.

[0110] It should be noted that the loss function is a function that evaluates the degree of difference between the predicted output of the model and the true label. It quantifies the degree of prediction error of the model and serves as the optimization objective during training. The model continuously adjusts its internal parameters to minimize the value of the loss function, thereby achieving better data fitting and generalization ability.

[0111] It should also be noted that the first type of loss function can include the cross-entropy loss function (CrossEntropy Loss, CE loss) for classification and the L2 loss for coordinate regression. Among them, CE loss can measure the degree of difference between two different probability distributions in the same random variable. When the two probability distributions are closer, the cross-entropy loss is smaller, indicating that the model prediction structure is more accurate. The mean square error (Mean Square Error, MSE), also known as L2 loss, refers to the average of the squares of the differences between the model prediction values and the sample true values.

[0112] The second type of loss function can include the binary cross-entropy loss function (Binary Cross Entropy Loss, BCE loss) and the Dice coefficient loss function (Dice Coefficient Loss, simply referred to as Dice Loss). Among them, BCE loss, also known as the binary cross-entropy loss function, is usually used as the loss function in binary classification problems to evaluate the difference between the output of the binary classification model and the actual label, and helps the model learn the correct classification. The formula of BCE loss is shown in formula (1):

[0113]

[0114] Among them, N represents the number of samples; y i represents the actual label of the i-th sample, taking values of 0 or 1; represents the predicted value of the i-th sample, with the value range between (0,1).

[0115] Dice Loss is a loss function widely used in image segmentation tasks, especially suitable for pixel-level binary classification or multi-classification tasks. The Dice coefficient is an index to measure the similarity between two sets, and its value ranges from 0 to 1. The closer the value is to 1, the higher the degree of overlap, and vice versa, the lower the degree of overlap.

[0116] The formula for the Dice coefficient is shown in Formula (2), and the formula for the Dice Loss is shown in Formula (3):

[0117]

[0118] where N represents the number of samples; y i represents the actual label of the i-th sample; represents the predicted value of the i-th sample.

[0119] It should be noted that the network weights can be optimized by adjusting model hyperparameters such as the learning rate, batch size, etc., or strategies such as early stopping. Among them, model hyperparameters refer to the parameters manually set by developers during the machine learning or deep learning process, and these parameters will not be automatically adjusted by the training data. Early stopping means that when the performance of the model on the training set is getting better and better, but the performance on the validation set starts to deteriorate, the early stopping mechanism will force the training to stop. This can avoid overfitting of the model on the training set, so as to obtain better generalization performance on unseen data sets.

[0120] In some embodiments, after obtaining the changed area where the sample device has changed, post-processing is performed on the changed area.

[0121] It should be noted that the post-processing may include performing morphological operations on the results output by the model, such as algorithms for image processing such as denoising, opening operation, and closing operation.

[0122] As Figure 5 shown, to train the initial detection model, first obtain a sample data set, perform data augmentation on the sample data set, and preprocess each pair of sample images (reference sample image and sample image to be tested) in it; then input the pair of sample images after data augmentation and preprocessing into the initial detection model to output the changed area where the sample device has changed; then perform post-processing on the output changed area, and then perform loss calculation and backpropagation to calculate the gradient, and update the weights of the model.

[0123] It should be noted that through the backpropagation algorithm, the gradient of the loss function with respect to the model parameters is calculated. An optimization algorithm (such as gradient descent or its variants) is used to update the parameters of the model to minimize the loss function. Repeat forward propagation, loss calculation, backpropagation, and parameter update until the loss function reaches a predetermined number of training iterations, stop the iteration, and obtain the trained target detection model.

[0124] In this way, after training the initial detection model through the above steps S201 to S203, a target detection model that can accurately detect the posture and / or state of the device to be tested is obtained.

[0125] In the embodiments of the present disclosure, real-time considerations can also be given to the trained object detection model, that is, the object detection model is optimized, including model quantization, pruning, etc., to adapt to embedded devices or edge computing environments. A lightweight Transformer structure can be used, or hardware acceleration, such as a Graphics Processing Unit (GPU), a Tensor Processing Unit (TPU), etc., to improve the real-time performance of the object detection model.

[0126] Furthermore, security and privacy protection can also be carried out. Specifically, consider the privacy protection measures of the object detection model in on-site deployment to ensure the secure processing of the image data collected by the camera, such as encryption.

[0127] It should also be noted that during use, model update and maintenance can also be carried out. Specifically, a regular update strategy for the model can be set to adapt to scenario changes and improve detection performance. Exemplarily, an online learning mechanism can be introduced so that the model can continuously optimize itself according to the data during actual operation.

[0128] It should also be noted that the trained object detection model is subjected to on-site testing and optimization. It is tested in actual scenarios such as substations to verify the performance of the model in the actual environment, and according to the results of on-site testing, it may be necessary to fine-tune the model parameters or improve the algorithm.

[0129] The embodiments of the present disclosure provide a detection method. Based on the Vision Transformer model, this detection method can better capture the key features of the posture and / or state of the device to be measured in the input picture and has a stronger representation learning ability. With the generalization of the deep learning model, this technology can resist the influence of environmental changes such as light and weather, and improve the stability and reliability of the detection of the posture and / or state of the device to be measured. Due to the use of the Transformer model, this technology has strong generalization when dealing with different types of devices and scenarios, and is not only applicable to the change detection of specific devices, but also applicable to the change detection of other types of substation devices. This technology can output the position where the device changes (i.e., the target area) in real time, providing timely and accurate change information about the posture and / or state of the device for the operation and maintenance personnel, which can reduce the operation and maintenance costs and improve the security and reliability of the power system.

[0130] In another embodiment of the present disclosure, refer to Figure 6 , which shows a schematic structural diagram of a detection device provided by the embodiments of the present disclosure. As Figure 6 shown, the detection device 30 includes:

[0131] An acquisition unit 301, configured to acquire a pair of images, where the pair of images includes a reference image and a to-be-tested image; the pair of images is obtained by photographing at least one to-be-tested device at the same position and different times;

[0132] A detection unit 302, configured to detect the pair of images by using a target detection model to determine a target area where the to-be-tested device has changed; the target detection model is obtained by training an initial detection model by using at least one set of sample image pairs; the change of the to-be-tested device includes: the posture of the to-be-tested device changes, and / or, the state of the to-be-tested device changes; the initial detection model is a Transformer model.

[0133] In some embodiments, as Figure 6 shown, the detection device 30 may further include an alarm unit 303, configured to determine whether the target area meets a preset condition; and if the target area meets the preset condition, give an alarm; and if the target area does not meet the preset condition, do not give an alarm.

[0134] In some embodiments, the preset condition is: the area of the target area is greater than a preset threshold, and / or, a target to-be-tested device among at least one to-be-tested device has changed.

[0135] In some embodiments, as Figure 6 shown, the detection device 30 may further include a model training unit 304, configured to obtain a sample data set; the sample data set includes at least one set of sample image pairs, and the sample image pair includes a reference sample image and a to-be-tested sample image; and label at least one sample device in each sample image pair to obtain label information; the label information is used to indicate that the sample device has changed; and use at least one set of sample image pairs to train the initial detection model to obtain the target detection model.

[0136] In some embodiments, the initial detection model includes a first feature extraction module, a second feature extraction module, a difference module, an upsampling module, and a decoding module; both the first feature extraction module and the second feature extraction module are Transformer blocks; the first feature extraction module is configured to extract feature information from the reference sample image in the sample image pair to obtain first feature information; the second feature extraction module is configured to extract feature information from the to-be-tested sample image in the sample image pair to obtain second feature information; the difference module is configured to perform difference processing on the first feature information and the second feature information to obtain difference information; the upsampling module is configured to perform upsampling on the difference information to obtain upsampling information; the decoding module is configured to decode the upsampling information to obtain a changed area where the sample device has changed.

[0137] In some embodiments, the changed region is represented by a target bounding box or a segmentation mask; if the changed region is represented by a target bounding box, a first type of loss function is used in the process of training the initial detection model; if the changed region is represented by a segmentation mask, a second type of loss function is used in the process of training the initial detection model; wherein, the first type of loss function and the second type of loss function are different.

[0138] In some embodiments, as Figure 6 shown, the detection device 30 may further include a data augmentation unit 305, configured to perform data augmentation on the sample data set after obtaining the sample data set, so as to increase the data volume of the sample data set.

[0139] In some embodiments, in a pair of pictures, the reference picture and the picture to be measured are separated by a preset number of frames, and the reference picture is a historical picture and the picture to be measured is the current picture.

[0140] It should be noted that the detection device 30 provided in the embodiments of the present disclosure is used to implement the detection method in the foregoing embodiments. For the details not disclosed in this embodiment, please refer to the description of the foregoing embodiments for understanding, and will not be elaborated here.

[0141] It can be understood that, in this embodiment, a "unit" may be a part of a circuit, a part of a processor, a part of a program or software, etc. Of course, it may also be a module or non-modular. Moreover, the components in this embodiment may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software function module.

[0142] If the integrated unit is implemented in the form of a software function module and is not sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment essentially or the part that contributes to the prior art or all or part of this technical solution may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method described in this embodiment. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

[0143] Therefore, the embodiments of the present disclosure provide a computer storage medium, which stores a computer program. When the computer program is executed by at least one processor, the steps of the detection method described in any one of the foregoing embodiments are implemented.

[0144] The embodiments of the present disclosure further provide a computer program product, which includes a computer program. When the computer program is executed by at least one processor, the steps of the detection method described in any one of the foregoing embodiments are implemented.

[0145] Based on the above computer storage medium and computer program product, refer to Figure 7 , which shows a schematic structural diagram of an electronic device provided by the embodiments of the present disclosure. Figure 1 As Figure 7 shown, the electronic device 40 may include: a communication interface 401, a memory 402, and a processor 403; each component is coupled together through a bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, in Figure 7 all kinds of buses are labeled as the bus system 404. Among them, the communication interface 401 is used for receiving and sending signals during the process of receiving and sending information with other external network elements;

[0146] The memory 402 is used to store a computer program that can run on the processor 403;

[0147] The processor 403 is used to execute the detection method described in any one of the foregoing when running the computer program.

[0148] It will be appreciated that the memory 402 in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct rambus random access memory (DRRAM). The memory 402 of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0149] The processor 403 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 403 or the instructions in the form of software. The above-mentioned processor 403 may be a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present disclosure can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 402, and the processor 403 reads the information in the memory 402 and combines its hardware to complete the steps of the above method.

[0150] It can be understood that these embodiments described herein can be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit can be implemented in one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, other electronic units for performing the functions described in the present disclosure, or a combination thereof.

[0151] For software implementation, the technologies described herein can be implemented by modules (such as procedures, functions, etc.) that execute the functions described herein. The software code can be stored in a memory and executed by a processor. The memory can be implemented inside or outside the processor.

[0152] In a further embodiment of the present disclosure, refer to Figure 8, which shows a schematic structure of an electronic device provided by an embodiment of the present disclosure Figure 2 . As Figure 8 shown, the electronic device 40 includes the detection device 30 described in any one of the foregoing embodiments.

[0153] In the embodiment of the present disclosure, for the electronic device 40, it is possible to efficiently detect the attitude and / or state change of the device to be measured, improve the stability, accuracy and adaptability of the detection; in addition, the operation and maintenance cost can be reduced, and the safety and reliability of the power system can be improved.

[0154] The above is only a preferred embodiment of the present disclosure, and is not used to limit the protection scope of the present disclosure.

[0155] It should be noted that in the present disclosure, the terms "include", "comprise" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the presence of another identical element in the process, method, article or device including the element.

[0156] The serial numbers of the above embodiments of the present disclosure are only for description and do not represent the advantages and disadvantages of the embodiments.

[0157] The methods disclosed in several method embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments.

[0158] The features disclosed in several product embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new product embodiments.

[0159] The features disclosed in several method or device embodiments provided by the present disclosure can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.

[0160] The above is only the specific implementation manner of the present disclosure, but the protection scope of the present disclosure is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present disclosure can easily think of changes or substitutions, which should be covered by the protection scope of the present disclosure.

Claims

1. A detection method, characterized in that, The detection method includes: Obtaining a pair of images, where the pair of images includes a reference image and a to-be-detected image; the pair of images is obtained by photographing at least one to-be-detected device at the same position and different times; Using a target detection model to detect the pair of images to determine the target area where the to-be-detected device has changed; the target detection model is obtained by training an initial detection model using at least one set of sample image pairs; the change of the to-be-detected device includes: the posture of the to-be-detected device changes, and / or, the state of the to-be-detected device changes; the initial detection model is a Transformer model.

2. The method according to claim 1, wherein The method further includes: Judging whether the target area meets a preset condition; If the target area meets the preset condition, an alarm is issued; If the target area does not meet the preset condition, no alarm is issued.

3. The method according to claim 2, wherein The preset condition is: the area of the target area is greater than a preset threshold, and / or, the target to-be-detected device among the at least one to-be-detected device has changed.

4. The method according to claim 1, wherein The method further includes: Obtaining a sample data set; the sample data set includes the at least one set of sample image pairs, and the sample image pair includes a reference sample image and a to-be-detected sample image; Labeling at least one sample device in each sample image pair to obtain labeling information; the labeling information is used to indicate that the sample device has changed; Using the at least one set of sample image pairs to train the initial detection model to obtain the target detection model.

5. The method according to claim 4, wherein The initial detection model includes a first feature extraction module, a second feature extraction module, a difference module, an upsampling module, and a decoding module; both the first feature extraction module and the second feature extraction module are Transformer blocks; The first feature extraction module is used to extract features from the reference sample image in the sample image pair to obtain first feature information; The second feature extraction module is used to extract features from the to-be-detected sample image in the sample image pair to obtain second feature information; The difference module is used to perform difference processing on the first feature information and the second feature information to obtain difference information; The upsampling module is used to perform upsampling on the difference information to obtain upsampling information; The decoding module is used to decode the upsampling information to obtain the changed area where the sample device has changed.

6. The method according to claim 5, characterized in that, The changed area is represented by a target box or a segmentation mask; If the changed area is represented by the target box, a first type of loss function is used during the process of training the initial detection model; If the changed area is represented by the segmentation mask, a second type of loss function is used during the process of training the initial detection model; Wherein, the first type of loss function and the second type of loss function are different.

7. The method according to claim 4, characterized in that, After obtaining the sample data set, the method further includes: Performing data augmentation on the sample data set to increase the data volume of the sample data set.

8. The method according to any one of claims 1 to 7, wherein In the pair of pictures, the reference picture and the picture to be measured are separated by a preset number of frames, and the reference picture is a historical picture while the picture to be measured is the current picture.

9. A detection device, characterized in that, The detection device includes: An acquisition unit, configured to acquire a pair of pictures, where the pair of pictures includes a reference picture and a picture to be measured; the pair of pictures is obtained by taking pictures of at least one device to be measured at the same position and different times; A detection unit, configured to detect the pair of pictures by using a target detection model to determine a target area where the device to be measured has changed; the target detection model is obtained by training an initial detection model by using at least one set of sample picture pairs; the change of the device to be measured includes: the posture of the device to be measured changes, and / or, the state of the device to be measured changes; the initial detection model is a Transformer model.

10. An electronic device, characterized in that, The electronic device includes a memory and a processor; The memory is configured to store a computer program that can run on the processor; The processor is configured to execute the steps of the method according to any one of claims 1 to 8 when running the computer program.

11. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and when the computer program is executed by at least one processor, the steps of the method according to any one of claims 1 to 8 are implemented.