A vision-based wiring detection method and system for current transformer error testing

By analyzing the wiring process of current transformers using visual inspection methods and deep learning models, the problem of traditional manual inspection being unable to capture dynamic contact timing was solved, achieving high-precision judgment of wiring anomalies and reliability of test data.

CN122493230APending Publication Date: 2026-07-31HUBEI ELECTRIC POWER CO JINGZHOU POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HUBEI ELECTRIC POWER CO JINGZHOU POWER SUPPLY CO
Filing Date
2026-05-12
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Traditional manual visual inspection methods cannot accurately capture the dynamic contact sequence and instantaneous connection stability during the wiring process of current transformers, resulting in wiring tests failing to fully reflect the compliance of the wiring process and affecting the reliability of error tests.

Method used

A vision-based wiring detection method is adopted. The wiring operation process is captured and recorded by a target detection model and a multi-target tracking model to generate wiring operation event records. Feature vectors are extracted using a 3D-CNN convolutional model, and potential wiring risks are judged by a long short-term memory autoencoder model.

Benefits of technology

It enables continuous and stable tracking of multiple key targets in the wiring operation, accurately constructs interactive timing relationships, improves the accuracy of wiring anomaly judgment, and ensures the reliability of the wiring process and the authenticity of test data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122493230A_ABST
    Figure CN122493230A_ABST
Patent Text Reader

Abstract

This invention provides a vision-based wiring detection method and system for current transformer error testing. Based on target detection and multi-target tracking algorithms, it achieves continuous and stable tracking of multiple key targets (such as wires, tools, and terminals) across video frames during wiring operations, effectively handling complex scenarios such as brief obstructions and target deformation. This method accurately constructs the interactive timing relationship between "operator's hand - tool - wire - terminal," automatically generating structured key event sequences such as "contact start" and "stable connection," providing a reliable and high-precision timing benchmark for subsequent in-depth behavioral analysis and improving the accuracy of subsequent wiring anomaly detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of current transformer technology, and more specifically, to a wiring detection method and system for vision-based current transformer error testing. Background Technology

[0002] As critical metering and protection devices in power systems, the accuracy of instrument transformer error test results directly affects the fairness of power metering and the safe and stable operation of the power system. Wiring is a fundamental prerequisite for instrument transformer error testing; the correctness of the wiring, the reliability of the connection, and the compliance with operational procedures directly impact the authenticity of the test data. Therefore, effectively detecting the wiring status of instrument transformer error tests is a crucial step in ensuring the smooth conduct of the tests.

[0003] Traditional methods for testing instrument transformer wiring often involve manual visual inspection. After completing the wiring operation, the operator visually inspects the terminal connections of the instrument transformer under test, the standard instrument transformer, the measuring instrument, and the load device according to the preset test wiring diagram. This includes checking whether the color and specifications of the wires are consistent with the requirements, and whether the connection between the wires and the terminals is secure and whether the wiring position meets the specifications.

[0004] However, manual visual inspection can only judge the final static state after the wiring is completed. It cannot accurately capture the dynamic contact sequence of the wires and terminals, the instantaneous connection stability, and the continuity of the operation during the wiring process. Therefore, it is difficult to identify hidden problems such as poor instantaneous contact and incorrect wiring sequence that may exist during the wiring process. Ultimately, the wiring test cannot fully reflect the compliance of the entire wiring process, which affects the reliability of the error test. Summary of the Invention

[0005] This invention addresses the technical problems existing in the prior art by providing a vision-based wiring detection method and system for current transformer error testing, which can overcome the problems in the background art.

[0006] According to a first aspect of the present invention, a wiring detection method for a vision-based current transformer error test is provided, comprising: Step S1: Record the current transformer wiring process to obtain the original current transformer wiring video. Step S2: Input the original current transformer wiring video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target. Generate wiring operation event records based on the target tracking trajectory. Extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record. Step S3: Extract features from the keyframe video using a 3D-CNN convolutional model to obtain the transformer wiring feature vector; Step S4: Input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether the current transformer wiring abnormality has occurred based on the wiring potential risk probability.

[0007] According to a second aspect of the present invention, a wiring detection system for testing current transformer errors based on vision is provided, comprising: The acquisition module is used to capture and record the current transformer wiring operation process to obtain the original current transformer wiring video. The target detection module is used to input the original current transformer wiring video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target, generate wiring operation event records based on the target tracking trajectory, and extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record. The feature extraction module is used to extract features from the key frame video of the current transformer wiring using a 3D-CNN convolutional model to obtain the current transformer wiring feature vector. The determination module is used to input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether a current transformer wiring abnormality has occurred based on the wiring potential risk probability.

[0008] This invention provides a vision-based wiring detection method and system for current transformer error testing. Based on target detection and multi-target tracking algorithms, it achieves continuous and stable tracking of multiple key targets (such as wires, tools, and terminals) across video frames during wiring operations, effectively handling complex scenarios such as brief obstructions and target deformation. This method accurately constructs the interactive timing relationship between "operator's hand - tool - wire - terminal," automatically generating structured key event sequences such as "contact start" and "stable connection," providing a reliable and high-precision timing benchmark for subsequent in-depth behavioral analysis and improving the accuracy of subsequent wiring anomaly detection. Attached Figure Description

[0009] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0010] Figure 1 A flowchart of a vision-based wiring detection method for current transformer error testing is provided as an embodiment of the present invention; Figure 2The flowchart illustrates the steps of obtaining the wiring feature vector of a current transformer in a vision-based wiring detection method for current transformer error testing, as proposed in one embodiment of the present invention. Figure 3 This is a framework diagram of an improved long short-term memory network autoencoder model for vision-based mutual inductor error testing, proposed in one embodiment of the present invention. Figure 4 This is a structural block diagram of a vision-based wiring detection system for testing instrument transformer errors, as proposed in one embodiment of the present invention. Detailed Implementation

[0011] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. In addition, the technical features of the various embodiments or individual embodiments provided by the present invention can be arbitrarily combined with each other to form feasible technical solutions. Such combinations are not constrained by the order of steps and / or structural composition patterns, but must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.

[0012] Figure 1 The diagram illustrates a wiring detection method for a vision-based current transformer error test according to an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: Step S1: Record the current transformer wiring process to obtain the original current transformer wiring video.

[0013] In the preparation stage of the transformer error test, the wiring items are first prepared. The models and specifications of the transformer under test, standard transformer, measuring instruments (such as error calibration devices), and load devices are verified. It is confirmed that the primary and secondary side terminals are clearly marked and structurally intact, without any physical defects that would affect the reliability of the electrical connection. Simultaneously, copper core wires with a cross-sectional area matching the circuit current, corresponding terminal blocks, and tools such as screwdrivers and wire strippers are prepared. An industrial camera is set up in the test area and its parameters are configured. Subsequently, under the full camera recording, the wiring operation of the entire test circuit is performed. The operation includes cleaning the surfaces of the terminals of the transformer under test, standard transformer, measuring instruments, and load devices with sandpaper or a special cleaning agent to remove the oxide layer. Wire strippers are used to strip the insulation layer from both ends of the wires to a length matching the terminal wiring depth. Then, according to the test wiring diagram, the wires are correctly connected to the corresponding terminals: including the series connection of the primary circuit of the transformer under test and the standard transformer, the connection of the voltage and current signal lines from the secondary circuit to the measuring instruments, and the connection of the load device in the secondary circuit. Use a screwdriver to tighten all terminal screws to the appropriate torque (refer to the equipment manual or relevant technical specifications) to ensure that the wires are securely fixed without excessive damage.

[0014] The entire wiring process of the aforementioned test circuit was continuously recorded by a color industrial camera fixed on a stable bracket. The camera was positioned to provide a complete, unobstructed view of all key equipment, including the current transformer under test, the standard current transformer, measuring instruments, load devices, and their interconnected secondary terminal blocks. An appropriate aperture (e.g., F4.0-F8.0 to ensure sufficient depth of field) was set according to the ambient lighting conditions, and the focal length was adjusted to ensure that terminal markings, wire connections, and tool operation details of all key components within the operating area were clearly visible within the depth of field. The camera was configured to capture video sequences at a resolution of 1920×1080 pixels at a rate of 25 frames per second. Recording began when the operator first entered the operating area or touched the tools / wires and continued until all terminal screws were tightened and preliminarily checked. The result was a complete video recording of the entire wiring process of the current transformer error test circuit, from power input to measurement signal output, and the final wiring status. This video clearly shows all key electrical connection points, serving as raw input data for subsequent visual analysis algorithms to determine the correctness and reliability of the wiring. The current transformer wiring video is parsed into a current transformer wiring video sequence, which is { , ,..., ,..., }, where each image frame The size is 1920×1080 pixels. The t-th frame of the video of the current transformer wiring is represented, where T represents the total number of video frames.

[0015] Step S2: Input the original current transformer wiring video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target. Generate wiring operation event records based on the target tracking trajectories. Extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record.

[0016] Understandably, in step S2, the current transformer wiring video sequence obtained in step S1 is first processed. , ,..., ,..., Preprocessing is performed to adapt to the input requirements of the target detection model. Specifically, 1920×1080 pixel images are read frame by frame from the original current transformer wiring video sequence. First, the image is scaled to the required input size of 640×640 pixels for the object detection model. Then, the pixel values ​​are normalized, and the preprocessed image frame sequence will be used as input to the improved YOLOv8 model.

[0017] The improved YOLOv8 model adopts an Anchor-Free detection paradigm, and its overall architecture consists of three parts connected sequentially: a backbone network, a neck network, and a detection head.

[0018] The backbone network is responsible for extracting multi-level features from the input image. Its processing begins with a 6×6 convolutional layer (i.e., a focusing layer) with a stride of 2. This layer performs preliminary downsampling and channel expansion on the 640×640×3 (640×640 is the height and width of the image, and 3 is the number of RGB channels) image frames from the current transformer wiring video, outputting a 320×320×64 feature map. Subsequently, the feature map goes through four core processing stages (Stage 1 to Stage 4). Specifically, Stage 1 contains one C2f module with 64 bottleneck channels, outputting a feature map with a spatial size of 160×160 and 128 channels; Stage 2 contains two C2f modules with 128 bottleneck channels, outputting a feature map with a spatial size of 80×80×256; Stage 3 contains two C2f modules with 256 bottleneck channels and an SKAttention attention mechanism module. After the first C2f module, an SKAttention attention mechanism module is embedded, which takes the feature map (40×40×512) output by the first C2f module as input. Stage 3 outputs an enhanced feature map of the same size and passes it to the second C2f module in this stage. The second C2f module outputs a feature map with a spatial size of 40×40×512. Stage 4 contains a bottleneck layer C2f module with 512 channels. The input of this C2f module is the 40×40×512 feature map output from Stage 3. In the bottleneck layer of this C2f module, the original standard 3×3 convolutional layer is replaced with a deformable convolutional (DCN) module, so that the convolutional kernel can adaptively adjust the sampling position to better handle the deformable targets in the feature map of this stage. Finally, the output feature map with a spatial size of 20×20×512 is sent to the neck network as the final feature of the backbone network.

[0019] The neck network includes a top-down path and a bottom-up path. The top-down path involves upsampling the feature map (20×20×512) output from the backbone network Stage4 by a factor of 2, then concatenating it along the channel dimension with the output feature map (40×40×512) from Stage3. The concatenated feature map is then processed by a C2f module (512 channels) to obtain the P4 feature map (40×40×512). This process is repeated: the P4 feature map is upsampled and concatenated with the output feature map (80×80×256) from Stage2, then processed by a C2f module (256 channels) to obtain the P3 feature map (80×80×256); the P3 feature map is then upsampled and concatenated with the output feature map (160×160×128) from Stage1, then processed by a C2f module (128 channels) to obtain the P2 feature map (160×160×128). Bottom-up approach: The P2 feature map is downsampled to an 80×80 spatial size through a convolutional layer with a stride of 2, concatenated with the P3 feature map, and then fused using the C2f module (256 channels) to obtain the fused P3' feature map. The fused result is downsampled again to 40×40, concatenated with the P4 feature map, and then fused using the C2f module (512 channels) to obtain the fused P4' feature map. Simultaneously, the original output feature map (20×20×512) of the backbone network Stage4 is designated as the P5 feature map.

[0020] Finally, the neck network outputs feature maps at four scales: P2, P3', P4', and P5, which are used as input for the subsequent detection head.

[0021] The detection head employs a decoupled head structure, independently handling classification and regression tasks. This structure is symmetrically applied to the four scale feature maps output by the neck network. For each input feature map (e.g., P2: 160×160×128), each detection head processes it through two parallel branches: the classification branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, where the number of output channels of the last 1×1 convolutional layer equals the total number of target categories (specific target types include "operator's right hand", "operator's left hand", "red wire", "blue wire", "screwdriver tip", "standard transformer terminal", "transformer terminal under test", etc.); the regression branch consists of two 3×3 convolutional layers and one 1×1 convolutional layer, with the last 1×1 convolutional layer outputting 4 channels, corresponding to the center point coordinates, width, and height offset of the bounding box.

[0022] The image frame sequence of the current transformer wiring video is input into the improved YOLOv8 model. For the t-th frame, the detection result of the t-th frame is obtained. , ={ , ,},in This represents the attribute information of the detection box of the i-th target detected in the t-th frame of the image, including the normalized center coordinates of the detection box and the width and height of the detection box. Let be the category label of the i-th target detected in the t-th frame of the image. Let be the confidence score of the i-th target detected in the t-th frame of the image. Set the confidence threshold to 0.6 to filter out low-quality detection boxes with confidence scores lower than the confidence threshold.

[0023] For each detection result generated by YOLOv8 in frame t According to its bounding box The corresponding image region is cropped from the original video frame and then scaled (normalized) to a target image patch of 256×128×3. This image patch corresponds to the category label. The target image patch is identified as an object (e.g., a red wire, a screwdriver tip). Subsequently, a pre-trained deep learning model (such as a publicly available ReID network) is used to perform forward propagation on the image patch, extracting a fixed-dimensional feature descriptor vector. This feature descriptor encapsulates the appearance information of the target.

[0024] YOLOv8 detection results The input is fed into the standard DeepSort multi-target tracking model, which assigns and maintains a unique and persistent identifier (ID) for each detected target (such as "red wire" or "screwdriver") in the image frame sequence of the current transformer wiring video, thereby achieving stable tracking across frames.

[0025] The DeepSort multi-target tracking algorithm is a cyclical process of "box-trajectory" matching. First, the algorithm creates and maintains a data structure called the target tracking trajectory for each target that needs to be tracked (including trajectory identifier ID, lifecycle state, motion state, and appearance history).

[0026] It should be noted that each tracking trajectory is identified by a unique ID and continuously records its motion state and appearance history. Motion State: A state vector maintained by the Kalman filter, typically containing the target's center point coordinates in the image, the width and height of the target detection box, and the rate of change of the target's position. By combining the current motion state, the Kalman filter can effectively predict the target's bounding box (target position and size) in the next frame. Appearance History: A fixed-capacity queue (e.g., storing feature descriptors from the most recent 100 frames) stores the feature descriptors (color and texture vectors of target image patches) corresponding to successful past matches of the tracking trajectory. This historical record is used to calculate appearance similarity, enhancing the robustness of the matching. Lifecycle State: A state variable (e.g., "Unconfirmed," "Confirmed," "Lost") manages the trajectory's lifespan (e.g., a new trajectory needs to match for multiple consecutive frames before becoming "Confirmed" to avoid noise interference; "Lost" trajectories are deleted after several consecutive unmatched frames).

[0027] Subsequently, in each frame, the algorithm performs the following matching step: calculates the Mahalanobis distance between all YOLOv8 object detection boxes and the "predicted boxes" of all existing tracking trajectories in the current frame. This measures the consistency of their motion paths. Appearance matching: Calculates the minimum cosine distance between the feature descriptor of each object detection box and the feature descriptors of the appearance history across all tracking trajectories. To measure the similarity in their appearance.

[0028] distance of movement Distance from appearance The comprehensive cost of association is calculated by combining the weighted summations:

[0029] Where C represents the overall associated cost. This is the weight for the movement distance, with a default value of 0.6. To mobilize military strength, For visual distance.

[0030] For a successfully matched detection box and trajectory pair: update the motion state and appearance history of the corresponding trajectory with the detection box. If the trajectory was originally in an "unconfirmed" state, its state will change to "confirmed" after it has successfully matched for a preset number of frames (e.g., 3 frames). For an unmatched detection box, initialize it as a new tracking trajectory and set its lifecycle state to "unconfirmed". For an unmatched tracking trajectory, mark its state as "temporarily lost". If a "temporarily lost" trajectory fails to match successfully within a preset number of frames (e.g., 30 frames), delete the tracking trajectory.

[0031] After acquiring a stable tracking trajectory containing a unique ID and location sequence, event logs for wiring operations are generated based on this trajectory data. When the intersection ratio (CRR) between the bounding box of a conductor ID and the bounding box of a terminal ID is consistently greater than 0.7 for more than 10 frames, it is recorded as a "Stable Connection" event; when the CRR changes from greater than 0.7 to less than 0.3, it is recorded as a "Disconnection" event. When the CRR between the bounding box of a conductor ID and the bounding box of a terminal ID first changes from less than 0.3 to greater than 0.7, the event type is recorded as "Contact Begins". After the "Stable Connection" state, if the CRR changes from greater than 0.7 to less than 0.3, the event type for the frame in which the change occurs is recorded as "Disconnection Begins". After the "Disconnection Begins" event, if the aforementioned CRR remains less than 0.3 for more than 10 frames, it is recorded as "Disconnection Completed".

[0032] Each event record data ,in This represents the start time of the k-th event. This represents the end time of the k-th event. The subject of the operation for the k-th event (e.g., wire ID). The object of operation for the k-th event (e.g., terminal ID). The event type for the k-th event (including four types: "Contact Start", "Stable Connection", "Disconnection Start", and "Disconnection Complete"). All events are stored in chronological order as a JSON file to form a complete wiring operation timing reference.

[0033] Set the time buffer parameter δ for video cropping. δ is usually 10 frames, which is used to preserve the appropriate operational context before and after critical events.

[0034] Event log data Perform the following cropping operation: First, locate the ( )th segment of the original video. -δ) frames, then extract up to the () frame. +δ) frames, generating the corresponding key video segments. When When -δ is less than 1, the capture starts from frame 1; when When +δ exceeds the total number of video frames T, the frame is truncated up to the Tth frame. The final time interval is then [ -δ, The video shows the keyframes of the current transformer wiring with +δ].

[0035] Step S3: Extract features from the key frame video of the transformer wiring using a 3D-CNN convolutional model to obtain the transformer wiring feature vector.

[0036] See Figure 2After acquiring the keyframe video of the current transformer wiring operation generated in step S2, a standard C3D convolutional neural network model is used as the spatiotemporal feature extractor. The 3D convolutional kernel of this model can directly capture the spatiotemporal patterns of dynamic changes such as tool movement, hand movements and component contact during the wiring operation by synchronously sliding the spatial dimension (height, width) and temporal dimension (continuous frame sequence) of the video clip.

[0037] For time intervals of [ -δ, The keyframe video of the current transformer wiring (+δ]) is extracted continuously with a length of [...]. (Default is 16) frames of subsequence to generate fixed-length video clips that meet the input requirements of the C3D network. When capturing data, the core time of the event is usually used (e.g., ...). The segment is truncated around the center. If the original segment length is insufficient... Frames are then extended to a fixed length by methods such as cyclic padding or zero padding.

[0038] video clips Each frame of the image is scaled to a fixed spatial size of 112×112 pixels, and the continuous... Frames are stacked along the channel dimension to form a 5-dimensional video tensor, which serves as an input sample for the network. Its shape is (1, 3, ...). ,112,112), which correspond to batch size, number of RGB channels, number of time frames, height and width, respectively.

[0039] The 3D-CNN convolutional model employs the C3D network. The C3D network uses a sequential structure consisting of alternating convolutional and pooling layers for spatiotemporal feature extraction. Its specific data flow is as follows: The network input is a tensor of size (3, 16, 112, 112), representing 3 RGB channels, 16 frames of temporal sequence, and a spatial size of 112 × 112 pixels. The first layer uses a 3D convolutional operation with a kernel size of (3, 3, 3) and a stride of (1, 1, 1) to perform feature transformation, expanding the number of output channels to 64 while maintaining the feature map size at (64, 16, 112, 112). This is followed by a 3D max-pooling layer with a kernel size of (1, 2, 2) and a stride of (1, 2, 2). This layer downsamples only along the spatial dimension, reducing the feature map size to (64, 16, 56, 56). The middle part of the network continues to process features through several sets of 3D convolutional layers (all with a kernel size of 3×3×3) interleaved with 3D max pooling layers, gradually increasing the number of channels and reducing the spatial size of the feature map. The final convolutional layer outputs a 512-channel feature map with a size of (512, 2, 4, 4). Finally, it passes through a 3D max pooling layer with a kernel size of (2, 2, 2) and a stride of (2, 2, 2), simultaneously compressing both the temporal and spatial dimensions to obtain a final output feature map with a size of (512, 1, 2, 2), completing the hierarchical feature extraction of the entire network.

[0040] The final output feature map has a size of (512, 1, 2, 2). Finally, the output feature map is processed by global average pooling. Converted to 512-dimensional feature vectors , Indicates video clip The characteristic vector of the current transformer wiring.

[0041] Step S4: Input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether a wiring abnormality has occurred based on the wiring potential risk probability.

[0042] Among them, each key video segment The corresponding current transformer wiring feature vector was generated. To construct the input suitable for subsequent sequence models (LSTM autoencoders), these discrete feature vectors need to be organized into a time series. Specifically, each event record in the time series reference file generated in step S2... timestamp (Event start frame number), all current transformer wiring characteristic vectors Arranged in ascending order of time, forming an ordered sequence of feature vectors. , Let F represent the wiring feature vector of the k-th current transformer, where K represents the total number of current transformer wiring feature vectors. This sequence F characterizes the dynamically changing spatiotemporal feature pattern throughout the wiring operation and will serve as the input to the improved LSTM autoencoder model.

[0043] See Figure 3 The improved long short-term memory network autoencoder model adopts an encoder-decoder symmetric architecture and enhances model performance by introducing a cross-layer connection mechanism.

[0044] The encoder receives a feature vector sequence F of dimension (K, 512) and its structure consists of two stacked LSTM layers. The first encoder LSTM has a hidden state dimension of 128, takes the feature vector sequence F as input, and outputs a 128-dimensional hidden state sequence (the first hidden state sequence) containing sequence details, while retaining the 128-dimensional final hidden state at the final time step. The second encoder LSTM has a hidden state dimension of 64, takes the 128-dimensional hidden state sequence from the first layer as input, and outputs a 64-dimensional context vector (retaining only the hidden state at the final time step), serving as a condensed representation of the entire sequence.

[0045] Cross-layer connection module: After linearly projecting the 128-dimensional first hidden state of the first layer encoder LSTM to 64 dimensions, it is added to the initial hidden state of the first layer decoder LSTM initialized by the 64-dimensional context vector, and together they serve as the final initial hidden state of the first layer L decoder STM; at the same time, the final hidden state (context vector) of the second layer encoder LSTM is directly used as the initial hidden state of the second layer decoder LSTM.

[0046] After initialization as described above, the decoder begins operation. The first-layer LSTM decoder has a hidden state dimension of 64. Based on the initial hidden state after cross-layer connections, it processes the "sequence generation input" step-by-step (the first step is a 512-dimensional initial identifier vector with all zeros, and subsequent steps are the reconstruction output of the previous time step), outputting a 64-dimensional hidden state sequence. This sequence is then passed as input to the second-layer LSTM decoder, which has a hidden state dimension of 128, outputting a 128-dimensional feature sequence. Finally, a fully connected layer maps the 128-dimensional features back to 512 dimensions step-by-step, generating a reconstructed sequence that perfectly matches the dimensions of the original input sequence. This completes the entire encoding-reconstruction process. The design effectively improves gradient flow through cross-layer connections, promotes feature reuse, and enhances the accuracy of sequence reconstruction.

[0047] The training of the Long Short-Term Memory Network Autoencoder Model employs an improved joint loss function. Simultaneously, the reconstruction error and feature distribution difference are optimized, as shown in the formula:

[0048] in, Denotes the joint loss function. Represents the reconstruction loss, used to measure the reconstruction sequence. With input sequence The mean square error; KL The divergence term is used to measure the distribution of input features. With reconstructed feature distribution Differences This is the weighting coefficient, usually taken as 0.1.

[0049] It should be noted that the input feature distribution and reconstructed feature distribution These are derived from the feature vector sequence F and the reconstructed sequence, respectively. The probability distribution is obtained by estimating all the feature vectors of the input sequence F, and is calculated by calculating the empirical mean vector of all its feature vectors. Covariance Matrix To parameterize the distribution , where the mean vector The covariance matrix reflects the global average of the features in the sequence. Describes the correlation between the various dimensions of the features; reconstructs the sequence. Feature distribution Calculate its mean vector in the same way. Covariance Matrix Obtained. Through comparison With P The KL divergence can measure the consistency of the distribution of the original sequence and the reconstructed sequence in the high-dimensional feature space, thereby constraining the model to maintain the statistical properties of the features during the reconstruction process.

[0050] Training was performed using the Adam optimizer with a learning rate of 0.001, iterating for 300 epochs until the loss converged. After the Long Short-Term Memory (LSTM) network autoencoder model was trained, for the input sequence... The reconstructed sequence is obtained through forward propagation. And calculate the final joint loss function value. .

[0051] Wiring potential risk probability The reconstruction error is controlled by the Sigmoid function. Mapping to the (0,1) interval yields:

[0052] in, For potential risk probability, ()express Activation function This represents the final value of the joint loss function.

[0053] This probability value monotonically depends on the reconstruction error, when Approaching 0 A value close to 0 indicates a normal sequence; when... When it increases A value close to 1 indicates an increased risk of anomalies. This probability output can be directly used to set a threshold to trigger an early warning, or as an input feature for subsequent classification models.

[0054] The Long Short-Term Memory (LSTM) autoencoder model ultimately outputs the probability of potential wiring risks. This probability serves as the basis for determining whether to perform a current transformer error test. Specifically, when the probability of potential wiring risks exceeds a preset threshold, it indicates a high probability of an abnormal wiring condition. In this case, the current transformer error test will not be triggered, and a wiring inspection and reference to event log data are recommended. The system manually verifies the transformer wiring video; conversely, when the probability is less than or equal to a preset threshold, it indicates that the wiring is normal, and the system will perform the transformer error test normally. This threshold-based binary decision mechanism enables automatic screening of test conditions, ensuring that subsequent tests are only performed when the wiring is reliable.

[0055] See Figure 4 This paper illustrates a vision-based wiring detection system for current transformer error testing according to an embodiment of the present invention. The system includes: The acquisition module 401 is used to capture and record the current transformer wiring operation process to acquire the original current transformer wiring video. The target detection module 402 is used to input the current transformer video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target, generate wiring operation event records based on the target tracking trajectory, and extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record. Feature extraction module 403 is used to extract features from the key frame video of the current transformer wiring using a 3D-CNN convolution model to obtain the current transformer wiring feature vector. The determination module 404 is used to input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether a wiring abnormality has occurred based on the wiring potential risk probability.

[0056] It is understood that the wiring detection system for current transformer error testing based on vision provided by the present invention corresponds to the wiring detection method for current transformer error testing based on vision provided in the foregoing embodiments. The relevant technical features of the wiring detection system for current transformer error testing based on vision can be referred to the relevant technical features of the wiring detection method for current transformer error testing based on vision, and will not be repeated here.

[0057] The wiring detection method and system for current transformer error testing based on vision provided in this invention have the following beneficial effects: (1) The application of the DeepSort multi-target tracking algorithm enables continuous and stable tracking of multiple key targets (such as wires, tools, and terminals) across video frames during wiring operations. By fusing the target's motion information (Kalman filter prediction) and appearance information (ReID feature extraction), the algorithm can effectively handle complex scenarios such as brief occlusion and target deformation, assigning a unique and consistent ID to each target. The beneficial effect of this is that it can accurately construct the interaction time sequence relationship between "operator's hand-tool-wire-terminal", automatically generate structured key event sequences such as "contact start" and "stable connection", and provide a reliable and high-precision time sequence benchmark for subsequent deep behavior analysis, fundamentally avoiding logical misjudgments caused by target loss or ID switching.

[0058] (2) The improved YOLOv8 target detection model introduces an attention mechanism into the deep stages of its backbone network, enabling the model to adaptively adjust its receptive field and focus on the image regions most relevant to the wiring task (such as the screwdriver tip held by the hand, or the terminal hole to be connected), significantly improving its feature extraction capability for small, occluded targets. Simultaneously, deformable convolution (DCN) is used instead of standard convolution in the C2f module, allowing the model to better adapt to the non-rigid deformation of targets (such as bent wires, tools at different angles) during the wiring process. The beneficial effects of these improvements are that they significantly improve the accuracy and robustness of target detection in complex and dynamic industrial wiring scenarios, effectively reduce false positives and false negatives, and lay a solid perceptual foundation for high-precision multi-target tracking.

[0059] (3) The improved Long Short-Term Memory (LSTM) autoencoder model, through its encoder-decoder architecture and cross-layer connection design, can effectively learn the deep spatiotemporal patterns of normal wiring operation sequences. The introduced joint loss function (combining reconstruction error and KL divergence of feature distribution) not only focuses on the point-to-point reconstruction accuracy of the sequence, but also constrains the consistency of the overall feature distribution, making the model more sensitive to anomalies. Its beneficial effect is that the model can learn a highly condensed "normal operation" reference benchmark from the dynamic video features of wiring operations, thereby enabling it to accurately quantify the deviation of the current wiring sequence from the normal pattern (i.e., the probability of potential wiring risks) in an unsupervised or semi-supervised manner. This achieves intelligent and quantitative assessment of potential risks in the wiring process, providing a scientific and reliable basis for decision-making on whether to conduct subsequent error experiments.

[0060] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0061] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0062] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0063] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1The function specified in one or more boxes.

[0064] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0065] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.

[0066] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.

Claims

1. A method for detecting connection of a visual-based error test of a mutual inductor, characterized in that, include: Step S1: Record the current transformer wiring process to obtain the original current transformer wiring video. Step S2: Input the original current transformer wiring video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target. Generate wiring operation event records based on the target tracking trajectory. Extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record. Step S3: Extract features from the keyframe video using a 3D-CNN convolutional model to obtain the transformer wiring feature vector; Step S4: Input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether the current transformer wiring abnormality has occurred based on the wiring potential risk probability.

2. The visual-based connection detection method for error testing of a mutual inductor according to claim 1, wherein, Step S1 involves recording the current transformer wiring process to obtain the original current transformer wiring video, including: In the transformer test loop wiring operation, through the shooting recording, a raw transformer wiring video recording the complete transformer error test loop wiring process from the power input to the measurement signal output and the final wiring state is generated, the raw transformer wiring video can clearly present all key electrical connection points, the raw transformer wiring video is analyzed into a transformer wiring video sequence, the transformer wiring video sequence is { , ,..., ,..., } wherein the size of each image frame is 1920x1080 pixels, represents the t-th image frame of the raw transformer wiring video, and T represents the total number of video frames.

3. The wiring detection method for current transformer error testing based on vision according to claim 2, characterized in that, Step S2 involves inputting the original transformer video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory for each target, including: Each image frame in the current transformer wiring video sequence is preprocessed, including size scaling and pixel value normalization, to obtain a preprocessed image frame sequence. Each image frame in the image frame sequence is input into the target detection model to obtain the target detection result. , represents the target detection result of the i-th target in the t-th image frame; Based on the detection results of each target Based on each target bounding box, the corresponding image area is cropped from the original current transformer wiring video, and the image area is scaled to 256×128×3 to form a target image block; Forward propagation is performed on each target image patch based on a pre-trained deep learning model to extract a feature descriptor vector of fixed dimension, which represents the appearance information of the target. All target detection results The data is input into a multi-target tracking model to obtain the target tracking trajectory for each target.

4. The wiring detection method for current transformer error testing based on vision according to claim 3, characterized in that, The target detection model is an improved YOLOv8 model, which includes a backbone network, a neck network, and multiple detection heads. The backbone network is used to extract multi-level features of each image frame; The neck network is used to upsample and downsample multi-level features and fuse them to obtain feature maps at multiple scales. Each detection head independently performs classification and regression tasks on the feature maps at each scale to obtain the target detection results for each image frame. , ={ , , },in This represents the attribute information of the detection box of the i-th target detected in the t-th frame of the image, including the normalized center coordinates of the detection box and the width and height of the detection box. Let be the category label of the i-th target detected in the t-th frame of the image. Let be the confidence score of the i-th target detected in the t-th frame of the image; The multi-target tracking model is the DeepSort model, which is based on the input target detection results. For each target that needs to be tracked, a data structure called the target tracking trajectory is created and maintained. The target tracking trajectory includes trajectory identifier ID, life cycle status, target motion status and target appearance information. In each image frame, the Mahalanobis distance between all target detection boxes and the "predicted boxes" of all existing tracking trajectories in the current image frame is calculated. , as the distance traveled; Calculate the minimum cosine distance between the feature descriptor vector of each object detection box and the feature descriptor vectors of the object's appearance in all tracking trajectories. , as the apparent distance; distance of movement Distance from appearance The comprehensive cost of association is calculated by combining the weighted summations: ; Where C represents the overall associated cost. This is the weight for the movement distance, with a default value of 0.

6. For the distance traveled, For visual distance; Based on the comprehensive association cost, each detected target in the current image frame is matched with the existing tracking trajectory target to generate a stable tracking trajectory containing a unique trajectory identifier ID and a position sequence.

5. The wiring detection method for vision-based current transformer error testing according to claim 1, 3, or 4, characterized in that, In step S2, a wiring operation event record is generated based on the target tracking trajectory. Based on each event record data, a corresponding keyframe video is extracted from the original current transformer wiring video, including: Based on the target tracking trajectory, event logs for wiring operations are generated, with each event log containing data. ,in This represents the start time of the k-th event. This represents the end time of the k-th event. The operator for the k-th event, For the operation object of the k-th event, The event type for the k-th event includes four types: "Contact Start", "Stable Connection", "Disconnection Start", and "Disconnection Complete". All events are pressed The timing sequence is stored as a JSON format file to form a complete timing reference for wiring operations; Set the time buffer parameter δ for video cropping, and record data for each event. Perform the following cropping operations: The video of the original current transformer wiring was located at the ( ) -δ) frames, extracted up to the () +δ) frames, generate the corresponding key video segments, and obtain the time interval as [ -δ, Keyframe video of the current transformer wiring with +δ]; Among them, when When -δ is less than 1, the capture starts from frame 1; when When +δ exceeds the total number of video frames T, the video is truncated up to frame T.

6. The wiring detection method for current transformer error testing based on vision according to claim 1, characterized in that, Step S3 involves extracting features from the keyframe video of the transformer wiring using a 3D-CNN convolutional model to obtain a transformer wiring feature vector, including: For time intervals of [ -δ, The keyframe video of the current transformer wiring (+δ]) is extracted continuously with a length of [...]. Subsequences of frames are used to generate fixed-length video clips that meet the input requirements of the C3D convolutional neural network model. ; video clips Each image frame is scaled to a fixed spatial size, and the continuous Frames are stacked along the channel dimension and combined into a video tensor, which serves as an input sample for the C3D convolutional neural network model. The C3D convolutional neural network model employs a sequential structure consisting of alternating convolutional and pooling layers to extract the spatiotemporal feature vector of each image frame. , Indicates video clip The characteristic vector of the current transformer wiring.

7. The wiring detection method for current transformer error testing based on vision according to claim 6, characterized in that, Step S4 involves inputting the current transformer wiring feature vector into a long short-term memory network autoencoder model and outputting the probability of potential wiring risks, including: Discrete current transformer wiring characteristic vector Organized into an ordered sequence of current transformer wiring characteristic vectors , Let K represent the wiring characteristic vector of the k-th current transformer, where K represents the total number of wiring characteristic vectors of current transformers. Transformer wiring characteristic vector sequence An improved long short-term memory network autoencoder model is input, which includes an encoder part, a cross-layer connection module, a decoder part, and a fully connected layer. The encoder section comprises two stacked LSTM layers. The input to the first encoder LSTM layer is a sequence of transformer wiring feature vectors. The output is a 128-dimensional first hidden state sequence including sequence details; the input of the second layer encoder LSTM is a 128-dimensional first hidden state sequence, and the output is a 64-dimensional context vector that retains only the hidden state at the final time step. The cross-layer connection module projects the 128-dimensional first hidden state sequence output by the first layer encoder LSTM onto 64 dimensions, adds it to the initial hidden state of the first layer decoder LSTM initialized by the 64-dimensional context vector, and uses them together as the final initial hidden state of the first layer decoder LSTM; at the same time, it directly uses the context vector of the second layer encoder LSTM as the initial hidden state of the second layer decoder LSTM. The first-layer decoder LSTM, based on the 64-dimensional final initial hidden state, processes the "sequence generation input" step by step in time and outputs a 64-dimensional second hidden state sequence. This second hidden state sequence is passed as input to the second-layer decoder LSTM and outputs a 128-dimensional feature sequence. Finally, a fully connected layer maps the 128-dimensional feature sequence back to 512 dimensions step by step, generating a current transformer wiring feature vector sequence that corresponds to the original input. Reconstruction sequence with perfect dimension matching ; Based on the characteristic vector sequence of current transformer wiring And the reconstructed sequence, calculate the joint loss function. ; Joint loss function Mapped to the probability of potential wiring risks .

8. The wiring detection method for current transformer error testing based on vision according to claim 7, characterized in that, The feature vector sequence based on the current transformer wiring And the reconstructed sequence, calculate the joint loss function. ,include: ; in, Denotes the joint loss function. Represents the reconstruction loss, used to measure the reconstruction sequence. Wiring characteristic sequence with input transformer The mean square error; KL This is the divergence term, used to measure the distribution of the input transformer wiring characteristics. With reconstructed feature distribution differences These are weighting coefficients; The joint loss function Mapped to the probability of potential wiring risks ,include: ; in, This represents the probability of potential risks associated with wiring. ()express Activation function This represents the final value of the joint loss function.

9. The wiring detection method for current transformer error testing based on vision according to claim 1, characterized in that, Determining whether a wiring abnormality has occurred based on the probability of potential wiring risks includes: When the probability of potential wiring risk is greater than a preset threshold, it indicates that the current wiring status is abnormal and the current transformer error test will not be triggered; conversely, when the probability of potential wiring risk is less than or equal to the preset threshold, it indicates that the current wiring status is normal and the current transformer error test will be performed normally.

10. A wiring detection system for current transformer error testing based on vision, characterized in that, include: The acquisition module is used to capture and record the current transformer wiring operation process to obtain the original current transformer wiring video. The target detection module is used to input the original current transformer wiring video into the target detection model and the multi-target tracking model to obtain the target tracking trajectory of each target, generate wiring operation event records based on the target tracking trajectory, and extract the corresponding key frame video from the original current transformer wiring video based on the data of each event record. The feature extraction module is used to extract features from the key frame video of the current transformer wiring using a 3D-CNN convolutional model to obtain the current transformer wiring feature vector. The determination module is used to input the current transformer wiring feature vector into the long short-term memory network autoencoder model, output the wiring potential risk probability, and determine whether a current transformer wiring abnormality has occurred based on the wiring potential risk probability.