Automatic meter reading identification method for inspection robot vehicle of transformer substation
By using improved DOA optimization algorithm and DDPG reinforcement learning algorithm for path planning, combined with image processing technology, the problems of low path planning efficiency and insufficient recognition accuracy of inspection robots in substations are solved, realizing efficient and safe inspection and automatic meter reading.
Patent Information
- Application Number
- CN202510922780.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-04
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-04
AI Technical Summary
Existing inspection robot systems in substations have problems such as low path planning efficiency, poor safety, and insufficient instrument recognition accuracy. They are difficult to adapt to complex environments and dynamic changes, and are prone to collision accidents and misidentification.
We employ an improved DOA optimization algorithm and DDPG reinforcement learning algorithm for path planning, combine visual and textual feature image processing, use an NLM model for image denoising, a SegVG model for feature extraction, and an MHAFF model for digit recognition.
It improves inspection efficiency and recognition accuracy, enhances system security, and ensures efficient and safe operation of inspection robots in substations.
Smart Images

Figure BDA0005483674130000021 
Figure BDA0005483674130000022 
Figure BDA0005483674130000023
Abstract
Description
Technical Field
[0001] The present invention relates to the field of inspection, and in particular to an automatic meter reading and identification method for a substation inspection robot vehicle. Background Art
[0002] With the transformation of the global energy structure and the continuous advancement of smart grid construction, substations, as key hubs in the power system, are growing in size and complexity. This undoubtedly places higher demands on substation operation and maintenance, especially inspection work. Traditional inspection methods rely primarily on manual labor, which is prone to low efficiency, high labor intensity, and numerous safety hazards.
[0003] To address the shortcomings of traditional inspection methods, inspection robot technology has emerged and demonstrates significant potential in substation inspections. Inspection robots can replace manual labor in repetitive and dangerous inspection tasks, improving efficiency and safety while reducing operational costs. However, existing inspection robot systems still face several challenges, limiting their scope and effectiveness.
[0004] Some inspection robots use rule-based path planning methods, which struggle to adapt to the complexity and dynamic changes of substation environments. This results in inefficient pathfinding and makes them prone to getting stuck in local optimal solutions, failing to find the global optimal path. Furthermore, these methods fail to fully consider safety factors such as distance to obstacles, driving speed, and steering angle, increasing the risk of collisions, compromising inspection safety and potentially causing equipment damage and casualties.
[0005] At the same time, the existing meter digital image recognition algorithm is more sensitive to factors such as lighting changes, obstructions and background interference, resulting in insufficient recognition accuracy and prone to digital misrecognition or unrecognition, which in turn affects the accuracy of meter reading results and may even lead to misjudgments and accidents. Summary of the Invention
[0006] Purpose of the invention: The present invention proposes an efficient, safe and accurate automatic meter reading and identification method for substation inspection robot vehicles.
[0007] Technical solution: The automatic meter reading and identification method for a substation inspection vehicle disclosed in the present invention includes the following steps:
[0008] S1. Construct a three-dimensional grid map of the substation environment;
[0009] S2. Establish a multi-objective function for global path planning of the inspection robot vehicle;
[0010] S3. Use the improved DOA optimization algorithm and the improved DDPG reinforcement learning algorithm to perform global and local path planning to find the optimal path for the inspection robot;
[0011] S4, DDPG algorithm introduces the improved GOA algorithm into the reward and punishment function to optimize the optimization path of the safety time distance model;
[0012] S5, combining visual and text features, using the NLM model for image denoising and the SegVG model for image feature extraction;
[0013] S6. Use the improved MHAFF model to perform multi-level feature extraction and adaptive fusion on the located dial numbers to achieve accurate recognition of the numbers.
[0014] Among them, S2 establishes a multi-objective function F for the global path planning of the inspection robot vehicle total for:
[0015] F total =w1·F length +w2·F collision +w3·F smooth
[0016] Among them, F length represents the total length of the path; F collision Indicates the minimum distance between the path and the obstacle; F smooth represents the steering angle of adjacent path segments;
[0017]
[0018] Among them, p i+1 、p i Represents the three-dimensional coordinates of two different nodes; Nr represents the total number of nodes;
[0019]
[0020] Among them, d obs,i represents the distance from the path to the obstacle; k collision represents the attenuation coefficient;
[0021]
[0022] Where Δθ i Indicates the horizontal steering angle; Δφ i Indicates the vertical climb angle.
[0023] Furthermore, S3 uses an improved DOA algorithm to solve the global optimal path for the multi-objective function, specifically including:
[0024] Randomly generate ND individuals in the search space, each of which represents a feasible path candidate solution.
[0025] X i =X l +rand×(X u*X l ),i=1,2,…,ND
[0026] Among them, X i represents the i-th feasible path candidate solution in the population; ND represents the population size, that is, the number of feasible path candidate solutions; rand represents a random number between [0,1]; X u 、X l They represent the upper and lower bounds of the search space, that is, the feasible range of candidate solutions;
[0027] The DOA algorithm is improved by introducing the Tent chaotic map to generate a uniform initial population:
[0028]
[0029] Among them, ɑ DOA Represents a random number between [0,1];
[0030] The initial population matrix is represented by X as follows:
[0031]
[0032] Among them, x i,j represents the i-th feasible path candidate solution in the j-th dimension; Dim represents the dimension of the problem, that is, the number of decision variables in the optimization problem;
[0033] The DOA optimization algorithm enters the exploration phase and searches for potential optimal feasible path solutions through global search;
[0034] Set the maximum number of iterations in the exploration phase When the current iteration number t is less than When , the algorithm focuses on the exploration phase of global search;
[0035] At the beginning of each iteration, each inspection robot remembers the best path found in the previous iteration in its group and resets its position information to the position information of the best path in the group at the beginning of each iteration. The specific formula is as follows:
[0036]
[0037] in, represents the feasible path candidate solution at the t+1th iteration; represents the current best path solution of the qth group in the tth iteration;
[0038] In each iteration, each inspection robot forgets part of the previously explored path information and self-organizes a new path based on the best path information in the current iteration. The specific formula is as follows:
[0039]
[0040] in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension in the qth group at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; Indicates the maximum number of iterations in the exploration phase; It means exploring the forgotten dimension;
[0041] In each iteration, each inspection robot will randomly obtain a portion of the location information on the forgotten dimension from other inspection robots to help the inspection robots explore new areas and find feasible paths that may have been overlooked. The specific formula is as follows:
[0042]
[0043] in, denotes the mth feasible path candidate solution of the jth dimension in the t+1th and tth iterations respectively; m denotes the randomly selected individual number, m∈[1,N];
[0044] During the development phase, the current iteration number range is: Next, before each iteration, the best candidate path in the previous iteration of the entire population is updated to the latest feasible path candidate solution;
[0045]
[0046] in, represents the i-th feasible path candidate solution at the t+1th iteration; represents the best feasible path candidate solution of the entire population at the tth iteration;
[0047] Continue to update each feasible path candidate solution in the forgotten dimension,
[0048]
[0049] in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; Indicates the forgotten dimension during the development phase;
[0050] The positive greedy selection mechanism is introduced to improve the algorithm. The current best individual position is selected according to the fitness value of the candidate solution, that is, the best feasible path for the inspection robot. The specific formula is as follows:
[0051]
[0052] Among them, F(X i ) represents the fitness value of the i-th feasible path solution in the population; It represents the fitness value of the feasible path solution in the jth dimension at the tth iteration. That is, through the judgment of the fitness value, when the fitness of the i-th feasible path solution selected in the population is smaller than the fitness value of the feasible path solution in the jth dimension at the tth iteration, the feasible path solution in the jth dimension will be replaced as the current better path solution.
[0053]
[0054] Among them, F(X best ) represents the fitness value of the optimal path solution of the inspection robot; represents the fitness value of the optimal path solution of the inspection robot vehicle at the tth iteration;
[0055] The forward greedy selection mechanism is used to determine the optimal solution, and the feasible path with the minimum objective function value is selected as the global optimal path solution of the inspection robot. At the same time, it is determined whether the iterative optimization meets the maximum iteration condition. If the condition is satisfied, the iteration stops; otherwise, the loop continues until the optimal solution is obtained, that is, the optimal path solution for the inspection robot vehicle.
[0056] Furthermore, S3 uses an improved DDPG reinforcement learning algorithm for local path planning, specifically including:
[0057] Randomly initialize the policy network and value network and their corresponding target network, and initialize the experience replay buffer;
[0058] Perform state-space design:
[0059]
[0060] in, Indicates the coordinate difference between the inspection robot's current position and the obstacle; Indicates the coordinate difference between the current position of the inspection robot and the target point; v current Indicates the linear speed of the inspection robot; Indicates the energy consumption of the inspection robot's current action;
[0061]
[0062] Among them, C obs_i (x obs_i ,y obs_i ) represents the current position of the i-th obstacle; C goal (x goal ,y goal ) represents the target point coordinates; C agent (x agent ,y agent ) represents the current coordinates of the inspection robot;
[0063] Perform action space design:
[0064]
[0065] in, Represents the output of the Actor network, i.e., the distance the inspection robot moves in the x and y directions;
[0066] Inspection robot performs action a t , move to a new position in the environment and observe the reward r and the next state S t+1 , and whether it is a terminal state, and the experience generated by this interaction (S t ,a t ,r,S t+1 ) is stored in the experience replay buffer R;
[0067] Design the reward function: by the collision penalty factor r collision , target gravity reward and punishment factor, destination reward factor r destination , energy consumption penalty r energy , speed reward factor r velocity And the safety interval model reward composition, the specific formula is as follows:
[0068] r(S,a)=∑γ·r
[0069] Among them, γ represents the weight, that is, the weight of the reward factor of each part of the reward function; r represents each type of reward function, and the specific reward and punishment factors are as follows:
[0070] Consider the inspection robot collision penalty factor r collision When the inspection robot collides with an obstacle, a penalty will be given. The specific formula is as follows:
[0071]
[0072] Among them, d obs represents the distance from the inspection vehicle to the nearest obstacle; τ represents the attenuation coefficient; Π collisionrepresents the collision indicator function; κ c ,η c represents the weight coefficient, and κ c Much larger than η c ;
[0073] Considering the inspection robot vehicle target gravity reward and penalty factor r gravitation , according to the target point, the gravitational potential field of the inspection robot is rewarded or penalized. Whenever the distance between the inspection robot and the target point shortens, a positive reward is given, otherwise a penalty is given.
[0074]
[0075] Among them, D i represents the Euclidean distance between the inspection robot and the target point at the i-th moment; D i-1 K represents the Euclidean distance between the inspection vehicle and the target point at the i-1th moment; alt represents the weight coefficient; r alt represents the gravitational coefficient;
[0076] Consider the reward factor r of the inspection robot car arriving at the destination destination , a reward is given when the inspection robot reaches the target point. The specific formula is as follows:
[0077]
[0078] Among them, d prev Indicates the distance from the previous moment to the target point; d current Indicates the current distance; d init represents the initial distance; π reached Arrival target indicator; η d 、 represents the weight coefficient;
[0079] Consider the energy consumption penalty r of the inspection robot energy , the specific formula is as follows:
[0080]
[0081] in, Indicates the energy consumption of the current action; t enery Indicates the time that has been consumed; T enery max Indicates the maximum allowed time; λ1 and λ2 indicate weight coefficients;
[0082] Consider the inspection robot speed reward factor r velocity , the value of the reward factor in the unit sampling period is positively correlated with the speed, and the specific formula is as follows:
[0083]
[0084] Among them, v current Indicates the linear speed of the inspection vehicle; v max Maximum permissible speed; v target Indicates the recommended speed under the current state, which can be adjusted dynamically; a v ,b v Represents the weight coefficient.
[0085] Furthermore, the safety time interval model is introduced to improve the reward and punishment function:
[0086] The traditional collision time TTC is expressed as follows:
[0087]
[0088] Among them, D rel Indicates the relative longitudinal distance between the inspection vehicle and the obstacle; v rel Indicates the relative longitudinal speed between the inspection robot and the obstacle;
[0089] Emergency braking safety time interval:
[0090]
[0091] Where d1 represents the braking distance of the inspection vehicle; v ego Indicates the speed of the inspection vehicle; v f represents the speed of the obstacle; t r Indicates the delay time of the emergency obstacle avoidance system; a ego Indicates the absolute value of the inspection vehicle's braking deceleration; d0 indicates the distance between the inspection vehicle and the obstacle in front after the vehicle stops due to the braking obstacle avoidance operation, d0 = 0.1m;
[0092]
[0093] Wherein, T1 represents the braking critical TTC value;
[0094] The time interval conditions that should be met for emergency braking to avoid obstacles are:
[0095]
[0096] Emergency turn safety time interval:
[0097]
[0098] Where d2 represents the emergency turning distance of the inspection robot; t c Indicates the delay time of emergency steering;
[0099]
[0100] Among them, T2 represents the turning critical TTC value;
[0101] The time interval conditions that should be met for emergency braking and steering are:
[0102]
[0103] The process of environmental interaction and network update is repeated, and the movement trajectory of the inspection robot in the local area is continuously adjusted until the maximum number of iterations is reached. The inspection robot can stably find the local optimal path within a certain period of time.
[0104] Furthermore, S4 specifically includes:
[0105] Design objective function minmize f(x TTC ):
[0106]
[0107] Among them, f(x TTC ) represents the objective function value; x TTC Represents the decision variable vector, including the speed of the inspection robot, braking deceleration, etc.; Represents the weight coefficient, which is used to balance the impact of different factors on the objective function;
[0108] The constructed objective function min mize f(x TTC ) and decision variables x TTC Input the improved GOA optimization algorithm to find the best solution that meets the safety time headway conditions;
[0109] Randomly generate an initial population of goats:
[0110]
[0111] in, represents the position vector of the i-th candidate solution individual, that is, a d-dimensional vector in the search space; d represents the dimension, that is, the number of decision variables; i represents the number of candidate solution individuals, i = 1, 2, ..., NR;
[0112] The initial population is generated as follows:
[0113]
[0114] Where UB and LB represent the upper and lower bounds of the search space, respectively; rand(d) represents a d-dimensional random vector with a value range between [0,1]. At the same time, the fitness value of each goat position is calculated, which represents the objective function value;
[0115] The new position update formula is:
[0116]
[0117] in, Indicates the position of the i-th candidate solution individual at the t+1-th iteration; represents the position of the i-th candidate solution individual at iteration t; α GOA represents the exploration coefficient, which controls the intensity of random motion; R Guass represents a random variable drawn from the Gaussian distribution N(0,1);
[0118] Adaptive inertia weight is introduced to optimize the exploration coefficient of the GOA optimization algorithm:
[0119]
[0120] Among them, k GOA Represents a parameter that controls the upper and lower limits of the weight; μ GOA Represents the control search smoothness parameter; Indicates the maximum number of iterations; T GOA Indicates the current iteration number;
[0121] The position update formula is:
[0122]
[0123] in, represents the current optimal solution; β GOA represents the development coefficient;
[0124] The Cauchy-Gauss mutation strategy is introduced for improvement. The specific formula is as follows:
[0125] β GOA =βC GOA [1+η GOA 1·cauchy(0,δ GOA 2 )+η GOA 2·Gauss(0,δ GOA 2 )]
[0126]
[0127] Among them, βC GOA represents the original development coefficient value in the range [0,1]; δ GOA represents the standard deviation; cauchy(0,δ GOA ) and Gauss(0,δ GOA ) represents a random variable based on Cauchy distribution and Gaussian distribution; and Represents the adaptive regulation variable, which is dynamically adjusted according to the population evolution state during the iteration process;
[0128] The position update formula is:
[0129]
[0130] Among them, J GOA represents the jump coefficient; represents a candidate solution individual randomly selected from the population;
[0131] Reset formula:
[0132]
[0133] Stopping criteria
[0134]
[0135] in, represents the fitness of the best solution at the t+1th iteration; represents the fitness of the best solution at the tth iteration;
[0136] If the difference is less than the preset threshold ε' GOA , the algorithm is considered to have converged to a good enough solution and the iteration can be stopped;
[0137]
[0138] Among them, the average value of the square of the difference between the fitness value of all candidate solution individuals in the population and the fitness value of the current best solution is calculated;
[0139] If the average value is less than the preset threshold δ' GOA , it is considered that the fitness values of all candidate solution individuals in the population are close enough to the current optimal solution, and the iteration can be stopped.
[0140] Furthermore, S5 uses the NLM model for image denoising, including:
[0141] Take the substation instrument image u NLM Pixel i in NLM Generate a pixel block of the reference window size as the center, and generate candidate pixel blocks in the search window in a sliding window manner. The weight is determined by the similarity between the non-local pixels of the substation instrument image.
[0142] Calculate pixel i NLM and j NLM The Euclidean distance of :
[0143]
[0144] Among them, a NLMrepresents the standard deviation of the Gaussian kernel; d NLM (i NLM ,j NLM ) represents the function of Euclidean distance, which indicates the attenuation of weight; and Represents pixel i NLM The size of the center is N NLM *N NLM The square neighborhood of pixel j NLM The size of the center is N NLM *N NLM square neighborhood of u(N iNLM ) represents N iNLM Gray value of express Gray value of u NLM represents the instrument image taken by the inspection robot, that is, the image containing noise;
[0145] The weight coefficient is:
[0146]
[0147] And satisfy
[0148]
[0149] Among them, w(i NLM ,j NLM ) is the weight used to evaluate pixel j NLM The degree of influence on pixel i; C(i NLM ) is the normalization constant, i.e., the normalization constant; i NLM ,j NLM ∈X NLM , X NLM is the pixel domain; h NLM represents the filtering parameter, which controls the attenuation degree of the exponential function;
[0150] Introducing a dynamic parameter mapping adjustment strategy to automatically adjust a according to the noise level of the local image block NLM and h NLM The value of , thus more effectively suppressing noise while maintaining image details,
[0151] First, estimate the local noise and calculate the variance of the grayscale value of each reference window:
[0152]
[0153] Among them, σ NLM 2 is the gray value variance; n NLM 2is the total number of pixels in the neighborhood window; N NLM is the reference window; μ NLM is the grayscale mean of the window;
[0154] Map the dynamic parameters according to the gray value variance:
[0155]
[0156] Among them, h NLM ' represents the filtering parameters after dynamic mapping; a NLM ' represents the standard deviation of the Gaussian kernel after dynamic mapping; α NLM , β NLM represents the adjustment coefficient; Indicates the maximum noise variance of the entire image;
[0157] The formula for the weight coefficient is updated as follows:
[0158]
[0159] Use the weight to process the pixel i NLM The surrounding pixel values are weighted averaged to obtain the denoised image U(i NLM ):
[0160]
[0161] Among them, U(i NLM ) is the image after denoising; u(j NLM ) is pixel j NLM The original pixel value.
[0162] Furthermore, S5 uses the SegVG model for image feature extraction:
[0163] For the denoised image U, DETR and BERT are used as the visual and text backbone networks to extract image and text features respectively;
[0164] The denoised instrument image is input into the pre-trained DETR model, and its ResNet and Transformer encoders are used to extract the visual features of the image. The pre-trained ResNet model is used to extract the 2D feature map of the image, and a 1×1 convolutional layer is used to reduce the number of channels of the feature map to obtain the feature map I', which is then flattened into a one-dimensional vector Z v ;
[0165] Vector Z v Add position encoding and input it into the encoder layer of the DETR model for feature extraction to obtain the final visual feature vector
[0166]
[0167] in, is the original visual feature of the i1+1th layer; is the original visual feature of layer i1; It is the i1th layer in the DETR backbone network; i1 is the total number of layers; v1 represents vision;
[0168] Use the pre-trained BERT model to process the text and obtain the text feature vector Z t :
[0169] Segment the input text and add special tags [CLS] and [SEP] at the beginning and end of the text respectively;
[0170] Use the embedding layer of the pre-trained BERT model to convert the text into a language feature vector Z t ;
[0171] The language feature vector Z t Input the encoder layer of the BERT model in sequence for feature extraction to obtain the final text feature vector Z t :
[0172]
[0173] in, is the original text feature of the i1+1th layer; is the original text feature of the i1th layer; t is the representative text; It is the 2i1+1 layer in the DETR backbone network; It is the 2i1th layer in the DETR backbone network;
[0174] The Triple Alignment module is introduced to iteratively update query, text, and visual features through a three-way attention mechanism to eliminate the domain differences between them:
[0175] Z o Initialized by a learnable embedding, including a regression query and multiple segmentation queries;
[0176] A triple multi-head attention mechanism is used to triangulate the query, text, and visual features so that they share the same feature space and alleviate domain differences. The formula is:
[0177]
[0178] Among them, Tri-MHA is the triangle attention mechanism; Z o is the original query feature; Z' o is the updated query feature; Z t' is the updated text feature; is the updated visual feature; o represents the query;
[0179] The updated features are merged back into the original features, and the formula is:
[0180]
[0181] The aligned text feature vector Z t and visual feature vector The concatenation forms a multimodal feature vector, which is then input into the Transformer encoder layer for feature fusion to obtain the encoder output containing text and visual information. The encoder output is then input into the decoder. The decoder uses the bbox2seg scheme to convert the target box annotation into a segmentation mask.
[0182] Furthermore, S6 includes:
[0183] The digital region image block U1 output by SegVG is cropped and input into the MHAFF module to meet the input requirements of CNN and Transformer;
[0184] Use CNN to extract local features of the image Capture local information such as edges and textures of numbers; use Transformer to extract global features of images Capture the positional relationship and overall layout of the numbers in the entire image;
[0185] Among them, X1 is the feature matrix extracted by the Transformer branch; Y is the feature matrix extracted by the CNN branch;
[0186] A multi-head attention mechanism is used to fuse the features extracted by CNN and Transformer to obtain a more comprehensive feature representation;
[0187] Classify the fused features and identify the numbers in the image.
[0188] Furthermore, a multi-head attention mechanism is used to fuse the features extracted by CNN and Transformer. The steps are as follows:
[0189] (1) Generate Q, K, V matrices:
[0190] Q=X1W Q ,K=YW K ,V=X1W V
[0191] Among them, Q is the query matrix, which is used to find relevant information; K is the key matrix, which is used to store information; V is the value matrix, which contains the information to be found WQ 、W K 、W V It is a learnable weight matrix used to map the input sequence to the Query, Key, and Value matrices;
[0192] Calculate the attention score:
[0193]
[0194] Among them, Attention(Q,K,V) is to calculate the attention score, which represents the similarity between Query and Key; is a scaling factor used to prevent gradient explosion; softmax is a normalization function that converts the attention score into a probability distribution;
[0195] (2) Calculating multi-head attention:
[0196] MHA(Q,K,V)=Concat(head1,head2,...,head h )W O
[0197] Among them, MHA (Q, K, V) is a multi-head attention mechanism, which inputs Query, Key and Value into h independent attention heads respectively and outputs the fused features; head1, head2, ..., head h There are h independent attention heads, each responsible for finding different information; W O is a learnable weight matrix used to linearly combine the outputs of h attention heads;
[0198] (3) Generate fused feature vector:
[0199] Z=f(MHA)
[0200] Where Z is the fused feature vector; f(·) is a nonlinear function;
[0201] The fused features are classified to identify the numbers in the image. The specific steps are as follows:
[0202] (1) Use the fully connected layer to map the fused features to the classification results:
[0203] Z'=AZ+b
[0204] Where Z' is the classification result vector, each element represents the probability of belonging to a certain digital category; A is the weight matrix used to map the fused feature vector to the classification result vector; b is the bias vector;
[0205] (2) The output is converted into a probability distribution through the softmax function, thereby realizing the classification and recognition of the numbers in the image:
[0206]
[0207] Where P(cow=k|Z') is the probability that sample i belongs to digital category k; k is the category index; C1 is the total number of categories; and e is the base of the natural logarithm.
[0208] Beneficial effects: Compared with the prior art, the present invention significantly improves inspection efficiency and recognition accuracy while enhancing system security. BRIEF DESCRIPTION OF THE DRAWINGS
[0209] Figure 1 is a flow chart of a method according to an embodiment of the present invention;
[0210] Figure 2 Flowchart of the improved DOA optimization algorithm according to an embodiment of the present invention;
[0211] Figure 3 This is the improved DDPG reinforcement learning algorithm local path optimization flowchart;
[0212] Figure 4 This is a flow chart of the improved GOA optimization algorithm according to an embodiment of the present invention. DETAILED DESCRIPTION
[0213] like Figure 1-4 As shown, this embodiment provides a method for automatic meter reading recognition of a substation inspection robot vehicle. The method is divided into two parts: path planning of the inspection robot vehicle and automatic meter reading image recognition.
[0214] The specific process of path planning includes: the inspection robot uses lidar and visual sensors to build a three-dimensional grid map of the substation; a multi-objective function for the inspection robot's global path is established, taking into account factors such as path length, obstacle avoidance, steering angle and energy consumption, and an improved dream optimization (DOA) algorithm is used for global path planning to find the optimal path from the starting point to the end point; then, the inspection robot uses an improved DDPG reinforcement learning algorithm to optimize the local path based on the global path and real-time environmental information; the DDPG algorithm introduces an improved goat optimization (GOA) algorithm in the reward and punishment function to optimize the safe time-distance model to ensure the safety of the inspection robot's driving, and selects the best action based on the current state to reduce energy consumption and time costs.
[0215] The specific process of automatic meter reading image recognition includes: the substation inspection robot reaches the destination through the optimal path discovered by path planning, uses a camera to capture the meter image, and adopts the improved NLM algorithm to remove image noise; then uses the SegVG model to extract and fuse the multi-scale features of the image to generate the target area segmentation mask and accurately locate the meter reading dial area; finally, the MHAFF model performs multi-level feature extraction and adaptive fusion on the located dial numbers to achieve accurate recognition of the numbers, thereby completing the automatic meter reading task.
[0216] The specific implementation process is as follows:
[0217] (1): Construct a 3D grid map of the substation environment;
[0218] The inspection robot uses its onboard A1M8 LiDAR to acquire real-time point cloud data of the substation environment. Its built-in sensors also detect obstacles in the working environment in real time. Specifically, the A1M8 LiDAR on the inspection robot calculates the time difference between the transmitted and received signals to determine the distance between the laser beam and obstacles. It then performs a 360-degree laser rangefinder scan to acquire point cloud data of the substation's surroundings.
[0219] The inspection robot integrates the data of LiDAR and visual sensors to build a 3D point cloud map, and uses visual information to correct the distortion between LiDAR frames. Then, the map is optimized and the accumulated error is corrected through loop detection to ensure the accuracy of the map. Figure 1 The system then divides the substation scene into local areas, constructs local maps, and fuses them together with high precision to create a complete map. During the inspection process, new data is continuously collected to update the map. Based on the semi-static environment of the substation, the system detects and updates information about changing objects, providing accurate environmental information for inspection robot path planning and automatic meter reading.
[0220] When constructing a detailed three-dimensional grid map of the substation, the substation area to be measured is divided into a×b×c l×l rectangular grids according to specific requirements. Each rectangular grid is marked with row and column numbers and corresponding angle information in the sensor coordinate system.
[0221] Set the positions of Nr data points detected by the lidar to: p i =(x i ,y i ,z i ), i∈[1,Nr], and project it into a×b×c rectangular grid to accurately reflect the substation environment; p i Represents the location information of the i-th data point; x i 、y i 、z iRepresent the x-axis coordinate, y-axis coordinate, and z-axis coordinate of the i-th data point in sequence;
[0222] Calculate the probability P of each small grid being occupied i , the formula is as follows:
[0223]
[0224] The distance Dr from the sensor's coordinate origin to the center of each small grid is calculated as follows:
[0225]
[0226] Where l represents the length of a grid; multiple small grid information stacks form a three-dimensional grid map;
[0227] The grid status is determined as follows:
[0228] ① when When , the grid is determined to be in the occupied state, and p i =1;
[0229] ② when When , the grid is determined to be unoccupied, and p i =0.
[0230] The grid information of each substation is fused to obtain a complete three-dimensional grid map of the substation environment.
[0231] (2): Multi-objective function for global path planning of inspection robot;
[0232] The path planning problem for a patrol robot can be described as finding a path from a starting point to a destination in three-dimensional space. This path should satisfy the following requirements: shortest path length, obstacle avoidance, minimal steering angle variation, and low energy consumption. A multi-objective function is established to optimize the global path of the patrol robot, taking into account factors such as path length, obstacle avoidance, steering angle, and energy consumption. A multi-objective function F is established for the global path planning of the patrol robot. total for:
[0233] F total =w1·F length +w2·F collision +w3·F smooth
[0234] Among them, F length represents the total length of the path; F collision Indicates the minimum distance between the path and the obstacle; F smooth Indicates the steering angle between adjacent path segments.
[0235] The total path length F length Defined as the sum of the Euclidean distances between adjacent nodes, the shortest path should directly reduce travel time and energy consumption. The specific formula is as follows:
[0236]
[0237] Among them, p i+1 、p i Represents the three-dimensional coordinates of two different nodes; Nr represents the total number of nodes.
[0238] An exponential penalty function is introduced to strengthen the penalty for close obstacles, ensuring a safe distance between the inspection robot and obstacles, and calculating the minimum distance between the path and obstacles. The specific formula is as follows:
[0239]
[0240] Among them, d obs,i represents the distance from the path to the obstacle; k collision Represents the attenuation coefficient.
[0241] By calculating the steering angles of adjacent path segments (including horizontal turning angles and vertical climbing angles), reducing the number of steering times and angle changes of the inspection robot can reduce mechanical wear and energy consumption. The smoothness cost is defined as follows:
[0242]
[0243] Where Δθ i Indicates the horizontal steering angle; Δφ i Indicates the vertical climb angle.
[0244] (3): Improved dream optimization (DOA) algorithm to solve the global optimal path;
[0245] In the present invention, an improved dream optimization (DOA) algorithm is used to optimize the multi-objective function F of global path planning. total Solve and find the optimal global path for the inspection robot. Set the population size to ND and the maximum number of iterations to Maximum number of iterations in the exploration phase And set the number of forgotten dimensions k p (exploration phase) and k r (Development phase).
[0246] (301) Initialization phase
[0247] In the initialization phase, ND individuals are first randomly generated in the search space, each of which represents a feasible path candidate solution.
[0248] X i =Xl +rand×(X u -X l ),i=1,2,…,ND
[0249] Among them, X i represents the i-th feasible path candidate solution in the population; ND represents the population size, that is, the number of feasible path candidate solutions; rand represents a random number between [0,1]; X u 、X l They represent the upper and lower bounds of the search space, that is, the feasible range of candidate solutions.
[0250] In this invention, the Tent chaotic map is introduced to generate a uniform initial population to improve the DOA algorithm, so as to better assist the algorithm in subsequent optimization and improve the optimization accuracy of the algorithm. The specific formula is as follows:
[0251]
[0252] Among them, α DOA Represents a random number between [0,1].
[0253] The initial population matrix is represented by X as follows:
[0254]
[0255] Among them, x i,j represents the i-th feasible path candidate solution in the j-th dimension; Dim represents the dimension of the problem, that is, the number of decision variables in the optimization problem.
[0256] (302) Exploration phase
[0257] The DOA optimization algorithm enters the exploration phase and searches for the potential optimal feasible path solution through global search. In order to ensure the scope of global search, the maximum number of iterations in the exploration phase is set. When the current iteration number t is less than When , the algorithm focuses on the exploration phase of global search. Each time the inspection robot has different ability to retain and forget path information during the global feasible path exploration process, the population is divided into 5 groups according to the difference in the memory ability of the inspection robots. Each group has a different number of forgetting dimensions, set as k q ,q=1,2,3,4,5, where q represents the ordinal number of the group. Each iteration is considered as a continuous exploration and optimization of feasible candidate paths. By adjusting the memory capacity of the inspection robot, a balance is achieved between exploring new paths and utilizing existing information.
[0258] (1) Memory strategies
[0259] At the beginning of each iteration, each inspection robot remembers the best path found in the previous iteration in its group and resets its position information to the position information of the best path in the group at the beginning of each iteration. This prevents the inspection robot from repeatedly exploring the same area and helps it find a feasible path more quickly. The specific formula is as follows:
[0260]
[0261] in, represents the feasible path candidate solution at the t+1th iteration; represents the current best path solution of the qth group in the tth iteration.
[0262] (2) Forgetting and supplementing strategies
[0263] In each iteration, each inspection robot forgets some of the previously explored path information and self-organizes a new path based on the best path information in the current iteration. The specific formula is as follows:
[0264]
[0265] in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension in the qth group at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; represents the maximum number of iterations in the exploration phase, Indicates exploring the forgotten dimension.
[0266]
[0267] Among them, randi(a,b) represents an integer randomly selected in the range; k q It represents the number of forgotten dimensions of the qth group in the exploration phase; Dim represents the problem dimension.
[0268] (3) Dream Sharing Strategy
[0269] In each iteration, each inspection robot randomly obtains a portion of the location information in the forgotten dimension from other inspection robots. This helps the inspection robots explore new areas and find feasible paths that may have been overlooked. The specific formula is as follows:
[0270]
[0271] in, They represent the mth feasible path candidate solution of the jth dimension in the t+1th and tth iterations respectively; m represents the randomly selected individual number, m∈[1,N].
[0272] (303) Development stage
[0273] (1) Memory strategies
[0274] During the development phase, the current iteration number range is: Next, before each iteration, the best candidate path in the previous iteration of the entire population is updated to the latest feasible path candidate solution.
[0275]
[0276] in, represents the i-th feasible path candidate solution at the t+1th iteration; It represents the best feasible path candidate solution of the entire population at the tth iteration.
[0277] (2) Forgetting and supplementing strategies
[0278] Continue to update each feasible path candidate solution in the forgotten dimension.
[0279]
[0280] in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; Indicates that the dimension is forgotten during the development phase.
[0281]
[0282] Among them, k r Indicates the number of forgotten dimensions during the development phase.
[0283] (304) Evaluation and Output
[0284] In this paper, a forward greedy selection mechanism is introduced to improve the algorithm. The current optimal individual position, i.e., the optimal feasible path for the inspection robot, is selected based on the fitness value of the candidate solution. The forward greedy selection mechanism is used to update the feasible path solution, improving the efficiency of the global search. The specific formula is as follows:
[0285]
[0286] Among them, F(X i ) represents the fitness value of the i-th feasible path solution in the population; It represents the fitness value of the feasible path solution in the jth dimension at the tth iteration. That is, through the fitness value judgment, when the fitness value of the i-th feasible path solution selected in the population is smaller than the fitness value of the feasible path solution in the jth dimension at the tth iteration, the feasible path solution in the jth dimension will be replaced as the current better path solution.
[0287]
[0288] Among them, F(X best ) represents the fitness value of the optimal path solution of the inspection robot; It represents the fitness value of the optimal path solution of the inspection robot vehicle in the tth iteration.
[0289] The forward greedy selection mechanism is used to determine the optimal solution, and the feasible path with the minimum objective function value is selected as the global optimal path solution for the inspection robot. At the same time, it is determined whether the iterative optimization meets the maximum iteration condition. If the condition is satisfied, the iteration stops; otherwise, the loop continues until the optimal solution is obtained, that is, the optimal path solution for the inspection robot vehicle.
[0290] (4): Improved DDPG reinforcement learning algorithm for local path optimization
[0291] In this paper, an improved DDPG reinforcement learning algorithm is used to further optimize the global path locally. This local path optimization is performed based on real-time local environmental information, further optimizing the inspection robot's path, enabling it to more efficiently avoid obstacles during meter reading tasks, reducing energy consumption and time costs.
[0292] (401) Initialization
[0293] Randomly initialize the policy network (Actor network) and the value network (Critic network) and their corresponding target network. The policy Actor network is initialized according to the input state S t (such as the current position of the inspection robot, the position of the target point, the surrounding obstacle information, etc.), output the speed and steering angle of the inspection robot in the local area and other actions a t ; The value critic network is used to evaluate the value of the current state-action pair.
[0294] Initialize the experience replay buffer and set a fixed-size experience replay buffer R to store the experience generated by the interaction between the inspection robot and the environment, including the current state S t 、Action a t , reward r, next state S t+1And whether it is in the termination state and other information.
[0295] (402) Interacting with the environment
[0296] 1) State space design:
[0297] The inspection vehicle travels in the substation environment in the general direction determined by global path planning. When local path optimization is required, the current local environment state is obtained. The state in the state space in this invention is the relationship between the inspection vehicle and obstacles in the current environment, the linear speed of the inspection vehicle, and the energy consumption of the inspection vehicle's current movement. The specific formula is as follows:
[0298]
[0299] in, Indicates the coordinate difference between the inspection robot's current position and the obstacle; Indicates the coordinate difference between the current position of the inspection robot and the target point; v current Indicates the linear speed of the inspection robot; Indicates the energy consumption of the inspection robot's current action.
[0300]
[0301] Among them, C obs_i (x obs_i ,y obs_i ) represents the current position of the i-th obstacle; C goal (x goal ,y goal ) represents the target point coordinates; C agent (x agent ,y agent ) represents the current coordinates of the inspection robot.
[0302] 2) Action Space Design
[0303] The action space of deep reinforcement learning is to select the next action to be performed based on the current state of the inspection robot. The specific setting formula of the action space is as follows:
[0304]
[0305] in, Represents the output of the Actor network, which is the distance the inspection robot moves in the x and y directions.
[0306] Inspection robot performs action a t , move to a new position in the environment and observe the reward r and the next state S t+1 , and whether it is the termination state. At the same time, the experience generated by this interaction (St ,a t ,r,S t+1 ) is stored in the experience replay buffer R.
[0307] (403) Furthermore, the reward function design for improving the DDPG optimization algorithm is specifically composed as follows:
[0308] In complex scenarios, planned paths may still have problems such as redundant points, low pathfinding efficiency, frequent turns, and excessive pathfinding nodes. In this invention patent, the reward and penalty function considers the following types of incentives: safe distance reward (turning and braking), speed reward, target gravity reward, destination reward, collision penalty, and energy consumption penalty.
[0309] Assume that the reward and penalty function is composed of the collision penalty factor r collision , target gravity reward and punishment factor, destination reward factor r destination , energy consumption penalty r energy , speed reward factor r velocity And the safety interval model reward composition, the specific formula is as follows:
[0310] r(S,a)=∑γ·r
[0311] Among them, γ represents the weight, that is, the weight of the reward factor of each part of the reward function; r represents each type of reward function. The specific reward and punishment factors are as follows:
[0312] (1) Considering the inspection robot collision penalty factor r collision When the inspection robot collides with an obstacle, a penalty will be given. The specific formula is as follows:
[0313]
[0314] Among them, d obs represents the distance from the inspection vehicle to the nearest obstacle; τ represents the attenuation coefficient; Π collision represents the collision indicator function (1 when a collision occurs); κ c ,η c represents the weight coefficient, and κ c Much larger than η c .
[0315] (2) Considering the inspection robot vehicle target gravity reward and punishment factor r gravitation , rewards and penalties are set for the gravitational potential field of the inspection robot car according to the target point, which is related to the distance between the inspection robot car and the target point. Whenever the distance between the inspection robot car and the target point becomes shorter, a positive reward is given, otherwise a penalty is given.
[0316]
[0317] Among them, D i represents the Euclidean distance between the inspection robot and the target point at the i-th moment; D i-1 K represents the Euclidean distance between the inspection vehicle and the target point at the i-1th moment; alt represents the weight coefficient; r alt represents the gravitational coefficient.
[0318] (3) Consider the reward factor r of the inspection robot reaching the destination destination , a reward is given when the inspection robot reaches the target point. The specific formula is as follows:
[0319]
[0320] Among them, d prev Indicates the distance from the previous moment to the target point; d current Indicates the current distance; d init represents the initial distance; reached Arrival target indicator; η d 、 Represents the weight coefficient.
[0321] (4) Considering the energy consumption penalty r of the inspection robot energy , the specific formula is as follows:
[0322]
[0323] in, Indicates the energy consumption of the current action; t enery Indicates the time that has been consumed; T enery max represents the maximum allowed time; λ1 and λ2 represent weight coefficients.
[0324] (5) Consider the inspection robot speed reward factor r velocity , the value of the reward factor in the unit sampling period is positively correlated with the speed, and the specific formula is as follows:
[0325]
[0326] Among them, v current Indicates the linear speed of the inspection vehicle; v max Maximum permissible speed; v target Indicates the recommended speed under the current state, which can be adjusted dynamically; a v ,b v Represents the weight coefficient.
[0327] (6) In the present invention, a safety time interval model is introduced to improve the reward and punishment function, and an improved goat (GOA) optimization algorithm is used to find the optimal solution for the safety time interval condition. The specific implementation process is as follows:
[0328] The specific formula of the traditional collision time TTC is as follows:
[0329]
[0330] Among them, D rel Indicates the relative longitudinal distance between the inspection vehicle and the obstacle; v rel Indicates the relative longitudinal speed between the inspection robot and the obstacle.
[0331] 1) Emergency braking safety time interval:
[0332]
[0333] Where d1 represents the braking distance of the inspection vehicle; v ego Indicates the speed of the inspection vehicle; v f represents the speed of the obstacle; t r Indicates the delay time of the emergency obstacle avoidance system; a ego It represents the absolute value of the braking deceleration of the inspection robot vehicle; d0 represents the distance between the inspection robot vehicle and the obstacle in front after it stops through the braking obstacle avoidance operation d0 = 0.1m.
[0334]
[0335] Wherein, T1 represents the braking critical TTC value.
[0336] The time interval conditions that should be met for emergency braking to avoid obstacles are:
[0337]
[0338] 2) Emergency turn safety time interval:
[0339]
[0340] Where d2 represents the emergency turning distance of the inspection robot; t c Indicates the delay time for emergency steering.
[0341]
[0342] Wherein, T2 represents the turning critical TTC value.
[0343] The time interval conditions that should be met for emergency braking and steering are:
[0344]
[0345] (404) In the application, the safe distance model is used to evaluate the safe distance between the inspection robot and obstacles to ensure the safety of the inspection robot. The improved goat optimization (GOA) algorithm is used to optimize the safe distance model, simulating the exploration, development, and jumping behaviors of goats. This can effectively search the solution space and find the optimal solution that meets the safe distance condition. The specific process is as follows:
[0346] Taking into account the safety time distance model, the emergency braking, emergency steering and the corresponding delay time factors, the objective function min mize f(x TTC ), the specific formula is as follows:
[0347]
[0348] Among them, f(x TTC ) represents the objective function value; x TTC Represents the decision variable vector, including the speed of the inspection robot, braking deceleration, etc.; Represents the weight coefficient, which is used to balance the impact of different factors on the objective function.
[0349] The constructed objective function min mize f(x TTC ) and decision variables x TTC Input the improved GOA optimization algorithm, and continuously adjust the decision variable x by simulating the exploration, development, and jumping behaviors of goats. TTC , so that the objective function value f(x TTC ) as small as possible. The GOA algorithm can effectively search the solution space and find the optimal solution that meets the safety time-distance condition, thereby ensuring the safety of the inspection robot vehicle.
[0350] (1) Population initialization
[0351] An initial population of goats is randomly generated, and each goat is represented as a d-dimensional vector in the search space, that is, a combination of decision variables of a safety time interval model, and its position is randomly generated within the given upper and lower bounds.
[0352]
[0353] in, represents the position vector of the i-th candidate solution individual, that is, a d-dimensional vector in the search space; d represents the dimension, that is, the number of decision variables; i represents the number of candidate solution individuals, i = 1, 2,…, NR.
[0354] The formula for generating the initial population is as follows:
[0355]
[0356] Where UB and LB represent the upper and lower bounds of the search space, respectively; rand(d) represents a d-dimensional random vector with a value in the range [0, 1]. The fitness value of each goat position is also calculated, representing the objective function value.
[0357] (2) Exploration stage
[0358] Simulate the foraging behavior of goats and explore the search space by random movement to find potential optimization areas. Each goat explores the search space by random movement, and the update formula for its new position is:
[0359]
[0360] in, Indicates the position of the i-th candidate solution individual at the t+1-th iteration; represents the position of the i-th candidate solution individual at iteration t; α GOA represents the exploration coefficient, which controls the intensity of random motion; R Guass represents a random variable drawn from the Gaussian distribution N(0,1).
[0361] Exploration coefficient α GOA To control the intensity of random motion and to perform a more detailed search in the global area, the present invention introduces an adaptive inertia weight to optimize the exploration coefficient of the GOA optimization algorithm. The specific formula is as follows:
[0362]
[0363] Among them, k GOA Represents a parameter that controls the upper and lower limits of the weight; μ GOA Represents the control search smoothness parameter; Indicates the maximum number of iterations; T GOA Indicates the current iteration number.
[0364] (3) Development stage
[0365] Simulate the behavior of a goat moving towards the optimal position and improve the solution by moving closer to the current optimal position. The goat gradually moves towards the current optimal solution to refine the quality of the solution. The position update formula is:
[0366]
[0367] in, represents the current optimal solution; β GOA represents the development coefficient.
[0368] Development coefficient β GOAIn order to control the step size of approaching the current optimal position in the development stage, the present invention introduces the Cauchy-Gauss mutation strategy for improvement. The specific formula is as follows:
[0369] β GOA =βC GOA [1+η GOA 1·cauchy(0,δ GOA 2 )+η GOA 2·Gauss(0,δ GOA 2 )]
[0370]
[0371] Among them, βC GOA represents the original development coefficient value in the range [0,1]; δ GOA represents the standard deviation; cauchy(0,δ GOA ) and Gauss(0,δ GOA ) represents a random variable based on Cauchy distribution and Gaussian distribution; and Represents the adaptive regulation variable, which is dynamically adjusted according to the population evolution state during the iteration process.
[0372] (4) Jumping strategy
[0373] Update the goat's position and simulate its behavior of jumping to a new position to escape the local optimal solution. The jumping mechanism helps the goat escape the local optimal solution. The position update formula is:
[0374]
[0375] Among them, J GOA represents the jump coefficient; represents a candidate solution individual randomly selected from the population.
[0376] (5) Parasite avoidance and resolution screening
[0377] Reset the positions of low-quality solutions to simulate goats avoiding areas infected by parasites. For goats with fitness values in the lowest 20% of the population, reset their positions to randomly generated new positions to maintain the diversity and robustness of the population. The reset formula is:
[0378]
[0379] (6) Stopping criteria
[0380] 1)
[0381] in, represents the fitness of the best solution at the t+1th iteration; represents the fitness of the best solution at the tth iteration.
[0382] If the difference is less than the preset threshold ε' GOA , the algorithm is considered to have converged to a good enough solution and the iteration can be stopped.
[0383] 2)
[0384] The average of the squares of the differences between the fitness values of all candidate solution individuals in the population and the fitness value of the current best solution is calculated.
[0385] If the average value is less than the preset threshold δ' GOA , it is considered that the fitness values of all candidate solution individuals in the population are close enough to the current optimal solution, and the iteration can be stopped.
[0386] Through the improved GOA optimization algorithm, the optimized decision variables can be used to calculate the optimal path that meets the safe time-distance conditions, and apply it to the path planning of the inspection robot vehicle to ensure the safety of the inspection robot vehicle during driving.
[0387] (405) Network Update
[0388] In the present invention, the optimal global path solution obtained by the above steps is placed in the experience replay buffer R, and the network update is started. First, a batch of experiences, namely one of the global path solutions, is randomly sampled from the experience replay buffer.
[0389] The DDPG algorithm uses the Actor-Critic framework and combines it with deep networks and deterministic policy gradients for improvement. In its architecture, DDPG uses a dual neural network structure of online network and target network.
[0390] Using the sampling experience, the value network parameters are updated by minimizing the mean square error between the Q value output by the value network and the target Q value. The critic network is updated and the target network is used to estimate the Q value function of the state-action of the inspection robot at the next moment. The Q value function is: Q w' (S i+1 ,π θ' (S i+1 )), use the Actor target network to approximate the next action value π θ' (S t+1 ), thereby obtaining the target value of the Q value function in the current state. The specific formula is as follows:
[0391] y i =r i +γQw' (S i+1 ,π θ' (S i+1 ))
[0392] Among them, r i represents the reward and punishment factor obtained by the inspection robot at time i; γ represents the discount factor, which is used to measure the importance of future rewards; Q w' represents the evaluation function of the current strategy, w' represents the parameter vector of the Q function; S i+1 represents the state at time i+1; π θ' (S i+1 ) indicates that in state S i+1 The probability distribution of actions selected according to the policy π with parameters θ'.
[0393] Use the Critic training network to output the Q-value function Q of the current state-action w (S i ,a i ), a i represents the action taken at time i. The target of the Critic network is defined as y i -Q w (S i ,a i ), the parameters of the Critic network are updated by minimizing the loss value. The loss function of the Critic network during update is expressed as:
[0394]
[0395] Among them, a i =π θ (S i )+ε, where ε represents the exploration noise on the behavior strategy.
[0396] The Actor target network and training network provide the strategy for the next state and the strategy for the current state respectively. Combined with the Q value function of the Critic training network, the policy gradient of the Actor during parameter update can be obtained. The specific formula is as follows:
[0397]
[0398] For the update of the target network parameters w' and θ', the DDPG soft update mechanism slowly updates the weights of the target network:
[0399] w'←ξw+(1-ξ)w'
[0400] θ'←ξθ+(1-ξ)θ'
[0401] Among them, ξ represents the update coefficient; w,θ represents the weights of the current Actor network and Critic network; w',θ' represents the target network weights of the current Actor network and Critic network.
[0402] (406) Iterative Optimization
[0403] The above process of interacting with the environment and updating the network is repeated, continuously adjusting the trajectory of the inspection robot within the local area until the maximum number of iterations is reached, and the inspection robot stably finds the local optimal path within a certain period of time. When the training end condition is met, the algorithm terminates, and the trained policy network and value network are obtained. They guide the inspection robot in optimizing the local path in the actual environment, and ultimately obtain the optimal path optimized based on the global path planning.
[0404] (5) In the present invention, the inspection robot vehicle combines the global path planning obtained by the improved DOA optimization algorithm and the local path optimization obtained by the improved DDPG optimization algorithm. The inspection robot vehicle can efficiently and safely perform automatic meter reading tasks in the substation environment and successfully reach the destination where the meter is located.
[0405] Furthermore, by acquiring instrument images from a substation inspection robot, the NLM model improved by the dynamic parameter mapping adjustment strategy is used to perform noise reduction processing on the collected substation instrument images to remove noise from the original images. The specific process is as follows:
[0406] Due to the complex environment of substations, images are susceptible to interference from various factors, generating noise that reduces image clarity and affects subsequent instrument recognition and data reading. Therefore, this paper uses an improved non-local means (NLM) filtering algorithm to denoise the instrument images captured by the inspection robot through the camera.
[0407] The NLM filtering algorithm uses the large number of repetitive structures contained in the image to perform denoising. It calculates the similarity between non-local pixels to determine the weights, and then restores the image by weighted average. The specific steps are:
[0408] (501) Determine the generation window
[0409] Take the substation instrument image u NLM Pixel i in NLM Generate a pixel block of the reference window size for the center, and generate candidate pixel blocks within the search window in a sliding window manner;
[0410] (502) Calculate weights
[0411] The weights are determined by the similarity between non-local pixels of the substation instrument image. The steps are as follows:
[0412] Calculate pixel iNLM and j NLM The Euclidean distance is:
[0413]
[0414] Among them, a NLM represents the standard deviation of the Gaussian kernel; d NLM (i NLM ,j NLM ) represents the function of Euclidean distance, which indicates the attenuation of weight; and Represents pixel i NLM The size of the center is N NLM *N NLM The square neighborhood of pixel j NLM The size of the center is N NLM *N NLM square neighborhood of u(N iNLM ) represents N iNLM Gray value of express Gray value of u NLM represents the instrument image taken by the inspection robot, that is, the image containing noise;
[0415] Therefore, the formula for the weight coefficient is:
[0416]
[0417] And meet the conditions
[0418]
[0419] Among them, w(i NLM ,j NLM ) is the weight used to evaluate pixel j NLM The degree of influence on pixel i; C(i NLM ) is the normalization constant, i.e., the normalization constant; i NLM ,j NLM ∈X NLM , X NLM is the pixel domain; h NLM represents the filtering parameter, which controls the attenuation degree of the exponential function;
[0420] In addition, considering the Gaussian kernel standard deviation a NLM and filter parameter h NLM The fixed setting of a may not be sufficient to meet the robustness requirements of various noise types. The present invention introduces a dynamic parameter mapping adjustment strategy to automatically adjust a according to the noise level of the local image block. NLM and h NLMThe value of , thereby more effectively suppressing noise while maintaining image details, the formula is:
[0421] First, estimate the local noise and calculate the variance of the grayscale value of each reference window. The formula is:
[0422]
[0423] Among them, σ NLM 2 is the gray value variance; n NLM 2 is the total number of pixels in the neighborhood window; N NLM is the reference window; μ NLM is the grayscale mean of the window;
[0424] Then the dynamic parameters are mapped according to the gray value variance, and the formula is:
[0425]
[0426] Among them, h NLM ' represents the filtering parameters after dynamic mapping; a NLM ' represents the standard deviation of the Gaussian kernel after dynamic mapping; α NLM , β NLM represents the adjustment coefficient; Indicates the maximum noise variance of the entire image;
[0427] Therefore, the formula for the weight coefficient is updated as follows:
[0428]
[0429] (503) Weighted average
[0430] Use the weight to process the pixel i NLM The surrounding pixel values are weighted averaged to obtain the denoised image U(i NLM ), the formula is:
[0431]
[0432] Among them, U(i NLM ) is the image after denoising; u(j NLM ) is pixel j NLM The original pixel value of
[0433] (6) The SegVG model is used to extract and fuse multi-scale features from the denoised image to generate the final target area segmentation mask, thereby achieving accurate positioning of the instrument reading dial.
[0434] During substation inspections, precise positioning of meter dials is a key step in achieving automated meter reading. Accurately locating the dial area provides valid regional information for subsequent digital recognition, avoiding interference from invalid areas and improving the efficiency and accuracy of the entire system.
[0435] Therefore, this paper uses an improved SegVG model to process the denoised image and achieve accurate positioning of the instrument reading dial. The model extracts and fuses multi-scale features to generate the final target region segmentation mask, thereby achieving accurate positioning of the instrument reading dial. The specific steps are as follows:
[0436] (601) Feature extraction stage
[0437] For the denoised image U, DETR and BERT are used as the visual and text backbone networks, respectively, to extract image and text features. Visual feature extraction can capture rich information from the image, helping the system identify the location of the instrument dial. Text features can use the labels, logos, or descriptive text on the instrument to more accurately identify and distinguish different types of instruments, thereby improving the accuracy of dial image positioning. The specific steps are as follows:
[0438] 1) Visual feature extraction
[0439] The denoised instrument image is input into the pre-trained DETR model, and its ResNet and Transformer encoders are used to extract the visual features of the image. The specific steps are as follows:
[0440] Resize the input denoised image U to 640×640, maintaining the original aspect ratio. The shorter sides are zero-padded to make them the same length as the longer sides, ensuring that the image does not lose important information or become distorted due to size issues during subsequent processing;
[0441] The pre-trained ResNet model is used to extract the 2D feature map of the image, and the number of channels of the feature map is reduced using a 1×1 convolution layer to obtain the feature map I'. The feature map I' is then flattened into a one-dimensional vector Z v ;
[0442] Vector Z v Add position encoding and input it into the encoder layer of the DETR model for feature extraction to obtain the final visual feature vector The formula is:
[0443]
[0444] in, is the original visual feature of the i1+1th layer; is the original visual feature of layer i1; It is the i1th layer in the DETR backbone network; i1 is the total number of layers; v1 represents vision;
[0445] 2) Text feature extraction
[0446] Use the pre-trained BERT model to process the text and obtain the text feature vector Z t , the specific steps are:
[0447] The input text is segmented and special tags [CLS] and [SEP] are added at the beginning and end of the text respectively.
[0448] Use the embedding layer of the pre-trained BERT model to convert the text into a language feature vector Z t ;
[0449] The language feature vector Z t Input the encoder layer of the BERT model in sequence for feature extraction to obtain the final text feature vector Z t , the formula is:
[0450]
[0451] in, is the original text feature of the i1+1th layer; is the original text feature of the i1th layer; t is the representative text; It is the 2i1+1 layer in the DETR backbone network; It is the 2i1th layer in the DETR backbone network;
[0452] (602) Feature Alignment Stage
[0453] The Triple Alignment module is introduced to iteratively update query, text, and visual features through a three-way attention mechanism to eliminate the domain differences between them. The specific steps are as follows:
[0454] Z o Initialized by a learnable embedding, including a regression query and multiple segmentation queries;
[0455] The triple multi-head attention mechanism (Tri-MHA) is used to triangulate the query, text, and visual features so that they share the same feature space and alleviate domain differences. The formula is:
[0456]
[0457] Among them, Tri-MHA is the triangle attention mechanism; Z o is the original query feature; Z' o is the updated query feature; Z t' is the updated text feature; is the updated visual feature; o represents the query;
[0458] Finally, the updated features are merged back into the original features, and the formula is:
[0459]
[0460] (603) Feature Fusion
[0461] The aligned text feature vector Z t and visual feature vector The concatenation forms a multimodal feature vector, which is then fed into the Transformer encoder layer for feature fusion to obtain the encoder output containing text and visual information;
[0462] (604)Target Positioning
[0463] The output of the encoder is input to the decoder, which uses the bbox2seg scheme to convert the target box annotation into a segmentation mask. The specific steps are:
[0464] 1) Regression query
[0465] The multimodal features output by the encoder are used to regress the target bounding box using a regression query vector. The regression query vector is fed into the Transformer decoder layer and interacts with the encoder output. An MLP network is then used to decode the regression query vector into a target bounding box to determine the position of the dial in the image.
[0466] 2) Split query
[0467] The target region is segmented using a segmentation query vector with different learnable position encodings. The segmentation query vector is fed into the Transformer decoder layer and interacts with the encoder output. The segmentation query vector is then repeated N times. v The algorithm then concatenates the image and visual feature vectors and inputs them into the MLP network for decoding to obtain the segmentation mask of the target area. By marking pixels within the dial area as foreground (value 1) and pixels outside the dial area as background (value 0), the dial area can be accurately distinguished from the background and other irrelevant areas, thus achieving pixel-level positioning of the dial area.
[0468] 3) Finally output the located dial image U1;
[0469] (7) Using the MHAFF model, the accurate recognition of dial numbers is achieved based on multi-head attention feature fusion;
[0470] Due to the complex environment of substations, the located instrument images often have problems such as uneven lighting, partial occlusion, and background interference. Therefore, this paper adopts an improved Multi-Head Attention Feature Fusion (MHAFF) model to perform multi-level feature extraction and adaptive fusion on the denoised instrument images, thereby improving the recognition accuracy in complex scenes. The specific steps are as follows:
[0471] (701) Data Preprocessing
[0472] The digital region image block U1 output by SegVG is cropped and input into the MHAFF module to meet the input requirements of CNN and Transformer.
[0473] (702) Dual-branch feature extraction:
[0474] Use CNN to extract local features of the image Capture local information such as edges and textures of numbers; use Transformer to extract global features of images Capture the positional relationship and overall layout of the numbers in the entire image.
[0475] Among them, X1 is the feature matrix extracted by the Transformer branch; Y is the feature matrix extracted by the CNN branch;
[0476] (703) Feature Fusion
[0477] A multi-head attention mechanism is used to fuse the features extracted by CNN and Transformer to obtain a more comprehensive feature representation. The specific steps are as follows:
[0478] (1) Generate Q, K, V matrices, the formula is:
[0479] Q=X1W Q ,K=YW K ,V=X1W V
[0480] Among them, Q is the query matrix, which is used to find relevant information; K is the key matrix, which is used to store information; V is the value matrix, which contains the information to be found W Q 、W K 、W V It is a learnable weight matrix used to map the input sequence to the Query, Key, and Value matrices;
[0481] Calculate the attention score, the formula is:
[0482]
[0483] Among them, Attention(Q,K,V) is to calculate the attention score, which represents the similarity between Query and Key; is a scaling factor used to prevent gradient explosion; softmax is a normalization function that converts the attention score into a probability distribution;
[0484] (2) Calculate the multi-head attention, the formula is:
[0485] MHA(Q,K,V)=Concat(head1,head2,...,head h )W O
[0486] Among them, MHA (Q, K, V) is a multi-head attention mechanism, which inputs Query, Key and Value into h independent attention heads respectively and outputs the fused features; head1, head2, ..., head h There are h independent attention heads, each responsible for finding different information; W O is a learnable weight matrix used to linearly combine the outputs of h attention heads;
[0487] (3) Generate the fused feature vector, the formula is:
[0488] Z=f(MHA)
[0489] Where Z is the fused feature vector; f(·) is a nonlinear function;
[0490] (704) Output Classification
[0491] The fused features are classified to identify the numbers in the image. The specific steps are as follows:
[0492] (1) Use the fully connected layer to map the fused features to the classification results. The formula is:
[0493] Z'=AZ+b
[0494] Where Z' is the classification result vector, each element represents the probability of belonging to a certain digital category; A is the weight matrix used to map the fused feature vector to the classification result vector; b is the bias vector;
[0495] (2) The output is converted into a probability distribution through the softmax function, thereby realizing the classification and recognition of the numbers in the image. The formula is:
[0496]
[0497] Where P(cow=k|Z') is the probability that sample i belongs to digital category k; k is the category index; C1 is the total number of categories; and e is the base of the natural logarithm.
Claims
1. A method for automatic meter reading and identification of a substation inspection vehicle, characterized in that: The following steps are involved: S1. Construct a three-dimensional grid map of the substation environment; S2. Establish a multi-objective function for global path planning of the inspection robot vehicle; S3. Use the improved DOA optimization algorithm and the improved DDPG reinforcement learning algorithm to perform global and local path planning to find the optimal path for the inspection robot; S4, DDPG algorithm introduces the improved GOA algorithm into the reward and punishment function to optimize the optimization path of the safety time distance model; S5, combining visual and text features, using the NLM model for image denoising and the SegVG model for image feature extraction; S6. Use the improved MHAFF model to perform multi-level feature extraction and adaptive fusion on the located dial numbers to achieve accurate recognition of the numbers.
2. The automatic meter reading and identification method for a substation inspection vehicle according to claim 1 is characterized in that: In S2, a multi-objective function F is established for the global path planning of the inspection robot vehicle. total for: F total =w1·F length +w2·F collision +w3·F smooth Among them, F length represents the total length of the path; F collision Indicates the minimum distance between the path and the obstacle; F smooth represents the steering angle of adjacent path segments; Among them, p i+1 、p i Represents the three-dimensional coordinates of two different nodes; Nr represents the total number of nodes; Among them, d obs,i represents the distance from the path to the obstacle; k collision represents the attenuation coefficient; Where Δθ i Indicates the horizontal steering angle; Δφ i Indicates the vertical climb angle.
3. The automatic meter reading and identification method for a substation inspection vehicle according to claim 2, characterized in that: S3 uses an improved DOA algorithm to solve the global optimal path for multi-objective functions, specifically including: Randomly generate ND individuals in the search space, each of which represents a feasible path candidate solution. X i =X l +rand×(X u -X l ),i=1,2,…,ND Among them, X i represents the i-th feasible path candidate solution in the population; ND represents the population size, that is, the number of feasible path candidate solutions; rand represents a random number between [0,1]; X u 、X l They represent the upper and lower bounds of the search space, that is, the feasible range of candidate solutions; The DOA algorithm is improved by introducing the Tent chaotic map to generate a uniform initial population: Among them, α DOA Represents a random number between [0,1]; The initial population matrix is represented by X as follows: Among them, x i,j represents the i-th feasible path candidate solution in the j-th dimension; Dim represents the dimension of the problem, that is, the number of decision variables in the optimization problem; The DOA optimization algorithm enters the exploration phase and searches for potential optimal feasible path solutions through global search; Set the maximum number of iterations in the exploration phase When the current iteration number t is less than When , the algorithm focuses on the exploration phase of global search; At the beginning of each iteration, each inspection robot remembers the best path found in the previous iteration in its group and resets its position information to the position information of the best path in the group at the beginning of each iteration. The specific formula is as follows: in, represents the feasible path candidate solution at the t+1th iteration; represents the current best path solution of the qth group in the tth iteration; In each iteration, each inspection robot forgets part of the previously explored path information and self-organizes a new path based on the best path information in the current iteration. The specific formula is as follows: in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension in the qth group at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; Indicates the maximum number of iterations in the exploration phase; It means exploring the forgotten dimension; In each iteration, each inspection robot will randomly obtain a portion of the location information on the forgotten dimension from other inspection robots to help the inspection robots explore new areas and find feasible paths that may have been overlooked. The specific formula is as follows: in, denotes the mth feasible path candidate solution of the jth dimension in the t+1th and tth iterations respectively; m denotes the randomly selected individual number, m∈[1,N]; During the development phase, the current iteration number range is: Next, before each iteration, the best candidate path in the previous iteration of the entire population is updated to the latest feasible path candidate solution; in, represents the i-th feasible path candidate solution at the t+1th iteration; represents the best feasible path candidate solution of the entire population at the tth iteration; Continue to update each feasible path candidate solution in the forgotten dimension, in, represents the i-th feasible path candidate solution of the j-th dimension in the t+1-th iteration; represents the best feasible path solution in the jth dimension at the tth iteration; x l,j 、x u,j Respectively represent the upper and lower bounds of the search space in the jth dimension; Indicates the maximum number of iterations; Indicates the forgotten dimension during the development phase; The positive greedy selection mechanism is introduced to improve the algorithm. The current best individual position is selected according to the fitness value of the candidate solution, that is, the best feasible path for the inspection robot. The specific formula is as follows: Among them, F(X i ) represents the fitness value of the i-th feasible path solution in the population; It represents the fitness value of the feasible path solution in the jth dimension at the tth iteration. That is, through the judgment of the fitness value, when the fitness of the i-th feasible path solution selected in the population is smaller than the fitness value of the feasible path solution in the jth dimension at the tth iteration, the feasible path solution in the jth dimension will be replaced as the current better path solution. Among them, F(X best ) represents the fitness value of the optimal path solution of the inspection robot; represents the fitness value of the optimal path solution of the inspection robot vehicle at the tth iteration; The forward greedy selection mechanism is used to determine the optimal solution, and the feasible path with the minimum objective function value is selected as the global optimal path solution of the inspection robot. At the same time, it is determined whether the iterative optimization meets the maximum iteration condition. If the condition is satisfied, the iteration stops; otherwise, the loop continues until the optimal solution is obtained, that is, the optimal path solution for the inspection robot vehicle.
4. The automatic meter reading and identification method for a substation inspection vehicle according to claim 3 is characterized in that: S3 uses an improved DDPG reinforcement learning algorithm for local path planning, specifically including: Randomly initialize the policy network and value network and their corresponding target network, and initialize the experience replay buffer; Perform state-space design: in, Indicates the coordinate difference between the inspection robot's current position and the obstacle; Indicates the coordinate difference between the current position of the inspection robot and the target point; v current Indicates the linear speed of the inspection robot; Indicates the energy consumption of the inspection robot's current action; Among them, C obs_i (x obs_i ,y obs_i ) represents the current position of the i-th obstacle; C goal (x goal ,y goal ) represents the target point coordinates; C agent (x agent ,y agent ) represents the current coordinates of the inspection robot; Perform action space design: in, Represents the output of the Actor network, i.e., the distance the inspection robot moves in the x and y directions; Inspection robot performs action a t , move to a new position in the environment and observe the reward r and the next state S t +1 , and whether it is a terminal state, and the experience generated by this interaction (S t ,a t ,r,S t+1 ) is stored in the experience replay buffer R; Design the reward function: by the collision penalty factor r collision , target gravity reward and punishment factor, destination reward factor r destination , energy consumption penalty r energy , speed reward factor r velocity And the safety interval model reward composition, the specific formula is as follows: r(S,a)=∑γ·r Among them, γ represents the weight, that is, the weight of the reward factor of each part of the reward function; r represents each type of reward function, and the specific reward and punishment factors are as follows: Consider the inspection robot collision penalty factor r collision When the inspection robot collides with an obstacle, a penalty will be given. The specific formula is as follows: Among them, d obs represents the distance from the inspection vehicle to the nearest obstacle; τ represents the attenuation coefficient; Π collision represents the collision indicator function; κ c ,η c represents the weight coefficient, and κ c Much larger than η c ; Considering the inspection robot vehicle target gravity reward and penalty factor r gravitation , according to the target point, the gravitational potential field of the inspection robot is rewarded or penalized. Whenever the distance between the inspection robot and the target point shortens, a positive reward is given, otherwise a penalty is given. Among them, D i represents the Euclidean distance between the inspection robot and the target point at the i-th moment; D i-1 K represents the Euclidean distance between the inspection vehicle and the target point at the i-1th moment; alt represents the weight coefficient; r alt represents the gravitational coefficient; Consider the reward factor r of the inspection robot car arriving at the destination destination , a reward is given when the inspection robot reaches the target point. The specific formula is as follows: Among them, d prev Indicates the distance from the previous moment to the target point; d current Indicates the current distance; d init represents the initial distance; π reached Arrival target indicator; η d 、 represents the weight coefficient; Consider the energy consumption penalty r of the inspection robot energy , the specific formula is as follows: in, Indicates the energy consumption of the current action; t enery Indicates the time that has been consumed; T enery max Indicates the maximum allowed time; λ1 and λ2 indicate weight coefficients; Consider the inspection robot speed reward factor r velocity , the value of the reward factor in the unit sampling period is positively correlated with the speed, and the specific formula is as follows: Among them, v current Indicates the linear speed of the inspection vehicle; v max Maximum permissible speed; v target Indicates the recommended speed under the current state, which can be adjusted dynamically; a v ,b v Represents the weight coefficient.
5. The automatic meter reading and identification method for a substation inspection vehicle according to claim 4 is characterized in that: The reward and punishment function is improved by introducing a safety time interval model: The traditional collision time TTC is expressed as follows: Among them, D rel Indicates the relative longitudinal distance between the inspection vehicle and the obstacle; v rel Indicates the relative longitudinal speed between the inspection robot and the obstacle; Emergency braking safety time interval: Where d1 represents the braking distance of the inspection vehicle; v ego Indicates the speed of the inspection vehicle; v f represents the speed of the obstacle; t r Indicates the delay time of the emergency obstacle avoidance system; a ego Indicates the absolute value of the inspection vehicle's braking deceleration; d0 indicates the distance between the inspection vehicle and the obstacle in front after the vehicle stops due to the braking obstacle avoidance operation, d0 = 0.1m; Wherein, T1 represents the braking critical TTC value; The time interval conditions that should be met for emergency braking to avoid obstacles are: Emergency turn safety time interval: Where d2 represents the emergency turning distance of the inspection robot; t c Indicates the delay time of emergency steering; Among them, T2 represents the turning critical TTC value; The time interval conditions that should be met for emergency braking and steering are: The process of environmental interaction and network update is repeated, and the movement trajectory of the inspection robot in the local area is continuously adjusted until the maximum number of iterations is reached. The inspection robot can stably find the local optimal path within a certain period of time.
6. The automatic meter reading and identification method for a substation inspection vehicle according to claim 5, characterized in that: S4 specifically includes: Design objective function minmize f(x TTC ): Among them, f(x TTC ) represents the objective function value; x TTC Represents the decision variable vector, including the speed of the inspection robot, braking deceleration, etc.; Represents the weight coefficient, which is used to balance the impact of different factors on the objective function; The constructed objective function minmize f(x TTC ) and decision variables x TTC Input the improved GOA optimization algorithm to find the best solution that meets the safety time headway conditions; Randomly generate an initial population of goats: in, represents the position vector of the i-th candidate solution individual, that is, a d-dimensional vector in the search space; d represents the dimension, that is, the number of decision variables; i represents the number of candidate solution individuals, i = 1, 2, ..., NR; The initial population is generated as follows: Where UB and LB represent the upper and lower bounds of the search space, respectively; rand(d) represents a d-dimensional random vector with a value range between [0,1]. At the same time, the fitness value of each goat position is calculated, which represents the objective function value; The new position update formula is: in, Indicates the position of the i-th candidate solution individual at the t+1-th iteration; represents the position of the i-th candidate solution individual at iteration t; α GOA represents the exploration coefficient, which controls the intensity of random motion; R Guass represents a random variable drawn from the Gaussian distribution N(0,1); Adaptive inertia weight is introduced to optimize the exploration coefficient of the GOA optimization algorithm: Among them, k GOA Represents a parameter that controls the upper and lower limits of the weight; μ GOA Represents the control search smoothness parameter; Indicates the maximum number of iterations; T GOA Indicates the current iteration number; The position update formula is: in, represents the current optimal solution; β GOA represents the development coefficient; The Cauchy-Gauss mutation strategy is introduced for improvement. The specific formula is as follows: b GOA =βC GOA [1+n GOA 1·cauchy(0,δ GOA 2 )+η GOA 2·Gauss(0,δ GOA 2 )] Among them, βC GOA represents the original development coefficient value in the range [0,1]; δ GOA represents the standard deviation; cauchy(0,δ GOA ) and Gauss(0,δ GOA ) represents a random variable based on Cauchy distribution and Gaussian distribution; and Represents the adaptive regulation variable, which is dynamically adjusted according to the population evolution state during the iteration process; The position update formula is: Among them, J GOA represents the jump coefficient; represents a candidate solution individual randomly selected from the population; Reset formula: Stopping criteria in, represents the fitness of the best solution at the t+1th iteration; represents the fitness of the best solution at the tth iteration; If the difference is less than the preset threshold ε' GOA , the algorithm is considered to have converged to a good enough solution and the iteration can be stopped; Among them, the average value of the square of the difference between the fitness value of all candidate solution individuals in the population and the fitness value of the current best solution is calculated; If the average value is less than the preset threshold δ' GOA , it is considered that the fitness values of all candidate solution individuals in the population are close enough to the current optimal solution, and the iteration can be stopped.
7. The automatic meter reading and identification method for a substation inspection vehicle according to claim 6, characterized in that: S5 uses the NLM model for image denoising, including: Take the substation instrument image u NLM Pixel i in NLM Generate a pixel block of the reference window size as the center, and generate candidate pixel blocks in the search window in a sliding window manner. The weight is determined by the similarity between the non-local pixels of the substation instrument image. Calculate pixel i NLM and j NLM The Euclidean distance of : Among them, a NLM represents the standard deviation of the Gaussian kernel; d NLM (i NLM ,j NLM ) represents the function of Euclidean distance, which indicates the attenuation of weight; and Represents pixel i NLM The size of the center is N NLM *N NLM The square neighborhood of pixel j NLM The size of the center is N NLM *N NLM square neighborhood of u(N iNLM ) represents N iNLM Gray value of express Gray value of u NLM represents the instrument image taken by the inspection robot, that is, the image containing noise; The weight coefficient is: And satisfy Among them, w(i NLM ,j NLM ) is the weight used to evaluate pixel j NLM The degree of influence on pixel i; C(i NLM ) is the normalization constant, i.e., the normalization constant; i NLM ,j NLM ∈X NLM , X NLM is the pixel domain; h NLM represents the filtering parameter, which controls the attenuation degree of the exponential function; Introducing a dynamic parameter mapping adjustment strategy to automatically adjust a according to the noise level of the local image block NLM and h NLM The value of , thus more effectively suppressing noise while maintaining image details, First, estimate the local noise and calculate the variance of the grayscale value of each reference window: Among them, σ NLM 2 is the gray value variance; n NLM 2 is the total number of pixels in the neighborhood window; N NLM is the reference window; μ NLM is the grayscale mean of the window; Map the dynamic parameters according to the gray value variance: Among them, h NLM ' represents the filtering parameters after dynamic mapping; a NLM ' represents the standard deviation of the Gaussian kernel after dynamic mapping; α NLM , β NLM represents the adjustment coefficient; Indicates the maximum noise variance of the entire image; The formula for the weight coefficient is updated as follows: Use the weight to process the pixel i NLM The surrounding pixel values are weighted averaged to obtain the denoised image U(i NLM ): Among them, U(i NLM ) is the image after denoising; u(j NLM ) is pixel j NLM The original pixel value.
8. The automatic meter reading and identification method for a substation inspection vehicle according to claim 7, characterized in that: S5 uses the SegVG model for image feature extraction: For the denoised image U, DETR and BERT are used as the visual and text backbone networks to extract image and text features respectively; The denoised instrument image is input into the pre-trained DETR model, and its ResNet and Transformer encoders are used to extract the visual features of the image. The pre-trained ResNet model is used to extract the 2D feature map of the image, and a 1×1 convolutional layer is used to reduce the number of channels of the feature map to obtain the feature map I', which is then flattened into a one-dimensional vector Z v ; Vector Z v Add position encoding and input it into the encoder layer of the DETR model for feature extraction to obtain the final visual feature vector in, is the original visual feature of the i1+1th layer; is the original visual feature of layer i1; It is the i1th layer in the DETR backbone network; i1 is the total number of layers; v1 represents vision; Use the pre-trained BERT model to process the text and obtain the text feature vector Z t : Segment the input text and add special tags [CLS] and [SEP] at the beginning and end of the text respectively; Use the embedding layer of the pre-trained BERT model to convert the text into a language feature vector Z t ; The language feature vector Z t Input the encoder layer of the BERT model in sequence for feature extraction to obtain the final text feature vector Z t : in, is the original text feature of the i1+1th layer; is the original text feature of the i1th layer; t is the representative text; It is the 2i1+1 layer in the DETR backbone network; It is the 2i1th layer in the DETR backbone network; The Triple Alignment module is introduced to iteratively update query, text, and visual features through a three-way attention mechanism to eliminate the domain differences between them: Z o Initialized by a learnable embedding, including a regression query and multiple segmentation queries; A triple multi-head attention mechanism is used to triangulate the query, text, and visual features so that they share the same feature space and alleviate domain differences. The formula is: Among them, Tri-MHA is the triangle attention mechanism; Z o is the original query feature; Z' o is the updated query feature; Z t ' is the updated text feature; is the updated visual feature; o represents the query; The updated features are merged back into the original features, and the formula is: The aligned text feature vector Z t and visual feature vector The concatenation forms a multimodal feature vector, which is then input into the Transformer encoder layer for feature fusion to obtain the encoder output containing text and visual information. The encoder output is then input into the decoder. The decoder uses the bbox2seg scheme to convert the target box annotation into a segmentation mask.
9. The automatic meter reading and identification method for a substation inspection vehicle according to claim 8, characterized in that: S6 includes: The digital region image block U1 output by SegVG is cropped and input into the MHAFF module to meet the input requirements of CNN and Transformer; Use CNN to extract local features of the image Capture local information such as edges and textures of numbers; use Transformer to extract global features of images Capture the positional relationship and overall layout of the numbers in the entire image; Among them, X1 is the feature matrix extracted by the Transformer branch; Y is the feature matrix extracted by the CNN branch; A multi-head attention mechanism is used to fuse the features extracted by CNN and Transformer to obtain a more comprehensive feature representation; Classify the fused features and identify the numbers in the image.
10. The automatic meter reading and identification method for a substation inspection vehicle according to claim 9, characterized in that: The multi-head attention mechanism is used to fuse the features extracted by CNN and Transformer. The steps are as follows: (1) Generate Q, K, V matrices: Q=X1W Q ,K=YW K ,V=X1W V Among them, Q is the query matrix, which is used to find relevant information; K is the key matrix, which is used to store information; V is the value matrix, which contains the information to be found W Q 、W K 、W V It is a learnable weight matrix used to map the input sequence to the Query, Key, and Value matrices; Calculate the attention score: Among them, Attention(Q,K,V) is to calculate the attention score, which represents the similarity between Query and Key; is a scaling factor used to prevent gradient explosion; softmax is a normalization function that converts the attention score into a probability distribution; (2) Calculating multi-head attention: MHA(Q,K,V)=Concat(head1,head2,...,head h )WO Among them, MHA (Q, K, V) is a multi-head attention mechanism, which inputs Query, Key and Value into h independent attention heads respectively and outputs the fused features; head1, head2, ..., head h There are h independent attention heads, each responsible for finding different information; W O is a learnable weight matrix used to linearly combine the outputs of h attention heads; (3) Generate fused feature vector: Z=f(MHA) Where Z is the fused feature vector; f(·) is a nonlinear function; The fused features are classified to identify the numbers in the image. The specific steps are as follows: (1) Use the fully connected layer to map the fused features to the classification results: Z'=AZ+b Where Z' is the classification result vector, each element represents the probability of belonging to a certain digital category; A is the weight matrix used to map the fused feature vector to the classification result vector; b is the bias vector; (2) The output is converted into a probability distribution through the softmax function, thereby realizing the classification and recognition of the numbers in the image: Where P(cow=k|Z') is the probability that sample i belongs to digital category k; k is the category index; C1 is the total number of categories; and e is the base of the natural logarithm.
Citation Information
Patent Citations
Substation inspection unmanned aerial vehicle path planning method based on fusion A*- grey wolf algorithm
CN119642839A
Method and device for planning global path of unmanned vehicle
WO2021135554A1
Multi-robot trajectory planning method
WO2022241808A1
Drivable area detection and autonomous obstacle avoidance method for unmanned transportation device for deep, confined spaces
WO2023123642A1