Dynamic following method for unmanned aerial vehicle routing inspection fan blade based on brain-like vision
Through the dual-channel processing mechanism of the brain-like visual system, combined with the V1, V2 and V4 visual cortex models, the dynamic detection and following problems of wind turbine blades by drones in a wind turbine yaw environment are solved, and efficient wind turbine blade inspection and monitoring are achieved.
Patent Information
- Application Number
- CN202510798419.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-09-23
AI Technical Summary
When inspecting wind turbine blades, drones face the problems of limited computing resources and the complex dynamic environment of the wind turbine yaw, making it difficult to achieve accurate moving target perception and stable tracking and positioning.
A drone inspection method based on brain-like vision is adopted. Through the V1 primary visual cortex model, V2 visual cortex model and V4 visual cortex model combined with the brain-like tracking model, the dual-pathway processing mechanism of the brain-like vision system is utilized to simulate the processing method of the human visual system to achieve dynamic detection and tracking of wind turbine blades.
It improves the computing efficiency and resource utilization of drones in wind turbine blade inspections, enhances their adaptability to complex environments and their ability to identify target features, and ensures stable tracking and monitoring accuracy of wind turbine blades.
Smart Images

Figure CN120689367A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of brain-like science, visual algorithms, pulse neural networks, and target detection and tracking technology, and in particular to an intelligent inspection method for following wind turbine blades. Background Art
[0002] As a key component in wind turbines' wind energy conversion, blades are prone to damage and even the destruction of the entire turbine when operated in desert, mountain, and marine environments for extended periods. To effectively ensure blade safety while in service, health monitoring technologies—using drones to monitor blade condition in real time and identify and provide early warning of damage—are becoming a research hotspot in the wind power industry. During drone inspections of wind turbines, real-time tracking of blades when the turbine yaws presents numerous challenges. Firstly, drones face strict limitations on hardware weight and energy consumption during flight, and their onboard edge computing platforms typically have limited computing power. This results in slow computational speeds when handling complex image recognition and target tracking tasks during inspections, making it difficult for drones to quickly respond to real-time changes in the blades. Furthermore, the position of the blades changes in real time. During operation, the blades not only rotate as the turbine yaws, but also experience complex swinging and twisting motions due to changes in wind speed and direction. This makes it difficult for drones to continuously and stably track the blades during inspections. Furthermore, when blades rotate rapidly or are obscured by other blades or the tower, the drone's visual system may temporarily lose track of the target. Re-locking on to the blade requires additional time and computing resources, further exacerbating the challenge of real-time tracking of blades in wind turbine yaw conditions. Therefore, utilizing limited computing resources to accurately perceive moving targets and ensure stable tracking and positioning of wind turbine blades is a pressing technical challenge.
[0003] While existing target detection and tracking methods perform well in feature-rich real-world scenarios, their effectiveness tends to diminish in complex, resource-constrained environments. Furthermore, while traditional target detection algorithms are widely used in wind turbine blade detection, and brain-inspired networks offer biological advantages and transfer research value in processing motion information. For moving object perception and key target tracking tasks, the visual cortex processing mechanism can efficiently and accurately capture and follow targets. However, due to the limited computing resources of drone edge platforms and the complex dynamic environment of wind turbine yaw, existing algorithms still have significant room for improvement in real-time target detection and tracking of blade pose changes during wind turbine yaw. Summary of the Invention
[0004] In view of this, in order to solve the dynamic detection and tracking problems caused by the yaw of wind turbines when drones inspect wind turbine blades, the purpose of the present invention is to provide a dynamic following method for drone inspection of wind turbine blades based on brain-like vision, so as to realize dynamic detection and accurate following of the posture of wind turbine blades, effectively reduce the real-time calculation complexity when the posture of wind turbine blades changes, and improve the stable following accuracy of drones for wind turbine blades.
[0005] The present invention provides a method for dynamically following wind turbine blades inspected by a drone based on brain-like vision, comprising the following steps:
[0006] 1) UAV captures images of wind turbine blades;
[0007] 2) The image of the wind turbine blade collected at time t is represented by the brightness distribution function I(x, y, t), where x and y represent the two-dimensional spatial coordinates. The following scale transformation is performed on I(x, y, t):
[0008]
[0009] In the formula, * represents the convolution operation; the three scale images obtained by formula (1) are input into the V1 primary visual cortex model, and the V1 primary visual cortex model is transformed into the V1 primary visual cortex model through the three-dimensional Gaussian kernel function f r (x, y, t) simulates the selection of input signals by simple cells in the primary visual cortex of V1. The selection process is described as follows:
[0010]
[0011] Where * represents the discrete convolution operation, the subscript r represents different spatiotemporal analysis scales, and σ v1simple Radius for simple cell detection;
[0012] Under the conditions of spatial coordinates (x, y), scale r and orientation k, the simple cell response is modeled as the third-order partial differential of a Gaussian function at that orientation:
[0013]
[0014] In the formula, ! represents the factorial operation, α v1lin is the scale factor, X represents the x-coordinate component of the spatial coordinate, Y represents the y-coordinate component of the spatial coordinate, and T represents the time scale component. and They represent the x, y, and t components of the three-dimensional orientation vector of the spatiotemporal filter under orientation k. Formula (3) is expressed in the following vector form:
[0015] L r =α v1lin Mb r (4)
[0016] Where, Lr and b r They are the activation part and differential component in scale r, and M is the transformation matrix composed of coefficients;
[0017] To L kr Perform dynamic adjustments to generate a stable response S kr :
[0018]
[0019] Where, the conversion coefficient α filt→rate,r Implement the mapping from the basic filter response to the pulse emission rate, α v1semi is the half-activation constant, σ v1norm is the Gaussian window radius, α v1rect is the proportionality coefficient;
[0020] Integrate the responses of neighboring simple cells S kr The activation intensity of complex cells in the V1 primary visual cortex is obtained, and the mathematical expression is:
[0021]
[0022] Where, α v1comp is the proportionality coefficient;
[0023] When the activation intensity C kr When (x, y, t) exceeds the set activation threshold, the complex cells in the primary visual cortex of V1 begin to discharge. The complex cells in the primary visual cortex of V1 convert the average discharge rate into a spike train. The intermediate temporal lobe model of the MT receives the spike train output by the complex cells in the primary visual cortex of V1. The intermediate temporal lobe MT contains component direction-selective neurons (CDS) and pattern direction-selective neurons (PDS). The spike train is summarized in the component direction-selective neurons (CDS). The connection weights from the primary visual cortex of V1 to the component direction-selective neurons (CDS) are configured as follows:
[0024]
[0025] In the formula, the change matrix M and the bias vector b r Keeping consistent with formula (4), The value is Here the ′ in the upper right corner indicates the partial derivative calculation; is a unit vector of any space-time orientation, where the ' in the upper right corner represents a vector, and Unit vectors in the spatial coordinates x, y and time t directions respectively; Represents the unit vector in the specific spatial coordinates X, Y and time T direction respectively; interpolation weight represents the neural connection strength between the kth V1 complex cell and the component direction selective neuron CDS in the intermediate temporal lobe area; the interpolation weight is calculated by the following method To tune:
[0026] Define the direction of MT neurons as Speed is According to the direction and speed of the MT neuron, the corresponding unit vector spherical rectangular coordinates are calculated and the weight is generated:
[0027]
[0028] Where o1, o2, and o3 represent triple loop calculations and satisfy o1! + o2! + o3! = 6, x′ i , y′ i , t′ i are the coordinate values after factorial adjustment, used to calculate the weight; i∈0,1,…7;
[0029] Tune the direction and speed of MT neurons to be consistent with the spatiotemporal direction of V1 neurons, so x i =cos(N i,1 )cos(N i,2 ), y i =cos(N i,1 )sin(N i,2 ), t i = sin(N i,1 ), and o1! +o2! +o3! = 6, sequentially select the interpolation weight ω in the range of j∈1,2,…,7 i =[ω i,1 ,ω i,2 ,...,ω i,j ];
[0030]
[0031] Among them, pinv is the pseudo-inverse operation, D is the three-dimensional spherical coordinate matrix set of the V1 filter;
[0032] The pattern-selective neurons PDS in the intermediate temporal lobe of MT directly receive input from the speed-direction tuned CDS, thereby inheriting its spatiotemporal selectivity from the θ-sense CDS. CDS Direction-selective CDS neurons (x CDS ,y CDS ) to have θ PDS Direction-selective PDS neurons (x PDS ,y PDS ) is constructed according to the following rules:
[0033]
[0034] Where Δθ = θ PDS -θ CDS , Δx=x PDS -x CDS , Δy=y PDS -y CDS , Gaussian integration range σ PDS,pool = 3 pixels and normalized scaling factor α CDS→PDS =1.8992; when the direction difference satisfies |Δθ|>π / 2, the calculated connection strength becomes negative, and the neural signal will be directed to the inhibitory interneuron cluster for processing:
[0035]
[0036] Where σ PDS,tuned,dir <45°, σ PDS,tuned,loc =2;
[0037] The motion feature vector of each spatial position (x, y) is calculated by the following formula
[0038]
[0039] Where r θ (x, y) is the spike rate of the PDS neuron with θ direction selectivity located at (x, y), e θ is the unit direction vector in the θ direction;
[0040] Integrate the motion feature vector of each position at time t Get the motion feature vector The intensity of the impulse response I ’ (x,y,t);
[0041] 3) Input the wind turbine blade image captured by the drone at time t into the V2 visual cortex model;
[0042] The V2 visual cortex model extracts the texture features of the leaves through even-symmetric Gabor filters and odd-symmetric Gabor filters. The texture feature expression is as follows:
[0043]
[0044] Where I(x,y) represents the input wind turbine blade image, (x,y) represents the coordinates of the pixel, and the function expressions of the even-symmetric Gabor filter and the odd-symmetric Gabor filter are as follows:
[0045]
[0046] The filter patch coordinates are defined as (x, y), where X = xcosθ + ysinθ, Y = -xsinθ + ycosθ, θ is the direction. In an even-symmetric Gabor filter, there are four sensing directions: θ = 0, π / 4, π / 2, and π / 4. In an odd-symmetric Gabor filter, there are eight sensing directions: θ = 0, ±π / 4, ±π / 2, ±3π / 4, and π. s is the scale, η is the effective width, μ is the wavelength, and γ is the compensation factor.
[0047] The V2 visual cortex model extracts the color features of the leaves through the mapping matrix. The color feature expression is as follows:
[0048] V2 color (x, y, c) = Map (R (x, y), G (x, y), B (x, y), c) (20)
[0049] Among them, R(x,y), G(x,y), B(x,y) represent the RGB color channel values of the image respectively, c represents the color name index, and Map() is defined as the conversion operator from the three primary color space to the probability distribution of 11 standard colors; in order to ensure consistency with V2 gabor Keep the feature dimensions consistent at the same scale V2 gabor (x, y, 12), and for grayscale images, V2 gabor Set to 0;
[0050] The extracted texture features are used as the real part of the complex number, the extracted color features are used as the imaginary part of the complex number, and the texture features and color features are combined to obtain a complex feature map;
[0051] 4) Input the obtained complex feature map into the V4 visual cortex model and divide the complex feature map into n s ×n s = 4×4 grid cells Σ;
[0052] For grid cell Σ, the V4 visual cortex model responds to texture features as follows:
[0053]
[0054] Among them, is the normalization factor, δ x , δ y is the displacement deviation;
[0055] For grid cell Σ, the V4 visual cortex model responds to color features as follows:
[0056]
[0057] 5) Inputting the response of the V4 visual cortex model into a brain-like tracking model to achieve target tracking; the brain-like tracking model includes the inferior frontal cortex (IT) and the prefrontal cortex (PFC);
[0058] ① The inferior frontal cortex IT pools the input V4 response within its receptive field and obtains the inferior frontal cortex IT response r IT , the expression is as follows:
[0059] r IT =exp(-β||X V4 -P V4 || 2 ) (twenty four)
[0060] Where β is the sharpness of the tuning coefficient, X V4 is the visual input of the inferior frontal cortex IT, and P is the stored prototype;
[0061] Will respond to r IT Perform the following conversion:
[0062]
[0063] Where T represents the matrix transpose operation, β=1 / 2σ 2 =1 / 2;
[0064] The activation response of the V4 visual cortex model to the visual input X is denoted as V4 X (x, y, k), the storage prototype mode is recorded as V4 P (x, y, k), traverse the feature space according to the index parameter k, and integrate the multi-channel feature responses through the mean pooling operation. The expression is as follows:
[0065]
[0066] Where K is the total number of feature responses of all channels in the feature space;
[0067] ② The IT(x, y) obtained by the inferior frontal cortex IT is input into the prefrontal cortex PFC, which then obtains the coordinates of the tracked target at time t+1. The connection between the inferior frontal cortex IT and the prefrontal cortex PFC is as follows:
[0068]
[0069] Where W is the connection weight of the neural network; the position coordinates of the tracked target are calculated as follows:
[0070] At time step t, a linear function is used to calculate IT t+1 (x, y), calculated as follows:
[0071]
[0072] Since the dot product in the frequency domain is equivalent to the convolution in the time domain, formula (28) is converted to the frequency domain and the convolution operation is accelerated by fast Fourier transform, and the calculation is as follows:
[0073]
[0074] in, represents the fast Fourier transform function, ⊙ represents the element-wise dot product;
[0075] In order to estimate the neural connection weight W, the prefrontal cortex PFC response map to the target position is modeled as follows:
[0076]
[0077] Among them, η s is the scale parameter; (x0, y0) is the coordinate of the target perception guidance point G, and (x0, y0) is obtained by the following strategy:
[0078] G(x0, y0)=Org(I′(x, y, t)) (31)
[0079] Wherein, Org represents the target perception guide point extraction strategy, I′(x, y, t) is the impulse response intensity obtained in step 2);
[0080] The neural connection weight W is expressed as follows:
[0081]
[0082] In the (t+1)th frame, the target position coordinates The calculation is determined by maximizing the PFC response graph at this moment:
[0083]
[0084] in Represents the inverse FFT function. According to the spatial domain and frequency domain, the brain-like tracking model is updated as follows. The weight parameter is updated in the tth frame as follows:
[0085]
[0086] Where α is the tuning parameter, represents the response sequence model of V4 visual cortex cells, and is the neural weight calculated by formula (32).
[0087] Furthermore, the brain-like vision-based UAV inspection wind turbine blade dynamic following method also includes accelerating the extraction of blade texture features by the V2 visual cortex model through the following method:
[0088] Orthogonalize the even-symmetric Gabor filter and the odd-symmetric Gabor filter, and calculate the orthogonal Gabor response map at each pixel (x, y). The calculation formula is as follows:
[0089]
[0090] Where G x (x, s(η, μ)) and G y (y, s(η, μ)) are the components of the orthogonal Gabor response in the x and y directions; let Θ(x, y, s(η, μ)) and A(x, y, s(η, μ)) denote the direction and magnitude of the Gabor gradient at pixel (x, y), respectively, as follows:
[0091]
[0092] When Θ(x, y, s(η, μ)) is within the corresponding direction range, the amplitude A(x, y, s(η, μ)) is set to the approximate response of the V2 visual cortex cell, which is calculated as follows:
[0093]
[0094] Beneficial effects of the present invention:
[0095] 1. The brain-inspired tracking model designed in this paper, based on a dual-pathway processing mechanism of brain-inspired vision, draws on the processing methods of the human visual system and can better adapt to different environments and target characteristics. In the field of wind turbine blade inspection, this method can better handle wind turbine blades of different models and installation locations, as well as various complex meteorological conditions and terrain environments, thereby increasing the scope and flexibility of inspection work.
[0096] 2. Compared with traditional deep learning models, the V1 primary visual cortex model in the present invention has higher computing efficiency and lower resource consumption. It can quickly perceive and locate the posture of wind turbine blades under the limited hardware configuration and computing power of the drone edge computing platform, thereby more effectively monitoring the operating status of wind turbine blades.
[0097] 3. When facing challenges in complex environments, the combination of the V2 visual cortex and the V4 visual cortex in the present invention can better understand and identify the target's shape, color and other characteristic information; and combined with the guided learning mechanism from V1-MT to IT and PFC, it enhances the adaptability and prediction ability of the brain-like tracking model to target movement. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] Figure 1 This is the spatiotemporal filtered spectrum of the V1 primary visual cortex model.
[0099] Figure 2 Schematic diagram of the receptive field of MT cells.
[0100] Figure 3 Schematic diagram of the motion perception method from V1 to MT.
[0101] Figure 4 A diagram of the motion perception and target tracking process based on the dual pathways of the brain-like visual cortex.
[0102] Figure 5 This is a flow chart of the dynamic following method for UAV inspection of wind turbine blades based on brain-like vision. DETAILED DESCRIPTION
[0103] The present invention will be further described below with reference to the accompanying drawings and examples.
[0104] In this embodiment, the method for dynamically following wind turbine blades inspected by a drone based on brain-like vision includes the following steps:
[0105] 1) UAV collects images of wind turbine blades.
[0106] 2) The image of the wind turbine blade collected at time t is represented by the brightness distribution function I(x, y, t), where x and y represent the two-dimensional spatial coordinates. The following scale transformation is performed on I(x, y, t):
[0107]
[0108] Where * represents a convolution operation. The multi-scale strategy provides feature representations at multiple resolutions for each spatial location, enhancing the model's robustness in detecting moving objects of varying sizes.
[0109] The three scale images obtained by formula (1) are input into the V1 primary visual cortex model, and the V1 primary visual cortex model is transformed into the V1 primary visual cortex model through the three-dimensional Gaussian kernel function f r (x, y, t) simulates the selection of input signals by simple cells in the primary visual cortex of V1. The selection process is described as follows:
[0110]
[0111] Where * represents the discrete convolution operation, the subscript r represents different spatiotemporal analysis scales, and σ v1simple =1.25 is the simple cell detection radius. Seven groups of linear receptive fields with specific spatiotemporal orientations are configured in the V1 simple cells, showing uniform distribution characteristics in the three-dimensional space of time, space and frequency. The unit vector represents the spatial orientation k = 1, 2, 3, ..., 7. Under the conditions of spatial coordinates (x, y), scale r and orientation k, the simple cell response is modeled as the third-order partial differential of the Gaussian function at this orientation:
[0112]
[0113] In the formula, ! represents the factorial operation, α v1lin =6.6084 is the scale factor, X represents the x-coordinate component of the spatial coordinate, Y represents the y-coordinate component of the spatial coordinate, and T represents the time scale component. and They represent the x, y, and t components of the three-dimensional orientation vector of the spatiotemporal filter under orientation k. Formula (3) is expressed in the following vector form:
[0114] L r =α v1lin Mb r (4)
[0115] Where, L r and b r are the activation part and the differential component in scale r, respectively, and M is the transformation matrix composed of 7×7 coefficients.
[0116] In the calculation process of formula (4), when the processing area (x, y) is located at the edge of the image, the activation intensity L of the simple cell will increase sharply. Therefore, the scale adjustment factor is introduced to adjust L kr Dynamic adjustment is performed, combining half-wave rectification with normalization operation under the constraint of large-scale Gaussian window function to produce a stable response S kr :
[0117]
[0118] Where, the conversion coefficient α filt→rate,r Implement the mapping from the basic filter response to the pulse emission rate, setting the reference value to 15Hz. v1semi =0.1 is the half-activation constant, σ v1comp =1.6 pixels is the Gaussian window radius, α v1rect =1.9263 is the proportional coefficient.
[0119] Integrate the responses of neighboring simple cells S kr The activation intensity of complex cells in the V1 primary visual cortex is obtained, and the mathematical expression is:
[0120]
[0121] Where, α v1comp =0.1 is the proportional coefficient. kr(x, y, t) describes the activation strength of a complex cell at a specific spatial location (x, y) and time t. It is a weighted sum of the responses of neighboring simple cells, with weights determined by a Gaussian function. This integration reflects the complex cell's integrated processing of the inputs from neighboring simple cells.
[0122] When the activation intensity C kr When (x, y, t) exceeds the set activation threshold, complex cells in the primary visual cortex of V1 begin firing. The final output neuronal average firing rate conforms to the characteristics of a Poisson process, with the value being the sum of spikes per unit time. This information, processed by the V1 visual tape, is then transmitted to the middle temporal lobe (MT) in the form of a spike train.
[0123] The linear receptive fields and weights of V1 responses satisfy two fundamental constraints: First, the sum of V1 responses (as a function of the spatiotemporal frequency of the stimulus) should be roughly constant over the desired frequency range. This function corresponds to the spatiotemporal frequency sensitivity of the entire system, known as the tiling property. Second, the set of V1 neurons contains receptive fields tuned to a predetermined set of spatiotemporal directions, but the responses of V1 neurons to any intermediate spatiotemporal direction can be precisely interpolated from this set. In other words, the information represented by V1 neurons does not depend on the specific choice of spatiotemporal direction, known as the interpolation property. These two constraints are actually independent of each other.
[0124] Both constraints are satisfied by the linear receptive field of the directional derivatives of the differentiable function g(x, y, t). For example, in the unit vector The first derivative in the direction can be written as:
[0125]
[0126] where {u x ,u y ,u t} is a unit vector This equation shows that the partial derivative in any direction is a linear combination of the partial derivatives in the x, y, and t directions. In other words, the first-order derivative in any direction can be interpolated as a linear combination of a fixed set of three derivatives. The flattening property of the first-order derivatives is easy to see in the frequency domain. Computing the partial derivatives (in space or time) is equivalent to multiplying a linear ramp function in the frequency domain. For example, the Fourier transform of the partial derivative of g(x, y, t) with respect to x is:
[0127]
[0128] Here i is an imaginary number, G(ω x ,ω y ,ω t) is the Fourier transform of g(x, y, t). The power spectrum of the derivatives in the (x, y, t) direction is summed up to get:
[0129]
[0130] In the formula is the radial frequency. Choosing the right g(x, y, t) can yield almost any desired spatiotemporal frequency coverage. For example, G(ω x ,ω y ,ω t )=1 / ω r It can uniformly cover the entire spatiotemporal frequency domain. The tiling and interpolation properties extend to higher-order (separable or directional) derivatives, which have narrower directional tuning curves than first-order derivatives and require a larger set to tile and interpolate all spatiotemporal directional rotations. In particular, in three-dimensional space (x, y, t), the size of a complete set of N-order directional derivatives is (N+1)(N+2) / 2. The interpolation property also extends to squared directional derivatives, requiring a larger set. The linear receptive field of V1 in the current model is the third-order derivative of a spatiotemporal Gaussian function. *The derivative order is chosen to match the typical directional bandwidth (orientation) of V1 neurons, and receptive fields of varying sizes are generated by adjusting the radius of the underlying Gaussian. The complete V1 receptive field encompasses 28 spatiotemporal directions, uniformly distributed on the surface of a sphere in the spatiotemporal frequency domain. The tiling region is approximated as a spherical ring with a bandwidth of approximately 1.5 octaves and reduced sensitivity near the time-frequency axis, which can be optimized while satisfying the constraints of the V1 model.
[0131] from Figure 1 As can be seen in the figure, the spatiotemporal filter spectrum of the V1 neuron model is displayed. Each direction-selective receptive field corresponds to a sphere, showing its response characteristics in the three-dimensional spatiotemporal frequency domain. The green and blue areas on the sphere represent the response strength of different frequencies and directions, respectively, where green corresponds to high-frequency response and blue corresponds to low-frequency response. The model in the figure is V1MT 38 The model and the lightweight V1MT7 model realize the selective response of the primary visual cortex to direction and speed, significantly reducing computational complexity and power consumption while retaining core perception capabilities.
[0132] The complex cells in the V1 primary visual cortex convert the average discharge rate into a pulse train. The MT intermediate temporal lobe model receives the pulse train output by the complex cells in the V1 primary visual cortex. The intermediate temporal lobe MT contains component direction-selective neurons (CDS) and pattern direction-selective neurons (PDS). CDS neurons also have spatiotemporal selectivity. In this embodiment, 24 CDS neurons are configured at each spatial position (covering 8 directions × 3 speed combinations), the direction interval step size is 45°, and the speed gradient is set to 1.5 / 0.125 / 9 pixels per frame. At the same time, 8 PDS neurons are deployed at each position for global motion analysis. Neuronal dynamics are implemented using the Izhikevich model.
[0133] The spike train is summarized in the component direction selective neuron CDS, and the connection weights from the V1 primary visual cortex to the component direction selective neuron CDS are configured as follows:
[0134]
[0135] In the formula, the change matrix M and the bias vector b r Keeping consistent with formula (4), The value is Here the ′ in the upper right corner indicates the partial derivative calculation; is a unit vector of any space-time orientation, where the ' in the upper right corner represents a vector, and Unit vectors in the spatial coordinates x, y and time t directions respectively; Represents the unit vector in the specific spatial coordinates X, Y and time T direction respectively; interpolation weight represents the neural connection strength between the kth V1 complex cell and the component direction selective neuron CDS. This connection only occurs between neurons with the same spatial position (x, y). The interpolation weight is adjusted by the following method. To tune:
[0136] Define the direction of MT neurons as Speed is According to the direction and speed of the MT neuron, the corresponding unit vector spherical rectangular coordinates are calculated and the weight is generated:
[0137]
[0138] Where o1, o2, and o3 represent triple loop calculations and satisfy o1! + o2! + o3! = 6, x′ i , y′ i , t′ i are the coordinate values after factorial adjustment, used to calculate the weight; i∈0,1,…7;
[0139] Tune the direction and speed of MT neurons to be consistent with the spatiotemporal direction of V1 neurons, so x i =cos(N i,1 )cos(N i,2 ), y i =cos(N i,1 )sin(N i,2 ), t i = sin(N i,1 ), and o1! +o2! +o3! = 6, sequentially select the interpolation weight ω in the range of j∈1,2,…,7 i =[ω i,1 ,ω i,2 ,...,ω i,j ];
[0140]
[0141] Where pinv is the pseudo-inverse operation, and D is the set of 7×2 coordinate matrices of the three-dimensional spherical surface of the V1 filter. In order to adjust the MT CDS unit to different speeds, is the only parameter that needs to be adjusted. The CDS unit receives projections from V1 complex cells at all three spatiotemporal resolutions. Since the interpolated weights can be positive or negative, it is necessary to pass on projections with negative weights to the inhibitory neuron population. If This will be used to generate excitatory projections from V1 complex cells to the MT inhibitory population, which will send a one-to-one connection back to the pool of MT CDS cells. Finally, the MT neurons subtract the responses of V1 neurons that are not near their preferred velocity plane. This is done by subtracting the mean of the weights from each weight (thus producing a set of weights with a mean of zero). These seven zero-valued weights correspond to the final linear receptive field weights of the MT neurons. The resulting spatiotemporal frequency response function is smooth and depends only on the Euclidean distance to the plane.
[0142] The pattern-selective neurons PDS in the intermediate temporal lobe of MT directly receive input from the speed-direction tuned CDS, thereby inheriting its spatiotemporal selectivity from the θ-sense CDS. CDS Direction-selective CDS neurons (x CDS ,y CDS ) to have θ PDS Direction-selective PDS neurons (x PDS ,y PDS ) is constructed according to the following rules:
[0143]
[0144] Where Δθ = θ PDS -θ CDS , Δx=xPDS -x CDS , Δy=y PDS -y CDS , Gaussian integration range σ PDS,pool = 3 pixels and normalized scaling factor α CDS→PDS =1.8992; when the direction difference satisfies |Δθ|>π / 2, the calculated connection strength becomes negative, and the neural signal will be directed to the inhibitory interneuron cluster for processing:
[0145]
[0146] Where σ PDS,tuned,dir <45°, ensuring that only one of the eight functional units is active. PDS,tuned,loc = 2 pixels, the inhibitory neuron network forms a feedback connection pathway with the pattern selection neurons.
[0147] The motion feature vector of each spatial position (x, y) is calculated by the following formula
[0148]
[0149] Where r θ (x,y) is the spike rate of the PDS neuron with θ direction selectivity located at (x,y), e θ is the unit direction vector in the direction of θ.
[0150] Integrate the motion feature vector of each position at time t Get the motion feature vector The intensity of the impulse response I ’ (x,y,t).
[0151] 3) Input the wind turbine blade image I(x,y) collected by the drone at time t into the V2 visual cortex model.
[0152] In scene classification and saliency detection research, joint color and texture information has been shown to be a key factor. To effectively represent colored objects, the V2 visual cortex model integrates color and texture information.
[0153] The V2 visual cortex model extracts the texture features of the leaves through even-symmetric Gabor filters and odd-symmetric Gabor filters. The texture feature expression is as follows:
[0154]
[0155] Where I(x,y) represents the input wind turbine blade image, (x,y) represents the coordinates of the pixel, and the function expressions of the even-symmetric Gabor filter and the odd-symmetric Gabor filter are as follows:
[0156]
[0157] The filter patch coordinates are defined as (x, y): X = xcosθ + ysinθ, Y = -xsinθ + ycosθ, where θ represents the direction. Even-symmetric Gabor filters have four receptive directions: θ = 0, π / 4, π / 2, and π / 4; odd-symmetric Gabor filters have eight receptive directions: θ = 0, ±π / 4, ±π / 2, ±3π / 4, and π. s represents the scale, η represents the effective width, μ represents the wavelength, and γ represents the compensation factor. Specifically, the Gabor filters are arranged in a pyramidal pattern, with sizes ranging from 7×7 to 15×15 pixels in steps of 2 pixels to simulate the receptive fields of V2 visual cortical cells. This yields 60 different types of V2 receptive fields.
[0158] The V2 visual cortex model extracts the color features of the leaves through the mapping matrix. The color feature expression is as follows:
[0159] V2 color (x,y,c)=Map(R(x,y),G(x,y),B(x,y),c) (20)
[0160] Among them, R(x,y), G(x,y), B(x,y) represent the RGB color channel values of the image respectively, c represents the color name index, and Map() is defined as a conversion operator from the three primary color space to the probability distribution of 11 standard colors (black, brown, green, pink, red, yellow, blue, gray, orange, purple and white); in order to ensure consistency with V2 gabor Keep the feature dimensions consistent at the same scale V2 gabor (x, y, 12), and for grayscale images, V2 gabor Set to 0.
[0161] The extracted texture features are used as the real part of the complex number, the extracted color features are used as the imaginary part of the complex number, and the texture features are combined with the color features to obtain a complex feature map. In this method, the color target is represented by 60 complex feature maps, of which 60 texture feature maps are obtained by convolution operation of multi-scale Gabor filter and input image, and used as the real part of the complex number. In addition, 12 color feature maps are copied 5 times (corresponding to 5 scales) as the imaginary part of the complex number; by combining texture features with color features and using them as the real part and imaginary part of the complex feature map respectively, this method can maintain a balance between different types of features. However, 60 convolution calculations have a serious impact on the real-time performance of the model framework. This embodiment speeds up the extraction of leaf texture features by the V2 visual cortex model by the following method:
[0162] Orthogonalize the even-symmetric Gabor filter and the odd-symmetric Gabor filter, and calculate the orthogonal Gabor response map at each pixel (x, y). The calculation formula is as follows:
[0163]
[0164] Where G x (x, s(η, μ)) and G y (y, s(η, μ)) is the component of the orthogonal Gabor response in the x and y directions, G x (x, s(η, μ)) and G y (y, s(η, μ)) contains five different scales; let Θ(x, y, s(η, μ)) and A(x, y, s(η, μ)) represent the direction and amplitude of the Gabor gradient at the pixel (x, y), respectively, and the expressions are as follows:
[0165]
[0166] When Θ(x, y, s(η, μ)) is within the corresponding direction range, the amplitude A(x, y, s(η, μ)) is set to the approximate response of the V2 visual cortex cell, which is calculated as follows:
[0167]
[0168] 4) Input the obtained complex feature map into the V4 visual cortex model and divide the complex feature map into n s ×n s = 4×4 grid cells Σ;
[0169] For grid cell Σ, the V4 visual cortex model responds to texture features as follows:
[0170]
[0171] Among them, is the normalization factor, δ x , δ y is the displacement deviation;
[0172] For grid cell Σ, the V4 visual cortex model responds to color features as follows:
[0173]
[0174] 5) Inputting the response of the V4 visual cortex model into a brain-like tracking model to achieve target tracking; the brain-like tracking model includes the inferior frontal cortex IT and the prefrontal cortex PFC.
[0175] ① The inferior frontal cortex IT pools the input V4 response within its receptive field and obtains the inferior frontal cortex IT response r IT , the expression is as follows:
[0176] r IT =exp(-β||X V4 -P V4 || 2 ) (28)
[0177] Where β is the sharpness of the tuning coefficient; X V4 is the visual input of the inferior frontal cortex IT, that is, the output of the V4 layer under the visual input X; P is the stored prototype.
[0178] At runtime, the response map of IT cells is calculated at all locations according to formula (28) and for each V4 population. According to the above formula, the response of IT cells can be regarded as a kernel method based on the radial basis function. In the case where the radial basis function is a standard normal function (i.e., β = 1 / 2σ 2 =1 / 2), the response r IT Perform the following conversion:
[0179]
[0180] Where T represents the matrix transpose operation. Here, X T X and P T P is the autocorrelation coefficient, and its value remains almost unchanged, with less influence on the PFC cortex, while X T P is mainly in the linear region of the exponential function. This method uses linear mapping instead of the highly complex radial basis function to calculate the activation distribution of IT neurons. The activation response of the V4 visual cortex model to the visual input X is denoted as V4 X (x, y, k), the storage prototype mode is recorded as V4 P (x, y, k), traverse the 60-dimensional feature space according to the index parameter k, and integrate the multi-channel feature responses through the mean pooling operation to obtain scale and rotation invariant features. The expression is as follows:
[0181]
[0182] Where K is the total number of feature responses of all channels in the feature space.
[0183] ② The IT(x, y) obtained by the inferior frontal cortex IT is input into the prefrontal cortex PFC, which then obtains the coordinates of the tracked target at time t+1. The connection between the inferior frontal cortex IT and the prefrontal cortex PFC is as follows:
[0184]
[0185] Where W is the connection weight of the neural network; the position coordinates of the tracked target are calculated as follows:
[0186] At time step t, a linear function is used to calculate IT t+1 (x, y), calculated as follows:
[0187]
[0188] Since the dot product in the frequency domain is equivalent to the convolution in the time domain, formula (28) is converted to the frequency domain and the convolution operation is accelerated by fast Fourier transform, and the calculation is as follows:
[0189]
[0190] in, represents the fast Fourier transform function, and ⊙ represents the element-wise dot product.
[0191] In order to estimate the neural connection weight W, the prefrontal cortex PFC response map to the target position is modeled as follows:
[0192]
[0193] Among them, η s is the scale parameter; (x0, y0) is the coordinate of the target perception guidance point G, and (x0, y0) is obtained by the following strategy:
[0194] G(x0, y0)=Org(I′(x, y, t)) (35)
[0195] Wherein, Org represents the target perception guide point extraction strategy, I′(x, y, t) is the impulse response intensity obtained in step 2), and the strategy is as follows: in the process of the MT intermediate temporal lobe area perceiving the tracked target, the channel that preferentially responds to images moving to the right at a speed of more than 1.5 pixels per frame is called the medium-speed channel, and the channel that preferentially responds to images moving to the right at a speed of less than 1.5 pixels is called the low-speed channel; when the medium-speed channel has many and scattered connected areas, the vacancies in the low-speed channel are used to search for the target, and the vacant areas in the low-speed channel are used as a reference to improve the accuracy of the target search in the medium-speed channel; when the low-speed channel has many connected areas, the target is searched through the main response of the medium-speed channel, and the main response of the medium-speed channel is used to more effectively screen out the area that truly contains the target with the assistance of the numerous connected areas of the low-speed channel.
[0196] The neural connection weight W is expressed as follows:
[0197]
[0198] In the (t+1)th frame, the target position coordinates The calculation is determined by maximizing the PFC response graph at this moment:
[0199]
[0200] in Represents the inverse FFT function. According to the spatial domain and frequency domain, the brain-like tracking model is updated as follows. The weight parameter is updated in the tth frame as follows:
[0201]
[0202] Where α is the tuning parameter, represents the response sequence model of V4 visual cortex cells, and is the neural weight calculated by formula (36).
[0203] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A dynamic tracking method for wind turbine blade inspection using a drone based on brain-inspired vision, characterized by: The following steps are involved: 1) UAV captures images of wind turbine blades; 2) The image of the wind turbine blade collected at time t is represented by the brightness distribution function I(x, y, t), where x and y represent the two-dimensional spatial coordinates. The following scale transformation is performed on I(x, y, t): In the formula, * represents the convolution operation; the three scale images obtained by formula (1) are input into the V1 primary visual cortex model, and the V1 primary visual cortex model is transformed into the V1 primary visual cortex model through the three-dimensional Gaussian kernel function f r (x, y, t) simulates the selection of input signals by simple cells in the primary visual cortex of V1. The selection process is described as follows: Where * represents the discrete convolution operation, the subscript r represents different spatiotemporal analysis scales, and σ v1simple Radius for simple cell detection; Under the conditions of spatial coordinates (x, y), scale r and orientation k, the simple cell response is modeled as the third-order partial differential of a Gaussian function at that orientation: In the formula, ! represents the factorial operation, α v1lin is the scale factor, X represents the x-coordinate component of the spatial coordinate, Y represents the y-coordinate component of the spatial coordinate, and T represents the time scale component; and They represent the x, y, and t components of the three-dimensional orientation vector of the spatiotemporal filter under orientation k. Formula (3) is expressed in the following vector form: L r =a v1lin Mb r (4) Where, L r and b r They are the activation part and differential component in scale r, and M is the transformation matrix composed of coefficients; To L kr Perform dynamic adjustments to generate a stable response S kr : Where, the conversion coefficient α filt→rate,r Implement the mapping from the basic filter response to the pulse emission rate, α v1semi is the half-activation constant, σ v1norm is the Gaussian window radius, α v1rect is the proportionality coefficient; Integrate the responses of neighboring simple cells S kr The activation intensity of complex cells in the V1 primary visual cortex is obtained, and the mathematical expression is: Where, α v1comp is the proportionality coefficient; When the activation intensity C kr When (x, y, t) is greater than the set activation threshold, the complex cells in the V1 primary visual cortex begin to discharge; the complex cells in the V1 primary visual cortex convert the average discharge rate into a pulse train, and the MT intermediate temporal lobe model receives the pulse train output by the complex cells in the V1 primary visual cortex. The intermediate temporal lobe MT contains component direction-selective neurons CDS and pattern direction-selective neurons PDS. The spike train is summarized in the component direction selective neuron CDS, and the connection weights from the V1 primary visual cortex to the component direction selective neuron CDS are configured as follows: In the formula, the change matrix M and the bias vector b r Keeping consistent with formula (4), The value is Here the ′ in the upper right corner indicates the partial derivative calculation; is a unit vector of any space-time orientation, where the ' in the upper right corner represents a vector, and Unit vectors in the spatial coordinates x, y and time t directions respectively; Represents the unit vector in the specific spatial coordinates X, Y and time T direction respectively; interpolation weight represents the neural connection strength between the kth V1 complex cell and the component direction selective neuron CDS in the intermediate temporal lobe area; the interpolation weight is calculated by the following method To tune: Define the direction of MT neurons as Speed is According to the direction and speed of the MT neuron, the corresponding unit vector spherical rectangular coordinates are calculated and the weight is generated: Where o1, o2, and o3 represent triple loop calculations and satisfy o1! + o2! + o3! = 6, x′ i , y′ i , t′ i are the coordinate values after factorial adjustment, used to calculate the weight; i∈0,1,…7; Tune the direction and speed of MT neurons to be consistent with the spatiotemporal direction of V1 neurons, so x i =cos(N i,1 )cos(N i,2 ), y i =cos(N i,1 )sin(N i,2 ), t i = sin(N i,1 ), and o1! +o2! +o3! = 6, sequentially select the interpolation weight ω in the range of j∈1,2,…,7 i =[ω i,1 ,ω i,2 ,...,ω i,j ]; Among them, pinv is the pseudo-inverse operation, D is the three-dimensional spherical coordinate matrix set of the V1 filter; The pattern-selective neurons PDS in the intermediate temporal lobe of MT directly receive input from the speed-direction tuned CDS, thereby inheriting its spatiotemporal selectivity from the θ-sense CDS. CDS Direction-selective CDS neurons (x CDS ,y CDS ) to have θ PDS Direction-selective PDS neurons (x PDS, y PDS ) is constructed according to the following rules: Where Δθ = θ PDS -θ CDS , Δx=x PDS -x CDS , Δy=y PDS -y CDS , Gaussian integration range σ PDS,pool = 3 pixels and normalized scaling factor α CDS→PDS =1.8992; when the direction difference satisfies |Δθ|>π / 2, the calculated connection strength becomes negative, and the neural signal will be directed to the inhibitory interneuron cluster for processing: where, σ PDS,tuned,dir < 45°, σ PDS,tuned,loc = 2; The motion feature vector of each spatial position (x, y) is calculated by the following formula Where r θ (x,y) is the spike rate of the PDS neuron with θ direction selectivity located at (x,y), e θ is the unit direction vector in the θ direction; Integrate the motion feature vector of each position at time t Get the motion feature vector The intensity of the impulse response I'(x,y,t); 3) Input the wind turbine blade image captured by the drone at time t into the V2 visual cortex model; The V2 visual cortex model extracts the texture features of the leaves through even-symmetric Gabor filters and odd-symmetric Gabor filters. The texture feature expression is as follows: Where I(x,y) represents the input wind turbine blade image, (x,y) represents the coordinates of the pixel, and the function expressions of the even-symmetric Gabor filter and the odd-symmetric Gabor filter are as follows: The filter patch coordinates are defined as (x, y), where X = xcosθ + ysinθ, Y = -xsinθ + ycosθ, θ is the direction. In an even-symmetric Gabor filter, there are four sensing directions: θ = 0, π / 4, π / 2, and π / 4. In an odd-symmetric Gabor filter, there are eight sensing directions: θ = 0, ±π / 4, ±π / 2, ±3π / 4, and π. s is the scale, η is the effective width, μ is the wavelength, and γ is the compensation factor. The V2 visual cortex model extracts the color features of the leaves through the mapping matrix. The color feature expression is as follows: V2 color (x,y,c)=Map(R(x,y),G(x,y),B(x,y),c) (20) Among them, R(x,y), G(x,y), B(x,y) represent the RGB color channel values of the image respectively, c represents the color name index, and Map() is defined as the conversion operator from the three primary color space to the probability distribution of 11 standard colors; in order to ensure consistency with V2 gabor Keep the feature dimensions consistent at the same scale V2 gabor (x, y, 12), and for grayscale images, V2 gabor Set to 0; The extracted texture features are used as the real part of the complex number, the extracted color features are used as the imaginary part of the complex number, and the texture features and color features are combined to obtain a complex feature map; 4) Input the obtained complex feature map into the V4 visual cortex model and divide the complex feature map into n s ×n s = 4×4 grid cells Σ; For grid cell Σ, the V4 visual cortex model responds to texture features as follows: Among them, is the normalization factor, δ x , δ y is the displacement deviation; For grid cell Σ, the V4 visual cortex model responds to color features as follows: 5) Inputting the response of the V4 visual cortex model into a brain-like tracking model to achieve target tracking; the brain-like tracking model includes the inferior frontal cortex (IT) and the prefrontal cortex (PFC); ① The inferior frontal cortex IT pools the input V4 response within its receptive field and obtains the inferior frontal cortex IT response r IT , the expression is as follows: r IT =exp(-β||X V4 -P V4 || 2 ) (24) Where β is the sharpness of the tuning coefficient, X V4 is the visual input of the inferior frontal cortex IT, and P is the stored prototype; Will respond to r IT Perform the following conversion: Where T represents the matrix transpose operation, β=1 / 2σ 2 =1 / 2; The activation response of the V4 visual cortex model to the visual input X is denoted as V4 X (x, y, k), the storage prototype mode is recorded as V4 P (x, y, k), traverse the feature space according to the index parameter k, and integrate the multi-channel feature responses through the mean pooling operation. The expression is as follows: Where K is the total number of feature responses of all channels in the feature space; ② The IT(x,y) obtained by the inferior frontal cortex IT is input into the prefrontal cortex PFC, and the prefrontal cortex PFC obtains the coordinates of the tracked target at time t+1. The connection between the inferior frontal cortex IT and the prefrontal cortex PFC is as follows: Where W is the connection weight of the neural network; the position coordinates of the tracked target are calculated as follows: At time step t, a linear function is used to calculate IT t+1 (x, y), calculated as follows: Since the dot product in the frequency domain is equivalent to the convolution in the time domain, formula (28) is converted to the frequency domain and the convolution operation is accelerated by fast Fourier transform, and the calculation is as follows: in, represents the fast Fourier transform function, ⊙ represents the element-wise dot product; In order to estimate the neural connection weight W, the prefrontal cortex PFC response map to the target position is modeled as follows: Among them, η s is the scale parameter; (x0, y0) is the coordinate of the target perception guidance point G, and (x0, y0) is obtained by the following strategy: G(x0, y0)=Org(I′(x, y, t)) (31) Wherein, Org represents the target perception guide point extraction strategy, I′(x, y, t) is the impulse response intensity obtained in step 2); The neural connection weight W is expressed as follows: In the (t+1)th frame, the target position coordinates The calculation is determined by maximizing the PFC response graph at this moment: in Represents the inverse FFT function. According to the spatial domain and frequency domain, the brain-like tracking model is updated as follows. The weight parameter is updated in the tth frame as follows: Among them, α is the tuning parameter, ρ is the compensation parameter, represents the response sequence model of V4 visual cortex cells, and is the neural weight calculated by formula (32).
2. The brain-inspired vision-based UAV inspection method for wind turbine blades according to claim 1 is characterized by: It also includes the following methods to accelerate the extraction of leaf texture features by the V2 visual cortex model: Orthogonalize the even-symmetric Gabor filter and the odd-symmetric Gabor filter, and calculate the orthogonal Gabor response map at each pixel (x, y). The calculation formula is as follows: Where G x (x, s(η, μ)) and G y (y, s(η, μ)) are the components of the orthogonal Gabor response in the x and y directions; let Θ(x, y, s(η, μ)) and A(x, y, s(η, μ)) denote the direction and magnitude of the Gabor gradient at pixel (x, y), respectively, as follows: When Θ(x, y, s(η, μ)) is within the corresponding direction range, the amplitude A(x, y, s(η, μ)) is set to the approximate response of the V2 visual cortex cell, which is calculated as follows: