Salient contour matching-based method for target measurement in severe imaging environment
By collecting images with a binocular camera and combining local and global background light estimation models with a stacked denoising autoencoder network, the problem of insufficient target measurement accuracy in harsh imaging environments is solved, and accurate target size measurement in turbid media is achieved.
Patent Information
- Application Number
- PCT/CN2024/137286
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2024-12-06
- Publication Date
- 2025-10-02
AI Technical Summary
In harsh imaging environments, existing target measurement methods are unable to effectively extract target features, resulting in insufficient measurement accuracy. In particular, in turbid media, detailed information such as the texture of the target surface is severely lost, making it impossible to accurately perceive three-dimensional information and measure dimensions.
A method based on salient contour matching is adopted. Images are collected by binocular cameras, and a background light estimation model with local and global joint constraints is combined. A stacked denoising autoencoder network is used for image restoration and target detection. A contour matching descriptor is constructed for 3D reconstruction, scattering effects are removed, and key dimensions are measured.
It achieves accurate measurement of targets in harsh environments, can effectively retain contour information, adapt to environmental changes, and provide stable target size measurement results.
Smart Images

Figure CN2024137286_02102025_PF_FP_ABST
Abstract
Description
Target measurement method in harsh imaging environment based on salient contour matching Technical Field
[0001] The invention relates to a target measurement method in a harsh imaging environment based on significant contour matching, and belongs to the technical fields of image processing and computer vision. Background Art
[0002] Object measurement refers to measuring the size, shape, and position of an object. Object measurement is a key branch of computer vision, involving the measurement of objects in images or multidimensional data. It is used to extract information from these images or multidimensional data to aid decision-making. Object measurement plays a crucial role in computer vision. It provides the foundation for many applications, such as autonomous driving, robotic navigation, and industrial quality inspection. Through object measurement, computers can understand and interpret the content of images, enabling more intelligent decision-making and behavior.
[0003] Detection in harsh environments is an important part of target measurement, but the severe backscattering effect in harsh imaging environments can lead to strong background noise in optical images, making target features obscured and difficult to effectively extract, posing a huge challenge to perception in harsh environments. There are a large number of suspended particles such as soil and gravel in the natural environment, and the information in harsh optical images is severely attenuated, significantly shortening the visible distance. In turbid media, a large number of scattering particles will not only reduce the contrast of target imaging, but will also introduce a large amount of noise, resulting in a large loss of detailed information such as the texture of the target surface, making it impossible to provide reliable color, texture and other feature information for stereo matching, resulting in failure of three-dimensional information perception. Traditional target measurement methods are unable to measure the size of targets in harsh imaging environments. Summary of the Invention
[0004] The present invention provides a target measurement method in a harsh imaging environment based on significant contour matching, which solves the problem of insufficient accuracy of existing target measurement results in harsh imaging environments.
[0005] In order to solve the above technical problems, the technical solution adopted by the present invention is:
[0006] The target measurement method in a harsh imaging environment based on salient contour matching specifically includes the following steps:
[0007] Step 1) Acquire the target binocular image: Pre-calibrate the binocular camera using Zhang’s calibration method to obtain the camera’s internal and external parameters, and use the calibrated binocular camera to simultaneously acquire the target’s left eye image I in a harsh imaging environment. l With the right eye image I r ;
[0008] Step 2) Based on the local and global joint constraint harsh environment imaging restoration, the left eye image and the right eye image are restored separately, and a global and local joint constraint background light estimation model is established to remove the scattering effect of the medium in the imaging environment and obtain the restored left eye image. With right eye image
[0009] Step 3) Target detection based on reconstruction residuals: The restored left and right images are processed separately, and the original images are learned by constructing a stacked denoising autoencoder network. According to the residual between the network reconstructed image and the original image, the target positioning prediction map of the left and right images is obtained. l With O r ;
[0010] Step 4) Target size measurement based on contour matching: The contours of the target in the left and right images are extracted respectively, and the shape context feature vector and the local appearance feature vector are combined to construct a feature matching descriptor of the contour points. The two sets of contours are stereo matched by minimizing the matching cost. Combined with the calibrated internal and external parameters, the contours are 3D reconstructed and the key dimensions are measured.
[0011] Furthermore, the step 2) specifically includes the following steps:
[0012] 21) Input original image I s ,s∈{l,r};
[0013] 22) Determine I s The best global background light color value is obtained by setting a set of background light candidate coordinates and a set of fog lines in the RGB space, fitting the image intensity value through the candidate values of the background light and fog lines, and accurately estimating the global background light by finding the fitting parameters closest to the intensity value of the image itself, thus determining the best global background light color value. in, Respectively represent the red channel value, green channel value and blue channel value of the global background light;
[0014] 23) Using contrast perception to adaptively determine I s The local background light estimation image block size is used to estimate the local background light in, Respectively represent the red channel value, green channel value and blue channel value of the local background light;
[0015] 24) Combine the global background light color value and the local background light to obtain the background light optimization map; The ratio between channels is the basis for the local background light value. Make corrections to obtain the background light optimization map with global and local joint constraints in, They represent the red channel value, green channel value, and blue channel value of the local background light in the background light optimization image respectively; the specific calculation method is:
[0016] Among them, avg(·) means averaging the values in the matrix;
[0017] The Gaussian kernel function is used to filter the local background light image to obtain a smooth local background light image. in, They represent the red channel value, green channel value, and blue channel value of the local background light in the smoothed local background light image respectively;
[0018] 25) Use the background light optimization map and the fog line model to calculate the transmittance Remove the scattering effect of the medium in the imaging environment to obtain the restored image
[0019] Furthermore, the specific steps of local background light estimation in step 23) are:
[0020] 231) Use contrast coding CCI to detect image content: Among them, Ω i (x)∈I s represents a square image block with x as the center coordinate and size (2i+1)×(2i+1), i∈{1,2,...,7}; σ represents the image block Ω i (x) Standard deviation of the three internal channel values;
[0021] 232) Adaptively determine the image block size S for local background light calculation based on the CCI value r : Among them, m p is the image block size adjustment coefficient;
[0022] 233) For each color channel c = {R, G, B}, the local background light The calculation method is: first find the image I s In the image block Ω c Minimum intensity map in That is the dark channel image; then calculate Middle image block ψ=2Ω c The maximum intensity value within is taken as the background light image, and its mathematical expression is Where x represents the coordinate in the background light image, y represents The spatial coordinates in , z represents the spatial coordinates in the input image I.
[0023] Further, in step 25), the transmittance is calculated And get the restored image The specific steps include:
[0024] 251) Assume that the initial transmittance value of pixel x is: in, H represents the total number of fog lines;
[0025] 252) Derived the lower bound of transmittance t LB (x): Add a lower bound constraint to the transmittance of each pixel to obtain the lower bound value of the transmittance of each pixel
[0026] 253) Establishing a function to optimize the regularization project By minimizing the objective function, the optimized transmittance map is obtained Where σ(x) is The standard deviation calculated for each fog line, N x represents the four neighboring pixels of pixel x, λ represents the constant parameter that weighs the data term and smoothness, and its value is 0.1;
[0027] 254) Accurately estimate the background light And the optimized transmittance map Substitute the bad imaging model to obtain the restored image without bad scattering blur
[0028] Furthermore, the target detection based on reconstruction residual in step 3) adopts an image reconstruction model based on a stacked denoising autoencoder; the model is used to learn and reconstruct the image blocks of the original image, detect the scale-invariant features in the reconstructed image, serve as the constraint condition for the self-training of the stacked denoising autoencoder network, and use the residual map between the reconstructed image and the original image as the target positioning result; the stacked denoising autoencoder network SDAE adopted has a symmetrical structure, consisting of 4 denoising autoencoders DAE, with 4 encoder layers and 4 decoder layers, which are respectively used to extract the features of the sample data and reconstruct the data according to the features; the parameter weights of each encoder and decoder are limited to the calculation of the current layer, and the encoding and decoding processes are performed layer by layer, that is, the output of the i-th layer will be further processed as the input of the i+1-th layer.
[0029] Furthermore, the step 3) specifically includes the following steps:
[0030] 31) The restored left eye image and right eye image Input image reconstruction model based on stacked denoising autoencoder;
[0031] 32) Randomly select several image blocks from the input image as network input, use the deep belief network DBN for training, use the generated network parameters as the initialization parameters of the SDAE network, and continuously update the weights and bias values by training each denoising autoencoder;
[0032] 33) Use the Gaussian difference model to extract the scale-invariant features of the original image and the reconstructed image respectively, calculate the difference between the scale-invariant features as the convergence cost, and complete the training process of the SDAE network by minimizing the reconstruction cost;
[0033] 34) Calculate the residual between the reconstructed image and the original image as the target positioning area, use the trained SDAE network parameters to perform sparse reconstruction on the input original image, and generate the target positioning prediction map in the scene in, Represents the reconstructed image.
[0034] Furthermore, the initialization method of the SDAE network parameters in step 32) is as follows: the network node parameters are initialized by randomly selecting values from Gaussian distribution data with a mean of 0 and a variance of 0.1; assuming that the sample X is defined as X v =(v1,v2,...,v m ), after encoding by the restricted Boltzmann machine RBM, we get the n-dimensional sample Y h =(h1,h2,...,h n ), the encoding process follows the following rules:
[0035] a. For a given X v =(v1,v2,...,v m ), h i Represents a certain network node, then the value of the i-th feature of the encoded sample, that is, the probability that the i-th element of the hidden layer is 1 is Among them, v is the set of visual layer training samples, c j is the bias of the hidden node, h i ={0,1} is the encoded sample Y h The elements in are used to control the connection weight between visible node i and hidden node j;
[0036] b. Assume b i is the deviation of the visible node, then the probability that the i-th element in the reverse reconstructed visible unit takes the value of 1 is expressed as: Among them, v′ iis the element in the sample X after reverse reconstruction;
[0037] c. Update the connection weight w according to the following rules ij , the bias c of the hidden node j Deviation b from visible nodes i :
[0038] Where h′ i is the reverse reconstructed sample Y h Elements in .
[0039] Furthermore, the training process of the SDAE network in step 33) is specifically as follows:
[0040] 331) Assuming that X represents the input vector, Y represents the hidden layer vector, and Z represents the output layer vector, the training process of DAE is as follows: through a random mapping converter Add random noise to the input data to obtain lossy input
[0041] 332) Lossy Input Map the encoder of DAE to the hidden layer and obtain the vector Y of the hidden layer: Where W is the connection weight matrix from the input layer to the hidden layer, p represents the offset vector of the hidden layer, and S(·) represents the nonlinear activation function. The decoding layer of DAE is used to map the vector Y of the hidden layer to the output layer, and the output vector Z is: Z = f d (Y) = S(W′Y+p′), where W′ represents the connection weight matrix from the hidden layer to the output layer, p′ represents the offset vector of the output layer, and the weight matrix and offset vector of the decoder and encoder are transposed to each other, that is: W′ = W T 、p′=p T ;
[0042] 333) Construct a 4-octave pyramid, where each octave contains five or six intervals to obtain images in different scale spaces. The specific construction process is as follows:
[0043] Double the original image as the first layer of the first octave, perform Gaussian blur on the first layer image, and use the blurred image as the second layer image of the first octave. The Gaussian convolution function is defined as follows:
[0044] Where x and y represent the coordinates of each position in the convolution kernel, x0 and y0 are the means of x and y respectively, δ represents the standard deviation, and the parameter δ is selected as a fixed value of 1.6. The parameter δ is multiplied by the scale factor k to obtain the new Gaussian blur parameter δ. Gaussian blur is performed on the second layer image, and the result is set as the third layer. Repeat the above steps to obtain the image of the Lth layer; the last three images of the first octave layer are used as the initial images of the next octave, and they are downsampled to generate the first layer image of the second octave layer;
[0045] Following the above rules, all octave images are constructed in sequence to complete the modeling of the Gaussian pyramid;
[0046] Here, f g (·) represents the DOG Gaussian difference pyramid process, and N L Represents the maturity of the Gaussian pyramid, then the new loss function is constructed as follows:
[0047] 334) All parameters of the model are continuously adjusted by gradient descent method to obtain the minimum reconstruction error. The update rule is defined as:
[0048] Among them, η represents the learning rate in the update.
[0049] Furthermore, step 4) specifically includes the following steps:
[0050] 41) Input target positioning prediction map {O l ,O r} and its corresponding restored binocular image The images are downsampled and then the discrete cosine transform is used to remove the noise points in the high-frequency components of the images. Alpha-Shape is used to fit the boundaries of the separated discrete points of the target edge to obtain the contour lines of the target prediction area in the left and right images respectively.
[0051] 42) Uniformly sample the target contour lines in the two images to obtain the set of target contour points in the left and right images.
[0052] 43) Solve the shape context feature vector and local appearance feature vector of each discrete contour point in the left image and the right image in turn, and record them as: and Combine the two to construct a feature descriptor for each contour point:
[0053] Constructing shape context feature vector: taking any point in the contour point set With this point as the center, a polar coordinate system is established. The coordinate system is evenly divided into r distance layers and v angle layers. The radial length of each distance layer is The tangential angle of each angle layer is The polar coordinate system is divided into r×v regions, and the number of contour points falling in different regions of the polar coordinate system is counted. After normalization, the shape context feature vector S of the contour point is formed. sc , and the dimension of this vector is 1×(r×v);
[0054] Construct the local appearance feature vector of the edge point: refer to the SIFT feature description method to restore the binocular image Extract the pixel points corresponding to each contour point, build a scale-determined neighborhood with the edge point as the center, and divide the area into 4×4 sub-regions. Then, divide each sub-region into 8 directional regions from 0° to 360°, with each region being 45°. Subsequently, the gradient size of the neighborhood pixel points in each sub-region is weighted using a Gaussian weighting function. The gradient histogram in the 8 directions is statistically calculated and normalized to form a 4×4×8=128-dimensional gradient distribution feature vector S gd ;
[0055] 44) Construct the matching cost of the contour line C = (1-β)C S +βC A , where C S Used to describe the shape similarity of edge points: Indicates the dimension of the descriptor, the latter constraint Represents the similarity of the local appearance distribution characteristics of different points on the contour in the image. β = 0.1 is the weight coefficient of the two constraints, which is used to adjust the influence of the two constraints on the matching cost. After obtaining the matching cost of the contour point, the Hungarian matching strategy is used to obtain the matching result of the same-name point with the minimum matching cost. M′=min(M,N);
[0056] 45) After obtaining the set of point pairs with the same name in the binocular image, the three-dimensional spatial coordinates P of the contour point are solved according to the bad spatial point positioning algorithm based on the light refraction model of the binocular vision system. 3D ={P1,P2,…,P M′}, realize the 3D reconstruction of the 3D contour line of the target to be measured and calculate the key size information of the turbid and harsh target.
[0057] Furthermore, the specific steps for calculating the key size information of the turbid and bad target in step 45) are as follows:
[0058] 451) Use statistical filtering method to further eliminate point cloud P 3DOutliers generated by matching errors are removed to improve the accuracy of point cloud; the retained main point cloud The final three-dimensional contour point cloud of the target to be measured, the same-name point pairs corresponding to the removed outliers will also be removed, and the valid matching point pair set will be retained.
[0059] 452) Construct the minimum bounding rectangle of the target area to be measured in the left image, and calculate the coordinates of the midpoints on the four sides of the bounding rectangle in a clockwise direction
[0060] 453) From the filtered set of points with the same name Extract the effective contour points of the left image as candidate points to be measured;
[0061] 454) Calculate the effective contour points in sequence The midpoints of the four sides with the minimum circumscribed moment The coordinates of the four closest points are used as the coordinates of the key dimension measurement points
[0062] 455) From the contour line 3D point cloud collection Extract Corresponding three-dimensional coordinates Calculate the three-dimensional distance between two points by offsetting to obtain the key size information L{L1,L2} of the target:
[0063] Compared with the prior art, the present invention has the following beneficial effects:
[0064] (1) The background light estimation result obtained by the present invention based on the joint constraint of global background light and local background light can not only ensure the reliability of the average level of background light, but also retain the detailed changes of local background light, which is closer to the actual light distribution.
[0065] (2) The present invention can accurately predict the areas of various types of harsh targets. This is due to the new loss function constructed in the algorithm based on the scale-invariant property, which effectively retains the very critical contour information in the target detection results.
[0066] (3) The present invention can accurately measure the key dimensions of different targets in harsh environments, has excellent stability in harsh scenes, provides an effective solution to the problem of target size measurement in harsh environments, better adapts to environmental changes, and has higher application potential in practical engineering applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] FIG1 is a flow chart of a target measurement method in a harsh imaging environment based on salient contour matching according to the present invention;
[0068] Figure 2 is a flowchart of background light estimation with global and local joint constraints;
[0069] FIG3 is a block diagram of a target detection method in a harsh environment based on reconstruction of an invariant feature constrained autoencoder;
[0070] FIG4 is a schematic diagram of the structure of a Gaussian difference pyramid;
[0071] Figure 5 is a schematic diagram of the measurement process of a target in a harsh environment based on contour matching. DETAILED DESCRIPTION
[0072] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0073] As shown in FIG1 , the present invention proposes a target measurement method in a harsh imaging environment based on salient contour matching, comprising the following steps:
[0074] Step 1: Capture the target binocular image. Use Zhang’s calibration method to calibrate the binocular camera in advance to obtain the internal and external parameters of the camera. Use the calibrated binocular camera to simultaneously capture the target’s left eye image I in a harsh imaging environment. l With the right eye image I r .
[0075] Step 2: Restoration of harsh environment imaging based on local and global joint constraints. The left and right eye images are restored separately. A background light estimation model with global and local joint constraints is established to remove the scattering effect of the medium in the imaging environment and obtain the restored left eye image. With right eye image As shown in Figure 2.
[0076] 21) Input original image I s ,s∈{l,r}.
[0077] 22) Estimation of global background light based on fog line model to determine I s The optimal global background light color value is obtained to ensure the average accuracy of background light estimation; then, local background light details are distinguished based on contrast perception to avoid poor effects of restored images appearing too bright or faded due to inaccurate background light; finally, the global and local background lights are jointly constrained to obtain the optimized background light estimation map.
[0078] A clear image without fog can be represented by hundreds of different RGB color clusters. When the image is affected by fog scattering, the pixels in the same color cluster are disturbed by fog elements of different concentrations at different depths, which will shift in the RGB color space and form a line passing through the background light. Similar pixels at different depths are selected and marked with lines of different colors. The pixels selected from each group of images are mapped to the RGB three-dimensional color space. Each group of pixels selected from the harsh image shows a similar linear distribution, indicating that the fog line model is also valid in harsh conditions. Accurately estimate the global background light based on the fog line model. Set a set of background light candidate coordinates and a set of fog lines in the RGB space, fit the image intensity value through the candidate values of the background light and the fog line, and achieve accurate estimation of the global background light by finding the fitting parameters closest to the intensity value of the image itself, and determine the optimal global background light color value. in, They represent the red channel value, green channel value, and blue channel value of the global background light respectively.
[0079] 23) In view of the characteristics of imaging in harsh environments, based on determining the global background light, contrast perception is used to further optimize the local details of the background light and estimate the local background light. in, Respectively represent the red channel value, green channel value and blue channel value of the local background light. The specific steps are as follows:
[0080] a. Use contrast perception to adaptively determine the image block size for local background light estimation. For simple areas, larger image blocks should be used to enhance the dehazing effect. For high-contrast areas with complex content (i.e., target areas), smaller image blocks should be used to estimate background light. To adaptively obtain the image block size for local background light calculation, contrast coding (CCI) is used to detect image content:
[0081] Among them, Ω i (x)∈I represents a square image block with x as the center coordinate and size (2i+1)×(2i+1), i∈{1,2,...,7}. σ represents the image block Ω i (x) Standard deviation of the three internal channel values.
[0082] b. CCI is composed of i, which represents the size of the image block Ω with the smallest standard deviation among the image blocks centered on pixel x. The CCI value reflects the contrast of the image: high-contrast areas have low CCI values, while low-contrast areas have high CCI values. Therefore, based on the CCI value, it can be inferred that the area belongs to the target area and the image block size for local background light calculation is adaptively determined:
[0083] Among them, m p is the image block size adjustment coefficient.
[0084] c. Use contrast perception to adaptively obtain the local estimated image block size and perform local background light For each color channel c = {R, G, B}, the local background light is calculated as:
[0085] First, find the image I c In the image block Ω c Minimum intensity map in (ie dark channel image), and then calculate Middle image block ψ=2Ω c The maximum intensity value within is taken as the background light image, and its mathematical expression is as follows:
[0086] Where x represents the coordinate in the background light image, y represents The spatial coordinates in the foggy image I, z represents the spatial coordinates in the image block Ω c Calculated from the image patch size calculated from the local background light.
[0087] 24) The global background light vector can effectively represent the uniform background light value of most areas in the harsh image, but it cannot reflect the changes in the illumination details of the background area. The local background light has a stronger adaptability to the diversity of illumination, but is easily affected by the abnormal illumination area, resulting in an overall deviation in the background light estimation. Therefore, the global background light vector is used to represent the uniform background light value of most areas in the harsh image, but it cannot reflect the changes in the illumination details of the background area. The local background light vector has a stronger adaptability to the diversity of illumination, but is easily affected by the abnormal illumination area, resulting in an overall deviation in the background light estimation. The ratio between channels is the basis for the local background light value. Make corrections to obtain the background light optimization map with global and local joint constraints The specific calculation method is as follows.
[0088] Here, avg(·) means averaging the values in the matrix.
[0089] The background light estimation result obtained by jointly constraining the global background light and the local background light can ensure the reliability of the average level of background light while retaining the details of the local background light, which is closer to the actual light distribution. Finally, the Gaussian kernel function is used to filter the local background light image to make the acquired background light image smoother, thereby reducing the sudden and discontinuous halo phenomenon and obtaining a smooth local background light image. in, Respectively represent the red channel value, green channel value, and blue channel value of the local background light.
[0090] 25) Use the background light optimization map and the fog line model to calculate the transmittance Remove the scattering effect of the medium in the imaging environment to obtain the restored image
[0091] a. Define r(x)=||I in the fog line model B (x)-B ∞ (x)|| is used to represent the distance between the pixel and the background light origin. According to the harsh imaging model, r(x) can also be expressed by the transmittance t(x) as: r(x) = t(x)||J(x)-B ∞ ||, t(x)∈[0,1]. When the transmittance value reaches the maximum, the maximum value of r(x) can be obtained, that is: r max (x)=||J(x)-B ∞ ||. Then the transmittance can be expressed according to the above conditions as: Since the fog line may contain fog-free pixels, the maximum distance r in the fog line H is max (x) needs to be further amended to read:
[0092] Then the initial transmittance value of pixel x can be expressed as:
[0093] in, H represents the total number of fog lines.
[0094] b. Since the fog line contains different numbers of fog-free pixels, there are noise and estimation errors in the initial transmittance. In order to improve the accuracy of the transmittance estimation and ensure the edge consistency between the restored image and the original image, the regularization method is used to optimize the initial transmittance. Since the amplitude of the fog-free image must be no less than 0 (i.e., J ≥ 0), the lower bound of the transmittance t can be derived. LB (x):
[0095] Add a lower bound constraint to the transmittance of each pixel to obtain the lower bound value of the transmittance of each pixel
[0096] c. The transmittance corresponding to the continuous depth region should also be continuous, so the transmittance map should be smooth in the non-boundary region. Therefore, the following optimization regularization objective function is established, and the optimized transmittance map is obtained by minimizing the objective function.
[0097] In the formula, the first term is the data term, and σ(x) in the data term is The standard deviation calculated for each fog line is used to avoid large estimation deviations in transmittance. The latter term is the smoothing term, and N in the smoothing term is x Represents the four neighboring pixels of pixel x. By minimizing the differences between adjacent image blocks, it eliminates interference from fine edges in the image, better preserves the main edges of the image, and improves the smoothness of the transmittance map. λ represents a parameter that balances data terms and smoothness. In this invention, λ is set to 0.1.
[0098] d. Accurately estimate the background light With the optimized transmittance Substitute the bad imaging model to obtain the restored image without bad scattering blur Its mathematical expression is as follows.
[0099] Step 3: Target detection based on reconstruction residuals. The restored left and right images are processed separately, and the original images are learned by building a stacked denoising autoencoder network. According to the residual between the network reconstructed image and the original image, the target positioning prediction map of the left and right images is obtained. l With O r .
[0100] As shown in Figure 3, a stacked denoising autoencoder (SDAE) image reconstruction model is used to learn and reconstruct image blocks from the original image. Scale-invariant features in the reconstructed image are detected and used as constraints for the self-training of the stacked denoising autoencoder network. The residual image between the reconstructed image and the original image is used as the target detection result. The stacked denoising autoencoder network (SDAE) has a symmetrical structure and consists of four denoising autoencoders (DAEs). Four encoder layers and four decoder layers are used to extract features from the sample data and reconstruct the data based on these features. The parameter weights of each encoder and decoder are limited to the calculation of the current layer, and the encoding and decoding process is performed layer by layer. That is, the output of layer i is used as the input of layer i+1 for further processing.
[0101] 31) The restored left eye image and right eye image Enter the image reconstruction model based on stacked denoising autoencoders.
[0102] 32) Randomly selected image patches from the original image are used as network inputs for training using a deep belief network (DBN). The generated network parameters are used as initialization parameters for the SDAE network. The weights and bias values are continuously updated through training of each denoising autoencoder. A Difference of Gaussian model is used to extract scale-invariant features from the original and reconstructed images, respectively. The difference between the scale-invariant features is calculated as the convergence cost, and the SDAE network training process is completed by minimizing the reconstruction cost. Finally, the residual difference between the reconstructed and original images is calculated as the target region.
[0103] In the unsupervised training process, the network node parameters are initialized by randomly selecting values from Gaussian distribution data with mean 0 and variance 0.1. Assume that the sample X is defined as X v =(v1,v2,...,v m ), after being encoded by the Boltzmann machine (RBM), we get the n-dimensional sample Y h =(h1,h2,...,h n ), the encoding process follows the following rules:
[0104] a. For a given X v =(v1,v2,...,v m ), h i Represents a certain network node, then the probability that the value of the i-th feature of the encoded sample (i.e., the i-th element of the hidden layer) is 1 is:
[0105] Among them, v is the set of visual layer training samples, c j is the bias of the hidden node, h i ={0,1} is the encoded sample Y h The elements in are used to control the connection weight w between visible node i and hidden node j ij .
[0106] b. Assume b i is the deviation of the visible node, then the probability that the i-th element in the reverse reconstructed visible unit takes the value 1 can be expressed as:
[0107] Among them, v′ i is the element in the sample X after reverse reconstruction.
[0108] c. Update the connection weight w according to the following rules ij , the bias c of the hidden node j Deviation b from visible nodes i :
[0109] Where h′ i is the reverse reconstructed sample Yh Elements in .
[0110] Since DBN and SDAE have the same network structure, the network parameters generated by DBN training will be directly transferred to the SDAE network as initialization parameters. In this invention, the number and size of input image blocks are set to 10,000 blocks and 7*7 pixels respectively. The specific settings of the DBN network and SDAE network are as follows:
[0111] The DBN network parameters are: the number of visual nodes is 147; the number of hidden nodes is L1=256, L2=128, L3=64, L4=32; the learning rate is 0.1; the regularization coefficient is 0.0002; the initial momentum is 0.5; the momentum after 5 iterations is 0.9; the number of iterations is 100;
[0112] In the SDAE network parameters, the number of visual nodes is 147; the number of hidden nodes is Encoder{L1=256, L2=128, L3=64, L4=32}, Decoder{L1=32, L2=64, L3=128, L4=256}; the learning rate is 0.01; the regularization coefficient is 0.0001; and the number of iterations is 100.
[0113] 33) The training process of the SDAE network is as follows:
[0114] a. Assuming X represents the input vector, Y represents the hidden layer vector, and Z represents the output layer vector, the training process of DAE is as follows: through a random mapping converter Add random noise to the input data to obtain lossy input
[0115] b. Lossy input Map the encoder of DAE to the hidden layer and obtain the vector Y of the hidden layer:
[0116] Where W is the connection weight matrix from the input layer to the hidden layer, p represents the offset vector of the hidden layer, and S(·) represents the nonlinear activation function. The decoding layer of DAE is used to map the hidden layer vector Y to the output layer, and the output vector Z is: Z = f d (Y) = S (W′Y + p′)
[0117] Where W′ represents the connection weight matrix from the hidden layer to the output layer, p′ represents the offset vector of the output layer, and the weight matrix and offset vector of the decoder and encoder are transposed to each other, that is: W′=W T 、p′=p T .
[0118] c. To make X and Z as similar as possible, the present invention constructs a Difference of Gaussian pyramid (DOG) to extract scale-invariant features between the original and reconstructed images. Based on the DOG between the original and reconstructed images, a new loss function L(W, p; X, Z) is designed. As shown in Figure 4, a four-octave pyramid is constructed, where each octave contains approximately five or six intervals, to obtain images at different scales. The specific construction steps are as follows:
[0119] Double the original image as the first layer of the first octave. Gaussian blur the first layer image and use the blurred image as the second layer image of the first octave. The function of Gaussian convolution is defined as follows:
[0120] Where x and y represent the coordinates of each location in the convolution kernel, and δ represents the standard deviation, which is fixed at 1.6. Multiply δ by the scaling factor k to obtain the new Gaussian blur parameter δ. Gaussian blur the second layer image, and set the result as the third layer. Repeat the above steps to obtain the Lth layer image. The last three images of the first octave layer are used as the initial images for the next octave, and these are downsampled to generate the first layer image of the second octave layer.
[0121] Following the above rules, all octave images are constructed in sequence to complete the modeling of the Gaussian pyramid.
[0122] Here, f g (·) represents the DOG Gaussian difference pyramid process, and N L Represents the maturity of the Gaussian pyramid, then the new loss function constructor is as follows:
[0123] Feedback and fine-tuning of the SDAE network through the loss function constructed by scale-invariant characteristics can significantly improve the edge detection effect of the target area and facilitate target detection and segmentation.
[0124] d. Finally, all parameters of the model are continuously adjusted by gradient descent to obtain the minimum reconstruction error. The update rule is defined as:
[0125] Among them, η represents the learning rate in the update.
[0126] The parameters with the lowest convergence cost between the original image and the reconstructed image are selected as the network parameters. When the DOG feature difference between the original image and the reconstructed image is the largest, the loss function constructed based on the scale-invariant feature will reach the minimum. Selecting the network parameters at this time for image reconstruction will effectively highlight the target boundary and facilitate accurate target detection.
[0127] 34) Use the trained SDAE network parameters to sparsely reconstruct the input original image, and use the residual image between the reconstructed image and the original image to generate the target positioning prediction map in the scene in, Represents the reconstructed image.
[0128] Step 4: Target size measurement based on contour matching. The contours of the target are extracted from the left and right images respectively. The shape context feature vector and the local appearance feature vector are combined to construct a feature matching descriptor for the contour points. The two sets of contours are stereo matched by minimizing the matching cost. Combined with the calibrated internal and external parameters, the contours are 3D reconstructed and key dimensions are measured, as shown in Figure 5.
[0129] 41) In harsh environments, target details are severely lost. The Robert operator has low computational complexity and high sensitivity to image details, making it effective for detecting target edges. Images captured in turbid media contain a large amount of scattered noise. Therefore, if edge detection is performed directly on the harsh original image, the target's outline points will be submerged in the background noise, making it difficult to separate.
[0130] Therefore, the input target positioning prediction map {O l ,O r} and its corresponding restored binocular image binocular image The images are downsampled respectively, and then discrete cosine transform is used to remove noise points in the high-frequency components of the image to achieve the separation of target edge points and background noise.
[0131] Alpha-Shape is an algorithm for characterizing the boundaries of point sets. It can be used to extract the boundary contours of any two-dimensional unordered point set and reconstruct a two-dimensional shape that contains all the unordered points. The basic idea is to set a circle with a radius of α in the two-dimensional plane and roll it around N unordered points. During the rolling process, when the ball can only contain two points, these two points are considered boundary points. By connecting the two points in sequence, the boundary line of the point set can be obtained. Therefore, Alpha-Shape is used to fit the boundaries of the discrete points on the edge of the target separated in the binocular image to obtain the contour line of the target area, thus achieving target area segmentation.
[0132] 42) Matching homonymous points is a key step in 3D information perception. However, in harsh environments, the scattering of light by a large number of suspended particles results in a significant loss of detailed information such as the target's surface texture, making it impossible to provide effective matching features for 3D dimension measurement. In this case, the contours of the harsh target are well preserved and exhibit significant similarity in the binocular image, providing effective matching features for the harsh target under test. Therefore, this paper proposes a target contour matching method to achieve target dimension measurement in harsh environments.
[0133] First, the contour line of the target is uniformly sampled to obtain the set of contour points of the target to be measured in the left and right images. Then, the shape context feature vector is combined with the local appearance feature vector to construct a feature descriptor for each contour point.
[0134] 43) Solve the shape context feature vector and local appearance feature vector of each discrete contour point in the left image and the right image in turn, and record them as: and The two are combined to construct a feature descriptor for each contour point.
[0135] The shape context feature vector can effectively express the shape features of a point set by counting the positions of any point in the point set and its surrounding vector boundary points. With this point as the center, a polar coordinate system is established. The coordinate system is evenly divided into r distance layers and v angle layers. The radial length of each distance layer is The tangential angle of each angle layer is The polar coordinate system is divided into r×v regions, and the number of contour points falling in different regions of the polar coordinate system is counted. After normalization, the shape context feature vector S of the contour point is formed. sc , and the dimension of this vector is 1×(r×v).
[0136] The shape context feature vector is essentially a statistical distribution density of global points around the target point from different directions and distances. Its purpose is to describe the shape characteristics of discrete points of a two-dimensional contour. It expresses the global information of the position distribution between contour points, but it cannot express the local appearance characteristics of edge points in the two-dimensional image space. In order to better obtain the characteristics of edge points, the present invention extracts the pixel points corresponding to each edge point in the grayscale image, and draws on the SIFT feature description method to construct the local appearance feature vector of the edge point according to the gradient distribution in the neighborhood. First, a neighborhood with a determined scale is constructed with the edge point as the center, and the area is divided into 4×4 sub-regions. Each sub-region is then evenly divided into 8 directional regions from 0° to 360°, and each region is 45°. Subsequently, the gradient size of the neighborhood pixel points in each sub-region is weighted using a Gaussian weighting function, and the gradient histograms in the 8 directions are statistically counted and normalized to form a 4×4×8=128-dimensional gradient distribution feature vector S gd .
[0137] 44) The purpose of contour matching in binocular images is to find the discrete point set P of the contour in the left image. L The i-th point The discrete point set P of the contour line in the right image R Mapping points in And the same-name points and Two conditions should be met: (1) and The descriptors between them should be similar; (2) and Based on the above criteria, the present invention transforms the matching of contour lines into a process of minimizing the minimum matching energy function.
[0138] Construct the matching cost of the contour line C = (1-β)C S +βC A Among them, the previous constraint C S ∈[0,1], used to describe the shape similarity of edge points, K S = r×v represents the dimension of the descriptor, then C S The χ between the discrete point set descriptors of the left and right image contours is 2 Distance, that is, the feature difference matrix between two point sets, is mathematically expressed as follows:
[0139] The next constraint C A It represents the similarity of the local appearance distribution characteristics of different points on the contour in the image. Its mathematical definition is as shown in formula (5.10), and its numerical range is [0,1].
[0140] Among them, β=0.1 is the weight coefficient of the two constraints, which is used to adjust the influence of the two constraints on the matching cost.
[0141] After obtaining the matching cost of the contour points, the Hungarian matching strategy is used to obtain the matching result of the same-name points with the minimum matching cost. M′=min(M,N).
[0142] 45) After obtaining the set of point pairs with the same name in the binocular image, the three-dimensional spatial coordinates P of the contour point are solved according to the bad spatial point positioning algorithm based on the light refraction model of the binocular vision system. 3D ={P1,P2,·…P M′}, realize the 3D reconstruction of the 3D contour line of the target to be measured and calculate the key size information of the turbid and harsh target. The specific steps are as follows:
[0143] a. Use statistical filtering method to further eliminate point cloud P 3D The outliers generated by the matching error are removed to improve the accuracy of the point cloud. The final three-dimensional contour point cloud of the target to be measured, the same-name point pairs corresponding to the removed outliers will also be removed, and the valid matching point pair set will be retained.
[0144] b. Construct the minimum bounding rectangle of the target area to be measured in the left image, and calculate the coordinates of the midpoints on the four sides of the bounding rectangle in a clockwise direction
[0145] c. From the filtered set of points with the same name Extract the effective contour points of the left image as candidate points to be measured.
[0146] d. Calculate the effective contour points in sequence The midpoints of the four sides with the minimum circumscribed moment The coordinates of the four closest points are used as the coordinates of the key dimension measurement points
[0147] e. From the contour line 3D point cloud collection Extract Corresponding three-dimensional coordinates The three-dimensional distance between two points is calculated by offset to obtain the key size information L{L1,L2} of the target.
[0148] The above method can effectively solve the problem of target measurement in harsh imaging environments.
[0149] A computer-readable storage medium storing one or more programs, wherein the one or more programs include instructions that, when executed by a computing device, cause the computing device to perform a target measurement method in a harsh imaging environment based on salient contour matching.
[0150] A computing device includes one or more processors, one or more memories, and one or more programs, wherein the one or more programs are stored in the one or more memories and configured to be executed by the one or more processors, and the one or more programs include instructions for executing a target measurement method in a harsh imaging environment based on salient contour matching.
[0151] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0152] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.
[0153] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0155] The above are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention are included in the scope of the claims of the present invention to be approved.
Claims
1. A target measurement method in a harsh imaging environment based on salient contour matching, characterized by: The method specifically comprises the following steps: Step 1) Acquire the target binocular image: Pre-calibrate the binocular camera using Zhang’s calibration method to obtain the camera’s internal and external parameters, and use the calibrated binocular camera to simultaneously acquire the target’s left eye image I in a harsh imaging environment. l With the right eye image I r ; Step 2) Based on the local and global joint constraint harsh environment imaging restoration, the left eye image and the right eye image are restored separately, and a global and local joint constraint background light estimation model is established to remove the scattering effect of the medium in the imaging environment and obtain the restored left eye image. With right eye image Step 3) Target detection based on reconstruction residuals: The restored left and right images are processed separately, and the original images are learned by constructing a stacked denoising autoencoder network. According to the residual between the network reconstructed image and the original image, the target positioning prediction map of the left and right images is obtained. l With O r ; Step 4) Target size measurement based on contour matching: The contours of the target in the left and right images are extracted respectively, and the shape context feature vector and the local appearance feature vector are combined to construct a feature matching descriptor of the contour points. The two sets of contours are stereo matched by minimizing the matching cost. Combined with the calibrated internal and external parameters, the contours are 3D reconstructed and the key dimensions are measured.
2. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 1, characterized in that: The step 2) specifically includes the following steps: 21) Input original image I s ,s∈{l,r}; 22) Determine I s The best global background light color value is obtained by setting a set of background light candidate coordinates and a set of fog lines in the RGB space, fitting the image intensity value through the candidate values of the background light and fog lines, and accurately estimating the global background light by finding the fitting parameters closest to the intensity value of the image itself, thus determining the best global background light color value. in, Respectively represent the red channel value, green channel value and blue channel value of the global background light; 23) Using contrast perception to adaptively determine I s The local background light estimation image block size is used to estimate the local background light in, Respectively represent the red channel value, green channel value and blue channel value of the local background light; 24) Combine the global background light color value and the local background light to obtain the background light optimization map; The ratio between channels is the basis for the local background light value. Make corrections to obtain the background light optimization map with global and local joint constraints in, They represent the red channel value, green channel value, and blue channel value of the local background light in the background light optimization image respectively; the specific calculation method is: Among them, avg(·) means averaging the values in the matrix; The Gaussian kernel function is used to filter the local background light image to obtain a smooth local background light image. in, They represent the red channel value, green channel value, and blue channel value of the local background light in the smoothed local background light image respectively; 25) Use the background light optimization map and the fog line model to calculate the transmittance Remove the scattering effect of the medium in the imaging environment to obtain the restored image 3. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 2, characterized in that: The specific steps of local background light estimation in step 23) are: 231) Use contrast coding CCI to detect image content: Among them, Ω i (x)∈I s represents a square image block with x as the center coordinate and size (2i+1)×(2i+1), i∈{1,2,...,7}; σ represents the image block Ω i (x) Standard deviation of the three internal channel values; 232) Adaptively determine the image block size S for local background light calculation based on the CCI value r : Among them, m p is the image block size adjustment coefficient; 233) For each color channel c = {R, G, B}, the local background light The calculation method is: first find the image I s In the image block Ω c Minimum intensity map in That is the dark channel image; then calculate Middle image block ψ=2Ω c The maximum intensity value within is taken as the background light image, and its mathematical expression is Where x represents the coordinate in the background light image, y represents The spatial coordinates in , z represents the spatial coordinates in the input image I.
4. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 2, characterized in that: Calculate the transmittance in step 25) And get the restored image The specific steps include: 251) Assume that the initial transmittance value of pixel x is: in, H represents the total number of fog lines; 252) Derived the lower bound of transmittance t LB (x): Add a lower bound constraint to the transmittance of each pixel to obtain the lower bound value of the transmittance of each pixel 253) Establishing a function to optimize the regularization project By minimizing the objective function, the optimized transmittance map is obtained Where σ(x) is The standard deviation calculated for each fog line, N x represents the four neighboring pixels of pixel x, λ represents the constant parameter that weighs the data term and smoothness, and its value is 0.1; 254) Accurately estimate the background light And the optimized transmittance map Substitute the bad imaging model to obtain the restored image without bad scattering blur 5. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 1, characterized in that: The target detection based on reconstruction residual in step 3) adopts an image reconstruction model based on a stacked denoising autoencoder; the model is used to learn and reconstruct the image blocks of the original image, detect the scale-invariant features in the reconstructed image, serve as the constraint condition for the self-training of the stacked denoising autoencoder network, and use the residual map between the reconstructed image and the original image as the target positioning result; the stacked denoising autoencoder network SDAE adopted has a symmetrical structure, consisting of 4 denoising autoencoders DAE, with 4 encoder layers and 4 decoder layers, which are respectively used to extract the features of the sample data and reconstruct the data according to the features; the parameter weights of each encoder and decoder are limited to the calculation of the current layer, and the encoding and decoding processes are performed layer by layer, that is, the output of the i-th layer will be further processed as the input of the i+1-th layer.
6. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 5, characterized in that: The step 3) specifically includes the following steps: 31) The restored left eye image and right eye image Input image reconstruction model based on stacked denoising autoencoder; 32) Randomly select several image blocks from the input image as network input, use the deep belief network DBN for training, use the generated network parameters as the initialization parameters of the SDAE network, and continuously update the weights and bias values by training each denoising autoencoder; 33) Use the Gaussian difference model to extract the scale-invariant features of the original image and the reconstructed image respectively, calculate the difference between the scale-invariant features as the convergence cost, and complete the training process of the SDAE network by minimizing the reconstruction cost; 34) Calculate the residual between the reconstructed image and the original image as the target positioning area, use the trained SDAE network parameters to perform sparse reconstruction on the input original image, and generate the target positioning prediction map in the scene in, Represents the reconstructed image.
7. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 6, characterized in that: The initialization method of the SDAE network parameters in step 32) is as follows: initialize the network node parameters by randomly selecting values from Gaussian distribution data with a mean of 0 and a variance of 0.1; assuming that the sample X is defined as X v =(v1,v2,...,v m ), after encoding by the restricted Boltzmann machine RBM, we get the n-dimensional sample Y h =(h1,h2,...,h n ), the encoding process follows the following rules: a. For a given X v =(v1,v2,...,v m ), h i Represents a certain network node, then the value of the i-th feature of the encoded sample, that is, the probability that the i-th element of the hidden layer is 1 is Among them, v is the set of visual layer training samples, c j is the bias of the hidden node, h i ={0,1} is the encoded sample Y h The elements in are used to control the connection weight between visible node i and hidden node j; b. Assume b i is the deviation of the visible node, then the probability that the i-th element in the reverse reconstructed visible unit takes the value of 1 is expressed as: Among them, v′ i is the element in the sample X after reverse reconstruction; c. Update the connection weight w according to the following rules ij , the bias c of the hidden node j Deviation b from visible nodes i : Where h′ i is the reverse reconstructed sample Y h Elements in .
8. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 6, characterized in that: The training process of the SDAE network in step 33) is specifically as follows: 331) Assuming that X represents the input vector, Y represents the hidden layer vector, and Z represents the output layer vector, the training process of DAE is as follows: through a random mapping converter Add random noise to the input data to obtain lossy input 332) Lossy Input Map the encoder of DAE to the hidden layer and obtain the vector Y of the hidden layer: Where W is the connection weight matrix from the input layer to the hidden layer, p represents the offset vector of the hidden layer, and S(·) represents the nonlinear activation function. The decoding layer of DAE is used to map the vector Y of the hidden layer to the output layer, and the output vector Z is: Z = f d (Y) = S(W′Y+p′), where W′ represents the connection weight matrix from the hidden layer to the output layer, p′ represents the offset vector of the output layer, and the weight matrix and offset vector of the decoder and encoder are transposed to each other, that is: W′ = W T 、p′=p T ; 333) Construct a 4-octave pyramid, where each octave contains five or six intervals to obtain images in different scale spaces. The specific construction process is as follows: Double the original image as the first layer of the first octave, perform Gaussian blur on the first layer image, and use the blurred image as the second layer image of the first octave. The Gaussian convolution function is defined as follows: Where x and y represent the coordinates of each position in the convolution kernel, x0 and y0 are the means of x and y respectively, δ represents the standard deviation, and the parameter δ is selected as a fixed value of 1.
6. The parameter δ is multiplied by the scale factor k to obtain the new Gaussian blur parameter δ. Gaussian blur is performed on the second layer image, and the result is set as the third layer. Repeat the above steps to obtain the image of the Lth layer; the last three images of the first octave layer are used as the initial images of the next octave, and they are downsampled to generate the first layer image of the second octave layer; Following the above rules, all octave images are constructed in sequence to complete the modeling of the Gaussian pyramid; Here, f g (·) represents the DOG Gaussian difference pyramid process, and N L Represents the maturity of the Gaussian pyramid, then the new loss function is constructed as follows: 334) All parameters of the model are continuously adjusted by gradient descent method to obtain the minimum reconstruction error. The update rule is defined as: Among them, η represents the learning rate in the update.
9. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 1, characterized in that: Step 4) specifically includes the following steps: 41) Input target positioning prediction map {O l ,O r } and its corresponding restored binocular image The images are downsampled and then the discrete cosine transform is used to remove the noise points in the high-frequency components of the images. Alpha-Shape is used to fit the boundaries of the separated discrete points of the target edge to obtain the contour lines of the target prediction area in the left and right images respectively. 42) Uniformly sample the target contour lines in the two images to obtain the set of target contour points in the left and right images. 43) Solve the shape context feature vector and local appearance feature vector of each discrete contour point in the left image and the right image in turn, and record them as: and Combine the two to construct a feature descriptor for each contour point: Constructing shape context feature vector: taking any point in the contour point set With this point as the center, a polar coordinate system is established. The coordinate system is evenly divided into r distance layers and v angle layers. The radial length of each distance layer is The tangential angle of each angle layer is The polar coordinate system is divided into r×v regions, and the number of contour points falling in different regions of the polar coordinate system is counted. After normalization, the shape context feature vector S of the contour point is formed. sc , and the dimension of this vector is 1×(r×v); Construct the local appearance feature vector of the edge point: refer to the SIFT feature description method to restore the binocular image Extract the pixel points corresponding to each contour point, build a scale-determined neighborhood with the edge point as the center, and divide the area into 4×4 sub-regions. Then, divide each sub-region into 8 directional regions from 0° to 360°, with each region being 45°. Subsequently, the gradient size of the neighborhood pixel points in each sub-region is weighted using a Gaussian weighting function. The gradient histogram in the 8 directions is statistically calculated and normalized to form a 4×4×8=128-dimensional gradient distribution feature vector S gd ; 44) Construct the matching cost of the contour line C = (1-β)C S +βC A , where C S Used to describe the shape similarity of edge points: K S = r×v represents the dimension of the descriptor, and the latter constraint Represents the similarity of the local appearance distribution characteristics of different points on the contour in the image. β = 0.1 is the weight coefficient of the two constraints, which is used to adjust the influence of the two constraints on the matching cost. After obtaining the matching cost of the contour point, the Hungarian matching strategy is used to obtain the matching result of the same-name point with the minimum matching cost. M′=min(M,N); 45) After obtaining the set of point pairs with the same name in the binocular image, the three-dimensional spatial coordinates P of the contour point are solved according to the bad spatial point positioning algorithm based on the light refraction model of the binocular vision system. 3D ={P1,P2,…,P M′ }, realize the 3D reconstruction of the 3D contour line of the target to be measured and calculate the key size information of the turbid and harsh target.
10. The target measurement method in a harsh imaging environment based on salient contour matching according to claim 9, characterized in that: The specific steps for calculating the key size information of the turbid and bad target in step 45) are as follows: 451) Use statistical filtering method to further eliminate point cloud P 3D Outliers generated by matching errors are removed to improve the accuracy of point cloud; the retained main point cloud The final three-dimensional contour point cloud of the target to be measured, the same-name point pairs corresponding to the removed outliers will also be removed, and the valid matching point pair set will be retained. 452) Construct the minimum bounding rectangle of the target area to be measured in the left image, and calculate the coordinates of the midpoints on the four sides of the bounding rectangle in a clockwise direction 453) From the filtered set of points with the same name Extract the effective contour points of the left image as candidate points to be measured; 454) Calculate the effective contour points in sequence The midpoints of the four sides with the minimum circumscribed moment The coordinates of the four closest points are used as the coordinates of the key dimension measurement points 455) From the contour line 3D point cloud collection Extract Corresponding three-dimensional coordinates Calculate the three-dimensional distance between two points by offsetting to obtain the key size information L{L1,L2} of the target:
Citation Information
Patent Citations
A target object size measurement method of sparse point cloud
CN109903327A
Target size measurement method based on multi-view binocular visual perception
CN114255286A
Dam water seepage area measurement method based on binocular remote sensing image saliency analysis
CN116758026A
Target measurement method in severe imaging environment based on significant contour matching
CN118505598A
Monocular world meshing
US20230177708A1
Cited By
Ore body virtual sectioning and drilling data generation method and ore body contour line intelligent generation method
CN120997246A
Video synchronous acquisition method and device based on optical far image
CN121000964A
Image recognition method and system for screening excellent single plants of tea trees
CN121033681A
Weld joint parameter anti-interference measurement method
CN121112903A
Low-porosity anode carbon block porosity real-time estimation system and low-porosity anode carbon block finished product thereof
CN121113828A