Projectile pose measuring and calculating system based on multistage feature fusion neural network and visual three-dimensional model fitting

By using multi-stage feature fusion neural network and visual three-dimensional model fitting technology in the projectile detection system, the problems of insufficient detection accuracy and poor anti-interference ability of high-speed moving projectiles are solved, and high-precision projectile position measurement and stronger anti-interference performance are achieved.

CN119942522APending Publication Date: 2025-05-06NORTHWESTERN POLYTECHNICAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411964017.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art has problems of insufficient accuracy and poor anti-interference ability when detecting high-speed moving projectiles. Especially in the complex environment of muzzles, it is difficult for traditional methods to accurately measure the position of the projectile.

Method used

The projectile position estimation system based on multi-level feature fusion neural network and visual three-dimensional model fitting is adopted. The projectile emission image is obtained through the image acquisition device, and combined with the target detection module and the attitude estimation module, the precise measurement of the projectile position and attitude is achieved.

Benefits of technology

The detection accuracy is improved, and the projectile muzzle status parameters can be accurately measured under the high-speed movement of the projectile and the muzzle smoke blocking. The measurement error is less than 5%, the measuring range is less than 10°, and it has stronger anti-interference performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942522A_ABST
    Figure CN119942522A_ABST
Patent Text Reader

Abstract

The invention discloses a projectile pose measuring and calculating system based on a multilevel feature fusion neural network and visual three-dimensional model fitting, which captures a high-speed projectile flight image on the basis of a deep learning technology, realizes the detection of a target image through a target detection technology based on the multilevel feature fusion of the neural network, and realizes the measurement and calculation of the projectile pose. And on the basis, a posture reconstruction resolving technology based on monocular vision and a projectile model is used for realizing the resolving work of the projectile state, so that the accurate measurement of the projectile muzzle state is realized, the test range is not less than 10 degrees, and the test error is less than 5%. According to the method, the detection precision is improved, the bullet muzzle state parameters can be accurately measured, solid data support is provided for artillery performance evaluation and improvement, and the method has higher anti-interference performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of target detection, and in particular relates to a projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting. Background Art

[0002] Traditional target detection and posture estimation algorithms often use template matching, geometric model fitting, etc., which currently show certain effectiveness in certain specific scenarios, but are already lagging behind in terms of accuracy and adaptability. Currently, common methods for target detection based on deep learning include the two-stage target detection algorithm R-CNN and the single-stage target detection algorithm YOLO, which have significantly improved accuracy and adaptability. In addition, there are ViT and SwinTransformer, which break through the CNN structure and integrate the Transformer architecture, but they all rely on sufficient data sets and may miss detection when detecting "smears" and deformations caused by high-speed movement. Currently, posture estimation algorithms based on deep learning include OpenPose and HRNet based on CNN structure, which can achieve relatively accurate key point detection, but they also rely on sufficient data sets, resulting in poor anti-interference ability of the algorithm to changes in environmental factors.

[0003] The velocity and posture of a projectile are key factors in determining its range and accuracy. With the increasing requirements for the first-shot hit rate and shooting accuracy of artillery, solving the posture of a projectile is extremely important for analyzing the subsequent projectile trajectory and landing point, as well as for the calibration of the launch device. Considering that the muzzle will vibrate violently when the artillery is fired, and strong electromagnetic interference will be generated around the barrel, laser, radar and other measurement methods may produce large errors. Conventional target detection methods focus on real-time performance and adopt a lightweight model strategy. When faced with image blur caused by the high-speed movement of the projectile, the projectile cannot be accurately located, resulting in serious damage to the detection accuracy and difficulty in meeting the high-precision measurement requirements. The traditional reference point-based posture solution method cannot effectively deal with smoke occlusion, high-speed movement of the projectile and changes in shooting perspective in the complex environment of the muzzle. The reference point is frequently lost, resulting in huge deviations in the posture solution results and poor reliability. In the past, model training required a large amount of labeled data, and the acquisition and labeling costs were extremely high. In addition, it was easy to cause model overfitting due to limited data, weakening the model's generalization ability and affecting the accuracy and stability of projectile detection and recognition. Summary of the invention

[0004] In order to overcome the shortcomings of the prior art, the present invention provides a projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting, which uses deep learning technology as a basis to capture high-speed projectile flight images, and detects target images through target detection technology based on multi-level feature fusion of neural network. On this basis, the posture reconstruction and solution technology based on monocular vision and projectile model is used to realize the solution of projectile state, aiming to achieve accurate measurement of projectile muzzle state, test range ≮10°, test error <5%. The present invention improves the detection accuracy, can accurately measure projectile muzzle state parameters, provides solid data support for artillery performance evaluation and improvement, and has stronger anti-interference performance.

[0005] The technical solution adopted by the present invention to solve the technical problem is as follows:

[0006] A projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting, including image acquisition equipment, target detection module and posture estimation module;

[0007] The image acquisition device is used to obtain projectile launch images, using a single or multiple cameras, and the hardware information provided by the image acquisition device itself is used in the posture solution of the posture estimation module;

[0008] The target detection module adopts a target detection method based on multi-scale feature fusion. According to the large field of view image, the multi-scale target detection network is used to fuse the characteristics of multi-level semantic information to achieve small-size target detection of projectiles under a large field of view; finally, the projectile position information is obtained by solving the calculation to achieve the initial velocity detection of the projectile muzzle;

[0009] The posture estimation module uses a monocular vision posture estimation method based on a 3D projectile model to complete the measurement; constructs a 3D discretized model of the projectile, and at the same time constructs an imaging plane in three-dimensional space, performs projection calculation on the model, and obtains the model projection corresponding to the imaging plane; combines the neural network model to perform threshold segmentation processing on the actual object image to obtain the binary imaging plane contour of the actual object, and obtains the projection with the highest similarity through scaling and comparison calculation of the contour and the projection of the three-dimensional space projectile model, and combines the projection matrix with the scaling parameters to obtain the specific posture parameters of the object.

[0010] Preferably, the target detection method based on multi-scale feature fusion is specifically as follows:

[0011] Step 1-1: Use feature pyramid and path aggregation network to effectively fuse features of different scales through top-down and bottom-up paths, as well as lateral connections and feature aggregation, and integrate high-level semantic information into the feature map while retaining high-resolution information at the bottom level;

[0012] Step 1-2: Use motion region detection and motion blur recovery algorithm to deblur the motion image; add noise factors to the transfer function of the Wiener filter model to perform blur recovery; identify the hierarchical quality based on the recovery, and the determination of the hierarchical quality is determined by the detection and tracking results; the accuracy of the detection and tracking results is calculated using the vertices of the regression box and the real labels in the training data, and the quality judgment loss function is shown in the following formula:

[0013]

[0014] Among them, α is the weight parameter, (x i ,y i ) is the vertex coordinate of the regression box, (x g ,y g ) are vertex coordinates of the true value;

[0015] Step 1-3: After completing the image restoration, use the target detection method based on multi-scale feature fusion to complete the projectile detection. The steps for constructing the target detection network model based on multi-scale feature fusion are as follows:

[0016] ① Design the backbone network as a lightweight network;

[0017] Depthwise separable convolution is used instead of standard convolution, and the convolution process is divided into two steps: channel-by-channel convolution and point-by-point 1×1 convolution. The unit of the lightweight network is a reverse residual structure, and the reverse residual block structure is wide in the middle and narrow on both sides: 1×1 convolution is used for dimensionality increase, 3×3 depthwise separable convolution is used for feature extraction, and finally 1×1 convolution is used for dimensionality reduction. The lightweight network contains 5 convolution modules composed of reverse residual blocks, and the output feature maps of each convolution module are shown in Figure 2. Indicated as C1, C2, C3, C4, C5; set the input image size of the network to 416×416, then the dimensions of the feature maps C1 to C5 are: 208×208×16, 104×104×24, 52×52×32, 26×26×96, 13×13×320, respectively. Compared with the original image, the size of the C1 layer is reduced by 2 times, the size of the C2 layer is reduced by 4 times, the size of the C3 layer is reduced by 8 times, the size of the C4 layer is reduced by 16 times, and the size of the C5 layer is reduced by 32 times;

[0018] ② Perform spatial pyramid pooling on the feature map obtained in step ① to increase the receptive field of the network;

[0019] Use three pooling layers of different scales: 5×5, 9×9, and 13×13. The outputs of the three pooling layers are channel-joined, and then a 1×1 convolution is performed to reduce the feature dimension. The feature map is denoted as C6.

[0020] ③ Feature fusion based on C2, C3, C4, and C6 network layers;

[0021] C2, C3, C4, and C6 are convolved with a kernel of 1×1 respectively; C6 is upsampled and then concatenated with the C4 layer after the convolution operation, and the resulting feature map is denoted as P1; P1 is upsampled again, C2 is downsampled, and concatenated with the C3 layer after the convolution operation, and the resulting feature map is denoted as P2, which is then sent to detection head 1 after the convolution operation; P2 is downsampled, and concatenated with P1 to obtain P3, which is sent to detection head 2 after the convolution operation; finally, P3 is downsampled, concatenated with C6, and the result is sent to detection head 3 after the convolution operation;

[0022] ④The target detection network uses a focal loss function, including the bounding box loss CIOU2 , category loss loss cls2 And confidence loss loss confidence2 , the formula is as follows:

[0023] loss=loss CIOU2 +loss cls2 +loss confidence2

[0024] The bounding box loss is:

[0025]

[0026] Where d is the Euclidean distance between the center points of the two bounding boxes, c is the diagonal distance of the union of the two bounding boxes, and υ is a parameter to measure the consistency of the aspect ratio, which is calculated as follows:

[0027]

[0028] Where:

[0029] w gt ——the width of the object’s ground-truth bounding box;

[0030] h gt ——The height of the target's true bounding box;

[0031] w——the width of the prediction box;

[0032] h – the height of the prediction box;

[0033] α is the weight parameter, calculated as follows:

[0034]

[0035] Category loss cls2 The coefficient is introduced to adjust the weight of difficult and easy samples, which is defined as follows:

[0036]

[0037] Among them, s 2 is the number of grids in the image, B is the number of prediction boxes, and I i,j no _ obj is an indicator function that takes a value of 1 when there is no target in the current prediction box, α t is the weight parameter, γ is the adjustment factor; p is the probability value predicted by the model, indicating the probability that a grid cell does not contain the target object.

[0038] Preferably, the posture estimation module is specifically:

[0039] The image segmentation model in the posture estimation module uses a V-Net-like architecture, adopts an end-to-end training method, and uses nonlinear transformation and histogram matching to perform data enhancement;

[0040] According to the obtained projectile edge contour, for the projectile pose estimation task of a single-frame moment image, a monocular vision pose estimation method based on a 3D projectile model is used to complete the measurement of the projectile pose of the projectile out of the chamber; combined with the projectile 3D model, an imaging plane is constructed in three-dimensional space, and the model is projected and calculated by simulating the camera imaging principle to obtain the projectile model projection corresponding to the imaging plane. The specific principle is as follows:

[0041] According to the image of the rigid body target captured by the camera, in the camera coordinate system O c -X c Y c Z c In the world coordinate system O, the rotation matrix and translation vector are used to represent the position and posture of the target; w -X w Y w Z w In the 3D coordinate system, the target's motion is represented by the product of the rotation matrix and the coordinate vector, which is equivalent to the re-expression of the point position in another coordinate system. In the 3D coordinate system, rotation is divided into rotations around the X, Y, and Z axes. The rotation matrix R has a characteristic that the inverse of the rotation matrix R is its own transposed matrix. The translation vector T is the offset between the relative positions of the target after movement.

[0042] Suppose the coordinate of a point in space in the world coordinate system is P, and the coordinate of this point mapped to the camera coordinate system by the shooting camera is P', then:

[0043] P'=R(PT)

[0044] The rotation matrix R is represented by a 3×3 matrix, and the translation vector T is represented by a 1×3 three-dimensional vector; the point in the world coordinate system is first rotated and translated to the camera coordinate system, then mapped to the image physical coordinate system o-xyz through the intrinsic parameter, and finally transferred to the image pixel coordinate system o-uvz through the distance and pixel ratio;

[0045]

[0046] In the formula, the point (u, v) is the point (X w ,Y w ,Z w ) is the projected pixel coordinate in the image, Z c Represents a point in the world coordinate system (X w ,Y w ,Z w ) is the Z-axis value in the camera coordinate system; k and l represent the actual length of each pixel unit in the x-axis and y-axis respectively, in millimeters, f represents the focal length of the camera, and t represents its translation vector; f x Represents the focal length of the camera on the x-axis, f y represents the focal length of the camera on the y-axis, (u0, v0) is the principal point of the image pixel coordinate system; where u0, v0, f x and f y These are the four internal parameters determined by the camera resolution and focal length. The data in the rotation angle and translation vector corresponding to the rotation matrix are the external parameters related to the camera's world position.

[0047] According to the above principle, the three-dimensional model of the projectile is projected on the specified simulation imaging plane to obtain the corresponding simulation imaging projection with the standard posture, that is, each posture angle is 0. The projection is fitted using the model constructed by the characteristic triangle method to obtain the characteristic triangle of the imaging projection of the projectile model under the standard posture;

[0048] Based on the principle of the characteristic triangle method to obtain the projectile's posture, according to the geometric characteristics of the projectile, the initial state is that the projectile's spatial characteristic triangle is on the horizontal plane, the triangle ΔB0C0D0 is an isosceles triangle, OD0 is the perpendicular bisector of ΔB0C0D0, pointing to the top of the imaging plane as the X-axis; the Y-axis is the direction of the plumb line, and inward is positive; the Z-axis coincides with B0C0, and forms a right-hand system with the X-axis and the Y-axis. The triangle ΔBCD represents the actual position of the projectile in space, and the triangle ΔB'C'D' is the projection of ΔBCD on the horizontal plane. In the horizontal plane, α represents the yaw angle, β represents the pitch angle, and γ represents the roll angle. The projectile is a rotating body, and the roll angle defaults to 0;

[0049] Yaw angle:

[0050] Pitch Angle:

[0051] The image of the bullet detection result is subjected to threshold segmentation to obtain the binary image plane contour of the actual bullet. The edge detection is used to extract the target contour features. After the image is filtered, the edge detection operator is used to find the pixel points with grayscale mutation in the image.

[0052] The contour is fitted by the characteristic triangle method, and the characteristic triangle of the physical imaging contour is used to scale and compare the characteristic triangle of the projection of the three-dimensional standard posture projectile model to obtain the projection matrix and scaling parameters between the two, that is, the estimated posture parameters of the projectile;

[0053] The posture solution technology based on optimal consistency is adopted to further optimize the accuracy of the basic posture solution; the characteristic point of the projectile model is D(x wd ,y wd ,z wd ),C(x wc ,y wc ,z wc ),B(x wb ,y wb ,z wb ); According to the relationship between the world coordinate system and the camera coordinate system, the relationship can be obtained:

[0054]

[0055] Where:

[0056] T wc : The coordinates of the origin of the world coordinate system in the camera coordinate system

[0057] R wc : Rotation matrix

[0058] At the same time, if the camera focal length is f, then the image coordinate system (x i ,y i ) and the camera coordinate system is:

[0059]

[0060] Where: (x i ,y i ) The distance from the image coordinate system to the target point can be obtained by intersection, i = B, C, D; let the three sides of the triangle be represented by respectively; the actual lengths of the three sides are:

[0061]

[0062] Construct the optimal function:

[0063] φ(S)=Min((L DB -L1) 2 +(L BC -L2) 2 +(L DC -L3) 2 )

[0064] At the initial value S 0Performing first-order Taylor expansion everywhere, we get the linear equation:

[0065] φ(S)=φ(S 0 )+Bδ S

[0066] Where: B is the optimal function at the initial value S 0 The first-order partial derivative at δ m is the limit error correction number; each item in S is corrected accordingly:

[0067]

[0068] Solution steps:

[0069] (1) Extract feature point image coordinates; set the limit error δ min , the initial value of the attitude is obtained according to the characteristic triangle method.

[0070] (2) Using the least squares solution, we can obtain the correction number δ for each parameter: s .

[0071] (3) If δ s Greater than δ min , repeat 2) and 3) to continue iterating, otherwise end.

[0072] The beneficial effects of the present invention are as follows:

[0073] The target detection algorithm of the present invention is based on the multi-level feature fusion of neural network: the test samples are generated by using simulation data and real bullet images, and the projectiles are stably detected with an average accuracy of 96%. The projectile posture solution method based on the 3D projectile model is studied: the test samples are generated by using the static real bullet model image combined with the measurement data, and the projectile 3D model reconstruction and state parameter solution are completed (the lower side in the figure below is the depth distribution effect of the 3D model), with a maximum range of 15 degrees (pitch angle) and 12 degrees (yaw angle), and the average measurement error is 4.7%.

[0074] Compared with other methods, the present invention improves the detection accuracy. When the projectile moves at high speed and is blocked by muzzle smoke, the projectile muzzle state parameters can be accurately measured. The projectile velocity and axial swing angle and other attitude angle measurement errors are less than 5%, and the range is ≮10°, providing solid data support for artillery performance evaluation and improvement. At the same time, the present invention has stronger anti-interference performance under the same conditions. The attitude solution technology based on the 3D projectile model successfully avoids the problem of reference point loss, effectively copes with complex working conditions, ensures that the attitude solution results are stable and reliable, and greatly improves the anti-interference ability of the measurement system. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] Figure 1 It is a schematic diagram of the system of the present invention;

[0076] Figure 2 It is a structural diagram of the detection network model;

[0077] Figure 3 It is a feature pyramid network structure diagram;

[0078] Figure 4 It is a path aggregation network structure diagram;

[0079] Figure 5 It is the structure diagram of the spatial pyramid pooling module;

[0080] Figure 6 It is the bottleneck layer structure diagram;

[0081] Figure 7 It is the CSP module structure diagram;

[0082] Figure 8 It is a diagram of the segmentation network model structure;

[0083] Fig. 9 It is a schematic diagram of the camera coordinate system and the world coordinate system;

[0084] Fig.10 It is a schematic diagram of the feature fitting model;

[0085] Fig.11 is the test result graph;

[0086] Fig.12 It is the depth distribution map of the 3D model;

[0087] Fig.13 This is the effect diagram of posture estimation. DETAILED DESCRIPTION

[0088] The present invention is further described below in conjunction with the accompanying drawings and embodiments.

[0089] Based on detection, the present invention performs 3D posture solution of the projectile target based on depth analysis, eliminates environmental interference, and obtains accurate projectile posture information, aiming to improve the accuracy of projectile detection and posture estimation, accurately capture high-speed moving projectile targets, and overcome the interference of complex environmental factors during the detection process.

[0090] In order to achieve the above object, the present invention adopts the following technical solutions:

[0091] The present invention provides a projectile posture measurement system based on a multi-level feature fusion neural network and a visual three-dimensional model fitting, comprising the following steps:

[0092] (1) Deploy image acquisition equipment in the mission area to collect images;

[0093] (2) The acquired image is passed to the target detection module to detect the projectile target, and the sub-image acquired by the detection is passed to the posture estimation module.

[0094] (3) The posture estimation module obtains image depth information based on the detected image, performs 3D posture calculation on the projectile target, and obtains the axial swing angle of the projectile target.

[0095] Figure 1 A schematic diagram of a projectile posture measurement system based on a multi-level feature fusion neural network and visual three-dimensional model fitting provided by the present invention.

[0096] The system includes image acquisition equipment, target detection module, and posture estimation module. The modules are as follows:

[0097] 1. Image acquisition equipment.

[0098] The image acquisition device is used to obtain the projectile launch image, using a single or multiple cameras. The hardware information provided by the image acquisition device itself can be used for the attitude solution of the attitude estimation module.

[0099] 2. Object detection module.

[0100] The target detection module adopts a target detection method based on multi-scale feature fusion. Based on the large field of view image, the multi-scale target detection network is used to fuse the characteristics of multi-level semantic information to achieve small-size target detection of projectiles under a large field of view. Finally, the projectile position information is obtained through calculation to achieve the detection of the initial velocity of the projectile muzzle.

[0101] The target detection method based on multi-scale feature fusion is a single-stage target detection model based on deep learning. It has the ability to fuse low-level high-resolution information and high-level semantic information, and has a better detection effect on small-sized targets in a large field of view. The target detection method based on multi-scale feature fusion adopts feature pyramid and path aggregation network. Through top-down and bottom-up paths, as well as lateral connections and feature aggregation, it effectively fuses features of different scales together, integrates high-level semantic information into the feature map, and helps to identify complex targets. At the same time, it retains the high-resolution information of the bottom level, which helps to locate and identify small targets, thereby improving the accuracy and robustness of small target detection under a large field of view.

[0102] At the same time, the optical imaging detection of high-speed projectiles mainly adopts visible light camera array technology, but the "smear" and blur problems that appear during the high-speed movement of the shooting device or the target are the biggest obstacles to target detection and recognition. To solve this problem, the present invention uses motion area detection and motion blur recovery algorithm to perform motion image deblurring.

[0103] Motion blur restoration is achieved based on the Wiener filter method. Conventional Wiener filter image restoration methods can obtain basically satisfactory results in situations where the requirements are not high. However, this method uses a constant to replace the power spectrum ratio of noise and image, and fails to make full use of the information of the image and noise itself, so it is difficult to obtain high-quality image restoration effects. In this project, noise factors are added to the transfer function of the Wiener filter model to perform blur restoration.

[0104] Being able to identify the quality of the hierarchy after restoration is a prerequisite for information enhancement and baseline hierarchy selection. From the perspective of detection and tracking, the judgment of hierarchy quality should be determined by the detection and tracking results. The accuracy of the detection and tracking results is mainly calculated using the vertices of the regression box and the true labels in the training data. A quality judgment loss function is proposed as shown in the following formula:

[0105]

[0106] On the basis of completing image restoration, the present invention uses a target detection method based on multi-scale feature fusion to complete projectile detection. This method has the common advantages of a single-stage detection feature model and can complete the object detection task in a single forward propagation. The main advantage of the single-stage detection method is its fast speed, because it does not require additional candidate region generation and classification steps. This method effectively improves the detection accuracy and speed by introducing a multi-scale feature fusion mechanism and an improved loss function. The target detection method based on multi-scale fusion adopts a multi-scale feature fusion mechanism, which improves the detection performance of small targets by fusing feature maps of different scales, while enhancing the positioning accuracy of large targets. The detection network model structure diagram is shown in the figure. Figure 2 shown.

[0107] The steps to build the detection model are as follows:

[0108] ① Design the backbone network as a lightweight network. Use depthwise separable convolution instead of standard convolution, and divide the convolution process into two steps: channel-by-channel convolution and point-by-point 1×1 convolution. Although the convolution process is expanded to two steps, the lightweight convolution method reduces computational redundancy, and the overall computational workload is about 1 / 9 of the standard volume, which improves the computing speed. The unit of the lightweight network is an inverse residual structure. The inverse residual block structure is wide in the middle and narrow on both sides: 1×1 convolution is used for dimensionality increase, 3×3 depthwise separable convolution is used for feature extraction, and finally 1×1 convolution is used for dimensionality reduction. The lightweight network contains 5 convolution modules composed of inverse residual blocks, and the output feature maps of each convolution module are represented as C1, C2, C3, C4, and C5 respectively. Set the input image size of the network to 416×416, then the dimensions of the C1 to C5 feature maps are 208×208×16, 104×104×24, 52×52×32, 26×26×96, and 13×13×320, respectively. Compared with the original image, the size of the C1 layer is reduced by 2 times, the size of the C2 layer is reduced by 4 times, the size of the C3 layer is reduced by 8 times, the size of the C4 layer is reduced by 16 times, and the size of the C5 layer is reduced by 32 times.

[0109] ② Perform spatial pyramid pooling on the feature map obtained in the first step to increase the receptive field of the network. Use three pooling layers of different scales: 5×5, 9×9, and 13×13. The outputs of the three pooling layers are concatenated through channels to obtain richer target features. Then, a 1×1 convolution is performed to reduce the dimension of the feature. The feature map is recorded as C6.

[0110] ③ In order to improve the detection capability of the lightweight detection network for objects of different scales, feature fusion is performed based on the C2, C3, C4, and C6 network layers. Since C1 is large in size, it is not considered here to save computational costs. First, C2, C3, C4, and C6 are convolved with a kernel of 1×1. This step changes the number of channels while ensuring that the area size of the feature map of each layer remains unchanged, so as to facilitate subsequent fusion operations. The feature pyramid network (FPN) enhances the semantic information of features from top to bottom, but ignores the enhancement of positioning information. A top-down and bottom-up double pyramid structure is used for feature fusion, and parameters of different detection layers are aggregated from different backbone layers, which will enhance the semantic information of high layers and the strong positioning information of low layers at the same time. C6 is upsampled and then concatenated with the C4 layer after the convolution operation. The resulting feature map is denoted as P1. P1 is upsampled again, C2 is downsampled, and concatenated with the C3 layer after the convolution operation. The resulting feature map is denoted as P2, and then sent to the detection head 1 after the convolution operation. P2 is downsampled and concatenated with P1 to obtain P3, which is sent to detection head 2 after convolution operation. Finally, P3 is downsampled and concatenated with C6, and the result is sent to detection head 3 after convolution operation.

[0111] ④ To solve the problem of inter-class imbalance, the target detection network uses a focal loss function, which mainly includes bounding box loss, category loss and confidence loss. The formula is as follows:

[0112] loss=loss CIOU2 +loss cls2 +loss confidence2

[0113] The bounding box loss is:

[0114]

[0115] Where d is the Euclidean distance between the center points of the two bounding boxes, c is the diagonal distance of the union of the two bounding boxes, and v is a parameter to measure the consistency of the aspect ratio, which is calculated as follows:

[0116]

[0117] Where:

[0118] w gt ——the width of the object’s ground-truth bounding box;

[0119] h gt ——The height of the target's true bounding box;

[0120] w——the width of the prediction box;

[0121] h – the height of the prediction box;

[0122] α is the weight parameter, calculated as follows:

[0123]

[0124] Category loss cls2 The coefficient is introduced to adjust the weight of difficult and easy samples, which is defined as follows:

[0125]

[0126] Specifically, this method combines the Feature Pyramid Network (FPN) and the Path Aggregation Network (PAN), and uses the CSP (CrossStage Partial) module to reduce the amount of network computation. In the Feature Pyramid Network, feature maps at different levels are fused through upsampling and lateral connections to generate a series of feature maps with different resolutions. These feature maps can capture target information at different scales. The structure of the Feature Pyramid Network is as follows: Figure 3As shown in the figure, the specific process is to pass features layer by layer through upsampling operations starting from the high-level feature map of the network through a top-down path. High-level feature maps have higher semantic information and help to identify complex targets. At the same time, through lateral links, at each scale, the upsampled high-level feature map is added to the low-level feature map whose channels are reduced by 1×1 convolution. The low-level feature map retains more spatial detail information, which helps to locate and identify small targets. This fusion method combines the semantic information of high-level features with the spatial information of low-level features. By adding a feature pyramid network, features of different scales can be effectively integrated, making the network more flexible and accurate in processing targets of different sizes and scales, capturing richer contextual information, and improving the detection performance of small-sized targets in a large field of view.

[0127] The path aggregation network further optimizes these feature maps. While retaining the high-level semantic information of the top-down path, it also enhances the feature representation by adding the bottom-up path. Its structure is as follows Figure 4 As shown in the figure, starting from the low-level feature map, the features are transferred layer by layer through downsampling, which aggregates the information of the low-level features and helps to better locate the target. The specific process is feature map N i Through 3×3 convolution, the size is reduced to the same level as the high-level feature map P in the feature pyramid network. i Consistent, then the feature map N i With the feature map P i+1 Splicing. The path aggregation network further enhances the richness and diversity of feature representation through top-down and bottom-up bidirectional feature aggregation, and can further improve the detection performance of small-sized objects in a large field of view.

[0128] The spatial pyramid pooling module has the characteristics of multi-scale feature fusion and can enhance the effectiveness and efficiency of feature extraction through pooling operations at different scales. The structure is as follows: Figure 5As shown in Figure 1. The spatial pyramid pooling module performs multi-scale pooling operations on the input feature map (e.g., 5×5, 9×9, and 13×13 pooling kernels), connects the pooled feature maps together, and forms a feature map containing information of different scales. This new feature map contains multi-scale context information, which can capture the global context information in the input feature map, help improve the effect of feature extraction, and enhance the network's perception of objects of different sizes. Assuming that the pooling kernel sizes used are 5×5, 9×9, and 13×13, respectively, the specific workflow of this module is to first perform a 5×5 maximum pooling operation on the input feature map, and then perform a 5×5 maximum pooling operation on the feature map after the first pooling. The two pooling operations are equivalent to a 9×9 pooling operation. Finally, the feature map after the second pooling is subjected to a 5×5 maximum pooling operation. The three pooling operations are equivalent to a 13×13 pooling operation. The original input feature map and the feature map after the three pooling operations are connected in the channel dimension to form a new feature map. At this time, the number of channels becomes four times the original, while the height and width remain unchanged. In this process, the module uses three 5×5 pooling kernels instead of three pooling kernels of different sizes, reducing redundant calculations and improving computational efficiency.

[0129] The CSP module has the characteristics of efficient feature fusion and channel compression, and can achieve effective feature extraction and compression through bottleneck layers and feature connections. Among them, the bottleneck layer is a residual structure containing two convolutional layers, such as Figure 6 As shown in the figure, the specific process is that the bottleneck layer first reduces the number of channels of the input feature map to half of the original number through 1×1 convolution. This process can reduce the amount of calculation and improve the calculation efficiency. Then, the bottleneck layer extracts features from the compressed feature map through 3×3 convolution. Finally, the input feature map and the feature map after 1×1 convolution and depth convolution are connected in the channel dimension, and the number of channels of the feature map is restored to the original number through another 1×1 convolution to obtain the output feature map. In this process, the bottleneck layer realizes channel compression and expansion through the combination of 1×1 convolution and 3×3 convolution, thereby effectively extracting features and enhancing feature expression capabilities while reducing the amount of calculation, improving the computational efficiency of the network, and reducing the risk of overfitting.

[0130] In the CSP module, the input feature map is divided into two parts A and B after passing through the first convolution layer. A is processed by multiple bottleneck layers to obtain A1, and B is concatenated with A1 in the channel dimension and passed through the second convolution layer to obtain the final output. In this process, the number of channels of the input data is transformed from c1 to 2×self.c through the first convolution layer, and then the input data is divided into two branches for processing. One of the branches is directly passed to the output, and the other branch is processed by multiple bottleneck layer modules. Finally, the features of different branches are concatenated in the channel dimension to achieve feature fusion, and the number of channels of the feature map after a series of operations is transformed from (2+n)×self.c to c2 through the second convolution layer. This process enhances the richness and diversity of feature representation through branch processing and feature fusion, generates more representational output, and helps to improve the performance of small target detection. The structure can be seen. Figure 7 .

[0131] In summary, the present invention adopts a target detection method based on multi-scale feature fusion, based on the large field of view images spliced ​​by the array camera, and uses the characteristics of multi-scale target detection network to fuse multi-level semantic information to achieve small-size target detection of projectiles under a large field of view. Finally, the projectile position information is obtained through calculation, which helps the subsequent process to realize the calculation of the projectile muzzle motion state parameters.

[0132] 3. Posture estimation module.

[0133] For the object pose estimation of a single-frame moment image, the present invention uses a monocular vision pose estimation method based on a 3D projectile model to complete the measurement. Compared with the traditional moving object pose estimation method based on setting reference points, this method avoids the reference point occlusion, high-speed motion and reference point loss caused by shooting angle, constructs a 3D discretized model of the projectile, and constructs an imaging plane in the three-dimensional space, performs projection calculation on the model, and obtains the model projection corresponding to the imaging plane.

[0134] On the other hand, the threshold segmentation process is performed on the actual object image by combining the neural network model to obtain the binary image plane contour of the actual object. The projection with the three-dimensional projectile model projection is scaled and compared to obtain the projection with the highest similarity. The specific pose parameters of the object can be obtained by combining its projection matrix and scaling parameters. This image segmentation model uses a V-Net-like architecture, which is a fully convolutional volume data segmentation neural network. The segmentation network model structure is as follows: Figure 8 As shown in the figure. It adopts an end-to-end training method, including a new objective function for optimization during training. At the same time, it can handle the strong imbalance between background and non-background very well. In order to solve the problem of limited data volume, nonlinear transformation and histogram matching are used for data enhancement.

[0135] According to the obtained projectile edge contour, for the projectile pose estimation task of a single-frame moment image, the present invention uses a monocular visual pose estimation method based on a 3D projectile model to complete the measurement of the projectile pose of the projectile out of the barrel. Compared with the traditional moving object pose estimation method based on setting reference points, this method avoids the reference point occlusion caused by smoke during firing, the reference point loss caused by the high-speed movement of the projectile and the shooting angle, and takes advantage of the relatively simple geometric appearance of the artillery projectile to construct a 3D discretized model of the projectile.

[0136] Combined with the 3D model of the projectile, an imaging plane is constructed in the three-dimensional space, and the model is projected and calculated by simulating the camera imaging principle to obtain the projection of the projectile model corresponding to the imaging plane. The specific principle is as follows:

[0137] Imaging under ideal imaging conditions can generally be described by a linear model. The pixel coordinates are linearly related to the coordinates of the three-dimensional world point. The pinhole imaging linear model contains an internal parameter matrix and an external parameter matrix. There is a certain pixel deviation between the real image and the result under ideal imaging conditions. The distortion of the camera can be described by a distortion model. The linear model and the distortion model together constitute a nonlinear model.

[0138] like Fig. 9 As shown, according to the image of the rigid body target captured by the camera, in the camera coordinate system O c -X c Y c Z c In the world coordinate system O, the position and attitude of the target can be represented by the rotation matrix and translation vector. w -X w Y w Z w In , the movement of the target can also be expressed by the product of the rotation matrix and the coordinate vector, which is equivalent to re-expressing the point position in another coordinate system. In a three-dimensional coordinate system, rotation can be divided into rotations around the X, Y, and Z axes. The rotation matrix R has a characteristic that its inverse is its transposed matrix, and the translation vector T is the offset between the relative positions of the target after movement.

[0139] Suppose the coordinate of a point in space in the world coordinate system is P, and the coordinate of this point mapped to the camera coordinate system by the shooting camera is P', then:

[0140] P'=R(PT)

[0141] The rotation matrix R can be represented by a 3×3 matrix, and the translation vector T can be represented by a 1×3 three-dimensional vector. The rotation matrix R can also be represented by three rotation angles, plus the three elements of the translation vector T, a total of 6 data, called camera extrinsic parameters. In fact, due to the existence of focal length, the point in the world coordinate system is first rotated and translated to the camera coordinate system, then mapped to the image physical coordinate system o-xyz through the intrinsic parameter, and finally transferred to the image pixel coordinate system o-uvz through the distance and pixel ratio.

[0142]

[0143] In the above formula, the point (u, v) is the point (X w ,Y w ,Z w ) is the projected pixel coordinate in the image, Z c Represents a point in the world coordinate system (X w ,Y w ,Z w ) is the Z-axis value in the camera coordinate system; k and l represent the actual length of each pixel unit in the x-axis and y-axis respectively, in millimeters, f represents the focal length of the camera, R represents the matrix of the world point rotation to the camera, and t represents its translation vector; f x Represents the focal length of the camera on the x-axis, f y represents the focal length of the camera on the y-axis, and (u0,v0) is the principal point of the image pixel coordinate system. x and f y These are the four internal parameters determined by the camera resolution and focal length. The data in the rotation angle and translation vector corresponding to the rotation matrix are the external parameters related to the camera's world position.

[0144] According to the above principle, the three-dimensional model of the projectile can be projected on the specified simulation imaging plane to obtain the corresponding simulation imaging projection of the standard posture (that is, each posture angle is 0). The projection is fitted using the characteristic triangle method to obtain the characteristic triangle of the imaging projection of the projectile model under the standard posture.

[0145] Based on the principle of the characteristic triangle method to obtain the projectile's attitude, according to the geometric characteristics of the projectile, the initial state is that the projectile's spatial characteristic triangle is on the horizontal plane, the triangle ΔB0C0D0 is an isosceles triangle, OD0 is the perpendicular bisector of ΔB0C0D0, pointing to the top of the imaging plane as the X-axis; the Y-axis is the direction of the plumb line, and the inward direction is positive; the Z-axis coincides with B0C0 and forms a right-hand system with the X-axis and the Y-axis. The triangle ΔBCD represents the actual position of the projectile in space, and the triangle ΔB'C'D' is the projection of ΔBCD on the horizontal plane. In the horizontal plane, α represents the yaw angle, β represents the pitch angle, and γ represents the roll angle. The projectile is a rotating body, and the roll angle defaults to 0. Fig.10 Schematic diagram of the feature fitting model.

[0146] Yaw angle:

[0147] Pitch Angle:

[0148] On the other hand, the image of the projectile detection result is subjected to threshold segmentation to obtain the binary image plane contour of the actual projectile. Edge detection extracts the target contour features. After the image is filtered, the edge detection operator is used to find the pixel points where the grayscale changes suddenly in the image. In the image, the grayscale value changes evenly and slowly in the target area, but it will suddenly change along the vertical direction of the boundary at the edge position of the target. A suitable gradient operator is used to find the location of the grayscale value mutation in the image. Based on this judgment method, the boundary contour of the target in the image can be determined. Such gradient operators are generally determined by calculating the spatial derivative of the grayscale value in the image.

[0149] The contour is fitted by the characteristic triangle method, and the characteristic triangle of the real imaging contour is scaled and compared with the projection characteristic triangle of the three-dimensional standard posture projectile model to obtain the projection matrix and scaling parameters between the two, and the pose estimation parameters of the projectile can be obtained.

[0150] In order to ensure the maximum accuracy of the test results, the present invention further processes the above solution process, adopts the posture solution technology based on the optimal consistency, and further optimizes the accuracy of the basic posture solution. As shown in the figure above, the characteristic point of the projectile model is D(x wd ,y wd ,z wd ),C(x wc ,y wc ,z wc ),B(x wb ,y wb ,z wb ). According to the relationship between the world coordinate system and the camera coordinate system, we can get the relationship:

[0151]

[0152] Where:

[0153] T wc : The coordinates of the origin of the world coordinate system in the camera coordinate system

[0154] R wc : Rotation matrix

[0155] At the same time, if the camera focal length is f, then the image coordinate system (x i ,y i ) and the camera coordinate system is:

[0156]

[0157] Where: (x i ,y i ) The distance from the image coordinate system to the target point can be obtained by intersection, i = B, C, D. Let the three sides of the triangle be represented by ; the actual lengths of the three sides are:

[0158]

[0159] Construct the optimal function:

[0160] φ(S)=Min((L DB -L1) 2 +(L BC -L2) 2 +(L DC -L3) 2 )

[0161] At the initial value S 0 Performing first-order Taylor expansion everywhere, we get the linear equation:

[0162] φ(S)=φ(S 0 )+Bδ S

[0163] Where: B is the optimal function at the initial value S 0 The first-order partial derivative at δ m is the limit error correction number. Each item in S is corrected accordingly:

[0164]

[0165] Solution steps:

[0166] (1) Extract feature point image coordinates; set the limit error δ min , the initial value of the attitude is obtained according to the characteristic triangle method.

[0167] (2) Using the least squares solution, we can obtain the correction number δ for each parameter: s .

[0168] (3) If δ s Greater than δ min , repeat 2) and 3) to continue iterating, otherwise end.

[0169] The posture solution obtained through this scheme can effectively guarantee the accuracy of the test results of this system and provide a reliable data basis for the overall error analysis.

[0170] The present invention uses a target detection algorithm based on multi-level feature fusion of a neural network, generates test samples through simulation data and live ammunition images, and stably detects projectiles with an average accuracy of 96%. Fig.11 The detection results of the algorithm are demonstrated, and through multi-level feature fusion, it has a stable detection effect on projectile targets on terrains of different scales.

[0171] At the same time, the present invention adopts a projectile posture solution method based on a 3D projectile model, uses a static real-bullet model image combined with measurement data to generate a test sample, completes the projectile 3D model reconstruction and state parameter solution, with a maximum range of 15 degrees (pitch angle) and 12 degrees (yaw angle); the average measurement error is 4.7%. Fig.12 The depth distribution map of the 3D model is displayed, which gives full play to the advantages of 3D model reconstruction and enhances the anti-interference ability under various environmental factors.

[0172] Fig.13 The actual effect of the posture estimation of the present invention is demonstrated, which fully integrates the depth information and effectively improves the detection effect.

Claims

1. A projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting, characterized in that: It includes image acquisition equipment, target detection module and posture estimation module; The image acquisition device is used to obtain projectile launch images, using a single or multiple cameras, and the hardware information provided by the image acquisition device itself is used in the posture solution of the posture estimation module; The target detection module adopts a target detection method based on multi-scale feature fusion. According to the large field of view image, the multi-scale target detection network is used to fuse the characteristics of multi-level semantic information to achieve small-size target detection of projectiles under a large field of view; finally, the projectile position information is obtained by solving the calculation to achieve the initial velocity detection of the projectile muzzle; The posture estimation module uses a monocular vision posture estimation method based on a 3D projectile model to complete the measurement; constructs a 3D discretized model of the projectile, and at the same time constructs an imaging plane in three-dimensional space, performs projection calculation on the model, and obtains the model projection corresponding to the imaging plane; combines the neural network model to perform threshold segmentation processing on the actual object image to obtain the binary imaging plane contour of the actual object, and obtains the projection with the highest similarity through scaling and comparison calculation of the contour and the projection of the three-dimensional space projectile model, and combines the projection matrix with the scaling parameters to obtain the specific posture parameters of the object.

2. According to claim 1, a projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting is characterized in that: The target detection method based on multi-scale feature fusion is specifically as follows: Step 1-1: Use feature pyramid and path aggregation network to effectively fuse features of different scales through top-down and bottom-up paths, as well as lateral connections and feature aggregation, and integrate high-level semantic information into the feature map while retaining high-resolution information at the bottom level; Step 1-2: Deblur the moving image using motion region detection and motion blur recovery algorithms; add noise factors to the transfer function of the Wiener filter model to perform blur recovery; The hierarchical quality is identified based on the restored image. The hierarchical quality is determined by the detection and tracking results. The accuracy of the detection and tracking results is calculated using the vertices of the regression box and the true labels in the training data. The quality judgment loss function is shown in the following formula: Among them, α is the weight parameter, (x i ,y i ) is the vertex coordinate of the regression box, (x g ,y g ) are vertex coordinates of the true value; Step 1-3: After completing the image restoration, use the target detection method based on multi-scale feature fusion to complete the projectile detection. The steps for constructing the target detection network model based on multi-scale feature fusion are as follows: ① Design the backbone network as a lightweight network; Depthwise separable convolution is used instead of standard convolution, and the convolution process is divided into two steps: channel-by-channel convolution and point-by-point 1×1 convolution. The unit of the lightweight network is a reverse residual structure, and the reverse residual block structure is wide in the middle and narrow on both sides: 1×1 convolution is used for dimensionality increase, 3×3 depthwise separable convolution is used for feature extraction, and finally 1×1 convolution is used for dimensionality reduction. The lightweight network contains 5 convolution modules composed of reverse residual blocks, and the output feature maps of each convolution module are shown in Figure 2. Indicated as C1, C2, C3, C4, C5; set the input image size of the network to 416×416, then the dimensions of the feature maps C1 to C5 are: 208×208×16, 104×104×24, 52×52×32, 26×26×96, 13×13×320, respectively. Compared with the original image, the size of the C1 layer is reduced by 2 times, the size of the C2 layer is reduced by 4 times, the size of the C3 layer is reduced by 8 times, the size of the C4 layer is reduced by 16 times, and the size of the C5 layer is reduced by 32 times; ② Perform spatial pyramid pooling on the feature map obtained in step ① to increase the receptive field of the network; Use three pooling layers of different scales: 5×5, 9×9, and 13×13. The outputs of the three pooling layers are channel-joined, and then a 1×1 convolution is performed to reduce the feature dimension. The feature map is denoted as C6. ③ Feature fusion based on C2, C3, C4, and C6 network layers; C2, C3, C4, and C6 are convolved with a kernel of 1×1 respectively; C6 is upsampled and then concatenated with the C4 layer after the convolution operation, and the resulting feature map is denoted as P1; P1 is upsampled again, C2 is downsampled, and concatenated with the C3 layer after the convolution operation, and the resulting feature map is denoted as P2, which is then sent to detection head 1 after the convolution operation; P2 is downsampled, and concatenated with P1 to obtain P3, which is sent to detection head 2 after the convolution operation; finally, P3 is downsampled, concatenated with C6, and the result is sent to detection head 3 after the convolution operation; ④The target detection network uses a focal loss function, including the bounding box loss CIOU2 , category loss loss cls2 And confidence loss loss confidence2 , the formula is as follows: loss=loss CIOU2 +loss cls2 +loss confidence2 The bounding box loss is: Where d is the Euclidean distance between the center points of the two bounding boxes, c is the diagonal distance of the union of the two bounding boxes, and υ is a parameter to measure the consistency of the aspect ratio, which is calculated as follows: Where: w gt ——the width of the object’s ground-truth bounding box; h gt ——The height of the target's true bounding box; w——the width of the prediction box; h – the height of the prediction box; α is the weight parameter, calculated as follows: Category loss cls2 The coefficient is introduced to adjust the weight of difficult and easy samples, which is defined as follows: Among them, s 2 is the number of grids in the image, B is the number of prediction boxes, and I i,j no_obj is an indicator function that takes a value of 1 when there is no target in the current prediction box, α t is the weight parameter, γ is the adjustment factor; p is the probability value predicted by the model, indicating the probability that a grid cell does not contain the target object.

3. According to claim 2, a projectile posture measurement system based on multi-level feature fusion neural network and visual three-dimensional model fitting is characterized in that: The posture estimation module is specifically: The image segmentation model in the posture estimation module uses a V-Net-like architecture, adopts an end-to-end training method, and uses nonlinear transformation and histogram matching to perform data enhancement; According to the obtained projectile edge contour, for the projectile pose estimation task of a single-frame moment image, a monocular vision pose estimation method based on a 3D projectile model is used to complete the measurement of the projectile pose of the projectile out of the chamber; combined with the projectile 3D model, an imaging plane is constructed in three-dimensional space, and the model is projected and calculated by simulating the camera imaging principle to obtain the projectile model projection corresponding to the imaging plane. The specific principle is as follows: According to the image of the rigid body target captured by the camera, in the camera coordinate system O c -X c Y c Z c In the world coordinate system O, the rotation matrix and translation vector are used to represent the position and posture of the target; w -X w Y w Z w In the 3D coordinate system, the target's motion is represented by the product of the rotation matrix and the coordinate vector, which is equivalent to the re-expression of the point position in another coordinate system. In the 3D coordinate system, rotation is divided into rotations around the X, Y, and Z axes. The rotation matrix R has a characteristic that the inverse of the rotation matrix R is its own transposed matrix. The translation vector T is the offset between the relative positions of the target after movement. Suppose the coordinate of a point in space in the world coordinate system is P, and the coordinate of this point mapped to the camera coordinate system by the shooting camera is P', then: P'=R(PT) The rotation matrix R is represented by a 3×3 matrix, and the translation vector T is represented by a 1×3 three-dimensional vector; the point in the world coordinate system is first rotated and translated to the camera coordinate system, then mapped to the image physical coordinate system o-xyz through the intrinsic parameter, and finally transferred to the image pixel coordinate system o-uvz through the distance and pixel ratio; In the formula, the point (u, v) is the point (X w ,Y w ,Z w ) is the projected pixel coordinate in the image, Z c Represents a point in the world coordinate system (X w ,Y w ,Z w ) is the Z-axis value in the camera coordinate system; k and l represent the actual length of each pixel unit in the x-axis and y-axis respectively, in millimeters, f represents the focal length of the camera, and t represents its translation vector; f x Represents the focal length of the camera on the x-axis, f y represents the focal length of the camera on the y-axis, (u0, v0) is the principal point of the image pixel coordinate system; where u0, v0, f x and f y These are the four internal parameters determined by the camera resolution and focal length. The data in the rotation angle and translation vector corresponding to the rotation matrix are the external parameters related to the camera's world position. According to the above principle, the three-dimensional model of the projectile is projected on the specified simulation imaging plane to obtain the corresponding simulation imaging projection with the standard posture, that is, each posture angle is 0. The projection is fitted using the model constructed by the characteristic triangle method to obtain the characteristic triangle of the imaging projection of the projectile model under the standard posture; Based on the principle of the characteristic triangle method for obtaining the projectile's posture, according to the geometric characteristics of the projectile, the initial state is that the projectile's spatial characteristic triangle is on the horizontal plane, the triangle ΔB0C0D0 is an isosceles triangle, OD0 is the perpendicular bisector of ΔB0C0D0, pointing to the top of the imaging plane as the X-axis; the Y-axis is the direction of the plumb line, and the inward direction is positive; the Z-axis coincides with B0C0, and forms a right-hand system with the X-axis and the Y-axis; the triangle ΔBCD represents the actual position of the projectile in space, and the triangle ΔB'C'D' is the projection of ΔBCD on the horizontal plane. In the horizontal plane, α represents the yaw angle, β represents the pitch angle, and γ represents the roll angle. The projectile is a rotating body, and the roll angle defaults to 0; Yaw angle: Pitch Angle: The image of the bullet detection result is subjected to threshold segmentation to obtain the binary image plane contour of the actual bullet. The edge detection is used to extract the target contour features. After the image is filtered, the edge detection operator is used to find the pixel points with grayscale mutation in the image. The contour is fitted by the characteristic triangle method, and the characteristic triangle of the physical imaging contour is used to scale and compare the characteristic triangle of the projection of the three-dimensional standard posture projectile model to obtain the projection matrix and scaling parameters between the two, that is, the estimated posture parameters of the projectile; The posture solution technology based on optimal consistency is adopted to further optimize the accuracy of the basic posture solution; the characteristic point of the projectile model is D(x wd ,y wd ,z wd ),C(x wc ,y wc ,z wc ),B(x wb ,y wb ,z wb ); According to the relationship between the world coordinate system and the camera coordinate system, the relationship can be obtained: Where: T wc : The coordinates of the origin of the world coordinate system in the camera coordinate system R wc : Rotation matrix At the same time, if the camera focal length is f, then the image coordinate system (x i ,y i ) and the camera coordinate system is: Where: (x i ,y i ) The distance from the image coordinate system to the target point can be obtained by intersection, i = B, C, D; let the three sides of the triangle be represented by respectively; the actual lengths of the three sides are: Construct the optimal function: φ(S)=Min((L DB -L1) 2 +(L BC -L2) 2 +(L DC -L3) 2 ) At the initial value S 0 Performing first-order Taylor expansion everywhere, we get the linear equation: φ(S)=φ(S 0 )+Bδ S Where: B is the optimal function at the initial value S 0 The first-order partial derivative at δ m is the limit error correction number; each item in S is corrected accordingly: Solution steps: (1) Extract feature point image coordinates; set the limit error δ min , obtain the initial value of the attitude according to the characteristic triangle method; (2) Using the least squares solution, we can obtain the correction number δ for each parameter: s ; (3) If δ s Greater than δ min , repeat 2) and 3) to continue iterating, otherwise end.

Citation Information

Cited By

  • Aircraft classification method and system based on fuzzy direction prior

    CN120808322A