A time domain confidence dynamic target image recognition method mounted on a sight
Patent Information
- Application Number
- CN202610854314.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-13
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-06-13
AI Technical Summary
然而,在实际应用中,上述传统方案存在两项显著缺陷:首先,高动态平台在运行瞬间产生的剧烈后坐力抖动,导致图像成像面产生严重的非线性透视投影畸变,使前后帧画面无法对齐,从而引发大规模的误检与漏检;其次,常规的目标追踪算法多依赖于硬阈值的多帧连续匹配逻辑,一旦目标在时域流中遭遇短时物理遮挡,算法的匹配链路便会瞬间中断
[0062]本发明通过构建双路径的并行交验机制,将瞬态运动突变特征与深度语义模式相融合,有效解决了传统单帧图像识别方法在高频振动、间歇性局部物理遮挡及多变光照环境下极易出现的误检与丢帧问题,显著提升了动态特征点捕获的时域连续性,利用初始帧基准图形的同心多级几何约束,自适应构建归一化的虚拟物理坐标系,实现了标定过程的完全自动化与零人工先验依赖;同时配合基于霍夫变换特征动态更新的单应性纠偏机制,能够实时对冲采集设备与目标物体之间的任意三维透视投影畸变,保证了多时段、多视角的亚像素级对齐精度。
Smart Images

Figure CN122368870B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and more specifically, to a temporal confidence dynamic target image recognition method mounted on a sight. Background Technology
[0002] With the rapid development of digital image processing technology, video stream pattern recognition systems mounted on high-dynamic motion platforms have been widely applied. In these high-dynamic application scenarios, image acquisition equipment typically faces severe mechanical vibrations, sudden pose shifts, and complex and ever-changing environmental lighting interference. Under such extreme working conditions, high-precision, continuous dynamic target recognition and spatial trajectory tracking face enormous challenges.
[0003] Currently, existing image recognition methods typically employ adjacent frame difference methods or conventional deep learning object detection algorithms to extract newly added feature points from video streams. However, in practical applications, these traditional approaches suffer from two significant drawbacks: First, the severe recoil jitter generated by high-dynamic platforms during operation causes severe nonlinear perspective projection distortion on the image imaging plane, making it impossible to align consecutive frames and leading to large-scale false positives and false negatives. Second, conventional object tracking algorithms often rely on multi-frame continuous matching logic with hard thresholds. If the target encounters a brief physical occlusion in the temporal stream, the matching chain of the algorithm is instantly interrupted. After the target is reproduced, the system must re-execute the accumulation calculation, resulting in extremely poor feature capture continuity in high-dynamic video streams and making it impossible to achieve precise spatiotemporal mapping at the sub-pixel level. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art, embodiments of the present invention provide a time-domain confidence dynamic target image recognition method mounted on a sight to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for time-domain confidence dynamic target image recognition mounted on a sight, comprising the following steps:
[0006] S1. Use image acquisition equipment to acquire real-time video stream of the scene to be observed; use a pre-constructed deep convolutional neural network to identify the reference graphic contour and its central features in the initial frame of the video stream; based on the geometric center and edge constraints of the reference graphic contour, construct a normalized virtual pixel coordinate system through two-dimensional coordinate transformation to achieve reference benchmark calibration without manual label dependence.
[0007] S2. Input the real-time video stream in parallel to the temporal dynamic difference path and the static semantic path;
[0008] In the temporal dynamic differential path, the pixel-level grayscale instantaneous jump region between the current frame and the reference frame after adaptive motion compensation is calculated and used as the target point to be determined;
[0009] In the static semantic path, a lightweight convolutional neural network is used to extract local features of the target point and calculate the semantic matching score between the target point and the prior morphological feature library.
[0010] S3. Establish a target state monitoring sequence and perform cross-frame trajectory tracking of the target point to be determined; construct a temporal confidence evolution model, introduce a dynamic forgetting factor to evaluate the temporal continuity of the target point to be determined using a nonlinear sliding window; when the target encounters instantaneous physical occlusion that causes local detection loss, perform nonlinear confidence decay and use state prediction to maintain the tracking link; when the accumulated confidence reaches the target, the target point to be determined is confirmed as a valid new feature point.
[0011] S4. Monitor the geometric deformation of the reference graphic contour in the video stream in real time, extract the perspective distortion parameters under the current imaging plane by using geometric parametric fitting, and dynamically update the homography mapping matrix of the current frame relative to the initial frame to compensate for the pose offset of the image acquisition device under high dynamic transient disturbances; use the updated homography mapping matrix to perform spatial correction on the effective newly added feature points and map them to the virtual pixel coordinate system.
[0012] S5. Perform density-based spatial clustering analysis on the effective new feature points mapped to the virtual pixel coordinate system, and remove non-systematic noise points that deviate from the main cluster; calculate the statistical centroid of the feature cluster, and solve the spatial geometric distribution envelope and directional standard deviation of the feature cluster, and output the quantitative analysis results.
[0013] Preferably, as a preferred embodiment of the time-domain confidence dynamic target image recognition method mounted on a sight according to the present invention, S1 specifically includes the following:
[0014] A deep convolutional neural network is used to perform multi-scale feature extraction and image segmentation on the initial frame image, outputting mask data of the main body region of the baseline image; an edge detection operator is used to obtain the set of feature points of the outermost closed contour. ,in, Indicates the first The pixel coordinates of each contour point in the original image coordinate system;
[0015] Based on closed contour feature point set The geometric center point of the reference figure is determined by using the least squares method to fit a geometric circle. And extract the key feature points inside the benchmark image. Calculate the spatial positional correlation vector between the geometric center point and the internal key feature points. This serves as a self-calibration constraint for global coordinate mapping; the self-calibration constraint is used to verify the self-calibration constraint in the event of nonlinear deformation in subsequent images. The rate of change of the modulus is used to evaluate the calibration confidence of the global coordinate mapping matrix;
[0016] Using the geometric center of the fitted reference figure Using the origin as the coordinate system origin and the major axis of the reference graphic's outline as the horizontal axis, a scale factor is established based on the ratio between the inherent physical dimensions of the reference graphic and the pixel dimensions. The scale factor The calculation formula is expressed as follows: ;in, The reference graphic has a pre-defined known physical dimension along its major axis. For the closed contour feature point set Extracted long axis pixel span;
[0017] The pixel coordinate system of the original image is mapped to a normalized virtual pixel coordinate system through translation, rotation, and scaling transformations. The transformation relationship is expressed as follows:
[0018]
[0019] in, It is a normalized mapping matrix that includes rotation, translation, and scaling parameters. For coordinates in the normalized virtual pixel coordinate system, These are the pixel coordinates in the original image;
[0020] The normalized mapping matrix The specific matrix parameter topology is described below:
[0021]
[0022] in, The scale factor is mentioned above. The angle between the major axis of the reference graphic outline and the horizontal axis of the original image. The geometric center point Pixel coordinates in the original image coordinate system.
[0023] Preferably, as a preferred embodiment of the time-domain confidence dynamic target image recognition method mounted on a sight according to the present invention, S2 specifically includes the following:
[0024] Get the current frame in the video stream With reference frame Background feature point pairs are used to calculate the global motion transformation matrix based on the random sampling consensus algorithm. For the reference frame Perform an affine transformation to align the pose shift caused by mechanical vibration, and obtain the compensated reference frame. ;in, Represents the real-time frame index of the video stream;
[0025] For the current frame With the compensated reference frame Perform weighted difference operation to calculate grayscale variation. ;in, The illumination-adaptive weighting coefficient is calculated using the following formula: ; To exclude the set of background pixel coordinates after excluding dynamic edge regions, These are the pixel coordinates in the original image;
[0026] Grayscale variation image Adaptive threshold binarization is performed, pixel clusters are extracted through connectivity analysis, and geometric feature filtering of pixel clusters is combined with prior feature size bounding boxes to obtain a set of undetermined target points that meet the area constraint conditions. ,in, Indicates the first The centroid position and bounding box parameters of the undetermined target point;
[0027] For each undetermined target point Crop local feature windows to the center The input is fed into a pre-trained lightweight convolutional neural network, which uses multiple convolutional kernels to extract multi-dimensional morphological features, including edge gradients, hollowness, and contrast distribution, and calculates the output local feature window using the Softmax function. semantic matching score ,in, Indicates the index of the target point to be determined. This is the feature mapping function of a convolutional neural network. For network learnable parameters, and They represent the first The local feature window corresponding to the nth target point and the nth target point The semantic matching score calculated for each target point; Used to quantitatively characterize the probability that a target point of unknown shape conforms to prior morphological features in terms of geometric appearance;
[0028] Score the semantic matching degree Greater than the preset first threshold Target points are retained, and a set of high-scoring target points is constructed. and set the high-scoring target points The set of undetermined target points output by the time-domain dynamic difference path Perform cross-comparison of local spatial neighborhoods. The specific cross-comparison logic is expressed as follows:
[0029] for Any high-scoring target point in With sets Any undetermined target point in Calculate the spatial Euclidean distance between the centroid coordinates of the two objects. and the intersection-union ratio of their bounding boxes. If and only if the following condition is met: and At that time, determine the high-scoring target point. and the target point to be determined Those belonging to the same spatial mutation source are output to the set of new candidate targets that simultaneously satisfy both temporal motion mutation features and static morphological semantic features; among them, The preset maximum neighborhood tolerance distance, This is the preset lower limit threshold for geometric overlap.
[0030] Preferably, as a preferred embodiment of the time-domain confidence dynamic target image recognition method mounted on a sight according to the present invention, S3 specifically includes the following:
[0031] Assign a unique global identifier to each target in the candidate new target set, and establish a status monitoring sequence. ,in, Indicates the record target at the 1st The frame's center pixel coordinates, bounding box size, and semantic matching score A time-domain confidence evolution model is constructed, specifying the initial confidence level when the target first enters the state monitoring sequence. The current confidence level of the target The evolutionary logic is as follows:
[0032] When the target is in the current number When a frame is successfully associated, the confidence level is positively accumulated, and the specific mathematical expression is as follows: ;in, This is the confidence level accumulation coefficient. For the first Frame semantic matching score, This represents the upper limit of confidence.
[0033] When the target is not detected in the current frame, the confidence decay mechanism is enabled: ;in, As a dynamic forgetting factor; when Below the extinction threshold When the target is determined to be dead, it is removed from the status monitoring sequence;
[0034] Predicting the target in the next frame using the Kalman filter algorithm The target is located at the expected position in the target image. Within a preset search radius, the detection point with the highest overlap with the predicted position is found to achieve cross-frame association. The Euclidean distance between consecutive frames is calculated as the spatial offset. When a target encounters momentary physical occlusion and fails to find a successfully associated detection point within the search radius, the expected position predicted by the Kalman filter algorithm is directly used as the virtual pixel coordinates of the target in the current frame to maintain the tracking link until the confidence decays below the extinction threshold. until;
[0035] The dynamic forgetting factor It is based on the target's current spatial offset. Dynamic correction is performed, and the specific formula is shown below: ,in, As the baseline forgetting factor, The spatial offset of the current frame. For scale parameters;
[0036] Set the time observation window length to For each frame, a multi-dimensional consistency evaluation is performed on the targets within the sequence. The undetermined target point is determined as a valid new feature point if and only if both of the following two feature constraints are met simultaneously. The determination logic is as follows:
[0037] Spatial stability characteristics: Verify that the average spatial displacement of the target within the time observation window satisfies the mean constraint. ,in, For the first Spatial offset of a frame A preset spatial displacement deviation threshold is used to eliminate random high-frequency ionizing noise in the environment;
[0038] Semantic persistence features: Verify that the average semantic score of the target within the time observation window satisfies the lower bound constraint: ,in, Indicates the first Frame semantic matching score, The minimum threshold for scoring semantic matching.
[0039] Preferably, as a preferred embodiment of the time-domain confidence dynamic target image recognition method mounted on a sight according to the present invention, S4 specifically includes the following:
[0040] In the current calibration frame In this process, local edge detection is performed based on the baseline graphic region determined by S1. The Canny operator is used to extract the edge pixel set of each level of concentric loops within the baseline graphic. Outlier noise points are removed by statistical filtering to obtain the cleaned geometric edge point set. ;in, This is the index of the current calibration frame;
[0041] Using the random Hough transform on the geometric edge point set Perform multi-ellipse fitting to extract ellipse constraint parameters under the current imaging plane, including the center pixel coordinates. Long axis short axis and rotation angle The perspective distortion rate in the current imaging plane is defined using the ratio of the major and minor axes. The calculation formula is expressed as follows: ;in, and These represent the parameters of the major axis of the ellipse fitted from the current frame. With minor axis parameters The maximum and minimum values between;
[0042] The ellipse constraint parameters extracted in the current frame are compared with the reference geometric parameters in the initial frame S1 to extract the coordinates of the endpoints and center pixel of the major and minor axes of the ellipse. Construct a set of corresponding point pairs for concentric features; before calculating the dynamic homography matrix, based on the aforementioned perspective distortion rate... The hierarchical adaptive verification logic is executed, and the specific determination is as follows:
[0043] 1) When When a slight deformation of the reference image is detected, a "weak mechanical vibration" status label is output, and entry into the homography matrix calculation is allowed; among which, The preset first distortion threshold;
[0044] 2) When When perspective distortion is detected in the reference image, a "high-frequency mechanical vibration" status label is output, and entry into homography matrix calculation is allowed; among which, This is the preset second distortion threshold;
[0045] 3) When When the reference image is determined to have severe distortion, a "Severe Distortion" status label is output, and entry into the homography matrix calculation is allowed; among which, This is the preset distortion tolerance threshold.
[0046] 4) The current frame homography estimation is deemed to have failed when any of the following triggering conditions are met: Condition 1, the perspective distortion rate... Condition 2, the rotation tilt angle mentioned in the current frame The absolute deviation of the rotation angle relative to the initial frame reference angle exceeds the preset angular deformation threshold. Condition 3: The cleaned set of geometric edge points The total number of feature points participating in the fitting is lower than the preset minimum number of points threshold. When homography estimation fails, the homography matrix of the previous valid calibration frame is directly retained and reused. As the mapping parameter of the current calibration frame, it sends a pose reset self-test prompt to the terminal;
[0047] After passing the above-mentioned hierarchical adaptive verification, using the correspondence of at least four sets of point pairs, the dynamic homography matrix of the current calibration frame relative to the initial reference frame is recalculated using a least-squares optimization algorithm. ,in, This is the index of the current calibration frame;
[0048] Using the updated dynamic homography matrix Sub-pixel level spatial correction is performed on the effective newly added feature points. The corrected coordinate mapping logic is expressed through a two-dimensional homogeneous coordinate transformation as follows:
[0049]
[0050] in, These are the real-time detected homogeneous pixel coordinates in the current calibration frame. This is the inverse of the pose offset matrix of the current calibration frame, used to inversely project the pixel coordinates of the current calibration frame back to the reference pixel plane of the initial frame. The global static normalized mapping matrix described in S1 includes rotation, translation, and scaling parameters. These are the final physical homogeneous coordinates mapped to the normalized virtual pixel coordinate system.
[0051] Preferably, as a preferred embodiment of the time-domain confidence dynamic target image recognition method mounted on a sight according to the present invention, S5 specifically includes the following:
[0052] After removing outliers from spatial clustering, the geometric center of the feature cluster is calculated as the statistical centroid. The specific calculation formula is as follows: , ,in, This represents the total number of valid new feature points included in the calculation after removing outlier feature points. Indicates the first The physical coordinates of each valid newly added feature point in the normalized virtual pixel coordinate system;
[0053] Extract the mapping coordinates of the reference graphic center feature determined in S1 in the normalized virtual pixel coordinate system. and statistical centroid The relative offset vector is constructed as follows: The relative offset vector The modulus length and orientation quantitatively characterize systematic pose assembly deviations;
[0054] Solve for the spatial geometric distribution envelope of the feature cluster, which is obtained by finding the minimum circumscribed circle radius that contains all valid new feature points. For quantitative characterization, the specific calculation formula is as follows: ;in, This indicates that the maximum value is calculated by iterating through all valid new feature points involved in the calculation, where the minimum circumscribed circle radius is... Used to quantitatively assess the overall characteristic scattering density of a high-dynamic platform during the observation period;
[0055] Calculate the standard deviation of the feature clusters in the horizontal direction respectively. Standard deviation in the vertical direction The specific calculation formula is as follows:
[0056] ,
[0057] Based on the relative offset vector Minimum circumscribed circle radius with directional standard deviation , It outputs a quantitative analysis report on the spatial distribution pattern of feature clusters to the terminal;
[0058] The specific logic for generating the quantitative analysis report is as follows: by analyzing the horizontal standard deviation... Standard deviation in the vertical direction The ratio relationship is used to automatically determine the main cause of error in high dynamic platforms: when the following conditions are met... When the horizontal dimension disturbance is determined to be the primary cause, then when the following conditions are met... At that time, the vertical dimension disturbance was determined to be the main cause; and targeted platform consistency calibration suggestions were output to the terminal; among them, The threshold for determining directional advantage.
[0059] On the other hand, the present invention provides a computer device including a memory and a processor, wherein the memory stores a computer program, wherein when the computer program is executed by the processor, it implements the steps of a time-domain confidence dynamic target image recognition method mounted on a sight as described above.
[0060] On the other hand, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements a time-domain confidence dynamic target image recognition method mounted on a sight as described above.
[0061] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0062] This invention constructs a parallel verification mechanism with dual paths, fusing transient motion mutation features with deep semantic patterns. This effectively solves the problems of false detection and frame loss that are prone to occur in traditional single-frame image recognition methods under high-frequency vibration, intermittent local physical occlusion, and variable lighting conditions. It significantly improves the temporal continuity of dynamic feature point capture. By utilizing the concentric multi-level geometric constraints of the initial frame reference image, a normalized virtual physical coordinate system is adaptively constructed, realizing complete automation and zero human prior dependence in the calibration process. At the same time, in conjunction with the homography correction mechanism based on Hough transform features for dynamic updating, it can offset arbitrary three-dimensional perspective projection distortion between the acquisition device and the target object in real time, ensuring sub-pixel level alignment accuracy across multiple time periods and multiple viewpoints. Attached Figure Description
[0063] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0064] Figure 1 This is a flowchart of the method of the present invention.
[0065] Figure 2 This is a schematic diagram of the topology of the parallel convergence of temporal dynamic difference and static semantic dual-track and the cross-comparison of spatial neighborhood in an embodiment of the present invention. Detailed Implementation
[0066] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0067] In the description of this application, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0068] In the description of this application, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0069] Example 1
[0070] This embodiment provides, for example Figure 1 The method for temporal confidence dynamic target image recognition mounted on a sight includes the following steps:
[0071] S1. Use image acquisition equipment to acquire real-time video stream of the scene to be observed; use a pre-constructed deep convolutional neural network to identify the reference graphic contour and its central features in the initial frame of the video stream; based on the geometric center and edge constraints of the reference graphic contour, construct a normalized virtual pixel coordinate system through two-dimensional coordinate transformation to achieve reference benchmark calibration without manual label dependence.
[0072] S2. Input the real-time video stream in parallel to the temporal dynamic difference path and the static semantic path;
[0073] In the temporal dynamic differential path, the pixel-level grayscale instantaneous jump region between the current frame and the reference frame after adaptive motion compensation is calculated and used as the target point to be determined;
[0074] In the static semantic path, a lightweight convolutional neural network is used to extract local features of the target point and calculate the semantic matching score between the target point and the prior morphological feature library.
[0075] S3. Establish a target state monitoring sequence and perform cross-frame trajectory tracking of the target point to be determined; construct a temporal confidence evolution model, introduce a dynamic forgetting factor to evaluate the temporal continuity of the target point to be determined using a nonlinear sliding window; when the target encounters instantaneous physical occlusion that causes local detection loss, perform nonlinear confidence decay and use state prediction to maintain the tracking link; when the accumulated confidence reaches the target, the target point to be determined is confirmed as a valid new feature point.
[0076] S4. Monitor the geometric deformation of the reference graphic contour in the video stream in real time, extract the perspective distortion parameters under the current imaging plane by using geometric parametric fitting, and dynamically update the homography mapping matrix of the current frame relative to the initial frame to compensate for the pose offset of the image acquisition device under high dynamic transient disturbances; use the updated homography mapping matrix to perform spatial correction on the effective newly added feature points and map them to the virtual pixel coordinate system.
[0077] S5. Perform density-based spatial clustering analysis on the effective new feature points mapped to the virtual pixel coordinate system, and remove non-systematic noise points that deviate from the main cluster; calculate the statistical centroid of the feature cluster, and solve the spatial geometric distribution envelope and directional standard deviation of the feature cluster, and output the quantitative analysis results.
[0078] Preferably, in step S1, a real-time video stream of the scene to be observed is acquired using an image acquisition device; a reference graphic contour and its center features in the initial frame of the video stream are identified through a pre-constructed deep convolutional neural network; based on the geometric center and edge constraints of the reference graphic contour, a normalized virtual pixel coordinate system is constructed through two-dimensional coordinate transformation to achieve reference benchmark calibration without manual label dependence, specifically including the following:
[0079] A deep convolutional neural network is used to perform multi-scale feature extraction and image segmentation on the initial frame image, outputting mask data of the main body region of the baseline image; an edge detection operator is used to obtain the set of feature points of the outermost closed contour. ,in, Indicates the first The pixel coordinates of each contour point in the original image coordinate system;
[0080] Based on closed contour feature point set The geometric center point of the reference figure is determined by using the least squares method to fit a geometric circle. And extract the key feature points inside the benchmark image. Calculate the spatial positional correlation vector between the geometric center point and the internal key feature points. This serves as a self-calibration constraint for global coordinate mapping; the self-calibration constraint is used to verify the self-calibration constraint in the event of nonlinear deformation in subsequent images. The rate of change of the modulus is used to evaluate the calibration confidence of the global coordinate mapping matrix;
[0081] Using the geometric center of the fitted reference figure Using the origin as the coordinate system origin and the major axis of the reference graphic's outline as the horizontal axis, a scale factor is established based on the ratio between the inherent physical dimensions of the reference graphic and the pixel dimensions. The scale factor The calculation formula is expressed as follows: ;in, The reference graphic has a pre-defined known physical dimension along its major axis. For the closed contour feature point set Extracted long axis pixel span;
[0082] The pixel coordinate system of the original image is mapped to a normalized virtual pixel coordinate system through translation, rotation, and scaling transformations. The transformation relationship is expressed as follows:
[0083]
[0084] in, It is a normalized mapping matrix that includes rotation, translation, and scaling parameters. For coordinates in the normalized virtual pixel coordinate system, These are the pixel coordinates in the original image;
[0085] The normalized mapping matrix The specific matrix parameter topology is described below:
[0086]
[0087] in, The scale factor is mentioned above. The angle between the major axis of the reference graphic outline and the horizontal axis of the original image. The geometric center point Pixel coordinates in the original image coordinate system.
[0088] Preferably, in step S2, the real-time video stream is input in parallel to the temporal dynamic difference path and the static semantic path; in the temporal dynamic difference path, the pixel-level grayscale instantaneous jump region between the current frame and the reference frame after adaptive motion compensation is calculated as the target point to be determined; in the static semantic path, a lightweight convolutional neural network is used to extract local features of the target point to be determined, and the semantic matching degree score between the target point to be determined and the prior morphological feature library is calculated, specifically including the following:
[0089] Get the current frame in the video stream With reference frame Background feature point pairs are used to calculate the global motion transformation matrix based on the random sampling consensus algorithm. For the reference frame Perform an affine transformation to align the pose shift caused by mechanical vibration, and obtain the compensated reference frame. ;in, Represents the real-time frame index of the video stream;
[0090] For the current frame With the compensated reference frame Perform weighted difference operation to calculate grayscale variation. ; for the grayscale variation image Adaptive threshold binarization is performed, pixel clusters are extracted through connectivity analysis, and geometric feature filtering of pixel clusters is combined with prior feature size bounding boxes to obtain a set of undetermined target points that meet the area constraint conditions. ,in, Indicates the first The centroid position and bounding box parameters of the unknown target point. This refers to the adaptive weighting coefficient for illumination, which is dynamically adjusted based on ambient light intensity.
[0091] The illumination adaptive weighting coefficient It is a dynamic mapping based on the average grayscale ratio of the global background pixels in the current frame and the reference frame, and the adjustment formula is expressed as follows: ;in, To exclude the set of background pixel coordinates after excluding dynamic edge regions, These are the pixel coordinates in the original image;
[0092] For each undetermined target point Crop local feature windows to the center The input is fed into a pre-trained lightweight convolutional neural network, which uses multiple convolutional kernels to extract multi-dimensional morphological features, including edge gradients, hollowness, and contrast distribution, and calculates the output local feature window using the Softmax function. semantic matching score ,in, Indicates the index of the target point to be determined. This is the feature mapping function of a convolutional neural network. For network learnable parameters, and They represent the first The local feature window corresponding to the nth target point and the nth target point The semantic matching score calculated for each target point; Used to quantitatively characterize the probability that a target point conforms to prior morphological features in terms of geometric shape, and to score semantic matching degree. Greater than the preset first threshold Target points are retained, and a set of high-scoring target points is constructed. and set the high-scoring target points The set of undetermined target points output by the time-domain dynamic difference path Perform cross-comparison of local spatial neighborhoods, such as Figure 2 As shown, the specific cross-comparison logic is expressed as follows:
[0093] for Any high-scoring target point in With sets Any undetermined target point in Calculate the spatial Euclidean distance between the centroid coordinates of the two objects. and the intersection-union ratio of their bounding boxes. If and only if the following condition is met: and At that time, determine the high-scoring target point. and the target point to be determined Trajectories belonging to the same spatial mutation source are identified, and these points are output to a set of new candidate targets that simultaneously satisfy both temporal motion mutation features and static morphological semantic features; among them, The preset maximum neighborhood tolerance distance, This is the preset lower limit threshold for geometric overlap.
[0094] It should be specifically noted that in S2, the illumination adaptive weighting coefficient The preferred dynamic adjustment range The first threshold in the static semantic path The preferred setting range is This is to ensure sensitivity to the morphological characteristics of micropores / spots during the initial screening stage.
[0095] Preferably, in step S3, a target state monitoring sequence is established, and cross-frame trajectory tracking is performed on the target point to be determined; a temporal confidence evolution model is constructed, and a dynamic forgetting factor is introduced to evaluate the temporal continuity of the target point through a nonlinear sliding window. When the target encounters instantaneous physical occlusion, resulting in local detection loss, nonlinear confidence decay is performed, and state prediction is used to maintain the tracking link; when the accumulated confidence reaches the target level, the target point to be determined is confirmed as a valid new feature point, specifically including the following:
[0096] Assign a unique global identifier to each target in the candidate new target set, and establish a status monitoring sequence. ,in, Indicates the record target at the 1st The center pixel coordinates, bounding box size, and semantic matching score of the frame A time-domain confidence evolution model is constructed, specifying the initial confidence level when the target first enters the state monitoring sequence. The current confidence level of the target The evolutionary logic is as follows:
[0097] When the target is in the current number When a frame is successfully associated, the confidence level is positively accumulated, and the specific mathematical expression is as follows: ;in, This is the confidence level accumulation coefficient. For the first Frame semantic matching score, This represents the upper limit of confidence.
[0098] When the target is not detected in the current frame, the confidence decay mechanism is enabled: ;in, As a dynamic forgetting factor; when Below the extinction threshold When the target is determined to be dead, it is removed from the status monitoring sequence;
[0099] Predicting the target in the next frame using the Kalman filter algorithm The target is located at the expected position in the target image. Within a preset search radius, the detection point with the highest overlap with the predicted position is found to achieve cross-frame association. The Euclidean distance between consecutive frames is calculated as the spatial offset. When a target encounters momentary physical occlusion and fails to find a successfully associated detection point within the search radius, the expected position predicted by the Kalman filter algorithm is directly used as the virtual pixel coordinates of the target in the current frame to maintain the tracking link until the confidence decays below the extinction threshold. until;
[0100] The dynamic forgetting factor It is dynamically corrected based on the target's current spatial offset, as shown in the following formula: ,in, As the baseline forgetting factor, The spatial offset of the current frame. The scale parameter is used; a nonlinear coupling between spatial displacement and time confidence is established through an exponential decay function, when the target spatial offset... When it increases, the forgetting factor The confidence level decreases rapidly as it decays, thus eliminating drift noise.
[0101] Set the time observation window length to For each frame, a multi-dimensional consistency evaluation is performed on the targets within the sequence. The undetermined target point is determined as a valid new feature point if and only if both of the following two feature constraints are met simultaneously. The determination logic is as follows:
[0102] Spatial stability characteristics: Verify that the average spatial displacement of the target within the time observation window satisfies the mean constraint. ,in, For the first Spatial offset of a frame A preset spatial displacement deviation threshold is used to eliminate random high-frequency ionizing noise in the environment;
[0103] Semantic persistence features: Verify that the average semantic score of the target within the time observation window satisfies the lower bound constraint: ,in, Indicates the first Frame semantic matching score, The minimum threshold for semantic matching score ensures that the target maintains high confidence in its prior geometric features throughout the entire observation period;
[0104] It should be specifically noted that in S3, the upper confidence level is... Set to 100, positive accumulation coefficient The preferred value is 15, the extinction threshold. Preferably 10; baseline forgetting factor The preferred value is 0.92, scale parameter This is used to normalize the displacement according to the image resolution, with a preferred value range of [value range missing]. The preferred time observation window length for G-frames is... Frame, determination threshold The preferred value is 0.80.
[0105] Preferably, in step S4, the geometric deformation of the reference graphic contour in the video stream is monitored in real time, and the perspective distortion parameters under the current imaging plane are extracted using geometric parametric fitting. The homography mapping matrix of the current frame relative to the initial frame is dynamically updated to compensate for the pose shift of the image acquisition device under high dynamic transient disturbances. The updated homography mapping matrix is used to perform spatial correction on the effectively newly added feature points and map them to the virtual pixel coordinate system. Specifically, this includes the following:
[0106] In the current calibration frame In this process, local edge detection is performed based on the baseline graphic region determined by S1. The Canny operator is used to extract the edge pixel set of each level of concentric loops within the baseline graphic. Outlier noise points are removed by statistical filtering to obtain the cleaned geometric edge point set. ;in, This is the index of the current calibration frame;
[0107] Using the random Hough transform on the geometric edge point set Perform multi-ellipse fitting to extract ellipse constraint parameters under the current imaging plane, including the center pixel coordinates. Long axis short axis and rotation angle The perspective distortion rate in the current imaging plane is defined using the ratio of the major and minor axes. Its mathematical formula is expressed as follows: ;in, and These represent the parameters of the major axis of the ellipse fitted from the current frame. With minor axis parameters The maximum and minimum values between; the perspective distortion rate It is used to quantitatively characterize the relative perspective distortion between the image acquisition device and the target object. By taking the ratio of the maximum and minimum values of the major and minor axes, the interference of the ellipse rotation direction on the distortion direction determination is eliminated, and it serves as the pre-verification basis for updating the homography matrix of the current frame.
[0108] The ellipse constraint parameters extracted in the current frame are compared with the reference geometric parameters in the initial frame S1 to extract the coordinates of the endpoints and center pixel of the major and minor axes of the ellipse. Construct a set of corresponding point pairs for concentric features; before calculating the dynamic homography matrix, based on the aforementioned perspective distortion rate... The hierarchical adaptive verification logic is executed, and the specific determination is as follows:
[0109] 1) When When the system determines that the current image acquisition device is minimally affected by high-dynamic transient disturbances and that the reference image undergoes slight deformation, it outputs a "weak mechanical vibration" status label and allows entry into homography matrix calculation; among which, The preset first distortion threshold is preferably set to 1.05;
[0110] 2) When When a spatial relative pose shift occurs between the image acquisition device and the target object, and the reference image exhibits "perspective distortion," a "high-frequency mechanical vibration" status label is output, and homography matrix calculation is allowed; among these, The preset second distortion threshold is preferably set to 1.20;
[0111] 3) When When the image acquisition device is subjected to a severe multi-degree-of-freedom impact, and the reference image is "severely distorted," a "severe distortion" status label is output, and a high-dynamic pose compensation mechanism is activated, allowing entry into homography matrix calculation; among which, The preset distortion tolerance threshold is preferably set to 1.30;
[0112] 4) The current frame homography estimation is deemed to have failed when any of the following triggering conditions are met: Condition 1, the perspective distortion rate... Condition 2, the rotation tilt angle mentioned in the current frame The absolute deviation of the rotation angle relative to the initial frame reference angle exceeds the preset angular deformation threshold. Condition 3: The cleaned set of geometric edge points The total number of feature points participating in the fitting is lower than the preset minimum number of points threshold. When homography estimation fails, the homography matrix of the previous valid calibration frame is directly retained and reused. As the mapping parameter of the current calibration frame, it sends a pose reset self-test prompt to the terminal;
[0113] After passing the above-mentioned hierarchical adaptive verification, using the correspondence of at least four sets of point pairs, the dynamic homography matrix of the current calibration frame relative to the initial reference frame is recalculated using a least-squares optimization algorithm. ,in, This is the index of the current calibration frame;
[0114] Using the updated dynamic homography matrix Sub-pixel level spatial correction is performed on the effectively newly added feature points. The corrected coordinate mapping logic is expressed through a two-dimensional homogeneous coordinate transformation as follows:
[0115]
[0116] in, These are the real-time detected homogeneous pixel coordinates in the current calibration frame. This is the inverse of the pose offset matrix of the current calibration frame, used to inversely project the pixel coordinates of the current calibration frame back to the reference pixel plane of the initial frame. The global static normalized mapping matrix described in S1 includes rotation, translation, and scaling parameters. These are the final physical homogeneous coordinates mapped to the normalized virtual pixel coordinate system.
[0117] Preferably, in step S5, density-based spatial clustering analysis is performed on the effective newly added feature points mapped to the virtual pixel coordinate system to remove non-systematic noise points that deviate from the main cluster; the statistical centroid of the feature cluster is calculated, and the spatial geometric distribution envelope and directional standard deviation of the feature cluster are solved, and the quantitative analysis results are output, specifically including the following:
[0118] After removing outliers from spatial clustering, the geometric center of the feature cluster is calculated as the statistical centroid. The specific calculation formula is as follows: , ,in, This represents the total number of valid new feature points included in the calculation after removing outlier feature points. Indicates the first The physical coordinates of each valid newly added feature point in the normalized virtual pixel coordinate system;
[0119] Extract the coordinates of the reference graphic center feature determined in S1 in the virtual pixel coordinate system. , with statistical centroid Constructing relative offset vectors The relative offset vector The modulus length and orientation quantitatively characterize systematic pose assembly deviations;
[0120] Solve for the spatial geometric distribution envelope of the feature cluster, which is obtained by finding the minimum circumscribed circle radius that contains all valid new feature points. For quantitative characterization, the specific calculation formula is expressed as follows: ;in, This indicates that the maximum value is calculated by iterating through all valid new feature points involved in the calculation, where the minimum circumscribed circle radius is... Used to quantitatively assess the overall characteristic scattering density of a high-dynamic platform during the observation period;
[0121] Calculate the standard deviation of the feature clusters in the horizontal direction respectively. Standard deviation in the vertical direction This is used to evaluate the observation stability of the image acquisition device on different degrees of freedom of a high dynamic platform. The specific calculation formula is as follows:
[0122] ,
[0123] Based on the relative offset vector Minimum circumscribed circle radius with directional standard deviation , It outputs a quantitative analysis report on the spatial distribution pattern of feature clusters to the terminal;
[0124] The specific logic for generating the quantitative analysis report is as follows: by analyzing the horizontal standard deviation... Standard deviation in the vertical direction The ratio relationship is used to automatically determine the main cause of error in high dynamic platforms: when the following conditions are met... When the "horizontal dimension perturbation" of the high-dynamic platform is determined to be the main cause of instability; when the condition is met... At that time, the "vertical dimension perturbation" of the high-dynamic platform was determined to be the main cause of instability; and targeted platform consistency calibration suggestions were output to the terminal; among them, A preset directional dominance threshold of 1 or higher is used to determine the significance of scattering differences between the two dimensions.
[0125] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0126] This embodiment also provides a storage medium storing a computer program that, when executed by a processor, implements the steps of implementing a time-domain confidence dynamic target image recognition method mounted on a sight as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0127] Example 2
[0128] The following is another embodiment of the present invention, which provides a time-domain confidence dynamic target image recognition method mounted on a sight. In order to verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.
[0129] This experiment aims to verify the effectiveness of a temporal confidence dynamic target image recognition method mounted on a sight. Through data acquisition, calibration, dynamic and static dual-track detection, temporal confidence state update, dynamic homography correction, and multi-point clustering statistics, it improves the ability to capture effective new feature points early and quantify systematic pose assembly deviations in high-dynamic and high-transient disturbance scenarios. The experiment uses real-time video streams of photoelectric observation under simulated high-frequency mechanical vibration and intermittent occlusion environments, including feature parameters, temporal continuity scores, spatial offsets, and normalized physical coordinates under different interference states. By analyzing the conformity between the model's abrupt feature point prediction results and actual known labels, the accuracy and robustness of the system in identifying, tracking, correcting, and predicting target features are verified.
[0130] The simulation experiment steps are implemented according to the content of the time-domain confidence dynamic target image recognition method mounted on a sight provided in Example 1, and the specific steps include:
[0131] By using image acquisition equipment to acquire real-time video streams of the scene to be observed, the outline of the reference graphic and its central features in the initial frame of the video stream are identified, and a normalized virtual pixel coordinate system is constructed through two-dimensional coordinate transformation to achieve reference calibration without manual label dependence.
[0132] The real-time video stream is input in parallel to the temporal dynamic difference path and the static semantic path. Through weighted difference operation and lightweight convolutional neural network feature extraction, transient motion change regions are captured in real time, and the semantic matching score of the local feature window is output. The set of target points to be determined is selected by comprehensive screening.
[0133] Global identifiers are assigned to candidate new targets, and a state monitoring sequence is established to construct a time-domain confidence evolution model. A dynamic forgetting factor is introduced to evaluate the time-domain continuity of the target points using a nonlinear sliding window. When the target encounters instantaneous physical occlusion or local detection loss, nonlinear confidence decay is performed, and Kalman filtering is used for state prediction to maintain the tracking link. When the accumulated confidence reaches the target level, the target points are confirmed as valid new feature points.
[0134] Real-time monitoring of the geometric deformation of the reference graphic contour in the video stream, extraction of perspective distortion parameters under the current imaging plane using random Hough transform, dynamic updating of the dynamic homography mapping matrix of the current frame relative to the initial frame, sub-pixel level spatial correction of the newly added feature points after weighting using the inverse matrix, pose offset under dynamic transient disturbances, and accurate mapping to the normalized virtual pixel coordinate system.
[0135] Density-based spatial clustering analysis is performed on the effective new feature points mapped to the virtual pixel coordinate system to remove non-systematic free noise points that deviate from the main cluster. The statistical centroid of the cleaned feature cluster is calculated, and the standard deviations in the horizontal and vertical directions are solved. Finally, a quantitative analysis report of the spatial distribution pattern is output to the terminal in combination with the relative offset vector.
[0136] The specific data from the above simulation experiments are shown in Table 1. Table 1 is the data record table for the simulation experiments of this invention: Table 1
[0137] Experimental Analysis:
[0138] By comparing the conformity of the model's prediction results, weighting nodes, and actual introduced physical disturbance labels at different time periods, the accuracy and robustness of the present invention in identifying and tracking dynamic mutation targets were verified.
[0139] Experimental results show that, during the 10-15 minute high-frequency mechanical vibration period, the perspective deformation ratio was successfully extracted using Hough transform multi-ellipse fitting. The constraint parameters and dynamic homography matrix successfully offset the device pose offset; during the 15-20 minute period of severe disturbance and occlusion, the temporal confidence evolution model withstood the pressure of detection loss, achieving a cumulative confidence of 95% and successfully completing spatial weighting for three newly added effective feature points; during the 25-30 minute convergence period, DBSCAN clustering analysis successfully removed non-systematic outlier noise drifting with airflow and accurately calculated the statistical centroid relative offset vector. and directional standard deviation;
[0140] In summary, this invention can solve the technical problems of target recognition interruption and large spatial mapping deviation caused by jitter distortion and short-term occlusion in high dynamic image acquisition environment, and improve the continuity of target recognition.
[0141] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. A method for time-domain confidence-based dynamic target image recognition mounted on a sight, characterized in that: Includes the following steps: S1. Use image acquisition equipment to acquire real-time video stream of the scene to be observed; use a pre-constructed deep convolutional neural network to identify the reference graphic contour and its central features in the initial frame of the video stream; based on the geometric center and edge constraints of the reference graphic contour, construct a normalized virtual pixel coordinate system through two-dimensional coordinate transformation to achieve reference benchmark calibration without manual label dependence. S2. Input the real-time video stream in parallel to the temporal dynamic difference path and the static semantic path; In the temporal dynamic differential path, the pixel-level grayscale instantaneous jump region between the current frame and the reference frame after adaptive motion compensation is calculated and used as the target point to be determined; In the static semantic path, a lightweight convolutional neural network is used to extract local features of the target point and calculate the semantic matching score between the target point and the prior morphological feature library. S3. Establish a target state monitoring sequence and perform cross-frame trajectory tracking of the target point to be determined; construct a temporal confidence evolution model, introduce a dynamic forgetting factor to evaluate the temporal continuity of the target point to be determined using a nonlinear sliding window; when the target encounters instantaneous physical occlusion that causes local detection loss, perform nonlinear confidence decay and use state prediction to maintain the tracking link; when the accumulated confidence reaches the target, the target point to be determined is confirmed as a valid new feature point. Specifically, S3 includes the following: Assign a unique global identifier to each target in the candidate new target set, and establish a status monitoring sequence. ,in, Indicates the record target at the 1st The frame's center pixel coordinates, bounding box size, and semantic matching score A time-domain confidence evolution model is constructed, specifying the initial confidence level when the target first enters the state monitoring sequence. The current confidence level of the target The evolutionary logic is as follows: When the target is in the current number When a frame is successfully associated, the confidence level is positively accumulated, and the specific mathematical expression is as follows: ;in, This is the confidence level accumulation coefficient. For the first Frame semantic matching score, This represents the upper limit of confidence. When the target is not detected in the current frame, the confidence decay mechanism is enabled: ;in, As a dynamic forgetting factor; when Below the extinction threshold When the target is determined to be dead, it is removed from the status monitoring sequence; In S3, during cross-frame trajectory tracking and when encountering physical occlusion, spatial offset is introduced to perform link maintenance and forgetting factor correction: Predicting the target in the next frame using the Kalman filter algorithm The target is located at the expected position in the target image. Within a preset search radius, the detection point with the highest overlap with the predicted position is found to achieve cross-frame association. The Euclidean distance between consecutive frames is calculated as the spatial offset. When a target encounters momentary physical occlusion and fails to find a successfully associated detection point within the search radius, the expected position predicted by the Kalman filter algorithm is directly used as the virtual pixel coordinates of the target in the current frame to maintain the tracking link until the confidence decays below the extinction threshold. until; The dynamic forgetting factor It is based on the target's current spatial offset. Dynamic correction is performed, and the specific formula is shown below: in, As the baseline forgetting factor, The spatial offset of the current frame. For scale parameters; S4. Monitor the geometric deformation of the reference graphic contour in the video stream in real time, extract the perspective distortion parameters under the current imaging plane by using geometric parametric fitting, and dynamically update the homography mapping matrix of the current frame relative to the initial frame to compensate for the pose offset of the image acquisition device under high dynamic transient disturbances; use the updated homography mapping matrix to perform spatial correction on the effective newly added feature points and map them to the virtual pixel coordinate system. S5. Perform density-based spatial clustering analysis on the effective new feature points mapped to the virtual pixel coordinate system, and remove non-systematic noise points that deviate from the main cluster; calculate the statistical centroid of the feature cluster, and solve the spatial geometric distribution envelope and directional standard deviation of the feature cluster, and output the quantitative analysis results.
2. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 1, characterized in that: Specifically, S1 includes the following: A deep convolutional neural network is used to perform multi-scale feature extraction and image segmentation on the initial frame image, and the mask data of the main region of the benchmark graphic is output. The outermost closed contour feature point set is obtained by edge detection operator. ,in, Indicates the first The pixel coordinates of each contour point in the original image coordinate system; Based on closed contour feature point set The geometric center point of the reference figure is determined by using the least squares method to fit a geometric circle. And extract the key feature points inside the benchmark image. Calculate the spatial positional correlation vector between the geometric center point and the internal key feature points. , serving as a self-calibration constraint for global coordinate mapping; Using the geometric center of the fitted reference figure Using the origin as the coordinate system origin and the major axis of the reference graphic's outline as the horizontal axis, a scale factor is established based on the ratio between the inherent physical dimensions of the reference graphic and the pixel dimensions. The scale factor The calculation formula is expressed as follows: ;in, The reference graphic has a pre-defined known physical dimension along its major axis. For the closed contour feature point set Extracted long axis pixel span; The pixel coordinate system of the original image is mapped to a normalized virtual pixel coordinate system through translation, rotation, and scaling transformations. The transformation relationship is expressed as follows: ; in, It is a normalized mapping matrix that includes rotation, translation, and scaling parameters. For coordinates in the normalized virtual pixel coordinate system, These are the pixel coordinates in the original image; The normalized mapping matrix The specific matrix parameter topology is described below: ; in, As a scale factor, The angle between the major axis of the reference graphic outline and the horizontal axis of the original image. The pixel coordinates of the geometric center point in the original image coordinate system.
3. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 1, characterized in that: Specifically, S2 includes the following: Get the current frame in the video stream With reference frame Background feature point pairs are used to calculate the global motion transformation matrix based on the random sampling consensus algorithm. For the reference frame Perform an affine transformation to align the pose shift caused by mechanical vibration, and obtain the compensated reference frame. ;in, Represents the real-time frame index of the video stream; For the current frame With the compensated reference frame Perform weighted difference operation to calculate grayscale variation. ;in, The adaptive weighting coefficient for illumination is calculated using the following formula: ; To exclude the set of background pixel coordinates after excluding dynamic edge regions, These are the pixel coordinates in the original image; Grayscale variation image Adaptive threshold binarization is performed, pixel clusters are extracted through connectivity analysis, and geometric feature filtering of pixel clusters is combined with prior feature size bounding boxes to obtain a set of undetermined target points that meet the area constraint conditions. ,in, Indicates the first The centroid position and bounding box parameters of the undetermined target point; For each undetermined target point Crop local feature windows to the center The input is fed into a pre-trained lightweight convolutional neural network, which uses multiple convolutional kernels to extract multi-dimensional morphological features, including edge gradients, hollowness, and contrast distribution, and calculates the output local feature window using the Softmax function. semantic matching score ,in, Indicates the index of the target point to be determined. This is the feature mapping function of a convolutional neural network. For network learnable parameters, and They represent the first The local feature window corresponding to the nth target point and the nth target point The semantic matching score calculated for each target point; It is used to quantitatively characterize the probability that a target point of unknown shape conforms to prior morphological features in terms of geometric appearance.
4. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 3, characterized in that: In step S2, the semantic matching score is calculated and output. The following spatial neighborhood cross-comparison steps are also included: Score the semantic matching degree Greater than the preset first threshold Target points are retained, and a set of high-scoring target points is constructed. and set the high-scoring target points The set of undetermined target points output by the time-domain dynamic difference path Perform cross-comparison of local spatial neighborhoods. The specific cross-comparison logic is expressed as follows: for Any high-scoring target point in With sets Any undetermined target point in Calculate the spatial Euclidean distance between the centroid coordinates of the two objects. and the intersection-union ratio of their bounding boxes. If and only if the following condition is met: and At that time, determine the high-scoring target point. and the target point to be determined Those belonging to the same spatial mutation source are output to the set of new candidate targets that simultaneously satisfy both temporal motion mutation features and static morphological semantic features; among them, The preset maximum neighborhood tolerance distance, This is the preset lower limit threshold for geometric overlap.
5. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 1, characterized in that: In step S3, the undetermined target point is confirmed as a valid new feature point, and the evaluation and judgment logic is as follows: Set the time observation window length to For each frame, a multi-dimensional consistency evaluation of targets within the sequence is performed. A pending target point is designated as a valid new feature point if and only if both of the following two feature constraints are simultaneously met: Spatial stability characteristics: Verify that the average spatial displacement of the target within the time observation window satisfies the mean constraint: ,in, For the first Spatial offset of a frame A preset spatial displacement deviation threshold is used to eliminate random high-frequency ionizing noise in the environment; Semantic persistence features: Verify that the average semantic score of the target within the time observation window satisfies the lower bound constraint: ,in, Indicates the first Frame semantic matching score, The minimum threshold for semantic matching score.
6. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 1, characterized in that: Specifically, S4 includes the following: In the current calibration frame In this process, local edge detection is performed based on the baseline graphic region determined by S1. The Canny operator is used to extract the edge pixel set of each level of concentric loops within the baseline graphic. Outlier noise points are removed by statistical filtering to obtain the cleaned geometric edge point set. ;in, This is the index of the current calibration frame; Using the random Hough transform on the geometric edge point set Perform multi-ellipse fitting to extract ellipse constraint parameters under the current imaging plane, including the center pixel coordinates. Long axis short axis and rotation angle The perspective distortion rate in the current imaging plane is defined using the ratio of the major and minor axes. The calculation formula is expressed as follows: ;in, and These represent the parameters of the major axis of the ellipse fitted from the current frame. With minor axis parameters The maximum and minimum values between.
7. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 6, characterized in that: Based on the aforementioned perspective distortion rate Perform hierarchical adaptive verification and reconstruct the dynamic homography matrix for spatial correction: According to the aforementioned perspective distortion rate Three vibration state intervals are divided for adaptive verification, and the specific determination is as follows: 1) When When a slight deformation of the reference image is detected, a "weak mechanical vibration" status label is output, and entry into the homography matrix calculation is allowed; among which, The preset first distortion threshold; 2) When When perspective distortion is detected in the reference image, a "high-frequency mechanical vibration" status label is output, and entry into homography matrix calculation is allowed; among which, This is the preset second distortion threshold; 3) When When the reference image is determined to have severe distortion, a "Severe Distortion" status label is output, and entry into the homography matrix calculation is allowed; among which, This is the preset distortion tolerance threshold. 4) The homography estimation of the current frame is deemed to have failed when any of the following triggering conditions are met: Condition 1, the perspective distortion rate ; Condition 2, the rotation tilt angle mentioned in the current frame The absolute deviation of the rotation angle relative to the initial frame reference angle exceeds the preset angular deformation threshold. ; Condition 3: The cleaned set of geometric edge points The total number of feature points participating in the fitting is lower than the preset minimum number of points threshold. ; When homography estimation is deemed to have failed, the homography matrix of the previous valid calibration frame is directly retained and reused. As the mapping parameter of the current calibration frame, it sends a pose reset self-test prompt to the terminal; After passing the above-mentioned hierarchical adaptive verification, using the correspondence of at least four sets of point pairs, the dynamic homography matrix of the current calibration frame relative to the initial reference frame is recalculated using a least-squares optimization algorithm. ,in, This is the index of the current calibration frame; Using the updated dynamic homography matrix Sub-pixel level spatial correction is performed on the effective newly added feature points. The corrected coordinate mapping logic is expressed through a two-dimensional homogeneous coordinate transformation as follows: ; in, These are the real-time detected homogeneous pixel coordinates in the current calibration frame. This is the inverse of the pose offset matrix of the current calibration frame, used to inversely project the pixel coordinates of the current calibration frame back to the reference pixel plane of the initial frame. The global static normalized mapping matrix described in S1 includes rotation, translation, and scaling parameters. These are the final physical homogeneous coordinates mapped to the normalized virtual pixel coordinate system.
8. The method for time-domain confidence dynamic target image recognition mounted on a sight according to claim 1, characterized in that: Specifically, S5 includes the following: After removing outliers from spatial clustering, the geometric center of the feature cluster is calculated as the statistical centroid. The specific calculation formula is as follows: , ,in, This represents the total number of valid new feature points included in the calculation after removing outlier feature points. Indicates the first The physical coordinates of each valid newly added feature point in the normalized virtual pixel coordinate system; Extract the mapping coordinates of the reference graphic center feature determined in S1 in the normalized virtual pixel coordinate system. and statistical centroid The relative offset vector is constructed as follows: The relative offset vector The modulus length and orientation quantitatively characterize systematic pose assembly deviations; Solve for the spatial geometric distribution envelope of the feature cluster, which is obtained by finding the minimum circumscribed circle radius that contains all valid new feature points. For quantitative characterization, the specific calculation formula is as follows: ;in, This indicates that the maximum value is obtained by iterating through all valid new feature points involved in the calculation. Calculate the standard deviation of the feature clusters in the horizontal direction respectively. Standard deviation in the vertical direction The specific calculation formula is as follows: , ; Based on the relative offset vector Minimum circumscribed circle radius with directional standard deviation , It outputs a quantitative analysis report on the spatial distribution pattern of feature clusters to the terminal; The specific logic for generating the quantitative analysis report is as follows: by analyzing the horizontal standard deviation... Standard deviation in the vertical direction The ratio relationship is used to automatically determine the main cause of error in high dynamic platforms: when the following conditions are met... When the horizontal dimension disturbance is determined to be the primary cause, then when the following conditions are met... At that time, the vertical dimension disturbance was determined to be the main cause; and targeted platform consistency calibration suggestions were output to the terminal; among them, The threshold for determining directional advantage.
Citation Information
Patent Citations
IMU (Inertial Measurement Unit)-assisted moving target detection method under dynamic condition
CN119273715A
Image segmentation and dynamic target identification method based on artificial intelligence
CN120689620A