Real-time image processing method and system for dynamic target recognition in video stream
By constructing a dual-channel feature extraction model to solve the visual reference center and dynamic target centroid in parallel, and combining Kalman filtering and optical flow methods, the problem of visual blind spots under high dynamic visual transitions is solved, realizing full-link quantitative tracing and accurate attribution of alignment status, and solving the problem of stability assessment of handheld optical imaging terminals.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 北京易悦科技有限公司
- Filing Date
- 2026-01-19
- Publication Date
- 2026-06-26
AI Technical Summary
Existing technologies cannot effectively handle the visual blind spots of handheld optical imaging terminals in highly dynamic visual transition scenarios, resulting in a discrepancy between the operator's subjective visual confirmation and the objective operational results, and making it impossible to accurately assess handheld stability.
A dual-channel feature extraction model based on attention mechanism and color space is constructed. By solving the visual reference center and dynamic target centroid in parallel, an alignment deviation vector flow is generated. Then, Kalman filtering and optical flow are used to generate virtual trajectories in the visual blind zone to repair the spatiotemporal causal chain break.
It fills the data sampling vacuum during high-dynamic visual transitions, establishes a high signal-to-noise ratio data foundation, accurately identifies the source of operational deviations, provides a visualized and quantitative basis for handheld stability assessment, and eliminates the deviation between the operator's subjective visual confirmation and the objective operation results.
Smart Images

Figure CN121937940B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically, to a real-time image processing method and system for dynamic target recognition in video streams. Background Technology
[0002] With the rapid development of computer vision and optoelectronic sensing technologies, real-time video stream processing has become a core technology in fields such as intelligent monitoring, industrial automated inspection, and high-precision skills training. In these applications, accurately extracting target features from continuous video frames and performing state analysis is crucial for achieving intelligent assisted decision-making. Existing image processing technologies primarily focus on improving target detection accuracy and processing speed in general scenarios, using deep learning models to perform feature mining and semantic understanding of video data to meet routine monitoring and recognition needs.
[0003] Existing technologies have made significant strides in video stream feature extraction and target recognition. For example, Chinese patent CN110569702B discloses a method and apparatus for processing video streams. It deploys multiple feature extraction models to extract features from each image frame in the video stream, obtaining feature vectors. These feature vectors are then combined to construct a feature matrix, which is processed using a pre-trained convolutional neural network model. This approach aims to reduce model annotation complexity and uncover the feature correlations in temporal video data. Furthermore, Chinese patent application CN109670488A discloses a method and system for detecting typical dynamic targets in video data. This scheme employs a distributed deep neural network computing platform framework. By constructing an image convolutional target segmentation model, it preprocesses and decodes the real-time video stream, achieving high-accuracy recognition and segmentation of typical objects such as people and vehicles in the current environment. This addresses, to some extent, the problems of weak traditional image recognition capabilities and poor real-time computation.
[0004] However, most of the aforementioned existing technologies are based on the ideal assumption that visual information is continuous and features are stable. They often prove inadequate when faced with the extreme scenarios of "high-dynamic visual transitions" in handheld optical imaging terminals (such as high-precision sniper aiming, long-focal-length industrial flaw detection, or microsurgical endoscopes). In actual high-precision operations, operators frequently encounter a situation where subjective visual confirmation indicates that the visual reference center precisely coincides with the distant target of interest at the moment of triggering the operation (such as pulling the trigger or pressing the shutter), but the final operational result shows a significant deviation. The underlying technical problem is that the physical impact (such as recoil or mechanical vibration) generated at the moment of triggering the operation causes the optical system to experience a "visual blind zone" of hundreds of milliseconds. During this period, the image sensor experiences motion blur due to rapid displacement, rendering conventional gradient-based optical flow methods ineffective and creating a "vacuum period" for data sampling. Existing general image processing algorithms lack the ability to differentiate the imaging degradation characteristics of dual-source heterogeneous features (relatively static visual reference markers and dynamically changing distant targets of interest). They not only fail to maintain logical trajectory continuity across this physical blind spot, but also ignore the deconstruction of frequency domain features. This results in the inability to decouple the operator's inherent physiological tremors from the instantaneous operational jitters caused by psychological tension or stiff movements in the time-series signal, thus causing the handheld stability assessment to fall into a technical dilemma of being unable to accurately attribute causes. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of existing technologies, this invention provides a real-time image processing method and system for dynamic target recognition in video streams. It constructs a dual-channel feature extraction model based on an attention mechanism and color space, calculates the visual reference center and the centroid of the dynamic target in parallel, and utilizes a mechanism combining Kalman filtering and optical flow to generate virtual trajectories within blind zones when highly dynamic visual transitions are detected. This invention effectively fills the data sampling vacuum period caused by physical impact at the moment of triggering, repairs the spatiotemporal causal chain break in discontinuous visual information streams, and achieves full-link quantitative tracing and accurate attribution of visual alignment behavior from physiological micro-movements to the moment of triggering.
[0006] To achieve the above objectives, the present invention provides the following technical solution:
[0007] Real-time image processing methods for dynamic target recognition in video streams include:
[0008] A video stream buffer queue is established for a handheld optical imaging terminal to acquire monocular raw video frames. A dual-channel feature extraction model based on attention mechanism and color space is constructed. The visual reference center coordinates and dynamic target centroid coordinates are solved in parallel from the monocular raw video frames through the dual-channel feature extraction model, and the target state vector is generated.
[0009] An alignment deviation vector flow is generated based on the coordinates of the visual reference center and the centroid coordinates of the dynamic target. The physiological tremor characteristics and operational shaking characteristics of the operator are deconstructed based on the alignment deviation vector flow to generate visual alignment situation data.
[0010] The system monitors the target state vector in real time and switches to the visual blind zone state when a highly dynamic visual transition is detected, generating a dynamic virtual trajectory of the target within the blind zone. After exiting the visual blind zone state, the system generates a full-process alignment trajectory based on the dynamic virtual trajectory of the target and spatially maps and fuses the full-process alignment trajectory with the visual alignment situation data to generate visual guidance instructions.
[0011] The method for calculating the coordinates of the visual reference center includes:
[0012] The original monocular video frame is converted into an HSV color space image, and morphological filtering is performed on the HSV color space image to generate a standardized HSV preprocessed image frame.
[0013] Threshold segmentation of the HSV color geometric subspace is performed on the standardized HSV preprocessed image frame to extract the visual reference identifier mask. Connectivity analysis and geometric moment calculation are performed on the visual reference identifier mask to output the coordinates of the visual reference center.
[0014] The method for calculating the dynamic target centroid coordinates includes:
[0015] Standardized HSV preprocessed image frames are input into a lightweight convolutional neural network with a CBAM attention mechanism. Pyramid feature extraction is performed to locate distant targets of interest. The target bounding box and target confidence score output by the lightweight convolutional neural network are obtained. Based on the target bounding box of the distant target of interest, the dynamic target centroid coordinates and estimated depth distance are calculated. The distant target of interest includes a target.
[0016] The target state vector is obtained by encapsulating the dynamic target centroid coordinates, estimated depth distance, and target confidence score.
[0017] The method for generating the alignment deviation vector flow includes:
[0018] Using the coordinates of the visual reference center as the origin and the coordinates of the dynamic target centroid in the target state vector as the moving point, the alignment deviation distance and alignment deviation phase are calculated. Combined with the acquisition timestamps of the original monocular video frames, a time sequence is constructed to generate the alignment deviation vector stream.
[0019] The deconstruction methods for the physiological tremor features and operational tremor features include:
[0020] The alignment deviation vector flow within the preset sliding time window is subjected to fast Fourier transform for spectral decomposition, separating the low-frequency energy component from the high-frequency abrupt component. The low-frequency energy component is used as the operator's physiological tremor characteristic, and the high-frequency abrupt component is used as the operator's operational jitter characteristic.
[0021] The method for generating the visual alignment situation data includes:
[0022] The tremor classification result is determined based on the energy ratio of the physiological tremor feature set and the operational tremor feature set;
[0023] The physiological tremor feature set and the operational jitter feature set are input into a pre-trained LSTM time series analysis model to identify the operational intention stage, which includes the intentional precision aiming stage and the moment of triggering the operation. In the intentional precision aiming stage, the handheld stability score is calculated based on the proportion of time that the dynamic target centroid coordinates fall within the circle centered on the visual reference center coordinates. The operational intention stage, the handheld stability score, and the tremor classification results are encapsulated into visual alignment situation data.
[0024] The method for detecting high dynamic visual transitions includes:
[0025] The global motion vector is calculated by sparse optical flow for two adjacent normalized HSV preprocessed image frames. When the target confidence score is less than or equal to the preset loss judgment threshold or the magnitude of the global motion vector exceeds the preset inter-frame displacement threshold, a high dynamic visual transition is detected.
[0026] The method for generating the dynamic target virtual trajectory includes:
[0027] In response to the visual blind spot state, the target state vector and global motion vector from the previous moment are used as inputs to the Kalman filter to perform inertial prediction and background compensation, and to calculate the dynamic target virtual trajectory within the blind spot.
[0028] The method for generating the alignment trajectory throughout the entire process includes:
[0029] The real detection trajectory, which is formed by the virtual trajectory of the dynamic target and the centroid coordinates of the dynamic target in the target state vector, is time-stamped and smoothly stitched together to form a full-process alignment trajectory; the real detection trajectory includes two parts: the real detection trajectory before entering the visual blind zone state and the real detection trajectory after exiting the visual blind zone state.
[0030] A real-time image processing system for dynamic target recognition in video streams, used to implement the aforementioned real-time image processing method for dynamic target recognition in video streams, the system comprising:
[0031] Dual-channel feature extraction module: used to establish a video stream buffer queue for handheld optical imaging terminal, acquire monocular raw video frames, construct a dual-channel feature extraction model based on attention mechanism and color space, and solve the visual reference center coordinates and dynamic target centroid coordinates in parallel from monocular raw video frames through the dual-channel feature extraction model, and generate target state vector;
[0032] Visual alignment posture module: Generates alignment deviation vector flow based on the coordinates of the visual reference center and the centroid coordinates of the dynamic target, and deconstructs the operator's physiological tremor characteristics and operational jitter characteristics based on the alignment deviation vector flow to generate visual alignment posture data;
[0033] Visual guidance module: It is used to monitor the target state vector in real time. When a high dynamic visual transition is detected, it switches to the visual blind zone state and generates a dynamic target virtual trajectory within the blind zone. After exiting the visual blind zone state, it generates a full-process alignment trajectory based on the dynamic target virtual trajectory, and performs spatial mapping and fusion with the visual alignment situation data to generate visual guidance instructions.
[0034] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0035] This invention effectively overcomes the heterogeneity problem of imaging degradation characteristics between static visual reference markers and dynamic distant targets of interest under uncontrolled lighting and complex backgrounds by constructing a dual-channel feature extraction model based on attention mechanisms and color space. It also solves the feature extraction failure problem caused by traditional single-channel algorithms when processing two types of targets with significantly different features, establishing a high signal-to-noise ratio data foundation. Based on the deconstruction mechanism of alignment deviation vector flow, this invention can deeply decouple the operator's inherent physiological tremor and sudden operational jitter in the time domain signal, breaking through the limitation of traditional methods that rely solely on spatial geometric distance while ignoring frequency domain features, leading to temporal logic aliasing. This allows for the identification of the source of operational deviation as... The invention accurately identifies whether the problem stems from "physical limitations" or "skill errors." More importantly, it addresses the highly dynamic visual transitions and blind spots caused by physical impact at the moment of triggering the operation. By generating a virtual trajectory of the dynamic target within the blind spot and constructing a full-process alignment trajectory, the invention successfully fills the data sampling vacuum caused by motion blur in the image sensor due to rapid displacement, repairs the break in the spatiotemporal causal chain in the discontinuous visual information flow, and thus objectively reproduces the true alignment state at the moment of triggering. This eliminates the deviation between the operator's subjective visual confirmation and the objective operation result, providing a visually quantitative basis with temporal continuity and logical completeness for handheld stability assessment. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart of a real-time image processing method for dynamic target recognition in a video stream provided in an embodiment of the present invention;
[0038] Figure 2 This is a schematic diagram of a dual-channel feature extraction architecture provided in an embodiment of the present invention;
[0039] Figure 3 This is a schematic diagram of the imaging field of view of a shooting training scenario provided in an embodiment of the present invention;
[0040] Figure 4 A flowchart illustrating the method for determining tremor classification results provided in an embodiment of the present invention;
[0041] Figure 5 A time series diagram of the alignment trajectory throughout the entire process provided in an embodiment of the present invention;
[0042] Figure 6 This is a functional block diagram of a real-time image processing system for dynamic target recognition in a video stream, provided in an embodiment of the present invention. Detailed Implementation
[0043] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0044] Example 1
[0045] Please see Figure 1 As shown, this embodiment provides a real-time image processing method for dynamic target recognition in a video stream, including:
[0046] Step S10: Obtain the original monocular video frame, construct a dual-channel feature extraction model based on attention mechanism and color space, and solve the coordinates of the visual reference center and the centroid of the dynamic target in parallel from the original monocular video frame through the dual-channel feature extraction model, and generate the target state vector.
[0047] The core task of step S10 is to simultaneously extract two types of coordinate information with significant feature differences from the monocular raw video frames acquired by the handheld optical imaging terminal: one type is the relatively static coordinates of the visual reference center, and the other type is the dynamically changing centroid coordinates of the distant target of interest, i.e., the dynamic target centroid coordinates. The dual-channel feature extraction model refers to two independent feature processing channels built for the two types of coordinate information, which can run in parallel. Please refer to... Figure 2 , Figure 2 The diagram illustrates the data flow and logical composition of the dual-channel feature extraction architecture. For example... Figure 2 As shown, this architecture uses a "standardized HSV preprocessed image frame" as a shared input node and logically divides into two paths: a processing path labeled "first channel" and a processing path labeled "second channel," which operate in parallel. Specifically, the first channel deploys a "geometric segmentation algorithm based on the HSV color space." After data flows through this channel, it is directed to the "visual reference center coordinates" at the output end. This process is specifically used to extract the center position of the inherent visual reference markers on the optical components inside the handheld optical imaging terminal. The second channel deploys a "lightweight convolutional neural network with embedded CBAM attention mechanism." After data flows through this channel, it is directed to the "dynamic target centroid coordinates" at the output end. This process is specifically used to detect and locate the centroid position of distant targets of interest in the field of view. The two channels share the preprocessed standardized HSV preprocessed image frame as input, but each uses completely different algorithmic logic for feature extraction. The design motivation for this dual-channel parallel architecture stems from the fundamental differences in image features between visual reference markers and distant targets of interest: visual reference markers typically have fixed tonal characteristics and remain relatively stationary relative to the imaging sensor, making them suitable for fast segmentation methods based on color thresholds; distant targets of interest, on the other hand, may have complex texture features, uncertain color distributions, and dynamic motion characteristics relative to the field of view, requiring the semantic understanding capabilities of deep learning networks for recognition. Separating the two algorithmic logics into independent channels allows each channel to be specifically optimized for the characteristics of its target, avoiding the performance trade-offs that arise when a single algorithm handles two types of targets with vastly different features.
[0048] Further, step S10 includes:
[0049] Step S11: Establish a video stream buffer queue for the handheld optical imaging terminal, acquire monocular raw video frames with acquisition timestamps, convert the monocular raw video frames into HSV color space images, perform morphological filtering on the HSV color space images, and generate standardized HSV preprocessed image frames.
[0050] A handheld optical imaging terminal refers to an optical imaging device installed at the back end of an operating platform, possessing real-time image acquisition and transmission capabilities. For example, in a shooting training scenario, the handheld optical imaging terminal could be a miniature camera module mounted on the eyepiece of a sight, which acquires the image inside the sight through optical coupling; in an ecological photography scenario, the handheld optical imaging terminal could be the video output interface of an electronic viewfinder of a DSLR camera; in an industrial inspection scenario, the handheld optical imaging terminal could be the digital image output port of a handheld infrared thermal imager. The video stream buffer queue is a first-in, first-out (FIFO) data structure used to temporarily store multiple consecutively acquired image frames, balancing the temporal differences between the image acquisition rate and subsequent processing rate, ensuring that keyframe data is not lost when the processing unit load fluctuates. The handheld optical imaging terminal's image acquisition unit reads monocular raw video frames in real time. Each monocular raw video frame is marked with a unique acquisition timestamp when entering the buffer queue. This timestamp records the precise moment of image acquisition, providing a time reference for the subsequent step S21 to construct the temporal sequence.
[0051] After entering the processing flow, the raw monocular video frames undergo color space conversion. The raw image format used is typically the RGB color model, which describes pixel color by superimposing three components: red, green, and blue channels. In the RGB model, color information and brightness information are highly coupled, and the RGB values of the same object will change significantly under different lighting conditions. For example, a red target may show a high red channel reading under strong light and a low red channel reading under shadow, although the human eye can still identify them as the same red target. This coupling characteristic makes the target segmentation algorithm based on RGB thresholds unstable in environments with changing lighting. Each frame of the raw monocular video is mapped to an HSV color space image through a nonlinear transformation. The HSV color space describes color using three independent dimensions: hue, saturation, and lightness. The hue dimension represents the basic type of color, and its value is distributed along a circle. The saturation dimension represents the purity of the color; a higher value indicates a more vivid color, and a lower value indicates a color closer to gray. The lightness dimension represents the brightness of the color; a higher value indicates a brighter color, and a lower value indicates a darker color. The RGB to HSV conversion follows the standard color space mapping formula. After the conversion, changes in illumination mainly affect the value of the lightness dimension, while the hue and saturation dimensions remain relatively stable. This characteristic allows the subsequent step S12 to achieve robustness to illumination changes by locking the threshold range of hue and saturation and widening the threshold range of lightness when extracting visual reference icons for specific hues.
[0052] HSV color space images require morphological filtering before further processing. A morphological opening operation is performed on the HSV color space image, consisting of an erosion operation and a dilation operation in sequence. The erosion operation uses a structuring element of a preset size to slide across the image, replacing the pixel values within the area covered by the structuring element with the minimum value of that area. This operation removes isolated bright spots smaller than the structuring element; for example, noise caused by the refraction of light by airborne dust particles and random bright spots formed by thermal noise from the image sensor are eliminated by the erosion operation. The dilation operation then uses the same structuring element to slide across the eroded image, replacing the pixel values within the area covered by the structuring element with the maximum value of that area. The dilation operation restores the target edges that were cut off by the erosion operation, bringing the target contour back to near its original size, while filling in small pores inside the target. The size of the structuring element is determined based on the relationship between the typical diameter of the noise and the minimum feature size of the target: the side length of the structuring element should be greater than the typical noise diameter to ensure effective noise removal, while it should be less than half the minimum feature size of the target to avoid over-erosion of the target edges. For example, when the diameter of a typical noise point is about three pixels and the line width of a visual reference identifier is about ten pixels, the side length of the structuring element can be set to five pixels. The image output after morphological opening is a standardized HSV preprocessed image frame. This image frame has eliminated random noise interference and has smooth edge contours, providing high-quality input data for threshold segmentation in step S12 and neural network inference in step S13.
[0053] Step S11 performs color space conversion and morphological filtering as a unified preprocessing step, allowing subsequent dual-channel processing to share the same standardized data and avoiding the redundant computational overhead caused by separate preprocessing for each channel. The synchronized timestamps ensure that the visual reference center coordinates and the dynamic target centroid coordinates have a precise temporal correspondence when constructing the time sequence in subsequent step S21. This correspondence is a prerequisite for calculating the alignment deviation vector flow. Without the preprocessing in step S11, the color thresholding segmentation in step S12 would directly apply to the noisy RGB image. Changes in illumination would cause drastic fluctuations in the RGB values of the visual reference marker, making it difficult to balance different illumination conditions and resulting in numerous false detection areas in the segmentation results. Furthermore, the neural network input in step S13 would contain random noise, requiring the network to consume additional computational resources to learn noise patterns, thus reducing both inference speed and detection accuracy.
[0054] Step S12: Perform threshold segmentation of HSV color geometric subspace on the standardized HSV preprocessed image frame, extract the visual reference identifier mask, perform connected component analysis and geometric moment calculation on the visual reference identifier mask, and output the coordinates of the visual reference center.
[0055] Visual reference markers refer to geometrical markers that are fixedly installed inside the optical system of a handheld optical imaging terminal and are captured along with the image. For example, see [link to relevant documentation]. Figure 3 , Figure 3 This is a schematic diagram of the imaging field of view for a shooting training scenario provided in an embodiment of this application. In the shooting training scenario, the visual reference marker can be the red crosshair on the reticle of a scope; in the photographic assistance scenario, the visual reference marker can be the autofocus frame in the center of the electronic viewfinder; in the medical endoscopy scenario, the visual reference marker can be the surgical instrument action point marker superimposed on the center of the image. The design of the visual reference marker typically uses a color with a significant tonal difference from the scene background to facilitate separation by image processing algorithms.
[0056] Thresholding of the HSV color geometric subspace is performed on standardized HSV preprocessed image frames. The HSV color space can be viewed as a cylindrical geometric structure, with the circumference corresponding to the hue dimension, the radius corresponding to the saturation dimension, and the height corresponding to the lightness dimension. The color geometric subspace refers to the closed region within this cylinder enclosed by threshold boundaries. Based on the color characteristics of pre-calibrated visual reference markers, six parameters are set: lower hue threshold, upper hue threshold, lower saturation threshold, upper saturation threshold, lower lightness threshold, and upper lightness threshold. These six parameters define a fan-shaped truncated space within the HSV cylinder. The calibration method for the six threshold parameters is as follows: under controlled lighting conditions, a standard image containing the visual reference markers is acquired; the HSV three-dimensional distribution of pixels in the visual reference marker region is extracted; and the boundary values of the distribution range are taken and appropriately expanded as threshold parameters. Iterate through each pixel of the standardized HSV preprocessed image frame and determine whether the hue value, saturation value, and brightness value of the pixel fall within the corresponding threshold range. If all three conditions are met, mark the pixel as foreground (set the value to one), otherwise mark it as background (set the value to zero), thereby generating a binarized visual reference mask.
[0057] After the visual reference icon mask is generated, connected component analysis and geometric moment calculation are required. Connected component analysis refers to the operation of aggregating adjacent foreground pixels in the mask into independent regions, using the eight-neighbor connectivity criterion, meaning that each pixel and its eight neighboring pixels in the surrounding directions are considered potentially connected. After connected component analysis, a list of all independent connected regions in the mask can be obtained, with each connected region corresponding to a potential visual reference icon component. For a cross-shaped visual reference icon, connected component analysis usually yields a primary cross-shaped connected region. The first-order geometric moment is calculated for this connected region to determine its centroid position. The calculation of the first-order geometric moment follows the centroid formula. Since the visual reference icon mask is a binary image, the weight of each foreground pixel is one. The centroid's x-coordinate is equal to the sum of the x-coordinates of all foreground pixels in the connected region divided by the total number of foreground pixels, and the centroid's y-coordinate is equal to the sum of the y-coordinates of all foreground pixels in the connected region divided by the total number of foreground pixels. The resulting coordinates are the geometric centroid of the connected region. Because handheld optical imaging terminals experience high-frequency micro-vibrations during operation, the projection of physical scale lines onto the imaging sensor produces sub-pixel-level inter-frame drift. Directly using preset fixed coordinates cannot reflect this dynamic change. The centroid coordinates obtained through geometric moment calculation can adaptively track the actual position of the visual reference marker in the image; this position is the visual reference center coordinate, denoted as C. ref Its representation is a pair of horizontal and vertical coordinate values in a two-dimensional image coordinate system.
[0058] Step S12 establishes a baseline integrity judgment mechanism to handle abnormal situations such as occlusion or overexposure of the visual reference marker. The total number of foreground pixels in the visual reference marker mask is calculated as the mask area. A baseline integrity threshold is set, determined based on the typical pixel area of the visual reference marker under normal imaging conditions. For example, if the pixel area of the visual reference marker under normal conditions is approximately 500 pixels, the baseline integrity threshold can be set to 200 pixels, meaning that an anomaly is determined when the detected area is less than 40% of the normal value. When the mask area is less than the baseline integrity threshold, possible reasons include strong light causing the scale lines to merge with the background and become indistinguishable, occlusions entering the field of view and covering the scale lines, and the threshold parameter failing under extreme lighting conditions. In this case, the visual reference marker detection for the current frame is determined to have failed, and the visual reference center coordinates cached in the previous frame are forcibly used as the output value for the current frame. This fault-tolerance mechanism ensures that the subsequent step S21 always has a valid coordinate origin when calculating the alignment deviation vector, avoiding interruption of the entire data stream due to a single-frame detection failure. The cache update strategy is as follows: the newly calculated visual reference center coordinates are written to the cache only when the current frame is successfully detected, replacing the old value; the cached value remains unchanged when detection fails. Step S12 uses a color threshold-based segmentation method instead of a deep learning method to extract visual reference icons because the features of visual reference icons are highly deterministic. The hue, shape, and approximate position of the visual reference icons in the image are all known prior information. Using a parameterized threshold segmentation method can fully utilize these prior constraints to obtain extremely high segmentation accuracy and extremely low computational latency. In contrast, deep learning methods require additional computational resources for network inference, which can cause unnecessary computational redundancy for targets with well-defined features.
[0059] Step S13: Input the standardized HSV preprocessed image frame into a lightweight convolutional neural network with CBAM attention mechanism, perform pyramid feature extraction to locate the distant interest target, obtain the target bounding box and target confidence score output by the lightweight convolutional neural network, calculate the dynamic target centroid coordinates and estimated depth distance based on the target bounding box of the distant interest target, and encapsulate the dynamic target centroid coordinates, estimated depth distance and target confidence score into a target state vector.
[0060] A remote target of interest refers to an object located within the field of view of a handheld optical imaging terminal, which the operator intends to visually align with. For example, see [link to example]. Figure 3In shooting training scenarios, the distant target of interest is the target itself; in ecological photography scenarios, the distant target of interest is flying birds or running animals; in industrial inspection scenarios, the distant target of interest is an abnormally hot spot on a power transmission tower; and in surveying scenarios, the distant target of interest is a feature corner point of a distant building. Common characteristics of distant targets of interest are that they occupy only a small pixel area in a wide field of view, may shift in position over time, and their appearance is significantly affected by ambient lighting and background interference.
[0061] Standardized HSV preprocessed image frames are used as input and fed into an improved YOLOv8-Nano deep neural network for inference. YOLOv8-Nano is a lightweight version of the YOLO series of object detection networks. Its network architecture has undergone parameter compression and structural optimization for edge computing devices, achieving low inference latency while maintaining high detection accuracy. The choice of a lightweight convolutional neural network is based on the computational resource constraints of handheld optical imaging terminals: the embedded processing units typically equipped in handheld devices are limited in terms of computing power and power consumption. Using a network with too many parameters would lead to inference latency exceeding real-time requirements or excessive power consumption resulting in insufficient device battery life.
[0062] To address the challenge of detecting distant targets of interest, which occupy a very small pixel area in a wide field of view, the system embeds a CBAM module into the backbone feature extraction layer of the YOLOv8-Nano network. CBAM, short for Convolutional Block Attention Module, consists of cascaded channel attention subunits and spatial attention subunits. The mechanism of the channel attention subunit is as follows: The feature map of the neural network consists of multiple channels, each encoding a certain feature response of the input image. For example, some channels may be sensitive to edge features, some to texture features, and some to specific shapes. The channel attention subunit compresses the spatial dimension of the feature map through global average pooling and global max pooling, and then learns the importance weights of each channel through a multilayer perceptron, finally weighting each channel in the form of a weight vector. This mechanism enables the network to adaptively enhance the channel responses containing target discriminative features and suppress the channel responses containing background noise. The spatial attention subunit operates as follows: It performs average pooling and max pooling along the channel dimension on the channel-weighted feature map, generating two single-channel feature maps. These maps are then concatenated and passed through a convolutional layer to generate a spatial attention weight map. Each position in this weight map corresponds to the importance weight of the corresponding spatial position in the original feature map. The spatial attention subunit enables the network to focus on the spatial region where the target is located in the feature map, suppressing interference from the background. For example, when a distant target of interest occupies only a small area in the center of the image while there is background motion such as swaying leaves or flowing clouds, the spatial attention mechanism will assign high weights to the target region and low weights to the background region. The placement of the CBAM module in the backbone network is determined based on the hierarchical structure of the feature pyramid: the CBAM module is embedded in the transition layer between shallow feature maps (high resolution, low semantics) and deep feature maps (low resolution, high semantics), allowing the attention mechanism to simultaneously apply to detailed and semantic features, enhancing the detection capability of distant, small targets.
[0063] Pyramid feature extraction refers to the mechanism by which a network extracts target information from feature maps of different scales. During forward propagation, the neural network generates a sequence of feature maps with progressively decreasing resolution through multiple downsampling operations, forming a pyramid-like hierarchical structure. High-resolution feature maps preserve the spatial details of the input image, making them suitable for detecting small targets; low-resolution feature maps have a larger receptive field and richer semantic information, making them suitable for detecting large targets. YOLOv8-Nano employs a path aggregation network structure that fuses feature maps from different levels, enabling the detection head to utilize multi-scale information simultaneously. After feature extraction and fusion, the network outputs the target's bounding box and target confidence score through the detection head. The bounding box describes the target's position and extent in the image as a rectangular region, containing four parameters: the top-left horizontal coordinate, the top-left vertical coordinate, width, and height. The target confidence score is a value between 0 and 1, representing the network's judgment on the validity of the current detection result; a higher value indicates that the network is more confident in detecting a real target rather than a false positive. The confidence score is calculated by the network's loss function, comprehensively considering both the probability of target existence and the accuracy of bounding box localization.
[0064] The dynamic target centroid coordinates and estimated depth distance are calculated based on the target bounding box output by the network. The dynamic target centroid coordinates are denoted as C. obj The calculation method involves taking the coordinates of the geometric center point of the bounding box, specifically the upper left corner's x-coordinate plus half the width, and the upper left corner's y-coordinate plus half the height. This simplification assumes the target is a uniformly distributed rigid body, with its geometric center approximately equal to its center of mass. The estimated depth distance is calculated based on the geometric principle of similar triangles. The system pre-sets the target's physical size parameters; for example, in a shooting training scenario, the true width of the target is pre-set, and in a surveying scenario, the true height of the prism is pre-set. Let the target's true width be Wreal, the bounding box's pixel width in the image be Wpixel, and the equivalent focal length of the handheld optical imaging terminal be fpixel (in pixels). Then, the estimated depth distance Dest between the target and the imaging device can be calculated using the formula Dest = (fpixel × Wreal) / Wpixel. The derivation of this formula is based on the fact that the target, the lens optical center, and the imaging sensor form similar triangles, and the ratio of the target's true size to its pixel size is equal to the ratio of the target distance to the focal length. The equivalent focal length parameter is pre-determined and fixed in the system through a calibration procedure. The calibration method involves placing a calibration object of known size at a known distance, acquiring the image, and then calculating the focal length value in reverse.
[0065] The coordinates of the centroid of the dynamic target are C. objThe estimated depth distance (Dest) and target confidence score are encapsulated into a target state vector for output. The target state vector is a data structure containing multi-dimensional information, designed to provide a complete target description for subsequent steps. Step S13 and Step S12 form a dual-channel parallel processing architecture. Both steps share the standardized HSV preprocessed image frame output from Step S11 as input, but each employs an algorithm optimized for its target: Step S12 uses low-computation thresholding to process visually clear reference markers, while Step S13 uses a neural network with strong semantic understanding to process distant interest targets with complex features. This division of labor allows the dual channels to simultaneously output two types of coordinate information within a single processing cycle, meeting real-time requirements while maintaining detection accuracy. If the detection of two types of targets were merged into a single channel, either a complex algorithm capable of simultaneously handling two types of targets with vastly different features would be required, leading to a sharp increase in computation, or a trade-off in algorithm parameters would result in a decrease in detection accuracy for both types of targets. The dual-channel architecture avoids these dilemmas, allowing each channel to reach peak performance in its focused task.
[0066] Step S10 constructs a dual-channel feature extraction model based on an attention mechanism and color space to parallelly calculate the coordinates of the visual reference center and the centroid coordinates of the dynamic target from the original monocular video frames. This step first uses HSV color space conversion to achieve dimensional decoupling of color and brightness information, and combines morphological opening operations to remove random noise and repair edges, effectively overcoming the interference of outdoor lighting fluctuations and imaging noise on detection stability, and providing standardized data with high signal-to-noise ratio for subsequent processing. On this basis, the first channel uses the color prior of the visual reference mark, and through color geometric subspace threshold segmentation and geometric moment calculation, achieves sub-pixel-level tracking of the mark position with extremely low computational latency, and ensures the continuity and stability of the coordinate origin by means of a benchmark integrity judgment mechanism. At the same time, the second channel embeds the CBAM attention mechanism in a lightweight convolutional neural network, and enhances the target feature response and suppresses background dynamic noise through adaptive weighting of channels and spatial dimensions. Combined with multi-scale fusion of feature pyramids, it significantly improves the recall rate and localization accuracy of distant small targets in a wide field of view. This dual-channel parallel architecture, which separates and processes heterogeneous targets, effectively avoids the performance trade-offs that occur when a single algorithm considers two types of targets with vastly different characteristics. It achieves temporal synchronization and spatial alignment of the two types of coordinate data while meeting real-time requirements, laying a precise data foundation for subsequent temporal micro-motion trajectory analysis and motion compensation processing.
[0067] Step S20: Generate alignment deviation vector flow based on the coordinates of the visual reference center and the centroid coordinates of the dynamic target; deconstruct the operator's physiological tremor characteristics and operational jitter characteristics based on the alignment deviation vector flow to generate visual alignment situation data.
[0068] The core task of step S20 is to transform the basic alignment coordinate data output from step S10—namely, the coordinates of the visual reference center and the centroid of the dynamic target—into an alignment deviation vector stream with temporal characteristics. Through joint analysis in the frequency and temporal domains, the overall tremor signal generated during the operator's hand-held operation is deconstructed into two independent dimensions: physiological tremor characteristics and operational jitter characteristics. Ultimately, this generates visual alignment posture data that can quantify visual alignment behavior. The alignment deviation vector stream refers to a data sequence composed of two-dimensional deviation vectors arranged in chronological order, with the visual reference center coordinates as the origin and the dynamic target centroid coordinates as the moving point. This sequence completely records the micro-movement trajectory of the distant target of interest relative to the visual reference center during visual alignment. Physiological tremor characteristics refer to the signal component formed by the inherent frequency micro-vibrations of the operator's skeletal muscles when maintaining a static posture. This component reflects the operator's physical condition and fitness level. Operational jitter characteristics refer to the instantaneous high-frequency disturbance signal component caused by psychological tension, movement errors, or external impacts at the moment of triggering the operation. This component reflects the operator's movement standardization and skill proficiency. Visual alignment posture data is a composite data structure that encapsulates the operational intent stage, handheld stability score, and tremor classification results, providing a quantitative basis for the generation of visual guidance instructions in subsequent step S30. The design motivation for step S20 stems from the fact that traditional handheld stability assessment methods can only obtain the result data after the trigger operation is completed, lacking the ability to quantitatively analyze the dynamic alignment process before the trigger operation. This makes it impossible for assessors to distinguish whether operational deviations are due to insufficient physical fitness or a lack of motor skill. By using frequency domain decomposition to separate the overall tremor signal into two independent dimensions—physiological and operational—the assessment feedback can be specifically directed towards improvement, significantly improving the attribution accuracy of handheld stability assessment.
[0069] Further, step S20 includes:
[0070] Step S21, see Figure 4 Using the coordinates of the visual reference center as the origin and the coordinates of the dynamic target centroid in the target state vector as the moving point, the alignment deviation distance and alignment deviation phase are calculated. Combined with the acquisition timestamps of the original monocular video frames, a time sequence is constructed to generate the alignment deviation vector stream.
[0071] Visual reference center coordinates C ref The geometric centroid position of the internal visual reference marker of the handheld optical imaging terminal's optical system in the current frame image is represented as a pair of horizontal and vertical coordinates in a two-dimensional image coordinate system, denoted as C. ref =(x ref ,y ref ), where x ref Let y be the x-coordinate of the visual reference center.ref The vertical coordinate is the coordinate of the visual reference center. The coordinates of the dynamic target's centroid are C. obj The geometric center position of the distant target of interest in the current frame image is represented by pairs of horizontal and vertical coordinates in a two-dimensional image coordinate system, denoted as C. obj =(x obj ,y obj ), where x obj Let y be the x-coordinate of the centroid of the dynamic target. obj Let be the ordinate of the centroid of the dynamic target. Using the visual reference center coordinates C... ref Establishing a local coordinate system with the origin at point C allows subsequent calculations of the deviation to directly reflect the offset of the distant target of interest relative to the operator's visual reference. ref The reason it's not the image coordinate system origin is because the image coordinate system origin is located at the top left corner of the image, and it has no direct geometric relationship with the operator's visual alignment behavior; while the visual reference center coordinates C... ref The reference position for visual alignment when the operator performs visual alignment using a handheld optical imaging terminal is used as the origin so that the deviation vector can directly correspond to the operator's subjective alignment error.
[0072] The alignment deviation vector is calculated using vector subtraction, starting from the visual reference center coordinates C. ref Pointing to the centroid coordinates C of the dynamic target obj The direction is defined as the positive direction of the deviation vector, and the alignment deviation vector V can be expressed as C. obj -C ref Its lateral component V x =x obj -x ref Its longitudinal component V y =y obj -y ref The magnitude of the alignment deviation vector V is the alignment deviation distance, calculated using the Euclidean distance formula. The physical meaning of the alignment deviation distance is the pixel distance between the centroid of the distant target of interest and the visual reference center; a larger distance indicates a more significant visual alignment deviation by the operator. The argument of the alignment deviation vector V is the alignment deviation phase θ, calculated using the arctangent function. The calculation needs to be based on V. x With V yThe sign of the argument determines the quadrant in which it falls. The physical meaning of the alignment deviation phase θ is the angle between the deviation direction and the positive direction of the horizontal axis of the image coordinate system. This phase value can reflect the operator's alignment deviation tendency. For example, in a shooting training scenario, if an operator's alignment deviation phase is consistently concentrated around ±90 degrees, it indicates that there is a systematic deviation in the vertical direction of their visual alignment, which may be related to their breathing rhythm; if the alignment deviation phase is concentrated around 0 degrees or 180 degrees, it indicates that there is a systematic deviation in the horizontal direction of their visual alignment, which may be related to their shooting posture.
[0073] Step S21 combines the alignment deviation distance and alignment deviation phase θ calculated for each frame with the acquisition timestamp marked in step S11 to construct a triplet data set, and pushes it into the first-in-first-out data queue in chronological order to form an alignment deviation vector stream. The data structure of the alignment deviation vector stream can be represented as a time series set, where each element is a triplet containing three components: acquisition timestamp, alignment deviation distance, and alignment deviation phase. The acquisition timestamp is marked in step S11 when the monocular raw video frame enters the video stream buffer queue, ensuring a precise correspondence between the visual reference center coordinates and the dynamic target centroid coordinates in the time dimension. The alignment deviation value at a single moment only reflects the visual alignment state at that instant and lacks the completeness of information for behavioral analysis. Through the construction of the time series, the alignment deviation vector stream completely records the micro-motion trajectory of the operator over several seconds before triggering the operation, enabling subsequent steps to extract tremor patterns and behavioral trends from the time dimension. For example, in an ecological photography scenario, the photographer needs to continuously track the birds in flight before pressing the shutter. The alignment deviation vector flow records the relative position change between the center focus frame of the camera viewfinder and the subject during the tracking process. In a medical endoscopy scenario, the alignment deviation vector flow records the relative position change between the center of the endoscope image and the target lesion during the hovering and positioning process of the doctor's endoscope-holding hand.
[0074] Step S22, see Figure 4 The alignment deviation vector flow within the preset sliding time window is subjected to fast Fourier transform for spectral decomposition to separate low-frequency energy components and high-frequency abrupt change components. The low-frequency energy components are used as the operator's physiological tremor characteristics, and the high-frequency abrupt change components are used as the operator's operational jerking characteristics. Physiological tremor feature sets and operational jerking feature sets are constructed respectively. The tremor classification result is determined according to the energy ratio of the physiological tremor feature set and the operational jerking feature set.
[0075] Step S22 performs a Fast Fourier Transform (FFT) on the alignment deviation vector flow within a preset sliding time window. The sliding time window refers to a data sampling interval that slides continuously along the time axis at a fixed length; the data within the window is used for frequency domain analysis. The length Tw of the sliding time window is determined based on the typical period of human physiological tremors and the frequency domain resolution requirements: the window length should cover at least several complete physiological tremor periods to ensure that low-frequency components can be effectively extracted, while the window length should not be too long to avoid introducing too much historical data, which would cause the analysis results to lag in response to the current state. For example, the typical frequency range of physiological tremors generated by human skeletal muscles when maintaining a static posture is 8 Hz to 12 Hz, corresponding to a period of approximately 0.08 seconds to 0.125 seconds. Setting the sliding time window length to 2 to 3 seconds can cover approximately 16 to 36 complete periods, meeting the sampling sufficiency requirements of frequency domain analysis. The Fast Fourier Transform (FFT) is an efficient algorithm for converting time-domain signals into frequency-domain signals. Its output is spectral data, where each frequency component in the spectrum corresponds to an amplitude and a phase, and the square of the amplitude represents the energy proportion of that frequency component in the original signal.
[0076] Fast Fourier Transform (FFT) is applied to the alignment deviation distance sequence and alignment deviation phase sequence in the alignment deviation vector flow, respectively. The time-domain waveform of the alignment deviation distance sequence reflects the variation of the deviation amplitude over time; the low-frequency components in its spectrum correspond to the slowly changing fundamental deviation trend, while the high-frequency components correspond to the rapidly changing instantaneous disturbances. The time-domain waveform of the alignment deviation phase sequence reflects the variation of the deviation direction over time; the periodic components in its spectrum reveal the regular oscillation pattern of the deviation direction. After spectral decomposition, a physiological frequency threshold range needs to be set to distinguish between physiological tremor components and operational jerking components. The lower limit fL and upper limit fH of the physiological frequency threshold range are determined based on human physiological characteristics: the physiological tremor frequency of human skeletal muscles is affected by muscle type, fatigue level, and individual differences, with a typical frequency range of 8 Hz to 12 Hz; the frequency of chest rise and fall caused by respiration is approximately 0.2 Hz to 0.5 Hz; and the frequency of weak pulsation caused by heartbeat is approximately 1 Hz to 1.5 Hz. Taking all the above factors into account, the physiological frequency threshold range can be set as a lower limit fL of approximately 0.2 Hz and an upper limit fH of approximately 15 Hz. The specific values can be calibrated according to the application scenario and the characteristics of the operator group.
[0077] During spectrum separation, spectral components whose frequencies fall within the physiological frequency threshold range and whose amplitudes exceed the noise floor level are extracted as low-frequency energy components. The noise floor level refers to the background energy generated in the spectrum by non-target factors such as thermal noise and quantization errors of the imaging sensor, which can be obtained by collecting blank data and calculating its spectral mean while the device is stationary. The low-frequency energy components reflect the periodic influence of physiological processes such as skeletal muscle, respiration, and heartbeat on visual alignment behavior, and are defined as physiological tremor characteristics. Spectral components whose frequencies exceed the upper limit fH of the physiological frequency threshold range and whose amplitudes exceed a preset multiple of the noise floor level are extracted as high-frequency abrupt change components. The preset multiple is determined based on the requirement that high-frequency abrupt change components be significantly different from random noise, typically set to 3 to 5 times the noise floor level as the judgment threshold. High-frequency abrupt change components reflect the instantaneous disturbances caused by operator errors, psychological tension, or external shocks at the moment of triggering the operation, and are defined as operational jitter characteristics. For example, in a shooting training scenario, if the shooter pulls the trigger and the gun body experiences a momentary lateral shift due to improper finger force, this shift will manifest as a high-amplitude, high-frequency spike pulse in the alignment deviation distance sequence. In a surveying scenario, if the surveyor presses the measurement key and the rangefinder experiences a momentary dip due to excessive finger pressure, this dip will manifest as a step abrupt change in phase in the alignment deviation phase sequence.
[0078] The physiological tremor feature set and the operational jerking feature set are two data structures that store corresponding spectral component information respectively. The physiological tremor feature set contains the frequency value, amplitude, and phase of each frequency component falling within the physiological frequency threshold range; the operational jerking feature set contains the frequency value, amplitude, and phase of each frequency component exceeding the upper limit of the physiological frequency threshold range. The tremor classification result is determined based on the energy ratio of the physiological tremor feature set and the operational jerking feature set. The energy ratio is calculated as follows: summing the squares of the amplitudes of all frequency components in the physiological tremor feature set and the operational jerking feature set respectively, to obtain the total energy of physiological tremor (Ephysio) and the total energy of operational jerking (Eopera); the energy ratio of physiological tremor (Rphysio) = Ephysio / (Ephysio + Eopera), and the energy ratio of operational jerking (Ropera) = Eopera / (Ephysio + Eopera). The tremor classification result is determined based on the relative magnitude of the energy proportions: if Rphysio is significantly greater than Ropera, the tremor within the current window is determined to be mainly due to physiological factors; if Ropera is significantly greater than Rphysio, the tremor within the current window is determined to be mainly due to operational factors; if the two are close, it is determined to be a mixed tremor. For example, the classification threshold is set to perform a single attribution when the absolute value of the energy proportion difference is greater than a preset threshold Rth; otherwise, it is determined to be a mixed type. Rth can be set to 0.3 to 0.4. Step S22 uses frequency domain analysis instead of time domain analysis for tremor feature extraction because physiological tremors and operational jerks are superimposed on the time axis in the time domain signal and are difficult to separate directly. Frequency domain transformation expands the signal according to frequency components, allowing tremor components with different frequency characteristics to be spatially separated in the spectrum, thus enabling component extraction through simple frequency band filtering. The Fast Fourier Transform (FFT), as a mature signal processing algorithm, has advantages such as high computational efficiency and controllable frequency domain resolution, meeting the performance requirements of real-time processing.
[0079] Step S23: Input the physiological tremor feature set and the operational jitter feature set into the pre-trained LSTM time series analysis model to identify the operational intention stage. The operational intention stage includes the intentional precision aiming stage and the moment of triggering the operation. In the intentional precision aiming stage, calculate the handheld stability score based on the proportion of time that the dynamic target centroid coordinates fall within the precision judgment circle centered on the visual reference center coordinates. Encapsulate the operational intention stage, the handheld stability score, and the tremor classification results into visual alignment situation data.
[0080] LSTM (Long Short-Term Memory) time series analysis models are time series classification models built upon the Long Short-Term Memory (LSTM) network architecture. LSTM is a special type of recurrent neural network whose core structure includes three gating units: a forget gate, an input gate, and an output gate. This gating mechanism enables selective memorization and forgetting of temporal information. The forget gate determines which information from the previous time step needs to be discarded, the input gate determines which information from the current time step needs to be written into the memory state, and the output gate determines which information from the current memory state needs to be output to the next layer. This gating mechanism allows LSTM networks to retain key information over long sequences, overcoming the gradient vanishing problem in traditional recurrent neural networks when processing long sequences. The network structure of the LSTM time series analysis model includes an input layer, hidden layers, and an output layer. The input layer receives physiological tremor feature sets and operational jitter feature sets as input vectors, and the dimension of the input vector is equal to the total number of frequency component parameters in the two feature sets. The hidden layer consists of multiple stacked LSTM unit layers. The number of layers and the number of units per layer are determined based on the model complexity and the size of the training data. For example, it can be set to two to three LSTM unit layers, with each layer containing sixty-four to one hundred and twenty-eight LSTM units. The output layer is a fully connected layer with a classification activation function, and the output dimension is equal to the number of categories in the operation intention stage.
[0081] The LSTM time series analysis model is trained using supervised learning. The training dataset consists of alignment deviation vector flow samples labeled with the operational intent stage. The operational intent stage labels are obtained by professional evaluators reviewing video recordings and labeling each time window with the corresponding operational intent stage category based on the operator's behavioral characteristics. The operational intent stage categories include two main types: intentional precision aiming stage and triggering operation moment. The behavioral characteristics of the intentional precision aiming stage are: the alignment deviation distance sequence of the alignment deviation vector flow shows a convergent trend, i.e., the deviation distance gradually decreases; the amplitudes of each frequency component of the physiological tremor feature set remain relatively stable; and the energy of the operational jitter feature set remains at a low level. The behavioral characteristics of the triggering operation moment are: a step abrupt change in the alignment deviation vector flow, i.e., a large jump in deviation distance or deviation phase within a short period of time, followed by a high-amplitude peak in the operational jitter feature set energy. The training process uses the cross-entropy loss function to measure the difference between the model output and the true label, and employs the Adam optimizer for parameter updates. The initial learning rate can be set to 0.001, and it decays according to a preset decay coefficient after a preset number of rounds. For example, it can be set to decay to 0.9 times the original value every ten rounds. After training is complete, the model parameters are stored in the system for inference deployment.
[0082] During the inference process of the LSTM time series analysis model, the model receives the physiological tremor feature set and the operational jitter feature set within the current sliding time window as input, and outputs the classification result of the operator's operational intention stage at the current moment. The classification result is represented in the form of a probability distribution, and the category with the highest probability value is taken as the final operational intention stage. The memory mechanism of the LSTM network allows the model to use the feature information of historical time steps to assist in the judgment of the current moment. For example, if the model observes that the alignment deviation distance has been continuously converging in the previous time step and there is a sudden step increase in the alignment deviation distance at the current moment, the model can combine the historical convergence trend and the current abrupt change feature to comprehensively judge that the current moment is the trigger moment for the operation.
[0083] The handheld stability score is calculated based on the percentage of time during which the centroid coordinates of the dynamic target fall within the accuracy judgment circle centered on the visual reference center during the intentional precision aiming phase. The radius of the accuracy judgment circle (Runit) is set according to the accuracy requirements of the application scenario: in shooting training scenarios, the radius can be set to half the projected size of the target's bullseye area on the imaging plane; in surveying scenarios, the radius can be set to the projected size of the measurement accuracy tolerance error on the imaging plane; in medical endoscopy scenarios, the radius can be set to the radius of the safe operating area of the target lesion. The method for determining whether the centroid coordinates of the dynamic target fall within the accuracy judgment circle is as follows: calculate the alignment deviation distance r of the current frame. If r is less than or equal to the radius Runit of the accuracy judgment circle, the centroid coordinates of the dynamic target in that frame are determined to fall within the accuracy judgment circle; otherwise, they are determined to fall outside the accuracy judgment circle. The handheld stability score (Sscore) is calculated as: Sscore = Tin / Taim, where Tin is the cumulative time during the intentional precision aiming phase where the centroid coordinates of the dynamic target fall within the accuracy judgment circle, and Taim is the total duration of the intentional precision aiming phase. The accumulation method for Tin is as follows: traverse all frames within the intentional precision aiming phase, count the number of frames falling within the precision judgment circle, and divide the count by the video frame rate to obtain the duration. The calculation method for Taim is as follows: determine the start and end times of the intentional precision aiming phase based on the operational intent phase recognition results, and calculate the time difference between the start and end times. The handheld stability score Sscore ranges from zero to one. The closer the value is to one, the higher the proportion of time the operator keeps the distant target of interest near the visual reference center during the intentional precision aiming phase, reflecting better handheld stability.
[0084] The visual alignment situation data is a composite data structure that encapsulates the operational intent stage, handheld stability score, and tremor classification results. The operational intent stage field stores the current operational intent stage category output by the LSTM time-series analysis model; the handheld stability score field stores the calculated Sscore value; and the tremor classification result field stores the energy ratio and classification result of physiological and operational tremors output in step S22. The encapsulation design of the visual alignment situation data aims to integrate multi-dimensional evaluation results into a unified data interface, providing structured input for the generation of visual guidance instructions in the subsequent step S30.
[0085] Step S20 transforms discrete spatial coordinates into visual alignment posture data containing behavioral semantics by constructing a temporal alignment deviation vector flow and combining frequency domain decomposition and LSTM temporal analysis. This overcomes the technical bottleneck of traditional assessments, which can only obtain the result after triggering but cannot trace the process before triggering, and achieves quantitative decomposition of the dynamic aiming process before triggering. Specifically, the fast Fourier transform accurately decouples the originally aliased temporal tremor signal in the frequency domain into physiological tremor (low frequency) representing physical fitness and operational jitter (high frequency) representing skill level, thereby objectively distinguishing whether the cause of the operational deviation is "physical limitation" or "skill deficiency," and achieving precise assessment attribution. At the same time, the LSTM model is used to identify the operational intention stage and generate a handheld stability score, establishing a quantifiable and unified evaluation standard for visual alignment behavior, transforming handheld stability assessment from a passive judgment of "only looking at the result" to an active diagnosis of "process analysis."
[0086] Step S30: Monitor the target state vector in real time. When a high dynamic visual transition is detected, switch to the visual blind zone state and generate a dynamic target virtual trajectory within the blind zone. After exiting the visual blind zone state, generate a full-process alignment trajectory based on the dynamic target virtual trajectory. Then, spatially map and fuse the full-process alignment trajectory with the visual alignment situation data to generate a visual guidance instruction.
[0087] Step S30 predicts and repairs the target loss caused by the high-dynamic visual transition that occurs at the moment of triggering the operation, and transforms the visual alignment situation data generated in step S20 into intuitive and understandable visual guidance instructions. High-dynamic visual transition refers to the phenomenon that a handheld optical imaging terminal undergoes a drastic displacement or attitude change in a very short time, resulting in a significant misalignment of image content between adjacent frames, making it impossible for the target detection algorithm to effectively identify distant targets of interest in the current frame. For example, in a shooting training scenario, the recoil generated at the moment of triggering the operation will cause the handheld optical imaging terminal to vibrate violently within tens of milliseconds, resulting in severe motion blur in the image, and the lightweight convolutional neural network in step S13 will be unable to output an effective target bounding box; in an ecological photography scenario, the photographer's finger exerts force at the moment of pressing the shutter, causing the camera to momentarily drop, the viewfinder image to shift and jump, and the subject may briefly move out of the frame boundary; in a medical endoscopy scenario, the doctor's hand holding the endoscope shakes momentarily due to tension when pressing the electrocoagulation pedal, causing the endoscope image to shake violently, and the detection confidence of the target lesion drops sharply. The visual blind zone state refers to a special operating mode that the system enters after detecting a highly dynamic visual transition. In this mode, the system stops using the real-time detection output of step S13 and instead uses a motion model-based prediction mechanism to calculate the position trajectory of the distant target of interest during the visual loss period. The dynamic target virtual trajectory refers to the sequence of centroid positions of the distant target of interest calculated by the system using the Kalman filter algorithm during the visual blind zone state. This trajectory is not directly obtained from the image but is calculated based on the target's motion state before loss and the global motion vector of the imaging device. The full-process alignment trajectory refers to the complete position sequence formed by aligning the dynamic target virtual trajectory during the visual blind zone state with the real detection trajectory in normal visual tracking mode through timestamp alignment and smooth stitching. This sequence covers the entire time interval from the start of visual alignment to the completion of the trigger operation, eliminating data breaks caused by highly dynamic visual transitions. The visual guidance instruction is the final output of step S30, containing an alignment trajectory heatmap and textual correction suggestions, used to visually present the operator's visual alignment behavior characteristics and improvement directions to the evaluator on the evaluation terminal. The design motivation for step S30 stems from the fact that traditional video-assisted evaluation systems are prone to losing target tracking at the moment of triggering an operation, resulting in a break in the alignment trajectory before and after the triggering operation. Evaluators cannot know the precise alignment state at the moment of triggering, and therefore cannot distinguish whether the operational deviation occurred during the visual alignment phase before the triggering operation or during the action execution phase at the moment of triggering. By using motion model prediction to fill the data gaps during the visual loss period, the alignment trajectory throughout the process has temporal continuity, allowing the evaluation analysis to accurately pinpoint the time and cause of the deviation.
[0088] Further, step S30 includes:
[0089] Step S31: Monitor the target confidence score in the target state vector in real time, perform sparse optical flow method to calculate the global motion vector for two adjacent standardized HSV preprocessed image frames, and determine that a high dynamic visual transition has been detected when the target confidence score is less than or equal to the preset loss judgment threshold or the magnitude of the global motion vector exceeds the preset inter-frame displacement threshold, and switch the system state to the visual blind zone state.
[0090] Step S31 continuously monitors the target confidence score in the target state vector output in step S13. The target confidence score is a quantitative judgment of the effectiveness of the detection result of the current frame by the lightweight convolutional neural network in step S13. Its value ranges from zero to one. The higher the value, the more confident the network is in detecting a real distant target of interest rather than a false detection or missed detection. When the distant target of interest becomes blurred and loses texture due to motion blur, the feature response extracted by the network weakens, and the output target confidence score decreases accordingly. When the distant target of interest completely moves out of the field of view or is covered by occlusion, the network may output an empty detection result. At this time, the target confidence score is zero or close to zero. Step S31 sets a loss decision threshold Tconf as the boundary for judging the effectiveness of target detection. The method for determining the loss decision threshold Tconf is as follows: under controlled experimental conditions, a test image set containing different degrees of motion blur is collected, the target confidence score output by the network under each degree of blur is recorded, and the lower limit of the confidence score corresponding to the manual determination that the target can still be identified is statistically calculated. This lower limit is used as the loss decision threshold Tconf. For example, if the confidence score corresponding to the point where the target cannot be identified manually is approximately 0.3 to 0.4, then the loss determination threshold Tconf can be set to 0.35. When the target confidence score is less than or equal to the loss determination threshold Tconf, the target detection result of the current frame is determined to be unreliable, and the target may be lost due to motion blur or moving out of the field of view.
[0091] Step S31 simultaneously performs sparse optical flow calculation on two adjacent normalized HSV preprocessed image frames to calculate the global motion vector. Sparse optical flow is a motion estimation algorithm based on feature point tracking. Its basic principle is: detect several corner points with significant features in the previous frame as tracking targets, search for the corresponding positions of these corner points in the next frame, and calculate the motion vector based on the positional differences of the corner points in the two frames. The specific implementation process of sparse optical flow is as follows: a corner detection algorithm is used to detect feature corner points in the previous normalized HSV preprocessed image frame. Corner detection is usually based on the eigenvalue analysis of the autocorrelation matrix of the image gradient, selecting pixel positions with larger eigenvalues as corner points. The upper limit of the number of detected corner points can be set to Nmax, and an exemplary Nmax can be set to one hundred to two hundred. For each detected corner point, within a local search window centered on the corner point position in the next frame, a pyramid iterative optical flow algorithm is used to calculate its corresponding position in the next frame. This algorithm is based on the optical flow constraint equation and multi-scale analysis of the image pyramid, and can handle large-amplitude displacements. After calculation, the displacement vector of each corner point is obtained, which is a two-dimensional vector pointing from the position of the previous frame to the position of the next frame. The calculation of the global motion vector Vglobal adopts a robust statistical method: median filtering or random sampling consistency estimation is performed on the displacement vectors of all corner points to eliminate abnormal displacement vectors caused by local moving objects, and the average of the remaining displacement vectors is taken as the global motion vector Vglobal. The global motion vector Vglobal reflects the overall displacement of the imaging device between two frames. Its magnitude |Vglobal| represents the magnitude of the inter-frame displacement, and its direction represents the direction of the displacement. For example, in a shooting training scenario, the recoil at the moment of triggering the operation causes the handheld optical imaging terminal to move violently backward and upward, and the global motion vector will have a large magnitude and point backward and upward; during normal visual alignment, the handheld device only has slight vibration, and the magnitude of the global motion vector is smaller.
[0092] An inter-frame displacement threshold Tdisp is set as the boundary for determining the motion amplitude of a high-dynamic visual transition. The method for determining the inter-frame displacement threshold Tdisp is as follows: Under controlled experimental conditions, video samples containing normal hand-held tremors and violent shaking during a triggered operation are collected. The global motion vector magnitude distribution is calculated for both cases. The midpoint between the upper limit of the normal hand-held tremor magnitude distribution and the lower limit of the triggered violent shaking magnitude distribution is taken as the inter-frame displacement threshold Tdisp. For example, if the upper limit of the inter-frame displacement magnitude for normal hand-held tremors is approximately ten pixels, and the lower limit for the inter-frame displacement magnitude for triggered violent shaking is approximately thirty pixels, then the inter-frame displacement threshold Tdisp can be set to twenty pixels. When the magnitude |Vglobal| of the global motion vector exceeds the inter-frame displacement threshold Tdisp, violent imaging device motion is determined to have occurred. Even if the target confidence score remains within an acceptable range, image quality may degrade due to motion blur, and the detection results of subsequent frames may be unreliable.
[0093] A dual-condition judgment logic is employed to determine whether a high-dynamic visual transition has been detected: Target detection is deemed to have failed when the target confidence score is less than or equal to the loss determination threshold Tconf; or the global motion vector magnitude |Vglobal| exceeds the inter-frame displacement threshold Tdisp, indicating that the imaging device has undergone drastic motion. Meeting either condition determines a high-dynamic visual transition has been detected, and the system immediately switches its operating state from visual tracking mode to visual blind zone mode. The technical consideration for using a dual-condition judgment instead of a single condition is that the decrease in target confidence score may lag behind the occurrence of imaging device motion. When drastic motion has just begun, the detection result of the previous frame may still be valid, but subsequent frames are no longer reliable. Relying solely on the confidence score would delay the state switching timing. The calculation of the global motion vector can detect anomalies in the frame immediately following the motion, thus triggering the state switch earlier and reducing the processing of invalid frames. The logical OR relationship between the two conditions ensures the timeliness of the state switch: the confidence score captures target-level detection failures, while the global motion vector captures image-level motion anomalies; the two complement each other, covering different types of visual loss scenarios. The synergy of the two-layer criteria ensures that the triggering of the visual blind zone state considers both the output quality of the target detection algorithm and the motion features of the original image, improving the comprehensiveness and robustness of the state switching judgment. It should be noted that the calculation of the global motion vector is continuously performed throughout the system's operation. During the visual blind zone state, the global motion vector continuously output in step S31 serves as the necessary input for background compensation in step S32. Without step S31, the subsequent Kalman filter prediction in step S32 would lack a start signal, and the system would be unable to determine when to stop using the real-time detection output of step S13 and instead activate the prediction mechanism. During highly dynamic visual transitions, it might continue to use low-confidence erroneous detection results for subsequent analysis, leading to abrupt changes in the alignment deviation vector stream and contaminating the temporal analysis results of step S20.
[0094] Step S32: In response to the visual blind spot state, the target state vector and global motion vector of the previous moment are called as inputs to the Kalman filter to perform inertial prediction and background compensation, and to calculate the dynamic target virtual trajectory in the blind spot.
[0095] Step S32, in response to the visual blind zone state being triggered, calls the target state vector output in the last frame before entering the visual blind zone state in step S13 as the initial state of the Kalman filter. The target state vector contains the dynamic target centroid coordinates C. obj The system consists of three components: estimated depth distance (Dest), target confidence score, and dynamic target centroid coordinates (C). objThe estimated depth distance, Dest, is used to construct the state vector of the Kalman filter. The Kalman filter is a recursive state estimation algorithm. Its basic principle is to predict the current state based on the system's dynamic model, and then use observations to correct the prediction, achieving a statistically optimal estimate between the predicted and observed values. The state vector X of the Kalman filter is defined as a four-dimensional vector containing the target's position and velocity, where X equals the abscissa x in vector form. obj y-axis obj Lateral velocity v x Longitudinal velocity v y The position component in the state vector is initialized by the dynamic target centroid coordinates of the last frame before entering the visual blind zone state, and the velocity component is calculated by the difference of the dynamic target centroid coordinates of several frames before entering the visual blind zone state. The difference calculation method is: take the difference of the dynamic target centroid coordinates of the two most recent frames and divide it by the inter-frame time interval. The inter-frame time interval is determined by the difference of the acquisition timestamps of adjacent frames.
[0096] The state transition model of the Kalman filter adopts the assumption of uniform motion, and the state transition matrix F is a 4×4 matrix, with the following specific form:
[0097] ;
[0098] in, Let F be the time interval. The state transition matrix F represents the predicted relationship from the previous state to the current state: position equals the previous position plus the velocity multiplied by the time interval, while the velocity remains unchanged. Time interval The time stamp is determined by the difference between the acquisition timestamps of adjacent frames, and is consistent with the acquisition timestamp marked in step S11. The assumption of uniform motion is a reasonable approximation under the condition that the duration of the visual blind zone state is short, because high-dynamic visual transitions typically last only tens to hundreds of milliseconds, and the acceleration changes of distant targets of interest during this period have a relatively limited impact on position prediction. The prediction step of the Kalman filter is performed according to the standard Kalman prediction formula: Predicted state X pred for ,in, This is a state estimate from the previous moment. Predicting covariance. for: ,in, Q is the covariance of the previous moment, and Q is the process noise covariance. The process noise covariance Q reflects the uncertainty of the state transition model. Its diagonal elements are set according to the typical random disturbance amplitude of the target motion. For example, the process noise of the position component can be set as the square of the standard deviation of the random displacement that the target may have in a single frame time, and the process noise of the velocity component can be set as the square of the standard deviation of the velocity change caused by the random acceleration that the target may have in a single frame time.
[0099] Step S32 introduces the global motion vector Vglobal calculated in step S31 as the observation input to perform background compensation correction on the inertial prediction result. The reason for background compensation is that the inertial prediction of the Kalman filter is based on the motion law of the target in the world coordinate system, while the centroid coordinates of the dynamic target are represented in the image coordinate system. When the imaging device moves, even if the target is stationary in the world coordinate system, its coordinates in the image coordinate system will change. Therefore, the predicted position of the target in the image coordinate system obtained by inertial prediction needs to be compensated in reverse according to the motion of the imaging device. The specific implementation method of background compensation is to add the global motion vector Vglobal to the target position obtained by inertial prediction to obtain the target position prediction value after background compensation. This superimposed motion compensation ensures that the predicted centroid coordinates of the dynamic target correctly reflect the position of the target in the image coordinate system. The Kalman filter continues to perform the prediction step during the visual blind zone state, and outputs a predicted centroid coordinate of the dynamic target in each frame. These predicted coordinates are arranged in chronological order to form the virtual trajectory of the dynamic target. The data structure of the dynamic target virtual trajectory is consistent with the time series structure of the alignment deviation vector stream in step S21. Each element contains the acquisition timestamp and the predicted dynamic target centroid coordinates, so that the timestamps can be aligned and stitched in subsequent step S33. The duration of the visual blind zone state depends on the recovery speed of the high-dynamic visual transition. In each frame, the system detects whether the target confidence score has risen above the loss determination threshold Tconf and whether the magnitude of the global motion vector has dropped below the inter-frame displacement threshold Tdisp. If so, the system determines that the visual detection has recovered stably, exits the visual blind zone state, and resumes using the real-time detection output of step S13.
[0100] The reason why step S32 uses a Kalman filter combined with background compensation instead of simple inertial extrapolation is that simple inertial extrapolation only predicts the position based on the target's historical motion state, without considering the influence of the imaging device's own motion on the image coordinate system. In high-dynamic visual transition scenarios where the imaging device moves violently, the prediction error of inertial extrapolation will accumulate rapidly, and the predicted position after several frames may deviate significantly from the target's true position. Introducing a global motion vector as a background compensation amount is equivalent to superimposing the influence of the imaging device's motion into the output of the prediction model, so that the inertial prediction result based on the target's own motion law can correctly reflect the target's actual position in the image coordinate system shifted by the camera motion, significantly reducing the accumulation rate of prediction error. The covariance propagation mechanism of the Kalman filter can also quantify the uncertainty of the prediction result. The prediction covariance increases with the number of prediction steps, reflecting the decrease in reliability of long-term prediction. This uncertainty information can be used as a weighted reference in the subsequent step S33 when stitching the trajectory. In step S13, the target state vector output in the last frame before entering the visual blind zone provides the Kalman filter with initial position and velocity estimates, ensuring that the predicted trajectory smoothly continues from the true detection position. If step S32 is missing, there will be no target position data during the visual blind spot state. Subsequent step S33 will not be able to form an alignment trajectory covering the entire process. The alignment trajectory heatmap in step S34 will have a data blank area near the moment of triggering the operation. The evaluator will not be able to know the precise alignment status at the moment of triggering the operation, and the goal of full-link visualization of visual alignment behavior will not be achieved.
[0101] Step S33: When the target confidence score is greater than the loss judgment threshold and the magnitude of the global motion vector is less than the inter-frame displacement threshold Tdisp, exit the visual blind zone state, and perform time stamp alignment and smooth stitching on the real detection trajectory formed by the dynamic target virtual trajectory and the dynamic target centroid coordinates in the target state vector to form the full-process alignment trajectory.
[0102] The execution trigger condition for step S33 is that the target confidence score is greater than the loss determination threshold Tconf and the magnitude of the global motion vector is less than the inter-frame displacement threshold Tdisp, indicating that the target detection algorithm in step S13 has resumed normal operation and the output dynamic target centroid coordinates are reliable again. The system exits the visual blind zone state, resumes using the real-time detection output of step S13, and simultaneously triggers the trajectory stitching process in step S33. The true detection trajectory refers to the position sequence composed of the dynamic target centroid coordinates output by step S13 arranged in chronological order within the time interval outside the visual blind zone state, including the true detection trajectory before entering the visual blind zone state and the true detection trajectory after exiting the visual blind zone state. The dynamic target virtual trajectory refers to the position sequence composed of the predicted dynamic target centroid coordinates output by the Kalman filter in step S32 arranged in chronological order during the duration of the visual blind zone state.
[0103] The purpose of timestamp alignment is to ensure that the three tracks are correctly connected on the timeline. Please refer to [link / reference]. Figure 5 , Figure 5 This is a time-series schematic diagram of the alignment trajectory throughout the entire process provided in an embodiment of this application. For example... Figure 5 As shown in the figure, the horizontal axis represents the passage of time, and the vertical axis represents the change of position, intuitively demonstrating that the alignment trajectory throughout the process is composed of line segments connected in different states in time sequence. The specific implementation method of timestamp alignment is as follows: using the acquisition timestamp marked in step S11 as a unified time reference, the actual detection trajectory before entering the visual blind zone state is arranged in ascending order of timestamps as the first segment, that is... Figure 5 The first segment of the trajectory, appearing as a solid line, is located before the "entering the blind zone" node; the second segment is the virtual trajectory of the dynamic target, arranged in ascending order of timestamps. Figure 5 The second segment of the trajectory, which is located within the dashed box area of the "visual blind spot state" and appears as a dashed line, is considered the third segment. The actual detected trajectory after exiting the visual blind spot state is then arranged in ascending order of timestamps. Figure 5 The third trajectory segment, located after the "exit blind zone" node and restored to a solid line form, is the third segment of the trajectory. The three trajectories are connected end-to-end in chronological order. Timestamp alignment requires handling the overlap of boundary frames: the last true detection frame at the moment of entering the visual blind zone corresponds to the same timestamp as the first predicted frame of the dynamic target's virtual trajectory; however, the first true detection frame at the moment of exiting the visual blind zone may have similar timestamps to the last predicted frame of the dynamic target's virtual trajectory. The boundary frame handling strategy is to retain the true detection result at the boundary as a connection point, corresponding to... Figure 5 The solid dots marked "Entering the Blind Zone" are used to indicate this; the actual detection results are retained as connection points at the exit boundary, corresponding to... Figure 5 The solid circle marked "Exit Blind Zone" indicates the target; the virtual trajectory of the dynamic target only fills the blank space between the two boundaries.
[0104] The purpose of smooth stitching is to eliminate potential position jumps at the junction of the actual detected trajectory and the virtual trajectory of the dynamic target. Position jumps occur because the Kalman filter prediction, based on an inertial model and background compensation, deviates from the target's actual motion. There may be a difference between the actual detected position at the moment of exiting the visual blind zone and the predicted position in the last frame. The specific implementation method of smooth stitching is as follows: a transition interval is set near the exit boundary. The length of the transition interval, Ntrans, can be set to several frames, for example, five to ten frames. Within the transition interval, linear interpolation is used to correct the predicted trajectory: let the predicted position at the exit boundary be P'pred, the actual detected position be Preal, and the position difference between them be ΔP = Preal - P'pred; for the predicted position of the i-th frame (where i ranges from 1 to Ntrans) traversing backward from the exit boundary within the transition interval, the correction amount is... The corrected position is the original predicted position plus the correction amount. This correction method results in a larger correction amount for frames closer to the exit boundary (the correction amount for the Ntrans frame is ΔP, perfectly aligned with the actual detection position), and a smaller correction amount for frames farther from the exit boundary (the correction amount for the first frame is ΔP / Ntrans), achieving a gradual transition from the predicted trajectory to the actual detection position. This linear gradual correction method allows the end of the predicted trajectory to smoothly transition to the actual detection position, eliminating positional abruptness at the junction. A similar smoothing process is used at the entry boundary: let the actual detection position at the entry boundary be Penter, and the predicted position of the first frame of the dynamic target virtual trajectory be Pfirst. If there is a difference between the two, a linear gradual transition is performed from Penter to several subsequent frames. The fully aligned trajectory is a complete position sequence formed after timestamp alignment and smooth stitching. Its data structure is a list of dynamic target centroid coordinates arranged in ascending order of timestamps, covering the entire time interval from the start of visual alignment to recovery after the trigger operation is completed. The time range of the alignment trajectory throughout the entire process is defined by the acquisition timestamp of the first frame and the acquisition timestamp of the last frame. There are no data gaps caused by visual loss in between, thus achieving the temporal continuity of the alignment trajectory.
[0105] Step S33 uses smooth stitching instead of direct stitching because direct stitching introduces a positional jump at the junction. This jump appears as an abrupt trajectory change in the alignment trajectory heatmap generated in subsequent step S34, which could be misinterpreted as abnormal operator behavior and interfere with the evaluator's judgment. Smooth stitching eliminates the jump through linear gradient correction, resulting in a continuous and smooth alignment trajectory throughout the process. This accurately reflects the target's movement trend during the period of visual loss, rather than introducing artificial positional abrupt changes. Simultaneously, smooth stitching preserves the overall trend characteristics of the predicted trajectory when correcting it, adjusting only local positional deviations to ensure that the predicted trajectory's direction and velocity characteristics are maintained. These characteristics are valuable for evaluating the alignment status at the moment of triggering the operation. The dynamic target virtual trajectory fills the data gap during the visual blind spot period and is seamlessly stitched with the real detection trajectory to form a temporally continuous alignment trajectory. After exiting the visual blind spot state, the output of valid dynamic target centroid coordinates is restored. These coordinates are used as the target position for smooth stitching, ensuring a natural connection between the end of the stitched trajectory and the subsequent real detection trajectory. If step S33 is missing, the virtual trajectory of the dynamic target and the real detection trajectory will remain separate. In the subsequent step S34, multiple independent trajectories need to be processed separately when generating the alignment trajectory heatmap. The connection between the trajectories cannot be visualized, making it difficult for the evaluator to form an overall understanding of the entire visual alignment process.
[0106] Step S34: Using the coordinates of the visual reference center as the origin, map the alignment trajectory of the entire process to generate an alignment trajectory heatmap, generate correction suggestions based on the visual alignment status data, and combine the alignment trajectory heatmap and correction suggestions into a visual guidance instruction and render it to the evaluation terminal.
[0107] Step S34 uses the visual reference center coordinates C output in step S12. ref Establish a two-dimensional plane coordinate system with the origin, and map the entire alignment trajectory onto this coordinate system. The method for establishing the coordinate system is as follows: using the visual reference center coordinates C... ref The position is taken as the origin of the coordinate system, with the positive direction of the horizontal axis aligned with the positive direction of the horizontal axis of the image coordinate system, and the positive direction of the vertical axis aligned with the positive direction of the vertical axis of the image coordinate system. The centroid coordinate C of each dynamic target in the alignment trajectory is set throughout the entire process. obj All are converted to offset coordinates relative to the origin, and the offset coordinates are equal to C. obj -C ref This coordinate transformation allows the alignment trajectory throughout the process to directly represent the offset of the distant target of interest relative to the visual reference center, establishing a direct geometric relationship with the operator's subjective visual alignment behavior.
[0108] The alignment trajectory heatmap is generated based on the dwell time statistics of each position point in the entire alignment trajectory. Dwell time refers to the cumulative duration for which the centroid of the distant target of interest stays near a certain spatial position; a longer dwell time indicates that the operator maintains visual alignment at that position for a longer period. The specific method for generating the alignment trajectory heatmap is as follows: a two-dimensional plane with the visual reference center coordinates as the origin is divided into several grid cells. The size of each grid cell can be set to a number of pixels square, for example, five pixels by five pixels. Each position point in the entire alignment trajectory is traversed, and the grid cell it falls into is determined based on its offset coordinates. The dwell count of that grid cell is incremented by one. After traversal, the dwell count of each grid cell reflects the cumulative number of frames of the centroid of the distant target of interest in that area. The dwell count is divided by the video frame rate to obtain the dwell time. The dwell time is normalized to the range of 0 to 1 and mapped to color values to generate the heatmap. The color mapping uses a gradient from blue to red, with areas of short dwell time appearing as blue or cool colors, and areas of long dwell time appearing as red or warm colors. The highest temperature region in the alignment trajectory heatmap represents the center of convergence of the operator's visual alignment behavior, and the offset of the highest temperature region relative to the origin reflects whether the operator has a systematic alignment deviation.
[0109] The generation of correction suggestions is based on a comprehensive analysis of the spatial distribution characteristics of the alignment trajectory heatmap and the visual alignment posture data output in step S23. The logic for generating correction suggestions includes the following judgment rules: calculate the centroid position of the highest temperature region in the alignment trajectory heatmap; if the offset distance of this centroid position relative to the origin exceeds the preset system deviation threshold Tsys, a systematic alignment deviation is determined, and corresponding posture adjustment suggestions are generated. The method for determining the system deviation threshold Tsys is as follows: set an acceptable upper limit for system deviation based on the accuracy requirements of the application scenario. For example, in a shooting training scenario, it can be set to half the projected size of the target's nine-ring area boundary on the imaging plane. The direction of the systematic alignment deviation can be further subdivided: if the centroid position is mainly biased upwards or downwards, it may be related to the operator's breathing rhythm or gun-holding height, generating suggestions for breathing control or gun-holding posture adjustment; if the centroid position is mainly biased to the left or right, it may be related to the operator's gun-holding posture or visual habits, generating posture adjustment suggestions. Analyze the tremor classification results in the visual alignment posture data output in step S23: If the proportion of physiological tremor energy is significantly higher than the proportion of operational tremor energy, it indicates that the operator's instability mainly stems from physical fitness factors, and suggestions are generated to strengthen physical training or adjust breathing rhythm; if the proportion of operational tremor energy is significantly higher than the proportion of physiological tremor energy, it indicates that the operator's instability mainly stems from motor skill factors, and suggestions are generated to standardize triggering actions or relax the mental state. Analyze the handheld stability score Sscore in the visual alignment posture data: If the Sscore is lower than the preset stability pass threshold Spass, suggestions are generated to strengthen stability training. The stability pass threshold Spass is determined by statistically setting it based on the skill distribution of the operator group. For example, it can be set to 0.6 to 0.7, indicating that during the intentional precision aiming phase, the target must be kept within the accuracy judgment circle for 60% to 70% of the time to be considered qualified.
[0110] The synthesis of visual guidance instructions overlays an alignment trajectory heatmap and corrective suggestions onto the original video frame to create a composite display interface. The overlay method is as follows: the alignment trajectory heatmap is displayed semi-transparently on one side of the video frame, with its origin aligned with the visual reference center coordinates of the current frame. When the evaluation terminal plays back the video, the origin of the heatmap updates synchronously with the visual reference center coordinates of each frame, allowing evaluators to visually see the distribution of the alignment trajectory relative to the visual reference markers. Corrective suggestions are displayed on the other side or bottom of the video frame, presented in a list format. The handheld stability score is displayed in a corner of the video frame. The synthesized visual guidance instructions are transmitted in real-time to the evaluation terminal via a network interface. The evaluation terminal can be a standalone display screen, tablet, or computer terminal, allowing evaluators to view the operator's visual alignment status and the system-generated evaluation results from a distance. The visual alignment status data includes the operational intent stage, handheld stability score, and tremor classification results. Based on this quantitative data, targeted corrective suggestions are generated, achieving a semantic transformation from abstract data to specific guidance. The entire alignment trajectory covers the complete time interval of visual alignment. Converting this into an alignment trajectory heatmap allows for spatial distribution visualization, enabling assessors to readily grasp the operator's alignment behavior characteristics. The visual reference center coordinates serve as the origin of the heatmap coordinate system, ensuring a correct correspondence between the heatmap's spatial location and the operator's visual reference. Assessors can directly determine the direction and degree of alignment deviation relative to the visual reference when viewing the heatmap. Without step S34, the quantitative data and entire trajectory generated in steps S20 and S33 would exist as abstract numerical values. Assessors would need professional data analysis skills to interpret this information, hindering the rapid formation of an intuitive understanding of the operator's visual alignment behavior and significantly reducing the guidance efficiency of handheld stability assessment.
[0111] Step S30 addresses the target tracking disruption caused by highly dynamic visual transitions (such as severe jitter at the moment of triggering), and transforms the abstract data generated in previous steps into intuitive visual guidance instructions. Step S30 establishes a complete closed loop from blind zone monitoring to trajectory reconstruction and then to visual mapping. Specifically, it keenly captures the moment of visual loss through dual criteria (target confidence and global motion vector), and uses a Kalman filter algorithm with background motion compensation to accurately calculate the virtual trajectory of the dynamic target within the "visual blind zone," successfully filling the key data gap at the moment of triggering and achieving temporal continuity of the alignment trajectory throughout the process. This not only allows evaluators to fully review the actual alignment state at the moment of triggering (thus accurately distinguishing whether the deviation stems from deviation during the aiming phase or from an action error at the moment of triggering), but also transforms the tedious numerical sequence into an intuitive alignment trajectory heatmap and targeted correction suggestions, achieving a qualitative leap in evaluation feedback from "blind men and the elephant" to "panoramic perspective," significantly improving the real-time and scientific nature of guidance and correction.
[0112] Example 2
[0113] This embodiment, based on Embodiment 1, provides a real-time image processing system for dynamic target recognition in video streams, such as... Figure 6 As shown, it includes:
[0114] Dual-channel feature extraction module: used to establish a video stream buffer queue for handheld optical imaging terminal, acquire monocular raw video frames, construct a dual-channel feature extraction model based on attention mechanism and color space, and solve the visual reference center coordinates and dynamic target centroid coordinates in parallel from monocular raw video frames through the dual-channel feature extraction model, and generate target state vector;
[0115] Visual alignment posture module: Generates alignment deviation vector flow based on the coordinates of the visual reference center and the centroid coordinates of the dynamic target, and deconstructs the operator's physiological tremor characteristics and operational jitter characteristics based on the alignment deviation vector flow to generate visual alignment posture data;
[0116] Visual guidance module: It is used to monitor the target state vector in real time. When a high dynamic visual transition is detected, it switches to the visual blind zone state and generates a dynamic target virtual trajectory within the blind zone. After exiting the visual blind zone state, it generates a full-process alignment trajectory based on the dynamic target virtual trajectory, and performs spatial mapping and fusion with the visual alignment situation data to generate visual guidance instructions.
[0117] Furthermore, in the dual-channel feature extraction module, the method for calculating the coordinates of the visual reference center includes:
[0118] The original monocular video frame is converted into an HSV color space image, and morphological filtering is performed on the HSV color space image to generate a standardized HSV preprocessed image frame.
[0119] Threshold segmentation of the HSV color geometric subspace is performed on the standardized HSV preprocessed image frame to extract the visual reference identifier mask. Connectivity analysis and geometric moment calculation are performed on the visual reference identifier mask to output the coordinates of the visual reference center.
[0120] The method for calculating the dynamic target centroid coordinates includes:
[0121] Standardized HSV preprocessed image frames are input into a lightweight convolutional neural network with a CBAM attention mechanism. Pyramid feature extraction is performed to locate distant targets of interest. The target bounding box and target confidence score output by the lightweight convolutional neural network are obtained. Based on the target bounding box of the distant target of interest, the dynamic target centroid coordinates and estimated depth distance are calculated. The distant target of interest includes a target.
[0122] The target state vector is obtained by encapsulating the dynamic target centroid coordinates, estimated depth distance, and target confidence score.
[0123] Furthermore, in the visual alignment posture module, the deconstruction method of the physiological tremor features and operational jerking features includes:
[0124] The alignment deviation vector flow within the preset sliding time window is subjected to fast Fourier transform for spectral decomposition, separating the low-frequency energy component from the high-frequency abrupt component. The low-frequency energy component is used as the operator's physiological tremor characteristic, and the high-frequency abrupt component is used as the operator's operational jitter characteristic.
[0125] The method for generating the visual alignment situation data includes:
[0126] The tremor classification result is determined based on the energy ratio of the physiological tremor feature set and the operational tremor feature set;
[0127] The physiological tremor feature set and the operational jitter feature set are input into a pre-trained LSTM time series analysis model to identify the operational intention stage, which includes the intentional precision aiming stage and the moment of triggering the operation. In the intentional precision aiming stage, the handheld stability score is calculated based on the proportion of time that the dynamic target centroid coordinates fall within the circle centered on the visual reference center coordinates. The operational intention stage, the handheld stability score, and the tremor classification results are encapsulated into visual alignment situation data.
[0128] The methods and systems of this application may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this application are not limited to the order specifically described above, unless otherwise specifically stated.
[0129] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.
[0130] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A real-time image processing method for dynamic target recognition in a video stream, characterized in that, The method includes: A video stream buffer queue is established for a handheld optical imaging terminal to acquire monocular raw video frames. A dual-channel feature extraction model based on attention mechanism and color space is constructed. The visual reference center coordinates and dynamic target centroid coordinates are solved in parallel from the monocular raw video frames through the dual-channel feature extraction model, and the target state vector is generated. Using the visual reference center coordinates as the origin and the dynamic target centroid coordinates in the target state vector as the moving point, the alignment deviation distance and alignment deviation phase are calculated. A time-series sequence is constructed by combining the acquisition timestamps of the original monocular video frames, generating an alignment deviation vector stream. A Fast Fourier Transform is performed on the alignment deviation vector stream within a preset sliding time window to perform spectral decomposition, separating low-frequency energy components and high-frequency abrupt change components. The low-frequency energy components are used as the operator's physiological tremor characteristics, and the high-frequency abrupt change components are used as the operator's operational jitter characteristics. The tremor classification result is determined based on the energy ratio of the physiological tremor feature set and the operational jitter feature set. The physiological tremor feature set and the operational jitter feature set are input into a pre-trained LSTM time-series analysis model to identify the operation intention stage, which includes the intentional precision aiming stage and the triggering operation instant. In the intentional precision aiming stage, the handheld stability score is calculated based on the proportion of time the dynamic target centroid coordinates fall within the precision judgment circle centered on the visual reference center coordinates. The operation intention stage, handheld stability score, and tremor classification result are encapsulated into visual alignment situation data. The system monitors the target state vector in real time and switches to the visual blind zone state when a highly dynamic visual transition is detected, generating a dynamic virtual trajectory of the target within the blind zone. After exiting the visual blind zone state, the system generates a full-process alignment trajectory based on the dynamic virtual trajectory of the target and spatially maps and fuses the full-process alignment trajectory with the visual alignment situation data to generate visual guidance instructions.
2. The real-time image processing method for dynamic target recognition in a video stream according to claim 1, characterized in that, The method for calculating the coordinates of the visual reference center includes: The original monocular video frame is converted into an HSV color space image, and morphological filtering is performed on the HSV color space image to generate a standardized HSV preprocessed image frame. Threshold segmentation of the HSV color geometric subspace is performed on the standardized HSV preprocessed image frame to extract the visual reference identifier mask. Connectivity analysis and geometric moment calculation are performed on the visual reference identifier mask to output the coordinates of the visual reference center.
3. The real-time image processing method for dynamic target recognition in a video stream according to claim 2, characterized in that, The method for calculating the dynamic target centroid coordinates includes: Standardized HSV preprocessed image frames are input into a lightweight convolutional neural network with a CBAM attention mechanism. Pyramid feature extraction is performed to locate distant targets of interest. The target bounding box and target confidence score output by the lightweight convolutional neural network are obtained. Based on the target bounding box of the distant target of interest, the dynamic target centroid coordinates and estimated depth distance are calculated. The distant target of interest includes a target. The target state vector is obtained by encapsulating the dynamic target centroid coordinates, estimated depth distance, and target confidence score.
4. The real-time image processing method for dynamic target recognition in a video stream according to claim 3, characterized in that, The method for detecting high dynamic visual transitions includes: The global motion vector is calculated by sparse optical flow for two adjacent normalized HSV preprocessed image frames. When the target confidence score is less than or equal to the preset loss judgment threshold or the magnitude of the global motion vector exceeds the preset inter-frame displacement threshold, a high dynamic visual transition is detected.
5. The real-time image processing method for dynamic target recognition in a video stream according to claim 4, characterized in that, The method for generating the dynamic target virtual trajectory includes: In response to the visual blind spot state, the target state vector and global motion vector from the previous moment are used as inputs to the Kalman filter to perform inertial prediction and background compensation, and to calculate the dynamic target virtual trajectory within the blind spot.
6. The real-time image processing method for dynamic target recognition in a video stream according to claim 5, characterized in that, The method for generating the alignment trajectory throughout the entire process includes: The real detection trajectory, which is formed by the virtual trajectory of the dynamic target and the centroid coordinates of the dynamic target in the target state vector, is time-stamped and smoothly stitched together to form a full-process alignment trajectory; the real detection trajectory includes two parts: the real detection trajectory before entering the visual blind zone state and the real detection trajectory after exiting the visual blind zone state.
7. A real-time image processing system for dynamic target recognition in a video stream, used to implement the real-time image processing method for dynamic target recognition in a video stream as described in any one of claims 1-6, characterized in that, The system includes: Dual-channel feature extraction module: used to establish a video stream buffer queue for handheld optical imaging terminal, acquire monocular raw video frames, construct a dual-channel feature extraction model based on attention mechanism and color space, and solve the visual reference center coordinates and dynamic target centroid coordinates in parallel from monocular raw video frames through the dual-channel feature extraction model, and generate target state vector; Visual alignment posture module: Generates alignment deviation vector flow based on the coordinates of the visual reference center and the centroid coordinates of the dynamic target, and deconstructs the operator's physiological tremor characteristics and operational jitter characteristics based on the alignment deviation vector flow to generate visual alignment posture data; Visual guidance module: It is used to monitor the target state vector in real time. When a high dynamic visual transition is detected, it switches to the visual blind zone state and generates a dynamic target virtual trajectory within the blind zone. After exiting the visual blind zone state, it generates a full-process alignment trajectory based on the dynamic target virtual trajectory, and performs spatial mapping and fusion with the visual alignment situation data to generate visual guidance instructions.
Citation Information
Patent Citations
A video data typical dynamic target detection method and system
CN109670488A
Video stream processing methods and devices
CN110569702B
Digital anti-shake method and digital anti-shake system for inspection video of comprehensive pipe gallery overhead rail robot
CN111935392A
Illegal action recognition method and device in video, medium and program product
CN121121867A