Fuzzy target auxiliary recognition method based on body feature and wearable visual device

By acquiring first-person perspective video frames in real time on a head-mounted device worn by law enforcement officers, constructing a confidence interval model of body features and dynamically assigning weights, the problem of identifying ambiguous targets in dynamic law enforcement scenarios is solved, achieving accurate and seamless auxiliary identification, and improving law enforcement efficiency and identification reliability.

CN122454597APending Publication Date: 2026-07-24FUTURE MAN (XIAMEN) ARTIFICIAL INTELLIGENCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUTURE MAN (XIAMEN) ARTIFICIAL INTELLIGENCE CO LTD
Filing Date
2026-04-02
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify ambiguous targets in dynamic law enforcement scenarios, especially when law enforcement officers wear head-mounted devices and their hands are restricted. Existing identification methods cannot adapt to the uncertainty of feature extraction in first-person perspectives and complex environments, resulting in poor identification reliability.

Method used

A fuzzy target-assisted recognition method based on body features is adopted. Wearable vision devices worn on the heads of law enforcement officers are used to collect first-person perspective video frames in real time, construct a confidence interval model of multi-dimensional body features, dynamically allocate feature weights, and perform fuzzy matching to generate augmented reality auxiliary information.

Benefits of technology

It enables real-time and accurate identification of ambiguous targets in dynamic law enforcement scenarios, reduces operational interference, improves law enforcement efficiency, enhances the robustness and real-time performance of identification, and meets the needs of seamless human-computer interaction.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

The application provides a fuzzy target auxiliary identification method based on body feature and a wearable visual device, through the wearable visual device worn on the head of a law enforcement personnel, real-time collection of continuous video frame images obtained in the first person perspective of the law enforcement personnel, human body detection and segmentation on a to-be-identified target contained in the continuous video frame images, extraction of multi-dimensional body feature of the to-be-identified target, acquisition of an environment parameter set of the continuous video frame images, dynamic allocation of feature weight sets of the multi-dimensional body feature in a matching process, fuzzy matching of the to-be-identified target and candidate targets in a preset database, and generation of auxiliary identification information on a display unit of the wearable visual device, which solves the problem of real-time identification of fuzzy targets in a dynamic law enforcement scene, has the advantages of being capable of assisting the law enforcement personnel in real-time and accurately identifying fuzzy targets, reducing operation interference, and improving law enforcement efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of wearable computing, artificial intelligence visual analysis, and public safety assistance technologies. Specifically, it relates to a method for fuzzy target assistance based on body features and a wearable vision device using this method. Background Technology

[0002] In the fields of public safety and law enforcement, law enforcement officers often need to quickly identify and assess suspicious persons in dynamic and complex environments. Traditional identification methods mainly rely on surveillance video captured by fixed cameras, using biometric identification technologies such as facial recognition and gait recognition to confirm identities.

[0003] However, such methods have revealed multiple technical bottlenecks in frontline law enforcement scenarios:

[0004] Most existing identification methods are based on fixed-view surveillance equipment, such as access control cameras and PTZ cameras. The images captured by these devices are from a third-person perspective, and the shooting angle and distance are relatively fixed, which cannot meet the needs of law enforcement officers to observe targets in real time from a first-person perspective while they are on the move. At the same time, law enforcement officers' hands are often occupied by police equipment or gear when performing their duties, making it difficult to operate additional acquisition or identification equipment. As a result, existing technologies lack real-time solutions for dynamic law enforcement perspectives.

[0005] In the field of fuzzy target recognition, existing technologies typically employ deterministic feature values ​​for matching, such as extracting height, limb proportions, and gait parameters as single values ​​and calculating distances with samples in a database. However, in complex environments such as low resolution, long distance, backlighting, and occlusion, a single deterministic value cannot characterize the uncertainty inherent in feature extraction, leading to poor reliability of the matching results. Existing technologies do not consider quantitative modeling of the confidence level of feature extraction and cannot dynamically adjust the matching tolerance based on image quality.

[0006] In conclusion, achieving accurate assisted identification of blurred targets with missing facial features or poor image quality, under conditions where law enforcement officers wear head-mounted devices and their hands are restricted, has become an urgent challenge for current technological development. Summary of the Invention

[0007] The purpose of this invention is to provide a method for assisting in the identification of fuzzy targets based on body features and a wearable vision device, which can accurately assist law enforcement officers in identifying fuzzy targets and improve law enforcement efficiency.

[0008] The technical solution of the present invention is as follows: A method for fuzzy target-assisted recognition based on body posture features, the method being applied to a head-mounted visual acquisition device worn by law enforcement officers at dynamic law enforcement scenes, comprising the following steps: Step S1: Using a wearable vision device worn on the head of law enforcement officers, continuously capture video frame images from the first-person perspective of the law enforcement officers in real time; Step S2: In response to the law enforcement officer's gaze focus or target lock command, perform human detection and segmentation on the target to be identified contained in the continuous video frame images, and extract multi-dimensional body features of the target to be identified; the multi-dimensional body features include at least height parameters, limb proportion parameters, and gait parameters; wherein, for each dimension of body feature, based on the imaging quality parameters, shooting angle parameters, and occlusion degree parameters of the current frame image, construct a confidence interval model corresponding to the dimension of body feature to characterize the reliability of the feature under the current observation conditions; Step S3: Obtain the environmental parameter set of the continuous video frame images; the environmental parameter set includes at least image resolution, light intensity, and dynamic shooting distance between the wearable vision device and the target to be identified; Step S4: Based on the environmental parameter set, dynamically allocate the feature weight set of the multi-dimensional body features in the matching process; wherein, when the dynamic shooting distance is greater than a preset distance threshold, increase the weight of the height parameter and limb proportion parameter, and decrease the weight of the gait parameter; when the dynamic shooting distance is less than or equal to the preset distance threshold and the continuous video frame images contain side view images, increase the weight of the gait parameter. Step S5: Based on the confidence interval models corresponding to each of the multi-dimensional body features and the feature weight set, perform fuzzy matching between the target to be identified and the candidate targets in the preset database, and calculate the comprehensive matching similarity; the fuzzy matching includes: for each dimension of body feature, calculating the interval overlap between the feature value and its confidence interval of the target to be identified and the corresponding feature value of the candidate target to obtain the single feature matching degree; and weighting and fusing each single feature matching degree according to the feature weight set to obtain the comprehensive matching similarity. Step S6: Based on the comprehensive matching similarity, generate auxiliary recognition information on the display unit of the wearable vision device.

[0009] Preferably, the construction of the confidence interval model in step S2 specifically includes: The error distribution range of feature extraction is determined based on the imaging quality parameters, which include at least image resolution and signal-to-noise ratio. The human body projection deformation coefficient is determined based on the shooting angle parameters, which are determined according to the relative azimuth angle between the wearable vision device and the target to be identified. The visible feature proportion factor is determined based on the occlusion degree parameter, and the occlusion degree parameter is determined based on the integrity proportion of the target to be identified in the image. The error distribution range, the deformation coefficient, and the proportion factor are used as parameters to input a preset uncertainty quantification model, and the confidence interval width of each dimension of body features is output. The confidence interval width is negatively correlated with the imaging quality parameter, positively correlated with the absolute value of the angle of deviation of the shooting angle parameter, and negatively correlated with the occlusion degree parameter.

[0010] Preferably, obtaining the dynamic shooting distance in step S3 specifically includes: The wearable vision device measures the straight-line distance between itself and the target to be identified in real time using a depth sensor built into it; or, Based on the proportion of human pixels of the target to be identified in the continuous video frame images, and combined with the preset average human height parameters, the dynamic shooting distance is calculated by a monocular visual ranging algorithm. Specifically, when the pixel percentage of the target to be identified in the image is lower than a first threshold, the dynamic shooting distance is determined to be greater than a preset distance threshold; when the pixel percentage is higher than a second threshold, the dynamic shooting distance is determined to be less than or equal to the preset distance threshold.

[0011] Preferably, the dynamic allocation of the feature weight set in step S4 further includes: The illumination intensity of the continuous video frame images is obtained. When the illumination intensity is lower than a preset illumination threshold, the weight of the limb proportion parameter and the gait parameter is increased, and the weight of the height parameter is decreased. The image resolution of the continuous video frame image is obtained. When the image resolution is lower than a preset resolution threshold, the weight of the limb proportion parameter is increased, the weight of the gait parameter is decreased, and the confidence interval width of the confidence interval model is expanded to a first preset multiple. The viewpoint type contained in the continuous video frame images is obtained. When the viewpoint type is a frontal view, the weights of the height parameter and the limb proportion parameter are increased; when the viewpoint type is a side view, the weights of the gait parameter are increased.

[0012] Preferably, step S2, which involves extracting the multi-dimensional body features of the target to be identified, further includes: Perform time series analysis on the continuous video frame images to extract the gait period of the target to be identified; Within the gait cycle, the joint angle change curve, stride change sequence, and gait frequency parameters of the target to be identified are extracted and together constitute the gait parameters; The joint angle change curves include at least the knee joint angle change curve and the hip joint angle change curve.

[0013] Preferably, in step S2, in response to the law enforcement officer's gaze focus or target locking command, human detection and segmentation are performed on the target to be identified contained in the continuous video frame images, specifically including: The eye-tracking module built into the wearable vision device obtains the coordinates of the law enforcement officer's gaze focus in the image coordinate system. A dynamic region of interest is constructed centered on the coordinates of the line of sight focus. A human detection algorithm is executed within the dynamic region of interest, prioritizing the detection and segmentation of targets to be identified located in the line of sight of the law enforcement officer.

[0014] Preferably, step S5 involves calculating the overlap between the feature values ​​and confidence intervals of the target to be identified and the corresponding feature values ​​of the candidate targets, specifically including: The feature value of the target to be identified is represented as a confidence interval centered on the detected value and with the width output by the confidence interval model as the radius; The corresponding feature values ​​of the candidate targets are represented as predetermined values ​​in a pre-defined database; The degree of overlap between the confidence interval and the determined value is calculated. When the determined value falls within the confidence interval, the single feature matching degree is a first preset value. When the determined value is outside the confidence interval, the single feature matching degree decreases linearly as the distance between the determined value and the boundary of the confidence interval increases.

[0015] Preferably, in step S5, the preset database stores a candidate target body feature library associated with the current law enforcement task, and the candidate target body feature library is synchronized in real time from the background command center to the local storage of the wearable vision device; the method further includes: When the wearable vision device establishes a communication connection with the back-end command center, it receives a task-level candidate target list issued by the back-end command center. The task-level candidate target list contains the body feature data of at least one candidate target associated with the current law enforcement area. The task-level candidate target list is stored in the local cache of the wearable vision device for real-time comparison of the fuzzy matching.

[0016] Preferably, step S6, which generates auxiliary recognition information on the display unit of the wearable vision device, specifically includes: When the overall matching similarity exceeds a preset recognition threshold, the identity information of the successfully matched candidate target and the matching confidence level are superimposed and displayed in the field of vision of the law enforcement personnel in an augmented reality manner; The augmented reality display includes: anchoring the auxiliary identification information to the spatial location of the target to be identified through the optical see-through display unit of the wearable vision device, so that the auxiliary identification information maintains a constant relative positional relationship with the target to be identified as the law enforcement officer's head moves.

[0017] The present invention also proposes a wearable vision device, comprising: Headset housing; An image acquisition unit is mounted on the head-mounted housing and is used to acquire continuous video frame images from a first-person perspective. An eye-tracking unit is used to obtain the coordinates of the wearer's gaze focus. A depth sensing unit is used to measure the distance between the wearable vision device and the target to be identified; An optical see-through display unit is used to overlay auxiliary identification information onto the wearer's field of vision in an augmented reality manner; A processor is connected to the image acquisition unit, the eye tracking unit, the depth sensing unit, and the optical perspective display unit, respectively, and the processor is configured to perform the steps of the method according to any one of claims 1 to 9; A memory, connected to the processor, is used to store a preset database and computer programs. As can be seen from the above, the present invention provides a method and wearable vision device for fuzzy target recognition based on body features. By acquiring first-person perspective video frames in real time, constructing a confidence interval model of body features, dynamically allocating feature weights, performing fuzzy matching, and generating augmented reality auxiliary information, it solves the problem of real-time recognition of fuzzy targets in dynamic law enforcement scenarios. It has the advantages of being able to assist law enforcement officers in recognizing fuzzy targets in real time and accurately, reducing operational interference, and improving law enforcement efficiency.

[0018] Compared with the prior art, the present invention has the following beneficial effects: (1) This invention achieves real-time target recognition from a first-person dynamic perspective by deploying the recognition method on a head-mounted visual acquisition device worn by frontline law enforcement officers. This method responds to the officer's gaze focus or target locking command for target detection, reducing computational resource consumption and meeting the needs of law enforcement scenarios where hands are occupied. By constructing a confidence interval model, the uncertainty of feature extraction is quantified, enabling the matching process to reflect the reliability of different features in the current scenario. By dynamically allocating feature weights based on environmental parameter sets, the matching strategy can adapt to changes in shooting distance, viewing angle, etc. The fuzzy matching method using interval overlap calculation and weighted fusion solves the problem of poor reliability of deterministic matching in fuzzy scenarios. Finally, auxiliary recognition information is generated through the display unit of the wearable visual device, achieving real-time feedback of the recognition results. Overall, this method solves the technical problem of real-time, reliable, and seamless auxiliary recognition of fuzzy targets by frontline law enforcement officers in dynamic scenarios.

[0019] (2) The present invention determines the error distribution range by imaging quality parameters, determines the projection deformation coefficient by shooting angle parameters, and determines the visible feature proportion factor by occlusion degree parameters. The three are input into the uncertainty quantification model, so that the confidence interval width is negatively correlated with imaging quality, positively correlated with shooting angle deviation, and negatively correlated with occlusion degree. This allows the confidence interval to adaptively widen in cases of image blurring, angle deviation, and target occlusion, avoiding misjudgment due to feature value deviation and improving the robustness of matching.

[0020] (3) The dynamic shooting distance acquisition method of the present invention includes direct measurement by a depth sensor or indirect calculation by combining human pixel ratio with monocular vision ranging algorithm, so that even in the absence of a depth sensor or sensor failure, distance parameters can still be obtained through image information, ensuring the redundancy and reliability of the system; at the same time, the pixel ratio-based determination method has a small computational load and is suitable for real-time operation on the low-power platform of wearable devices.

[0021] (4) When the light intensity is lower than the preset illuminance threshold, the reliability of the height parameter decreases (due to shadow interference), while the limb proportion and gait parameters are less affected, so the weight is adjusted; when the image resolution is lower than the threshold, gait detail information is lost, so the weight of the limb proportion parameter is increased and the confidence interval is widened, realizing refined adaptation to environmental changes, enabling the matching strategy to be dynamically optimized according to specific conditions such as light intensity and resolution, further improving the recognition accuracy in harsh environments.

[0022] (5) The gait parameter extraction method of the present invention includes gait cycle analysis, joint angle change curve, stride change sequence and stride frequency parameter, which refines gait features into multi-dimensional parameters with time series characteristics. Compared with a single gait energy map or simple stride parameter, it can more comprehensively characterize the unique walking pattern of an individual and improve the distinguishability and recognition accuracy of gait features.

[0023] (6) The present invention obtains the coordinates of the gaze focus through the eye-tracking module, constructs a dynamic region of interest, and prioritizes the detection of targets within the region. It utilizes the visual attention of law enforcement officers as prior information, which greatly narrows the search range of target detection, reduces computational overhead, and enables the system to quickly respond to the target that the user is concerned about, thereby improving the efficiency and real-time performance of human-machine collaboration.

[0024] (7) The present invention represents the target feature to be identified as a confidence interval and the candidate target feature as a definite value. The matching degree is determined by calculating the situation where the definite value falls into the interval and the degree of deviation, thus realizing the quantification of fuzzy matching: when the candidate target feature value falls within the confidence interval, a high matching degree is given; when it falls outside the interval, the matching degree decreases with the increase of distance. This conforms to the physical intuition of "approximate matching" in fuzzy scenarios and is more reasonable than the traditional point-to-point distance calculation.

[0025] (8) This invention distributes a task-level candidate target list in real time through the background command center and stores it in a local high-speed cache. This technical feature enables the real-time association between the candidate target library and the current law enforcement task in dynamic law enforcement scenarios, avoiding network latency and bandwidth dependence of real-time comparison in the cloud, ensuring the real-time performance and reliability of on-site identification, and meeting the requirements of law enforcement data security and controllability.

[0026] (9) When the match is successful, the present invention further overlays the identity information and the matching confidence in an augmented reality manner and anchors it to the spatial position of the target. The relative position remains unchanged as the law enforcement officer’s head moves, realizing the spatial fusion of the recognition result and the physical world. This allows law enforcement officers to obtain information without looking down at the device, which meets the front-line law enforcement’s need for “unobtrusive” human-computer interaction and avoids the problem of information loss due to head movement. Detailed Implementation

[0027] The technical solutions of the present invention will be clearly and completely described below with reference to specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0028] Example 1 (Method) Traditional identification methods primarily rely on fixed-view monitoring equipment, which is unsuitable for law enforcement officers' dynamic first-person perspective enforcement scenarios and is inconvenient to operate when their hands are occupied. In fuzzy target recognition, existing technologies use deterministic feature value matching, failing to quantify the uncertainty of feature extraction, resulting in poor matching reliability in complex environments. Furthermore, statically preset feature weights cannot dynamically adapt to environmental changes, limiting recognition accuracy. The display method of recognition results is distracting for law enforcement officers and lacks integration with augmented reality headsets.

[0029] To address this, this embodiment proposes a fuzzy target-assisted recognition method based on body posture features. This method is applied to head-mounted visual acquisition devices worn by law enforcement officers at dynamic law enforcement scenes, and includes: Step S1: Real-time acquisition of continuous video frame images from the first-person perspective of the law enforcement officer by using a wearable vision device worn on the head of the officer. Step S2: In response to the law enforcement officer's gaze focus or target lock command, perform human detection and segmentation on the target to be identified contained in the continuous video frame images, and extract the multi-dimensional body features of the target to be identified; the multi-dimensional body features include at least height parameters, limb proportion parameters, and gait parameters; wherein, for each dimension of body feature, based on the imaging quality parameters, shooting angle parameters, and occlusion degree parameters of the current frame image, construct the confidence interval model corresponding to the dimension of body feature to characterize the reliability of the feature under the current observation conditions; Step S3: Obtain the environmental parameter set of the continuous video frame images; the environmental parameter set includes at least the image resolution, light intensity, and dynamic shooting distance between the wearable vision device and the target to be identified; Step S4: Based on the environmental parameter set, dynamically allocate the feature weight set of the multi-dimensional body features in the matching process; wherein, when the dynamic shooting distance is greater than the preset distance threshold, increase the weight of the height parameter and the limb proportion parameter, and decrease the weight of the gait parameter; when the dynamic shooting distance is less than or equal to the preset distance threshold and the continuous video frame image contains a side view image, increase the weight of the gait parameter. Step S5: Based on the confidence interval model corresponding to each of the multi-dimensional body features and the feature weight set, perform fuzzy matching between the target to be identified and the candidate targets in the preset database, and calculate the comprehensive matching similarity; the fuzzy matching includes: for each dimension of body feature, calculating the interval overlap between the feature value and its confidence interval of the target to be identified and the corresponding feature value of the candidate target to obtain the single feature matching degree; and weighting and fusing the single feature matching degrees according to the feature weight set to obtain the comprehensive matching similarity. Step S6: Based on the comprehensive matching similarity, generate auxiliary recognition information on the display unit of the wearable vision device.

[0030] For ease of understanding, the following explains some key terms in this embodiment: Wearable vision devices are devices that can be worn on the head by law enforcement officers and integrate image acquisition, processing, and display functions, such as AI glasses. These devices aim to provide law enforcement officers with real-time, convenient visual information acquisition and interaction capabilities, and are particularly suitable for dynamic law enforcement scenarios.

[0031] First-person perspective refers to the acquisition of video or images by the image acquisition unit from the perspective of the wearer (i.e., law enforcement officer). This perspective can realistically reflect the on-site observation of law enforcement officers, facilitating subsequent analysis and auxiliary identification.

[0032] A gaze focus or target lock instruction refers to a clear indication by law enforcement officers, through eye tracking, gestures, voice, or other interactive methods, of the area of ​​image they are currently focusing on or a specific target. This instruction guides the system to prioritize targets of interest to law enforcement officers, thereby improving recognition efficiency.

[0033] A confidence interval model refers to constructing an interval range for each dimension of body features, based on observation conditions such as image quality, shooting angle, and degree of occlusion, to quantify the uncertainty or reliability of the feature extraction result. This model can reflect the fluctuation range of feature values ​​under different observation conditions.

[0034] The feature weight set refers to the relative importance coefficients assigned to different dimensions of body features during fuzzy matching. These weights are dynamically adjusted to reflect the degree of contribution of different features to the recognition results under specific environmental conditions.

[0035] Assisted identification information refers to information presented to law enforcement officers on the display unit of a wearable vision device regarding the identity, matching degree, or other relevant prompts of the target to be identified. This information is intended to assist law enforcement officers in making on-site decisions.

[0036] The specific implementation method is described in detail below with reference to each step of this embodiment: In step S1, a wearable vision device worn on the head of the law enforcement officer captures continuous video frame images from the officer's first-person perspective in real time. This wearable vision device can be AI glasses, or a head-mounted device such as glasses that integrates a small camera, such as a wide-angle camera or a standard focal length camera, configured to continuously record video streams within the officer's field of vision. These video frame images are transmitted to a processor inside the device for further processing.

[0037] In step S2, in response to the law enforcement officer's gaze focus or target locking command, human detection and segmentation are performed on the target to be identified contained in the continuous video frame images, and multi-dimensional body features of the target to be identified are extracted. These multi-dimensional body features include at least height parameters, limb proportion parameters, and gait parameters. Specifically, the law enforcement officer can indicate the target of their interest through voice commands or head posture. After receiving the command, the system runs a general human detection algorithm across the entire video frame image to identify all possible human targets and accurately segment the target indicated by the command. Subsequently, static body features such as height and limb length proportions are extracted from the segmented human body region, and gait features are obtained by analyzing changes in human posture in continuous frame images; for example, the average stride length and average stride frequency can be simply calculated. For each dimension of body feature, a confidence interval model corresponding to that dimension of body feature is constructed based on the imaging quality parameters, shooting angle parameters, and occlusion degree parameters of the current frame image to characterize the reliability of the feature under the current observation conditions. For example, a lookup table can be preset to determine a fixed confidence interval width based on image resolution, shooting angle (such as front or back) and occlusion ratio (such as no occlusion or partial occlusion). When the resolution is low, the angle is skewed, or the occlusion is severe, the confidence interval is set to be wider.

[0038] In step S3, an environmental parameter set for the continuous video frame images is acquired. This environmental parameter set includes at least image resolution, illumination intensity, and the dynamic shooting distance between the wearable vision device and the target to be identified. Image resolution and illumination intensity can be obtained directly from the metadata of the video frame images or through image analysis algorithms. The dynamic shooting distance can be obtained in various ways; for example, law enforcement officers can manually input an estimated distance, or the system can make a rough estimate based on the size of the target in the image and a preset average human body size.

[0039] In step S4, the feature weight set of the multi-dimensional body posture features is dynamically allocated in the matching process based on the environmental parameter set. Specifically, when the dynamic shooting distance is greater than a preset distance threshold (e.g., when the target is far away), the system increases the weight of height and limb proportion parameters while decreasing the weight of gait parameters because the details of gait features are difficult to capture accurately. When the dynamic shooting distance is less than or equal to the preset distance threshold and the continuous video frames contain side-view images (e.g., when the target is close and side features are obvious), the system increases the weight of gait parameters to utilize their discriminative power at close range and from the side. This weight allocation can be implemented based on a preset rule set or conditional judgment logic.

[0040] In step S5, based on the confidence interval models corresponding to each of the multi-dimensional body features and the feature weight set, the target to be identified is fuzzily matched with candidate targets in a preset database to calculate the comprehensive matching similarity. The preset database can be a database stored locally on the wearable vision device, containing a set of pre-recorded body feature data for candidate targets. The specific process of fuzzy matching includes: for each dimension of body feature, calculating the interval overlap between the feature value of the target to be identified and its confidence interval and the corresponding feature value of the candidate target to obtain the single feature matching degree. For example, it can be simply determined whether the feature value of the target to be identified falls within a preset error range of the candidate target's feature value; if it falls within the range, the matching degree is high; otherwise, it is low. Subsequently, the single feature matching degrees are weighted and fused according to the feature weight set to obtain the comprehensive matching similarity.

[0041] In step S6, auxiliary identification information is generated on the display unit of the wearable vision device based on the comprehensive matching similarity. When the comprehensive matching similarity reaches a preset level, for example, exceeding a preset threshold, the system can display a "potential match" prompt or the candidate target's number on the display unit. This display unit can be a small display screen on the device, which law enforcement officers need to view by shifting their gaze.

[0042] This embodiment achieves reliable identification of ambiguous targets by acquiring first-person perspective images using a head-mounted vision device at dynamic law enforcement scenes and combining confidence interval modeling of multi-dimensional body features with an adaptive weight allocation mechanism for environmental parameters. This method effectively copes with complex and changing environmental conditions, overcomes the limitations of traditional recognition technologies in low-quality images and dynamic scenes, and improves the identification efficiency and on-site decision-making ability of law enforcement officers when their hands are occupied by generating auxiliary recognition information on wearable devices.

[0043] This embodiment further proposes to construct the confidence interval model, specifically including: The error distribution range of feature extraction is determined based on the imaging quality parameters, which include at least image resolution and signal-to-noise ratio. The human body projection deformation coefficient is determined based on the shooting angle parameters, which are determined according to the relative azimuth angle between the wearable vision device and the target to be identified. The visible feature proportion factor is determined based on the occlusion degree parameter, and the occlusion degree parameter is determined based on the integrity proportion of the target to be identified in the image. The error distribution range, the deformation coefficient, and the proportion factor are used as parameters to input a preset uncertainty quantification model, and the confidence interval width of each dimension of body features is output. The confidence interval width is negatively correlated with the imaging quality parameter, positively correlated with the absolute value of the angle of deviation of the shooting angle parameter, and negatively correlated with the occlusion degree parameter.

[0044] Image quality parameters, such as image resolution and signal-to-noise ratio (SNR), directly affect the accuracy of extracting body features from images. When the image resolution is low or the SNR is poor, the error in feature point localization increases, leading to increased uncertainty in feature values. To address this, a mapping relationship between image quality parameters and the distribution range of feature extraction errors can be established beforehand, for example, through fitting experimental data or modeling based on image processing theory. For instance, depending on the image resolution, the pixel-level error of feature points can be modeled as different Gaussian or uniform distributions, thereby quantifying the error range.

[0045] The shooting angle parameters, especially the relative azimuth angle between the wearable vision device and the target to be identified, cause distortion in the projection of the human body onto the two-dimensional image, thus affecting the accurate measurement of body features such as height and limb proportions. For example, when shooting from a side or pitch angle, the visual height of the human body changes significantly. This distortion coefficient aims to quantify the impact of this projection distortion on feature values. The projection distortion factor for different body features can be calculated based on the relative azimuth angle using the principles of geometric projection. For example, for the height parameter, its projection scaling ratio on the image plane can be calculated based on the shooting angle.

[0046] The occlusion parameter, which represents the proportion of the target's integrity in the image, directly affects the visibility and extractability of body features. When a target is partially occluded, certain body features (such as a complete gait cycle or limb joints) may not be fully observable, thus reducing their reliability. The visible feature proportion factor measures the proportion of a specific body feature that can be effectively observed in the current image. For example, if the lower body of the target is occluded, the visible feature proportion factor of the gait parameter will significantly decrease, or even become zero. This factor can be estimated by comparing the image segmentation results with a pre-defined human skeleton model.

[0047] To comprehensively evaluate the impact of various observation conditions on the reliability of body features, this embodiment employs a pre-defined uncertainty quantification model. This model receives the aforementioned determined error distribution range, deformation coefficient, and visible feature proportion factor as input. This model can be a machine learning-based regression model that learns the patterns of feature uncertainty under different observation conditions through a large amount of labeled data; or it can be a composite function model built based on expert knowledge and statistical principles, which weights and combines various input parameters to output a comprehensive uncertainty measure.

[0048] The core output of the uncertainty quantification model is the confidence interval width of each dimension of body posture features. This width directly reflects the reliability of the extracted posture feature values ​​under the current observation conditions. A larger width indicates higher uncertainty and less reliable feature values; a smaller width indicates lower uncertainty and more reliable feature values. For example, when imaging quality parameters (such as resolution and signal-to-noise ratio) increase, the feature extraction error decreases, and the confidence interval width should decrease accordingly. When the absolute value of the shooting angle deviating from the standard viewing angle (such as frontal or side view) increases, the distortion of the human body projection intensifies, increasing feature uncertainty, and the confidence interval width should increase accordingly. When occlusion parameters (such as integrity ratio) decrease, visible feature information decreases, feature reliability declines, and the confidence interval width should increase accordingly.

[0049] This embodiment further proposes a specific method for obtaining dynamic shooting distance. This method includes two main approaches: One approach involves using a depth sensor built into the wearable vision device to measure the straight-line distance between the device and the target in real time. Depth sensors typically employ active detection technologies, such as time-of-flight (ToF) or structured light, to accurately calculate the distance to the target object by emitting specific signals (like infrared light) and measuring their round-trip time or analyzing their deformation patterns. This direct measurement method offers high accuracy and real-time performance, making it particularly suitable for short- or medium-range ranging scenarios.

[0050] Another approach is to calculate the dynamic shooting distance based on the proportion of human pixels in the target image within consecutive video frames, combined with preset average human height parameters, using a monocular vision ranging algorithm. Monocular vision ranging estimates distance by utilizing the perspective relationship between the size of the target object in the image and its actual size. Specifically, the system first performs human detection on the target image and calculates its pixel area or height within the image. Then, based on preset average human height parameters and internal parameters such as the camera's focal length and pixel size, it uses geometric perspective principles or a pre-trained deep learning model to estimate the distance between the target and the wearable vision device. This method provides an effective alternative when a dedicated depth sensor is unavailable or sensor limitations exist.

[0051] When the pixel percentage of the target in the image is below a first threshold, the system determines that the dynamic shooting distance is greater than a preset distance threshold. This usually means that the target is far away and appears as a small area in the image. Conversely, when the pixel percentage is above a second threshold, the system determines that the dynamic shooting distance is less than or equal to the preset distance threshold. This indicates that the target is close and occupies a large area in the image. By setting different first and second thresholds, the boundaries of near and far distances can be flexibly defined, providing a fast and coarse distance judgment for subsequent feature weight allocation.

[0052] Through the above technical solution, this embodiment can flexibly select or integrate two ranging methods—direct depth sensor measurement and monocular vision ranging algorithm—based on the complexity and diversity of actual law enforcement scenarios, thereby ensuring the accuracy and robustness of dynamic shooting distance acquisition in different scenarios. For example, in close-range scenarios with sufficient lighting and clear targets, depth sensors can provide high-precision data; while in long-range scenarios with insufficient lighting, blurred targets, or limited sensors, monocular vision ranging can still provide effective distance estimation. In particular, quickly determining the distance range by the proportion of human body pixels further improves the real-time performance and efficiency of ranging, avoiding complex and precise ranging calculations in all situations. This accurate and efficient distance acquisition mechanism provides a reliable basis for the dynamic weight allocation of multi-dimensional body features in the subsequent matching process. For example, when the distance is far, the weight of height and limb proportion parameters can be increased, while the weight of gait parameters can be decreased, and vice versa.

[0053] This embodiment further proposes the following considerations when dynamically allocating the feature weight set: The illumination intensity of the continuous video frame images is obtained. When the illumination intensity is lower than a preset illumination threshold, the weight of the limb proportion parameter and the gait parameter is increased, and the weight of the height parameter is decreased. The image resolution of the continuous video frame image is obtained. When the image resolution is lower than a preset resolution threshold, the weight of the limb proportion parameter is increased, the weight of the gait parameter is decreased, and the confidence interval width of the confidence interval model is expanded to a first preset multiple. The viewpoint type contained in the continuous video frame images is obtained. When the viewpoint type is a frontal view, the weights of the height parameter and the limb proportion parameter are increased; when the viewpoint type is a side view, the weights of the gait parameter are increased.

[0054] Specifically, regarding the acquisition of illumination intensity from the continuous video frame images, the ambient light intensity can be directly measured using the light sensor built into the wearable vision device, or image processing analysis can be performed on the video frame images, such as calculating the average brightness value of the image and performing histogram analysis, to quantify the current scene's illumination conditions. A preset illumination threshold is a lower limit of brightness set based on actual application scenarios and experience, used to determine whether the current illumination is sufficient. When the detected illumination intensity is lower than this preset threshold, it indicates insufficient ambient light, and image quality may degrade. In this case, to improve recognition accuracy, the system dynamically adjusts the feature weights, specifically increasing the weights of limb proportion parameters and gait parameters while decreasing the weight of height parameters. This is because under low light conditions, the overall outline of the target (affecting the accurate measurement of height parameters) may become blurred, while the relative periodic features of limb proportions and gait can still maintain their relative stability to a certain extent, thus providing a more reliable recognition basis in low-light environments.

[0055] Regarding the acquisition of image resolution for the continuous video frames, the wearable vision device can acquire the image resolution of the currently captured video frame in real time. A preset resolution threshold is a lower limit set based on the system's requirements for feature extraction accuracy. When the image resolution is lower than this preset threshold, it means that image details are severely lost, significantly impacting the extraction of fine features (such as joint angle changes in gait parameters). In this case, the system increases the weight of limb proportion parameters and decreases the weight of gait parameters, because limb proportion, as a relative feature, can still maintain its effectiveness to a certain extent, while gait parameter extraction is easily distorted at low resolution. Simultaneously, to reflect the increased uncertainty in feature extraction due to image quality degradation, the confidence interval width of the confidence interval model is expanded to a first preset multiple to encompass possible errors over a wider range, avoiding misjudgments caused by low-quality data. This first preset multiple can be set based on experimental data or experience.

[0056] Regarding the acquisition of the viewpoint types contained in the continuous video frame images, the viewpoint of the target to be identified relative to the wearable vision device can be determined by performing human pose estimation or target detection algorithms on the video frame images, combined with geometric relationship analysis. Viewpoint types can be categorized as frontal viewpoint, side viewpoint, etc. When determined to be a frontal viewpoint, the target to be identified is directly or approximately directly facing the camera. In this case, the extraction of height and limb proportion parameters is relatively stable and accurate, thus increasing their weight. When determined to be a side viewpoint, the target to be identified is sideways or approximately sideways to the camera. In this case, the periodicity and dynamism of gait features (such as stride length, cadence, and joint angle changes) are more pronounced and easier to extract and analyze, thus increasing the weight of gait parameters.

[0057] Through the above technical solution, this embodiment, when dynamically allocating the weights of multi-dimensional body features, not only considers the dynamic shooting distance but also comprehensively considers various environmental factors such as light intensity, image resolution, and shooting angle. When the light intensity is low, by increasing the weights of limb proportion parameters and gait parameters and decreasing the weight of height parameters, it is possible to effectively utilize features that are relatively more recognizable in low-light environments for matching. When the image resolution is low, by increasing the weight of limb proportion parameters and decreasing the weight of gait parameters, while expanding the confidence interval width, it is possible to adapt to the uncertainty of feature extraction caused by the decline in image quality and avoid misjudgment due to ambiguous information. In addition, according to different shooting angles, such as emphasizing height and limb proportions in a frontal view and gait parameters in a side view, the feature weight allocation becomes more refined and reasonable. This multi-dimensional, adaptive weight allocation strategy significantly improves the accuracy and robustness of identifying ambiguous targets in complex and dynamic law enforcement scenes, ensuring that it can still effectively assist law enforcement officers in target identification under various adverse observation conditions.

[0058] In step S2, when extracting the multi-dimensional body features of the target to be identified, the method further includes performing time series analysis on the continuous video frame images to extract the gait cycle of the target to be identified. Within the gait cycle, the joint angle change curves, stride length change sequences, and gait frequency parameters of the target to be identified are extracted, which together constitute the gait parameters. The joint angle change curves include at least the knee joint angle change curve and the hip joint angle change curve.

[0059] Specifically, performing time-series analysis on the continuous video frame images to extract the gait cycle of the target to be identified is the foundation of refined gait analysis. The gait cycle defines a complete gait loop, such as the time point from when one foot lands to when the same foot lands again. Accurate extraction of the gait cycle is crucial for subsequent calculation of refined gait parameters. In practice, various time-series analysis methods can be used to detect the gait cycle. For example, the periodic changes of key human body points (such as ankles, knees, and hips) in continuous frames can be detected by analyzing their motion trajectories. Commonly used methods include period detection based on motion energy, which calculates the motion energy of key human body points in the vertical direction; when the motion energy reaches a local peak or trough, it may correspond to a specific stage of the gait cycle. Alternatively, autocorrelation analysis can be performed on the motion signals of key points; the peak of the autocorrelation function at a specific lag is the gait cycle. Fourier transform-based methods can also be used to convert the motion signals of key points to the frequency domain, identify the dominant frequency, and thus infer the gait cycle. During implementation, issues such as motion noise, viewpoint changes, and occlusion need to be addressed. It may be necessary to combine Kalman filtering or smoothing algorithms to preprocess the key point trajectory to ensure the accuracy of periodic detection.

[0060] Within the gait cycle, the joint angle change curves, stride length change sequences, and gait frequency parameters of the target to be identified are extracted, which together constitute the gait parameters. These parameters are key indicators describing an individual's walking characteristics. Refining their extraction allows for a more comprehensive and detailed characterization of an individual's gait features, improving the accuracy and discriminative power of identification.

[0061] The extraction of joint angle variation curves involves identifying key points on the human body (such as the hip, knee, and ankle) and then calculating the angle changes of the corresponding joints (such as the knee and hip joints) within each gait cycle based on the coordinates of these key points. For example, the knee joint angle can be calculated using the vector angle between the femur and tibia. The sequence of these angle changes over time constitutes the joint angle variation curve.

[0062] Extracting a stride variation sequence refers to stride length, typically defined as the distance between two landings of the same-side foot or the distance between the two feet when the opposite-side foot lands. In continuous video frames, stride length can be obtained by tracking the positions of key points on the feet and calculating their displacement on the ground projection. Since stride length may vary slightly within different time periods, extracting a stride variation sequence better reflects dynamic characteristics.

[0063] Step frequency extraction refers to the number of steps taken per unit of time. After extracting the gait cycle, step frequency can be calculated directly by the reciprocal of the gait cycle, or by counting the number of gait cycles detected within a certain time window.

[0064] It is important to note that joint angle calculation requires accurate key point detection and skeleton estimation, stride calculation requires consideration of perspective distortion and ground plane estimation, and cadence calculation requires stable period detection.

[0065] The joint angle variation curves include at least the knee joint angle variation curve and the hip joint angle variation curve. The knee and hip joints are the two most important joints in human lower limb movement, and their angle variation patterns significantly distinguish individual gait characteristics. Explicitly specifying the angle variation curves of these two joints emphasizes the focus on core gait features. Using the coordinates of key points in the human skeleton (such as the hip, knee, and ankle), geometric methods are used to calculate the angles of the knee and hip joints in consecutive frames. For example, the knee joint angle can be determined by the angle between the femur (hip-knee line) and the tibia (knee-ankle line); the hip joint angle can be determined by the angle between the trunk (shoulder-hip line) and the femur (hip-knee line). These angle variation curves over time reflect the flexion and extension patterns of the lower limbs during walking. In practical applications, angle calculation needs to consider joint movement in three-dimensional space, which may require three-dimensional pose estimation or approximation using projective geometry.

[0066] By performing time-series analysis on continuous video frames and refining the gait cycle of the target to be identified, this embodiment provides an accurate time reference for subsequent gait feature extraction. Based on this, joint angle variation curves (especially knee and hip joint angle variation curves), stride length variation sequences, and gait frequency parameters are further extracted to construct richer and more refined gait parameters. This meticulous description of gait features significantly enhances the discriminative power and robustness of gait features, especially in dynamic law enforcement scenarios, where it can more accurately capture the unique movement patterns of individual targets when facing blurred, distant, or partially occluded targets. Compared to schemes that only extract general gait parameters, this scheme effectively improves the accuracy and reliability of assisted identification of blurred targets by providing more distinctive gait features. Furthermore, this embodiment proposes a specific method for human detection and segmentation of the target to be identified contained in the continuous video frame images in response to the law enforcement officer's gaze focus or target locking command. The method includes: obtaining the coordinates of the law enforcement officer's gaze focus in the image coordinate system through the eye-tracking module built into the wearable vision device; constructing a dynamic region of interest centered on the gaze focus coordinates; and executing a human detection algorithm within the dynamic region of interest to prioritize the detection and segmentation of the target to be identified located in the direction of the law enforcement officer's gaze.

[0067] The wearable vision device incorporates an eye-tracking module, a sensor system integrated within the device. Its primary function is to monitor law enforcement officers' eye movements in real time and accurately calculate the coordinates of their gaze point within the current video frame. This module typically employs an infrared emitter and receiver, combined with a high-resolution miniature camera, to infer the direction of the gaze by analyzing changes in corneal reflection and pupil center position. The acquired gaze focus coordinates form the basis for determining the area of ​​interest in subsequent processing, ensuring the system accurately understands the officer's intent.

[0068] Based on this, a dynamic region of interest (ROI) is constructed centered on the coordinates of the gaze focus. A ROI is a local area with a specific shape and size defined within the current video frame, centered on the gaze focus coordinates of the law enforcement officer. The shape of this region can be rectangular, circular, or elliptical, and its size can be adaptively adjusted according to the actual application scenario, the stability of the law enforcement officer's gaze, the target distance, or a preset detection range. The purpose of constructing a ROI is to limit subsequent image processing to the area of ​​actual interest to the law enforcement officer, thereby effectively reducing unnecessary computation and improving processing efficiency.

[0069] A human detection algorithm is executed within the dynamically defined region of interest (ROI), prioritizing the detection and segmentation of targets within the law enforcement officer's line of sight. Within the constructed ROI, the system invokes a pre-defined human detection algorithm to identify and locate any potential human targets. Subsequently, a human segmentation algorithm is used to perform pixel-level precision separation of the detected human targets. By limiting human detection and segmentation to the ROI, the system prioritizes targets within the law enforcement officer's line of sight, avoiding ineffective processing of targets outside the region or those of non-interest. This prioritization ensures that, in complex multi-target scenarios, the system can efficiently and accurately pinpoint the targets of genuine interest to the law enforcement officer.

[0070] In step S5, the overlap between the feature values ​​and confidence intervals of the target to be identified and the corresponding feature values ​​of the candidate targets is calculated, specifically including: The feature value of the target to be identified is represented as a confidence interval centered on the detected value and with the width output by the confidence interval model as the radius; The corresponding feature values ​​of the candidate targets are represented as predetermined values ​​in a pre-defined database; Calculate the degree of overlap between the confidence interval and the determined value. When the determined value falls within the confidence interval, the single feature matching degree is a first preset value. When the determined value is outside the confidence interval, the single feature matching degree decreases linearly as the distance between the determined value and the boundary of the confidence interval increases.

[0071] In dynamic law enforcement scenarios, image data acquired in real-time by wearable vision devices is often affected by various environmental factors such as lighting, shooting angle, and occlusion, leading to inherent uncertainties in the extracted physical features of the target being identified. To accurately reflect this uncertainty, this embodiment treats the feature value of the target (e.g., the preliminary height measurement obtained through image processing algorithms) as the central estimate of that feature. Simultaneously, combined with the confidence interval model constructed in step S2, this model outputs a quantified range of uncertainty, i.e., the confidence interval width, based on current imaging quality parameters, shooting angle parameters, and occlusion degree parameters. This width is used as the radius to construct an interval around the central estimate, thereby representing the feature value of the target being identified as a confidence interval with a certain range. This representation effectively captures the inherent ambiguity of real-time observation data, providing a more realistic and robust input for subsequent fuzzy matching.

[0072] Unlike the real-time collected features of the target to be identified, the physical features of candidate targets stored in the pre-set database are usually the result of multiple measurements, calibrations, or manual input, possessing high accuracy and certainty. For example, information such as the suspect's height and limb proportions recorded in the database are typically considered precise, error-free known quantities. Therefore, during matching, the corresponding feature value of the candidate target is treated as a single, precise, and definite value, rather than an interval. This approach reflects the reliability of the database information and contrasts with the ambiguity of the target's features, providing a clear reference point for subsequent interval overlap calculations.

[0073] After representing the features of the target to be identified as confidence intervals and the features of the candidate target as deterministic values, the core task is to quantify the similarity between the two. This embodiment achieves this by calculating the degree of overlap between the confidence intervals and the deterministic values. Here, "degree of overlap" does not simply refer to the geometric intersection of intervals, but rather to a generalized matching metric. It aims to evaluate the extent to which the deterministic values ​​of the candidate targets match the fuzzy features of the target to be identified.

[0074] When the certainty value of a candidate target lies entirely within the confidence interval of the target to be identified, it indicates that the candidate target highly matches the target to be identified in that specific body feature. That is, under the current observational uncertainty, the candidate target features are compatible with the detected target features. In this ideal situation, this embodiment sets the matching degree of this single feature to a first preset value. This first preset value is typically a high fixed value (e.g., 1 or close to 1), representing a strong matching relationship to highlight the importance of this high degree of similarity. This ensures that the highest single-feature matching score can be obtained when the precise information in the database perfectly matches the ambiguity range of real-time observations.

[0075] In many real-world scenarios, the determined value of a candidate target may not precisely fall within the confidence interval of the target to be identified. To avoid simply classifying such cases as "mismatches," this embodiment introduces a more refined matching measurement mechanism. When the determined value is outside the confidence interval, the system calculates the distance between the determined value and the nearest boundary of the confidence interval. The single-feature matching degree decreases linearly as this distance increases. This means that even if the determined value is not within the interval, if it is close to the interval boundary, it can still obtain a certain matching score, but the score will gradually decrease with increasing distance until it reaches a preset minimum matching degree or zero. This linearly decreasing mechanism makes fuzzy matching more flexible and robust, capable of capturing near-matching situations and avoiding misjudgments or missed judgments that may be caused by hard threshold judgments.

[0076] This embodiment further proposes that the preset database in step S5 stores a candidate target body feature library associated with the current law enforcement task, and the candidate target body feature library is synchronized in real time from the background command center to the local storage of the wearable vision device; the method also includes: when the wearable vision device establishes a communication connection with the background command center, receiving a task-level candidate target list issued by the background command center, the task-level candidate target list containing body feature data of at least one candidate target associated with the current law enforcement area; storing the task-level candidate target list in the local cache of the wearable vision device for real-time comparison of the fuzzy matching.

[0077] Specifically, the preset database does not indiscriminately store all possible candidate target information. Instead, it pre-screens and stores physical characteristic data of candidate targets that are highly relevant to the specific needs of the current law enforcement task. For example, if the law enforcement task targets suspects in a specific area, the database only contains physical characteristic information of suspects in that area. This associative storage aims to reduce interference from invalid data and improve the efficiency and accuracy of matching.

[0078] The candidate target body feature database is synchronized in real time from the back-end command center to the local storage of the wearable vision device. As the core of the entire law enforcement system, the back-end command center is responsible for managing and maintaining the body feature data of all candidate targets. When a new law enforcement task is initiated or existing task information is updated, the command center, based on the task's relevance, transmits and updates the corresponding candidate target body feature database in real time via network connection to the local storage of the wearable vision device worn by law enforcement officers. This real-time synchronization mechanism ensures that the database used by law enforcement officers on-site is always up-to-date and highly relevant, and even in the event of a network connection interruption, the locally stored database can still support offline matching.

[0079] When the wearable vision device establishes a communication connection with the back-end command center, the wearable vision device receives a task-level candidate target list issued by the back-end command center. The wearable vision device and the back-end command center establish a communication connection via a wireless network (such as Wi-Fi, 4G / 5G, etc.). Once the connection is established, the command center will actively or passively send a "task-level candidate target list" to the wearable vision device. This list is a set of candidate targets highly relevant to the current law enforcement area, selected from massive amounts of data based on factors such as the geographical location, time, and event type of the current law enforcement task. For example, if law enforcement officers enter a specific block, the command center will issue physical characteristic data of key individuals of interest within that block.

[0080] The task-level candidate target list contains the physical characteristic data of at least one candidate target associated with the current law enforcement area. This list is not merely a target ID, but directly includes detailed physical characteristic data of these candidate targets, such as height, limb proportions, and gait parameters. This data is preprocessed and standardized, and can be directly used for fuzzy matching. By directly distributing candidate target data associated with the current law enforcement area, it ensures that law enforcement officers can quickly and accurately obtain the most needed target information for comparison when performing tasks in a specific area, avoiding inefficient searches across a large database.

[0081] The task-level candidate target list is stored in the local cache of the wearable vision device. The received task-level candidate target list is stored in the local cache (such as RAM or high-speed flash memory) within the wearable vision device. The cache features fast read / write speeds and low access latency, making it ideal for storing data that requires frequent access and comparison. This storage method allows the fuzzy matching process to be performed locally, without frequent network access to remote servers, thus significantly improving matching speed and response time.

[0082] The task-level candidate target list is used for real-time comparison in the fuzzy matching. The task-level candidate target list, stored in a local cache, primarily serves to provide a fast, real-time comparison data source for the fuzzy matching in step S5. After the wearable vision device extracts the multi-dimensional body features of the target to be identified, it can directly compare them with the task-level candidate target list in the local cache to calculate the comprehensive matching similarity.

[0083] Preferably, this embodiment also proposes generating auxiliary identification information on the display unit of the wearable vision device, specifically including: when the comprehensive matching similarity exceeds a preset identification threshold, superimposing the successfully matched candidate target identity information and matching confidence level on the law enforcement officer's field of vision in an augmented reality manner.

[0084] The augmented reality display includes: anchoring the auxiliary identification information to the spatial location of the target to be identified through the optical see-through display unit of the wearable vision device, so that the auxiliary identification information maintains a constant relative positional relationship with the target to be identified as the law enforcement officer's head moves.

[0085] When the overall matching similarity calculated by the system reaches or exceeds a preset recognition threshold, it indicates that the target to be identified has a high degree of matching with a candidate target in the database. At this time, the system will trigger the display of auxiliary identification information. This recognition threshold can be flexibly configured according to the needs of actual law enforcement scenarios. For example, a higher threshold can be set in scenarios where the accuracy of identification is extremely important, while a lower threshold can be appropriately set in scenarios where a broad screening is required. The displayed information includes the identity information of the successfully matched candidate target, such as its name, number, or other key identifiers related to the case, as well as the confidence level of this match. This confidence level can be presented intuitively in the form of percentage, level, or color to indicate the reliability of the matching result. This information will be presented to law enforcement officers in an augmented reality (AR) manner, that is, by overlaying virtual information onto a real-world view.

[0086] The augmented reality display is achieved through an optical see-through display unit on a wearable vision device. This unit allows law enforcement officers to directly observe the real environment while simultaneously projecting virtual auxiliary identification information into their field of vision. To ensure the intuitiveness and usability of the information, the auxiliary identification information is precisely anchored to the spatial position of the target in the real world. This means the system continuously tracks the three-dimensional position of the target and the officer's head posture. Through real-time spatial positioning and attitude estimation techniques, such as visual inertial odometry (VIO) or simultaneous localization and mapping (SLAM) algorithms, combined with continuous tracking of the target object, the system can calculate the correct projection position of the auxiliary identification information in the officer's field of vision. Therefore, regardless of how the officer turns their head or changes their perspective, the auxiliary identification information remains "attached" to the target, maintaining its relative position. This spatial anchoring technology ensures a high degree of correlation between the information and the target, greatly enhancing the intuitiveness and readability of the information.

[0087] Through the aforementioned technical solution, when the comprehensive matching similarity reaches a preset recognition threshold, the system can directly overlay and display the identity information of successfully matched candidate targets and their matching confidence levels in an augmented reality manner within the law enforcement officer's field of vision. Furthermore, through the optical see-through display unit of a wearable vision device, this auxiliary identification information is precisely anchored to the spatial location of the target to be identified, ensuring that its relative position to the target remains unchanged as the officer's head moves. This allows law enforcement officers to intuitively and in real-time obtain key identification information without taking their eyes off the target, greatly improving the efficiency and accuracy of information acquisition. This immersive display method effectively avoids the problem of information being disconnected from the target in traditional display modes, significantly reducing the cognitive load on law enforcement officers and enhancing their situational awareness in dynamic law enforcement scenarios. This enables them to make judgments and decisions faster and more accurately, improving the efficiency and safety of overall law enforcement operations.

[0088] Example 2 (Device) In dynamic law enforcement scenarios, law enforcement officers face the challenge of recognizing blurry targets from a first-person perspective. Existing technologies, due to their reliance on fixed-viewpoint devices, lack of uncertainty quantification mechanisms, static feature weights, and distracting interaction methods, struggle to meet the demands for seamless, real-time recognition. This embodiment integrates a head-mounted casing with an image acquisition unit, an eye-tracking unit, a depth sensing unit, an optical vision display unit, a processor, and a memory to form a wearable architecture specifically designed for dynamic law enforcement. This enables reliable recognition and seamless interaction of blurry targets even when law enforcement officers' hands are occupied, thereby improving recognition accuracy and on-site decision-making efficiency.

[0089] A wearable vision device according to this embodiment includes: Headset housing; An image acquisition unit, mounted on the head-mounted housing, is used to acquire continuous video frame images from a first-person perspective. An eye-tracking unit is used to obtain the coordinates of the wearer's gaze focus. A depth sensing unit is used to measure the distance between the wearable vision device and the target to be identified; An optical see-through display unit is used to overlay auxiliary identification information onto the wearer's field of vision in an augmented reality manner; The processor is connected to the image acquisition unit, the eye tracking unit, the depth sensing unit, and the optical perspective display unit, respectively, and the processor is configured to perform the steps of the above method; The memory, connected to the processor, is used to store a preset database and computer programs.

[0090] The head-mounted casing is made of lightweight materials, ensuring wearing comfort and long-term stability. The image acquisition unit is configured with a wide-angle camera to capture continuous video frames within the law enforcement officer's field of vision in real time. Its optical parameters are automatically adjusted according to ambient light intensity to ensure image quality under low-light conditions. The eye-tracking unit monitors the wearer's eye movements through an infrared sensor array, accurately acquiring the coordinates of the gaze focus and using these coordinates as the input source for target locking commands, eliminating the need for manual operation by law enforcement officers. The depth sensing unit integrates a time-of-flight sensor to continuously measure the straight-line distance between the wearable vision device and the target to be identified, providing key parameters for subsequent dynamic allocation of feature weights. The optical vision display unit uses waveguide-based augmented reality technology to anchor auxiliary identification information to the spatial location of the target to be identified, ensuring that the information remains in a constant relative position as the law enforcement officer's head moves, guaranteeing seamless interaction. The processor is optimized as a low-power embedded chip, efficiently executing the steps of the above methods, including environmental parameter set acquisition, confidence interval model construction, and fuzzy matching calculation. The memory uses a high-speed cache architecture to store a pre-set database and computer programs, supporting local real-time comparison of task-level candidate target lists.

[0091] Through the aforementioned technical solution, this wearable vision device effectively addresses the reliability issues in dynamic law enforcement caused by image quality fluctuations, changing viewing angles, and hand-occupancy. Specifically, the collaborative operation of the eye-tracking unit and the image acquisition unit enables automatic target locking; the dynamic shooting distance data provided by the depth sensing unit drives the adaptive adjustment of the feature weight set; and the optical perspective display unit ensures that auxiliary recognition information is bound to the physical spatial location in real time, preventing law enforcement officers from shifting their gaze. Overall, this device significantly improves the accuracy and real-time performance of blurry target recognition in complex environments, while meeting the stringent requirements of frontline law enforcement for seamless, real-time human-computer interaction.

[0092] The above technical solution will be explained in detail below with a more specific example: In dynamic law enforcement missions, officers wear head-mounted visual acquisition devices to patrol densely populated public areas. One officer is instructed to identify a suspicious person whose physical characteristics match a specific description, but whose facial features may be difficult to discern due to distance, obstruction, or insufficient lighting.

[0093] First, wearable vision devices worn by law enforcement officers continuously capture a series of video frames from their first-person perspective. This differs from the third-person perspective images captured by traditional fixed monitoring equipment, ensuring that law enforcement officers can obtain real-time visual information about their areas of interest at dynamic law enforcement scenes.

[0094] When law enforcement officers focus their gaze on a target matching the initial description (e.g., by acquiring the coordinates of the gaze focus through a built-in eye-tracking module and constructing a dynamic region of interest at the center of those coordinates, prioritizing the detection of human bodies within that region), the wearable vision device immediately performs human detection and segmentation on the target. Subsequently, the system extracts multi-dimensional postural features of the target from consecutive video frames, including height parameters, limb proportion parameters, and gait parameters. For example, gait parameter extraction involves performing time-series analysis on consecutive video frames to obtain the gait cycle of the target, and further extracting joint angle change curves (such as knee and hip joint angle changes), stride length change sequences, and cadence parameters within that cycle.

[0095] While extracting these body posture features, the system constructs a confidence interval model for each dimension of body posture feature based on the imaging quality parameters (such as image resolution and signal-to-noise ratio), shooting angle parameters (such as the relative azimuth angle between the wearable vision device and the target to be identified), and occlusion degree parameters (such as the integrity ratio of the target to be identified in the image) of the current frame image. For example, if the target to be identified is far away, resulting in low image resolution, or if part of the target is occluded, the confidence interval width of its height parameter or limb proportion parameter will be expanded accordingly to characterize the uncertainty of the feature under the current observation conditions.

[0096] Simultaneously, the system acquires a set of environmental parameters from continuous video frames in real time, including image resolution, illumination intensity, and the dynamic shooting distance between the wearable vision device and the target to be identified. The dynamic shooting distance can be measured directly by the built-in depth sensor or calculated using a monocular vision ranging algorithm based on the percentage of human pixels representing the target in the image, combined with preset average human height parameters. For example, when the percentage of pixels representing the target in the image is lower than a first threshold, the system determines that the dynamic shooting distance is greater than a preset distance threshold.

[0097] Based on these environmental parameter sets, the system dynamically allocates feature weights for multi-dimensional body posture features during the matching process. For example, when the dynamic shooting distance is greater than a preset distance threshold, the system increases the weights of height and limb proportion parameters while decreasing the weight of gait parameters, because the reliability of gait feature extraction at long distances is generally low. Conversely, when the dynamic shooting distance is less than or equal to the preset distance threshold and consecutive video frames contain side-view images, the system increases the weight of gait parameters, because gait features are more significant and reliable at close-range side views. Furthermore, if the illumination intensity is lower than a preset illuminance threshold, the system increases the weights of limb proportion and gait parameters while decreasing the weight of height parameters; if the image resolution is lower than a preset resolution threshold, the system increases the weight of limb proportion parameters while decreasing the weight of gait parameters, and expands the confidence interval width of the confidence interval model to a first preset multiple.

[0098] Next, based on the confidence interval models corresponding to the multi-dimensional body features of the target to be identified and the dynamically allocated feature weight set, the system performs fuzzy matching between the target and candidate targets in a preset database. This preset database stores a library of candidate target body features associated with the current law enforcement task and is synchronized in real-time from the back-end command center to the local storage of the wearable vision device. The fuzzy matching process includes: for each dimension of body feature, calculating the overlap between the feature value of the target to be identified (represented as a confidence interval centered on the detected value and with a radius equal to the width output by the confidence interval model) and the corresponding feature value of the candidate target (represented as a definite value in the preset database) to obtain a single feature matching degree. For example, when the definite value of the candidate target falls within the confidence interval of the target to be identified, the single feature matching degree is the first preset value; when the definite value is outside the confidence interval, the single feature matching degree decreases linearly as the distance between the definite value and the boundary of the confidence interval increases. Finally, the system weights and fuses each single feature matching degree according to the dynamically allocated feature weight set to calculate the comprehensive matching similarity.

[0099] Finally, based on the calculated comprehensive matching similarity, the system generates auxiliary recognition information on the optical see-through display unit of the wearable vision device. When the comprehensive matching similarity exceeds a preset recognition threshold, the system overlays the identity information of the successfully matched candidate target and the matching confidence level onto the law enforcement officer's field of vision in an augmented reality manner. This auxiliary recognition information is anchored to the actual spatial location of the target to be identified, ensuring that its relative position to the target remains unchanged as the officer's head moves. This "unobtrusive and real-time" augmented reality display method avoids the distraction of law enforcement officers looking down at the device, significantly improving law enforcement efficiency and safety, and solving the problem that traditional recognition result display methods do not meet the needs of front-line law enforcement.

[0100] The above description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. For those skilled in the art, the present invention can have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for fuzzy target-assisted recognition based on body posture features, characterized in that, The method is applied to head-mounted visual acquisition devices worn by law enforcement officers at dynamic law enforcement scenes, and includes the following steps: Step S1: Using a wearable vision device worn on the head of law enforcement officers, continuously capture video frame images from the first-person perspective of the law enforcement officers in real time; Step S2: In response to the law enforcement officer's gaze focus or target lock command, perform human detection and segmentation on the target to be identified contained in the continuous video frame images, and extract multi-dimensional body features of the target to be identified; the multi-dimensional body features include at least height parameters, limb proportion parameters, and gait parameters; wherein, for each dimension of body feature, based on the imaging quality parameters, shooting angle parameters, and occlusion degree parameters of the current frame image, construct a confidence interval model corresponding to the dimension of body feature to characterize the reliability of the feature under the current observation conditions; Step S3: Obtain the environmental parameter set of the continuous video frame images; the environmental parameter set includes at least image resolution, light intensity, and dynamic shooting distance between the wearable vision device and the target to be identified; Step S4: Based on the environmental parameter set, dynamically allocate the feature weight set of the multi-dimensional body features in the matching process; wherein, when the dynamic shooting distance is greater than a preset distance threshold, increase the weight of the height parameter and limb proportion parameter, and decrease the weight of the gait parameter; when the dynamic shooting distance is less than or equal to the preset distance threshold and the continuous video frame images contain side view images, increase the weight of the gait parameter. Step S5: Based on the confidence interval models corresponding to each of the multi-dimensional body features and the feature weight set, perform fuzzy matching between the target to be identified and the candidate targets in the preset database, and calculate the comprehensive matching similarity; the fuzzy matching includes: for each dimension of body feature, calculating the interval overlap between the feature value and its confidence interval of the target to be identified and the corresponding feature value of the candidate target to obtain the single feature matching degree; and weighting and fusing each single feature matching degree according to the feature weight set to obtain the comprehensive matching similarity. Step S6: Based on the comprehensive matching similarity, generate auxiliary recognition information on the display unit of the wearable vision device.

2. The method according to claim 1, characterized in that, The construction of the confidence interval model in step S2 specifically includes: The error distribution range of feature extraction is determined based on the imaging quality parameters, which include at least image resolution and signal-to-noise ratio. The human body projection deformation coefficient is determined based on the shooting angle parameters, which are determined according to the relative azimuth angle between the wearable vision device and the target to be identified. The visible feature proportion factor is determined based on the occlusion degree parameter, and the occlusion degree parameter is determined based on the integrity proportion of the target to be identified in the image. The error distribution range, the deformation coefficient, and the proportion factor are used as parameters to input a preset uncertainty quantification model, and the confidence interval width of each dimension of body features is output. The confidence interval width is negatively correlated with the imaging quality parameter, positively correlated with the absolute value of the angle of deviation of the shooting angle parameter, and negatively correlated with the occlusion degree parameter.

3. The method according to claim 1, characterized in that, Obtaining the dynamic shooting distance in step S3 specifically includes: The wearable vision device measures the straight-line distance between itself and the target to be identified in real time using a depth sensor built into it; or, Based on the proportion of human pixels of the target to be identified in the continuous video frame images, and combined with the preset average human height parameters, the dynamic shooting distance is calculated by a monocular visual ranging algorithm. Specifically, when the pixel percentage of the target to be identified in the image is lower than a first threshold, the dynamic shooting distance is determined to be greater than a preset distance threshold; when the pixel percentage is higher than a second threshold, the dynamic shooting distance is determined to be less than or equal to the preset distance threshold.

4. The method according to claim 1, characterized in that, The step S4 of dynamically allocating the feature weight set further includes: The illumination intensity of the continuous video frame images is obtained. When the illumination intensity is lower than a preset illumination threshold, the weight of the limb proportion parameter and the gait parameter is increased, and the weight of the height parameter is decreased. The image resolution of the continuous video frame image is obtained. When the image resolution is lower than a preset resolution threshold, the weight of the limb proportion parameter is increased, the weight of the gait parameter is decreased, and the confidence interval width of the confidence interval model is expanded to a first preset multiple. The viewpoint type contained in the continuous video frame images is obtained. When the viewpoint type is a frontal view, the weights of the height parameter and the limb proportion parameter are increased; when the viewpoint type is a side view, the weights of the gait parameter are increased.

5. The method according to claim 1, characterized in that, Step S2, which involves extracting the multi-dimensional body features of the target to be identified, further includes: Perform time series analysis on the continuous video frame images to extract the gait period of the target to be identified; Within the gait cycle, the joint angle change curve, stride change sequence, and gait frequency parameters of the target to be identified are extracted and together constitute the gait parameters; The joint angle change curves include at least the knee joint angle change curve and the hip joint angle change curve.

6. The method according to claim 1, characterized in that, In step S2, in response to the law enforcement officer's gaze focus or target locking command, human detection and segmentation are performed on the target to be identified contained in the continuous video frame images, specifically including: The eye-tracking module built into the wearable vision device obtains the coordinates of the law enforcement officer's gaze focus in the image coordinate system. A dynamic region of interest is constructed centered on the coordinates of the line of sight focus. A human detection algorithm is executed within the dynamic region of interest, prioritizing the detection and segmentation of targets to be identified located in the line of sight of the law enforcement officer.

7. The method according to claim 1, characterized in that, Step S5 involves calculating the overlap between the feature values ​​and confidence intervals of the target to be identified and the corresponding feature values ​​of the candidate targets. Specifically, this includes: The feature value of the target to be identified is represented as a confidence interval centered on the detected value and with the width output by the confidence interval model as the radius; The corresponding feature values ​​of the candidate targets are represented as predetermined values ​​in a pre-defined database; The degree of overlap between the confidence interval and the determined value is calculated. When the determined value falls within the confidence interval, the single feature matching degree is a first preset value. When the determined value is outside the confidence interval, the single feature matching degree decreases linearly as the distance between the determined value and the boundary of the confidence interval increases.

8. The method according to claim 1, characterized in that, In step S5, the preset database stores a candidate target body feature library associated with the current law enforcement task, and the candidate target body feature library is synchronized in real time from the background command center to the local storage of the wearable vision device; the method further includes: When the wearable vision device establishes a communication connection with the back-end command center, it receives a task-level candidate target list issued by the back-end command center. The task-level candidate target list contains the body feature data of at least one candidate target associated with the current law enforcement area. The task-level candidate target list is stored in the local cache of the wearable vision device for real-time comparison of the fuzzy matching.

9. The method according to claim 1, characterized in that, Step S6, which generates auxiliary recognition information on the display unit of the wearable vision device, specifically includes: When the overall matching similarity exceeds a preset recognition threshold, the identity information of the successfully matched candidate target and the matching confidence level are superimposed and displayed in the field of vision of the law enforcement personnel in an augmented reality manner; The augmented reality display includes: anchoring the auxiliary identification information to the spatial location of the target to be identified through the optical see-through display unit of the wearable vision device, so that the auxiliary identification information maintains a constant relative positional relationship with the target to be identified as the law enforcement officer's head moves.

10. A wearable vision device, characterized in that, include: Headset housing; An image acquisition unit is mounted on the head-mounted housing and is used to acquire continuous video frame images from a first-person perspective. An eye-tracking unit is used to obtain the coordinates of the wearer's gaze focus. A depth sensing unit is used to measure the distance between the wearable vision device and the target to be identified; An optical see-through display unit is used to overlay auxiliary identification information onto the wearer's field of vision in an augmented reality manner; A processor is connected to the image acquisition unit, the eye tracking unit, the depth sensing unit, and the optical perspective display unit, respectively, and the processor is configured to perform the steps of the method according to any one of claims 1 to 9; The memory, connected to the processor, is used to store a preset database and computer programs.