A method and apparatus for image feature recognition testing
By acquiring and preprocessing images in real time, combined with target tracking and locking and dual-database feature evaluation, the problems of loss of field of view and delayed judgment response of dynamic targets in the test scenario of the image recognition system were solved, and efficient monitoring and rapid early warning of dynamic targets were achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG DONGFANG XINDA INFORMATION TECH CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-26
AI Technical Summary
Existing image recognition systems, in experimental scenarios, rely on fixed-angle image acquisition devices and a single local feature library, resulting in the loss of dynamic target field of view and delayed response in safety status determination, making it difficult to meet the needs of efficient monitoring and rapid early warning of dynamic targets.
By employing real-time image acquisition and preprocessing, combined with target tracking and locking, and hierarchical cross-matching evaluation using dual standard feature libraries, stable focusing and accurate feature extraction of dynamic targets are achieved, generating feedback signals.
The dynamic tracking and locking mechanism ensures that dynamic targets remain continuously within the monitoring field of view, improving the accuracy and response speed of feature recognition, solving the problems of lost field of view and delayed judgment response, and meeting the needs of accurate monitoring and rapid early warning in experimental scenarios.
Smart Images

Figure CN122090143A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image feature recognition technology, and in particular to an image feature recognition experimental method and apparatus. Background Technology
[0002] With the rapid iteration of artificial intelligence technology, AI image recognition technology has become a core supporting technology in fields such as security and prevention, industrial inspection, and test scenario monitoring. Through intelligent analysis and processing of visual information, it can achieve automatic identification and status determination of target objects, playing a key role in improving scenario security and optimizing detection efficiency. In particular, in various test scenarios, the demand for accurate identification of dynamic targets and safety status determination continues to rise, driving the continuous exploration and upgrading of related technologies.
[0003] Currently, most mainstream image recognition systems rely on fixed-angle image acquisition devices to acquire visual information and combine them with matching algorithms based on a single local feature library to achieve target recognition. While this technology can meet basic monitoring needs in static scenarios, its limitations are becoming increasingly apparent in experimental scenarios where the target object moves dynamically and environmental interference factors are complex.
[0004] Specifically, existing technologies rely on fixed acquisition equipment and a single feature library. When faced with dynamically moving targets in experimental scenarios, they are prone to problems such as loss of target field of view and inaccurate feature capture, resulting in a lag in the determination of the target's safety status and making it difficult to meet the needs of efficient monitoring and rapid early warning of dynamic targets in experimental scenarios. Summary of the Invention
[0005] In view of this, this application provides an image feature recognition test method and apparatus, which solves the technical problems of loss of dynamic target field of view and delayed response of safety status determination caused by existing image recognition systems relying on fixed-angle image acquisition equipment and a single local feature library in test scenarios.
[0006] This application provides an image feature recognition experimental method and apparatus, which adopts the following technical solution: An image feature recognition experimental method, comprising: Real-time images of the test scene are acquired and preprocessed to obtain continuous preprocessed images; Target tracking and locking are performed on the preprocessed images of the consecutive frames to obtain the locked target detection images; Feature extraction is performed on the locked target detection image to obtain target feature data; Based on a pre-defined dual-database standard feature library, the target feature data is used to perform hierarchical cross-matching evaluation to obtain the security status determination result of the target object. Based on the security status determination result, a corresponding feedback signal is generated.
[0007] By adopting the above technical solution, real-time acquisition and preprocessing are used to ensure image quality, target tracking and locking are combined to achieve stable focusing on the target, and then feature extraction is used to obtain the core representation of the target. Relying on dual-database hierarchical cross-matching, matching efficiency and comprehensiveness are taken into account. Finally, feedback signals are generated based on the judgment results, realizing closed-loop control of the entire process from image acquisition to result response, and improving the accuracy and intelligence level of image feature recognition experiments.
[0008] Preferably, the real-time acquisition and preprocessing of real-time images of the test scene to obtain consecutive preprocessed images includes: Real-time acquisition of images of the test scene; Gaussian filtering is applied to the real-time images of the test scene to obtain denoised images of consecutive frames; Image enhancement is performed on the denoised images of the consecutive frames to obtain pre-processed images of the consecutive frames.
[0009] By adopting the above technical solution, the original scene image is first acquired in real time to capture complete target information. Then, Gaussian filtering is used to effectively suppress Gaussian noise and preserve the target outline. Combined with image enhancement, the difference between the target and the background is strengthened and the feature recognition is improved, providing high-quality image data support for subsequent target detection, tracking and locking, and reducing the interference of noise and low contrast on the experimental results.
[0010] Preferably, the step of performing target tracking and locking on the consecutive preprocessed images to obtain locked target detection images includes: Target detection is performed on the preprocessed images of the consecutive frames to obtain the target detection parameters of the target object; Using the target detection parameters, the target motion parameter set is determined; Trajectory prediction is performed using the target motion parameter set to obtain the trajectory prediction result; The trajectory prediction results drive the image acquisition device to perform target locking operations, resulting in a locked target detection image.
[0011] By adopting the above technical solution, the target is first accurately located and its core parameters are obtained through target detection. Then, the motion state set is calculated based on the parameters to capture the target's motion pattern. Combined with trajectory prediction, the target's future position is predicted. Finally, the device is driven to perform a locking operation, realizing continuous control of the target from static positioning to dynamic tracking, ensuring that the locked image focuses on the target and meets the feature extraction requirements.
[0012] Preferably, the target detection parameters include the coordinates of the bounding rectangle of the target object, and the step of using the target detection parameters to determine the target motion parameter set includes: Determine the sequence of center coordinates of the bounding rectangle of the target object; Extract the acquisition timestamps of the consecutive preprocessed images; The target's movement speed is determined using the frame center coordinate sequence and the corresponding acquisition timestamp; A target coordinate sequence is constructed using the frame center coordinate sequence and the corresponding acquisition timestamp; Based on the target's motion speed, determine the target's resultant velocity; By integrating the target coordinate sequence, the target motion velocity, and the target resultant velocity, a target motion parameter set is obtained.
[0013] By adopting the above technical solution, the target's temporal position is represented by the sequence of center coordinates derived from the coordinates of the outer rectangular box. By combining the collected timestamps with the associated position and time dimension, the motion velocity and resultant velocity are calculated to quantify the target's motion state. Finally, a complete set of motion parameters is integrated to provide comprehensive and accurate motion feature input for subsequent trajectory prediction, ensuring the reliability of trajectory prediction.
[0014] Preferably, the step of using the target motion parameter set to perform trajectory prediction and obtain trajectory prediction results includes: Sort the target coordinate sequence and the corresponding target motion velocity; The sorted target coordinate sequence and target motion velocity are standardized to obtain the target input vector; The target input vector is used to input a preset trajectory prediction model to perform trajectory prediction, and the trajectory prediction result is obtained.
[0015] By adopting the above technical solution, the coordinate sequence and motion velocity are sorted to ensure the consistency of data time sequence. After standardization to eliminate the influence of differences in dimensions and numerical ranges and to unify the data scale, the target input vector is fed into the pre-trained model to capture the temporal dependency relationship, thereby realizing the accurate prediction of the target's future trajectory and providing a scientific basis for subsequent equipment locking operations.
[0016] Preferably, the trajectory prediction result includes the target trajectory prediction coordinates, and the step of driving the image acquisition device to perform a target locking operation based on the trajectory prediction result to obtain a locked target detection image includes: Calculate the first offset between the predicted coordinates of the target trajectory and the image center coordinates of the consecutive preprocessed images; Based on the first offset, the rotation speed of the gimbal is dynamically adjusted according to the target combined velocity; The rotation direction of the gimbal is determined based on the target's movement speed, and the rotation is executed until the predicted coordinates of the target trajectory coincide with the coordinates of the image center. Real-time images of the target are acquired by an image acquisition device mounted on the rotated gimbal. Calculate the second offset between the predicted coordinates of the target trajectory and the actual coordinates of the real-time image of the target; Based on the second offset, the PID control algorithm is used to fine-tune and lock the gimbal, and the image acquisition device after fine-tuning and locking acquires the target detection image.
[0017] By adopting the above technical solution, the first offset is calculated to achieve coarse adjustment of the gimbal and quickly reduce the deviation between the target and the center of the image. Combined with the resultant velocity, the rotation speed is dynamically adjusted to adapt to the target movement. Then, the second offset and PID algorithm are used for fine adjustment to eliminate the deviation between the predicted and actual positions. This achieves hierarchical locking from coarse adjustment to fine adjustment, ensuring that the target is stably in the center of the image and improving the clarity and stability of the target detection image.
[0018] Preferably, the preset dual-database standard feature library includes a local feature library and a cloud-based feature library. The step of performing hierarchical cross-matching evaluation using the target feature data based on the preset dual-database standard feature library to obtain the security status determination result of the target object includes: Calculate the first matching degree between the target feature data and the associated local feature data in the local feature library; If the first matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target; If the first matching degree is less than the second preset matching degree threshold, the target object is determined to be an unsafe target; When the first matching degree is greater than or equal to the second preset matching degree threshold and less than the first preset matching degree threshold, the second matching degree between the target feature data and the cloud feature data associated in the cloud feature library is calculated. The first matching degree and the second matching degree are combined with a preset weighting factor to perform a weighted calculation to obtain the comprehensive matching degree; When the overall matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target; If the overall matching degree is less than the first preset matching degree threshold, the target object is determined to be an unsafe target; The first preset matching threshold is greater than the second preset matching threshold.
[0019] By adopting the above technical solution, the initial matching is first completed quickly through the local feature library to ensure the efficiency of the test. Then, for fuzzy matching scenarios, the cloud feature library is linked to supplement the matching and improve the coverage. The two-dimensional matching results are fused by weighted factors, and the matching interval is divided by dual thresholds. This achieves the determination of the target safety status while taking into account both efficiency and accuracy, and reduces the probability of misjudgment and missed judgment.
[0020] Preferably, generating the corresponding feedback signal based on the security status determination result includes: If the safety status determination result is an unsafe target, an alarm feedback signal is generated and the corresponding alarm action is executed. When the security status determination result is a security target, a security feedback signal is generated and the corresponding security action is executed.
[0021] By adopting the above technical solutions, audible and visual alarms, cloud push notifications, and linked protection are triggered for unsafe targets to achieve rapid risk handling; prompt signals are generated for safe targets, records are archived, and linked clearance is implemented to ensure efficient passage. This achieves differentiated automated response based on judgment results, taking into account both safety protection and control efficiency in the test scenario.
[0022] Preferably, the method further includes performing environmental adaptation correction on the first preset matching degree threshold based on a preset correction period, wherein the environmental adaptation correction specifically includes: Perform grayscale conversion on the preprocessed image of the current frame to obtain the target grayscale image; Extract the light intensity and background complexity of the target grayscale image; The environment adaptation coefficient is calculated using the light intensity and the background complexity. The environmental adaptation coefficient and the preset dynamic correction coefficient are multiplied to obtain a new first preset matching degree threshold. Based on the new first preset matching degree threshold, the process jumps to the step of determining the target object as a safe target when the first matching degree is greater than or equal to the first preset matching degree threshold.
[0023] By adopting the above technical solution, the light intensity and background complexity are extracted after the image grayscale conversion, the environmental adaptation coefficient is calculated to quantify the degree of environmental interference, and the threshold is adjusted by combining the dynamic correction coefficient to achieve environmental adaptive optimization of the first preset matching degree threshold. This avoids the judgment deviation caused by changes in illumination and background complexity, and improves the robustness of the experimental method in complex environments.
[0024] An image feature recognition experimental device, comprising: The acquisition module is used to acquire real-time images of the test scene and perform preprocessing to obtain continuous preprocessed images. The tracking module is used to perform target tracking and locking on the continuous frame preprocessed images to obtain the locked target detection images; The extraction module is used to extract features from the locked target detection image to obtain target feature data; The evaluation module is used to perform hierarchical cross-matching evaluation based on the target feature data using a preset dual-database standard feature library to obtain the security status judgment result of the target object; The execution module is used to generate a corresponding feedback signal based on the security status determination result.
[0025] By adopting the above technical solution, the acquisition module ensures the quality of image preprocessing, the tracking module achieves stable target locking, the extraction module obtains core feature data, the evaluation module completes accurate matching judgment, and the execution module generates corresponding feedback signals. The cooperation of each module realizes the full-process automation of image feature recognition experiment, ensuring the method operates efficiently and stably, and improving the standardization and repeatability of the experiment.
[0026] In summary, this application includes at least one of the following beneficial technical effects: This application provides an image feature recognition experimental method and apparatus. It acquires images of the experimental scene in real time and preprocesses them to obtain clear, continuous frame images. Based on these, a tracking and locking operation is performed on the dynamic target to ensure that the target remains within the effective monitoring field of view. Then, precise target feature data is extracted from the locked target detection image, and hierarchical cross-matching evaluation is performed using a dual-library standard feature library (including local and cloud-based libraries). Finally, a corresponding feedback signal is generated based on the evaluation results. Addressing the problem of dynamic target field of view loss caused by fixed-angle acquisition devices in existing technologies, the dynamic tracking and locking mechanism can adjust the image acquisition angle in real time to ensure that the dynamic target remains continuously within the monitoring field of view, avoiding field of view shift and target loss under fixed-angle acquisition, and effectively solving the problem of dynamic... To address the shortcomings of target field of view loss and the lag in security status determination caused by existing technologies relying on a single local feature library, this application adopts a dual-library hierarchical cross-matching evaluation mechanism. It prioritizes rapid matching using the local feature library, only invoking the cloud-based feature library for secondary verification when the matching degree is in the intermediate range. This ensures matching accuracy while reducing unnecessary cloud interaction overhead through hierarchical logic, significantly improving matching efficiency and effectively shortening the response time for security status determination, thus solving the problem of delayed determination response. Furthermore, the preprocessing stage's denoising and enhancement of the image further improves the accuracy of feature extraction. This application's efficient closed loop from image acquisition to status determination better meets the needs of accurate monitoring and rapid early warning of dynamic targets in experimental scenarios. Attached Figure Description
[0027] Figure 1This is a flowchart of the steps of an image feature recognition experimental method provided in an embodiment of this application.
[0028] Figure 2 This is a structural block diagram of an image feature recognition experimental device provided in an embodiment of this application. Detailed Implementation
[0029] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application. It should be noted that in the optional embodiments of this application, the object information and other related data involved require the permission or consent of the object when the embodiments of this application are applied to specific products or technologies, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. That is to say, if the embodiments of this application involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant departments, and in compliance with the relevant laws, regulations, and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information requires the consent of the individual. If sensitive information is involved, the separate consent of the information subject is required, and the embodiments also need to be implemented with the authorization and consent of the object.
[0030] Please see Figure 1 This application provides an image feature recognition experimental method, comprising: Image feature recognition experimental method is a series of technical processes used in experimental scenarios to identify and determine the safety status of target objects through image acquisition, preprocessing, target tracking, feature extraction, dual-library matching evaluation, and feedback. Real-time images of the experimental scenario are raw image data captured in real time by image acquisition devices (such as gimbal cameras), reflecting the environment, target object dynamics, and location information within the experimental scenario. Preprocessing involves denoising and enhancing the acquired raw real-time images to eliminate interference and improve image quality, providing suitable data for subsequent target detection and feature extraction. Continuous frame preprocessed images are a collection of multiple frames generated sequentially after preprocessing, reflecting the dynamic trajectory of the target object in the experimental scenario. Target tracking and locking, based on continuous frame preprocessed images, involves locating the target position, predicting its motion trajectory, and driving the image acquisition device to follow the target until it is stably focused and the shooting angle is fixed. The locked target detection image is captured by the image acquisition device after target tracking and locking is completed, with the target clearly centered in the frame. The focused image is used for subsequent feature extraction. Feature extraction is the extraction of data information representing the unique attributes (such as contour and texture) of the target from the locked target detection image using image processing algorithms (such as gradient extraction and feature vector construction). Target feature data is the structured data obtained after feature extraction (usually constructed as feature vectors), which is the core basis for distinguishing different target objects and conducting matching evaluation. The preset dual-library standard feature library is a pre-built dual feature database containing a local feature library and a cloud feature library, storing standard feature data of various security targets and providing a reference for matching evaluation. The hierarchical cross-matching evaluation is a hierarchical matching logic that first performs preliminary matching through the local feature library, then decides whether to link the cloud feature library for secondary matching based on the matching results, and finally merges the matching results to complete the evaluation. The security status judgment result is the final judgment conclusion based on the hierarchical cross-matching evaluation result, which is whether the target object is a "safe target" or an "unsafe target". The feedback signal is a structured signal generated based on the security status judgment result, divided into alarm feedback signal and security feedback signal, used to trigger the corresponding action execution.
[0031] Step 101: Acquire real-time images of the test scene and perform preprocessing to obtain continuous frame preprocessed images.
[0032] In this embodiment, considering that the target objects in the test scenario are mostly in a dynamic moving state and are easily affected by environmental noise, light changes and other factors affecting image quality, real-time acquisition of test scenario images can ensure the acquisition of complete temporal data of target motion, meeting the requirements of dynamic target tracking for continuous frame images. At the same time, preprocessing the acquired real-time images can effectively filter out redundant noise, optimize image clarity and contrast, and avoid noise interference and image blurring problems affecting the accuracy of subsequent target feature extraction. This provides high-quality data support for the next step of performing target tracking and locking on continuous frame preprocessed images, ensuring that the target position can be quickly captured and the motion trajectory locked during target tracking. This avoids tracking deviations caused by poor image quality or discontinuous data from the source, thereby laying a solid foundation for the efficient implementation of feature extraction, matching evaluation and safety status determination, and helping to solve the problems of dynamic target field of view loss and delayed judgment response in the prior art.
[0033] Further, step 101 may include the following sub-steps: S11. Real-time acquisition of real-time images of the test scene; S12. Perform Gaussian filtering on the real-time image of the test scene to obtain a continuous frame denoised image; S13. Perform image enhancement on the denoised images of consecutive frames to obtain pre-processed images of consecutive frames.
[0034] Gaussian filtering is a linear smoothing filtering algorithm based on the Gaussian function. It eliminates Gaussian noise in an image by weighted averaging of image pixels and their neighboring pixels, while preserving the target contour information. Continuous frame denoising is a collection of images generated sequentially in time after Gaussian filtering. Compared to the original image, noise interference is suppressed, and the image is clearer. Image enhancement is a process that improves the visual effect and feature recognition of an image by adjusting parameters such as image contrast, brightness, and sharpness, or by using algorithms to enhance the difference between target features and the background.
[0035] In this embodiment, S11 is first executed using an image acquisition device with a dynamically adjustable viewing angle deployed in the test scene to capture dynamic images of the test scene in real time at a preset frame rate, obtaining raw image data containing the complete motion sequence of the target object; then S12 is executed, using a 5×5 Gaussian kernel to perform convolution operation on each frame of the raw image, and filtering out Gaussian noise and random interference points in the image by weighted averaging, reducing the interference of noise on subsequent target detection, and obtaining continuous frame denoised images; then S13 is executed, using a histogram equalization algorithm to adjust the brightness and contrast distribution of the denoised image, enhancing the visual distinction between the target area and the background, improving the recognizability of image details, and finally obtaining continuous frame preprocessed images, which can provide a clear and stable visual basis for target tracking and locking in step 102, avoiding target capture deviation caused by poor image quality, and ensuring the accuracy of subsequent feature extraction and matching evaluation from the source.
[0036] Step 102: Perform target tracking and locking on the preprocessed images of consecutive frames to obtain the locked target detection images.
[0037] In this embodiment, based on the clear continuous frame preprocessed image output in step 101, a tracking and locking operation is performed on the movement characteristics of dynamic targets in the test scenario. The position changes and movement trends of the target are captured in real time, and the image acquisition device is synchronously linked to adjust the viewing angle and posture to ensure that the target is always within the acquisition field of view and maintains clear imaging. This effectively solves the problem that existing fixed-angle acquisition devices are difficult to follow dynamic targets and are prone to losing field of view. At the same time, by accurately locking the target area, irrelevant background interference in the image can be filtered out to obtain a detection image focused on the target itself. This provides an accurate target carrier for feature extraction, avoids background redundancy information from affecting the accuracy of feature extraction, and simplifies the amount of subsequent data processing, indirectly improving the overall process efficiency. This provides strong support for rapidly advancing feature extraction and matching evaluation and improving the problem of delayed response in safety status determination.
[0038] Furthermore, step 102 may include the following sub-steps: S21. Perform target detection on the preprocessed images of consecutive frames to obtain the target detection parameters of the target object; S22. Using target detection parameters, determine the target motion parameter set; S23. Use the target motion parameter set to predict the trajectory and obtain the trajectory prediction result; S24. Based on the trajectory prediction results, drive the image acquisition device to perform target locking operation and obtain the locked target detection image.
[0039] Target detection is the process of identifying target objects from consecutive preprocessed images using target detection algorithms (such as YOLO and SSD), locating target positions, and obtaining parameters such as target size. Target detection parameters are the set of parameters characterizing the attributes of the target object obtained after target detection. The core parameters include the coordinates of the target's bounding box, and may also include data such as target size and confidence level. The target motion parameter set is the set of parameters characterizing the target's motion state calculated based on the target detection parameters and image acquisition timestamps, including target coordinate sequences, motion speed, and resultant velocity. Trajectory prediction is the process of using the target motion parameter set to capture the target's motion patterns through a preset trajectory prediction model and predicting the target's position changes over a future period. The trajectory prediction result is the future position information of the target obtained after trajectory prediction, the core of which is the target trajectory prediction coordinates, and may also include data such as prediction timestamps and prediction frame numbers. Image acquisition equipment is the equipment used to acquire images of the test scene (such as cameras), usually mounted on a gimbal, which can rotate with the gimbal to adjust the shooting angle. Target locking operation is based on the trajectory prediction result, driving the gimbal and image acquisition equipment to adjust their posture so that the target is always in the center of the image, achieving stable tracking and fixing the shooting angle.
[0040] In this embodiment, based on the clear continuous frame preprocessed image obtained in step 101, a lightweight target detection algorithm is used to scan each frame of the image during step S21 to accurately identify the target object and output its bounding box coordinates, target category, and other core detection parameters. Then, in step S22, the moving speed, direction, and position sequence of the target object are calculated by combining the detection parameters of multiple consecutive frames with the corresponding image acquisition timestamps, and integrated to form a target motion parameter set, which fully represents the dynamic features of the target. Subsequently, in step S23, based on the motion parameter set, a time-series prediction model is used to predict the short-term trajectory of the target's position at the next moment, and the trajectory prediction result is obtained. Finally, in step S24, the pan-tilt unit of the image acquisition device is controlled in conjunction with the trajectory prediction result to dynamically adjust the device's viewing angle and attitude, and to track the target's motion trajectory in real time, ensuring that the target is always in the center area of the image, completing target locking and acquiring the target detection image after locking. This effectively avoids the problem of target field of view loss caused by fixed-angle acquisition and provides a high-quality image carrier focused on the target for accurate feature extraction in step 103.
[0041] Furthermore, the target detection parameters include the coordinates of the bounding rectangle of the target object, and S22 may include the following sub-steps: S221. Determine the sequence of center coordinates of the bounding rectangle of the target object. S222, Extract the acquisition timestamps of consecutive preprocessed frames; S223. Determine the target's movement speed using the frame center coordinate sequence and the corresponding acquisition timestamp; S224. Construct the target coordinate sequence using the frame center coordinate sequence and the corresponding acquisition timestamp; S225. Determine the resultant velocity of the target based on its motion velocity; S226. Integrate the target coordinate sequence, target motion velocity, and target resultant velocity to obtain the target motion parameter set.
[0042] The bounding box coordinates are the coordinates of the four vertices of the smallest rectangle enclosing the edge of the target object in object detection (usually with the top left corner of the image as the origin, representing the position and size of the rectangle), used to locate the target's position in the image; the box center coordinate sequence is a coordinate sequence formed by calculating the box center coordinates of each frame based on the target's bounding box coordinates, arranged in the order of image acquisition time, reflecting the temporal change of the target's position; the acquisition timestamp is the absolute time (such as a Unix timestamp, unit: seconds) recorded when the image acquisition device captures each preprocessed image frame, used to associate the target's position with time and calculate the motion velocity; the target's motion velocity is based on the box center coordinates. The target coordinate sequence is calculated using the target sequence and corresponding acquisition timestamps, which includes horizontal and vertical velocity components. The target coordinate sequence is a structured sequence formed by binding the center coordinates of the target bounding box in each frame with the corresponding acquisition timestamp and organizing them in chronological order, providing basic data for trajectory prediction. The target resultant velocity is the velocity obtained by synthesizing the instantaneous velocity components in the horizontal and vertical directions of the target, representing the overall speed and actual movement trend of the target. The target motion parameter set is a collection that integrates the target coordinate sequence, target motion velocity, target resultant velocity, and other parameters, comprehensively reflecting the target motion state and providing complete input data for trajectory prediction.
[0043] In this embodiment of the application, the coordinates of the bounding rectangle of the target object obtained in S21 (let the coordinates of the upper left corner of the bounding rectangle in a single frame image be...) The coordinates of the lower right corner are When executing S221, the center coordinates of the target box in each frame are calculated using a formula, specifically: ;
[0044] in, For the first The center coordinates (two-dimensional pixel coordinates, unit: pixels) of the bounding rectangle of the target object in the frame image. For the first The x-axis pixel coordinates (horizontal pixel position, unit: pixels) of the top-left corner of the bounding rectangle of the target object in the frame image. For the first The x-axis pixel coordinates (horizontal pixel position, unit: pixels) of the bottom right corner of the bounding rectangle of the target object in the frame image. For the first The y-axis pixel coordinates (vertical pixel position, unit: pixels) of the top-left corner of the bounding rectangle of the target object in the frame image. For the first The y-axis pixel coordinates (vertical pixel position, unit: pixels) of the lower right corner of the bounding rectangle of the target object in the frame image are arranged in frame time sequence to form the center coordinate sequence of the box; Then, step S222 is executed to extract the acquisition timestamp corresponding to each preprocessed image frame. Establish a one-to-one correspondence between timestamps and the center coordinates of the bounding box; Next, S223 calculates the target's velocity using the difference between the center coordinates and timestamps of two adjacent frames, using the following formula: ;
[0045] in, For the first Frame to the Instantaneous motion speed of the target object within a frame time period (unit: pixels / second). For the first The x-axis component of the center coordinates of the frame target bounding box (horizontal pixel position, unit: pixels). For the first The y-axis component of the center coordinates of the frame target bounding box (vertical pixel position, unit: pixels). For the first The x-axis component of the center coordinates of the frame target bounding box (horizontal pixel position, unit: pixels). For the first The y-axis component of the center coordinates of the frame target bounding box (vertical pixel position, unit: pixels). For the first The acquisition timestamp of the frame preprocessed image (absolute time, in seconds, such as Unix timestamp). For the first The acquisition timestamp (absolute time, in seconds) of the preprocessed frame image; Then S224 sets the frame center coordinates for each frame. With corresponding timestamp Binding, arranging the target coordinate sequence in chronological order ( (Number of consecutive frames); Then execute S225, based on the motion speed calculated from multiple frames. Decompose the velocity components in the x-direction and y-direction velocity components Calculate the average velocity in the x and y directions respectively, specifically:
[0046] in, for The average velocity of the target object in the x-axis direction within a frame image (unit: pixels / second). for The average velocity of the target object in the y-axis direction within a frame image (unit: pixels / second). For the first Frame to the Instantaneous velocity component of the target object in the x-axis direction (unit: pixels / second). For the first Frame to the Instantaneous velocity component of the target object in the y-axis direction (unit: pixels / second); The target resultant velocity is then obtained using the following formula:
[0047] in, for The average velocity of the target object within a frame image (characterizing the overall speed of the target's motion, unit: pixels / second) is used to characterize the overall motion trend of the target object; Finally, S226 is executed to integrate the constructed target coordinate sequence, frame-by-frame motion velocity, and target resultant velocity to form a complete target motion parameter set. This provides accurate motion feature data support for trajectory prediction in S23, ensuring the accuracy of trajectory prediction and thus guaranteeing the effectiveness of target tracking and locking.
[0048] Furthermore, S23 may include the following sub-steps: S231. Sort the target coordinate sequence and the corresponding target motion velocity; S232. Standardize the sorted target coordinate sequence and target motion velocity to obtain the target input vector; S233. Use the target input vector to input the preset trajectory prediction model to perform trajectory prediction and obtain the trajectory prediction result.
[0049] Standardization involves using algorithms such as Z-score to process the sorted target coordinate sequence and target motion velocity, eliminating the differences in the dimensions and numerical ranges of different parameters, and performing a preprocessing operation to bring the data to the same scale. The target input vector is a multi-dimensional vector (e.g., 3×m dimension, where m is the number of effective frames) formed by concatenating the standardized target coordinate components and motion velocity in a temporal sequence, adapting it to the input format of the pre-set trajectory prediction model. The pre-set trajectory prediction model is a model (e.g., a lightweight LSTM model) that has been pre-trained using massive amounts of target motion temporal data from experimental scenarios, and can capture the temporal dependencies of target motion to achieve trajectory prediction.
[0050] In this embodiment, based on the target coordinate sequence, frame-by-frame motion velocity, and target resultant velocity obtained in S22, S231 is first executed, using the timestamps of each frame acquisition. Based on the target coordinate sequence and corresponding frame-by-frame motion velocity in ascending order. Synchronous sorting is performed to remove discrete frame data with temporal anomalies, ensuring a one-to-one temporal correspondence between the sorted coordinate sequence and the velocity data, resulting in a sorted time-series dataset. ,in The number of valid frames after sorting ( ≤ (dimensionless), providing ordered and regular basic data for subsequent standardization and model input; Then, step S232 is executed, using the Z-score standardization method to standardize the sorted target coordinate sequence components and target motion velocity, eliminating the influence of dimensional differences and numerical ranges. The standardization formula for the x-component of the coordinate is:
[0051] The formula for standardizing the y-component of the coordinate is:
[0052] The standardized formula for motion speed is:
[0053] in, , The first The standardized (dimensionless) values of the x and y components of the frame center coordinates. , For the sorted number Frame center coordinates x and y components (unit: pixels). , These are the mean values (in pixels) of the x and y components of the center coordinates of all frames after sorting. , These are the standard deviations (in pixels) of the x and y components of the center coordinates of all frames after sorting. For the first The standardized value of frame motion velocity (dimensionless). For the sorted number Frame motion speed (unit: pixels / second). This represents the average motion speed of all frames after sorting (unit: pixels / second). The standard deviation of motion speed for all frames after sorting (unit: pixels / second).
[0054] Then, each frame was normalized. , , By concatenating the vectors sequentially, a target input vector of dimension 3×m is constructed. ; Next, S233 is executed. The pre-set trajectory prediction model uses a lightweight Long Short-Term Memory (LSTM) network model. This model has been pre-trained using massive amounts of target motion time-series data from experimental scenarios. The target input vector is input into the model, and the hidden layers of the model capture the temporal dependencies and velocity change patterns of the target motion, outputting the future... frame( The target center coordinates are predicted for the preset number of prediction frames (dimensionless). This is the trajectory prediction result, where... For the first The predicted center coordinates of the frame (in pixels). For the first The frame prediction timestamp (unit: seconds) relies on ordered and standardized data input and pre-trained models to ensure the timeliness and accuracy of trajectory prediction, providing a reliable position prediction basis for the S24 drive equipment to lock onto the target.
[0055] Furthermore, the trajectory prediction result includes the target trajectory prediction coordinates, and S24 may include the following sub-steps: S241. Calculate the first offset between the predicted coordinates of the target trajectory and the coordinates of the image center of the preprocessed images of consecutive frames; S242. Based on the first offset, dynamically adjust the rotation speed of the gimbal according to the target velocity; S243. Determine the rotation direction of the gimbal based on the target's movement speed and execute the rotation until the predicted coordinates of the target trajectory coincide with the coordinates of the image center. S244. Acquire real-time images of the target using an image acquisition device mounted on the rotated pan-tilt platform; S245. Calculate the second offset between the predicted coordinates of the target trajectory and the actual coordinates of the real-time image of the target; S246. Based on the second offset, the PID control algorithm is used to fine-tune and lock the gimbal, and the locked target detection image is acquired through the image acquisition device after fine-tuning and locking.
[0056] The predicted target trajectory coordinates are the center coordinates (two-dimensional pixel coordinates, unit: pixels) of the predicted target in a future frame (or at a certain moment) from the trajectory prediction results, serving as the target position for gimbal adjustment. The image center coordinates are the geometric center coordinates of the preprocessed images of consecutive frames, determined by the image resolution (horizontal resolution W / 2, vertical resolution H / 2), serving as the reference position for gimbal target locking. The first offset is the difference between the predicted target trajectory coordinates and the image center coordinates in the horizontal and vertical directions, representing the deviation between the predicted target position and the image center, providing a basis for gimbal coarse adjustment. The gimbal is a rotatable platform for mounting image acquisition equipment, which can adjust its horizontal and vertical rotation speed and direction according to control commands to achieve viewpoint adjustment and target tracking of the image acquisition equipment. The rotation speed is the rate at which the gimbal drives the image acquisition equipment to rotate (unit: degrees / second), dynamically adjusted based on the first offset and the target's combined velocity. The system ensures that the target's movement speed is matched; the rotation direction is the orientation of the gimbal rotation (horizontal left / right, vertical up / down), determined based on the target's movement speed direction and the sign of the first offset, bringing the target's predicted coordinates closer to the image center; the real-time target image is the target image captured in real-time by the image acquisition device after the gimbal rotates according to the trajectory prediction result, reflecting the target's current actual position; the second offset is the difference between the target's trajectory predicted coordinates and the target's actual coordinates in the real-time image, representing the deviation between the predicted position and the true position, providing a basis for gimbal fine-tuning; the PID control algorithm is a proportional-integral-derivative control algorithm, which calculates the fine-tuning amount through the proportional, integral, and derivative stages to achieve precise control of the gimbal's attitude and eliminate the second offset; fine-tuning lock is based on the fine-tuning angle calculated by the PID control algorithm, adjusting the gimbal attitude to make the target's actual coordinates completely coincide with the image center coordinates, achieving stable target locking.
[0057] In this embodiment, based on the trajectory prediction result obtained in S23, S241 is first executed, assuming the target trajectory prediction coordinates are... (Unit: pixels), where, For the first The center coordinates of the target trajectory prediction in the frame (two-dimensional pixel coordinates, representing the predicted position of the target at a future time, unit: pixels). For the first The horizontal (x-axis) component of the frame target trajectory prediction coordinates (unit: pixels). For the first Vertical (y-axis) component of the frame target trajectory prediction coordinates (unit: pixels). The number of valid historical frames after sorting (dimensionless). ≤ , (Number of original consecutive frames) The index of the predicted frame (dimensionless, value 1~). , k is the preset number of future prediction frames). The image center coordinates of the consecutive preprocessed images are: (Unit: pixels, determined by image resolution, e.g., when the resolution is W×H) ),in, The center coordinates (two-dimensional pixel coordinates, serving as the reference position for gimbal locking, unit: pixels) of the preprocessed images in consecutive frames. The horizontal (x-axis) component of the image center coordinates (unit: pixels). Let be the vertical (y-axis) component of the image center coordinates (in pixels), W be the horizontal resolution of the preprocessed image (total pixels), and H be the vertical resolution of the preprocessed image (total pixels). Using the formula... , Calculate the first offset, where, This is the first horizontal offset (unit: pixels). The first offset in the vertical direction (unit: pixels) represents the target's position to the right / left and below / above the image center, respectively. Then, step S242 is executed to dynamically adjust the gimbal rotation speed based on the first offset and the target's combined velocity v. The formula for the horizontal rotation speed is... The formula for vertical rotational speed is: ,in, The horizontal rotation speed of the gimbal (unit: degrees / second). The vertical rotation speed of the gimbal (unit: degrees / second). , This is a preset proportionality coefficient (dimensionless, ranging from 0.1 to 1.0). The maximum resultant velocity of the target in the test scenario (unit: pixels / second, determined by scene preset value) is used. This dynamic adjustment logic can avoid loss of field of view caused by mismatch between the gimbal rotation speed and the target movement speed. Then, S243 is executed to determine the gimbal rotation direction based on the sign of the first offset. If >0, the gimbal will rotate horizontally to the right. If <0, then turn horizontally to the left. If >0, the gimbal will rotate vertically downwards. If the value is less than 0, rotate vertically upwards, continuing until... =0 and =0, meaning the predicted coordinates of the target trajectory coincide with the coordinates of the image center. Then, execute S244 to acquire a real-time image of the target using the image acquisition device on the rotated pan-tilt platform, obtaining the target's actual position at the current moment. Then execute S245, assuming the actual coordinates of the target's real-time image are... (Unit: pixels), where, The actual center coordinates of the target in the real-time image (two-dimensional pixel coordinates, representing the target's true position at the current moment, unit: pixels). The horizontal (x-axis) component (unit: pixels) of the actual coordinates of the target in the real-time image. The vertical (y-axis) component (in pixels) of the actual coordinates of the target real-time image. This is expressed by the formula... , Calculate the second offset, where This is the second horizontal offset (unit: pixels). The second offset in the vertical direction (unit: pixels) is used to characterize the deviation between the predicted coordinates and the actual coordinates. Finally, S246 is executed, and the gimbal is fine-tuned and locked based on the second offset using a PID control algorithm. The formula for the fine-tuning angle in the horizontal direction is: The formula for fine-tuning the vertical angle is: ,in For fine-tuning the horizontal angle of the gimbal (unit: degrees). For fine-tuning the vertical angle of the gimbal (unit: degrees). , This is the proportionality constant (dimensionless). , The integral coefficient is dimensionless. , The differential coefficient (dimensionless) is used to eliminate the second offset through the proportional, integral, and differential steps of the PID algorithm, so that the actual coordinates of the real-time target image completely coincide with the coordinates of the image center. Finally, the locked target detection image is acquired by the fine-tuned and locked image acquisition device, providing a stable and focused target image for subsequent feature extraction.
[0058] Step 103: Extract features from the locked target detection image to obtain target feature data.
[0059] In this embodiment, based on the target detection image after locking the focused target area output in step 102, the feature extraction type corresponding to the current test scenario is first identified. If it is a face recognition scenario, 68 facial key points of the target are extracted through a Haar feature classifier, including key position information such as the corners of the eyes and the tip of the nose, and a 128-dimensional feature vector is generated. If it is a color recognition scenario, the target area in the image is cropped, the RGB values of the area are analyzed, and the main color data with the highest proportion is selected as the core feature. If it is an object state recognition scenario, the object contour is extracted through an edge detection algorithm, the contour size (length × width × height) and shape similarity parameters are calculated, and finally the target core feature data of the corresponding type is output. This process relies on the locked target detection image to ensure the accuracy of the feature extraction area and avoids interference from background redundant information. The generated feature data can accurately match the recognition requirements of different test scenarios, providing a highly recognizable data foundation for subsequent hierarchical cross-matching evaluation based on a dual-library standard feature library. At the same time, the targeted feature extraction strategy can also effectively improve the efficiency of subsequent matching operations and help solve the technical problem of delayed response in security status determination.
[0060] Step 104: Based on the preset dual-database standard feature library, perform hierarchical cross-matching evaluation using target feature data to obtain the security status judgment result of the target object.
[0061] In this embodiment, based on the core feature data of the corresponding type of target output in step 103, a hierarchical cross-matching evaluation is carried out using a preset dual-database standard feature library containing both local and cloud feature libraries. The target feature data is first matched quickly with the associated features in the local feature library to obtain a first matching degree. If the first matching degree meets the security judgment threshold, the security status result is directly output. If the matching degree is in the unverified range, the cloud feature library is further called for a second cross-matching. The matching results of the local and cloud libraries are combined to generate a comprehensive matching degree. Finally, the security status judgment result of the target object is output based on the comprehensive matching degree. This hierarchical matching strategy not only utilizes the low latency advantage of the local feature library to ensure rapid response in basic scenarios, but also improves the judgment accuracy in complex scenarios through the supplementary verification of the cloud feature library. It effectively solves the problem of response lag and insufficient judgment accuracy caused by the existing single local feature library matching. At the same time, the cross-matching logic can also reduce invalid cloud interactions and further optimize the overall process efficiency.
[0062] Furthermore, the preset dual-database standard feature library includes a local feature library and a cloud-based feature library. Step 104 may include the following sub-steps: S31. Calculate the first matching degree between the target feature data and the associated local feature data in the local feature library; S32. When the first matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target. S33. When the first matching degree is less than the second preset matching degree threshold, the target object is determined to be an unsafe target. S34. When the first matching degree is greater than or equal to the second preset matching degree threshold and less than the first preset matching degree threshold, the second matching degree between the target feature data and the cloud feature data associated in the cloud feature library is calculated. S35. Using the first matching degree and the second matching degree, and combining them with a preset weighting factor, a weighted calculation is performed to obtain the comprehensive matching degree; S36. When the overall matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target. S37. When the overall matching degree is less than the first preset matching degree threshold, the target object is determined to be an unsafe target. The first preset matching threshold is greater than the second preset matching threshold.
[0063] The local feature library is a feature database pre-stored on local devices (such as edge computing devices), storing standard feature vector sets of common security targets in the test scenario, with advantages of fast matching and low latency; the cloud feature library is a feature database deployed on a cloud server, storing a wider range and more comprehensive standard feature vector sets of security targets, supporting real-time updates, and making up for the shortcomings of the local feature library in terms of coverage; the first matching degree is calculated using a similarity algorithm (such as cosine similarity) to calculate the degree of matching between the target feature data and the associated feature data in the local feature library (value range [0,1]), with a higher value indicating a higher matching degree; the first preset matching degree threshold is a matching degree judgment threshold pre-determined through test data (value range [0.8,1]), serving as the core critical value for determining whether a target is a security target, with a value higher than the second preset matching degree threshold; the second preset The matching degree threshold is a low matching degree judgment threshold (range [0.4, 0.6]) pre-calibrated using experimental data. It serves as a critical value for initially determining that the target is an unsafe target and is used to divide the matching result range. The second matching degree is calculated using the same similarity algorithm as the first matching degree when the first matching degree is between the two thresholds. It calculates the degree of matching between the target feature data and the associated feature data in the cloud feature library (range [0, 1]). The preset weighting factor is a pre-set coefficient (including the local matching weighting factor α and the cloud matching weighting factor β) used to allocate the weights of the first and second matching degrees. It satisfies α + β = 1, with the cloud factor having a higher weight. The comprehensive matching degree is the result obtained by weighting and summing the first and second matching degrees using the preset weighting factor (range [0, 1]). It combines the local and cloud matching results to improve the accuracy of the judgment.
[0064] In this embodiment, based on the target detection image obtained after the aforementioned steps, target feature data is extracted using a feature extraction algorithm and constructed into a target feature vector. (Dimension is d×1, where d is the feature dimension, dimensionless). The local feature library pre-stores a set of standard feature vectors for common security targets in the test scenario, while the cloud feature library pre-stores a set of standard feature vectors for security targets that covers a wider range and is more comprehensive. The cloud feature library supports real-time updates. First, S31 is executed, and the cosine similarity algorithm is used to calculate the target feature vector and the local feature vector in the local feature library with the highest correlation. The first degree of matching between them is calculated using the following formula: ,in The first matching degree (value range [0,1], dimensionless, the larger the value, the higher the matching degree). This is the dot product operation between the target feature vector and the local feature vector. Let the magnitude of the target feature vector be denoted as . The modulus of the local feature vector is given, and then S32 is executed to set the first preset matching degree threshold as... (Values range [0.8, 1], dimensionless, determined by experimental data), when ≥ If the target object is determined to be a safe target, no further cloud matching operation is required. Then, S33 is executed to set the second preset matching threshold. (Values range [0.4, 0.6], dimensionless, determined by experimental data), and satisfy the following conditions: > ,when < When the target object is determined to be an unsafe target, if the first matching degree satisfies ≤ < Then, execute S34, which retrieves the cloud feature vector associated with the target feature vector from the cloud feature library via the network communication module. The same cosine similarity algorithm is used to calculate the second matching degree, and the corresponding calculation formula is as follows: ,in The second matching degree (value range [0,1], dimensionless). These are the standard feature vectors associated within the cloud-based feature library. This is the dot product operation between the target feature vector and the cloud feature vector. The feature vector in the cloud is the magnitude. Then, S35 is executed, introducing a preset weighting factor to weight and fuse the first and second matching degrees to calculate the comprehensive matching degree. The corresponding calculation formula is as follows: ,in The overall matching degree (value range [0,1], dimensionless). This is a weighting factor for local matching (value range [0.3, 0.5], dimensionless). The cloud-based matching degree weighting factor (value range [0.5, 0.7], dimensionless) satisfies the following conditions: The higher weighting factor value in the cloud is because the data in the cloud feature library is more comprehensive and updated more promptly, making the matching results more valuable. Then, steps S36 and S37 are executed to compare the overall matching degree with the first preset matching degree threshold. When comparing, ≥ When the target object is determined to be a security target, < When a target is determined to be an unsafe target, this dual-database hierarchical matching strategy combines the efficiency of local matching with the comprehensiveness of cloud matching, improving the accuracy and reliability of target determination while ensuring matching efficiency.
[0065] Furthermore, step 104 may also include the following sub-steps: S38. Based on the preset correction period, perform environmental adaptation correction on the first preset matching degree threshold.
[0066] The specific environmental adaptation correction is as follows: Perform grayscale conversion on the preprocessed image of the current frame to obtain the target grayscale image; Extract the light intensity and background complexity of the target grayscale image; The environment adaptability coefficient is calculated using light intensity and background complexity; The environmental adaptation coefficient and the preset dynamic correction coefficient are multiplied to obtain a new first preset matching degree threshold. Based on the new first preset matching degree threshold, the process jumps to the step of determining the target object as a safe target when the first matching degree is greater than or equal to the first preset matching degree threshold.
[0067] The preset correction cycle is a pre-defined time interval (range 10-60 seconds) that triggers the correction of the first preset matching degree threshold. It is determined by the device's computing power and the frequency of scene environment changes, and periodically adapts to environmental changes. Environment adaptation correction is an operation that dynamically adjusts the first preset matching degree threshold based on the current scene environment parameters (light intensity, background complexity) to avoid environmental interference causing judgment bias and improve robustness. The target grayscale image is a single-channel image obtained by grayscale conversion (weighted average method) of the current frame's preprocessed image, containing only grayscale information (range [0,255]), used to extract environmental feature parameters. Light intensity is a parameter characterizing the current experimental scene's lighting conditions, represented by the average grayscale value of the target grayscale image, reflecting the overall brightness of the image. Background complexity... The complexity is a parameter characterizing the degree of background interference in the current test scene. It is represented by the variance of the gradient magnitude of the target grayscale image. The larger the variance, the more complex the background texture and the stronger the interference. The environment adaptation coefficient is a coefficient obtained by weighted summation based on the normalized light intensity and background complexity (value range [0,1]), which quantifies the degree of influence of the environment on feature matching. The preset dynamic correction coefficient is a correction coefficient calibrated in advance through test data (value range [0.9,1.1]), which is used to adjust the correction magnitude of the environment adaptation coefficient on the original threshold to ensure the rationality of the threshold. The new first preset matching degree threshold is a threshold obtained by multiplying the original first preset matching degree threshold, the environment adaptation coefficient, and the preset dynamic correction coefficient. It is adapted to the current environment and replaces the original threshold for judgment.
[0068] In this embodiment, to improve the accuracy and robustness of target determination under different environments, step S38 is executed. First, a preset correction period is set to τ (ranging from 10 to 60 seconds, determined by the device's computing power and the frequency of scene environment changes). A threshold environment adaptation correction process is triggered every τ seconds. First, the preprocessed image of the current frame is converted to grayscale. Then, a weighted average method is used to convert the RGB color image to the target grayscale image. The corresponding calculation formula is as follows: ,in Coordinates in the target grayscale image The gray value at the location (range [0, 255], dimensionless). , , The preprocessed image is in coordinates The pixel values of the red, green, and blue channels at the location (value range [0, 255], dimensionless). Given image pixel coordinates (unit: pixels), after grayscale conversion, extract two environmental feature parameters from the target grayscale image: light intensity and background complexity. Light intensity... The average grayscale value of the target grayscale image is used for characterization, and the calculation formula is as follows: W and H represent the horizontal and vertical resolutions of the image (in pixels), respectively. The background complexity C is represented by the gradient variance of the grayscale image. The horizontal gradient is first calculated using the Sobel operator. with vertical gradient Then calculate the gradient magnitude. Finally, the background complexity is obtained. That is, the variance of the gradient magnitude (dimensionless), followed by the light intensity. With background complexity After normalization, we get , ,in , These represent the minimum and maximum light intensity in the test scenario, respectively. , These represent the minimum and maximum background complexity in the test scenario, respectively. , The normalized environmental characteristic parameters (value range [0,1], dimensionless) are then used to calculate the environmental adaptability coefficient. The calculation formula is: ,in , which is the light intensity weighting factor (value range [0.6, 0.8], dimensionless). The background complexity weighting factor (value range [0.2, 0.4], dimensionless) satisfies the following conditions: The higher weighting factor for light intensity is because lighting conditions have a more significant impact on feature matching. A preset dynamic correction coefficient is then introduced. (Value range [0.9, 1.1], dimensionless, calibrated by experimental data), the original first preset matching degree threshold is calculated by multiplying the environmental adaptation coefficient and the preset dynamic correction coefficient. After correction, a new first preset matching degree threshold is obtained, calculated using the following formula: ,in, The first preset matching threshold after correction (value range [0.7,1], dimensionless). The original first preset matching threshold (value range [0.8, 1], dimensionless) is used when the ambient light is dim ( Smaller) or more complex background ( When it is relatively large, The value decreases. The corresponding reduction is made to avoid misjudging unsafe targets due to decreased feature matching accuracy caused by environmental interference. This is especially important when the ambient light is sufficient and the background is simple. The value increases, To ensure the rigor of security target determination, the criteria are improved accordingly, and finally, based on the new first preset matching degree threshold... Then, the process jumps to execute the operation in step S32, which states that "if the first matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target." This enables dynamic environmental adaptation of the threshold and improves the adaptability and reliability of the entire target determination process.
[0069] Step 105: Generate corresponding feedback signals based on the safety status determination results.
[0070] In this embodiment, based on the target object safety status determination result output in step 104, when the determination result is a safe target, a safety feedback signal is generated to trigger corresponding safety actions, such as synchronously recording target status information and maintaining the current monitoring mode; when the determination result is an unsafe target, an alarm feedback signal is generated to trigger corresponding alarm actions, such as activating audible and visual warnings and pushing abnormal information to associated terminals. This feedback mechanism relies on the accurate determination result obtained from the previous hierarchical cross-matching evaluation to ensure the timeliness and effectiveness of the feedback signal, echoing the efficient closed loop of the entire process from image acquisition to status determination, and further ensuring the rapid response capability of dynamic target monitoring in experimental scenarios.
[0071] Furthermore, step 105 may include the following sub-steps: S41. When the safety status determination result is an unsafe target, an alarm feedback signal is generated and the corresponding alarm action is executed. S42. When the safety status determination result is a safety target, a safety feedback signal is generated and the corresponding safety action is executed.
[0072] Alarm feedback signals are structured signals generated when a target is determined to be unsafe. They include target feature data, real-time location coordinates, and determination timestamps, and are divided into local trigger signals and cloud synchronization signals. Alarm actions are a series of protective operations triggered by alarm feedback signals, including local audible and visual alarms, high-definition capture, data storage, cloud-based early warning push notifications, and activation of linked protective equipment (such as closing access control). Security feedback signals are structured signals generated when a target is determined to be safe. They include target identity information, matching degree data, and passage time, and are used for local alerts and cloud archiving. Security actions are a series of operations triggered by security feedback signals, including local de-alarm, PTZ reset, restoration of normal monitoring mode, cloud-based record archiving, and linked access control opening (if necessary).
[0073] In this embodiment, based on the target object safety status determination result obtained in step 104, the feedback and action control process corresponding to step 105 is executed. When the determination result is an unsafe target, the system immediately triggers the alarm feedback signal generation module to generate a structured alarm feedback signal containing target feature data, real-time location coordinates, and determination timestamp. This signal is divided into a local trigger signal and a cloud synchronization signal. The local trigger signal executes an alarm action through the device's built-in audible and visual alarm module, i.e., the red warning light flashes at a frequency of 2Hz, and the buzzer continuously emits an alarm prompt sound at a value of 80dB, simultaneously driving the image acquisition device on the pan-tilt-zoom platform to perform continuous high-definition capture of 5 frames. The captured images are bound and stored in the local storage module along with the target feature data. The cloud synchronization signal is pushed to the remote monitoring platform in real time through a wireless communication module (such as 4G / 5G). The platform interface automatically pops up an alarm pop-up window and marks the target location and feature information, while simultaneously linking the linkage protection equipment in the test scenario, such as closing the area access control. The system activates infrared security devices to prevent unsafe targets from entering the controlled area. When a target is determined to be safe, the system generates a security feedback signal containing the target's identity information, matching degree data, and passage time. This signal is transmitted to the local display module, displaying a "Safe Target" message and matching degree value on the device's screen. Simultaneously, it is uploaded to the cloud server for safe target passage record archiving, facilitating subsequent traceability and querying. Corresponding security actions are executed: if a warning was previously issued, it automatically de-warns, stops the audible and visual alarm module, and drives the pan-tilt unit to reset to the preset patrol position, restoring normal image acquisition and target monitoring modes. For controlled scenarios requiring access permissions, it can also automatically open the access channel in conjunction with the access control system without manual intervention. The entire feedback and action control process is automated based on the judgment result, ensuring timely warning and handling of unsafe targets while achieving efficient passage and recording of safe targets, thus improving the efficiency and intelligence level of scenario control.
[0074] Please see Figure 2 This application provides an image feature recognition experimental device, comprising: The acquisition module 201 is used to acquire real-time images of the test scene and perform preprocessing to obtain continuous frame preprocessed images; Tracking module 202 is used to perform target tracking and locking on consecutive preprocessed images to obtain locked target detection images; Extraction module 203 is used to extract features from the locked target detection image to obtain target feature data; Evaluation module 204 is used to perform hierarchical cross-matching evaluation based on a preset dual-database standard feature library and target feature data to obtain the security status judgment result of the target object; The execution module 205 is used to generate corresponding feedback signals based on the safety status determination results.
[0075] Since the above is a device corresponding to one image feature recognition experimental method, and its implementation principle is the same as that of an image feature recognition experimental method, for the sake of convenience and brevity, those skilled in the art can clearly understand that the specific working process of the device and module described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0076] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0077] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection between devices or units through some interfaces, and may be electrical, mechanical, or other forms.
[0078] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0079] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0080] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0081] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this disclosure.
[0082] Furthermore, the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article, unless otherwise specified, generally indicates that the preceding and following related objects have an "or" relationship.
[0083] The embodiments described in this specific implementation are preferred embodiments of this application and are not intended to limit the scope of protection of this application. Identical components are represented by the same reference numerals. Therefore, all equivalent changes made to the structure, shape, and principle of this application should be covered within the scope of protection of this application.
Claims
1. An image feature recognition experimental method, characterized in that, include: Real-time images of the test scene are acquired and preprocessed to obtain continuous preprocessed images; Target tracking and locking are performed on the preprocessed images of the consecutive frames to obtain the locked target detection images; Feature extraction is performed on the locked target detection image to obtain target feature data; Based on a pre-defined dual-database standard feature library, the target feature data is used to perform hierarchical cross-matching evaluation to obtain the security status determination result of the target object. Based on the security status determination result, a corresponding feedback signal is generated.
2. The image feature recognition experimental method according to claim 1, characterized in that, The real-time acquisition and preprocessing of real-time images of the test scene to obtain consecutive preprocessed images includes: Real-time acquisition of images of the test scene; Gaussian filtering is applied to the real-time images of the test scene to obtain denoised images of consecutive frames; Image enhancement is performed on the denoised images of the consecutive frames to obtain pre-processed images of the consecutive frames.
3. The image feature recognition experimental method according to claim 1, characterized in that, The step of performing target tracking and locking on the preprocessed images of the consecutive frames to obtain the locked target detection image includes: Target detection is performed on the preprocessed images of the consecutive frames to obtain the target detection parameters of the target object; Using the target detection parameters, the target motion parameter set is determined; Trajectory prediction is performed using the target motion parameter set to obtain the trajectory prediction result; The trajectory prediction results drive the image acquisition device to perform target locking operations, resulting in a locked target detection image.
4. The image feature recognition experimental method according to claim 3, characterized in that, The target detection parameters include the coordinates of the bounding rectangle of the target object. Determining the target motion parameter set using the target detection parameters includes: Determine the sequence of center coordinates of the bounding rectangle of the target object; Extract the acquisition timestamps of the consecutive preprocessed images; The target's movement speed is determined using the frame center coordinate sequence and the corresponding acquisition timestamp; A target coordinate sequence is constructed using the frame center coordinate sequence and the corresponding acquisition timestamp; Based on the target's motion speed, determine the target's resultant velocity; By integrating the target coordinate sequence, the target motion velocity, and the target resultant velocity, a target motion parameter set is obtained.
5. The image feature recognition experimental method according to claim 4, characterized in that, The process of using the target motion parameter set to perform trajectory prediction and obtain trajectory prediction results includes: Sort the target coordinate sequence and the corresponding target motion velocity; The sorted target coordinate sequence and target motion velocity are standardized to obtain the target input vector; The target input vector is used to input a preset trajectory prediction model to perform trajectory prediction, and the trajectory prediction result is obtained.
6. The image feature recognition experimental method according to claim 4, characterized in that, The trajectory prediction result includes the target trajectory prediction coordinates. The step of driving the image acquisition device based on the trajectory prediction result to perform a target locking operation to obtain a locked target detection image includes: Calculate the first offset between the predicted coordinates of the target trajectory and the image center coordinates of the consecutive preprocessed images; Based on the first offset, the rotation speed of the gimbal is dynamically adjusted according to the target combined velocity; The rotation direction of the gimbal is determined based on the target's movement speed, and the rotation is executed until the predicted coordinates of the target trajectory coincide with the coordinates of the image center. Real-time images of the target are acquired by an image acquisition device mounted on the rotated gimbal. Calculate the second offset between the predicted coordinates of the target trajectory and the actual coordinates of the real-time image of the target; Based on the second offset, the PID control algorithm is used to fine-tune and lock the gimbal, and the image acquisition device after fine-tuning and locking acquires the target detection image.
7. The image feature recognition experimental method according to any one of claims 1-6, characterized in that, The preset dual-database standard feature library includes a local feature library and a cloud-based feature library. The step of using the target feature data to perform hierarchical cross-matching evaluation based on the preset dual-database standard feature library to obtain the security status determination result of the target object includes: Calculate the first matching degree between the target feature data and the associated local feature data in the local feature library; If the first matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target; If the first matching degree is less than the second preset matching degree threshold, the target object is determined to be an unsafe target; When the first matching degree is greater than or equal to the second preset matching degree threshold and less than the first preset matching degree threshold, the second matching degree between the target feature data and the cloud feature data associated in the cloud feature library is calculated. The first matching degree and the second matching degree are combined with a preset weighting factor to perform a weighted calculation to obtain the comprehensive matching degree; When the overall matching degree is greater than or equal to the first preset matching degree threshold, the target object is determined to be a safe target; If the overall matching degree is less than the first preset matching degree threshold, the target object is determined to be an unsafe target; The first preset matching threshold is greater than the second preset matching threshold.
8. The image feature recognition experimental method according to claim 7, characterized in that, The generation of the corresponding feedback signal based on the security status determination result includes: If the safety status determination result is an unsafe target, an alarm feedback signal is generated and the corresponding alarm action is executed. When the security status determination result is a security target, a security feedback signal is generated and the corresponding security action is executed.
9. The image feature recognition experimental method according to claim 7, characterized in that, It also includes performing environmental adaptation correction on the first preset matching degree threshold based on a preset correction period, wherein the environmental adaptation correction specifically includes: Perform grayscale conversion on the preprocessed image of the current frame to obtain the target grayscale image; Extract the light intensity and background complexity of the target grayscale image; The environment adaptation coefficient is calculated using the light intensity and the background complexity. The environmental adaptation coefficient and the preset dynamic correction coefficient are multiplied to obtain a new first preset matching degree threshold. Based on the new first preset matching degree threshold, the process jumps to the step of determining the target object as a safe target when the first matching degree is greater than or equal to the first preset matching degree threshold.
10. An image feature recognition experimental device, characterized in that, include: The acquisition module is used to acquire real-time images of the test scene and perform preprocessing to obtain continuous preprocessed images. The tracking module is used to perform target tracking and locking on the continuous frame preprocessed images to obtain the locked target detection image; The extraction module is used to extract features from the locked target detection image to obtain target feature data; The evaluation module is used to perform hierarchical cross-matching evaluation based on the target feature data using a preset dual-database standard feature library to obtain the security status judgment result of the target object; The execution module is used to generate a corresponding feedback signal based on the security status determination result.