A mobile analysis camera end side target identification and abnormal alarm method
Patent Information
- Application Number
- CN202611251169.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-25
AI Technical Summary
当目标被短暂遮挡或检测模型产生误检时,传统跟踪器极易发生轨迹跳变并持续累积误差,最终将虚警信号上报至地面集控中心
1.本申请通过引入巷道基准轨迹向量作为空间先验知识,从根本上改变了传统视觉跟踪仅依赖图像表观特征的局限。计算目标框中心点位移矢量与巷道基准轨迹向量的点积,能够基于向量投影理论精准区分沿巷道方向的有效运动与横向干扰或背景噪声。进一步构建视觉跳变、目标面积变化率与载具物理振动峰值的三重复合校验机制,能够准确识别因振动模糊或光照突变引起的检测模型误检,并通过动态修正置信度权重及时削弱异常帧对全局轨迹的干扰。有效解决了井下恶劣环境导致的高误报率难题,使目标跟踪轨迹在干扰环境下依然保持高度的连续性和准确性,提升了端侧设备输出报警信息的有效性,降低了安全监管的漏报风险和人力复核成本。
Smart Images

Figure CN122821170A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of visual measurement technology, and in particular to a method for target recognition and anomaly alarm at the end of a mobile analysis camera. Background Technology
[0002] Inspection operations in coal mine roadways and tunnels primarily rely on manual foot patrols or fixed monitoring equipment. Manual inspections are inefficient and pose safety hazards, while fixed monitoring equipment has limited coverage and cannot achieve dynamic monitoring of the entire roadway cross-section. With the development of edge computing technology, some inspection vehicles are now equipped with cameras, but due to the harsh underground environment and drastic changes in lighting, the image recognition accuracy of existing equipment is easily and significantly reduced, failing to meet the needs of real-time early warning.
[0003] Existing edge-side target recognition solutions typically deploy general target detection algorithms directly on embedded devices. However, the underground equipment generates continuous mechanical vibrations during operation, leading to image blurring and irrational target bounding box drift. Simultaneously, edge-side devices have extremely limited computing power, power, and storage resources, and the underground network conditions are highly unstable. Existing algorithms do not fully consider the coupled effects of vibration interference and resource constraints, often resulting in equipment shutdown due to overheating, frequency throttling, or power depletion, thus missing critical anomalies and severely impacting the reliability of the monitoring system.
[0004] In target tracking, existing solutions mostly rely on conventional feature point matching or intersection-union-ratio (IVR) correlation, lacking spatial geometric constraints for specific roadway orientations. When the target is briefly obscured or the detection model produces false detections, traditional trackers are prone to trajectory jumps and continuous error accumulation, ultimately reporting false alarm signals to the ground control center. This not only reduces the reliability of alarm information but also increases the workload of back-end personnel in filtering out invalid alarms. Summary of the Invention
[0005] The embodiments of this application provide a method for target identification and anomaly alarm at the edge of a mobile analysis camera, achieving highly reliable, low-power, and self-consistent mobile target identification and anomaly alarm in harsh environments such as coal mine roadways. To achieve the above objectives, this application adopts the following technical solution: A method for target identification and anomaly alarm at the edge of a mobile analysis camera, the method comprising: Upon power-on, the system reads the remaining battery power, chip temperature, pre-stored tunnel baseline trajectory vector, current real-time network bandwidth, and preset threshold for the currently effective trajectory matching coefficient. It then extracts the average brightness of historical image frames to obtain illumination intensity parameters and acquires real-time vibration data: if a vibration sensor is installed, it collects vehicle vibration waveform data; otherwise, it calculates the global motion vector by matching feature points on multiple consecutive frames of original images, separating the high-frequency jitter component as equivalent vibration waveform data. These parameters are then summarized to obtain the device status set. The original images are acquired and superimposed with the real-time vibration data to obtain an image sequence with vibration markers; Electronic image stabilization is performed on the image sequence to obtain a stabilized image; Feature extraction and bounding box labeling are performed on the stabilized image, and the bounding boxes and category labels are output. The trajectory matching coefficient is obtained by calculating the dot product of the displacement vector of the target box center point and the reference trajectory vector of the roadway; When the trajectory matching coefficient is greater than a preset threshold, the target bounding box and category label of multiple consecutive frames are associated to obtain a continuous tracking trajectory; Based on the remaining battery power, chip temperature, peak vibration waveform data, and light intensity parameters in the device status set, adjust the image acquisition frame rate and feature extraction frequency; When the duration of the continuous tracking trajectory exceeds a preset duration, a stable image of the corresponding time period is captured to obtain a time-series image sequence and structured alarm data; When the network is interrupted, the structured alarm data and time-series image sequence are sorted by location number and stored in a local storage queue; When the network is restored, data is extracted from the local storage queue according to the alarm level and sent to the backend.
[0006] In some possible implementations, the step of associating the target bounding box and category label across multiple consecutive frames to obtain a continuous tracking trajectory when the trajectory matching coefficient is greater than a preset threshold includes: Extract the displacement vector of the center point of the target box in two adjacent frames of the continuous tracking trajectory, calculate the second dot product of the displacement vector and the reference trajectory vector of the roadway, and generate the directional consistency coefficient of adjacent frames. When the directional consistency coefficient of the adjacent frames is lower than the preset directional threshold, it is determined that the target box has undergone a trajectory change, and a change marker is generated; The jump markers are superimposed on the continuous tracking trajectory to generate a tracking sequence with jump markers; When the number of consecutive occurrences of the jump marker in the tracking sequence with the jump marker exceeds a preset threshold, the association relationship of the target box is released, and the corresponding continuous tracking trajectory is terminated.
[0007] In some possible implementations, generating the tracking sequence with jump markers includes: Using the tracking sequence with jump markers as input, the area change rate of the target boxes is statistically analyzed to generate area fluctuation parameters; The area fluctuation parameter is compared with a preset area fluctuation threshold; when the area fluctuation parameter is greater than the preset area fluctuation threshold, an area anomaly marker is generated. Perform a logical AND operation between the area anomaly marker and the jump marker to generate a composite misjudgment identifier; When the composite misjudgment identifier is valid, the peak value of the most recently acquired vibration waveform data is retrieved from the device status set to generate a vibration interference factor; The confidence level of the original category label is attenuated and corrected based on the vibration interference factor to generate a corrected single-frame recognition result; Based on the confidence level of the corrected single-frame recognition result, the target association operation or alarm trigger determination is re-executed to obtain the corrected tracking sequence with jump marker.
[0008] In some possible implementations, generating the corrected single-frame recognition result includes: Read the chip temperature from the device status set. When the chip temperature is higher than a preset high temperature threshold, reduce the baseline value of the feature extraction frequency and generate a frequency reduction compensation command. In response to the frequency reduction compensation command, a frame interval identifier is inserted into the image sequence with vibration markers to generate a sparsely sampled image sequence; Electronic image stabilization is performed on the sparsely sampled image sequence to generate a frequency-reduced, image-stabilized image; The process of inputting feature extraction and target box calibration of the down-frequency stabilized image is used to output the down-frequency target box and category label as the corrected single-frame recognition result.
[0009] In some possible implementations, after obtaining the down-concentration target bounding box and category label as the corrected single-frame recognition result, the method further includes: The continuous tracking trajectory is updated based on the corrected single-frame recognition result to obtain the temperature-adaptive tracking trajectory. Historical image frames corresponding to the temperature adaptation tracking trajectory are extracted, and the average brightness of the image frames is extracted to generate illumination intensity parameters. The environmental interference coefficient is calculated based on the light intensity parameter and the peak value of the vibration waveform data. When the environmental interference coefficient is greater than the preset interference threshold, the image acquisition frame rate is reduced, and target detection is performed with a computational complexity lower than that of full-resolution feature extraction: the input image is downsampled and features are extracted, or the sliding step size of the feature extraction operator is increased, and the detection result that meets the preset response threshold is output, and a low-power detection result record is generated. When the environmental interference coefficient falls below the interference threshold, full frame rate acquisition and complete feature extraction are resumed, the low-power detection result record is called to supplement the target box calibration, and the temperature adaptation tracking trajectory is updated.
[0010] In some possible implementations, adjusting the image acquisition frame rate and feature extraction frequency based on the remaining battery power, chip temperature, vibration waveform data peak value, and light intensity parameters in the device status set includes: The remaining battery power in the device status set and the illumination intensity parameters extracted from historical image frames are obtained. When the remaining battery power is lower than a preset low battery threshold, the current continuous tracking trajectory is locked, and a trajectory freeze command is generated. In response to the trajectory freeze command, a stock trajectory maintenance mode is generated; In the existing trajectory maintenance mode, the positions of existing target boxes in the continuous tracking trajectory are updated to increase the confidence threshold for new target detection, and temporary tracking trajectories are created for new targets that meet the high threshold. When the maintenance tracking trajectory or the temporary tracking trajectory meets the corresponding duration determination condition, the stable image of the corresponding time period is captured to obtain the time-series image sequence and structured alarm data, and the structured alarm data is written into the local storage queue to obtain the standby trigger condition; Based on the computational load under the existing trajectory maintenance mode and the standby triggering conditions, the image acquisition frame rate and the feature extraction frequency are reduced.
[0011] In some possible implementations, after reducing the image acquisition frame rate and the feature extraction frequency, the method further includes: Read the occupied capacity of the local storage queue and generate storage remaining parameters; When the remaining capacity parameter is lower than the preset capacity threshold, the time-series image sequences in the local storage queue are traversed, the duration information of each segment is extracted, and a duration list is generated. The duration list is sorted according to the alarm level, and the longest time-series image sequence is selected to generate a segment identifier to be deleted. Based on the identifier of the segment to be deleted, the frame rate of the segment to be deleted is compressed to generate a compressed time-series image sequence; The compressed time-series image sequence is used to replace the original time-series image sequence, and the storage margin parameter of the local storage queue is updated to obtain alarm data with optimized storage space.
[0012] In some possible implementations, upon network recovery, data is retrieved from the local storage queue according to the alarm level and sent to the backend, including: Read the category labels from the structured alarm data, and generate an immediate alarm identifier when the category label is a specific hazard type; Write the real-time alarm identifier into the retransmission task package to generate a retransmission task package with the identifier; The identified retransmission task packets are extracted first to generate an emergency retransmission data stream; Read the real-time network bandwidth value from the device status set and generate available bandwidth parameters; Adjust the sending rate of the emergency retransmission data stream according to the available bandwidth parameters, generate a rate-adapted retransmission stream, and send the rate-adapted retransmission stream to the background.
[0013] In some possible implementations, after sending the rate-adapted supplementary stream to the backend, the method further includes: Receive alarm handling receipts returned from the backend, parse the handling conclusion field in the receipts, and when the handling conclusion field indicates a false alarm, extract the trajectory matching coefficients of all frames in the corresponding continuous tracking trajectory to generate a false alarm sample set; The difference between the trajectory matching coefficients in the false alarm sample set and the current preset threshold is calculated to generate a threshold correction amount; The preset threshold is updated using the threshold correction amount to generate a new preset threshold; Write the new preset threshold into the device state set, replacing the original preset threshold; When calculating the dot product of the displacement vector of the target box center point and the reference trajectory vector of the roadway to obtain the trajectory matching coefficient in the next calculation, the new preset threshold is called to perform a comparison operation to generate an optimized continuous tracking trajectory.
[0014] In some possible implementations, the step of performing feature extraction and bounding box labeling on the stabilized image, and outputting the bounding box and category label, includes: Edge feature enhancement is performed on the stabilized image, and linear contour features are extracted from the stabilized image to generate a contour feature vector; Calculate the similarity between the contour feature vector and the roadway reference trajectory vector to generate a contour matching score; When the contour matching degree is lower than a preset contour threshold, it is determined that the current field of view deviates from the roadway reference direction, and a viewpoint offset signal is generated; In response to the viewpoint offset signal, if the camera is equipped with a motorized pan-tilt unit, the pan-tilt unit is triggered to perform angle correction and generate a corrected and stabilized image; if the camera is installed with a fixed viewpoint, the stabilized image is adaptively cropped to retain the image area consistent with the roadway reference direction, or viewpoint offset alarm information is output to the backend. The corrected or cropped stabilized image is input into the feature extraction and bounding box labeling process, and the corrected bounding box and category label are output. The position information in the continuous tracking trajectory is updated using the corrected target box to generate a viewpoint calibration tracking trajectory.
[0015] As can be seen from the above technical solution, this application has the following beneficial effects: 1. This application fundamentally changes the limitation of traditional visual tracking, which relies solely on image appearance features, by introducing the roadway reference trajectory vector as spatial prior knowledge. Calculating the dot product of the target bounding box center point displacement vector and the roadway reference trajectory vector enables accurate differentiation between effective motion along the roadway direction and lateral interference or background noise based on vector projection theory. Furthermore, a triple-replication verification mechanism is constructed, incorporating visual jumps, target area change rate, and vehicle physical vibration peak values. This accurately identifies false detections caused by vibration blurring or sudden changes in illumination, and dynamically adjusts the confidence weights to promptly reduce the interference of abnormal frames on the global trajectory. This effectively solves the problem of high false alarm rates caused by harsh underground environments, ensuring that the target tracking trajectory maintains high continuity and accuracy even under interference conditions. It improves the effectiveness of alarm information output by end-side equipment, reduces the risk of missed alarms in safety supervision, and lowers the cost of manual verification.
[0016] 2. This application constructs a closed-loop survival strategy integrating perception, computing, and communication for scenarios with severely limited resources. Addressing the coupled constraints of computing power, power consumption, and thermal power consumption on the device side, this application dynamically performs sparse sampling frequency reduction based on chip temperature, automatically switches to the existing trajectory maintenance mode based on remaining power, and incorporates an intelligent storage space compression mechanism to ensure that the device can maintain core tracking capabilities without crashing or stopping before power is exhausted. To address the practical pain point of frequent underground network interruptions, a structured local temporary storage and network outage retransmission strategy based on point number and alarm level is designed, and rate adaptation is performed based on real-time bandwidth after network recovery. False alarm receipts returned from the backend are used to automatically correct the terminal-side decision threshold, forming a data-driven closed-loop self-evolution capability. These multi-dimensional system-level optimizations significantly extend the effective working time of the equipment in extreme environments, reduce operation and maintenance costs, and truly realize the unmanned, intelligent, and autonomous operation of mobile inspection devices, providing solid technical support for the construction of smart mines. Attached Figure Description
[0017] The invention will now be further described with reference to the accompanying drawings.
[0018] Figure 1 An overall flowchart provided for embodiments of this application; Figure 2 This is a flowchart of transition detection and termination provided in an embodiment of this application; Figure 3 A flowchart for correcting composite misjudgments provided in the embodiments of this application; Figure 4 A flowchart illustrating high-temperature frequency reduction and environmental interference provided in the embodiments of this application; Figure 5 The flowcharts for low-power maintenance, storage optimization, retransmission adaptation, and viewing angle correction provided in the embodiments of this application are as follows. Detailed Implementation
[0019] The terms "first," "second," and "third," etc., used in this application specification, claims, and drawings are used to distinguish different objects, not to limit a specific order.
[0020] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0021] Research has revealed that existing end-side identification solutions lack the ability to withstand interference from harsh environments such as downhole vibrations and sudden changes in lighting. Target tracking is prone to drift and jumps, resulting in a persistently high false alarm rate. Furthermore, existing solutions do not dynamically manage the power, temperature, and network outage conditions of end-side equipment, often leading to shutdowns and missed alarms due to overheating or power depletion, making it difficult to meet the autonomous operation requirements of mobile inspections.
[0022] To address the aforementioned problems, this application provides a method for target identification and anomaly alarm at the edge of a mobile analysis camera: Example 1 This embodiment provides a method for target recognition and anomaly alarm using a mobile analytical camera, applicable to mobile inspection scenarios in underground or subsurface spaces such as coal mine roadways and tunnels. In this scenario, the analytical camera is fixedly mounted on an inspection vehicle, such as a track-mounted inspection robot, a rubber-tired vehicle, or a monorail crane, and travels along the roadway with the vehicle, performing real-time visual monitoring of equipment status, personnel activity, and environmental anomalies within the roadway. Figures 1-5 As shown, the method in this embodiment includes the following steps.
[0023] Device self-test and initialization.
[0024] After the camera is powered on, it first performs a self-test and initialization. Specifically, a roadway reference trajectory vector is pre-written into the device's built-in storage unit. This roadway reference trajectory vector is obtained by pre-calibrating the roadway direction; it is essentially a dimensionless unit direction vector that describes the expected direction of the roadway extension in the two-dimensional image plane. The calibration method is as follows: during device deployment, the camera is positioned facing the roadway extension direction, the direction of the roadway centerline in the image is recorded, and this direction is normalized to a unit vector before being stored in the storage unit. For curved roadways, calibration can be performed segmentally along the roadway direction, with each segment corresponding to an independent reference vector. The device automatically switches the currently used reference vector based on its own positioning information, such as odometer readings. The odometer reading is mapped to the image plane direction vector via the camera's intrinsic parameter matrix and extrinsic parameter transformation table. For zoom lenses, the pre-stored reference vector is rotated and scaled according to the ratio of the current focal length to the focal length during calibration. The vehicle yaw angle is read from the odometer heading channel, and the reference vector is rotated by the same angle according to the yaw angle before being used in the dot product calculation. The internal and external parameter transformation tables and focal length ratio compensation coefficients are written into the storage unit once during the equipment deployment and calibration phase.
[0025] The remaining battery power, chip temperature, and tunnel baseline trajectory vector read above, along with the real-time acquired vibration waveform data, the current real-time network bandwidth value, and the currently effective trajectory matching coefficient preset threshold, are summarized to construct a device status set. This set is a dynamic variable group in the running memory, containing seven fields: remaining battery power (percentage), chip temperature (°C), tunnel baseline trajectory vector (two-dimensional unit vector), real-time vibration waveform data (three-axis acceleration value), current network bandwidth (bps), trajectory matching coefficient preset threshold (pixels), and environmental interference coefficient (dimensionless). During the device's initial power-on self-test phase, before image brightness and vibration peak values are acquired, the environmental interference coefficient E is temporarily set to a neutral initial value of 0.5. After the brightness L of the first frame image and the first vibration peak value V are obtained, the values are refreshed according to the formula in Example 5, and subsequently updated with each frame. Each field is updated in real time as the device operates, and is used by subsequent processing steps.
[0026] Original image acquisition and vibration labeling.
[0027] After the equipment completes its self-test, the camera continuously captures raw images of the tunnel at an initially set frame rate, which is pre-configured based on the equipment's computing power. For example, on a typical computing platform, this frame rate can be set to capture 15 to 30 frames per second. At the same time, vibration sensors installed on the vehicle, such as MEMS accelerometers, collect vibration waveform data generated by the vehicle during its movement in real time. This data is expressed in units of gravitational acceleration, recording the instantaneous acceleration values of the vehicle in three axes (forward and backward, left and right, and up and down).
[0028] This application overlays the raw image captured in each frame with the vibration waveform data corresponding to the time of image capture, generating an image sequence with vibration tags. Here, "overlay" refers to adding a vibration data field to the metadata area of the image data. This field contains the vibration amplitude and dominant frequency information along three axes at the time of image capture. Specifically, the vibration data is written into the EXIF private tag field of the image frame. The image sensor and MEMS accelerometer share the SoC master clock. At the frame-ready interrupt trigger moment, the most recent sampling snapshot is read from the vibration FIFO and written into the metadata of that frame, with the timestamp aligned to the millisecond level. The metadata of each image frame also includes the timestamp of the capture time to ensure time alignment between the image and the vibration data. Through this method, each image frame not only contains visual content but also carries physical vibration information synchronized with its time, providing a data foundation for subsequent electronic image stabilization and interference detection.
[0029] Electronic image stabilization.
[0030] After obtaining the image sequence with vibration markers, electronic image stabilization is performed on the image sequence to eliminate or reduce image blur and positional jitter caused by vehicle vibration, resulting in a stabilized image.
[0031] The specific processing logic of electronic image stabilization is as follows: For each frame in the image sequence, feature points are first extracted using a feature point detection algorithm, such as FAST corner detection. These feature points are then matched with the feature points of the previous frame to calculate the global motion vector between the two frames. This global motion vector includes the smooth displacement caused by the vehicle's active movement, i.e., the scanning motion and the random jitter component caused by vibration. To accurately separate these two components, this application uses the aforementioned vibration marker data as an auxiliary criterion: the vibration waveform data is low-pass filtered, for example, using a Butterworth filter with a relatively low cutoff frequency, to separate the low-frequency component (corresponding to active motion) and the high-frequency component (corresponding to random jitter). An inverse compensation vector is calculated based on the high-frequency jitter component, and an affine transformation is performed on the current frame image to correct the image content in spatial position. The corrected image is the stabilized image. This process is performed sequentially for each frame, generating a continuous sequence of stabilized images.
[0032] Target detection and bounding box calibration.
[0033] After obtaining the stabilized image, feature extraction and bounding box labeling are performed. This operation uses a pre-trained lightweight object detection model. This model is built on a deep separable convolutional neural network, with its backbone consisting of multiple stacked deep separable convolutional modules. Each module contains a deep convolutional layer (3×3 kernel size, stride 1 or 2), a batch normalization layer, a ReLU activation function, and a pointwise convolutional layer (1×1 kernel size). The network structure is divided into three scale feature extraction layers, outputting feature maps at different spatial resolutions for detecting large, medium, and small targets. During training, a training dataset containing tens of thousands of labeled tunnel scene images is used. Each target in each image is labeled with a bounding box indicating its location (top-left and bottom-right pixel coordinates) and category label, such as "personnel," "mining truck," or "equipment malfunction." During training, a focus loss function is used for classification, and a smoothed L1 loss function is used for regression. Multiple iterations of training are performed using a stochastic gradient descent optimizer until the total loss function value converges to a stable range. After training, the model weights are fixed and deployed to the inference engine of the edge device. During inference, the model performs forward computation on the input stabilized image and outputs a set of bounding boxes and their corresponding class labels and confidence scores. The bounding boxes are represented by the center point coordinates, width, and height; the class labels are one of the predefined classes; and the confidence scores are probability values between 0 and 1.
[0034] Trajectory matching coefficient calculation.
[0035] After obtaining the target bounding box in the current frame, it is necessary to determine whether the bounding box corresponds to a target moving along the lane direction. For each target bounding box detected in the current frame, it is associated with the target bounding box of the same target in the previous frame (skipping the previous frame if there is no previous frame for the first detection), and the displacement vector of the target bounding box center point on the image plane is calculated. The x-coordinate component of this displacement vector is equal to the x-coordinate of the target bounding box center point in the current frame minus the x-coordinate of the target bounding box center point in the previous frame; the y-coordinate component is processed similarly. The displacement vector is expressed in pixels.
[0036] Calculate the dot product (inner product) of the displacement vector and the roadway reference trajectory vector to obtain the trajectory matching coefficient. Let the displacement vector be... The reference trajectory vector of the tunnel is ( If the vector is a unit vector, then the formula for calculating the dot product is:
[0037] This formula originates from vector projection theory. The physical meaning of the dot product is the projected length of the displacement vector along the reference direction of the roadway. When the target moves along the roadway direction, the projected length is larger and positive; when the target moves laterally or in the opposite direction, the projected length is smaller or even negative. Dot product value The unit is pixels, and its size reflects the salience of the target's movement along the lane direction. To eliminate the pixel displacement differences caused by the target's distance, the dot product result is... Divide by the number of pixels diagonal of the current target bounding box The normalized projection ratio is obtained. (Dimensionless), trajectory matching coefficients are based on Participate in subsequent judgment; the preset threshold can be set to the normal projection ratio. Between 30% and 60% of the mean, this calculation is the core step in introducing spatial geometric constraints into target tracking and judgment, enabling the system to filter valid targets based on the direction of motion and avoid misjudging background noise or lateral interference as targets.
[0038] Establishment of continuous tracking trajectory.
[0039] When the calculated trajectory matching coefficient is greater than a preset threshold, the target movement direction corresponding to the current target box is determined to be sufficiently consistent with the roadway reference direction, and is therefore considered a valid target. The preset threshold is set based on the following: during the initial deployment of the equipment, trajectory matching coefficient data for all target boxes are collected over a period of normal operation, such as several hours, and their distribution characteristics are statistically analyzed. Targets moving normally along the roadway typically have trajectory matching coefficients distributed over a large positive range; while false detections caused by noise or interference tend to have trajectory matching coefficients concentrated near zero or in the negative range. Based on the statistical results, the preset threshold is set near the boundary between the normal and noise distributions, for example, between 30% and 60% of the average normal trajectory matching coefficient, to ensure that most real targets pass the judgment while filtering out most noise. This threshold can be dynamically optimized during operation through background feedback.
[0040] For the target bounding boxes that pass the determination, the target bounding boxes and category labels belonging to the same target in multiple consecutive frames are associated to form a continuous tracking trajectory for that target. The association method is as follows: For each target bounding box in the current frame, the feature similarity (based on the cosine similarity of the depth feature vectors of the image content within the target bounding box) and the positional proximity (the reciprocal of the Euclidean distance between the center points of the target bounding boxes) between it and all tracked target bounding boxes in the previous frame are calculated. The two are weighted and summed to obtain a comprehensive matching score, where the feature similarity weight is 0.6 and the positional proximity weight is 0.4, and the sum of the two weights is 1. During deployment, this can be adjusted according to the scenario, and the existing trajectory with the highest matching score is selected for association. If a target bounding box in the current frame cannot be matched with any existing trajectory, a new tracking trajectory is created for it. Each tracking trajectory includes a target bounding box position sequence, a category label sequence, a timestamp sequence, and a trajectory matching coefficient sequence.
[0041] Dynamic frame rate and feature extraction frequency adjustment.
[0042] During continuous operation, the remaining power is constantly consumed, and the chip temperature gradually rises. This application dynamically adjusts the image acquisition frame rate and feature extraction frequency based on the remaining power and chip temperature in the device status set. When the remaining power is above a certain high level and the chip temperature is below a certain low level, the device operates at a higher frame rate and frequency; when the power decreases or the temperature increases, the device appropriately reduces the frame rate and frequency. Specific adjustment strategies are detailed in Embodiments Six and Four.
[0043] Alarm triggering and data interception.
[0044] When the duration of a continuous tracking trajectory exceeds a preset duration, the target is determined to constitute a persistent event, triggering an alarm process. The preset duration is set based on the different timeliness requirements of different types of monitoring targets. For example, for targets like "personnel intrusion," a rapid response is required, and the preset duration can be set to several seconds; for targets like "abnormal equipment status," to avoid false alarms caused by transient noise, the preset duration can be set to more than ten seconds. This value is configured by administrators according to the actual scenario during equipment deployment and can be adjusted via backend commands during operation.
[0045] Upon alarm triggering, image frames corresponding to the continuous tracking trajectory period are extracted from the stored stabilized image sequence and packaged into a time-series image sequence, such as an MP4 file encoded with H.264. Simultaneously, structured information such as the target bounding box sequence, category label sequence, and trajectory matching coefficient sequence corresponding to the continuous tracking trajectory are summarized to generate structured alarm data, such as a JSON-formatted text file. The time-series image sequence provides visual evidence, while the structured alarm data facilitates automated analysis in the background.
[0046] Local storage when offline.
[0047] When a network interruption is detected, the device sorts the structured alarm data and time-series image sequences by location number and stores them in a local storage queue. The location number is a unique identifier for the device's installation location, uniformly assigned by the backend management system during device deployment. Its format is, for example, "lane number-mileage-device serial number," ensuring global uniqueness. Sorting and storing data by location number facilitates orderly management and rapid retrieval during subsequent data retransmissions, avoiding data corruption.
[0048] Re-transmit via network.
[0049] During network outages, in addition to writing structured alarm data to the local queue, the endpoint simultaneously activates audible and visual alerts. Once the network is restored, high-level alarm data is prioritized for retransmission based on alarm level. When the network is restored, data is retrieved from the local storage queue according to alarm level and sent to the backend. Alarm levels are pre-set based on the degree of danger posed by category labels; for example, "personnel entering a dangerous area" is a high level, and "minor equipment displacement" is a low level. High-level alarm data is sent first to ensure that the most urgent information is delivered immediately.
[0050] Through the complete process described above, this application achieves autonomous and intelligent operation of target recognition and anomaly alarm at the mobile analysis camera end in alleyway scenarios with limited resources, unstable networks, and harsh environments.
[0051] Example 2 During target tracking, occlusion, sudden changes in lighting, or false detection by the detection model may cause abnormal jumps in the tracking trajectory, i.e., a large abrupt change in the position of the target box center point that does not conform to the laws of physical motion. To address this, after establishing a continuous tracking trajectory, this application further extracts the displacement vector of the target box center point in two adjacent frames of the trajectory, calculates the second dot product of this displacement vector and the roadway reference trajectory vector, and uses this second dot product as the directional consistency coefficient between adjacent frames. This coefficient uses the same formula as the trajectory matching coefficient, but its application differs: the former is used for the "admission" determination of the target box in a single frame, while the latter is used for monitoring the motion continuity of the tracked target between adjacent frames.
[0052] When the directional consistency coefficient of adjacent frames is lower than the preset directional threshold, it indicates that the target's motion direction has deviated significantly from the roadway reference direction in a very short time. This does not conform to the normal physical motion law of the target, so it is determined that a trajectory jump has occurred and a jump marker is generated.
[0053] The preset direction threshold is set based on the following: When a vehicle travels in a tunnel, the target's direction of motion in the image will fluctuate within a certain range due to the vehicle's own acceleration, deceleration, and turning. This fluctuation amplitude can be estimated using the vehicle's kinematic model. For example, the vehicle's maximum lateral acceleration and maximum rate of change of heading angle determine the limit of the target's direction change between adjacent frames. This limit is multiplied by a certain safety factor, such as 120% to 150%, as the preset direction threshold. In the specific mapping, the vehicle's maximum lateral acceleration and maximum heading angular velocity are projected onto the image plane angular velocity using the camera's focal length and extrinsic parameters. Multiplying this by the single-frame time yields the limit of pixel direction change between adjacent frames. This limit is the baseline value before the safety factor calculation, allowing for normal direction fluctuations while also identifying abnormal abrupt changes. Specific values can be calibrated based on the vehicle parameters during equipment deployment.
[0054] A jump marker is superimposed onto the corresponding frame of the continuous tracking trajectory to generate a tracking sequence with jump markers. The number of consecutive occurrences of the jump marker in this sequence is then counted. When the number of consecutive occurrences exceeds a preset threshold, it indicates that the trajectory has become abnormal in multiple consecutive frames and has lost its reliable tracking basis. At this point, the association between the target bounding box and the trajectory is severed, and the tracking trajectory is terminated. The preset threshold is set based on the following: if a target is briefly occluded, for example, by another target momentarily occluding it for one or two frames, its directional consistency coefficient may be temporarily low, but it will return to normal after the occlusion ends. Considering that the duration of occlusion in the tunnel is usually short, for example, no more than a few frames, the preset threshold is set to a relatively small value, such as three to five times, which can tolerate false detections caused by brief occlusions while promptly terminating continuously drifting false trajectories.
[0055] This mechanism identifies and tracks drift early, avoiding the accumulation of false alarms, and prevents false termination caused by single-frame noise through multiple consecutive conditions.
[0056] Example 3 Some jump markers may be false detections caused by image blurring due to vehicle vibration, rather than actual target motion anomalies. Therefore, this application uses a tracking sequence with jump markers as input to statistically analyze the rate of change of the target bounding box area. The target bounding box area is equal to its width multiplied by its height, expressed in squared pixels. Let the... The area of the frame target box is , No. The area of the same target box in the frame is Then the rate of change of area The calculation formula is:
[0057] This formula is a dimensionless ratio, representing the fluctuation range of the area relative to the previous frame. During normal tracking, when the target moves along the channel at a constant or variable speed, the rate of area change should remain within a small range, for example, no more than 20% to 30%. When the rate of area change suddenly increases, for example, exceeding 50%, it indicates a drastic change in the size of the target bounding box, which usually corresponds to false detections by the detection model.
[0058] This application first compares the area fluctuation parameter with a preset area fluctuation threshold. When the area fluctuation parameter is greater than the preset area fluctuation threshold, an area anomaly marker is generated. Then, a logical AND operation is performed between the area anomaly marker and the jump marker. The specific judgment rule is: when a frame simultaneously meets both the conditions of "existence of jump marker" and "valid area anomaly marker", the composite false positive marker is set to valid. The preset area fluctuation threshold is set based on the fact that the target box area change rate under normal tracking conditions is usually no more than 20%~30%. 1.5 times this range (i.e., 30%~45%) is used as the threshold to balance normal motion fluctuations and false detection recognition requirements. At this time, the system judges that the jump change may be a false detection caused by external interference.
[0059] To further verify this, the peak value of the most recently acquired vibration waveform data from the equipment status set was retrieved, and the maximum value of the composite acceleration along the three axes was taken. If this peak value is high, for example, exceeding a certain empirical threshold, it indicates that the vehicle has experienced severe vibration, further supporting the inference of false detection. The confidence level of the original category label is attenuated and corrected based on the vibration interference factor: first, the vibration interference factor is normalized to the 0-1 range; then, the original confidence level is multiplied by (1 - the normalized vibration interference factor) to obtain the corrected confidence level, which is then written into the single-frame recognition result. For example, when the vibration peak value reaches the threshold for severe vehicle vibration, the normalized interference factor is 0.8, the original confidence level is 0.9, and the corrected confidence level is reduced to 0.18.
[0060] The trajectory matching coefficient is calculated by the dot product of the target box displacement vector and the roadway reference trajectory vector. It belongs to the geometric motion feature and is not affected by the category confidence. This application re-inputs the corrected single-frame recognition result into the target association process: if the corrected confidence is lower than the preset association threshold (e.g., increased from the default 0.5 to 0.7), the association between the target box of the current frame and the existing tracking trajectory is canceled, the jump mark corresponding to the canceled frame is synchronously cleared, the continuous jump count is only accumulated on the frames that have not been canceled, or the current frame is directly determined not to constitute an alarm trigger condition, thereby reducing the impact of vibration false detection on the tracking result from the data source level, and obtaining the corrected tracking sequence with jump mark.
[0061] Example 4 As the edge device operates continuously, the chip temperature gradually rises. When the chip temperature exceeds the preset high-temperature threshold, maintaining the original feature extraction frequency may lead to overheating, frequency throttling, or even damage. The preset high-temperature threshold is set based on the upper limit of the safe operating temperature specified in the chip datasheet, typically 85 degrees Celsius. The preset high-temperature threshold is set to 80% to 90% of this upper limit to allow for a safety margin. For example, if the upper limit is 85 degrees Celsius, the threshold can be set to 70% to 75 degrees Celsius.
[0062] When the chip temperature exceeds the threshold, a frequency reduction compensation command is generated. In response to this command, frame interval markers are inserted into the image sequence marked with vibration. Specifically, a skip marker is placed every preset number of frames, such as every one or two frames. Subsequent processing modules skip feature extraction and bounding box mapping for these frames, retaining only image acquisition and storage. After frame interval insertion, the originally dense image sequence becomes a sparsely sampled image sequence.
[0063] The same electronic image stabilization operation as in Example 1 was performed on the sparsely sampled image sequence to obtain a down-frequency stabilized image. This down-frequency stabilized image was then input into a feature extraction and bounding box labeling process to obtain the down-frequency bounding boxes and category labels, which were output as the corrected single-frame recognition result.
[0064] While frequency reduction decreases the number of recognition attempts, image acquisition continues, and unselected frames are cached. Once the temperature drops, the cached images can be used for retrospective analysis. Specifically, when the chip temperature falls below a preset high-temperature threshold by 5°C, retrospection is triggered. The most recent 30 frequency-reduced, stabilized images are extracted from the cache, and feature extraction and target bounding box calibration are re-executed. The retrospective results are then merged into the end of the existing tracking trajectory according to the timestamp, thus preventing permanent information loss. This mechanism forms a closed-loop control of "temperature load," ensuring continuous equipment operation without downtime.
[0065] Example 5 After obtaining the down-concentrated target bounding box and category label, the continuous tracking trajectory is updated based on these recognition results to obtain the temperature-adaptive tracking trajectory. Subsequently, the stabilized image frames corresponding to this trajectory are extracted, and the average brightness of these image frames is extracted as the illumination intensity parameter. The average brightness is calculated as follows: convert the image to grayscale, calculate the arithmetic mean of the grayscale values of all pixels, and then normalize it to the range of 0 to 1 (0 represents all black and 1 represents all white).
[0066] Simultaneously, the peak value of the aforementioned vibration waveform data was obtained. (Normalized to 0 to 1). Regarding the light intensity parameter... Peak values of the normalized vibration waveform data Perform a weighted summation operation to generate the environmental interference coefficient. :
[0067] in and These are weighting coefficients, and their sum is 1. The rationale behind this formula is: insufficient lighting ( Small) and violent vibration ( Both large and small objects can interfere with target recognition, and their impact on recognition accuracy can be balanced by weighting coefficients. In typical alleyway scenarios, insufficient lighting and severe vibration have roughly the same impact on recognition accuracy, therefore... and Each factor can be assigned a weight of 0.5. If a particular factor is more prominent in the actual scenario, the weight can be adjusted through on-site calibration.
[0068] When the environmental interference coefficient When the interference exceeds the preset threshold, it indicates that environmental conditions strongly interfere with recognition accuracy. In this case, a frequency reduction command is generated instead of completely pausing recognition. In response to the frequency reduction command, the image acquisition frame rate is reduced, and the feature extraction process adopts an adapted low-power inference method. If the lightweight detection model supports an early exit mechanism, only the first half of the network layer is used to output coarse-grained detection results for high-confidence targets. If the model does not support an early exit structure, the input image resolution is reduced to 1 / 2 to 1 / 4 of the original resolution before being input into the complete model for inference, or a pre-stored ultra-lightweight detection model with fewer parameters is switched to output detection results for high-confidence targets, generating low-power detection result records. Reducing the input resolution can reduce the amount of multiplication and addition operations by more than 70%, with computational power consumption comparable to only using a portion of the network layers. This avoids the accuracy loss caused by modifying the model structure. This approach significantly reduces computational power consumption and heat generation while retaining basic anomaly detection capabilities, preventing complete missed detection of sudden dangerous events.
[0069] Example 6 This embodiment, based on the first embodiment described above, provides a detailed explanation of the defined low-power inventory trajectory maintenance mode, such as... Figure 5 As shown in the upper part.
[0070] When the remaining battery power falls below a preset low battery threshold, the device enters a low battery state. The preset low battery threshold is set based on the battery discharge curve. When the battery power drops below a certain percentage, such as 20% to 30%, the battery voltage decreases rapidly. If the device continues to operate at full load, the remaining working time may be insufficient for a short period, such as several tens of minutes. This battery power value is set as the low battery threshold to reserve sufficient power for data saving and shutdown preparation.
[0071] At this point, a trajectory freeze command is generated, locking all currently tracking continuous trajectories. In response to the trajectory freeze command, the device enters a stock trajectory maintenance mode, rather than completely stopping new target detection: on the one hand, the locked continuous tracking trajectories are updated using a Kalman filter-based method, a process involving only simple matrix operations, with computational consumption far lower than complete target detection; on the other hand, the confidence threshold for new target detection is increased (e.g., from the default 0.5 to 0.8), and temporary tracking trajectories are created only for new targets with a confidence level not lower than this higher threshold. Simultaneously, the preset duration of the temporary tracking trajectory is shortened to half that of the original continuous tracking trajectory (e.g., from 10 seconds to 5 seconds), minimizing the computational overhead of new target detection while ensuring no sudden high-risk targets are missed.
[0072] When the maintenance tracking trajectory or temporary tracking trajectory meets the corresponding duration judgment condition, the alarm interception process is triggered. Stabilized images for the corresponding time period are captured to obtain a time-series image sequence and structured alarm data. The structured alarm data is then written to the local storage queue to obtain the standby trigger condition. Subsequently, based on the current computing load, the image acquisition frame rate and feature extraction frequency are further reduced to 1 / 3 to 1 / 2 of the original frequency to extend the effective working time of the device in low-power conditions, ensuring that at least the alarm reporting for the currently active trajectory can be completed.
[0073] Example 7 In low-power mode, storage space may also be limited. This application reads the occupied capacity of the local storage queue and calculates the remaining storage parameter, which is the percentage of remaining capacity to the total capacity. When this parameter is lower than a preset capacity threshold, for example, when the remaining capacity is less than 10% of the total capacity, space optimization is initiated. The preset capacity threshold is set based on reserving a certain amount of space for writing new alarm data to avoid data loss due to space exhaustion; it is typically set to 5% to 15% of the total capacity.
[0074] Traverse all time-series image sequences in the local storage queue and extract the duration information of each segment. Sort the segments according to alarm level: first, sort by alarm level in ascending order (lower level first), and if the levels are the same, sort by duration in descending order (longer duration first). Select the segment ranked first and generate a segment identifier to be deleted.
[0075] Frame rate compression is performed on the segment: A subset of frames is uniformly extracted from the video, for example, discarding one frame every other frame. The retained frames are re-encoded into an independent I-frame sequence (keyframe sequence), independent of the P / B frame reference relationship of the discarded frames, to avoid video decoding failure, generating a compressed time-series image sequence. The amount of data after compression is significantly reduced. The compressed segment replaces the original segment, updating the storage margin parameter. The above process is repeated until the storage margin parameter is restored to above the preset capacity threshold.
[0076] This strategy prioritizes compressing low-level, long-duration segments, freeing up space while preserving high-value evidence to the maximum extent possible.
[0077] Example 8 After the network is restored, the category labels in the structured alarm data are read. When the category label belongs to a predefined specific hazard type, such as "personnel intrusion", "abnormal high temperature", or "gas leak", an immediate alarm indicator is generated.
[0078] The identifier is written into the retransmission task packet, generating a retransmission task packet with the identifier. In the retransmission queue, all task packets carrying the immediate alarm identifier are extracted first to form an emergency retransmission data stream.
[0079] Subsequently, the real-time value of the current network bandwidth is read (obtained by querying the system network interface), and an available bandwidth parameter (in bits per second) is generated. The sending rate is adjusted based on this parameter: if bandwidth is ample, for example, exceeding a certain threshold, the original rate is maintained; if bandwidth is limited, the sending rate is reduced to avoid congestion and packet loss. Rate adjustment can employ simple proportional control, such as setting the sending rate to a certain percentage of the available bandwidth, for example, 80%, leaving a margin.
[0080] After rate adaptation, a rate-adapted supplementary stream is generated and sent to the backend. This mechanism ensures that the most urgent data is delivered reliably and with priority.
[0081] Example 9 The backend administrator reviews the received alarms. If they are determined to be false alarms, they mark them in the backend system and generate an alarm handling receipt, which includes a handling conclusion field ("false alarm" or "confirmed"). After receiving the receipt, the device parses the handling conclusion field.
[0082] When the handling conclusion field returned by the backend indicates a false alarm, the trajectory matching coefficients of all frames in the corresponding continuous tracking trajectory are extracted and added to the false alarm sample pool of the same type. The false alarm sample pool is managed using a sliding window mechanism, retaining only the most recent 10-20 valid false alarm samples. Expired samples exceeding the window size are automatically removed to avoid early scene samples affecting the later adaptive effect.
[0083] The threshold update process is triggered only when the number of false alarm samples in the sample pool is ≥5. First, the average value of all trajectory matching coefficients in the sample set is calculated. 50% of the difference between the average value and the current preset threshold is used as the threshold correction amount to avoid significant threshold drift caused by a single false alarm or a small number of samples. After generating a new preset threshold using the threshold correction amount, upper and lower limits are set for the new threshold: its value must not exceed the range of 50% to 150% of the initial preset threshold. This range is taken from the continuous debugging statistics of the same model of equipment in multiple roadways. The conflict rate between thresholds exceeding this range and manual review conclusions increases significantly. Therefore, it serves as a hard constraint to prevent threshold drift failure and to prevent the threshold from being adjusted to a completely invalid range.
[0084] The false alarm rate on the device side is calculated as the ratio of the number of false alarms reported by the backend to the number of false alarms reported by the device side; the device side only maintains a counter. The false alarm rate is calculated by the backend periodically issuing detailed comparison tables; the device side does not automatically count false alarms. If the false alarm rate increases by more than 20% within 24 hours after a threshold adjustment, or if the false alarm rate exceeds 10% after three consecutive adjustments, the system automatically reverts to the previous valid threshold and triggers a manual review reminder. Finally, the verified new preset threshold is written to the device status set to replace the original threshold. During the next trajectory matching judgment, the new threshold is called to perform the comparison operation and generate an optimized continuous tracking trajectory. As the running time increases, this mechanism can continuously reduce the system's false alarm rate while avoiding degradation in the adaptive process.
[0085] Example 10 During vehicle movement, various factors may cause the camera's field of view to deviate from the roadway's reference direction. This application performs viewpoint deviation detection before performing target detection on the stabilized image. Edge feature enhancement is applied to the stabilized image, for example, using the Canny operator to extract linear contour features (mainly coal wall lines and track lines on both sides of the roadway). After extraction, segment length filtering is applied, retaining only segments with a length greater than 15% of the image width for subsequent contour vector weighted averaging, filtering out non-parallel short contour noise such as equipment, pipelines, and personnel. The direction vectors of these contour segments are statistically analyzed to generate contour feature vectors. This vector is a weighted average of the direction vectors of all contour line segments, with the weights being the line segment lengths.
[0086] Calculate the contour feature vector and the roadway reference trajectory vector The cosine of the included angle is used as the contour matching degree. :
[0087] This formula is derived from the cosine of the angle between vectors. The value ranges from -1 to 1. The closer it is to 1, the more consistent the contour direction is with the roadway direction.
[0088] when If the angle is below a preset contour threshold, such as 0.8, a field of view deviation is determined. The preset contour threshold is set based on the fact that, under normal direct view, the angle between the tunnel contour direction and the reference direction is usually within a small range, such as no more than ten degrees, corresponding to a cosine value of approximately 0.985 or higher. Considering certain measurement errors and local tunnel curvature, setting the threshold between 0.85 and 0.9 is more appropriate.
[0089] In response to the viewpoint shift signal, if the camera is equipped with a motorized pan-tilt unit, it triggers the pan-tilt unit to perform horizontal and / or vertical rotation correction. If the camera is installed with a fixed viewpoint, it adaptively crops the stabilized image based on the viewpoint shift signal, retaining the image area consistent with the roadway reference direction, or outputs viewpoint shift alarm information to the backend to prompt maintenance personnel to adjust the equipment installation posture. For fixed-viewpoint cameras, the size of the adaptively cropped area is determined based on the deviation angle between the contour feature vector and the roadway reference trajectory vector. The larger the deviation angle, the smaller the cropped area, ensuring that the effective monitoring area always conforms to the roadway extension direction. If a target bounding box with a confidence level of not less than 0.7 is detected outside the cropped area, it automatically reverts to full-image inference and reports an alarm, preventing the loss of high-risk targets due to cropping. After correction, a corrected stabilized image is obtained.
[0090] The corrected and stabilized image is re-input into the feature extraction and bounding box calibration process, outputting corrected bounding boxes and category labels. The corrected bounding boxes are then used to update the position information in the continuous tracking trajectory, generating a viewpoint calibration tracking trajectory. This integrated hardware and software closed-loop system ensures that the camera always monitors from the optimal viewpoint.
Claims
1. A method for target identification and anomaly alarm at the edge of a mobile analytical camera, characterized in that, The method includes: Upon power-on, the system reads the remaining battery power, chip temperature, pre-stored tunnel baseline trajectory vector, current real-time network bandwidth, and preset threshold for the currently effective trajectory matching coefficient. It then extracts the average brightness of historical image frames to obtain illumination intensity parameters and acquires real-time vibration data: if a vibration sensor is installed, it collects vehicle vibration waveform data; otherwise, it calculates the global motion vector by matching feature points on multiple consecutive frames of original images, separating the high-frequency jitter component as equivalent vibration waveform data. These parameters are then summarized to obtain the device status set. The original images are acquired and superimposed with the real-time vibration data to obtain an image sequence with vibration markers; Electronic image stabilization is performed on the image sequence to obtain a stabilized image; Feature extraction and bounding box labeling are performed on the stabilized image, and the bounding boxes and category labels are output. The trajectory matching coefficient is obtained by calculating the dot product of the displacement vector of the target box center point and the reference trajectory vector of the roadway; When the trajectory matching coefficient is greater than a preset threshold, the target bounding box and category label of multiple consecutive frames are associated to obtain a continuous tracking trajectory; Based on the remaining battery power, chip temperature, peak vibration waveform data, and light intensity parameters in the device status set, adjust the image acquisition frame rate and feature extraction frequency; When the duration of the continuous tracking trajectory exceeds a preset duration, a stable image of the corresponding time period is captured to obtain a time-series image sequence and structured alarm data; When the network is interrupted, the structured alarm data and time-series image sequence are sorted by location number and stored in a local storage queue; When the network is restored, data is extracted from the local storage queue according to the alarm level and sent to the backend.
2. The method according to claim 1, characterized in that, The step of obtaining a continuous tracking trajectory by associating the target bounding box and category label of multiple consecutive frames when the trajectory matching coefficient is greater than a preset threshold includes: Extract the displacement vector of the center point of the target box in two adjacent frames of the continuous tracking trajectory, calculate the second dot product of the displacement vector and the reference trajectory vector of the roadway, and generate the directional consistency coefficient of adjacent frames. When the directional consistency coefficient of the adjacent frames is lower than the preset directional threshold, it is determined that the target box has undergone a trajectory change, and a change marker is generated; The jump markers are superimposed on the continuous tracking trajectory to generate a tracking sequence with jump markers; When the number of consecutive occurrences of the jump marker in the tracking sequence with the jump marker exceeds a preset threshold, the association relationship of the target box is released, and the corresponding continuous tracking trajectory is terminated.
3. The method according to claim 2, characterized in that, The generation of the tracking sequence with jump markers includes: Using the tracking sequence with jump markers as input, the area change rate of the target boxes is statistically analyzed to generate area fluctuation parameters; The area fluctuation parameter is compared with a preset area fluctuation threshold; when the area fluctuation parameter is greater than the preset area fluctuation threshold, an area anomaly marker is generated. Perform a logical AND operation between the area anomaly marker and the jump marker to generate a composite misjudgment identifier; When the composite misjudgment identifier is valid, the peak value of the most recently acquired vibration waveform data is retrieved from the device status set to generate a vibration interference factor; The confidence level of the original category label is attenuated and corrected based on the vibration interference factor to generate a corrected single-frame recognition result; Based on the confidence level of the corrected single-frame recognition result, the target association operation or alarm trigger determination is re-executed to obtain the corrected tracking sequence with jump marker.
4. The method according to claim 3, characterized in that, The generation of the corrected single-frame recognition result includes: Read the chip temperature from the device status set. When the chip temperature is higher than a preset high temperature threshold, reduce the baseline value of the feature extraction frequency and generate a frequency reduction compensation command. In response to the frequency reduction compensation command, a frame interval identifier is inserted into the image sequence with vibration markers to generate a sparsely sampled image sequence; Electronic image stabilization is performed on the sparsely sampled image sequence to generate a frequency-reduced, image-stabilized image; The process of inputting feature extraction and target box calibration of the down-frequency stabilized image is used to output the down-frequency target box and category label as the corrected single-frame recognition result.
5. The method according to claim 4, characterized in that, After obtaining the down-concentration target bounding box and category label as the corrected single-frame recognition result, the method further includes: The continuous tracking trajectory is updated based on the corrected single-frame recognition result to obtain the temperature-adaptive tracking trajectory. Historical image frames corresponding to the temperature adaptation tracking trajectory are extracted, and the average brightness of the image frames is extracted to generate illumination intensity parameters. The environmental interference coefficient is calculated based on the light intensity parameter and the peak value of the vibration waveform data. When the environmental interference coefficient is greater than the preset interference threshold, the image acquisition frame rate is reduced, and target detection is performed with a computational complexity lower than that of full-resolution feature extraction: the input image is downsampled and features are extracted, or the sliding step size of the feature extraction operator is increased, and the detection result that meets the preset response threshold is output, and a low-power detection result record is generated. When the environmental interference coefficient falls below the interference threshold, full frame rate acquisition and complete feature extraction are resumed, the low-power detection result record is called to supplement the target box calibration, and the temperature adaptation tracking trajectory is updated.
6. The method according to claim 1, characterized in that, The step of adjusting the image acquisition frame rate and feature extraction frequency based on the remaining battery power, chip temperature, vibration waveform data peak value, and light intensity parameters in the device status set includes: The remaining battery power in the device status set and the illumination intensity parameters extracted from historical image frames are obtained. When the remaining battery power is lower than a preset low battery threshold, the current continuous tracking trajectory is locked, and a trajectory freeze command is generated. In response to the trajectory freeze command, a stock trajectory maintenance mode is generated; In the existing trajectory maintenance mode, the positions of existing target boxes in the continuous tracking trajectory are updated to increase the confidence threshold for new target detection, and temporary tracking trajectories are created for new targets that meet the high threshold. When the maintenance tracking trajectory or the temporary tracking trajectory meets the corresponding duration determination condition, the stable image of the corresponding time period is captured to obtain the time-series image sequence and structured alarm data, and the structured alarm data is written into the local storage queue to obtain the standby trigger condition; Based on the computational load under the existing trajectory maintenance mode and the standby triggering conditions, the image acquisition frame rate and the feature extraction frequency are reduced.
7. The method according to claim 6, characterized in that, After reducing the image acquisition frame rate and the feature extraction frequency, the method further includes: Read the occupied capacity of the local storage queue and generate storage remaining parameters; When the remaining capacity parameter is lower than the preset capacity threshold, the time-series image sequences in the local storage queue are traversed, the duration information of each segment is extracted, and a duration list is generated. The duration list is sorted according to the alarm level, and the longest time-series image sequence is selected to generate a segment identifier to be deleted. Based on the identifier of the segment to be deleted, the frame rate of the segment to be deleted is compressed to generate a compressed time-series image sequence; The compressed time-series image sequence is used to replace the original time-series image sequence, and the storage margin parameter of the local storage queue is updated to obtain alarm data with optimized storage space.
8. The method according to claim 1, characterized in that, When the network is restored, data is extracted from the local storage queue according to the alarm level and sent to the backend, including: Read the category labels from the structured alarm data, and generate an immediate alarm identifier when the category label is a specific hazard type; Write the real-time alarm identifier into the retransmission task package to generate a retransmission task package with the identifier; The identified retransmission task packets are extracted first to generate an emergency retransmission data stream; Read the real-time network bandwidth value from the device status set and generate available bandwidth parameters; Adjust the sending rate of the emergency retransmission data stream according to the available bandwidth parameters, generate a rate-adapted retransmission stream, and send the rate-adapted retransmission stream to the background.
9. The method according to claim 8, characterized in that, After sending the rate-adapted supplementary stream to the backend, the process further includes: Receive alarm handling receipts returned from the backend, parse the handling conclusion field in the receipts, and when the handling conclusion field indicates a false alarm, extract the trajectory matching coefficients of all frames in the corresponding continuous tracking trajectory to generate a false alarm sample set; The difference between the trajectory matching coefficients in the false alarm sample set and the current preset threshold is calculated to generate a threshold correction amount; The preset threshold is updated using the threshold correction amount to generate a new preset threshold; Write the new preset threshold into the device state set, replacing the original preset threshold; When calculating the dot product of the displacement vector of the target box center point and the reference trajectory vector of the roadway to obtain the trajectory matching coefficient in the next calculation, the new preset threshold is called to perform a comparison operation to generate an optimized continuous tracking trajectory.
10. The method according to claim 1, characterized in that, The step of performing feature extraction and bounding box labeling on the stabilized image, and outputting the bounding box and category label, includes: Edge feature enhancement is performed on the stabilized image, and linear contour features are extracted from the stabilized image to generate a contour feature vector; Calculate the similarity between the contour feature vector and the roadway reference trajectory vector to generate a contour matching score; When the contour matching degree is lower than a preset contour threshold, it is determined that the current field of view deviates from the roadway reference direction, and a viewpoint offset signal is generated; In response to the viewpoint offset signal, if the camera is equipped with a motorized pan-tilt unit, the pan-tilt unit is triggered to perform angle correction and generate a corrected and stabilized image; if the camera is installed with a fixed viewpoint, the stabilized image is adaptively cropped to retain the image area consistent with the roadway reference direction, or viewpoint offset alarm information is output to the backend. The corrected or cropped stabilized image is input into the feature extraction and bounding box labeling process, and the corrected bounding box and category label are output. The position information in the continuous tracking trajectory is updated using the corrected target box to generate a viewpoint calibration tracking trajectory.