Lightweight real-time target detection method and system for automobile data recorder

By using a lightweight target detection model to separate the foreground and background and analyze traffic scenes in dashcam video frames, the limitations of resource consumption and processing time in real-time target detection in existing technologies are solved. This enables efficient and accurate hazard identification and alarm in complex traffic environments, improving driving safety and system adaptability.

CN121767964AInactive Publication Date: 2026-03-31SHENZHEN MEITONG VIDEO TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-03-31
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing real-time target detection methods are difficult to achieve efficient processing on devices with limited resources. In particular, they lack accuracy and speed in complex traffic environments, are unable to capture potential dangers in a timely manner, and lack the ability to adapt to dynamic traffic environments, thus affecting driving safety.

Method used

A lightweight target detection model is used to separate the foreground and background regions of real-time video frames. Combined with key target information, traffic scene constraints are analyzed, motion change trends are calculated, and a buzzer alarm and visual prompts are triggered when dangerous conditions are detected.

Benefits of technology

It improves the real-time performance and accuracy of dashcams in complex traffic environments, ensures timely identification of hazards, enhances safety and information transmission efficiency during driving, and strengthens the scalability and adaptability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767964A_ABST
    Figure CN121767964A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle-mounted images, in particular to a lightweight real-time target detection method and system for an automobile data recorder. The method comprises the following steps: collecting a real-time video frame from an automobile data recorder camera, zooming the real-time video frame to a fixed size to obtain a standardized video frame, combining a preset lightweight target detection model, dividing the standardized video frame into a foreground region image and a background region image, and detecting key target information, key target dynamic data and real-time frames are displayed in an enhanced mode, and when a detection result reaches a dangerous condition threshold value, video buzzing alarm and visual prompt are triggered immediately. According to the invention, the analysis of the key target dynamic data is realized, the change trend in the traffic scene can be accurately captured, the information transmission efficiency is improved, and the reaction time assistance caused by information delay is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of vehicle imaging technology, and in particular to a lightweight real-time target detection method and system for dashcams. Background Technology

[0002] Existing real-time object detection methods often rely on complex deep learning models, making them difficult to implement on resource-constrained devices, especially in driving situations. The accuracy and speed of object detection directly impact driving safety. Traditional methods perform poorly in high-speed and complex environments, prone to false alarms and missed detections, affecting the judgment of the surrounding environment. Furthermore, the processing of standardized real-time video frames also faces threats to real-time performance, failing to adapt quickly to changing traffic conditions. Existing systems typically lack adaptability to dynamic traffic environments, particularly in high-density traffic or high-glucose topography, where object detection accuracy significantly decreases. Existing methods fail to effectively integrate multiple environmental factors, resulting in the inability to capture potential hazards in a timely manner at critical moments. The lack of comprehensive analysis of traffic scenarios affects the monitoring and judgment of key target dynamics. Overall, the shortcomings of existing technologies in terms of flexibility and real-time performance limit the practical application effectiveness of dashcams. Summary of the Invention

[0003] Therefore, it is necessary to provide a lightweight real-time target detection method and system for dashcams to solve at least one of the above-mentioned technical problems.

[0004] To achieve the above objectives, a lightweight real-time target detection method for dashcams includes the following steps: Step S1: Capture real-time video frames from the dashcam camera; scale the real-time video frames to a fixed input size to obtain standardized video frames; Step S2: Divide the standardized video frames into foreground region images and background region images using a preset lightweight target detection model; jointly detect key target information using the foreground region images and background region images; Step S3: Analyze traffic scenario constraints through key target information; calculate the motion change trend based on traffic scenario constraints and key target information to obtain dynamic evolution data of key targets; Step S4: Based on the dynamic evolution data of key targets and the real-time video frames of the dashcam, the system displays the data. When the detection result meets the danger threshold, a buzzer alarm and visual prompt are immediately triggered.

[0005] The present invention also provides a lightweight real-time target detection system for a dashcam, for performing the lightweight real-time target detection method for a dashcam as described above, the lightweight real-time target detection system for a dashcam comprising: The data processing module is used to acquire real-time video frames from the dashcam camera; and to scale the real-time video frames to a fixed input size to obtain standardized video frames. The joint detection module is used to distinguish standardized video frames into foreground region images and background region images using a preset lightweight target detection model; and to jointly detect key target information using the foreground region images and background region images. The trend evolution module is used to analyze traffic scenario constraints through key target information; and to calculate the motion change trend based on traffic scenario constraints and key target information to obtain dynamic evolution data of key targets. The alarm judgment module is used to overlay and display the dynamic evolution data of key targets with the real-time video frames of the dashcam. When the detection result meets the danger threshold, it immediately triggers a buzzer alarm and a visual prompt.

[0006] The beneficial effects of this invention are as follows: On the one hand, the standardized processing of real-time video frames ensures the high efficiency, consistency, and reliability of data input, effectively distinguishes between foreground and background areas, enhances the extraction of key target information, significantly reduces the demand for computing resources, ensures real-time processing under the limited hardware conditions of the dashcam, improves the system's response speed and operational stability, and, based on the analysis of dynamic data of key targets, can accurately capture the changing trends in traffic scenarios, monitor the behavior of key targets in real time, provide a scientific basis for traffic safety, and the acquisition of dynamic data further promotes the development of intelligent transportation systems and improves the focus and positioning level of dashcams in complex traffic environments.

[0007] On the other hand, it enables timely identification of dangerous conditions, ensuring that alarms and visual prompts are triggered when critical conditions are met, minimizing the occurrence of potential accidents and improving driving safety. The system's interconnectivity makes the monitoring and alarm mechanisms complement each other, enhancing alertness and safety protection during driving. The dynamic display function enables insight into traffic conditions, improves the efficiency of information transmission, reduces reaction time caused by information delays, and the real-time alarm enhances the practicality and reliability of the dashcam.

[0008] On the other hand, it not only improves the performance of dashcams in target detection and dynamic analysis, but also promotes the advancement of intelligent transportation technology, reduces the limitations of traditional target detection methods in terms of resource consumption and processing time, enhances the scalability and application flexibility of the system, adapts to the needs of different traffic environments, enhances the competitiveness of dashcams, improves efficiency and standardization, and promotes safer driving and more advanced traffic management. Attached Figure Description

[0009] Figure 1This is a flowchart illustrating the steps of a lightweight real-time target detection method for dashcams. Figure 2 The PR curve of the model's performance; Figure 3 This is a screenshot of the scene test results; The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0010] The technical method of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0011] Furthermore, the accompanying drawings are merely illustrative of the invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted. Some block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.

[0012] It should be understood that although the terms "first," "second," etc., may be used herein to describe various units, these units should not be limited by these terms. These terms are used merely to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, a first unit may be referred to as a second unit, and similarly, a second unit may be referred to as a first unit. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0013] To achieve the above objectives, please refer to Figures 1 to 3 A lightweight real-time target detection method for dashcams includes the following steps: Step S1: Capture real-time video frames from the dashcam camera; scale the real-time video frames to a fixed input size to obtain standardized video frames; In one implementation of this invention, video frame data is collected in real time from the front-facing camera of the dashcam. The collected video frames are preprocessed and scaled to a fixed input size, such as a uniformly scaled resolution of 640×480 or 480×272, to obtain standardized video frames.

[0014] In one implementation of this invention, scaling can be performed using bilinear interpolation during video frame preprocessing to minimize the loss of image details while reducing the size. Simultaneously, image normalization can be performed during scaling to map pixel values ​​to between 0 and 1.

[0015] Step S2: Divide the standardized video frames into foreground region images and background region images using a preset lightweight target detection model; jointly detect key target information using the foreground region images and background region images; In one implementation of this invention, a preset lightweight target detection model is used to divide a standardized video frame into a foreground region image and a background region image. By detecting the feature differences between the foreground region image and the background region image, the key target category and its location are identified to obtain key target information.

[0016] In one implementation of this invention, when dividing the foreground region and the background region, the moving subject in the foreground region is extracted, and the relatively static background region is identified using the background difference method. Then, these two types of data are sent to different branches of the lightweight detection model for processing. The foreground branch focuses on detecting moving targets such as vehicles, pedestrians and non-motorized vehicles, while the background branch focuses on identifying static targets such as traffic signs and road boundaries. The detection results are uniformly output as key target information through a fusion mechanism.

[0017] Step S3: Analyze traffic scenario constraints through key target information; calculate the motion change trend based on traffic scenario constraints and key target information to obtain dynamic evolution data of key targets; In one implementation of this invention, based on key target information, constraint analysis is performed on road boundaries, traffic signs and traffic lights in a traffic scene, and the temporal motion change trend of the target is calculated by combining the target's speed and direction, thereby obtaining dynamic evolution data of the key target.

[0018] In one implementation of this invention, when performing traffic scene constraint analysis, the spatial location of the key target is used to determine whether it is located within the drivable area of ​​the road, the traffic signal status is used to determine whether the target is in a controlled environment, and the information of traffic signs is combined to further limit the target's behavioral range. Finally, when predicting the motion change trend, a time series analysis method is used to fit the speed and direction data in consecutive frames to obtain the motion evolution trend of the key target in the future.

[0019] Step S4: Based on the dynamic evolution data of key targets and the real-time video frames of the dashcam, the system displays the data. When the detection result meets the danger threshold, a buzzer alarm and visual prompt are immediately triggered.

[0020] In one implementation of this invention, the dynamic evolution data of key targets is superimposed on the real-time video frames of the dashcam. When the detection result meets the danger threshold, such as when the distance to the vehicle in front is less than a set safety value, a buzzer alarm is immediately triggered or a visual prompt is displayed on the video frame.

[0021] In one implementation of this invention, when overlaying the data, the dynamic evolution data of the key target can be presented on the video frame by means of bounding boxes, highlighted trajectory lines or text labels. At the same time, a danger threshold judgment logic is set. For example, when the deceleration of the target vehicle is detected to exceed a preset threshold and the relative distance between the vehicle and the target vehicle is less than the safe braking distance, a buzzer alarm is triggered and an "emergency braking" prompt is marked on the video screen.

[0022] Preferably, step S2, which uses a preset lightweight target detection model to distinguish standardized video frames into foreground region images and background region images, includes: Featured target images in standardized video frames are detected using a pre-defined lightweight target detection model; The frame difference between the current frame and the previous frame in the feature target image is calculated, and the target motion information is captured by the frame difference calculation result; Dense optical flow is calculated based on the frame difference results, and the motion amplitude of each pixel is extracted. Based on the target motion information and the motion amplitude of each pixel, the feature target image is divided into a moving region image and a static environment region image. The moving region image is used as the foreground region image, and the static environment region image is used as the background region image.

[0023] In one implementation of this invention, standardized video frames are input into a lightweight target detection model to detect salient targets such as vehicles, pedestrians, non-motorized vehicles, and traffic signs, and generate corresponding feature target images.

[0024] In another embodiment of the present invention, multi-scale features are extracted using the improved CSPNet backbone during the detection process, and cross-scale feature fusion is performed using a lightweight PANet. Finally, the YOLOv3 multi-scale detection head outputs the bounding boxes and class labels of salient targets, and the feature target image is cropped according to the position of the detection box.

[0025] In another embodiment of the present invention, pixel-by-pixel difference operation is performed on the feature target images of the current frame and the previous frame to obtain a frame difference image. The rough motion information of the target is captured through the brightness change area of ​​the frame difference. Before the frame difference calculation, the feature target image is processed by grayscale and Gaussian filtering to reduce the interference of illumination changes and noise. The difference area is highlighted by binarization operation, thereby accurately capturing the approximate outline and displacement direction of the moving target.

[0026] In another embodiment of the present invention, the frame difference result is used as an initial constraint to perform dense optical flow calculation on two adjacent frames of feature target images, thereby obtaining the motion amplitude of each pixel.

[0027] In another embodiment of the present invention, the pixel displacement vector is estimated by multi-scale image pyramid during the optical flow estimation process, and the initial optical flow field is constrained by the frame difference result. Finally, the motion amplitude and direction of each pixel are output. For example, in the detected pedestrian salient target area, dense optical flow shows that the head and legs pixels have different amplitudes, of which the leg pixels have larger motion amplitudes, indicating that the pedestrian is walking.

[0028] In another embodiment of the present invention, the coarse motion region obtained from the frame difference result is superimposed and analyzed with the motion amplitude extracted by the dense optical flow. When the motion amplitude of a pixel exceeds a preset threshold, it is classified as a significant motion region; otherwise, it is classified as a static environment region.

[0029] In another embodiment of the present invention, an adaptive threshold segmentation method is used in the segmentation process. First, the median and standard deviation are calculated based on the overall motion amplitude distribution, and then the segmentation threshold is dynamically set. This enables stable differentiation between significant motion regions and static environment regions under different scenes and lighting conditions. The significant motion regions are output separately as foreground region images, while the static environment regions are output as background region images.

[0030] In another embodiment of the present invention, the foreground region and the background region are respectively marked with binary masks during output and superimposed on the original normalized video frame to achieve visual differentiation.

[0031] Preferably, the joint detection of key target information by the foreground region image and the background region image in step S2 includes: Identify the image gradient direction of the foreground region image; Multi-scale convolution is performed on the foreground region image data, and the multi-scale features of the foreground region image are enhanced by combining the image gradient direction to obtain foreground structural features; Identify background aberrations in background region images; Spatial coordinate information of background texture features is extracted separately to obtain foreground coordinate data and background coordinate data; Geometric registration is performed between foreground and background coordinate data, and time series correction is performed by combining timestamps to generate time series correction data; The differences in the temporal correction data are calculated pixel by pixel, and the local differences are enhanced by multi-scale convolutional mapping to generate difference mapping features; Key target information in video frames is standardized based on difference mapping feature mapping.

[0032] In one implementation of this invention, edge detection is performed on the foreground region image, and the gradient direction distribution of each pixel is calculated to describe the contour information of the foreground region.

[0033] In another embodiment of the present invention, the Sobel operator is used to calculate the gradients in the horizontal and vertical directions respectively in edge detection, and the gradient direction angle of each pixel is obtained by the arctangent function to generate a complete gradient direction image.

[0034] In another embodiment of the present invention, the foreground region image is input into a convolutional layer to extract multi-scale feature maps, and then combined with gradient direction information to enhance the response of significant edges to obtain foreground structural features.

[0035] In another embodiment of the present invention, different sizes of convolution kernels (e.g., 3×3, 5×5, 7×7) are used to extract local and global features during multi-scale convolution, and gradient direction weights are then superimposed in the feature channels to highlight key foreground information such as vehicle outlines and pedestrian edges.

[0036] In another embodiment of the present invention, texture analysis is performed on the background region image and abnormal perturbation patterns are detected therein to identify unstable background features. During texture analysis, features such as energy, contrast and entropy are extracted using the gray-level co-occurrence matrix, and abnormal fluctuation amplitude is detected by combining a time-series sliding window. Regions exceeding the threshold are marked as abnormal background perturbations.

[0037] In another embodiment of the present invention, the spatial coordinates of key points are extracted from foreground structural features and background abnormal disturbances, and saved as foreground coordinate data and background coordinate data respectively. The Harris corner detection method is used to extract high response points in the foreground and background images, and the pixel coordinates of these corner points are recorded as spatial reference data.

[0038] In another embodiment of the present invention, the coordinate data of the foreground and background are geometrically registered and corrected in combination with the timestamps of the video frames to generate time-series correction data. Affine transformation is used to align the coordinates of the foreground and background, and linear interpolation is used to correct the timestamp differences.

[0039] In another embodiment of the present invention, pixel-by-pixel difference is performed on the temporally corrected foreground and background data to obtain a difference feature map. Then, local differences are enhanced by multi-scale convolution to generate difference mapping features. After difference, 3×3 and 5×5 convolution kernels are applied to extract detailed differences, and a channel attention mechanism is combined to highlight the change area, thereby improving the detection effect of small targets.

[0040] In another embodiment of the present invention, the difference mapping features are fused with the standardized video frames to locate and label key target information, output the target category and location, remove overlapping boxes during the fusion process, and filter low-confidence targets by thresholding, thereby obtaining optimized key target detection results.

[0041] Preferably, identifying background anomalous perturbations in the background region image includes: The pixel gray levels of the background region image are statistically analyzed, and the gray level gradient distribution of the background region image is calculated based on the pixel gray levels. Detect the texture direction of the background region image based on the gray-level gradient distribution; Decompose the background texture features of the background region image by texture direction; Sparse coding of features is reconstructed based on background texture features; The background texture features and the reconstructed features are sparsely encoded and differentially compared to generate the background reconstruction error; The abnormal fluctuation amplitude of the background reconstruction error is calculated by combining the time series sliding window. Background abnormal disturbances are determined based on preset fluctuation amplitude thresholds and background texture features.

[0042] In one implementation of this invention, the image is converted into a grayscale image, each pixel records its grayscale value, and each pixel is traversed using image processing software or a programming language (such as Python's OpenCV library) to calculate the grayscale difference between adjacent pixels, thereby obtaining the gradient values ​​of the image in the horizontal and vertical directions. A grayscale gradient distribution map is then generated to reflect the strength and direction of grayscale changes in the image.

[0043] In another embodiment of the present invention, when calculating the gray-level gradient, the Sobel operator or the Prewitt operator can be used to perform convolution operations on the horizontal and vertical directions respectively to obtain the gradient magnitude matrix and the gradient direction matrix. The gradient magnitude matrix is ​​then normalized to the range of 0-255. Gaussian filtering can be performed on the image before calculation to reduce the interference of noise on the gradient calculation, thereby obtaining a smoother and more stable gradient distribution.

[0044] In another embodiment of the present invention, assuming there is a surveillance background image with a size of 640×480 pixels, after converting it to a grayscale image, the horizontal gradient is calculated using a 3×3 Sobel operator. and vertical gradient Then through the formula ; Calculate the gradient magnitude to obtain the gradient value distribution map for each pixel.

[0045] In another embodiment of the present invention, the gradient direction matrix obtained in step 1 is used to classify the gradient direction of each pixel into a certain angle interval (such as dividing 0°-180° into 18 intervals of 10°), and the number of pixels in each interval is counted to obtain the texture direction distribution of the image.

[0046] In another embodiment of the present invention, local texture direction statistics are performed on different regions, the image is divided into several small blocks (such as 32×32 pixels), and the pixel direction distribution in each block is statistically analyzed to obtain a local texture direction map, thereby enabling the detection of local abnormal direction disturbances.

[0047] In another embodiment of the present invention, when calculating the local texture direction, a weighting coefficient can be added. For example, the weight of pixels with large gradient magnitudes can be increased. For example, for a background image with a resolution of 640×480, it can be divided into 20×15 small blocks of 32×32 pixels. The histogram of the pixel gradient direction of each small block is calculated to obtain the main texture direction of each small block. If the main texture direction of most small blocks is consistent, the background texture direction is stable.

[0048] In another embodiment of the present invention, each pixel is filtered or convolved along its corresponding texture direction according to the texture direction information to extract texture features with consistent direction and form a complete background texture feature map. Gabor filters are used to perform multi-directional and multi-scale convolution on the image to extract texture features with different directions and frequencies, and these features are combined to form a background texture feature vector.

[0049] In another embodiment of the present invention, the background texture feature vector is sparsely represented by the learning dictionary matrix to generate a sparse encoding of the reconstructed features for each pixel or image block. The feature vector is normalized to ensure that each feature component is within the same numerical range. Then, the dictionary matrix and sparse coefficients are iteratively optimized to obtain the converged reconstructed sparse encoding.

[0050] In another embodiment of the present invention, a block-based processing method is adopted, dividing the image into several small blocks, each block is sparsely encoded separately, and finally the sparse encoding of the reconstruction features of all blocks is combined to form the reconstruction encoding of the entire image. For each pixel or image block, the original background texture features are subtracted from the sparsely encoded reconstructed features element by element to obtain a difference matrix. The value of the matrix is ​​the background reconstruction error. The absolute value or square value of the difference matrix of each pixel is calculated, and the entire error matrix is ​​normalized to obtain an error map in the range of 0-1, which more intuitively reflects the abnormal area.

[0051] In another embodiment of the present invention, the background reconstruction error maps of consecutive frames are arranged in chronological order, and a sliding window (e.g., 5 frames) is defined. The mean and variance of the error of each pixel are calculated within the window to obtain the pixel-level abnormal fluctuation amplitude. In order to reduce the impact of sudden noise, a weighted sliding window can be used to assign smaller weights to frames that are far from the center frame, and the weighted variance is calculated as an indicator of abnormal fluctuation amplitude.

[0052] In another embodiment of the present invention, global statistics are performed outside the window, and the maximum or average value of the fluctuation amplitude of the entire image is used as the overall abnormality index of the background. The abnormal fluctuation amplitude is compared with a preset threshold, and the direction and frequency information in the background texture features are combined to determine whether the pixel or region is an abnormal disturbance. An abnormality marker map is generated, and a threshold dynamic adjustment strategy is set, such as automatically adjusting the threshold based on the overall fluctuation average of the most recent sliding window.

[0053] In another embodiment of the present invention, by combining regional connectivity analysis, scattered abnormal pixels are aggregated into complete abnormal regions, thereby improving detection accuracy and outputting labeled contours. Assuming the threshold is set to 0.2, pixels exceeding the threshold in the abnormal fluctuation amplitude map are marked as red regions. By combining texture direction information, adjacent red pixels are connected into complete abnormal regions, and finally, a background abnormal perturbation map is output.

[0054] Preferably, determining background anomalies based on a preset fluctuation amplitude threshold and background texture features includes: Multiple local sub-regions are extracted from the background image, and the mean and standard deviation of grayscale values ​​of each sub-region are calculated to obtain local statistical distribution data. By comparing the local statistical distribution data with the global grayscale mean and global standard deviation of the background region image, the global comparison difference data is analyzed. Identify sub-regions with excessively large differences in the mean grayscale fluctuation in global comparative difference data, calculate the maximum fluctuation value of the sub-region, and identify candidate fluctuation amplitudes; The dynamic threshold reference is calculated by combining the candidate fluctuation amplitude with the texture direction consistency index of the background region image. The fluctuation amplitude threshold is set according to the dynamic threshold reference, and the background texture features are compared together to generate preliminary abnormal disturbance labeling data; Background anomalous disturbances were identified based on preliminary anomalous disturbance labeling data.

[0055] In one implementation of this invention, the input background region image is divided into several local sub-regions of fixed size (e.g., 32×32 pixels), and the pixel grayscale values ​​of each sub-region are statistically analyzed to calculate the grayscale mean and standard deviation, thereby forming local statistical distribution data for each sub-region.

[0056] In another embodiment of the present invention, when dividing the sub-regions, an overlapping division strategy is used, that is, adjacent sub-regions have some overlapping areas. Overlapping calculation can enhance the smoothness and continuity of local statistics, thereby reducing the noise impact of local anomalies. When calculating the gray mean for each sub-region, extreme pixels (such as the highest and lowest 5% of pixels) are removed first. For a 640×480 pixel background image, it is divided into 32×32 pixel sub-regions, forming a total of 300 sub-regions. The gray mean and standard deviation of each sub-region are calculated to obtain a 300×2 local statistical distribution table.

[0057] In another embodiment of the present invention, the global grayscale mean and global standard deviation of the entire background region image are calculated, and the local statistical distribution data of each sub-region are compared with the global mean and standard deviation to obtain a difference value matrix.

[0058] In another embodiment of the present invention, the Z-score method can be used to subtract the global mean from the mean of each sub-region and then divide by the global standard deviation to obtain a standardized difference value. The larger the value, the more significant the gray-scale fluctuation of the sub-region. At the same time, the gray-scale standard deviation is compared and analyzed to form a comprehensive difference index matrix, which combines the mean difference and the standard deviation difference.

[0059] In another embodiment of the present invention, assuming the global grayscale mean is 120 and the standard deviation is 15, and the grayscale mean of a certain sub-region is 160 and the standard deviation is 20, then the mean Z-score of the sub-region is (160-120) / 15≈2.67 and the standard deviation Z-score is (20-15) / 15≈0.33, and the difference in mean is significant.

[0060] In another embodiment of the present invention, sub-regions in the difference value matrix that are greater than a preset threshold are marked as sub-regions with significant gray-level fluctuations. The maximum difference value or maximum fluctuation value of the gray-level values ​​of these sub-regions is calculated as a candidate fluctuation amplitude. When calculating the maximum fluctuation value, the maximum and minimum values ​​of the pixel gray-level values ​​within each significant sub-region are calculated to obtain a more accurate fluctuation amplitude reference.

[0061] In another embodiment of the present invention, the candidate fluctuation amplitudes are sorted, and the top N% are selected as key areas of focus to improve the accuracy of subsequent dynamic threshold references. Based on the candidate fluctuation amplitudes, the texture direction consistency index (such as gradient direction variance) of each sub-region is extracted. The fluctuation amplitude and texture direction consistency are weighted and calculated to obtain the dynamic threshold reference.

[0062] In another embodiment of the present invention, the gradient direction matrix is ​​used to calculate the pixel direction variance within each sub-region. The smaller the variance, the higher the direction consistency. The fluctuation amplitude is multiplied by the direction consistency coefficient to form a dynamic threshold reference. The dynamic threshold reference is normalized and mapped to a preset threshold range, such as the range of 0.1 to 0.3, as the actual fluctuation threshold reference.

[0063] In another embodiment of the present invention, a dynamic threshold reference is used as the fluctuation amplitude threshold of the sub-region. The actual grayscale fluctuation of each sub-region is compared with the threshold. At the same time, combined with texture features (such as direction consistency and frequency features), preliminary abnormal disturbance labeling data is generated, and regions with texture abnormalities above the threshold are marked as potential anomalies.

[0064] In another embodiment of the present invention, weights are set for different sub-regions. For example, the weight of regions with low texture direction consistency is reduced. At the same time, multiple feature indicators can be combined into a comprehensive score to judge anomalies. Binarization processing is used to generate a 0-1 matrix from the preliminary anomaly perturbation marking data, where 1 represents a potential abnormal region and 0 represents a normal region.

[0065] In another embodiment of the present invention, spatial connectivity analysis is combined to determine the final background abnormal disturbance region, remove isolated noise points and form a complete abnormal region, and perform smoothing or dilation processing on the abnormal region. For consecutive frames, temporal consistency analysis can be performed. If a certain region is marked as abnormal in multiple frames, the final background abnormal disturbance is confirmed.

[0066] Preferably, step S3, which involves analyzing traffic scenario constraints using key target information, includes: Analyze the spatial location information of key targets in standardized video frames; Identify road boundary information in standardized video frames; By combining road boundary information and spatial location information, spatial constraints are matched to obtain road boundary constraints; Extract traffic information from standardized video frames; Traffic rule constraints are analyzed based on preset traffic rule data and traffic information. By integrating road boundary constraints and traffic rule constraints, traffic scenario constraints can be identified.

[0067] In one implementation of this invention, such as Figure 3 As shown, key traffic targets (such as vehicles, pedestrians, and bicycles) are identified, and the center coordinates, bounding box size, and orientation information of each target in the image are extracted to obtain the spatial location information of the key targets.

[0068] In another embodiment of the present invention, the bounding box coordinates of the key target are mapped to a unified coordinate system, and the same target in consecutive frames is tracked. For example, three cars are detected in a 640×480 pixel video image, and their bounding box center coordinates are extracted as (120,200), (300,250), and (500,400), respectively, and the bounding box width and height are (60,40), (70,50), and (80,60), respectively, and normalized to the 0-1 coordinate system.

[0069] In another embodiment of the present invention, edge detection or semantic segmentation is performed on standardized video frames to identify the boundary lines and lane line information of the road area and generate a road boundary map.

[0070] In another embodiment of the present invention, during the boundary recognition process, the video frame is mapped to a bird's-eye view by combining perspective transformation to more accurately determine the road boundary and lane position, and the detected road boundary line is fitted in multiple segments (such as straight line or polynomial fitting) to obtain a continuous road boundary curve.

[0071] In another embodiment of the present invention, it is determined whether the target is located inside or at the edge of the road boundary to form a road boundary constraint. The shortest distance from the center point of the target to the left and right road boundaries is calculated and a threshold is set. If the target exceeds the threshold, it is marked as violating the road boundary constraint. The data is then combined with continuous frame information for averaging or filtering.

[0072] In another embodiment of the present invention, traffic flow information, including vehicle speed, acceleration, relative distance and traffic signal status, is extracted from standardized video frames to generate a traffic information data table. The motion vectors of vehicles or pedestrians are calculated by combining optical flow method or inter-frame difference method, and the actual speed is converted by combining frame rate. The traffic information of multiple frames is statistically analyzed to calculate parameters such as average speed and vehicle distance distribution.

[0073] In another embodiment of the present invention, preset traffic rule data (such as speed limits, no parking, and lane driving rules) are used, combined with vehicle speed, position, and distance information, to make a traffic rule compliance judgment for each target, generate traffic rule constraint data, encode different traffic rules into logical judgment conditions, such as speed thresholds, lane position thresholds, and minimum safe distance thresholds, and compare each target in turn to mark whether the rules are violated.

[0074] In another embodiment of the present invention, the rule threshold is dynamically adjusted according to the time period or traffic signal status. For example, the speed threshold is set to 0 during the red light period and judged according to the speed limit rule during the green light period. For example, if a vehicle is traveling at 40 km / h and the speed limit is 30 km / h, the traffic rule constraint is marked as a violation, and another vehicle is in the lane and traveling at 25 km / h, which is marked as compliant.

[0075] In another embodiment of the present invention, road boundary constraint data and traffic rule constraint data are fused together. By combining spatial location, motion state and constraint markers, traffic scenario constraints of each key target are identified. Different weights are assigned to road boundary constraints and traffic rule constraints to form a comprehensive constraint score, and low-scoring targets are marked as potential anomalies.

[0076] Preferably, step S3, which calculates the motion change trend based on traffic scenario constraints and key target information, includes: Extract the velocity direction data of key targets in consecutive frames to construct the motion sequence of key targets; Fitting motion constraint trajectories based on traffic scenario constraints and key target motion sequences; Predict the target trajectory based on the key target motion sequence; Project the motion constraint trajectory onto the target motion trajectory and identify the contact points between the motion constraint trajectory and the target motion trajectory; Based on the analysis of contact-related points, the influence of traffic scenario constraints on the motion sequence of key targets is analyzed, and the trend of motion change is predicted based on the influence of constraints.

[0077] In one implementation of this invention, for each key target in a series of video frames, the center coordinates of the target in each frame are extracted. The velocity vector is obtained by calculating the difference between the center coordinates of adjacent frames and dividing it by the inter-frame time. The magnitude and direction of the velocity vectors are then arranged in chronological order to form a motion sequence of the key target.

[0078] In another embodiment of the present invention, Kalman filtering or weighted averaging is performed on the velocity vectors of consecutive frames to reduce the impact of target jitter or detection error on the motion sequence. The motion sequence is sampled or down-framed to reduce the amount of computation while maintaining the main motion change trend.

[0079] In another embodiment of the present invention, a motion trajectory constrained by the traffic scene is generated based on the motion sequence of the key target and combined with traffic scene constraint information (such as lane boundaries, traffic rules, and adjacent vehicle constraints), which is called a motion constraint trajectory.

[0080] In another embodiment of the present invention, constraints are set for the target position at each time point, such as not exceeding the lane boundary and not violating the minimum vehicle distance. The trajectory of consecutive frames is smoothed, such as by using weighted smoothing or Gaussian filtering, so that the fitted trajectory is smoother and conforms to the traffic scene constraints.

[0081] In another embodiment of the present invention, the motion trajectory of the key target is predicted in several future frames using the motion sequence of the key target, forming the target motion trajectory. In the prediction process, the historical speed change trend and acceleration change can be considered to obtain a motion trajectory that is more in line with reality. At the same time, random disturbances can be added to the prediction results to simulate the uncertainty of actual driving behavior. Different prediction models can be fused together, such as the weighted average of linear prediction and LSTM prediction results, to improve the prediction accuracy.

[0082] In another embodiment of the present invention, for example, if the velocity vector of the vehicle motion sequence is known to be (5, 5), (7, 8), (8, 7), (9, 8), the coordinates of the next frame are predicted to be (138, 236) using a linear prediction model, and then the coordinates of the next 3 frames are predicted to be (148, 244), (159, 252), and (171, 261), respectively, thereby generating a predicted motion trajectory.

[0083] In another embodiment of the present invention, the motion constraint trajectory and the target motion trajectory are projected and compared in time and space to calculate the closest point or intersection point between the two trajectories, identify the contact and association points of the trajectories, and mark the location of the key target affected by traffic scene constraints.

[0084] In another embodiment of the present invention, a threshold judgment is made on the distance between trajectory points. If the distance between trajectory points is less than the set threshold, it is considered that there is a contact association. Multiple contact points can be weighted to determine the most critical influencing point. The location of the contact association point is analyzed. The motion speed and direction near the contact point are compared with the constraint trajectory to determine the restriction or guidance effect of the constraint on the motion of the key target. The future motion prediction trajectory is adjusted according to the analysis results to generate a motion change trend prediction.

[0085] In another embodiment of the present invention, different influence weights are assigned to different types of constraints (such as lane boundaries, distance to the vehicle in front, and traffic signals), and the influence of the constraints on the target speed and direction is comprehensively calculated to form a weighted motion change trend. Combined with time series smoothing, the short-term abnormal influence is weakened, while the main trend guided by the constraints is preserved, so that the motion change trend is more in line with actual driving behavior.

[0086] Of particular importance is the analysis of the impact of traffic scenario constraints on the motion sequence of key targets based on contact-related point analysis, and the prediction of motion change trends based on the constraint impact, including: Extract the frame index position of the contact-associated point in the motion sequence of the key target to form a contact point index; Spatially overlay and map the contact point index data with traffic scenario constraints to identify the contact point scenario mapping data. Calculate the velocity and direction change rates of key targets based on scene mapping data of contact points; Predict dynamic changes in contact point data by measuring the rate of change of velocity and direction of key targets; By combining dynamic change data of contact points with traffic scenario constraints, the impact of contact point constraints can be identified. Extrapolating time series trends based on the influence of contact point constraints to predict local motion trends; The motion sequence of key targets is modified based on local motion trends to obtain the trend of motion changes.

[0087] In one implementation of this invention, based on the contact points between the motion constraint trajectory and the target motion trajectory, the frame number corresponding to each contact point in the key target motion sequence is extracted to form a contact point index list.

[0088] In another embodiment of the present invention, continuous contact points are merged, and contact points within several consecutive frames are regarded as a contact event, and an index list records the start frame and the end frame.

[0089] In another embodiment of the invention, the contact point index is sorted and deduplicated to ensure that each contact event appears only once, so as to accurately map the constraint effects.

[0090] In another embodiment of the present invention, the key target positions in the frame corresponding to the contact point index are superimposed on traffic scene constraint information (such as lane boundaries, front vehicle positions, and traffic signal positions) to generate scene mapping data corresponding to each contact point.

[0091] In another embodiment of the present invention, the contact point mapping data, including the coordinates of key targets, traffic constraint types and constraint strengths, is stored in the form of a structured table or matrix, and the spatial mapping results are visualized, for example, by marking the contact point locations and their constraint types on the frame image, and intuitively displaying the constraint-affected area.

[0092] In another embodiment of the present invention, the speed and direction of the key target before and after the contact frame are calculated using the scene mapping data of the contact point. The speed change rate is the difference between the current speed and the speed of the previous frame / time interval, and the direction change rate is the difference between the current motion direction and the direction of the previous frame / time interval.

[0093] In another embodiment of the invention, a sliding window is used to calculate the average rate of change for contact points in consecutive frames, smoothing instantaneous fluctuations and improving the stability of the rate of change data. Simultaneously, positive and negative changes are recorded to distinguish between acceleration / deceleration and left / right turns, providing detailed information for subsequent dynamic prediction. For example, in frame 4, the vehicle speed is (9,8), and in frame 5, the speed is (7,7), and the rate of change of speed is... , Rate of change of direction = .

[0094] In another embodiment of the present invention, the speed and direction changes of the contact point are predicted in the next few frames based on the speed and direction change rate, and dynamic change data of the contact point is generated. Constraints are added to the prediction results, such as not exceeding the lane boundary or the maximum acceleration limit, to ensure that the predicted dynamic change data conforms to the traffic scenario constraints. The prediction data is then processed by time smoothing to reduce the interference of short-term jitter on the motion trend.

[0095] In another embodiment of the present invention, the predicted dynamic change data of the contact point is compared with the traffic scene constraints to calculate whether the speed and direction are restricted by the constraints. If they are restricted, they are marked as constraint effects. The intensity of the constraint effects is quantified, such as speed reduction and direction offset, to form a constraint effect vector. Weights are assigned to different constraint types, and the influence of different constraints on the contact point is comprehensively calculated to form a constraint score.

[0096] In another embodiment of the present invention, the influence of contact point constraints is utilized, and combined with historical motion sequences, extrapolation is performed on several future frames to predict the motion trend of key targets within the local constraint area. Linear or nonlinear time series prediction is used to extrapolate the influence of constraints, thereby improving the accuracy of local trend prediction. The extrapolation amplitude is dynamically adjusted in combination with changes in constraint intensity, making the predicted motion trend more consistent with the actual scene.

[0097] In another embodiment of the present invention, the local motion trend is fused and corrected with the original key target motion sequence, the speed and direction of the corresponding frames in the original sequence are adjusted to generate the final motion change trend sequence, the corrected motion sequence is smoothed to eliminate short-term jitter, and at the same time ensures that the trend change is consistent with the constraint influence. The corrected trends of multiple contact points are superimposed to generate the overall key target motion change trend.

[0098] Preferably, step S4 includes the following steps: Step S41: Align the dynamic evolution data of key targets with the real-time video frames of the dashcam in terms of spatial and temporal coordinates to form superimposed fusion reference data; Step S42: Overlay the detection bounding boxes, category labels, and motion trend trajectories of key targets into the fused baseline data to generate enhanced video frames with semantic annotations; Step S43: Determine the threshold for dangerous conditions in the enhanced video frame. When the relative speed, relative distance, or motion trend of the key target exceeds the preset threshold, a dangerous event is marked, forming dangerous event detection data. Step S44: Convert the hazardous event detection data into trigger commands, and output a buzzer alarm signal in real time or overlay a visual prompt in the video frame.

[0099] In one implementation of this invention, the dynamic evolution data of key targets, including target position coordinates, motion trend sequence, speed and direction information, are aligned according to the timestamps of video frames. At the same time, the target coordinates are made consistent with the coordinate system of the image captured by the dashcam through perspective transformation or coordinate mapping to generate superimposed fusion reference data.

[0100] In another embodiment of the present invention, the timestamp is interpolated or the frame rate is matched. If the frame rate of the dynamic data of the key target is higher than the video frame rate, the high-frequency data is downsampled to ensure that each video frame corresponds to a set of dynamic data of the key target. During the spatial alignment process, calibration parameters or known reference points are used to convert the target coordinates from the world coordinate system to the image pixel coordinates to ensure that the key target is accurately superimposed in the video frame.

[0101] In another embodiment of the present invention, in each frame of fused reference data, the bounding box information of key targets is drawn on the video image, and the target category label (such as vehicle, pedestrian, bicycle) and motion trend trajectory are displayed at the same time to form an enhanced video frame with semantic annotation. Different colors or line styles are used to represent the speed or acceleration of the motion trend trajectory, which intuitively reflects the dynamic information of the target.

[0102] In another embodiment of the present invention, the size, font and position of the category label are set to ensure that the key target is not obscured. At the same time, a semi-transparent bounding box can be used to increase the visibility effect. The data of the key target in each enhanced video frame is analyzed. When the relative speed between the target and the vehicle's own position exceeds the set maximum speed threshold, or the relative distance is lower than the minimum safe distance threshold, or the movement trend shows a potential collision risk, it is marked as a dangerous event, and dangerous event detection data is generated.

[0103] In another embodiment of the present invention, a multi-frame continuous determination method is combined to smooth short-term fluctuations, avoid misjudgment caused by single-frame anomalies, and record the frame number, target category, location and speed information of the dangerous event. Different levels are set for different types of dangerous events, such as minor collision risk, moderate collision risk and high collision risk, and the level is marked in the data.

[0104] In another embodiment of the present invention, for example, if the distance between the vehicle and the vehicle in front is 5 pixels in frame 15, the preset minimum safe distance threshold is 10 pixels, and the speed vector (6, 6) exceeds the speed limit threshold of 5 pixels / frame, then a dangerous event is marked and the data {frame 15, vehicle, position (132, 236), speed (6, 6), danger level = high} is recorded.

[0105] In another embodiment of the present invention, based on the dangerous event detection data, the dangerous event is converted into a system trigger command and sent to the vehicle alarm module in real time, outputting a buzzer alarm or superimposing a red warning box and flashing indicator in the enhanced video frame to alert the driver of potential danger. The alarm intensity or visual prompting effect is set according to the danger level. For example, high-level danger uses continuous buzzing and flashing red box, medium-level danger uses intermittent buzzing and yellow border, and low-level danger uses prompt symbol.

[0106] The above description is merely a specific embodiment of the present invention, enabling those skilled in the art to understand or implement the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the present invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features of the invention herein.

Claims

1. A lightweight real-time target detection method for dashcams, characterized in that, Includes the following steps: Step S1: Capture real-time video frames from the dashcam camera; scale the real-time video frames to a fixed input size to obtain standardized video frames; Step S2: Divide the standardized video frames into foreground region images and background region images using a preset lightweight target detection model; jointly detect key target information using the foreground region images and background region images; Step S3: Analyze traffic scenario constraints through key target information; calculate the motion change trend based on traffic scenario constraints and key target information to obtain dynamic evolution data of key targets; Step S4: Based on the dynamic evolution data of key targets and the real-time video frames of the dashcam, the system displays the data. When the detection result meets the danger threshold, a buzzer alarm and visual prompt are immediately triggered.

2. The lightweight real-time target detection method for a dashcam according to claim 1, characterized in that, In step S2, the standardized video frames are divided into foreground region images and background region images using a preset lightweight object detection model, including: The target image in the standardized video frame is detected by a pre-set lightweight target detection model; The frame difference between the current frame and the previous frame in the feature target image is calculated, and the target motion information is captured by the frame difference calculation result; Dense optical flow is calculated based on the frame difference results, and the motion amplitude of each pixel is extracted. Based on the target motion information and the motion amplitude of each pixel, the feature target image is divided into a moving region image and a static environment region image. The moving region image is used as the foreground region image, and the static environment region image is used as the background region image.

3. The lightweight real-time target detection method for a dashcam according to claim 1, characterized in that, Step S2 involves jointly detecting key target information using foreground and background region images, including: Identify the image gradient direction of the foreground region image; Multi-scale convolution is performed on the foreground region image data, and the multi-scale features of the foreground region image are enhanced by combining the image gradient direction to obtain foreground structural features; Identify background aberrations in background region images; Spatial coordinate information of background texture features is extracted separately to obtain foreground coordinate data and background coordinate data; Geometric registration is performed between foreground and background coordinate data, and time series correction is performed by combining timestamps to generate time series correction data; The differences in the temporal correction data are calculated pixel by pixel, and the local differences are enhanced by multi-scale convolutional mapping to generate difference mapping features; Key target information in video frames is standardized based on difference mapping feature mapping.

4. The lightweight real-time target detection method for a dashcam according to claim 3, characterized in that, Identifying background anomalous perturbations in background region images includes: The pixel gray levels of the background region image are statistically analyzed, and the gray level gradient distribution of the background region image is calculated based on the pixel gray levels. Detect the texture direction of the background region image based on the gray-level gradient distribution; Decompose the background texture features of the background region image by texture direction; Sparse coding of features is reconstructed based on background texture features; The background texture features and the reconstructed features are sparsely encoded and differentially compared to generate the background reconstruction error; The abnormal fluctuation amplitude of the background reconstruction error is calculated by combining the time series sliding window. Background abnormal disturbances are determined based on preset fluctuation amplitude thresholds and background texture features.

5. The lightweight real-time target detection method for a dashcam according to claim 4, characterized in that, Background anomalies are determined based on preset fluctuation amplitude thresholds and background texture features, including: Multiple local sub-regions are extracted from the background image, and the mean and standard deviation of grayscale values ​​of each sub-region are calculated to obtain local statistical distribution data. By comparing the local statistical distribution data with the global grayscale mean and global standard deviation of the background region image, the global comparison difference data is analyzed. Identify sub-regions with excessively large differences in the mean grayscale fluctuation in global comparative difference data, calculate the maximum fluctuation value of the sub-region, and identify candidate fluctuation amplitudes; The dynamic threshold reference is calculated by combining the candidate fluctuation amplitude with the texture direction consistency index of the background region image. The fluctuation amplitude threshold is set according to the dynamic threshold reference, and the background texture features are compared together to generate preliminary abnormal disturbance labeling data; Background anomalous disturbances were identified based on preliminary anomalous disturbance labeling data.

6. The lightweight real-time target detection method for a dashcam according to claim 1, characterized in that, Step S3 involves analyzing traffic scenario constraints using key target information, including: Analyze the spatial location information of key targets in standardized video frames; Identify road boundary information in standardized video frames; By combining road boundary information and spatial location information, spatial constraints are matched to obtain road boundary constraints; Extract traffic information from standardized video frames; Traffic rule constraints are analyzed based on preset traffic rule data and traffic information. By integrating road boundary constraints and traffic rule constraints, traffic scenario constraints can be identified.

7. The lightweight real-time target detection method for a dashcam according to claim 1, characterized in that, Step S3, which calculates the motion change trend based on traffic scenario constraints and key target information, includes: Extract the velocity direction data of key targets in consecutive frames to construct the motion sequence of key targets; Fitting motion constraint trajectories based on traffic scenario constraints and key target motion sequences; Predict the target trajectory based on the key target motion sequence; Project the motion constraint trajectory onto the target motion trajectory and identify the contact points between the motion constraint trajectory and the target motion trajectory; Based on the analysis of contact-related points, the influence of traffic scenario constraints on the motion sequence of key targets is analyzed, and the trend of motion change is predicted based on the influence of constraints.

8. The lightweight real-time target detection method for a dashcam according to claim 7, characterized in that, Based on the analysis of contact-related points, the influence of traffic scenario constraints on the motion sequence of key targets is analyzed, and the trend of motion change is predicted based on the influence of constraints, including: Extract the frame index position of the contact-associated point in the motion sequence of the key target to form a contact point index; Spatially overlay and map the contact point index data with traffic scenario constraints to identify the contact point scenario mapping data. Calculate the velocity and direction change rates of key targets based on scene mapping data of contact points; Predict dynamic changes in contact point data by measuring the rate of change of velocity and direction of key targets; By combining dynamic change data of contact points with traffic scenario constraints, the impact of contact point constraints can be identified. Extrapolating time series trends based on the influence of contact point constraints to predict local motion trends; The motion sequence of key targets is modified based on local motion trends to obtain the trend of motion changes.

9. The lightweight real-time target detection method for a dashcam according to claim 1, characterized in that, Step S4 includes the following steps: Step S41: Align the dynamic evolution data of key targets with the real-time video frames of the dashcam in terms of spatial and temporal coordinates to form superimposed fusion reference data; Step S42: Overlay the detection bounding boxes, category labels, and motion trend trajectories of key targets into the fused baseline data to generate enhanced video frames with semantic annotations; Step S43: Determine the threshold for dangerous conditions in the enhanced video frame. When the relative speed, relative distance, or motion trend of the key target exceeds the preset threshold, a dangerous event is marked, forming dangerous event detection data. Step S44: Convert the hazardous event detection data into trigger commands, and output a buzzer alarm signal in real time or overlay a visual prompt in the video frame.

10. A lightweight real-time target detection system for a dashcam, characterized in that, For performing the lightweight real-time target detection method for a dashcam as described in claim 1, the lightweight real-time target detection system for a dashcam includes: The data processing module is used to acquire real-time video frames from the dashcam camera; and to scale the real-time video frames to a fixed input size to obtain standardized video frames. The joint detection module is used to distinguish standardized video frames into foreground region images and background region images using a preset lightweight target detection model; and to jointly detect key target information using the foreground region images and background region images. The trend evolution module is used to analyze traffic scenario constraints through key target information; and to calculate the motion change trend based on traffic scenario constraints and key target information to obtain dynamic evolution data of key targets. The alarm judgment module is used to overlay and display the dynamic evolution data of key targets with the real-time video frames of the dashcam. When the detection result meets the danger threshold, it immediately triggers a buzzer alarm and a visual prompt.

Citation Information

Cited By

  • Intelligent video acquisition and processing method for mine monitoring

    CN121982619A