Driving video recording method and system based on four-way monitoring
By adopting four-way monitoring technology in the driving monitoring system, synchronously collecting the vehicle's surrounding environment and driver's head attitude data, and combining the head attitude change characteristics for video grading and view synthesis, the problem that the existing system cannot fully cover and associate the driver's attention distribution, and efficient and complete video recording and dangerous scene recognition are achieved.
Patent Information
- Application Number
- CN202510210092.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing driving monitoring system cannot achieve full coverage of the vehicle's surrounding environment, and cannot effectively correlate the driver's visual attention distribution and environmental monitoring data, resulting in the possible loss of data in key scenarios.
The driving video recording method based on four-way monitoring is adopted, and the surrounding environment of the vehicle and the driver's head posture are synchronized through real-time video streams, and video grading processing and view synthesis are carried out in combination with the head posture change characteristics to achieve enhanced mapping of video content in areas not being watched by the driver and scene threat identification.
It achieves comprehensive coverage of the vehicle's surrounding environment, improves the integrity of video recording, can accurately identify and record potential dangerous scenarios, and ensures the security and reliability of critical data through multi-copy storage strategies.
Smart Images

Figure CN120147987A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image recognition and storage, and particularly to a driving video recording method and system based on four-way monitoring. Background Art
[0002] In existing driving monitoring systems, single or dual cameras are usually used to record the vehicle's surrounding environment. These systems mainly focus on the traffic conditions in front of and behind the vehicle, and record key information during driving through video recording. Some advanced driver assistance systems also introduce driver state monitoring functions, tracking features such as the driver's facial expressions and eye movements through in-vehicle cameras to determine fatigue driving or distracted attention.
[0003] However, the existing technology has obvious limitations. Firstly, single or dual cameras cannot achieve full coverage of the vehicle's surrounding environment, and visual blind spots are likely to occur. Secondly, although the attention state of the driver can be detected, the visual attention distribution of the driver cannot be effectively associated with the environmental monitoring data, resulting in the inability to perform differential processing according to actual safety requirements when storing video data. Especially when the driver's attention is distracted, key detail information in the corresponding period of video data may be lost due to the unified compression and storage strategy. Summary of the Invention
[0004] This application provides a driving video recording method and system based on four-way monitoring, which is used to solve the technical problem that key scenario data may be lost during the real-time recording of multiple videos due to the lack of consideration of the driver's attention distribution.
[0005] In a first aspect, the present application provides a driving video recording method based on four-way monitoring. The driving video recording method based on four-way monitoring includes: synchronously collecting the vehicle surrounding environment and the driver's head posture through a real-time video stream, performing time-corresponding processing on the video stream and the head posture data to obtain driving attention reference video data; performing video grading processing according to the head posture change characteristics in the driving attention reference video data, compressing and storing the video data in the driver's line of sight blind area to obtain attention-oriented graded video data; performing view synthesis processing on the attention-oriented graded video data, enhancing and mapping the video content in the area not gazed by the driver to obtain an attention-compensated panoramic image; using the attention-compensated panoramic image to perform scene threat recognition, prioritizing abnormal situations in the line of sight blind area to obtain scene threat level data; constructing an index for the video content based on the scene threat level data, highlighting the scenes when the driver's attention is distracted to obtain an attention-related retrieval structure; storing and classifying the data in the attention-related retrieval structure according to the driver's attention coverage, and protecting and storing the video data in the high-threat non-gazed area to obtain a hierarchically stored video data set.
[0006] In a second aspect, the present application provides a driving video recording system based on four-way monitoring. The driving video recording system based on four-way monitoring includes:
[0007] An acquisition module, configured to synchronously collect the vehicle surrounding environment and the driver's head posture through a real-time video stream, perform time-corresponding processing on the video stream and the head posture data to obtain driving attention reference video data;
[0008] A grading module, configured to perform video grading processing according to the head posture change characteristics in the driving attention reference video data, compress and store the video data in the driver's line of sight blind area to obtain attention-oriented graded video data;
[0009] A synthesis module, configured to perform view synthesis processing on the attention-oriented graded video data, enhance and map the video content in the area not gazed by the driver to obtain an attention-compensated panoramic image;
[0010] An identification module, configured to use the attention-compensated panoramic image to perform scene threat recognition, prioritize abnormal situations in the line of sight blind area to obtain scene threat level data;
[0011] A construction module, configured to construct an index for the video content based on the scene threat level data, highlight the scenes when the driver's attention is distracted to obtain an attention-related retrieval structure;
[0012] A storage module, configured to store and classify the data in the retrieval structure associated with the attention according to the degree of driver attention coverage, and protect and store the video data in the high-threat un-gazed area, so as to obtain a hierarchically stored video data set.
[0013] In the technical solution provided by this application, by establishing a linkage mechanism between driving attention and environmental monitoring, the intelligent processing and storage of video data are realized. Through the synchronous acquisition and time alignment processing of the vehicle's surrounding environment and the driver's head posture, the corresponding relationship between the driver's attention distribution and environmental changes is accurately captured; based on the importance analysis and compression storage strategy of the head rotation characteristics, the differential processing of video data is realized, so that the storage resources are more reasonably allocated; the visual enhancement processing is performed on the un-gazed area, effectively compensating for the blind area of the driver's visual attention and improving the integrity of the video recording; through the threat recognition and prioritization of the attention-compensated panoramic image, a scene threat assessment mechanism based on attention characteristics is established, enabling the system to accurately identify and record potential dangerous scenes; the retrieval marking mechanism associated with attention enables high-risk scenes to be quickly located and played back, facilitating post-event analysis and summary; the multi-copy storage strategy for high-threat un-gazed areas ensures the security and reliability of key data. At the algorithm level, this solution innovatively combines the deep learning object detection algorithm with the attention feature extraction algorithm. By real-time analyzing the driver's attention pattern, the video processing strategy is dynamically adjusted, enabling the system to more intelligently allocate computing and storage resources in the face of complex road environments. At the same time, based on the time-series associated data processing framework, the system can achieve the efficient processing and management of large-scale video data while maintaining a low computational complexity, providing a feasible technical solution for applications in actual road scenarios and improving the pertinence and effectiveness of driving records. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0015] Figure 1 It is a schematic diagram of an embodiment of the driving video recording method based on four-way monitoring in the embodiments of this application;
[0016] Figure 2 It is a schematic diagram of the hierarchical mapping relationship of video regions in the embodiments of this application;
[0017] Figure 3 It is a schematic diagram of the scene threat priority ranking in the embodiments of this application;
[0018] Figure 4 This is a schematic diagram of an embodiment of a driving video recording system based on four-way monitoring in an embodiment of the present application. Detailed implementation manners
[0019] The embodiments of the present application provide a driving video recording method and system based on four-way monitoring. Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order different from that shown or described here. In addition, the term "comprising" or "having" and any variation thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0020] For ease of understanding, the specific process of the embodiments of the present application will be described below. Please refer to Figure 1 An embodiment of the driving video recording method based on four-way monitoring in the embodiments of the present application includes:
[0021] Step S101: Synchronously collect the vehicle surrounding environment and the driver's head posture through the real-time video stream, perform time-corresponding processing on the video stream and the head posture data to obtain the driving attention reference video data;
[0022] Step S102: Perform video grading processing according to the head posture change characteristics in the driving attention reference video data, compress and store the video data in the driver's line-of-sight blind area to obtain the attention-oriented graded video data;
[0023] Step S103: Perform view synthesis processing on the attention-oriented graded video data, and perform enhanced mapping on the video content in the area not gazed by the driver to obtain the attention-compensated panoramic image;
[0024] Step S104: Use the attention-compensated panoramic image to perform scene threat recognition, prioritize the abnormal situations in the line-of-sight blind area to obtain the scene threat level data;
[0025] Step S105: Build an index for the video content based on the scene threat level data, and mark the scenes when the driver's attention is distracted as key points to obtain the attention-associated retrieval structure;
[0026] Step S106: Store and classify the data in the retrieval structure associated with attention according to the degree of driver attention coverage, and protect and store the video data in the high-threat un-gazed area to obtain a hierarchically stored video data set.
[0027] It can be understood that the execution entity of this application can be a driving video recording system based on four-way monitoring, or it can also be a terminal or a server, and specific limitations are not made here. In this embodiment of the application, the server is used as the execution entity for illustration.
[0028] Specifically, four high-definition cameras are respectively installed at the front, rear, left, and right positions of the vehicle. Each camera collects a video stream with a resolution of 1080P, and at the same time collects the head pose of the driver. The head pose collection quantifies the rotation angle and rotation speed of the driver's head, establishes a corresponding relationship between the head pose parameters and the video timestamp, and generates a time-series correlation matrix. Through the time-series correlation matrix, a mapping relationship between the video data and the attention data is established to form the driving attention reference video data. In the video grading process, the head pose change characteristics in the driving attention reference video data are analyzed, the residence time of the driver's fixation point in different regions is calculated, and the position of the blind spot of sight is determined. For the area that the driver continuously does not gaze at, the system determines it as the blind spot of sight, and compresses and stores the video data in this area. The compression method uses differential coding to encode and store the difference information between consecutive frames, thereby obtaining the attention-guided hierarchical video data.
[0029] Process the attention-guided hierarchical video data, splice the four-way video images spatially to form a 360-degree panoramic image. For the video content in the area not gazed at by the driver, perform contrast enhancement and detail enhancement processing to improve the visual saliency of these areas and generate an attention-compensated panoramic image. In the scene threat recognition stage, the system focuses on analyzing the blind spot of sight in the attention-compensated panoramic image to identify potential dangerous targets. The dangerous targets include situations such as a vehicle approaching rapidly, a pedestrian suddenly appearing, and a sudden change in road conditions. The system evaluates the threat level of these abnormal situations according to factors such as the movement speed, approaching distance, and collision risk of the target, and generates scene threat level data.
[0030] According to the scene threat level data, index and mark the video clips. The system records the scene characteristics when the driver's attention is distracted, including information such as the duration, degree, and line-of-sight deviation direction of attention distraction, establishes a multi-dimensional index structure, and forms a retrieval structure associated with attention. Finally, in the storage classification stage, the system hierarchically stores the video data according to the degree of driver attention coverage based on the data in the retrieval structure associated with attention. For the area with high threat and not gazed at by the driver, a multi-copy redundancy storage strategy is adopted to ensure the reliability of the data, and finally a hierarchically stored video data set is generated.
[0031] For example, when a vehicle is driving on an urban road, the system collects four-channel video data in real time. When the driver is looking ahead at the road, the head pose data shows that their line of sight is mainly concentrated within a 120-degree range in front. At this time, an electric bicycle approaching rapidly appears in the left blind spot. The system immediately enhances the video of this area and marks this scene as a high-threat level. The system records the complete scene information at this moment, including data such as the driver's line of sight direction, the movement trajectory of the electric bicycle, and the relative speed. These data are given the highest storage protection level to ensure traceability and analysis afterwards.
[0032] In the embodiments of the present application, by establishing a linkage mechanism between driving attention and environmental monitoring, the intelligent processing and storage of video data are realized. Through the synchronous acquisition and time alignment processing of the vehicle's surrounding environment and the driver's head pose, the corresponding relationship between the driver's attention distribution and environmental changes is accurately captured; based on the importance analysis and compressed storage strategy of head rotation features, the differential processing of video data is realized, and the storage resources are more reasonably allocated; the visual enhancement processing of the un-gazed area effectively compensates for the blind spot of the driver's visual attention and improves the integrity of video recording; through the threat recognition and prioritization of the attention-compensated panoramic image, a scene threat assessment mechanism based on attention features is established, enabling the system to accurately identify and record potential dangerous scenes; the attention-associated retrieval marking mechanism enables high-risk scenes to be quickly located and played back for easy post-event analysis and summary; the multi-copy storage strategy for high-threat un-gazed areas ensures the security and reliability of key data. At the algorithm level, this solution innovatively combines the deep learning object detection algorithm with the attention feature extraction algorithm. By real-time analyzing the driver's attention pattern, the video processing strategy is dynamically adjusted, enabling the system to more intelligently allocate computing and storage resources in the face of complex road environments. At the same time, based on the time-series associated data processing framework, the system can achieve the efficient processing and management of large-scale video data while maintaining a low computational complexity, providing a feasible technical solution for applications in actual road scenarios and improving the pertinence and effectiveness of driving records.
[0033] In a specific embodiment, the process of executing step S101 may specifically include the following steps:
[0034] (1) Collect the environmental video stream data of the front, rear, left, and right of the vehicle through four cameras, synchronously integrate the environmental video stream data according to the time stamp, and obtain multi-view environmental video data;
[0035] (2) Continuously sample the changes in the driver's head pose, convert the head rotation angle, rotation speed, and gaze duration into attention feature parameters, and obtain head pose feature data;
[0036] (3) Establish a corresponding relationship between the multi-view environmental video data and the head pose feature data according to the time stamps to form a time-series correlation matrix, and obtain video-attention mapping data;
[0037] (4) Perform spatial partitioning on the attention coverage area in the video-attention mapping data, divide the video screen into a fixation area and a blind area, and obtain spatial attention distribution data;
[0038] (5) Calculate the attention weights for the video content according to the spatial attention distribution data, and assign different data processing priorities to the fixation area and the blind area to obtain the attention weight allocation result;
[0039] (6) Perform spatio-temporal fusion on the attention weight allocation result and the video data, and mark the video data according to the attention distribution characteristics to obtain the driving attention reference video data.
[0040] Specifically, the four-way cameras collect video stream data of the vehicle's front, rear, left, and right environments. Each camera collects a video stream with a resolution of 1080P (1920x1080 pixels) and a frame rate of 30fps. Each camera generates a time stamp when collecting the video, accurate to the millisecond level. The video stream data is synchronized and integrated according to the time stamps, and the frames with a time stamp difference of no more than 10ms in the four video streams are combined to generate multi-view environmental video data.
[0041] The synchronous integration of the video adopts the time window method. A time window of 10ms is set, and the video frames within this window are considered to be synchronized. For frames not within the same time window, linear interpolation is used to generate intermediate frames to ensure that the four-way videos are completely aligned in time. Each frame in the multi-view environmental video data contains image data of four views and corresponding time stamp information. The collection of the driver's head pose data is completed by an in-vehicle camera, and the sampling frequency is 60Hz. The head pose change data contains three basic parameters: the head rotation angle α (the rotation angle in the horizontal plane, positive to the left, negative to the right, and the value range is -90° to 90°), the rotation angular velocity ω (unit: degrees / second), and the fixation duration t (unit: seconds). The expression of the attention feature parameter A is as follows:
[0042]
[0043] where: k 1 = 0.4, is the angle weight coefficient; k 2 = 0.3, is the angular velocity weight coefficient; k 3 = 0.3, is the duration weight coefficient; ω max = 90° / s, is the maximum angular velocity threshold, t ref = 2s, is the reference fixation time.
[0044] The establishment of the time - sequence correlation matrix adopts the following form:
[0045] M[i][j] = [t s ,α,ω,t,A,V 1 ,V 2 ,V 3 ,V 4
[0046] Where: t s is the timestamp; V 1 ,V 2 ,V 3 ,V 4 are the feature vectors of four - channel videos respectively.
[0047] Each video feature vector contains the position coordinates (x, y) and the eigenvalue f:
[0048] f = [B,D,L,V]
[0049] Where: B is the brightness value (0 - 255); D is the target detection result (0 / 1); L is the target distance (meters); V is the movement speed (meters per second).
[0050] The spatial division of the attention coverage area is based on the head rotation angle and the attention feature parameters. The field of view is divided into a fixation area (a 120 - degree fan - shaped area centered on the current head orientation) and a blind area (the remaining area). Each area has a spatial coordinate range and a corresponding attention weight value.
[0051] The calculation of the attention weight adopts cosine - function weighting:
[0052] W(θ)=A·cos(θα)
[0053] Where: θ is the azimuth angle of a certain point in the field of view.
[0054] For example: When the driver needs to turn left during driving, the four - channel cameras continuously collect environmental data. Assume that at the moment t = 1000ms, the driver's head starts to turn left, the angle gradually increases from 0° to 45°, the angular velocity is 15° / s, and the fixation on the left side lasts for 2 seconds. At this time:
[0055] Original data acquisition: The four - channel video streams generate frame data with timestamps, and the head pose parameters: α = 45°, ω = 15° / s, t = 2s.
[0056] Attention feature calculation:
[0057] The time - sequence correlation matrix generates a record:
[0058] M[i] = [1000, 45, 15, 2, 0.75, (x 1 , y 1 , f 1 ), (x 2 , y 2 , f 2 ), (x 3 , y 3 , f 3 ), (x 4 , y 4 , f 4 )]
[0059] Spatial division: Gaze area: [-15°, 105°], blind spot area: [-90°, -15°] and [105°, 90°].
[0060] Weight calculation: W left = 0.75·cos(0) = 0.75, W right = 0.75·cos(180) = -0.75; Finally, mark the weight information in the video data to form the driving attention benchmark video data.
[0061] In a specific embodiment, the process of executing step S102 may specifically include the following steps:
[0062] (1) Extract the head pose change features from the driving attention benchmark video data, mark the importance of the video area according to the head rotation frequency and gaze duration, and obtain the video area importance data;
[0063] (2) Establish a correspondence between the video area importance data and the video frame data, divide the video area into levels according to the driver's gaze frequency, and obtain the video area hierarchical mapping relationship;
[0064] (3) Locate the driver's line of sight blind area according to the video area hierarchical mapping relationship, mark the continuously un-gazed video area as the line of sight blind area, and obtain the line of sight blind area location data;
[0065] (4) Calculate the compression ratio of the video content in the line of sight blind area location data, determine the compression priority according to the blind area duration, and obtain the blind area compression strategy data;
[0066] (5) Associate the blind area compression strategy data with the video data, and perform block compression on the video data in the line of sight blind area according to the compression priority to obtain the compressed area data;
[0067] (6) Integrate the compressed area data and the non-compressed area data, and organize the data according to the video area importance to obtain the attention-oriented hierarchical video data.
[0068] Specifically, extracting the head pose change features from the driving attention benchmark video data involves three basic parameters: the head rotation angle, rotation speed, and continuous fixation time. By combining and analyzing these parameters, the importance weight of the video region is calculated.
[0069] The calculation formula for the importance of the video region is as follows:
[0070]
[0071] Where: R i is the region importance index; H f is the head rotation frequency; H max is the maximum head rotation frequency threshold; T d is the fixation duration; T max is the maximum fixation time threshold; V c is the degree of visual change; V max is the maximum visual change threshold; β 1 , β 2 , β 3 is the weight coefficient.
[0072] The compression ratio of the line-of-sight blind area is calculated using the following formula:
[0073]
[0074] Where: C r is the compression ratio; B t is the blind area duration; B max is the maximum blind area time threshold; S d is the blind area spatial scale; S max is the maximum spatial scale threshold; F c is the frame content complexity; F max is the maximum complexity threshold; δ 1 , δ 2 , δ 3 is the weight coefficient. The comprehensive score of the blind area compression strategy is as follows:
[0075]
[0076] Where: P s is the compression strategy score; U d is the data update rate; U max is the maximum update rate threshold; L c is the computational load; L max is the maximum load threshold; K t is the key frame retention rate; K max is the maximum retention rate threshold; σ 1 , σ 2, σ 3 is the weight coefficient.
[0077] By processing the driving attention benchmark video data, the blind spot data of the line of sight is located and compressed. The specific implementation process starts from the extraction of the head pose change features, marks the importance of the video stream data, establishes the mapping relationship between the fixation area and the blind area, and finally completes the compressed storage. The driving attention benchmark video data contains the head rotation frequency data, marked as H f , with the unit of times per second, recording the number of times the driver's head rotates within a unit time; the fixation duration data T d , with the unit of seconds, recording the continuous fixation time of the driver on a certain area; the visual change degree V c , which is used to quantify the degree of drastic change of the scene within the field of view. These data are associated through timestamps to form a continuous time-series data stream.
[0078] The regional importance index R calculated according to the formula i will be mapped to a value between 0 and 1 and divided into three levels: high, medium, and low according to the threshold. When R i > 0.8, it is marked as a high-importance area; 0.5 ≤ R i ≤ 0.8 is the medium-importance area; R i < 0.5 is the low-importance area. This classification directly determines the subsequent compression strategy. As Figure 2 shown, it is a schematic diagram of the video area classification mapping relationship in the embodiment of the present application, including four processing levels. Among them: the first layer shows the input of the driving attention benchmark video data and the process of extracting the head pose features; the second layer shows the calculation and classification process of the regional importance, and the area is divided into three levels: high importance (R> 0.8), medium importance (0.5 ≤ R ≤ 0.8) and low importance (R <0.5) according to the importance index R; the third layer represents the mapping relationship between different importance areas and the actual space areas, the high-importance area corresponds to the front and side areas, the medium-importance area corresponds to the side-rear area, and the low-importance area corresponds to the rear area; the fourth layer is the final storage strategy, and different compression methods are adopted for different areas, including lossless compression storage, low-loss compression storage and high compression storage. The arrows in the figure represent the flow of data processing and the logical relationship between each level.
[0079] For the determination of the blind spot of the line of sight, the dynamic time window method is adopted. The reference observation window is set to 2 seconds. Within this time window, if the fixation time of a certain area is less than 20% of the total window duration, then this area is marked as a blind spot. The blind spot data will record the duration B t , the spatial range S d and the content complexity F c and other characteristic parameters.
[0080] Compression ratio C r The calculation of takes into account three dimensions: the blind spot duration, the spatial scale, and the frame content complexity. Among them, the frame content complexity F c is obtained by calculating the inter-frame difference, which reflects the degree of change of the video content. The compression ratio is used to determine the encoding parameters, including the adjustment range of the quantization parameter QP value.
[0081] Compression strategy score P s As the final decision-making basis, it combines the data update rate U d (reflecting the scene change frequency), the computational load L c (indicating the processor occupancy), and the key frame retention rate K t (controlling the video quality benchmark). The scoring result determines the processing method for the video segment. High-score segments use lossless compression, while low-score segments use lossy compression.
[0082] In practical applications, taking a left turn scenario as an example: The driver's head starts to turn left from the straight-ahead position, and the turning frequency H f is 2 times per second, and the fixation duration T d in the left area reaches 1.5 seconds. At the same time, the degree of visual change V c in this area is large (due to the rapid change of the scene during the left turn). Substituting these parameters into the formula, the importance index R i of the left area is calculated to be 0.85, which is marked as a high-importance area. At this time, the right area is judged as a blind spot because it has not been fixated for a long time (B t exceeds 2 seconds), and its compression ratio C r is calculated to be 0.7, indicating that a higher degree of compression is required. The compression strategy score P s is 0.4. Therefore, a lossy compression method is used to store this area. Through the precise quantification and dynamic evaluation of the driver's visual behavior, differential storage of video data is achieved.
[0083] In a specific embodiment, the process of executing step S103 may specifically include the following steps:
[0084] (1) Extract the overlapping area between adjacent video frames from the attention-guided hierarchical video data, register and mark the feature points of the overlapping area to obtain the video frame splicing reference data;
[0085] (2) Locate the boundaries of the areas not fixated by the driver in the video frame splicing reference data, and perform feature matching on the overlapping areas of adjacent video frames to obtain the video frame splicing relationship data;
[0086] (3) Geometrically transform the video images according to the video frame stitching relationship data, reconstruct the video content of the un-gazed area according to the spatial position, and obtain the video image spatial mapping data;
[0087] (4) Adjust the brightness of the un-gazed area in the video image spatial mapping data, perform compensation calculations on the contrast and clarity of the video content, and obtain the video enhancement parameter data;
[0088] (5) Adjust the contrast of the un-gazed area according to the video enhancement parameter data, and fuse the compensated video content with the original video content to obtain the region-enhanced video data;
[0089] (6) Integrate the region-enhanced video data with the panoramic video data, synthesize the video content according to the spatial position relationship, and obtain the panoramic image with attention compensation.
[0090] Specifically, extract the overlapping area between adjacent video frames from the attention-guided hierarchical video data. The overlapping area is formed by the intersection of the viewing angles of four cameras. The field of view angle of each camera is 120 degrees, and there is a 30-degree overlapping area between adjacent cameras. Feature points are extracted within the overlapping area through the SIFT feature point detection algorithm. Each feature point contains information such as position coordinates, gradient direction, and local descriptor. For the extracted feature points, a two-way matching strategy is used for feature point registration. The specific process is to first select feature points in the first frame, then find the best matching points within the search window of the second frame, and then match back from the second frame to the first frame. Only the feature point pairs with consistent two-way matching are retained as reliable matching points. These matching points form the video frame stitching reference data.
[0091] Based on the video frame stitching reference data, locate the boundary of the un-gazed area by the driver. The un-gazed area is determined by the head pose data. When a certain area has not been gazed at for 3 consecutive seconds, it is marked as an un-gazed area. For the boundary of the un-gazed area, determine the boundary position through feature point density analysis, perform feature matching on the overlapping areas of adjacent video frames, and generate the video frame stitching relationship data.
[0092] The geometric transformation of the video image uses a projection transformation matrix to reconstruct the video content of the un-gazed area according to the spatial position. The transformation matrix is obtained by solving the corresponding relationship of feature points through the least squares method to ensure seamless stitching of the transformed image. The video content after geometric transformation forms the video image spatial mapping data. Enhance the brightness and contrast of the un-gazed area in the video image spatial mapping data. Calculate the brightness histogram of the un-gazed area to obtain the brightness distribution characteristics. Then adjust the brightness distribution through the histogram equalization method, and at the same time use the local contrast enhancement algorithm to enhance the image details. This process generates a set of video enhancement parameter data, including brightness gain coefficient, contrast adjustment parameter, etc.
[0093] Fuse the enhanced video of the un-gazed area with the original video content. The fusion process uses a weighted average method and employs a gradient weight in the boundary area to ensure a natural transition. Finally, align the spatial positions of the fusion result with the panoramic video to generate an attention-compensated panoramic image.
[0094] For example: When the driver is about to turn left at an intersection, the head is mainly facing the front left, and at this time, the area in the rear right becomes the un-gazed area. The system detects that there are 50 pairs of feature points in the overlapping area between the rear right camera and the rear camera in the four-way video, and these feature points are mainly distributed in edge areas such as lane lines and curbs. Determine the spatial correspondence relationship of the overlapping area through feature point matching.
[0095] During the image enhancement process, it is found that the average brightness value of the un-gazed area is low and the contrast is insufficient. Determine the brightness compensation parameters through histogram analysis and enhance the dark details to a visible level. Adopt a 20-pixel-wide gradient band in the boundary area, with the weight linearly changing from 0 to 1 to ensure a smooth transition between the enhanced area and the original area. In the finally generated panoramic image, the visual features of the un-gazed area reach a similar level to those of the gazed area, forming a driving scene monitoring record. By focusing on processing the un-gazed area, the integrity of the driving monitoring data is improved.
[0096] In a specific embodiment, the process of executing step S104 may specifically include the following steps:
[0097] (1) Extract the scene change information of the blind spot of sight from the attention-compensated panoramic image, calculate the dynamic features according to the difference degree of consecutive frames of the scene, and obtain the blind spot scene change data;
[0098] (2) Establish a spatio-temporal change matrix based on the blind spot scene change data, correlate and match the displacement and speed information of moving targets in the blind spot of sight, and obtain the target motion trend data;
[0099] (3) Cross-verify the target motion trend data with the change characteristics of the driver's head posture, and conduct a spatial overlap analysis of the target motion direction and the blind spot of gaze to obtain threat prediction data;
[0100] (4) Calculate the intervention urgency of the abnormal situation in the blind spot according to the threat prediction data, compare and analyze the scene change rate and the attention transfer time, and obtain threat level determination data;
[0101] (5) Establish a multi-dimensional scoring system for the threat level determination data, comprehensively evaluate the spatial position, change speed, and attention deviation degree of the abnormal situation, and obtain threat ranking standard data;
[0102] (6) Classify and grade the abnormal situations according to the threat ranking criteria data, and prioritize the scenarios according to the urgency of the threat level to obtain the scenario threat level data.
[0103] Specifically, as Figure 3 shown, it is a schematic diagram of the scenario threat priority ranking in the embodiment of the present application. It is divided into four main parts from top to bottom: the threat assessment part shows the processing flow from scenario change detection to intervention urgency calculation; the priority ranking part divides the threat level into four levels; finally, specific scenario examples are used to illustrate the determination criteria for different priorities, where the highest priority corresponds to scenarios with a collision warning time less than 3 seconds or approaching the blind area at high speed, the high priority corresponds to scenarios of quickly cutting into the blind area or sudden deceleration, the medium priority corresponds to scenarios of stable following or normal lane change, and the low priority corresponds to scenarios of distant targets or stationary scenes. The arrow indicates the data processing flow and logical relationship. Extract the scenario change information of the line-of-sight blind area from the panoramic image of attention compensation, and the calculation of the scenario change adopts the following formula:
[0104]
[0105] Where: G d is the scenario dynamic change score; J n is the inter-frame optical flow intensity; Y n is the target size change rate; Z n is the scenario depth change amount; ρ 1 , ρ 2 , ρ 3 are the weight coefficients; N is the number of consecutive frames.
[0106] Spatial overlap threat degree calculation formula:
[0107]
[0108] Where: X t is the spatial threat degree; Q d is the motion direction deviation; W s is the speed vector intensity; O v is the spatial overlap degree; χ 1 , χ 2 , χ 3 are the weight coefficients.
[0109] Intervention urgency calculation formula:
[0110]
[0111] Where: U r is the intervention urgency; P t is the predicted collision time; D c is the distance change rate; E tis the escape time window; ζ 1 , ζ 2 , ζ 3 are weight coefficients.
[0112] Extract the scene change information of the line-of-sight blind area from the panoramic image with attention compensation. The blind area scene change data contains multiple dimensions, and the inter-frame optical flow intensity J n reflects the intensity of object movement and is obtained by calculating the displacement vector of pixel points between adjacent frames; the target size change rate Y n represents the change speed of the object size in the image and is calculated by the area ratio of the target detection box; the scene depth change amount Z n is then obtained by calculating the change value of the depth information through binocular disparity.
[0113] During the establishment of the spatio-temporal change matrix, the displacement and velocity information of the moving target are quantized. Specifically, for each detected moving target, its position coordinates, movement direction vector, and velocity scalar in consecutive frames are recorded. These data are organized into a matrix form according to the time stamp, with each row representing the state at a time point and each column representing different feature dimensions, thus forming the target movement trend data.
[0114] The generation of threat prediction data requires the correlation analysis of the target movement trend and the driver's head posture. The spatial threat degree X t is calculated considering the movement direction deviation Q d (the angle between the target movement direction and the driver's line-of-sight direction), the velocity vector intensity W s (the magnitude of the target relative velocity), and the spatial overlap degree O v (the intersection degree of the target trajectory and the vehicle prediction path).
[0115] When calculating the intervention urgency, the predicted collision time P t is obtained by dividing the current distance by the relative velocity, and the distance change rate D c reflects the change trend of the approaching speed, and the escape time window E t represents the minimum reaction time required to avoid a collision. These parameters are combined to form the intervention urgency U r .
[0116] For example: When a vehicle is driving on the highway, a vehicle approaching rapidly appears in the right rear blind area. Through continuous frame analysis, it is detected that the optical flow intensity J n of the vehicle is large, indicating intense movement; the target size increases rapidly in a short time, and the Y n value is significant; the depth change amount Z n is also large, indicating that the distance is rapidly shortening. These data are used to calculate a high change score G d .
[0117] Meanwhile, the movement trajectory of the vehicle overlaps with the position of the current lane, and the spatial overlap degree O v is close to 1, and the angle between the movement direction and the current gaze direction of the driver (left front) is large, resulting in a high movement direction deviation Q d . The speed vector analysis shows that the relative speed difference is significant, generating a high threat prediction score. When calculating the intervention urgency, due to the short predicted collision time, large distance change rate, and small escape time window, a high intervention urgency U is finally obtained r .
[0118] In a specific embodiment, the process of executing step S105 may specifically include the following steps:
[0119] (1) Extract the timestamp and spatial position information from the scene threat level data, and associate and map the spatio-temporal features of the video segment with the threat level to obtain threat scene location data;
[0120] (2) Establish an attention distraction degree evaluation index according to the threat scene location data, and conduct an association analysis between the change frequency of the driver's head posture and the change law of the threat level to obtain attention fluctuation feature data;
[0121] (3) Perform time series segmentation on the attention fluctuation feature data, align the continuously occurring attention distraction intervals with the corresponding scenes in time to obtain an attention distraction scene sequence;
[0122] (4) Establish a multi-level index structure according to the attention distraction scene sequence, group and organize the scenes with different threat levels according to the attention distraction degree to obtain scene classification index data;
[0123] (5) Mark the high-threat scenes in the scene classification index data, and comprehensively score the attention distraction duration and the scene threat degree to obtain scene marking weight data;
[0124] (6) Integrate the scene marking weight data with the scene index structure, and optimize and sort the retrieval path according to the marking weight to obtain a retrieval structure associated with attention.
[0125] Specifically, time information including millisecond-level timestamps and spatial position information based on the vehicle coordinate system is extracted from the scene threat level data. The timestamp is accurate to 1 ms, and the spatial position is represented by a polar coordinate system with the vehicle center as the origin, including two parameters: distance r and azimuth angle θ. Each threat scene segment is assigned a unique spatio-temporal identifier and mapped to its corresponding threat level (divided into high, medium, and low levels) to form threat scene location data.
[0126] The attention distraction degree evaluation index is constructed based on the characteristics of the driver's head posture changes. The head posture change frequency is calculated by the number of head rotations per unit time, and at the same time, the angle range and duration of each rotation are recorded. Through the temporal comparison of the head posture change frequency and the change law of the threat level in the scene, the correlation analysis between the two is established. This analysis process focuses on the time correspondence between the moments of rapid head rotation and the moments of sudden change in the threat level, generating attention fluctuation characteristic data. When segmenting the attention fluctuation characteristic data in time series, the sliding window method is adopted, and the window length is set to 3 seconds. In each window, if the head posture change frequency exceeds a preset threshold (such as 2 times per second) and the duration exceeds 1 second, then this time period is marked as an attention distraction interval. Synchronize these attention distraction intervals with the scene data at the corresponding moments in time to establish an attention distraction scene sequence.
[0127] The establishment of the multi-level index structure adopts a tree structure. The first layer is classified according to the threat level (high, medium, low), the second layer is classified according to the attention distraction degree (severe, moderate, mild), and the third layer is organized in chronological order. The scenes with different threat levels are grouped according to the attention distraction degree. Each group of scenes contains complete spatio-temporal information and scene feature descriptions, forming scene classification index data. In the process of highlighting the high-threat scenes, two dimensions of the attention distraction duration and the scene threat degree are comprehensively considered. The calculation of the marking weight adopts the weighted average method, where the weight of the attention distraction duration is 0.6 and the weight of the scene threat degree is 0.4. Such a weight assignment highlights the danger of long-term attention distraction and also considers the threat degree of the scene itself, generating scene marking weight data.
[0128] Finally, integrate the scene marking weight data into the index structure and optimize the sorting of the retrieval path. The sorting rule gives priority to the scenes with high threat and high weight. These scenes are placed at the top-level nodes of the index tree for quick access. At the same time, establish the association relationship between the scenes in the index structure. Similar scenes are connected to each other through pointers to form a complete attention association retrieval structure.
[0129] For example: When the driver is driving on the highway, he frequently checks the navigation device. Through spatio-temporal data extraction, it is recorded that the driver looked down at the navigation 6 times within 10 seconds, and each time lasted about 0.5 seconds. At the same time, a vehicle is overtaking quickly from the right rear. The data processing process extracts all the information during this period from the scene threat level data: the time stamp shows that the overtaking process occurred from the 3rd second to the 7th second, and the spatial position data shows that the overtaking vehicle approached from a 45-degree angle from the right rear, and the closest distance was 2 meters.
[0130] This scene is determined by the system to be of a high threat level and is associated with the driver's attention distraction data. Analysis reveals that within the critical 4 seconds when overtaking occurred, the driver looked down at the navigation 3 times, each time lasting approximately 0.5 seconds, which is a typical state of attention distraction. The system labels this video clip as the "High Threat - Severe Distraction" category and assigns the highest access priority in the index structure.
[0131] In a specific embodiment, the process of performing step S106 may specifically include the following steps:
[0132] (1) Extract the driver's attention coverage area information from the attention - associated retrieval structure, perform an association calculation on the spatial range of the coverage area and the observation duration to obtain attention coverage intensity data;
[0133] (2) Perform partition annotation on the video clip according to the attention coverage intensity data, conduct a combined analysis of the areas with different coverage intensities and the corresponding threat levels to obtain area storage priority data;
[0134] (3) Perform hierarchical mapping on the area storage priority data, elevate the storage priority of the high - threat un - gazed areas to the protection level to obtain storage protection area data;
[0135] (4) Establish a data redundancy backup strategy based on the storage protection area data, store multiple copies of the video content of the high - threat un - gazed areas to obtain a data fault - tolerance protection strategy;
[0136] (5) Allocate storage space for the video data under the data fault - tolerance protection strategy, divide the data with different priorities according to the storage hierarchy to obtain hierarchical storage structure data;
[0137] (6) Organize and manage the video content in the hierarchical storage structure data according to the storage protection requirements, perform storage allocation on the data based on the attention coverage degree and threat level to obtain a hierarchical storage video data set.
[0138] Specifically, when extracting the driver's attention coverage area information from the attention - associated retrieval structure, determine the attention distribution within the field of view. In the vehicle coordinate system, the driver's field of view is divided into four main areas: 120 degrees in the front, 90 degrees on each side (left and right), and 60 degrees in the rear. For each area, record the time length and fixation frequency of the driver's line of sight stay to form initial attention coverage data. The calculation of attention coverage intensity comprehensively considers the spatial range and time dimension. The spatial range is defined by angles and distances, and the coverage intensity of each area is proportional to the time the line of sight stays in that area and inversely proportional to the angle of deviation of the line of sight from the center. This calculation method reflects the natural distribution characteristics of visual attention: more attention is paid to the central area, and the attention degree gradually decreases towards the edge area.
[0139] During the process of partitioning and annotating video clips, the attention coverage intensity data is combined with the threat level. The threat level is divided into three levels from high to low, and each level is further subdivided into three sub-levels: strong coverage, medium coverage, and weak coverage according to the attention coverage intensity. The combined analysis generates nine storage priority levels, and the areas with high threat and low coverage are given the highest storage priority. After the storage priority data is hierarchically mapped, different levels of storage policies are determined. For the areas with high threat and un-gazed, the highest-level protection storage policy is adopted, including lossless compression, multi-copy backup, and regular integrity checks. The medium-threat areas adopt a balanced storage policy, with moderate compression while ensuring quality. The low-threat areas adopt high compression ratio storage, mainly retaining key frame information.
[0140] The data redundancy backup policy is formulated based on the hierarchical storage requirements. For the video content in the areas with high threat and un-gazed, complete copies are saved at multiple physical storage locations simultaneously, and each copy contains complete metadata information. The backup policy includes two methods: real-time backup and regular incremental backup to ensure data reliability and recoverability. The storage space is allocated using a hierarchical architecture, dividing the storage media into a cache layer, a primary storage layer, and an archive storage layer. The data in high-threat scenarios is preferentially stored in the cache layer for quick access; the medium-threat scenarios are stored in the primary storage layer; and the low-threat scenarios are stored in the archive layer. Dynamic balance is maintained between different levels through data migration policies.
[0141] For example: In a lane-changing scenario, the driver's main attention is concentrated on the left front, and the continuous observation time reaches 3 seconds. At this time, a vehicle approaching rapidly appears in the right rear, which is determined as a high-threat scenario. The data processing flow calculates that the attention coverage intensity in the right rear area is low (because the driver's line of sight is mainly in the left front). After being combined with the high-threat level, the video data in this area is given the highest storage priority. The storage system immediately creates complete copies of this video clip at multiple storage nodes and saves the metadata containing information such as the threat level and attention coverage intensity. These data are preferentially stored in the cache layer and synchronized and backed up in the primary storage layer at the same time. This storage policy ensures the security and accessibility of key scenario data.
[0142] In a specific embodiment, the process of performing the step of partitioning and annotating video clips according to the attention coverage intensity data may specifically include the following steps:
[0143] (1) Extract the coverage duration value and the coverage area value from the attention coverage intensity data, perform numerical normalization processing on the duration value and the area value to obtain the normalized coverage intensity data;
[0144] (2) Divide the video space according to the coverage intensity normalization data, mark the video frames according to the distribution characteristics of the coverage intensity, and obtain the coverage area distribution data;
[0145] (3) Conduct a time series analysis on the coverage area distribution data, match the change trend of the coverage intensity within continuous time with the threat level, and obtain the regional threat correlation data;
[0146] (4) Combine and calculate the regional characteristics in the regional threat correlation data, calculate the priority index according to the product relationship between the coverage intensity and the threat level, and obtain the regional priority index data;
[0147] (5) Conduct a hierarchical threshold division on the regional priority index data, stratify the regions with different priority indexes according to the storage importance, and obtain the regional hierarchical mapping data;
[0148] (6) Sort and organize the regional hierarchical mapping data according to the storage priority, allocate and plan the storage resources according to the regional hierarchical results, and obtain the regional storage priority data.
[0149] Specifically, two key parameters are extracted from the attention coverage intensity data: the coverage duration value and the coverage area value. The coverage duration value is obtained by accumulating the fixation duration of the driver in a specific area, with the unit of seconds; the coverage area value is obtained by calculating the scanned area of the fixation points within the visual field, with the unit of square degrees (visual angle area). The numerical normalization process uses the maximum-minimum normalization method. For the duration value, 10 seconds is selected as the maximum standard value, and the actual coverage time is divided by 10 to obtain a normalized value between 0 and 1. For the area value, a standard visual field area of 120 degrees × 60 degrees is used as the benchmark, and the ratio of the actual coverage area to the standard area is calculated. These two normalized values are weighted and averaged to obtain the final coverage intensity normalization data.
[0150] The regional division of the video space is based on the coverage intensity normalization data. The entire visual field range is divided into grid-like regions, with each grid size of 10 degrees × 10 degrees. According to the distribution characteristics of the normalization data, each grid is assigned a corresponding coverage intensity value. Adjacent grids with similar coverage intensities are merged into larger regions to form the coverage area distribution data. The time series analysis of the coverage area distribution data uses the sliding window method, with the window size set to 3 seconds. Within each time window, the change trend of the coverage intensity is calculated, and at the same time, the threat level data for this time period is obtained. By performing time alignment and correlation analysis on these two sets of data, the regional threat correlation data is obtained.
[0151] The combined operation of regional features comprehensively considers the coverage intensity and threat level. The coverage intensity reflects the distribution of the driver's visual attention, and the threat level reflects the degree of danger of the scene. The priority index is calculated by the weighted product of the two, and the weight coefficient is determined according to the actual road safety requirements. This calculation method emphasizes that high-threat areas should be given high-priority treatment even if they are briefly ignored.
[0152] The grading threshold division of the priority index data adopts the dynamic threshold method. First, the distribution characteristics of the priority index in historical data are statistically analyzed to determine the key demarcation points. Then, according to these demarcation points, the regions are divided into different storage importance levels. Regions above the 90th percentile are divided into the highest importance level and need to be stored without loss; regions in the 80%-90% percentile are compressed with low loss; the remaining regions can be stored with a high compression ratio.
[0153] Taking a highway driving scenario as an example: When the driver is changing lanes, he mainly focuses on the left front area. The cumulative coverage duration of this area reaches 2.5 seconds, and the coverage area is approximately 30 degrees × 20 degrees. After normalization, the duration normalization value is 0.25, and the area normalization value is 0.083. At the same time, a vehicle approaching rapidly appears in the right rear blind spot, and the threat level is determined to be the highest level. The coverage intensity of this blind spot is almost zero, but due to the highest threat level, its priority index still reaches a relatively high level. During storage allocation, the video data of this area is divided into the highest storage importance level, stored in a lossless manner, and a backup is created. This method ensures the complete recording of critical threat scenarios, even if this area is not fully noticed by the driver.
[0154] In a specific embodiment, the process of performing the storage space allocation step for video data under the data fault tolerance protection strategy may specifically include the following steps:
[0155] (1) Extract the information on the number of copies of video data from the data fault tolerance protection strategy, group and sort the video data with different priorities according to the number of copies to obtain the data copy distribution data;
[0156] (2) Calculate the storage space requirements for each group of data according to the data copy distribution data, perform a multiplication operation on the file size of the video data and the number of copies to obtain the storage capacity requirement data;
[0157] (3) Perform a storage level division on the storage capacity requirement data, configure the storage space in layers according to the access speed and data importance to obtain the storage level configuration data;
[0158] (4) Establish a data allocation rule according to the storage level configuration data, map the video data with different priorities to the corresponding storage levels to obtain the data allocation rule data;
[0159] (5) Reserve storage space for the data distribution rule data, match the usage rate of each level of storage space with the data growth trend, and obtain the space reservation strategy data;
[0160] (6) Integrate the space reservation strategy data with the storage level configuration, plan the storage structure according to the data distribution characteristics of each level, and obtain the hierarchical storage structure data.
[0161] Specifically, first extract the copy number information of video data from the data fault tolerance protection strategy, including the copy number configuration corresponding to each priority level. The highest priority data is configured with 3 copies, the medium priority data is configured with 2 copies, and the low priority data is configured with 1 copy. These configuration information are sorted into a data copy distribution table, which records the space distribution characteristics of different priority data. When calculating the storage space requirement according to the data copy distribution data, the size of the original video data and the number of copies need to be considered. The size of video data is determined by the resolution, frame rate, and encoding method. The original video with a resolution of 1080P and a frame rate of 30fps, after being encoded by H.264, occupies about 150MB of storage space per minute. For an 8-hour driving record, including four video streams, the total basic storage requirement is about 432GB. Considering the number of copies, the high priority data requires three times the storage space, and the medium priority data requires twice the storage space.
[0162] The storage level is divided into a three-layer architecture: cache layer, main storage layer, and archive storage layer. The cache layer uses solid state drives with an access speed of more than 500MB / s; the main storage layer uses mechanical hard drives with an access speed of about 100MB / s; the archive storage layer uses large-capacity storage devices with an access speed of less than 50MB / s. Each layer is configured with different storage capacities and read / write bandwidths to form the storage level configuration data.
[0163] The data distribution rule is established based on the storage level configuration. High priority data is preferentially stored in the cache layer, and at the same time, one copy is reserved in the main storage layer and the archive layer respectively; medium priority data is stored in the main storage layer and a backup is reserved in the archive layer; low priority data is directly stored in the archive layer. This distribution strategy ensures the fast access and reliability of important data. The space reservation strategy takes into account the data growth trend. By analyzing the historical data growth rate, the storage requirement for a future period is predicted. The cache layer reserves 30% of the space for data buffering, the main storage layer reserves 25% of the space, and the archive layer reserves 20% of the space. This differential reservation strategy adapts to the data update characteristics of different levels.
[0164] Taking a four-hour urban road driving as an example: During the driving process, multiple high-risk scenarios were recorded, including sudden braking, emergency lane changes, etc. These high-priority scenario data account for about 15% of the total data volume. Three complete copies are created for each scenario and stored in three levels respectively. Medium-priority data accounts for 35%, including general lane changes, turning, etc. scenarios. Two copies are created and stored in the main storage and the archive layer respectively. The remaining 50% is low-priority data, and only a single copy is saved in the archive layer. The entire storage solution not only ensures the security of critical data but also realizes the rational utilization of storage resources. Differentiated storage is implemented according to the importance of data. Through the multi-copy mechanism and the hierarchical storage architecture, the requirements for data security, access performance, and storage cost are balanced. In practical applications, the storage policy will be dynamically adjusted according to the real-time monitored storage status to ensure the stable operation of the storage system.
[0165] The above describes the driving video recording method based on four-way monitoring in the embodiments of the present application. Next, the driving video recording system based on four-way monitoring in the embodiments of the present application will be described. Please refer to Figure 4 , an embodiment of the driving video recording system based on four-way monitoring in the embodiments of the present application includes:
[0166] An acquisition module 201, configured to synchronously acquire the vehicle surrounding environment and the driver's head posture through a real-time video stream, perform time-corresponding processing on the video stream and the head posture data, and obtain driving attention reference video data;
[0167] A grading module 202, configured to perform video grading processing according to the head posture change characteristics in the driving attention reference video data, compress and store the video data in the driver's line of sight blind area, and obtain attention-oriented graded video data;
[0168] A synthesis module 203, configured to perform view synthesis processing on the attention-oriented graded video data, perform enhanced mapping on the video content in the area not gazed by the driver, and obtain an attention-compensated panoramic image;
[0169] An identification module 204, configured to use the attention-compensated panoramic image to perform scene threat identification, prioritize abnormal situations in the line of sight blind area, and obtain scene threat level data;
[0170] A construction module 205, configured to build an index based on the scene threat level data for the video content, mark the scenes when the driver's attention is distracted, and obtain an attention-associated retrieval structure;
[0171] A storage module 206 is configured to store and classify the data in the retrieval structure associated with the attention according to the degree of driver attention coverage, and protect and store the video data in the high-threat un-gazed area, so as to obtain a hierarchically stored video data set.
[0172] Through the collaborative cooperation of the above-mentioned various components, by establishing a linkage mechanism between driving attention and environmental monitoring, the intelligent processing and storage of video data are realized. Through the synchronous acquisition and time alignment processing of the vehicle surrounding environment and the driver's head posture, the corresponding relationship between the driver's attention distribution and environmental changes is accurately captured; based on the importance analysis of head rotation features and the compression storage strategy, the differential processing of video data is realized, and the storage resources are more reasonably allocated; the visual enhancement processing of the un-gazed area effectively compensates for the blind area of the driver's visual attention and improves the integrity of video recording; through the threat recognition and prioritization of the attention-compensated panoramic image, a scene threat assessment mechanism based on attention features is established, enabling the system to accurately identify and record potential dangerous scenes; the retrieval marking mechanism associated with attention enables high-risk scenes to be quickly located and played back, facilitating post-event analysis and summary; the multi-copy storage strategy for high-threat un-gazed areas ensures the security and reliability of key data. At the algorithm level, this solution innovatively combines the deep learning object detection algorithm with the attention feature extraction algorithm, and dynamically adjusts the video processing strategy through the real-time analysis of the driver's attention pattern, enabling the system to more intelligently allocate computing and storage resources in the face of complex road environments. At the same time, based on the time-series associated data processing framework, the system can achieve the efficient processing and management of large-scale video data while maintaining a low computational complexity, providing a feasible technical solution for applications in actual road scenarios and improving the pertinence and effectiveness of driving records.
[0173] As mentioned above, the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.
Claims
1. A driving video recording method based on four-way monitoring, characterized in that: The driving video recording method based on four-way monitoring includes: The vehicle surrounding environment and the driver's head posture are synchronously collected through real-time video stream, and the video stream and head posture data are processed in time correspondence to obtain driving attention benchmark video data; Performing video classification processing according to the head posture change characteristics in the driving attention benchmark video data, compressing and storing the video data located in the driver's blind spot, and obtaining attention-oriented classified video data; Performing view synthesis processing on the attention-oriented hierarchical video data, enhancing and mapping the video content of the area not being looked at by the driver, and obtaining an attention-compensated panoramic image; Using the attention-compensated panoramic image to identify scene threats, prioritize abnormal situations in the blind spot, and obtain scene threat level data; Indexing and constructing video content based on the scene threat level data, focusing on marking scenes when the driver's attention is distracted, and obtaining an attention-related retrieval structure; The data in the attention-related retrieval structure are stored and classified according to the driver's attention coverage, and the video data of the high-threat unwatched area is protected and stored to obtain a hierarchically stored video data set.
2. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The vehicle surrounding environment and the driver's head posture are synchronously collected through real-time video stream, and the video stream and the head posture data are processed in time correspondence to obtain driving attention benchmark video data, including: The environmental video stream data of the front, rear, left, and right sides of the vehicle are collected by four cameras, and the environmental video stream data are synchronously integrated according to timestamps to obtain multi-view environmental video data; Continuously sample the driver's head posture changes, convert the head rotation angle, rotation speed, and gaze duration into attention feature parameters to obtain head posture feature data; Establishing a correspondence between the multi-view environment video data and the head posture feature data according to the timestamp to form a temporal correlation matrix to obtain video-attention mapping data; Performing spatial division on the attention coverage area in the video-attention mapping data, dividing the video screen into a fixation area and a blind area, and obtaining spatial attention distribution data; Calculate the attention weight of the video content according to the spatial attention distribution data, assign different data processing priorities to the fixation area and the blind area, and obtain an attention weight allocation result; The attention weight allocation result is spatially and temporally fused with the video data, and the video data is marked according to the attention distribution characteristics to obtain driving attention benchmark video data.
3. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The step of performing video classification processing according to the head posture change characteristics in the driving attention benchmark video data, compressing and storing the video data located in the driver's blind spot, and obtaining attention-oriented classified video data includes: Extracting head posture change features from the driving attention benchmark video data, marking the importance of the video area according to the head rotation frequency and the gaze duration, and obtaining video area importance data; Establishing a corresponding relationship between the video region importance data and the video frame data, and classifying the video regions according to the driver's gaze frequency to obtain a hierarchical mapping relationship of the video regions; Positioning the driver's blind spot according to the video area hierarchical mapping relationship, marking the continuous video area that is not watched as the blind spot, and obtaining the blind spot positioning data; Calculating the compression ratio of the video content in the blind spot positioning data, determining the compression priority according to the blind spot duration, and obtaining the blind spot compression strategy data; Associating the blind spot compression strategy data with the video data, and compressing the video data of the blind spot in blocks according to the compression priority to obtain compressed area data; The compressed region data is integrated with the non-compressed region data, and the data is organized according to the importance of the video regions to obtain attention-oriented hierarchical video data.
4. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The step of performing view synthesis processing on the attention-directed hierarchical video data, enhancing and mapping the video content of the area not watched by the driver, and obtaining a panoramic image with attention compensation includes: Extracting overlapping areas between adjacent video frames from the attention-guided hierarchical video data, registering and marking feature points in the overlapping areas, and obtaining video frame splicing benchmark data; Performing boundary positioning on the driver's non-attention area in the video frame splicing reference data, and performing feature matching on the overlapping areas of adjacent video frames to obtain video frame splicing relationship data; Performing geometric transformation on the video picture according to the video frame splicing relationship data, reconstructing the video content of the unobserved area according to the spatial position, and obtaining the video picture space mapping data; Adjusting the brightness of the non-watched area in the video picture space mapping data, and performing compensation calculation on the contrast and clarity of the video content to obtain video enhancement parameter data; According to the video enhancement parameter data, the contrast of the non-watched area is adjusted, and the compensated video content is merged with the original video content to obtain the regional enhanced video data; The regional enhanced video data is integrated with the panoramic video data, and the video content is synthesized according to the spatial position relationship to obtain an attention-compensated panoramic image.
5. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The method of using the attention-compensated panoramic image to identify scene threats, prioritizing abnormal situations in the blind spot, and obtaining scene threat level data includes: Extracting scene change information of the sight blind spot from the attention-compensated panoramic image, calculating dynamic features according to the difference between consecutive frames of the scene, and obtaining scene change data of the blind spot; Establishing a spatiotemporal change matrix based on the blind spot scene change data, correlating and matching the displacement and speed information of the moving target in the sight blind spot, and obtaining the target movement trend data; Cross-validate the target motion trend data with the driver's head posture change characteristics, perform spatial overlap analysis on the target motion direction and the blind spot of gaze, and obtain threat prediction data; Calculate the intervention urgency of the abnormal situation in the blind spot according to the threat prediction data, compare and analyze the scene change rate and the attention shift time, and obtain threat level determination data; A multi-dimensional scoring system is established for the threat level determination data, and the spatial location, change speed and degree of attention deviation of the abnormal situation are comprehensively evaluated to obtain threat ranking standard data; The threat ranking standard data is used to classify abnormal situations, and the scenarios are prioritized according to the urgency of the threat level to obtain scenario threat level data.
6. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The indexing and constructing of the video content based on the scene threat level data, focusing on marking the scene when the driver's attention is distracted, and obtaining the attention-related retrieval structure includes: Extracting timestamp and spatial location information from the scene threat level data, associating and mapping the spatiotemporal features of the video clips with the threat level, and obtaining threat scene location data; Establishing an attention distraction degree evaluation index based on the threat scene positioning data, correlating and analyzing the driver's head posture change frequency with the threat level change pattern, and obtaining attention fluctuation characteristic data; Performing time-series segmentation on the attention fluctuation characteristic data, and aligning the consecutive attention distraction intervals with the corresponding scenes in time to obtain an attention distraction scene sequence; Establishing a multi-level index structure according to the attention distraction scene sequence, grouping and organizing scenes of different threat levels according to the degree of attention distraction, and obtaining scene classification index data; High-threat scenes in the scene classification index data are marked with emphasis, and the duration of distraction and the degree of scene threat are comprehensively scored to obtain scene marking weight data; The scene tag weight data is integrated with the scene index structure, and the retrieval paths are optimized and sorted according to the tag weights to obtain an attention-associated retrieval structure.
7. The driving video recording method based on four-way monitoring according to claim 1 is characterized in that: The data in the attention-related retrieval structure is stored and classified according to the driver's attention coverage, and the video data of the high-threat unwatched area is protected and stored to obtain a hierarchically stored video data set, including: Extracting driver attention coverage area information from the attention-related retrieval structure, and correlating and calculating the spatial range of the coverage area with the observation duration to obtain attention coverage intensity data; The video clips are marked according to the attention coverage intensity data, and regions with different coverage intensities are combined and analyzed with corresponding threat levels to obtain regional storage priority data; Performing hierarchical mapping on the regional storage priority data, raising the storage priority of the high-threat unwatched area to a protection level, and obtaining storage protection area data; Establishing a data redundancy backup strategy based on the storage protection area data, storing multiple copies of the video content in the high-threat unwatched area, and obtaining a data fault-tolerant protection strategy; Allocating storage space for the video data under the data fault-tolerant protection strategy, dividing data of different priorities according to storage levels, and obtaining hierarchical storage structure data; The video content in the hierarchical storage structure data is organized and managed according to the storage protection requirements, and the data is stored and allocated according to the attention coverage degree and the threat level to obtain a hierarchically stored video data set.
8. The driving video recording method based on four-way monitoring according to claim 7 is characterized in that: The video clips are marked according to the attention coverage intensity data, and regions with different coverage intensities are combined and analyzed with corresponding threat levels to obtain regional storage priority data, including: Extracting the coverage duration value and the coverage area value from the attention coverage intensity data, performing numerical normalization processing on the duration value and the area value to obtain coverage intensity normalized data; Divide the video space into regions according to the coverage intensity normalization data, mark the regions of the video screen according to the distribution characteristics of the coverage intensity, and obtain coverage area distribution data; Performing time series analysis on the coverage area distribution data, matching the variation trend of coverage intensity in continuous time with the threat level, and obtaining regional threat correlation data; Combining the regional features in the regional threat association data, calculating the priority index according to the product relationship between the coverage intensity and the threat level, and obtaining regional priority index data; The regional priority index data is divided into hierarchical thresholds, and regions with different priority indexes are layered according to storage importance to obtain regional hierarchical mapping data; The regional classification mapping data is sorted and organized according to storage priority, and storage resources are allocated and planned according to the regional classification result to obtain regional storage priority data.
9. The driving video recording method based on four-way monitoring according to claim 7 is characterized in that: The storage space is allocated for the video data under the data fault-tolerant protection strategy, and data of different priorities are divided according to storage levels to obtain hierarchical storage structure data, including: Extracting the number of copies of the video data from the data fault-tolerant protection strategy, grouping and arranging the video data of different priorities according to the number of copies, and obtaining data copy distribution data; Calculate the storage space requirement of each group of data according to the data copy distribution data, multiply the file size of the video data by the number of copies to obtain storage capacity requirement data; Dividing the storage capacity demand data into storage levels, configuring the storage space in layers according to access speed and data importance, and obtaining storage level configuration data; Establishing data allocation rules according to the storage level configuration data, mapping video data of different priorities to corresponding storage levels, and obtaining data allocation rule data; Reserving storage space for the data allocation rule data, matching the usage rate of each level of storage space with the data growth trend, and obtaining space reservation strategy data; The space reservation strategy data is integrated with the storage level configuration, and the storage structure is planned according to the data distribution characteristics of each level to obtain hierarchical storage structure data.
10. A driving video recording system based on four-way monitoring, used to implement the driving video recording method based on four-way monitoring as described in any one of claims 1 to 7, characterized in that: The driving video recording system based on four-way monitoring includes: The acquisition module is used to synchronously acquire the vehicle's surrounding environment and the driver's head posture through real-time video stream, and perform time correspondence processing on the video stream and the head posture data to obtain driving attention benchmark video data; A grading module, used to perform video grading processing according to the head posture change characteristics in the driving attention benchmark video data, compress and store the video data located in the driver's blind spot, and obtain attention-oriented graded video data; A synthesis module, configured to perform view synthesis processing on the attention-directed hierarchical video data, enhance and map the video content of the area not being looked at by the driver, and obtain a panoramic image with attention compensation; An identification module, used to identify scene threats using the attention-compensated panoramic image, prioritize abnormal situations in the blind spot, and obtain scene threat level data; A construction module, used to index and construct the video content based on the scene threat level data, focus on marking the scene when the driver's attention is distracted, and obtain a retrieval structure associated with attention; The storage module is used to store and classify the data in the attention-related retrieval structure according to the driver's attention coverage, and to protect and store the video data of the high-threat unwatched area to obtain a hierarchically stored video data set.
Citation Information
Cited By
Multi-source traffic video data fusion management method and system
CN120953938A
Multi-source traffic video data fusion management method and system
CN120953938B
Data backup management method and system for automobile data recorder
CN121144112A
Data backup management method and system for automobile data recorder
CN121144112B
Method and system for recognizing object watched by pilot based on eye movement data
CN122347795A