Method and system for storing video data in a vehicle
By determining frame intervals based on object classes and instances in vehicle sensor data, the method reduces video data storage size in vehicles, leveraging existing data processing for driver assistance tasks.
Patent Information
- Application Number
- DE102024113015
- Authority / Receiving Office
- DE · DE
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-05-08
- Publication Date
- 2025-08-28
- Estimated Expiration
- 2044-05-08
AI Technical Summary
Modern vehicles face challenges in reducing the processing amount for compressing video data captured by vehicle cameras due to the extensive real-time data processing required for driving automation systems, making it difficult to implement additional compression methods.
A method that determines object classes and instances within vehicle sensor data to identify frame intervals for exclusion, generating object lists for similar frames, thereby reducing the number of frames stored and using existing data processing for driver assistance tasks.
Reduces the storage size of video data by reusing data processing for driver assistance functions, minimizing additional processing effort and memory requirements.
Smart Images

Figure 00000001_0000 
Figure 00000015_0000 
Figure 00000016_0000
Abstract
Description
FIELD OF TECHNOLOGY
[0001] The invention relates generally to storing video data acquired by vehicle cameras, and more particularly to reducing the storage size of the video data based on object detection performed by a vehicle configured to perform at least some functions of a driving automation system (DAS) implementing at least one driver assistance. BACKGROUND
[0002] Modern vehicles that implement at least some level of driving automation perform extensive processing on vehicle sensor data to implement DAS functions. The vehicle sensor data typically includes video data acquired by vehicle cameras, which can be stored in the vehicle to implement, for example, an event data recorder (EDR). Given the size of video data, video data can typically be compressed before storage to reduce the storage requirements of the video data. While various compression techniques for video data exist, these techniques typically require additional data processing. Because the DAS functions implemented by the vehicle already require extensive real-time data processing, providing additional processing to compress the video data can be challenging.
[0003] An object of the present disclosure is therefore to reduce the processing effort for compressing video data in a vehicle.
[0004] Document EP 3 739 503 A1 discloses an apparatus, a method, and a computer program product for video processing. The apparatus may comprise means for receiving data representing a first frame of video content comprising a plurality of frames, and for determining an object type and a position in the first frame for at least one object in a first frame. The apparatus may also comprise means for determining a number of frames N to be skipped based on a comparison of the type and position of the object in the first frame with the type and position of one or more objects in one or more previous frames, and for providing the N+1 frame, rather than the skipped frames, to means for applying the frame to an image model.
[0005] Document US 2021 / 0142068 A1 discloses a computer system for decimating video data, including a processor, a persistent storage system coupled to the processor, and a memory storing instructions that, when executed by the processor, cause the processor to decimate a batch of frames of video data by: receiving the batch of individual images of video data, mapping the frames of the batch by a feature extractor to corresponding feature vectors in a feature space, each of the feature vectors having a lower dimension than a corresponding one of the frames of the batch, selecting a set of dissimilar frames from the plurality of frames of video data based on dissimilarities between corresponding ones of the feature vectors, and storing the selected set of dissimilar frames in the persistent storage system,where the size of the selected set of dissimilar frames is smaller than the number of frames in the batch of frames of video data.,
[0006] Document US 2019 / 0208136 A1 discloses an optical system for a vehicle that can be configured with a plurality of camera sensors. Each camera sensor can be configured to generate corresponding image data for a corresponding field of view. The optical system is further configured with a plurality of image processing units coupled to the plurality of camera sensors. The image processing units are configured to compress the image data acquired by the camera sensors. A computer system is configured to store the compressed image data in a memory. The computer system is further configured with a vehicle control processor configured to control the vehicle based on the compressed image data. The optical system and the computer system can be communicatively coupled via a data bus. SUMMARY OF THE INVENTION
[0007] To achieve this goal, the present disclosure provides a method for storing video data comprising a plurality of frames and captured by one or more vehicle cameras installed in a vehicle configured to perform at least one driver assistance function. The method comprises determining one or more object classes and one or more object instances within vehicle sensor data. The vehicle sensor data comprises the video data captured by the one or more vehicle cameras, wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least one driver assistance function.The method further comprises determining a reduced plurality of frames based on the one or more object classes and the one or more object instances, the plurality of frames, and one or more frame spacings. Each frame spacing indicates a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames. Determining the reduced plurality of frames and the one or more frame spacings comprises generating the one or more frame spacings based on a change in object classes and object instances between consecutive frames of the plurality of frames compared to a change threshold. The change threshold indicates a percentage of object classes and object instances within a frame of the plurality of frames that corresponds to object classes and object instances within a previous frame of the plurality of frames.The method further comprises, for each frame of the plurality of frames not included in the reduced plurality of frames, generating a corresponding object list. Each object list identifies the one or more object classes and the one or more object instances determined within a corresponding frame, as well as their corresponding positions within the corresponding frame. Finally, the method comprises storing the reduced plurality of frames and the object lists.
[0008] The present disclosure further provides a vehicle control unit comprising at least one processing unit and a memory coupled to the at least one processing unit and configured to store machine-readable instructions. The machine-readable instructions cause the at least one processing unit to determine one or more object classes and one or more object instances within the vehicle sensor data. The vehicle sensor data includes the video data acquired by the one or more vehicle cameras, wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least the driver assistance.The machine-readable instructions further cause the at least one processing unit to determine a reduced plurality of frames based on the one or more object classes and the one or more object instances, the plurality of frames, and one or more frame spacings. Each frame spacing indicates a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames. The machine-readable instructions further cause the at least one processing unit to generate a corresponding object list for each frame of the plurality of frames not included in the reduced plurality of frames. Each object list identifies the one or more object classes and the one or more object instances determined within a corresponding frame, as well as their corresponding positions within the corresponding frame.Finally, the machine-readable instructions further cause the at least one processing unit to store the reduced plurality of frames and the object lists.
[0009] The present disclosure further provides a vehicle including a plurality of sensors and the vehicle control unit. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Examples of the present disclosure will be described with reference to the following accompanying drawings, in which like reference numerals refer to like components. The Fig. 1 shows a plurality of frames, a reduced plurality of frames, and frame spacing according to examples of the present disclosure. The Fig. 2 provides a flowchart of a method for storing video data comprising a plurality of frames captured by one or more vehicle cameras installed in a vehicle configured to perform at least one driver assistance, in accordance with examples of the present disclosure. The Fig. 3 shows the vehicle including a plurality of vehicle sensors according to examples of the present disclosure. The Fig. 4 shows a vehicle control unit according to examples of the present disclosure.
[0011] It should be understood that the above-referenced drawings are not intended to limit the present disclosure in any way. Rather, these drawings are provided to aid understanding of the present disclosure. Those skilled in the art will readily understand that aspects of the present invention illustrated in one drawing may be combined with aspects in another drawing or omitted without departing from the scope of the present disclosure. DETAILED DESCRIPTION
[0012] The present disclosure generally provides a method, a vehicle control unit, and a vehicle configured to store video data. The vehicle includes a plurality of vehicle sensors including one or more vehicle cameras. The one or more vehicle cameras are configured to acquire video data comprising a plurality of frames. The video data, and more generally, the vehicle sensor data acquired by the vehicle sensors, are processed as part of one or more visual perception tasks that provide the environmental perception required for the vehicle to implement one or more DAS features, such as a cruise control function or a controlled-access highway cruise control feature.As part of the one or more visual perception tasks, one or more object classes and one or more object instances are determined within the vehicle sensor data, and thus within the video data. Based on the one or more object classes and the one or more object instances, a reduced number of frames is determined. This means that one or more frame intervals are determined based on the one or more object classes and the one or more object instances.
[0013] The one or more frame distances indicate a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames. In other words, the one or more frame distances indicate a number of frames that are similar with regard to the one or more object classes and the one or more object instances determined as part of the one or more visual perception tasks. For example, the one or more object classes and the one or more object instances determined within a frame n and the subsequent k frames indicate that these frames are similar. The subsequent frame n + k + 1, on the other hand, can no longer be considered similar due to the one or more object classes and the one or more object instances determined within that frame.Accordingly, the frame spacing in this example is k and only frame n and frame n+k+1 are determined as part of the reduced plurality of frames and stored accordingly.
[0014] For the frames within the frame spacing, object lists are generated that identify the one or more object classes and the one or more object instances identified within each corresponding frame, as well as their corresponding positions. In other words, for each frame similar to frame n, the object list identifies the object classes and object instances identified in each frame, allowing reconstruction of these frames based on the stored frame n and the information contained in the object lists. The object lists are thus stored instead of the actual frames, reducing the amount of memory required to store the video data.
[0015] The concept of frame spacing and object lists is in the Fig. 1. The Fig. 1 shows a plurality of frames 110, which frames 111 n up to 111 n+5 comprises a reduced plurality of frames 120, which frames 111 n , 111 n+3 and 111 n+5 and the object lists 121 n+1 , 121 n+2 and 121 n+4 and two frame intervals d frame,1 and d frame,2 . It is understood that the indices of the reference symbols indicate, on the one hand, the relationship between the frames of the plurality of frames 110 and, on the other hand, the relationship between the frames and object lists of the reduced plurality of frames 120. That is, the object list 121 n+2 corresponds to frame 111 n+2 and thus identifies frame 111 n+2 by identifying the object classes and object instances within the frame 111 n+2 .
[0016] The frame spacing d frame,1indicates that the frames 111 n up to 111 n+2 are considered similar with regard to the specific object classes and object instances within these frames. Accordingly, the frame spacing d frame,1 that the frames 111 n+1 and 111 n+2 are excluded from the reduced plurality of frames 120. Consequently, the reduced plurality of frames 120 includes only the object lists 121 n+1 and 121 n+2 instead of frames 111 n+1 and 111 n+2 , whereby the memory size of the reduced plurality of frames 120 is reduced compared to the memory size of the plurality of frames 110. The same applies to the frame spacing d frame,2 and for the corresponding frame 111 n+4 and the object list 121 n+4 .
[0017] It is understood that the number of Fig. 1 is merely exemplary. Since vehicle cameras can capture video data at frame rates of multiple frames per second, such as 20 to 60 frames per second, both the plurality of frames 110 and the reduced plurality of frames 120 may comprise dozens or hundreds of frames or more, with the reduced plurality of frames 120 comprising significantly fewer frames than the plurality of frames 110 in view of the frame reduction explained above.
[0018] By determining frame spacing based on the specific object classes and object instances already identified during visual perception tasks to enable DAS functions, and by using these frame spacings to reduce the storage size of video data, the data processing required to compress the video data can be reduced compared to regular video compression methods used outside of vehicles. In other words, the storage size of the reduced plurality of frames is reduced by reusing data processing that has already been performed for another functionality of the vehicle.
[0019] This general concept will now be explained with reference to the accompanying drawings, where the Fig. 2 provides a flowchart of the method for storing video data captured by one or more vehicle cameras installed in a vehicle. Fig. Figure 3 shows the vehicle including the majority of vehicle sensors and the vehicle control unit. Fig. 4 finally shows an example of the vehicle control unit in more detail.
[0020] It is understood that the dashed boxes in the Fig. 2 represent optional steps of the method 200.
[0021] The method 200 is configured to store video data comprising a plurality of frames 110 captured by one or more vehicle cameras installed in a vehicle, such as the vehicle 300 of the Fig. 3 - are installed, which is designed to provide at least driver assistance.
[0022] With brief reference to the Fig. 3, "vehicle 300" and more generally the term "vehicle" in the context of the present disclosure refers to any type of motor vehicle configured for the transport of persons and / or cargo. The engine of the vehicle 300 may be any type of engine, such as an electric motor or an internal combustion engine. For example, the vehicle 300 may be configured as shown in Fig. 3, a passenger car. However, it is understood that the vehicle 300 may also be a bus, a truck, or any other type of vehicle that includes one or more vehicle sensors 310 and a vehicle control unit 400 that enable the vehicle 300 to perform at least one driver assistance function. In other words, the vehicle control unit 300 and one or more sensors 310 may be configured to enable the vehicle 300 to provide vehicle control functionality that can perform at least one driver assistance function, i.e., Level 1 of the automated driving classification defined in the SAE International J3016 standard.That is, the vehicle 300 may be configured to provide at least one DAS function that performs the sustained and Operational Design Domain (ODD)-specific execution of either the vehicle lateral motion control subtask or the vehicle longitudinal motion control subtask of the dynamic driving task (DDT) (but not both simultaneously) in anticipation of the driver executing the remaining DDT.
[0023] For the purposes of this disclosure, ODD refers to the operating conditions for which a particular DAS function is specifically designed, including, but not limited to, environmental, geographic, and time-of-day constraints and / or the required presence or absence of certain traffic or roadway features.
[0024] For the purposes of this disclosure, the DDT includes all real-time operational and tactical functions required to operate the vehicle 300 in road traffic, with the exception of strategic functions such as trip planning and selection of destinations and waypoints.
[0025] It is understood that the vehicle 300 may be configured to enable higher levels of automated driving, such as partially automated driving—i.e., Level 2 or higher of the automated driving classification defined in SAE International's J3016 standard.
[0026] It is understood that the vehicle 300 can be configured to perform DAS functions of various driving automation levels, ie in particular also DAS functions of lower levels of automated driving, wherein at least one DAS function of the vehicle 300 provides driver assistance as defined in the SAE International standard J3016.
[0027] The one or more sensors 310 are configured to capture vehicle sensor data indicative of the surroundings of the vehicle 300. Accordingly, the vehicle sensor data provides environmental awareness to the one or more vehicle control modules, and thus to the vehicle 300, to enable at least one DAS function that provides driver assistance. For example, the vehicle sensor data captured by the one or more sensors 310 may provide the vehicle 300 with information about the position and size of other vehicles, lane markings, or traffic signs. To this end, the one or more sensors 310 may be radar sensors configured to emit radio waves to determine a distance, angle, and speed of objects in the surroundings of the vehicle based on the reflected radio waves.The one or more sensors 310 may be Light Detection and Ranging (LIDAR) sensors configured to emit laser beams to determine a distance, angle, and speed of objects in the surroundings of the vehicle 300 based on the reflected laser beams. The one or more sensors 310 may be vehicle cameras that capture video data of the vehicle's surroundings. The one or more sensors 310 may be thermal imaging cameras that capture images of the surroundings of the vehicle 300 based on infrared radiation. It is understood that LIDAR sensors, radar sensors, or vehicle cameras are merely given as examples of sensor types of the one or more sensors 310. The one or more sensors 310 may also be, for example, ultrasonic sensors.The one or more sensors 310 may also be Global Navigation Satellite System (GNSS) sensors configured to receive position data, such as satellite signals, for determining the position of the vehicle 300. Generally speaking, the one or more sensors 310 may be any type of sensor capable of sensing vehicle sensor data indicative of the surroundings of the vehicle 300. Furthermore, the one or more sensors 310 may additionally be any type of sensor capable of sensing odometry data of the vehicle 300, such as speed and acceleration. This sensing capability may be integrated into the sensor types discussed above or provided by dedicated motion sensors. It is further understood that the one or more sensors 310 may include multiple sensors of different sensor types.Furthermore, the one or more sensors 310 of the same type may have different characteristics, for example, by being configured to collect sensor data in different ranges, such as a near range, a mid-range, and a far range. For example, the vehicle 300 may include three near-range radar sensors at a front and a rear of the vehicle 300, a mid-to-far range radar sensor at the rear of the vehicle 300, a lidar sensor at the front of the vehicle 300, a rear-facing camera at the rear of the vehicle 300, a forward-facing camera at the front of the vehicle, a forward-facing camera on the rearview mirror, and a near-to-mid-range rear-facing radar sensor in each door-mounted exterior rearview mirror. It is understood that the vehicle 300 may include more or fewer vehicle sensors than in the . Fig. 3 and explained in the example above.
[0028] In step 210, the method 200 determines one or more object classes and one or more object instances within the vehicle sensor data, which, as described above, includes the video data acquired by the one or more vehicle cameras 310. The determination of the one or more object classes and one or more object instances is performed as part of one or more visual perception tasks that enable at least one DAS function to provide driver assistance.
[0029] For the purposes of the present disclosure, "visual perception task" refers to any type of task that identifies one or more object classes and object instances, i.e., individual instances of the object classes, within the vehicle sensor data collected by the one or more sensors 310. For example, the visual perception task may identify, within the video data provided by the one or more vehicle cameras included in the vehicle 300, whether the vehicle 300 is on a controlled-access highway, a limited-access road, a major arterial road, a residential street, or a parking lot. In this example, the one or more object classes correspond to the type of road on which the vehicle 300 may be located.Furthermore, the visual perception task may, for example, within the vehicle sensor data provided by a LIDAR sensor and one or more vehicle cameras included in the vehicle 300, identify other vehicles and the type of vehicle, lane markings and the type of lane marking, traffic signs and the type of sign, vulnerable road users (VRUs), and traffic lights and the display state of the traffic lights. Accordingly, the one or more object classes may correspond to any possible road user, any possible road traffic control device, and any possible lane marking, as well as any other type of element encountered in the driving environment of the vehicle 300 that is relevant for enabling at least one DAS function that provides at least driver assistance.More generally, the visual perception task can thus be any perception task that determines the class of objects and the instances of the different classes of objects in the environment of the vehicle 300, wherein the objects relate both to a determination of the general environment of the vehicle 300 and to a determination of individual elements in the environment of the vehicle 300.
[0030] It is understood that step 210 may already be performed as part of implementations of visual perception tasks and / or one or more DAS functions. Thus, the method 200 reuses the processing of the vehicle sensor data, and more specifically, the video data included in the vehicle sensor data, to generate and store the reduced plurality of frames, as explained generally above and in more detail below. Accordingly, step 210 does not result in any additional processing within the vehicle control unit 400 and thus does not impact the processing resources of the vehicle control unit 400 or impair its real-time processing.
[0031] Although method 200 relates to compressing video data based on object classes and object instances within the video data, it is further understood that the determination of the object classes and object instances in step 210 may be based on all vehicle sensor data used by the one or more visual perception tasks to detect object classes and object instances. In other words, the determination of the object classes and object instances within the video data may take into account vehicle sensor data from additional vehicle sensors 310, which may make the determination of the object classes and object instances more robust. As explained below with regard to step 220, this, in turn, may improve the determination of correspondences between frames and may thereby result in smaller storage sizes of the video data.
[0032] In step 220, the method 200 determines the reduced plurality of frames 120 based on the one or more object classes and the one or more object instances determined in step 210, as well as the plurality of frames 110 and one or more frame spacings, such as the one specified in the Fig. 1 shown frame intervals d frame1 , d frame,2 . As explained above, each frame spacing indicates a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames.
[0033] The one or more frame spacings may be based on similarities between the frames 111 n up to 111 n+5of the plurality of frames 110 with respect to the one or more object classes and the one or more object instances determined in step 210. For this purpose, step 220 comprises a step 221 in which the method 200 generates the one or more frame distances based on a change in the object classes and the object instances between consecutive frames of the plurality of frames compared to a change threshold. The change threshold indicates a percentage of object classes and object instances within a frame, such as frame 111 n+1 , the plurality of frames 110, the object classes and object instances within a previous frame, such as frame 111 n , the majority of frames 110. Assuming an exemplary change threshold of 80% and ten detected object instances and their corresponding object classes in frame 111n can frame 111 n+1 with nine detected objects corresponding to the objects in frame 111 n correspond, still as frame 111 n be viewed accordingly. The same applies to frame 111 n+2 , which may contain eight detected objects corresponding to the objects in frame 111 n In contrast, frame 111 includes n+3 possibly only seven detected objects corresponding to the objects in frame 111 n and may therefore not be considered as frame 111 n accordingly. Accordingly, the corresponding frame spacing d frame,1 that the frames 111 n+1 and 111 n+2are excluded from the reduced plurality of frames 120. More generally, the frame distances are determined in step 221 based on the percentage of detected object instances and their corresponding object classes that correspond to each other across a number of consecutive frames, i.e., that differ only in their position within the respective frames. In this context, the change threshold can be viewed as the degree of correspondence across a number of consecutive frames, below which the frames can no longer be considered to correspond to each other.
[0034] The generation of the reduced plurality of frames in steps 220 and 221 may further be based on a target storage size of the reduced plurality of frames 120. To this end, the frame intervals generated as part of step 220 may be determined in a manner that changes the respective frame intervals. That is, if the frame intervals generated in step 220 are lengthened, the storage size of the reduced plurality 120 decreases because more frames of the plurality of frames 110 are excluded from the reduced plurality 120. Conversely, if the frame intervals generated as part of step 220 are shortened, the storage size of the reduced plurality 120 increases because fewer frames of the plurality of frames 110 are excluded from the reduced plurality 120.In step 221, the lengthening of the frame spacing and thus the reduction of the storage size of the reduced plurality of frames 120 can be achieved by varying the change threshold, i.e., the change threshold can be proportional to the target storage size. In other words, the degree of correspondence between frames used to generate the reduced plurality of frames 120 from the plurality of frames 110 can be reduced in order to exclude more frames of the plurality of frames 110 from the reduced plurality of frames 120.
[0035] In step 230, the method 200 generates a corresponding object list for each frame 111 of the plurality of frames 110 that is not included in the reduced plurality of frames 120. Each object list identifies the one or more object classes and the one or more object instances determined within a corresponding frame 111, as well as their corresponding positions within the corresponding frame. That is, the method 200 uses the determination of the object classes and object instances performed in step 210 to generate object lists for each frame 111 of the plurality of frames 110 that is not included in the reduced plurality of frames 120.The object lists define each excluded frame 111 in terms of the object instances and their corresponding classes encompassed in the nearest preceding frame 111 of the reduced plurality 120, as well as the positions of these object instances within each excluded frame 111. Take frame 111 as an example. n+2 the Fig. 1, then the corresponding object list 121 identifies n+2 which in frame 111 n+2 visible object instances and their corresponding classes by referring to the corresponding object instances and their corresponding classes in frame 111 n . Furthermore, the object list identifies 121 n+2 the position of these object instances and their corresponding classes in frame 111 n+2 . Consequently, the method 200 reduces the memory size of the reduced plurality of frames 120 compared to the plurality of frames 110 by determining object lists that, as in the Fig. 1, replace the excluded frames of the plurality of frames 110. The object lists 121 are generated based on the processing performed in step 210 of the method 200 and thus based on the processing performed by the vehicle control unit 400 in each case to provide the environmental awareness required for one or more DAS functions.
[0036] Frames 111 that have been excluded from the reduced plurality of frames 120 can be regenerated based on their corresponding object lists 121 and their references to the nearest preceding frame 111 included in the reduced plurality of frames 120. For example, frame 111 n+4 in the Fig. 1 excluded from the reduced majority of frames 120, but can be based on frame 111 n+3 and the object list 121 n+4 be regenerated.
[0037] In step 240, the method 200 may generate one or more reference images based on a frame of the reduced plurality of frames 120 that corresponds to a beginning of a respective frame interval. Each reference image may correspond to an object instance within the frame that corresponds to the beginning of the respective frame interval. That is, to enable the regeneration of the frames 111 excluded from the reduced plurality of frames 120, the method 200 may generate reference images of object instances included in multiple frames 111 of the plurality of frames 120.
[0038] It is understood that in implementations of the method 200 that implement the reference image generation of step 240, even the frames at the beginning and end of the frame intervals, i.e., the frames that are intended to be included in the reduced plurality of frames 120, may not be included in the reduced plurality of frames 120. Rather, in step 270, even for these frames, only object lists may ultimately be stored together with a library of reference images from which all frames of the plurality of frames 110 can be regenerated. With such exemplary implementations, even further reduced memory sizes can be achieved.
[0039] The method 200 may additionally include measures to protect the information contained in the reduced plurality of frames 120. For this purpose, the method 200 may include a step 250 in which the method 250 may encrypt the reduced plurality of frames 120 and the object lists 121. The reduced plurality of frames 120 and the object lists 121 may be encrypted using any encryption method suitable for preventing unauthorized access to the reduced plurality of frames 120, such as Advanced Encryption Standard (AES) 128 or AES-256.
[0040] The method 200 may additionally include measures to protect the privacy of persons visible in the reduced plurality of frames 120. For this purpose, the method 200 may include a step 260 in which the method 200 obscures the faces of persons visible in each frame of the reduced plurality of frames 120.
[0041] Finally, in step 270, the method 200 stores the reduced plurality of frames 120 and the object lists 121. In this context, it is understood that the object lists 121 can also be considered part of the reduced plurality of frames 120, in which the object lists 121 can replace frames 111 of the plurality of frames 120 that were excluded from the reduced plurality 120 based on the frame distances explained above. This concept is described in the Fig. 1. Furthermore, it is understood that the frames 111 of the plurality of frames 110 included in the reduced plurality of frames 120, such as the frames 111 n , 111 n+3 and 111 n+5 in the Fig. 1, can be stored unchanged or can be compressed by a suitable compression algorithm if this is necessary in view of the storage size requirements in the main memory.
[0042] Step 270 may include a step 271 in which the method 200 may store the one or more reference images generated in step 240. As explained above, in steps 270 and 271, frames 111 determined in step 220 to be included in the reduced plurality of frames 120, such as frames 111 n , 111 n+3 and 111 n+5 in the Fig. 1, in the form of reference images and object lists 121, if necessary in view of the memory size requirements.
[0043] The reduced plurality of frames 120 and the object lists 121 can then be used to regenerate the video data based on the references in the object lists 121 to object instances and their classes in frames 111 of the plurality of frames 110 included in the reduced plurality 120. During regeneration, for example, generative adversarial networks (GANs) or other deep learning techniques can be used to realistically regenerate frames 111 not included in the reduced plurality of frames 120.
[0044] In summary, the method 200 provides a way to store video data captured by vehicle cameras in a vehicle in a compressed manner that reuses the data processing performed by visual perception tasks and thus reduces the processing overhead associated with storing the compressed video data.
[0045] The Fig. 4 shows the vehicle control unit 400 configured to perform the method 100. The vehicle control unit 400 may include a processor 410, a graphics processing unit (GPU) 420, a vehicle processing system 430, a memory 440, a removable storage 450, a memory 460, a cellular interface 470, a global navigation satellite system (GNSS) interface 480, and a communications interface 490.
[0046] Processor 410 may be any type of single-core or multi-core processing unit using a reduced instruction set (RISC) or a complex instruction set (CISC). Example RISC processing units include ARM-based cores or RISC-V-based cores. Example CISC processing units include x86-based cores or x86-64-based cores. Processor 410 may execute instructions that cause vehicle control unit 400 to perform method 200. Processor 410 may be directly coupled to one of the components of vehicle control unit 400 or may be directly coupled to memory 430, GPU 420, and a device bus.
[0047] The GPU 420 may be any type of processing unit optimized for processing graphics-related instructions or, more generally, for parallel processing of instructions. As such, the GPU 420 may be configured to generate a display of information, such as ADAS information or telemetry data, to a driver of the vehicle—e.g., via a head-up display (HUD) or a display located in the driver's field of view. The GPU 420 may be coupled to the HUD and / or the display via connection 420C. The GPU 420 may further execute at least a portion of the method 100 to enable rapid parallel processing of instructions related to the method 100. It should be noted that in some embodiments, the processor 410 may determine that the GPU 420 does not need to execute instructions related to the method 200.The GPU 420 may be directly coupled to one of the components of the vehicle control unit 400 or may be directly coupled to the processor 410 and the memory 430. In some embodiments, the GPU 420 may also be coupled to the device bus.
[0048] The vehicle processing system 430 may be any type of system-on-chip configured to perform trillions of operations per second (TOPS) to enable the vehicle control unit 400 to implement one or more ADAS while driving. The vehicle processing system 430 may interface only with the processor 410 or may interface with other devices via the system bus. For example, the vehicle processing system 430 may execute the instructions related to the one or more vehicle sensor data processing modules and the one or more vehicle control modules.
[0049] Memory 440 may be any type of fast memory that enables the processor 410, GPU 420, and vehicle processing system 430 to store instructions for rapid retrieval during instruction processing, as well as to cache and buffer data. Memory 440 may be a unified memory coupled to the processor 410, GPU 420, and vehicle processing system 430 to enable allocation of memory 440 as needed to the processor 410, GPU 420, and vehicle processing system 430. Alternatively, the processor 410, GPU 420, and vehicle processing system 430 may be coupled to separate processor memory 440a, GPU memory 440b, and vehicle processing system memory 440c.
[0050] Removable storage 450 may be a storage device that is removably coupled to vehicle control unit 400. Examples include a digital versatile disc (DVD), a compact disc (CD), a universal serial bus (USB) storage device such as an external SSD or magnetic tape. It should be noted that removable storage 450 may store data such as instructions of method 200, vehicle sensor data, intermediate data, and / or vehicle control data, or the storage may be omitted.
[0051] Memory 460 may be a storage device that enables the storage of program instructions and other data. Memory 460 may be, for example, a hard disk drive (HDD), a solid state disk (SSD), or another type of non-volatile memory. For example, the instructions of method 100, vehicle sensor data, intermediate data, and / or vehicle control data may be stored in memory 460.
[0052] Removable storage 450 and memory 460 may be coupled to processor 410 via the system bus. The system bus may be any type of bus system that enables processor 410 and optionally GPU 420 and vehicle processing system 430 to communicate with the other devices of vehicle control unit 400. Bus 440 may be, for example, a Peripheral Component Interconnect Express (PCIe) bus or a Serial AT Attachment (SATA) bus.
[0053] The cellular interface 470 may be any type of interface that enables the vehicle control unit 400 to communicate over a cellular network, such as a 4G network or a 5G network.
[0054] The GNSS interface 480 may be any type of interface that enables the vehicle control unit 300 to receive position data provided by a satellite network, such as the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), or Galileo. The position data may be one of the types of vehicle sensor data within the scope of the present disclosure.
[0055] The communication interface 490 may enable the vehicle control unit 400 to interface with external devices, either directly or via a network. For example, the communication interface 480 may enable the vehicle control unit 400 to interface with a wired or wireless network such as Ethernet, Wi-Fi, a Controller Area Network (CAN) bus, or any bus system suitable for vehicles. For example, the vehicle control unit 400 may be coupled to the one or more vehicle sensors 310 to receive vehicle sensor data.
[0056] The vehicle control unit 400 may be integrated into the vehicle 300, e.g., underneath the vehicle cabin, under the dashboard, or in the trunk of the vehicle 300.
[0057] The invention can be explained in more detail using the following examples.
[0058] In one example, a method for storing video data comprising a plurality of frames and captured by one or more vehicle cameras installed in a vehicle configured to perform at least one driver assistance comprises: determining one or more object classes and one or more object instances within vehicle sensor data, wherein the vehicle sensor data comprises the video data captured by the one or more vehicle cameras, wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least the driver assistance, determining a reduced plurality of frames based on the one or more object classes, the one or more object instances, the plurality of frames, and one or more frame intervals,wherein each frame spacing indicates a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames, for each frame of the plurality of frames not included in the reduced plurality of frames, generating a corresponding object list, each object list identifying the one or more object classes and the one or more object instances determined within a corresponding frame and their corresponding positions within the corresponding frame, and storing the reduced plurality of frames and the object lists.
[0059] The example method may further comprise, for each frame pitch, generating one or more reference images based on a frame of the reduced plurality of frames that corresponds to a start of a respective frame pitch, wherein each reference image corresponds to an object instance within the frame that corresponds to the start of the respective frame pitch, wherein storing the reduced plurality of frames and the object lists may further comprise storing the one or more reference images.
[0060] The example method may further comprise generating the one or more frame distances based on a change in object classes and object instances between consecutive frames of the plurality of frames compared to a change threshold, wherein the change threshold may indicate a percentage of object classes and object instances within a frame of the plurality of frames that corresponds to object classes and object instances within a previous frame of the plurality of frames.
[0061] In the example method, generating the one or more frame distances may be further based on a target memory size of the reduced plurality of frames, wherein the change threshold may be proportional to the target memory size.
[0062] The example method may further include encrypting the reduced plurality of frames and the object lists.
[0063] The example method may further include blurring faces of people visible in each frame of the reduced plurality of frames.
[0064] In one example, a vehicle control unit comprises at least one processing unit and a memory coupled to the at least one processing unit and configured to store machine-readable instructions. The machine-readable instructions cause the at least one processing unit to: determine one or more object classes and one or more object instances within vehicle sensor data, wherein the vehicle sensor data comprises the video data acquired by the one or more vehicle cameras, wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least driver assistance, determine a reduced plurality of frames based on the one or more object classes and the one or more object instances, the plurality of frames, and one or more frame spacings,wherein each frame spacing indicates a number of frames between frames of the plurality of frames to be excluded from the reduced plurality of frames, for each frame of the plurality of frames not included in the reduced plurality of frames, generates a corresponding object list, each object list identifying the one or more object classes and the one or more object instances determined within a corresponding frame, as well as their corresponding positions within the corresponding frame, and stores the reduced plurality of frames and the object lists.
[0065] In the example vehicle control unit, the machine-readable instructions may further cause the at least one processing unit to perform one of the preceding example methods.
[0066] In one example, a vehicle includes a plurality of vehicle sensors and the foregoing example vehicle control unit.
[0067] The foregoing description has been provided to illustrate a method and system for storing video data in a vehicle. It should be understood that the description is in no way intended to limit the scope of the present disclosure to the precise embodiments discussed throughout the description. Rather, those skilled in the art will appreciate that the examples of the present disclosure may be combined, modified, or abbreviated without departing from the scope of the present disclosure as defined by the following claims. List of reference symbols 110 Multiple frames 111 frames 120 reduced majority of frames 121 Object list 200 procedures 210-271 Procedural steps 300 vehicles 310 vehicle sensor 400 vehicle control unit 410 CPU 420 GPU 420c connection 430 Vehicle Processing System 440 RAM 450 removable storage 460 memory 470 mobile radio interface 480 GNSS interface 490 Communication interface
Claims
[1] A method (200) for storing video data comprising a plurality of frames (110) captured by one or more vehicle cameras (310) installed in a vehicle (300) configured to perform at least one driver assistance, the method comprising: Determining (210) one or more object classes and one or more object instances within vehicle sensor data, wherein the vehicle sensor data comprises the video data acquired by the one or more vehicle cameras (310), wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least the driver assistance; Determining (220) a reduced plurality of frames (120) based on the one or more object classes, the one or more object instances, the plurality of frames (110) and one or more frame intervals (d frame,1 , d frame,2 ), where each frame interval (d frame,1 , d frame,2 ) a number of frames between frames (111 n -111 n+5 ) of the plurality of frames (110) to be excluded from the reduced plurality of frames (120), wherein determining (220) the reduced plurality of frames (120) and the one or more frame spacings (d frame,1 , d frame,2 ) generating (221) the one or more frame distances (d frame,1 , d frame,2) based on a change in object classes and object instances between consecutive frames of the plurality of frames (110) compared to a change threshold, wherein the change threshold comprises a percentage of object classes and object instances within a frame (111 n -111 n+5 ) of the plurality of frames (110) which indicates object classes and object instances within a previous frame (111 n -111 n+5 ) corresponds to the plurality of frames (120); for each frame (111 n -111 n+5 ) of the plurality of frames (110) not included in the reduced plurality of frames (120), generating (230) a corresponding object list (121 n+1 , 121 n+2 , 121 n+4 ), where each object list (121 n+1 , 121 n+2 , 121 n+4) the one or more object classes and the one or more object instances that are contained within a corresponding frame (111 n -111 n+5 ) and their corresponding positions within the corresponding frame (111 n -111 n+5 ) identified; and Saving (270) the reduced plurality of frames (120) and the object lists. [2] The method (200) of claim 1, further comprising: Generate (240), for each frame interval (d frame,1 , d frame,2 ), one or more reference images based on a frame (111 n , 111 n+3 ) of the reduced plurality of frames (120) corresponding to a start of a respective frame interval (d frame,1 , d frame,2 ), where each reference image corresponds to an object instance within the frame (111 n , 111 n+3 ) corresponding to the beginning of the respective frame interval (d frame,1 , d frame,2 ) corresponds, wherein the storing (270) of the reduced plurality of frames (120) and the object lists (121 n+1 , 121 n+2 , 121 n+4 ) further comprises storing (271) the one or more reference images. [3] Method (200) according to one of the preceding claims, wherein: generating (221) the one or more frame intervals (d frame,1 , d frame,2 ) is further based on a target memory size of the reduced plurality of frames (120), where the change threshold is proportional to the target memory size. [4] Method (200) according to one of the preceding claims, further comprising encrypting (250) the reduced plurality of frames (120) and the object lists (121 n+1 , 121 n+2 , 121 n+4 ) comprehensively. [5] Method (200) according to one of the preceding claims, further comprising obscuring (260) faces of persons appearing in each frame (111n , 111 n+3 , 111 n+5 ) of the reduced majority of frames (120) are visible. [6] Vehicle control unit (400) comprising: at least one processing unit (410, 420, 430); and a working memory (440, 450) coupled to the at least one processing unit (410, 420, 430) and configured to store machine-readable instructions, wherein the machine-readable instructions cause the at least one processing unit (410, 420, 430) to perform the following: Determining one or more object classes and one or more object instances within vehicle sensor data, wherein the vehicle sensor data comprises video data comprising a plurality of frames (110) and captured by one or more vehicle cameras (310), wherein the determination of the one or more object classes and the one or more object instances is performed as part of one or more visual perception tasks that enable at least one driver assistance in a vehicle (300); Determining a reduced plurality of frames (120) based on the one or more object classes, the one or more object instances, the plurality of frames (110) and one or more frame intervals (d frame,1 , d frame,2 ), where each frame interval (d frame,1 , d frame,2 ) a number of frames between frames (111 n -111 n+5) of the plurality of frames (110) to be excluded from the reduced plurality of frames (120), wherein determining (220) the reduced plurality of frames (120) and the one or more frame spacings (d frame,1 , d frame,2 ) generating (221) the one or more frame distances (d frame,1 , d frame,2 ) based on a change in object classes and object instances between consecutive frames of the plurality of frames (110) compared to a change threshold, wherein the change threshold comprises a percentage of object classes and object instances within a frame (111 n -111 n+5 ) of the plurality of frames (110) which indicates object classes and object instances within a previous frame (111 n -111 n+5 ) corresponds to the plurality of frames (120); for each frame (111 n -111 n+5) of the plurality of frames (110) not included in the reduced plurality of frames (120), generating a corresponding object list (121 n+1 , 121 n+2 , 121 n+4 ), where each object list (121 n+1 , 121 n+2 , 121 n+4 ) the one or more object classes and the one or more object instances determined within a corresponding frame (111n-111n+5), as well as their corresponding positions within the corresponding frame (111 n -111 n+5 ) identified; and Saving the reduced majority of frames (120) and the object lists (121 n+1 , 121 n+2 , 121 n+4 ). [7] The vehicle control unit (400) of claim 6, wherein the machine-readable instructions further cause the at least one processing unit (410, 420, 430) to execute the method (200) of any one of claims 2 to 5. [8] A vehicle (300) comprising a plurality of vehicle sensors (310) and the vehicle control unit (400) according to any one of claims 6 and 7.
Citation Information
Patent Citations
Video processing
EP3739503A1
High-speed image readout and processing
US20190208136A1
Methods and systems for real-time data reduction
US20210142068A1