Device and method for recording a video sequence in a vehicle

The device and method automate the conversion of vehicle-recorded video to portrait format, addressing the inconvenience of landscape footage by using computer vision and gesture recognition to optimize video sequences for social media platforms.

DE102024120899A1Pending Publication Date: 2026-01-29BAYERISCHE MOTOREN WERKE AG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
DE102024120899
Authority / Receiving Office
DE · DE
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-23
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing video recording systems in vehicles often produce landscape format footage, which is inconvenient for social media platforms requiring portrait format, necessitating cumbersome post-processing to adapt the videos for platforms like Instagram Reels, TikTok, or YouTube Shorts.

Method used

A device and method that utilizes an interior vehicle camera to detect occupants, automatically select and crop images to portrait format, and generate video sequences optimized for social media platforms, using computer vision and gesture recognition to simplify the recording process.

Benefits of technology

Enables easy production of vertical video formats suitable for social media, reducing the time and effort required for post-processing and ensuring optimal display on mobile devices without the need for manual editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

In a device and a method for recording a video sequence in a vehicle, at least one image depicting at least part of the vehicle interior is captured, and corresponding image data is generated. The image data is processed, and based on this data, at least one vehicle occupant is detected. At least one vehicle occupant is selected as the subject. An image format is selected. A video sequence containing images of the subject is recorded by the camera and output in the selected image format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] In a device and a method for recording a video sequence in a vehicle, at least one image depicting at least part of the vehicle interior is captured, and image data corresponding to the image is generated. The image data is processed, and based on this data, at least one vehicle occupant is detected.

[0002] Photos and videos of vehicle occupants can be taken using the vehicle's interior camera. However, it's not possible to record these videos in a desired format. For various platforms, such as Instagram Reels, TikTok, or YouTube Shorts, videos in portrait format are advantageous. Subject tracking would also be beneficial, especially when recording a vehicle occupant's performance.

[0003] Several software solutions exist for adapting videos to portrait format and using subject tracking. These programs allow you to crop videos to a vertical format and offer powerful tools for tracking objects within the video. The process typically begins with selecting the sequence settings to crop the video to portrait format (for example, 1080 x 1920 pixels). The video is then imported, and the desired crop is defined. By using transformation tools, the main subject of the video can be centered and adjusted within the frame. Some software solutions include dedicated tracking tools that allow you to mark and track an object or person in the video. The software analyzes the movements of the marked object and adjusts the camera accordingly, ensuring the subject remains in focus.Manual keyframes can also be set to precisely adjust the movements.

[0004] In addition to professional desktop applications, there are also mobile apps that provide basic functions for video cropping and subject tracking. While these are less precise, they are well-suited for quick and easy adjustments.

[0005] Document DE 10 2022 103 077 A1 discloses a system for taking a photograph or recording a video using cameras integrated into the vehicle.

[0006] Document WO 2023 / 031890 A1 discloses a method for cropping videos for display on screens that do not conform to standard video dimensions, including picture-in-picture (PIP) displays such as YouTube mini-videos. The cropping is performed by detecting at least one "object of interest".

[0007] However, it is time-consuming to edit videos recorded in a vehicle and then upload them to a platform.

[0008] The object of the invention is to provide a device and a method for recording a video sequence in a vehicle, by which a video sequence adapted for further use can be easily produced.

[0009] This problem is solved by a device having the features of claim 1 and by a method having the features of the independent method claim. Advantageous embodiments are specified in the dependent claims.

[0010] The device with the features of claim 1 allows the video sequence to be easily output in the desired image format. Based on the image data, the processing unit can detect a vehicle occupant and preferably select that occupant as the subject. Based on the detected vehicle occupants, the processing unit can also suggest a subject via an output unit, which is then confirmed or rejected by input from a vehicle occupant. Preferably, the image captured by the camera includes a portion of the vehicle's interior, encompassing at least a portion of the driver's seat and / or a portion of the front passenger seat, such that at least one first vehicle occupant sitting in the driver's seat, a second vehicle occupant sitting in the front passenger seat, and / or at least one vehicle occupant in a rear seat of the vehicle are visible in the image.

[0011] The selected image format can be portrait, resulting in a so-called vertical video. YouTube videos are typically recorded and edited in landscape format, while portrait format is more suitable for TikTok videos, Snapchat, and Instagram Stories. Vertical videos are optimized to fit the vertical screens of mobile devices without needing to be rotated, thus creating a better user experience on mobile devices. A striking, full-screen video can prevent users from scrolling past content in their Instagram feed. For example, videos on Instagram's IGTV can be up to ten minutes long for most accounts and up to an hour for larger accounts. In contrast, TikTok videos are limited to a maximum of 60 seconds.

[0012] Interior cameras in vehicles typically record images in landscape format. However, as just explained, this is inconvenient for various applications. Converting the footage afterward is often cumbersome and time-consuming. In particular, different platforms require different aspect ratios. Instagram Stories, IGTV, and TikTok require a 9:16 aspect ratio, while Snapchat needs a 4:5 aspect ratio.

[0013] The processing unit can also be designed to determine, based on the image data, the positions of vehicle occupants within the vehicle's interior. This allows for the simple identification, suggestion, and / or selection of an image section for generating the output video sequence.

[0014] Furthermore, the processing unit can be configured to control the camera and / or the processing of the image data in such a way that the generation of the video sequence is initiated by an input from a vehicle occupant via an input unit in the vehicle. This allows the start time of the video sequence to be easily determined.

[0015] Furthermore, a pre-defined purpose can be selected before or during the start of recording. Selecting the purpose allows for easy selection of an associated image format, aspect ratio, and / or resolution. This simplifies the recording of video sequences in a vehicle.

[0016] It is particularly advantageous if the video sequence is output in a vertical image format. This allows the output video sequence to be easily adapted for social media.

[0017] A cropped section can also be extracted from the overall image captured by the camera to generate the video sequence. This cropped section focuses on the driver, front passenger, and / or a rear-seat occupant. A section of the image captured by the camera is thus automatically selected and / or defined, preferably using a facial recognition algorithm, by the processing unit. This allows the video sequence to focus on a suitable subject. Preferably, the camera does not have variable magnification. It is particularly advantageous if the selected image section is suggested to the vehicle occupant. The system can also suggest and / or define an expansion of the image section to include another vehicle occupant or a change to the image section to focus on a different vehicle occupant.

[0018] The processing unit can also be equipped to perform a computer vision procedure for person detection and / or motion detection of a vehicle occupant and / or for defining an image section containing the subject. This allows the vehicle occupants, and in particular their heads, to be easily identified.

[0019] Furthermore, the processing unit can be equipped to utilize a computer vision method to dynamically crop the images captured by the camera to generate the video sequence with the subject. This allows for the easy creation and output of a high-quality video sequence.

[0020] Furthermore, the image can be cropped to focus on the driver if the driver is the sole occupant. This allows for a particularly simple selection of the motif. Preferably, the driver's head is chosen and / or specified as the motif.

[0021] It is advantageous if the processing unit is equipped to use a computer vision method to detect vehicle occupants. This enables simple and reliable identification of the vehicle occupants.

[0022] The processing unit can also be designed to select the subject based on the viewing direction of each vehicle occupant and / or based on an input to start the video sequence to be recorded. This allows for simple, automated subject selection.

[0023] Alternatively or additionally, the processing unit can be designed to determine, based on the image data for subject selection, which vehicle occupant has activated the function to record the video sequence using a gesture recognition method and / or a body pose prediction method. This allows a dynamic decision to be made, based on information about the actions of the driver and / or the passenger, as to which subject is automatically selected for recording the video sequence or suggested for selection by a vehicle occupant.

[0024] For example, the processing unit automatically selects the passenger as the subject if the passenger has their hand on a touchscreen or other input device of the vehicle through which the function to record the video sequence has been activated.

[0025] Alternatively or additionally, the processing unit can be designed to use a gesture recognition method and / or a body pose prediction method, based on the image data, to determine which vehicle occupant(s) exhibit unusual body activity and select these occupants as the subject. Essentially, the body pose of each occupant is predicted in 3D based on a 2D infrared image, resulting in a 3D skeleton with the important joints. The presence of an occupant's hand in a 3D area near the control element would enable the assignment of a subject, or the selection of one subject from several possibilities.Unusual physical activity of a vehicle occupant can be detected, in particular, if the occupant's movement increases or decreases immediately before or immediately after the video recording function is activated. The period immediately before and / or immediately after activation is within the range of + / - 10 seconds.

[0026] To record a video sequence, the processing unit can also suggest one or more subjects for selection. Furthermore, the processing unit can suggest several purposes, each specified by a social media platform, for selection. The selection can be made, in particular, via an input unit in the vehicle. By selecting a social media platform, the upload of the generated video sequence to a pre-configured user account can also occur automatically or after additional confirmation by a vehicle occupant.

[0027] The processing of image data by the processing unit can also be carried out using a neural network. This enables simple and efficient processing of the image data for recording a video sequence in the vehicle.

[0028] The camera can be integrated into the vehicle's interior rearview mirror or its housing, mounted above the driver and / or front passenger in the roof area, and / or in the B-pillar and / or A-pillar on the driver's and / or passenger's side, and / or on a central display unit of the vehicle. This allows for the classification of not only the driver and their hand(s), but also other objects inside the vehicle that may affect the driver, such as other people and / or animals, as well as potential and probable sources of interference. Furthermore, it is advantageous if the camera's field of view has a detection angle of 100° to 150°, and the camera is preferably a monocular camera and / or an RGB-IR camera.This allows the processing unit to perform image processing based on the captured IR images. Using IR images is advantageous because these images are independent of ambient light. Alternatively, the camera can also be a dedicated IR camera and / or a stereo camera (3D camera). The processing of the camera's image data can take place in a computer vision system provided by the processing unit. Using an RGB-IR camera combines the ability to be independent of ambient light with the ability to provide RGB information when visible light is available.

[0029] A second aspect concerns a method for recording a video sequence in a vehicle with the features of the dependent method claim. This method has the same advantages as the claimed device. In particular, the method can be further developed with the features of the dependent claims directed to the device.

[0030] Exemplary embodiments of the invention are explained in more detail below with reference to the figures. These show: Fig. 1. A perspective schematic representation of a vehicle cockpit; Fig. 2 in schematic representation a first image of the interior of the vehicle taken by an interior camera of the vehicle; Fig. 3. In schematic representation, a section of the image according to Fig. 1 with a first motif in a first image format; Fig. 4 in schematic representation a section of the image according to Fig. 1 with a second motif in a second image format, and Fig. 5. A schedule for recording a video sequence in the vehicle.

[0031] Fig. Figure 1 shows a perspective schematic representation of the cockpit 100 of a vehicle 102. A driver 104 is seated in a driver's seat 106 of the vehicle 102. A display 108 for showing information when the passenger airbag 146 of the vehicle 102 is deactivated is arranged, by way of example, in the area of ​​a center console 110 of the cockpit 100. The cockpit 100 also includes a central information display (CID) 112, a head-up display 114, and a graphic instrument cluster 116, which are arranged in a dashboard 118 of the vehicle 102. The dashboard 118 also includes air outlet nozzles 120 of an air conditioning system of the vehicle 102. Fig. Figure 1 shows, by way of example, air outlet nozzles 120 in the area of ​​the center console 110 and on one passenger side. The cockpit 100 also includes a so-called multifunction steering wheel 122 with a first right-hand section 124 with primary controls and a second left-hand section 125 with secondary controls, a gear selector 126, pedals 128, and an input unit 130 with a rotary dial and push-button function and / or a touch input panel. This input unit 130 is also referred to as the Ergo Commander.

[0032] A camera 134 is arranged in the upper area of ​​the windshield 132 of the vehicle 102. The camera 134 is arranged and oriented such that it captures images 200 depicting an area of ​​the interior of the vehicle 102 that includes at least part of the driver's seat 106 and the passenger seat 202. An example image 200 from the camera 134 is shown in Fig. Figure 2 shows that the camera 134 is specifically designed to capture several successive images as a video sequence, generate corresponding image data 200 for each image, and transmit this data to a control unit 136. The camera 134 has a field of view of 120 degrees. In other embodiments, the camera 134 can also have a field of view in the range of 100 degrees to 150 degrees or more.

[0033] The control unit 136 has data outputs 138 and data inputs 140, which serve to connect to other units of the vehicle 102, for example with other sensors, input and output units and control units 142 of assistance systems and / or safety systems.

[0034] The control unit 136 also includes a communication module 144, which is designed to establish a connection with a telecommunications network, in particular a mobile communications network.

[0035] Based on the image data from camera 134, the control unit 136 can detect at least one vehicle occupant as the subject. Furthermore, the control unit 136 can select an image format. The selection of the image format can be specified, in particular, by selecting and / or presetting the purpose of a video sequence to be generated. Alternatively or additionally, the image format can be defined and the defined image format selected by the control unit 136. Based on the image data, a video sequence containing images of the subject is generated using sequentially captured images from camera 134 and output in the selected image format.

[0036] The image format can specify the basic orientation (portrait or landscape), a specific aspect ratio and / or a resolution of the images in the video sequence.

[0037] The resulting video sequence can then be transmitted to a mobile device belonging to a vehicle occupant or directly to a platform and / or database, preferably a social media platform or a social media platform's database. This can be done, in particular, using the communication module 144 of the vehicle 102. Such a social media platform could be, for example, one of the following: Facebook, Facebook Messenger, Instagram, YouTube, XING, LinkedIn, Pinterest, Snapchat, TikTok, X (formerly Twitter), Threads, WhatsApp, Vimeo, Reddit, Twitch, Tumblr, WeChat, Douyin, QQ, Telegram, and others.

[0038] One or more possible images featuring one or more vehicle occupants can be suggested to the vehicle occupant(s) for selection. A section of the image containing the respective image can be displayed, in particular on a display unit 112, 114, or 116 of vehicle 102, and offered for selection. Furthermore, the vehicle occupant(s) can be offered a purpose for recording the video sequence or several options for selecting the purpose. In particular, the purpose can be displayed as uploading to a specific social media platform.

[0039] The specific image format is then stored and assigned to the selected purpose. The control unit 136 then outputs the video sequence in that specific image format during or after recording.

[0040] Furthermore, in other embodiments, the field of view of the camera 134 can have a detection angle in the range of 100° to 150°, wherein the camera 134 is preferably a monocular camera and / or an RGB-IR camera. This allows, in particular, image processing by the control unit 136 based on the captured IR images. The use of IR images is advantageous because these images are independent of ambient light. Alternatively, the camera 134 can also be a pure IR camera and / or a stereo camera (3D camera). This allows a suitable camera 134 to be used in the interior of the vehicle 102.

[0041] Fig. Figure 2 shows a schematic representation of an image 200 of the interior of vehicle 102 taken by the interior camera 134. The camera 134 is directed towards the driver's seat 106, a passenger seat 202, and a rear bench seat 204 of vehicle 102. In the Fig. In the situation shown in Figure 2, the driver 104 is sitting in the driver's seat 106 and the other occupant 206 is sitting in the passenger seat 202. Thus, in the present embodiment, the other occupant is the passenger 206.

[0042] On the rear seat 204 sits a child 210 in a child seat 212. The driver 104 wears a smartwatch on his right hand and holds a smartphone 218 in his right hand 150.

[0043] When processing the image data from the image 200 captured by camera 134, the control unit 136 detects at least the driver 104, the passenger 206, and the child 210 in the child seat as possible subjects. To record a video sequence, the control unit 136 can suggest one or more subjects for selection. Furthermore, the control unit 136 can suggest several purposes, each specified by a social media platform. The selection can be made, in particular, via a touchscreen 112. By selecting a social media platform, the upload of the generated video sequence to a pre-configured user account can also occur automatically or after additional confirmation by a vehicle occupant.

[0044] The selection of the subject or a suggested subject can also be determined by the control unit 136 based on input via an input unit 112 of the vehicle 102 for recording a video sequence. Alternatively or additionally, the control unit 136 can be configured to determine, based on the image data from the camera 134, which vehicle occupant 104, 208, or 210 has activated the function for recording the video sequence, using a gesture recognition method and / or a body pose prediction method. This allows for a dynamic decision regarding which subject is automatically selected for recording the video sequence or suggested for selection by a vehicle occupant 104, 208, or 210.

[0045] For example, the control unit 136 automatically selects the passenger as the subject if the passenger 208 has their hand on a touch display 112 or another input unit of the vehicle 102, via which the function to record the video sequence has been activated.

[0046] Alternatively or additionally, the control unit 136 can be configured, based on the image data for subject selection, to determine, using a gesture recognition method and / or a body pose prediction method, which vehicle occupant 104, 208, 210 and / or which vehicle occupants 104, 208, 210 exhibit unusual physical activity and to select these vehicle occupants 104, 208, 210 or these vehicle occupants 104, 208, 210 as the subject. Unusual physical activity is detected in particular when the movement of a vehicle occupant 104, 208, 210 increases or decreases immediately before or immediately after the activation of the function to record the video sequence. The period for immediately before and / or immediately after activation is in the range of + / - 10 seconds.

[0047] Fig. Figure 3 shows a schematic representation of a section of the image. Fig. 1. The first subject is the driver (104). The video sequence is then output in a selected image format with the specifications portrait and aspect ratio of 9:16, preferably transferred to a selected social media platform.

[0048] Fig. Figure 4 shows a schematic representation of a section of the image. Fig. 1. A second subject in a second image format. The image format is defined as a portrait format for Snapchat with an aspect ratio of 4:5. The second subject is the passenger 208 and the child 210 in the child seat 212. The video sequence is thus output with a selected image format with the specification portrait format and aspect ratio 4:5, preferably transferred to a selected social media platform.

[0049] Fig. Figure 5 shows a process flow for recording a video sequence in a vehicle 102. The process is started in step S100. In step S102, an image is captured with camera 134, and corresponding image data is transmitted to the control unit 136. In step S104, the control unit 136 detects at least one vehicle occupant as the subject. In step S106, the processing unit 136 selects an image format. Subsequently, in step S108, the processing unit 136 generates a video sequence of images of the subject from the image data of pictures taken successively with camera 134 and outputs it in the selected image format. The process ends in step S110.

[0050] In the based on the Fig. 1, Fig. 2, Fig. 3, Fig. 4 to Fig. In the embodiments described in section 5, the camera 134 and the control unit 136 form a device for recording a video sequence in a vehicle 102. Further embodiments are described in the Fig. 1, Fig. 2, Fig. 3, Fig. 4 to Fig. The five elements and features shown and mentioned in the preceding description can be part of this device. The control unit 136 is also referred to as the processing unit. Reference symbol list 100 Cockpit 102 vehicles 104 drivers 106 Driver's seat 108 ads 110 Center console 112, 114, 116 Display element 118 Dashboard 120 air outlet nozzles 122 Steering wheel 124 Controls on the right side of the steering wheel 125 Controls left steering wheel area 126 gear selector 128 Pedals 130 Input unit 132 Windscreen 134 Camera 136 Control unit 138 Data output 140 Data input 142 Airbag control unit 144 Communication module 146 Passenger airbag 150 right hand 152 left hand 200, 300, 400 image 202 Passenger seat 204 Rear seat 206 passengers 210 children 212 Child seat 218 Smartwatch 220 Smartphone S100 to S110 process steps QUOTES INCLUDED IN THE DESCRIPTION

[0000] This list of documents cited by the applicant was automatically generated and is included solely for the reader's convenience. The list is not part of the German patent or utility model application. The DPMA accepts no liability for any errors or omissions. Cited patent literature

[0000] DE 10 2022 103 077 A1

[0005] WO 2023 / 031890 A1

[0006]

Claims

[1] Device for recording a video sequence in a vehicle, with at least one camera (134) of the vehicle (102) which is designed to capture at least one image (200, 300, 400) with an image of at least one part of the vehicle interior and to generate image data corresponding to the image (200, 300, 400), with a processing unit (136) that is trained to process the image data, wherein the processing unit (136) is further trained to detect at least one vehicle occupant (104, 206, 210) as a subject from the image data, wherein the processing unit (136) is further configured to select an image format, wherein the processing unit (136) is further configured to generate a video sequence with images of the subject from image data of images taken successively with the aid of the camera (134) and to output it in the selected image format. [2] Device according to claim 1, wherein the processing unit is configured to determine, based on the image data, at which positions in the interior of the vehicle (102) vehicle occupants (104, 206, 210) are present. [3] Device according to one of the preceding claims, wherein the processing unit is configured to control the camera (102) and / or the processing of the image data such that the start to generate the video sequence is made by an input from a vehicle occupant (104, 206, 210) via an input unit (112) of the vehicle (102). [4] Device according to claim 3, wherein a selection of a preset stored purpose is made at or before the start of the recording. [5] Device according to claim 4, wherein, starting from the selected preset stored purpose, the image format and / or the resolution of the video sequence is determined. [6] Device according to one of the preceding claims, wherein the video sequence is output in a vertical image format. [7] Device according to one of the preceding claims, wherein the processing unit (136) is configured to produce a section of the image taken by the camera (134) to generate the video sequence in the form of a vertical crop with focus on the driver (104) and / or the passenger (206) and / or a vehicle occupant (210) on the rear seat (204). [8] Device according to one of the preceding claims, wherein the processing unit (138) is configured to perform a computer vision method for person detection and / or motion detection of a vehicle occupant and / or for defining an image section with the subject. [9] Device according to one of the preceding claims, wherein the processing unit (136) is configured to use a computer vision method to perform a dynamic cropping of the images taken by the camera to generate the video sequence with the subject. [10] Device according to one of the preceding claims, wherein the cut is made to fit the driver (104) when the driver (104) is the only occupant of the vehicle. [11] Device according to one of the preceding claims, wherein the processing unit (136) is configured to use a computer vision method to perform a recognition of the vehicle occupants (104, 206, 210). [12] Device according to one of the preceding claims, wherein the processing unit (136) is configured to perform the selection of the motif based on the viewing direction of each of the vehicle occupants (104, 206, 210) and / or based on the input to start the video sequence to be recorded. [13] Method Device for recording a video sequence in a vehicle, where at least one image (200, 300, 400) is captured showing at least part of the vehicle interior and corresponding image data is generated for the image (200, 300, 400), in which the image data is processed and, based on the image data, at least one vehicle occupant (104, 206, 210) is detected and selected as the subject, where an image format is selected, and in which a video sequence recorded by the camera (134) is captured with images of the subject and output in the selected image format.

Citation Information

Patent Citations

  • Method for recording a motor vehicle's driving situation

    DE102015008448B4

  • METHOD FOR RECORDING IN-VEHICLE HYPERLAPSE VIDEO AND SOCIAL NETWORKING CAPABILITY

    DE102016112606A1

  • System for generating a complete media file, logging device, central media storage device, media processing device and motor vehicle

    DE102021125792A1

  • Personalized content creation for autonomous vehicle rides

    US20180259958A1