Systems and methods for vehicle video memos

The in-vehicle computing system addresses the challenge of capturing scenic elements by using AI to automatically generate personalized travel diaries, efficiently storing and stitching relevant feeds, thus improving user experience and memory usage.

WO2025198580A1PCT designated stage Publication Date: 2025-09-25HARMAN INT IND INC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/020434
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-03-18
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Capturing interesting elements of a surrounding environment during a vehicle journey is challenging, particularly for the driver, as existing systems do not efficiently store and stitch relevant video and audio feeds without manual input, leading to inefficient memory usage and user effort in identifying and editing clips.

Method used

An in-vehicle computing system uses AI models to detect features of interest, automatically store and stitch relevant video and audio feeds, reducing storage demands by capturing only when the feature is visible, and allowing users to view or edit the video memo later.

Benefits of technology

This solution enables automatic generation of personalized travel diaries with reduced storage demands, enhancing user experience by capturing key moments without manual intervention and optimizing memory usage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024020434_25092025_PF_FP_ABST
    Figure US2024020434_25092025_PF_FP_ABST
Patent Text Reader

Abstract

Systems and methods for automatically generating video memos in a vehicle are provided herein. In an example, a method includes detecting a key moment trigger in a video feed captured with a camera of the vehicle and / or in an audio feed captured with a microphone of the vehicle. In response to detecting key moment trigger, the method includes adjusting one or more capture parameters of the camera, the microphone, and / or one or more additional cameras and / or microphones; saving a first segment of the video feed and a second segment of the audio feed in long-term memory along with third segments of any additional video feeds and / or audio feeds of the one or more additional cameras and / or microphones; automatically generating the video memo by stitching together aspects of the first segment, the second segment, and / or the third segments; and displaying the video memo on a display device when requested.
Need to check novelty before this filing date? Find Prior Art

Description

SYSTEMS AND METHODS FOR VEHICLE VIDEO MEMOSTECHNICAL FIELD

[0001] Embodiments of the subject matter disclosed herein relate to video memos generated for occupants of a vehicle.BACKGROUND

[0002] Vehicle journeys frequently include navigation through scenic environments. The occupants of the vehicle may see interesting elements of a surrounding environment, but capturing the interesting elements on a personal camera while the vehicle is in motion may be challenging or impossible, particularly for the driver of the vehicle.SUMMARY

[0003] The current disclosure at least partially addresses one or more of the above identified issues with a method for automatically generating a video memo with a computing system of a vehicle. The method may include detecting a key moment trigger in a video feed captured with a camera of the vehicle and / or in an audio feed captured with a microphone of the vehicle, wherein the key moment trigger indicates that the vehicle is within range of a feature of interest, and in response to detecting the key moment trigger: adjusting one or more capture parameters of the camera, the microphone, and / or one or more additional cameras and / or microphones of the vehicle; saving a first segment of the video feed and a second segment of the audio feed in long-term memory of the vehicle along with third segments of any additional video feeds and / or audio feeds of the one or more additional cameras and / or microphones of the vehicle; automatically generating the video memo by stitching together aspects of the first segment, the second segment, and / or the third segments; and displaying the video memo on a display device of the vehicle when requested and / or sending the video memo to an external device.

[0004] The above advantages and other advantages, and features of the present description will be readily apparent from the following Detailed Description when taken alone or in connection with the accompanying drawings. It should be understood that the summary above is provided to introduce in simplified form a selection of concepts that are further described in the detailed description. It is not meant to identify key or essential features of the claimed subject matter, the scope of which is defined uniquely by the claims that follow the detaileddescription. Furthermore, the claimed subject matter is not limited to implementations that solve any disadvantages noted above or in any part of this disclosure.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] Various aspects of this disclosure may be better understood upon reading the following detailed description and upon reference to the drawings in which:

[0006] FIG. 1 is a schematic diagram of a vehicle, according to one or more embodiments of the disclosure;

[0007] FIG. 2 shows a block diagram of an embodiment of a vehicle computing system including an in-vehicle video memo system, according to one or more embodiments of the disclosure;

[0008] FIG. 3 is a flow chart illustrating a method for customizing one or more models of the in-vehicle video memo system for deployment;

[0009] FIG. 4 is a flow chart illustrating a method for deploying the one or more customized models of the in-vehicle video memo system to detect a feature of interest; and

[0010] FIG. 5 is a flow chart illustrating a method for automatically generating a video memo in response to detecting the feature of interest.

[0011] The drawings illustrate specific aspects of the described systems and methods. Together with the following description, the drawings demonstrate and explain the structures, methods, and principles described herein. In the drawings, the size of components may be exaggerated or otherwise modified for clarity. Well-known structures, materials, or operations are not shown or described in detail to avoid obscuring aspects of the described components, systems, and methods.DETAILED DESCRIPTION

[0012] A vehicle may include various systems for monitoring internal and external environments of the vehicle, and communicating information to a driver of the vehicle based on the monitored internal and external environments. As one example. Advanced Driver Assistance Systems (ADAS) are developed to adapt and enhance operation of vehicles to increase quality of driving and the driving experience. An ADAS monitors an inside and outside environment of a vehicle using various types of sensors, such as camera, lidar, radar, and ultrasonic sensors. The sensor data may be processed by an in-vehicle computing system to extract information regarding the vehicle operator / occupant(s), and communicate the information to the operator and / or implement actions based on the information. The actionsmay include alerting the operator via an audio or visual notification, or controlling the throttle, brake, or steering of the vehicle.

[0013] The ADAS may rely on a Driver Monitoring System (DMS) and / or an Occupant Monitoring System (OMS). The DMS includes a set of advanced safety features that applies output from in-cabin sensors (including cameras) to detect and track a physical state (e.g., vehicle operator experiencing drowsiness) or mental state (e.g.. vehicle operator is distracted) of the operator / driver. Further, the DMS may alert the vehicle operator when the vehicle operator is experiencing physical states or mental states that are not suitable for operation of the vehicle. In comparison with the DMS, the OMS may monitor an entirety of an in-cabin environment. The OMS includes features such as seat belt detection, smart airbag deployment, left object warning, and forgotten child warning.

[0014] The ADAS, DMS, and OMS systems are typically oriented towards information of use to the operator (e.g., driver), for example, to increase a safety of the operator and occupants of the vehicle and / or to increase a driving performance of the operator. However, the sensors of the ADAS, DMS, and OMS systems (e.g.. inward and outward facing cameras, microphones) may capture features of interest in the environment surrounding the vehicle during the course of a journey, such as scenic sights (e g., waterfalls, mountains), unique architecture, historical sights, and the like. An operator of the vehicle and / or vehicle occupants may desire to access video clips captured by cameras of the ADAS, DMS, and OMS systems that include features of interest, but doing so may pose challenges. For example, the video feeds from the cameras of ADAS, DMS, and OMS systems may not be stored for any extended period of time; rather, the video may only be stored in a temporary buffer and after processing the video feeds for the tasks described above, the video may be deleted, or the video may be stored in an in-loop fashion (e.g., only the immediately previous five minutes’ worth of video may be stored and then written over with the next five minutes’ worth of video). In some examples, it may be possible for a user to command to the system to save a segment of video from one or more cameras, but doing so may result in the feature of interest being missed by the cameras (e.g., the feature of interest may no longer be visible by the time the user commands the system to start recording).

[0015] As such, to ensure a given feature of interest is captured and available for viewing by a user once the vehicle is stopped and / or has passed by the feature of interest, the video feed from each of the cameras may be stored in long-term memory. Because the system may not know when the vehicle passed near / through the feature of interest, an entirety of each the video feeds may be stored (e.g., from when the vehicle is turned on to when the vehicle is turned off),thereby demanding a relatively large amount of memory. Further, even if the video feeds are stored as clips (e.g., five-minute-long clips) with timestamps, a user may have to view multiple different clips to identify which clips include the feature, fast-forward through segments of clips that do not include the feature, etc. Further still, if the user wishes to document their own reaction to seeing the feature, the user may be forced to edit or stitch together multiple different video clips. Such processes may be inefficient and utilize undue processing and memory resources. Additionally, the positioning of the cameras, frame rate of the cameras, and so forth may be configured for the tasks typically associated with the ADAS, DMS, and OMS systems and thus may not optimally capture the feature of interest or the vehicle occupants.

[0016] Thus, when a user (such as a driver or passenger) drives for a long trip to a beautiful place, it may not be possible to capture all the scenic beauties and wow factors for addressing the user’s personal (emotional or social) interests on the way. The user may not be ready with a camera at the right time or the user may be busy focusing on the road for attentive driving. This is a pain point and at present there is no solution available which solves this problem in an automatic manner without explicit manual input / instructions from the user.

[0017] To address this issue, systems and methods are proposed herein to improve the user experience in a vehicle by providing an intelligent way of combining video feeds from the various available camera systems such as one or more interior cameras (e.g., dashcam, in-cabin camera) and one or more exterior cameras (e.g., 360° camera), to automatically create a personalized user experience with key moments specific to a user, including the user’s facial and / or voice expressions, in form of a travel vlog / digital travel diary (referred to herein as a video memo). With the aid of one or more artificial intelligence (Al) models (e.g., deep learning models such as convolutional neural networks), the in-vehicle computing system may process image data from the one or more cameras to determine if the vehicle is approaching or passing by a feature of interest and automatically store the feeds from each of the cameras, in-cabin microphones, and the like in response to determining that the vehicle is approaching or passing by a feature of interest. The various feeds may then be processed to identify images, video clips, and / or audio clips of interest (e.g., video clips where the feature of interest is visible, images of the user, audio clips of the user saying “wow” or of the music playing in the vehicle while passing by the feature of interest, etc.) that may then be stitched together to form a video memo. The user may view' the video memo on a display device of the vehicle, make edits if desired, and the video memo may be uploaded to the cloud, posted on a social media account of the user. etc.

[0018] In this way, a video memo capturing a feature of interest in the surrounding environment may be generated automatically without explicit user input. The embodiments disclosed herein provide for the generation of automatically personalized diaries during a vehicle journey by choosing, combining, and editing together any available video and audio data. The available data may be acquired from vehicle cameras (in and out facing), from the vehicle microphones, the vehicle's entertainment system, etc. The data may also be acquired from a personal drone which is connected to the vehicle or even from the road infrastructure. Video can also be obtained from another vehicle through vehicle-to-vehicle (V2V) communication, co-passengers camera streams, etc. By triggering the recording / saving of the various video and audio feeds in response to detecting the feature of interest, storage demands may be lowered relative to other solutions by only storing video and / or audio feeds while the vehicle is approaching, passing by, and driving away from the feature of interest (e.g., only while the feature of interest is actually visible to the cameras). Further, by processing the video and audio data to identify the relevant images, video clips, and audio clips and using only the relevant images, video clips, and audio clips, the storage demand of the final video memo may be further reduced.

[0019] Turning now to the figures, FIG. 1 schematically shows an exemplar}’ vehicle 100. The vehicle 100 includes a dashboard 102, a driver seat 104, a first passenger seat 106, a second passenger seat 108, and a third passenger seat 110. In other examples, the vehicle 100 may include more or fewer passenger seats. The driver seat 104 and the first passenger seat 106 are located in a front of the vehicle, proximate to the dashboard 102, and therefore may be referred to as front seats. The second passenger seat 108 and the third passenger seat 110 are located at a rear of the vehicle and may be referred to as back (or rear) seats.

[0020] Vehicle 100 includes a plurality of integrated speakers 114, which may be arranged around a periphery of the vehicle 100. In some embodiments, the integrated speakers 114 are electronically coupled to an electronic control system of the vehicle, such as to a computing system 120, via a wired connection. In other embodiments, the integrated speakers 114 may wirelessly communicate with the computing system 120. As an example, an audio file may be generated by computing system 120 or selected by an occupant of the vehicle 100, and the selected audio file may be played at one or more of the integrated speakers 1 14. In some examples, audio alerts may be generated by the computing system 120 and also may be played at the integrated speakers 114. In some embodiments, audio files, signals, and / or alerts may be selected by or generated for an occupant of vehicle 100, and may be played an integrated speaker 114 associated with and / or proximal to a seat of the occupant. Further, in someembodiments, audio files, signals, and / or alerts selected by or generated for the occupant may be played at an integrated speaker 114 associated with and / or proximal to a seat of one or more other occupants of vehicle 100.

[0021] The vehicle 100 includes a steering wheel 112 and a steering column 122, through which the driver may input steering commands for the vehicle 100. The vehicle 100 further includes one or more cameras 118, which may be directed at respective occupants of the vehicle. In the embodiment shown in FIG. 1, a camera 118 is positioned to the side of the driver seat 104, which may aid in monitoring the driver in profile. However, in other examples, the one or more cameras 118 may be positioned in other locations in the vehicle, such as on the steering column 122, directly in front of the driver seat 104 and / or to the side of or directly in front of passenger seats 106, 108. and 110. The one or more cameras 118 may be arranged such that all of the occupants are captured collectively by the one or more cameras 1 18.

[0022] Additionally, in the embodiment shown in FIG. 1, one or more cameras may be positioned on an exterior of the vehicle 100, which may capture images of a surrounding environment of vehicle 100 and aid in monitoring a position of vehicle 100 in a lane and / or monitor the position of the vehicle 100 relative to other vehicles and / or the surrounding environment. For example, the one or more cameras on the exterior of the vehicle (referred to as exterior cameras) may include camera 119. While camera 119 is depicted in FIG. 1 at a back end 170 of vehicle 100, it should be appreciated that one or more cameras may be positioned at other locations on the exterior of vehicle 100. For example, camera 119 positioned at back end 170 of vehicle 100 may be a first camera of the one or more exterior cameras; a second camera may be positioned at a front end 172 of vehicle 100; a third camera may be positioned at a right side 174 of vehicle 100; and a fourth camera may be positioned at a left side 176 of vehicle 100. In some examples, one or more of the exterior cameras may be located in side mirrors of vehicle 100. In some embodiments, a plurality of cameras may be arranged on each of one or more sides of vehicle 100. Additionally or alternatively, in some embodiments, the one or more cameras may be positioned at various locations inside vehicle 100 and directed outward from vehicle 100 at the surrounding environment, for example, through various windows of the vehicle. The one or more exterior cameras may be arranged such that a 360- degree view of the surrounding environment is captured collectively by the one or more cameras 119.

[0023] The vehicle 100 may further include a driver seat sensor 124 coupled to or within the driver seat 104 and a passenger seat sensor 126 coupled to or within the first passenger seat 106. The back seats may also include seat sensors, such as a passenger seat sensor 128 coupledto the second passenger seat 108 and a passenger seat sensor 130 coupled to the third passenger seat 110. The driver seat sensor 124 and the passenger seat sensor 126 may each include one or a plurality of sensors, such as a weight sensor, a pressure sensor, and one or more seat position sensors that output a measurement signal to the computing system 120. For example, the output of the weight sensor or pressure sensor may be used by the computing system 120 to determine whether or not the respective seat is occupied, and if occupied, a weight of a person occupying the seat. As another example, the output of the one or more seat position sensors may be used by the computing system 120 to determine one or more of an occupant of vehicle 100, a seat height, a longitudinal position with respect to the dashboard 102 and the back seats, and an angle (e.g., tilt) of a seat back of a corresponding seat.

[0024] In some examples, vehicle 100 further includes a driver seat motor 134 coupled to or positioned within the driver seat 104 and a passenger seat motor 138 coupled to or positioned within the first passenger seat 106. Although not shown, in some embodiments, the back seats may also include seat motors. The driver seat motor 134 may be used to adjust the seat position, including the seat height, the longitudinal seat position, and the angle of the seat back of the driver seat 104 and may include an adjustment input 136. For example, the adjustment input 136 may include one or more toggles, buttons, and switches. The passenger seat motor 138 may be used to adjust the seat position, including the seat height, the longitudinal seat position, and the angle of the seat back of the first passenger seat 106 and may include an adjustment input 140. the adjustment input 140 may include one or more toggles, buttons, and switches. Although not shown, in some embodiments, the back seats may be adjustable in a similar manner.

[0025] Computing system 120 may include a user interface (UI) 116. Computing system 120 may receive inputs via UI 116 as well as output information to UI 116. The user interface 116 may be included in a digital cockpit, for example, and may include a display and one or more input devices. The one or more input devices may include one or more touchscreens, knobs, dials, hard buttons, and soft buttons for receiving user input from a vehicle occupant. UI 116 may include a display screen on which information, images, videos, and the like may be displayed. The display screen may be a touchscreen, and an occupant of vehicle 100 may interact with UI 116 via control elements displayed on the display screen. In some embodiments, UI 116 may additionally or alternatively display or project text, images, videos, and the like on a windshield 113 of vehicle 100. Further, in some embodiments, different visual content may be generated and displayed on different display screens of the dashboard, or at different locations of windshield 113, such that the driver may view a first set of visual content,and a front seat passenger of vehicle 100 may view a second, different set of visual content. For example, the first set of visual content may be relevant to operating the vehicle, and the second set of visual content may not be relevant to operating the vehicle.

[0026] Computing system 120 may include one or more additional UIs 117, which may be positioned at locations accessible to back seat passengers of vehicle 100. In various embodiments, the one or more additional UIs 117 may comprise or be comprised by an RSE system. In other words, a first UI 117 may be positioned on a rear side of first passenger seat 106, such that a first rear passenger sitting in second passenger seat 108 may interact with the first UI 117. A second UI 117 may be positioned on a rear side of driver seat 104, such that a second rear passenger sitting in third passenger seat 110 may interact with the second UI 117. In vehicles including additional rear seats, additional UIs 117 may be arranged similarly such that each occupant of vehicle 100 may interact with a UI 117.

[0027] Each additional UI 117 may include the same or similar features included in UI 116, such as physical and / or virtual control elements, one or more display screens, etc. Additionally, each additional UI 117 may be configured to receive the same or different video and / or audio content from computing system 120 and display the content independently of other UI 117s. For example, computing system 120 may generate a first set of visual content on a first display screen associated with first UI 117; a second, different set of visual content on a second display screen associated with second UI 117; a third, different set of visual content on a third display screen associated with third UI 117; and so on. Similarly, a first occupant of vehicle 100 may interact with a first UI 117 to generate a first set of visual content on a first display screen associated with first UI 117; a second occupant of vehicle 100 may interact with a second UI 117 to generate a second, different set of visual content on a second display screen associated with second UI 117; a third occupant of vehicle 100 may interact with a third UI 117 to generate a third, different set of visual content on a third display screen associated with third UI 117; and so on. The sets of visual content may include associated audio content, which may be played at a respective speaker 114.

[0028] Computing system 120 includes a processor 142 configured to execute machine readable instructions stored in a memory 144. The processor 142 may be single core or multicore, and the programs executed by processor 142 may be configured for parallel or distributed processing. In some embodiments, the processor 142 is a microcontroller. The processor 142 may optionally include individual components that are distributed throughout two or more devices, which may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 142 may be virtualized and executed byremotely -accessible networked computing devices configured in a cloud computing configuration. For example, the computing system 120 may be communicatively coupled with a wireless network.

[0029] Computing system 120 may include a DMS 147, which may monitor a driver of the vehicle 100. For example, the camera 118 may be located at the front of the vehicle 100 (e.g., on dashboard 102) or on a side of the vehicle 100 next to the driver seat 104, and may be positioned to view a face of the driver. The DMS 147 may detect facial features of the driver. In some embodiments, the DMS 147 may be used to retrieve a driver profile of the driver based on the facial features, which may be used to customize a position of the driver seat 104, steering wheel 112, and / or other components or software of the vehicle 100, including customizing a video memo based on the driver and predefined preferences of the driver, as explained in more detail below.

[0030] Computing system 120 may include an OMS 148, which may monitor one or more passengers of the vehicle 100. For example, the camera 118 may be located at a side of the vehicle 100 next to one or more of passenger seats 106, 108, and 110, and may be positioned to view a face of a passenger of the vehicle.

[0031] Computing system 120 may include an ADAS 149, which may provide assistance to the driver based at least partially on the DMS 147. For example, ADAS 149 may receive data from DMS 147, and the ADAS 149 may process the data to provide the assistance to the driver. Computing system 120 and ADAS 149 may additionally include one or more artificial intelligence (Al) and / or machine learning (ML) / deep learning (DL) models, which may be used to perform various automated tasks related to image data captured by cameras 118 and 119. In particular, the AI / ML / DL models may be used to detect that the vehicle is within range of (e.g., approaching or passing by) a feature of interest to the driver or passenger of the vehicle and identify relevant images, video clips, and / or audio clips of captured video and audio feeds to include in a video memo, as explained in more detail below.

[0032] Computing system 120 may communicate with networked computing devices via short-range communication protocols, such as Bluetooth®. In some embodiments, the computing system 120 may include other electronic components capable of carrying out processing functions, such as a digital signal processor, a field-programmable gate array (FPGA), or a graphic board. In some embodiments, the processor 142 may include multiple electronic components capable of carrying out processing functions. For example, the processor 142 may include two or more electronic components selected from a plurality of possible electronic components, including a central processor, a digital signal processor, afield-programmable gate array, and a graphics board. In still further embodiments, the processor 142 may be configured as a graphical processing unit (GPU), including parallel computing architecture and parallel processing capabilities.

[0033] Further, the memory 144 may include any non-transitory tangible computer readable medium in which programming instructions are stored. As used herein, the term “tangible computer readable medium” is expressly defined to include any type of computer readable storage. The example methods described herein may be implemented using coded instruction (e.g., computer readable instructions) stored on a non-transitory computer readable medium such as a flash memory', a read-only memory' (ROM), a random-access memory (RAM), a cache, or any other storage media in which information is stored for any duration (e.g. for extended period time periods, permanently, brief instances, for temporarily buffering, and / or for caching of the information).

[0034] Computer memory of computer readable storage mediums as referenced herein may include volatile and non-volatile or removable and non-removable media for a storage of electronically formatted information, such as computer readable program instructions or modules of computer readable program instructions, data, etc. that may be stand-alone or as part of a computing device. Examples of computer memory' may include any other medium which can be used to store the desired electronic format of information and which can be accessed by the processor or processors or at least a portion of a computing device. In various embodiments, the memory 144 may include an SD memory card, an internal and / or external hard disk, USB memory device, or a similar modular memory'.

[0035] Further still, in some examples, the computing system 120 may include a plurality of sub-systems or modules tasks with performing specific functions related to performing image acquisition and analysis. As used herein, the terms “system.” “unit,” or “module” may include a hardware and / or software system that operates to perform one or more functions. For example, a module, unit, or system may include a computer processor, controller, or other logic-based device that performs operations based on instructions stored on a tangible and non- transitory computer readable storage medium, such as a computer memory. Alternatively, a module, unit, or system may include a hard-wired device that performs operations based on hard-wired logic of the device. Various modules or units shown in the attached figures may represent the hardware that operates based on software or hardwired instructions, the software that directs hardware to perform the operations, or a combination thereof.

[0036] Referring now to FIG. 2. a block diagram 200 shows an example of a vehicle computing system 202 of a vehicle, which may be a non-limiting example of computing system120 described above in reference to vehicle 100 of FIG.1. Vehicle computing system 202 may be communicatively coupled to one or more internal cameras 218 of the vehicle (e.g., the one or more cameras 118), and one or more external cameras 219 of the vehicle (e.g., the one or more exterior cameras, such as camera 119). Internal cameras 218 and external cameras 219 may provide image data to vehicle computing system 202.

[0037] Vehicle computing system 202 includes a processor 204 configured to execute machine readable instructions stored in non-transitory memory 206. Processor 204 may be single core or multi-core, and the programs executed thereon may be configured for parallel or distributed processing. In some embodiments, the processor 204 may optionally include individual components that are distributed throughout two or more devices, which may be remotely located and / or configured for coordinated processing. In some embodiments, one or more aspects of the processor 204 may be virtualized and executed by remotely-accessible networked computing devices configured in a cloud computing configuration.

[0038] Non-transitory memory7206 includes an in-vehicle video memo system 208, which may manage generation of visual and audio content to occupants of the vehicle, as described above in reference to FIG. 1. In particular, in-vehicle video memo system 208 may include a video Al engine 210 and an audio Al engine 211, which may each include one or more Al models, and instructions for implementing the one or more one or more Al models, as described in greater detail below. The one or more Al models may be deep learning (DL) neural network models. Video Al engine 210 and audio Al engine 211 may include trained and / or untrained neural networks and may further include various data or metadata pertaining to the one or more neural networks stored therein. Video Al engine 210 and audio Al engine 211 may also include other ty pes of statistical, rules-based, or other models. Video Al engine 210 and audio Al engine 211 may include instructions for implementing one or more gradient descent algorithms, applying one or more loss functions, and / or training routines, for use in adjusting parameters of one or more neural networks of video Al engine 210 and audio Al engine 211. Video Al engine 210 and audio Al engine 211 may include training datasets for the one or more Al models of video Al engine 210 and audio Al engine 211. Video Al engine 210 and audio Al engine 211 may include instructions for deploying one or more trained Al models. Additionally, in-vehicle video memo system 208 (including video Al engine 210 and audio Al engine 211) may include instructions that, when executed by processor 204, cause vehicle computing system 202 to conduct one or more of the steps of methods 300, 400, and 500, described in greater detail below in reference to FIGS. 3-5.

[0039] As explained in more detail below, video Al engine 210 may include a first model (e.g., a first DL model, such as a first convolutional neural network) and a second model (e.g., a second DL model, such as a second convolutional neural network). Audio Al engine 211 may include a third model (e.g., a third DL model, such as a third convolutional neural network). The first model may be trained to identify facial expressions of a user in a video feed, and in particular the first model may be trained to identify facial expressions of the user that indicate a positive reaction or otherwise identify facial expressions that indicate the user may be visualizing a feature of interest. The first model may be pre-trained in a user non-specific manner to identify positive facial expressions for all users / generic users, and then the first model may be trained in a user-specific manner, after installation in the vehicle computing system 202, to identify facial expressions of a specific user. It is to be appreciated that more than one instance of the first model may be stored as part of video Al engine 210, such that an instance of the first model may be generated and stored for each user that has a user profile on the vehicle computing system 202. The second model may be trained to identify features of interest in the environment surrounding the vehicle, such as nature-based features (e.g., mountains, waterfalls, animals), architectural features (e.g.. skyscrapers, certain houses, churches, etc ), and historical features (e.g., monuments, statues). In some examples, the second model may be pre-trained to identify generic features of interest and then may be trained or otherwise updated once installed and deployed in the vehicle computing system 202 to identify user-specific features of interest, such that multiple instances of the second model are stored as part of the video Al engine 210. The third model may be trained to detect voice expressions of a user in an audio feed, and in particular the third model may be trained to identify voice expressions of the user that indicate a positive reaction or otherwise identify voice expressions that indicate the user may be visualizing a feature of interest. The third model may be pretrained in a user non-specific manner to identify positive voice expressions for all users / generic users, and then the third model may be trained in a user-specific manner, after installation in the vehicle computing system 202, to identify voice expressions of a specific user. It is to be appreciated that more than one instance of the third model may be stored as part of audio Al engine 211, such that an instance of the third model may be generated and stored for each user that has a user profile on the vehicle computing system 202.

[0040] Vehicle computing system 202 may be coupled to various UI elements 220, which may include one or more speakers 222, one or more display devices 224, and one or more projectors 226. UI elements 220 may include non-limiting examples of UI 116 and / or UI 117of vehicle 100 of FIG. 1. As such, the one or more speakers 222 may be the same as or similar to speakers 114.

[0041] The one or more display devices 224 and the one or more projectors 226 may be used to display visual content to occupants of the vehicle. The one or more display devices 224 may include one or more display screens utilizing virtually any ty pe of technology7. The one or more projectors 226 may be configured to project visual content on one or more windows of the vehicle. While not depicted in FIG. 2, UI elements 220 may also include one or more user input devices, which may comprise one or more of a touchscreen, a keyboard, a trackpad, a microphone, a motion sensing camera, or other device configured to enable a user to interact with and manipulate data within vehicle computing system 202.

[0042] Non-transitory memory 206 further includes an ADAS module 212 that provides assistance to a driver of the vehicle, and a DMS module 240 and an QMS module 242 that monitor a vehicle operator and occupant(s), respectively. ADAS module 212 and / or DMS module 240 may detect specific states, attributes, and poses of the driver, and may notify the occupant using one or more UI elements 220. As one example, ADAS module 212 may interface with one or more UI elements 220 to output an alert via display devices 224 and / or speakers 222.

[0043] Vehicle computing system 202 may include an onboard navigation system 244. Onboard navigation system 244 may provide route information to the driver, including a current location of the vehicle and a destination of the vehicle. For example, when operating the vehicle, the driver may enter a destination into onboard navigation system 244. Onboard navigation system 244 may indicate one or more routes from the current location to the destination on a map displayed by onboard navigation system 244. In various embodiments, the map may be displayed on a screen of a display of onboard navigation system 244, such as a dashboard display. The driver may select a route of the one or more routes, and onboard navigation system 244 may provide instructions to the driver and / or indicate a progress of the vehicle towards the destination on the map. Onboard navigation system 244 may additionally calculate information such as an estimated distance to the destination, an estimated time of arrival at the destination based on a speed of the vehicle, indications of traffic on the route, and / or other information of use to the driver.

[0044] Onboard navigation system 244 may rely on a global positioning system (GPS) to determine the current location of the vehicle. In various embodiments, the current location, route, and / or other GPS-based information may be used by one or more models of video Al engine 210 and / or audio Al engine 21 1 to assist in identifying a feature of interest or assist inidentifying that a user may be reacting to viewing a feature of interest, or to location-gate generation of video memos. For example, certain locations may forbid video recording (e.g., military bases, archeological sites) and the GPS-based information may be used to disable video memo generation when the vehicle is at a restricted location. In still further examples, the GPS-based information may be used to disable video memo generation (and in particular disable feature of interest detection) when the vehicle is within range of a home and / or work location, to avoid unnecessary processing of video and audio feeds when the likelihood of a feature of interest being nearby is low.

[0045] Vehicle computing system 202 may include a communication module 250, which may manage wireless communication between vehicle computing system 202 and one or more wireless networks, such as wireless network 260. In various embodiments, wireless network 260 may include the Internet. Additionally, wireless network 260 may include one or more private networks. For example, communication module 250 may be used to send final video memos to a cloud-based database 262 for long-term storage.

[0046] One or more microphones 216 may be communicatively coupled to vehicle computing system 202 to receive voice commands from a user, to measure ambient noise in the vehicle, to determine whether audio from speakers of the vehicle is tuned in accordance with an acoustic environment of the vehicle, and so on. Vehicle computing system 202 may include a speech processing unit 214 configured to process voice commands, such as the voice commands received from the one or more microphones 216. It is to be appreciated that in some examples, the third model of audio Al engine 211 may be an instance of the speech processing unit 214, customized to detect voice expressions of the user indicative of view ing a feature of interest.

[0047] Vehicle computing system 202 may be configured to communicate with one or more external devices 270 located external to the vehicle via communication module 250 directly and / or via wireless netw ork 260. The external devices 270 may be located external to the vehicle and / or the external devices 270 may be temporarily housed in the vehicle, such as when the user is operating the external devices while operating the vehicle. In other words, the external devices 270 are not integral to the vehicle. External devices 270 may include a mobile device 272 (e g., connected via a Bluetooth®, NFC, WI-FI Direct®, or other wireless connection) or one or more alternate external devices 274, such as a drone, other vehicles, road infrastructure, and the like. When the one or more alternate external devices 274 includes a drone, the drone may be configured with one or more cameras to capture images / video of the vehicle and / or environment surrounding the vehicle and send the images / video to the vehiclecomputing system 202. (Wi-Fi Direct® is a registered trademark of Wi-Fi Alliance, Austin, Texas.)

[0048] Mobile device 272 may be a mobile phone, smart phone, wearable devices / sensors that may communicate with the vehicle computing system 202 via wired and / or wireless communication, or other portable electronic device(s). Other external devices include one or more external services 276, such as social media platforms. Still other external devices include one or more external storage devices 278, such as solid-state drives, pen drives, Universal Serial Bus (USB) drives, and so on. External devices 270 may communicate with vehicle computing system 202 either wirelessly or via connectors without departing from the scope of this disclosure. For example, external devices 270 may communicate with vehicle computing system 202 through communication module 250 over wireless network 260. a USB connection, a direct wired connection, a direct wireless connection, and / or other communication link.

[0049] It should be understood that vehicle computing system 202 shown in FIG. 2 is for illustration, not for limitation. Another appropriate vehicle computing system may include more, fewer, or different components.

[0050] Referring now to FIG. 3. a method 300 is shown for customizing models of a video Al engine (such as video Al engine 210) and audio Al engine (such as audio Al engine 211) of a vehicle for deployment in generating video memos. Method 300 may be performed by a processor of a vehicle computing system, such as processor 204 of FIG. 2, based on instructions stored in non-transitory memory of the vehicle computing system, such as memory 206.

[0051] At 302, method 300 includes establishing a user profile for a user. The user may be a driver of the vehicle or a passenger of the vehicle (or both depending on the user, e.g., sometimes the driver and sometimes a passenger). Establishing the user profile may include presenting a list of questions to the user and receiving answers to the questions from the user. The answers may be used to set video memo preferences for the user that may be saved as part of the user profile, such as what types of features are of interest to the user (e.g., nature features, architectural features, etc.), whether the user prefers video memos only be made at certain locations or on certain types of journeys (e.g., the user may indicate that video memos are to be made only when the vehicle is a certain distance from the user’s home or workplace), whether the user wants only images / video clips of the external environment included in the video memos (or whether user also wants images / video clips of the user and / or any other occupants of the vehicle included in the video memos), and so forth.

[0052] At 304. facial expressions of the user are captured with one or more intemal / interior cameras, such as internal cameras 218. During vehicle operation, the internalcameras may collect image data (e.g., internal video feeds) that may be used by the DMS and / or OMS to detect driver and occupant states. The image data / video feeds may include the user, particularly when the user is driving. The captured facial expressions in the image data / intemal video feeds may be used to update a first model of the video Al engine, as indicated at 306. As explained above, the first model may be pre-trained to identify facial expressions of generic users in image frames of a video feed. For example, the first model may be pre-trained to identify that a user is smiling, frowning, talking, etc. However, when the first model is pretrained, image frames and / or video clips from a plurality of different users may be used as input during the pre-training procedure. Because different users may have different versions of various facial expressions, the first model may not identify the particular facial expressions of the user as accurately as desired. Thus, the facial expressions of the user, captured with the one or more internal cameras, may be used to refine or fine-tune the first model so that a userspecific instance of the first model is generated.

[0053] At 308, method 300 includes updating a second model of the video Al engine based on the user profile. As explained above, the user profile may include user preferences of the types of features the user is interested in. The second model may be pre-trained to identify various features in image data captured from one or more external cameras, e.g., features in the environment surrounding the vehicle. The second model may thus be updated based on the user profile / user preferences to narrow the features that the second model identifies as features of interest and thus create a user-specific instance of the second model. However, in some examples, the second model may not be updated based on the user profile, and instead the user profile / user preferences may be used to determine whether a detected feature in the surrounding environment is of interest and should trigger creation of a video memo.

[0054] At 310, voice expressions of the user are captured with one or more microphones, such as microphones 216. During vehicle operation, the microphone(s) may collect audio data (e.g., audio feeds) that may be used by the DMS and / or OMS to detect driver and occupant states and / or used as user input (e.g., processed to detect user voice commands). The audio data / audio feeds may include speech (e.g., voice expressions) from the user. The captured voice expressions in the audio data / audio feeds may be used to update a third model of the audio Al engine, as indicated at 312. As explained above, the third model may be pre-trained to identify voice expressions of generic users in an audio feed. For example, the third model may be pretrained to identify' that a user is excited, in awe, disappointed, etc., and / or the third model may be pre-trained to identify specific phrases such as "‘look at that” or “wow.” However, when the third model is pre-trained, audio data from a plurality of different users may be used as inputduring the pre-training procedure. Because different users may have different versions of various voice expressions, the third model may not identify the particular voice expressions of the user as accurately as desired. Thus, the voice expressions of the user, captured with the one or more microphones, may be used to refine or fine-tune the third model so that a user-specific instance of the third model is generated.

[0055] At 314, the models of the video Al engine and the audio Al engine (e.g., the first model, the second model, and the third model) are deployed to generate a video memo when indicated. Additional details about generating a video memo are provided below. Briefly, the models of the video Al engine and the audio Al engine are deployed in order to detect if the vehicle is within range of a feature of interest, as explained below with respect to FIG. 4. If a feature of interest is detected, a video memo may be automatically generated, which may include identification of relevant images, video clips, and / or audio clips by the models of the video Al engine and the audio Al engine, as explained in more detail below with respect to FIG. 5.

[0056] At 316, the models of the video Al engine and / or the audio Al engine may be updated based on user commands and / or user modifications to generate video memos. For example, if the user knows that the vehicle is about to become in range of a particular feature of interest, the user may enter a command for the vehicle computing system to generate a video memo. In some examples, when the vehicle computing system initiates the process of generating a video memo, a notification may be output on a display device of the vehicle or via a speaker of the vehicle to alert the user that a video memo is going to be created. If the user does not wish to have a video memo created, the user may enter a command for the vehicle computing system to stop generating the video memo. Each of these commands may be used by the models to further refine what features, locations, and the like the user prefers for the creation of video memos. Additionally, once a video memo has been generated, the user may view the video memo on a display device of the vehicle and accept / store the video memo or reject / delete the video memo, or the user may make modifications to the video memo (e.g., update the music included in the video memo, remove images / video clips of the user from the video memo. etc.). Whether a user chooses to keep or delete a video memo or make changes to the video memo may also be used by the models to further refine what features, locations, and the like the user prefers for the creation of video memos as well what content the user prefers to be included in the video memos. Method 300 then ends.

[0057] Thus, method 300 provides for personalizing video memo creation for a user. The models disclosed herein may learn about the user’s choice of key scenes in nature and keyevents of interest via a set of questions, observing the user’s facial expressions during specific moments over time (a continuous learning process), and by listening to user’s conversation, and speech detection method, to understand the user’s voice expressions. The intonations used during conversation may assist the third model / audio Al engine to understand more about the user’s feelings at specific moments. The continuous learning process also identifies key speech indications for specific users. Further, the models may continue to learn by the user's manual edits to a video memo. For example, at the end of the trip, the vehicle computing system provides an option for the user to preview the created video memo. The user may manually edit and make the final video memo. This acts as a continuous learning process for the vehicle computing system to understand the user’s choice. As a result, the vehicle computing system / in-vehicle video memo system can identify key scenes and the cameras may be triggered automatically to capture video feeds. Users need not provide voice commands or press a capture button for the camera systems to capture such moments. On the other hand, the vehicle computing system may allow user to provide commands for capturing specific moments. This acts as a learning input for the models and helps to perform better in future. Further, different members of a family / group can have separate profiles stored in memory of the vehicle computing system and their preferences will be stored accordingly. Based on a priority7of the user, the vehicle computing system may follow the preference factors and capture moments as per the priority of the user. For example, the driver may be given highest priority but the vehicle computing system may automatically create video memos based on passenger preferences as well.

[0058] FIG. 4 illustrates a method 400 for automatically triggering generation of a video memo using models of a video Al engine (such as video Al engine 210) and audio Al engine (such as audio Al engine 211) of a vehicle. Method 400 may be performed by a processor of a vehicle computing system, such as processor 204 of FIG. 2, based on instructions stored in non-transitory memory' of the vehicle computing system, such as memory 206. Method 400 may be performed along with, or as part of, method 300. For example, method 400 may be executed at 314 of method 300 in order to deploy the first, second, and third models updated according to method 300 to generate a video memo.

[0059] At 402, method 400 includes capturing user facial expressions, user voice expressions, and an external video feed with an internal camera (e.g., internal cameras 218), a microphone (e.g., microphones 216), and an external camera (e.g., external cameras 219) of the vehicle, respectively. The user facial expressions may be captured in an internal video feed by the internal camera. The user voice expressions may be captured in an audio feed by themicrophone. The user may be the driver of the vehicle, or a passenger of the vehicle. For example, when the vehicle is turned on, internal video captured by the internal camera may be used to identify that the user is in the driver seat and retrieve the user preferences stored in the user profde for the user. Alternatively, when the vehicle is turned on, the user may select their user profde from a menu displayed on a display device of the vehicle and the user preferences stored in the user profde may be retrieved. Based on the user preferences, the vehicle computing system (and specifically the in-vehicle video memo system 208) may determine that the user has authorized video memo creation and may identify which instances of the first model, the second model, and / or the third model to use for key moment trigger detection based on the user and / or user preferences.

[0060] At 404, method 400 includes entering the internal video feed as input to the first model of the video Al engine (or the selected instance of the first model, wherein the selected instance of the first model is specific to the user). The internal video feed may include a plurality7of image frames, and each image frame may be entered as input individually to the first model. In other examples, only a selected set of image frames from the internal video feed may be entered as input, such as each fifth image frame or each tenth image frame of the internal video feed, which may reduce the processing demand of the first model. In still further examples, multiple successive image frames (e.g., five or ten image frames) may be concatenated and entered as input collectively to the first model. At 406, output from the first model is received that indicates whether a key moment trigger is detected in the internal video feed. The first model may generate the output for each image frame or group of image frames that is entered as input. The key moment trigger may include a user facial expression that indicates the user may be viewing, or about to view, a feature of interest. Thus, the key moment trigger in the internal video feed may include a facial expression that indicates a feature of interest is visible (and hence within range of the vehicle), such as facial expressions that indicate excitement (e.g., smiling, raised eyebrows). In some examples, the first model may also be trained to detect user poses that indicate a feature of interest is within range (e.g., pointing).

[0061] At 408, method 400 includes entering the external video feed as input to the second model of the video Al engine (which may be a user-specific instance of the second model, at least in some examples). Similar to the internal video feed, the external video feed may include a plurality7of image frames, and each image frame may be entered as input individually to the second model. In other examples, only a selected set of image frames from the external video feed may be entered as input, such as each fifth image frame or each tenth image frame of theexternal video feed, which may reduce the processing demand of the second model. In still further examples, multiple successive image frames (e.g., five or ten image frames) may be concatenated and entered as input collectively to the second model. At 410, output from the second model is received that indicates whether a key moment trigger is detected in the external video feed. The second model may generate the output for each image frame or group of image frames that is entered as input. The key moment trigger may include a feature of interest captured in the external video feed (and hence indicating that the vehicle is within range of a feature of interest). In some examples, the second model may be trained to output an indication of whether or not a feature of interest to the user is visible in the external video feed, and the indication that the feature of interest is visible may be the indication that a key moment trigger is detected. In other examples, the second model may be trained to output an indication of whether or not a feature of interest is visible in the external video feed and if so, an identification of the type of feature of interest (e.g., waterfall, skyscraper). In such examples, the video Al engine or in-vehicle video memo system may compare the type of the feature of interest to the user's preferences stored in the user profile to determine if the feature of interest is actually of interest to the user. If so, the detection of the feature of interest is the key moment trigger.

[0062] At 412, method 400 includes entering the audio feed as input to the third model of the audio Al engine (which may be a user-specific instance of the third model). The audio feed may be segmented into clips of a specified duration (e.g., 10-1000 ms) and each clip may be entered as input to the third model individually or as groups of clips. In other examples, every other clip, every fifth clip, or the like may be entered as input to the third model. At 414, output from the third model is received indicating if a key moment trigger is detected in the audio feed. The third model may generate the output for each audio clip or group of audio clips that is entered as input. The key moment trigger may include a user voice expression that indicates the user may be viewing, or about to view, a feature of interest. Thus, the key moment trigger in the audio feed may include a voice expression that indicates a feature of interest is visible to the user (and hence that the vehicle is within range of a feature of interest), such as voice expressions that indicate excitement and / or that the user is viewing the feature of interest (e.g.. specific intonations, specific words).

[0063] At 416, method 400 determines if a key moment trigger is detected in any one of the internal video feed, the external video feed, and the audio feed. For example, method 400 may determine if the user facial expressions captured in the internal video feed indicate the user is viewing a feature of interest, based on the output of the first model. Method 400 mayalso determine if the environment surrounding the vehicle, as captured in the external video feed, includes a feature of interest, based on the output of the second model. Further, method 400 may determine if the user voice expressions captured in the audio feed indicate the user is viewing a feature of interest, based on the output of the third model.

[0064] If a key moment trigger is detected based on the output of the first model, the output of the second model, and / or the output of the third model, method 400 proceeds to 418 to automatically generate a video memo, which is explained in more detail with respect to FIG. 5. If a key moment trigger is not detected based on the output of the first model, the output of the second model, and / or the output of the third model, method 400 proceeds to 420, which includes not storing the internal and external video feeds or audio feed, and not triggering any subsequent recording of any additional cameras or microphones. Rather, once it is determined that a key moment trigger is not detected over a duration of time, the internal and external video feeds and the audio feed over that duration of time may be deleted. For example, the internal and external video feeds and the audio feed may be stored temporarily in a buffer of the vehicle computing system. Once it is determined that a key moment trigger is not detected in any of the internal video feed, the external video feed, and the audio feed, the internal video feed, the external video feed, and the audio feed may be deleted from the buffer and not saved in any long-term storage. It is to be appreciated that the internal video feed, the external video feed, and the audio feed may not be deleted instantaneously but may be deleted once new internal video, external video, and audio is captured and stored in the buffer. Any cameras or microphones not currently recording are maintained in the current state.

[0065] At 422, method 400 determines if the vehicle has been turned off. If not, method 400 loops back to 402 to continue to capture the internal video feed, the external video feed, and the audio feed, enter the feeds as input to the respective models, and determine if a key moment trigger is detected in any of the feeds. If the vehicle is turned off, method 400 ends. It is to be appreciated that if a video memo is generated, the video memo is stored in long-term storage (whether as part of the vehicle computing system or on an external device / system, such as the cloud) until the user chooses to delete the video memo. While method 400 is described above as including the monitoring of one internal video feed, one external video feed, and one audio feed to detect a key moment trigger, in some examples, more than one external video feed, more than one internal video feed, and / or more than one audio feed may be monitored (e.g., entered as input to the second model, the first model, and the third model, respectively) to detect the key moment trigger. However, by only monitoring one external video feed, one internal video feed, and one audio feed, processing resources may be reduced.

[0066] It is to be appreciated that in some examples, the internal video feed may also be used by the DMS (e.g., DMS module 240) to determine a current state of the driver (also referred to as a driver state). For example, the internal video feed may be used by the DMS to determine if the driver is attentive (e.g., eyes on the road, alert), drowsy, or inattentive (e.g., looking at a phone). If the DMS determines the driver state meets a conditional state, such as if the driver state is drowsy or inattentive, the key moment trigger detection described above may be disabled until the driver state returns the attentive state, for example. In such examples, when the driver state meets the conditional state, the internal video feed is not entered as input to the first model, which may focus processing resources on monitoring the driver and outputting notifications to alert the driver of the conditional state. In this way, in some examples, the internal video feed may only be entered as input to the first model in response to an indication that the driver state does not meet the conditional state. Further, as will be explained in more detail below, in some examples capture parameters of the internal camera may be adjusted in response to detecting the key moment trigger, so that further image data captured with the internal camera is better suited to inclusion in the video memo, such as frame rate, focus, etc., and these adjustments to the capture parameters may affect driver state monitoring done by the DMS. To avoid disruption to the DMS while the driver is drowsy or otherwise inattentive, the video memo may not be generated when the driver state meets the conditional state.

[0067] Similarly, the external video feed may be used by the ADAS to monitor for a conditional vehicle state, such as road obstacles (e.g., pedestrians, other vehicles, animals), lane departure, or other vehicle states in order to alert the driver of the vehicle state. Thus, if the ADAS detects a road obstacle within a certain distance of the vehicle, detects lane departure, or other conditional vehicle states, the key moment trigger detection described above may be disabled until the road obstacle is off the road or the vehicle passes by the road obstacle, for example. In such examples, when a conditional vehicle state is detected (e.g., a road obstacle is detected), the external video feed is not entered as input to the second model, which may focus processing resources on monitoring the road and outputting notifications to alert the driver of the vehicle state. In this way. in some examples, the external video feed may only be entered as input to the second model in response to an indication that the vehicle state does not meet a conditional vehicle state (e.g., no road obstacles, no lane departures, etc.). Further, as will be explained in more detail below, in some examples capture parameters of the external camera may be adjusted in response to detecting the key moment trigger, so that further image data captured with the external camera is better suited to inclusion in the video memo, such asframe rate, focus, etc., and these adjustments to the capture parameters may affect vehicle state monitoring done by the ADAS. To avoid disruption to the ADAS while the vehicle is in the conditional state (e.g., nearing a road obstacle), the video memo may not be generated when the vehicle state meets the conditional vehicle state.

[0068] Turning now to FIG. 5, a method 500 for generating a video memo is presented. Method 500 may be performed by a processor of a vehicle computing system, such as processor 204 of FIG. 2, based on instructions stored in non-transitory memory of the vehicle computing system, such as memory 206. Method 500 may be performed along with, or as part of, method 400. For example, method 500 may be executed at 418 of method 400 in order to generate a video memo in response to detection of a key moment trigger via output from the first model, the second model, and / or the third model.

[0069] At 502, method 500 optionally includes adjusting one or more capture parameters of one or more cameras and / or microphones of the vehicle, in response to detecting the key moment trigger. For example, the key moment trigger detection described above with respect to FIG. 4 may rely on only a subset of available cameras and / or microphones in the vehicle. For example, only the feed from a front-facing external camera may be recorded and entered as input to the second model; the feed from a rear-facing external camera may not be stored or analyzed to detect a key moment trigger. However, once a key moment trigger is detected, the rear-facing external camera may be activated (if currently deactivated) and the feed from the rear-facing camera may be stored, as explained below at 504. Likewise, any other cameras or microphones of the vehicle not currently recording may be activated and their feeds stored. In some examples, the vehicle may be in communication with a drone including one or more cameras, and the drone may be commanded to activate its camera(s), fly above the vehicle, and send the video feed(s) from the drone camera(s) to the vehicle computing system in response to detecting the key moment trigger. In still further examples, the capture parameters of the external camera and / or the internal camera may be adjusted in response to detecting the key moment trigger. As explained above, the internal camera may be primarily used to detect driver states via the DMS, and thus may have capture parameters suited for detecting driver states, such as a particular frame rate and focus area (e.g., on the eyes of the driver). However, it may be desired to capture a w ider field of view of the driver and / or at a higher resolution to include in the video memo, and thus the frame rate and / or focus of the internal camera may be adjusted in response to detecting the key moment trigger.

[0070] At 504, the triggered camera and audio feeds (e.g., from the cameras and microphones of the vehicle, and any cameras on the drone in communication with the vehicle)are collected for a duration and stored in memory / long-term storage (e.g., longer term storage than the buffer explained above). In some examples, the camera and audio feeds may be collected for a predetermined amount of time following detection of the key moment trigger, such as five minutes or another suitable amount of time. The predetermined amount time may balance the duration for ensuring the feature of interest is sufficiently captured for inclusion in the video memo with avoiding undue memory storage demands for recording longer video / audio feeds. In other examples, the duration that the camera and audio feeds are collected may depend on the feature of interest, the distance from the feature of interest to the vehicle, and the visibility of the feature of interest. For example, the external video feeds may be monitored, via the second model, to determine if the feature of interest is visible in the feeds. Once the feature of interest is no longer visible in any of the external video feeds, and has not been visible for a certain amount of time (e.g., five seconds, ten seconds), the vehicle computing system may determine that the feature of interest is no longer visible and stop collecting the video and audio feeds. To further reduce memory demands, the external video feeds may be monitored via the second model to determine which external camera(s) is capturing the feature of interest and any external cameras that are not capturing the feature of interest may be deactivated or their feeds may not be stored in long-term storage. For example, if a first side-facing camera is currently capturing the feature of interest, a second side-facing camera on an opposite side of the vehicle may be deactivated or its feed not stored in long-term storage, unless the orientation of the vehicle changes.

[0071] Further, in examples where the key moment trigger is detected in the internal video feed or the audio feed but not in the external video feed (e.g., the video memo is generated only based on user expressions), the initial process of generating the video memo (e.g., adjusting the capture parameters and collecting the camera and audio feeds as explained above) may be performed, but the external video feed(s) may be evaluated, via the first model, to confirm if a feature of interest is indeed being captured. If a feature of interest is not detected in any of the external video feeds after a set amount of time following the key moment trigger being detected (e.g., 10-30 seconds), a notification may be output asking the user to confirm if the user would like to continue generating the video memo. If the user confirms that the video memo should be generated, the method may proceed as described below to complete the video memo. However, if the user indicates that a video memo should not be generated, the method may end and the video and audio feeds collected up to that point may be deleted. In doing so, erroneous generation of video memos based solely on user expressions may be avoided, as userexpressions indicative of a feature of interest being viewed may be occasionally confused with certain conversations occurring in the vehicle.

[0072] It is to be appreciated that at least some portion of the external video feed, the internal video feed, and the audio feed captured prior to the key moment trigger being detected may be stored in the buffer, and the external video feed, the internal video feed, and the audio feed stored in the buffer may be transferred to the long-term storage as well and included as part of the feeds used to generate the video memo. In doing so. any images, video clips, or audio clips captured as the feature of interest was first coming into view may be saved for inclusion in the video memo, which may be particularly beneficial for features of interest that are only briefly visible. Accordingly, once the key moment trigger is detected, the vehicle computing system may continue to obtain captured video and audio with the internal camera, the external camera, and the microphone to form an extended internal video feed (or a first segment of the internal video feed), an extended external video feed (or a second segment of the external video feed), and an extended audio feed (or a third segment of the audio feed), and the extended internal video feed, the extended external video feed, and the extended audio feed may be stored in memory, along with additional segments of video feed(s) and audio feed(s) collected after the key moment trigger detection.

[0073] At 506, the stored audio feed(s) are filtered to remove extraneous sounds, such as ambient noise, location service instructions (e.g., directions output by the onboard navigation system), etc., from the captured audio data. The filtering may be achieved by using active noise cancellation and audio source separation methods, which filters out unwanted noise such as users’ casual conversation, any alerts in the vehicle, navigation announcements, and other noise inside or outside the vehicle.

[0074] At 508, one or more relevant images, video clips, and / or audio clips are identified based on the output from the audio Al engine and the video Al engine (e.g., based on the output from the first model, the second model, and / or the third model). As the various video and audio feeds are collected, the video and audio feeds may be entered as input to the first model, the second model, and / or the third model. Based on the output from the models, images, video clips, and / or audio clips may be determined to be relevant for inclusion in the video memo, as those images, video clips, and / or audio clips may include the feature of interest, the user reacting to the feature of interest, other vehicle occupants reacting to the feature of interest, and so forth. For example, each of the feeds from the external cameras may be entered as input to the second model to determine which images (or series of images) in each feed includes the feature of interest. Each image or series of images that includes the feature of interest may betagged or otherwise identified as including the feature of interest (and therefore identified as relevant). The feed(s) from the internal camera(s) may be entered as input to the first model to determine which images (or series of images) in each feed includes the user (or another vehicle occupant) reacting to the feature of interest or otherw ise includes the user (or another vehicle occupant) displaying a positive reaction. Each image or series of images that includes the user or other vehicle occupant with a positive reaction may be tagged or otherwise identified as being relevant. A similar process may be applied for the (filtered or unfiltered) audio feed(s) from the microphone(s) in order to identify, via the third model, audio clips that include the user and / or another vehicle occupant vocally reacting to seeing the feature of interest. However, in some examples, the entirety of the filtered audio may be deemed relevant, as the user may wish to include only the music playing in the vehicle at the time the feature of interest was viewed in the video memo. In some examples, if the output from the first model indicates that insufficient images or video clips of the user were collected, the vehicle computing system may output a notification to the user to pose for a photograph or short video (e.g.. a selfie) and the internal camera may be controlled to capture the photograph or short video for inclusion in the video memo.

[0075] At 510, the relevant images, video clips, and / or audio clips are stitched together to create an initial video memo. In some examples, only the relevant images, video clips, and / or audio clips may be included in the initial video memo. Further, in some examples, any images, video clips, and / or audio clips that are not deemed relevant may be deleted or otherwise removed from the long-term storage. By identifying the relevant images, video clips, and / or audio clips and deleting the remaining (e.g., non-relevant) images, video clips, and / or audio clips, the memory footprint of the video memo may be further reduced. The relevant images, video clips, and / or audio clips may be stitched together in a suitable fashion, e.g.. such that the video memo includes all of the relevant images, video clips, and / or audio clips or such that the video memo includes only selected relevant images, video clips, and / or audio clips. The relevant images, video clips, and / or audio clips selected for inclusion in the video memo may be presented in temporal order, to emphasize the feature of interest over the user or vice versa, to create a desired artistic vision, etc.

[0076] At 512, the initial video memo may be processed to remove selected segments of the initial video memo that include other vehicles, people outside the vehicle, or other external features that the user may have indicated should be excluded from the video memo, to create a processed video memo. It is to be appreciated that the processing to remove the selectedsegments may be performed before the selected relevant images, video clips, and / or audio clips are stitched together to form the video memo, in some examples.

[0077] At 514, the processed video memo may be displayed on a display device of the vehicle (e.g., display device 224) when requested. For example, when the vehicle is stopped and placed into park, for example, a notification may be output on the display device that the video memo is available for viewing. The user may then request to view the video memo and the video memo may be displayed on the display device. In some examples, the user may request updates to the video memo, such as to delete segments of the video memo, reorder segments of the video memo, change the music played during the video memo, etc. Thus, at 516, method 500 may include updating the video memo based on user input (e.g., to a touch screen of the display device) to form a final video memo. At 518. the final video memo is sent to an external device(s), such as a cloud-based system, the user’s mobile device, a social media platform, etc., depending the preferences set by the user and stored in the user profile. Method 500 then ends.

[0078] Thus, during vehicle operation, the vehicle computing system, utilizing an in- vehicle video memo system that includes a video Al engine and an audio Al engine, may be configured to detect key moments of nature and key happenings around the vehicle as per predicted / estimated user’s preference. During those detected key moments, the system records videos from all available cameras outside facing, the driver facing camera and if available also drone camera that may provide a view of the vehicle and its occupants from various angles, for use to create an overall video memo with a more personalized feel. During those detected key moments, the system records audio from the mics and / or acquires audio from the recorded videos. For example, the videos may include the music that is played in the vehicle during those moments. The audio Al engine detects the audio of interest, and filters out the users talking, navigation announcements, or any other noise inside or outside the vehicle, so that the video used in the video memo may be noise free. The video Al engine identifies videos of interest outside and inside the car for embedding in the overall personalized video memo. For example, passengers’ expressions and gestures, incidents and special events, important and interesting characters, etc., may be included. The system "stitches” together the key moments videos and videos from the previous steps. It may also add filters, text, relevant audio clips, and / or background music to create a final video memo for a particular place. The video memo may be saved to the cloud, shared in social media, and to geographical map systems.

[0079] As such, unlike the existing solutions that record video throughout the journey inloop fashion, the camera and audio feed recording are triggered based on key scenes / events.As a result, video length is much smaller than usual camera systems and includes only key moments (photos and short videos) during the journey. These video streams are merged to create the final memo that provides an extended surround experience which is personalized with user’s expressions. At the end of the journey, the system provides an option for user to preview the video memo and the user can manually add / delete specific video clips or pictures. This is an optional step and the user can ignore go ahead with the system-created video memo. The final video memo gets saved in cloud for future use. Also, based on user’s prior consent, the video memo can be auto shared in social media, with friends / family, or to geographical maps as experiences in the places (pictures / videos). The information on social media accounts and contact of friend / family may be obtained from the user's profile.

[0080] During such video capture, surrounding vehicles, people may also be captured which is not desired. Video post processing before storing in cloud may filter out unwanted objects. Also, some areas may be prohibited from capturing images / videos by normal users (e.g., military area, archaeological site. etc.). The vehicle computer system gets such information based on location of those areas in the navigation system and will not trigger any recording in those areas.

[0081] The technical effect of generating a video memo during a journey in a vehicle in response to detection of a key moment trigger in any one of an external video feed, an internal video feed, and an audio feed is that the video memo may be created with video and / or audio stored / recorded in response to the key moment trigger and without having to record video and / or audio over the entire course of the journey, which may save memory. Another technical effect is that ADAS, DMS, and / or OMS sensors such as cameras and microphones may be leveraged to create the video memos, thereby minimizing extra equipment needed to create the video memos.

[0082] The disclosure also provides support for a method for automatically generating a video memo with a computing system of a vehicle, comprising: detecting a key moment trigger in a video feed captured with a camera of the vehicle and / or in an audio feed captured with a microphone of the vehicle, wherein the key moment trigger indicates that the vehicle is within range of a feature of interest, in response to detecting the key moment trigger: adjusting one or more capture parameters of the camera, the microphone, and / or one or more additional cameras and / or microphones of the vehicle, saving a first segment of the video feed and a second segment of the audio feed in long-term memory of the vehicle along with third segments of any additional video feeds and / or audio feeds of the one or more additional cameras and / or microphones of the vehicle, automatically generating the video memo by stitching togetheraspects of the first segment, the second segment, and / or the third segments, and displaying the video memo on a display device of the vehicle when requested and / or sending the video memo to an external device. In a first example of the method, the camera is a first camera of the vehicle and wherein adjusting one or more capture parameters comprises instructing a second camera of the vehicle to begin capturing video to generate a first additional video feed. In a second example of the method, optionally including the first example, detecting the key moment trigger comprises entering the video feed as input to a first model trained to detect facial expressions of an occupant of the vehicle and determine that the vehicle is within range of the feature of interest based on the detected facial expressions. In a third example of the method, optionally including one or both of the first and second examples, detecting the key moment trigger comprises entering the video feed as input to a second model trained to detect the feature of interest in the video feed. In a fourth example of the method, optionally including one or more or each of the first through third examples, detecting the key moment trigger comprises entering the audio feed as input to a third model trained to detect voice expressions of an occupant of the vehicle and determine that the vehicle is within range of the feature of interest based on the detected voice expressions. In a fifth example of the method, optionally including one or more or each of the first through fourth examples, automatically generating the video memo by stitching together aspects of the first segment, the second segment, and / or the third segments comprises identifying one or more relevant images, video clips, and / or audio clips in the first segment, the second segment, and / or the third segments based on output from an audio engine and a video engine of the computing system of the vehicle, and stitching together the identified one or more relevant images, video clips, and / or audio clips to form the video memo.

[0083] The disclosure also provides support for a system for a vehicle, comprising: a processor, and a non-transitory memory storing a first model, a second model, and a third model and instructions that when executed, cause the processor to: obtain a first video feed from a first, interior camera of the vehicle, a second video feed from a second, external camera of the vehicle, and an audio feed from a microphone of the vehicle, enter the first video feed as input to the first model, the second video feed as input to the second model, and the audio feed as input to the third model, in response to receiving an indication from one or more of the first model, the second model, and the third model that a key moment trigger has been detected indicative of the vehicle being within range of a feature of interest: continue to obtain captured video and audio with the first camera, the second camera, and the microphone to form an extended first video feed, an extended second video feed, and an extended audio feed, and storethe extended first video feed, the extended second video feed, and the extended audio feed in memory, identify one or more relevant images, video clips, and / or audio clips in the extended first video feed, the extended second video feed, and the extended audio feed based on output from first model, the second model, and / or the third model, stitch together the identified one or more relevant images, video clips, and / or audio clips to form a video memo, and display the video memo on a display device of the vehicle when requested and / or send the video memo to an external device. In a first example of the system, the instructions are further executable detect a driver state with at least the first video feed and only enter the first video feed as input to the first model when the driver state meets a conditional state. In a second example of the system, optionally including the first example, the instructions are further executable to, in response to receiving the indication that the key moment trigger has been detected, adjust one or more capture parameters of the first camera relative to the one or more capture parameters used for detecting the driver state. In a third example of the system, optionally including the first and / or second example, the instructions are further executable to cause the processor to adjust one or more capture parameters of the first camera, the second camera, and / or the microphone in response to receiving the indication that the key moment trigger has been detected. In a fourth example of the system, optionally including one or more or each of the first through third examples, the first video feed, the second video feed, and the audio feed are stored in a buffer prior to entering the first video feed as input to the first model, the second video feed as input to the second model, and the audio feed as input to the third model, and wherein the first video feed, the second video feed, and the audio feed are deleted from the buffer in response to receiving an indication from each of the first model, the second model, and the third model that the key moment trigger has not been detected.

[0084] When introducing elements of various embodiments of the present disclosure, the articles “a,” “an,” and '‘the” are intended to mean that there are one or more of the elements. The terms “first,” “second,” and the like, do not denote any order, quantity, or importance, but rather are used to distinguish one element from another. The terms “comprising,” “including,” and “having” are intended to be inclusive and mean that there may be additional elements other than the listed elements. As the terms “connected to.” “coupled to,” etc. are used herein, one object (e g., a material, element, structure, member, etc.) can be connected to or coupled to another object regardless of whether the one object is directly connected or coupled to the other object or whether there are one or more intervening objects between the one object and the other object. In addition, it should be understood that references to “one embodiment” or “anembodiment'’ of the present disclosure are not intended to be interpreted as excluding the existence of additional embodiments that also incorporate the recited features.

[0085] In addition to any previously indicated modification, numerous other variations and alternative arrangements may be devised by those skilled in the art without departing from the spirit and scope of this description, and appended claims are intended to cover such modifications and arrangements. Thus, while the information enhancement described above with particularity and detail in connection with what is presently deemed to be the most practical and preferred aspects, it will be apparent to those of ordinary skill in the art that numerous modifications, including, but not limited to, form, function, manner of operation and use may be made without departing from the principles and concepts set forth herein. Also, as used herein, the examples and embodiments, in all respects, are meant to be illustrative only and should not be construed to be limiting in any manner.

Claims

CLAIMS:

1. A method for automatically generating a video memo with a computing system of a vehicle, comprising: detecting a key moment trigger in a video feed captured with a camera of the vehicle and / or in an audio feed captured with a microphone of the vehicle, wherein the key moment trigger indicates that the vehicle is within range of a feature of interest; in response to detecting the key moment trigger: adjusting one or more capture parameters of the camera, the microphone, and / or one or more additional cameras and / or microphones of the vehicle; saving a first segment of the video feed and a second segment of the audio feed in long-term memory of the vehicle along with third segments of any additional video feeds and / or audio feeds of the one or more additional cameras and / or microphones of the vehicle; automatically generating the video memo by stitching together aspects of the first segment, the second segment, and / or the third segments; and displaying the video memo on a display device of the vehicle when requested and / or sending the video memo to an external device.

2. The method of claim 1. wherein the camera is a first camera of the vehicle and wherein adjusting one or more capture parameters comprises instructing a second camera of the vehicle to begin capturing video to generate a first additional video feed.

3. The method of claim 1 or 2, wherein detecting the key moment trigger comprises entering the video feed as input to a first model trained to detect facial expressions of an occupant of the vehicle and determine that the vehicle is within range of the feature of interest based on the detected facial expressions.

4. The method of any one of the previous claims, wherein detecting the key moment trigger comprises entering the video feed as input to a second model trained to detect the feature of interest in the video feed.

5. The method of any one of the previous claims, wherein detecting the key moment trigger comprises entering the audio feed as input to a third model trained to detect voice expressions of an occupant of the vehicle and determine that the vehicle is within range of the feature of interest based on the detected voice expressions.

6. The method of any one of the previous claims, wherein automatically generating the video memo by stitching together aspects of the first segment, the second segment, and / or the third segments comprises identifying one or more relevant images, video clips, and / or audio clips in the first segment, the second segment, and / or the third segments based on output from an audio engine and a video engine of the computing system of the vehicle, and stitching together the identified one or more relevant images, video clips, and / or audio clips to form the video memo.

7. A system for a vehicle, comprising: a processor; and a non-transitory memory storing a first model, a second model, and a third model and instructions that, when executed, cause the processor to: obtain a first video feed from a first, interior camera of the vehicle, a second video feed from a second, external camera of the vehicle, and an audio feed from a microphone of the vehicle; enter the first video feed as input to the first model, the second video feed as input to the second model, and the audio feed as input to the third model; in response to receiving an indication from one or more of the first model, the second model, and the third model that a key moment trigger has been detected indicative of the vehicle being within range of a feature of interest: continue to obtain captured video and audio with the first camera, the second camera, and the microphone to form an extended first video feed, an extended second video feed, and an extended audio feed, and store the extended first video feed, the extended second video feed, and the extended audio feed in memory; identify one or more relevant images, video clips, and / or audio clips in the extended first video feed, the extended second video feed, and the extended audio feed based on output from first model, the second model, and / or the third model: stitch together the identified one or more relevant images, video clips, and / or audio clips to form a video memo; anddisplay the video memo on a display device of the vehicle when requested and / or send the video memo to an external device.

8. The system of claim 7, wherein the instructions are further executable detect a driver state with at least the first video feed and only enter the first video feed as input to the first model when the driver state meets a conditional state.

9. The system of claim 8. wherein the instructions are further executable to, in response to receiving the indication that the key moment trigger has been detected, adjust one or more capture parameters of the first camera relative to the one or more capture parameters used for detecting the driver state.

10. The system of claim 7, wherein the instructions are further executable to cause the processor to adjust one or more capture parameters of the first camera, the second camera, and / or the microphone in response to receiving the indication that the key moment trigger has been detected.

11. The system of any one of claims 7-10, wherein the first video feed, the second video feed, and the audio feed are stored in a buffer prior to entering the first video feed as input to the first model, the second video feed as input to the second model, and the audio feed as input to the third model, and wherein the first video feed, the second video feed, and the audio feed are deleted from the buffer in response to receiving an indication from each of the first model, the second model, and the third model that the key moment trigger has not been detected.

Citation Information

Patent Citations

  • System and method for multimedia capture

    US20170251163A1