Image processing method and device, vehicle, medium and program product
By using a method that actively identifies and automatically synthesizes key event image fragments, the problems of complex operation and low information density in existing vehicle imaging systems are solved, enabling efficient and complete recording of key events and improving user experience.
Patent Information
- Application Number
- CN202511434279.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-01-13
AI Technical Summary
Existing in-vehicle imaging systems require users to manually edit videos when recording critical events, which is complex and time-consuming, and cannot achieve multi-view recording, resulting in low information density and poor user experience.
By actively identifying key events through vehicles, multiple image fragments related to key events are automatically acquired and synthesized. The multi-camera system is then used for automatic editing and synthesis to generate the target image.
It ensures the integrity and information density of key events, lowers the barrier to entry for users, optimizes the user experience, and saves storage space and computing resources.
Smart Images

Figure CN121334331A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing or smart cockpit, and more particularly to an image processing method, apparatus, vehicle, medium, and program product. Background Technology
[0002] With the development of vehicle-related fields, in-vehicle imaging systems, such as dashcams, have become indispensable electronic devices in modern cars. Their core function is to continuously record visual information during driving, aiming to provide key evidence for determining liability in traffic accidents. Summary of the Invention
[0003] This disclosure provides an image processing method, apparatus, vehicle, medium, and program product for vehicles to actively record key events, ensuring the integrity of the recorded key events, lowering the user's barrier to entry, and optimizing the user experience.
[0004] According to a first aspect of the present disclosure, an image processing method is provided, the method comprising: In response to a user's trigger command for image processing, at least two key events are identified, and image fragments related to the key events are acquired, wherein the image fragments are at least partial images captured by a camera connected in communication with the vehicle; In response to the end of image acquisition, the image fragments related to the at least two key events are combined into a single target image.
[0005] In this embodiment, the vehicle actively identifies key events and acquires image fragments related to the key events, thereby automatically synthesizing multiple image fragments related to key events into a target image. This greatly ensures the integrity of the recorded key events, lowers the user's barrier to entry, and optimizes the user experience.
[0006] In some possible implementations, image fragments related to the key event are acquired, including: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
[0007] In this embodiment, editing all images captured by the camera based on key events ensures data integrity and reliability. Furthermore, by editing all images captured by the camera, users can add or delete frames as needed to synthesize the target image, satisfying their custom editing requirements and offering high processing flexibility.
[0008] In some possible implementations, based on the key event, all images are clipped to obtain image segments related to the key event, including: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
[0009] In this embodiment, by determining the editing nodes based on key events and automatically executing the editing process based on the editing nodes, all valuable image segments can be automatically extracted from the original image data. This automates the image segment acquisition process, making it suitable for large-scale deployment and ensuring the consistency and accuracy of image processing, thereby improving the user experience.
[0010] In some possible implementations, the process of editing all the images based on the key event to obtain image segments related to the key event further includes: Determine the clip duration related to the key events; At least based on the clipping node, all the images are clipped to obtain image fragments related to the key event, including: Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
[0011] In this implementation, the editing duration and editing nodes are determined by key events. For events of different natures, segments with optimized durations can be generated, making the information content and viewing experience of each segment more suitable, thereby improving the usability of the generated content and the user experience.
[0012] In some possible implementations, image fragments related to the key event are acquired, including: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
[0013] In this embodiment, by recognizing key events, the camera can capture only image fragments related to the key events without editing them after capture. This results in fast acquisition of image fragments related to key events with low latency, and eliminates the need to save the complete driving video, thus greatly saving storage space.
[0014] In some possible implementations, combining the image fragments related to the at least two key events into a single target image includes: Post-processing is performed on image segments related to any of the at least two key events; Based on the post-processed image fragments related to any key event, the image fragments related to at least two key events are combined into a single target image.
[0015] In this embodiment, by post-processing one or more of the at least two image segments and then synthesizing them, one or more of the information density, image quality, and visual appeal of the synthesized target image can be significantly improved, thereby effectively enhancing the user experience.
[0016] In some possible implementations, image fragments related to any of the at least two key events are processed, including: Based on the visual features corresponding to the image fragments related to any key event, and / or the user-preset filter style, the parameters of the image fragments are adjusted; The parameters include at least one of the following: brightness, size, contrast, saturation, blur, watermark, cropping size, sharpness, exposure, highlights, shadows, hue, and color temperature.
[0017] In this embodiment, by adjusting the parameters of the image segments based on the visual features corresponding to the image segments related to any key event and / or the user-preset filter style, it is possible to achieve adaptive improvement of image quality based on visual features, and also to make the adjusted image effect conform to the user's personalized aesthetic based on the user-preset filter style, thus significantly improving the user's experience.
[0018] In some possible implementations, the cameras include multiple cameras, and the method further includes: Image fragments related to the at least two key events are captured by the multiple cameras.
[0019] In this embodiment, a single-view camera cannot see other perspectives, and a camera with viewpoint movement function cannot quickly switch angles. A multi-camera system can quickly and from multiple perspectives record events, providing comprehensive visual information, making the content of the final generated target image richer and improving the user experience.
[0020] In some possible implementations, image segments related to the at least two key events are captured by the plurality of cameras, including: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
[0021] In this embodiment, by employing multiple cameras to record image segments related to different key events, it is ensured that each key event is recorded from the optimal perspective, avoiding the need to activate all cameras unnecessarily, and significantly saving computing resources, memory bandwidth, and storage space. Simultaneously, it supports accurate recording of complex scenes, and its architecture is highly flexible. For example, when a new camera is added to the vehicle, the mapping relationship between the camera and the corresponding key event can be added to the decision rule base of the scheduling engine, without needing to refactor the entire system, thus enhancing the system's flexibility and scalability.
[0022] In some possible implementations, combining the image fragments related to the at least two key events into a single target image includes: The image segments related to the current key event and the new key event are processed by image transition and combined into a single target image.
[0023] In this embodiment, by performing screen transition processing on the image segments related to the current key event and the image segments related to the new key event, and merging them into a single target image, the visual discomfort and sense of discontinuity caused by sudden changes in device perspective, position, parameters, etc., are greatly reduced, making the playback effect of the synthesized target image smoother and improving the user experience.
[0024] In some possible implementations, the target image is a set of pictures or a video.
[0025] In this embodiment, two output formats for synthesized images, namely image sets and videos, are provided. These formats can be adapted to different image processing needs in different user scenarios, demonstrating strong flexibility and adaptability, and expanding the application scenarios of image processing methods.
[0026] In some possible implementations, the image segments are video segments, and combining the image segments related to the at least two key events into a single target image includes: Add special effects to video clips related to any of the at least two key events; Based on the image fragments related to any key event after adding special effects, the image fragments related to at least two key events are combined into a single target image.
[0027] In this embodiment, by adding special effects to video clips related to any key event, the final synthesized video has more artistic and technical effects, and can automatically generate more professional and exciting short films, which can greatly stimulate users' desire to share and create, and significantly improve the user experience.
[0028] In some possible implementations, special effects are added to video clips related to any of the at least two key events, including: Determine the event type corresponding to any of the key events; The video effects to be added are determined based on the event type corresponding to any of the key events; Add the video effects to any video clips related to the key event.
[0029] In this embodiment, by automatically identifying the event type corresponding to key events and using it to determine the corresponding special effects, special effects corresponding to different event types can be intelligently matched, which can ensure the efficiency of image synthesis and significantly improve the user experience.
[0030] In some possible implementations, the video effects to be added are determined based on the event type corresponding to any of the key events, including: Based on the event type corresponding to the arbitrary key event, and the correspondence between the video effect and the event type, determine the video effect corresponding to the arbitrary key event.
[0031] In this embodiment, video effects are determined by the correspondence relationship, which can quickly match the effects corresponding to different event types. Moreover, the correspondence relationship is easy to maintain and update, which can greatly reduce the complexity and risk of system iteration and has strong maintainability and scalability.
[0032] In some possible implementations, the method further includes: Add audio to the synthesized target image.
[0033] In this embodiment, by adding audio to the synthesized target image, more professional and exciting short videos can be automatically generated, which can greatly stimulate users' desire to share and create, meet users' personalized needs, and significantly improve the user experience.
[0034] In some possible implementations, the key event includes at least one of the following preset events: Arrive at the designated location; Arrive at a specific time; Identify unusual natural phenomena; It identifies passengers with preconceived emotions; User-triggered camera switching; and Image frame acquisition events triggered by preset time-lapse photography parameters.
[0035] In this embodiment, by providing a variety of preset events, users can record images by setting the location and time. The system can actively perceive the scene and mechanically record the images. It can integrate multi-dimensional scenes such as location, time, environment, emotion and interaction, and realize a three-dimensional narrative of road documentary, natural aesthetics and emotional moments.
[0036] In some possible implementations, the at least two key events include image frame acquisition events triggered by preset time-lapse photography parameters. Before acquiring image fragments related to the key events, the method further includes: Based on the preset time-lapse photography parameters, determine the inter-frame interval corresponding to the key frame; The inter-frame interval is sent to a camera that is connected to the vehicle for the camera to acquire images based on the inter-frame interval to obtain image fragments related to the key event.
[0037] In this embodiment, compared to continuous video recording, sending time-lapse photography parameters to the camera allows the camera to capture single images at intervals, obtaining image fragments related to time-lapse photography. This enables the system to perform long-term event monitoring at extremely low cost, saves a significant amount of storage space, operates with low power consumption, does not affect the normal use of the vehicle, and improves the user experience.
[0038] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: an external vehicle camera and an internal vehicle camera.
[0039] In this embodiment, both the vehicle's external camera and the vehicle's internal camera can be used to acquire image segments. The synthesized target image can achieve all-around image recording, significantly improving the user experience.
[0040] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: a camera that is fixedly connected to the vehicle and a third-party camera that communicates with the vehicle wirelessly.
[0041] In this embodiment, both the camera fixedly connected to the vehicle and the third-party camera that communicates with the vehicle wirelessly can be used to acquire image segments. By working together with the camera fixedly connected to the vehicle and the third-party camera, the physical boundary between the device and the vehicle is broken, and an open and collaborative visual perception network is constructed. This can further realize all-round image recording and significantly improve the flexibility of the function and the user experience.
[0042] In some possible implementations, the image acquisition is performed while the vehicle is in motion or stationary.
[0043] In this implementation, full-time coverage from driving to parking is achieved, ensuring that events occurring at any time and in any place can be fully recorded, realizing comprehensive video recording and significantly improving the flexibility of the function and user experience.
[0044] According to a second aspect of the present disclosure, an image processing apparatus is provided, the apparatus comprising: The first response module is configured to respond to a user's trigger command for image processing, identify at least two key events, and acquire image fragments related to the key events, wherein the image fragments are at least partial images captured by a camera connected in communication with the vehicle; The second response module is configured to combine image fragments related to the at least two key events into a single target image in response to the end of image acquisition.
[0045] In some possible implementations, the first response module is configured to: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
[0046] In some possible implementations, the first response module is configured to: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
[0047] In some possible implementations, the first response module is configured to: Determine the clip duration related to the key events; Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
[0048] In some possible implementations, the first response module is configured to: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
[0049] In some possible implementations, the cameras include multiple cameras, and the image processing device is further configured to: Image fragments related to the at least two key events are captured by the multiple cameras.
[0050] In some possible implementations, the image processing apparatus is further configured to: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
[0051] According to a third aspect of the present disclosure, a vehicle is provided, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to perform the image processing method described in the first aspect of the present disclosure.
[0052] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the image processing method described in the first aspect of the present disclosure.
[0053] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the image processing method described in the first aspect of the present disclosure.
[0054] The technical solutions provided by the embodiments of this disclosure may include the following beneficial effects: This disclosure identifies at least two key events in response to a user's trigger command for image processing, acquires image fragments related to the key events, wherein the image fragments are at least partial images captured by a camera communicatively connected to the vehicle; and, in response to the completion of image acquisition, combines the image fragments related to the at least two key events into a single target image. Thus, by having the vehicle actively identify key events and acquire image fragments related to them, and automatically synthesize multiple image fragments related to key events into a single target image, the integrity of the recorded key events is greatly ensured, the user's learning curve is lowered, and the user experience is optimized.
[0055] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0056] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0057] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0058] Figure 2This is an architectural diagram of an image processing system according to an exemplary embodiment.
[0059] Figure 3 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0060] Figure 4 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0061] Figure 5 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0062] Figure 6 This is a block diagram illustrating an image processing apparatus according to an exemplary embodiment.
[0063] Figure 7 This is a block diagram illustrating a vehicle according to an exemplary embodiment.
[0064] Figure 8 This is a block diagram illustrating a chip system according to an exemplary embodiment. Detailed Implementation
[0065] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0066] In related technologies, with the development of vehicle technology, the application scenarios of in-vehicle imaging systems can be expanded from simple accident evidence collection to multiple dimensions such as recording travel experiences, monitoring driving behavior, and improving driving safety.
[0067] For example, users can record their journey or external scenes using a car camera or mobile phone. Generally, a single camera is used for continuous shooting. The resulting images or videos are lengthy, have low information density, and have a fixed perspective. Users need to spend a lot of time reviewing and editing them to extract valuable clips, and multi-perspective recording is not possible.
[0068] For example, some vehicles are equipped with multi-camera systems, but these systems typically only record independent video streams simultaneously. If a user wants to combine highlights captured by different cameras at different times, such as a sunset ahead, a river view outside the side window, or the smiling faces of passengers inside the vehicle, into a coherent short film, it requires complex manual editing, splicing, and rendering using professional video editing software on a personal computer. This is technically challenging and complex for ordinary users.
[0069] In some exemplary scenarios, when a user drives from the city to a suburban park, the camera function is activated, and the vehicle's front-view camera is used to record road dynamics. When the vehicle enters the cross-river bridge, the user switches to the side window wide-angle camera to capture images of the river surface, and enables a blue filter to enhance the light and shadow on the river surface and increases exposure compensation to deal with backlighting. If a child in the back seat suddenly cheers during the trip, the system can immediately switch to the cabin camera and use the portrait beautification feature to capture a clear smile. After the data collection is completed, the system automatically synthesizes a 30-second multi-view video, which can fully present a three-dimensional narrative of road documentary, natural aesthetics and emotional moments.
[0070] Based on the above exemplary scenarios, an image processing method, device, vehicle, medium, and program product are proposed for vehicles to actively record key events, ensuring the integrity of the recorded key events, reducing the user's usage threshold, and optimizing the user experience.
[0071] Reference Figure 1 , Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 1 As shown, the image processing method may include the following steps.
[0072] In step S101, in response to the user's trigger command for image processing, at least two key events are identified, and image segments related to the key events are acquired. The image segments are at least partial images captured by a camera connected to the vehicle in communication. In step S102, in response to the end of image acquisition, the image fragments related to the at least two key events are combined into a single target image.
[0073] For example, the image processing method can be applied to a vehicle or a mobile terminal connected to the vehicle, such as a mobile phone, tablet computer, camera, or other terminal device.
[0074] For example, user-triggered commands for image processing can be initiated manually by the user. For instance, if a user actively clicks a button such as "Start Creative Shooting" within the system, the user's command for image processing is triggered.
[0075] For example, user-triggered image processing commands can also be triggered by voice commands. For instance, by outputting a voice command such as "Start creative shooting" in the system, the user's image processing command is triggered.
[0076] For example, user-triggered commands for image processing can be triggered passively or automatically. For instance, a user may pre-set to activate the function at a specified time or location, at which time or location the user's trigger command for image processing is activated. For example, the system may detect the occurrence of a pre-defined event, such as a vehicle collision, dangerous behavior, or key events like passing through a specific location, observing a special natural phenomenon, or experiencing a pre-defined emotional state if the user has activated the function.
[0077] For example, key events can include various event types, such as driving safety events, scene recording events, and user-triggered events. Driving safety key events can include vehicle collisions, sudden braking of the vehicle ahead, and the detection of a pedestrian suddenly appearing. Scene recording key events can include arriving at a preset location, arriving at a specific time, recognizing a special natural phenomenon, and recognizing a passenger exhibiting a preset emotion. User-triggered events can include user-triggered camera switching operations, image frame acquisition events triggered by user-preset time-lapse photography parameters, and user-manually controlled recording events.
[0078] For example, a camera that communicates with a vehicle may include a camera that is connected to the vehicle by wired or other means, which may include a camera that is fixedly connected to the vehicle or a third-party camera that communicates with the vehicle wirelessly.
[0079] For example, it can include the vehicle's front-facing dashcam, rear-facing, surround-view, side-view, and cabin cameras, as well as cameras of terminals that communicate with the vehicle, such as cameras of mobile phones and tablets.
[0080] For example, key events can be identified using image recognition algorithms such as OpenCV and YOLO series, or image recognition models, such as identifying special natural phenomena or identifying passengers' pre-set emotions; and / or, key events can be identified using user-preset specific parameters that do not require image recognition, such as preset locations, specific times, camera switching operations, preset time-lapse photography parameters, etc.
[0081] For example, an image segment is at least a portion of an image captured by a camera that is in communication with the vehicle; it may be the entire image or a portion of the image.
[0082] For example, when an image segment is all the images captured by a camera, the total number of images acquired by the camera in a critical event may only be those relevant to that critical event. For instance, a camera might only capture images relevant to the critical event in a critical event, while an image segment could be all the images captured by a camera connected to the vehicle in communication.
[0083] For example, when an image segment is a portion of the images clipped or extracted from all images captured by a camera, all images acquired by the camera in a critical event may include images related to that critical event. For instance, all images acquired by the camera in a critical event may include images related to the critical event, as well as images before and after the critical event; all images acquired by the camera in a critical event may be all images of the vehicle during this or multiple trips.
[0084] For example, the camera can capture all images during the duration of the critical event, including images corresponding to the time period before and / or after the critical event, such as video clips corresponding to the time period 10 seconds before and 5 seconds after the car accident.
[0085] For example, the camera can capture images of the entire journey of the vehicle, and by identifying key events in the images, image segments related to those key events can be determined.
[0086] For example, whether image acquisition has ended can be determined by setting conditions, including any of the following: the user actively ends the shooting, such as the user manually clicking or using voice control to end the shooting; the system determines that the vehicle has reached its destination and has been parked for a long time; the vehicle is turned off; other conditions, etc.
[0087] For example, image fragments can be pictures, multiple pictures, video frames, or video clips, and the target image can be a set of pictures or a video. Image fragments related to at least two key events are synthesized and stitched together according to set rules, such as based on chronological order or a set layout, to form a set of pictures or a complete video.
[0088] For example, after receiving a user's trigger command for image processing, the vehicle can identify key events in real time using preset rules or algorithms. For instance, it can acquire video streams in real time through cameras connected to the vehicle, and computer vision algorithms can analyze the video frames to identify key events; it can also identify key events by monitoring in real time whether set parameters are met.
[0089] For example, upon detecting a critical event, one or more cameras connected to the vehicle acquire image segments related to the critical event. Here, the raw image data acquired by the cameras can be stored in a designated area and timestamped. If all images acquired by the cameras in a critical event are related to that critical event, no editing is required; all images acquired by the cameras in the critical event can be used as the image segments corresponding to that critical event. If all images acquired by the cameras in a critical event include images related to that critical event, the critical event and its corresponding timestamp can be identified to edit the entire set of images, using a portion of the images acquired by the cameras in the critical event as the image segments corresponding to that critical event.
[0090] For example, in the at least two identified key events, video synthesis algorithms, image synthesis algorithms, or chronological stitching can be used to combine the image fragments related to the at least two key events into a video or image set. In some embodiments, a single image can also be synthesized.
[0091] For example, image clips captured by different cameras for the same key event can be composited into a single frame using a picture-in-picture or split-screen layout. For image clips captured for different key events, multiple image clips can be stitched together into a continuous video or image set according to the chronological order of their occurrence; or, based on multiple image clips, keyframes can be extracted from each image clip and arranged on a canvas to generate a single image.
[0092] This disclosure identifies at least two key events in response to a user's trigger command for image processing, acquires image fragments related to the key events, wherein the image fragments are at least partial images captured by a camera communicatively connected to the vehicle; and, in response to the completion of image acquisition, combines the image fragments related to the at least two key events into a single target image. Thus, by having the vehicle actively identify key events and acquire image fragments related to them, and automatically synthesize multiple image fragments related to key events into a single target image, the integrity of the recorded key events is greatly ensured, the user's learning curve is lowered, and the user experience is optimized.
[0093] In some possible implementations, image fragments related to the key event are acquired, including: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
[0094] For example, all the images captured by the camera could be all images acquired during a critical event. These images might all be related to that critical event, or they might include images related to that critical event. Alternatively, all the images captured by the camera could be all images acquired during a complete journey of the vehicle.
[0095] For example, all images captured by the camera, whether they are all images acquired during a key event or all images acquired during a complete journey of the vehicle, can be edited based on the identified key event.
[0096] In one example, the entire set of images captured by the camera represents all images acquired during a complete journey. The images corresponding to key events in the entire set of images can be identified, and the entire set of images can be edited to obtain image segments related to the key events.
[0097] In one example, all images captured by a camera during a key event may be related to that key event. Alternatively, the images can be edited according to the type corresponding to the key event to obtain image segments related to the key event.
[0098] For example, if the key event is sunrise, since the time from sunrise to sunrise is relatively long, a frame-by-frame editing method can be used to extract one frame or multiple key frames at intervals from the sunrise video to obtain image segments related to the key event.
[0099] In one example, a camera may capture all images of a key event, possibly including images related to that key event. Based on the key event, parameters such as clipping nodes and clipping durations can be determined, and all images can be clipped to obtain image segments related to the key event.
[0100] For example, the editing process can be performed before or after image acquisition ends; the editing process can be performed while the vehicle is moving or while the vehicle is stationary, without restriction.
[0101] In some exemplary scenarios, users may need to add image frames 2 seconds before and after a key event for screen transition processing. Users may also need to obtain image fragments corresponding to new events other than the identified key events from all images of the entire journey and use them to synthesize the target image. Or other user-defined scenarios, all images captured by the camera can be edited according to the key events. For example, the complete video library and recorded timestamps can be directly used for re-editing.
[0102] In this embodiment, editing all images captured by the camera based on key events ensures data integrity and reliability. Furthermore, by editing all images captured by the camera, users can add or delete frames as needed to synthesize the target image, satisfying their custom editing requirements and offering high processing flexibility.
[0103] In some possible implementations, based on the key event, all images are clipped to obtain image segments related to the key event, including: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
[0104] For example, a clip node is an anchor point on the timeline of the complete video that identifies the start and / or end of a clip operation. The anchor point can be a precise time point or frame number. A clip node can be a single anchor point or a set of anchor points.
[0105] For example, when the key event retains only one frame of image, the editing node is a single anchor point; when the key event retains multiple frames of image, the editing node is a set of anchor points; when the key event can retain a video segment or multiple frames of images continuously captured over a period of time, the editing node can be the start anchor point and end anchor point corresponding to the editing operation, or the editing node can be the start anchor point or end anchor point of the editing operation, and can be combined with the editing duration for editing.
[0106] For example, key events can be identified by pinpointing the critical points when a key event occurs or the time points corresponding to its duration, which can then be used as editing nodes. In some scenarios, editing nodes can be expanded forward and backward to include both the cause and effect of the event.
[0107] For example, the clipping nodes corresponding to all the images can be identified using a neural network model and set rules. For instance, if a rainbow is detected outside the car during driving, the time point from the appearance of the rainbow to its disappearance is identified as a clipping node; or, if a user is detected performing a camera switching operation, the trigger moment corresponding to the switching operation is identified as a clipping node.
[0108] For example, all images can be edited based at least on the edit nodes. This indicates that in addition to edit nodes, more dimensions of information can be introduced to optimize the editing process and make it more intelligent. For example, the edit duration can be determined based on the event type, with different edit durations for different event types. For instance, driving safety events might be edited by cutting 10 seconds before and after the event, while scene recording events might not cut the images before and after the event, i.e., only the segments corresponding to the key events are edited.
[0109] For example, a complete image file can be read by a video processing engine, and the corresponding position in the image file can be quickly located using the frame number or millisecond-level timecode represented by the clipping node. The located image segment is decoded and precisely cut according to the start node and / or end node. The cut image data is the image segment related to the key event, or the cut video data is re-encoded, and the encoded image data is the image segment related to the key event.
[0110] In this embodiment, by determining the editing nodes based on key events and automatically executing the editing process based on the editing nodes, all valuable image segments can be automatically extracted from the original image data. This automates the image segment acquisition process, making it suitable for large-scale deployment and ensuring the consistency and accuracy of image processing, thereby improving the user experience.
[0111] In some possible implementations, the process of editing all the images based on the key event to obtain image segments related to the key event further includes: Determine the clip duration related to the key events; At least based on the clipping node, all the images are clipped to obtain image fragments related to the key event, including: Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
[0112] For example, the clip duration can be expressed as the total time length of the image segments related to the key event, which can be used to describe the time length corresponding to the image segments.
[0113] For example, the editing duration can be a fixed duration preset by the system for key events, which can be set by event type or not. Users can customize or change the fixed duration. For example, the editing duration can be uniformly set to 10 seconds. For example, the editing duration for key events related to driving safety is 30 seconds, and the editing duration for key events related to scene recording is 15 seconds.
[0114] For example, the clipping duration can be determined by the duration of the key event. For instance, if the duration of the smiling face of the rear passenger is identified as 5 seconds, the clipping duration can be determined to be 5 seconds. Alternatively, based on preset rules, an image segment including the 2 seconds before and 2 seconds after the key event can be clipped, and the clipping duration can be determined to be 9 seconds.
[0115] For example, the editing duration can also be determined based on the duration range of the key event. For instance, the duration range can include 0-5s, 6-10s, 11-20s, 20-60s, and greater than 60s, and the corresponding editing durations can be set to 2s, 10s, 15s, 25s, and 30s, etc. This is just an example.
[0116] For example, if the duration of a smiling face from a rear passenger is identified as 8:00:15-8:00:23, corresponding to a duration of 8 seconds, then according to the aforementioned rules, the corresponding clip duration can be determined to be 10 seconds. The extra 2 seconds can be freely allocated to the image data corresponding to the period before or after the critical event.
[0117] For example, a clipping node can be the start or end anchor point of a clipping operation. Based on the clipping node and the clipping duration, a clipping operation can be performed to obtain image fragments related to the key event. For instance, if the smiling face of a rear passenger is detected to start at 8:00:15 and last for 3 seconds, the clipping start anchor point can be determined to be 8:00:15, the clipping duration to be 3 seconds, and the clipping operation can be performed.
[0118] For example, after identifying a key event, the corresponding clip duration and clipping node can be determined based on the key event. If the clip duration is longer than the duration corresponding to the key event, the total clip duration can be allocated to the period before and after the key event according to a fixed ratio rule, thereby determining the precise clipping range. The image processing engine can use the clipping range to cut out image segments from the complete image.
[0119] In this implementation, the editing duration and editing nodes are determined by key events. For events of different natures, segments with optimized durations can be generated, making the information content and viewing experience of each segment more suitable, thereby improving the usability of the generated content and the user experience.
[0120] In some possible implementations, image fragments related to the key event are acquired, including: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
[0121] For example, the camera can acquire a video stream in real time, and by monitoring the camera's video stream, the key events can be identified in real time. It is understood that the video stream acquired by the camera in real time may not be saved; for example, the video stream acquired when the shutter button is not pressed while taking a photo with a mobile phone.
[0122] For example, when a certain frame or sequence in a video stream is identified as meeting the criteria for determining a key event, the camera can be used to immediately take a picture. Based on the set rules, the shooting duration related to the key event is determined, and the camera is used to capture image segments related to the key event. The vehicle or mobile terminal can then obtain the image segments related to the key event captured by the camera.
[0123] For example, in response to the identification of the key event, the vehicle or mobile terminal can send a shooting command to the camera. The shooting command may include a shooting duration, for capturing and saving image segments related to the key event through the camera. After the shooting is completed, the vehicle or mobile terminal can obtain the image segments related to the key event captured by the camera; or, the vehicle or mobile terminal can obtain the video stream related to the key event captured by the camera in real time. After the key event shooting is completed, the complete video stream related to the key event obtained by the vehicle or mobile terminal is used as the image segments related to the key event.
[0124] In this embodiment, by recognizing key events, the camera can capture only image fragments related to the key events without editing them after capture. This results in fast acquisition of image fragments related to key events with low latency, and eliminates the need to save the complete driving video, thus greatly saving storage space.
[0125] In some possible implementations, combining the image fragments related to the at least two key events into a single target image includes: Post-processing is performed on image segments related to any of the at least two key events; Based on the post-processed image fragments related to any key event, the image fragments related to at least two key events are combined into a single target image.
[0126] For example, any key event is one or more of the at least two key events. Any key event can be some of the at least two key events or all of the at least two key events, without limitation.
[0127] For example, after obtaining the original key event image fragments and before finally synthesizing the target image, each individual image fragment can be processed or enhanced through post-processing to improve the quality, information content, and visual appeal of the individual image fragments.
[0128] For example, post-processing can be used to overlay information, such as embedding data like time, vehicle speed, geographical location, and event type into an image as subtitles. Post-processing can also enhance image quality, such as reducing noise, increasing brightness, and improving contrast in low-light or nighttime footage. Furthermore, post-processing can protect privacy, such as automatically blurring license plate numbers of other vehicles or pedestrian faces in a video clip.
[0129] For example, the post-processing engine reads multiple original image segments from a library of original image segments. It can then perform post-processing operations on one or more image segments in parallel according to preset rules, such as AI-automated recognition or user-preset post-processing parameters, generating post-processed image segments. The post-processed image segments are identical in content to the original image segments, but offer significant improvements in visual appeal and information content. The image processing engine can then stitch together one or more post-processed image segments to create a single target image.
[0130] In this embodiment, by post-processing one or more of the at least two image segments and then synthesizing them, one or more of the information density, image quality, and visual appeal of the synthesized target image can be significantly improved, thereby effectively enhancing the user experience.
[0131] In some possible implementations, image fragments related to any of the at least two key events are processed, including: Based on the visual features corresponding to the image fragments related to any key event, and / or the user-preset filter style, the parameters of the image fragments are adjusted; The parameters include at least one of the following: brightness, size, contrast, saturation, blur, watermark, cropping size, sharpness, exposure, highlights, shadows, hue, and color temperature.
[0132] For example, visual features are quantitative information extracted from the scene corresponding to the image fragment, used to describe the content attributes of the image fragment. Visual features corresponding to image fragments related to any key event can be identified using computer vision technology, image recognition technology, and other methods.
[0133] For example, visual features may include the overall brightness and exposure of the image, such as whether the image is too dark or too exposed; visual features may also include the color distribution of the image, such as whether the image has a color cast, for example, the lighting in a tunnel causing a yellowish tint; visual features may also include the scene type, such as whether it is a city road, a highway, or a natural landscape, and whether the natural landscape is a forest or a lake. Visual features may also include subject sharpness, such as whether key targets such as vehicles and pedestrians have motion blur.
[0134] For example, a user-preset filter style is a fixed combination of image parameters that the user selects in advance according to personal preferences. Filter styles can be used to give image segments a unified and unique visual style.
[0135] For example, a classic filter style allows for adjustments such as enhancing contrast and sharpness, ensuring accurate color reproduction, and achieving a professional documentary feel. A bright and vibrant filter style allows for significant increases in saturation, contrast, and brightness, making the sky bluer and the vegetation greener, perfect for sharing landscapes. A cinematic filter style allows for adjustments such as reducing shadow brightness, adding subtle vignetting, and adjusting the hue to a bluish-green tone, creating a cinematic feel.
[0136] For example, parameters may include one or more of the following: brightness, size, contrast, saturation, blur, watermark, crop size, sharpness, exposure, highlights, shadows, hue, and color temperature.
[0137] For example, brightness controls the overall brightness of the image; exposure mimics camera exposure, affecting the details of the brightest parts of the image; highlights and shadows adjust the details of the brightest and darkest areas of the image respectively; contrast controls the contrast between bright and dark areas of the image; saturation controls the vibrancy of colors in the image; hue controls the overall color tendency of the image, such as leaning towards red or green; color temperature controls the warm or cool feel of the image, such as leaning towards blue or yellow; sharpness or clarity enhances edge details, making object outlines clearer; blurring is used to blur parts or the entire image; size is the cropping size of an image or video frame, used to change video resolution or crop the image to highlight the subject; watermarks can be used to overlay logos, time, location, and other identifying information.
[0138] For example, visual features corresponding to image segments related to any key event can be identified, and the parameters of the image segments can be adjusted based on these visual features. A user-preset filter style can be obtained, and the parameters of the image segments can be adjusted based on this preset filter style. Alternatively, image segments related to any key event can be identified, and a user-preset filter style can be obtained. The parameters of the image segments can then be adjusted based on both the visual features and the user-preset filter style. For instance, after enhancing the image quality of an image segment based on visual features, the enhanced image segment can be adjusted using a user-preset filter style.
[0139] In this embodiment, by adjusting the parameters of the image segments based on the visual features corresponding to the image segments related to any key event and / or the user-preset filter style, it is possible to achieve adaptive improvement of image quality based on visual features, and also to make the adjusted image effect conform to the user's personalized aesthetic based on the user-preset filter style, thus significantly improving the user's experience.
[0140] In some possible implementations, the cameras include multiple cameras, and the method further includes: Image fragments related to the at least two key events are captured by the multiple cameras.
[0141] For example, the cameras may include multiple cameras, which together constitute a multi-view shooting system. For instance, the cameras may include a front-facing main camera, mounted behind the windshield, responsible for the field of view in the main forward direction, and serving as the primary camera for identifying key forward-facing events.
[0142] For example, cameras can include rear-view cameras, typically used for reversing imaging, but also for recording key events and scenery behind the vehicle. Cameras can also include surround-view cameras, mounted on the front and rear bumpers and side mirrors, providing a bird's-eye view, primarily used for low-speed parking, passing oncoming traffic in narrow roads, and recording minor scrapes and collisions. Cameras can also include in-cabin cameras to monitor the driver's status or the condition of passengers / items inside the vehicle. Finally, cameras can include terminal devices with camera capabilities, such as smartphones or tablets, connected to the vehicle.
[0143] For example, the cameras recording key events may include one or more cameras, which may capture image segments related to the at least two key events individually or simultaneously.
[0144] For example, when a key event is identified, image segments related to the key event can be captured simultaneously by multiple cameras. Image data at the same point in time can be acquired synchronously from multiple related cameras, resulting in image segments with overlapping timeframes and multiple perspectives. For example, when one key event is identified, image segments related to the key event can be captured by one camera, and when another key event is identified, image segments related to the other key event can be captured by another camera.
[0145] For example, if there are multiple cameras, acquiring image clips related to at least two key events transforms from a single path into a collaborative network. For instance, if a vehicle's processing unit identifies a key event, it can schedule a group of relevant cameras based on the event type and location. Image clips generated by each camera can be tagged with a unique event ID and / or timestamp to ensure they belong to the same event.
[0146] For example, multiple cameras scheduled for the same critical event can be triggered simultaneously or sequentially, without limitation. This can be set according to design requirements or customized by the user. For instance, if the critical event is the appearance of a rainbow, image segments related to the critical event can be recorded simultaneously by the front-view and rear-view cameras. Alternatively, image segments related to the rainbow can be recorded first by the front-view or side-view cameras, and then, after the vehicle has been driving for a period of time, image segments related to the rainbow can be recorded by the rear-view camera.
[0147] For example, when generating the final target image, multiple image segments from different perspectives can be presented synchronously, regardless of whether they overlap in time. For instance, when playing the main screen from the front-facing camera, image segments from multiple cameras can be played simultaneously using split-screen or picture-in-picture mode. Split-screen mode divides the display screen into one or more display areas, each capable of playing one image segment. Picture-in-picture mode plays one image segment on the main screen while another image segment is played simultaneously within a small window on the main screen.
[0148] In this embodiment, a single-view camera cannot see other perspectives, and a camera with viewpoint movement function cannot quickly switch angles. A multi-camera system can quickly and from multiple perspectives record events, providing comprehensive visual information, making the content of the final generated target image richer and improving the user experience.
[0149] In some possible implementations, image segments related to the at least two key events are captured by the plurality of cameras, including: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
[0150] For example, the first camera and the second camera are different cameras invoked to capture different key events. The first camera may be the camera invoked in response to the current key event, and the second camera may be the camera invoked in response to a new key event. The first camera may be one or more of the plurality of cameras, and the second camera may also be one or more of the plurality of cameras.
[0151] It should be noted that when both the first camera and the second camera are among a plurality of cameras, the first camera and the second camera are different cameras. When either the first camera or the second camera is among many of the plurality of cameras, the first camera and the second camera may have the same camera, but they are not exactly the same.
[0152] For example, the first camera is camera A, and the second camera is camera B; for example, the first camera is camera A, and the second camera is both camera A and camera B; for example, the first camera is both camera A and camera B, and the second camera is both camera B and camera C; for example, the first camera is both camera A and camera B, and the second camera is both camera C and camera D.
[0153] For example, the steps for obtaining a new key event are the same as or similar to those for identifying key events. The steps for obtaining image fragments related to the new key event through the second camera are also the same as or similar to those for obtaining image fragments related to key events. Please refer to the relevant descriptions in the foregoing embodiments, which will not be repeated here.
[0154] For example, when a key event is identified and recorded by the first camera, if a new key event is identified, and a second camera is determined based on the new key event, it is used to acquire image fragments related to the new key event. Image fragments from different cameras that record different events are ultimately used to synthesize the target image.
[0155] For example, when both the first camera and the second camera are among multiple cameras, and the first and second cameras are different cameras, the first camera can be turned off after it has finished acquiring image segments related to the current key event, and the second camera can be turned on to acquire image segments related to the new key event in response to a new key event. The timing of turning off the first camera and turning on the second camera is determined based on the corresponding key event and does not have a fixed order.
[0156] In this embodiment, by employing multiple cameras to record image segments related to different key events, it is ensured that each key event is recorded from the optimal perspective, avoiding the need to activate all cameras unnecessarily, and significantly saving computing resources, memory bandwidth, and storage space. Simultaneously, it supports accurate recording of complex scenes, and its architecture is highly flexible. For example, when a new camera is added to the vehicle, the mapping relationship between the camera and the corresponding key event can be added to the decision rule base of the scheduling engine, without needing to refactor the entire system, thus enhancing the system's flexibility and scalability.
[0157] In some possible implementations, if the times corresponding to the current key event and the new key event do not overlap, the image fragments related to the at least two key events are combined into a single target image, including: The image segments related to the current key event and the new key event are processed by image transition and combined into a single target image.
[0158] For example, the time periods corresponding to the current critical event and the new critical event do not overlap, meaning that the two critical events occur consecutively on the timeline or are spaced apart. The current critical event and the new critical event are not multiple perspectives of the same critical event, nor are they two critical events that occur simultaneously.
[0159] For example, screen transition processing involves visually smoothing the transitions between image segments related to the current key event and new key events. Screen transition processing can reduce the visual jumpiness caused by changes in perspective, position, and parameters when switching devices, resulting in a smoother and more natural playback effect for image-processed videos.
[0160] For example, image transition processing includes changes in transparency, such as gradually increasing the transparency of the last few frames of the current key event and gradually decreasing the transparency of the first few frames of the new key event.
[0161] For example, screen transition processing can include blurring. When the video playback frame rate is high, directional blurring can be applied to the frames between image segments related to the current key event and the new key event to simulate the dynamics of fast movement or switching, masking the sense of jumping, which is suitable for motion scenes.
[0162] For example, image transition processing may include intelligent frame interpolation, in which an AI model can be used to generate appropriate intermediate transition frames between image segments related to the current key event and new key events.
[0163] In this embodiment, by performing screen transition processing on the image segments related to the current key event and the image segments related to the new key event, and merging them into a single target image, the visual discomfort and sense of discontinuity caused by sudden changes in device perspective, position, parameters, etc., are greatly reduced, making the playback effect of the synthesized target image smoother and improving the user experience.
[0164] In some possible implementations, the target image is a set of pictures or a video.
[0165] For example, image fragments can be pictures, multiple pictures, video frames, or video clips, and the target image can be a set of pictures or a video. Image fragments related to at least two key events are synthesized and stitched together according to set rules, such as based on chronological order or a set layout, to form a set of pictures or a complete video.
[0166] It is understood that the image types corresponding to at least two key events can be the same or different, and there is no restriction here. For example, in two key events, the image fragment related to the first key event can be an image, and the image fragment related to the second key event can also be an image; or, the image fragment related to the first key event can be multiple images, and the image fragment related to the second key event can be a video clip.
[0167] In this embodiment, two output formats for synthesized images, namely image sets and videos, are provided. These formats can be adapted to different image processing needs in different user scenarios, demonstrating strong flexibility and adaptability, and expanding the application scenarios of image processing methods.
[0168] In some possible implementations, the image segments are video segments, and combining the image segments related to the at least two key events into a single target image includes: Add special effects to video clips related to any of the at least two key events; Based on the image fragments related to any key event after adding special effects, the image fragments related to at least two key events are combined into a single target image.
[0169] For example, video clips are dynamic video files, not static images. Dynamic video files are the basis for adding special effects. Special effects refer to dynamic visual elements or processing techniques added to enhance visual appeal, highlight key points, or create a specific atmosphere. Examples include slow motion, image shaking, flames, flowing clouds, and other AR element effects.
[0170] For example, the target special effect corresponding to any video segment related to a key event can be determined by the type or attribute of the key event, based on various special effect templates built into the system or updated online, and the target special effect can be added to the video segment related to the key event.
[0171] In this embodiment, by adding special effects to video clips related to any key event, the final synthesized video has more artistic and technical effects, and can automatically generate more professional and exciting short films, which can greatly stimulate users' desire to share and create, and significantly improve the user experience.
[0172] In some possible implementations, special effects are added to video clips related to any of the at least two key events, including: Determine the event type corresponding to any of the key events; The video effects to be added are determined based on the event type corresponding to any of the key events; Add the video effects to any video clips related to the key event.
[0173] For example, the event types corresponding to key events can include driving safety events, driving recorder events, user-triggered events, etc. Key events in the driving safety category can include sudden braking, forward collision, lane departure, pedestrian approach, etc. Key events in the scene recording category can include arriving at a preset location, arriving at a specific time, recognizing a special natural phenomenon, recognizing a passenger exhibiting a preset emotion, etc. Key events in the user-triggered category can include user-triggered camera switching operations, image frame acquisition events triggered by user-preset time-lapse photography parameters, and user-manually activated camera shooting events, etc.
[0174] For example, the system can pre-define the correspondence between the video effects and the event types. Based on the event type and the correspondence, the video effects to be added can be determined and used to add the video effects to video clips related to any key event. For instance, the correspondence can be stored through an event type-effect mapping table within the system. This mapping table is a rule base that defines which effect to apply to a certain type of event.
[0175] In this embodiment, by automatically identifying the event type corresponding to key events and using it to determine the corresponding special effects, special effects corresponding to different event types can be intelligently matched, which can ensure the efficiency of image synthesis and significantly improve the user experience.
[0176] In some possible implementations, the video effects to be added are determined based on the event type corresponding to any of the key events, including: Based on the event type corresponding to the arbitrary key event, and the correspondence between the video effect and the event type, determine the video effect corresponding to the arbitrary key event.
[0177] For example, the system can pre-define the correspondence between the video effects and the event types. Different event types can correspond to different video effects, and different event types can correspond to the same video effect. One event type can correspond to one or more video effects, which is not limited here.
[0178] For example, by establishing a correspondence between video effects and event types, when adding a new event type or a new effect, a new matching rule can be added to the correspondence based on developer configuration updates or user-defined needs, allowing the system to easily adapt to new user requirements. For instance, a pet event type could be added in the future, with "cute stickers" and "fun sound effects" configured for it.
[0179] In this embodiment, video effects are determined by the correspondence relationship, which can quickly match the effects corresponding to different event types. Moreover, the correspondence relationship is easy to maintain and update, which can greatly reduce the complexity and risk of system iteration and has strong maintainability and scalability.
[0180] In some possible implementations, the method further includes: Add audio to the synthesized target image.
[0181] For example, the composited target image refers to the final video file that has undergone all video processing steps. The composited target image contains all key event segments arranged in chronological order, post-processed, and with added visual effects.
[0182] For example, the added audio can be the original ambient sound that comes with the video clip; it can be music that the user has selected in advance from the system's built-in background music library; it can also be short sound effects that the user has selected in advance from the system's sound effects library, such as transition sound effects, prompt sounds, etc.; the added audio can also be music files that the user has uploaded or specified; it can also be intelligently generated voice narration, such as voice used to introduce events, generated through text-to-speech technology, etc.
[0183] For example, selected audio can also be technically processed to ensure it blends perfectly with the video. For instance, volume equalization can be used to ensure that background music does not drown out important original ambient sounds; fade-in / fade-out processing can be used to create smooth volume transitions at the beginning and end of audio or when switching between different music tracks, avoiding abrupt changes; and mixing processing can be used to blend multiple audio tracks, such as background music, sound effects, and original ambient sounds, together in proportion.
[0184] For example, an audio addition scheme can be determined based on a preset strategy or video content analysis. The type of event within the video can be identified; for instance, a scene recording type can be used to add soothing background music. The video processing engine then encapsulates the processed audio track with the video track to generate a new video file containing sound.
[0185] In this embodiment, by adding audio to the synthesized target image, more professional and exciting short videos can be automatically generated, which can greatly stimulate users' desire to share and create, meet users' personalized needs, and significantly improve the user experience.
[0186] In some possible implementations, the key event includes at least one of the following preset events: Arrive at the designated location; Arrive at a specific time; Identify unusual natural phenomena; It identifies passengers with preconceived emotions; User-triggered camera switching; and Image frame acquisition events triggered by preset time-lapse photography parameters.
[0187] For example, preset events are pre-defined conditions or rules that automatically trigger the recording of key events when the conditions are met. Preset events can be set by system defaults or user-defined settings, reflecting the configurability and proactivity of the system.
[0188] For example, arrival at a preset location can be triggered when the vehicle's actual location matches a user-defined point of interest. For instance, if the user has preset a "viewing platform" as a point of interest, the system will automatically trigger and record a stunning panoramic video when the vehicle reaches that location.
[0189] For example, the arrival of a specific time can be automatically triggered by the actual time and the user's preset schedule, such as the user setting a daily "sunset time" trigger, in which the system automatically records a time-lapse video or short video of the sunset process in the evening.
[0190] For example, special natural phenomena can be identified using cameras and AI models, such as rainbows, lightning, and snow. For instance, a vehicle can use computer vision algorithms to detect a rainbow in the sky ahead, automatically triggering recording and potentially matching saturation filters and music.
[0191] For example, identifying a passenger's pre-set emotion can be achieved by analyzing their facial expressions using in-cabin cameras and emotion recognition AI algorithms to determine their emotional state. For instance, if the system detects that the driver is singing or laughing, indicating a very pleasant mood, it will automatically trigger the recording of that footage.
[0192] For example, the user-triggered camera switching operation is a manual camera switching operation performed by the user through the user interface, such as the central control screen button or the voice command "switch to rear camera". This behavior itself is also a key event, representing the user's proactive need to switch perspectives for shooting.
[0193] For example, time-lapse photography is a special photographic technique that compresses and presents a long process by reducing the frame rate of the shot and then playing it back. It is widely used in many fields such as recording astronomical phenomena, changes in natural landscapes, and construction projects. By using time compression technology, a long natural process can be presented in a very short time by continuously shooting photos or video clips.
[0194] Time-lapse photography parameters can include image processing parameters used to control the time-lapse photography effect. These parameters may include the inter-frame interval of frame extraction and the total duration of the entire time-lapse photography plan. Users can customize the time-lapse photography parameters or use the default settings by configuring the shooting mode.
[0195] For example, an image frame acquisition event triggered by preset time-lapse photography parameters is a periodic shooting command. Time-lapse photography parameters can be used to command the camera to periodically acquire a still image. For instance, a one-hour sunset process can be condensed into a short film of about 30 seconds, showing the rapid movement of clouds and the swift setting of the sun, using time-lapse photography parameters.
[0196] In this embodiment, by providing a variety of preset events, users can record images by setting the location and time. The system can actively perceive the scene and mechanically record the images. It can integrate multi-dimensional scenes such as location, time, environment, emotion and interaction, and realize a three-dimensional narrative of road documentary, natural aesthetics and emotional moments.
[0197] In some possible implementations, the at least two key events include an image frame acquisition event triggered by preset time-lapse photography parameters, and the acquisition of image fragments related to the key events includes: acquiring all images captured by the camera; and, according to the preset time-lapse photography parameters, editing the all images to obtain the image fragments related to the key events.
[0198] In some possible implementations, the at least two key events include image frame acquisition events triggered by preset time-lapse photography parameters. Before acquiring image fragments related to the key events, the method further includes: Based on the preset time-lapse photography parameters, determine the inter-frame interval corresponding to the key frame; The inter-frame interval is sent to a camera that is connected to the vehicle for the camera to acquire images based on the inter-frame interval to obtain image fragments related to the key event.
[0199] For example, the inter-frame interval is used to define the time interval or number of frames at which keyframes are extracted from the original video stream. The inter-frame interval determines the degree of time compression. A larger inter-frame interval results in a shorter compressed video. For instance, if the original video is 30fps (30 frames per second), theoretically, one frame per second would need to be extracted to achieve a 30x delay.
[0200] For example, the user can customize or the system can automatically calculate the time compression factor, which represents how many times the actual duration of the event is compared to the final time-lapse video duration. The inter-frame interval can be calculated using the time compression factor. For instance, if a user intends to compress a 1-hour (3600-second) video into a 10-second video, the time compression factor would be 3600 / 10 = 360 times.
[0201] For example, a user can control one or more cameras to start the time-lapse photography function via voice or interface. The user can customize the time compression factor or the system can automatically calculate it. The time compression factor can determine the inter-frame interval corresponding to the key frame.
[0202] For example, the calculated inter-frame interval is sent to a camera connected to the vehicle. The camera can then capture images at intervals based on the received time-lapse photography parameters to obtain image segments related to the time-lapse photography. For instance, the camera can be in a low-power standby state when not acquiring image segments. Each time the clock completes an inter-frame interval, the camera can be woken up to capture a still image before entering sleep mode again.
[0203] For example, after the time-lapse shooting time has elapsed, all the captured photos can be combined into a single image set based on chronological order, resulting in a dynamic video file. This allows for the presentation of slowly changing natural processes or cityscapes in a fast-paced manner, creating visually striking and aesthetically pleasing works.
[0204] In this embodiment, compared to continuous video recording, sending time-lapse photography parameters to the camera allows the camera to capture single images at intervals, obtaining image fragments related to time-lapse photography. This enables the system to perform long-term event monitoring at extremely low cost, saves a significant amount of storage space, operates with low power consumption, does not affect the normal use of the vehicle, and improves the user experience.
[0205] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: an external vehicle camera and an internal vehicle camera.
[0206] For example, an external vehicle camera is a camera installed on the exterior of the vehicle body to perceive the surrounding environment, traffic conditions, and other road users. Examples include the vehicle's front main camera, rear camera, surround view camera, side / blind spot camera, and forward-facing telephoto camera.
[0207] For example, an in-vehicle camera refers to a camera installed inside the vehicle compartment to monitor the status of the driver and passengers, as well as the cabin environment. Examples include driver status monitoring cameras, occupant status monitoring cameras, panoramic cabin cameras, and shooting devices that communicate with the vehicle. These shooting devices can be mobile phones, cameras, drones, tablets, etc.
[0208] For example, the camera communicating with the vehicle may include only an external camera; the camera communicating with the vehicle may include only an internal camera; or the camera communicating with the vehicle may include both an external camera and an internal camera.
[0209] In this embodiment, both the vehicle's external camera and the vehicle's internal camera can be used to acquire image segments. The synthesized target image can achieve all-around image recording, significantly improving the user experience.
[0210] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: a camera that is fixedly connected to the vehicle and a third-party camera that communicates with the vehicle wirelessly.
[0211] For example, a camera permanently connected to a vehicle refers to a camera that is part of the vehicle's original factory configuration or official accessories and is permanently or semi-permanently connected to the vehicle's main controller via a physical cable or dedicated vehicle network. Examples include pre-installed cameras such as front ADAS cameras, surround view cameras, and in-cabin DMS cameras, as well as aftermarket cameras such as dashcams.
[0212] For example, a third-party camera that communicates with a vehicle wirelessly refers to a camera device that is not part of the vehicle's own system but can establish a temporary or long-term connection with the vehicle through a wireless communication protocol and transmit video data. Examples include passenger devices such as smartphones, tablets, cameras, and drones.
[0213] For example, the wireless communication method may include any of the following, such as Wi-Fi, Bluetooth, cellular networks, C-V2X, etc., without limitation.
[0214] In some exemplary scenarios, third-party cameras communicating with vehicles wirelessly may also include external or internal cameras of other vehicles in the vicinity equipped with vehicle-to-vehicle (V2V) technology. Third-party cameras communicating with vehicles wirelessly may also include roadside smart surveillance cameras, traffic light cameras, and so on, communicating with vehicles via V2I technology.
[0215] For example, the camera that communicates with the vehicle may include only a camera that is fixedly connected to the vehicle; the camera that communicates with the vehicle may include only a third-party camera that communicates with the vehicle wirelessly; the camera that communicates with the vehicle may include both a camera that is fixedly connected to the vehicle and a third-party camera that communicates with the vehicle wirelessly.
[0216] In some exemplary scenarios, rear passengers use their mobile phones to take pictures of the scenery outside the window, and the mobile phones share the video stream with the vehicle's infotainment system via Wi-Fi or Bluetooth.
[0217] In some exemplary scenarios, during multi-vehicle accidents, dashcam video from following vehicles is transmitted to the vehicle behind via V2V technology as evidence. Alternatively, a vehicle may receive footage from a traffic camera at an intersection ahead, showing another angle of the accident.
[0218] In this embodiment, both the camera fixedly connected to the vehicle and the third-party camera that communicates with the vehicle wirelessly can be used to acquire image segments. By working together with the camera fixedly connected to the vehicle and the third-party camera, the physical boundary between the device and the vehicle is broken, and an open and collaborative visual perception network is constructed. This can further realize all-round image recording and significantly improve the flexibility of the function and the user experience.
[0219] In some possible implementations, the image acquisition is performed while the vehicle is in motion or stationary.
[0220] For example, the driving process refers to the vehicle being in motion, including all dynamic conditions such as acceleration, deceleration, constant speed cruising, and turning. The stationary state refers to the vehicle being in a non-moving state, including but not limited to temporary parking and parking with the engine off.
[0221] For example, while the vehicle is in motion, the system operates in a high-power, high-performance mode, with all cameras and chips functioning. When the vehicle is stationary, the system may enter a low-power, event-triggered mode, with most of the system in hibernation. It can be woken up by specific signals such as vibration or by a low-power AI unit.
[0222] For example, while the vehicle is in motion, image fragments related to any key event can be dynamically captured and combined into a single target image. Even when the vehicle is stationary, image fragments related to any key event can be continuously captured and combined into a single target image.
[0223] In some exemplary scenarios, the vehicle can continuously record snow scenes while in motion. The vehicle can also continuously record sunrise scenes by using delayed shooting while temporarily stopped.
[0224] In this implementation, full-time coverage from driving to parking is achieved, ensuring that events occurring at any time and in any place can be fully recorded, realizing comprehensive video recording and significantly improving the flexibility of the function and user experience.
[0225] In time-lapse techniques, I-frames can be decoded independently, while P and B frames rely on the previous frame for decoding. Therefore, if the extracted frames are P or B frames, they cannot be directly decoded by the decoder. Currently, the solution is to fully decode all video frames in the video stream and then select the desired frames from the decoded video frames.
[0226] In some exemplary scenarios, time-lapse photography over Ethernet, by connecting to a camera via RTSP (Real Time Streaming Protocol), can acquire H.264 or H.265 video data transmitted by the camera. Each frame of the video data is then decoded to obtain decoded YUV data. The YUV data is then extracted intermittently and encoded to save as a time-lapse video. However, the complete video stream is massive, especially for high-resolution, high-frame-rate video, and the decoding process consumes significant computing resources, resulting in slow processing speed.
[0227] Please refer to Figure 2 Taking a vehicle executing the image processing method of this disclosure as an example, an architecture diagram of an image processing system is shown. The image processing system includes a vehicle infotainment system, which includes a multi-source camera data module and a camera system. The camera system can have various photographic functions such as recording, taking pictures, or time-lapse photography. The multi-source camera data module can be used to process the video data acquired by the camera system. The image processing system also includes at least one camera from a camera device, such as an action camera, drone, mobile phone, or other camera device. At least one camera device can be connected to the vehicle infotainment system via Ethernet or Wi-Fi.
[0228] For example, data from the vehicle's camera system can be transmitted to the multi-source camera data module of the vehicle's infotainment system to obtain information about the camera's capabilities, such as its name and supported functions. The multi-source camera data module of the infotainment system can then transmit video streams to the camera via the RTSP protocol and process the video streams.
[0229] In some exemplary shooting scenarios, when a user drives from the city to a suburban park, image processing is initiated, initially using the vehicle's front-view camera to record road dynamics; after the vehicle enters the cross-river bridge, the user manually switches to the side window wide-angle camera, activates the blue filter to enhance the light and shadow on the river surface, and increases exposure compensation to deal with backlighting; if a child in the back seat suddenly cheers during the journey, the user immediately switches to the cabin camera, calling up its portrait beautification features to capture a clear smile; after the image acquisition is completed, the system can automatically synthesize multi-view videos, fully presenting a three-dimensional narrative of road documentary, natural aesthetics, and emotional moments.
[0230] The vehicle's driving time is relatively long, allowing for smooth switching and real-time parameter adjustment of any heterogeneous camera during time-lapse photography. This is achieved through efficient I-frame extraction and decoding, and parallel post-processing during the switching wait period. This transforms in-vehicle time-lapse photography from a simple recording tool into a powerful multi-view creative platform, while ensuring the efficient and stable operation of the vehicle's infotainment system.
[0231] Please refer to Figure 3-5 The switching process of the camera equipment for time-lapse photography is shown, including the following steps S301-S303.
[0232] In step S301, the first camera device is activated to perform time-lapse photography, including the following steps S3011-S3013.
[0233] In step S3011, the user can pre-set the time compression factor and other parameters of the first camera device on the vehicle display screen and enable time-lapse photography.
[0234] The camera equipment and the vehicle infotainment system are connected via Ethernet, and the camera's capability parameters can be transmitted via the WebSocket protocol. When a camera is selected on the vehicle infotainment system, the user can learn about its capabilities, such as adjusting focus, filters, and watermarks. The user can then adjust the time-lapse photography parameters.
[0235] In step S3012, the inter-frame interval of the I-frame is determined according to the time compression factor, and the inter-frame interval and other parameters of the first camera device are sent to the first camera device.
[0236] The inter-frame interval of I-frames is calculated using the time compression factor, and the formula is as follows: ; ; Where Δt represents the inter-frame interval; R represents the time compression factor; T real Indicates the actual duration of the original video; T play Indicates the duration of the image processing video; Fplay F represents the video frame rate of the camera device. iframe This indicates the frame rate of time-lapse photography. For example, shooting a 1-hour (3600s) sunset and compressing it into a 30-second video: R = 3600s / 30s = 120 times. At 120x compression, if the camera's video frame rate is 30fps; Δt = 120 / 30 = 4 seconds. That is, there are approximately 30 video frames per second, with one I-frame set every 4 seconds.
[0237] After obtaining the inter-frame interval corresponding to the I-frame, the vehicle system transmits the inter-frame interval and other parameters to the camera device via the WebSocket protocol. The camera device can then obtain the video stream according to the time-lapse photography parameters transmitted by the vehicle system.
[0238] In step S3013, the first camera device acquires video stream data according to the inter-frame interval of I-frames and sends the video stream data to the vehicle's infotainment system for decoding and synthesis processing.
[0239] Here, in the vehicle's infotainment system, the multi-source camera data module obtains the video data stream transmitted by the first camera device. Since the first camera device obtains the video stream data according to the inter-frame interval, and the video data stream transmitted by the first camera device has been pre-set to I-frames based on the inter-frame interval, only I-frames can be decoded, thereby reducing CPU consumption.
[0240] After decoding the I-frame, post-processing can be performed to meet personalized needs such as switching camera devices and adjusting parameters during time-lapse photography.
[0241] When no instruction from the user to switch to the second camera device is received, the video stream from the first camera device is continuously acquired.
[0242] In step S302, while the first camera device is continuously recording, the user instructs the user to switch to the second camera device.
[0243] The time compression factor, filters, watermarks, and other photography parameters of the second camera device can be set before switching cameras or pre-set before activating the first camera device. The photography parameters of the second camera device are also transmitted to it by the vehicle's infotainment system via the WebSocket protocol, and the second camera device can acquire the video stream according to the parameters transmitted by the system.
[0244] In step S303, time-lapse photography is performed using a second photographic device.
[0245] In step S3031, the second camera device still determines the I-frame interval and other parameters through the camera parameters transmitted by the vehicle system, and then transmits the I-frame data to the vehicle system for video stream acquisition.
[0246] In step S3032, before the video stream transmitted by the second camera device stabilizes, the data stream of the first camera device undergoes a post-processing process.
[0247] For example, adding transparency, cropping, and other transition effects to the last few I-frames of data decoded by the first camera device.
[0248] In step S3033, after the video stream transmitted by the second camera device stabilizes, the video stream of the first camera device is disconnected, and the I-frames of the second camera device are decoded and synthesized.
[0249] In this embodiment, during time-lapse photography, the camera equipment can be switched and the photography parameters can be adjusted at will, which can better stimulate the user's creative ideas and optimize the user experience.
[0250] In this embodiment, during the waiting period for the I-frame of the second camera device, which is typically 50-300ms, the post-processing of the video stream of the first camera device can be completed in parallel. The switching process of the camera device has motion effects processing to achieve a smooth visual experience of screen switching.
[0251] In this embodiment, the computational load of filters, watermarks, etc., can be completed by the camera device. The camera device is decoupled from the vehicle system, and the performance consumption is mainly in the camera device. The processor load of the vehicle system is reduced, which can ensure the stability of the vehicle system.
[0252] Reference Figure 6 , Figure 6 This is a block diagram illustrating an image processing apparatus 600 according to an exemplary embodiment. (Refer to...) Figure 6 The image processing device 600 includes a first response module 601 and a second response module 602.
[0253] The first response module 601 is configured to respond to a user's trigger command for image processing, identify at least two key events, and acquire image fragments related to the key events, wherein the image fragments are at least a portion of images captured by a camera connected in communication with the vehicle; The second response module 602 is configured to combine image fragments related to the at least two key events into a single target image in response to the end of image acquisition.
[0254] In some possible implementations, the first response module 601 is configured to: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
[0255] In some possible implementations, the first response module 601 is configured to: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
[0256] In some possible implementations, the first response module 601 is configured to: Determine the clip duration related to the key events; Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
[0257] In some possible implementations, the first response module 601 is configured to: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
[0258] In some possible implementations, the second response module 602 is further configured to: Post-processing is performed on image segments related to any of the at least two key events; Based on the post-processed image fragments related to any key event, the image fragments related to at least two key events are combined into a single target image.
[0259] In some possible implementations, the second response module 602 is further configured to: Based on the visual features corresponding to the image fragments related to any key event, and / or the user-preset filter style, the parameters of the image fragments are adjusted; The parameters include at least one of the following: brightness, size, contrast, saturation, blur, watermark, cropping size, sharpness, exposure, highlights, shadows, hue, and color temperature.
[0260] In some possible implementations, the cameras include multiple cameras, and the image processing device 600 is further configured to: Image fragments related to the at least two key events are captured by the multiple cameras.
[0261] In some possible implementations, the image processing apparatus 600 is further configured to: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
[0262] In some possible implementations, where the times corresponding to the current critical event and the new critical event do not overlap, the second response module 602 is further configured to: The image segments related to the current key event and the new key event are processed by image transition and combined into a single target image.
[0263] In some possible implementations, the target image is a set of pictures or a video.
[0264] In some possible implementations, the image fragment is a video fragment, and the second response module 602 is further configured to: Add special effects to video clips related to any of the at least two key events; Based on the image fragments related to any key event after adding special effects, the image fragments related to at least two key events are combined into a single target image.
[0265] In some possible implementations, the second response module 602 is further configured to: Determine the event type corresponding to any of the key events; The video effects to be added are determined based on the event type corresponding to any of the key events; Add the video effects to any video clips related to the key event.
[0266] In some possible implementations, the second response module 602 is further configured to: Based on the event type corresponding to the arbitrary key event, and the correspondence between the video effect and the event type, determine the video effect corresponding to the arbitrary key event.
[0267] In some possible implementations, the image processing apparatus 600 is further configured to: Add audio to the synthesized target image.
[0268] In some possible implementations, the key event includes at least one of the following preset events: Arrive at the designated location; Arrive at a specific time; Identify unusual natural phenomena; It identifies passengers with preconceived emotions; User-triggered camera switching; and Image frame acquisition events triggered by preset time-lapse photography parameters.
[0269] In some possible implementations, the at least two key events include an image frame acquisition event triggered by preset time-lapse photography parameters, and the image processing device 600 is further configured to: Based on the preset time-lapse photography parameters, determine the inter-frame interval corresponding to the key frame; The inter-frame interval is sent to a camera that is connected to the vehicle for the camera to acquire images based on the inter-frame interval to obtain image fragments related to the key event.
[0270] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: an external vehicle camera and an internal vehicle camera.
[0271] In some possible implementations, the camera that is in communication with the vehicle includes at least one of the following: a camera that is fixedly connected to the vehicle and a third-party camera that communicates with the vehicle wirelessly.
[0272] In some possible implementations, the image acquisition is performed while the vehicle is in motion or stationary.
[0273] Regarding the image processing apparatus 600 in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the image processing method, and will not be elaborated upon here.
[0274] Based on the same inventive concept, this disclosure also provides a vehicle, comprising: processor; Memory used to store processor-executable instructions; The processor is configured to execute the image processing method described in this disclosure.
[0275] Based on the same inventive concept, this disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image processing method described in this disclosure.
[0276] Based on the same inventive concept, this disclosure also provides a computer program product, including a computer program that, when executed by a processor, implements the image processing method described in this disclosure.
[0277] Reference Figure 7 , Figure 7 This is a block diagram illustrating a vehicle 700 according to an exemplary embodiment. For example, vehicle 700 can be a hybrid vehicle, a non-hybrid vehicle, an electric vehicle, a fuel cell vehicle, or other types of vehicle. Vehicle 700 can be an autonomous vehicle, a semi-autonomous vehicle, or a non-autonomous vehicle.
[0278] like Figure 7 As shown, vehicle 700 may include various subsystems, such as infotainment system 710, perception system 720, decision control system 730, drive system 740, and computing platform 750. Vehicle 700 may also include more or fewer subsystems, and each subsystem may include multiple components. Furthermore, each subsystem and component of vehicle 700 can be interconnected via wired or wireless means.
[0279] In some embodiments, the infotainment system 710 may include a communication system, an entertainment system, and a navigation system, etc.
[0280] The perception system 720 may include several sensors for sensing information about the environment surrounding the vehicle 700. For example, the perception system 720 may include a global positioning system (which may be GPS, BeiDou, or other positioning systems), an inertial measurement unit (IMU), lidar, millimeter-wave radar, ultrasonic radar, and a camera device.
[0281] The decision control system 730 may include a computing system, a vehicle controller, a steering system, a throttle, and a braking system.
[0282] The drive system 740 may include components that provide powered motion to the vehicle 700. In one embodiment, the drive system 740 may include an engine, an energy source, a transmission system, and wheels. The engine may be one or a combination of internal combustion engines, electric motors, and compressed air engines. The engine is capable of converting energy provided by the energy source into mechanical energy.
[0283] Some or all of the functions of vehicle 700 are controlled by computing platform 750. Computing platform 750 may include at least one processor 751 and memory 752, and processor 751 may execute instructions 753 stored in memory 752.
[0284] Processor 751 can be any conventional processor, such as a commercially available CPU. Processors may also include graphics processing units (GPUs), field-programmable gate arrays (FPGAs), systems-on-chips (SoCs), application-specific integrated circuits (ASICs), or combinations thereof.
[0285] The memory 752 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0286] In addition to instruction 753, memory 752 can also store data, such as road maps, route information, vehicle position, direction, speed, and other data. The data stored in memory 752 can be used by computing platform 750.
[0287] In this embodiment of the disclosure, processor 751 may execute instruction 753 to complete all or part of the steps of the above-described image processing method.
[0288] Some embodiments of this disclosure also provide a chip system, such as Figure 8 As shown, the chip system includes at least one processor 801 and at least one interface circuit 802. The processor 801 and the interface circuit 802 are interconnected via lines. For example, the interface circuit 802 can be used to receive signals from other devices (e.g., the memory of an electronic device). As another example, the interface circuit 802 can be used to send signals to other devices (e.g., the processor 801). Exemplarily, the interface circuit 802 can read instructions stored in memory and send those instructions to the processor 801. When the instructions are executed by the processor 801, the image processing device can perform the steps in the above embodiments. Of course, the chip system may also include other discrete components, and some embodiments of this disclosure do not specifically limit this.
[0289] In some embodiments of this disclosure, the interface circuit 802 can acquire data, program instructions, and / or information from the internal storage area of the chip system; it can also acquire data, program instructions, and / or information from outside the chip system.
[0290] Optionally, the chip system may also include a memory for storing necessary computer programs and data.
[0291] Those skilled in the art will also understand that the various illustrative logical blocks and steps listed in the embodiments of this application can be implemented by electronic hardware, computer software, or a combination of both. Whether such functionality is implemented through hardware or software depends on the specific application and the overall system design requirements. Those skilled in the art can implement the described functionality using various methods for each specific application, but such implementation should not be construed as exceeding the scope of protection of the embodiments of this application.
[0292] It should be understood that, unless otherwise specifically indicated, features of various embodiments of this disclosure described herein can be combined with each other. As used herein, the term “and / or” includes any one of the relevant listed items and any combination of any two or more; similarly, “at least one of…” includes any one of the relevant listed items and any combination of any two or more.
[0293] Although terms such as “first,” “second,” and “third” may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. Rather, these terms are used only to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the examples described herein, the first component, part, region, layer, or section mentioned in the examples may also be referred to as the second component, part, region, layer, or section. Furthermore, the terms “first” and “second” are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as “first” or “second” may explicitly or implicitly include at least one of that feature. In the description herein, “a plurality” means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0294] Furthermore, the term “exemplary” is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as “exemplary” is not necessarily to be construed as advantageous compared to other aspects or designs. Rather, the use of the term “exemplary” is intended to present the concept in a concrete manner. As used herein, the term “or” is intended to mean an inclusive “or” rather than an exclusive “or.” That is, unless otherwise specified or clear from the context, “X applies A or B” is intended to mean any of the natural inclusive arrangements. That is, “X applies A or B” satisfies any of the foregoing instances if X applies A; X applies B; or both X applies A and B. Additionally, unless otherwise specified or clear from the context to refer to the singular form, the articles “a” and “an” as used in this application and the appended claims are generally understood to mean “one or more.”
[0295] Similarly, although this disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. This disclosure includes all such modifications and variations and is limited only by the scope of the claims. In particular, with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terminology used to describe such components is intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if structurally not equivalent to the disclosed structure. Furthermore, although specific features of this disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations, as may be desired and advantageous to any given or particular application. Moreover, with regard to the terms “comprising,” “owning,” “having,” “having,” or variations thereof as used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term “including.”
[0296] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the appended claims.
[0297] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, The method includes: In response to a user's trigger command for image processing, at least two key events are identified, and image fragments related to the key events are acquired, wherein the image fragments are at least partial images captured by a camera connected in communication with the vehicle; In response to the end of image acquisition, the image fragments related to the at least two key events are combined into a single target image.
2. The method according to claim 1, characterized in that, Obtain image fragments related to key events, including: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
3. The method according to claim 2, characterized in that, Based on the key event, all images are edited to obtain image fragments related to the key event, including: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
4. The method according to claim 3, characterized in that, Based on the key event, the images are edited to obtain image fragments related to the key event, and the process further includes: Determine the clip duration related to the key events; At least based on the clipping node, all the images are clipped to obtain image fragments related to the key event, including: Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
5. The method according to claim 1, characterized in that, Obtain image fragments related to key events, including: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
6. The method according to claim 1, characterized in that, Combining the image fragments related to the at least two key events into a single target image includes: Post-processing is performed on image segments related to any of the at least two key events; Based on the post-processed image fragments related to any key event, the image fragments related to at least two key events are combined into a single target image.
7. The method according to claim 6, characterized in that, Processing image fragments related to any of the at least two key events, including: Based on the visual features corresponding to the image fragments related to any key event, and / or the user-preset filter style, the parameters of the image fragments are adjusted; The parameters include at least one of the following: brightness, size, contrast, saturation, blur, watermark, cropping size, sharpness, exposure, highlights, shadows, hue, and color temperature.
8. The method according to claim 1, characterized in that, The camera includes multiple cameras, and before acquiring image fragments related to the key event, the method further includes: Image fragments related to the at least two key events are captured by the multiple cameras.
9. The method according to claim 8, characterized in that, Image fragments related to the at least two key events are captured by the multiple cameras, including: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
10. The method according to claim 9, characterized in that, If the times corresponding to the current key event and the new key event do not overlap, the image fragments related to the at least two key events are combined into a single target image, including: The image segments related to the current key event and the new key event are processed by image transition and combined into a single target image.
11. The method according to any one of claims 1-10, characterized in that, The target image is a set of pictures or a video.
12. The method according to claim 11, characterized in that, The image segments are video segments. Combining the image segments related to the at least two key events into a single target image includes: Add special effects to video clips related to any of the at least two key events; Based on the image fragments related to any key event after adding special effects, the image fragments related to at least two key events are combined into a single target image.
13. The method according to claim 12, characterized in that, Add special effects to video clips related to any of the at least two key events, including: Determine the event type corresponding to any of the key events; The video effects to be added are determined based on the event type corresponding to any of the key events; Add the video effects to any video clips related to the key event.
14. The method according to claim 13, characterized in that, Based on the event type corresponding to any of the key events, determine the video effects to be added, including: Based on the event type corresponding to the arbitrary key event, and the correspondence between the video effect and the event type, determine the video effect corresponding to the arbitrary key event.
15. The method according to any one of claims 1-10, characterized in that, The method further includes: Add audio to the synthesized target image.
16. The method according to any one of claims 1-10, characterized in that, The key events include at least one of the following preset events: Arrive at the designated location; Arrive at a specific time; Identify unusual natural phenomena; It detected that the passenger had a preconceived emotion; User-triggered camera switching; and Image frame acquisition events triggered by preset time-lapse photography parameters.
17. The method according to any one of claims 1-10, characterized in that, The at least two key events include image frame acquisition events triggered by preset time-lapse photography parameters. Before acquiring image fragments related to the key events, the method further includes: Based on the preset time-lapse photography parameters, determine the inter-frame interval corresponding to the key frame; The inter-frame interval is sent to a camera that is connected to the vehicle for the camera to acquire images based on the inter-frame interval to obtain image fragments related to the key event.
18. The method according to any one of claims 1-10, characterized in that, The camera that is in communication with the vehicle includes at least one of the following: an external vehicle camera and an internal vehicle camera.
19. The method according to any one of claims 1-10, characterized in that, The camera that communicates with the vehicle includes at least one of the following: a camera that is fixedly connected to the vehicle and a third-party camera that communicates with the vehicle wirelessly.
20. The method according to any one of claims 1-10, characterized in that, The images are captured while the vehicle is in motion or stationary.
21. An image processing apparatus, characterized in that, The device includes: The first response module is configured to respond to a user's trigger command for image processing, identify at least two key events, and acquire image fragments related to the key events, wherein the image fragments are at least partial images captured by a camera connected in communication with the vehicle; The second response module is configured to combine image fragments related to the at least two key events into a single target image in response to the end of image acquisition.
22. The apparatus according to claim 21, characterized in that, The first response module is configured as follows: Acquire all images captured by the camera; Based on the key event, all the images are edited to obtain image fragments related to the key event.
23. The apparatus according to claim 22, characterized in that, The first response module is configured as follows: Based on the key events, determine the clipping nodes in all the images; At least based on the clipping nodes, all the images are clipped to obtain image fragments related to the key event.
24. The apparatus according to claim 23, characterized in that, The first response module is configured as follows: Determine the clip duration related to the key events; Based on the clipping node and the clipping duration, image fragments related to the key event are obtained.
25. The apparatus according to claim 21, characterized in that, The first response module is configured as follows: In response to the identification of the key event, the camera captures image fragments related to the key event, thereby acquiring the image fragments related to the key event.
26. The apparatus according to claim 21, characterized in that, The cameras include multiple cameras, and before acquiring image segments related to the key event, the image processing device is further configured to: Image fragments related to the at least two key events are captured by the multiple cameras.
27. The apparatus according to claim 26, characterized in that, The image processing device is further configured to: The first camera was used to capture image fragments related to this critical event. In response to the identification of a new key event, the second camera corresponding to the new key event is determined, and image fragments related to the new key event are acquired through the second camera.
28. A vehicle, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to perform the image processing method according to any one of claims 1-20.
29. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image processing method according to any one of claims 1-20.
30. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the image processing method according to any one of claims 1-20.