Video generation method and device, electronic equipment and vehicle
By acquiring vehicle information and driving scenarios and using pre-trained video generation models to generate driving operation videos, the problems of long production cycles and high costs are solved, and rapid generation and reuse are achieved.
Patent Information
- Application Number
- CN202411950162.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2044-12-27
AI Technical Summary
In the existing technology, the production cycle of driving operation videos is long and cannot be reused, resulting in high costs.
By obtaining the target vehicle's model information and driving scenarios, a pre-trained target video generation model is used to generate driving operation videos, avoiding direct shooting or animation simulation.
It enables rapid generation of driving operation videos, reduces production costs, and supports reuse of different vehicles and scenes.
Smart Images

Figure CN119788932B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of vehicle control technology, and in particular to a video generation method, device, electronic equipment, and vehicle. Background Art
[0002] With the development of the automotive industry, driving operation videos can help users understand the details of the driving process. Because driving operations involve many scenarios and operational details, detailed explanations of driving operations through videos can help users understand driving services more easily. However, filming actual vehicles or creating driving operation videos through animation simulations can lead to long production cycles, high costs, and the inability to reuse driving operation videos for different models.
[0003] In view of this, how to avoid the long production cycle and non-reusability of driving operation videos has become a technical problem that needs to be solved urgently. Summary of the Invention
[0004] In view of this, the purpose of the present disclosure is to propose a video generation method, device, electronic device and vehicle to solve the problem in the prior art that driving operation videos have a long production cycle and are not reusable.
[0005] Based on the above objectives, the first aspect of the present disclosure provides a video generation method, the method comprising:
[0006] Obtain target vehicle model information and target driving scenario;
[0007] A pre-trained target video generation model is used to generate a target driving operation video of the target vehicle in the target driving scenario according to the target vehicle model information and the target driving scenario.
[0008] Based on the same inventive concept, the second aspect of the present disclosure provides a video generation device, comprising:
[0009] an acquisition module configured to acquire target vehicle model information and a target driving scenario of a target vehicle;
[0010] The generation module is configured to use a pre-trained target video generation model to generate a target driving operation video of the target vehicle in the target driving scenario according to the target vehicle model information and the target driving scenario.
[0011] Based on the same inventive concept, the third aspect of the present disclosure proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, wherein the processor implements the method described above when executing the computer program.
[0012] Based on the same inventive concept, a fourth aspect of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described above.
[0013] Based on the same inventive concept, a fifth aspect of the present disclosure provides a vehicle comprising the video generation apparatus of the second aspect or the electronic device of the third aspect or the storage medium of the fourth aspect.
[0014] As can be seen from the above, the present disclosure provides a video generation method, apparatus, electronic device and vehicle. The target vehicle model information and the target driving scene of a target vehicle are obtained. A target driving operation video of the target vehicle in the target driving scene is generated by using a pre-trained target video generation model according to the target vehicle model information and the target driving scene. In this way, the target driving operation video of the target vehicle can be generated by using the target video generation model, without the need to shoot or animate the target vehicle to make the target driving operation video, thereby reducing the production cycle of the target driving operation video. The vehicle model information of different vehicles is input into the target video generation model, which can generate driving operation videos of different vehicles, thereby realizing the reuse of the target video generation model and reducing the production cost of the driving operation video. BRIEF DESCRIPTION OF DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the present disclosure or the related art, the drawings needed to be used in the embodiments or related art descriptions will be briefly introduced. Obviously, the drawings in the following description are only embodiments of the present disclosure, and other drawings can be obtained by those skilled in the art without creative labor.
[0016] Figure 1 A flowchart of the video generation method of the embodiments of the present disclosure;
[0017] Figure 2 A flowchart of the AIGC-based intelligent driving operation video generation method of the embodiments of the present disclosure;
[0018] Figure 3 A structural schematic diagram of the video generation apparatus of the embodiments of the present disclosure;
[0019] Figure 4 A structural schematic diagram of the electronic device of the embodiments of the present disclosure. DETAILED DESCRIPTION
[0020] In order to make the purpose, technical solutions and advantages of the present disclosure clearer, the present disclosure will be further described in detail below with reference to specific embodiments and drawings.
[0021] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.
[0022] Based on the background information provided, vehicle intelligent driving technology is currently experiencing rapid development, primarily achieving Level 2 intelligent driving capabilities, with some models already equipped with Level 3 intelligent driving capabilities. Intelligent driving technology has significantly reduced traffic accidents, improved traffic efficiency, and provided a more comfortable driving experience. Because intelligent driving involves numerous scenarios and operational details, vehicle manufacturers (OEMs) need to produce intelligent driving operation videos for specific models equipped with different intelligent driving features, providing detailed explanations through videos to facilitate user understanding of intelligent driving services.
[0023] Currently, there are two methods for producing intelligent driving operation videos: one is for a filming team to film the actual vehicle, and the other is for a styling video team to create animation simulations. However, both methods involve long production cycles, high capital investment, and the fact that each vehicle model must be produced individually and cannot be reused.
[0024] The terms used in this disclosure are explained as follows:
[0025] Artificial Intelligence Generated Content (AIGC) refers to the use of artificial intelligence technology to generate various forms of content and data, including text, images, audio, video, etc.
[0026] Navigate on Autopilot (NOA) is an advanced driver assistance feature that allows vehicles to operate autonomously on highways, including lane keeping, automatic lane changing, autonomous cruise control, and navigation functions.
[0027] Adaptive Cruise Control (ACC) is an intelligent automatic control system that is mainly used to assist the driver in controlling the vehicle's speed and maintaining a safe distance from the vehicle in front.
[0028] The Highway Assist (HWA) system is an advanced system based on L2 autonomous driving technology. It integrates front and rear corner radars, front millimeter-wave radars, front-view cameras, domain controllers, panoramic surround-view cameras, and ultrasonic radars to achieve accurate detection of the surrounding environment, thereby providing assisted driving functions on highways.
[0029] Human Machine Interface (HMI) is the medium for interaction and information exchange between the system and the user, realizing the conversion between the internal form of information and the form acceptable to humans.
[0030] The Electronic Park Brake (EPB) system is an electronically controlled parking brake system. The electronic control unit receives the driver's instructions and controls the motor to drive the compression nut, thereby achieving braking.
[0031] The Autonomous Emergency Braking (AEB) system is a vehicle active safety technology, mainly composed of three core parts: a control module, a ranging module and a braking module.
[0032] The Electronic Stability Program (ESP) is primarily designed to enhance vehicle safety and handling. It monitors the vehicle's driving state and, in emergency situations, intervenes with braking to help maintain dynamic balance, preventing oversteer or understeer during cornering or emergency obstacle avoidance.
[0033] As mentioned above, how to avoid the long production cycle and non-reusability of driving operation videos has become an important research issue.
[0034] Based on the above description, if Figure 1 As shown, the video generation method proposed in this embodiment includes:
[0035] Step 101: Obtain target vehicle model information and target driving scenario of a target vehicle.
[0036] In a specific implementation, the target vehicle is the vehicle for which a driving operation video is to be produced. The target vehicle model information is the model information of the target vehicle. The target driving scenario is the intelligent driving scenario configured for the target vehicle.
[0037] The target vehicle model information can be a target vehicle model picture or a target vehicle model video. In the embodiment of the present disclosure, the target vehicle model information is preferably a target vehicle model picture. The target vehicle model information includes the target vehicle body information, interior information, exterior information, instrument information and host information.
[0038] For example, if the target vehicle's model information is model C and is configured with smart driving scenario a and smart driving scenario b, then the target vehicle model information is determined to be model C, and the target driving scenarios are smart driving scenario a and smart driving scenario b.
[0039] Step 102 , using a pre-trained target video generation model, based on the target vehicle model information and the target driving scene, generates a target driving operation video of the target vehicle in the target driving scene.
[0040] In specific implementation, the target vehicle model information and target driving scenario are input into a pre-trained target video generation model to generate a target driving operation video of the target vehicle in the target driving scenario. The target driving operation video is a video of the intelligent driving operation of the target vehicle, which contains the target vehicle model information (body information, interior information, exterior information, instrument information, and host information) and the configured target driving scenario.
[0041] For example, the target vehicle model information is model C, and the configured target driving scenarios are intelligent driving scenario a and intelligent driving scenario b. Model C, intelligent driving scenario a, and intelligent driving scenario b are input into the target video generation model to generate target driving operation videos of model C in intelligent driving scenario a and target driving operation videos of model C in intelligent driving scenario b.
[0042] In this way, simply by inputting the target vehicle model information and the target driving scenario, the target video generation model can be used to generate a target driving operation video, which can be used to guide vehicle users in understanding related intelligent driving operation services. The target video generation model can quickly generate the target driving operation video for the target vehicle, eliminating the need to film or animate the target vehicle. This has the advantages of a short production cycle and low capital expenditure. The target video generation model can be reused, allowing for batch production of intelligent driving operation videos with different models and different intelligent driving scenarios, meeting the growing demand for intelligent driving videos.
[0043] Through the above embodiment, the target vehicle model information and target driving scenario of the target vehicle are obtained. Using a pre-trained target video generation model, a target driving operation video of the target vehicle in the target driving scenario is generated based on the target vehicle model information and the target driving scenario. In this way, the target video generation model can be used to generate the target driving operation video of the target vehicle, eliminating the need to film or animate the target vehicle to produce the target driving operation video, thereby reducing the production cycle of the target driving operation video. By inputting the target video generation model with the model information of different vehicles, driving operation videos of different vehicles can be generated, enabling the reuse of the target video generation model, thereby reducing the production cost of the driving operation videos.
[0044] When generating a target driving operation video of a target vehicle, in order to make the target driving operation video more vivid, the target driving operation video includes the target vehicle model information of the target vehicle. The target driving operation video is generated based on the target vehicle model information and the target driving scene. The specific process is as follows:
[0045] In some embodiments, the target vehicle model information includes: target appearance information and target interior information; step 102 includes:
[0046] Step 1021 : Generate an initial driving operation video under the target driving scenario based on the target driving scenario using the target video generation model.
[0047] Step 1022 : Using the target video generation model, the target appearance information and the target interior information are added to the initial driving operation video to obtain a target driving operation video of the target vehicle in the target driving scene.
[0048] In a specific implementation, the target video generation model is used to generate an initial driving operation video under the target driving scene according to the target driving scene, wherein the initial driving operation video includes the target driving scene of the target vehicle.
[0049] For example, the target vehicle is configured with target driving scenarios of smart driving scenario a and smart driving scenario b. Smart driving scenario a and smart driving scenario b are input into the target video generation model to generate an initial driving operation video under smart driving scenario a and an initial driving operation video under smart driving scenario b.
[0050] The target appearance information (body information and exterior information) and target interior information (interior information, instrument information and host information) of the target vehicle are added to the initial driving operation video to obtain the target driving operation video of the target vehicle in the target driving scenario.
[0051] For example, the target vehicle model information is vehicle model C (exterior information C and interior information C), the exterior information C and the interior information C are added to the initial driving operation video to obtain the target driving operation video of vehicle model C in the intelligent driving scene a and the target driving operation video of vehicle model C in the intelligent driving scene b.
[0052] Through the above scheme, the initial driving operation video under the target driving scene can be generated by using the target video generation model. The target appearance information and the target interior information of the target vehicle are added to the initial driving operation video, so that the target driving operation video obtained contains the target vehicle model information and the target driving scene.
[0053] In order to be able to more accurately generate the initial driving operation video containing the target driving scene. The target driving scene of the target vehicle is judged, and the specific process is as follows:
[0054] In some embodiments, the target driving scene includes at least one of: ramp-in and ramp-out scene, lane change protection scene and distraction detection scene; step 1021 includes:
[0055] Step 1021A, judging the target driving scene of the target vehicle.
[0056] Step 1021B, in response to determining that the target driving scene is a ramp-in and ramp-out scene, determining a ramp-in and ramp-out video generation sub-model from the target video generation model, and generating an initial driving operation video of ramp-in and ramp-out by using the ramp-in and ramp-out video generation sub-model.
[0057] Step 1021C, in response to determining that the target driving scene is a lane change protection scene, determining a lane change protection video generation sub-model from the target video generation model, and generating an initial driving operation video of lane change protection by using the lane change protection video generation sub-model.
[0058] Step 1021D, in response to determining that the target driving scene is a distraction detection scene, determining a distraction detection video generation sub-model from the target video generation model, and generating an initial driving operation video of distraction detection by using the distraction detection video generation sub-model.
[0059] In specific implementation, the target video generation model includes: a ramp-in and ramp-out video generation sub-model, a lane change protection video generation sub-model and a distraction detection video generation sub-model.
[0060] When the target driving scenario configured for the target vehicle is an on-ramp or off-ramp scenario, the on-ramp or off-ramp video generation sub-model is used to generate the initial driving operation video for the on-ramp or off-ramp. When the target driving scenario configured for the target vehicle is a lane change protection scenario, the lane change protection video generation sub-model is used to generate the initial driving operation video for lane change protection. When the target driving scenario configured for the target vehicle is a distraction detection scenario, the distraction detection video generation sub-model is used to generate the initial driving operation video for distraction detection.
[0061] Through the above scheme, the target driving scene of the target vehicle is judged, and the initial driving operation video is generated using the sub-model corresponding to the target driving scene, avoiding the problem of directly using the target video generation model to generate the initial driving operation video and wasting resources.
[0062] In order to generate the target driving operation video more accurately, the target video generation model is pre-trained. The training process of the target video generation model is as follows:
[0063] In some embodiments, the training process of the target video generation model includes:
[0064] Step 201 : Obtain a sample video of a driving operation, extract sample driving information of the driving operation from the sample video, and use the sample driving information to train a general video generation model to obtain an initial video generation model.
[0065] In a specific implementation, the sample driving operation video is sample data of an intelligent driving operation video. There are multiple sample driving operation videos. The universal video generation model is a universal large model for generating videos.
[0066] The general video generation model is trained using the sample driving information in the sample videos to obtain an initial video generation model. In this way, the initial video generation model is trained to be a model specifically for producing intelligent driving operation videos. The initial video generation model is used to generate intelligent driving operation videos.
[0067] Step 202: Obtain sample vehicle model information and sample intelligent driving scenes, and use the sample vehicle model information and the sample intelligent driving scenes to train the initial video generation model to obtain a target video generation model.
[0068] In specific implementation, the sample vehicle model information is sample data of the vehicle model, and the sample intelligent driving scenario is sample data of the intelligent driving scenario.
[0069] The initial video generation model is trained using sample vehicle model information and sample intelligent driving scenarios to obtain the target video generation model. This model is then trained to produce intelligent driving operation videos of different vehicles. The target video generation model is then used to generate intelligent driving operation videos of the target vehicle in the target intelligent driving scenario.
[0070] For example, the sample vehicle model information includes: vehicle model A, vehicle model B, and vehicle model C, and the sample intelligent driving scenarios include: intelligent driving scenario a, intelligent driving scenario b, and intelligent driving scenario C. The initial video generation model is trained using the sample vehicle model information and the sample intelligent driving scenarios, and the target video generation model is obtained to be able to generate intelligent driving operation videos for the sample vehicle model information and the sample intelligent driving scenarios.
[0071] When the target vehicle model information is model C, and the target intelligent driving scenarios include intelligent driving scenarios a and b, model C, intelligent driving scenarios a, and intelligent driving scenarios b are input into the target video generation model to generate an intelligent driving operation video of the target vehicle. The intelligent driving operation video of the target vehicle includes: a first intelligent driving operation video of the target vehicle (including model C) in intelligent driving scenario a, and a second intelligent driving operation video of the target vehicle (including model C) in intelligent driving scenario b.
[0072] The above solution accurately extracts sample driving information from sample driving videos. This information is used to train a general video generation model, enabling the resulting initial video generation model to specifically produce intelligent driving operation videos. The initial video generation model is then trained using sample vehicle model information and sample intelligent driving scenarios, enabling the resulting target video generation model to produce intelligent driving operation videos for different vehicles.
[0073] In some embodiments, step 201 includes:
[0074] Step 2011: extract sample start conditions, sample scenes, sample operations, and sample exit conditions of intelligent driving from the sample video.
[0075] Step 2012: Train the universal video generation model using the sample start condition, the sample scene, the sample operation, and the sample exit condition to obtain an initial video generation model.
[0076] In specific implementation, the sample video of intelligent driving includes: sample start-up conditions, sample scenarios, sample operations and sample exit conditions of intelligent driving.
[0077] The universal video generation model is trained using sample startup conditions, sample scenes, sample operations, and sample exit conditions of intelligent driving, so that the universal video generation model learns the startup, scenes, operations, and exit of intelligent driving to obtain an initial video generation model.
[0078] In this way, the initial video generation model is trained as a model specifically for producing intelligent driving operation videos, and the initial video generation model is used to generate intelligent driving operation videos including startup, scenes, operations, and exits.
[0079] Through this approach, a universal video generation model is trained using sample startup conditions, sample scenarios, sample operations, and sample exit conditions for intelligent driving. The universal video generation model learns the startup, scenarios, operations, and exit conditions for intelligent driving. This trained initial video generation model is capable of generating intelligent driving operation videos that include the startup conditions, scenarios, operations, and exit conditions for intelligent driving.
[0080] In some embodiments, step 2012 includes:
[0081] Step 2012A: Send the sample start condition, the sample scene, the sample operation, and the sample exit condition to the universal video generation model to obtain a first simulation video of intelligent driving.
[0082] Step 2012B: compare the first simulation video with the pre-stored first test video to obtain a first comparison result.
[0083] Step 2012C: Based on the first comparison result, update the parameters of the universal video generation model to obtain an initial video generation model.
[0084] In specific implementations, sample start conditions, sample scenarios, sample operations, and sample exit conditions are input into the universal video generation model, which then outputs a first simulated video of intelligent driving. The first simulated video is a video of intelligent driving operations generated by the universal video generation model based on the sample videos during training.
[0085] The first test video is a standard video used to test the universal video generation model. The first simulation video and the first test video are compared to obtain a first comparison result, and a first loss function is determined based on the first comparison result. Based on the first loss function, the parameters of the universal video generation model are updated. When the first loss function approaches a first preset value, the updated universal video generation model is used as the initial video generation model.
[0086] Through this approach, the universal video generation model is trained using sample startup conditions, sample scenarios, sample operations, and sample exit conditions for intelligent driving, enabling the resulting initial video generation model to be specifically designed to produce intelligent driving operation videos. Furthermore, based on a first comparison result between a first simulated video output by the universal video generation model and a pre-stored first test video, the parameters of the universal video generation model can be accurately adjusted and updated, thereby further improving the accuracy of the trained initial video generation model.
[0087] In some embodiments, step 2012 includes:
[0088] Step 2012a: Use the sample starting conditions to train the universal video generation model to obtain a first video generation model.
[0089] Step 2012b: Use the sample scene to train the first video generation model to obtain a second video generation model.
[0090] Step 2012c: Use the sample operation to train the second video generation model to obtain a third video generation model.
[0091] Step 2012d: Use the sample exit condition to train the third video generation model to obtain an initial video generation model.
[0092] In a specific implementation, the sample activation conditions are used to train the universal video generation model to obtain a first video generation model. The trained first video generation model is a universal video generation model that has learned the intelligent driving activation conditions and can generate videos of intelligent driving activation.
[0093] The first video generation model is trained using sample scenarios to generate a second video generation model. The trained second video generation model is a general video generation model that has learned intelligent driving scenarios and can generate videos of intelligent driving startup and different scenarios.
[0094] The second video generation model is trained using the sample operations to generate a third video generation model. The trained third video generation model is a general video generation model that has learned intelligent driving operations and is capable of generating videos of intelligent driving startup, different scenarios, and corresponding operations.
[0095] The sample exit conditions are used to train the third video generation model to obtain an initial video generation model. The trained initial video generation model is a general video generation model that has learned intelligent driving scenarios and can generate videos of intelligent driving startup, different scenarios, corresponding operations, and exit.
[0096] In addition, step 2012 also includes: using the sample start condition to train the universal video generation model to obtain a first video generation model; using the sample scene to train the universal video generation model to obtain a second video generation model; using the sample operation to train the universal video generation model to obtain a third video generation model; using the sample exit condition to train the universal video generation model to obtain a fourth video generation model; combining the first video generation model, the second video generation model, the third video generation model and the fourth video generation model to obtain an initial video generation model.
[0097] Through the above scheme, the universal video generation model is trained sequentially using sample startup conditions, sample scenarios, sample operations, and sample exit conditions. This allows the universal video generation model to sequentially learn the startup conditions, scenarios, operations, and exit conditions for intelligent driving. This allows the trained initial video generation model to generate videos that include intelligent driving startup, different scenarios, and corresponding operations and exits, enabling the initial video generation model to specifically produce intelligent driving operation videos.
[0098] In some embodiments, the process of extracting the sample starting condition includes:
[0099] Step 20111, extracting the gear status, electronic parking brake system status, vehicle navigation status, driver's seat belt status, door status, hood status, trunk status, automatic emergency brake system status, vehicle electronic stability system status and vehicle speed status from the sample video.
[0100] Step 20112, based on the gear status, the electronic parking brake system status, the vehicle navigation status, the main driver's seat belt status, the door status, the hood status, the trunk status, the automatic emergency braking system status, the vehicle electronic stability system status and the vehicle speed status, determine the sample startup conditions for intelligent driving.
[0101] In specific implementation, the sample startup conditions for intelligent driving include: the gear status is forward gear, the EPB status is inactive, the vehicle navigation status is navigation on (that is, a starting point and an end point are set), the main driver's seat belt status is fastened, the door status is closed, the hood status and the trunk status are closed, the AEB status and the ESP status are on, and the vehicle speed status is within the preset driving speed range.
[0102] The preset driving speed range is the speed range of the vehicle under normal driving conditions, for example, the preset driving speed range may be 10 km / h-120 km / h.
[0103] Through the above scheme, based on the gear status, electronic parking brake system status, vehicle navigation status, main driver's seat belt status, door status, hood status, trunk status, automatic emergency braking system status, vehicle electronic stability system status and vehicle speed status, the sample startup conditions of intelligent driving can be comprehensively and accurately determined.
[0104] In some embodiments, the sample scene extraction process includes:
[0105] Step 20113: extracting the on-ramp and off-ramp scenes, lane change protection scenes, distraction detection scenes, entrance avoidance scenes, vehicle avoidance scenes, and intersection identification scenes from the sample video.
[0106] Step 20114, taking the on-ramp and off-ramp scenario, the lane change protection scenario, the distraction detection scenario, the avoidance entrance scenario, the vehicle avoidance scenario and the intersection identification scenario as sample scenarios for intelligent driving.
[0107] In practice, ramp-in / out scenarios are implemented by setting routes based on high-precision map navigation, enabling intelligent ramp-in / out. Specifically, ramp features are extracted from sample videos and then used to determine the ramp-in / out scenarios. These ramp-in / out scenarios include both ramp entry and ramp exit scenarios.
[0108] The lane change protection scenario monitors the status of the vehicle behind in real time and promptly terminates the lane change if a fast-moving vehicle approaches, reducing the risk of rear-end collisions. Specifically, lane change features are extracted from sample videos and the lane change protection scenario is determined based on these features.
[0109] The distraction detection scenario monitors the driver's status in real time. When the driver is distracted, fatigued, or the current operating environment does not support the system's continued control of the vehicle, different levels of prompts and alarms are issued. Specifically, driver characteristics are extracted from sample videos and the distraction detection scenario is determined based on these characteristics.
[0110] The merging entrance avoidance scenario uses navigation information to accurately identify the merging lane ahead and change lanes in advance to avoid the merging entrance. Specifically, merging entrance features are extracted from sample videos and the merging entrance avoidance scenario is determined based on these features.
[0111] The vehicle avoidance scenario involves real-time monitoring of surrounding vehicles. When a large truck is driving in an adjacent lane, intelligent avoidance is implemented to ensure safety in multiple dimensions. Specifically, the vehicle features of the surrounding vehicles are extracted from sample videos and the vehicle avoidance scenario is determined based on these features.
[0112] Intersection scene detection involves real-time monitoring of the presence of an intersection ahead and, when approaching an intersection, lane changes to ensure efficient and safe driving. Specifically, intersection features are extracted from sample videos and the intersection scene is determined based on these features.
[0113] Through the above scheme, the entry and exit ramp scenes, lane change protection scenes, distraction detection scenes, avoidance entrance scenes, vehicle avoidance scenes and intersection identification scenes are extracted from the sample videos, which can comprehensively and accurately determine the sample scenes of intelligent driving.
[0114] In order to train the general video generation model into a video generation model capable of generating intelligent driving operation videos, the general video generation model is trained using sample driving information. The specific process of extracting sample operations from the sample driving information is as follows:
[0115] In some embodiments, the sample driving information includes: sample scenes and sample operations; the sample scenes include: on-ramp and off-ramp scenes; the extraction process of the sample operations includes:
[0116] Step 2011A, in response to determining that the sample scene is the on-ramp and off-ramp scene, extracting the on-ramp and off-ramp operations corresponding to the on-ramp and off-ramp scene from the sample video, and determining that the on-ramp and off-ramp operations are: reducing the vehicle speed to a preset speed, turning on the turn signal and giving a first prompt.
[0117] In specific implementation, when the sample scenario is an on-ramp or off-ramp scenario, the on-ramp or off-ramp operations are as follows: reducing the vehicle speed to a preset speed, turning on the turn signal, and giving the first prompt.
[0118] The preset speed is the pre-set speed for on-ramp and off-ramp use. The vehicle enters and exits the ramp when the turn signal is activated. The first prompt is the on-ramp and off-ramp prompt, which can be provided by voice announcement, screen display, and interior lighting.
[0119] For example, if the preset speed is 20 km / h, the first prompt is a voice announcement. The specific on-ramp and off-ramp operation involves reducing the vehicle speed to 20 km / h, turning on the turn signal, and controlling the vehicle to enter or exit the ramp. The first voice announcement is then given to remind the driver that the on-ramp and off-ramp operation is in progress.
[0120] Through the above scheme, by extracting the on-ramp and off-ramp operations corresponding to the on-ramp and off-ramp scenes, the general video generation model is trained using the on-ramp and off-ramp scenes and the corresponding on-ramp and off-ramp operations, so that the obtained initial video generation model can generate videos containing on-ramp and off-ramp scenes and on-ramp and off-ramp operations.
[0121] In order to train the general video generation model into a video generation model capable of generating intelligent driving operation videos, the general video generation model is trained using sample driving information. The specific process of extracting sample operations from the sample driving information is as follows:
[0122] In some embodiments, the sample driving information includes: a sample scene and a sample operation; the sample scene includes: a lane changing protection scene; and the extraction process of the sample operation includes:
[0123] In step 2011B, in response to determining that the sample scene is the lane changing protection scene, the lane changing protection operation corresponding to the lane changing protection scene is extracted from the sample video, and the lane changing protection operation is determined to be: turning on the turn signal before lane changing and performing a second prompt, determining whether to stop lane changing by monitoring the state of the rear vehicle, and turning off the turn signal and controlling the steering wheel to return to a preset angle after determining to stop lane changing.
[0124] In particular implementation, when the sample scene is the lane changing protection scene, the lane changing protection operation specifically includes: turning on the turn signal before lane changing and performing a second prompt, determining whether to stop lane changing by monitoring the state of the rear vehicle, and turning off the turn signal and controlling the steering wheel to return to a preset angle after determining to stop lane changing.
[0125] The second prompt is a lane changing protection prompt, and the prompting manner of the second prompt includes: voice broadcast prompt, screen display prompt, and in-vehicle light prompt. The preset angle is a preset steering wheel angle for normal driving of the vehicle after stopping lane changing.
[0126] For example, the prompting manner of the second prompt is voice broadcast prompt. The lane changing protection operation specifically includes: turning on the turn signal before lane changing and performing a second voice broadcast prompt to remind the driver that the lane changing protection operation is being performed. The state of the rear vehicle is monitored in real time before lane changing, and when the state of the rear vehicle is that a rear vehicle is rapidly approaching, the lane changing is stopped and the turn signal is turned off, and the steering wheel is controlled to return to a preset angle for normal driving.
[0127] Through the above scheme, the lane changing protection operation corresponding to the lane changing protection scene is extracted, so that the general video generation model is trained using the lane changing protection scene and the corresponding lane changing protection operation, so that the obtained initial video generation model can generate a video containing the lane changing protection scene and the lane changing protection operation.
[0128] In order to train the general video generation model into a video generation model capable of generating intelligent driving operation videos, the general video generation model is trained using sample driving information. The extraction process of the sample operation in the sample driving information is as follows:
[0129] In some embodiments, the sample driving information includes: a sample scene and a sample operation; the sample scene includes: a distraction detection scene; and the extraction process of the sample operation includes:
[0130] Step 2011C, in response to determining that the sample scene is the distraction detection scene, extracting the distraction detection operation corresponding to the distraction detection scene from the sample video, and determining that the distraction detection operation is: after monitoring that the driver is in a distracted state, performing a third prompt; after monitoring that the driver is continuously distracted for a preset time period, controlling the vehicle to stop by reducing the vehicle speed and turning on the hazard lights.
[0131] In specific implementation, when the sample scenario is a distraction detection scenario, the distraction detection operation is specifically as follows: after detecting that the driver is in a distracted state, a third prompt is given; after detecting that the driver is continuously distracted for a preset period of time, the vehicle is controlled to stop by reducing the speed and turning on the hazard lights.
[0132] Among them, the third prompt is the lane change protection prompt, and the prompt methods of the third prompt include: voice broadcast prompt, screen display prompt and in-car light prompt.
[0133] For example, the third prompt is a voice announcement and interior lighting. The distraction detection operation is specifically as follows: real-time monitoring of the driver's state. When the driver is distracted (distracted, fatigued, or the current operating environment does not support the autonomous driving system to continue to control the vehicle), a third voice announcement and a third interior lighting prompt are performed. If the driver remains distracted for a preset period of time, reduces speed, and stops the vehicle at the roadside, the hazard lights are activated to alert vehicles behind.
[0134] Through the above scheme, by extracting the distraction detection operations corresponding to the distraction detection scenes, the general video generation model is trained using the distraction detection scenes and the corresponding distraction detection operations, so that the obtained initial video generation model can generate videos containing distraction detection scenes and distraction detection operations.
[0135] In some embodiments, the process of extracting the sample exit condition includes:
[0136] Step 2011a, extracting the electronic parking brake system status, vehicle navigation status, driver's seat belt status, door status, automatic emergency braking system status, vehicle electronic stability system status and driver's hands on the steering wheel status from the sample video.
[0137] Step 2011b, based on the electronic parking brake system status, the main driver's seat belt status, the door status, the automatic emergency braking system status, the vehicle body electronic stability system status and the driver's hands on the steering wheel status, determine the first sample exit condition of intelligent driving.
[0138] Step 2011c: Determine a second sample exit condition for intelligent driving based on the vehicle navigation status.
[0139] In specific implementation, the first sample exit condition of intelligent driving includes at least one of the following: the main driver's seat belt status is not fastened, the door status is open, the EPB status is activated, the AEB status and ESP status are off, and the driver's hands on the steering wheel are not in the state of holding.
[0140] When the vehicle meets at least one of the first sample exit conditions, the vehicle is controlled to exit the automatic driving mode, that is, the NOA function is completely exited, the adaptive cruise control system ACC and the highway driving assistance system HWA are also exited, and the driver is prompted through the human-machine interface HMI (for example, the on-board screen).
[0141] In addition, when the driver pulls out the cruise control handle once or presses the brake pedal, the NOA function can be exited and the driver can be prompted through the human-machine interface HMI (for example, the vehicle screen).
[0142] The second sample exit condition of intelligent driving includes: the vehicle navigation status is the navigation exit status.
[0143] When the vehicle meets the second sample exit condition, the vehicle is controlled to exit the automatic driving mode and switch to the highway assisted driving mode, that is, the NOA function exits and switches to the highway assisted driving system HWA.
[0144] This solution comprehensively and accurately determines the first sample exit condition for intelligent driving based on the status of the electronic parking brake system, the driver's seat belt, the door status, the automatic emergency braking system, the electronic stability control system, and the driver's hands-on steering wheel. When the first sample exit condition is met, the NOA function is completely exited. Based on the vehicle navigation status, the second sample exit condition for intelligent driving is accurately determined. When the second sample exit condition is met, the NOA function is exited and the system switches to the Highway Assist Driving System (HWA).
[0145] In some embodiments, step 202 includes:
[0146] Step 2021: Use the sample vehicle model information and the sample intelligent driving scene to train the initial video generation model to obtain a second simulation video of intelligent driving.
[0147] Step 2022: Compare the second simulation video with the pre-stored second test video to obtain a second comparison result.
[0148] Step 2023: Based on the second comparison result, update the parameters of the initial video generation model to obtain a target video generation model.
[0149] In specific implementations, sample vehicle model information and sample intelligent driving scenarios are input into the initial video generation model, which then outputs a second simulated video of intelligent driving. The second simulated video is the intelligent driving operation video generated by the initial video generation model based on the sample vehicle model information and sample intelligent driving scenarios during training.
[0150] The second test video is a standard video used to test the initial video generation model. The second simulated video and the second test video are compared to obtain a second comparison result, and a second loss function is determined based on the second comparison result. The parameters of the initial video generation model are updated based on the second loss function. When the second loss function approaches a second preset value, the updated initial video generation model is used as the target video generation model.
[0151] Through this approach, the initial video generation model is trained using sample vehicle model information and sample intelligent driving scenarios, enabling the resulting target video generation model to produce intelligent driving operation videos for different vehicles. Furthermore, based on a second comparison result between a second simulated video output by the initial video generation model and a pre-stored second test video, the parameters of the initial video generation model can be accurately adjusted and updated, thereby further improving the accuracy of the trained target video generation model.
[0152] Through the above embodiment, the target vehicle model information and target driving scenario of the target vehicle are obtained. Using a pre-trained target video generation model, a target driving operation video of the target vehicle in the target driving scenario is generated based on the target vehicle model information and the target driving scenario. In this way, the target video generation model can be used to generate the target driving operation video of the target vehicle, eliminating the need to film or animate the target vehicle to produce the target driving operation video, thereby reducing the production cycle of the target driving operation video. By inputting the target video generation model with the model information of different vehicles, driving operation videos of different vehicles can be generated, enabling the reuse of the target video generation model, thereby reducing the production cost of the driving operation videos.
[0153] It should be noted that the embodiments of the present disclosure may be further described in the following manner:
[0154] Figure 2 Flowchart of the method for generating intelligent driving operation video based on AIGC according to an embodiment of the present disclosure. Figure 2 As shown, the method for generating an intelligent driving operation video based on AIGC includes:
[0155] 1. AIGC general large model training stage
[0156] The intelligent driving operation video data (i.e., sample video) is input into the AIGC general large model (i.e., general video generation model), so that the general large model can learn to determine the intelligent driving service startup conditions, decompose the intelligent driving scenarios, learn operations in different intelligent driving scenarios, and learn to determine the intelligent driving service exit conditions based on the input video data.
[0157] Through the above learning, the AIGC general large model generates neural networks and algorithms for the intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations, and intelligent driving service exit conditions. The AIGC general large model is trained as the first version of the AIGC intelligent driving video dedicated large model (i.e., the initial video generation model) specifically for producing intelligent driving videos.
[0158] (1) Obtaining intelligent driving video
[0159] The intelligent driving platform feeds a large amount of intelligent driving video data into the AIGC general large model.
[0160] (2) Learning to determine the conditions for starting intelligent driving services
[0161] The AIGC general large model learns and determines the conditions for starting the intelligent driving service based on the input video data. For example, after learning, the AIGC large model determines that the following conditions must be met to start the intelligent driving service:
[0162] ① The gear is in forward gear, EPB is not activated, and the driver sets the starting and ending points through the car navigation;
[0163] ② The driver's seat belt is fastened;
[0164] ③The doors, hood and trunk are closed;
[0165] ④The vehicle's AEB and ESP function switches are turned on;
[0166] ⑤The vehicle speed is between 10-120km / h, etc.
[0167] (3) Decomposition of intelligent driving scenarios
[0168] The AIGC general large model decomposes intelligent driving scenarios based on the input video. For example, the intelligent driving scenarios decomposed by the AIGC general large model include:
[0169] ① Intelligent on / off ramp assistance: Set routes based on high-precision map navigation to achieve intelligent on / off ramps;
[0170] ② Intelligent Lane Change Protection: Real-time monitoring of the vehicle behind and timely termination of lane change when a fast-moving vehicle approaches from behind, reducing the risk of rear-end collisions.
[0171] ③ Distraction / fatigue detection: Real-time monitoring of the driver's status. When the driver is distracted, fatigued, or the current operating environment does not support the system to continue to control the vehicle, different levels of prompts and alarms will be issued;
[0172] ④ Intelligent merging entrance avoidance: Based on navigation information, the system accurately identifies the merging lane ahead and changes lanes in advance to avoid the merging entrance;
[0173] ⑤ Intelligent Truck Avoidance: Real-time monitoring of surrounding conditions allows for intelligent avoidance of large trucks traveling in adjacent lanes, ensuring safety in multiple dimensions.
[0174] ⑥ Intelligent identification of easily confused intersections: Real-time monitoring of surrounding conditions allows the driver to change lanes in advance when encountering easily confused intersections, ensuring efficient and safe driving.
[0175] (4) Learn different intelligent driving scenarios
[0176] The AIGC general large model learns operations in different intelligent driving scenarios based on input videos, such as:
[0177] ① Intelligent on / off ramp assistance: Before entering / exiting the ramp, the vehicle speed is reduced to 20km / h, the turn signal is turned on, and a voice announcement is made to remind the driver that the intelligent on / off ramp is in progress;
[0178] ② Intelligent Lane Change Protection: Activates the turn signal and provides a voice notification before changing lanes. It also monitors the status of vehicles behind it in real time before changing lanes. If a fast-moving vehicle approaches from behind, the driver immediately terminates the lane change, turns off the turn signal, and returns the steering wheel to the center position to continue driving.
[0179] ③ Distraction / fatigue detection: Real-time monitoring of the driver's status. When the driver is distracted, fatigued, or the current operating environment does not support the system to continue to control the vehicle, the driver can be reminded through the vehicle's interior lights or voice. If necessary, the vehicle will turn on the turn signal, reduce the speed and park by the roadside, and turn on the hazard lights to alert the vehicles behind.
[0180] (5) Learning to determine the exit conditions of intelligent driving services
[0181] The AIGC general large model learns from the input video and determines the exit conditions for the intelligent driving service. For example, after learning, the AIGC large model determines that the intelligent driving service will be exited if the following conditions are met:
[0182] ① If the driver's seat belt is not fastened, the door is open, EPB is activated, and AEB and ESP functions are not turned on, the NOA function will be completely exited without returning to Adaptive Cruise Control (ACC) and Highway Driving Assist (HWA), and the user will be prompted through the HMI;
[0183] ② If the NOA function is activated and the customer exits the in-car navigation system, provided the conditions for starting the intelligent driving service are met, NOA will be terminated and switched to Highway Driving Assist (HWA). The customer can also terminate NOA by pulling the cruise control handle outward once or pressing the brake pedal, and the user will be notified via the HMI.
[0184] ③ When the driver does not hold the steering wheel, the NOA function may also exit and the driver will be reminded.
[0185] (6) The first version of AIGC intelligent driving video dedicated large model
[0186] Through the above learning, the AIGC general large model generates neural networks and algorithms for intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations, and intelligent driving service exit conditions. The AIGC general large model is trained as the first version of the AIGC intelligent driving video dedicated large model specifically for producing intelligent driving videos.
[0187] 2. AIGC generates a large model for intelligent driving video reasoning experimental stage
[0188] The specific vehicle model design (including body, interior and exterior trim, instruments, host, etc.) and the desired intelligent driving scenario are input into the initial version of the AIGC intelligent driving operation video dedicated large model. The initial version of the dedicated large model uses the neural network and algorithm of the intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations and intelligent driving service exit conditions in the model to generate an intelligent driving operation video of an intelligent driving vehicle model with a specific shape and specific driving scenario.
[0189] Then, the parameters of the initial version of the dedicated large model are corrected according to the video problems, the neural network and learning algorithm of the initial version of the dedicated large model are optimized, and the dedicated large model is tuned to finally form the final version of the AIGC intelligent driving operation video dedicated large model (i.e., the target video generation model).
[0190] (1) The first version of the AIGC intelligent driving video dedicated large model obtains vehicle model images and intelligent driving scenarios
[0191] The specific vehicle model design (including body, interior and exterior trim, instrumentation, host, and other information) and the desired intelligent driving scenario are input into the initial version of the AIGC intelligent driving operation video dedicated large model for inference experiments.
[0192] (2) Generate a video explaining how to operate smart driving
[0193] The initial version of the AIGC intelligent driving video dedicated large model generates an intelligent driving operation instruction video based on the input styling diagram and intelligent driving scenario, combined with the neural network and algorithm of the intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations and intelligent driving service exit conditions in the model.
[0194] (3) Optimize the parameters of the initial version of the large model
[0195] In response to the problems with the generated videos, the parameters of the initial version of the dedicated large model were corrected, the neural network and learning algorithm of the initial version of the dedicated large model were optimized, and the dedicated large model was tuned.
[0196] (4) AIGC generates a large model specifically for intelligent driving videos
[0197] After repeated reasoning experiments, the large-model neural network and learning algorithm were deeply optimized, and finally the final version of AIGC was formed to generate a large-scale model dedicated to intelligent driving operation videos.
[0198] 3. Intelligent Driving Operation Video Generation Stage
[0199] The final version of the AIGC intelligent driving video dedicated large model generates an intelligent driving operation instruction video based on the input styling images and intelligent driving scenarios, combined with the neural network and algorithms of the intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations and intelligent driving service exit conditions in the dedicated model, to guide vehicle users to understand how to operate intelligent driving related services.
[0200] (1) AIGC intelligent driving video dedicated large model to obtain vehicle model images and intelligent driving scenes
[0201] The specific vehicle model design (i.e. target vehicle model information, including body, interior and exterior trim, instrumentation, host, etc.) and the desired intelligent driving scenario (i.e. target intelligent driving scenario) are input into the final AIGC intelligent driving operation video dedicated large model.
[0202] (2) Generate a video explaining how to operate smart driving
[0203] The final version of the AIGC intelligent driving video dedicated large model generates an intelligent driving operation instruction video based on the input styling images and intelligent driving scenarios, combined with the neural network and algorithms of the intelligent driving service startup conditions, intelligent driving scenarios, different intelligent driving scenario operations and intelligent driving service exit conditions in the dedicated model, to guide vehicle users to understand how to operate intelligent driving related services.
[0204] Through the above embodiment, simply inputting the target vehicle model information and the target intelligent driving scenario, the target video generation model can be used to generate intelligent driving operation videos to guide vehicle users in operating intelligent driving-related services. This has the advantages of a short production cycle, low capital expenditure, and the ability to reuse technology. It can be used to mass-produce intelligent driving operation videos with different models and different intelligent driving scenarios, meeting the growing demand for intelligent driving videos.
[0205] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.
[0206] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0207] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure further provides a video generating device.
[0208] refer to Figure 3 , the video generating device comprises:
[0209] An acquisition module 301 is configured to acquire target vehicle model information and a target driving scenario of a target vehicle;
[0210] The generation module 302 is configured to use a pre-trained target video generation model to generate a target driving operation video of the target vehicle in the target driving scenario according to the target vehicle model information and the target driving scenario.
[0211] In some embodiments, the target vehicle model information includes: target appearance information and target interior information;
[0212] The generating module 302 includes:
[0213] an initial video generating unit, configured to generate an initial driving operation video under the target driving scenario based on the target driving scenario using the target video generating model;
[0214] The target video generation unit is configured to use the target video generation model to add the target appearance information and the target interior information to the initial driving operation video to obtain a target driving operation video of the target vehicle in the target driving scene.
[0215] In some embodiments, the target driving scenario includes at least one of: an on-ramp or off-ramp scenario, a lane change protection scenario, and a distraction detection scenario;
[0216] The initial video generating unit includes:
[0217] a judgment processing subunit, configured to perform judgment processing on a target driving scene of the target vehicle;
[0218] A first initial video sub-unit is configured to, in response to determining that the target driving scene is an on-ramp or off-ramp scene, determine an on-ramp or off-ramp video generation sub-model from the target video generation model, and generate an initial driving operation video of the on-ramp or off-ramp using the on-ramp or off-ramp video generation sub-model;
[0219] a second initial video sub-unit, configured to, in response to determining that the target driving scene is a lane change protection scene, determine a lane change protection video generation sub-model from the target video generation model, and generate an initial driving operation video of lane change protection using the lane change protection video generation sub-model;
[0220] The third initial video sub-unit is configured to, in response to determining that the target driving scene is a distraction detection scene, determine a distraction detection video generation sub-model from the target video generation model, and use the distraction detection video generation sub-model to generate an initial driving operation video for distraction detection.
[0221] In some embodiments, the apparatus further comprises:
[0222] a first training module configured to obtain a sample video of a driving operation, extract sample driving information of the driving operation from the sample video, and train a general video generation model using the sample driving information to obtain an initial video generation model;
[0223] The second training module is configured to obtain sample vehicle model information and sample intelligent driving scenes, and use the sample vehicle model information and the sample intelligent driving scenes to train the initial video generation model to obtain a target video generation model.
[0224] In some embodiments, the sample driving information includes: sample scenes and sample operations; the sample scenes include: on-ramp and off-ramp scenes;
[0225] The first training module includes:
[0226] The on-ramp and off-ramp operation extraction unit is configured to, in response to determining that the sample scene is the on-ramp and off-ramp scene, extract the on-ramp and off-ramp operation corresponding to the on-ramp and off-ramp scene from the sample video, and determine that the on-ramp and off-ramp operation is: reducing the vehicle speed to a preset speed, turning on the turn signal and giving a first prompt.
[0227] In some embodiments, the sample driving information includes: sample scenarios and sample operations; the sample scenarios include: lane change protection scenarios;
[0228] The first training module includes:
[0229] The lane change protection operation extraction unit is configured to, in response to determining that the sample scene is the lane change protection scene, extract the lane change protection operation corresponding to the lane change protection scene from the sample video, and determine that the lane change protection operation is: turning on the turn signal and giving a second prompt before changing lanes, judging whether to stop changing lanes by monitoring the status of the vehicle behind, and turning off the turn signal and controlling the steering wheel to return to a preset angle after determining to stop changing lanes.
[0230] In some embodiments, the sample driving information includes: sample scenes and sample operations; the sample scenes include: distraction detection scenes;
[0231] The first training module includes:
[0232] The distraction detection operation extraction unit is configured to, in response to determining that the sample scene is the distraction detection scene, extract a distraction detection operation corresponding to the distraction detection scene from the sample video, and determine that the distraction detection operation is: after monitoring that the driver is in a distracted state, provide a third prompt; after monitoring that the driver is continuously distracted for a preset time period, control the vehicle to stop by reducing the vehicle speed and turning on the hazard lights.
[0233] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0234] The apparatus of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0235] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the video generation method described in any of the above embodiments is implemented.
[0236] Figure 410 is a schematic diagram showing a more specific hardware structure of an electronic device provided in this embodiment. The device may include: a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, the memory 1020, the input / output interface 1030, and the communication interface 1040 are communicatively connected to each other within the device via the bus 1050.
[0237] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0238] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 1020 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0239] The input / output interface 1030 is used to connect input / output modules to implement information input and output. The input / output modules can be configured as components within the device (not shown in the figure) or can be externally connected to the device to provide corresponding functions. Input devices may include a keyboard, mouse, touch screen, microphone, various sensors, etc., and output devices may include a display, speaker, vibrator, indicator light, etc.
[0240] The communication interface 1040 is used to connect to a communication module (not shown) to enable communication between the device and other devices. The communication module can communicate via a wired method (e.g., USB (Universal Serial Bus), network cable, etc.) or a wireless method (e.g., mobile network, WIFI (Wireless Fidelity), Bluetooth, etc.).
[0241] The bus 1050 comprises a path for transmitting information between the various components of the device (eg, the processor 1010 , the memory 1020 , the input / output interface 1030 , and the communication interface 1040 ).
[0242] It should be noted that although the above device only shows the processor 1010, the memory 1020, the input / output interface 1030, the communication interface 1040, and the bus 1050, in a specific implementation, the device may also include other components necessary for normal operation. In addition, it will be understood by those skilled in the art that the above device may only include the components necessary to implement the embodiments of this specification, and does not necessarily include all the components shown in the figure.
[0243] The electronic device of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0244] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable the computer to execute the video generation method described in any of the above embodiments.
[0245] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0246] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the video generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0247] Based on the same inventive concept, corresponding to any of the above-mentioned embodiments and methods, the present application also provides a vehicle, including the video generating device, or electronic device, or storage medium in the above-mentioned embodiments, and the vehicle equipment implements the video generating method described in any of the above embodiments.
[0248] The vehicle of the above embodiment is used to implement the video generation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0249] Based on the same inventive concept, the present application also provides a computer program product corresponding to the method of any of the above embodiments, comprising computer program instructions, which, when executed on a computer, cause the computer to perform the video generation method of any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which are not described here again.
[0250] It can be understood that, before using the technical solutions of various embodiments in the present disclosure, the type, use range, use scenario, etc. of the personal information involved will be informed to the user in a proper manner, and the authorization of the user will be obtained.
[0251] For example, in response to receiving the active request of the user, prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require the acquisition and use of personal information of the user. Thus, the user can voluntarily choose whether to provide personal information to the software or hardware such as electronic devices, application programs, servers or storage media that perform the technical solutions of the present disclosure according to the prompt information.
[0252] As an optional but not limited implementation manner, in response to accepting the active request of the user, the manner of sending prompt information to the user may, for example, be a pop-up window manner, and the prompt information may be presented in the form of text in the pop-up window. In addition, the pop-up window may also carry selection controls for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0253] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation of the present disclosure, and other ways that meet the relevant laws and regulations can also be applied to the implementation of the present disclosure.
[0254] Those skilled in the art will understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure is limited to these examples; under the idea of the present disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes of different aspects of the embodiments of the present disclosure as described above. In order to be brief, they are not provided in details.
[0255] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0256] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0257] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the present disclosure. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A video generation method, characterized in that: The method comprises: Obtain target vehicle model information and target driving scenario; A pre-trained target video generation model is used to generate a target driving operation video of the target vehicle in the target driving scenario according to the target vehicle model information and the target driving scenario.
2. The method according to claim 1, characterized in that The target vehicle model information includes: target appearance information and target interior information; The method of generating a target driving operation video of the target vehicle in the target driving scenario using a pre-trained target video generation model according to the target vehicle model information and the target driving scenario includes: generating an initial driving operation video under the target driving scenario using the target video generation model according to the target driving scenario; The target video generation model is used to add the target appearance information and the target interior information to the initial driving operation video to obtain a target driving operation video of the target vehicle in the target driving scene.
3. The method according to claim 2, characterized in that The target driving scenario includes at least one of: an on-ramp or off-ramp scenario, a lane change protection scenario, and a distraction detection scenario; The generating the target video model based on the target driving scene to generate an initial driving operation video under the target driving scene includes: Determining and processing a target driving scene of the target vehicle; In response to determining that the target driving scene is an on-ramp or off-ramp scene, determining an on-ramp or off-ramp video generation sub-model from the target video generation model, and generating an initial driving operation video of the on-ramp or off-ramp using the on-ramp or off-ramp video generation sub-model; In response to determining that the target driving scenario is a lane change protection scenario, determining a lane change protection video generation sub-model from the target video generation model, and generating an initial driving operation video for lane change protection using the lane change protection video generation sub-model; In response to determining that the target driving scene is a distraction detection scene, a distraction detection video generation sub-model is determined from the target video generation model, and an initial driving operation video for distraction detection is generated using the distraction detection video generation sub-model.
4. The method according to claim 1, wherein The training process of the target video generation model includes: Obtaining a sample video of a driving operation, extracting sample driving information of the driving operation from the sample video, and using the sample driving information to train a general video generation model to obtain an initial video generation model; Sample vehicle type information and sample intelligent driving scenes are obtained, and the initial video generation model is trained using the sample vehicle type information and the sample intelligent driving scenes to obtain a target video generation model.
5. The method according to claim 4, characterized in that The sample driving information includes: sample scenes and sample operations; the sample scenes include: on-ramp and off-ramp scenes; The extraction process of the sample operation includes: In response to determining that the sample scene is the on-ramp and off-ramp scene, an on-ramp and off-ramp operation corresponding to the on-ramp and off-ramp scene is extracted from the sample video, and the on-ramp and off-ramp operation is determined to be: reducing the vehicle speed to a preset speed, turning on the turn signal and giving a first prompt.
6. The method according to claim 4, characterized in that The sample driving information includes: sample scenes and sample operations; the sample scenes include: lane change protection scenes; The extraction process of the sample operation includes: In response to determining that the sample scene is the lane change protection scene, the lane change protection operation corresponding to the lane change protection scene is extracted from the sample video, and the lane change protection operation is determined to be: turning on the turn signal and giving a second prompt before changing lanes, judging whether to stop changing lanes by monitoring the status of the vehicle behind, and turning off the turn signal and controlling the steering wheel to return to a preset angle after determining to stop changing lanes.
7. The method according to claim 4, characterized in that The sample driving information includes: sample scenes and sample operations; the sample scenes include: distraction detection scenes; The extraction process of the sample operation includes: In response to determining that the sample scene is the distraction detection scene, a distraction detection operation corresponding to the distraction detection scene is extracted from the sample video, and the distraction detection operation is determined to be: after monitoring that the driver is in a distracted state, a third prompt is given; after monitoring that the driver is continuously distracted for a preset time period, the vehicle is controlled to stop by reducing the vehicle speed and turning on the hazard lights.
8. A video generating device, characterized in that: include: an acquisition module configured to acquire target vehicle model information and a target driving scenario of a target vehicle; The generation module is configured to use a pre-trained target video generation model to generate a target driving operation video of the target vehicle in the target driving scenario according to the target vehicle model information and the target driving scenario.
9. An electronic device, characterized in that: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.
10. A vehicle, characterized in that: Includes the video generating device according to claim 8 or the electronic device according to claim 9.
Citation Information
Patent Citations
Driving training video generation method and system based on virtual reality
CN111768673A
Automatic driving scene generation method and device, vehicle, equipment and storage medium
CN118569073A