Self-driving parade vehicle short video automatic generation method and system and vehicle thereof

By combining intelligent video devices and voice equipment with convolutional neural network algorithms, the automated shooting and editing of short videos of self-driving tours has been achieved, solving the problems of safety hazards and high production difficulty in existing technologies, and improving the convenience and efficiency of travel sharing.

CN116132611BActive Publication Date: 2026-01-02CHINA FAW CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310008736.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-04
Publication Date
2026-01-02
Estimated Expiration
2043-01-04

AI Technical Summary

Technical Problem

Existing technologies pose safety risks when shooting short videos during self-driving tours, and the production process is time-consuming and labor-intensive, making it difficult to meet the sharing needs of car owners.

Method used

It uses intelligent video devices and voice equipment for video recording, image capture, and facial recognition. It uses convolutional neural network algorithms for video processing, automatically adds background music, classifies and sorts short video clips, and provides user-customizable settings.

Benefits of technology

It lowers the barrier to entry for short video production, automates shooting and editing, improves driving safety and sharing efficiency, and makes it easier for car owners to record and share their travel experiences.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116132611B_ABST
    Figure CN116132611B_ABST
Patent Text Reader

Abstract

The application discloses a kind of self-driving tour car short video automatic generation method, system and vehicle thereof, specifically include: through intelligent video device video recording, image shooting and face recognition;Utilize intelligent voice equipment to carry out audio acquisition;Start intelligent recording function, merge video information and audio information, carry out short video shooting.The application has the following advantages compared with prior art: the application reduces the production threshold of self-driving tourism short video, provides a kind of method relying on artificial intelligence big data, automatically shoots beautiful scenery or character, after travel, automatically completes photo classification, video editing, and automatically matches background music, it is convenient for the record and sharing of car owner to tourism experience, better service car owner's life.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a short video automatic generation method, system and vehicle thereof, in particular to a self-driving travel vehicle short video automatic generation method, system and vehicle thereof. BACKGROUND

[0002] Nowadays, self-driving travel is becoming more and more popular, and car owners want to make short videos of beautiful scenery and characters on the road and share them on the Internet. During driving, beautiful scenery and good characters are often encountered, and manual shooting will affect driving safety; the production of short videos will also consume a lot of time of the car owner, and if the car owner manually shoots photos or videos during driving, it will cause driving safety problems; the production of short videos has certain technical threshold, and video editing and music configuration are required, and the production of short videos is also a time-consuming and laborious work.

[0003] In summary, the existing technology of self-driving travel short video shooting method cannot meet the requirements of people and needs to be improved. SUMMARY

[0004] The purpose of the present application is to provide a self-driving travel vehicle short video automatic generation method, system and vehicle thereof, to solve the defects of the prior art.

[0005] The present application provides the following solutions:

[0006] A self-driving travel vehicle short video automatic generation method, specifically comprising:

[0007] Video recording, image shooting and face recognition are performed by an intelligent video device;

[0008] Audio collection is performed by an intelligent voice device;

[0009] The intelligent recording function is started, video information and audio information are merged, and short video shooting is performed.

[0010] Further, the intelligent video device specifically includes a front camera, a driver monitoring system and a passenger monitoring system, and the intelligent video device performs video recording, image shooting and face recognition through a convolutional neural network algorithm.

[0011] Further, when the intelligent video device performs video recording and face recognition, a central control system is notified by triggering a system message, and the central control system drives the intelligent voice device to perform audio collection.

[0012] Further, background music is automatically matched when merging video information and audio information.

[0013] Further, the short video shooting specifically includes:

[0014] Classify short videos by time and place;

[0015] Select short video clips according to scenery and facial expressions, and score and rank short videos;

[0016] Select different background music according to different shooting locations;

[0017] Provide a merging service of background music and short videos, and provide user-defined settings and manual modifications.

[0018] Further, the short videos are classified by time and place, specifically: first-level classification is performed by time, and second-level classification is performed by place based on the first-level classification;

[0019] The short video clips are selected according to scenery and facial expressions, and the short videos are scored and ranked, specifically: the basic shooting clips are scored according to two dimensions of "scenic beauty" and "in-vehicle facial expression": scenic beauty accounts for 70%, and facial expression accounts for 30%.

[0020] A self-driving tour driving short video automatic generation system, specifically comprising:

[0021] A video image face recognition module for video recording, image shooting and face recognition through an intelligent video device;

[0022] An intelligent audio acquisition module for audio acquisition using an intelligent voice device;

[0023] A short video shooting and recording module for starting an intelligent recording function, merging video information and audio information, and shooting short videos.

[0024] An electronic device, comprising: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory stores a computer program, when the computer program is executed by the processor, the processor executes the steps of the method.

[0025] A computer readable storage medium storing a computer program executable by an electronic device, when the computer program runs on the electronic device, the electronic device executes the steps of the method.

[0026] A vehicle, specifically comprising:

[0027] An electronic device for implementing the method;

[0028] A processor, the processor runs a program, when the program runs, the steps of the method are performed on the data output from the electronic device;

[0029] A storage medium for storing a program, which, when executed, performs the steps of the method for data output from an electronic device.

[0030] Compared with the prior art, the present application has the following advantages: the present application reduces the production threshold of self-driving travel short videos, provides a method relying on artificial intelligence big data to automatically shoot beautiful scenery or characters, automatically completes photo classification, video editing after the trip, and automatically matches background music, which is convenient for the record and sharing of travel experience of the vehicle owner and better serves the life of the vehicle owner. BRIEF DESCRIPTION OF DRAWINGS

[0031] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the drawings needed to be used in the specific embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor on the basis of these drawings.

[0032] Figure 1 is a flowchart of the self-driving travel short video automatic generation method.

[0033] Figure 2 is an architecture diagram of the self-driving travel short video automatic generation system.

[0034] Figure 3 is a schematic diagram of the convolutional neural network.

[0035] Figure 4 is a flowchart of the intelligent video device calling the built-in algorithm for image recognition and face recognition.

[0036] Figure 5 is a flowchart of the self-driving travel short video automatic generation method in a specific application scenario.

[0037] Figure 6 is a schematic diagram of the video device for intelligent recording.

[0038] Figure 7 is a flowchart of the short video production.

[0039] Figure 8 is a structural schematic diagram of an electronic device. DETAILED DESCRIPTION

[0040] The technical solutions of the present application will be described clearly and completely in connection with the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0041] As shown in the flow of the short video automatic generation method of the self-driving tour vehicle, specifically comprising: Figure 1

[0042] Step S1, video recording, image shooting, beautiful scenery shooting and face recognition are performed by the intelligent video device;

[0043] Specifically, the intelligent video device includes a front camera, a driver monitoring system (DMS) and a passenger monitoring system (OMS). The intelligent video device performs video recording, image shooting, beautiful scenery shooting and face recognition through a convolutional neural network algorithm.

[0044] Specifically, when the intelligent video device performs video recording and face recognition, the system message is triggered to notify the central control system, and the central control system drives the intelligent voice device to perform audio acquisition. For example, when the sound in the vehicle is obviously human voice, it is determined as in-vehicle voice, and the intelligent voice device is specifically a vehicle-mounted microphone.

[0045] Specifically, background music is automatically matched when merging video information and audio information.

[0046] For example, the identity of the vehicle owner or passenger is recognized by the DMS / OMS, and the expression is recognized and automatically photographed. The DMS / OMS automatically captures special expression photos.

[0047] For example, the intelligent video device triggers a message to notify the central control system to call the GPS module to obtain the current geographic location information and the current time information.

[0048] Step S2, audio acquisition is performed by the intelligent voice device; for example, the vehicle owner or passenger sees interesting scenery, and triggers the photo by the voice of the vehicle owner or passenger.

[0049] Step S3, the (vehicle-mounted or independent) intelligent recording function is started, the video information and audio information obtained by the intelligent video device and the intelligent voice device are merged, and short video shooting is performed.

[0050] For example, the video information and audio information obtained by the intelligent video device and the intelligent voice device are merged to perform short video shooting.

[0051] Specifically, the short videos are classified by time and place;

[0052] ​Short video clips are selected based on scenery and facial expressions, and then scored and ranked.

[0053] Different background music is selected depending on the shooting location;

[0054] It offers a service to merge background music and short videos, and allows users to customize settings and make manual modifications.

[0055] For example, short videos can be categorized by time (days) and location. Specifically, a primary category is defined by time (days), and a secondary category is defined by location, with the location in the secondary category specifying a city or a specific tourist attraction.

[0056] The process involves selecting short video clips based on scenery and facial expressions, and then scoring and ranking the short videos.

[0057] The basic shooting clips were scored based on two dimensions: "scenery beauty" and "expressions of people inside the car". Scenery beauty accounted for 70%.

[0058] Human facial expressions account for 30%; for example, human facial expressions include: laughing, surprised, making faces, etc., and scores are given for each expression. The scores of video clips from a single location within a day are sorted in descending order, and a maximum of 3 video clips with the highest scores are selected.

[0059] For the purpose of simplicity, the method steps disclosed in the above embodiments are described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0060] Any flowchart or other description of a process or method can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing a particular logical function or process. Furthermore, the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed and implemented not in the order shown or discussed, including substantially simultaneously or in reverse order according to the functions involved, or by executing computer instructions and implementing corresponding functions according to program structures such as loops, branches, etc., as will naturally be understood by those skilled in the art when practicing embodiments of the invention.

[0061] like Figure 2 The self-driving tour short video automatic generation system shown includes:

[0062] The video image face recognition module is used for video recording, image shooting and face recognition through the intelligent video device.

[0063] The intelligent audio acquisition module is used for audio acquisition through the intelligent voice device.

[0064] The short video shooting recording module is used for starting the intelligent recording function, merging video information and audio information, and shooting a short video.

[0065] It is worth noting that, although only some basic function modules are disclosed in the embodiments of the present application, it does not mean that the composition of the system is limited to the above basic function modules, on the contrary, the meaning expressed by the embodiments is that on the basis of the above basic function modules, one or more function modules can be added by those skilled in the art in combination with the prior art to form infinite embodiments or technical solutions, that is, the system is open rather than closed, and the protection scope of the present application claim cannot be limited to the disclosed basic function modules because only individual basic function modules are disclosed in the embodiments. At the same time, in order to facilitate description, the above device is described as various units and modules. Of course, when the present application is implemented, the functions of the units and modules can be implemented in the same software and / or hardware.

[0066] The embodiments of the system described above are only illustrative, for example: the various function modules, units or subsystems in the system can or can not be physically separated, or can or can not be physical units, that is, they can be located in the same place or distributed to multiple different systems and their subsystems or modules. Those skilled in the art can select part or all of the function modules, units or subsystems to achieve the purpose of the embodiments of the present application according to actual needs, and those skilled in the art can understand and implement without creative labor.

[0067] As shown in the schematic diagram of the convolutional neural network: Figure 3 When an image is input, the algorithm of the convolutional neural network is used to identify the image, and the image is processed through: convolution layer->pooling layer->convolution layer->pooling layer->full connection layer. After layer-by-layer processing, when the final result meets the requirements, the output is the result of the beautiful scenery or facial expression, and the result is scored;

[0068] Input layer: the input layer is the input of the entire neural network, in the convolutional neural network for processing images, it generally represents a pixel matrix of an image, such as 28X28X1, 32X32X3

[0069] As the name suggests, a convolutional layer is the most important part of a convolutional neural network. Unlike fully connected layers, the input to each node in a convolutional layer is only a small piece of the input from the previous layer. Common sizes are 3x3 or 5x5, but the depth increases. Convolutional layers perform a more in-depth analysis of each small piece in the neural network to obtain more abstract features.

[0070] Pooling layers reduce features, meaning they can reduce the size of the matrix without changing its depth. The pooling operation is similar to converting a known high-resolution image into a low-resolution one. By using pooling layers, the number of nodes in the final fully connected layer can be further reduced, thereby reducing the overall number of parameters in the neural network.

[0071] Fully connected layers: We can view convolutional and pooling layers as a process of automatic image feature extraction. After feature extraction and flattening, fully connected layers are still needed to complete the classification task.

[0072] Softmax layer: Similar to that in a fully connected neural network, the Softmax layer can be used to determine the probability of the current sample belonging to different classes.

[0073] like Figure 4 The flowchart shown illustrates how the intelligent video device uses its built-in algorithm for image and face recognition. After the camera exposes a photo, it calls the built-in algorithm to identify whether the current photo is a "scenery photo" or a photo with a matching "facial expression". If it matches, the photo and its score are sent to the main controller, which scores the photo and stores it in the storage device for later use.

[0074] like Figure 5 The flowchart shown illustrates the method for automatically generating short videos of self-driving tours in a specific application scenario, and includes the following:

[0075] The system collects video and audio signals through the in-vehicle microphone (MIC), front-facing camera (with a smart algorithm for landscape recognition), and DMS / OMS (a smart algorithm for smile capture). The collected signals are then sent to the central control system. Simultaneously, the GPS positioning device is triggered to obtain the current geographical location information. By recognizing the voice commands of the driver (owner) or passengers, the system triggers the automatic shooting function to capture short videos. Multiple elements from the real world (such as scenery, people inside the car, in-vehicle voices, and captured video) are generated into multiple basic shooting clips. These clips, along with background music, are then ranked to automatically generate a short video. Users can then make appropriate manual modifications before the final short video is uploaded to the user's mobile phone.

[0076] likeFigure 6 The video device shown in the schematic diagram of intelligent recording, the method steps are as follows:

[0077] 1. Video signal, face recognition collection function:

[0078] 1.1) Front-view camera:

[0079] Automatic triggering: start the beautiful scenery recording mode, and the system automatically saves the photo when it is recognized that the current image is a beautiful scenery;

[0080] Voice triggering: the owner or passenger sees the interesting beautiful scenery, and the voice triggers the photo shooting;

[0081] 1.2) OMS\DMS:

[0082] Automatic trigger 1: OMS\DMS starts the face recording mode, and automatically takes a photo when a smiling face, surprise or astonishment is recognized;

[0083] Automatic trigger 2: when the front-view camera takes a photo, notify the DMS / OMS to take a photo of the corresponding owner or passenger;

[0084] 1.3) When the front camera takes a photo, the system message is automatically triggered to notify the central control system, and the central control drives the MIC device to collect the in-vehicle sound. If there is obvious human voice in the vehicle, the current sound segment is automatically saved;

[0085] 1.4) When the front camera takes a photo, the system message is triggered to notify the central control system, and the central control system calls the GPS module to obtain the position information at this time, and saves the position and time information at this time;

[0086] 1.5) Start intelligent recording once to form a shooting segment, including (scenery photo\in-vehicle figure photo\sound\time\position information);

[0087] 2. Recording steps:

[0088] 2.1) Owner voice: start the intelligent recording mode;

[0089] 2.2) In the intelligent recording mode: the front-view camera automatically takes a photo, or the driver (passenger) voice triggers the photo shooting;

[0090] 2.3) When the front-view camera takes a photo, notify the DMS / OMS to take a photo of the corresponding owner or passenger;

[0091] 2.4) DMS / OMS will automatically capture the special expression of the owner or passenger to take a photo;

[0092] 2.5) When the camera takes a photo, the GPS module will be called, and the photo will contain the photo position and time information;

[0093] 2.6) When the camera takes a photo, the central control system will also collect MIC sound, and if there is a person's voice, the corresponding sound segment will be saved;

[0094] 2.7) Shooting results: a basic shooting segment is generated each time shooting is performed;

[0095] Basic shooting segment: composed of an outside scenery shot, an inside person shot, an inside voice segment, time information, and location information;

[0096] 2.8) Car owner voice - end intelligent recording; ending recording produces one or more basic shooting segments;

[0097] In the prior art, there are professional audio and video, image merging software, such as ProCut Pro, and those skilled in the art can fully utilize the existing audio and video processing software and image processing tool software to realize the merging of audio, video, and images.

[0098] As shown in the flowchart of short video production, the method steps specifically include: Figure 7

[0099] 3. Short video production method:

[0100] 3.1) Classification:

[0101] Primary classification: classified in basic units of time days;

[0102] Secondary classification: on the basis of primary classification, secondary classification is performed according to location, and the location is specific to a city or a specific tourist attraction;

[0103] 3.2) Selection of shooting segments:

[0104] According to the two dimensions of "beautiful scenery" and "inside person expression", the basic shooting segments are comprehensively scored;

[0105] Beautiful scenery: scored according to scenery, accounting for 70%;

[0106] Person expression: scored according to smile, surprise, and ghost face, accounting for 30%;

[0107] The scoring of the video segment list of a place in a day is sorted in descending order, and at most 3 video segments with the highest scores are selected;

[0108] 3.3) Selection of background music:

[0109] Take location as the basic unit, and configure background music; select music with obvious regional characteristics as background music, such as selecting

[0110] ​The background music reflecting the grassland is selected; the background music reflecting the desert is selected in Xinjiang;

[0111] 3.4) Final short video: the selected basic video segment list and the corresponding background music are combined to form the final travel short video;

[0112] If the basic video segment is not satisfied:

[0113] 3.5) The user can edit the video segment, and the scenery photo, the portrait and the sound segment in the segment can be edited, modified or deleted;

[0114] 3.6) The user can add or delete the video segment through the software;

[0115] 3.7) Automatically uploaded to the mobile phone of the vehicle owner.

[0116] As shown in Figure 8 the application discloses the self-driving travel short video automatic generation method and system, and further discloses the corresponding electronic device, storage medium and vehicle:

[0117] An electronic device comprises a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the self-driving travel short video automatic generation method.

[0118] A computer readable storage medium stores a computer program executable by an electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the self-driving travel short video automatic generation method.

[0119] A vehicle specifically comprises:

[0120] An electronic device is used to implement the self-driving travel short video automatic generation method;

[0121] A processor runs a program, and when the program runs, the steps of the self-driving travel short video automatic generation method are executed for the data output from the electronic device;

[0122] A storage medium is used to store the program, and when the program runs, the steps of the self-driving travel short video automatic generation method are executed for the data output from the electronic device.

[0123] The processor mentioned above can be a general processor, including a central processing unit (CPU), a network processor (NP), etc.; can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component.

[0124] The communication bus mentioned above can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For the convenience of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or only one type of bus.

[0125] The electronic device includes a hardware layer, an operating system layer running on the hardware layer, and an application layer running on the operating system. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and a memory. The operating system can be any one or more computer operating systems that realize the control of the electronic device through a process, such as a Linux operating system, a Unix operating system, an Android operating system, an iOS operating system, or a windows operating system. In the embodiments of the present application, the electronic device can be a handheld device such as a smart phone or a tablet computer, or can be an electronic device such as a desktop computer or a portable computer, and is not particularly limited in the embodiments of the present application.

[0126] The execution subject of the electronic device control in the embodiment of the present application can be an electronic device, or a functional module capable of calling and executing a program in the electronic device. The electronic device can obtain the firmware corresponding to the storage medium, the firmware corresponding to the storage medium is provided by a supplier, and the firmware corresponding to different storage media can be the same or different, which is not limited herein. After obtaining the firmware corresponding to the storage medium, the electronic device can write the firmware corresponding to the storage medium into the storage medium, specifically, burn the firmware corresponding to the storage medium into the storage medium. The process of burning the firmware into the storage medium can be implemented by using the prior art, which is not described in detail in the embodiment of the present application.

[0127] The electronic device can also obtain the reset command corresponding to the storage medium, the reset command corresponding to the storage medium is provided by a supplier, and the reset command corresponding to different storage media can be the same or different, which is not limited herein.

[0128] At this time, the storage medium of the electronic device is the storage medium in which the corresponding firmware is written, and the electronic device can respond to the reset command corresponding to the storage medium in the storage medium in which the corresponding firmware is written, so that the electronic device resets the storage medium in which the corresponding firmware is written according to the reset command corresponding to the storage medium. The process of resetting the storage medium according to the reset command can be implemented by using the prior art, which is not described in detail in the embodiment of the present application.

[0129] Those skilled in the art can understand that, unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which the present application belongs. It should also be understood that terms such as those defined in general dictionaries should be understood as having meanings consistent with those in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless specifically defined.

[0130] It should be noted that some words are used in the specification and claims to refer to specific elements. Those skilled in the art should understand that different manufacturers and producers can use different names to refer to the same element. The specification and claims do not distinguish elements based on the difference in names, but based on the difference in function.

[0131] The technical features of the above embodiments can be combined in any way, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combination of the technical features does not exist contradictory, it should be considered as the scope of the present application.

[0132] Furthermore, those skilled in the art will recognize that references to various embodiments of the application are not intended to be bound by the combination of features in any one embodiment. For example, any one of the embodiments claimed in the claims can be used in any combination with embodiments of the application.

[0133] In the description of the specification, the description of the terms "one embodiment", "an example", "a specific example" and the like is intended to mean that the specific feature, structure, material or characteristic described in connection with the embodiment or example is included in at least one embodiment or example of the application. Illustrative expressions of the above terms in the specification do not necessarily refer to the same embodiment or example. Also, the specific features, structures, materials or characteristics described can be combined in any appropriate manner in one or more embodiments or examples.

[0134] In addition, the technical solutions among various embodiments of the application can be combined with each other, but it must be based on the fact that a person skilled in the art can realize it. When the combination of technical solutions contradicts each other or cannot be realized, it should be considered that the combination of technical solutions does not exist and is not within the protection scope of the application.

[0135] All features disclosed in the specification, or the steps of all methods or processes disclosed in the specification, can be combined in any manner, except where features or steps are mutually exclusive. Any of the features disclosed in the specification can be replaced by alternative features or equivalents having the same purpose, unless specifically stated otherwise. That is, each feature is merely one example of a range of equivalent or similar features unless specifically stated otherwise. Throughout the specification, like reference numerals indicate like elements.

[0136] Those skilled in the art will understand that the modules in the devices in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and furthermore can be divided into multiple sub-modules or sub-units or sub-components. All features disclosed in the specification (including the corresponding claims, abstract and drawings) and all processes or units of any method or device disclosed in this way can be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless explicitly stated otherwise, each feature disclosed in the specification (including the corresponding claims, abstract and drawings) can be replaced by an alternative feature providing the same, equivalent or similar purpose.

[0137] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions recorded in the above embodiments can be modified, or some or all of the technical features can be replaced by equivalents; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for automatically generating short videos of a self-driving parade, characterized in that, Specifically comprising: Video recording, image shooting and face recognition through intelligent video device; Audio collection through intelligent voice device; Starting intelligent recording function, merging video information and audio information, and shooting short video; The intelligent video device specifically comprises: front camera, driver monitoring system and passenger monitoring system, and the intelligent video device performs video recording, image shooting and face recognition through convolutional neural network algorithm. When an image is input, the image is identified through convolutional neural network algorithm, and the image sequentially passes through convolutional layer, pooling layer, convolutional layer, pooling layer and full connection layer, and after layer-by-layer processing, the output is the result of beautiful scenery or facial expression when the final result meets the requirement, and the result is scored. Front camera: When the beautiful scenery recording mode is started and it is identified that the current image is beautiful scenery, the system automatically saves the photo; When the driver or passenger sees interesting beautiful scenery, the photo is triggered by voice; Automatic trigger 1: when the OMS\DMS starts the face recording mode, the photo is automatically taken when a smiling face, surprise or astonishment is identified; Automatic trigger 2: when the front camera takes a photo, the DMS / OMS is notified to take a photo of the corresponding driver or passenger; When the front camera takes a photo, the system message is automatically triggered to notify the central control system, and the central control system drives the MIC device to collect the sound in the vehicle. If there is obvious human voice in the vehicle, the current sound segment is automatically saved. When the front camera takes a photo, the system message is triggered to notify the central control system, and the central control system calls the GPS module to obtain the position information at this time, and saves the position and time information at this time. Starting intelligent recording once forms a shooting segment, including scenery photo, in-vehicle figure photo, sound, time and position information.

2. The method of claim 1, wherein, When the intelligent video device performs video recording and face recognition, the system message is triggered to notify the central control system, and the central control system drives the intelligent voice device to collect audio.

3. The method of claim 1, wherein, When merging video information and audio information, background music is automatically matched.

4. The method of claim 1, wherein, The short video shooting specifically comprises: Classifying short videos according to time and place; Selecting short video segments according to scenery and facial expression, and scoring and sorting short videos; Selecting different background music according to different shooting places; Providing merging service of background music and short video, and providing user-defined setting and manual modification.

5. The method of claim 4, wherein, Classifying short videos according to time and place specifically comprises: classifying according to time as a basic unit, and classifying according to place as a secondary unit on the basis of the primary classification. Selecting short video segments according to scenery and facial expression, and scoring and sorting short videos specifically comprises: scoring the basic shooting segment according to two dimensions of "beautiful scenery" and "in-vehicle facial expression": beautiful scenery accounts for 70%, and facial expression accounts for 30%.

6. An automatic short video generation system for a self-driving parade vehicle, characterized by, The self-driving tour vehicle short video automatic generation system is used to execute the self-driving tour vehicle short video automatic generation method as claimed in any one of claims 1-5. The self-driving tour vehicle short video automatic generation system specifically comprises: The video image face recognition module records video, captures image and recognizes face through the intelligent video device; The intelligent audio acquisition module acquires audio through the intelligent voice device; The short video shooting recording module starts the intelligent recording function, merges video information and audio information, and shoots short video.

7. An electronic device, comprising: Comprise: A processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method in any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, It stores a computer program executable by the electronic device, and when the computer program runs on the electronic device, the electronic device executes the steps of the method in any one of claims 1 to 5.

9. A vehicle characterized by comprising: Specifically comprising: An electronic device for implementing the method in any one of claims 1 to 5; A processor, wherein the processor runs a program, and when the program runs, the steps of the method in any one of claims 1 to 5 are executed for data output from the electronic device; A storage medium for storing a program, and when the program runs, the steps of the method in any one of claims 1 to 5 are executed for data output from the electronic device.

Citation Information

Patent Citations

  • Short video generation method and device, equipment and storage medium

    CN112165585A

  • Video collection generation method and display device

    CN113973216A

  • Vehicle-mounted short video generation method and device

    CN114943964A

  • Travel video generation method and device based on cockpit, vehicle and storage medium

    CN115529423A