Video generation method, system, device and medium based on out-of-vehicle scene
By identifying and analyzing image information during driving, and using scene databases and multiple algorithms to generate target videos, the problem of drivers or passengers finding it difficult to capture and share beautiful scenery has been solved, achieving safe and convenient video generation.
Patent Information
- Application Number
- CN202110852611.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-07-27
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2041-07-27
AI Technical Summary
It is difficult for drivers or passengers to take and share beautiful scenery while the vehicle is in motion, and existing technologies pose safety hazards and are inconvenient to operate.
By acquiring video information captured during driving, the system identifies target scenes, determines the image segments to be processed, and analyzes and processes them to generate videos. It utilizes scene databases and keyness scoring models, combined with eye tracking, posture, and speech recognition algorithms to determine interest values and automatically generate target videos.
It enables safe and convenient shooting and generation of scenic videos without manual operation during driving, improving the accuracy and efficiency of video generation.
Smart Images

Figure CN115695906B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electronic equipment, in particular to a video generation method and system based on a scene outside a vehicle, a device and a medium. BACKGROUND
[0002] During the driving of a vehicle, occasionally, a beautiful scene is encountered which is interesting to the driver, passengers and other people in the vehicle and which they want to take a picture of.
[0003] However, for the driver, it is very dangerous to use a mobile phone, a camera or other camera device while driving. Such operation violates the road traffic regulations and brings great danger to pedestrians and other vehicles. In addition, although the passengers can use a mobile phone, the scenery is fleeting during driving, and the passengers may not have time to take out a mobile phone or a camera before missing it. This makes it difficult for the driver or the passengers to take pictures of the moments worth recording during driving and to share them with family and friends who are not present. SUMMARY
[0004] An object of the present application is to provide a video generation method and system based on a scene outside a vehicle, a device and a medium, so as to improve the efficiency and convenience of video generation.
[0005] An object of the present application is to provide a video generation method based on a scene outside a vehicle, which has the advantage of obtaining image information taken during driving, identifying the image information, determining a target scene, determining a to-be-processed image segment according to the target scene, analyzing and processing the to-be-processed image segment, and thus generating a target video, so as to achieve the purpose of efficiently and conveniently generating a video.
[0006] Another object of the present application is to provide a video generation method based on a scene outside a vehicle, which has the advantage of labeling a scene by means of internet platform resources or a large amount of manual input, determining a keyword label corresponding to the scene, generating a scene database according to the scene and the keyword label, so as to avoid the problem of scene recognition error and improve the accuracy of scene recognition.
[0007] Another object of the present application is to provide a video generation method based on a scene outside a vehicle, which has the advantage of querying the scene database, matching the image information, determining a scene in the image information, performing key degree scoring on the scene according to a preset key degree scoring model, wherein the key degree scoring model includes scoring the scene one by one according to a target recognition algorithm, a relative position of the scene and a duration of the scene, and determining a scene with a key score higher than a first threshold value in the scene as a target scene according to the scoring result of the key degree scoring, so as to accurately identify the scene information and improve the accuracy of video generation.
[0008] Another object of the present application is to provide a video generation method based on an outside scene, which has the advantage that the appearance time of the target scene is queried, the timestamp information corresponding to the target scene is determined, the interest value of the target person for the target scene is obtained through a line-of-sight tracking algorithm according to the timestamp information, and if it is determined that the interest value is greater than a first interest threshold, the image segment corresponding to the target scene is determined as a to-be-processed image segment, so that the image segment of interest to the target person is accurately obtained through the line-of-sight tracking algorithm, so as to achieve the purpose of improving the accuracy of video generation.
[0009] Another object of the present application is to provide a video generation method based on an outside scene, which has the advantage that the appearance time of the target scene is queried, the timestamp information corresponding to the target scene is determined, the interest value of the target person for the target scene is obtained through a line-of-sight tracking algorithm according to the timestamp information, and if it is determined that the interest value is greater than a first interest threshold, the image segment corresponding to the target scene is determined as a to-be-processed image segment, so that the image segment of interest to the target person is accurately obtained through the line-of-sight tracking algorithm, so as to achieve the purpose of improving the accuracy of video generation.
[0010] Another object of the present application is to provide a video generation method based on an outside scene, which has the advantage that the appearance time of the target scene is queried, the timestamp information corresponding to the target scene is determined, the interest value of the target person for the target scene is obtained through a line-of-sight tracking algorithm according to the timestamp information, and if it is determined that the interest value is greater than a first interest threshold, the image segment corresponding to the target scene is determined as a to-be-processed image segment, so that the image segment of interest to the target person is accurately obtained through the line-of-sight tracking algorithm, so as to achieve the purpose of improving the accuracy of video generation.
[0011] To achieve the above object,
[0012] In a first aspect, the embodiments of the present application provide a video generation method based on an outside scene, which comprises the following steps:
[0013] Obtaining image information shot during driving;
[0014] Identifying the image information and determining a target scene;
[0015] Determining a to-be-processed image segment according to the target scene;
[0016] Analyzing and processing the to-be-processed image segment to generate a target video.
[0017] In a second aspect, the embodiments of the present application provide a vehicle exterior scene-based video generation system, the vehicle wake-up system comprising: a processor, a memory, and a communication unit; the processor is in communication connection with the memory and the communication unit;
[0018] a processor in communication with the communication unit, and executing instructions stored in the memory;
[0019] the communication unit acquires image information captured during driving;
[0020] the memory is configured to store instructions which, when executed by the processor, cause the processor to perform steps comprising:
[0021] identifying the image information to determine a target scene, determining a to-be-processed image segment according to the target scene, and performing analysis and processing on the to-be-processed image segment to generate a target video.
[0022] In a third aspect, the embodiments of the present application provide an electronic device, comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs comprise instructions for performing steps in any method of the first aspect of the embodiments of the present application.
[0023] In a fourth aspect, the embodiments of the present application provide a computer readable storage medium, wherein the computer readable storage medium stores a computer program for electronic data exchange, and the computer program causes a computer to perform some or all of the steps described in any method of the second aspect of the embodiments of the present application.
[0024] In a fifth aspect, the embodiments of the present application provide a computer program product, which comprises a non-transitory computer readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps described in any method of the second aspect of the embodiments of the present application. The computer program product can be a software installation package.
[0025] It can be seen that in the embodiment of the present application, a video generation method, system, electronic device and computer storage medium based on an outside scene of a vehicle are provided. The method comprises: first acquiring image information shot in a driving process; then identifying the image information to determine a target scene; then determining a to-be-processed image segment according to the target scene; and then analyzing and processing the to-be-processed image segment to generate a target video. It can be seen that when a target person encounters a scene of interest in a driving process, a target video is generated by analyzing and processing an image segment, which is conducive to automatically shooting and generating a beautiful scene video in a driving process without manual operation of the target person, and is conducive to improving the safety and interest of the generation of the outside beautiful scene video and improving the efficiency and convenience of the generation of the video. BRIEF DESCRIPTION OF DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0027] Figure 1 is a flowchart of a video generation method based on an outside scene of a vehicle provided by an embodiment of the present application;
[0028] Figure 2 is a flowchart of another video generation method based on an outside scene of a vehicle provided by an embodiment of the present application;
[0029] Figure 3 is a flowchart of another video generation method based on an outside scene of a vehicle provided by an embodiment of the present application;
[0030] Figure 4 is a schematic diagram of a video generation system based on an outside scene of a vehicle provided by an embodiment of the present application;
[0031] Figure 5 is a structural schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0033] The terms "first", "second", and the like in the description and in the claims of the present application and above drawings are used for distinguishing between similar objects, not necessarily described in a particular order. Also, the terms "comprise", "comprising", and the like are intended to encompass non-exclusive inclusions. For example, processes, methods, articles, or apparatuses that comprise a list of steps or elements are not limited to the listed steps or elements, but can also comprise additional steps or elements not expressly listed, or can also comprise steps or elements inherent in such processes, methods, articles, or apparatuses.
[0034] Reference herein to "an embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment can be included in at least one embodiment of the application. The appearances of the phrase that an embodiment in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of one another. As will be apparent to those of ordinary skill in the art, embodiments described herein can be combinable with other embodiments.
[0035] The electronic device related to the embodiments of the present application can include various handheld devices, vehicle-mounted devices, wearable devices, computing devices or other processing devices connected to wireless modems, and various forms of user equipment (UE), mobile stations (MS), terminal devices, etc. with wireless communication functions.
[0036] The embodiments of the present application will be described in detail below.
[0037] Please refer to Figure 1 , Figure 1 is a flowchart of a video generation method based on an outside-the-vehicle scene provided by the embodiments of the present application, and the method comprises:
[0038] S101, an electronic device acquires image information shot in a driving process;
[0039] Wherein, when shooting the image in the driving process, the shooting position can be adjusted according to the gestures and gaze points of the target personnel, so as to obtain key images.
[0040] S102, the electronic device identifies the image information and determines a target scene;
[0041] Wherein, the image information comprises a plurality of continuous image information, and the target scene is included in the image information.
[0042] Wherein, the target scene is associated with a keyword tag and stored in a scene database.
[0043] S103, the electronic device determines a to-be-processed image segment according to the target scene;
[0044] The to-be-processed image segment is one or more image segments in the image information.
[0045] S104, the electronic device analyzes and processes the to-be-processed image segment to generate a target video.
[0046] The target video includes picture information, music information, and text information.
[0047] It can be seen that in the embodiments of the present application, a video generation method, system, electronic device, and computer storage medium based on an outside scene are provided. The method includes first acquiring image information captured during driving, then identifying the image information to determine a target scene, then determining a to-be-processed image segment according to the target scene, and then analyzing and processing the to-be-processed image segment to generate a target video. As can be seen, when a target person encounters a scene of interest during vehicle travel, a target video is generated by analyzing and processing an image segment, which is conducive to automatically capturing and generating a beautiful scene video during driving without manual operation by the target person, and is conducive to improving the safety and interest of the generation of an outside beautiful scene video and improving the efficiency and convenience of video generation.
[0048] In one possible example, before the electronic device acquires the image information captured during driving, the electronic device includes: the electronic device labels a scene by means of internet platform resource ingestion or a large amount of manual input to determine a keyword tag corresponding to the scene; and the electronic device generates a scene database according to the scene and the keyword tag.
[0049] The scene database is updated at a preset time to ensure the accuracy of the scene database.
[0050] The scene in the scene database is associated with a keyword tag, and the scene is used as a query identifier to query a preset mapping relationship to obtain a keyword tag corresponding to the scene. The mapping relationship includes a corresponding relationship between a scene and a keyword tag.
[0051] In a specific implementation, the electronic device labels a scene by means of internet platform resource ingestion and intelligent learning to determine a keyword tag corresponding to the scene, such as flowers, fireworks, streams, etc. The electronic device associates a scene with a keyword tag according to the determined scene and keyword tag to generate a scene database.
[0052] It can be seen that in the example, the scene database is established by labeling the scene through the Internet platform resource ingestion or a large number of manual input, which is beneficial to improve the accuracy and efficiency of the target scene determination.
[0053] In one possible example, the electronic device identifies the image information and determines the target scene, including: the electronic device queries the scene database, matches the image information, and determines the scene in the image information; the electronic device scores the scene according to a preset key score model; and the electronic device determines the scene with a key score higher than a first threshold as the target scene according to the scoring result of the key score.
[0054] The first threshold includes a pre-set score threshold.
[0055] The key score model includes scoring the scene according to a target recognition algorithm, a scene relative position, and a scene appearance duration.
[0056] Specifically, the key score model includes: determining a first-level scene set according to a target recognition algorithm; obtaining a scene relative position and a scene appearance duration corresponding to each first-level scene in the first-level scene set; and determining a key score result of the scene according to the scene relative position, the scene appearance duration, and a preset formula.
[0057] Specifically, the preset formula includes P=xt, where x represents a position score corresponding to the scene relative position, t represents the scene appearance duration, and P represents the key score of the scene.
[0058] Specifically, the position score includes that the closer the scene is to the center position of the image frame, the higher the position score is.
[0059] In a specific implementation, the electronic device queries the scene database, matches the image information, and determines the scene in the image information, including "clouds" and "trees"; the electronic device scores the scene according to a preset key score model, and determines that the score of "clouds" is 60 and the score of "trees" is 20; and the electronic device determines that "clouds" with a key score higher than a first threshold as the target scene according to the scoring result of the key score.
[0060] It can be seen that in the example, the key score model is used to quantify the key score of the scene, which is beneficial to improve the accuracy and convenience of the target scene determination, and further beneficial to improve the accuracy of the video generation.
[0061] In a possible example, the electronic device determines the image segment to be processed according to the target scene, including: the electronic device queries the appearance time of the target scene, and determines the timestamp information corresponding to the target scene; the electronic device acquires the interest value of the target person to the target scene according to the timestamp information through a line-of-sight tracking algorithm; and the electronic device determines that the image segment corresponding to the target scene is the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0062] The line-of-sight tracking algorithm includes determining the relationship between the line-of-sight dwell time of the target person and the target scene according to the line-of-sight movement information of the target person, and further determining the interest value of the target person to the target scene.
[0063] In a specific implementation, the electronic device queries the appearance time of the target scene "cloud" as 35 minutes, and determines the timestamp information corresponding to the target scene as "16:02-16:37". The electronic device acquires the interest value of the target person to the target scene as 70 through a line-of-sight tracking algorithm, because the target person's line-of-sight stays in "cloud" for a total of 10 minutes according to the timestamp information. The electronic device determines that the interest value 70 is greater than the first interest threshold, and further determines that the image segment corresponding to the target scene "cloud" is the image segment to be processed.
[0064] It can be seen that, in this example, the interest value of the target person to the target scene is acquired through the line-of-sight tracking algorithm, which is beneficial to improve the efficiency of determining the interest value of the target person to the target scene, and is beneficial to improve the accuracy and efficiency of video generation.
[0065] In a possible example, the electronic device determines the image segment to be processed according to the target scene, including: the electronic device queries the appearance time of the target scene, and determines the timestamp information corresponding to the target scene; the electronic device acquires the posture information of the target person within the appearance time; the electronic device determines the interest value of the target person to the target scene according to the posture information and the timestamp information; and the electronic device determines that the image segment corresponding to the target scene is the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0066] The method further includes: if the line-of-sight of the target person can be recognized, determining the interest value of the target person to the target scene according to the line-of-sight tracking algorithm, the posture information, and the timestamp information.
[0067] The longer the appearance time of the target scene is, the greater the interest value of the target scene is.
[0068] The posture information includes facial expression information and gesture information.
[0069] In a specific implementation, the electronic device queries an appearance time of the target scene "cloud" as 35 minutes, determines time stamp information corresponding to the target scene as "16:02-16:37", obtains posture information of the target person within the appearance time, finds that the target person points to the target scene at 16:10 and 16:23, and the target person has a relatively happy facial expression at 16:24, determines, according to the posture information and the time stamp information, that an interest value of the target person for the target scene is 70, and determines, when the interest value is greater than a first interest threshold, that an image segment corresponding to the target scene "cloud" is a to-be-processed image segment.
[0070] It can be seen that, in the example, when the target person wears an eye shield such as blue light glasses or sunglasses, the interest value is determined through the posture information, which is beneficial to determining the interest value when the line of sight of the target person cannot be tracked, and is beneficial to improving the efficiency of video generation in diversified scenes.
[0071] In one possible example, the electronic device determines, according to the target scene, a to-be-processed image segment, including: the electronic device queries an appearance time of the target scene, and determines time stamp information corresponding to the target scene; the electronic device obtains voice information of the target person within the appearance time; the electronic device determines, according to the voice information and the time stamp information, an interest value of the target person for the target scene through a voice recognition algorithm; and the electronic device determines, when the interest value is greater than a first interest threshold, that an image segment corresponding to the target scene is a to-be-processed image segment.
[0072] The voice information includes conversation voice of the target person and an in-vehicle person, and communication voice of the target person when the target person mentions a scene topic.
[0073] In a specific implementation, the electronic device queries an appearance time of the target scene "moon" as 62 minutes, determines time stamp information corresponding to the target scene as "20:05-20:52" and "21:46-22:01", obtains voice information of the target person within the appearance time, determines, according to the voice information and the time stamp information, that an interest value of the target person for the target scene "moon" is 80 through a voice recognition algorithm, and determines, when the interest value is greater than a first interest threshold, that an image segment corresponding to the target scene is a to-be-processed image segment.
[0074] It can be seen that in this example, the weak light scene in the vehicle is fully considered, which is conducive to meeting the interest value determination in multiple scenes and improving the accuracy of interest value determination, thereby improving the efficiency of video generation.
[0075] In one possible example, the electronic device determines a timestamp information set composed of multiple timestamps corresponding to each of the target scenes; the electronic device determines time intervals between the multiple timestamps corresponding to the target scene according to the timestamp information set; and the electronic device determines the image segment corresponding to the target scene according to the time intervals and the interest value.
[0076] The timestamp information set and the target scene have a mapping relationship, and the timestamp information set and the target scene correspond one-to-one.
[0077] The determination of the image segment corresponding to the target scene according to the time intervals and the interest value includes: determining an image segment time length threshold of the target image according to the interest value; if it is determined that the time interval between the multiple timestamps corresponding to the target scene is less than a second threshold, determining that the image segment between the initial timestamp and the end timestamp constitutes an image segment; and if it is determined that the time interval between the multiple timestamps corresponding to the target scene is greater than the second threshold, determining that the image segment within a certain time of each of the timestamps constitutes multiple image segments.
[0078] Specifically, the image segment time length threshold is greater than or equal to the time between the initial timestamp and the end timestamp, and / or the certain time.
[0079] In a specific implementation, the electronic device determines a timestamp information set composed of timestamps "20:05-20:52" and "21:46-22:01" corresponding to the target scene "moon"; the electronic device determines that the time interval between the multiple timestamps corresponding to the target scene "moon" is 54 minutes according to the timestamp information set; it is determined that the time interval between the multiple timestamps corresponding to the target scene "moon" is greater than a second threshold, and the electronic device determines that the image segment "20:15-20:37" constitutes an image segment 1 and the image segment "21:47-22:00" constitutes an image segment 2.
[0080] It can be seen that in this example, the image segment is determined according to the time interval and the interest value, which is conducive to improving the correlation between the image segment and the target person interest value, and improving the accuracy of video generation.
[0081] In one possible example, the electronic device analyzes and processes the to-be-processed image segment to generate a target video, including: the electronic device queries the scene database according to the target scene in the to-be-processed image segment to determine a target keyword; the electronic device acquires a segment duration of the to-be-processed image segment; the electronic device determines an emotion classification of the to-be-processed image segment through the posture information and the speech recognition algorithm; the electronic device determines music information and text information of the to-be-processed image segment according to the target keyword, the segment duration, and the emotion classification; and the electronic device generates a target video according to the music information and the text information.
[0082] In this example, the target keyword includes a keyword label corresponding to each target scene in the to-be-processed image segment.
[0083] In a specific implementation, the electronic device queries the scene database according to the target scene "moon" in the to-be-processed image segment to determine a target keyword; the electronic device acquires a segment duration of 3 minutes of the to-be-processed image segment, and then determines an emotion classification of the to-be-processed image segment as "missing" through the posture information and the speech recognition algorithm; the electronic device determines music information of the to-be-processed image segment as slow music 1 and determines text information as "spring tide water connects the sea level, and the moon on the sea shares the tide"; and the electronic device generates a target video according to the music information and the text information.
[0084] As can be seen, in this example, the video is generated according to the scenes in the driving process, which is conducive to improving the efficiency of video generation and facilitating the target person to share the scenes in the driving process.
[0085] Consistent with the above Figure 1 described embodiments, please refer to Figure 2 , Figure 2 is another flow diagram of a video generation method based on an outside scene of a vehicle provided by the embodiments of the present application; as shown in the figure, the video generation method based on the outside scene of the vehicle includes:
[0086] S201, the electronic device labels the scene through an internet platform resource ingestion or a large amount of manual input to determine a keyword label corresponding to the scene;
[0087] S202, the electronic device generates a scene database according to the scene and the keyword label;
[0088] S203, the electronic device acquires image information shot in a driving process;
[0089] S204, the electronic device queries the scene database, matches the image information, and determines the scene in the image information;
[0090] S205, the electronic device scores the scene according to a preset key degree scoring model;
[0091] S206, the electronic device determines the scene with a key degree score higher than a first threshold value as a target scene according to the scoring result of the key degree scoring;
[0092] S207, the electronic device determines a to-be-processed image segment according to the target scene;
[0093] S208, the electronic device analyzes and processes the to-be-processed image segment to generate a target video. It can be seen that in the embodiment of the application, a video generation method, system, electronic device and computer storage medium based on an out-of-vehicle scene are provided, the method comprising: first, acquiring image information shot in a driving process; then, identifying the image information to determine a target scene; then, determining a to-be-processed image segment according to the target scene; and then, analyzing and processing the to-be-processed image segment to generate a target video. It can be seen that when a target person encounters a scene of interest in a driving process, a target video is generated by analyzing and processing an image segment, which is conducive to automatically shooting and generating a beautiful scene video in a driving process without manual operation by the target person, and is conducive to improving the safety and interest of the out-of-vehicle beautiful scene video generation and improving the efficiency and convenience of the video generation.
[0094] In addition, the scenes are labeled by means of internet platform resource ingestion or a large amount of manual input, and then a scene database is established, and a key degree scoring model is further used to quantify the key degree of the scene, which is conducive to improving the accuracy and convenience of target scene determination, and further conducive to improving the accuracy of video generation.
[0095] In accordance with the above Figure 1 embodiment, please refer to Figure 3 , Figure 3 is a flowchart of another video generation method based on an out-of-vehicle scene provided by the embodiment of the application; as shown in the figure, the video generation method based on the out-of-vehicle scene comprises:
[0096] S301, the electronic device labels the scene by means of internet platform resource ingestion or a large amount of manual input to determine the key word label corresponding to the scene;
[0097] S302, the electronic device generates a scene database according to the scene and the key word label;
[0098] S303, the electronic device acquires image information photographed in a driving process;
[0099] S304, the electronic device queries the scene database, matches the image information, and determines a scene in the image information;
[0100] S305, the electronic device scores the scene according to a preset key degree scoring model;
[0101] S306, the electronic device determines a scene with a key degree score higher than a first threshold value as a target scene according to a scoring result of the key degree scoring;
[0102] S307, the electronic device queries an appearance time of the target scene and determines timestamp information corresponding to the target scene;
[0103] S308, the electronic device acquires an interest value of a target person to the target scene through a line-of-sight tracking algorithm according to the timestamp information;
[0104] S309, the electronic device determines an image segment corresponding to the target scene as a to-be-processed image segment if the interest value is greater than a first interest threshold value;
[0105] S310, the electronic device analyzes and processes the to-be-processed image segment to generate a target video.
[0106] It can be seen that in the embodiments of the present application, a video generation method, system, electronic device and computer storage medium based on an out-of-vehicle scene are provided, the method comprises: first acquiring image information photographed in a driving process, then identifying the image information, determining a target scene, then determining a to-be-processed image segment according to the target scene, and then analyzing and processing the to-be-processed image segment to generate a target video; it can be seen that when a target person encounters a scene of interest in a vehicle driving process, a target video is generated by analyzing and processing an image segment, which is beneficial to automatically photograph and generate a beautiful scene video in a driving process without manual operation of the target person, and is beneficial to improving the safety and interest of out-of-vehicle beautiful scene video generation, and is beneficial to improving the efficiency and convenience of video generation.
[0107] In addition, the interest value of the target person to the target scene is acquired through a line-of-sight tracking algorithm, which is beneficial to improving the efficiency of judging the interest value of the target person to the target scene, and is beneficial to improving the accuracy and efficiency of video generation.
[0108] The embodiments are consistent with the above Figure 1 , Figure 2 , Figure 3 The embodiments are consistent with the above Figure 4 ,Figure 4 is a schematic diagram of a vehicle exterior scene based video generation system 400 involved in an embodiment of the present application. The vehicle exterior scene based video generation system 400 is applied to an electronic device; the vehicle exterior scene based video generation system 400 comprises a processor 401, a memory 402, and a communication unit 403; the processor 401 is in communication connection with the memory 402 and the communication unit 403;
[0109] The processor 401 communicates with the communication unit 403 and executes instructions stored in the memory 402;
[0110] The communication unit 403 acquires image information captured during driving;
[0111] The memory 402 is configured to store instructions, which, when executed by the processor 401, cause the processor 401 to perform steps, the steps comprising:
[0112] Identifying the image information to determine a target scene; determining a to-be-processed image segment according to the target scene; and analyzing and processing the to-be-processed image segment to generate a target video.
[0113] It can be seen that in the embodiments of the present application, a vehicle exterior scene based video generation method, system, electronic device, and computer storage medium are provided, the method comprising first acquiring image information captured during driving, then identifying the image information to determine a target scene, then determining a to-be-processed image segment according to the target scene, and then analyzing and processing the to-be-processed image segment to generate a target video; it can be seen that when a target person encounters a scene of interest during vehicle driving, a target video is generated by analyzing and processing an image segment, which is conducive to automatically capturing and generating a beautiful scene video during driving without manual operation by the target person, and is conducive to improving the safety and interest of vehicle exterior beautiful scene video generation and improving the efficiency and convenience of video generation.
[0114] In one possible example, before the acquisition of the image information captured during driving, the processor 401 is specifically configured to perform the following steps: annotating a scene through internet platform resource ingestion or in a large number of manual input manner to determine a keyword label corresponding to the scene; and generating a scene database according to the scene and the keyword label.
[0115] In a possible example, the identifying the image information and determining the target scene, the processor 401 is specifically configured to perform the following steps: querying the scene database, matching the image information, and determining the scene in the image information; performing key degree scoring on the scene according to a preset key degree scoring model, wherein the key degree scoring model comprises scoring the scene according to a target recognition algorithm, a scene relative position, and a scene appearance time length; and determining a scene in which a key score is higher than a first threshold as the target scene according to a scoring result of the key degree scoring.
[0116] In a possible example, the determining the image segment to be processed according to the target scene, the processor 401 is specifically configured to perform the following steps: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining an interest value of a target person for the target scene by a line-of-sight tracking algorithm according to the timestamp information; and determining an image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0117] In a possible example, the determining the image segment to be processed according to the target scene, the processor 401 is specifically configured to perform the following steps: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining posture information of a target person in the appearance time, wherein the posture information comprises facial expression information and gesture information; determining an interest value of the target person for the target scene according to the posture information and the timestamp information; and determining an image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0118] In a possible example, the determining the image segment to be processed according to the target scene, the processor 401 is specifically configured to perform the following steps: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining voice information of a target person in the appearance time; determining an interest value of the target person for the target scene by a voice recognition algorithm according to the voice information and the timestamp information; and determining an image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0119] In a possible example, the target scene corresponds to a video clip, and the processor 401 is specifically configured to perform the following steps: determining a timestamp information set composed of a plurality of timestamp information corresponding to each target scene, wherein the timestamp information set and the target scene have a mapping relationship, and the timestamp information set and the target scene correspond one by one; determining a time interval between a plurality of timestamp information corresponding to the target scene according to the timestamp information set; and determining the video clip corresponding to the target scene according to the time interval and the interest value.
[0120] In a possible example, the target scene corresponds to a video clip, and the processor 401 is specifically configured to perform the following steps: determining a timestamp information set composed of a plurality of timestamp information corresponding to each target scene, wherein the timestamp information set and the target scene have a mapping relationship, and the timestamp information set and the target scene correspond one by one; determining a time interval between a plurality of timestamp information corresponding to the target scene according to the timestamp information set; and determining the video clip corresponding to the target scene according to the time interval and the interest value.
[0121] The above mainly introduces the scheme of the embodiments of the present application from the perspective of the execution process of the method. It can be understood that, in order to implement the above functions, the electronic device contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the unit and algorithm steps of each example described in the embodiments provided in the present text, the present application can be implemented in the form of hardware or the combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. The professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0122] The above mainly introduces the scheme of the embodiments of the present application from the perspective of the execution process of the method. It can be understood that, in order to implement the above functions, the electronic device contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the unit and algorithm steps of each example described in the present text, the present application can be implemented in the form of hardware or the combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. The professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. Figure 1 、 Figure 2 、 Figure 3 The above mainly introduces the scheme of the embodiments of the present application from the perspective of the execution process of the method. It can be understood that, in order to implement the above functions, the electronic device contains the hardware structure and / or software module corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the unit and algorithm steps of each example described in the present text, the present application can be implemented in the form of hardware or the combination of hardware and computer software. Whether a certain function is implemented in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. The professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application. Figure 5 , Figure 5 is a structural schematic diagram of an electronic device 500 provided by the embodiments of the present application, as shown in the figure, the electronic device 500 comprises a processor 510, a memory 520, a communication interface 530 and one or more programs 521, wherein the one or more programs 521 are stored in the above-mentioned memory 520 and are configured to be executed by the above-mentioned processor 510, and the one or more programs 521 comprise instructions for performing the following steps:
[0123] acquire image information captured during driving;
[0124] identify the image information to determine a target scene;
[0125] determine a to-be-processed image segment according to the target scene;
[0126] analyze and process the to-be-processed image segment to generate a target video.
[0127] It can be seen that in the embodiments of the present application, a video generation method, system, electronic device and computer storage medium based on an outside scene are provided, the method includes first acquiring image information captured during driving, then identifying the image information to determine a target scene, then determining a to-be-processed image segment according to the target scene, and then analyzing and processing the to-be-processed image segment to generate a target video. It can be seen that when a target person encounters a scene of interest during vehicle driving, a target video is generated by analyzing and processing an image segment, which is conducive to automatically capturing and generating a beautiful scene video during driving without manual operation by the target person, and is conducive to improving the safety and interest of the generation of an outside beautiful scene video and improving the efficiency and convenience of video generation.
[0128] In one possible example, before the acquiring of the image information captured during driving, the instructions in the program are specifically used to perform the following operation: annotating a scene through an Internet platform resource ingestion or in a large amount of manual input manner to determine a keyword label corresponding to the scene; and generating a scene database according to the scene and the keyword label.
[0129] In one possible example, the identifying of the image information to determine a target scene, the instructions in the program are specifically used to perform the following operation: querying the scene database, matching the image information to determine a scene in the image information; performing a key degree scoring on the scene according to a preset key degree scoring model, wherein the key degree scoring model includes scoring the scene one by one according to a target recognition algorithm, a scene relative position and a scene appearance duration; and determining a scene in which a key score is higher than a first threshold value as a target scene according to a scoring result of the key degree scoring.
[0130] In a possible example, the determining the image segment to be processed according to the target scene, the instructions in the program specifically perform the following operation: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining an interest value of the target person to the target scene according to the timestamp information and through a line-of-sight tracking algorithm; and determining the image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0131] In a possible example, the determining the image segment to be processed according to the target scene, the instructions in the program specifically perform the following operation: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining posture information of the target person within the appearance time, wherein the posture information includes facial expression information and gesture information; determining an interest value of the target person to the target scene according to the posture information and the timestamp information; and determining the image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0132] In a possible example, the determining the image segment to be processed according to the target scene, the instructions in the program specifically perform the following operation: querying an appearance time of the target scene, determining timestamp information corresponding to the target scene; obtaining voice information of the target person within the appearance time; determining an interest value of the target person to the target scene according to the voice information and the timestamp information through a voice recognition algorithm; and determining the image segment corresponding to the target scene as the image segment to be processed if it is determined that the interest value is greater than a first interest threshold.
[0133] In a possible example, the image segment corresponding to the target scene, the instructions in the program specifically perform the following operation: determining a timestamp information set composed of a plurality of timestamp information corresponding to each target scene, wherein the timestamp information set and the target scene have a mapping relationship, and the timestamp information set and the target scene correspond one by one; determining a time interval between a plurality of the timestamp information corresponding to the target scene according to the timestamp information set; and determining the image segment corresponding to the target scene according to the time interval and the interest value.
[0134] In a possible example, the analyzing and processing of the to-be-processed image segment to generate a target video, the instructions in the program are specifically used for performing the following operations: querying the scene database according to the target scene in the to-be-processed image segment to determine a target keyword, wherein the target keyword includes a keyword label corresponding to each target scene in the to-be-processed image segment; obtaining a segment duration of the to-be-processed image segment; determining an emotion classification of the to-be-processed image segment through the posture information and the speech recognition algorithm; determining music information and text information of the to-be-processed image segment according to the target keyword, the segment duration and the emotion classification; and generating a target video according to the music information and the text information.
[0135] The embodiments of the present application can divide the functional units of the electronic device according to the above method examples. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in the form of hardware or software functional unit. It should be noted that the division of the units in the embodiments of the present application is illustrative, and is only a logical function division. In actual implementation, another division mode can be used.
[0136] The embodiments of the present application also provide a computer storage medium, which stores a computer program for electronic data exchange. The computer program causes a computer to execute part or all of the steps of any method described in the above method embodiments. The computer includes an electronic device.
[0137] The embodiments of the present application also provide a computer program product. The computer program product includes a non-transitory computer readable storage medium storing a computer program. The computer program is operable to cause a computer to execute part or all of the steps of any method described in the above method embodiments. The computer program product can be a software installation package. The computer includes an electronic device.
[0138] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, certain steps can be performed in other order or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily necessary for the present application.
[0139] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0140] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely illustrative, and the division of the units can be changed according to actual needs. For example, the units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0141] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiments.
[0142] In addition, each functional unit in the various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0143] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0144] Those of ordinary skill in the art can understand that all or part of the steps of the various methods in the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable storage medium, which can include a flash disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, etc.
[0145] The above has introduced the embodiments of the present application in detail, and the principles and implementation manners of the present application are described by applying specific examples; the above embodiment description is only for helping to understand the method of the present application and its core idea; meanwhile, for the general technical personnel in the art, according to the idea of the present application, the specific implementation manner and application range will have changes; and in view of the above, the content of the specification should not be understood as the limitation of the present application.
Claims
1. A method for generating a video based on an out-of-vehicle scene, the method comprising: obtaining image information captured during driving; identifying the image information to determine a target scene; determining a to-be-processed image segment according to the target scene, wherein a time of appearance of the target scene is queried to determine timestamp information corresponding to the target scene, posture information of a target person within the time of appearance is obtained, the posture information including facial expression information and gesture information, an interest value of the target person for the target scene is determined according to the posture information and the timestamp information, and if the interest value is determined to be greater than a first interest threshold, an image segment corresponding to the target scene is determined to be a to-be-processed image segment; analyzing and processing the to-be-processed image segment to generate a target video, including querying a scene database according to the target scene in the to-be-processed image segment to determine a target keyword, the target keyword including a keyword label corresponding to each target scene in the to-be-processed image segment, obtaining a segment duration of the to-be-processed image segment, determining an emotion classification of the to-be-processed image segment through a posture information and voice recognition algorithm, determining music information and text information of the to-be-processed image segment according to the target keyword, the segment duration and the emotion classification, and generating a target video according to the music information and the text information.
2. The method of claim 1, wherein, Before the obtaining image information captured during driving, the method comprises: annotating scenes through internet platform resources or in a large number of manual input manner to determine keyword labels corresponding to the scenes; and generating the scene database according to the scenes and the keyword labels.
3. The method of claim 2, wherein, The identifying the image information to determine a target scene comprises: querying the scene database to match the image information to determine scenes in the image information; performing key degree scoring on the scenes according to a preset key degree scoring model, the key degree scoring model including scoring the scenes one by one according to a target recognition algorithm, scene relative position and scene appearance duration; and determining a scene in which a key score is higher than a first threshold as a target scene according to a scoring result of the key degree scoring.
4. The method of claim 3, wherein, The determining a to-be-processed image segment according to the target scene comprises: querying a time of appearance of the target scene to determine timestamp information corresponding to the target scene; determining an interest value of a target person for the target scene through a line-of-sight tracking algorithm according to the timestamp information; and if the interest value is determined to be greater than a first interest threshold, determining an image segment corresponding to the target scene to be a to-be-processed image segment.
5. The method of claim 3, wherein, The determining a to-be-processed image segment according to the target scene comprises: querying a time of appearance of the target scene to determine timestamp information corresponding to the target scene; obtaining voice information of a target person within the time of appearance; determining an interest value of the target person for the target scene through a voice recognition algorithm according to the voice information and the timestamp information. If it is determined that the interest value is greater than a first interest threshold, it is determined that the image segment corresponding to the target scene is a to-be-processed image segment.
6. The method according to claim 4 or 5, characterized in that, The image segment corresponding to the target scene includes: A timestamp information set composed of a plurality of timestamp information corresponding to each target scene is determined, wherein the timestamp information set and the target scene have a mapping relationship, and the timestamp information set and the target scene correspond one by one. According to the timestamp information set, the time interval between a plurality of timestamp information corresponding to the target scene is determined. According to the time interval and the interest value, the image segment corresponding to the target scene is determined.
7. An off-board scene based video generation system, the video generation system comprising: A processor, a memory, and a communication unit; The processor is in communication connection with the memory and the communication unit; The processor communicates with the communication unit and executes the instructions stored in the memory; The communication unit acquires image information shot during driving; The memory is configured to store instructions which, when executed by the processor, cause the processor to perform steps including: Identify the target scene from the image information, and determine a to-be-processed image segment according to the target scene, wherein the appearance time of the target scene is queried, and timestamp information corresponding to the target scene is determined; the posture information of the target person within the appearance time is acquired, wherein the posture information includes facial expression information and gesture information; the interest value of the target person to the target scene is determined according to the posture information and the timestamp information; if it is determined that the interest value is greater than a first interest threshold, it is determined that the image segment corresponding to the target scene is a to-be-processed image segment; the to-be-processed image segment is analyzed and processed to generate a target video, including: according to the target scene in the to-be-processed image segment, querying a scene database to determine a target keyword, wherein the target keyword includes a keyword tag corresponding to each target scene in the to-be-processed image segment; the segment duration of the to-be-processed image segment is acquired; the emotion classification of the to-be-processed image segment is determined through a posture information and voice recognition algorithm; the music information and the text information of the to-be-processed image segment are determined according to the target keyword, the segment duration, and the emotion classification; and a target video is generated according to the music information and the text information.
8. An electronic device, comprising: A processor, a memory, and a communication interface, and one or more programs stored in the memory and configured to be executed by the processor, the programs including instructions for performing steps in the method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, A computer program for electronic data exchange, wherein the computer program causes a computer to perform the method of any one of claims 1-6.
Citation Information
Patent Citations
Video processing methods and devices, electronic equipment and memory medium
CN109618184A
Video data processing method and device, computer readable medium and electronic equipment
CN110198432A
System and method for detecting human gaze and gesture in unconstrained environments
CN111989537A
Driving along-way landscape snapshot method and device and automobile center console
CN112954207A