A video generation method, apparatus, device, and medium
By extracting the target sequence of scene video and correcting the display method according to the target projection, new videos are generated, which solves the problems of high cost and poor authenticity in the prior art, and achieves low-cost and high-reality video generation.
Patent Information
- Application Number
- CN202211383007.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-07
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2042-11-07
AI Technical Summary
In the prior art, when acquiring video materials in specific scenarios, there are problems of high cost and poor authenticity, especially when generating videos through manual simulation and three-dimensional rendering, there are great safety risks and limited operation.
By obtaining the background image of the scene and videos of different time periods, the target sequence is extracted and the display mode of overlapping parts is corrected according to the projection area of the target to generate a new video.
It effectively saves the cost of video generation and improves the authenticity of generated videos, especially the display accuracy in target occlusion.
Smart Images

Figure CN115866328B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a video generation method, apparatus, device, and medium. Background Art
[0002] Video analysis technology is one of the important research contents in many fields such as intelligent transportation and smart cities. However, for videos in some specific scenarios, such as videos of vehicle collision events or motor vehicle non-yielding to pedestrians events at intersections, since the frequency of such events is low, it is difficult to obtain such videos. The following several methods have been proposed in the related art to obtain such videos:
[0003] The first method: Obtain the video by artificial simulation. For example, arrange dummy or dummy vehicle models in a test environment, and manually command the driver to drive the vehicle according to a set route to simulate collision or non-yielding to pedestrians events. This method requires artificial simulation for each scenario, with a large workload and potential safety hazards.
[0004] The second method: Obtain the video through 3D software modeling and rendering. Simulate the process of event occurrence by building a 3D simulation environment and 3D models of various objects in the environment. Although this video acquisition method can avoid safety hazards and can also handle multiple scenarios, the operation of simulating the video is limited by 3D rendering technology, and there are certain differences between the simulated video and the real scenario. Moreover, generating videos through 3D modeling requires a large investment cost. Summary of the Invention
[0005] Embodiments of the present invention provide a video generation method, apparatus, device, and medium, which are used to solve problems such as high cost and poor authenticity in existing video generation methods.
[0006] In a first aspect, the present application proposes a video generation method, including:
[0007] Obtain a background image of a scene, as well as a first video and a second video captured of the scene, and obtain a first target sequence and a second target sequence according to the first video and the second video, where the first target sequence is an image sequence of an area including a first target in the first video, and the second target sequence is an image sequence of an area including a second target in the second video;
[0008] Sequentially extract the target areas in the first target sequence and the second target sequence in the order of shooting time, and combine each extracted target area with the background image;
[0009] When there is an overlapping part in the target area of any frame image obtained by combination, according to the projections of the first target and the second target on the ground included in the any frame image in the set direction, correct the display mode of the overlapping part;
[0010] Combine the corrected images in the order of shooting time to form the video to be generated.
[0011] In some embodiments, the correcting the display mode of the overlapping part according to the projections of the first target and the second target on the ground included in the any frame image in the set direction includes:
[0012] Determine the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target;
[0013] If the first target is the occluded target, make the second target visible in the overlapping part;
[0014] If the second target is the occluded target, make the first target visible in the overlapping part.
[0015] In some embodiments, the determining the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target includes:
[0016] Determine a line passing through any point in the projection of the first target along a preset direction;
[0017] Determine the intersection point of the line and the second target, and judge whether the intersection point is above the any point in the preset direction;
[0018] If the intersection point is above the any point, determine that the second target is the occluded target;
[0019] If the intersection point is not above the any point, determine that the first target is the occluded target.
[0020] In some embodiments, the obtaining the background image of the scene includes:
[0021] Use an image that does not contain the first target and the second target in at least one frame of the third video obtained by shooting the scene as the background image; or,
[0022] Generate the background image based on a preset background generation algorithm according to at least one frame of the third video.
[0023] In some embodiments, the following method is used to determine the projection of the first target:
[0024] Obtain multiple sets of key points and multiple ground contact points included in the first target in any one of the frames of images; the distance between two key points included in each set of key points in the multiple sets of key points is used to characterize the size of the first target;
[0025] Determine the size and shape of the projection of the first target according to the multiple sets of key points, and determine the position of the projection of the first target in the ground in any one of the frames of images according to the multiple ground contact points.
[0026] In some embodiments, the forming the first target sequence from the regions including the first target in each frame of the first video includes:
[0027] Perform target contour detection on each frame of the first video through a preset target contour detection algorithm to obtain the regions including the first target in each frame of the first video;
[0028] Form the regions of the first target obtained by detection into a first target sequence according to the shooting time of at least one frame of image included in the first video.
[0029] In a second aspect, the present application provides a video generation device, and the device includes:
[0030] An acquisition unit, configured to acquire a background image of a scene, and a first video and a second video obtained by shooting the scene;
[0031] A processing unit, configured to perform:
[0032] Obtain a first target sequence and a second target sequence according to the first video and the second video, where the first target sequence is an image sequence of the regions including the first target in the first video, and the second target sequence is an image sequence of the regions including the second target in the second video;
[0033] Sequentially extract the target regions in the first target sequence and the second target sequence according to the shooting time order, and combine each extracted target region with the background image;
[0034] When there is an overlapping part in the target regions in any one of the combined frames of images, correct the display mode of the overlapping part according to the projections of the first target and the second target on the ground included in any one of the frames of images in a set direction;
[0035] Combine the corrected images into a video to be generated according to the shooting time order.
[0036] In some embodiments, when the processing unit corrects the display mode of the overlapping part according to the projections of the first target and the second target on the ground included in any frame of the image in a set direction, it is specifically configured to:
[0037] Determine the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target;
[0038] If the first target is the occluded target, set the second target to be visible in the overlapping part;
[0039] If the second target is the occluded target, set the first target to be visible in the overlapping part.
[0040] In some embodiments, when the processing unit determines the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target, it is specifically configured to:
[0041] Determine a line passing through any point in the projection of the first target along a preset direction;
[0042] Determine the intersection point of the line and the second target, and determine whether the intersection point is above the any point in the preset direction;
[0043] If the intersection point is above the any point, determine that the second target is the occluded target;
[0044] If the intersection point is not above the any point, determine that the first target is the occluded target.
[0045] In some embodiments, when the acquisition unit acquires the background image of the scene, it is specifically configured to:
[0046] Use an image that does not include the first target and the second target in at least one frame of the third video obtained by photographing the scene as the background image; or,
[0047] Generate the background image according to at least one frame of the image included in the third video based on a preset background generation algorithm.
[0048] In some embodiments, when the processing unit determines the projection of the first target, it is specifically configured to:
[0049] Obtain, through the acquisition unit, multiple sets of key points and multiple ground contact points included in the first target in any frame of the image; the distance between two key points included in each set of key points in the multiple sets of key points is used to characterize the size of the first target;
[0050] Determine the size and shape of the projection of the first target based on the multiple sets of key points, and determine the position of the projection of the first target in any frame of the image based on the multiple ground contact points.
[0051] In some embodiments, the processing unit is specifically configured to:
[0052] Perform target contour detection on each frame of the first video through a preset target contour detection algorithm to obtain the region including the first target in each frame of the first video;
[0053] According to the shooting time of at least one frame of the first video, form a first target sequence from the detected regions of the first target.
[0054] In a third aspect, an electronic device is provided. The electronic device includes a controller and a memory. The memory is used to store computer execution instructions, and the controller executes the computer execution instructions in the memory to perform the operation steps of any possible implementation method of the first aspect by using the hardware resources in the controller.
[0055] In a fourth aspect, a computer-readable storage medium is provided. Instructions are stored in the computer-readable storage medium, and when they run on a computer, the computer is caused to execute the methods of the above aspects.
[0056] The present application proposes a video generation method, which extracts the target regions included in each frame of a video to form a target image sequence, combines the extracted target image sequence with the background image of the shooting scene to generate a set of new images, and combines the generated set of new images into a new video. Compared with the prior art that uses 3D rendering, it can save the cost of video generation. Moreover, the present application also proposes that if there is a problem of mutual occlusion between targets in the generated new images, the occluded targets can be determined based on the projection regions of the targets, so as to determine the display mode of the generated new images, improving the authenticity of the newly generated video. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0058] Figure 1 It is a flowchart of a video generation method provided by an embodiment of the present application;
[0059] Figure 2Schematic diagram of a process for extracting a background image and a first target image sequence from a video provided by an embodiment of the present application;
[0060] Figure 3 Schematic diagram of a process for generating a new video provided by an embodiment of the present application;
[0061] Figure 4 Another schematic diagram of a process for generating a new video provided by an embodiment of the present application;
[0062] Figure 5 Schematic diagram of a situation with an overlapping target area provided by an embodiment of the present application;
[0063] Figure 6 Schematic diagram of a scenario for determining an occluded target based on a projection area provided by an embodiment of the present application;
[0064] Figure 7 Schematic diagram of the ground projection of a vehicle provided by an embodiment of the present application;
[0065] Figure 8 Schematic diagram of the ground contact points of a bicycle provided by an embodiment of the present application;
[0066] Figure 9 Schematic diagram of the ground contact points of a road cone provided by an embodiment of the present application;
[0067] Figure 10 Another flowchart of a video generation method provided by an embodiment of the present application;
[0068] Figure 11 Schematic diagram of the structure of a video generation device provided by an embodiment of the present application;
[0069] Figure 12 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0071] It should be noted that the terms "first", "second", etc. in this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described here can be implemented in an order other than those illustrated or described here. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are only examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.
[0072] In the related art, in order to increase video materials, generally, more video materials are obtained by deploying models in a set scene or by modeling and rendering through 3D software. These two methods have problems such as high security risks, high costs, and poor authenticity. To solve these problems, this application provides a video generation method, which extracts targets from the image frames included in the video of the set scene collected at different times to obtain a target image sequence. Based on the background image of the set scene, the obtained target image sequence is combined with the background image frame by frame to generate a new video, thereby increasing the video material library. Compared with the existing methods for generating videos, the video generation scheme proposed in this application effectively saves costs and improves the authenticity of the newly generated videos.
[0073] Next, a video generation method, device, equipment, and medium proposed in this application will be specifically introduced. In the following embodiments of this application, "and / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c can be single or multiple. The singular expression forms "a", "one kind", "the", "the above", "this", and "this one" are also intended to include expressions such as "one or more", unless the context clearly indicates otherwise. And, unless otherwise stated, the ordinal numbers such as "first", "second", etc. mentioned in the embodiments of this application are used to distinguish multiple objects and do not limit the order, time sequence, priority, or importance of multiple objects. For example, the first task execution device and the second task execution device are only used to distinguish different task execution devices, rather than indicating differences in the priority or importance of these two task execution devices, etc.
[0074] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of this application. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in another way. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in another way.
[0075] Next, the solution provided by this application will be specifically introduced. Refer to Figure 1 , which is a flowchart of a video generation method provided by an embodiment of this application. It should be noted that this application does not limit the execution subject of the video generation method. For example, it can be executed by terminal devices such as computers and mobile phones, or by electronic devices with computing functions such as servers, server clusters, processors, or processing chips, or by a computing platform in the cloud. Figure 1 The specific method flow shown includes:
[0076] 101. Obtain the background image of the scene, as well as the first video and the second video captured of the scene.
[0077] Among them, the first video and the second video are different videos captured of the scene. Optionally, the first video can be any video captured of the scene, that is, it can be a video captured of the scene by a camera within any time period. Optionally, the first video and the second video can be two videos captured of the scene by the camera at different time periods.
[0078] 102. Obtain a first target sequence according to the first video, and obtain a second target sequence according to the second video.
[0079] Among them, the first target sequence is an image sequence of the area in the first video that includes the first target, that is, the first target sequence includes multiple target areas, and any target area is an image of the area in any frame image of the first video that includes the first target. Among them, the second target sequence is an image sequence of the area in the second video that includes the second target, that is, the second target sequence includes multiple target areas, and any target area is an image of the area in any frame image of the second video that includes the second target.
[0080] Optionally, the first target and the second target can be of the same type of target. For example, they can both be vehicles. Or the first target and the second target can be of different types of targets. For example, the first target can be a person, and the second target can be a vehicle.
[0081] 103. Sequentially extract the target regions in the first target sequence and the second target sequence in the order of shooting time, and combine the target region extracted each time with the background image.
[0082] For example, the first target region in the first target sequence and the first target region in the second target sequence can be extracted, and the two extracted target regions are combined with the background image to obtain a new frame of image.
[0083] 104. When there is an overlapping part in the target regions in any frame of the combined image, correct the display mode of the overlapping part according to the projections of the first target and the second target on the ground included in any frame of image in the set direction.
[0084] In some possible cases, when combining two target regions and the background image respectively from the first target sequence and the second target sequence, there may be an overlapping part in the positions of the two target regions in the background image. In this case, the present application proposes that the layers where the two target regions are located can be determined according to the ground projections of the two target regions, and the display mode of the overlapping part in the image can be determined based on the layers of the two target regions. As an alternative, when there is an overlapping part in the target regions in any frame of the image obtained by combining the two target regions and the background image, the display mode of the overlapping part of the two target regions in any frame of the image can be corrected according to the projections of the two target regions on the ground included in any frame of the image in the set direction.
[0085] 105. Combine the corrected images in the order of shooting time into the video to be generated.
[0086] Based on the above solution, the present application proposes a video generation method, which extracts the target regions included in each frame of the video to form a target image sequence, combines the extracted target image sequence with the background image of the shooting scene to generate a group of new images, and combines the generated group of new images into a new video. Compared with the prior art which uses 3D rendering, it can save the cost of video generation. Moreover, the present application proposes that if there is a problem of mutual occlusion between targets in the generated new image, the occluded target can be determined based on the projection area of the target, so as to determine the display mode of the generated new image, improving the authenticity of the newly generated video.
[0087] Exemplarily, when correcting the display mode of the overlapping part of the first target and the second target in any frame of image, the occluded target among the first target and the second target can be determined according to the projections of the first target and the second target. If the first target is the occluded target, the second target is set to be visible in the overlapping part of any frame of image. If the second target is the occluded target, the first target is set to be visible in the overlapping part of any frame of image.
[0088] As an alternative, when judging the occluded target among two targets according to the projections of the two targets, a line passing through any point in the projection of the first target can be determined along a preset direction. Further, the intersection point of the line and the second target is determined, and it is judged whether the intersection point is above any point in the preset direction. If the intersection point is above any point, it can be determined that the second target is the occluded target. If the intersection point is not above any point, it can be determined that the first target is the occluded target.
[0089] In some embodiments, before combining the first target sequence, the second target sequence and the background image into a new video, the background image, the first target sequence and the second target sequence of the scene can be obtained from the video shot in the real scene first.
[0090] In one or more embodiments, when obtaining the first target sequence from the first video, the target contour detection algorithm can be used to perform target contour detection on each frame of image included in the first video to obtain the sub-images included in each frame of image. The sub-image included in any frame of image is the area where the first target exists in that frame of image. Further, according to the shooting time of at least one frame of image included in the first video, the multiple detected sub-images can be combined into the first target sequence in chronological order.
[0091] Optionally, the first target detected from each frame of image can be any one or more of objects such as pedestrians, vehicles, animals, specific items (such as packages or cable cars, etc.) or buildings existing in the scene during the period of shooting the first video.
[0092] As a possible implementation, when extracting the first target sequence from at least one frame image included in the first video, the contour of the first target existing in each frame image can be extracted and composed into the first target sequence in chronological order. Optionally, when obtaining the contour information of the first target, an image segmentation algorithm can be used to segment the target included in a certain frame image, or a contour tracking algorithm can also be used. Among them, the image segmentation algorithm can use image object segmentation algorithms such as Mask-RCNN, PolarMask, SOLO, etc. The contour tracking algorithm can define the contour of the first target as a polygon point set, and track the contour of the first target in the video by tracking the polygon point set in at least one frame image.
[0093] Above, the process of obtaining the target sequence is introduced by taking the extraction of the first target sequence from the first video as an example. Optionally, the steps of extracting the second target sequence from the second video can refer to the steps of extracting the first target sequence introduced in the above embodiments, and will not be introduced in detail. Next, the process of obtaining the background image of the scene will be introduced. Optionally, the background image can be an image when there are no targets (any target includes the first target and the second target) in the scene. Optionally, the background image can be obtained from the third video captured from the shooting scene. Among them, the third video can be any video captured from the shooting scene, and the third video can be different from the first video and the second video, that is, the time of shooting the third video can be different from the time of shooting the first video and the second video. Or the third video can also be the first video or the second video.
[0094] In some embodiments, when obtaining the background image from the third video, the image that does not include the first target and the second target in at least one frame image included in the third video can be used as the background image. For example, a video of a set time period captured by a camera can be extracted as the third video, such as a video of a certain time period in the early morning or at night. Since there are fewer or no targets in the scene during the set time period, the image that does not include the target in the third video can be used as the background image.
[0095] In other embodiments, when obtaining the background image, the background image can also be generated based on a pre-set background generation algorithm according to at least one frame image included in the third video. For example, a background modeling method based on mean shift can be used to generate the background image.
[0096] For example, at least one pixel point at a fixed position in at least one frame image included in the third video can be selected as a sample point, and a set of typical points can be selected from the at least one sample point. Optionally, the typical points can be obtained by sampling from the at least one sample point, or can be the local average of the at least one sample point. Further, the typical points can be iteratively calculated repeatedly until converging to a preset local extreme value to obtain a plurality of convergence points. The convergence points are divided into at least one point set according to the size order of the convergence points, and the difference between the convergence points included in each point set is less than a preset value. Still further, according to the weight corresponding to each point set, the weighted sum of all the convergence points is calculated as the background of the selected pixel point (that is to say, if the pixel point contains a target, the calculated weighted sum is the remaining background after removing the target contained in the pixel point).
[0097] Optionally, when obtaining the background image, a background modeling method based on Gaussian mixture or a background modeling method based on deep learning can also be adopted. Optionally, the background modeling method based on deep learning can adopt a deep learning background modeling method based on an autoencoder network. Or the background of a single-frame image can also be identified based on an image segmentation algorithm, and then the background image is determined by a multi-frame combination method. The multi-frame combination background image determination method can complement the missing partial area in the single-frame image caused by the occlusion of the target to the background.
[0098] To further understand the process of obtaining the background image and the first target sequence proposed in this application, reference can be made to Figure 2 , which is a process of extracting a background image and a first target sequence from a video proposed in an embodiment of this application. Among them, Figure 2 (a) in is the original video (which can be understood as the first video or the third video introduced in the above embodiment), and the vehicle shown in it is the first target. The original video shows the process of the first target running from point A to point B. Figure 2 (b) in is the background image extracted from the original video. It can be seen that (b) does not include any target. Specifically, the method of extracting Figure 2 (b) from Figure 2 (a) can be referred to the introduction in the above embodiment and will not be elaborated here. Figure 2 (c) in is the first target sequence extracted from the original video, where Figure 2 each target contour shown in (c) can be understood as the sub-image extracted from each frame image introduced in the above embodiment, and the first target sequence is composed of multiple sub-images. For the method of detecting the contour of the first target from each frame image to obtain the sub-image and forming the detected multiple sub-images into the first target sequence, reference can be made to the above introduction and will not be elaborated here.
[0099] The above takesFigure 2 Taking [a certain example] as an example, the process of extracting the background image and the target sequences from the video is introduced. Next, the process of generating a new video by combining the first target sequence, the second target sequence, and the background image will be specifically introduced.
[0100] In some scenarios, the first target sequence and the second target sequence can be respectively extracted from the first video and the second video, where the different videos can be obtained by shooting the same scene at different time periods, or can also be obtained by shooting different scenes. This application does not make any restrictions in this regard. For the convenience of subsequent description, it is assumed that the number of sub-images included in the first target sequence is the same as the number of sub-images included in the second target sequence.
[0101] As an example, refer to Figure 3 , which is a process provided by an embodiment of this application for combining the target sequences and the background image extracted from two videos into a new video. Among them, Figure 3 (a) in [the figure] shows the first video of a vehicle running from point A to point B. The vehicle included in the first video is the first target. Further, the first target in each frame image of the first video can be detected by using a target contour detection algorithm and composed into the first target sequence in chronological order. Figure 3 (b) in [the figure] shows the second video of a vehicle running from point C to point D. The vehicle included in the second video is the second target. Similarly, the detected second target is composed into the second target sequence. Figure 3 (c) in [the figure] shows the new video formed by combining the first target sequence, the second target sequence, and the background image. It can be seen that Figure 3 the new video generated by the process shown [in the figure] shows a violation event of a collision between vehicles.
[0102] As another example, refer to Figure 4 , which is another process provided by an embodiment of this application for combining the target sequences and the background image extracted from two videos into a new video. Among them, Figure 4 (a) in [the figure] shows the first video of a pedestrian running from point A to point B. The pedestrian included in the first video is the first target, and then the first target sequence can be extracted from the first video. Figure 4 (b) in [the figure] shows the second video of a vehicle running from point C to point D. The vehicle included in the second video is the second target, and the contours of the detected second target are composed into the second target sequence. Figure 4 (c) in [the figure] shows the new video formed by combining the first target sequence, the second target sequence, and the background image. It can be seen that Figure 4 the new video generated by the process shown [in the figure] shows a violation event of a vehicle not yielding to a pedestrian.
[0103] Based on the above scheme, it can be seen that the present application can generate videos of violation incidents based on existing video materials, which can be used for traffic or city video analysis and training of preset models, so as to issue warnings before real violation incidents occur.
[0104] In some possible cases, when the first target sequence, the second target sequence and the background image are combined into a new video, if there is an overlapping area between the first target and the second target in a certain frame of the new video, it is difficult to determine which one is the occluded target based only on the outlines of the two targets in the overlapping area. For example, you can refer to Figure 5 ,from Figure 5 It can be seen that there is an overlapping area between the two targets, but it is not known which target is in front and which target is blocked behind. Based on this, the embodiment of the present application provides a method for determining the blocked target based on the projection area. For details, please refer to the introduction in the above embodiment. In order to more intuitively understand the method for determining the blocked target proposed in this application, for example, please refer to Figure 6 . Figure 6 The projection area of the first target shown in FIG is quadrilateral ABCD, and the projection area of the second target is EFGH. Alternatively, a vertical line can be drawn through point A, and the intersection of the line and the projection area of the second target is determined to be above point A. The second target can be determined to be the obscured target.
[0105] In one possible implementation, when determining the projection area of the first target or the second target on the ground, the spatial coordinates of the target can be obtained through a 3D camera based on a monocular 3D detection algorithm, and then the position coordinates of each point in the target projected on the ground can be determined, thereby determining the projection area of the target.
[0106] In another possible implementation, the projected area of the target can be inferred based on the target's key points and the target's ground contact points. Continuing with the example of the first target, multiple sets of key points and multiple ground contact points can be obtained for the first target. The distances between each set of key points in the multiple sets of key points are used to represent the size of the first target. For example, a set of key points could be the most prominent points of a vehicle's two headlights. The ground contact points are the points where the first target connects to the ground, such as where the wheels touch the ground.
[0107] Furthermore, the size and shape of the projection area of the first target may be determined according to the multiple groups of key points, and the position of the projection area of the first target in the image may be determined according to the multiple ground contact points.
[0108] As an example, see Figure 7 , shows a schematic diagram of the ground projection of the vehicle, in Figure 7In the method, we can first determine the contact points of the four wheels (i.e. the ground contact points of the vehicle), and further, we can determine multiple key points of the vehicle, such as the left and right vertices and the left and right headlight points. Furthermore, we can infer the ground projection points other than the key points based on the contact points. For example, we can make an extension line of the connection line between the contact points of the front and rear wheels on the right side, and make a perpendicular line through the right headlight point, and use the intersection of the extension line and the perpendicular line as the projection point of the right headlight point on the ground. Similarly, we can determine the projection point of the left headlight on the ground, and then we can determine the projection area as Figure 7 The quadrilateral MNPQ in .
[0109] It is important to know that, in addition to the above Figure 7 In addition to the vehicles described above, the ground projection areas of other targets can also be determined using ground contact points and key points. For example, see Figure 8 , showing an example of a bicycle's ground contact points. Figure 9 , showing the ground contact point of the road cone as an example.
[0110] In order to further understand the solution provided by this application, the following is an introduction with reference to specific embodiments. Figure 10 , is a flow chart of a video generation method provided in an embodiment of the present application, specifically including:
[0111] 1001. Acquire a first target sequence from images included in a first video, acquire a second target sequence from images included in a second video, and acquire a background image from images included in a third video.
[0112] The first target sequence is a target sequence obtained by detecting the outline of the first target from the image included in the first video, and the second target sequence is a target sequence obtained by detecting the outline of the second target from the image included in the second video.
[0113] Optionally, the process of obtaining the first target sequence, the second target sequence and the background image may refer to the introduction in the above embodiment, which will not be described in detail here.
[0114] 1002 , combining the first target sequence and the second target sequence including the sub-images and the background image in chronological order to obtain a plurality of new images.
[0115] Optionally, a sub-image included in the first target sequence, a sub-image included in the second target sequence, and the background image may be combined into a new image.
[0116] 1003 : Determine whether an overlapping area exists between the first object and the second object in a first image of the plurality of new images.
[0117] 1004. Determine the occluded object from the second object and the second object according to the projection areas of the first object and the second object.
[0118] 1005. Adjust the first image and do not display the overlapping area of the occluded object in the first image.
[0119] 1006. Combine the adjusted multiple new images into a new video.
[0120] Based on the same concept as the above method, see Figure 11 , a video generation device 1100 provided by an embodiment of the present application. The device 1100 is used to execute each step in the above method. To avoid repetition, it will not be elaborated here. The device 1100 includes: an acquisition unit 1101 and a processing unit 1102.
[0121] The acquisition unit 1101 is used to acquire the background image of the scene, as well as the first video and the second video obtained by shooting the scene;
[0122] The processing unit 1102 is configured to execute:
[0123] Obtain a first object sequence and a second object sequence according to the first video and the second video, where the first object sequence is an image sequence of the area including the first object in the first video, and the second object sequence is an image sequence of the area including the second object in the second video;
[0124] Sequentially extract the target areas in the first object sequence and the second object sequence according to the shooting time sequence, and combine each extracted target area with the background image;
[0125] When there is an overlapping part in the target area of any frame image obtained by combination, correct the display mode of the overlapping part according to the projections of the first object and the second object on the ground included in the any frame image in the set direction;
[0126] Combine the corrected images in the shooting time sequence into the video to be generated.
[0127] In some embodiments, when the processing unit 1102 corrects the display mode of the overlapping part according to the projections of the first object and the second object on the ground included in the any frame image in the set direction, it is specifically used for:
[0128] Determine the occluded object among the first object and the second object according to the projection of the first object and the projection of the second object;
[0129] If the first object is the occluded object, make the second object visible in the overlapping part;
[0130] If the second target is an occluded target, make the first target visible in the overlapping part.
[0131] In some embodiments, when determining the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target, the processing unit 1102 is specifically configured to:
[0132] Determine a line passing through any point in the projection of the first target along a preset direction;
[0133] Determine the intersection point of the line and the second target, and determine whether the intersection point is above the any point in the preset direction;
[0134] If the intersection point is above the any point, determine that the second target is the occluded target;
[0135] If the intersection point is not above the any point, determine that the first target is the occluded target.
[0136] In some embodiments, when the acquisition unit 1101 acquires the background image of the scene, it is specifically configured to:
[0137] Use an image that does not include the first target and the second target in at least one frame of the third video obtained by photographing the scene as the background image; or,
[0138] Generate the background image according to at least one frame of the third video based on a preset background generation algorithm.
[0139] In some embodiments, when the processing unit 1102 determines the projection of the first target, it is specifically configured to:
[0140] Acquire multiple sets of key points and multiple ground contact points included in the first target in any frame of the image through the acquisition unit 1101; the distance between two key points included in each set of key points in the multiple sets of key points is used to characterize the size of the first target;
[0141] Determine the size and shape of the projection of the first target according to the multiple sets of key points, and determine the position of the projection of the first target in any frame of the image according to the multiple ground contact points.
[0142] In some embodiments, the processing unit 1102 is specifically configured to:
[0143] By using a pre - set target contour detection algorithm, perform target contour detection on each frame image of the first video to obtain the region including the first target in each frame image of the first video;
[0144] According to the shooting time of at least one frame image included in the first video, form a first target sequence with the detected regions of the first target.
[0145] Figure 12 FIG. shows a schematic structural diagram of the electronic device 1200 provided in an embodiment of the present application. The electronic device 1200 in the embodiment of the present application may further include a communication interface 1203. The communication interface 1203 is, for example, a network port, and the electronic device can transmit data through the communication interface 1203.
[0146] In the embodiment of the present application, the memory 1202 stores instructions executable by at least one controller 1201. By executing the instructions stored in the memory 1202, at least one controller 1201 can be used to execute each step in the above - mentioned method. For example, the controller 1201 can implement the functions of the acquisition unit 1101 and the processing unit 1102 described above. Figure 11 functions.
[0147] Among them, the controller 1201 is the control center of the electronic device. It can connect various parts of the entire electronic device through various interfaces and lines, and run or execute the instructions stored in the memory 1202 and call the data stored in the memory 1202. Optionally, the controller 1201 may include one or more processing units. The controller 1201 may integrate an application controller and a modulation - demodulation controller. Among them, the application controller mainly processes the operating system and application programs, etc., and the modulation - demodulation controller mainly processes wireless communication. It can be understood that the above - mentioned modulation - demodulation controller may not be integrated into the controller 1201. In some embodiments, the controller 1201 and the memory 1202 can be implemented on the same chip. In some embodiments, they can also be separately implemented on independent chips.
[0148] The controller 1201 can be a general - purpose controller, such as a central controller (English: Central Processing Unit, abbreviated as: CPU), a digital signal controller, an application - specific integrated circuit, a field - programmable gate array, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general - purpose controller can be a micro - controller or any conventional controller, etc. The steps executed by the data statistics platform disclosed in combination with the embodiments of the present application can be directly executed by the hardware controller, or executed by a combination of hardware and software modules in the controller.
[0149] The memory 1202, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. The memory 1202 may include at least one type of storage medium, for example, it may include flash memory, hard disk, multimedia card, card-type memory, random access memory (abbreviation: RAM), static random access memory (abbreviation: SRAM), programmable read-only memory (abbreviation: PROM), read-only memory (abbreviation: ROM), electrically erasable programmable read-only memory (abbreviation: EEPROM), magnetic memory, magnetic disk, optical disk, and so on. The memory 1202 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 1202 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.
[0150] By programming the design of the controller 1201, for example, the code corresponding to the training method of the neural network model introduced in the foregoing embodiments can be solidified into the chip, so that the chip can execute the steps of the foregoing neural network model training method when running. How to program the design of the controller 1201 is a well-known technology to those skilled in the art and will not be elaborated here.
[0151] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to the application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the controller of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the controller of the computer or other programmable data processing device generate means for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0153] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0154] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operational steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0155] Although the preferred embodiments of this application have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of this application.
[0156] Obviously, those skilled in the art can make various changes and variations to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and variations.
Claims
1. A video generation method, characterized in that, The method includes: Obtaining a background image of a scene, a first video and a second video captured of the scene, and obtaining a first target sequence and a second target sequence based on the first video and the second video, where the first target sequence is an image sequence of an area including a first target in the first video, and the second target sequence is an image sequence of an area including a second target in the second video; Sequentially extracting target areas in the first target sequence and the second target sequence in chronological order of shooting, and combining each extracted target area with the background image; When there is an overlapping part in the target area in any frame image obtained by combination, determining the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target; if the first target is the occluded target, setting the second target as visible in the overlapping part; if the second target is the occluded target, setting the first target as visible in the overlapping part; Combining the corrected images in chronological order of shooting into a video to be generated.
2. The method according to claim 1, wherein The determining the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target includes: Determining a line passing through any point in the projection of the first target along a preset direction; Determining the intersection point of the line and the second target, and judging whether the intersection point is above the any point in the preset direction; If the intersection point is above the any point, determining that the second target is the occluded target; If the intersection point is not above the any point, determining that the first target is the occluded target.
3. The method according to any one of claims 1-2, characterized in that The obtaining the background image of the scene includes: Taking an image that does not include the first target and the second target in at least one frame image included in a third video captured of the scene as the background image; or, Generating the background image based on a preset background generation algorithm according to at least one frame image included in the third video.
4. The method according to any one of claims 1-2, characterized in that, The projection of the first target is determined in the following manner: Obtaining multiple sets of key points and multiple ground contact points included in the first target in the any frame image; the distance between two key points included in each set of key points in the multiple sets of key points is used to characterize the size of the first target; Determining the size and shape of the projection of the first target according to the multiple sets of key points, and determining the position of the projection of the first target in the any frame image according to the multiple ground contact points.
5. The method according to any one of claims 1-2, characterized in that, The first target sequence is obtained in the following manner: Performing target contour detection on each frame image of the first video through a preset target contour detection algorithm to obtain an area including the first target in each frame image of the first video; Forming the detected areas of the first target into a first target sequence according to the shooting time of at least one frame image included in the first video.
6. A video generation device, characterized in that, The device includes: An acquisition unit, configured to acquire a background image of a scene, a first video and a second video captured of the scene; A processing unit, configured to execute: Obtain a first target sequence and a second target sequence based on the first video and the second video, where the first target sequence is an image sequence of the region including the first target in the first video, and the second target sequence is an image sequence of the region including the second target in the second video; Extract the target regions in the first target sequence and the second target sequence in order of shooting time, and combine each extracted target region with the background image; When there is an overlapping part in the target regions in any frame image obtained by combination, determine the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target; if the first target is the occluded target, set the second target to be visible in the overlapping part; if the second target is the occluded target, set the first target to be visible in the overlapping part; Combine the corrected images in order of shooting time into the video to be generated.
7. The device according to claim 6, characterized in that When determining the occluded target among the first target and the second target according to the projection of the first target and the projection of the second target, the processing unit is specifically configured to: Determine a line passing through any point in the projection of the first target along a preset direction; Determine the intersection point of the line and the second target, and judge whether the intersection point is above the any point in the preset direction; If the intersection point is above the any point, determine that the second target is the occluded target; If the intersection point is not above the any point, determine that the first target is the occluded target.
8. The device according to any one of claims 6-7, characterized in that, When determining the projection of the first target, the processing unit is specifically configured to: Obtain multiple groups of key points and multiple ground contact points included in the first target in the any frame image through the obtaining unit; the distance between the two key points included in each group of key points in the multiple groups of key points is used to characterize the size of the first target; Determine the size and shape of the projection of the first target according to the multiple groups of key points, and determine the position of the projection of the first target in the any frame image according to the multiple ground contact points.
Citation Information
Patent Citations
Target object position relation analysis method and device, storage medium and electronic equipment
CN113313075A