An image generation method and apparatus
By acquiring the to-process images of each camera in the target scene, segmenting and selecting the image area containing the complete foreground object, and projecting it with the background image to a three-dimensional scene model, the problem of missing moving objects in the image is solved and image quality is improved.
Patent Information
- Application Number
- CN202210665142.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-06-13
AI Technical Summary
In the prior art, during the motion of an object in the target scene, the moving object may cause only a partial image of the moving object to be processed, resulting in the missing object in the image browsed by the user, and the generated image quality is low.
By obtaining the images to be processed corresponding to the cameras in the target scene, performing image segmentation to determine the image area of the foreground object, selecting the image area containing the foreground object with the greatest integrity as the specified foreground image area, generating an image to be projected with the complete foreground object, and projecting the background image and the image to be projected to the target scene.
Ensure that at any moment, each foreground object in the target scene contains completeness in the images collected by the cameras of each set up, thereby improving the integrity of the foreground object in the generated image, avoiding missing situations, and improving image quality.
Smart Images

Figure CN115049539B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and in particular, to an image generation method and apparatus. Background Art
[0002] In related technologies, multiple cameras can be installed at different positions in a target scene, and images of the target scene from different shooting perspectives (which can be referred to as images to be processed) can be captured by the multiple cameras. Then, a three-dimensional scene model of the target scene is obtained, and the images to be processed are mapped to the above three-dimensional scene model. Based on the mapping result, a user can view images of the target scene from different monitoring perspectives.
[0003] However, for a moving object in a target scene, during the movement of the moving object, it may cause that only a partial image of the moving object is included in the image to be processed captured by the installed camera at a certain moment, and further cause the moving object to be missing in the images browsed by the user. That is, in related technologies, the quality of the generated images is low. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an image generation method and apparatus to improve the quality of the generated images.
[0005] The specific technical solutions are as follows:
[0006] In a first aspect, to achieve the above purpose, the embodiments of this application disclose an image generation method, and the method includes:
[0007] Obtain the images to be processed corresponding to each installed camera in the target scene; wherein, the image to be processed corresponding to one installed camera is obtained based on the image of the target scene captured by this installed camera; each installed camera satisfies a preset constraint condition, and the constraint condition is: at any moment, for each foreground object in the target scene, at least one of the images captured by each installed camera includes the complete foreground object;
[0008] Based on image segmentation of each image to be processed, respectively determine the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object;
[0009] Determine, from each foreground image area, the image area with the highest degree of completeness of the included foreground object as the designated foreground image area;
[0010] For each pixel coordinate of the foreground object at a specified monitoring perspective, determine the pixel value of the pixel point corresponding to this pixel coordinate in the designated foreground image area as the pixel value of this pixel coordinate, and obtain the image to be projected of the foreground object at the specified monitoring perspective;
[0011] Obtain an image containing the background in the target scene collected by the installed cameras corresponding to the specified foreground image area as the background image;
[0012] Project the image of the background image at the specified monitoring angle and the image to be projected onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene at the specified monitoring angle.
[0013] Optionally, before obtaining the images to be processed corresponding to each installed camera in the target scene, the method further includes:
[0014] Obtain the current installation parameters of each installed camera;
[0015] Determine whether each installed camera satisfies a preset constraint condition based on the current installation parameters;
[0016] If each installed camera does not satisfy the constraint condition based on the current installation parameters, send a control instruction to the control device so that the control device adjusts the installation parameters of each installed camera according to the control instruction, so that each installed camera satisfies the constraint condition based on the adjusted installation parameters.
[0017] Optionally, for each pixel coordinate of the foreground object at the specified monitoring angle, determining the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate to obtain the image to be projected of the foreground object at the specified monitoring angle includes:
[0018] For each pixel coordinate of the foreground object at the specified monitoring angle, based on the first conversion relationship between the image coordinates in the specified foreground image area and the image coordinates of the imaging at the specified monitoring angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate to obtain the image to be projected of the foreground object at the specified monitoring angle.
[0019] Optionally, before for each pixel coordinate of the foreground object at the specified monitoring angle, determining the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate to obtain the image to be projected of the foreground object at the specified monitoring angle, the method further includes:
[0020] Based on each foreground image area and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determine the three-dimensional pose plane of the foreground object in the target scene;
[0021] For each pixel coordinate of the foreground object at a specified monitoring perspective, determining the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtaining the image to be projected of the foreground object at the specified monitoring perspective, includes:
[0022] For each pixel coordinate of the foreground object at a specified monitoring perspective, based on the third conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the fourth conversion relationship between the image coordinates imaged at the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane, determining the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtaining the image to be projected of the foreground object at the specified monitoring perspective.
[0023] Optionally, the determining the three-dimensional pose plane of the foreground object in the target scene based on each foreground image area and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene includes:
[0024] Determining two mounted cameras corresponding to the image to be processed that contain the foreground object as the specified mounted cameras;
[0025] For each specified mounted camera, determining the detection frame containing the foreground object in the image to be processed corresponding to the specified mounted camera;
[0026] Based on the second conversion relationship between the image coordinates corresponding to the specified mounted camera and the spatial coordinates, determining the three-dimensional point corresponding to the specified point in the detection frame in the target scene;
[0027] Determining the plane containing the optical center of the specified mounted camera and the corresponding three-dimensional point as the reference plane corresponding to the specified mounted camera;
[0028] Determining the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0029] Optionally, the determining the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene includes:
[0030] Taking the intersection line of each reference plane as the rotation axis, and rotating the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0031] Optionally, before determining, for each pixel coordinate of the foreground object at a specified monitoring perspective, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region as the pixel value of the pixel coordinate to obtain the image to be projected of the foreground object at the specified monitoring perspective, the method further includes:
[0032] Determine the edge pixel points of the foreground object in the image to be processed to which the specified foreground image region belongs;
[0033] For each edge pixel point, determine the corresponding pixel coordinate of the edge pixel point at the specified monitoring perspective as the edge pixel coordinate;
[0034] Determine the pixel coordinates included in the region with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object at the specified monitoring perspective.
[0035] Optionally, the determining, based on image segmentation of each image to be processed, of the image region occupied by the foreground object in each image to be processed as the foreground image region corresponding to the foreground object includes:
[0036] Based on image segmentation of each image to be processed, obtain the image regions occupied by each foreground object in each image to be processed as the image regions to be processed;
[0037] Based on the image similarity between the image regions to be processed, determine the image regions to be processed belonging to the same foreground object as the foreground image region corresponding to the foreground object.
[0038] Optionally, the obtaining of the images to be processed corresponding to the cameras installed in the target scene includes:
[0039] Obtain the images of the target scene collected by each installed camera as the initial images;
[0040] Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of each installed camera to obtain each image to be processed.
[0041] Optionally, the obtaining of the image including the background in the target scene collected by the camera corresponding to the specified foreground image region as the background image includes:
[0042] Obtain the image to be filled; wherein, the image to be filled is obtained by deleting the specified foreground image region from the image to be processed to which it belongs;
[0043] Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain an image of the background in the target scene captured by the specified camera; wherein, the preset image is an image of the background in the target scene captured in advance.
[0044] In a second aspect, to achieve the above object, an embodiment of the present application discloses an image generation device, the device includes:
[0045] An image to be processed acquisition module, configured to acquire an image to be processed corresponding to each installed camera in the target scene; wherein, the image to be processed corresponding to one installed camera is obtained based on the image of the target scene captured by this installed camera; each of the installed cameras satisfies a preset constraint condition, and the constraint condition is: at any moment, for each foreground object in the target scene, at least one of the images captured by each installed camera contains the complete foreground object;
[0046] A foreground image area determination module, configured to respectively determine the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object based on image segmentation of each image to be processed;
[0047] A foreground image area selection module, configured to determine, from each foreground image area, the image area with the highest degree of integrity of the contained foreground object as the specified foreground image area;
[0048] A to-be-projected image generation module, configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, to obtain the to-be-projected image of the foreground object at the specified monitoring view angle;
[0049] A background image acquisition module, configured to acquire an image of the background in the target scene captured by the installed camera corresponding to the specified foreground image area as the background image;
[0050] A target image generation module, configured to project the image of the background image at the specified monitoring view angle and the to-be-projected image onto the three-dimensional scene model of the target scene in sequence, to obtain the target image of the target scene at the specified monitoring view angle.
[0051] Optionally, the device further includes:
[0052] An installation parameter acquisition module, configured to acquire the current installation parameters of each installed camera before the image to be processed acquisition module executes acquiring the image to be processed corresponding to each installed camera in the target scene;
[0053] A judgment module, configured to judge whether each erected camera satisfies a preset constraint condition based on the current erection parameters;
[0054] A control instruction sending module, configured to, if each erected camera does not satisfy the constraint condition based on the current erection parameters, send a control instruction to a control device, so that the control device adjusts the erection parameters of each erected camera according to the control instruction, so that each erected camera satisfies the constraint condition based on the adjusted erection parameters.
[0055] Optionally, the projected image generation module is specifically configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a first conversion relationship between the image coordinates in a specified foreground image area and the image coordinates of the image formed at the specified monitoring view angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the projected image of the foreground object at the specified monitoring view angle.
[0056] Optionally, the device further includes:
[0057] A three-dimensional pose plane determination module, configured to, before the projected image generation module executes, for each pixel coordinate of the foreground object at a specified monitoring view angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the projected image of the foreground object at the specified monitoring view angle, execute based on each foreground image area, and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, and determine the three-dimensional pose plane of the foreground object in the target scene;
[0058] The projected image generation module is specifically configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a third conversion relationship between the image coordinates in a specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and a fourth conversion relationship between the image coordinates of the image formed at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the projected image of the foreground object at the specified monitoring view angle.
[0059] Optionally, the three-dimensional pose plane determination module is specifically configured to determine two erected cameras corresponding to the image to be processed that contain the foreground object as the specified erected cameras;
[0060] For each specified erected camera, determine the detection frame containing the foreground object in the image to be processed corresponding to the specified erected camera;
[0061] Based on the second conversion relationship between the image coordinates corresponding to the specified mounted camera and the spatial coordinates, determine the three-dimensional point corresponding to the specified point in the detection frame in the target scene;
[0062] Determine the plane containing the optical center of the specified mounted camera and the corresponding three-dimensional point as the reference plane corresponding to the specified mounted camera;
[0063] Determine the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0064] Optionally, the three-dimensional pose plane determination module is specifically configured to use the intersection line of each reference plane as the rotation axis and rotate the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0065] Optionally, the device further includes:
[0066] An edge pixel point determination module, configured to determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image region before the to-be-projected image generation module executes, for each pixel coordinate of the foreground object in the specified monitoring view, to determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region as the pixel value of the pixel coordinate, so as to obtain the to-be-projected image of the foreground object in the specified monitoring view;
[0067] An edge pixel point mapping module, configured to, for each edge pixel point, determine the corresponding pixel coordinate of the edge pixel point in the specified monitoring view as the edge pixel coordinate;
[0068] A pixel coordinate determination module, configured to determine the pixel coordinates included in the region bounded by the edge pixel coordinates as the pixel coordinates of the foreground object in the specified monitoring view.
[0069] Optionally, the foreground image region determination module is specifically configured to, based on image segmentation of each to-be-processed image, obtain the image regions occupied by each foreground object in each to-be-processed image as the to-be-processed image regions;
[0070] Based on the image similarity between each to-be-processed image region, determine the to-be-processed image regions belonging to the same foreground object as the foreground image region corresponding to the foreground object.
[0071] Optionally, the to-be-processed image acquisition module is specifically configured to acquire the images of the target scene collected by each mounted camera as the initial images;
[0072] Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of the installed cameras, to obtain each image to be processed.
[0073] Optionally, the background image acquisition module is specifically configured to acquire an image to be filled; wherein, the image to be filled is obtained by deleting the specified foreground image area from the image to be processed to which it belongs.
[0074] Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain an image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area; wherein, the preset image is a pre-collected image containing the background in the target scene.
[0075] An embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus;
[0076] The memory is used to store a computer program;
[0077] When the processor is used to execute the program stored on the memory, it implements the steps of the image generation method described in any one of the above.
[0078] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the image generation method described in any one of the above.
[0079] An embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the image generation method described in any one of the above.
[0080] An image generation method provided by an embodiment of the present application obtains to-be-processed images corresponding to each installed camera in a target scene; a to-be-processed image corresponding to an installed camera is obtained based on an image of the target scene collected by the installed camera; each installed camera satisfies that at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera contains the complete foreground object; based on image segmentation of each to-be-processed image, the image regions occupied by the foreground objects in each to-be-processed image are respectively determined as the foreground image regions corresponding to the foreground objects; from each foreground image region, the image region with the highest integrity of the contained foreground object is determined as the specified foreground image region; for each pixel coordinate of the foreground object under a specified monitoring perspective, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region is determined as the pixel value of the pixel coordinate, and a to-be-projected image of the foreground object under the specified monitoring perspective is obtained; an image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image region is obtained as the background image; the image of the background image and the to-be-projected image under the specified monitoring perspective are sequentially projected onto the three-dimensional scene model of the target scene to obtain the target image of the target scene under the specified monitoring perspective.
[0081] Based on the above processing, at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera in the target scene contains the complete foreground object. Correspondingly, the specified foreground image region has the highest integrity of the contained foreground object, that is, the specified foreground image region contains the complete foreground object. Furthermore, the to-be-projected image obtained based on the specified foreground image region also contains the complete foreground object. Mapping the background image and the to-be-projected image under the specified monitoring perspective to the three-dimensional scene model of the target scene can map the to-be-projected image containing the complete foreground object to the three-dimensional scene model, which can, to a certain extent, avoid the situation that the foreground object in the generated image is missing, and further improve the quality of the generated image.
[0082] Of course, it is not necessary for any product or method implementing the present application to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0084] Figure 1a An image of a target scene provided by an embodiment of the present application;
[0085] Figure 1b An image of another target scenario provided by the embodiment of the present application;
[0086] Figure 2 A flowchart of an image generation method provided by the embodiment of the present application;
[0087] Figure 3 A schematic diagram of a camera installation method provided by the embodiment of the present application;
[0088] Figure 4 A schematic diagram of another camera installation method provided by the embodiment of the present application;
[0089] Figure 5 A flowchart of another image generation method provided by the embodiment of the present application;
[0090] Figure 6 A flowchart of another image generation method provided by the embodiment of the present application;
[0091] Figure 7 A comparison diagram of an initial image and an image to be processed provided by the embodiment of the present application;
[0092] Figure 8 A flowchart of another image generation method provided by the embodiment of the present application;
[0093] Figure 9 A schematic diagram of the principle of foreground image region extraction provided by the embodiment of the present application;
[0094] Figure 10 A schematic diagram of the principle of foreground image region matching provided by the embodiment of the present application;
[0095] Figure 11 A flowchart of another image generation method provided by the embodiment of the present application;
[0096] Figure 12 A flowchart of another image generation method provided by the embodiment of the present application;
[0097] Figure 13 A schematic diagram of the principle of mapping a specified foreground image region to a specified monitoring perspective provided by the embodiment of the present application;
[0098] Figure 14 A flowchart of another image generation method provided by the embodiment of the present application;
[0099] Figure 15 A flowchart of another image generation method provided by the embodiment of the present application;
[0100] Figure 16Schematic diagram of the principle for determining a three-dimensional pose plane provided by an embodiment of the present application;
[0101] Figure 17 Another schematic diagram of the principle for mapping a specified foreground image area to a specified monitoring perspective provided by an embodiment of the present application;
[0102] Figure 18 Flowchart of another image generation method provided by an embodiment of the present application;
[0103] Figure 19 Comparison diagram of a to-be-processed image, a to-be-filled image, and a background image provided by an embodiment of the present application;
[0104] Figure 20 Comparison diagram of a three-dimensional scene model and a target image of a target scene from a specified monitoring perspective provided by an embodiment of the present application;
[0105] Figure 21 Comparison diagram of a three-dimensional scene model provided by an embodiment of the present application;
[0106] Figure 22 Flowchart of a method for obtaining a three-dimensional scene model provided by an embodiment of the present application;
[0107] Figure 23 Flowchart of another image generation method provided by an embodiment of the present application;
[0108] Figure 24 Flowchart of another image generation method provided by an embodiment of the present application;
[0109] Figure 25 Comparison diagram of a target image of a target scene from a specified monitoring perspective provided by an embodiment of the present application;
[0110] Figure 26 Structural diagram of an image generation device provided by an embodiment of the present application;
[0111] Figure 27 Structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0112] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0113] In the related art, for a moving object in a target scenario, during the movement of the moving object, it may cause that the image to be processed captured by the erected camera at a certain moment may only contain a partial image of the moving object, thereby resulting in the missing of the moving object in the image browsed by the user. That is, in the related art, the quality of the generated image is low.
[0114] Exemplarily, refer to Figure 1a , Figure 1a , which is an image of a target scenario provided by an embodiment of the present application. The target scenario is a parking lot, and the buildings therein are static objects. Since the positions of the static objects in the target scenario are fixed, the image captured by the erected camera contains a relatively complete image of the static objects. Therefore, this image has a good registration effect on the static objects, and the image obtained by mapping this image to the three-dimensional scene model will basically not be missing.
[0115] Refer to Figure 1b , Figure 1b , which is another image of a target scenario provided by an embodiment of the present application. The target scenario is a factory, where the tables, chairs, ceiling, floor wall columns, and the workbench of the staff are static objects, and the staff are moving objects. Since the positions of the static objects in the target scenario are fixed, the image captured by the erected camera contains a relatively complete image of the static objects. Therefore, this image has a good registration effect on the static objects, and the image obtained by mapping this image to the three-dimensional scene model will basically not be missing.
[0116] However, for a moving object in a target scenario, during the movement of the moving object, it may cause that at a certain moment, the image to be processed captured by the erected camera may only contain a partial image of the moving object, making the moving object missing in the image browsed by the user. That is, in the related art, the quality of the generated image is low.
[0117] To solve the above problems, an embodiment of the present application provides an image generation method, which is applied to an electronic device. The electronic device can obtain the images to be processed corresponding to each erected camera in the target scenario, and process the images to be processed based on the method provided by the embodiment of the present application to obtain a target image of the target scenario from a specified monitoring perspective. Subsequently, the user can view the image of the target scenario from the specified monitoring perspective. For example, for a monitoring video scenario, a monitoring video containing the target scenario from the specified monitoring perspective can be obtained, and the user can view the monitoring video of the target scenario from the specified monitoring perspective through the electronic device to achieve the effect of all-round monitoring.
[0118] Refer to Figure 2 , Figure 2The flowchart of an image generation method provided by an embodiment of this application. This method may include the following steps:
[0119] S201: Obtain the images to be processed corresponding to each installed camera in the target scene.
[0120] Among them, the image to be processed corresponding to one installed camera is obtained based on the image of the target scene collected by this installed camera. Each installed camera satisfies that at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera contains the complete foreground object.
[0121] S202: Based on image segmentation of each image to be processed, respectively determine the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object.
[0122] S203: From each foreground image area, determine the image area with the highest degree of integrity of the contained foreground object as the specified foreground image area.
[0123] S204: For each pixel coordinate of the foreground object at the specified monitoring perspective, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, and obtain the image to be projected of the foreground object at the specified monitoring perspective.
[0124] S205: Obtain the image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image.
[0125] S206: Project the image of the background image at the specified monitoring perspective and the image to be projected onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene at the specified monitoring perspective.
[0126] Based on the image generation method provided by the embodiment of this application, at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera in the target scene contains the complete foreground object. Correspondingly, the degree of integrity of the foreground object contained in the specified foreground image area is the highest, that is, the specified foreground image area contains the complete foreground object. Furthermore, the image to be projected obtained based on the specified foreground image area also contains the complete foreground object. Mapping the background image and the image to be projected at the specified monitoring perspective to the three-dimensional scene model of the target scene can map the image to be projected containing the complete foreground object to the three-dimensional scene model, which can avoid the situation of missing foreground objects in the generated image to a certain extent, and thus improve the quality of the generated image.
[0127] For step S201, the target scenario can be scenarios such as factories and streets. Multiple mounted cameras are installed in the target scenario, and each mounted camera is used to capture images of the target scenario from different shooting perspectives. For the surveillance video scenario, for each mounted camera, the electronic device can obtain the video image captured by the mounted camera at the current moment, and obtain the image to be processed based on the video image.
[0128] Each mounted camera is a mounted camera that matches the specified monitoring perspective. That a mounted camera matches the specified monitoring perspective means that there is an overlapping area between the shooting perspective of the mounted camera and the specified monitoring perspective.
[0129] The specified monitoring perspective can be any monitoring perspective selected by the user when viewing the image of the target scenario. When the specified monitoring perspective changes, for example, when the user switches the monitoring perspective, the electronic device can determine the mounted camera that matches the switched monitoring perspective. Subsequently, based on the method provided in the embodiments of the present application, the image of the target scenario under the switched monitoring perspective can be determined.
[0130] In at least one embodiment of the present application, the images in the video streams captured by the above-mentioned multiple mounted cameras can all be processed in the above manner to obtain the target images, and the multiple processed target images can form a more complete and clear high-quality video stream.
[0131] In one embodiment, when installing multiple mounted cameras in the target scenario, when installing the multiple mounted cameras, the installation parameters of the mounted cameras can be adjusted so that the multiple mounted cameras meet the following constraints:
[0132] At any moment, for each foreground object in the target scenario, at least one of the images captured by the mounted cameras contains the complete foreground object.
[0133] Based on this, the complete image of the foreground object can be obtained, the integrity of the foreground object in the generated target image can be improved, and thus the quality of the generated image can be improved.
[0134] The installation parameters of the mounted camera can include: the installation height, pitch angle, field of view angle of the mounted camera, the height of the foreground object in the target scenario, and the distance between the mounted cameras, etc.
[0135] Exemplarily, see Figure 3 , Figure 3 is a schematic diagram of a camera installation method provided by an embodiment of the present application. Among them, different triangles represent different mounted cameras. The area corresponding to a triangle and a dotted line represents the range of the field of view of the mounted camera, and each mounted camera can be installed in a straight line (i.e., Figure 3 the upper straight line in
[0136] Each installed camera includes an installed camera facing left and an installed camera facing right. When the specified monitoring view includes the left area in the target scene, the installed camera facing left is enabled to acquire an image; when the specified monitoring view includes the right area in the target scene, the installed camera facing right is enabled to acquire an image.
[0137] The installation height of the installed camera is 2 to 6 times the height of the foreground object in the target scene, so as to ensure that the installed camera can capture a complete and clear image of the foreground object in the target scene. When the installed camera is installed on the ceiling or support frame in the target scene, the installation height of the installed camera can be determined by technicians through manual measurement. Since the heights of different foreground objects are different, the height of the foreground object can be set according to actual needs. For example, when the foreground object is a person, the height of the foreground object can be set to 2 meters, but it is not limited to this.
[0138] The initial pose of the installed camera (for example, the pitch angle of the camera) cannot exceed the adjustable range of the installed camera itself. Moreover, the field of view of the installed camera should cover the area it shoots in the target scene, so that the installed camera can capture a complete image of the foreground object in the target scene and there is no dead angle that cannot be photographed.
[0139] For two adjacent installed cameras facing the same direction, the overlapping area of the fields of view of the two installed cameras accounts for 10% to 30% of the field of view of a single installed camera, and the included angle between the optical axes of the two installed cameras is small. For example, the included angle between the optical axes of the two installed cameras is less than 30 degrees.
[0140] See Figure 4 , Figure 4 In, C1 and C2 represent two adjacent installed cameras facing the same direction. The installation heights, field of view angles and pitch angles of the two installed cameras are the same. The distance between the two installed cameras needs to satisfy the following formula:
[0141] b≤(H - h)cot(α - θ)-Hcot(α + θ)(1)
[0142] b represents the distance between the two installed cameras, H represents the installation height of the installed camera C1, h represents the height of the foreground object in the target scene, 2θ represents the field of view angle of the installed camera C1, and α represents the pitch angle of the installed camera C1.
[0143] Installing in the above manner can ensure that the foreground object moving in the target scene is captured by at least one installed camera with a complete image at each moment. Furthermore, the integrity of the foreground object in the generated target image is also relatively high, which can further improve the quality of the generated image.
[0144] In one embodiment, based on Figure 2 and referring to Figure 5 , before step S201, the method may further include the following steps:
[0145] S207: Obtain the current installation parameters of each installed camera.
[0146] S208: Determine whether each installed camera satisfies a preset constraint condition based on the current installation parameters.
[0147] The constraint condition is that at any moment, for each foreground object in the target scene, at least one of the images captured by each installed camera contains the complete foreground object.
[0148] S209: If each installed camera does not satisfy the constraint condition based on the current installation parameters, send a control instruction to the control device so that the control device adjusts the installation parameters of each installed camera according to the control instruction, so that each installed camera satisfies the constraint condition based on the adjusted installation parameters.
[0149] The installed cameras can be installed in the target scene through the control device, and the control device can communicate with the electronic device.
[0150] The electronic device can obtain the current installation parameters of each installed camera and determine whether each installed camera satisfies a preset constraint condition based on the current installation parameters. For example, for two adjacent installed cameras facing the same direction, determine whether the installation height, pitch angle, field of view angle of the two installed cameras, the height of the foreground object in the target scene, and the distance between the two installed cameras satisfy the above formula (1). When the installation parameters of the installed cameras satisfy the above formula (1), it can be determined that each installed camera satisfies the constraint condition.
[0151] If each installed camera satisfies the constraint condition based on the current installation parameters of each installed camera, the electronic device can obtain the images of the target scene captured by each installed camera (i.e., the initial images), and obtain the image to be processed based on the obtained initial images.
[0152] If each installed camera does not satisfy the constraint condition based on the current installation parameters of each installed camera, the electronic device can send a control instruction to the control device. When the control device receives the control instruction, it can adjust the installation parameters of the installed camera. For example, the control device can control the installed camera to move to adjust the installation height of the installed camera, or the control device can control the installed camera to rotate to adjust the pitch angle of the installed camera, etc., so that each installed camera satisfies the constraint condition based on the adjusted installation parameters of each installed camera.
[0153] When each installed camera meets the constraint conditions, the electronic device can obtain the images of the target scene collected by each installed camera (i.e., the initial images), and based on the obtained initial images, obtain the images to be processed.
[0154] In one implementation, for each installed camera, the electronic device can directly obtain the initial image of the target scene collected by the installed camera as the image to be processed corresponding to the installed camera.
[0155] In another implementation, for each installed camera, the initial image captured by the installed camera may be distorted, and there may be color differences between the initial images captured by each installed camera. To improve the quality of the target image of the generated target scene from a specified monitoring perspective, the electronic device can perform image correction on the initial images captured by the installed cameras to obtain the images to be processed.
[0156] Correspondingly, on the basis of Figure 2 , referring to Figure 6 , step S201 may include the following steps:
[0157] S2011: Obtain the images of the target scene collected by each installed camera as the initial images.
[0158] S2012: Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of each installed camera to obtain each image to be processed.
[0159] The electronic device can obtain the initial images of the target scene collected by each installed camera. For each pixel point in the initial image, calculate the product of the pixel value of the pixel point and the color correction matrix (English: Color Correction Matrix, abbreviated: CCM) to obtain the corrected pixel value of the pixel point, and thus can obtain the image after color difference correction (i.e., the image to be processed).
[0160] Exemplarily, the electronic device can perform color difference correction on the initial image according to the following formula:
[0161] P 2 = C × P 1 (2)
[0162] C represents the color correction matrix, P 2 represents the matrix containing the pixel values of each pixel point in the initial image, and P 1 represents the matrix containing the pixel values of each pixel point in the image to be processed.
[0163] In one implementation, P can be P 1 or P 2. R represents the pixel value of the red channel of the pixel in the corresponding image, G represents the pixel value of the green channel of the pixel in the corresponding image, and B represents the pixel value of the blue channel of the pixel in the corresponding image.
[0164] r1 to r3 represent red compensation values, g1 to g3 represent green compensation values, b1 to b3 represent blue compensation values, and c1 to c3 represent preset color compensation coefficients.
[0165] In another implementation, P can be P 1 or P 2 . R represents the pixel value of the red channel of the pixel in the corresponding image, G represents the pixel value of the green channel of the pixel in the corresponding image, and B represents the pixel value of the blue channel of the pixel in the corresponding image.
[0166] r1 to r3 represent red compensation values, g1 to g3 represent green compensation values, and b1 to b3 represent blue compensation values.
[0167] The values in the color difference correction matrix can be set by technicians according to requirements. For example, when it is necessary to make the image tend to be red, r1 to r3 can be set to larger values. Or, when it is necessary to make the image tend to be green, g1 to g3 can be set to larger values.
[0168] Or, the electronic device can pre-obtain the pixel values of the images captured by any two installed cameras, and calculate the values in the color difference correction matrix based on the pixel values of the images captured by the two installed cameras.
[0169] For each installed camera, if the installed camera is a pinhole camera, the distortion coefficients of the corresponding distortion type of the installed camera include: K1, K2, P1, and P2. In the corresponding distortion type of the pinhole camera, K1 and K2 represent radial distortion, and radial distortion is generated during the conversion of the camera coordinate system of the installed camera to the image coordinate system of the initial image. P1 and P2 represent tangential distortion, and tangential distortion is caused by the lens of the installed camera not being completely parallel to the initial image. If the installed camera is a fish-eye camera, the distortion coefficients of the corresponding distortion type of the installed camera include: K1, K2, K3, and K4.
[0170] For each pixel in the initial image, the electronic device can calculate the corrected coordinates of the pixel according to the distortion coefficients of the corresponding distortion type of the installed camera and the coordinates of the pixel in the initial image, and thus obtain the image to be processed after distortion correction.
[0171] See Figure 7 , Figure 7The image on the left is the initial image captured by the mounted camera. The floor in this initial image is distorted. The image on the right is the image to be processed obtained after distortion correction, and there is no distortion in the floor of the image to be processed.
[0172] Based on the above processing, the electronic device performs distortion correction and / or chromatic aberration correction on the initial image captured by the mounted camera, which can, to a certain extent, avoid distortion or chromatic aberration in the finally generated target image, and can further improve the quality of the generated image.
[0173] Regarding step S202, the foreground object is an object that may move in the target scene. For example, the foreground object can be a vehicle, a pedestrian, etc. The electronic device can determine the object belonging to the preset foreground object type in the image to be processed as the foreground object. The preset foreground object type can include vehicles, pedestrians, etc.
[0174] For the surveillance video scene, after obtaining the image to be processed based on the initial image captured by the mounted camera at the current moment, the electronic device can determine the position of each object in the image to be processed obtained at the current moment (which can be called the first position), and the position of each object in the image to be processed obtained at the previous moment (which can be called the second position), and determine the object with different first and second positions. The determined object is the moving object in the current target scene and is used as the foreground object.
[0175] In one embodiment, on the basis of Figure 2 referring to Figure 8 , step S202 may include the following steps:
[0176] S2021: Based on image segmentation of each image to be processed, obtain the image area occupied by each foreground object in each image to be processed as the image area to be processed.
[0177] S2022: Based on the image similarity between the image areas to be processed, determine the image areas to be processed belonging to the same foreground object as the foreground image area corresponding to the foreground object.
[0178] For each image to be processed, the electronic device can perform image segmentation on the image to be processed based on the target detection algorithm, obtain the image area occupied by each foreground object in the image to be processed, and extract and determine the image area to obtain the image area to be processed.
[0179] The object detection algorithm can be Mask-RCNN (Mask Regions with Convolutional Neural Network), or the object detection algorithm can also be YOLOv3 (You only look once-v3, an end-to-end object detection algorithm based on deep learning), but it is not limited thereto.
[0180] Exemplarily, refer to Figure 9 , Figure 9 In the left image in, the image to be processed is shown, and the image on the right represents the region of the image to be processed extracted from the image to be processed. The pedestrian in the image to be processed is the foreground object included in the image to be processed. In the image to be processed, for each foreground object, the black rectangle containing the foreground object is the detection frame of the foreground object, and this rectangle can be the minimum bounding rectangle of the foreground object. The electronic device can perform instance segmentation according to the edge of the foreground object to obtain Figure 9 3 regions of the image to be processed shown in the right image in.
[0181] Then, the electronic device can extract the image features of each region of the image to be processed, and based on the similarity between the image features of every two regions of the image to be processed, perform clustering processing on each region of the image to be processed to obtain the regions of the image to be processed belonging to the same foreground object as the foreground image region corresponding to the foreground object.
[0182] For example, the electronic device can calculate the similarity of the image features of every two regions of the image to be processed to obtain a similarity matrix containing each similarity. Then, the electronic device can perform decomposition processing on the similarity matrix based on the RNMF (Robust Nonnegative Matrix Factorization) algorithm to obtain the regions of the image to be processed belonging to the same foreground object as the foreground image region corresponding to the foreground object.
[0183] Exemplarily, refer to Figure 10 , Figure 10 The images shown are the images to be processed obtained based on different mounted cameras. Based on Figure 10 the left image in, regions of the image to be processed containing the foreground object represented by ID1 (which can be called image region 1) and regions of the image to be processed containing the foreground object represented by ID2 (which can be called image region 2) can be obtained. Based on Figure 10 the right image in, regions of the image to be processed containing the foreground object represented by ID1 (which can be called image region 3) and regions of the image to be processed containing the foreground object represented by ID2 (which can be called image region 4) can be obtained.
[0184] Furthermore, based on the image similarity between the four image regions to be processed, the matching relationships of different instances can be obtained. The matching relationships include: Image region 1 and image region 3 are the foreground image regions of the foreground object represented by ID1, and image region 2 and image region 4 are the foreground image regions of the foreground object represented by ID2.
[0185] Based on the above processing, the foreground image regions containing the foreground object can be extracted from the image to be processed. Then, the specified foreground image regions can be determined from each foreground image region. The specified foreground image region has the highest integrity of the foreground object it contains, that is, the specified foreground image region contains the complete foreground object. Furthermore, the projected image to be obtained based on the specified foreground image region also contains the complete foreground object. Mapping the background image and the projected image to be obtained under the specified monitoring view angle to the three-dimensional scene model of the target scene can map the projected image containing the complete foreground object to the three-dimensional scene model, which can, to a certain extent, avoid the situation where the foreground object in the generated image is missing. Furthermore, the quality of the generated image can be improved.
[0186] Regarding step S203, since there are multiple images to be processed obtained by the electronic device, there may also be multiple foreground image regions of the foreground object.
[0187] Since the foreground object is a moving object, for each foreground object, during the movement of the foreground object, it may move from the shooting view angle of one installed camera to the shooting view angle of another installed camera. Then, the integrity of the foreground object in each foreground image region obtained based on different installed cameras may be different. To avoid, to a certain extent, the foreground object in the generated target image being missing and improve the quality of the generated target image, the electronic device can determine, from each foreground image region of the foreground object, the image region with the highest integrity of the foreground object it contains as the specified foreground image region of the foreground object.
[0188] Regarding step S204, in one implementation, for each foreground object, the electronic device can determine the coordinates of each pixel point (which can be called the first pixel point) in the specified foreground image region of the foreground object.
[0189] For each first pixel point, the electronic device can perform coordinate transformation on the first pixel point based on the first transformation relationship between the image coordinates in the specified foreground image and the image coordinates of the image formed by the specified monitoring view angle, to obtain the pixel coordinates (which can be called the first pixel coordinates) corresponding to the first pixel point under the specified monitoring view angle.
[0190] The electronic device can use the pixel value of the first pixel point as the pixel value of the corresponding first pixel coordinates, and the projected image to be obtained of the foreground object under the specified monitoring view angle can be obtained.
[0191] For each foreground object, the size of the image area imaged from the specified monitoring perspective may be different from the specified foreground image area of the foreground object. For example, if the image area imaged from the specified monitoring perspective is larger than the specified foreground image area of the foreground object, directly mapping the specified foreground image area of the foreground object to the specified monitoring perspective may cause some pixel coordinates in the image area imaged from the specified monitoring perspective to not be able to match corresponding pixel points in the specified foreground image area, and thus the pixel values of such pixel coordinates cannot be determined, resulting in pixel missing problems in the generated target image.
[0192] To avoid, to a certain extent, the problem that some pixel coordinates in the image area imaged from the specified monitoring perspective cannot match corresponding pixel points in the specified foreground image area, resulting in pixel missing problems in the generated target image and improving the quality of the generated target image, based on Figure 2 and referring to Figure 11 before step S204, the method may further include the following steps:
[0193] S211: Determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area.
[0194] S212: For each edge pixel point, determine the corresponding pixel coordinates of the edge pixel point in the specified monitoring perspective as the edge pixel coordinates.
[0195] S213: Determine the pixel coordinates included in the area with the edge pixel coordinates as the edge as the pixel coordinates of the foreground object in the specified monitoring perspective.
[0196] For each foreground object, the electronic device can determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area, and the determined edge pixel points can represent the contour of the foreground object.
[0197] For each edge pixel point, the electronic device can perform coordinate transformation on the edge pixel point based on the first transformation relationship between the image coordinates in the specified foreground image area and the image coordinates of the image imaged from the specified monitoring perspective to obtain the corresponding pixel coordinates of the edge pixel point in the specified monitoring perspective, that is, obtain the edge pixel coordinates of the edge pixel point in the specified monitoring perspective, and the area composed of the determined edge pixel coordinates can represent the contour of the foreground object in the specified monitoring perspective.
[0198] Furthermore, the electronic device can determine the area with the edge pixel coordinates as the edge. This area is the area occupied by the foreground object in the specified monitoring perspective, and the pixel coordinates within this area are the pixel coordinates of the foreground object in the specified monitoring perspective.
[0199] In another implementation, based on Figure 2 , referring to Figure 12 , step S204 may include the following steps:
[0200] S2041: For each pixel coordinate of the foreground object under the specified monitoring perspective, based on the first conversion relationship between the image coordinates in the specified foreground image region and the image coordinates formed by the specified monitoring perspective, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image region as the pixel value of this pixel coordinate, and obtain the image to be projected of the foreground object under the specified monitoring perspective.
[0201] For each foreground object, the electronic device may obtain each pixel coordinate (which may be referred to as the second pixel coordinate) of the foreground object under the specified monitoring perspective, and each second pixel coordinate is the pixel coordinate in the region with the edge pixel coordinate as the edge.
[0202] For each second pixel coordinate, the electronic device may perform coordinate conversion on this second pixel coordinate based on the first conversion relationship between the image coordinates in the specified foreground image and the image coordinates formed by the specified monitoring perspective, and obtain the coordinate of the pixel point (which may be referred to as the second pixel point) corresponding to this second pixel coordinate in the specified foreground image region of this foreground object.
[0203] Furthermore, the electronic device may use the pixel value of this second pixel point as the pixel value of the corresponding second pixel coordinate, and may obtain the image to be projected of the foreground object under the specified monitoring perspective.
[0204] Exemplarily, referring to Figure 13 , the camera set up for collecting the background image is the set-up camera, the monitoring camera is a virtual camera, and the perspective represented by the monitoring camera is the specified monitoring perspective. The internal parameter of the monitoring camera is denoted as K, the external parameter is denoted as [R|t], the internal parameter of the set-up camera is denoted as K′, and the external parameter is denoted as [R′|t′].
[0205] For each second pixel coordinate under the specified monitoring perspective, this second pixel coordinate is denoted as g 2 representing the spatial depth value of this second pixel coordinate, the electronic device may, based on the internal parameter, external parameter of the monitoring camera and the following formula, obtain the coordinate of the corresponding three-dimensional point (the model point represented by the five-pointed star in Figure 13 ) in the target scene, denoted as
[0206]
[0207] Then, the electronic device can obtain the coordinates of the second pixel point corresponding to the three-dimensional point in the background image based on the internal parameters, external parameters of the erected camera, and the following formula, denoted as g′ 2 which represents the spatial depth value of the second pixel point.
[0208]
[0209] Based on the above processing, the integrity of the foreground object included in the specified foreground image area is the largest, that is, the specified foreground image area contains the complete foreground object. Furthermore, the to-be-projected image obtained based on the specified foreground image area also contains the complete foreground object. Subsequently, mapping the background image and the to-be-projected image under the specified monitoring view angle to the three-dimensional scene model of the target scene can map the to-be-projected image containing the complete foreground object to the three-dimensional scene model, which can avoid the situation of the foreground object in the generated image being missing to a certain extent, and thus improve the quality of the generated image.
[0210] In one embodiment, on the basis of Figure 2 refer to Figure 14 , before step S204, the method may further include the following steps:
[0211] S210: Based on each foreground image area and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determine the three-dimensional pose plane of the foreground object in the target scene.
[0212] Correspondingly, step S204 may include the following steps:
[0213] S2042: For each pixel coordinate of the foreground object under the specified monitoring view angle, based on the third conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the fourth conversion relationship between the image coordinates of the image captured under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the to-be-projected image of the foreground object under the specified monitoring view angle.
[0214] In the related art, only static objects (such as buildings) in the target scene are modeled, and moving foreground objects (such as pedestrians) in the target scene are not modeled. When mapping the image to be processed to the three-dimensional scene model, the part corresponding to the foreground object in the image to be processed will be mapped to the model part corresponding to the background of the foreground object in the target scene in the three-dimensional scene model, resulting in distortion of the foreground object in the image browsed by the user.
[0215] To avoid distortion of foreground objects in the target image of the generated target scene from a specified monitoring perspective, for each foreground object, the electronic device can determine the three-dimensional pose plane of the foreground object in the target scene, and subsequently, through the three-dimensional pose plane of the foreground object, map the foreground object to the three-dimensional scene model of the target scene.
[0216] In one embodiment, based on 12, referring to Figure 15 , step S210 may include the following steps:
[0217] S2101: Determine two mounted cameras corresponding to the to-be-processed image containing the foreground object as the specified mounted cameras.
[0218] S2102: For each specified mounted camera, determine the detection box containing the foreground object in the to-be-processed image corresponding to the specified mounted camera.
[0219] S2103: Based on the second conversion relationship between the image coordinates and the spatial coordinates corresponding to the specified mounted camera, determine the three-dimensional points corresponding to the specified points in the detection box in the target scene.
[0220] S2104: Determine the plane containing the optical center of the specified mounted camera and the corresponding three-dimensional points as the reference plane corresponding to the specified mounted camera.
[0221] S2105: Determine the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0222] For each foreground object, the to-be-processed image corresponding to the mounted camera containing the foreground object means that the mounted camera can capture the foreground object. The electronic device can determine two mounted cameras corresponding to the to-be-processed image containing the foreground object as the specified mounted cameras. For example, the electronic device can determine two adjacent mounted cameras that can capture the foreground object as the specified mounted cameras.
[0223] For each specified mounted camera, the electronic device can determine the detection box containing the foreground object in the to-be-processed image corresponding to the specified mounted camera. For example, for Figure 9 the embodiment of, the black rectangle containing a foreground object is the detection box of the foreground object. The specified points of the detection box may include: two points (for example, vertices) on the lower edge of the detection box. The lower edge of the detection box represents the position where the foreground object contacts the ground in the target scene. For example, if the foreground object is a pedestrian, the lower edge of the detection box represents the position where the pedestrian's feet are in the target scene.
[0224] Then, the electronic device can perform coordinate transformation on the specified points in the detection frame according to the second transformation relationship between the image coordinates corresponding to the specified mounted camera and the spatial coordinates of the target scene, so as to obtain the coordinates of the corresponding three-dimensional points in the target scene. The electronic device can determine a plane (i.e., the reference plane) that includes the determined three-dimensional points and the optical center of the specified mounted camera. The image coordinates corresponding to the specified mounted camera are the image coordinates of the image to be processed obtained based on the specified mounted camera.
[0225] When two specified mounted cameras simultaneously capture a point in space, according to the principle of multi-view geometry, the coordinates of the point in space can be accurately restored. The two reference planes extend along the directions of the perspectives of the two specified mounted cameras, and the intersection line of the two reference planes can be obtained. This intersection line represents the position where the foreground object contacts the ground in the target scene. For example, if the foreground object is a pedestrian, the lower edge of the detection frame represents the position where the pedestrian's feet are located in the target scene. The electronic device can determine a plane that passes through the intersection line of each reference plane and forms a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0226] The specified angle can be any angle between 0 degrees and 90 degrees, and the specific value of the specified angle can be determined based on actual needs. For example, when the foreground object is a pedestrian, the pedestrian usually stands vertically on the ground, so the specified angle can be 90 degrees.
[0227] In one embodiment, step S2105 may include the following steps: taking the intersection line of each reference plane as the rotation axis, and rotating the initial object plane that passes through the rotation axis and is parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0228] Exemplarily, referring to Figure 16 , Figure 16 Camera 1 and Camera 2 in are the specified mounted cameras. The optical center of Camera 1 and the two vertices of the lower edge of the detection frame corresponding to Camera 1 form a reference plane, and the optical center of Camera 2 and the two vertices of the lower edge of the detection frame corresponding to Camera 2 form another reference plane. The two points in the intersection line of each reference plane (i.e., the rotation axis of the object plane) are represented as A(x1, y1, z1) and B(x2, y2, z2). Based on the coordinates of point A and point B, the straight-line equation of this intersection line (i.e., AB) can be obtained as:
[0229] Based on this straight-line equation, the coordinates of any point in this intersection line can be obtained, denoted as M(x0, y0, z0), and the direction vector of this intersection line, denoted as S = [m, n, p].
[0230] Based on the direction vector of the intersection line and the coordinates of point M, the plane equation of the initial object plane passing through the intersection line and parallel to the horizontal plane can be obtained as: ax + by + c = 0, where a = n, b = -m, c = -(nx 0 -my 0 ).
[0231] The plane equation of the three-dimensional pose plane of the foreground object can be expressed as: a'x + b'y + c'z + d' = 0.
[0232] Among them, R(s(θ)) represents the Rodriguez transformation of rotating a specified angle around the rotation axis, θ represents the specified angle, and S represents the straight-line equation of the rotation axis, represents the zero vector.
[0233] For each foreground object, after obtaining the three-dimensional pose plane of the foreground object, the electronic device can map the foreground image area of the foreground object to the specified monitoring perspective through the three-dimensional pose plane of the foreground object to obtain the projected image to be projected of the foreground object under the specified monitoring perspective.
[0234] In one implementation, for each pixel coordinate (which can be called the third pixel coordinate) of the foreground object under the specified monitoring perspective, the electronic device can perform coordinate transformation on the third pixel coordinate based on the fourth conversion relationship between the image coordinates imaged under the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane to obtain the coordinates of the three-dimensional point corresponding to the third pixel coordinate in the three-dimensional pose plane of the foreground object.
[0235] Then, the electronic device can perform coordinate transformation on the determined three-dimensional point based on the third conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane to obtain the pixel point (which can be called the third pixel point) corresponding to the three-dimensional point in the specified foreground image area, that is, the third pixel point corresponding to the third pixel coordinate in the specified foreground image. The electronic device can use the pixel value of the determined third pixel point as the pixel value of the third pixel coordinate, and the projected image to be projected of the foreground object under the specified monitoring perspective can be obtained.
[0236] Exemplarily, refer to Figure 17 , the estimated object plane is the three-dimensional pose plane of the foreground object. The erected camera is the erected camera corresponding to the image to be processed to which the specified foreground image area of the foreground object belongs. The monitoring camera is a virtual camera, and the perspective represented by the monitoring camera is the specified monitoring perspective.
[0237] For each third pixel coordinate of the foreground object under the specified monitoring perspective, the third pixel coordinate can be denoted as The spatial depth value of the third pixel coordinate is denoted as g1 , the spatial depth value represents the distance between the monitoring camera and the three-dimensional point corresponding to the third pixel coordinate in the target scene.
[0238] The internal parameters of the monitoring camera are denoted as K, the external parameters are denoted as [R|t], the internal parameters of the installed camera are denoted as K′, and the external parameters are denoted as [R′|t′]. The electronic device can obtain the three-dimensional point ( Figure 17 the point represented by the pentagram in
[0239]
[0240]
[0241] corresponding to the third pixel coordinate in the three-dimensional pose plane of the foreground object) based on the internal parameters, external parameters of the monitoring camera and the following formula, denoted as The spatial depth value of the third pixel point is denoted as g′ 1 , the spatial depth value represents the distance between the installed camera and the three-dimensional point corresponding to the third pixel point in the target scene.
[0242]
[0243] The electronic device can obtain the third pixel point in the specified foreground image area of the pixel value, as the third pixel coordinate under the specified monitoring view of the pixel value, and can obtain the image to be projected of the foreground object under the specified monitoring view.
[0244] Based on the above processing, the three-dimensional pose plane of the foreground object can be determined. Furthermore, the specified foreground image area can be projected onto the specified monitoring view through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring view. Furthermore, the background image and the image to be projected under the specified monitoring view are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent, and can avoid the distortion of the foreground object in the generated image to a certain extent. Furthermore, the quality of the generated image is improved.
[0245] In addition, in order to avoid, to a certain extent, that some pixel coordinates in the image area imaged under the specified monitoring view cannot match the corresponding pixel points in the specified foreground image area, resulting in pixel missing problems in the generated target image, and improve the quality of the generated target image.
[0246] For each foreground object, the electronic device can determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area. For each edge pixel point, the electronic device can perform coordinate transformation on the edge pixel point based on the third transformation relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, so as to obtain the coordinates of the three-dimensional point corresponding to the edge pixel point in the three-dimensional pose plane of the foreground object. Then, the electronic device can perform coordinate transformation on the determined three-dimensional point based on the fourth transformation relationship between the image coordinates imaged under the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane, so as to obtain the pixel coordinates corresponding to the three-dimensional point under the specified monitoring perspective, that is, obtain the edge pixel coordinates of the edge pixel point under the specified monitoring perspective. The area composed of the determined edge pixel coordinates can represent the contour of the foreground object under the specified monitoring perspective.
[0247] Furthermore, the electronic device can determine the area with the edge pixel coordinates as the edge. This area is the area occupied by the foreground object under the specified monitoring perspective, and the pixel coordinates within this area are the pixel coordinates of the foreground object under the specified monitoring perspective.
[0248] For step S205, in one implementation, the electronic device can obtain the image containing the background in the target scene pre-collected by the erected camera corresponding to the specified foreground image area as the background image.
[0249] Since the pre-collected image and the specified foreground image area are not collected at the same time, and the illumination in the target scene at different times is different, there may be a color difference between the pre-collected image and the specified foreground image area. Directly obtaining the pre-collected image as the background image may cause a color difference between the background and the foreground object in the finally generated target image.
[0250] In another implementation, in order to avoid, to a certain extent, the color difference between the background and the foreground object in the finally generated target image and improve the quality of the generated target image, on the Figure 2 basis, referring to Figure 18 , step S205 may include the following steps:
[0251] S2051: Obtain the image to be filled.
[0252] The image to be filled is obtained by deleting the specified foreground image area from the to-be-processed image to which it belongs.
[0253] S2052: Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, so as to obtain the image containing the background in the target scene pre-collected by the erected camera corresponding to the specified foreground image area.
[0254] Among them, the preset image is an image that is pre-acquired and contains the background in the target scene.
[0255] The electronic device can determine the image after deleting the specified foreground image from the image to be processed as the image to be filled, and the position corresponding to the specified foreground image area in the image to be filled is the area to be filled. The specified foreground image area and the image to be filled belong to the same image to be processed, and there is no color difference caused by different illuminations between the specified foreground image area and the image to be filled.
[0256] Then, the electronic device can obtain the preset image that is pre-acquired and contains the background in the target scene, and determine the part corresponding to the specified foreground image area in the preset image. The electronic device can fill the part corresponding to the specified foreground image area in the preset image into the area to be filled in the image to be filled based on the Poisson filling algorithm, and can obtain the background image. There is no color difference caused by different illuminations between the obtained background image and the specified foreground image area, which can avoid the color difference between the background and the foreground object in the finally generated target image to a certain extent and improve the quality of the generated target image.
[0257] See Figure 19 , Figure 19 In the figure, the left image is the image to be processed, the middle image is the image to be filled obtained after extracting the specified foreground image area, and the right image is the background image obtained by filling.
[0258] After the electronic device extracts the specified foreground image area from the left image to be processed in Figure 19 , it obtains the image to be filled shown in the middle image in Figure 19 , where the black area is the area to be filled. The electronic device fills the area to be filled in the image to be filled based on the preset image and the Poisson filling algorithm, and obtains the background image shown in the right image in Figure 19 .
[0259] After obtaining the background image, the electronic device can also map the background image to the specified monitoring perspective to obtain the image of the background in the target scene under the specified monitoring perspective.
[0260] In one implementation, for each pixel coordinate in the image region imaged from a specified monitoring perspective, the electronic device can determine the three-dimensional point corresponding to the pixel coordinate in the target scene based on the conversion relationship between the image coordinates imaged from the specified monitoring perspective and the spatial coordinates of the target scene. Then, the electronic device can determine the pixel point corresponding to the three-dimensional point in the background image based on the conversion relationship between the image coordinates of the background image and the spatial coordinates of the target scene, and obtain the pixel point corresponding to the pixel coordinate in the background image. The electronic device can use the pixel value of the pixel point corresponding to the pixel coordinate in the background image as the pixel value of the pixel coordinate to obtain the image of the background image from the specified monitoring perspective.
[0261] For step S206, the electronic device can first project the image of the background image from the specified monitoring perspective onto the three-dimensional scene model of the target scene, and then project the images to be projected of each foreground object onto the three-dimensional scene model to obtain the target image of the target scene from the specified monitoring perspective. Projecting the background image onto the three-dimensional scene model first can, to a certain extent, avoid the problem that the background image obscures the image to be projected, resulting in misalignment and missing of foreground objects in the generated target image, and improve the quality of the generated target image.
[0262] In one implementation, the electronic device can use the background image as a texture and add the background image to the corresponding position in the three-dimensional scene model in the way of texture mapping. Then, the electronic device can use the images to be projected as textures and add the images to be projected to the corresponding positions in the three-dimensional scene model in the way of texture mapping to obtain the target image of the target scene from the specified monitoring perspective.
[0263] In addition, after obtaining the target image of the target scene from the specified monitoring perspective, the electronic device can also render according to the target image to display the target image, and the user can view the target image of the target scene from the specified monitoring perspective.
[0264] Exemplarily, see Figure 20 , Figure 20 In the left image in, the three-dimensional scene model of the target scene is shown. In the right image, the left part inside the detection box is the background part of the three-dimensional model of the target scene, and the right detection box is the target image of the target scene from the specified monitoring perspective. It can be seen that there are no problems such as distortion, missing, and misalignment of the moving foreground objects in the target image obtained by the image generation method provided by the embodiments of the present application, that is, the quality of the generated image can be improved.
[0265] In one embodiment, the electronic device can also obtain a three-dimensional scene model of the target scene. The three-dimensional scene model can be a file in formats such as osgb format, s3c format, max format, fbx format, and obj format. The three-dimensional scene model contains data such as point cloud, mesh, and texture. The point cloud represents the three-dimensional points corresponding to each position in the target scene, and each three-dimensional point can be represented by three-dimensional coordinates. Based on the spatial relationship between the positions corresponding to the respective three-dimensional points in the target scene, the respective three-dimensional points are connected to obtain a plurality of meshes formed by the respective three-dimensional points. Based on the material of the objects in the target scene, texture mapping is performed on each of the meshes formed by the respective three-dimensional points to obtain the three-dimensional scene model of the target scene.
[0266] See Figure 21 , Figure 21 In the figure on the left, the image is a three-dimensional point cloud model, the image in the middle represents a three-dimensional mesh model formed by connecting the respective three-dimensional points in the three-dimensional point cloud model, and the image on the right represents the three-dimensional scene model obtained by performing texture mapping on each of the meshes in the three-dimensional mesh model.
[0267] See Figure 22 , Figure 22 This is a flowchart of a method for obtaining a three-dimensional scene model provided by an embodiment of the present application.
[0268] Technicians can conduct on-site surveys of the target scene to obtain the scene data of the target scene. For example, the size of the target scene, the static objects included in the target scene, the positions of the respective static objects, and the distances between the positions of the respective static objects.
[0269] Then, select a modeling scheme, that is, determine the method of three-dimensional modeling. For example, NURBS (Non-Uniform Rational B-Splines) modeling, polygon modeling, etc. Furthermore, perform three-dimensional scanning, that is, use a three-dimensional scanning device (such as a radar scanning device, a laser scanning device, etc.) to perform three-dimensional scanning on the target scene to obtain a three-dimensional point cloud model of the target scene.
[0270] Furthermore, technicians perform manual modeling, that is, based on the spatial relationship between the positions corresponding to the respective three-dimensional points in the target scene, connect the respective three-dimensional points in the three-dimensional point cloud model to obtain a three-dimensional mesh model. Based on the material of the objects in the target scene, texture mapping is performed on each of the meshes in the three-dimensional mesh model to obtain the three-dimensional scene model of the target scene. Technicians can also perform model verification, that is, detect whether the three-dimensional scene model conforms to the target scene. For example, detect whether the positional relationship between the respective static objects is correct.
[0271] See Figure 23 , Figure 23The flowchart of an image generation method provided by an embodiment of this application.
[0272] For the monitoring video scenario, the electronic device can obtain the image after video decoding, that is, the electronic device obtains the initial image containing the target scenario collected by the camera set up at the current moment. The electronic device can correct image distortion and color, that is, the electronic device performs color difference correction on each initial image according to the specified color difference correction parameters, and / or performs distortion correction on each initial image according to the distortion types of the cameras set up, to obtain each image to be processed.
[0273] The electronic device can detect moving objects in the image (i.e., the foreground object in the foregoing embodiment) and segment the corresponding instances to obtain the foreground layer of the camera set up and the background layer of the camera set up. That is, the electronic device performs image segmentation on each image to be processed, obtains the area of the image to be processed containing the foreground object, and determines the image to be filled.
[0274] For the foreground layer of the camera set up, the electronic device can match the object instances between cameras and establish a matching relationship. The matching relationship includes the object ID and the detection frame, and the foreground objects with the same ID are the same foreground object. That is, the electronic device determines the foreground image areas belonging to the same foreground object based on the image similarity between the areas of each image to be processed.
[0275] The electronic device estimates the three-dimensional pose plane of the object to obtain plane parameters. That is, for each foreground object, the electronic device determines the three-dimensional pose plane of the foreground object in the target scenario based on the foreground image areas of the foreground object and the second conversion relationships between the image coordinates of each image to be processed and the spatial coordinates of the target scenario. The plane parameters are the plane equations of the three-dimensional pose plane.
[0276] The electronic device can select the object with the highest segmentation integrity and map it to the perspective of the monitoring camera using the three-dimensional pose plane to obtain the foreground layer after correction by the monitoring camera. That is, for each foreground object, the electronic device selects the image with the maximum integrity of the foreground object to obtain the specified foreground image area of the foreground object, and maps the specified foreground image area of the foreground object to the specified monitoring perspective through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring perspective.
[0277] For the background layer of the camera set up, the electronic device can perform background hole filling to obtain the background layer after correction by the monitoring camera. That is, the electronic device fills the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain the background image of the background of the target scenario collected by the camera set up corresponding to the specified foreground image area.
[0278] The electronic device can map the foreground layer after calibration of the monitoring camera and the background layer after calibration of the erected camera to the three-dimensional scene model of the target scene to obtain the two-dimensional base map to be processed. That is, the electronic device projects the background image at the specified monitoring view angle and the image to be projected onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene at the specified monitoring view angle.
[0279] Based on the above processing, the specified foreground image area contains the complete foreground object. Furthermore, the image to be projected obtained based on the specified foreground image area also contains the complete foreground object. Mapping the background image and the image to be projected at the specified monitoring view angle to the three-dimensional scene model of the target scene can also map the image to be projected containing the complete foreground object to the three-dimensional scene model, which can avoid the situation that the foreground object in the generated image is missing to a certain extent. And the three-dimensional pose plane is the plane where the foreground object is located in the target scene. The specified foreground image area is projected onto the specified monitoring view angle through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object at the specified monitoring view angle. Furthermore, the background image and the image to be projected at the specified monitoring view angle are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene and can avoid the situation that the foreground object in the generated image is distorted to a certain extent. Furthermore, the quality of the generated image is improved.
[0280] See Figure 24 , Figure 24 which is a flowchart of an image generation method provided by an embodiment of the present application. Camera erection refers to installing multiple erected cameras in the target scene according to the specified camera erection method. Camera bitstream acquisition is performed through the erected cameras, that is, the initial image of the target scene at the current moment is captured through the erected cameras, the image to be processed is obtained based on the acquired initial image, and the specified foreground image area of the foreground object is obtained based on the image to be processed.
[0281] Three-dimensional scene modeling refers to performing three-dimensional modeling on the target scene to obtain the three-dimensional scene model of the target scene.
[0282] An electronic device can obtain calibration parameters inside and outside the system and fuse the parameters. The calibration parameters inside and outside the system refer to the internal and external parameters of the installed camera, as well as the internal and external parameters of the monitoring camera, etc. Fusing the parameters means that the electronic device performs image fusion on the image to be processed obtained based on the installed camera and the monitoring camera based on parameters such as the internal and external parameters of the installed camera and the internal and external parameters of the monitoring camera to obtain a fused image. That is, the electronic device maps the specified foreground image area of the foreground object to the specified monitoring view of the monitoring camera based on parameters such as the internal and external parameters of the installed camera and the internal and external parameters of the monitoring camera to obtain an image to be projected. The electronic device sequentially maps the image of the background image containing the background in the target scene at the specified monitoring view and the image to be projected to the three-dimensional scene model to obtain the target image of the target scene at the specified monitoring view (i.e., the fused image of the installed camera + the monitoring camera).
[0283] Then, the electronic device can perform rendering and display to obtain the fused image of the monitoring camera. That is, it renders according to the target image of the target scene at the specified monitoring view to display the target image.
[0284] For the monitoring video scene, the installed camera collects the initial image containing the target scene in real time and stores the collected initial image at a specified storage location. The monitoring camera view video stream address is the address of the storage location of the initial image. The electronic device can obtain the initial image collected by the installed camera at each moment and process the initial image based on the method provided in the embodiments of the present application to obtain the target image of the target scene at different monitoring views, which can realize real-time and all-round monitoring of the target scene.
[0285] See Figure 25 , Figure 25 which is a comparison diagram of a target image provided by the embodiments of the present application. Figure 25 The left image in Figure 25 is the target image of the target scene at the specified monitoring view generated based on the related technology, and the right image is the target image of the target scene at the specified monitoring view generated based on the method provided by the embodiments of the present application. Figure 25 In the left image, there are obvious distortion, dislocation, and missing problems with the pedestrians,
[0286] while in the right image in
[0287] there are no distortion, dislocation, and missing problems with the pedestrians. It can be seen that the image method provided by the embodiments of the present application can improve the quality of the generated images.Figure 2 corresponds to the method embodiment, see Figure 26 , Figure 26 is a structural diagram of an image generation device provided by an embodiment of the present application. The device includes:
[0288] A to-be-processed image acquisition module 2601, configured to acquire to-be-processed images corresponding to each installed camera in a target scene; wherein, a to-be-processed image corresponding to one installed camera is obtained based on an image of the target scene collected by the installed camera; the installed cameras satisfy a preset constraint condition, and the constraint condition is: at any moment, for each foreground object in the target scene, at least one of the images collected by the installed cameras contains the complete foreground object;
[0289] A foreground image area determination module 2602, configured to respectively determine, based on image segmentation of each to-be-processed image, an image area occupied by a foreground object in each to-be-processed image as a foreground image area corresponding to the foreground object;
[0290] A foreground image area selection module 2603, configured to determine, from each foreground image area, an image area with the highest degree of completeness of the contained foreground object as a specified foreground image area;
[0291] A to-be-projected image generation module 2604, configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, determine a pixel value of a pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain a to-be-projected image of the foreground object at the specified monitoring view angle;
[0292] A background image acquisition module 2605, configured to acquire an image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area as a background image;
[0293] A target image generation module 2606, configured to project the image of the background image at the specified monitoring view angle and the to-be-projected image onto a three-dimensional scene model of the target scene in sequence, and obtain a target image of the target scene at the specified monitoring view angle.
[0294] Optionally, the device further includes:
[0295] An installation parameter acquisition module, configured to acquire current installation parameters of the installed cameras before the to-be-processed image acquisition module 2601 executes acquiring to-be-processed images corresponding to each installed camera in the target scene;
[0296] A judgment module, configured to judge whether the installed cameras satisfy a preset constraint condition based on the current installation parameters;
[0297] A control instruction sending module, configured to send a control instruction to a control device if, based on current erection parameters, each erection camera does not meet the constraint condition, so that the control device adjusts the erection parameters of each erection camera according to the control instruction, such that each erection camera meets the constraint condition based on the adjusted erection parameters.
[0298] Optionally, the to-be-projected image generation module 2604 is specifically configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a first conversion relationship between the image coordinates in a specified foreground image area and the image coordinates of the image formed at the specified monitoring view angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, so as to obtain the to-be-projected image of the foreground object at the specified monitoring view angle.
[0299] Optionally, the apparatus further includes:
[0300] A three-dimensional pose plane determination module, configured to, before the to-be-projected image generation module 2604 executes, for each pixel coordinate of the foreground object at a specified monitoring view angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, so as to obtain the to-be-projected image of the foreground object at the specified monitoring view angle, execute based on each foreground image area and each second conversion relationship between the image coordinates of each to-be-processed image and the spatial coordinates of the target scene, to determine the three-dimensional pose plane of the foreground object in the target scene;
[0301] The to-be-projected image generation module 2604 is specifically configured to, for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a third conversion relationship between the image coordinates in a specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and a fourth conversion relationship between the image coordinates of the image formed at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, so as to obtain the to-be-projected image of the foreground object at the specified monitoring view angle.
[0302] Optionally, the three-dimensional pose plane determination module is specifically configured to determine two erection cameras corresponding to a to-be-processed image containing the foreground object as specified erection cameras;
[0303] For each specified erection camera, determine a detection frame containing the foreground object in the to-be-processed image corresponding to the specified erection camera;
[0304] Based on the second conversion relationship between the image coordinates corresponding to the specified installed camera and the spatial coordinates, determine the three-dimensional point corresponding to the specified point in the detection frame in the target scene;
[0305] Determine the plane containing the optical center of the specified installed camera and the corresponding three-dimensional point as the reference plane corresponding to the specified installed camera;
[0306] Determine the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0307] Optionally, the three-dimensional pose plane determination module is specifically configured to use the intersection line of each reference plane as the rotation axis and rotate the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0308] Optionally, the device further includes:
[0309] An edge pixel point determination module, configured to determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image region before the to-be-projected image generation module 2604 executes to determine the pixel value of the pixel coordinate corresponding to the pixel coordinate in the specified foreground image region for each pixel coordinate of the foreground object in the specified monitoring view as the pixel value of the pixel coordinate, so as to obtain the to-be-projected image of the foreground object in the specified monitoring view;
[0310] An edge pixel point mapping module, configured to determine, for each edge pixel point, the corresponding pixel coordinate of the edge pixel point in the specified monitoring view as the edge pixel coordinate;
[0311] A pixel coordinate determination module, configured to determine the pixel coordinates included in the region with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object in the specified monitoring view.
[0312] Optionally, the foreground image region determination module 2602 is specifically configured to perform image segmentation on each to-be-processed image to obtain the image regions occupied by each foreground object in each to-be-processed image as the to-be-processed image regions;
[0313] Based on the image similarity between the to-be-processed image regions, determine the to-be-processed image regions belonging to the same foreground object as the foreground image region corresponding to the foreground object.
[0314] Optionally, the to-be-processed image acquisition module 2601 is specifically configured to acquire the images of the target scene collected by each installed camera as the initial images;
[0315] Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of the installed cameras, to obtain each image to be processed.
[0316] Optionally, the background image acquisition module 2605 is specifically configured to acquire an image to be filled; wherein, the image to be filled is obtained by deleting the specified foreground image area from the image to be processed to which it belongs;
[0317] Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain an image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area; wherein, the preset image is an image collected in advance containing the background in the target scene.
[0318] Based on the image generation device provided in the embodiments of the present application, at any moment, for each foreground object in the target scene, at least one of the images collected by the installed cameras in the target scene contains the complete foreground object. Correspondingly, the completeness of the foreground object included in the specified foreground image area is the largest, that is, the specified foreground image area contains the complete foreground object. Furthermore, the image to be projected obtained based on the specified foreground image area also contains the complete foreground object. Mapping the background image and the image to be projected under the specified monitoring view to the three-dimensional scene model of the target scene can map the image to be projected containing the complete foreground object to the three-dimensional scene model, which can avoid the situation of missing foreground objects in the generated image to a certain extent, and thus improve the quality of the generated image.
[0319] The embodiments of the present application also provide an electronic device, as Figure 27 shown, including a processor 2701, a communication interface 2702, a memory 2703, and a communication bus 2704. Among them, the processor 2701, the communication interface 2702, and the memory 2703 complete mutual communication through the communication bus 2704,
[0320] The memory 2703 is used to store a computer program;
[0321] When the processor 2701 is used to execute the program stored on the memory 2703, the following steps are implemented:
[0322] Obtain the images to be processed corresponding to each installed camera in the target scenario; among them, the image to be processed corresponding to one installed camera is obtained based on the image of the target scenario collected by this installed camera; the installed cameras satisfy a preset constraint condition, and the constraint condition is: at any moment, for each foreground object in the target scenario, at least one of the images collected by the installed cameras contains the complete foreground object;
[0323] Based on performing image segmentation on each image to be processed, respectively determine the image regions occupied by the foreground objects in each image to be processed as the foreground image regions corresponding to the foreground objects;
[0324] From each foreground image region, determine the image region with the highest degree of completeness of the contained foreground object as the specified foreground image region;
[0325] For each pixel coordinate of the foreground object under the specified monitoring perspective, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image region as the pixel value of this pixel coordinate, and obtain the image to be projected of the foreground object under the specified monitoring perspective;
[0326] Obtain the image containing the background in the target scenario collected by the installed camera corresponding to the specified foreground image region as the background image;
[0327] Successively project the image of the background image under the specified monitoring perspective and the image to be projected onto the three-dimensional scene model of the target scenario to obtain the target image of the target scenario under the specified monitoring perspective.
[0328] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity, only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.
[0329] The communication interface is used for communication between the above electronic device and other devices.
[0330] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0331] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0332] In another embodiment provided by the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above image generation methods are implemented.
[0333] In another embodiment provided by the present application, a computer program product containing instructions is further provided. When it runs on a computer, the computer is caused to execute any of the image generation methods in the above embodiments.
[0334] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0335] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0336] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the apparatus, electronic device, computer-readable storage medium and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the corresponding parts of the method embodiments for the relevant content.
[0337] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. An image generation method, characterized in that, the method includes: obtaining the to-be-processed images corresponding to each installed camera in the target scene; wherein, the to-be-processed image corresponding to one installed camera is obtained based on the image of the target scene collected by this installed camera; each installed camera satisfies a preset constraint condition, and the constraint condition is: at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera contains the complete foreground object; respectively determining, based on image segmentation of each to-be-processed image, the image area occupied by the foreground object in each to-be-processed image as the foreground image area corresponding to the foreground object; determining, from each foreground image area, the image area with the highest degree of integrity of the contained foreground object as the specified foreground image area; for each pixel coordinate of the foreground object at the specified monitoring perspective, determining the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, to obtain the to-be-projected image of the foreground object at the specified monitoring perspective; obtaining the image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image; sequentially projecting the image of the background image at the specified monitoring perspective and the to-be-projected image onto the three-dimensional scene model of the target scene to obtain the target image of the target scene at the specified monitoring perspective.
2. The method according to claim 1, characterized in that, before the obtaining the to-be-processed images corresponding to each installed camera in the target scene, the method further includes: obtaining the current installation parameters of each installed camera; judging whether each installed camera satisfies a preset constraint condition based on the current installation parameters; if each installed camera does not satisfy the constraint condition based on the current installation parameters, sending a control instruction to the control device so that the control device adjusts the installation parameters of each installed camera according to the control instruction, such that each installed camera satisfies the constraint condition based on the adjusted installation parameters.
3. The method according to claim 1, characterized in that, the for each pixel coordinate of the foreground object at the specified monitoring perspective, determining the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, to obtain the to-be-projected image of the foreground object at the specified monitoring perspective includes: for each pixel coordinate of the foreground object at the specified monitoring perspective, based on the first conversion relationship between the image coordinates in the specified foreground image area and the image coordinates of the imaging at the specified monitoring perspective, determining the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, to obtain the to-be-projected image of the foreground object at the specified monitoring perspective.
4. The method according to claim 1, characterized in that, Before obtaining the projected image of the foreground object at the specified monitoring view angle by determining the pixel value of the pixel point corresponding to each pixel coordinate of the foreground object at the specified monitoring view angle in the specified foreground image area as the pixel value of the pixel coordinate, the method further includes: Based on each foreground image area and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determining the three-dimensional pose plane of the foreground object in the target scene; The step of obtaining the projected image of the foreground object at the specified monitoring view angle by determining the pixel value of the pixel point corresponding to each pixel coordinate of the foreground object at the specified monitoring view angle in the specified foreground image area as the pixel value of the pixel coordinate includes: For each pixel coordinate of the foreground object at the specified monitoring view angle, based on the third conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the fourth conversion relationship between the image coordinates of the image captured at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determining the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, to obtain the projected image of the foreground object at the specified monitoring view angle.
5. The method according to claim 4, wherein, The step of determining the three-dimensional pose plane of the foreground object in the target scene based on each foreground image area and each second conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene includes: Determining two mounted cameras corresponding to the image to be processed that contain the foreground object as the specified mounted cameras; For each specified mounted camera, determining the detection frame containing the foreground object in the image to be processed corresponding to the specified mounted camera; Based on the second conversion relationship between the image coordinates corresponding to the specified mounted camera and the spatial coordinates, determining the three-dimensional point corresponding to the specified point in the detection frame in the target scene; Determining the plane containing the optical center of the specified mounted camera and the corresponding three-dimensional point as the reference plane corresponding to the specified mounted camera; Determining the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
6. The method according to claim 5, wherein, The step of determining the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene includes: Taking the intersection line of each reference plane as the rotation axis, and rotating the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
7. The method according to claim 1, wherein, Before obtaining the projected image of the foreground object at the specified monitoring angle by determining the pixel value of the pixel point corresponding to each pixel coordinate of the foreground object at the specified monitoring angle in the specified foreground image area as the pixel value of the pixel coordinate, the method further includes: Determining the edge pixel points of the foreground object in the to-be-processed image to which the specified foreground image area belongs; For each edge pixel point, determining the corresponding pixel coordinate of the edge pixel point at the specified monitoring angle as the edge pixel coordinate; Determining the pixel coordinates included in the area with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object at the specified monitoring angle.
8. The method according to claim 1, wherein, the determining, based on image segmentation of each to-be-processed image, of the image area occupied by the foreground object in each to-be-processed image as the foreground image area corresponding to the foreground object includes: Based on image segmentation of each to-be-processed image, obtaining the image areas occupied by each foreground object in each to-be-processed image as the to-be-processed image areas; Based on the image similarity between the to-be-processed image areas, determining the to-be-processed image areas belonging to the same foreground object as the foreground image area corresponding to the foreground object.
9. The method according to claim 1, wherein, the obtaining of the to-be-processed images corresponding to the cameras installed in the target scene includes: Obtaining the images of the target scene collected by each installed camera as the initial images; Performing chromatic aberration correction on each initial image according to the specified chromatic aberration correction parameters, and / or performing distortion correction on each initial image according to the distortion types of the installed cameras to obtain each to-be-processed image.
10. The method according to claim 1, wherein, the obtaining of the image including the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image includes: Obtaining the to-be-filled image; wherein, the to-be-filled image is obtained by deleting the specified foreground image area from the to-be-processed image to which it belongs; Filling the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the to-be-filled image to obtain the image including the background in the target scene collected by the installed camera corresponding to the specified foreground image area; wherein, the preset image is the image including the background in the target scene collected in advance.
11. An image generation device, wherein, the device includes: A to-be-processed image acquisition module, configured to acquire to-be-processed images corresponding to the cameras installed in the target scene; wherein, the to-be-processed image corresponding to one installed camera is obtained based on the image of the target scene collected by the installed camera; the installed cameras satisfy a preset constraint condition, and the constraint condition is that at any moment, for each foreground object in the target scene, at least one of the images collected by the installed cameras includes the complete foreground object; Foreground image region determination module, configured to respectively determine, based on image segmentation of each image to be processed, the image region occupied by the foreground object in each image to be processed as the foreground image region corresponding to the foreground object; Foreground image region selection module, configured to determine, from each foreground image region, the image region with the highest integrity of the foreground object contained therein as the designated foreground image region; Image to be projected generation module, configured to, for each pixel coordinate of the foreground object at a designated monitoring perspective, determine the pixel value of the pixel point corresponding to the pixel coordinate in the designated foreground image region as the pixel value of the pixel coordinate, and obtain the image to be projected of the foreground object at the designated monitoring perspective; Background image acquisition module, configured to acquire, as the background image, the image containing the background in the target scene captured by the installed camera corresponding to the designated foreground image region; Target image generation module, configured to sequentially project the image of the background image at the designated monitoring perspective and the image to be projected onto the three-dimensional scene model of the target scene to obtain the target image of the target scene at the designated monitoring perspective.
Citation Information
Patent Citations
Live broadcast image synthesis method and device, terminal equipment and readable storage medium
CN113837979A