An image generation method and apparatus
By determining the three-dimensional pose plane of the foreground object in the target scene and performing image projection, combined with background image mapping, the problem of distortion of moving objects is solved and the image generation quality is improved.
Patent Information
- Application Number
- CN202210662274.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-13
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-06-13
AI Technical Summary
In the prior art, insufficient modeling of moving objects in the target scene results in low generated image quality and distortion of moving objects when mapped to a three-dimensional scene model.
By obtaining the images to be processed by each set up camera in the target scene, performing image segmentation to determine the foreground object area, using the three-dimensional pose plane to project the image of the foreground object under the specified monitoring angle to the three-dimensional scene model, and mapping it with the background image to avoid distortion of the foreground object.
Improve the quality of the generated image, avoid distortion of moving objects, and ensure the integrity and clarity of the image.
Smart Images

Figure CN115147262B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technologies, and particularly to an image generation method and apparatus. Background Art
[0002] In related technologies, multiple cameras can be installed at different positions in a target scene, and images of the target scene from different shooting perspectives (which can be referred to as images to be processed) are captured by these multiple cameras. Then, a three-dimensional scene model of the target scene is obtained, and the images to be processed are mapped to the above three-dimensional scene model. Based on the mapping result, a user can view images of the target scene from different monitoring perspectives.
[0003] However, in related technologies, only static objects (such as buildings) in the target scene are modeled, and moving objects (such as pedestrians) in the target scene are not modeled. When mapping the images to be processed to the three-dimensional scene model, the parts corresponding to the moving objects in the images to be processed are mapped to the model parts corresponding to the background of the moving objects in the target scene in the three-dimensional scene model, resulting in distortion of the moving objects in the images viewed by the user. That is, in related technologies, the quality of the generated images is low. Summary of the Invention
[0004] The purpose of the embodiments of this application is to provide an image generation method and apparatus to improve the quality of the generated images. The specific technical solutions are as follows:
[0005] In a first aspect, to achieve the above objective, the embodiments of this application disclose an image generation method, and the method includes:
[0006] Obtain the images to be processed corresponding to each installed camera in the target scene; where the image to be processed corresponding to one installed camera is obtained based on the image of the target scene collected by this installed camera;
[0007] Based on performing image segmentation on each image to be processed, respectively determine the image regions occupied by the foreground objects in each image to be processed as the foreground image regions corresponding to the foreground objects;
[0008] Based on each foreground image region and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determine the three-dimensional pose plane of the foreground object in the target scene;
[0009] For each pixel coordinate of the foreground object under a specified monitoring view angle, based on the second conversion relationship between the image coordinates in the specified foreground image region and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates imaged under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region as the pixel value of the pixel coordinate, and obtain the image to be projected of the foreground object under the specified monitoring view angle;
[0010] Obtain an image containing the background in the target scene collected by a set-up camera corresponding to the specified foreground image region as the background image;
[0011] Project the image of the background image under the specified monitoring view angle and the image to be projected onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene under the specified monitoring view angle.
[0012] Optionally, the determining of the three-dimensional pose plane of the foreground object in the target scene based on each foreground image region and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene includes:
[0013] Determine two set-up cameras corresponding to the image to be processed that contain the foreground object as the specified set-up cameras;
[0014] For each specified set-up camera, determine the detection frame containing the foreground object in the image to be processed corresponding to the specified set-up camera;
[0015] Based on the first conversion relationship between the image coordinates corresponding to the specified set-up camera and the spatial coordinates, determine the three-dimensional points corresponding to the specified points in the detection frame in the target scene;
[0016] Determine the plane containing the optical center of the specified set-up camera and the corresponding three-dimensional points as the reference plane corresponding to the specified set-up camera;
[0017] Determine the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0018] Optionally, the determining of the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene includes:
[0019] Using the intersection line of each reference plane as the rotation axis, rotate the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0020] Optionally, before determining, for each pixel coordinate of the foreground object under a specified monitoring view angle, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates imaged under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the image to be projected of the foreground object under the specified monitoring view angle, the method further includes:
[0021] Determine the edge pixel points of the foreground object in the image to be processed to which the specified foreground image area belongs;
[0022] For each edge pixel point, based on the second conversion relationship and the third conversion relationship, determine the corresponding pixel coordinate of the edge pixel point under the specified monitoring view angle as the edge pixel coordinate;
[0023] Determine the pixel coordinates included in the area with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object under the specified monitoring view angle.
[0024] Optionally, before determining, for each pixel coordinate of the foreground object under a specified monitoring view angle, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates imaged under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the image to be projected of the foreground object under the specified monitoring view angle, the method further includes:
[0025] Determine, from each foreground image area, the image area with the highest integrity of the included foreground object as the specified foreground image area.
[0026] Optionally, the determining, by performing image segmentation on each image to be processed, of the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object includes:
[0027] Perform image segmentation on each image to be processed to obtain the image areas occupied by each foreground object in each image to be processed as the image areas to be processed;
[0028] Based on the image similarity between the image areas to be processed, determine the image areas to be processed belonging to the same foreground object as the foreground image area corresponding to the foreground object.
[0029] Optionally, obtaining the to-be-processed images corresponding to each installed camera in the target scenario includes:
[0030] Obtaining the images of the target scenario collected by each installed camera as initial images;
[0031] Performing color difference correction on each initial image according to specified color difference correction parameters, and / or performing distortion correction on each initial image according to the distortion types of each installed camera to obtain each to-be-processed image.
[0032] Optionally, obtaining the image including the background in the target scenario collected by the installed camera corresponding to the specified foreground image region as the background image includes:
[0033] Obtaining the to-be-filled image; wherein, the to-be-filled image is obtained by deleting the specified foreground image region from the to-be-processed image to which it belongs;
[0034] Filling the part of the preset image corresponding to the specified foreground image region into the position corresponding to the specified foreground image region in the to-be-filled image to obtain the image including the background in the target scenario collected by the installed camera corresponding to the specified foreground image region; wherein, the preset image is an image collected in advance including the background in the target scenario.
[0035] Optionally, at any moment, for each foreground object in the target scenario, at least one of the images collected by each installed camera includes the complete foreground object.
[0036] In a second aspect, to achieve the above object, an embodiment of the present application discloses an image generation device, and the device includes:
[0037] A to-be-processed image acquisition module, configured to acquire to-be-processed images corresponding to each installed camera in the target scenario; wherein, the to-be-processed image corresponding to one installed camera is obtained based on the image of the target scenario collected by this installed camera;
[0038] A foreground image region determination module, configured to respectively determine the image regions occupied by foreground objects in each to-be-processed image as the foreground image regions corresponding to the foreground objects based on image segmentation of each to-be-processed image;
[0039] A three-dimensional pose plane determination module, configured to determine the three-dimensional pose plane of the foreground object in the target scenario based on each foreground image region and each first conversion relationship between the image coordinates of each to-be-processed image and the spatial coordinates of the target scenario;
[0040] A projected image generation module, configured to, for each pixel coordinate of the foreground object at a specified monitoring perspective, based on a second conversion relationship between the image coordinates in a specified foreground image region and the three-dimensional coordinates in the three-dimensional pose plane, and a third conversion relationship between the image coordinates imaged at the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region as the pixel value of the pixel coordinate, so as to obtain a projected image of the foreground object at the specified monitoring perspective;
[0041] A background image acquisition module, configured to acquire an image containing the background in the target scene collected by a setup camera corresponding to a specified foreground image region as a background image;
[0042] A target image generation module, configured to project the image of the background image at the specified monitoring perspective and the projected image to the three-dimensional scene model of the target scene in sequence, so as to obtain a target image of the target scene at the specified monitoring perspective.
[0043] Optionally, the three-dimensional pose plane determination module is specifically configured to determine two setup cameras corresponding to a to-be-processed image containing the foreground object as specified setup cameras;
[0044] For each specified setup camera, determine a detection frame containing the foreground object in the to-be-processed image corresponding to the specified setup camera;
[0045] Based on a first conversion relationship between the image coordinates corresponding to the specified setup camera and the spatial coordinates, determine a three-dimensional point corresponding to a specified point in the detection frame in the target scene;
[0046] Determine a plane containing the optical center of the specified setup camera and the corresponding three-dimensional point as a reference plane corresponding to the specified setup camera;
[0047] Determine a plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0048] Optionally, the three-dimensional pose plane determination module is specifically configured to rotate an initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle with the intersection line of each reference plane as the rotation axis, so as to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0049] Optionally, the apparatus further includes:
[0050] An edge pixel point determination module, configured to, before the to-be-projected image generation module executes for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a second conversion relationship between the image coordinates in a specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and a third conversion relationship between the image coordinates imaged at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the to-be-projected image of the foreground object at the specified monitoring view angle, determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area;
[0051] An edge pixel point mapping module, configured to, for each edge pixel point, based on the second conversion relationship and the third conversion relationship, determine the corresponding pixel coordinate of the edge pixel point at the specified monitoring view angle as the edge pixel coordinate;
[0052] A pixel coordinate determination module, configured to determine the pixel coordinates included in the area with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object at the specified monitoring view angle.
[0053] Optionally, the apparatus further includes:
[0054] A foreground image area selection module, configured to, before the to-be-projected image generation module executes for each pixel coordinate of the foreground object at a specified monitoring view angle, based on a second conversion relationship between the image coordinates in a specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and a third conversion relationship between the image coordinates imaged at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the to-be-projected image of the foreground object at the specified monitoring view angle, determine, from each foreground image area, the image area with the highest integrity of the included foreground object as the specified foreground image area.
[0055] Optionally, the foreground image area determination module is specifically configured to, based on image segmentation of each to-be-processed image, obtain the image areas occupied by each foreground object in each to-be-processed image as the to-be-processed image areas;
[0056] Based on the image similarity between each to-be-processed image area, determine the to-be-processed image areas belonging to the same foreground object as the foreground image areas corresponding to the foreground object.
[0057] Optionally, the to-be-processed image acquisition module is specifically configured to acquire the images of the target scene collected by each installed camera as the initial images;
[0058] Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of the installed cameras, to obtain each image to be processed.
[0059] Optionally, the background image acquisition module is specifically configured to acquire an image to be filled; wherein, the image to be filled is obtained by deleting the specified foreground image area from the image to be processed to which it belongs;
[0060] Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain an image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area; wherein, the preset image is an image collected in advance containing the background in the target scene.
[0061] Optionally, at any moment, for each foreground object in the target scene, at least one of the images collected by the installed cameras contains the complete foreground object.
[0062] The embodiment of the present application further provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0063] The memory is used to store a computer program;
[0064] When the processor is used to execute the program stored on the memory, it implements the steps of the image generation method described in any one of the above.
[0065] The embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the image generation method described in any one of the above.
[0066] The embodiment of the present application further provides a computer program product containing instructions, which when running on a computer, causes the computer to execute the image generation method described in any one of the above.
[0067] An image generation method provided by an embodiment of the present application obtains to-be-processed images corresponding to each installed camera in a target scene; the to-be-processed image corresponding to an installed camera is obtained based on the image of the target scene collected by the installed camera; based on performing image segmentation on each to-be-processed image, the image regions occupied by foreground objects in each to-be-processed image are respectively determined as foreground image regions corresponding to the foreground objects; based on each foreground image region and each first conversion relationship between the image coordinates of each to-be-processed image and the spatial coordinates of the target scene, a three-dimensional pose plane of the foreground object in the target scene is determined; for each pixel coordinate of the foreground object under a specified monitoring perspective, based on a second conversion relationship between the image coordinates in the specified foreground image region and the three-dimensional coordinates in the three-dimensional pose plane, and a third conversion relationship between the image coordinates of the image formed under the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region is determined as the pixel value of the pixel coordinate, and a to-be-projected image of the foreground object under the specified monitoring perspective is obtained; an image collected by the installed camera corresponding to the specified foreground image region and including the background in the target scene is obtained as a background image; the image of the background image and the to-be-projected image under the specified monitoring perspective are sequentially projected onto the three-dimensional scene model of the target scene, and a target image of the target scene under the specified monitoring perspective is obtained.
[0068] Based on the above processing, the three-dimensional pose plane is the plane where the foreground object is located in the target scene. The specified foreground image region is projected onto the specified monitoring perspective through the three-dimensional pose plane of the foreground object, and a to-be-projected image of the foreground object under the specified monitoring perspective is obtained. Furthermore, the background image and the to-be-projected image under the specified monitoring perspective are mapped onto the three-dimensional scene model of the target scene. That is, the foreground object is mapped onto the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent, and can avoid the distortion of the foreground object in the generated image to a certain extent. Furthermore, the quality of the generated image is improved.
[0069] Of course, any product or method implementing the present application does not necessarily need to achieve all the above-mentioned advantages at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other embodiments can also be obtained based on these drawings.
[0071] Figure 1a An image of a target scene provided by an embodiment of the present application;
[0072] Figure 1b Another image of the target scenario provided by the embodiment of the present application;
[0073] Figure 2 Flowchart of an image generation method provided by the embodiment of the present application;
[0074] Figure 3 Flowchart of another image generation method provided by the embodiment of the present application;
[0075] Figure 4 Comparison diagram of an initial image and an image to be processed provided by the embodiment of the present application;
[0076] Figure 5 Flowchart of another image generation method provided by the embodiment of the present application;
[0077] Figure 6 Schematic diagram of the principle for extracting the foreground image region provided by the embodiment of the present application;
[0078] Figure 7 Schematic diagram of the principle for matching the foreground image region provided by the embodiment of the present application;
[0079] Figure 8 Flowchart of another image generation method provided by the embodiment of the present application;
[0080] Figure 9 Schematic diagram of the principle for determining the three-dimensional pose plane provided by the embodiment of the present application;
[0081] Figure 10 Flowchart of another image generation method provided by the embodiment of the present application;
[0082] Figure 11 Schematic diagram of the principle for mapping a specified foreground image region to a specified monitoring view angle provided by the embodiment of the present application;
[0083] Figure 12 Flowchart of another image generation method provided by the embodiment of the present application;
[0084] Figure 13 Comparison diagram of an image to be processed, an image to be filled, and a background image provided by the embodiment of the present application;
[0085] Figure 14 Schematic diagram of the principle for mapping a background image to a specified monitoring view angle provided by the embodiment of the present application;
[0086] Figure 15 Comparison diagram of a three-dimensional scene model and a target image of the target scenario at a specified monitoring view angle provided by the embodiment of the present application;
[0087] Figure 16 A comparison diagram of a three-dimensional scene model provided by an embodiment of the present application;
[0088] Figure 17 A flowchart of a method for obtaining a three-dimensional scene model provided by an embodiment of the present application;
[0089] Figure 18 A schematic diagram of a camera installation method provided by an embodiment of the present application;
[0090] Figure 19 A schematic diagram of another camera installation method provided by an embodiment of the present application;
[0091] Figure 20 A flowchart of another image generation method provided by an embodiment of the present application;
[0092] Figure 21 A flowchart of another image generation method provided by an embodiment of the present application;
[0093] Figure 22 A comparison diagram of a target image of a target scene under a specified monitoring view provided by an embodiment of the present application;
[0094] Figure 23 A structural diagram of an image generation device provided by an embodiment of the present application;
[0095] Figure 24 A structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0096] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0097] In the related art, when generating images of a target scene from various monitoring views, only static objects (such as buildings) in the target scene are modeled, and moving objects (such as pedestrians) in the target scene are not modeled. When mapping the images captured by the installed cameras to the three-dimensional scene model, the parts corresponding to the moving objects in the images are mapped to the model parts corresponding to the backgrounds of the moving objects in the target scene in the three-dimensional scene model, resulting in distortion of the moving objects in the images viewed by the user. That is, in the related art, the quality of the generated images is low.
[0098] Exemplarily, refer to Figure 1a , Figure 1aAn image of a target scenario provided by an embodiment of the present application. The target scenario is a parking lot, and the buildings therein are static objects. Since in the related art, static objects in the target scenario are modeled, when mapping the image captured by the erected camera to the three-dimensional scene model, the corresponding part of the static object in the image can be mapped to the three-dimensional model of the static object in the three-dimensional scene model. Therefore, the image has a good registration effect on the static object, and the image obtained by mapping the image to the three-dimensional scene model basically does not distort.
[0099] See Figure 1b , Figure 1b Another image of a target scenario provided by an embodiment of the present application. The target scenario is a factory, where the tables, chairs, ceiling, floor wall columns, and the workbench of the staff are static objects, and the staff are moving objects. Since in the related art, only static objects in the target scenario are modeled, when mapping the image captured by the erected camera to the three-dimensional scene model, the corresponding part of the static object in the image can be mapped to the three-dimensional model of the static object in the three-dimensional scene model. Therefore, the image has a good registration effect on the static object, and the image obtained by mapping the image to the three-dimensional scene model basically does not distort.
[0100] However, in the related art, moving objects in the target scenario are not modeled. When mapping the image captured by the erected camera to the three-dimensional scene model, the corresponding part of the moving object in the image will be mapped to the model part corresponding to the background of the moving object in the target scenario in the three-dimensional scene model, resulting in distortion of the moving object in the image viewed by the user. That is, in the related art, the quality of the generated image is low.
[0101] To solve the above problems, an embodiment of the present application provides an image generation method applied to an electronic device. The electronic device can obtain the to-be-processed images corresponding to each erected camera in the target scenario, and process the to-be-processed images based on the method provided by the embodiment of the present application to obtain the target image of the target scenario from the specified monitoring perspective. Subsequently, the user can view the image of the target scenario from the specified monitoring perspective. For example, for a surveillance video scenario, a surveillance video containing the target scenario from the specified monitoring perspective can be obtained, and the user can view the surveillance video of the target scenario from the specified monitoring perspective through the electronic device to achieve the effect of all-round monitoring.
[0102] See Figure 2 , Figure 2 A flowchart of an image generation method provided by an embodiment of the present application. The method may include the following steps:
[0103] S201: Obtain the to-be-processed images corresponding to each erected camera in the target scenario.
[0104] Among them, the image to be processed corresponding to one mounted camera is obtained based on the image of the target scene collected by the mounted camera.
[0105] S202: Based on image segmentation of each image to be processed, respectively determine the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object.
[0106] S203: Based on each foreground image area and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determine the three-dimensional pose plane of the foreground object in the target scene.
[0107] S204: For each pixel coordinate of the foreground object under a specified monitoring view angle, based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates of the image formed under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate, and obtain the image to be projected of the foreground object under the specified monitoring view angle.
[0108] S205: Obtain the image containing the background in the target scene collected by the mounted camera corresponding to the specified foreground image area as the background image.
[0109] S206: Project the image of the background image and the image to be projected under the specified monitoring view angle onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene under the specified monitoring view angle.
[0110] Based on the image generation method provided in the embodiments of the present application, the three-dimensional pose plane is the plane where the foreground object is located in the target scene. The specified foreground image area is projected onto the specified monitoring view angle through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring view angle. Furthermore, the background image and the image to be projected under the specified monitoring view angle are mapped onto the three-dimensional scene model of the target scene. That is, the foreground object is mapped onto the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent, and can avoid the distortion of the foreground object in the generated image to a certain extent. Furthermore, the quality of the generated image is improved.
[0111] In at least one embodiment of the present application, the images in the video streams captured by multiple mounted cameras can all be processed in the above manner to obtain the target image. The images in the video stream captured in a continuous time period, after the above processing, the obtained target images can form a more complete and clear high-quality video stream.
[0112] For step S201, the target scenario can be scenarios such as factories, streets, etc. Multiple installed cameras are installed in the target scenario, and each installed camera is used to capture images of the target scenario from different shooting perspectives. For the surveillance video scenario, for each installed camera, the electronic device can obtain the video image captured by the installed camera at the current moment, and obtain the image to be processed based on the video image.
[0113] Each installed camera is an installed camera that matches the specified monitoring perspective. That an installed camera matches the specified monitoring perspective means that there is an overlapping area between the shooting perspective of the installed camera and the specified monitoring perspective.
[0114] The specified monitoring perspective can be any monitoring perspective selected by the user when viewing the image of the target scenario. When the specified monitoring perspective changes, for example, when the user switches the monitoring perspective, the electronic device can determine the installed camera that matches the switched monitoring perspective. Subsequently, based on the method provided in the embodiments of the present application, the image of the target scenario under the switched monitoring perspective can be determined.
[0115] In one implementation, for each installed camera, the electronic device can directly obtain the image of the target scenario captured by the installed camera (i.e., the initial image) as the image to be processed corresponding to the installed camera.
[0116] In another implementation, for each installed camera, the initial image captured by the installed camera may be distorted, and there may be color differences between the initial images captured by each installed camera. To improve the quality of the target image of the target scenario under the specified monitoring perspective, the electronic device can perform image correction on the initial image captured by the installed camera to obtain the image to be processed.
[0117] Correspondingly, on the basis of Figure 2 , referring to Figure 3 , step S201 may include the following steps:
[0118] S2011: Obtain the images of the target scenario captured by each installed camera as the initial images.
[0119] S2012: Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion type of each installed camera to obtain each image to be processed.
[0120] The electronic device can obtain the initial images of the target scenario captured by each installed camera. For each pixel point in the initial image, calculate the product of the pixel value of the pixel point and the color correction matrix (English: Color Correction Matrix, abbreviation: CCM) to obtain the corrected pixel value of the pixel point, and thus the image after color difference correction (i.e., the image to be processed) can be obtained.
[0121] Exemplarily, an electronic device can perform color difference correction on an initial image according to the following formula:
[0122] P2 = C × P1 (1)
[0123] C represents a color difference correction matrix, P2 represents a matrix containing the pixel values of each pixel point in the initial image, and P1 represents a matrix containing the pixel values of each pixel point in the image to be processed.
[0124] In one implementation, P can be P1 or P2. R represents the pixel value of the red channel of the pixel point in the corresponding image, G represents the pixel value of the green channel of the pixel point in the corresponding image, and B represents the pixel value of the blue channel of the pixel point in the corresponding image.
[0125] r1 to r3 represent red compensation values, g1 to g3 represent green compensation values, b1 to b3 represent blue compensation values, and c1 to c3 represent preset color compensation coefficients.
[0126] In another implementation, P can be P1 or P2. R represents the pixel value of the red channel of the pixel point in the corresponding image, G represents the pixel value of the green channel of the pixel point in the corresponding image, and B represents the pixel value of the blue channel of the pixel point in the corresponding image.
[0127] r1 to r3 represent red compensation values, g1 to g3 represent green compensation values, and b1 to b3 represent blue compensation values.
[0128] The values in the color difference correction matrix can be set by technicians according to requirements. For example, when it is necessary to make the image tend to be red, r1 to r3 can be set to larger values, or when it is necessary to make the image tend to be green, g1 to g3 can be set to larger values.
[0129] Alternatively, the electronic device can pre-acquire the pixel values of the images captured by any two mounted cameras, and calculate the values in the color difference correction matrix based on the pixel values of the images captured by the two mounted cameras.
[0130] For each installed camera, if the installed camera is a pinhole camera, the distortion coefficients of the distortion type corresponding to the installed camera include: K1, K2, P1, and P2. In the distortion type corresponding to the pinhole camera, K1 and K2 represent radial distortion, which is generated during the process of converting the camera coordinate system of the installed camera to the image coordinate system of the initial image, and P1 and P2 represent tangential distortion, which is caused by the lens of the installed camera not being completely parallel to the initial image. If the installed camera is a fisheye camera, the distortion coefficients of the distortion type corresponding to the installed camera include: K1, K2, K3, and K4.
[0131] For each pixel point in the initial image, the electronic device can calculate the corrected coordinates of the pixel point according to the distortion coefficients of the distortion type corresponding to the installed camera and the coordinates of the pixel point in the initial image, and thus can obtain the image to be processed after distortion correction.
[0132] See Figure 4 , Figure 4 In the figure on the left, the image is the initial image captured by the installed camera, and the floor in the initial image is distorted. The image on the right is the image to be processed obtained by distortion correction, and the floor in the image to be processed is not distorted.
[0133] Based on the above processing, the electronic device performs distortion correction and / or chromatic aberration correction on the initial image captured by the installed camera, which can, to a certain extent, avoid distortion or chromatic aberration in the finally generated target image and can further improve the quality of the generated image.
[0134] Regarding step S202, the foreground object is an object that may move in the target scene. For example, the foreground object can be a vehicle, a pedestrian, etc. The electronic device can determine the object belonging to the preset foreground object type in the image to be processed as the foreground object. The preset foreground object type can include vehicles, pedestrians, etc.
[0135] Regarding the surveillance video scene, after obtaining the image to be processed based on the initial image captured by the installed camera at the current moment, the electronic device can determine the positions of each object in the image to be processed obtained at the current moment (which can be called the first position), and the positions of each object in the image to be processed obtained at the previous moment (which can be called the second position), and determine the objects with different first positions and second positions. The determined objects are the moving objects in the current target scene and are used as the foreground objects.
[0136] In one embodiment, on the basis of Figure 2 , see Figure 5 , step S202 may include the following steps:
[0137] S2021: Based on image segmentation of each image to be processed, obtain the image regions occupied by each foreground object in each image to be processed as the image regions to be processed.
[0138] S2022: Based on the image similarity between the image regions to be processed, determine the image regions to be processed that belong to the same foreground object as the foreground image regions corresponding to the foreground object.
[0139] For each image to be processed, the electronic device can perform image segmentation on the image to be processed based on the object detection algorithm, obtain the image regions occupied by each foreground object in the image to be processed, and extract and determine the image regions to obtain the image regions to be processed.
[0140] The object detection algorithm can be Mask-RCNN (Mask Regions with Convolutional Neural Network), or the object detection algorithm can also be YOLOv3 (You only look once-v3, an end-to-end object detection algorithm based on deep learning), but it is not limited to this.
[0141] Exemplarily, refer to Figure 6 , Figure 6 In, the left image is the image to be processed, and the right image represents the image regions to be processed extracted from the image to be processed. The pedestrian in the image to be processed is the foreground object included in the image to be processed. In the image to be processed, for each foreground object, the black rectangle containing the foreground object is the detection box of the foreground object, and the rectangle can be the minimum bounding rectangle of the foreground object. The electronic device can perform instance segmentation according to the edge of the foreground object to obtain Figure 6 the 3 image regions to be processed shown in the right image in.
[0142] Then, the electronic device can extract the image features of each image region to be processed, perform clustering processing on the image regions to be processed based on the similarity between the image features of every two image regions to be processed, and obtain the image regions to be processed that belong to the same foreground object as the foreground image regions corresponding to the foreground object.
[0143] For example, the electronic device can calculate the similarity of the image features of every two image regions to be processed to obtain a similarity matrix containing each similarity. Then, the electronic device can perform decomposition processing on the similarity matrix based on the RNMF (Robust Nonnegative Matrix Factorization) algorithm to obtain the image regions to be processed that belong to the same foreground object as the foreground image regions corresponding to the foreground object.
[0144] Exemplarily, refer to Figure 7 , Figure 7 The images to be processed shown are based on different mounted cameras. Based on Figure 7 the left image in, the image region to be processed containing the foreground object represented by ID1 (which can be called image region 1) and the image region to be processed containing the foreground object represented by ID2 (which can be called image region 2) can be obtained. Based on Figure 7 the right image in, the image region to be processed containing the foreground object represented by ID1 (which can be called image region 3) and the image region to be processed containing the foreground object represented by ID2 (which can be called image region 4) can be obtained.
[0145] Furthermore, based on the image similarity between the 4 image regions to be processed, the matching relationships of different instances can be obtained. The matching relationships include: Image region 1 and image region 3 are the foreground image regions of the foreground object represented by ID1, and image region 2 and image region 4 are the foreground image regions of the foreground object represented by ID2.
[0146] Based on the above processing, the foreground image region containing the foreground object can be extracted from the image to be processed, and then the three-dimensional pose plane of the foreground object can be determined. Furthermore, the specified foreground image region is projected onto the specified monitoring view through the three-dimensional pose plane of the foreground object, and the image to be projected of the foreground object under the specified monitoring view is obtained. Furthermore, the background image and the image to be projected under the specified monitoring view are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent, and can avoid the distortion of the foreground object in the generated image to a certain extent. Furthermore, the quality of the generated image is improved.
[0147] Regarding step S203, for each foreground object, the three-dimensional pose plane of the foreground object represents the plane where the foreground object is located in the target scene, and the three-dimensional pose plane of the foreground object can reflect the position of the foreground object in the target scene.
[0148] In one embodiment, on the basis of Figure 2 , refer to Figure 8 , step S203 may include the following steps:
[0149] S2031: Determine the two mounted cameras corresponding to the image to be processed that contain the foreground object as the specified mounted cameras.
[0150] S2032: For each specified mounted camera, determine the detection frame containing the foreground object in the image to be processed corresponding to the specified mounted camera.
[0151] S2033: Determine the three-dimensional point corresponding to the specified point in the detection frame in the target scene based on the first conversion relationship between the image coordinates and the spatial coordinates of the camera corresponding to the specified setup.
[0152] S2034: Determine the plane containing the optical center of the specified setup camera and the corresponding three-dimensional point as the reference plane corresponding to the specified setup camera.
[0153] S2035: Determine the plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0154] For each foreground object, the to-be-processed image corresponding to the setup camera containing the foreground object means that the setup camera can capture the foreground object. The electronic device can determine two setup cameras corresponding to the to-be-processed image containing the foreground object as the specified setup cameras. For example, the electronic device can determine two adjacent setup cameras that can capture the foreground object as the specified setup cameras.
[0155] For each specified setup camera, the electronic device can determine the detection frame containing the foreground object in the to-be-processed image corresponding to the specified setup camera. For example, in the Figure 6 embodiment, the black rectangle containing a foreground object is the detection frame of the foreground object. The specified points of the detection frame can include: two points (e.g., vertices) on the lower edge of the detection frame. The lower edge of the detection frame represents the position where the foreground object contacts the ground in the target scene. For example, if the foreground object is a pedestrian, the lower edge of the detection frame represents the position where the pedestrian's feet are in the target scene.
[0156] Then, the electronic device can perform coordinate transformation on the specified points in the detection frame according to the first conversion relationship between the image coordinates corresponding to the specified setup camera and the spatial coordinates of the target scene to obtain the coordinates of the corresponding three-dimensional points in the target scene. The electronic device can determine the plane (i.e., the reference plane) containing the determined three-dimensional points and the optical center of the specified setup camera. The image coordinates corresponding to the specified setup camera are: the image coordinates of the to-be-processed image obtained based on the specified setup camera.
[0157] When two specified mounted cameras simultaneously capture a point in space, according to the principles of multi-view geometry, the coordinates of the point in space can be accurately restored. By extending the two reference planes along the directions of the perspectives of the two specified mounted cameras, the intersection line of the two reference planes can be obtained. This intersection line represents the position where the foreground object contacts the ground in the target scene. For example, if the foreground object is a pedestrian, the lower edge of the detection frame represents the position of the pedestrian's feet in the target scene. The electronic device can determine a plane that passes through the intersection line of each reference plane and forms a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0158] The specified angle can be any angle between 0 degrees and 90 degrees, and the specific value of the specified angle can be determined based on actual requirements. For example, when the foreground object is a pedestrian, the pedestrian usually stands vertically on the ground, so the specified angle can be 90 degrees.
[0159] In one embodiment, step S2035 may include the following steps: taking the intersection line of each reference plane as the rotation axis, and rotating the initial object plane that passes through the rotation axis and is parallel to the horizontal plane by a specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0160] Exemplarily, referring to Figure 9 , Figure 9 in which Camera 1 and Camera 2 are the specified mounted cameras. The optical center of Camera 1 and the two vertices of the lower edge of the detection frame corresponding to Camera 1 form a reference plane, and the optical center of Camera 2 and the two vertices of the lower edge of the detection frame corresponding to Camera 2 form another reference plane. Two points on the intersection line (i.e., the object plane rotation axis) of each reference plane are denoted as A(x1, y1, z1) and B(x2, y2, z2). Based on the coordinates of point A and point B, the straight line equation of this intersection line (i.e., AB) can be obtained as:
[0161] Based on this straight line equation, the coordinates of any point on this intersection line can be obtained, denoted as M(x0, y0, z0), and the direction vector of this intersection line, denoted as S = [m, n, p].
[0162] Based on the direction vector of this intersection line and the coordinates of point M, the plane equation of the initial object plane that passes through this intersection line and is parallel to the horizontal plane can be obtained as: ax + by + c = 0, where a = n, b = -m, and c = -(nx0 - my0).
[0163] The plane equation of the three-dimensional pose plane of the foreground object can be expressed as: a′x + b′y + c′z + d′ = 0.
[0164] Where, R(s(θ)) represents the Rodriguez transform that rotates a specified angle around the rotation axis, θ represents the specified angle, and S represents the line equation of the rotation axis. represents a zero vector.
[0165] Based on the above processing, the three-dimensional pose plane of the foreground object can be determined. Furthermore, the specified foreground image region can be projected onto the specified monitoring view through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring view. Furthermore, the background image and the image to be projected under the specified monitoring view are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene and can, to a certain extent, avoid the distortion of the foreground object in the generated image. Furthermore, the quality of the generated image is improved.
[0166] For each foreground object, after obtaining the three-dimensional pose plane of the foreground object, the electronic device can map the foreground image region of the foreground object to the specified monitoring view through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring view.
[0167] Since the images to be processed obtained by the electronic device are multiple, the foreground image regions of the foreground object may also be multiple. In one implementation, for each foreground object, the electronic device can select one foreground image region from the multiple foreground image regions of the foreground object to obtain the specified foreground image region of the foreground object.
[0168] In another implementation, before step S204, the method may further include the following steps: determining, from each foreground image region, the image region with the highest integrity of the foreground object contained therein as the specified foreground image region.
[0169] Since the foreground object is a moving object, for each foreground object, during the movement of the foreground object, it may move from the shooting view of one installed camera to the shooting view of another installed camera. Then, the integrity of the foreground object in each foreground image region obtained based on different installed cameras may be different. To avoid, to a certain extent, problems such as misalignment and missing of the foreground object in the generated target image and improve the quality of the generated target image, the electronic device can determine, from each foreground image region of the foreground object, the image region with the highest integrity of the foreground object contained therein as the specified foreground image region of the foreground object.
[0170] For step S204, in one implementation, for each foreground object, the electronic device can obtain the coordinates of each pixel point (which can be referred to as the first pixel point) included in the specified foreground image area of the foreground object in the specified foreground image area.
[0171] Then, for each first pixel point in the specified foreground image area, the electronic device performs coordinate conversion on the first pixel point based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the coordinates of the three-dimensional point corresponding to the first pixel point in the three-dimensional pose plane of the foreground object.
[0172] Furthermore, the electronic device can perform coordinate conversion on the determined three-dimensional point based on the third conversion relationship between the image coordinates of the image captured from the specified monitoring perspective and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the pixel coordinates (which can be referred to as the first pixel coordinates) corresponding to the three-dimensional point from the specified monitoring perspective, that is, the first pixel coordinates corresponding to the first pixel point from the specified monitoring perspective. The electronic device can use the pixel value of the first pixel point as the pixel value of the corresponding first pixel coordinates, and can obtain the image to be projected of the foreground object from the specified monitoring perspective.
[0173] For each foreground object, the size of the image area captured from the specified monitoring perspective and the specified foreground image area of the foreground object may be different. For example, the image area captured from the specified monitoring perspective is larger than the specified foreground image area of the foreground object. Directly projecting the specified foreground image area of the foreground object to the specified monitoring perspective may cause some pixel coordinates in the image area captured from the specified monitoring perspective to not be able to match the corresponding pixel points in the specified foreground image area, and thus the pixel values of this type of pixel coordinates cannot be determined, resulting in the problem of pixel loss in the generated target image.
[0174] In another implementation, in order to avoid to a certain extent the problem that some pixel coordinates in the image area captured from the specified monitoring perspective cannot match the corresponding pixel points in the specified foreground image area, resulting in pixel loss in the generated target image and improving the quality of the generated target image, on the Figure 2 basis, referring to Figure 10 , before step S204, the method may further include the following steps:
[0175] S207: Determine the edge pixel points of the foreground object in the image to be processed where the specified foreground image area belongs.
[0176] S208: For each edge pixel point, based on the second conversion relationship and the third conversion relationship, determine the corresponding pixel coordinates of the edge pixel point from the specified monitoring perspective as the edge pixel coordinates.
[0177] S209: Determine the pixel coordinates included in the area with the edge pixel coordinates as the edge, and use them as the pixel coordinates of the foreground object under the specified monitoring view.
[0178] For each foreground object, the electronic device can determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area. The determined edge pixel points can represent the contour of the foreground object.
[0179] For each edge pixel point, the electronic device performs coordinate transformation on the edge pixel point based on the second transformation relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the coordinates of the three-dimensional point corresponding to the edge pixel point in the three-dimensional pose plane of the foreground object. Then, the electronic device can perform coordinate transformation on the determined three-dimensional point based on the third transformation relationship between the image coordinates imaged under the specified monitoring view and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the pixel coordinates corresponding to the three-dimensional point under the specified monitoring view, that is, to obtain the edge pixel coordinates of the edge pixel point under the specified monitoring view. The area composed of the determined edge pixel coordinates can represent the contour of the foreground object under the specified monitoring view.
[0180] Furthermore, the electronic device can determine the area with the edge pixel coordinates as the edge. This area is the area occupied by the foreground object under the specified monitoring view, and the pixel coordinates within this area are the pixel coordinates of the foreground object under the specified monitoring view.
[0181] For each pixel coordinate (which can be called the second pixel coordinate) of the foreground object under the specified monitoring view, the electronic device can perform coordinate transformation on the second pixel coordinate based on the third transformation relationship, to obtain the coordinates of the three-dimensional point corresponding to the second pixel coordinate in the three-dimensional pose plane of the foreground object. The second pixel coordinates of the foreground object under the specified monitoring view are also the pixel coordinates in the area with the edge pixel coordinates as the edge.
[0182] Then, the electronic device can perform coordinate transformation on the determined three-dimensional point based on the second transformation relationship, to obtain the pixel point (which can be called the second pixel point) corresponding to the three-dimensional point in the specified foreground image area, that is, the second pixel point corresponding to the second pixel coordinate in the specified foreground image area. The electronic device can use the pixel value of the determined second pixel point as the pixel value of the second pixel coordinate, and can obtain the to-be-projected image of the foreground object under the specified monitoring view.
[0183] Exemplarily, refer to Figure 11, the estimated object plane is the three-dimensional pose plane of the foreground object. The installed camera is the installed camera corresponding to the image to be processed to which the specified foreground image area of the foreground object belongs. The monitoring camera is a virtual camera, and the viewing angle represented by the monitoring camera is the specified monitoring viewing angle.
[0184] For each second pixel coordinate of the foreground object at the specified monitoring viewing angle, the second pixel coordinate can be denoted as The spatial depth value of the second pixel coordinate is denoted as g1, and the spatial depth value represents the distance between the monitoring camera and the three-dimensional point corresponding to the second pixel coordinate in the target scene.
[0185] The internal parameters of the monitoring camera are denoted as K, the external parameters are denoted as [R|t], the internal parameters of the installed camera are denoted as K′, and the external parameters are denoted as [R′|t′]. The electronic device can obtain the coordinates of the three-dimensional point ( Figure 11 the point represented by the pentagram in) corresponding to the second pixel coordinate in the three-dimensional pose plane of the foreground object based on the internal parameters, external parameters of the monitoring camera and the following formula, and denote it as
[0186]
[0187]
[0188] Then, the electronic device can obtain the second pixel point corresponding to the second pixel coordinate in the specified foreground image area based on the internal parameters, external parameters of the installed camera and the following formula, and denote it as The spatial depth value of the second pixel point is denoted as g′1, and the spatial depth value represents the distance between the installed camera and the three-dimensional point corresponding to the second pixel point in the target scene.
[0189]
[0190] The electronic device can obtain the second pixel point in the specified foreground image area of the pixel value, as the pixel value of the second pixel coordinate at the specified monitoring viewing angle, and can obtain the image to be projected of the foreground object at the specified monitoring viewing angle.
[0191] For step S205, in one implementation, the electronic device can obtain the image containing the background in the target scene pre-captured by the installed camera corresponding to the specified foreground image area as the background image.
[0192] Since the pre - acquired image and the specified foreground image region are not acquired at the same time, and the illumination in the target scene at different times is different, there may be a color difference between the pre - acquired image and the specified foreground image region. Directly obtaining the pre - acquired image as the background image may cause a color difference between the background and the foreground object in the finally generated target image.
[0193] In another implementation, in order to avoid, to a certain extent, the color difference between the background and the foreground object in the finally generated target image and improve the quality of the generated target image, on the basis of Figure 2 , referring to Figure 12 , step S205 may include the following steps:
[0194] S2051: Obtain the image to be filled.
[0195] Among them, the image to be filled is obtained by deleting the specified foreground image region from the to - be - processed image to which it belongs.
[0196] S2052: Fill the part of the preset image corresponding to the specified foreground image region into the position corresponding to the specified foreground image region in the image to be filled, to obtain an image of the background in the target scene collected by the camera corresponding to the specified foreground image region.
[0197] Among them, the preset image is a pre - acquired image containing the background in the target scene.
[0198] The electronic device can determine the image obtained by deleting the specified foreground image from the to - be - processed image as the image to be filled, and the position corresponding to the specified foreground image region in the image to be filled is the region to be filled. The specified foreground image region and the image to be filled belong to the same to - be - processed image, and there is no color difference between the specified foreground image region and the image to be filled due to different illuminations.
[0199] Then, the electronic device can obtain the preset image pre - acquired containing the background in the target scene and determine the part corresponding to the specified foreground image region in the preset image. The electronic device can fill the part of the preset image corresponding to the specified foreground image region into the region to be filled in the image to be filled based on the Poisson filling algorithm, and can obtain the background image. There is no color difference between the obtained background image and the specified foreground image region, which can, to a certain extent, avoid the color difference between the background and the foreground object in the finally generated target image and improve the quality of the generated target image.
[0200] Referring to Figure 13 , Figure 13 The left image in
[0201] After the electronic device extracts a specified foreground image area from the to-be-processed image on the left side in Figure 13 it obtains the to-be-filled image shown in the middle image, where the black area is the to-be-filled area. The electronic device fills the to-be-filled area in the to-be-filled image based on a preset image and a Poisson filling algorithm, and obtains Figure 13 the background image shown in the image on the right side in Figure 13 .
[0202] After obtaining the background image, the electronic device can also map the background image to a specified monitoring perspective to obtain an image of the background in the target scene under the specified monitoring perspective.
[0203] In one implementation, for each pixel coordinate in the image area imaged under the specified monitoring perspective, the electronic device can determine the three-dimensional point corresponding to the pixel coordinate in the target scene based on the conversion relationship between the image coordinates imaged under the specified monitoring perspective and the spatial coordinates of the target scene. Then, the electronic device can determine the pixel point corresponding to the three-dimensional point in the background image based on the conversion relationship between the image coordinates of the background image and the spatial coordinates of the target scene, and obtain the pixel point corresponding to the pixel coordinate in the background image. The electronic device can use the pixel value of the pixel point corresponding to the pixel coordinate in the background image as the pixel value of the pixel coordinate to obtain an image of the background image under the specified monitoring perspective.
[0204] Exemplarily, referring to Figure 14 , the erected camera is the camera for collecting the background image, the monitoring camera is a virtual camera, and the perspective represented by the monitoring camera is the specified monitoring perspective. The internal parameter of the monitoring camera is denoted as K, the external parameter is denoted as [R|t], the internal parameter of the erected camera is denoted as K′, and the external parameter is denoted as [R′|t′].
[0205] For each pixel coordinate under the specified monitoring perspective, the pixel coordinate is denoted as g2 represents the spatial depth value of the pixel coordinate. The electronic device can obtain the coordinates of the three-dimensional point ( Figure 14 the model point represented by the pentagram in
[0206]
[0207] corresponding to the pixel coordinate in the target scene based on the internal parameter, external parameter of the monitoring camera and the following formula, and denote it as g′2 represents the spatial depth value of the pixel point.
[0208]
[0209] For step S206, the electronic device can first project the image of the background image at the specified monitoring perspective onto the three-dimensional scene model of the target scene, and then project the images to be projected of each foreground object onto the three-dimensional scene model to obtain the target image of the target scene at the specified monitoring perspective. Projecting the background image onto the three-dimensional scene model first can, to a certain extent, avoid the problem that the background image obscures the image to be projected, resulting in misalignment and missing foreground objects in the generated target image, and improve the quality of the generated target image.
[0210] In one implementation, the electronic device can use the background image as a texture and add the background image to the corresponding position in the three-dimensional scene model in the way of texture mapping. Then, the electronic device can use the image to be projected as a texture and add each image to be projected to the corresponding position in the three-dimensional scene model in the way of texture mapping to obtain the target image of the target scene at the specified monitoring perspective.
[0211] In addition, after obtaining the target image of the target scene at the specified monitoring perspective, the electronic device can also render according to the target image to display the target image, and the user can view the target image of the target scene at the specified monitoring perspective.
[0212] Exemplarily, see Figure 15 , Figure 15 In the figure, the left image is the three-dimensional scene model of the target scene, and the left detection box in the right image is the background part of the three-dimensional model of the target scene, and the right detection box is the target image of the target scene at the specified monitoring perspective. It can be seen that there are no problems such as distortion, missing, and misalignment of the moving foreground objects in the target image obtained by the image generation method provided in the embodiments of the present application, that is, the quality of the generated image can be improved.
[0213] In one embodiment, the electronic device can also obtain the three-dimensional scene model of the target scene. The three-dimensional scene model can be a file in formats such as osgb format, s3c format, max format, fbx format, and obj format. The three-dimensional scene model contains data such as point cloud, mesh, and texture. The point cloud represents the three-dimensional points corresponding to each position in the target scene, and each three-dimensional point can be represented by three-dimensional coordinates. Based on the spatial relationship between the positions corresponding to the three-dimensional points in the target scene, the three-dimensional points are connected to obtain multiple meshes composed of the three-dimensional points. Based on the materials of the objects in the target scene, texture mapping is performed on each mesh composed of the three-dimensional points to obtain the three-dimensional scene model of the target scene.
[0214] See Figure 16 , Figure 16The left image in the middle is a three-dimensional point cloud model, the middle image represents a three-dimensional mesh model formed by connecting each three-dimensional point in the three-dimensional point cloud model, and the right image represents a three-dimensional scene model obtained by performing texture mapping on each mesh in the three-dimensional mesh model.
[0215] See Figure 17 , Figure 17 which is a flowchart of a method for obtaining a three-dimensional scene model provided by an embodiment of the present application.
[0216] Technicians can conduct on-site surveys of the target scene to obtain scene data of the target scene. For example, the size of the target scene, the static objects included in the target scene, the positions of the static objects, and the distances between the positions of the static objects.
[0217] Then, select a modeling solution, that is, determine the method of three-dimensional modeling. For example, NURBS (Non-Uniform Rational B-Splines) modeling, polygon modeling, etc. Furthermore, perform three-dimensional scanning, that is, use a three-dimensional scanning device (such as a radar scanning device, a laser scanning device, etc.) to perform three-dimensional scanning on the target scene to obtain a three-dimensional point cloud model of the target scene.
[0218] Furthermore, technicians perform manual modeling, that is, based on the spatial relationship between the positions corresponding to each three-dimensional point in the target scene, connect each three-dimensional point in the three-dimensional point cloud model to obtain a three-dimensional mesh model. Based on the materials of the objects in the target scene, perform texture mapping on each mesh in the three-dimensional mesh model to obtain a three-dimensional scene model of the target scene. Technicians can also perform model verification, that is, detect whether the three-dimensional scene model conforms to the target scene. For example, detect whether the positional relationship between static objects is correct.
[0219] In one embodiment, when installing multiple mounted cameras in the target scene, when installing multiple mounted cameras, the mounting parameters of the mounted cameras can be adjusted so that the multiple mounted cameras meet the following constraints:
[0220] At any moment, for each foreground object in the target scene, at least one of the images captured by each mounted camera contains the complete foreground object.
[0221] Based on this, a complete image of the foreground object can be obtained, improving the integrity of the foreground object in the generated target image, and thus improving the quality of the generated image.
[0222] Exemplarily, see Figure 18 , Figure 18Schematic diagram of a camera mounting method provided by an embodiment of this application. Among them, different triangles represent different mounted cameras. A triangle corresponding to a dotted area represents the range of the field of view of the mounted camera. Each mounted camera can be mounted in a straight line (i.e., Figure 18 the upper straight line in
[0223] Each mounted camera includes a mounted camera facing left and a mounted camera facing right. When the specified monitoring perspective includes the left area in the target scene, the mounted camera facing left is enabled to obtain an image; when the specified monitoring perspective includes the right area in the target scene, the mounted camera facing right is enabled to obtain an image.
[0224] The mounting height of the mounted camera is 2 to 6 times the height of the foreground object in the target scene, so as to ensure that the mounted camera can capture a complete and clear image of the foreground object in the target scene. When the mounted camera is installed on the ceiling or support frame in the target scene, the mounting height of the mounted camera can be determined by technicians through manual measurement. Since the heights of different foreground objects are different, the height of the foreground object can be set according to actual needs. For example, when the foreground object is a person, the height of the foreground object can be set to 2 meters, but it is not limited to this.
[0225] The initial posture of the mounted camera (for example, the pitch angle of the camera) cannot exceed the adjustable range of the mounted camera itself. And, the field of view of the mounted camera should cover the area it shoots in the target scene, so that the mounted camera can capture a complete image of the foreground object in the target scene and there is no dead angle that cannot be shot.
[0226] For two adjacent mounted cameras facing the same direction, the overlapping area of the fields of view of the two mounted cameras accounts for 10% to 30% of the field of view of a single mounted camera, and the included angle between the optical axes of the two mounted cameras is small. For example, the included angle between the optical axes of the two mounted cameras is less than 30 degrees.
[0227] See Figure 19 , Figure 19 In
[0228] b ≤ (H - h)cot(α - θ) - Hcot(α + θ) (7)
[0229] b represents the distance between the two mounted cameras, H represents the mounting height of the mounted camera C1, h represents the height of the foreground object in the target scene, 2θ represents the field of view angle of the mounted camera C1, and α represents the pitch angle of the mounted camera C1.
[0230] When installed in the above - mentioned manner, it can ensure that at each moment, the foreground object moving in the target scene is photographed with a complete image by at least one installed camera. Furthermore, the integrity of the foreground object in the generated target image is relatively high, which can further improve the quality of the generated image.
[0231] In one embodiment, the installed camera can be installed in the target scene through a control device, and the control device can communicate with an electronic device.
[0232] The electronic device can obtain the current installation parameters of each installed camera in the target scene and determine whether each installed camera meets the preset constraint conditions based on the current installation parameters. The installation parameters of the installed camera can include: the installation height of the installed camera, the pitch angle, the field - of - view angle, the height of the foreground object in the target scene, and the distance between the installed cameras, etc.
[0233] For example, for two adjacent installed cameras facing the same direction, it is determined whether the overlapping area of the fields of view of the two installed cameras accounts for 10% to 30% of the field of view of a single installed camera, and whether the installation height, pitch angle, field - of - view angle of the two installed cameras, the height of the foreground object in the target scene, and the distance between the two installed cameras meet the above formula (7).
[0234] If the above - mentioned constraint conditions are met based on the current installation parameters of each installed camera, the electronic device can obtain the initial images of the target scene collected by each installed camera, and then obtain the image to be processed based on the obtained initial images.
[0235] If the above - mentioned constraint conditions are not met based on the current installation parameters of each installed camera, the electronic device can send a control instruction to the control device. When receiving the control instruction, the control device can adjust the installation parameters of the installed camera. For example, the control device can control the installed camera to move to adjust the installation height of the installed camera, or the control device can control the installed camera to rotate to adjust the pitch angle of the installed camera, etc., so that each installed camera meets the constraint conditions based on the adjusted installation parameters of each installed camera.
[0236] When each installed camera meets the constraint conditions, the electronic device can obtain the initial images of the target scene collected by each installed camera, and then obtain the image to be processed based on the obtained initial images.
[0237] See Figure 20 , Figure 20 which is a flowchart of an image generation method provided by an embodiment of this application.
[0238] For the surveillance video scenario, the electronic device can obtain the image after video decoding, that is, the electronic device obtains the initial image containing the target scenario captured by the camera set up at the current moment. The electronic device can correct image distortion and color, that is, the electronic device performs color difference correction on each initial image according to the specified color difference correction parameters, and / or performs distortion correction on each initial image according to the distortion types of the cameras set up, to obtain each image to be processed.
[0239] The electronic device can detect moving objects (i.e., the foreground objects in the foregoing embodiments) in the image, and segment the corresponding instances to obtain the foreground layer and the background layer of the camera set up. That is, the electronic device performs image segmentation on each image to be processed, obtains the image region to be processed containing the foreground object, and determines the image to be filled.
[0240] For the foreground layer of the camera set up, the electronic device can match the object instances between cameras and establish a matching relationship. The matching relationship includes the object ID and the detection box, and the foreground objects with the same ID are the same foreground object. That is, the electronic device determines the foreground image regions belonging to the same foreground object based on the image similarity between the image regions to be processed.
[0241] The electronic device estimates the three-dimensional pose plane of the object to obtain the plane parameters. That is, for each foreground object, the electronic device determines the three-dimensional pose plane of the foreground object in the target scenario based on the foreground image regions of the foreground object and the first conversion relationships between the image coordinates of each image to be processed and the spatial coordinates of the target scenario. The plane parameters are the plane equations of the three-dimensional pose plane.
[0242] The electronic device can select the object with the highest segmentation integrity and map it to the perspective of the monitoring camera using the three-dimensional pose plane to obtain the corrected foreground layer of the monitoring camera. That is, for each foreground object, the electronic device selects the image with the largest integrity of the foreground object to obtain the specified foreground image region of the foreground object, and maps the specified foreground image region of the foreground object to the specified monitoring perspective through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring perspective.
[0243] For the background layer of the camera set up, the electronic device can perform background hole filling to obtain the corrected background layer of the monitoring camera. That is, the electronic device fills the part of the preset image corresponding to the specified foreground image region into the position corresponding to the specified foreground image region in the image to be filled to obtain the background image containing the background in the target scenario captured by the camera set up corresponding to the specified foreground image region.
[0244] The electronic device can map the foreground layer after calibration of the monitoring camera and the background layer after calibration of the erected camera to the three-dimensional scene model of the target scene to obtain a two-dimensional base map to be processed. That is, the electronic device projects the image of the background image at the specified monitoring angle and the image to be projected onto the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene at the specified monitoring angle.
[0245] Based on the above processing, the three-dimensional pose plane is the plane where the foreground object is located in the target scene. The specified foreground image area is projected onto the specified monitoring angle through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object at the specified monitoring angle. Furthermore, the background image and the image to be projected at the specified monitoring angle are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent and avoid the distortion of the foreground object in the generated image. Furthermore, the quality of the generated image is improved.
[0246] See Figure 21 , Figure 21 FIG. is a flowchart of an image generation method provided by an embodiment of the present application. Camera erection refers to installing multiple erected cameras in the target scene according to the specified camera erection method. Camera bitstream acquisition is performed through the erected cameras, that is, the initial image of the target scene at the current moment is captured by the erected cameras, the image to be processed is obtained based on the acquired initial image, and the specified foreground image area of the foreground object is obtained based on the image to be processed.
[0247] Three-dimensional scene modeling refers to performing three-dimensional modeling on the target scene to obtain the three-dimensional scene model of the target scene.
[0248] The electronic device can obtain the internal and external calibration parameters of the system and fuse the parameters. The internal and external calibration parameters of the system refer to the internal and external parameters of the erected cameras, as well as the internal and external parameters of the monitoring cameras, etc. Fusing the parameters means that the electronic device performs image fusion on the image to be processed obtained based on the erected cameras and the monitoring cameras based on the internal and external parameters of the erected cameras, the internal and external parameters of the monitoring cameras, etc. to obtain the fused image. That is, the electronic device maps the specified foreground image area of the foreground object to the specified monitoring angle of the monitoring camera through the three-dimensional pose plane of the foreground object based on the internal and external parameters of the erected cameras, the internal and external parameters of the monitoring cameras, etc. The electronic device projects the image of the background image containing the background in the target scene at the specified monitoring angle and the image to be projected onto the three-dimensional scene model in sequence to obtain the target image of the target scene at the specified monitoring angle (i.e., the fused image of the erected camera + the monitoring camera).
[0249] Then, the electronic device can perform rendering and display to obtain the fused image of the monitoring camera. That is, it renders according to the target image of the target scene from the specified monitoring perspective to display the target image.
[0250] For the monitoring video scene, cameras are set up to collect the initial images containing the target scene in real time, and the collected initial images are stored at a specified storage location. The monitoring camera perspective video stream address is the address of the storage location of the initial image. The electronic device can obtain the initial images collected by the set-up cameras at each moment. Based on the method provided in the embodiments of the present application, the initial images are processed to obtain images of the target scene from different monitoring perspectives, which can achieve real-time and all-round monitoring of the target scene.
[0251] See Figure 22 , Figure 22 which is a comparison diagram of a target image provided by the embodiments of the present application. Figure 22 In the figure, the left image is the target image of the target scene from the specified monitoring perspective generated based on the related technology, and the right image is the target image of the target scene from the specified monitoring perspective generated based on the method provided by the embodiments of the present application. Figure 22 In the left image, there are obvious distortion, dislocation, and missing problems with the pedestrians. Figure 22 In the right image in the figure, there are no problems of distortion, dislocation, and missing with the pedestrians. It can be seen that the image method provided by the embodiments of the present application can improve the quality of the generated images.
[0252] In at least one embodiment of the present application, the images in the video streams captured by multiple set-up cameras can all be processed in the above manner to obtain target images. The images in the video stream captured in a continuous time period, after the above processing, the obtained target images can form a more complete and clear high-quality video stream.
[0253] Corresponding to the method embodiment of Figure 2 , see Figure 23 , Figure 23 which is a structural diagram of an image generation device provided by the embodiments of the present application. The device includes:
[0254] A to-be-processed image acquisition module 2301, configured to acquire the to-be-processed images corresponding to each set-up camera in the target scene; wherein, the to-be-processed image corresponding to one set-up camera is obtained based on the image of the target scene collected by the set-up camera.
[0255] A foreground image region determination module 2302, configured to respectively determine the image regions occupied by the foreground objects in each to-be-processed image as the foreground image regions corresponding to the foreground objects based on image segmentation of each to-be-processed image.
[0256] A three-dimensional pose plane determination module 2303, configured to determine a three-dimensional pose plane of the foreground object in the target scene based on each foreground image region and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene;
[0257] A projected image generation module 2304, configured to, for each pixel coordinate of the foreground object under a specified monitoring view angle, determine the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image region based on a second conversion relationship between the image coordinates in the specified foreground image region and the three-dimensional coordinates in the three-dimensional pose plane, and a third conversion relationship between the image coordinates imaged under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, and use the determined pixel value as the pixel value of the pixel coordinate, so as to obtain a projected image of the foreground object under the specified monitoring view angle;
[0258] A background image acquisition module 2305, configured to acquire an image containing the background in the target scene collected by an installed camera corresponding to the specified foreground image region as a background image;
[0259] A target image generation module 2306, configured to project the image of the background image under the specified monitoring view angle and the projected image onto a three-dimensional scene model of the target scene in sequence, so as to obtain a target image of the target scene under the specified monitoring view angle.
[0260] Optionally, the three-dimensional pose plane determination module 2303 is specifically configured to determine two installed cameras corresponding to the image to be processed that contain the foreground object as specified installed cameras;
[0261] For each specified installed camera, determine a detection frame containing the foreground object in the image to be processed corresponding to the specified installed camera;
[0262] Based on the first conversion relationship between the image coordinates corresponding to the specified installed camera and the spatial coordinates, determine the three-dimensional point corresponding to the specified point in the detection frame in the target scene;
[0263] Determine a plane containing the optical center of the specified installed camera and the corresponding three-dimensional point as a reference plane corresponding to the specified installed camera;
[0264] Determine a plane passing through the intersection line of each reference plane and making a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
[0265] Optionally, the three-dimensional pose plane determination module 2303 is specifically configured to use the intersection line of each reference plane as the rotation axis, and rotate the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
[0266] Optionally, the device further includes:
[0267] The edge pixel point determination module is configured to, before the projected image generation module 2304 executes to determine the pixel value of the pixel point corresponding to each pixel coordinate of the foreground object in the specified monitoring view in the specified foreground image area based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates imaged in the specified monitoring view and the three-dimensional coordinates in the three-dimensional pose plane, and use the determined pixel value as the pixel value of the pixel coordinate to obtain the projected image of the foreground object in the specified monitoring view, determine the edge pixel points of the foreground object in the to-be-processed image belonging to the specified foreground image area;
[0268] The edge pixel point mapping module is configured to, for each edge pixel point, determine the corresponding pixel coordinate of the edge pixel point in the specified monitoring view based on the second conversion relationship and the third conversion relationship as the edge pixel coordinate;
[0269] The pixel coordinate determination module is configured to determine the pixel coordinates included in the area with the edge pixel coordinates as the edge as the pixel coordinates of the foreground object in the specified monitoring view.
[0270] Optionally, the device further includes:
[0271] The foreground image area selection module is configured to, before the projected image generation module 2304 executes to determine the pixel value of the pixel point corresponding to each pixel coordinate of the foreground object in the specified monitoring view based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates imaged in the specified monitoring view and the three-dimensional coordinates in the three-dimensional pose plane, and use the determined pixel value as the pixel value of the pixel coordinate to obtain the projected image of the foreground object in the specified monitoring view, determine, from each foreground image area, the image area with the largest integrity of the included foreground object as the specified foreground image area.
[0272] Optionally, the foreground image region determination module 2302 is specifically configured to perform image segmentation on each image to be processed, obtain the image regions occupied by each foreground object in each image to be processed as the regions of the images to be processed, and determine the regions of the images to be processed that belong to the same foreground object based on the image similarity between the regions of the images to be processed as the foreground image regions corresponding to the foreground object.
[0273] Based on the image similarity between the regions of the images to be processed, determine the regions of the images to be processed that belong to the same foreground object as the foreground image regions corresponding to the foreground object.
[0274] Optionally, the image acquisition module 2301 to be processed is specifically configured to acquire the images of the target scene collected by each installed camera as the initial images, perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of each installed camera to obtain each image to be processed.
[0275] Perform color difference correction on each initial image according to the specified color difference correction parameters, and / or perform distortion correction on each initial image according to the distortion types of each installed camera to obtain each image to be processed.
[0276] Optionally, the background image acquisition module 2305 is specifically configured to acquire an image to be filled, where the image to be filled is obtained by deleting the specified foreground image region from the image to be processed, and fill the part of the preset image corresponding to the specified foreground image region to the position corresponding to the specified foreground image region in the image to be filled to obtain an image of the background in the target scene collected by the installed camera corresponding to the specified foreground image region, where the preset image is an image of the background in the target scene collected in advance.
[0277] Fill the part of the preset image corresponding to the specified foreground image region to the position corresponding to the specified foreground image region in the image to be filled to obtain an image of the background in the target scene collected by the installed camera corresponding to the specified foreground image region, where the preset image is an image of the background in the target scene collected in advance.
[0278] Optionally, at any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera contains the complete foreground object.
[0279] Based on the image generation device provided in the embodiment of the present application, the three-dimensional pose plane is the plane where the foreground object is located in the target scene. The specified foreground image region is projected onto the specified monitoring view through the three-dimensional pose plane of the foreground object to obtain the image to be projected of the foreground object under the specified monitoring view. Then, the background image and the image to be projected under the specified monitoring view are mapped to the three-dimensional scene model of the target scene. That is, the foreground object is mapped to the three-dimensional pose plane where it is located in the target scene, which can reflect the position of the foreground object in the target scene to a certain extent, avoid the distortion of the foreground object in the generated image, and improve the quality of the generated image.
[0280] The embodiment of the present application also provides an electronic device, such as Figure 24As shown, it includes a processor 2401, a communication interface 2402, a memory 2403, and a communication bus 2404. Among them, the processor 2401, the communication interface 2402, and the memory 2403 complete mutual communication through the communication bus 2404.
[0281] The memory 2403 is used to store computer programs.
[0282] When the processor 2401 is used to execute the program stored on the memory 2403, the following steps are implemented:
[0283] Obtain the to-be-processed images corresponding to each installed camera in the target scene; among them, the to-be-processed image corresponding to one installed camera is obtained based on the image of the target scene collected by this installed camera.
[0284] Based on image segmentation of each to-be-processed image, respectively determine the image area occupied by the foreground object in each to-be-processed image as the foreground image area corresponding to the foreground object.
[0285] Based on each foreground image area and each first conversion relationship between the image coordinates of each to-be-processed image and the spatial coordinates of the target scene, determine the three-dimensional pose plane of the foreground object in the target scene.
[0286] For each pixel coordinate of the foreground object under a specified monitoring view angle, based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates of the image formed under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, and obtain the to-be-projected image of the foreground object under the specified monitoring view angle.
[0287] Obtain the image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image.
[0288] Project the image of the background image under the specified monitoring view angle and the to-be-projected image to the three-dimensional scene model of the target scene in sequence to obtain the target image of the target scene under the specified monitoring view angle.
[0289] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0290] The communication interface is used for communication between the above electronic device and other devices.
[0291] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0292] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0293] In another embodiment provided by this application, a computer-readable storage medium is also provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the steps of any of the above image generation methods are implemented.
[0294] In another embodiment provided by this application, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute any of the image generation methods in the above embodiments.
[0295] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0296] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0297] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, electronic device, computer-readable storage medium, and computer program product, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0298] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. An image generation method, characterized in that, The method includes: Obtaining the images to be processed corresponding to each installed camera in the target scene; wherein, the image to be processed corresponding to one installed camera is obtained based on the image of the target scene collected by this installed camera; Based on performing image segmentation on each image to be processed, respectively determining the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object; Based on each foreground image area and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene, determining the three-dimensional pose plane of the foreground object in the target scene; For each pixel coordinate of the foreground object under the specified monitoring view angle, based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates of the image formed under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, determining the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate, and obtaining the image to be projected of the foreground object under the specified monitoring view angle; the specified foreground image area is one of each foreground image area; Obtaining the image containing the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image; Sequentially projecting the image of the background image under the specified monitoring view angle and the image to be projected onto the three-dimensional scene model of the target scene to obtain the target image of the target scene under the specified monitoring view angle.
2. The method according to claim 1, characterized in that, The determining the three-dimensional pose plane of the foreground object in the target scene based on each foreground image area and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene includes: Determining two installed cameras whose corresponding images to be processed contain the foreground object as the specified installed cameras; For each specified installed camera, determining the detection frame containing the foreground object in the image to be processed corresponding to this specified installed camera; Based on the first conversion relationship between the image coordinates corresponding to this specified installed camera and the spatial coordinates, determining the three-dimensional point corresponding to the specified point in the detection frame in the target scene; Determining the plane containing the optical center of this specified installed camera and the corresponding three-dimensional point as the reference plane corresponding to this specified installed camera; Determining the plane passing through the intersection line of each reference plane and forming a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene.
3. The method according to claim 2, wherein The determining the plane passing through the intersection line of each reference plane and forming a specified angle with the horizontal plane as the three-dimensional pose plane of the foreground object in the target scene includes: Taking the intersection line of each reference plane as the rotation axis and rotating the initial object plane passing through the rotation axis and parallel to the horizontal plane by the specified angle to obtain the three-dimensional pose plane of the foreground object in the target scene.
4. The method according to claim 1, characterized in that Before obtaining the projected image of the foreground object at the specified monitoring view angle, by determining, for each pixel coordinate of the foreground object at the specified monitoring view angle, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane and the third conversion relationship between the image coordinates imaged at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, the method further includes: Determining the edge pixel points of the foreground object in the to-be-processed image to which the specified foreground image area belongs; For each edge pixel point, determining the corresponding pixel coordinate of the edge pixel point at the specified monitoring view angle based on the second conversion relationship and the third conversion relationship as the edge pixel coordinate; Determining the pixel coordinates included in the area with the edge pixel coordinate as the edge as the pixel coordinates of the foreground object at the specified monitoring view angle.
5. The method according to claim 1, wherein Before obtaining the projected image of the foreground object at the specified monitoring view angle, by determining, for each pixel coordinate of the foreground object at the specified monitoring view angle, the pixel value of the pixel point corresponding to the pixel coordinate in the specified foreground image area as the pixel value of the pixel coordinate based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane and the third conversion relationship between the image coordinates imaged at the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, the method further includes: Determining, from each foreground image area, the image area with the highest integrity of the foreground object included therein as the specified foreground image area.
6. The method according to claim 5, wherein The determining, by performing image segmentation on each to-be-processed image, of the image area occupied by the foreground object in each to-be-processed image as the foreground image area corresponding to the foreground object includes: Performing image segmentation on each to-be-processed image to obtain the image areas occupied by each foreground object in each to-be-processed image as the to-be-processed image areas; Determining, based on the image similarity between the to-be-processed image areas, the to-be-processed image areas belonging to the same foreground object as the foreground image area corresponding to the foreground object.
7. The method according to claim 1, characterized in that The obtaining of the to-be-processed images corresponding to each installed camera in the target scene includes: Obtaining the images of the target scene collected by each installed camera as the initial images; Performing chromatic aberration correction on each initial image according to the specified chromatic aberration correction parameters, and / or performing distortion correction on each initial image according to the distortion types of the installed cameras to obtain each to-be-processed image.
8. The method according to claim 1, wherein The obtaining of the image including the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image includes: Obtaining the to-be-filled image; wherein the to-be-filled image is obtained by deleting the specified foreground image area from the to-be-processed image to which it belongs. Fill the part of the preset image corresponding to the specified foreground image area into the position corresponding to the specified foreground image area in the image to be filled, to obtain an image of the background in the target scene collected by the installed camera corresponding to the specified foreground image area; wherein, the preset image is an image collected in advance containing the background in the target scene.
9. The method according to claim 1, wherein At any moment, for each foreground object in the target scene, at least one of the images collected by each installed camera contains the complete foreground object.
10. An image generation device, characterized in that, The device includes: An image to be processed acquisition module, configured to acquire images to be processed corresponding to each installed camera in the target scene; wherein, an image to be processed corresponding to an installed camera is obtained based on the image of the target scene collected by this installed camera. A foreground image area determination module, configured to respectively determine the image area occupied by the foreground object in each image to be processed as the foreground image area corresponding to the foreground object based on image segmentation of each image to be processed. A three-dimensional pose plane determination module, configured to determine the three-dimensional pose plane of the foreground object in the target scene based on each foreground image area and each first conversion relationship between the image coordinates of each image to be processed and the spatial coordinates of the target scene. An image to be projected generation module, configured to, for each pixel coordinate of the foreground object under a specified monitoring view angle, determine the pixel value of the pixel point corresponding to this pixel coordinate in the specified foreground image area as the pixel value of this pixel coordinate based on the second conversion relationship between the image coordinates in the specified foreground image area and the three-dimensional coordinates in the three-dimensional pose plane, and the third conversion relationship between the image coordinates of the image formed under the specified monitoring view angle and the three-dimensional coordinates in the three-dimensional pose plane, to obtain the image to be projected of the foreground object under the specified monitoring view angle; the specified foreground image area is one of each foreground image area. A background image acquisition module, configured to acquire an image of the background in the target scene collected by the installed camera corresponding to the specified foreground image area as the background image. A target image generation module, configured to project the image of the background under the specified monitoring view angle and the image to be projected onto the three-dimensional scene model of the target scene in sequence, to obtain the target image of the target scene under the specified monitoring view angle.
Citation Information
Patent Citations
Live broadcast image synthesis method and device, terminal equipment and readable storage medium
CN113837979A
Panoramic image synthesis device, panoramic image synthesis method and panoramic image synthesis program
US20220180475A1