Image fusion method, device, storage medium and electronic device
By determining the calibration parameters and pixel point projection coordinates of the target camera device in a three-dimensional scene, and removing pixel points that are beyond the range and blocked, the high-quality integration of the three-dimensional virtual scene and real-time video is achieved, and the problem of difficulty in fusion in the prior art is solved.
Patent Information
- Application Number
- CN202411440090.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2044-10-15
AI Technical Summary
In the prior art, it is difficult to integrate three-dimensional virtual scenes with real-time videos with high quality.
By determining the calibration parameters of the target imaging device in the three-dimensional scene and the projection coordinates of the collected target image pixel points, the pixel points that exceed the projection range of the imaging device and are blocked are eliminated, the target pixel points are obtained, and the target pixel points are rendered into the three-dimensional scene.
It realizes the perfect integration of video images and three-dimensional scenes, solving the problem of high-quality integration of three-dimensional virtual scenes and real-time videos.
Smart Images

Figure CN118967899B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of communications, and more particularly, to an image fusion method, apparatus, storage medium, and electronic device. Background Art
[0002] In the related art, in the process of achieving high-quality fusion display between a three-dimensional virtual scene and a real-time video, in most cases, it is very difficult for a three-dimensional scene modeled manually to be perfectly fused with a real video stream.
[0003] It can be seen that there is a problem in the related art that it is relatively difficult to achieve high-quality fusion between a three-dimensional virtual scene and a real-time video.
[0004] In view of the above problems existing in the related art, no effective solution has been proposed yet. Summary of the Invention
[0005] The embodiments of the present invention provide an image fusion method, apparatus, storage medium, and electronic device, so as to at least solve the problem in the related art that it is relatively difficult to achieve high-quality fusion between a three-dimensional virtual scene and a real-time video.
[0006] According to an embodiment of the present invention, there is provided an image fusion method, including: determining calibration parameters of a target imaging device in a three-dimensional scene; determining projection coordinates obtained by projecting each pixel point included in a target image captured by the target imaging device into the three-dimensional scene; based on the projection coordinates and the calibration parameters, removing first pixel points that exceed the projection range of the target imaging device among the pixel points to obtain second pixel points; removing third pixel points that are blocked in the three-dimensional scene among the second pixel points to obtain target pixel points; and rendering the target pixel points into the three-dimensional scene to fuse the target image captured by the target imaging device into the three-dimensional scene.
[0007] According to another embodiment of the present invention, there is provided an image fusion apparatus, including: a first determination module, configured to determine calibration parameters of a target imaging device in a three-dimensional scene; a second determination module, configured to determine projection coordinates obtained by projecting each pixel point included in a target image captured by the target imaging device into the three-dimensional scene; a first removal module, configured to remove first pixel points that exceed the projection range of the target imaging device among the pixel points based on the projection coordinates and the calibration parameters to obtain second pixel points; a second removal module, configured to remove third pixel points that are blocked in the three-dimensional scene among the second pixel points to obtain target pixel points; and a rendering module, configured to render the target pixel points into the three-dimensional scene to fuse the target image captured by the target imaging device into the three-dimensional scene.
[0008] According to another embodiment of the present invention, there is also provided a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any one of the above method embodiments when running.
[0009] According to another embodiment of the present invention, there is also provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0010] According to another embodiment of the present invention, there is also provided a computer program product, including a computer program, wherein the steps of the method described in each embodiment of the present application are implemented when the computer program is executed by a processor.
[0011] Through the present invention, in the case where the calibration parameters of the target imaging device in the three-dimensional scene and the projection coordinates corresponding to the pixel points of the acquired target image are determined, the pixel points blocked by the imaging range of the target imaging device and the pixel points beyond the imaging range are removed according to the calibration parameters and the projection coordinates to obtain target pixel points, and the target pixel points are rendered into the three-dimensional scene, so that the target image can be fused into the three-dimensional scene. On the premise of maximizing the retention of the original video picture, removing the pixel points blocked and outside the projection range can achieve the perfect fusion of the video picture and the three-dimensional scene. Therefore, the problem of relatively difficult high-quality fusion of the three-dimensional virtual scene and the real-time video in the related art can be solved, and the perfect fusion effect is achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 is a hardware structure block diagram of a mobile terminal for the image fusion method according to an embodiment of the present invention;
[0013] Figure 2 is a flowchart of the image fusion method according to an embodiment of the present invention;
[0014] Figure 3 is a schematic structural diagram of a target frustum according to an embodiment of the present invention;
[0015] Figure 4 is a schematic diagram of a camera coordinate system according to an embodiment of the present invention;
[0016] Figure 5 is a schematic diagram of a plane space coordinate system according to an embodiment of the present invention;
[0017] Figure 6 is a schematic diagram of depth data and projection distance according to an embodiment of the present invention;
[0018] Figure 7Flowchart of the image fusion method according to a specific embodiment of the present invention;
[0019] Figure 8 Effect diagram of removing occluded images according to an embodiment of the present invention;
[0020] Figure 9 Schematic diagram of the fused image according to an embodiment of the present invention;
[0021] Figure 10 Block diagram of the structure of the image fusion device according to an embodiment of the present invention. Specific implementation mode
[0022] In the following, embodiments of the present invention will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0023] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence.
[0024] The method embodiments provided in the embodiments of the present application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 Is the hardware block diagram of a mobile terminal of an image fusion method according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that, Figure 1 The structure shown in the figure is only for illustration and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0025] The memory 104 can be used to store computer programs, for example, software programs and modules of application software, such as the computer program corresponding to the image fusion method in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above-mentioned method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely provided with respect to the processor 102, and these remote memories can be connected to the mobile terminal through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and their combinations.
[0026] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0027] In this embodiment, a method for image fusion is provided. Figure 2 It is a flowchart of the image fusion method according to the embodiments of the present invention, as Figure 2 shown, and the process includes the following steps:
[0028] Step S202, determining the calibration parameters of the target camera device in the three-dimensional scene;
[0029] Step S204, determining the projection coordinates obtained by projecting each pixel point included in the target image collected by the target camera device into the three-dimensional scene;
[0030] Step S206, removing the first pixel points that exceed the projection range of the target camera device among the pixel points based on the projection coordinates and the calibration parameters to obtain second pixel points;
[0031] Step S208, removing the third pixel points that are blocked in the three-dimensional scene among the second pixel points to obtain target pixel points;
[0032] Step S210, rendering the target pixel points into the three-dimensional scene to fuse the target image collected by the target camera device into the three-dimensional scene.
[0033] In the above embodiments, the target imaging device may be a camera or a surveillance camera, but is not limited thereto. The calibration parameters in the three-dimensional scene may include: the coordinates CX, CY, CZ of the target imaging device, the rotation angles A, B, C of the target imaging device, the horizontal field of view FOV of the target imaging device, the maximum projection distance Maxdistance of the target imaging device, and the coordinates WP(X W , Y W , Z W ) of any point in the three-dimensional scene. The projection coordinates may be the coordinates corresponding to the target image pixel points collected by the target imaging device in the three-dimensional scene.
[0034] In the above embodiments, since the range captured by the camera at a certain angle (i.e., the projection range of the above target imaging device) is limited, the existence of image pixel points outside the shooting range will affect the quality of image fusion. Therefore, pixel points outside the projection range of the target imaging device (i.e., the above first pixel points) can be removed according to the calibration parameters and the projection coordinates to obtain second pixel points.
[0035] In the above embodiments, there are still pixel point information corresponding to the occluded objects (i.e., the above third pixel points) in the second pixel points. Rendering the target pixel points obtained after removing the third pixel points into the three-dimensional scene can achieve the fusion of the target image and the three-dimensional scene. Among them, the rendering may be two-dimensional rendering or three-dimensional rendering, but is not limited thereto.
[0036] Through the present invention, in the case of determining the calibration parameters of the target imaging device in the three-dimensional scene and the projection coordinates corresponding to the target image pixel points collected, the pixel points occluded by the target imaging device's shooting range and the pixel points outside the shooting range are removed according to the calibration parameters and the projection coordinates to obtain target pixel points. Rendering the target pixel points into the three-dimensional scene can fuse the target image into the three-dimensional scene. On the premise of maximizing the retention of the original video picture, removing the occluded and out-of-projection-range pixel points can achieve the perfect fusion of the video picture and the three-dimensional scene. Therefore, the problem of difficult high-quality fusion of three-dimensional virtual scenes and real-time videos in the related art can be solved, and thus the effect of perfect fusion is achieved.
[0037] Optionally, the execution subject of the above steps may be a processor, but is not limited thereto.
[0038] In an exemplary embodiment, removing the first pixel points included in the pixel points that exceed the projection range of the target imaging device based on the projection coordinates and the calibration parameters includes: determining the camera coordinates of the target imaging device in the three-dimensional scene included in the calibration parameters; determining a first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; when the first distance is greater than a preset distance, determining the pixel point corresponding to the first projection coordinate as the first pixel point; and removing the first pixel points from the pixel points included in the target image.
[0039] In the above embodiment, the coordinates CX, CY, CZ of the target imaging device (i.e., the above camera coordinates) can be obtained by projecting the coordinates of the actual target imaging device in the three-dimensional scene. The camera X-direction vector CX can be expressed as (CX1, CX2, CX3), the camera Y-direction vector CY can be expressed as (CY1, CY2, CY3), and the camera Z-direction vector can be expressed as (CZ1, CZ2, CZ3). Among them, the coordinates of the actual target imaging device can be represented by CP(X C ,Y C ,Z C ).
[0040] In the above embodiment, the preset distance can be the maximum projection distance D of the camera, which can be understood as the projection distance of the maximum perspective distance Maxdistance of the camera in the three-dimensional scene and can be calculated by the following formula: D = Sqrt(CRP·CRP), where CRP is the position of any point in the three-dimensional scene relative to the camera and can be calculated by the following formula: CRP = WP - CP.
[0041] In the above embodiment, the distance (i.e., the above first distance) between each projection coordinate (i.e., each of the above first projection coordinates) formed by projecting each pixel point of the target image into the three-dimensional scene and the camera coordinates is determined, and the pixel points (i.e., the above first pixel points) with the first distance greater than the maximum projection distance D of the camera are removed from the pixel points included in the target image.
[0042] In an exemplary embodiment, removing the first pixel points included in the pixel points that exceed the projection range of the target imaging device based on the projection coordinates and the calibration parameters includes: determining a target viewing frustum of the target imaging device in the three-dimensional scene based on the calibration parameters, where the target viewing frustum is used to represent the shooting range of the target imaging device; and removing the first pixel points located outside the target viewing frustum in the three-dimensional scene based on the projection coordinates.
[0043] In the above embodiments, the target frustum can be understood as the shooting range of the target imaging device. The target frustum can be determined according to the calibration parameters of the target imaging device. After obtaining the target frustum, the first pixel points outside the target frustum can be removed according to the projection coordinates.
[0044] In an exemplary embodiment, determining the target frustum of the target imaging device in the three-dimensional scene based on the calibration parameters includes: determining the maximum projection distance of the target imaging device included in the calibration parameters, the horizontal field of view angle of the target imaging device, and determining the horizontal component of the camera coordinates of the target imaging device in the horizontal direction, the vertical component in the vertical direction, and the vertical component in the vertical direction in the three-dimensional scene; determining the first product of the maximum projection distance and a first constant; determining the first ratio of the horizontal field of view angle and a second constant; determining the tangent value of the first ratio; determining the second product of the first product and the tangent value; determining the second product as the bottom length of the target frustum; determining the third product of the bottom length and the aspect ratio of the target frustum determined in advance as the bottom width of the target frustum; determining the fourth product of the vertical component and the maximum projection distance; determining the fifth product of the vertical component, the bottom width, and a third constant; determining the sixth product of the horizontal component, the bottom width, and a fourth constant; determining the first sum value of the fourth product, the fifth product, and the sixth product as the first side vector of the target side of the target frustum; determining the seventh product of the vertical component, the bottom width, and a fifth constant; determining the eighth product of the horizontal component, the bottom width, and a sixth constant; determining the second sum value of the fourth product, the seventh product, and the eighth product as the second side vector of the target side; determining the target frustum based on the bottom length, bottom width, the first side vector, and the second side vector.
[0045] In the above embodiments, the length and width of the bottom surface of the target frustum can be calculated by the principle of spatial vectors, as Figure 3 shown Figure 3 is a schematic structural diagram of the target frustum. The bottom length L of the target frustum can be calculated by the following formula: L = MaxDistacne * 2.0 * (Tan(FOV / 2.0)), and the bottom width W of the target frustum can be calculated by the following formula: W = L * ResYX. Where, MaxDistacne represents the maximum projection distance, and FOV represents the horizontal field of view angle. The first constant can be 2, and the second constant can be 2.
[0046] In the above embodiments, the left side vectors a (i.e., the first side vector above) and b (i.e., the second side vector above) of the target frustum can be calculated by the following formula: a = CZ * MaxDistacne + CY * 0.5W + CX * (-0.5L), b = CZ * MaxDistacne + CY * (-0.5W) + CX * (-0.5L), where CZ represents the vertical component in the vertical direction of the target imaging device in the three-dimensional scene, CY represents the vertical component of the target imaging device in the three-dimensional scene, and CX represents the horizontal component in the horizontal direction of the target imaging device in the three-dimensional scene. The third constant can be 0.5, the fourth constant can be -0.5, the fifth constant can be -0.5, and the sixth constant can be -0.5.
[0047] It should be noted that the above first constant, second constant, third constant, fourth constant, fifth constant, and sixth constant are only exemplary descriptions. The first constant, second constant, third constant, fourth constant, fifth constant, and sixth constant can also take other values, and the present invention does not limit this.
[0048] In an exemplary embodiment, determining the target frustum based on the bottom length, bottom width, the first side vector, and the second side vector includes: determining an initial frustum based on the bottom length, bottom width, the first side vector, and the second side vector; converting the initial frustum to the display plane space to obtain a display frustum; and removing the pixel points outside the target range included in the display frustum to obtain the target frustum.
[0049] In the above embodiments, the initial frustum can be understood as the frustum determined in the three-dimensional scene according to the calibration parameters of the camera, and the display frustum can be understood as the frustum obtained by converting the initial frustum from the three-dimensional scene space to the plane space (i.e., the above display plane space). First, the world space, that is, the three-dimensional scene space, can be transferred to the camera space, and then from the camera space to the display plane space. For the schematic diagram of the coordinate system of the camera space, please refer to the appendix Figure 4 , as Figure 4 shown. The origin of the camera coordinate system is at the center point, the direction of increasing X-axis is to the right, and the direction of increasing Y-axis is upward. For the schematic diagram of the coordinate system of the display plane space, please refer to the appendix Figure 5 , as Figure 5 shown. In the display plane space, the direction of increasing Y-axis is to the right, and the direction of increasing X-axis is downward.
[0050] In the above embodiments, an initial frustum can be determined based on the bottom length, bottom width, first side vector, and second side vector of the frustum, and a display frustum can be obtained by transforming the initial frustum into the planar space. The perspective projection of obtaining the display frustum from the initial frustum can be expressed as PP = ((1, 0, 0) • CRP • CX + (0, 1, 0) • CRP • CY) / ((0, 0, 1) • CRP • CZ * Tan(FOV / 2) * (1, ResYX)). After obtaining the display frustum, the perspective coordinates of the display frustum can be mapped to the target range, such as the range of 0 - 1, and the result Result = PP * (0.5, -0.5) + (0.5, 0.5). Output 1 for the result within the range of [0, 1] in the Result result, and exclude 0 for the outside of the range (i.e., the pixel points outside the above target range), then the target frustum can be obtained.
[0051] In an exemplary embodiment, the culling of the first pixel points located outside the target frustum in the three - dimensional scene based on the projection coordinates includes: determining the camera coordinates of the target camera device in the three - dimensional scene included in the calibration parameters; determining the first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; determining the first angle between the pixel point corresponding to the projection coordinates and the target side of the target frustum; determining the second distance between the target point included in the bottom surface of the target frustum and the camera coordinates based on the first angle; in the case where the first distance is greater than the second distance, determining that the pixel point corresponding to the projection coordinates is the first pixel point; and culling the first pixel points from the pixel points included in the target image.
[0052] In the above embodiments, the unit normal of the left side of the frustum can be expressed as LFNormal, and the angle between the line connecting any pixel point and the camera and the left side of the frustum (i.e., the above - mentioned target side) can be expressed as α (i.e., the above - mentioned first angle). Among them, LFNormal can be calculated from the first side vector a and the second side vector a of the target frustum, and the formula is as follows: LFNormal = a × b; the first angle α can be calculated through the following formula: α = ArcCos(LFNormal • CRP).
[0053] In the above embodiments, when the distance FB (i.e., the above - mentioned second distance) between a point on the bottom surface of the target frustum and the camera in the projection coordinates is greater than the maximum projection distance D of the camera (i.e., the above - mentioned first distance), the first pixel points outside the target frustum are culled. Among them, the distance FB between a point on the target frustum and the camera in the projection coordinates can be calculated through the following formula: FB = Maxdistance / cos(Fov / 2 - α).
[0054] In an exemplary embodiment, removing the third pixel points included in the second pixel points that are occluded in the three-dimensional scene includes: performing the following operations for each fourth pixel point included in the second pixel points to determine the third pixel points from the second pixel points: outputting the depth data of the target image as texture data, determining a first depth of the second pixel points based on the data of the second pixel points included in the texture data in a target channel, determining a second projection coordinate corresponding to the second pixel points included in the projection coordinates, determining a relative vector between the second projection coordinate and the camera coordinate of the target imaging device in the three-dimensional scene included in the calibration parameters, projecting the relative vector onto the vertical direction to obtain a second depth, and when the second depth is greater than the first depth, determining the fourth pixel point as the third pixel point; removing the third pixel points from the target image.
[0055] In the above embodiment, the depth data of the target image can be understood as the distance between a certain pixel point on the target image and the target imaging device. Figure 6 It is a schematic diagram of depth data and projection distance according to an embodiment of the present invention. As Figure 6 shown, the target imaging device is facing a wall, there is a circular stone behind the wall, and there is a cube stone in front of the wall. Then the depth data of the circular stone is the distance between the wall and the target imaging device, and the depth data of the cube stone is the projection H of the distance from the target imaging device to the cube stone on the z-axis of the target imaging device.
[0056] In the above embodiment, the fourth pixel point can be any pixel point in the second pixel points, and the target channel can be the R channel. After calculating the depth data of the target image, the depth data in the scene is output as texture data, and the R channel is written into a 32-bit 2D rendering target (RenderTarget2D) according to the texture data to obtain the first depth of the second pixel points in the texture data. Determine the projection distance (i.e., the above second depth) of the relative vector between the projection coordinate corresponding to the second pixel points in the three-dimensional scene (i.e., the above second projection coordinate) and the camera coordinate corresponding to the target imaging device in the three-dimensional scene on the Z axis (i.e., the above vertical direction), and remove the pixel points with the second depth greater than the first depth (i.e., the above third pixel points) from the target image.
[0057] In an exemplary embodiment, rendering the target pixel points into the three-dimensional scene includes: determining the scene color tone and scene brightness of the three-dimensional scene; adjusting the target color tone and target brightness of the target pixel points based on the scene color tone and the scene brightness; and rendering the target pixel points with the adjusted target color tone and target brightness into the three-dimensional scene.
[0058] In the above embodiments, since each pixel is formed by superimposing three colors of R / G / B (i.e., the above scene color tone), it is necessary to first determine each pixel P(R, G, B, A) in the three-dimensional scene during the process of rendering colors. Multiply the pixel P by the custom parameter C (xr, yg, zb), and multiply the self-luminous intensity of the pixel (i.e., the above scene brightness) by the custom parameter L. Exposing the parameters xr, yg, zb, and L can adjust the RGB and the self-luminous intensity. Among them, R, G, and B respectively represent the three color channels of red, green, and blue, and A represents transparency. During the process of adjusting the target color tone and target brightness of the target pixel, it can be adjusted by manually touching and selecting parameters on the display interface, or by inputting parameters, or it can be automatically adjusted by a machine.
[0059] The following will explain the image fusion method in combination with specific embodiments:
[0060] Figure 7 It is a flowchart of the image fusion method according to a specific embodiment of the present invention, as Figure 7 shown, and this process includes the following steps:
[0061] Step S702, start calibration;
[0062] Step S704, calculate the necessary parameters CX, CY, CZ, L, W, a, b, LFNormal;
[0063] Step S706, convert world coordinates to screen coordinates;
[0064] Step S708, remove non-projected areas;
[0065] Step S710, distance / cone clipping;
[0066] Step S712, occlusion culling;
[0067] Step S714, customize the color tone and brightness.
[0068] In the above embodiment, the parameters of the calibrated camera are obtained, which may include: camera coordinates CP (Xc, Yc, Zc), camera rotation angles A, B, C, camera horizontal field of view FOV, camera maximum projection distance Maxdistance, and coordinates WP (Xw, Yw, Zw) of any point in the three-dimensional scene. The calibration parameters also include: camera X direction vector CX=(CX1, CX2, CX3), camera Y direction vector CY=(CY1, CY2, CY3), camera Z direction vector CZ=(CZ1, CZ2, CZ3), the length L and width W of the bottom surface of the cone, the left side vectors a and b of the cone, and the left side normal LFNormal of the cone, which can be calculated using the above calibrated camera parameters.
[0069] In the aforementioned embodiment, the coordinate system needs to be converted when the real image is displayed in the two-dimensional display interface, so the coordinate system of the target image in the world space is converted to the coordinate system of the camera space, and then the coordinate system of the camera space is projected to the two-dimensional plane space (i.e., the screen space in the figure). Since the projection distance of the target camera device is limited, the pixels in the non-projection range and the pixels beyond the projection range according to the maximum projection distance are removed, and then the pixels of the target image that are blocked in the scene are removed, and the remaining pixels in the three-dimensional scene are rendered and the color tone and brightness are customized, so that the perfect fusion of the image can be achieved.
[0070] In the above-mentioned embodiment, the effect diagram of removing the obstructed target image pixels can be referred to as Figure 8 ,like Figure 8 As shown in (a), there is a white wall in the scene. The left side of the white wall is the effect image without removing the occluded pixels. The occluded image is still attached to the model to form a penetrating image, causing a messy display. Figure 8 (b) in the figure shows the effect of turning on occlusion culling. Figure 8 In (a), the pixels that penetrate the white wall are removed to avoid the cluttered display of the scene. Through the depth detection method, the video screen blocked by the scene can be removed, and the redundant pixels outside the frustum can be removed to improve the visualization effect. The picture can also be adjusted in real time to enhance the video picture so that it can be better integrated into the scene. The image after image fusion can be referenced Figure 9 In the process of image fusion, you only need to calibrate the camera posture to fuse the ordinary video image with the three-dimensional scene. It is easy to operate and still makes the image rich and beautiful.
[0071] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present invention.
[0072] In this embodiment, an image fusion device is further provided. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0073] Figure 10 is a structural block diagram of an image fusion device according to an embodiment of the present invention. As Figure 10 shown, the device includes:
[0074] A first determination module 1002, configured to determine calibration parameters of a target imaging device in a three-dimensional scene;
[0075] A second determination module 1004, configured to determine projection coordinates obtained by projecting each pixel point included in a target image captured by the target imaging device into the three-dimensional scene;
[0076] A first rejection module 1006, configured to reject first pixel points that exceed the projection range of the target imaging device among the pixel points based on the projection coordinates and the calibration parameters, so as to obtain second pixel points;
[0077] A second rejection module 1008, configured to reject third pixel points that are blocked in the three-dimensional scene among the second pixel points, so as to obtain target pixel points;
[0078] A rendering module 1010, configured to render the target pixel points into the three-dimensional scene, so as to fuse the target image captured by the target imaging device into the three-dimensional scene.
[0079] In an exemplary embodiment, the first rejection module 1006 may reject the first pixel points included in the pixel points that exceed the projection range of the target imaging device based on the projection coordinates and the calibration parameters in the following manner: determining the camera coordinates of the target imaging device in the three-dimensional scene included in the calibration parameters; determining the first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; in the case where the first distance is greater than a preset distance, determining the pixel point corresponding to the first projection coordinate as the first pixel point; and rejecting the first pixel point from the pixel points included in the target image.
[0080] In an exemplary embodiment, the first rejection module 1006 may reject the first pixel points included in the pixel points that exceed the projection range of the target imaging device based on the projection coordinates and the calibration parameters in the following manner: determining a target frustum of the target imaging device in the three-dimensional scene based on the calibration parameters, where the target frustum is used to represent the shooting range of the target imaging device; and rejecting the first pixel points located outside the target frustum in the three-dimensional scene based on the projection coordinates.
[0081] In an exemplary embodiment, the first culling module 1006 may implement the determination of the target frustum of the target imaging device in the three-dimensional scene based on the calibration parameters in the following manner: determining the maximum projection distance of the target imaging device included in the calibration parameters, the horizontal field of view angle of the target imaging device, and determining the horizontal component of the camera coordinates of the target imaging device in the horizontal direction, the vertical component in the vertical direction, and the vertical component in the perpendicular direction in the three-dimensional scene; determining a first product of the maximum projection distance and a first constant; determining a first ratio of the horizontal field of view angle and a second constant; determining the tangent value of the first ratio; determining a second product of the first product and the tangent value; determining the second product as the bottom length of the target frustum; determining a third product of the bottom length and the aspect ratio of the target frustum determined in advance as the bottom width of the target frustum; determining a fourth product of the vertical component and the maximum projection distance; determining a fifth product of the vertical component, the bottom width, and a third constant; determining a sixth product of the horizontal component, the bottom width, and a fourth constant; determining a first sum value of the fourth product, the fifth product, and the sixth product as the first edge vector of the target side of the target frustum; determining a seventh product of the vertical component, the bottom width, and a fifth constant; determining an eighth product of the horizontal component, the bottom width, and a sixth constant; determining a second sum value of the fourth product, the seventh product, and the eighth product as the second edge vector of the target side; and determining the target frustum based on the bottom length, the bottom width, the first edge vector, and the second edge vector.
[0082] In an exemplary embodiment, the first culling module 1006 may implement the determination of the target frustum based on the bottom length, the bottom width, the first edge vector, and the second edge vector in the following manner: determining an initial frustum based on the bottom length, the bottom width, the first edge vector, and the second edge vector; converting the initial frustum to the display plane space to obtain a display frustum; and culling the pixel points outside the target range included in the display frustum to obtain the target frustum.
[0083] In an exemplary embodiment, the first culling module 1006 may implement culling of the first pixel points located outside the target frustum in the three-dimensional scene based on the projection coordinates in the following manner: determining the camera coordinates of the target imaging device in the three-dimensional scene included in the calibration parameters; determining the first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; determining the first angle between the pixel point corresponding to the projection coordinates and the target side of the target frustum; determining the second distance between the target point included in the bottom surface of the target frustum and the camera coordinates based on the first angle; in the case where the first distance is greater than the second distance, determining that the pixel point corresponding to the projection coordinates is the first pixel point; and culling the first pixel point from the pixel points included in the target image.
[0084] In an exemplary embodiment, the second culling module 1008 may implement culling of the third pixel points occluded in the three-dimensional scene included in the second pixel points in the following manner: performing the following operations for each fourth pixel point included in the second pixel points to determine the third pixel points from the second pixel points: outputting the depth data of the target image as texture data, determining the first depth of the second pixel points based on the data of the second pixel points in the target channel included in the texture data, determining the second projection coordinates corresponding to the second pixel points included in the projection coordinates, determining the relative vector between the second projection coordinates and the camera coordinates of the target imaging device in the three-dimensional scene included in the calibration parameters, projecting the relative vector onto the vertical direction to obtain the second depth, and in the case where the second depth is greater than the first depth, determining that the fourth pixel point is the third pixel point; and culling the third pixel points from the target image.
[0085] In an exemplary embodiment, the rendering module 1010 may implement rendering the target pixel points into the three-dimensional scene in the following manner: determining the scene color tone and scene brightness of the three-dimensional scene; adjusting the target color tone and target brightness of the target pixel points based on the scene color tone and the scene brightness; and rendering the target pixel points with the adjusted target color tone and target brightness into the three-dimensional scene.
[0086] It should be noted that the above-mentioned various modules may be implemented by software or hardware. For the latter, it may be implemented in the following manner, but not limited thereto: the above-mentioned modules are all located in the same processor; or, the above-mentioned various modules are respectively located in different processors in any combination form.
[0087] Embodiments of the present invention also provide a computer-readable storage medium storing a computer program, where the computer program is configured to execute the steps in any of the above method embodiments when running.
[0088] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as USB flash drives, read-only memory (ROM for short), random access memory (RAM for short), mobile hard disks, magnetic disks, or optical discs that can store computer programs.
[0089] Embodiments of the present invention also provide an electronic device including a memory and a processor, where the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above method embodiments.
[0090] In an exemplary embodiment, the above electronic device may further include a transmission device and an input / output device, where the transmission device is connected to the above processor, and the input / output device is connected to the above processor.
[0091] Embodiments of the present invention also provide a computer program product including a computer program, where the steps of the methods in various embodiments of the present application are implemented when the computer program is executed by a processor.
[0092] Specific examples in this embodiment may refer to the examples described in the above embodiments and exemplary embodiments, and will not be repeated here.
[0093] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device, so that they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. Thus, the present invention is not limited to any specific combination of hardware and software.
[0094] The above are only preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. An image fusion method, characterized in that: include: Determining calibration parameters of a target camera device in a three-dimensional scene; Determine a projection coordinate obtained by projecting each pixel point included in the target image captured by the target camera device onto the three-dimensional scene; Based on the projection coordinates and the calibration parameters, first pixel points included in the pixel points and exceeding the projection range of the target camera device are removed to obtain second pixel points; Eliminate the third pixel points that are blocked in the three-dimensional scene and are included in the second pixel points to obtain a target pixel point; Rendering the target pixel points into the three-dimensional scene to fuse the target image captured by the target camera device into the three-dimensional scene; Eliminating the first pixel points included in the pixel points that exceed the projection range of the target camera device based on the projection coordinates and the calibration parameters includes: determining the target viewing cone of the target camera device in the three-dimensional scene based on the calibration parameters, wherein the target viewing cone is used to represent the shooting range of the target camera device; eliminating the first pixel points located outside the target viewing cone in the three-dimensional scene based on the projection coordinates; wherein determining the target viewing cone of the target camera device in the three-dimensional scene based on the calibration parameters includes: determining the maximum viewing distance of the target camera device and the horizontal field of view of the target camera device included in the calibration parameters, and determining the horizontal component of the camera coordinates of the target camera device in the three-dimensional scene in the horizontal direction, the vertical component in the vertical direction, and the vertical component in the vertical direction; determining the first product of the maximum viewing distance and a first constant; determining the first ratio of the horizontal field of view angle to a second constant; determining the tangent value of the first ratio; and determining Determine the second product of the first product and the tangent value; determine the second product as the base length of the target visual cone; determine the third product of the base length and the predetermined aspect ratio of the target visual cone as the base width of the target visual cone; determine the fourth product of the vertical component and the maximum viewing distance; determine the fifth product of the vertical component, the base width and a third constant; determine the sixth product of the horizontal component, the base width and a fourth constant; determine the first sum of the fourth product, the fifth product and the sixth product as the first edge vector of the target side of the target visual cone; determine the seventh product of the vertical component, the base width and a fifth constant; determine the eighth product of the horizontal component, the base width and a sixth constant; determine the second sum of the fourth product, the seventh product and the eighth product as the second edge vector of the target side; determine the target visual cone based on the base length, the base width, the first edge vector and the second edge vector; The target view frustum is a view frustum obtained by removing the pixel points outside the target range included in the display view frustum, the display view frustum is a view frustum converted from the initial view frustum to a plane space, the initial view frustum is a view frustum determined in a three-dimensional scene based on the calibration parameters of the target camera device, and the transmission projection of the display view frustum obtained from the initial view frustum is expressed by the following formula: PP = ((1,0,0)*CRP·CX+(0,1,0)CRP·CY) / ((0,0,1)CRP· CZ*Tan(FOV / 2)*(1,ResYX)), wherein CRP represents the position of any point in the three-dimensional scene relative to the target camera device, CX represents the horizontal component of the target camera device in the horizontal direction of the three-dimensional scene, CY represents the vertical component of the target camera device in the three-dimensional scene, CZ represents the vertical component of the target camera device in the vertical direction of the three-dimensional scene, FOV represents the horizontal field of view of the target camera device, and ResYX represents the resolution ratio of the target camera device.
2. The method according to claim 1, characterized in that Eliminating the first pixel points included in the pixel points and exceeding the projection range of the target camera device based on the projection coordinates and the calibration parameters includes: Determining the camera coordinates of the target camera device included in the calibration parameters in the three-dimensional scene; Determine a first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; When the first distance is greater than a preset distance, determining the pixel point corresponding to the first projection coordinate as the first pixel point; The first pixel point is eliminated from the pixel points included in the target image.
3. The method according to claim 1, characterized in that Determining the target view frustum based on the bottom length, the bottom width, the first edge vector, and the second edge vector comprises: Determine an initial view frustum based on the bottom length, the bottom width, the first edge vector, and the second edge vector; Converting the initial viewing cone to a display plane space to obtain a display viewing cone; Pixel points outside the target range included in the display viewing cone are removed to obtain the target viewing cone.
4. The method according to claim 1, characterized in that: Determining based on the projection coordinates that the first pixel point located outside the target viewing frustum in the three-dimensional scene is eliminated includes: Determining the camera coordinates of the target camera device included in the calibration parameters in the three-dimensional scene; determining a first distance between each first projection coordinate included in the projection coordinates and the camera coordinates; Determine a first angle between a pixel point corresponding to the projection coordinate and a target side surface of the target viewing cone; Determine a second distance between a target point included in a bottom surface of the target viewing cone and the camera coordinates based on the first angle; When the first distance is greater than the second distance, determining the pixel point corresponding to the projection coordinate as the first pixel point; The first pixel point is eliminated from the pixel points included in the target image.
5. The method according to claim 1, characterized in that Eliminating the third pixel points included in the second pixel points and blocked in the three-dimensional scene includes: The following operations are performed for each fourth pixel point included in the second pixel points to determine the third pixel point from the second pixel points: output the depth data of the target image as texture data, determine the first depth of the second pixel point based on the data of the second pixel point in the target channel included in the texture data, determine the second projection coordinates corresponding to the second pixel point included in the projection coordinates, determine the relative vector between the second projection coordinates and the camera coordinates of the target camera device included in the calibration parameters in the three-dimensional scene, project the relative vector to the vertical direction to obtain a second depth, and when the second depth is greater than the first depth, determine the fourth pixel point as the third pixel point; The third pixel is eliminated from the target image.
6. The method according to claim 1, characterized in that Rendering the target pixel point into the three-dimensional scene includes: Determining a scene hue and a scene brightness of the three-dimensional scene; Adjusting a target hue and a target brightness of the target pixel based on the scene hue and the scene brightness; The target pixel points with the target hue and the target brightness adjusted are rendered into the three-dimensional scene.
7. An image fusion device, characterized in that: include: A first determination module is used for calibrating parameters of a target camera device in a three-dimensional scene; A second determination module is used to determine the projection coordinates of each pixel point included in the target image captured by the target camera device projected onto the three-dimensional scene; A first elimination module is used to eliminate first pixel points included in the pixel points and exceeding the projection range of the target camera device based on the projection coordinates and the calibration parameters to obtain second pixel points; A second elimination module is used to eliminate the third pixel points included in the second pixel points and blocked in the three-dimensional scene to obtain the target pixel points; A rendering module, used for rendering the target pixel points into the three-dimensional scene, so as to fuse the target image captured by the target camera device into the three-dimensional scene; The first elimination module eliminates the first pixel points that are included in the pixel points and are beyond the projection range of the target camera device based on the projection coordinates and the calibration parameters in the following manner: determining a target viewing cone of the target camera device in the three-dimensional scene based on the calibration parameters, wherein the target viewing cone is used to represent the shooting range of the target camera device; eliminating the first pixel points that are outside the target viewing cone in the three-dimensional scene based on the projection coordinates; wherein the first elimination module determines the target viewing cone of the target camera device in the three-dimensional scene based on the calibration parameters in the following manner: determining the maximum viewing distance of the target camera device and the horizontal field of view of the target camera device included in the calibration parameters, and determining the horizontal component in the horizontal direction, the vertical component in the vertical direction, and the vertical component in the vertical direction of the camera coordinates of the target camera device in the three-dimensional scene; determining a first product of the maximum viewing distance and a first constant; determining a first ratio of the horizontal field of view angle to a second constant; determining Determine the tangent value of the first ratio; determine the second product of the first product and the tangent value; determine the second product as the base length of the target visual cone; determine the third product of the base length and the predetermined aspect ratio of the target visual cone as the base width of the target visual cone; determine the fourth product of the vertical component and the maximum viewing distance; determine the fifth product of the vertical component, the base width and a third constant; determine the sixth product of the horizontal component, the base width and a fourth constant; determine the first sum of the fourth product, the fifth product and the sixth product as the first edge vector of the target side of the target visual cone; determine the seventh product of the vertical component, the base width and a fifth constant; determine the eighth product of the horizontal component, the base width and a sixth constant; determine the second sum of the fourth product, the seventh product and the eighth product as the second edge vector of the target side; determine the target visual cone based on the base length, the base width, the first edge vector and the second edge vector; The target view frustum is a view frustum obtained by removing the pixel points outside the target range included in the display view frustum, the display view frustum is a view frustum converted from the initial view frustum to a plane space, the initial view frustum is a view frustum determined in a three-dimensional scene based on the calibration parameters of the target camera device, and the transmission projection of the display view frustum obtained from the initial view frustum is expressed by the following formula: PP = ((1,0,0)*CRP·CX+(0,1,0)CRP·CY) / ((0,0,1)CRP· CZ*Tan(FOV / 2)*(1,ResYX)), wherein CRP represents the position of any point in the three-dimensional scene relative to the target camera device, CX represents the horizontal component of the target camera device in the horizontal direction of the three-dimensional scene, CY represents the vertical component of the target camera device in the three-dimensional scene, CZ represents the vertical component of the target camera device in the vertical direction of the three-dimensional scene, FOV represents the horizontal field of view of the target camera device, and ResYX represents the resolution ratio of the target camera device.
8. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 1 to 6 when executed.
9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Method and apparatus for identifying foreground target obejct
CN105225230A
Three-dimensional rendering method and device, equipment and medium
CN117237502A
Three-dimensional scene rendering-based viewpoint-independent multi-channel video fusion method and system
CN117560578A