A video fusion method and device, computer equipment and storage medium
Patent Information
- Application Number
- CN202111660173.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-31
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2041-12-31
AI Technical Summary
[0017]本发明实施例通过将视频帧图片以及与其匹配的场景图片分别划分成相同数量的平面多边形,对场景区域中的目标平面多边形内的目标点,确定目标线段端点,计算目标线段端点的平面坐标值和深度值信息,根据目标线段端点的深度值信息,确定目标线段端点在视频帧图片中的纹理坐标值,并根据目标线段端点的平面坐标值和纹理坐标值,确定目标点的纹理坐标值,以根据目标点的纹理坐标值进行渲染。解决了现有技术中的视频融合方式,容易产生纹理畸变,影响视频融合的效果的问题,实现了按照划分区域进行纹理映射,保证了视频融合的效果。
Smart Images

Figure CN116416356B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and in particular to a method, apparatus, computer device, and storage medium for video fusion. Background Technology
[0002] With the development of virtual reality technology, video fusion technology has become an important means of fusing bidirectional information from models and videos. Video fusion technology refers to the process of fusing one or more videos of a scene or model captured by video acquisition devices with a related virtual scene to generate a new virtual scene or model about that scene.
[0003] In video fusion technology, monitoring and shooting devices need to acquire real-time video, and scene shooting devices render the video footage onto a model or scene. During video fusion, the video frame images in 1a and... Figure 1b For example, in the scene images, such as Figure 1c As shown, when mapping the texture of region 2 in the video frame image to the scene image, the sparse triangles can easily cause texture distortion, affecting the video fusion effect. Summary of the Invention
[0004] This invention provides a method, apparatus, computer device, and storage medium for video fusion, which enables texture mapping according to divided regions, thereby ensuring the effectiveness of video fusion.
[0005] In a first aspect, embodiments of the present invention provide a video fusion method, the method comprising:
[0006] Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons.
[0007] For target points within the target planar polygon in the scene area, determine the endpoints of the target line segments, and calculate the planar coordinates and depth information of the endpoints of the target line segments;
[0008] Based on the depth information of the endpoints of the target line segment, calculate the texture coordinate values that match the endpoints of the target line segment with the video frame image;
[0009] Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, the texture coordinates of the target point are calculated to match the video frame image, and then the video is rendered based on the texture coordinates of the target point to achieve video fusion.
[0010] Secondly, embodiments of the present invention also provide a video fusion apparatus, the apparatus comprising:
[0011] The planar region division module is used to acquire video frame images and scene regions that match the video frame images, and to divide the video frame images and scene regions into a number of corresponding planar polygons.
[0012] The depth value calculation module is used to determine the endpoints of target line segments for target points within the target planar polygon in the scene area, and to calculate the planar coordinates and depth value information of the endpoints of the target line segments.
[0013] The texture coordinate calculation module is used to calculate the texture coordinate values of the endpoints of the target line segment and the video frame image based on the depth value information of the endpoints of the target line segment.
[0014] The target pixel texture coordinate value calculation module is used to calculate the texture coordinate value of the target point and the video frame image based on the planar coordinate value and texture coordinate value of the endpoint of the target line segment, so as to render according to the texture coordinate value of the target point and realize video fusion.
[0015] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video fusion method as described in any of the embodiments of the present invention.
[0016] Fourthly, embodiments of the present invention also provide a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform the video fusion method as described in any of the embodiments of the present invention.
[0017] This invention addresses the problem of texture distortion in existing video fusion methods, which negatively impacts fusion quality. By dividing a video frame image and its matching scene image into an equal number of planar polygons, and by performing texture mapping based on region division, the invention ensures effective video fusion. Attached Figure Description
[0018] Figure 1a This is a schematic diagram of a video frame image in the background art of this invention;
[0019] Figure 1b This is a schematic diagram of a scene area in the background art of this invention;
[0020] Figure 1c This is a schematic diagram illustrating the effect of texture mapping in the background technology of this invention;
[0021] Figure 1d This is a flowchart of a video fusion method according to Embodiment 1 of the present invention;
[0022] Figure 2a This is a flowchart of a video fusion method according to Embodiment 2 of the present invention;
[0023] Figure 2b This is a schematic diagram illustrating the relationship between points in a scene area and pixels in a video frame image, according to Embodiment 2 of the present invention.
[0024] Figure 2c This is a schematic diagram illustrating the relationship between the depth values of two vertices corresponding to a diagonal and the endpoint of the diagonal in Embodiment 2 of the present invention.
[0025] Figure 2d This is a schematic diagram of a target point and the endpoints of a target line segment in a scene area according to Embodiment 2 of the present invention;
[0026] Figure 2e This is a schematic diagram of pixels in a video frame image that match the target point and the endpoint of the target line segment in Embodiment 2 of the present invention;
[0027] Figure 3 This is a schematic diagram of the structure of a video fusion device according to Embodiment 3 of the present invention;
[0028] Figure 4 This is a schematic diagram of the structure of a computer device according to Embodiment 4 of the present invention. Detailed Implementation
[0029] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and not intended to limit it. Furthermore, it should be noted that, for ease of description, the accompanying drawings show only the parts relevant to the present invention, and not all of the structures.
[0030] Example 1
[0031] Figure 1d This is a flowchart of a video fusion method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of video fusion of video frame images captured by video shooting equipment with virtual scenes. The method can be executed by a video fusion device, which can be implemented by software and / or hardware, and is generally integrated into a computer device for use in conjunction with the video shooting equipment.
[0032] like Figure 1d As shown, the technical solution of this embodiment of the invention specifically includes the following steps:
[0033] S110. Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons.
[0034] In this embodiment of the invention, a scene model needs to be created before performing planar region division and texture mapping. Preferably, highly mobile interfering factors such as vehicles and pedestrians are removed from the scene model beforehand. This is because even small changes in the scene model can cause inconsistencies between the information depth values in the video frame and the scene region, resulting in poor video fusion or even distortion. These issues can only be resolved by modifying or reconstructing the scene model, which is costly. Therefore, removing highly mobile interfering factors from the scene model ensures effective video fusion. This embodiment does not limit the methods used to create the scene model or remove interfering factors.
[0035] Since the scene model is pre-built, the scene region can be obtained based on the scene model during video fusion. For example... Figure 1b As shown, highly mobile interference factors have been removed from the scene area. A video frame image is a single frame from a video captured by a video recording device of the actual scene, such as... Figure 1a As shown, the captured image of the video frame matches the scene area.
[0036] Preferably, the planar polygons can be set as quadrilaterals. The number of planar polygons divided in the scene area and the video frame image must be the same and correspond to each other, so as to... Figure 1a and Figure 1b For example, the video frame image and the scene region are each divided into four corresponding quadrilaterals. In this embodiment of the invention, the video frame image and the scene region are each divided into multiple different planar polygons. For each planar polygon, the corresponding texture in the video frame image is mapped onto the planar polygon corresponding to the scene region, thereby achieving the effect of video fusion.
[0037] S120. For target points within the target planar polygon in the scene area, determine the endpoints of the target line segments, and calculate the planar coordinates and depth information of the endpoints of the target line segments.
[0038] In this embodiment of the invention, texture mapping is performed sequentially on each planar polygon in the scene area. The target point is a point in the target planar polygon. The purpose of texture mapping is to determine the texture coordinates of the target point in the scene area within the video frame image, thereby realizing the perspective texture mapping transformation between the planar polygons in the scene area and the planar polygons in the video frame image.
[0039] The endpoint of the target line segment is the endpoint of the target line segment containing the target point. The target line segment is a line segment formed by drawing parallel lines through the target point along the coordinate axes of the two-dimensional coordinate system within the coordinate system containing the target planar polygon. The endpoint of the target line segment can be the intersection of the parallel line drawn through the target point and a side of the polygon, or it can be the intersection of the parallel line drawn through the target point and a diagonal of the polygon; this embodiment does not impose any restrictions on this.
[0040] Since the coordinates of the vertices of the target polygon can be predetermined within its coordinate system, taking the intersection of the target line segment endpoint and the diagonal of the polygon as an example, the ordinate of the target line segment endpoint can be determined because it lies on the parallel line. Simultaneously, the target line segment endpoint is also on the diagonal of the polygon, and the coordinates of the two vertices corresponding to the diagonal are known. Furthermore, a linear relationship exists between these two vertices. Therefore, based on the ordinate of the target line segment endpoint and the linear relationship between them, the abscissa of the target line segment endpoint can be obtained. Thus, the planar coordinates of the target line segment endpoint can be derived.
[0041] For the depth information of the endpoint of the target line segment, the depth value relationship between the two vertices of the diagonal where the endpoint of the target line segment is located and the intersection of the diagonal can be determined in advance. Since there is a linear relationship between the endpoint of the target line segment and the two vertices corresponding to the diagonal in the scene area, the depth information of the endpoint of the target line segment can be obtained, that is, the depth value relationship between the endpoint of the target line segment and the intersection of the diagonal.
[0042] S130. Based on the depth information of the endpoints of the target line segment, calculate the texture coordinate values that match the endpoints of the target line segment with the video frame image.
[0043] The texture coordinates of the endpoint of the target line segment that match the video frame image are the coordinates of the pixel in the video frame image that matches the endpoint of the target line segment. In this embodiment of the invention, the depth relationship between the two vertices of the diagonal where the endpoint of the target line segment is located and the intersection point of the diagonal is predetermined, i.e., the depth information of the two vertices. In the scene region, there is a depth relationship between the endpoint of the target line segment and the two vertices of the diagonal. In the video frame image, there is also a depth relationship between the pixel corresponding to the endpoint of the target line segment and the two vertices of the diagonal of the video frame image. The depth relationship in the scene region is the same as the depth relationship in the video frame image. Since the pixel corresponding to the endpoint of the target line segment in the video frame image and the two vertices of the diagonal of the video frame image are both located on the diagonal of the video frame image, there is also a linear relationship between the pixel corresponding to the endpoint of the target line segment in the video frame image and the two vertices of the diagonal of the video frame image. Since the pixel coordinates of the two vertices of the diagonal of the video frame image are predetermined, the pixel coordinates of the pixel corresponding to the endpoint of the target line segment in the video frame image, i.e., the texture coordinates, can be obtained.
[0044] S140. Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, calculate the texture coordinates of the target point that match the video frame image, and then render based on the texture coordinates of the target point to achieve video fusion.
[0045] The texture coordinates that match the target point in the video frame image are the coordinates of the pixels in the video frame image that match the target point. After obtaining the texture coordinates of the endpoints of the two target line segments that match the target point, both the target point and the endpoints of the two target line segments lie on the target line segments in the scene area. Therefore, there is a linear relationship between the target point and the endpoints of the two target line segments. Based on the planar coordinates of the endpoints of the target line segments and the planar coordinates of the target point, the linear relationship between the target point and the endpoints of the two target line segments can be determined. Based on the texture coordinates of the endpoints of the two target line segments and the linear relationship between the target point and the endpoints of the two target line segments, the texture coordinates that match the target point in the video frame image can be determined.
[0046] After obtaining the texture coordinates of the target point, the pixel in the video frame that matches the target point can be determined based on the texture coordinates of the target point, and the target point can be rendered based on the pixel information of that pixel, thereby achieving video fusion.
[0047] The technical solution of this embodiment divides the scene region and its matching video frame image into the same number of planar polygons. For target points within the target planar polygons in the scene region, the endpoints of target line segments are determined. The planar coordinates and depth information of the target line segment endpoints are calculated. Based on the depth information of the target line segment endpoints, the texture coordinates of the target line segment endpoints in the video frame image are determined. Finally, based on the planar coordinates and texture coordinates of the target line segment endpoints, the texture coordinates of the target point are determined, and rendering is performed based on the texture coordinates of the target point. This solves the problem of texture distortion easily occurring in existing video fusion methods, which affects the video fusion effect. It achieves texture mapping according to the divided regions, ensuring the effect of video fusion.
[0048] Example 2
[0049] Figure 2a This is a flowchart of a video fusion method provided in Embodiment 2 of the present invention. Based on the above embodiments, the present invention further specifies the process of calculating the planar coordinate values and depth values of the endpoints of the target line segment, and the process of calculating the texture coordinate values of the endpoints of the target line segment and the video frame image. It also adds a step of performing two-dimensional coordinate transformation on the target planar polygon, and a step of calculating the depth value relationship between the vertices and diagonal intersections of the target planar polygon.
[0050] Correspondingly, such as Figure 2a As shown, the technical solution of this embodiment of the invention specifically includes the following steps:
[0051] S210. Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons.
[0052] In this embodiment of the invention, the scene region and the video frame image are respectively divided into planar polygons, and each planar polygon in the scene region is sequentially texture-mapped according to the planar polygon of its corresponding video frame image to achieve video fusion.
[0053] S220. Based on the target planar polygon in the scene area and the relationship between the planar polygons in the video frame image that matches the target planar polygon, determine the depth value relationship between the intersection of the vertices and the diagonals of the target planar polygon.
[0054] Figure 2b A schematic diagram illustrating the relationship between points in a scene region and pixels in a video frame image is provided, such as... Figure 2bAs shown, the origin is the scene camera used for screen rendering, and the screen is used to display the final merged image. The distance between the scene camera and the screen is N. It can be seen that any point q(x, z) on line segment ab corresponds to a point q'(x', z') on the screen. The relationship between q(x, z) and q'(x', z') is x / x' = z / N. Furthermore, there is a linear relationship between points on line segment ab and points a and b. Ultimately, it can be concluded that points in the scene region and pixels in the video frame image have a linear relationship with z / N. Since N is a preset value, points in the scene region and pixels in the video frame image also have a linear relationship with the depth value of the points in the scene.
[0055] by Figure 1b Taking the diagonal P1P3 of region 2 in the scene as an example, the intersection of diagonal P1P3 and diagonal P2P4 is P0. Figure 2c A schematic diagram is provided showing the relationship between the depth values of two corresponding vertices on a diagonal and the endpoints of the diagonal, such as... Figure 2c As shown, the horizontal distances from vertices P1 and P3 to the diagonal intersection point P0 in the scene are L1 and L3, respectively. When vertices P1 and P3 are projected onto the screen, their distances from the diagonal intersection point P0 are L1' and L3', respectively. L1' and L3' are the horizontal pixel distances from the pixel corresponding to vertex P1 and the pixel corresponding to vertex P3 in the video frame image to the pixel corresponding to the diagonal intersection point P0, respectively. The depth values of vertices P1, P3, and the diagonal intersection point P0 are Z1, Z3, and Z, respectively.
[0056] according to Figure 2c It can be known that:
[0057]
[0058] Therefore, we can deduce that:
[0059]
[0060] Since the depth values of points on the diagonal P1P3 change linearly, the relationship between the depth values of vertices P1 and P3 and the intersection point P0 of the diagonal can be determined as follows:
[0061]
[0062] Therefore, it can be deduced that:
[0063]
[0064] Similarly, the depth relationships between vertex P2, vertex P4, and the intersection point P0 of the diagonals can be calculated.
[0065] Optionally, based on the target planar polygon in the scene area and the relationship between planar polygons in the video frame image that matches the target planar polygon, the depth value relationship between the vertices and diagonal intersections of the target planar polygon can be determined. This can include calculating the depth value relationship between the vertices and diagonal intersections of the target planar polygon using the following formula:
[0066]
[0067]
[0068]
[0069]
[0070] Where Z is the depth value of the diagonal intersection point, Z1, Z2, Z3, and Z4 are the depth values of the four vertices of the target planar polygon, L1', L2', L3', and L4' are the pixel distances from the four vertices of the planar polygon in the video frame image matching the target planar polygon to the diagonal intersection point, |P1P|, |P2P|, |P3P|, and |P4P| are the spatial coordinate distances between the four vertices of the target planar polygon and the diagonal intersection point, and |P1P3| and |P2P4| are the spatial coordinate distances between the two vertices corresponding to the two diagonals.
[0071] In this embodiment of the invention, for ease of subsequent calculation, the depth values of the intersection points of vertices P1, P2, P3, and P4 with the diagonal are represented as K1Z, K2Z, K3Z, and K4Z.
[0072] S230. For the target planar polygon in the scene area, establish a two-dimensional coordinate system on the plane where the target planar polygon is located, and perform two-dimensional coordinate transformation on the target planar polygon.
[0073] In this embodiment of the invention, for each point in the target planar polygon of the scene area, its spatial coordinates are converted into two-dimensional coordinates. For example, using... Figure 1b Taking region 2 as an example, we can take P2 as the origin. For the t-axis direction, do Then by Establish a coordinate system w'. For any point P' in region 2, it can be represented as... Then the coordinates of P' are (a, b).
[0074] It should be noted that the above is only one way to perform two-dimensional coordinate transformation on the target planar polygon. This embodiment does not restrict the way the origin and coordinate axes of the two-dimensional coordinate system are set.
[0075] S240. For a target point within a target plane polygon in the scene area, draw a parallel line through the target point along the coordinate axis of the two-dimensional coordinate system. The parallel line intersects with the adjacent edge of the target plane polygon to obtain the endpoint of the target line segment. The adjacent edge includes the diagonal and / or side of the target plane polygon.
[0076] In this embodiment of the invention, texture coordinates are calculated for all pixels within the target planar polygon using the S240-S290 method, thereby rendering the target planar polygon. Each planar polygon in the scene area is rendered using the same method, ultimately achieving video fusion.
[0077] In this embodiment of the invention, a parallel line can be drawn through the target point along the horizontal axis of the two-dimensional coordinate system, or a horizontal line can be drawn through the target point along the vertical axis of the two-dimensional coordinate system. This embodiment does not impose any restrictions on this.
[0078] Figure 2d This provides a schematic diagram of pixels in a video frame image that match the target point and the endpoints of the target line segment, such as... Figure 2d As shown, the coordinates of the target point are P(s, t). In this embodiment, we take drawing a parallel line along the horizontal axis of the two-dimensional coordinate system through the target point as an example. The parallel line intersects the diagonal P1P3 to obtain P. L The intersection of the diagonal P2P4 and P is obtained by P. R P L P R That is, the target line segment. P L P R The endpoints of the target line segment.
[0079] It should be noted that this embodiment is only one way to determine the endpoints of the target line segment. Optionally, a parallel line can be drawn through the target point, intersecting the diagonal and sides of the target planar polygon, resulting in multiple intersection points. The two intersection points that are closest to the target point horizontally are then taken as the endpoints of the target line segment. For example, Figure 2d In the diagram, the horizontal line intersects with points P1P2, P2P3, P3P4, P1P4, P1P3, and P2P4 respectively. Calculate the coordinates of each intersection point on the S-axis. Compare the S-axis coordinates of each intersection point with the S-axis coordinates of point P. The intersection point with the closest difference to negative 0 is taken as P. L The intersection point whose difference is closest to positive 0 is taken as P. R This embodiment does not impose any restrictions on the method for determining the endpoints of the target line segment.
[0080] S250. Calculate the planar coordinates of the endpoints of the target line segment.
[0081] by Figure 2d For example, due to P LThe intersection point is obtained by drawing a line parallel to the x-axis of the two-dimensional coordinate system through the target point and intersecting the diagonal P1P3. Therefore, P L The coordinates of the T-axis are the same as those of the target point, that is, P. L The coordinate value of the T-axis is t, according to P L The linear relationship between P1 and P3 can be obtained as follows:
[0082]
[0083] It can be deduced that:
[0084]
[0085] Therefore, we can obtain P. L coordinates (s) L Similarly, P can be obtained (t). R The coordinate values.
[0086] S260. Based on the relationship between the endpoints of the target line segment and their adjacent sides, and the depth relationship between the vertices and diagonal intersections of the target planar polygon, calculate the depth relationship between the endpoints of the target line segment and the diagonal intersections of the target planar polygon.
[0087] In this embodiment of the invention, since the depth value relationship between the vertices of the target plane polygon and the intersection of the diagonals has been obtained in S220, and since the depth values of the points on the same diagonal have a linear relationship, the depth value relationship between the endpoints of the target line segment and the intersection of the diagonals can be calculated.
[0088] by Figure 2d For example, the depth value of the diagonal intersection point P is Z, the depth value of vertex P1 is K1Z, and the depth value of vertex P3 is K3Z. Through P... L The linear relationship between P1 and P3 is used to obtain P. L The depth value of the point is K L Z. Similarly, P can be obtained. R The depth value is K R Z.
[0089] S270. Based on the depth relationship between the endpoint of the target line segment and the intersection of the diagonals of the target planar polygon, and the linear relationship of the edges in the video frame image that matches the adjacent edge of the endpoint of the target line segment, calculate the texture coordinate value of the endpoint of the target line segment in the video frame image.
[0090] For example, Figure 2eThis is a schematic diagram of pixels matching the target point and the endpoints of the target line segment in a video frame image according to Embodiment 2 of the present invention. In the video frame image, the vertices of the region matching the target planar polygon in the scene area are P'1(s1', t1'), P'2(s2', t2'), P'3(s3', t3'), and P'4(s4', t4'). The pixel matching the target point is P', and the pixel matching the endpoints of the target line segment is P'. L '(s L ',t L '), P R '(s R ',t R ').
[0091] like Figure 2e As shown, after obtaining the depth values of the endpoints of the target line segment, and mapping them to the video frame image after projection, we can see that:
[0092]
[0093] Furthermore, due to:
[0094]
[0095] Therefore, P can be calculated. L '(s L ',t L Similarly, we can obtain P. R '(s R ',t R '). P L '(s L ',t L '), P R '(s R ',t R ') is P L P R Texture coordinates in a video frame image.
[0096] S280. Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, and the relationship between the planar coordinates of the target point and the planar coordinates of the endpoints of the target line segment, calculate the texture coordinates of the target point that match the video frame image.
[0097] In this embodiment of the invention, the planar coordinates of the endpoints of the target line segment and the target point in the scene area, as well as the texture coordinates of the endpoints of the target line segment in the video frame image, have been determined. Since there is a linear relationship between the endpoints of the target line segment and the target point in the scene area, the texture coordinates of the target point can be calculated based on the linear relationship between the endpoints of the target line segment and the target point in the scene area, and the texture coordinates of the endpoints of the target line segment.
[0098] S290. Render based on the texture coordinates of the target point to achieve video fusion.
[0099] In this embodiment of the invention, after obtaining the texture coordinate value of the target point, image rendering can be performed based on the pixel information of the pixel points in the video frame image that match the texture coordinate value, thereby achieving video fusion.
[0100] Example 3
[0101] Figure 3 This is a schematic diagram of a video fusion device provided in Embodiment 3 of the present invention. This device is generally integrated into a computer device and used in conjunction with a video shooting device. The device includes: a planar polygon division module 310, a depth value calculation module 320, a texture coordinate value calculation module 330, and a target point texture coordinate value calculation module 340. Wherein:
[0102] The planar polygon division module 310 is used to acquire video frame images and scene regions that match the video frame images, and to divide the video frame images and scene regions into a number of corresponding planar polygons.
[0103] The depth value calculation module is used to determine the endpoints of target line segments for target points within the target planar polygon in the scene area, and to calculate the planar coordinates and depth value information of the endpoints of the target line segments.
[0104] The texture coordinate calculation module is used to calculate the texture coordinate values of the endpoints of the target line segment and the video frame image based on the depth value information of the endpoints of the target line segment.
[0105] The target point texture coordinate value calculation module is used to calculate the texture coordinate value of the target point and the video frame image based on the planar coordinate value and texture coordinate value of the endpoint of the target line segment, so as to render according to the texture coordinate value of the target point and realize video fusion.
[0106] The technical solution of this embodiment divides the scene region and its matching video frame image into the same number of planar polygons. For target points within the target planar polygons in the scene region, the endpoints of target line segments are determined. The planar coordinates and depth information of the target line segment endpoints are calculated. Based on the depth information of the target line segment endpoints, the texture coordinates of the target line segment endpoints in the video frame image are determined. Finally, based on the planar coordinates and texture coordinates of the target line segment endpoints, the texture coordinates of the target point are determined, and rendering is performed based on the texture coordinates of the target point. This solves the problem of texture distortion easily occurring in existing video fusion methods, which affects the video fusion effect. It achieves texture mapping according to the divided regions, ensuring the effect of video fusion.
[0107] Based on the above embodiments, the device further includes:
[0108] The 2D coordinate transformation module is used to establish a 2D coordinate system on the plane containing the target planar polygon in the scene area and perform 2D coordinate transformation on the target planar polygon.
[0109] Based on the above embodiments, the device further includes:
[0110] The depth value relationship calculation module is used to determine the depth value relationship between the vertices and diagonal intersections of the target planar polygon based on the relationship between the target planar polygon in the scene area and the planar polygons in the video frame image that matches the target planar polygon.
[0111] Based on the above embodiments, the depth value relationship calculation module is specifically used for:
[0112] The depth relationship between the vertices and the intersection points of the diagonals of the target planar polygon can be calculated using the following formula:
[0113]
[0114]
[0115]
[0116]
[0117] Where Z is the depth value of the diagonal intersection point, Z1, Z2, Z3, and Z4 are the depth values of the four vertices of the target planar polygon, L1', L2', L3', and L4' are the pixel distances from the four vertices of the planar polygon in the video frame image matching the target planar polygon to the diagonal intersection point, |P1P|, |P2P|, |P3P|, and |P4P| are the spatial coordinate distances between the four vertices of the target planar polygon and the diagonal intersection point, and |P1P3| and |P2P4| are the spatial coordinate distances between the two vertices corresponding to the two diagonals.
[0118] Based on the above embodiments, the depth value information calculation module 320 includes:
[0119] The target line segment endpoint determination unit is used to draw a parallel line through the target point along the coordinate axis direction of the two-dimensional coordinate system. The parallel line intersects with the adjacent side of the target plane polygon to obtain the endpoint of the target line segment. The adjacent side includes the diagonal and / or edge of the target plane polygon.
[0120] The planar coordinate value calculation unit is used to calculate the planar coordinate values of the endpoints of the target line segment;
[0121] The depth value relationship calculation unit is used to calculate the depth value relationship between the endpoint of the target line segment and the intersection point of the diagonal of the target plane polygon based on the relationship between the endpoint of the target line segment and its adjacent side, and the depth value relationship between the vertex and the intersection point of the diagonal of the target plane polygon.
[0122] Based on the above embodiments, the texture coordinate value calculation module 330 includes:
[0123] The texture coordinate calculation unit is used to calculate the texture coordinate value of the target line segment endpoint in the video frame image based on the depth value relationship between the intersection point of the target line segment endpoint and the diagonal of the target planar polygon, and the linear relationship of the edges in the video frame image that match the adjacent edge of the target line segment endpoint.
[0124] Based on the above embodiments, the target point texture coordinate value calculation module 340 includes:
[0125] The target point texture coordinate value calculation unit is used to calculate the texture coordinate value of the target point matching the video frame image based on the planar coordinate value and texture coordinate value of the endpoint of the target line segment, as well as the relationship between the planar coordinate value of the target point and the planar coordinate value of the endpoint of the target line segment.
[0126] The video fusion apparatus provided in the embodiments of the present invention can execute the video fusion method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0127] Example 4
[0128] Figure 4 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention, as shown below. Figure 4 As shown, the computer device includes a processor 70, a memory 71, an input device 72, and an output device 73; the number of processors 70 in the computer device can be one or more. Figure 4 Taking a processor 70 as an example; the processor 70, memory 71, input device 72, and output device 73 in a computer device can be connected via a bus or other means. Figure 4 Taking the example of a connection between China and Israel via a bus.
[0129] The memory 71, as a computer-readable storage medium, can be used to store software programs, computer-executable programs, and modules, such as modules corresponding to the video fusion method in this embodiment of the invention (e.g., the planar polygon division module 310, depth value calculation module 320, texture coordinate value calculation module 330, and target point texture coordinate value calculation module 340 in the video fusion apparatus). The processor 70 executes various functional applications and data processing of the computer device by running the software programs, instructions, and modules stored in the memory 71, thereby implementing the aforementioned video fusion method. The method includes:
[0130] Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons.
[0131] For target points within the target planar polygon in the scene area, determine the endpoints of the target line segments, and calculate the planar coordinates and depth information of the endpoints of the target line segments;
[0132] Based on the depth information of the endpoints of the target line segment, calculate the texture coordinate values that match the endpoints of the target line segment with the video frame image;
[0133] Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, the texture coordinates of the target point are calculated to match the video frame image, and then the video is rendered based on the texture coordinates of the target point to achieve video fusion.
[0134] The memory 71 may primarily include a program storage area and a data storage area. The program storage area may store the operating system and at least one application program required for a given function; the data storage area may store data created based on terminal usage. Furthermore, the memory 71 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory, or other non-volatile solid-state storage device. In some instances, the memory 71 may further include memory remotely located relative to the processor 70, which can be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0135] Input device 72 can be used to receive input digital or character information, and to generate key signal inputs related to user settings and function control of the computer device. Output device 73 may include display devices such as a display screen.
[0136] Example 5
[0137] Embodiment 5 of the present invention also provides a storage medium containing computer-executable instructions, which, when executed by a computer processor, are used to perform a video fusion method, the method comprising:
[0138] Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons.
[0139] For target points within the target planar polygon in the scene area, determine the endpoints of the target line segments, and calculate the planar coordinates and depth information of the endpoints of the target line segments;
[0140] Based on the depth information of the endpoints of the target line segment, calculate the texture coordinate values that match the endpoints of the target line segment with the video frame image;
[0141] Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, the texture coordinates of the target point are calculated to match the video frame image, and then the video is rendered based on the texture coordinates of the target point to achieve video fusion.
[0142] Of course, the computer-executable instructions provided in the embodiments of the present invention are not limited to the method operations described above, but can also perform related operations in the video fusion method provided in any embodiment of the present invention.
[0143] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0144] It is worth noting that in the embodiments of the above-mentioned video fusion device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.
[0145] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.
Claims
1. A video fusion method, characterized in that, include: Obtain video frame images and scene regions that match the video frame images, and divide the video frame images and scene regions into a number of corresponding planar polygons. For a target planar polygon in the scene area, establish a two-dimensional coordinate system on the plane where the target planar polygon is located, and perform two-dimensional coordinate transformation on the target planar polygon; For a target point within a target planar polygon in the scene area, determine the endpoint of the target line segment, including: drawing a parallel line through the target point along the coordinate axis of the two-dimensional coordinate system, the parallel line intersecting with the adjacent edge of the target planar polygon to obtain the endpoint of the target line segment, the adjacent edge including the diagonal and / or side of the target planar polygon; and calculating the planar coordinate value and depth value information of the endpoint of the target line segment, including: calculating the planar coordinate value of the endpoint of the target line segment; and calculating the depth value relationship between the endpoint of the target line segment and the intersection point of the diagonal of the target planar polygon based on the relationship between the endpoint of the target line segment and its adjacent edge, and the depth value relationship between the intersection point of the vertex and the diagonal of the target planar polygon. Based on the depth value information of the endpoint of the target line segment, calculate the texture coordinate value of the endpoint of the target line segment and the video frame image that match it. This includes: calculating the texture coordinate value of the endpoint of the target line segment in the video frame image based on the depth value relationship between the endpoint of the target line segment and the intersection of the diagonals of the target planar polygon, and the linear relationship of the edges in the video frame image that match the adjacent edge of the endpoint of the target line segment. Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, the texture coordinates of the target point are calculated to match the video frame image, and then the video is rendered based on the texture coordinates of the target point to achieve video fusion.
2. The method according to claim 1, characterized in that, Before performing a two-dimensional coordinate transformation on the target planar polygon, the following is also included: Based on the relationship between the target planar polygon in the scene area and the planar polygons in the video frame image that matches the target planar polygon, determine the depth value relationship between the vertices and the intersection points of the diagonals of the target planar polygon.
3. The method according to claim 2, characterized in that, Based on the relationship between the target planar polygon in the scene area and the planar polygons in the video frame images that match the target planar polygon, determine the depth value relationship between the vertices and diagonal intersections of the target planar polygon, including: The depth relationship between the vertices and the intersection points of the diagonals of the target planar polygon can be calculated using the following formula: ; ; ; ; Where Z is the depth value of the intersection of the diagonals. , , as well as These are the depth values of the four vertices of the target planar polygon. , , , These are the pixel distances from the four vertices of the planar polygon in the video frame image that matches the target planar polygon to the intersection of its diagonals. , , , These are the spatial coordinate distances between the four vertices of the target planar polygon and the intersection of its diagonals. , These represent the spatial coordinate distances between the two vertices corresponding to the two diagonals.
4. The method according to claim 1, characterized in that, Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, calculate the texture coordinates of the target point that match the video frame image, including: Based on the planar coordinates and texture coordinates of the endpoints of the target line segment, and the relationship between the planar coordinates of the target point and the planar coordinates of the endpoints of the target line segment, the texture coordinates of the target point that match the video frame image are calculated.
5. A video fusion apparatus, characterized in that, include: The planar polygon division module is used to acquire video frame images and scene regions that match the video frame images, and to divide the video frame images and scene regions into a number of corresponding planar polygons. The 2D coordinate transformation module is used to establish a 2D coordinate system on the plane where the target planar polygon is located for the target planar polygon in the scene area, and to perform 2D coordinate transformation on the target planar polygon; The depth value calculation module is used to determine the endpoints of target line segments for target points within the target planar polygon in the scene area, and to calculate the planar coordinates and depth value information of the endpoints of the target line segments. The depth value calculation module includes: The target line segment endpoint determination unit is used to draw a parallel line through the target point along the coordinate axis direction of the two-dimensional coordinate system. The parallel line intersects with the adjacent side of the target plane polygon to obtain the endpoint of the target line segment. The adjacent side includes the diagonal and / or edge of the target plane polygon. The planar coordinate value calculation unit is used to calculate the planar coordinate values of the endpoints of the target line segment; The depth value relationship calculation unit is used to calculate the depth value relationship between the endpoints of the target line segment and the intersection of the diagonals of the target plane polygon based on the relationship between the endpoints of the target line segment and their adjacent sides, and the depth value relationship between the vertices and the intersection points of the diagonals of the target plane polygon. The texture coordinate calculation module is used to calculate the texture coordinate values of the endpoints of the target line segment and the video frame image based on the depth value information of the endpoints of the target line segment. The texture coordinate value calculation module includes: The texture coordinate calculation unit is used to calculate the texture coordinate value of the target line segment endpoint in the video frame image based on the depth value relationship between the intersection point of the target line segment endpoint and the diagonal of the target planar polygon, and the linear relationship of the edges in the video frame image that matches the adjacent edge of the target line segment endpoint. The target point texture coordinate value calculation module is used to calculate the texture coordinate value of the target point and the video frame image based on the planar coordinate value and texture coordinate value of the endpoint of the target line segment, so as to render according to the texture coordinate value of the target point and realize video fusion.
6. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the video fusion method as described in any one of claims 1-4.
7. A storage medium containing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the video fusion method as described in any one of claims 1-4.
Citation Information
Patent Citations
Two-dimensional video and three-dimensional scene fusion method and device, equipment and storage medium
CN112184922A
Texture paving method and device and storage medium
CN113689536A