Cross-device rendering method, device, equipment and medium
By introducing the intermediate camera model, the images of different camera devices are converted into unified intermediate images, and a camera-independent three-dimensional Gaussian sputtering model is constructed, which solves the problem of poor generalization ability in the prior art and realizes a wider camera device adaptability and simplified model construction process.
Patent Information
- Application Number
- CN202510466352.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The prior art has poor generalization capabilities when processing images of different camera models, and when extended to new camera models, it is necessary to re-derive the explicit gradient flow, resulting in increased model construction complexity, limiting the generalization capabilities of 3D-GS.
By introducing an intermediate camera model, the original images from different camera devices are converted into intermediate images, and a camera-independent three-dimensional Gaussian sputtering model is constructed to achieve cross-device rendering.
Enhanced generalization capabilities of 3D-GS, making it suitable for a wide variety of camera devices, simplifies the model building process, and supports a variety of flexible rendering requirements.
Smart Images

Figure CN119991919A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a cross-device rendering method, apparatus, device and medium. Background Art
[0002] Realistic and real-time 3D reconstruction and rendering are the focus of academic research and industry, and are applied to virtual reality (VR), augmented reality (AR) and various robotic tasks. 3D-GS (3D Gaussian Splatting) and its variants have made significant progress in these areas. 3D-GS demonstrates state-of-the-art rendering quality and speed by optimizing a set of 3D Gaussians to effectively model the scene.
[0003] However, related technologies are usually designed for pinhole camera models, which makes it difficult to process images from different camera models, resulting in poor generalization ability. Moreover, when expanding to a new camera model, a new explicit gradient flow specific to the camera model needs to be re-derived. As the complexity of the camera model increases, the complexity of model construction will also increase significantly, and the 3D-GS model construction process becomes more cumbersome, which seriously limits its generalization ability. Summary of the invention
[0004] In order to overcome the problems existing in the related art, the present disclosure provides a cross-device rendering method, apparatus, device and medium. The technical solution of the present disclosure is as follows: According to a first aspect of an embodiment of the present disclosure, a cross-device rendering method is provided, including: Get raw images from different camera devices; According to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image, the original image is converted into an intermediate image, wherein the intermediate image carries a corresponding intermediate pose; Constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; According to the rendering direction of the target intermediate image, the target intermediate image is rendered using the three-dimensional Gaussian sputtering model, and the target intermediate image is inversely mapped to the inverse mapping area corresponding to the target original image to obtain the target original image.
[0005] Optionally, it also includes: Initializing a learnable camera embedding by the number of raw camera models; the number of the raw camera models is consistent with the number of categories of the camera devices; Splicing the learnable camera embedding and the target original image along the image channel direction to obtain a spliced image; Inputting the stitched image into an optimization network to output a pixel-level target transformation and a restored image, wherein the optimization network is used to optimize the target original image; An optimized target original image is obtained through the target transformation, the restored image and the target original image.
[0006] Optionally, it also includes: Filtering out redundant three-dimensional Gaussians in the three-dimensional Gaussian sputtering model by using the inverse mapping area and the rendering direction to obtain a filtered three-dimensional Gaussian sputtering model, wherein the redundant three-dimensional Gaussians are three-dimensional Gaussians that do not participate in generating the target original image; Obtaining a target intermediate image by rendering in the rendering direction using the three-dimensional Gaussian sputtering model includes: Based on the filtered three-dimensional Gaussian sputtering model, assigning effective three-dimensional Gaussians to corresponding rendering directions to obtain three-dimensional Gaussians corresponding to each rendering direction, wherein the effective three-dimensional Gaussians are three-dimensional Gaussians that participate in rendering the intermediate image in the rendering direction; The target intermediate image is obtained by rendering the three-dimensional Gaussian corresponding to each rendering direction using the three-dimensional Gaussian sputtering model.
[0007] Optionally, converting the original image into an intermediate image according to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image includes: Determining, according to the camera intrinsic parameters of the intermediate camera model, a latitude rotation and a longitude rotation of the intermediate image relative to an original pose of the original image; Determine a rotation matrix of the intermediate image relative to the original image through the latitude rotation and the longitude rotation; Determining an intermediate pose of the intermediate image by using the rotation matrix and the original pose; The original image is converted into the intermediate image through the rotation matrix, the back projection function corresponding to the original camera model, and the projection function of the intermediate camera model.
[0008] Optionally, the intermediate image includes a reference intermediate image with texture mapping and an intermediate image to be optimized without texture mapping, and further includes: Determining, according to the intermediate camera model, an adaptive mapping direction of the intermediate image to be optimized relative to the reference intermediate image, and image transformation parameters corresponding to each adaptive mapping direction; Determining a texture region of the reference intermediate image, and determining an intermediate pose corresponding to the reference intermediate image; Determining adaptive mapping images in various adaptive mapping directions according to the texture area of the reference intermediate image and the image transformation parameters; Replacing the intermediate image to be optimized with the adaptive mapping image to obtain an updated intermediate image to be optimized; Determining an updated intermediate pose corresponding to the intermediate image to be optimized according to the intermediate pose corresponding to the reference intermediate image and the adaptive mapping direction; According to the intermediate image and the corresponding intermediate posture, a three-dimensional Gaussian sputtering model is constructed, including: A three-dimensional Gaussian sputtering model is constructed according to the reference intermediate image and its corresponding intermediate pose and the updated intermediate image to be optimized and its corresponding intermediate pose.
[0009] Optionally, constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose includes: Constructing a first initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; Acquire a first initial target intermediate image having the same position and posture as the intermediate image through the first initial three-dimensional Gaussian sputtering model; Determining a first loss using the intermediate image and the first initial target intermediate image; adjusting parameters of the first initial three-dimensional Gaussian sputtering model according to the first loss; According to the above steps, the parameters of the first initial three-dimensional Gaussian sputtering model are adjusted multiple times to obtain a three-dimensional Gaussian sputtering model.
[0010] Optionally, constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose includes: Constructing a second initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; Acquire a second initial target original image having the same posture as the original image through the second initial three-dimensional Gaussian sputtering model; Obtaining an initial restored image and an initial target transformation corresponding to the second initial target original image through an initial optimization network; Determining an optimized second initial target original image according to the initial restored image, the initial target transformation and the second initial target original image; Determining a distortion loss according to the initial restored image and the original image; Determining a second loss by using the optimized second initial target original image and the original image; According to the distortion loss and the second loss, respectively adjusting the parameters of the initial optimization network and the second initial three-dimensional Gaussian sputtering model; According to the above steps, the parameters of the initial optimized network and the second initial three-dimensional Gaussian sputtering model are adjusted multiple times to obtain the optimized network and the three-dimensional Gaussian sputtering model.
[0011] According to a second aspect of an embodiment of the present disclosure, a cross-device rendering apparatus is provided, including: Acquisition module, used to acquire raw images from different camera devices; A mapping module, used to convert the original image into an intermediate image according to a mapping relationship between an original camera model corresponding to the original image and an intermediate camera model, wherein the intermediate image carries a corresponding intermediate pose; A construction module, used for constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; The determination module is used to render the target intermediate image using the three-dimensional Gaussian sputtering model according to the rendering direction of the target intermediate image, and inversely map the target intermediate image to the inverse mapping area corresponding to the target original image to obtain the target original image.
[0012] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the cross-device rendering method described in the first aspect are implemented.
[0013] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the cross-device rendering method described in the first aspect are implemented.
[0014] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps of the cross-device rendering method described in the first aspect are implemented.
[0015] The present disclosure can process raw images from different camera devices by introducing an intermediate camera model, thereby enhancing the generalization capability of 3D-GS and making it applicable to more types of camera devices. By using the intermediate camera model as a bridge, the tedious process of deriving gradient flows separately for each camera model is avoided, thereby simplifying the entire model building process. The method can render according to the rendering direction of the target intermediate image, and inversely map the rendering result to the inverse mapping area of the target raw image to obtain the target raw image, which can flexibly support a variety of rendering requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings required for use in the description of the embodiments of the present disclosure will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0017] Figure 1 is a schematic diagram of steps of a cross-device rendering method shown in an embodiment of the present disclosure; Figure 2 is a schematic diagram of a cross-device rendering system shown in an embodiment of the present disclosure; Figure 3 is a schematic diagram of a cross-device rendering apparatus shown in an embodiment of the present disclosure; Figure 4 It is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0019] The terms "first", "second", etc. in the specification and claims of the present disclosure are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable under appropriate circumstances, so that the embodiments of the present disclosure can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally of the same type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship. In order to solve the problems existing in the related art, the present disclosure provides a cross-device rendering method, which can construct a camera-independent three-dimensional Gaussian sputtering model to efficiently process multiple camera source images.
[0020] Figure 1 Schematic diagram of a cross-device rendering method shown in an embodiment of the present disclosure. Figure 1 As shown, the method may specifically include the following steps: Step S11: Acquire original images from different camera devices.
[0021] Camera devices may include multiple camera types, such as fisheye cameras, panoramic cameras, etc. Raw images from different camera devices refer to images captured by different camera devices. Different camera types of camera devices have different characteristics of captured images. When different camera devices capture the same object, different information can be captured in one captured image.
[0022] Step S12: according to the mapping relationship between the original camera model corresponding to the original image and the intermediate camera model, the original image is converted into an intermediate image, wherein the intermediate image carries the corresponding intermediate posture.
[0023] When converting any original image into an intermediate image, the original camera model corresponding to the original image is determined. The original camera model is related to the camera device used to shoot the original image. The original camera model corresponding to the original image can be determined by the camera type of the camera device. For example, if an original image is shot by a fisheye camera, then the original image corresponds to the fisheye camera model.
[0024] The intermediate camera model is the camera model corresponding to the intermediate image. The intermediate camera model is unified. When any original image is converted to an intermediate image, the same intermediate camera model is used. The intermediate camera model can be set according to the actual situation.
[0025] An original image can be converted into one or more intermediate images. The intermediate images have a unified image format and can provide a standardized representation. Converting original images captured by different camera devices with different viewing angles and projection rules into a unified intermediate image can ensure the consistency of processing.
[0026] Each intermediate image has its own corresponding intermediate position information, which can indicate the spatial positioning of the intermediate image and provide the spatial information and direction information of the intermediate image in three-dimensional space.
[0027] Step S13: constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate posture.
[0028] After obtaining the intermediate image and its corresponding intermediate pose, the intermediate image and its corresponding intermediate pose are used to construct a three-dimensional Gaussian sputtering model. A unified three-dimensional Gaussian sputtering model is established through the original images from different camera devices. The three-dimensional Gaussian sputtering model is represented as a set of 3D Gaussian points.
[0029] Step S14: according to the rendering direction of the target intermediate image, the target intermediate image is rendered using the three-dimensional Gaussian sputtering model, and the target intermediate image is inversely mapped to the inverse mapping area corresponding to the target original image to obtain the target original image.
[0030] Rendering direction refers to the process of determining the angle and direction from which the image is generated during image rendering.
[0031] The inverse mapping region refers to the reverse mapping of the target intermediate image to a specific region in the target original image. The inverse mapping region is the portion of the target original image that corresponds to the target intermediate image. The inverse mapping region can ensure that the rendered target intermediate image is aligned with the final target original image to avoid distortion of the target original image.
[0032] The rendering direction and the inverse mapping area are pre-calculated based on the mapping relationship between the original camera model and the intermediate camera model. First, the two-dimensional plane corresponding to the final target original image can be determined. After the two-dimensional plane is determined, the mapping relationship between the original camera model and the intermediate camera model can be used to determine which intermediate images of the rendering direction the two-dimensional plane has a mapping relationship with, and which area of the two-dimensional plane these intermediate images have a mapping relationship with. A target original image can be obtained from one or more target intermediate images through inverse mapping.
[0033] Rendering is performed using a three-dimensional Gaussian sputtering model according to a rendering direction to obtain a target intermediate image, and the target intermediate image is obtained by rendering using a three-dimensional Gaussian sputtering model.
[0034] After obtaining the target intermediate image, in order to convert the target intermediate image back to the space of the original image, inverse mapping is required. Inverse mapping refers to the mapping from the target intermediate image to the target original image. By mapping the target intermediate image back to the inverse mapping area corresponding to the target original image, the target intermediate image is accurately converted back to the original image space. After converting back to the original image space, the final target original image is obtained. The inverse mapping mechanism can ensure that the final target original image can maintain the geometric shape and camera characteristics of the original image.
[0035] By adopting the embodiments of the present disclosure, raw images are obtained from different camera devices, and the raw images from these different sources are converted into a unified intermediate image through the mapping relationship between the original camera model and the intermediate camera model, which ensures that the raw images of different devices have consistency and a unified processing framework during the processing process, solves the problem that images taken by different camera devices have different viewing angles, field angles, distortions, etc., and ensures the efficiency of cross-device rendering, so that a unified and original camera model-independent three-dimensional Gaussian sputtering model can be constructed based on these intermediate images and corresponding intermediate poses. During the rendering process, according to the rendering direction of the target intermediate image, a unified three-dimensional Gaussian sputtering model is used to render the target intermediate image. The rendered target intermediate image will be mapped to the inverse mapping area corresponding to the target original image through the inverse mapping technology, and the target intermediate image can be restored to the target original image.
[0036] Among them, in an optional embodiment, according to the mapping relationship between the original camera model corresponding to the original image and the intermediate camera model, the original image is converted into the intermediate image, including: determining the latitude rotation and longitude rotation of the intermediate image relative to the original posture of the original image according to the camera intrinsic parameters of the intermediate camera model; determining the rotation matrix of the intermediate image relative to the original image through the latitude rotation and the longitude rotation; determining the intermediate posture of the intermediate image through the rotation matrix and the original posture; converting the original image into the intermediate image through the rotation matrix, the back projection function corresponding to the original camera model, and the projection function of the intermediate camera model.
[0037] The original camera model is the camera model used when shooting the original image, which contains the camera's internal parameters and external parameters. The internal parameters define the camera's internal characteristics, which determine how the camera projects points in the three-dimensional world onto the two-dimensional image plane. Specifically, they can include information such as focal length, principal point coordinates, and distortion coefficients. The external parameters describe the camera's position and direction in the world coordinate system, and determine how the camera shoots the same scene from different angles and positions. The external parameters can be understood as the camera's position information.
[0038] The intermediate camera model is a standardized camera model used to construct a three-dimensional Gaussian sputtering model. The internal parameters of the intermediate camera model can be expressed as follows:
[0039] In the above formula, represents the width of the reference intermediate image, Indicates the height of the reference intermediate image.
[0040] Through the internal parameters of the intermediate camera model, six fixed rendering directions that are perpendicular to each other can be determined, namely front, back, left, right, top, and bottom. The width of the intermediate image corresponding to each rendering direction is and high The intermediate images in the six rendering directions can be stitched into a cube, and each face of the cube is a two-dimensional plane corresponding to the intermediate image. The intermediate image can be the pinhole image corresponding to the pinhole camera.
[0041] The rendering direction corresponding to the internal parameters of the intermediate camera model can be determined as a fixed mapping direction, and the fixed mapping direction is used to convert the original image into the intermediate image. The fixed mapping direction corresponds to an intermediate image.
[0042] The original position information corresponding to the original image represents the position and orientation of the original camera model in the three-dimensional space. The position information may include the position of the camera in the three-dimensional space and the rotation angle of the camera.
[0043] It is necessary to determine the rotation transformation relationship between the original image and the intermediate image, so as to convert the original image to a specific fixed mapping direction. The intermediate image corresponding to each fixed mapping direction has a specific rotation transformation relationship relative to the original image, and the rotation transformation relationship changes the observation angle and viewing angle of the original image. A corresponding rotation transformation relationship can be determined for each fixed mapping direction.
[0044] Specifically, the latitude rotation and longitude rotation of each intermediate image relative to the original position can be determined. Latitude rotation and longitude rotation represent different dimensions of rotation. Latitude rotation refers to the angle of rotation along the horizontal direction. For example, the original image is rotated a certain angle in the horizontal direction from the camera position of the original image to obtain different latitude rotation angles. Longitude rotation refers to the angle of rotation along the vertical direction. For example, the original image is rotated along the vertical direction of the image from the camera position of the original image to obtain different longitude rotation angles.
[0045] A rotation matrix can be determined by latitude rotation and longitude rotation. The rotation matrix is used to represent the rotation transformation relationship of the intermediate image relative to the original image. The rotation matrix is used to describe the relative rotation relationship of the original image to the intermediate image in three-dimensional space.
[0046] After obtaining the rotation matrix, the intermediate position corresponding to the intermediate image is determined by combining the original position corresponding to the original image. The intermediate position can be obtained by the following formula:
[0047] In the above formula, Indicates the intermediate position corresponding to the intermediate image, Indicates the original position information corresponding to the original image, represents the rotation matrix of the intermediate image relative to the original image, where represents the latitude rotation, Indicates longitude rotation.
[0048] The rotation matrix is used to transform the coordinates of the original camera coordinate system to the coordinates of the target coordinate system. The rotation matrix can be used to rotate points in the original camera coordinate system, specifically, pixels in the original image, to the intermediate image coordinate system. Therefore, after obtaining the rotation matrix of any fixed rendering direction, the original image is converted into the intermediate image of the fixed rendering direction according to the rotation matrix, the back projection function corresponding to the original camera model, and the projection function of the intermediate camera model. The intermediate image can be obtained specifically by the following formula:
[0049] In the above formula, represents the projection function of the intermediate camera model, represents the back-projection function of the original camera model, represents the original image, Represents the original image To the intermediate image The mapping of Represents the rotation matrix of the intermediate image relative to the original image.
[0050] Among them, bilinear interpolation can be used to fill the blank areas of the intermediate image. The blank areas of the intermediate image refer to the areas of the intermediate image that are not covered by the data of the original image due to the viewing angle or projection method during the image conversion process, which appear as blank or information-free areas in the image.
[0051] By using the embodiments of the present disclosure, the coordinate systems of the original image and the intermediate image are accurately aligned by calculating the rotation matrix, so that the conversion from the original image to the intermediate image becomes accurate. By combining the back projection and projection functions, it can be further ensured that the points in the three-dimensional space can be correctly mapped to the two-dimensional image plane to avoid distortion and information loss. According to the camera of the intermediate camera model, Among them, in an optional embodiment, the intermediate image includes a baseline intermediate image with texture mapping and an intermediate image to be optimized without texture mapping, and also includes: determining the adaptive mapping direction of the intermediate image to be optimized relative to the baseline intermediate image according to the intermediate camera model, and the image transformation parameters corresponding to each adaptive mapping direction; determining the texture area of the baseline intermediate image, and determining the intermediate pose corresponding to the baseline intermediate image; determining the adaptive mapping image of each adaptive mapping direction according to the texture area of the baseline intermediate image and the image transformation parameters; replacing the intermediate image to be optimized with the adaptive mapping image to obtain an updated intermediate image to be optimized; determining the updated intermediate pose corresponding to the intermediate pose corresponding to the baseline intermediate image and the adaptive mapping direction; constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose, including: constructing a three-dimensional Gaussian sputtering model according to the baseline intermediate image and its corresponding intermediate pose and the updated intermediate image to be optimized and its corresponding intermediate pose.
[0052] After mapping an original image to one or more intermediate images, the intermediate images obtained have different fields of view of the original camera model, resulting in intermediate images with less texture areas or even no texture in the intermediate images obtained based on a fixed mapping direction. The intermediate images can be classified according to the texture areas of each intermediate image, including reference intermediate images with texture mapping and intermediate images to be optimized without texture mapping. A texture threshold can be set to classify the intermediate images based on the comparison result of the texture area of the intermediate image and the texture threshold.
[0053] After determining the reference intermediate image and the intermediate image to be optimized, the intermediate image to be optimized needs to be updated according to the reference intermediate image, specifically, the texture area of the reference intermediate image is processed and converted to the intermediate image to be optimized, thereby updating the intermediate image to be optimized.
[0054] The texture region of the reference intermediate image and the corresponding intermediate position are determined, and the texture region can be represented by the width and height of the circumscribed rectangle of the texture region. After the texture region of the reference intermediate image is determined, the texture region of the reference intermediate image is transformed according to the relative position relationship between the reference intermediate image and the intermediate image to be optimized.
[0055] Determine the adaptive mapping direction of the reference intermediate image and the intermediate image to be optimized, and the adaptive mapping direction can represent the relative position relationship between the reference intermediate image and the intermediate image to be optimized. Specifically, when the camera intrinsic parameters of the intermediate camera model can determine 6 fixed mapping directions that are perpendicular to each other, the adaptive mapping directions of the reference intermediate image and the intermediate image to be optimized can be divided into three categories: front-back, left-right, and top-bottom.
[0056] Corresponding image transformation parameters can be formulated for each type of adaptive mapping direction. The image transformation parameters are used to represent the transformation of the texture region between the reference intermediate image and the intermediate image to be optimized.
[0057] The adaptive mapping image represents the result of converting the texture area in the reference original image to the intermediate image to be optimized, and the texture area in the adaptive mapping image is adjacent to the image boundary. After the adaptive mapping image replaces the intermediate image to be optimized, the intermediate image to be optimized is updated. The intermediate position corresponding to the updated optimized intermediate image can be determined by the intermediate position corresponding to the reference intermediate image and the adaptive mapping direction.
[0058] After the intermediate image to be optimized is updated, a three-dimensional Gaussian sputtering model is constructed by using the updated intermediate image to be optimized and its corresponding intermediate position information, as well as the reference intermediate image and its corresponding intermediate position information.
[0059] Specifically, when the camera intrinsic parameters of the intermediate camera model can determine six fixed mapping directions that are perpendicular to each other, the intermediate image to be optimized in each type of adaptive mapping direction can be updated in the following manner.
[0060] (1) The adaptive mapping direction is the front-to-back direction When the adaptive mapping direction between the reference intermediate image and the intermediate image to be optimized is the front-to-back direction, the intermediate position information corresponding to the intermediate image to be optimized need not be updated, and only the adaptive mapping image of the intermediate image to be optimized needs to be determined. The adaptive mapping image can be determined by the texture area obtained by the transformation, and the transformation of the texture area is determined based on the image transformation parameters. The image transformation parameters in the front-to-back direction may not modify the size of the texture area in the reference intermediate image, and the texture area of the intermediate image to be optimized can be updated by the following formula:
[0061] In the above formula, Represents the width of the bounding rectangle of the texture area corresponding to the reference intermediate image, Represents the height of the bounding rectangle of the texture area corresponding to the reference intermediate image, Indicates the updated width of the intermediate image to be optimized. Indicates the height of the updated intermediate image to be optimized.
[0062] (2) The adaptive mapping direction is left and right When the adaptive mapping direction between the reference intermediate image and the intermediate image to be optimized is the left-right direction, the intermediate position and texture area corresponding to the intermediate image to be optimized needs to be updated.
[0063] The intermediate position corresponding to the intermediate image to be optimized can be determined by the following formula.
[0064]
[0065] in, represents the longitude rotation between the reference intermediate image and the intermediate image to be optimized, Indicates the original position information corresponding to the original image corresponding to the reference intermediate image and the intermediate image to be optimized, represents the latitudinal rotation of the reference intermediate image relative to the original image pose, Represents the rotation matrix of the intermediate image to be optimized relative to the original image.
[0066] The longitude rotation between the reference intermediate image and the intermediate image to be optimized can be determined by the following formula:
[0067] In the above formula, Represents the width of the bounding rectangle of the texture area corresponding to the reference intermediate image, Represents the focal length of the middle camera on the x-axis of the image plane.
[0068] The texture area of the adaptive mapping image corresponding to the intermediate image to be optimized can be determined by the following formula:
[0069] In the above formula, Indicates the width of the bounding rectangle of the texture area corresponding to the reference intermediate image, Represents the height of the bounding rectangle of the texture area corresponding to the reference intermediate image, Indicates the updated width of the intermediate image to be optimized. Indicates the height of the updated intermediate image to be optimized. represents the width of the reference intermediate image, Represents the focal length of the middle camera on the x-axis of the image plane.
[0070] (3) The adaptive mapping direction is up and down When the adaptive mapping direction between the reference intermediate image and the intermediate image to be optimized is the up-down direction, the intermediate position and texture area corresponding to the intermediate image to be optimized needs to be updated.
[0071] The intermediate position corresponding to the intermediate image to be optimized can be determined by the following formula.
[0072]
[0073] in, represents the latitudinal rotation between the reference intermediate image and the intermediate image to be optimized, Indicates the original position information corresponding to the original image corresponding to the reference intermediate image and the intermediate image to be optimized, represents the longitude rotation of the reference intermediate image relative to the original image pose, Represents the rotation matrix of the intermediate image to be optimized relative to the original image.
[0074] The latitude rotation between the reference intermediate image and the intermediate image to be optimized can be determined by the following formula:
[0075] In the above formula, Represents the height of the bounding rectangle of the texture area corresponding to the reference intermediate image, Represents the focal length of the middle camera in the y-axis direction of the image plane.
[0076] The texture area of the adaptive mapping image corresponding to the intermediate image to be optimized can be determined by the following formula:
[0077] In the above formula, Indicates the width of the bounding rectangle of the texture area corresponding to the reference intermediate image, Represents the height of the bounding rectangle of the texture area corresponding to the reference intermediate image, Indicates the updated width of the intermediate image to be optimized. Indicates the height of the updated intermediate image to be optimized. represents the height of the reference intermediate image, Represents the focal length of the middle camera in the y-axis direction of the image plane.
[0078] By adopting the embodiments of the present disclosure, the intermediate image to be optimized can be updated through the calculation of adaptive mapping direction and image transformation parameters, so that the texture-free area of the intermediate image to be optimized can be filled and aligned with the texture area of the reference intermediate image. This can reduce the texture-free area in the intermediate image and optimize the image quality, making the final image smoother and more detailed.
[0079] Among them, in an optional embodiment, it also includes: filtering out redundant three-dimensional Gaussians in the three-dimensional Gaussian sputtering model through the inverse mapping area and the rendering direction to obtain a filtered three-dimensional Gaussian sputtering model, and the redundant three-dimensional Gaussians are three-dimensional Gaussians that do not participate in generating the target original image; through the rendering direction, using the three-dimensional Gaussian sputtering model to render a target intermediate image, including: based on the filtered three-dimensional Gaussian sputtering model, allocating valid three-dimensional Gaussians to corresponding rendering directions to obtain three-dimensional Gaussians corresponding to each rendering direction, and the valid three-dimensional Gaussians are three-dimensional Gaussians that participate in rendering the intermediate image of the rendering direction; through the three-dimensional Gaussians corresponding to each rendering direction, using the three-dimensional Gaussian sputtering model to render the target intermediate image.
[0080] According to the mapping relationship between the original camera model and the intermediate camera model, the correspondence between the intermediate image and the original image is determined. According to the correspondence between the intermediate image and the original image, the rendering direction and the inverse mapping area are pre-calculated. The rendering direction is used to obtain the target intermediate image through the three-dimensional Gaussian sputtering model, and the inverse mapping area is used to obtain the target original image through the target intermediate image.
[0081] The three-dimensional Gaussian sputtering model can be represented by a series of three-dimensional Gaussians. Redundant three-dimensional Gaussians can be determined through the inverse mapping area and the rendering direction. For example, when the target intermediate image is obtained based on the rendering direction through the three-dimensional Gaussian sputtering model, the three-dimensional Gaussians that do not participate in rendering the target intermediate image can be determined. When the target intermediate image is inversely mapped to the target original image, the pixels in the target intermediate image that will not be inversely mapped to the inverse mapping area can be determined, thereby further determining the three-dimensional Gaussians that do not participate in the inverse mapping. These three-dimensional Gaussians that do not participate in rendering the target intermediate image and the three-dimensional Gaussians that do not participate in the inverse mapping can be determined as redundant three-dimensional Gaussians.
[0082] The redundant three-dimensional Gaussian is filtered from the three-dimensional Gaussian sputtering model to obtain a filtered three-dimensional Gaussian sputtering model.
[0083] The effective 3D Gaussians in the filtered 3D Gaussian sputtering model are assigned to the corresponding rendering directions, and each rendering direction will have a set of 3D Gaussians. The effective 3D Gaussians represent the 3D Gaussians that participate in rendering to generate the target intermediate image, and these 3D Gaussians can have a substantial impact on the result of the target original image.
[0084] In each rendering direction, rendering calculation is performed based on the corresponding three-dimensional Gaussian to generate different parts of the target intermediate image. Each effective three-dimensional Gaussian affects a certain area in the target intermediate image. The contributions of all these effective Gaussians are combined to form the final rendering effect and obtain the target intermediate image.
[0085] By adopting the embodiments of the present disclosure, by filtering redundant three-dimensional Gaussians and retaining only the three-dimensional Gaussians that contribute to the target original image, the amount of calculation and memory consumption can be significantly reduced. After filtering the redundant three-dimensional Gaussians, the remaining three-dimensional Gaussians will be assigned to different rendering directions. The three-dimensional Gaussians assigned to different rendering directions will only participate in the calculation in the reverse direction of the rendering, which can further reduce the computational complexity and ensure that computing resources are only used for the parts that truly affect the rendering results. During the rendering process, by assigning Gaussians to corresponding rendering directions, it can be ensured that the three-dimensional Gaussians used in each rendering direction are closely related to the rendering requirements of that direction. By reducing the amount of calculation in each rendering direction, the overall rendering time will be significantly shortened.
[0086] Among them, in an optional embodiment, it also includes: initializing the learnable camera embedding according to the number of original camera models; the number of the original camera models is consistent with the number of categories of the camera equipment; splicing the learnable camera embedding and the target original image along the image channel direction to obtain a spliced image; inputting the spliced image into an optimization network, outputting a pixel-level target transformation and a restored image, the optimization network is used to optimize the target original image; and obtaining an optimized target original image through the target transformation, the restored image and the target original image.
[0087] In cross-device rendering, the original image may come from different original camera models, such as fisheye cameras, panoramic cameras, pinhole cameras, etc. Different original camera models have different field of view angles, imaging methods and other characteristics. Therefore, a learnable camera embedding is initialized for each original camera model. The learnable camera embedding is a vector obtained through training. In this embodiment, the vector can be a vector of length 64. The learnable camera embedding contains the characteristics and information of the camera model. By learning the learnable camera embedding, the optimization network can understand the characteristics of different original camera models, thereby improving accuracy and consistency.
[0088] The target original image obtained by inverse mapping may have problems of distortion, noise or loss of details, which can be processed by the optimization network. The optimization network is a neural network used to process the target original image, repair the distortion, noise or inaccurate areas of the target original image, and thus restore a clearer and more realistic target original image.
[0089] Each learnable camera embedding is stitched with the target original image. Specifically, the stitching can be performed along the channel direction of the image. The learnable camera embedding is added to the features of each pixel or region of the target original image to obtain a stitched image. The channel direction of the image represents the depth direction of the image. The stitched image is the input of the optimization network. After the stitched image is input into the optimization network, the optimization network optimizes the target original image based on the stitched image.
[0090] The output of the optimization network includes pixel-level target transformation and restored image. Pixel-level target transformation means that the optimization network can adjust the color, brightness or details of each pixel of the target original image through learning, reduce distortion and enhance details. Through pixel-level transformation, the optimization network can generate a restored image, which refers to the optimized target original image with richer details and less distortion.
[0091] The pixel-level target transformation and restoration image output by the optimization network and the target original image obtained by inverse mapping are used to obtain the final optimized target original image, which can be specifically achieved by the following formula:
[0092] In the above formula, represents the target original image for optimization, represents the pixel-level target transformation, represents the restored image, represents the target original image obtained by inverse mapping, Represents the inverse mapping splicing edge area of the intermediate image in the target original image, for The complementary region of .
[0093] By using the embodiments of the present disclosure, the learnable camera embedding vector is initialized according to the number of original camera models, and the particularity of each camera device can be modeled, so that the characteristics of each camera model can be obtained through learning, so that it is possible to identify and perform effective image optimization according to the characteristics of different camera models. By splicing the camera embedding and the target original image along the image channel direction, the spliced image can contain both the pixel information of the target original image and the feature information related to the camera model. By combining the target transformation, the restored image and the target original image, an optimized target original image is generated, so that the optimized target original image finally obtained can effectively reduce distortion, improve image quality, and restore more realistic and delicate image details.
[0094] The three-dimensional Gaussian sputtering model is obtained through iterative training. The training process of the three-dimensional Gaussian sputtering model is introduced below.
[0095] (1) In the early stages of training the 3D Gaussian sputtering model Among them, in an optional embodiment, a three-dimensional Gaussian sputtering model is constructed according to the intermediate image and the corresponding intermediate posture, including: constructing a first initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate posture; obtaining a first initial target intermediate image with the same posture as the intermediate image through the first initial three-dimensional Gaussian sputtering model; determining a first loss through the intermediate image and the first initial target intermediate image; adjusting the parameters of the first initial three-dimensional Gaussian sputtering model through the first loss; and adjusting the parameters of the first initial three-dimensional Gaussian sputtering model multiple times according to the above steps to obtain a three-dimensional Gaussian sputtering model.
[0096] In the early stage of training the 3D Gaussian sputtering model, the difference between the intermediate image obtained by mapping the original image and the first initial target intermediate image rendered by the 3D Gaussian sputtering model can be considered to update the 3D Gaussian sputtering model. The intermediate image obtained by mapping the original image is the true value, and the first initial target intermediate image rendered by the 3D Gaussian sputtering model is the predicted value.
[0097] First, determine the intermediate image and the intermediate position information corresponding to the intermediate image. For specific steps, refer to the above step S12.
[0098] According to the intermediate image and its corresponding intermediate position data, a first initial three-dimensional Gaussian sputtering model is constructed. The first initial three-dimensional Gaussian sputtering model is a preliminary and inaccurate representation, but provides a starting point for subsequent optimization.
[0099] The first initial target intermediate image is rendered by using the first initial three-dimensional Gaussian sputtering model, and the first initial target intermediate image has the same position as the intermediate image.
[0100] The difference between the intermediate image and the first initial target intermediate image is calculated, and the difference can be expressed as the first loss, which can be determined by the following formula:
[0101] in, represents the image structure loss between the intermediate image and the first initial target intermediate image, represents the pixel-level loss between the intermediate image and the first initial target intermediate image, is the weight coefficient.
[0102] According to the calculated loss, the parameters of the first initial three-dimensional Gaussian sputtering model are adjusted, thereby completing an iteration of optimization for the first initial three-dimensional Gaussian sputtering model.
[0103] The three-dimensional Gaussian sputtering model is obtained by adjusting the parameters of the first initial three-dimensional Gaussian sputtering model multiple times until the loss reaches an acceptable threshold or the model performance is no longer significantly improved.
[0104] Using the embodiments of the present disclosure, a first initial three-dimensional Gaussian sputtering model is firstly quickly constructed according to the intermediate image and the corresponding intermediate posture, which provides a basis for subsequent model optimization and adjustment. Through the first initial three-dimensional Gaussian sputtering model, a first initial target intermediate image with the same posture as the intermediate image can be obtained. By comparing the intermediate image and the first initial target intermediate image, a first loss is determined, which quantifies the difference between the image generated by the model and the real image, and is the key to subsequent model optimization. The parameters of the initial three-dimensional Gaussian sputtering model are adjusted multiple times. This iterative optimization method can gradually approach the optimal solution and improve the accuracy and robustness of the model.
[0105] (2) In the later stage of training the 3D Gaussian sputtering model Among them, in an optional embodiment, a three-dimensional Gaussian sputtering model is constructed according to the intermediate image and the corresponding intermediate posture, including: constructing a second initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate posture; obtaining a second initial target original image with the same posture as the original image through the second initial three-dimensional Gaussian sputtering model; obtaining an initial restored image and an initial target transformation corresponding to the second initial target original image through an initial optimization network; determining an optimized second initial target original image according to the initial restored image, the initial target transformation and the second initial target original image; determining a distortion loss according to the initial restored image and the original image; determining a second loss through the optimized second initial target original image and the original image; adjusting the parameters of the initial optimization network and the second initial three-dimensional Gaussian sputtering model according to the distortion loss and the second loss respectively; adjusting the parameters of the initial optimization network and the second initial three-dimensional Gaussian sputtering model multiple times according to the above steps to obtain the optimization network and the three-dimensional Gaussian sputtering model.
[0106] In the later stage of training the three-dimensional Gaussian sputtering model, the trained three-dimensional Gaussian sputtering model can be determined as the second initial three-dimensional Gaussian sputtering model. The parameters of the second initial three-dimensional Gaussian sputtering model can be adjusted by the difference between the original image and the second initial target original image finally obtained. The original image is an image taken by the original camera, which is a real value; the second initial target original image is an image obtained by the second initial three-dimensional Gaussian sputtering model, which is a predicted value.
[0107] When the initial optimization network is used to optimize the second initial target original image, the initial optimization network can be jointly trained on the basis of training the second initial three-dimensional Gaussian sputtering model, and the parameters of the initial optimization network can be adjusted in each iteration to obtain the optimized network.
[0108] First, a second initial three-dimensional Gaussian sputtering model is constructed through the intermediate image and the corresponding intermediate pose. The intermediate image is obtained by mapping the original image.
[0109] A second initial target original image with the same bit rate as the original image is generated through a second initial three-dimensional Gaussian sputtering model. The generation of the second initial target original image involves rendering a second target intermediate image through the second initial three-dimensional Gaussian sputtering model, and obtaining the second target original image through inverse mapping of the second target intermediate image.
[0110] Subsequently, an initial restored image and an initial target transformation corresponding to the second initial target original image are obtained through an initial optimization network; an optimized second initial target original image is determined according to the initial restored image, the initial target transformation and the second initial target original image. The optimized second target original image is used to determine a second loss, and the second loss reflects the difference between the real image and the predicted image. The real image refers to the original image, and the predicted image refers to the optimized second target original image.
[0111] In order to focus on optimizing the distortion of the splicing edge of the inverse mapping area of the intermediate image in the second target original image, a distortion loss can be set. The following formula represents the specific way of determining the distortion loss through the initial restored image and the original image:
[0112] In the above formula, represents the target original image for optimization, represents the pixel-level target transformation, represents the restored image, represents the target original image obtained by inverse mapping, Represents the inverse mapping splicing edge area of the intermediate image in the target original image, for The complementary region of .
[0113] A second loss is determined by optimizing the second initial target original image and the original image.
[0114] Finally, the total loss is determined based on the second loss and the distortion loss. The total loss is expressed by the following formula:
[0115] in, represents the second loss, represents the pixel-level loss between the original image and the optimized second initial target original image, represents the image structure loss between the original image and the optimized second initial target original image, is the weight coefficient, is the distortion loss.
[0116] According to the calculated total loss, the parameters of the second initial three-dimensional Gaussian sputtering model and the initial optimization network are adjusted, thereby completing an iteration of adjusting the second initial three-dimensional Gaussian sputtering model and the initial optimization network.
[0117] The three-dimensional Gaussian sputtering model and the initial network are obtained by adjusting the parameters of the first initial three-dimensional Gaussian sputtering model and the initial optimized network multiple times until the total loss reaches an acceptable threshold or the model performance is no longer significantly improved.
[0118] By adopting the embodiments of the present disclosure, by specifically setting the distortion loss to focus on optimizing the splicing edge, the distortion of the splicing edge in the second initial target original image can be reduced in a targeted manner, so that the spliced image is more natural and the transition is smoother. In the later stage of training, the rendering quality of the second initial three-dimensional Gaussian sputtering model has been initially improved, and the model has begun to have a certain ability to render more complex scenes. By jointly optimizing the initial optimization network, the realism and details of the rendering can be greatly improved.
[0119] In order to evaluate the cross-device rendering performance of the present invention, a cross-device rendering evaluation dataset was constructed. The present invention adopts the commonly used performance evaluation criteria for image rendering: Peak Signal-to-Noise Ratio (PSNR), Structural Similarity Index (SSIM), and Learned Perceptual Image Patch Similarity (LPIPS). The test environment of the present invention is Ubuntu 20.04, equipped with an Intel Gold 6330 series CPU with a processing frequency of 2.00GHz, and an NVIDIA GTX4090 image processor with a core frequency of 2235MHz and a video memory capacity of 24GB.
[0120] The present disclosure is tested on a cross-device dataset, and the quantitative results are shown in Table 1. The present disclosure exceeds the advanced 3D-GS method and achieves the most advanced rendering quality, demonstrating its superior performance in cross-device rendering tasks and its strong generalization ability. It is worth noting that the present disclosure can rely entirely on a unified intermediate camera model based on a simple pinhole camera. This model can be easily integrated with the pinhole camera-based methods widely used in the field of vision, further enhancing the scalability of the present disclosure. In addition, the present disclosure can achieve a rendering speed of 0.022085 seconds on the NVIDIA GTX 4090 image processor, meeting the requirements of real-time rendering.
[0121] Table 1: Quantitative test results of the present disclosure in a cross-device dataset
[0122] Based on the same technical concept, the present disclosure provides a cross-device rendering system, which includes an adaptive mapping module, a pre-filtering and allocation module, and a camera-based optimization module. Figure 2 As shown, Figure 2 It is a schematic diagram of a cross-device rendering system shown in an embodiment of the present disclosure.
[0123] The adaptive mapping module is responsible for converting the original images from different cameras into a unified intermediate image. Based on the intermediate image, a universal 3D Gaussian sputtering model is constructed and trained to achieve effective generalization processing of multiple camera source images.
[0124] The pre-filtering and allocation module is used to render the target intermediate images based on the unified three-dimensional Gaussian sputtering model, and reversely map these intermediate images back to the original images. In order to improve the rendering efficiency, the present disclosure proposes a pre-filtering and allocation module to optimize the processing process and avoid processing all three-dimensional Gaussian points each time rendering.
[0125] The camera-based optimization module is used to obtain the original image of the inverse mapping. Due to the loss in the mapping process, the generated original image may be rough. Considering that the mapping loss is closely related to the camera model, the original image needs to be optimized based on the camera model to reduce image distortion and improve the final effect.
[0126] Based on the same technical concept, the present disclosure provides a cross-device rendering apparatus. Figure 3 Schematic diagram of a cross-device rendering apparatus shown in an embodiment of the present disclosure. Figure 3 As shown, the device comprises: An acquisition module 310 is used to acquire original images from different camera devices; A mapping module 320, configured to convert the original image into an intermediate image according to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image, wherein the intermediate image carries a corresponding intermediate pose; A construction module 330, configured to construct a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; The determination module 340 is used to render the target intermediate image using the three-dimensional Gaussian sputtering model according to the rendering direction of the target intermediate image, and inversely map the target intermediate image to the inverse mapping area corresponding to the target original image to obtain the target original image.
[0127] The present disclosure also provides an electronic device, referring to Figure 4 , Figure 4 is a schematic diagram of an electronic device shown in an embodiment of the present disclosure. Figure 4 As shown, the electronic device 400 includes: a memory 410 and a processor 420. The memory 410 and the processor 420 are connected via a bus communication. A computer program is stored in the memory 410. The computer program can be run on the processor 420 to implement the steps in the cross-device rendering method disclosed in the embodiment of the present disclosure.
[0128] The embodiment of the present disclosure further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the cross-device rendering method disclosed in the embodiment of the present disclosure are implemented.
[0129] The embodiments of the present disclosure also provide a computer program product, including a computer program. When the computer program is executed by a processor, the steps in the cross-device rendering method disclosed in the embodiments of the present disclosure are implemented.
[0130] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.
[0131] It should be understood by those skilled in the art that the embodiments of the present disclosure may be provided as methods, devices or computer program products. Therefore, the embodiments of the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present disclosure may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0132] The embodiments of the present disclosure are described with reference to the flowcharts and / or block diagrams of the methods, apparatuses, electronic devices, and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0133] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0134] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device so that a series of operating steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable terminal device to implement the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0135] Although some embodiments of the present disclosure have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the present disclosure.
[0136] The above is a detailed introduction to a cross-device rendering method, apparatus, device and medium provided by the present disclosure. Specific examples are used in this article to illustrate the principles and implementation methods of the present disclosure. The description of the above embodiments is only used to help understand the method of the present disclosure and its core idea. At the same time, for those skilled in the art, according to the ideas of the present disclosure, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present disclosure.
Claims
1. A cross-device rendering method, characterized in that: include: Get raw images from different camera devices; According to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image, the original image is converted into an intermediate image, wherein the intermediate image carries a corresponding intermediate pose; Constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; According to the rendering direction of the target intermediate image, the target intermediate image is rendered using the three-dimensional Gaussian sputtering model, and the target intermediate image is inversely mapped to the inverse mapping area corresponding to the target original image to obtain the target original image.
2. The method according to claim 1, characterized in that Also includes: Initializing a learnable camera embedding by the number of raw camera models; the number of the raw camera models is consistent with the number of categories of the camera devices; Splicing the learnable camera embedding and the target original image along the image channel direction to obtain a spliced image; Inputting the stitched image into an optimization network to output a pixel-level target transformation and a restored image, wherein the optimization network is used to optimize the target original image; An optimized target original image is obtained through the target transformation, the restored image and the target original image.
3. The method according to claim 1, characterized in that Also includes: Filtering out redundant three-dimensional Gaussians in the three-dimensional Gaussian sputtering model by using the inverse mapping area and the rendering direction to obtain a filtered three-dimensional Gaussian sputtering model, wherein the redundant three-dimensional Gaussians are three-dimensional Gaussians that do not participate in generating the target original image; Obtaining a target intermediate image by rendering in the rendering direction using the three-dimensional Gaussian sputtering model includes: Based on the filtered three-dimensional Gaussian sputtering model, assigning effective three-dimensional Gaussians to corresponding rendering directions to obtain three-dimensional Gaussians corresponding to each rendering direction, wherein the effective three-dimensional Gaussians are three-dimensional Gaussians that participate in rendering the intermediate image in the rendering direction; The target intermediate image is obtained by rendering the three-dimensional Gaussian corresponding to each rendering direction using the three-dimensional Gaussian sputtering model.
4. The method according to claim 1, characterized in that: According to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image, converting the original image into an intermediate image includes: Determining, according to the camera intrinsic parameters of the intermediate camera model, a latitude rotation and a longitude rotation of the intermediate image relative to an original pose of the original image; Determine a rotation matrix of the intermediate image relative to the original image through the latitude rotation and the longitude rotation; Determining an intermediate pose of the intermediate image by using the rotation matrix and the original pose; The original image is converted into the intermediate image through the rotation matrix, the back projection function corresponding to the original camera model, and the projection function of the intermediate camera model.
5. The method according to claim 4, characterized in that The intermediate image includes a reference intermediate image with texture mapping and an intermediate image to be optimized without texture mapping, and further includes: Determining, according to the intermediate camera model, an adaptive mapping direction of the intermediate image to be optimized relative to the reference intermediate image, and image transformation parameters corresponding to each adaptive mapping direction; Determining a texture region of the reference intermediate image, and determining an intermediate pose corresponding to the reference intermediate image; Determining adaptive mapping images in various adaptive mapping directions according to the texture area of the reference intermediate image and the image transformation parameters; Replacing the intermediate image to be optimized with the adaptive mapping image to obtain an updated intermediate image to be optimized; Determining an updated intermediate pose corresponding to the intermediate image to be optimized according to the intermediate pose corresponding to the reference intermediate image and the adaptive mapping direction; According to the intermediate image and the corresponding intermediate posture, a three-dimensional Gaussian sputtering model is constructed, including: A three-dimensional Gaussian sputtering model is constructed according to the reference intermediate image and its corresponding intermediate pose and the updated intermediate image to be optimized and its corresponding intermediate pose.
6. The method according to any one of claims 1 to 5, characterized in that: According to the intermediate image and the corresponding intermediate posture, a three-dimensional Gaussian sputtering model is constructed, including: Constructing a first initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; Acquire a first initial target intermediate image having the same position and posture as the intermediate image through the first initial three-dimensional Gaussian sputtering model; Determining a first loss using the intermediate image and the first initial target intermediate image; adjusting parameters of the first initial three-dimensional Gaussian sputtering model according to the first loss; According to the above steps, the parameters of the first initial three-dimensional Gaussian sputtering model are adjusted multiple times to obtain a three-dimensional Gaussian sputtering model.
7. The method according to claim 2, characterized in that According to the intermediate image and the corresponding intermediate posture, a three-dimensional Gaussian sputtering model is constructed, including: Constructing a second initial three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; Acquire a second initial target original image having the same posture as the original image through the second initial three-dimensional Gaussian sputtering model; Obtaining an initial restored image and an initial target transformation corresponding to the second initial target original image through an initial optimization network; Determining an optimized second initial target original image according to the initial restored image, the initial target transformation and the second initial target original image; Determining a distortion loss according to the initial restored image and the original image; Determining a second loss by using the optimized second initial target original image and the original image; According to the distortion loss and the second loss, respectively adjusting the parameters of the initial optimization network and the second initial three-dimensional Gaussian sputtering model; According to the above steps, the parameters of the initial optimized network and the second initial three-dimensional Gaussian sputtering model are adjusted multiple times to obtain the optimized network and the three-dimensional Gaussian sputtering model.
8. A cross-device rendering apparatus, characterized in that: include: Acquisition module, used to acquire raw images from different camera devices; A mapping module, configured to convert the original image into an intermediate image according to a mapping relationship between an original camera model and an intermediate camera model corresponding to the original image, wherein the intermediate image carries a corresponding intermediate posture; A construction module, used for constructing a three-dimensional Gaussian sputtering model according to the intermediate image and the corresponding intermediate pose; A determination module is used to render a target intermediate image using the three-dimensional Gaussian sputtering model according to a rendering direction of the target intermediate image, and inversely map the target intermediate image to an inverse mapping area corresponding to a target original image to obtain the target original image.
9. An electronic device, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the computer program is executed by the processor, the steps of the cross-device rendering method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the cross-device rendering method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Three-dimensional scene reconstruction method and electronic equipment
CN118365805A
Three-dimensional rendering method and device, storage medium and program product
CN119251368A
Image rendering method and device based on Gaussian splashing, equipment, storage medium and program product
CN119295638A
New view angle synthesis method and system based on generative adversarial strategy and Gaussian sputtering
CN119379548A
Three-dimensional map construction method and device, computer equipment and storage medium
CN119722974A