3D rendering model training method, electronic equipment, medium and computer product
By determining and processing target anchor points containing benchmark features and residual features in the 3D rendering model, generating and fusing views to determine model adjustment parameters, the problem that 3D Gaussian models cannot circumvent rendering losses is solved, achieving higher rendering quality and lower rendering losses.
Patent Information
- Application Number
- CN202510092386.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-05-30
AI Technical Summary
The 3D Gaussian sputtering model cannot effectively avoid rendering losses during the rendering process, making it difficult for the rendering quality to reach an ideal state.
By determining the target anchor point containing the reference features and residual features and inputting it into the preset initial 3D rendering model, generating the reference view and residual view, fusing the view to generate the target view, determining the model adjustment parameters based on the view, and adjusting the initial model to generate the target 3D rendering model.
Effectively avoid rendering losses, improve rendering quality, and reduce distortion generated during rendering.
Smart Images

Figure CN120070698A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method for training a 3D rendering model, an electronic device, a storage medium, and a computer program product. Background Art
[0002] With the continuous development of 3D rendering technology, the rendering quality of 3D rendering technology has become increasingly refined.
[0003] In related technologies, when technicians perform 3D rendering and storage work, they usually use a 3D Gaussian sputtering model to perform 3D rendering operations on anchor points, so as to arrange a series of Gaussian points with specific attributes that match the anchor points in three-dimensional space, thereby constructing an approximate expression of a complex scene, and then realizing the conversion of a virtual scene into a visual image.
[0004] However, the optimization strategies adopted by the 3D Gaussian sputtering model usually focus on adjusting the parameters of the Gaussian points themselves. In this way, this adjustment method cannot effectively avoid the rendering loss problem that occurs during the rendering process, and thus the final rendering quality is difficult to reach an ideal state. Summary of the Invention
[0005] The main purpose of this application is to provide a method for training a 3D rendering model, an electronic device, a storage medium, and a computer program product, aiming to solve the technical problem that the 3D Gaussian model in related technologies cannot avoid rendering loss.
[0006] To achieve the above purpose, this application proposes a method for training a 3D rendering model, including:
[0007] Determine a target anchor point to be rendered, where the target anchor point includes a reference feature and a residual feature;
[0008] Input the target anchor point into a preset initial 3D rendering model, and generate a reference view by the initial 3D rendering model learning and reasoning on the reference feature, and generate a residual view by the initial 3D rendering model learning and reasoning on the residual feature;
[0009] Fuse the reference view and the residual view to obtain a target view, and determine a model adjustment parameter according to the reference view, the residual view, the target view, and a preset standard view;
[0010] Adjust the initial 3D rendering model according to the model adjustment parameter to generate a target 3D rendering model.
[0011] In one embodiment, the step of generating a reference view by the initial 3D rendering model learning and reasoning on the reference feature includes:
[0012] Performing learning inference on the basis of the reference features through the initial 3D rendering model to determine the reference Gaussian point attribute information of the reference Gaussian point corresponding to the target anchor point;
[0013] Determining the coordinate position information included in the target anchor point, and determining the reference Gaussian point position information of the reference Gaussian point according to the coordinate position information;
[0014] Performing rendering according to the reference Gaussian point attribute information and the reference Gaussian point position information to generate a reference view.
[0015] In one embodiment, the step of generating a residual view by performing learning inference on the residual features through the initial 3D rendering model includes:
[0016] Performing learning inference on the basis of the residual features through the initial 3D rendering model to determine the residual Gaussian point attribute information of the residual Gaussian point corresponding to the target anchor point;
[0017] Performing rendering according to the residual Gaussian point attribute information and the reference Gaussian point position information to generate a residual view.
[0018] In one embodiment, the step of determining the model adjustment parameters according to the reference view, the residual view, the target view and a preset standard view includes:
[0019] Determining a first rendering distortion parameter according to the target view and the preset standard view, determining a second rendering distortion parameter according to the reference view and the standard view, and determining a third rendering distortion parameter according to the residual view, the standard view and the reference view;
[0020] Determining a target loss function based on the first rendering distortion parameter, the second rendering distortion parameter and the third rendering distortion parameter;
[0021] Determining the model adjustment parameters corresponding to the initial 3D rendering model according to the target loss function.
[0022] In one embodiment, after the step of determining the target anchor point to be rendered, the method further includes:
[0023] Extracting the coordinate position information included in the target anchor point;
[0024] Performing encoding processing on the coordinate position information to generate a target bitstream, so as to compress the coordinate position information through the target bitstream.
[0025] In one embodiment, the step of performing encoding processing on the coordinate position information to generate a target bitstream includes:
[0026] Determine the anchor point coordinates with the smallest value included in the coordinate position information;
[0027] Determine the quantization step based on the anchor point coordinates, and divide the space where the target anchor point is located based on the quantization step to determine the octree structure position information corresponding to the coordinate position information;
[0028] Perform encoding processing on the octree structure position information to generate a target bitstream.
[0029] In one embodiment, the step of performing encoding processing on the octree structure position information to generate a target bitstream includes:
[0030] Determine multiple ancestor nodes that match the target anchor point, and use the multiple ancestor nodes as context information;
[0031] Insert the octree structure position information into the context information to obtain a low-level feature through conversion, and integrate the low-level feature to determine the occupancy code probability distribution corresponding to the octree structure position information;
[0032] Extract the occupancy code included in the octree structure position information, and perform entropy encoding processing on the occupancy code based on the occupancy code probability distribution to generate a target bitstream.
[0033] In addition, to achieve the above object, the present application further provides an electronic device, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the 3D rendering model training method as described above.
[0034] In addition, to achieve the above object, the present application further provides a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, it implements the steps of the 3D rendering model training method as described above.
[0035] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps of the 3D rendering model training method as described above.
[0036] The training method of the 3D rendering model provided by the embodiment of the present application determines a target anchor point to be rendered, where the target anchor point includes a reference feature and a residual feature; inputs the target anchor point into a preset initial 3D rendering model, and the initial 3D rendering model performs learning and inference on the reference feature to generate a reference view, and performs learning and inference on the residual feature to generate a residual view; fuses the reference view and the residual view to obtain a target view, and determines a model adjustment parameter according to the reference view, the residual view, the target view, and a preset standard view; adjusts the initial 3D rendering model according to the model adjustment parameter to generate a target 3D rendering model.
[0037] In this way, the present application solves the technical problem that the 3D Gaussian model in the related art cannot avoid rendering loss. That is, the present application collects a target anchor point including a reference feature and a residual feature in a three-dimensional space that needs to be rendered, and inputs the target anchor point into a preset initial 3D rendering model. The initial 3D rendering model processes the target anchor point to generate a reference view and a residual view that match the target anchor point based on the reference feature and the residual feature, and uses the residual view to enhance the reference view to obtain a target view. Thus, a model adjustment parameter for adjusting the initial 3D rendering model is determined according to the reference view, the residual view, the target view, and a preset standard view. Finally, the initial 3D rendering model is adjusted according to the model adjustment parameter, so as to achieve the technical effect of obtaining a 3D rendering model that can avoid rendering loss, and further reduce the rendering loss in the target view obtained through the 3D rendering model. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0040] Figure 1 It is a schematic flowchart provided by Embodiment 1 of the training method of the 3D rendering model of the present application;
[0041] Figure 2 It is a detailed flowchart of the training method of the 3D rendering model of the present application;
[0042] Figure 3Schematic diagram of the occupancy code probability distribution prediction process involved in the second embodiment of the training method of the 3D rendering model of the present application;
[0043] Figure 4 Schematic diagram of the device structure of the hardware operating environment involved in the training method of the 3D rendering model in the embodiment of the present application.
[0044] The implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0045] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0046] In order to better understand the technical solutions of the present application, the following will be described in detail with reference to the accompanying drawings of the specification and specific implementation manners.
[0047] In this embodiment, for the convenience of description, the following takes an electronic device internally configured with a model training module, or a mobile terminal, a data storage control terminal, a PC and other terminals connected to an electronic control unit supporting the electronic device as an execution subject for elaboration.
[0048] Based on the above-mentioned electronic device, the overall concept of the training method of the 3D rendering model of the present application is proposed here.
[0049] With the continuous development of 3D rendering technology, the rendering quality of 3D rendering technology is becoming more and more refined. In the related art, when technicians perform 3D rendering and storage work, they usually use a 3D Gaussian sputtering model to perform 3D rendering operations on the anchor points, so as to arrange a series of Gaussian points with specific attributes matching the anchor points in the three-dimensional space, thereby constructing an approximate expression of a complex scene, and then realizing the conversion of the virtual scene into a visual image. However, the optimization strategies adopted by the 3D Gaussian sputtering model usually focus on adjusting the parameters of the Gaussian points themselves. In this way, this adjustment method cannot effectively avoid the rendering loss problem that occurs during the rendering process, and thus the final rendering quality is difficult to reach an ideal state.
[0050] In view of the above phenomenon, the present application provides a training method for a 3D rendering model. The training method for the 3D rendering model includes: determining a target anchor point to be rendered, where the target anchor point includes a reference feature and a residual feature; inputting the target anchor point into a preset initial 3D rendering model, and through the initial 3D rendering model, learning and inferring the reference feature to generate a reference view, and through the initial 3D rendering model, learning and inferring the residual feature to generate a residual view; fusing the reference view and the residual view to obtain a target view, and determining a model adjustment parameter according to the reference view, the residual view, the target view, and a preset standard view; adjusting the initial 3D rendering model according to the model adjustment parameter to generate a target 3D rendering model.
[0051] In this way, the present application solves the technical problem in the related art that the 3D Gaussian model cannot avoid rendering loss. That is, the present application collects a target anchor point including a reference feature and a residual feature in a three-dimensional space that needs to be rendered, and inputs the target anchor point into a preset initial 3D rendering model. The initial 3D rendering model processes the target anchor point to generate a reference view and a residual view that match the target anchor point based on the reference feature and the residual feature, and uses the residual view to enhance the reference view to obtain a target view. Then, according to the reference view, the residual view, the target view, and a preset standard view, a model adjustment parameter for adjusting the initial 3D rendering model is determined. Finally, the initial 3D rendering model is adjusted according to the model adjustment parameter, thereby achieving the technical effect of obtaining a 3D rendering model that can avoid rendering loss, and further reducing the rendering loss in the target view obtained through the 3D rendering model.
[0052] Based on the overall concept of the training method for the 3D rendering model of the present application, an embodiment of the present application provides a training method for a 3D rendering model. Refer to Figure 1 , Figure 1 which is a schematic flowchart of the first embodiment of the training method for the 3D rendering model of the present application. In this embodiment, the training method for the 3D rendering model includes steps S10 to S40:
[0053] Step S10: Determine a target anchor point to be rendered, where the target anchor point includes a reference feature and a residual feature;
[0054] Step S20: Input the target anchor point into a preset initial 3D rendering model, and through the initial 3D rendering model, learn and infer the reference feature to generate a reference view, and through the initial 3D rendering model, learn and infer the residual feature to generate a residual view;
[0055] Step S30: Fuse the reference view and the residual view to obtain a target view, and determine model adjustment parameters according to the reference view, the residual view, the target view, and a preset standard view;
[0056] Step S40: Adjust the initial 3D rendering model according to the model adjustment parameters to generate a target 3D rendering model.
[0057] It should be noted that the target anchor contains a reference feature f, a residual feature e, attribute parameters, and coordinate position information x, where the attribute parameters include an offset parameter o and a scaling parameter s.
[0058] In this embodiment, the electronic device first determines a 3D scene to be rendered, and determines target anchors each having a residual feature and a reference feature included in the 3D scene. Then, the electronic device inputs the target anchors into a preset model training module, and the model training module inputs the target anchors into a preset initial 3D rendering model. The initial 3D rendering model extracts the reference feature included in the target anchor and performs learning and reasoning based on the reference feature to generate a reference view. At the same time, the initial 3D rendering model extracts the residual feature included in the target anchor and performs learning and reasoning based on the residual feature to generate a residual view. Then, the model training module fuses the residual view and the reference view to enhance the reference view through the residual view, thereby obtaining a target view that finally completes the rendering. At the same time, the model training module reads the storage module configured by the electronic device to obtain a preset standard view without distortion. The model training module then determines model adjustment parameters according to the target view, the reference view, the residual view, and the standard view. Finally, the model training module adjusts the initial 3D rendering model according to the model adjustment parameters to obtain a target 3D rendering model.
[0059] Exemplarily, for example, please refer to Figure 2 , Figure 2 is a detailed flowchart of the training method of the 3D rendering model of the present application. As Figure 2 shown, the electronic device first obtains a 3D scene to be rendered, and determines a plurality of target anchors anchor including a reference feature f, a residual feature e, attribute parameters, and coordinate position information x included in the 3D scene. Then, the electronic device inputs the plurality of target anchors anchor into the model training module, and the model training module calls a preset initial 3D rendering model to perform learning, reasoning, and Gaussian sputtering processing based on the reference feature f included in the target anchor anchor to render a reference view I corresponding to the target anchor anchor baisc, meanwhile, the model training module invokes the 3D rendering model to perform learning inference and Gaussian sputtering processing based on the residual feature e included in the target anchor anchor, so as to render the residual view I corresponding to the target anchor anchor res , after that, the model training module uses the reference view I baisc and the residual view I res are fused to obtain the target view I final , meanwhile, the model training module reads the above storage module to obtain the preset label image gt without distortion, and uses the reference view I baisc , the residual view I res , the target view I final and the label image gt are compared to determine the model adjustment parameters for adjusting the initial 3D rendering model. Finally, the model training module adjusts the initial 3D rendering model according to the model adjustment parameters, so as to obtain a reference view I rendered based on the reference feature f by using the residual feature e included in the target anchor anchor baisc is enhanced to reduce the distortion generated during the rendering process of the target 3D rendering model.
[0060] It should be noted that the initial 3D rendering model is a 3D-GS model. In addition, the specific process of adjusting the 3D rendering model according to the model adjustment instruction is a prior art, so it will not be elaborated here.
[0061] In this way, the present application solves the technical problem that the 3D Gaussian model in the related art cannot avoid rendering losses, that is, the present application collects target anchors including reference features and residual features in the three-dimensional space to be rendered, and inputs the target anchors into the preset initial 3D rendering model. The initial 3D rendering model processes the target anchors to generate a reference view and a residual view that match the target anchors based on the reference features and residual features, and uses the residual view to enhance the reference view to obtain the target view. Then, according to the reference view, the residual view, the target view and the preset standard view, the model adjustment parameters for adjusting the initial 3D rendering model are determined. Finally, the initial 3D rendering model is adjusted according to the model adjustment parameters, so as to achieve the technical effect of a 3D rendering model that can avoid rendering losses, and further reduce the rendering losses in the target view obtained through the 3D rendering model.
[0062] In a feasible implementation manner, the step of "generating a reference view by performing learning inference on the reference feature through the initial 3D rendering model" in the above step S20 may specifically include steps S201 to S203:
[0063] Step S201: performing learning and reasoning based on the reference features through the initial 3D rendering model to determine reference Gaussian point attribute information of the reference Gaussian point corresponding to the target anchor point;
[0064] Step S202: determining the coordinate position information included in the target anchor point, and determining the reference Gauss point position information of the reference Gauss point according to the coordinate position information;
[0065] Step S203: Rendering is performed according to the reference Gauss point attribute information and the reference Gauss point position information to generate a reference view.
[0066] For example, after the model training module inputs the target anchor point anchor into the initial 3D rendering model, the initial 3D rendering model first extracts the reference feature f contained in the target anchor point anchor. Due to the reference feature f in the target anchor point anchor, the reference Gaussian point g that matches the target anchor point anchor is extracted. b The corresponding benchmark Gaussian point attribute information is associated. Therefore, the initial 3D rendering model uses MLP learning and reasoning to learn the benchmark feature f, and the benchmark Gaussian point g b The corresponding included color c b , Opacitya b , scaling information b and rotation information r b The initial 3D rendering model then extracts the coordinate position information x contained in the target anchor point anchor and the offset parameter o and scaling parameter s contained in the target anchor point anchor. The initial 3D rendering model then calculates the reference Gaussian point g according to the offset parameter o, scaling parameter s, and coordinate position information x. b Finally, the initial 3D rendering model performs Gaussian sputtering on the attribute information of the benchmark Gaussian points and the spatial position μ of the benchmark Gaussian points to render the benchmark view I baisc .
[0067] It should be noted that the specific process of MLP learning and reasoning is an existing technology, so it will not be repeated here.
[0068] In a feasible implementation manner, the step of "generating a residual view by learning and reasoning the residual features through the initial 3D rendering model" in the above step S20 includes steps S204 to S205:
[0069] Step S204: performing learning and reasoning based on the residual features through the initial 3D rendering model to determine residual Gaussian point attribute information of the residual Gaussian point corresponding to the target anchor point;
[0070] Step S205: Render according to the residual Gaussian point attribute information and the reference Gaussian point position information to generate a residual view.
[0071] Exemplarily, for example, after the model training module inputs the target anchor into the initial 3D rendering model, the initial 3D rendering model can also first extract the residual feature e contained in the target anchor. The initial 3D rendering model performs MLP learning inference through the residual feature e to determine the residual Gaussian point g that matches the target anchor r Correspondingly, it includes the opacity a r 、scaling information t r 、rotation information r r 、color c r in the residual Gaussian point attribute information. After that, the initial 3D rendering model performs Gaussian sputtering processing on the residual Gaussian point attribute information and the above-mentioned reference Gaussian point spatial position μ to render and obtain the residual view I res .
[0072] It can be understood that since the positions of the residual Gaussian point g r and the above-mentioned reference Gaussian point g b are the same, therefore, the initial 3D rendering model does not need to re-infer the position information of the residual Gaussian point g r .
[0073] In a feasible implementation manner, the step of "determining the model adjustment parameters according to the reference view, the residual view, the target view, and a preset standard view" in the above step S30 may specifically include steps S301 to S303:
[0074] Step S301: Determine a first rendering distortion parameter according to the target view and the preset standard view, and determine a second rendering distortion parameter according to the reference view and the standard view, and determine a third rendering distortion parameter according to the residual view, the standard view, and the reference view;
[0075] Step S302: Determine a target loss function based on the first rendering distortion parameter, the second rendering distortion parameter, and the third rendering distortion parameter;
[0076] Step S303: Determine the model adjustment parameters corresponding to the initial 3D rendering model according to the target loss function.
[0077] Exemplarily, for example, the model training module fuses the reference view I baisc and the residual view I res to obtain the target view I finalAfter that, the above storage module can be read first to obtain the standard view gt without distortion, and then the model training module will use the target view I final and the standard view gt for comparison to determine the target view I final and the first rendering distortion parameter L 1 (I final , gt). At the same time, the model training module will compare the reference view I baisc with the standard view gt to determine the second rendering distortion parameter L baisc (I 1 , gt). At the same time, the model training module will compare the reference view I basic , the standard view gt, and the residual view I baisc simultaneously to determine the third rendering distortion parameter L res corresponding to the residual between the residual view I res and the reference view I baisc , the standard view gt, i.e., L 1 (I res , (gt - I basic )). After that, the model training module determines the reference setting weight λ baisc corresponding to the reference view I baisc , and determines the residual setting weight λ res corresponding to the residual view I res . The model training module calculates the first rendering distortion parameter L baisc (I res , gt), the second rendering distortion parameter L 1 (I final , gt), and the third rendering distortion parameter L 1 (I basic , (gt - I 1 )) according to the reference setting weight λ res and the residual setting weight λ basic to determine the loss function L D for the constrained rendering distortion part in the initial 3D rendering model:
[0078] L D = L 1 (I final , gt) + λ basic L 1 (I basic , gt) + λ res L 1 (I res , (gt - I basic ));
[0079] Finally, the model training module determines model adjustment parameters for adjusting the initial 3D rendering model based on the loss function L D to determine the model adjustment parameters for adjusting the initial 3D rendering model.
[0080] In this embodiment, the electronic device first determines the 3D scene to be rendered and determines the target anchor points included in the 3D scene, each having residual features and reference features. Then, the electronic device inputs the target anchor points into a preset model training module, and the model training module inputs the target anchor points into a preset initial 3D rendering model. The initial 3D rendering model extracts the reference features included in the target anchor points and performs learning and inference based on the reference features to generate a reference view. At the same time, the initial 3D rendering model extracts the residual features included in the target anchor points and performs learning and inference based on the residual features to generate a residual view. Then, the model training module fuses the residual view and the reference view to enhance the reference view through the residual view, thereby obtaining the target view that finally completes the rendering. At the same time, the model training module reads the storage module configured by the electronic device to obtain a preset standard view without distortion. The model training module then determines the model adjustment parameters according to the target view, the reference view, the residual view, and the standard view. Finally, the model training module adjusts the initial 3D rendering model according to the model adjustment parameters to obtain the target 3D rendering model.
[0081] In this way, the present application solves the technical problem that the 3D Gaussian model in the related art cannot avoid rendering losses. That is, the present application collects target anchor points including reference features and residual features in the three-dimensional space to be rendered, and inputs the target anchor points into a preset initial 3D rendering model. The initial 3D rendering model processes the target anchor points to generate a reference view and a residual view that match the target anchor points based on the reference features and the residual features, and uses the residual view to enhance the reference view to obtain the target view. Thus, the model adjustment parameters for adjusting the initial 3D rendering model are determined according to the reference view, the residual view, the target view, and the preset standard view. Finally, the initial 3D rendering model is adjusted according to the model adjustment parameters, thereby achieving the technical effect of obtaining a 3D rendering model that can avoid rendering losses, and further reducing the rendering losses in the target view obtained through the 3D rendering model.
[0082] Based on the first embodiment of the present application, the second embodiment of the present application is proposed here. In the second embodiment of the present application, the same or similar content as in the above embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, after the above step S10, the training method of the 3D rendering model of the present application may further include steps A10 to A20:
[0083] Step A10: Extract the coordinate position information included in the target anchor points;
[0084] Step A20: Perform encoding processing on the coordinate position information to generate a target bitstream, so as to compress the coordinate position information through the target bitstream.
[0085] In this embodiment, after the electronic device obtains the target anchor point, if it is necessary to compress the position information of the target anchor point, it first extracts the coordinate position information included in the target anchor point. Then, the electronic device converts the coordinate position information to encode it into octree structure position information. The electronic device then performs encoding processing on the octree structure position information to generate a target bitstream, thereby compressing the coordinate position information of the target anchor point through the target bitstream.
[0086] Exemplarily, for example, please refer to Figure 2 and Figure 3 , where Figure 3 is a schematic diagram of the occupancy code probability distribution prediction process involved in the second embodiment of the 3D rendering model training method of this application. As Figure 2 shown, after the electronic device obtains the target anchor point anchor, if it is necessary to compress and store the coordinate position information x included in the target anchor point anchor, the electronic device first reads the coordinate position information x included in the target anchor point anchor. Then, as Figure 3 shown, the electronic device converts the coordinate position information x to convert the coordinate position information x into octree structure position information, so as to calculate the occupancy code probability distribution Q(m i ) corresponding to the target anchor point anchor through the occupancy code m, depth l, and octant position h included in each tree node of the octree structure position information. The electronic device then performs entropy encoding on the occupancy code m based on the obtained occupancy code probability distribution Q(m i ) to generate a target bitstream, thereby completing the storage of the coordinate position information x through the target bitstream.
[0087] It should be noted that in the related art, when compressing the target anchor point, usually only the attribute information included in the target anchor point can be compressed, and the complex coordinate position information cannot be compressed. Thus, the uncompressed coordinate position information will consume more storage space. Thus, in this application, by converting the coordinate position information into an octree structure to constrain the discrete coordinate distribution, and at the same time, using the ancestor node matching the target anchor point as context information to calculate the occupancy code probability distribution matching the target anchor point, and then encoding and compressing the occupancy code based on the obtained occupancy code probability distribution, the difficult-to-compress coordinate position information can be encoded and compressed, reducing the storage space consumed by the coordinate position information.
[0088] In a feasible implementation manner, the above step A20 may specifically include steps A201 to A203:
[0089] Step A201: Determine the anchor coordinate with the smallest value included in the coordinate position information;
[0090] Step A202: Determine the quantization step size based on the anchor coordinate, and divide the space where the target anchor is located based on the quantization step size to determine the octree structure position information corresponding to the coordinate position information;
[0091] Step A203: Perform encoding processing on the octree structure position information to generate a target bitstream.
[0092] Exemplarily, for example, after the electronic device extracts the coordinate position information x of the target anchor anchor, it first determines the anchor coordinate offset with the smallest value among all the coordinate values included in the coordinate position information x, and performs a difference calculation based on the coordinate position information x and the anchor coordinate offset to determine the quantization function corresponding to the difference result, thereby determining the quantization step size Q anchor , and the electronic device thus divides the space where the target anchor anchor is located recursively according to the quantization step size Q anchor , so as to determine the length corresponding to the smallest octant block of the quantization step size Q anchor :
[0093] x q = round((x - offset) / Q anchor ); where round represents the quantization function of rounding, and x q is the quantization result of the coordinate position information x;
[0094] The electronic device thus determines the occupancy code m, depth l, and octant block position h corresponding to the target anchor anchor according to the quantization result x of the coordinate position information x q , combines the occupancy code m, depth l, and octant block position h to obtain the octree structure position information. Finally, the electronic device performs entropy encoding on the octree structure position information to generate a target bitstream, thereby completing the storage of the coordinate position information x with the target bitstream.
[0095] In a feasible implementation manner, the above step A203 may specifically include steps A2031 to A2033:
[0096] Step A2031: Determine a plurality of ancestor nodes that match the target anchor, and use the plurality of ancestor nodes as context information;
[0097] Step A2032: Insert the octree structure position information into the context information to obtain a low-level feature after conversion, and integrate the low-level feature to determine the occupancy code probability distribution corresponding to the octree structure position information;
[0098] Step A2033: Extract the occupancy codes included in the octree structure position information, and perform entropy encoding processing on the occupancy codes based on the occupancy code probability distribution to generate a target bitstream.
[0099] Exemplarily, for example, after the electronic device determines the octree structure position information corresponding to the target anchor, it first determines k ancestor nodes n that have been encoded and match the target anchor ancestor , and uses the multiple ancestor nodes n ancestor as context information {mΦ m , lΦ l , hΦ h}. After that, the electronic device embeds the depth l and the octant position h of the target anchor into the context information {mΦ m , lΦ l , hΦ h} to convert the difficult-to-compress coordinate information into learnable low-level features. Then, after integrating the features, a lightweight encoding network as shown in Figure 3 is used for further feature extraction to predict the occupancy code probability distribution Q(m i ) corresponding to the target anchor:
[0100]
[0101] Finally, the electronic device inputs the occupancy code probability distribution Q(m i ) and the above-mentioned occupancy code m into the encoder, so that the encoder performs entropy encoding on the occupancy code m in the node information based on the occupancy code distribution probability Q(m i ) to generate a target bitstream.
[0102] In this way, this application constrains the discrete coordinate distribution by converting the coordinate position information into an octree structure. At the same time, by using the ancestor nodes matching the target anchor as context information to infer the occupancy code probability distribution matching the target anchor, and then encoding and compressing the occupancy code based on the obtained occupancy code probability distribution, it is possible to encode and compress the difficult-to-compress coordinate position information and reduce the storage space consumed by the coordinate position information.
[0103] This application provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the training method of the 3D rendering model in the first embodiment above.
[0104] Next, refer toFigure 4 which shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, an electronic device internally configured with a model training module, or a mobile terminal, a data storage control terminal, a PC, etc. connected to an electronic control unit supporting the electronic device. Figure 4 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0105] As Figure 4 shown, the electronic device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the electronic device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows an electronic device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or had alternatively.
[0106] Particularly, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the method of the embodiments disclosed in the present application are performed.
[0107] The electronic device provided by this application adopts the training method of the 3D rendering model in the above embodiments, which can solve the technical problem that the 3D Gaussian model in the related art cannot avoid rendering losses. Compared with the prior art, the beneficial effects of the electronic device provided by this application are the same as those of the training method of the 3D rendering model provided by the above embodiments, and other technical features in this electronic device are the same as those disclosed in the method of the previous embodiment, which will not be elaborated here.
[0108] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0109] As mentioned above, only the specific implementation manners of this application are described, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claimed rights.
[0110] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the training method of the 3D rendering model in the above embodiments.
[0111] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems or devices, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0112] The above computer-readable storage medium may be included in an electronic device; or may exist independently without being assembled into the electronic device.
[0113] The above computer-readable storage medium stores one or more programs, which when executed by the electronic device, cause the electronic device to: determine a target anchor point to be rendered, where the target anchor point includes a reference feature and a residual feature; input the target anchor point into a preset initial 3D rendering model, and generate a reference view by learning and reasoning the reference feature through the initial 3D rendering model, and generate a residual view by learning and reasoning the residual feature through the initial 3D rendering model; fuse the reference view and the residual view to obtain a target view, and determine model adjustment parameters according to the reference view, the residual view, the target view, and a preset standard view; and adjust the initial 3D rendering model according to the model adjustment parameters to generate a target 3D rendering model.
[0114] Computer program code for performing the operations of the present application may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (for example, by connecting through an Internet service provider using the Internet).
[0115] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0116] The modules described in the embodiments of the present application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.
[0117] The readable storage medium provided by the present application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned 3D rendering model training method, which can solve the technical problem that the 3D Gaussian model in the related art cannot avoid rendering loss. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as those of the 3D rendering model training method provided by the above embodiments, and will not be elaborated here.
[0118] The present application also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the steps of the 3D rendering model training method as described above.
[0119] The computer program product provided by the present application can solve the technical problem that the 3D Gaussian model in the related art cannot avoid rendering loss. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the 3D rendering model training method provided by the above embodiments, and will not be elaborated here.
[0120] The above are only some embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A training method for a 3D rendering model, characterized in that: The training method of the 3D rendering model includes: Determine a target anchor point to be rendered, wherein the target anchor point includes a reference feature and a residual feature; Inputting the target anchor point into a preset initial 3D rendering model, performing learning and reasoning on the reference feature through the initial 3D rendering model to generate a reference view, and performing learning and reasoning on the residual feature through the initial 3D rendering model to generate a residual view; fusing the reference view and the residual view to obtain a target view, and determining a model adjustment parameter according to the reference view, the residual view, the target view and a preset standard view; The initial 3D rendering model is adjusted according to the model adjustment parameters to generate a target 3D rendering model.
2. The training method for a 3D rendering model according to claim 1, characterized in that: The step of generating a reference view by learning and reasoning the reference feature through the initial 3D rendering model comprises: Performing learning and reasoning based on the reference features by the initial 3D rendering model to determine reference Gauss point attribute information of the reference Gauss point corresponding to the target anchor point; Determine the coordinate position information included in the target anchor point, and determine the reference Gauss point position information of the reference Gauss point according to the coordinate position information; Rendering is performed according to the reference Gauss point attribute information and the reference Gauss point position information to generate a reference view.
3. The training method of a 3D rendering model as claimed in claim 2, characterized in that: The step of generating a residual view by learning and reasoning the residual features through the initial 3D rendering model comprises: Performing learning and reasoning based on the residual features by the initial 3D rendering model to determine residual Gaussian point attribute information of the residual Gaussian point corresponding to the target anchor point; Rendering is performed according to the residual Gaussian point attribute information and the reference Gaussian point position information to generate a residual view.
4. The training method for a 3D rendering model according to claim 3, wherein: The step of determining the model adjustment parameters according to the reference view, the residual view, the target view and the preset standard view includes: Determining a first rendering distortion parameter according to the target view and a preset standard view, and determining a second rendering distortion parameter according to the reference view and the standard view, and determining a third rendering distortion parameter according to the residual view, the standard view, and the reference view; Determining a target loss function based on the first rendering distortion parameter, the second rendering distortion parameter, and the third rendering distortion parameter; The model adjustment parameters corresponding to the initial 3D rendering model are determined according to the target loss function.
5. The training method of a 3D rendering model according to claim 1, characterized in that: After the step of determining the target anchor point to be rendered, the method further includes: Extracting coordinate position information contained in the target anchor point; The coordinate position information is encoded to generate a target code stream, so as to compress the coordinate position information through the target code stream.
6. The 3D rendering model training method according to claim 5, characterized in that: The step of encoding the coordinate position information to generate a target code stream includes: Determine the anchor point coordinates with the smallest value contained in the coordinate position information; Determine a quantization step size according to the anchor point coordinates, and divide the space where the target anchor point is located based on the quantization step size to determine the octree structure position information corresponding to the coordinate position information; The octree structure position information is encoded to generate a target bit stream.
7. The training method of a 3D rendering model according to claim 6, characterized in that: The step of encoding the octree structure position information to generate a target bitstream comprises: Determine multiple ancestor nodes that match the target anchor point, and use the multiple ancestor nodes as context information; Inserting the octree structure position information into the context information to convert into low-level features, and integrating the low-level features to determine the occupancy code probability distribution corresponding to the octree structure position information; The occupied code contained in the octree structure position information is extracted, and the occupied code is entropy encoded based on the occupied code probability distribution to generate a target code stream.
8. An electronic device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the training method for a 3D rendering model as described in any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the training method for a 3D rendering model as described in any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the training method for a 3D rendering model according to any one of claims 1 to 7 are implemented.