Depth supervision Gaussian splash model determination method and device, equipment and storage medium

By introducing a deep supervision mechanism in the Gaussian splattering model, using depth images and segmentation models to construct a depth loss function, the shortcomings of the existing models in reconstructing depth details are solved, and the better detail reconstruction effect is achieved.

CN120032053APending Publication Date: 2025-05-23Hefei Xinming Intelligent Technology Co., Ltd.
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510110423.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

The existing Gaussian splatter model is difficult to effectively reflect the depth details in the image when reconstructing 3D scenes, resulting in poor detail reconstruction results.

Method used

By introducing a deep supervision mechanism, the image and depth images are acquired, the image is segmented based on the segmentation model is divided, and the depth loss function is constructed, and the Gaussian splashing model is trained to improve the model's attention to depth details.

Benefits of technology

The Gaussian splatter model's attention to detail areas with orders of magnitude depth is significantly improved, thereby improving the detail effect of 3D scene reconstruction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032053A_ABST
    Figure CN120032053A_ABST
Patent Text Reader

Abstract

The invention provides a depth supervision Gaussian splash model determination method and device, equipment and a storage medium, and relates to the technical field of computers, and the method comprises the steps: obtaining an image of a reconstruction target and a depth image corresponding to the image, carrying out the segmentation of the image based on a segmentation model, obtaining the depth data of different objects in the image, training a Gaussian splash model based on the image and the depth image, obtaining a trained Gaussian splash model under the condition that a target loss function value constructed by a depth loss function, an L1 loss function and an SSIM loss function reaches a preset threshold value, and supervising the training process in the Gaussian splash model training process based on the depth data. The attention of the Gaussian splash model to the detail area with the small depth order of magnitude is improved, and then the reconstruction effect on details is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of the times, 3D scene applications are becoming more and more widespread. In related technologies, the Gaussian splashing method is usually used to complete 3D scene reconstruction. Gaussian splashing usually uses multi-view images as input and supervision information. However, it is difficult to reflect the depth in the image only using multi-view images as supervision information.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention

[0004] The present invention provides a method, device, equipment and storage medium for determining a deep supervised Gaussian splash model, which at least improves the Gaussian splash model's attention to detail areas with a smaller depth order, thereby improving the reconstruction effect on the details.

[0005] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.

[0006] According to one aspect of the present disclosure, a method for determining a deep supervised Gaussian splash model is provided, comprising:

[0007] Obtain an image of the reconstructed target and a depth image corresponding to the image;

[0008] Segment the image based on the segmentation model to obtain the depth data of different objects in the image;

[0009] Train the Gaussian splash model based on images and depth images;

[0010] When the target loss function value corresponding to the Gaussian splash model reaches the preset threshold, the trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0011] In one embodiment of the present disclosure, the method further includes:

[0012] Determine the depth mean and depth standard deviation of different objects in the image based on the depth data of different objects;

[0013] Determine the depth mean and depth standard deviation of different predicted objects in the predicted image based on the predicted depth data;

[0014] A depth loss function is constructed based on the depth mean and depth standard deviation of different predicted objects and the depth mean and depth standard deviation of different objects in the image.

[0015] In one embodiment of the present disclosure, the depth loss function includes:

[0016]

[0017] Among them, L depth is the depth loss function value, x is the predicted depth data of the object predicted by the Gaussian splash model, is the depth data of the object, n is the total number of objects, U(i) is the mean of the object depth data, and std(i) is the standard deviation of the object depth data.

[0018] In one embodiment of the present disclosure, segmenting an image based on a segmentation model to obtain depth data of different objects in the image includes:

[0019] Segment the image based on the segmentation model to obtain masks of different objects in the image;

[0020] Depth data of different objects are determined based on the masks of different objects.

[0021] In one embodiment of the present disclosure, a Gaussian splash model is trained based on an image and a depth image, including:

[0022] Initialize the point cloud of the image based on the structure-from-motion algorithm SFM;

[0023] The Gaussian splash model is trained based on the initialized point cloud.

[0024] In one embodiment of the present disclosure, the method further includes:

[0025] When the target loss function value does not reach a preset threshold, the parameters of the Gaussian splash model are adjusted based on the target loss function value.

[0026] In one embodiment of the present disclosure, the method further includes:

[0027] The target perspective is input into the trained Gaussian splash model to obtain the reconstructed target prediction image under the target perspective.

[0028] According to another aspect of the present disclosure, a device for determining a deep supervised Gaussian splash model is provided, comprising:

[0029] An acquisition module is used to acquire an image of a reconstructed target and a depth image corresponding to the image;

[0030] A segmentation module is used to segment the image based on the segmentation model to obtain the depth data of different objects in the image;

[0031] A training module for training the Gaussian splash model based on images and depth images;

[0032] The judgment module is used to obtain the trained Gaussian splash model when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold. The target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0033] In one embodiment of the present disclosure, the device further includes:

[0034] A first determination module, configured to determine depth mean values ​​and depth standard deviations of different objects in an image based on depth data of different objects;

[0035] A second determination module, configured to determine depth means and depth standard deviations of different predicted objects in the predicted image based on the predicted depth data;

[0036] A construction module is used to construct a depth loss function based on the depth mean and depth standard deviation of different predicted objects and the depth mean and depth standard deviation of different objects in the image.

[0037] In one embodiment of the present disclosure, the depth loss function includes:

[0038]

[0039] Among them, L depth is the depth loss function value, x is the predicted depth data of the object predicted by the Gaussian splash model, is the depth data of the object, n is the total number of objects, U(i) is the mean of the object depth data, and std(i) is the standard deviation of the object depth data.

[0040] In one embodiment of the present disclosure, the segmentation module includes:

[0041] A segmentation unit, used to segment the image based on the segmentation model to obtain masks of different objects in the image;

[0042] The determination unit is used to determine the depth data of different objects based on the masks of different objects.

[0043] In one embodiment of the present disclosure, the training module includes:

[0044] An initialization unit, used to initialize the point cloud of the image based on the structure from motion algorithm SFM;

[0045] The training unit is used to train the Gaussian splash model based on the initialized point cloud.

[0046] In one embodiment of the present disclosure, the device further includes:

[0047] The adjustment module is used to adjust the parameters of the Gaussian splash model based on the target loss function value when the target loss function value does not reach a preset threshold.

[0048] In one embodiment of the present disclosure, the device further includes:

[0049] The input module is used to input the target perspective into the trained Gaussian splash model to obtain a reconstructed target prediction image under the target perspective.

[0050] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned deep supervised Gaussian splash model determination methods by executing the executable instructions.

[0051] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned deep supervised Gaussian splash model determination methods is implemented.

[0052] According to another aspect of the present disclosure, a computer program product is provided, which includes a computer program or computer instructions, and the computer program or computer instructions are loaded and executed by a processor to enable a computer to implement any of the above-mentioned deep supervised Gaussian splash model determination methods.

[0053] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splatter model is trained based on the image and the depth image, and a trained Gaussian splatter model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splatter model, the training process is supervised based on the depth data, thereby improving the Gaussian splatter model's attention to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0054] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.

[0056] Figure 1 A schematic diagram of the structure of a deep supervised Gaussian splash model determination system in an embodiment of the present disclosure is shown.

[0057] Figure 2 A flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0058] Figure 3 A flow chart of another method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0059] Figure 4 A flow chart of another method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0060] Figure 5 A flow chart of another method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0061] Figure 6 A flow chart of another method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0062] Figure 7 A flow chart of another method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0063] Figure 8 A schematic diagram of a device for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown.

[0064] Fig. 9 A structural block diagram of an electronic device in an embodiment of the present disclosure is shown.

[0065] Fig.10 A schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION

[0066] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0067] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0068] It should be understood that the various steps described in the method embodiments of the present disclosure may be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0069] It should be noted that the concepts such as "first" and "second" mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0070] It should be noted that the modifications of "one" and "plurality" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".

[0071] It should be pointed out that, in the absence of conflict, the embodiments of the present disclosure and the technical features therein may be combined with each other.

[0072] Figure 1 A schematic diagram of the structure of a deep supervised Gaussian splash model determination system in an embodiment of the present disclosure is shown. The system can be applied to the deep supervised Gaussian splash model determination method and the deep supervised Gaussian splash model determination device in each embodiment of the present disclosure.

[0073] like Figure 1 As shown, the deep supervised Gaussian splash model determination system 10 may include a training device 101 and an application device 102 .

[0074] In some embodiments, in some embodiments, the training device 101 and the application device 102 may be located on different devices, and the training device 101 may be a module on an electronic device with data acquisition capability. The application device 102 may be a module on an electronic device with data processing capability such as a computer or a server. The training device 101 and the application device 102 may be located on the same device, such as the training device 101 and the application device 102 may be an input module and a processing module on a computer or a mobile phone.

[0075] The training device 101 and the application device 102 are connected to each other via a network, which may be a wired network or a wireless network.

[0076] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.

[0077] The following describes a situation where the training device 101 and the application device 102 are located on two different devices.

[0078] The training device 101 may be located on a terminal device, which may be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, wearable devices, augmented reality devices, virtual reality devices, etc.

[0079] Optionally, the client of the application installed in different terminal devices is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile client, a PC client, etc.

[0080] The application device 102 may be located on a server, which may be a server that provides various services, such as a background management server that provides support for devices operated by users using terminal devices. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.

[0081] Optionally, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), as well as big data and artificial intelligence platforms.

[0082] Those skilled in the art will know that Figure 1 The number of training devices 101 and application devices 102 in the figure is only for illustration, and any number of training devices 101 and application devices 102 may be provided according to actual needs. The embodiments of the present disclosure are not limited to this.

[0083] Figure 2 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 2 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0084] S210, obtaining an image of a reconstructed target and a depth image corresponding to the image.

[0085] In some embodiments, the reconstruction target may be any non-planar target from which images may be acquired.

[0086] In some embodiments, the depth image may include depth parameters corresponding to the image.

[0087] In some embodiments, a depth camera may be used to acquire an image and a depth image corresponding to the image.

[0088] In some embodiments, the image may be an RGB image.

[0089] S220, segmenting the image based on the segmentation model to obtain depth data of different objects in the image.

[0090] In some embodiments, the segmentation model may include a model for performing regional segmentation on different objects in the image. Exemplarily, the segmentation model may include segment anything.

[0091] S230, training a Gaussian splash model based on the image and the depth image.

[0092] In some embodiments, the images used to train the Gaussian splash model may include images and labels or annotations corresponding to the images. During the training process, the Gaussian splash model will learn the features of the images and optimize its own parameters based on the labels or annotations to improve the capabilities of the model.

[0093] In some embodiments, before S230, the method may further include:

[0094] The image and the depth image corresponding to the image are preprocessed by denoising and smoothing.

[0095] S240, when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0096] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splash model is trained based on the image and the depth image, and a trained Gaussian splash model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splash model, the training process is supervised based on the depth data, thereby improving the attention of the Gaussian splash model to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0097] Figure 3 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 3 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0098] S310, obtaining an image of a reconstructed target and a depth image corresponding to the image;

[0099] S320, segmenting the image based on the segmentation model to obtain depth data of different objects in the image;

[0100] S330, training a Gaussian splash model based on the image and the depth image;

[0101] S340: Determine depth mean values ​​and depth standard deviations of different objects in the image based on the depth data of different objects.

[0102] In some embodiments, the depth mean may include a summed average value of a plurality of depth data corresponding to any object. The depth standard deviation may include a standard deviation of a plurality of depth data corresponding to any object.

[0103] S350: Determine depth means and depth standard deviations of different predicted objects in the predicted image based on the predicted depth data.

[0104] In some embodiments, after the image and the depth image are input into the Gaussian splash model, the image and the depth image predicted by the Gaussian splash model can be obtained. The predicted depth data can be determined according to the image and the depth image predicted by the Gaussian splash model.

[0105] In some embodiments, after determining the predicted depth data, the method for determining the depth mean and depth standard deviation of different predicted objects is the same as the method for determining the depth mean and standard deviation of the object in the image in the above embodiment, and will not be repeated here.

[0106] S360, constructing a depth loss function based on the depth means and depth standard deviations of different predicted objects and the depth means and depth standard deviations of different objects in the image.

[0107] In some embodiments, the depth loss function may include:

[0108]

[0109] Among them, L depth is the depth loss function value, x is the predicted depth data of the object predicted by the Gaussian splash model, is the depth data of the object, n is the total number of objects, U(i) is the mean of the object depth data, and std(i) is the standard deviation of the object depth data.

[0110] S370, when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0111] In some embodiments, L 1 The loss function can be used to reflect the pixel difference between the predicted image and the real image, and the SSIM loss function can be used to reflect the visual impact and structural similarity between the predicted image and the real image.

[0112] In some embodiments, L 1 The weighted average of the loss function value, SSIM loss function value, and depth loss function value is taken as the target loss function value.

[0113] In some embodiments, the preset threshold may be determined by a user, which is not limited in the present disclosure.

[0114] In some embodiments, L 1 The loss function is mixed with the SSIM loss function, and then the weighted average of the mixed loss function value and the depth loss function value is taken as the target loss function value.

[0115] Exemplarily, the mixed loss function value may be:

[0116] L=αL 1 (β,γ)+(1-α)(1-SSIM(β,γ))

[0117] Among them, β represents the image of the object, γ represents the image of the predicted object, α is a weight parameter between 0 and 1, and L 1 Indicates L 1 Loss function value, SSIM represents the SSIM loss function value.

[0118] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splash model is trained based on the image and the depth image, and a trained Gaussian splash model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splash model, the training process is supervised based on the depth data, thereby improving the attention of the Gaussian splash model to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0119] Figure 4 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 4 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0120] S410, obtaining an image of a reconstructed target and a depth image corresponding to the image;

[0121] S420, segmenting the image based on the segmentation model to obtain masks of different objects in the image.

[0122] In some embodiments, the image mask is a matrix or image of the same size as the original image, and is used to indicate or identify pixels or regions in a specific area in the original image.

[0123] S430, determining depth data of different objects based on the masks of different objects.

[0124] In some embodiments, after the masks of different objects are determined, depth data corresponding to each object may be determined based on the different objects.

[0125] S440, training a Gaussian splash model based on the image and the depth image;

[0126] S450, when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0127] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splash model is trained based on the image and the depth image, and a trained Gaussian splash model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splash model, the training process is supervised based on the depth data, thereby improving the attention of the Gaussian splash model to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0128] Figure 5 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 5 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0129] S510, obtaining an image of a reconstructed target and a depth image corresponding to the image;

[0130] S520, segmenting the image based on the segmentation model to obtain depth data of different objects in the image;

[0131] S530, initializing the point cloud of the image based on the structure from motion algorithm SFM;

[0132] In some embodiments, it is an important technology in the field of computer vision, which infers the three-dimensional structure of the scene and the motion parameters of the camera by analyzing the image sequences captured by the camera at different perspectives.

[0133] S540, training a Gaussian splash model based on the initialized point cloud;

[0134] In some embodiments, the point cloud of the depth image may also be initialized based on the SFM algorithm, and then the Gaussian splatter model may be trained based on the initialized point cloud and the point cloud corresponding to the depth image.

[0135] S550, when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0136] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splatter model is trained based on the image and the depth image, and a trained Gaussian splatter model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splatter model, the training process is supervised based on the depth data, thereby improving the Gaussian splatter model's attention to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0137] Figure 6 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 6 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0138] S610, obtaining an image of a reconstructed target and a depth image corresponding to the image;

[0139] S620, segmenting the image based on the segmentation model to obtain depth data of different objects in the image;

[0140] S630, training a Gaussian splash model based on the image and the depth image;

[0141] S640: When the target loss function value does not reach a preset threshold, adjust the parameters of the Gaussian splash model based on the target loss function value.

[0142] In some embodiments, the network weights may be updated based on the back-propagation algorithm to calculate the gradient of the target loss function with respect to the network parameters, thereby reducing the loss function.

[0143] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splatter model is trained based on the image and the depth image, and a trained Gaussian splatter model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splatter model, the training process is supervised based on the depth data, thereby improving the Gaussian splatter model's attention to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0144] Figure 7 FIG. 2 is a flow chart of a method for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure. Figure 7 As shown, the method for determining the deep supervised Gaussian splash model in the embodiment of the present disclosure may include:

[0145] S710, obtaining an image of a reconstructed target and a depth image corresponding to the image;

[0146] S720, segmenting the image based on the segmentation model to obtain depth data of different objects in the image;

[0147] S730, training a Gaussian splash model based on the image and the depth image;

[0148] S740, when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained, and the target loss function value is L 1 The loss function, the SSIM loss function and the depth loss function are determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model;

[0149] S750, input the target perspective into the trained Gaussian splash model to obtain a reconstructed target prediction image under the target perspective.

[0150] In some embodiments, in addition to inputting the target perspective into the trained Gaussian splash model, camera parameters under the target perspective may also be input into the Gaussian splash model.

[0151] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splash model is trained based on the image and the depth image, and a trained Gaussian splash model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splash model, the training process is supervised based on the depth data, thereby improving the attention of the Gaussian splash model to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0152] Based on the same inventive concept, the present disclosure also provides a device for determining a deep supervised Gaussian splash model, such as the following embodiment. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.

[0153] Figure 8 A schematic diagram of a device for determining a deep supervised Gaussian splash model in an embodiment of the present disclosure is shown. Figure 8 As shown, the deep supervised Gaussian splash model determination device 800 includes:

[0154] An acquisition module 810 is used to acquire an image of a reconstructed target and a depth image corresponding to the image;

[0155] A segmentation module 820 is used to segment the image based on the segmentation model to obtain depth data of different objects in the image;

[0156] A training module 830, for training a Gaussian splash model based on the image and the depth image;

[0157] The judgment module 840 is used to obtain the trained Gaussian splash model when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold value. The target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0158] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splatter model is trained based on the image and the depth image, and a trained Gaussian splatter model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splatter model, the training process is supervised based on the depth data, thereby improving the Gaussian splatter model's attention to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0159] In some embodiments, the apparatus further comprises:

[0160] A first determination module, configured to determine depth mean values ​​and depth standard deviations of different objects in an image based on depth data of different objects;

[0161] A second determination module, configured to determine depth means and depth standard deviations of different predicted objects in the predicted image based on the predicted depth data;

[0162] A construction module is used to construct a depth loss function based on the depth mean and depth standard deviation of different predicted objects and the depth mean and depth standard deviation of different objects in the image.

[0163] In some embodiments, the depth loss function includes:

[0164]

[0165] Among them, L depth is the depth loss function value, x is the predicted depth data of the object predicted by the Gaussian splash model, is the depth data of the object, n is the total number of objects, U(i) is the mean of the object depth data, and std(i) is the standard deviation of the object depth data.

[0166] In some embodiments, the segmentation module includes:

[0167] A segmentation unit, used to segment the image based on the segmentation model to obtain masks of different objects in the image;

[0168] The determination unit is used to determine the depth data of different objects based on the masks of different objects.

[0169] In some embodiments, the training module includes:

[0170] An initialization unit, used to initialize the point cloud of the image based on the structure from motion algorithm SFM;

[0171] The training unit is used to train the Gaussian splash model based on the initialized point cloud.

[0172] In some embodiments, the apparatus further comprises:

[0173] The adjustment module is used to adjust the parameters of the Gaussian splash model based on the target loss function value when the target loss function value does not reach a preset threshold.

[0174] In some embodiments, the apparatus further comprises:

[0175] The input module is used to input the target perspective into the trained Gaussian splash model to obtain a reconstructed target prediction image under the target perspective.

[0176] In the disclosed embodiment, an image of a reconstruction target and a depth image corresponding to the image are obtained, the image is segmented based on a segmentation model to obtain depth data of different objects in the image, a Gaussian splash model is trained based on the image and the depth image, and a trained Gaussian splash model is obtained when the target loss function value constructed by the depth loss function, the L1 loss function and the SSIM loss function reaches a preset threshold. During the training of the Gaussian splash model, the training process is supervised based on the depth data, thereby improving the attention of the Gaussian splash model to detail areas with a smaller depth order, thereby improving the reconstruction effect on details.

[0177] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".

[0178] The present disclosure provides an electronic device, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any of the above-mentioned deep supervised Gaussian splash model determination methods by executing the executable instructions.

[0179] For example, refer to Fig. 9 The electronic device 900 according to this embodiment of the present disclosure is described. Fig. 9 The electronic device 900 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0180] like Fig. 9 As shown, the electronic device 900 is in the form of a general computing device. The components of the electronic device 900 may include but are not limited to: at least one processing unit 910, at least one storage unit 920, and a bus 930 connecting different system components (including the storage unit 920 and the processing unit 910).

[0181] The storage unit stores program codes, which can be executed by the processing unit 910, so that the processing unit 910 performs the steps of various exemplary embodiments of the present disclosure described in the above “Exemplary Method” section of this specification. For example, the processing unit 910 can perform the following steps of the above method embodiment:

[0182] Obtain an image of the reconstructed target and a depth image corresponding to the image;

[0183] Segment the image based on the segmentation model to obtain the depth data of different objects in the image;

[0184] Train the Gaussian splash model based on images and depth images;

[0185] When the target loss function value corresponding to the Gaussian splash model reaches the preset threshold, the trained Gaussian splash model is obtained, and the target loss function value is L 1 The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

[0186] The storage unit 920 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 9201 and / or a cache storage unit 9202 , and may further include a read-only storage unit (ROM) 9203 .

[0187] The storage unit 920 may also include a program / utility 9204 having a set (at least one) of program modules 9205, such program modules 9205 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.

[0188] Bus 930 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.

[0189] The electronic device 900 may also communicate with one or more external devices 940 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 900, and / or may communicate with any device that enables the electronic device 900 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 950. Furthermore, the electronic device 900 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 960. As shown, the network adapter 960 communicates with other modules of the electronic device 900 via a bus 930. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 900, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0190] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0191] In the disclosed exemplary embodiments, a computer-readable storage medium is also provided, which may be a readable signal medium or a readable storage medium. Fig.10 A schematic diagram of a computer-readable storage medium in an embodiment of the present disclosure is shown. Fig.10 As shown, the computer-readable storage medium 1000 stores a program product capable of implementing the above method of the present disclosure.

[0192] In some possible implementations, various aspects of the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a terminal device, the program code is used to enable the terminal device to execute the steps of various exemplary implementations of the present disclosure described in the above "Specific Implementation Methods" section of this specification.

[0193] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0194] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0195] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.

[0196] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages ​​such as Java, C++, etc., and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0197] The embodiments of the present disclosure provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. A processor of a computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the display method provided in various optional ways in any embodiment of the present disclosure.

[0198] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.

[0199] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.

[0200] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope of the present disclosure is indicated by the appended claims.

Claims

1. A method for determining a deep supervised Gaussian splash model, characterized in that: include: Acquire an image of a reconstructed target and a depth image corresponding to the image; Segmenting the image based on the segmentation model to obtain depth data of different objects in the image; Training a Gaussian splash model based on the image and the depth image; When the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, a trained Gaussian splash model is obtained. The target loss function value is determined by the L1 loss function, the SSIM loss function and the depth loss function. The depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

2. The method according to claim 1, characterized in that: The method further comprises: Determine depth mean and depth standard deviation of different objects in the image based on the depth data of the different objects; Determining depth means and depth standard deviations of different predicted objects in the predicted image based on the predicted depth data; The depth loss function is constructed based on the depth means and depth standard deviations of the different predicted objects and the depth means and depth standard deviations of different objects in the image.

3. The method according to claim 1, characterized in that The depth loss function includes: Among them, L depth is the depth loss function value, x is the predicted depth data of the object predicted by the Gaussian splash model, is the depth data of the object, n is the total number of objects, U(i) is the mean of the object depth data, and std(i) is the standard deviation of the object depth data.

4. The method according to claim 1, characterized in that: The step of segmenting the image based on the segmentation model to obtain depth data of different objects in the image includes: Segmenting the image based on the segmentation model to obtain masks of different objects in the image; Depth data of the different objects are determined based on the masks of the different objects.

5. The method according to claim 1, characterized in that The Gaussian splash model is trained based on the image and the depth image, including: Initializing the point cloud of the image based on the structure from motion algorithm SFM; The Gaussian splash model is trained based on the initialized point cloud.

6. The method according to claim 1, characterized in that The method further comprises: When the target loss function value does not reach a preset threshold, the parameters of the Gaussian splash model are adjusted based on the target loss function value.

7. The method according to claim 1, characterized in that The method further comprises: The target perspective is input into the trained Gaussian splash model to obtain a reconstructed target prediction image under the target perspective.

8. A device for determining a deep supervised Gaussian splash model, characterized in that: include: An acquisition module, used to acquire an image of a reconstructed target and a depth image corresponding to the image; A segmentation module, used to segment the image based on a segmentation model to obtain depth data of different objects in the image; A training module, used for training a Gaussian splash model based on the image and the depth image; A judgment module is used to obtain a trained Gaussian splash model when the target loss function value corresponding to the Gaussian splash model reaches a preset threshold, wherein the target loss function value is determined by the L1 loss function, the SSIM loss function and the depth loss function, and the depth loss function is determined by the depth data and the predicted depth data of the object predicted by the Gaussian splash model.

9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the deep supervised Gaussian splash model method described in any one of claims 1 to 7 by executing the executable instructions.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the deep supervised Gaussian splash model method described in any one of claims 1 to 7 is implemented.