Training method, three-dimensional reconstruction method, device and medium of network model
By constructing a network model that includes feature extraction, deblurring, and reconstruction modules, the blurring problem during human movement was solved, the reconstruction effect was improved, the accurate restoration of the virtual human hand was achieved, and the training and inference time was reduced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING CENTURY TAL EDUCATION TECH CO LTD
- Filing Date
- 2023-03-10
- Publication Date
- 2026-07-24
Smart Images

Figure CN116309158B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and in particular to a method for training a network model, a method for three-dimensional reconstruction, an apparatus, a device, and a medium. Background Technology
[0002] With the rise of the metaverse, virtual human technology has become increasingly mature. Human body reconstruction is a crucial step in the creation of virtual humans, and the reconstruction results directly affect the driving performance of the virtual human. Currently, motion blur frequently occurs during human movement, often affecting reconstruction accuracy. For areas of the human body that require frequent movement and are therefore more prone to motion blur, the reconstruction results are relatively poor. Summary of the Invention
[0003] To address the aforementioned technical problems, this disclosure provides a method for training a network model, a method for 3D reconstruction, an apparatus, a device, and a medium that can effectively solve the motion blur problem and improve the reconstruction effect in motion blur scenes.
[0004] According to one aspect of this disclosure, a method for training a network model is provided, the network model including a feature extraction module, a deblurring module, and a reconstruction module, the method comprising:
[0005] Acquire a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image;
[0006] The feature extraction module is used to extract feature information from the blurred sample image;
[0007] The feature information is deblurred using the deblurring module to obtain a deblurred image of the blurred sample image, and a first loss is calculated based on the deblurred image and the clear sample image.
[0008] The feature information is reconstructed using the reconstruction module to obtain reconstruction prediction data of the blurred sample image, and a second loss is calculated based on the reconstruction prediction data and the reconstructed sample data.
[0009] The network model parameters are optimized based on the first loss and the second loss until convergence, resulting in a trained network model.
[0010] According to another aspect of this disclosure, a three-dimensional reconstruction method is provided, the method comprising:
[0011] Obtain a blurred image;
[0012] The feature extraction module in the network model trained using the above-described network model training method extracts the feature information of the blurred image.
[0013] The reconstruction module in the network model generates a reconstructed image of the blurred image based on the feature information.
[0014] According to another aspect of this disclosure, a training apparatus for a network model is provided, the network model including a feature extraction module, a deblurring module, and a reconstruction module, the apparatus comprising:
[0015] The first acquisition unit is used to acquire a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image.
[0016] The first extraction unit is used to extract feature information of the blurred sample image using the feature extraction module;
[0017] The deblurring processing unit is used to deblurr the feature information using the deblurring module to obtain a deblurred image of the blurred sample image, and to calculate a first loss based on the deblurred image and the clear sample image.
[0018] The first reconstruction unit is used to reconstruct the feature information using the reconstruction module to obtain reconstruction prediction data of the blurred sample image, and to calculate a second loss based on the reconstruction prediction data and the reconstructed sample data.
[0019] The training unit is used to optimize the model parameters of the network model based on the first loss and the second loss until convergence, so as to obtain the trained network model.
[0020] According to another aspect of this disclosure, a three-dimensional reconstruction apparatus is provided, the apparatus comprising:
[0021] The second acquisition unit is used to acquire the blurred image;
[0022] The second extraction unit is used to extract the feature information of the blurred image using the feature extraction module in the network model trained by the above method.
[0023] The second reconstruction unit is used to generate a reconstructed image of the blurred image based on the feature information through the reconstruction module in the network model.
[0024] According to another aspect of this disclosure, an electronic device is provided, the electronic device comprising: a processor; and a memory storing a program, wherein the program includes instructions that, when executed by the processor, cause the processor to perform a training method according to the network model described above, or to perform a three-dimensional reconstruction method according to the three-dimensional reconstruction method described above.
[0025] According to another aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions for causing the computer to perform a training method based on a network model, or to perform the three-dimensional reconstruction method described above.
[0026] According to another aspect of this disclosure, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the above-described network model training method or the above-described three-dimensional reconstruction method.
[0027] The technical solution provided in this disclosure has the following advantages compared with the prior art:
[0028] The network model includes a feature extraction module, a deblurring module, and a reconstruction module. The method includes: acquiring blurred sample images, corresponding sharp sample images, and reconstructed sample data corresponding to the sharp sample images; extracting feature information from the blurred sample images using the feature extraction module; deblurring the feature information using the deblurring module, and calculating a first loss based on the obtained deblurred image and sharp sample image; reconstructing the feature information using the reconstruction module, and calculating a second loss based on the obtained reconstruction prediction data and reconstructed sample data; optimizing the model parameters of the network model based on the first and second losses until convergence, resulting in a trained network model. The method provided in this disclosure can effectively solve the motion blur problem and improve the reconstruction effect in motion-blurred scenes. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0030] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0031] Figure 1 A flowchart illustrating the training method for the network model provided in this embodiment of the disclosure;
[0032] Figure 2 A schematic diagram of the network model provided in the embodiments of this disclosure;
[0033] Figure 3 for Figure 1 The diagram shows a detailed flowchart of S140 in the training method of the network model.
[0034] Figure 4A flowchart of the three-dimensional reconstruction method provided in the embodiments of this disclosure;
[0035] Figure 5 A schematic diagram of the structure of a training device for a network model provided in an embodiment of this disclosure;
[0036] Figure 6 This is a schematic diagram of the structure of the three-dimensional reconstruction device provided in the embodiments of this disclosure;
[0037] Figure 7 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0038] To better understand the above-described objects, features, and advantages of this disclosure, embodiments of the disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of the disclosure are shown in the drawings, it should be understood that the disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0039] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0040] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0041] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0042] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0043] With the rise of the metaverse, virtual human technology has matured significantly. Human body reconstruction, a crucial step in the creation of driveable virtual humans, directly impacts the final driving effect. Motion blurring frequently occurs during human movement, especially in the hand area. This is because the hand occupies a smaller portion of the image compared to the body, and as a vital interface for virtual human interaction, the hand requires frequent movement to transmit information. Therefore, hand motion blur is significantly more likely to occur than in other areas of the body. Effectively addressing the blurring issue in human hands during movement presents a major challenge for achieving high-quality, driveable virtual human technology.
[0044] Currently, deep learning methods are often used to solve motion blur problems. Existing deep learning methods can solve this problem through two approaches: one approach is to increase the number of blurred training samples through data augmentation. Although this approach can improve the blurring problem, the deblurring effect is limited. The other approach is to train a deblurring model, first use the deblurring model to deblur the data, and then reconstruct the deblurred data. That is, deblurring first and then reconstructing. Although this approach can deblur well, it requires training a separate deblurring model, which increases the time consumption of the inference process.
[0045] To address the aforementioned technical problems, this disclosure provides a method for training a network model. The method constructs a network model including a feature extraction module, a deblurring module, and a reconstruction module. The deblurring and reconstruction modules share the feature information output by the feature extraction module, enabling the network model to perform both deblurring and reconstruction tasks simultaneously. Finally, the network model is trained based on the outputs of the deblurring and reconstruction modules, performing joint training (multi-task training), which effectively reduces training time. Simultaneously, the reconstruction module also possesses a certain deblurring function. For subsequent 3D reconstruction inference tasks involving blurred images, the reconstruction module can effectively solve the motion blur problem without requiring a separately trained deblurring model for pre-reconstruction deblurring. This effectively reduces inference time and provides accurate reconstruction of real-world movements for virtual humans, making it highly practical. A detailed explanation is provided below using at least one embodiment, taking a human hand as an example.
[0046] Specifically, the network model training method and the 3D reconstruction method can be executed by a server or a terminal. In one feasible application scenario, the server executes the network model training method to obtain a trained network model, and the terminal executes the 3D reconstruction method. The terminal obtains the trained network model from the server and uses the reconstruction model within the network model to perform deblurring and 3D reconstruction processing on the blurred image. This blurred image can be obtained by the terminal itself, or it can be obtained from another device. The executing entities of the network model training method and the 3D reconstruction method can be the same or different. In another application scenario, the server trains the network model. Further, the server uses the trained network model to perform deblurring and 3D reconstruction processing on the blurred image. The way the server obtains the blurred image can be similar to the way the terminal obtains the blurred image as described above, and will not be repeated here. In yet another application scenario, the terminal trains the network model. Further, the terminal uses the trained network model to perform deblurring and 3D reconstruction processing on the blurred image. It is understood that the network model training method and 3D reconstruction method provided in this disclosure are not limited to the possible scenarios described above. The following one or more embodiments are described in detail using server-side network model training and 3D reconstruction methods as examples.
[0047] Figure 1 A flowchart of a network model training method provided in this disclosure embodiment, applied to a server, specifically includes, as follows: Figure 1 The following steps S110 to S150 are shown:
[0048] The network model includes a feature extraction module, a deblurring module, and a reconstruction module.
[0049] For example, see Figure 2 , Figure 2This is a schematic diagram of the network model provided in this embodiment. The network model includes a feature extraction module, a deblurring module, and a reconstruction module. The network model can be considered an improved reconstruction model. The original reconstruction model includes a feature extraction module and a reconstruction module. Based on the original reconstruction model, a deblurring branch is added to restructure the model, resulting in the improved reconstruction model, which is the network model. In the network model, the deblurring module and the reconstruction module share the feature extraction module. During the training phase of the network model, it can simultaneously complete the 3D reconstruction task and the deblurring task. The feature extraction module (encoder) can use a Human Keypoint Detection Network (HRNet) to extract feature information from the blurred image. For example, HRNet32 can be used as the encoder, and HRNet can output feature information at different feature scales. The deblurring module (Deblur) is used to correct the feature information of the blurred image, that is, to deblur the blurred image and output a clear image. The reconstruction module is used to predict reconstructed data (3D mesh) based on feature information, that is, to perform three-dimensional reconstruction processing on blurred images and output reconstructed images. The reconstruction module also includes a perception submodule and a reconstruction submodule. The perception submodule can use a multilayer perceptron (MLP) to calculate parameter information, and the reconstruction submodule can use a mesh generation model to predict reconstructed data based on parameter information. For example, the mesh generation model can be a parametric model (MANO) for the hand. The reconstructed data can be the reconstructed image of the blurred image, or it may be the key points in the blurred image.
[0050] S110. Obtain a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image.
[0051] Understandably, taking the human hand region as an example, a clear sample image can be a clear image of the hand, a blurry sample image can be a blurry image of the hand, and the reconstructed sample data can be a reconstructed sample image of a clear sample image, or parameter data and keypoint data related to reconstruction. The reconstructed sample data can be understood as real data. Real data and the predicted data output by the grid model are used to train the grid model. Specifically, multiple blurry sample images, multiple clear sample images, and multiple reconstructed sample data can be combined into a first dataset for training the network model. It is understood that each blurry sample data has a corresponding clear sample data, and each clear sample data has a corresponding reconstructed sample data.
[0052] Optionally, the above S110 can be implemented through the following steps:
[0053] A reconstructed sample set is obtained, which includes clear sample images and reconstructed sample data corresponding to the clear sample images; the clear sample images are blurred using a constructed blur kernel to obtain blurred sample images.
[0054] Understandably, the reconstructed sample set refers to the second dataset used to train the original reconstruction model. The second dataset includes sharp sample images and the corresponding reconstructed sample data. In the sample preprocessing stage, the sharp sample images are blurred using a constructed blur kernel to obtain blurred sample images corresponding to the sharp sample images. This is used to obtain the first dataset based on the second dataset. In other words, motion-blurred data sample pairs are constructed by constructing a blur kernel, and each sample pair includes a sharp sample image and a blurred sample image.
[0055] S120. Use the feature extraction module to extract the feature information of the blurred sample image.
[0056] Understandably, based on the above S110, the blurred sample image is input into the feature extraction module to extract feature information of the blurred sample image at different feature scales. Feature information at different feature scales can also be understood as feature maps of different resolutions. Specifically, taking four types of resolution feature maps as an example, the feature maps of the blurred sample image are extracted to obtain the first type of resolution feature map. The first type of resolution feature map is downsampled to obtain the second type of resolution feature map. The second type of resolution feature map is downsampled again to obtain the third type of resolution feature map, and so on, until the fourth type of resolution feature map is obtained. The feature information of the blurred sample image can be obtained based on the four types of resolution feature maps.
[0057] S130. The feature information is deblurred using the deblurring module to obtain a deblurred image of the blurred sample image, and a first loss is calculated based on the deblurred image and the clear sample image.
[0058] Understandably, based on the above steps S110 and S120, the feature information is input into the deblurring module for deblurring processing to obtain a deblurred image of the blurred sample image. The deblurred image can be understood as the predicted image output by the deblurring module after correcting the feature information. The deblurred image and the blurred sample image can have the same size, for example, 256*256*3. After obtaining the deblurred image, it is compared with the clear sample image, and the first loss of the deblurring module is calculated. The clear sample image is the true image.
[0059] The feature information includes first feature information at a first feature scale.
[0060] Optionally, in S130 above, the deblurring module is used to deblur the feature information to obtain the deblurred image of the blurred sample image. This can be achieved through the following steps:
[0061] The first feature information is deblurred using the deblurring module to obtain the deblurred image of the blurred sample image.
[0062] Understandably, the first feature scale can be 64*64*32, which can be obtained by downsampling the encoder by 4 times. Here, the encoder is based on feature information directly extracted from the blurred sample image. Subsequently, the first feature information is input into the deblurring module for deblurring processing to obtain the deblurred image of the blurred sample image. The specific deblurring processing method is not limited.
[0063] S140. The feature information is reconstructed using the reconstruction module to obtain the reconstruction prediction data of the blurred sample image, and a second loss is calculated based on the reconstruction prediction data and the reconstruction sample data.
[0064] Understandably, based on the above S110 and S120, feature information is input into the reconstruction module. The reconstruction module predicts the 3D mesh of the blurred sample image based on the feature information, and obtains the reconstruction prediction data of the blurred sample image. Subsequently, the reconstruction prediction data is compared with the reconstruction sample data, which serves as the real data, and the second loss of the reconstruction module is calculated. Here, the reconstruction prediction data can be a reconstructed image with the same size of 256*256*3.
[0065] The feature information includes second feature information at the second feature scale.
[0066] Optionally, in S140 above, the reconstruction module is used to reconstruct the feature information to obtain the reconstruction prediction data of the blurred sample image, which is specifically achieved through the following steps:
[0067] The reconstruction module is used to reconstruct the second feature information to obtain the reconstruction prediction data of the blurred sample image.
[0068] Understandably, the second feature scale can be 8*8*256. The second feature information can be obtained by downsampling the encoder by 32 times. The second feature information is then input into the reconstruction module for three-dimensional reconstruction processing to obtain the deblurred image of the blurred sample image. The specific reconstruction processing method is not limited.
[0069] The reconstruction module includes a perception submodule and a reconstruction submodule. The reconstruction prediction data of the blurred sample image includes the parameter prediction data and the key point prediction data.
[0070] Optionally, the reconstruction module above outputs reconstruction prediction data, which can be achieved through the following steps:
[0071] The perception submodule calculates parameters based on the second feature information to obtain parameter prediction data; the reconstruction submodule then reconstructs the parameter prediction data to obtain key point prediction data.
[0072] Understandably, the second feature information is input into the perception submodule for parameter prediction, yielding parameter prediction data. This data includes at least one of pose, shape, and camera parameters. The camera parameters (R, T, S) include intrinsic and extrinsic parameters, with the extrinsic parameter representing the camera's pose. The parameter prediction data is then input into the reconstruction submodule for 3D reconstruction. This directly yields keypoint prediction data, denoted as `render`, or a reconstructed prediction image is obtained, from which keypoint prediction data is calculated.
[0073] The reconstructed sample data includes parameter sample data, two-dimensional keypoint sample data, and three-dimensional keypoint sample data, and the keypoint prediction data includes two-dimensional keypoint prediction data and three-dimensional keypoint prediction data.
[0074] Optional, such as Figure 3 As shown, Figure 1 The diagram illustrates the detailed process of S140 in the training method of the network model. See also... Figure 3 In S140 above, the second loss is calculated based on the reconstruction prediction data and the reconstruction sample data, specifically including, for example... Figure 3 The following steps S1401 to S1404 are shown:
[0075] S1401. Calculate the parameterized model loss based on the parameter sample data and the parameter prediction data.
[0076] Understandably, the parameterized model loss of the perception submodule is calculated based on the labeled real data parameter sample data and the parameter prediction data output by the module. The formula for calculating the parameterized model loss is shown in formula (1):
[0077]
[0078] In the formula: L params This represents the loss of the parameterized model, where the parameter sample data is (θ, β, C) and the parameter prediction data is... Where θ is the attitude parameter, β is the shape parameter, and C is the camera parameter.
[0079] S1402. Calculate the two-dimensional keypoint loss based on the two-dimensional keypoint sample data and the two-dimensional keypoint prediction data.
[0080] Understandably, based on the two-dimensional keypoint sample data as real data and the two-dimensional keypoint prediction data output by the module, the two-dimensional keypoint loss related to the reconstruction sub-module is calculated, and the calculation formula is shown in formula (2):
[0081]
[0082] In the formula, L 2D joints The loss is for two-dimensional keypoints, where N is the number of samples. For two-dimensional keypoint prediction data, x i This is two-dimensional keypoint sample data, where i ranges from 0 to N, and x... i Let represent the sample data of the i-th two-dimensional key point.
[0083] S1403. Calculate the 3D keypoint loss based on the 3D keypoint sample data and the 3D keypoint prediction data.
[0084] Understandably, based on the 3D keypoint sample data as real data and the 3D keypoint prediction data output by the module, the 3D keypoint loss related to the reconstruction submodule is calculated, and the calculation formula is shown in formula (3):
[0085]
[0086] In the formula, L 3D ioints Indicates 3D keypoint loss. X represents the predicted 3D keypoints. i This represents the 3D keypoint label.
[0087] S1404. Calculate the sum of the parameterized model loss, the two-dimensional keypoint loss, and the three-dimensional keypoint loss to obtain the second loss.
[0088] Understandably, the sum of the parametric model loss, the 2D keypoint loss, and the 3D keypoint loss is calculated to obtain the second loss used to update the network parameters of the reconstruction module. The calculation formula is shown in formula (4):
[0089] L reconstruct =L 2D joints +L 3D joints +L params (4)
[0090] In the formula, L reconstruct This is the second loss for the reconstruction module.
[0091] Understandably, by calculating the two-dimensional keypoint loss and the three-dimensional keypoint loss, the network parameters of key reconstruction sub-modules can be quickly estimated and learned.
[0092] S150. Optimize the model parameters of the network model based on the first loss and the second loss until convergence, to obtain the trained network model.
[0093] Understandably, based on S130 and S140 above, the sum of the first and second losses is calculated to obtain the total loss. The network parameters of the entire network model are continuously optimized using the total loss until the calculated total loss is less than or equal to the preset loss value, at which point the network model converges to obtain the trained network model. Understandably, during the entire network parameter optimization process, the network parameters of the feature extraction module, deblurring module, and reconstruction module are updated synchronously.
[0094] This disclosure provides a method for training a network model. In the training phase, a deblurring module is added. By jointly training the network model through the deblurring module and the reconstruction module, the feature extraction module can extract effective feature information applicable to fuzzy reconstruction scenarios, and the reconstruction module has the ability to deblur during the reconstruction process, while not increasing the model training time to a certain extent.
[0095] Based on the above embodiments, Figure 4 The flowchart of the 3D reconstruction method provided in this embodiment is applied to a server. After the network model is trained, 3D reconstruction is performed based on the feature extraction module and the reconstruction module, specifically including, as follows: Figure 4 The following steps S410 to S430 are shown:
[0096] S410, Obtain the blurred image.
[0097] Understandably, a blurred image is obtained. The blurred image can be an RGB image, and the blur type of the blurred image is not limited. The blurred image can also be preprocessed to meet the input conditions of the feature extraction module.
[0098] S420. The feature extraction module in the network model trained using the network model training method extracts the feature information of the blurred image.
[0099] Understandably, based on the above S410, the feature extraction module trained by the above network model training method is obtained, the blurred image is input into the feature extraction module, and the feature information of the blurred image is extracted. In application, only the second feature information applicable to the reconstruction module can be retained.
[0100] S430. Based on the feature information, the reconstruction module in the network model generates a reconstructed image of the blurred image.
[0101] Understandably, based on the above S420, the second feature information is input into the trained reconstruction module. During the reconstruction process, the reconstruction module will simultaneously perform deblurring and reconstruction processing on the second feature information. In other words, the trained reconstruction module has both deblurring and reconstruction functions, generating a reconstructed image of the blurred image.
[0102] This disclosure provides a three-dimensional reconstruction method that uses a trained feature extraction module and a reconstruction module to deblur and reconstruct a blurred image. It eliminates the need to explicitly obtain the deblurred image. After feature extraction, a reconstruction module can effectively improve motion blur and obtain more accurate reconstruction results in motion blur scenarios, thereby enabling virtual humans to accurately reproduce real-world actions.
[0103] Figure 5 This is a schematic diagram of the structure of a training device for a network model provided in an embodiment of this disclosure. The network model includes a feature extraction module, a deblurring module, and a reconstruction module, such as... Figure 5 As shown, the apparatus 500 provided in this embodiment may include a first acquisition unit 510, a first extraction unit 520, a deblurring processing unit 530, a first reconstruction unit 540, and a training unit 550, wherein:
[0104] The first acquisition unit 510 is used to acquire a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image.
[0105] The first extraction unit 520 is used to extract feature information of the blurred sample image using the feature extraction module;
[0106] The deblurring processing unit 530 is used to perform deblurring processing on the feature information using the deblurring module to obtain a deblurred image of the blurred sample image, and to calculate a first loss based on the deblurred image and the clear sample image;
[0107] The first reconstruction unit 540 is used to reconstruct the feature information using the reconstruction module to obtain reconstruction prediction data of the blurred sample image, and to calculate a second loss based on the reconstruction prediction data and the reconstruction sample data.
[0108] Training unit 550 is used to optimize the model parameters of the network model based on the first loss and the second loss until convergence, so as to obtain the trained network model.
[0109] In one implementation, the first acquisition unit 510 is used for:
[0110] Obtain a reconstruction sample set, which includes clear sample images and corresponding reconstruction sample data of the clear sample images;
[0111] The clear sample image is blurred by using a constructed blur kernel to obtain a blurred sample image.
[0112] In one implementation, the feature information includes first feature information at a first feature scale.
[0113] In one implementation, the deblurring unit 530 is used for:
[0114] The first feature information is deblurred using the deblurring module to obtain the deblurred image of the blurred sample image.
[0115] In one implementation, the feature information includes second feature information at a second feature scale.
[0116] In one implementation, the first reconstruction unit 540 is used for:
[0117] The reconstruction module is used to reconstruct the second feature information to obtain the reconstruction prediction data of the blurred sample image.
[0118] In one implementation, the reconstruction module includes a perception submodule and a reconstruction submodule.
[0119] In one implementation, the first reconstruction unit 540 is used for:
[0120] The perception submodule is used to calculate parameters based on the second feature information to obtain parameter prediction data.
[0121] The reconstruction submodule is used to reconstruct the parameter prediction data to obtain key point prediction data.
[0122] The reconstruction prediction data of the blurred sample image includes the parameter prediction data and the key point prediction data.
[0123] In one embodiment, the reconstructed sample data includes parameter sample data, two-dimensional keypoint sample data, and three-dimensional keypoint sample data, and the keypoint prediction data includes two-dimensional keypoint prediction data and three-dimensional keypoint prediction data.
[0124] In one implementation, the first reconstruction unit 540 is used for:
[0125] Calculate the parameterized model loss based on the parameter sample data and the parameter prediction data;
[0126] Calculate the two-dimensional keypoint loss based on the two-dimensional keypoint sample data and the two-dimensional keypoint prediction data;
[0127] The 3D keypoint loss is calculated based on the 3D keypoint sample data and the 3D keypoint prediction data.
[0128] The sum of the parameterized model loss, the two-dimensional keypoint loss, and the three-dimensional keypoint loss is calculated to obtain the second loss.
[0129] The device provided in this embodiment has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0130] Figure 6 This is a schematic diagram of the structure of the three-dimensional reconstruction device provided in the embodiments of this disclosure, as shown below. Figure 6 As shown, the apparatus 600 provided in this embodiment may include a second acquisition unit 610, a second extraction unit 620, and a second reconstruction unit 630, wherein:
[0131] The second acquisition unit 610 is used to acquire the blurred image;
[0132] The second extraction unit 620 is used to extract feature information of the blurred image from the feature extraction module in the network model trained by the above-mentioned network model training method.
[0133] The second reconstruction unit 630 is used to generate a reconstructed image of the blurred image based on the feature information through the reconstruction module in the network model.
[0134] The device provided in this embodiment has the same implementation principle and technical effect as the aforementioned method embodiment. For the sake of brevity, any parts not mentioned in the device embodiment can be referred to the corresponding content in the aforementioned method embodiment.
[0135] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.
[0136] Exemplary embodiments of this disclosure also provide a computer program product, including a computer program, wherein, when executed by a processor of a computer, the computer program is used to cause the computer to perform a method according to an embodiment of this disclosure.
[0137] refer to Figure 7The present invention describes a structural block diagram of an electronic device 700 that can serve as a server or client of the present disclosure, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0138] like Figure 7 As shown, the electronic device 700 includes a computing unit 701, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. The RAM 703 may also store various programs and data required for the operation of the device 700. The computing unit 701, ROM 702, and RAM 703 are interconnected via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0139] Multiple components in electronic device 700 are connected to I / O interface 705, including: input unit 706, output unit 707, storage unit 708, and communication unit 709. Input unit 706 can be any type of device capable of inputting information to electronic device 700. Input unit 706 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 707 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 704 may include, but is not limited to, disk and optical disk. Communication unit 709 allows electronic device 700 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0140] The computing unit 701 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 performs the various methods and processes described above. For example, in some embodiments, the method for training a network model or the method for training a reconstructed model can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 700 via ROM 702 and / or communication unit 709. In some embodiments, the computing unit 701 can be configured by any other suitable means (e.g., by means of firmware) to perform the method for training a network model or the method for training a reconstructed model.
[0141] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0142] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0143] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0146] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0147] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for training a network model, characterized in that, The network model includes a feature extraction module, a deblurring module, and a reconstruction module; the method includes: Acquire a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image, wherein the reconstructed sample data includes parameter sample data, two-dimensional key point sample data, and three-dimensional key point sample data; The feature extraction module is used to extract feature information from the blurred sample image; The feature information is deblurred using the deblurring module to obtain a deblurred image of the blurred sample image, and a first loss is calculated based on the deblurred image and the clear sample image. The feature information is reconstructed using the reconstruction module to obtain reconstruction prediction data of the blurred sample image, and a second loss is calculated based on the reconstruction prediction data and the reconstructed sample data; wherein, the reconstruction prediction data of the blurred sample image includes parameter prediction data and key point prediction data, the parameter prediction data includes at least one of pose parameters, shape parameters and camera parameters, and the key point prediction data includes two-dimensional key point prediction data and three-dimensional key point prediction data; The network model parameters are optimized based on the first loss and the second loss until convergence, resulting in a trained network model. This allows the trained reconstruction module to perform deblurring and 3D reconstruction on the feature information output by the trained feature extraction module, generating a reconstructed image of the blurred image.
2. The method according to claim 1, characterized in that, The acquisition of the blurred sample image, the corresponding clear sample image, and the reconstructed sample data corresponding to the clear sample image includes: Obtain a reconstruction sample set, which includes clear sample images and corresponding reconstruction sample data of the clear sample images; The clear sample image is blurred by using a constructed blur kernel to obtain a blurred sample image.
3. The method according to claim 1, characterized in that, The feature information includes first feature information at a first feature scale. The step of using the deblurring module to deblurr the feature information to obtain a deblurred image of the blurred sample image includes: The first feature information is deblurred using the deblurring module to obtain the deblurred image of the blurred sample image.
4. The method according to claim 1, characterized in that, The feature information includes second feature information at a second feature scale. The reconstruction module performs reconstruction processing on the feature information to obtain the reconstruction prediction data of the blurred sample image, including: The reconstruction module is used to reconstruct the second feature information to obtain the reconstruction prediction data of the blurred sample image.
5. The method according to claim 4, characterized in that, The reconstruction module includes a perception submodule and a reconstruction submodule. The reconstruction module is used to reconstruct the second feature information to obtain the reconstruction prediction data of the blurred sample image, including: The perception submodule is used to calculate parameters based on the second feature information to obtain parameter prediction data. The reconstruction submodule is used to reconstruct the parameter prediction data to obtain key point prediction data.
6. The method according to claim 5, characterized in that, The calculation of the second loss based on the reconstruction prediction data and the reconstruction sample data includes: Calculate the parameterized model loss based on the parameter sample data and the parameter prediction data; Calculate the two-dimensional keypoint loss based on the two-dimensional keypoint sample data and the two-dimensional keypoint prediction data; The 3D keypoint loss is calculated based on the 3D keypoint sample data and the 3D keypoint prediction data. The sum of the parameterized model loss, the two-dimensional keypoint loss, and the three-dimensional keypoint loss is calculated to obtain the second loss.
7. A three-dimensional reconstruction method, characterized in that, The method includes: Obtain a blurred image; The feature extraction module in the network model trained using the training method of any one of claims 1-6 extracts the feature information of the blurred image; The reconstruction module in the network model generates a reconstructed image of the blurred image based on the feature information.
8. A training device for a network model, characterized in that, The network model includes a feature extraction module, a deblurring module, and a reconstruction module; the device includes: The first acquisition unit is used to acquire a blurred sample image, a clear sample image corresponding to the blurred sample image, and reconstructed sample data corresponding to the clear sample image, wherein the reconstructed sample data includes parameter sample data, two-dimensional key point sample data, and three-dimensional key point sample data. The first extraction unit is used to extract feature information of the blurred sample image using the feature extraction module; The deblurring processing unit is used to deblurr the feature information using the deblurring module to obtain a deblurred image of the blurred sample image, and to calculate a first loss based on the deblurred image and the clear sample image. The first reconstruction unit is used to perform three-dimensional reconstruction processing on the feature information using the reconstruction module to obtain reconstruction prediction data of the blurred sample image, and calculate a second loss based on the reconstruction prediction data and the reconstruction sample data; wherein, the reconstruction prediction data of the blurred sample image includes parameter prediction data and key point prediction data, the parameter prediction data includes at least one of pose parameters, shape parameters and camera parameters, and the key point prediction data includes two-dimensional key point prediction data and three-dimensional key point prediction data; The training unit is used to optimize the model parameters of the network model based on the first loss and the second loss until convergence, so as to obtain the trained network model, so that the trained reconstruction module can perform deblurring and three-dimensional reconstruction processing on the feature information output by the trained feature extraction module to generate a reconstructed image of the blurred image.
9. A three-dimensional reconstruction device, characterized in that, The device includes: The second acquisition unit is used to acquire the blurred image; The second extraction unit is used to extract feature information of the blurred image using the feature extraction module in the network model trained by any one of claims 1-6. The second reconstruction unit is used to generate a reconstructed image of the blurred image based on the feature information through the reconstruction module in the network model.
10. An electronic device, characterized in that, The electronic device includes: Processor; and Stored program memory, The program includes instructions that, when executed by the processor, cause the processor to perform the training method of the network model according to any one of claims 1 to 6, or to perform the three-dimensional reconstruction method according to claim 7.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the training method of the network model according to any one of claims 1 to 6, or to execute the three-dimensional reconstruction method according to claim 7.