Training method of generative model, image processing method, device and equipment

By mapping three-dimensional voxel features to four-dimensional space and using combined training of multiple generative models, the problem of low reconstruction accuracy of generative models is solved, and the reconstruction accuracy of scene information and the robustness of the model are improved.

CN120807772APending Publication Date: 2025-10-17XIAOMI EV TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510798859.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

The existing generative models have low accuracy in reconstructing scene information over a period of time and are unable to meet the requirements of diversity and complexity of scene information.

Method used

The voxel features in three-dimensional space are mapped to four-dimensional space, and the feature reconstruction capability is improved by combining and training multiple generative models and utilizing a supervised mechanism, including the latent space feature encoding and decoding capabilities of the encoder and decoder.

Benefits of technology

It significantly improves the accuracy of four-dimensional feature reconstruction in generative models, enhances the ability to reconstruct dynamic scene information, reduces training costs, and improves the robustness and reliability of assisted driving, robot control, and extended reality models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807772A_ABST
    Figure CN120807772A_ABST
Patent Text Reader

Abstract

The invention relates to a training method of a generative model, an image processing method, a device and equipment, and belongs to the technical field of artificial intelligence. The method comprises the following steps: mapping first sample features of voxels of M frames in a three-dimensional space to a four-dimensional space to generate first four-dimensional features; the first four-dimensional features are processed through a first generation model, second four-dimensional features are generated, and the second four-dimensional features carry M frames of first reconstruction features of voxels in the three-dimensional space; generating a second reconstruction feature of the voxel of each frame in the three-dimensional space through the first generation model and the second generation model; and training the first generation model based on the first sample feature, the first reconstruction feature and the second reconstruction feature. Therefore, the training of the first generative model can be supervised by using the second generative model, and the first generative model can learn the single-frame three-dimensional feature reconstruction capability of the second generative model in the training process, so that the reconstruction precision of the dynamic scene information of the first generative model is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of artificial intelligence, and particularly relates to a training method of a generation model, an image processing method and device, electronic equipment and program product. BACKGROUND

[0002] At present, the cost of real scene information collection is high, and the scene coverage is limited. The reconstruction accuracy of the scene information of a period of time of the generation model in the related technology is low, which is difficult to meet the diversity and complexity requirements of the scene information. SUMMARY

[0003] The present disclosure provides a training method of a generation model, an image processing method and device, electronic equipment, chip, storage medium and program product to at least solve the problem that the reconstruction accuracy of the scene information of a period of time of the generation model in the related technology is low. The technical solutions of the present disclosure are as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, a training method of a generation model is provided, comprising: mapping a first sample feature of M frames of voxels in a three-dimensional space to a four-dimensional space to generate a first four-dimensional feature, M being an integer greater than 1, the four-dimensional space being constructed based on three spatial dimensions and a time dimension; processing the first four-dimensional feature through a first generation model to generate a second four-dimensional feature in the four-dimensional space, the second four-dimensional feature carrying a first reconstruction feature of the M frames of voxels in the three-dimensional space; generating a second reconstruction feature of voxels in the three-dimensional space of each frame through the first generation model and a second generation model, the second generation model being used to reconstruct the feature of voxels in the three-dimensional space of any frame; and training the first generation model based on the first sample feature, the first reconstruction feature and the second reconstruction feature.

[0005] The training method of the generation model provided by the embodiments of the present disclosure can supervise the training of the first generation model by using the second generation model. The first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model in the training process, so that the trained first generation model and the second generation model keep consistent in reconstructing the three-dimensional feature of the same frame, and the four-dimensional feature reconstruction accuracy of the first generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0006] In some possible implementation manners, the generating the second reconstruction feature of each frame in the three-dimensional space by the first generation model and the second generation model comprises: projecting the first sample feature corresponding to any frame to a two-dimensional space by a second encoder of the second generation model to generate a third reconstruction feature of an image in the two-dimensional space of the any frame, the two-dimensional space being constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension; performing fusion processing on the third reconstruction features corresponding to each frame to generate a first projection feature; and decoding the first projection feature by a first decoder of the first generation model to generate a third four-dimensional feature in the four-dimensional space, the third four-dimensional feature carrying the second reconstruction feature corresponding to each frame.

[0007] Thus, the training of the first encoder of the first generation model can be supervised by the second encoder of the second generation model, and the first encoder can learn the latent space feature coding capability of the second encoder during the training, so that the trained first encoder and the second encoder keep consistent in the reconstruction of the latent space feature of the same frame, the coding accuracy of the latent space feature of the first encoder is significantly improved, and the reconstruction accuracy of the four-dimensional feature of the first generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0008] In some possible implementation manners, the generating the second reconstruction feature of each frame in the three-dimensional space by the first generation model and the second generation model comprises: projecting the first sample feature corresponding to any frame to a two-dimensional space by a second encoder of the second generation model to generate a third reconstruction feature of an image in the two-dimensional space of the any frame, the two-dimensional space being constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension; performing fusion processing on the third reconstruction features corresponding to each frame to generate a first projection feature; and decoding the first projection feature by a first decoder of the first generation model to generate a third four-dimensional feature in the four-dimensional space, the third four-dimensional feature carrying the second reconstruction feature corresponding to each frame.

[0009] Thus, the training of the first encoder of the first generation model can be supervised by the second encoder of the second generation model, and the first encoder can learn the latent space feature coding capability of the second encoder during the training, so that the trained first encoder and the second encoder keep consistent in the reconstruction of the latent space feature of the same frame, the coding accuracy of the latent space feature of the first encoder is significantly improved, and the reconstruction accuracy of the four-dimensional feature of the first generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0010] In some possible implementation manners, the method further includes: determining initial parameters of the first generation model based on at least one of model parameters of the trained second generation model and model parameters of a third generation model; and wherein the third generation model is used to reconstruct a fourth four-dimensional feature in the four-dimensional space, and the fourth four-dimensional feature carries fifth reconstruction features of voxels in the three-dimensional space of N frames, N being an integer greater than 1 and less than M.

[0011] In this way, the first generation model can be trained based on the trained second generation model and / or the third generation model with a relatively small number of parallel processing frames, so as to avoid repeated training of the generation model, help to accelerate the convergence speed of the first generation model, and further improve the training efficiency of the first generation model.

[0012] In some possible implementation manners, the method further includes: mapping second sample features of K frames of voxels in the three-dimensional space to the four-dimensional space to generate a fifth four-dimensional feature, K being an integer greater than M; processing the fifth four-dimensional feature by a fourth generation model to generate a sixth four-dimensional feature in the four-dimensional space, the sixth four-dimensional feature carrying sixth reconstruction features of voxels in the three-dimensional space of the K frames; generating, by the first generation model and the fourth generation model, a seventh four-dimensional feature in the four-dimensional space, the seventh four-dimensional feature carrying seventh reconstruction features of voxels in the three-dimensional space of each frame; and training the fourth generation model based on the second sample features, the sixth reconstruction features, and the seventh reconstruction features.

[0013] In this way, the first generation model can be used to supervise the training of the fourth generation model, and the fourth generation model can learn the four-dimensional feature reconstruction capability of the first generation model in the training process, so that the trained fourth generation model is consistent with the first generation model in reconstructing the three-dimensional features of the same frame, and the four-dimensional feature reconstruction accuracy of the fourth generation model is significantly improved, that is, the reconstruction accuracy of dynamic scene information of the fourth generation model is significantly improved.

[0014] In some possible implementation manners, the generating, by the first generation model and the fourth generation model, of the seventh four-dimensional feature in the four-dimensional space includes: projecting, by a first encoder of the first generation model, the fifth four-dimensional feature to a two-dimensional space to generate a third projection feature, the two-dimensional space being constructed based on at least one of any two spatial dimensions or any spatial dimension and a time dimension; and decoding, by a fourth decoder of the fourth generation model, the third projection feature to generate the seventh four-dimensional feature.

[0015] Or, the fifth four-dimensional feature is projected to a two-dimensional space by a fourth encoder of the fourth generative model to generate a fourth projected feature; and the fourth projected feature is decoded by a first decoder of the first generative model to generate the seventh four-dimensional feature.

[0016] Thus, the training of the fourth encoder of the fourth generative model can be supervised by the first encoder of the first generative model, and the fourth encoder can learn the latent space feature encoding capability of the first encoder during the training, so that the trained first encoder and fourth encoder keep consistent with the reconstructed latent space features of the same frame, significantly improving the latent space feature encoding accuracy of the fourth encoder, and further significantly improving the four-dimensional feature reconstruction accuracy of the fourth generative model, i.e., significantly improving the reconstruction accuracy of the dynamic scene information of the fourth generative model.

[0017] Or, the training of the fourth decoder of the fourth generative model can be supervised by the first decoder of the first generative model, and the fourth decoder can learn the latent space feature decoding capability of the first decoder during the training, so that the trained fourth decoder and the first decoder keep consistent with the reconstructed three-dimensional features of the same frame, significantly improving the latent space feature decoding accuracy of the fourth decoder, and further significantly improving the four-dimensional feature reconstruction accuracy of the fourth generative model, i.e., significantly improving the reconstruction accuracy of the dynamic scene information of the fourth generative model.

[0018] According to a second aspect of the embodiments of the present disclosure, an image processing method is provided, including: performing denoising processing on a noisy image based on a preset generation condition to generate a target projected feature of an image in a two-dimensional space, the two-dimensional space being constructed based on at least one of any two spatial dimensions or any spatial dimension and a time dimension; and processing the target projected feature by a first generative model to generate a target video.

[0019] The target video is determined based on a target four-dimensional feature in a four-dimensional space, the four-dimensional space being constructed based on three spatial dimensions and a time dimension, the target four-dimensional feature carrying a target reconstruction feature of voxels in a three-dimensional space of each frame, and the target four-dimensional feature being obtained by processing the target projected feature by the first generative model.

[0020] The first generative model is trained based on a second reconstruction feature of voxels in a three-dimensional space of each frame, the second reconstruction feature being generated by the first generative model and a second generative model, and the second generative model being used to reconstruct a feature of voxels in a three-dimensional space of any frame.

[0021] The image processing method provided by the embodiment of the present disclosure can perform denoising processing on the noise image based on a preset generation condition, generate target projection features of the image in a two-dimensional space, and process the target projection features through a first generation model to generate a target video, thereby significantly improving the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improving the video reconstruction accuracy of the first generation model, i.e., significantly improving the reconstruction accuracy of dynamic scene information of the first generation model.

[0022] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target vehicle, a driving instruction for the target vehicle, road information, and obstacle information.

[0023] After the target video is generated, at least one of the following operations is performed:

[0024] The target video is used as a simulation driving video.

[0025] The simulation driving video is used to train and / or test an auxiliary driving model.

[0026] The simulation driving video is sent to at least one of a terminal device and the target vehicle, and the simulation driving video is used for visual display of at least one of the terminal device and the target vehicle.

[0027] The simulation driving video is used to train and / or test a traffic planning model.

[0028] Therefore, the noise image can be denoised based on the preset generation condition, the target projection features of the image in the two-dimensional space are generated, and the target projection features are processed through the first generation model to generate the simulation driving video, thereby significantly improving the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improving the driving video reconstruction accuracy of the first generation model, i.e., significantly improving the reconstruction accuracy of dynamic driving scene information of the first generation model.

[0029] In addition, the simulation driving video generated by the first generation model can be used to train and / or test the auxiliary driving model, which helps to reduce the collection cost of the real driving video and enhances the adaptability of the auxiliary driving model to various driving scenes, i.e., improves the robustness and reliability of the auxiliary driving model.

[0030] In addition, the simulation driving video can be sent to at least one of the terminal device and the target vehicle to display the simulation driving video on at least one of the terminal device and the target vehicle, so that the user can intuitively understand the simulation driving situation of the target vehicle, thereby optimizing the user experience.

[0031] In addition, the simulation driving video can be sent to at least one of the terminal device and the target vehicle, so that the simulation driving video is displayed on at least one of the terminal device and the target vehicle, and the user can intuitively understand the simulation driving situation of the target vehicle, thereby optimizing the user experience.

[0032] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target robot, a control instruction for the target robot, road information, and obstacle information;

[0033] After the target video is generated, at least one of the following operations is performed:

[0034] The target video is used as a simulation robot action video;

[0035] The simulation robot action video is used to train and / or test a robot control model.

[0036] The simulation robot action video is sent to at least one of the terminal device and the target robot, and the simulation robot action video is used for visual display on at least one of the terminal device and the target robot.

[0037] Therefore, the noise image can be denoised based on the preset generation condition, the target projection feature of the image in the two-dimensional space is generated, and the target projection feature is processed by the first generation model to generate the simulation robot action video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the robot action video reconstruction accuracy of the first generation model, that is, the dynamic robot scene information reconstruction accuracy of the first generation model.

[0038] In addition, the simulation robot action video generated by the first generation model can be used to train and / or test the robot control model, which helps to reduce the acquisition cost of the real robot action video and enhances the adaptability of the robot control model to various robot scenes, that is, improves the robustness and reliability of the robot control model.

[0039] In addition, the simulation robot action video can be sent to at least one of the terminal device and the target robot, so that the simulation robot action video is displayed on at least one of the terminal device and the target robot, and the user can intuitively understand the simulation action situation of the target robot, thereby optimizing the user experience.

[0040] In some possible implementation manners, the generation condition comprises at least one of simulation environment information and object behavior information of an extended reality scene;

[0041] After the target video is generated, at least one of the following operations is performed:

[0042] The target video is taken as an extended reality video.

[0043] The extended reality video is sent to an extended reality device, and the extended reality video is used for visual display by the extended reality device.

[0044] Thus, the noise image can be denoised based on the preset generation condition, the target projection feature of the image in the two-dimensional space is generated, and the target projection feature is processed by the first generation model to generate the new extended reality video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the extended reality video reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of the dynamic XR scene information of the first generation model is significantly improved.

[0045] In addition, the new extended reality video can be sent to the extended reality device to display the new extended reality video on the extended reality device.

[0046] According to a third aspect of the embodiments of the present disclosure, a training device of a generation model is provided, comprising: a mapping module configured to map first sample features of M frames of voxels in a three-dimensional space to a four-dimensional space to generate first four-dimensional features, M being an integer greater than 1, and the four-dimensional space being constructed based on three spatial dimensions and a time dimension; a first processing module configured to process the first four-dimensional features by a first generation model to generate second four-dimensional features in the four-dimensional space, the second four-dimensional features carrying first reconstruction features of the M frames of voxels in the three-dimensional space; a second processing module configured to generate second reconstruction features of voxels in the three-dimensional space for each frame by the first generation model and a second generation model, the second generation model being used to reconstruct features of voxels in the three-dimensional space for any frame; and a training module configured to train the first generation model based on the first sample features, the first reconstruction features, and the second reconstruction features.

[0047] The training device of the generation model provided by the embodiments of the present disclosure can supervise the training of the first generation model by the second generation model, and the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model in the training process, so that the trained first generation model and the second generation model keep consistent in reconstructing the three-dimensional features of the same frame, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0048] In some possible implementation manners, the second processing module is further configured to: project the first sample feature corresponding to any frame to a two-dimensional space by a second encoder of the second generation model, to generate third reconstruction features of images of the any frame in the two-dimensional space, the two-dimensional space being constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; perform fusion processing on the third reconstruction features corresponding to each frame to generate first projection features; and decode the first projection features by a first decoder of the first generation model to generate third four-dimensional features in the four-dimensional space, the third four-dimensional features carrying the second reconstruction features corresponding to each frame.

[0049] In this way, the training of the first encoder of the first generation model can be supervised by the second encoder of the second generation model, and the first encoder can learn the latent space feature coding capability of the second encoder during the training, so that the trained first encoder and the second encoder keep consistent in reconstructing the latent space features of the same frame, the coding accuracy of the latent space features of the first encoder is significantly improved, and the reconstruction accuracy of the four-dimensional features of the first generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0050] In some possible implementation manners, the second processing module is further configured to: project the first four-dimensional features to a two-dimensional space by a first encoder of the first generation model to generate second projection features, the two-dimensional space being constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; extract fourth reconstruction features of images of each frame in the two-dimensional space from the second projection features; and decode the fourth reconstruction features corresponding to any frame by a second decoder of the second generation model to generate the second reconstruction features corresponding to the any frame.

[0051] In this way, the training of the first decoder of the first generation model can be supervised by the second decoder of the second generation model, and the first decoder can learn the latent space feature coding capability of the second decoder during the training, so that the trained first decoder and the second decoder keep consistent in reconstructing the three-dimensional features of the same frame, the decoding accuracy of the latent space features of the first decoder is significantly improved, and the reconstruction accuracy of the four-dimensional features of the first generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the first generation model is significantly improved.

[0052] In some possible embodiments, the training module is further configured to: determine the initial parameters of the first generative model based on at least one of the model parameters of the trained second generative model and the model parameters of the trained third generative model; wherein the third generative model is used to reconstruct the fourth four-dimensional feature in the four-dimensional space, and the fourth four-dimensional feature carries the fifth reconstructed feature of the voxel of N frames in the three-dimensional space, where N is an integer greater than 1 and less than M.

[0053] Therefore, the first generation model can be trained based on the trained second generation model and / or the third generation model with a relatively small number of parallel processing frames, avoiding repeated training of the generation model, helping to speed up the convergence speed of the first generation model, and thus improving the training efficiency of the first generation model.

[0054] In some possible embodiments, the training module is further configured to: map the second sample feature of the voxel of the K frame in the three-dimensional space to the four-dimensional space to generate a fifth four-dimensional feature, where K is an integer greater than M; process the fifth four-dimensional feature through the fourth generation model to generate a sixth four-dimensional feature in the four-dimensional space, and the sixth four-dimensional feature carries the sixth reconstruction feature of the voxel of the K frame in the three-dimensional space; generate the seventh four-dimensional feature in the four-dimensional space through the first generation model and the fourth generation model, and the seventh four-dimensional feature carries the seventh reconstruction feature of the voxel of each frame in the three-dimensional space; and train the fourth generation model based on the second sample feature, the sixth reconstruction feature and the seventh reconstruction feature.

[0055] Therefore, the first generative model can be used to supervise the training of the fourth generative model. During the training process, the fourth generative model can learn the four-dimensional feature reconstruction capability of the first generative model, so that the trained fourth generative model and the first generative model reconstruct the three-dimensional features of the same frame consistent, which significantly improves the four-dimensional feature reconstruction accuracy of the fourth generative model, that is, significantly improves the reconstruction accuracy of the dynamic scene information of the fourth generative model.

[0056] In some possible embodiments, the training module is further configured to: project the fifth four-dimensional feature into a two-dimensional space through the first encoder of the first generation model to generate a third projection feature, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension and the time dimension; and decode the third projection feature through the fourth decoder of the fourth generation model to generate the seventh four-dimensional feature.

[0057] Alternatively, the fifth four-dimensional feature is projected into a two-dimensional space through the fourth encoder of the fourth generation model to generate a fourth projection feature; and the fourth projection feature is decoded through the first decoder of the first generation model to generate the seventh four-dimensional feature.

[0058] Thus, the first encoder of the first generation model can be used to supervise the training of the fourth encoder of the fourth generation model, and the fourth encoder can learn the latent space feature encoding capability of the first encoder during the training, so that the trained first encoder and the fourth encoder keep consistent for the reconstruction of the latent space feature of the same frame, significantly improving the latent space feature encoding accuracy of the fourth encoder, and further significantly improving the four-dimensional feature reconstruction accuracy of the fourth generation model, that is, significantly improving the reconstruction accuracy of the dynamic scene information of the fourth generation model.

[0059] And / or, the first decoder of the first generation model can be used to supervise the training of the fourth decoder of the fourth generation model, and the fourth decoder can learn the latent space feature decoding capability of the first decoder during the training, so that the trained fourth decoder and the first decoder keep consistent for the reconstruction of the three-dimensional feature of the same frame, significantly improving the latent space feature decoding accuracy of the fourth decoder, and further significantly improving the four-dimensional feature reconstruction accuracy of the fourth generation model, that is, significantly improving the reconstruction accuracy of the dynamic scene information of the fourth generation model.

[0060] According to a fourth aspect of the embodiments of the present disclosure, an image processing apparatus is provided, comprising: a first processing module configured to perform denoising processing on a noisy image based on a preset generation condition to generate target projection features of an image in a two-dimensional space, the two-dimensional space being constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension; and a second processing module configured to process the target projection features by a first generation model to generate a target video.

[0061] The target video is determined based on target four-dimensional features in a four-dimensional space, the four-dimensional space being constructed based on three spatial dimensions and a time dimension, the target four-dimensional features carrying target reconstruction features of voxels in a three-dimensional space of each frame, and the target four-dimensional features being obtained by processing the target projection features by the first generation model.

[0062] The first generation model is trained based on second reconstruction features of voxels in a three-dimensional space of each frame, the second reconstruction features being generated by the first generation model and a second generation model, and the second generation model being used to reconstruct features of voxels in a three-dimensional space of any frame.

[0063] The image processing apparatus provided by the embodiments of the present disclosure can perform denoising processing on a noisy image based on a preset generation condition to generate target projection features of an image in a two-dimensional space, and process the target projection features by a first generation model to generate a target video, significantly improving the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improving the video reconstruction accuracy of the first generation model, that is, significantly improving the reconstruction accuracy of the dynamic scene information of the first generation model.

[0064] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target vehicle, a driving instruction for the target vehicle, road information, and obstacle information;

[0065] After the target video is generated, the second processing module is further configured to perform at least one of the following operations:

[0066] displaying the target video as a simulation driving video;

[0067] training and / or testing an auxiliary driving model based on the simulation driving video;

[0068] sending the simulation driving video to at least one of a terminal device and the target vehicle, so that the simulation driving video is visually displayed on the at least one of the terminal device and the target vehicle.

[0069] training and / or testing a traffic planning model based on the simulation driving video.

[0070] Therefore, the noise image can be denoised based on the preset generation condition, the target projection feature of the image in the two-dimensional space is generated, and the target projection feature is processed by the first generation model to generate the simulation driving video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the driving video reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of the dynamic driving scene information of the first generation model.

[0071] In addition, the simulation driving video generated by the first generation model can be used to train and / or test the auxiliary driving model, which helps to reduce the collection cost of the real driving video and enhances the adaptability of the auxiliary driving model to various driving scenes, that is, improves the robustness and reliability of the auxiliary driving model.

[0072] In addition, the simulation driving video can be sent to at least one of a terminal device and a target vehicle, so that the simulation driving video is displayed on the at least one of the terminal device and the target vehicle, and thus a user can intuitively understand the simulation driving situation of the target vehicle, and the user experience is optimized.

[0073] In addition, the simulation driving video can be sent to at least one of a terminal device and a target vehicle, so that the simulation driving video is displayed on the at least one of the terminal device and the target vehicle, and thus a user can intuitively understand the simulation driving situation of the target vehicle, and the user experience is optimized.

[0074] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target robot, a control instruction for the target robot, road information, and obstacle information.

[0075] After the target video is generated, the second processing module is further configured to perform at least one of the following operations:

[0076] taking the target video as a simulated robot action video;

[0077] training and / or testing a robot control model based on the simulated robot action video;

[0078] sending the simulated robot action video to at least one of a terminal device and the target robot, so that the at least one of the terminal device and the target robot visually displays the simulated robot action video.

[0079] Therefore, the noise image can be denoised based on the preset generation condition, the target projection feature of the image in the two-dimensional space is generated, and the target projection feature is processed by the first generation model to generate the simulated robot action video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the robot action video reconstruction accuracy of the first generation model, that is, the dynamic robot scene information reconstruction accuracy of the first generation model.

[0080] In addition, the simulated robot action video generated by the first generation model can be used to train and / or test the robot control model, which helps to reduce the acquisition cost of the real robot action video and enhances the adaptability of the robot control model to various robot scenes, that is, improves the robustness and reliability of the robot control model.

[0081] In addition, the simulated robot action video can be sent to at least one of a terminal device and the target robot to display the simulated robot action video on the at least one of the terminal device and the target robot, so that the user can intuitively understand the simulated action of the target robot, and the user experience is optimized.

[0082] In some possible implementation manners, the generation condition comprises at least one of simulated environment information and object behavior information of an extended reality scene.

[0083] After the target video is generated, the second processing module is further configured to perform at least one of the following operations:

[0084] taking the target video as an extended reality video;

[0085] transmitting the extended reality video to an extended reality device for visualization by the extended reality device.

[0086] Thus, the noise image can be denoised based on the preset generation condition to generate the target projection feature of the image in the two-dimensional space, and the target projection feature is processed by the first generation model to generate the new extended reality video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the extended reality video reconstruction accuracy of the first generation model, i.e., the reconstruction accuracy of the dynamic XR scene information of the first generation model.

[0087] In addition, the new extended reality video can be transmitted to the extended reality device to display the new extended reality video on the extended reality device.

[0088] According to a fifth aspect of the embodiments of the present disclosure, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the program, the steps of the training method of the generation model according to the first aspect of the present disclosure are implemented, and / or the steps of the image processing method according to the second aspect of the present disclosure are implemented.

[0089] According to a sixth aspect of the embodiments of the present disclosure, a computer readable storage medium is provided, which stores computer program instructions, and when the processor executes the program instructions, the steps of the training method of the generation model according to the first aspect of the present disclosure are implemented, and / or the steps of the image processing method according to the second aspect of the present disclosure are implemented.

[0090] According to a seventh aspect of the embodiments of the present disclosure, a chip is provided, which includes an interface circuit and a processing circuit coupled with each other, the interface circuit is configured to input or output signals, and the processing circuit is configured to implement the steps of the training method of the generation model according to the first aspect of the present disclosure, and / or implement the steps of the image processing method according to the second aspect of the present disclosure.

[0091] According to an eighth aspect of the embodiments of the present disclosure, a computer program product is provided, which includes a computer program, and when the processor executes the computer program, the steps of the training method of the generation model according to the first aspect of the present disclosure are implemented, and / or the steps of the image processing method according to the second aspect of the present disclosure are implemented.

[0092] The technical scheme provided by the embodiments of the present disclosure at least brings the following beneficial effects: the first sample feature of a voxel in a three-dimensional space of M frames is mapped to a four-dimensional space to generate a first four-dimensional feature, M is an integer greater than 1, the four-dimensional space is constructed based on three spatial dimensions and a time dimension, a second four-dimensional feature in the four-dimensional space is generated by processing the first four-dimensional feature through a first generation model, the second four-dimensional feature carries a first reconstructed feature of the voxel in the three-dimensional space of the M frames, the second reconstructed feature of the voxel in the three-dimensional space of each frame is generated through the first generation model and a second generation model, the second generation model is used to reconstruct the feature of the voxel in the three-dimensional space of any frame, and the first generation model is trained based on the first sample feature, the first reconstructed feature and the second reconstructed feature. Thus, the second generation model can be used to supervise the training of the first generation model, the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model in the training process, the trained first generation model and the second generation model keep consistent with the reconstructed three-dimensional features of the same frame, and the four-dimensional feature reconstruction accuracy of the first generation model, i.e., the reconstruction accuracy of dynamic scene information of the first generation model, is significantly improved.

[0093] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS

[0094] The accompanying drawings incorporated in and forming a part of the specification illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the disclosure without imposing undue limitations on the disclosure.

[0095] Figure 1 FIG. 1 is a flowchart of a training method of a generation model according to an example embodiment.

[0096] Figure 2 FIG. 2 is a flowchart of a training method of a generation model according to another example embodiment.

[0097] Figure 3 FIG. 3 is a flowchart of a training method of a generation model according to another example embodiment.

[0098] Figure 4 FIG. 4 is a flowchart of an image processing method according to an example embodiment.

[0099] Figure 5 FIG. 5 is a schematic diagram of a video generation model according to an example embodiment.

[0100] Figure 6 FIG. 6 is a structural schematic diagram of a training device of a generation model according to an example embodiment.

[0101] Figure 7 is a structural schematic diagram of an image processing device according to an exemplary embodiment.

[0102] Figure 8 is a structural schematic diagram of an electronic device according to an exemplary embodiment.

[0103] Figure 9 is a structural schematic diagram of a chip according to an exemplary embodiment. DETAILED DESCRIPTION

[0104] In order for those skilled in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings.

[0105] It should be noted that the terms "first", "second" and the like in the specification and claims of the present disclosure and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all the embodiments consistent with the present disclosure. Rather, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0106] The training method of a generation model, the image processing method, the device, the electronic equipment, the chip, the storage medium and the program product of the embodiments of the present disclosure will be described below with reference to the drawings.

[0107] Figure 1 is a flowchart of a training method of a generation model according to an exemplary embodiment, as shown in Figure 1 The training method of a generation model of the embodiments of the present disclosure includes the following steps.

[0108] S101, map the first sample feature of M frames of voxels in a three-dimensional space to a four-dimensional space to generate a first four-dimensional feature, M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension.

[0109] It should be noted that the execution subject of the training method of the generation model of the embodiments of the present disclosure is an electronic device, such as a vehicle terminal, a vehicle controller, a server, an ISP (Image Signal Processor, image processing chip), etc. The training method of the generation model of the embodiments of the present disclosure can be executed by the training device of the generation model of the embodiments of the present disclosure, and the training device of the generation model of the embodiments of the present disclosure can be configured in any electronic device to execute the training method of the generation model of the embodiments of the present disclosure.

[0110] A voxel is a short form of Volume Pixel, which is a basic unit of discretization of information in three-dimensional scale, representing a volume unit of information in three-dimensional space. Three-dimensional space is constructed based on three spatial dimensions, which are not limited too much, such as including x, y, and z. For example, four-dimensional space is constructed based on three spatial dimensions x, y, z and time dimension t. The first reconstruction feature of a voxel in three-dimensional space of any frame can be referred to as a three-dimensional feature, and the first four-dimensional feature carries the first sample feature of the voxel in three-dimensional space of M frames.

[0111] Any type of sample feature is not limited too much, such as including semantic information, depth information, position information, semantic information at a certain depth information, etc. of each voxel.

[0112] Any type of four-dimensional feature is not limited too much, such as including semantic information, depth information, position information, semantic information at a certain depth information, etc. of each voxel in three-dimensional space of each frame.

[0113] The first sample feature of the voxel in three-dimensional space of M frames is mapped to four-dimensional space to generate the first four-dimensional feature, which can be realized by any feature mapping method in related technologies, which is not limited too much here.

[0114] Optionally, the first sample feature of the voxel in three-dimensional space of M frames is mapped to four-dimensional space to generate the first four-dimensional feature, which includes the first sample feature of the jth voxel in three-dimensional space of the ith frame as the first mapping feature of the four-dimensional subspace determined by the ith frame and the jth voxel, and the first four-dimensional feature is generated based on the first mapping feature of each four-dimensional subspace, i is a positive integer not greater than M, j is a positive integer not greater than Q, and Q is the number of voxels.

[0115] For example, four-dimensional space is constructed based on three spatial dimensions x, y, z and time dimension t, i.e. the coordinates of the four-dimensional subspace are (x, y, z, t), if the coordinates of the jth voxel in three-dimensional space of the ith frame are (x ij ,y ij ,z ij ), then the coordinates of the four-dimensional subspace determined by the ith frame and the jth voxel are (x ij ,y ij ,z ij ,i), and the first sample feature of the jth voxel in three-dimensional space of the ith frame can be taken as the first mapping feature of the four-dimensional subspace (x ij ,y ij ,z ij ,i), and the first four-dimensional feature is generated based on the first mapping feature of each four-dimensional subspace.

[0116] S102, processing the first four-dimensional feature through the first generative model to generate a second four-dimensional feature in a four-dimensional space, the second four-dimensional feature carrying the first reconstruction feature of the voxels of the M frames in a three-dimensional space.

[0117] It should be noted that the first generative model is used to reconstruct the second four-dimensional feature and also used to reconstruct the M frames of three-dimensional features. Any type of generative model is not limited, such as diffusion model, GAN (Generative Adversarial Networks), VAE (Variational Autoencoders), etc. It should be noted that the VAE includes an encoder and a decoder.

[0118] Optionally, the first generative model includes a first encoder and a first decoder. For example, any type of encoder is a VAE encoder, and any type of decoder is a VAE decoder.

[0119] Optionally, processing the first four-dimensional feature through the first generative model to generate a second four-dimensional feature in a four-dimensional space includes projecting the first four-dimensional feature to a two-dimensional space through a first encoder of the first generative model to generate a second projection feature, the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension, and decoding the second projection feature through a first decoder of the first generative model to generate the second four-dimensional feature.

[0120] In this embodiment, the latent space refers to a two-dimensional space. The latent space is also known as the hidden space.

[0121] It should be noted that the number of two-dimensional spaces is not limited, such as there can be 6 two-dimensional spaces, two-dimensional space 1 is constructed based on spatial dimensions x and y, two-dimensional space 2 is constructed based on spatial dimensions x and z, two-dimensional space 3 is constructed based on spatial dimensions y and z, two-dimensional space 4 is constructed based on spatial dimension x and time dimension t, two-dimensional space 5 is constructed based on spatial dimension y and time dimension t, and two-dimensional space 6 is constructed based on spatial dimension z and time dimension t. Thus, the four-dimensional feature can be compressed into 6 two-dimensional spaces, which can highly compress the four-dimensional feature while preserving geometric details, helping to improve scene reconstruction accuracy and greatly reduce memory and computing overhead.

[0122] Projecting the first four-dimensional feature to a two-dimensional space to generate a second projection feature can be implemented by any feature projection method in related technologies, which is not limited here.

[0123] S103, generating, by the first generation model and the second generation model, a second reconstruction feature of voxels of each frame in three-dimensional space, the second generation model being used for reconstructing the feature of voxels of any frame in three-dimensional space.

[0124] It should be noted that the second generation model is used for processing the original feature of voxels of any frame in three-dimensional space to generate the reconstructed feature of voxels of any frame in three-dimensional space, that is, the second generation model is used for generating the reconstructed three-dimensional feature of a frame based on the original three-dimensional feature of the frame, that is, for reconstructing the three-dimensional feature of a single frame. The training content of the second generation model can be implemented by using any training method of a generation model for reconstructing the three-dimensional feature of a single frame in related technologies, which is not limited here.

[0125] At present, the cost of real scene information collection is high, the scene coverage is limited, and the reconstruction accuracy of the generation model in related technologies for a period of scene information is low, which is difficult to meet the diversity and complexity requirements of scene information.

[0126] The generation of a period of scene information is also called dynamic scene generation, and the scene information includes three-dimensional features, four-dimensional features, images, videos, etc.

[0127] In the present disclosure, the second reconstruction feature is generated by the first generation model and the trained second generation model, that is, the four-dimensional feature reconstruction capability of the first generation model and the single-frame three-dimensional feature reconstruction capability of the second generation model are used together to generate the second reconstruction feature to train the first generation model. The second generation model has good single-frame three-dimensional feature reconstruction capability, so that the training of the first generation model can be supervised by the second generation model. In the training process, the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model, so that the trained first generation model and the second generation model keep consistent in reconstructing the three-dimensional feature of the same frame, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of the dynamic scene information of the first generation model.

[0128] In addition, the four-dimensional features in multiple scenes can be reconstructed by the first generation model, which can meet the diversity and complexity requirements of scene information.

[0129] In addition, using the second generation model to supervise the training of the first generation model also helps to speed up the convergence speed of the first generation model, and thus improves the training efficiency of the first generation model.

[0130] In addition, no matter how many frames (such as 16 frames, 32 frames) of three-dimensional features the first generation model is used to reconstruct, the training of the first generation model can be supervised by the second generation model, so that the trained first generation models of different frame numbers keep consistent with the three-dimensional features reconstructed by the second generation model for the same frame, which helps to improve the training stability of the first generation models of different frame numbers.

[0131] Optionally, the second reconstruction features of the voxels of each frame in the three-dimensional space are generated through the first generation model and the second generation model, including processing the first sample features corresponding to any frame through a second subnetwork of the second generation model to generate a first intermediate processing result corresponding to any frame, performing fusion processing on the first intermediate processing results corresponding to each frame to generate a first total intermediate processing result, and processing the first fusion result through a first subnetwork of the first generation model to generate a third four-dimensional feature in the four-dimensional space, the third four-dimensional feature carrying the second reconstruction features corresponding to each frame.

[0132] Optionally, the second reconstruction features of the voxels of each frame in the three-dimensional space are generated through the first generation model and the second generation model, including processing the first four-dimensional feature through a first subnetwork of the first generation model to generate a second total intermediate processing result, extracting the second intermediate processing result corresponding to each frame from the second total intermediate processing result, processing the second intermediate processing result corresponding to any frame through a second subnetwork of the second generation model to generate the second reconstruction feature corresponding to any frame.

[0133] S104, training the first generation model based on the first sample features, the first reconstruction features, and the second reconstruction features.

[0134] It should be noted that the first generation model is trained based on the first sample features, the first reconstruction features, and the second reconstruction features, and any model training method in the related art can be used to achieve this, which is not limited here.

[0135] For example, taking M=16 as an example, the first sample features corresponding to the first 16 frames are mapped to the four-dimensional space to generate the first four-dimensional feature.

[0136] The first four-dimensional feature is projected to the two-dimensional space 1 through the first encoder of the first generation model to generate the second projection feature in the two-dimensional space 1. The first four-dimensional feature is projected to the two-dimensional space 2 through the first encoder to generate the second projection feature in the two-dimensional space 2. The first four-dimensional feature is projected to the two-dimensional space 3 through the first encoder to generate the second projection feature in the two-dimensional space 3. The first four-dimensional feature is projected to the two-dimensional space 4 through the first encoder to generate the second projection feature in the two-dimensional space 4. The first four-dimensional feature is projected to the two-dimensional space 5 through the first encoder to generate the second projection feature in the two-dimensional space 5. The first four-dimensional feature is projected to the two-dimensional space 6 through the first encoder to generate the second projection feature in the two-dimensional space 6.

[0137] The second projection features in the two-dimensional spaces 1 to 6 are decoded through the first decoder of the first generation model to generate the second four-dimensional feature in the four-dimensional space, the second four-dimensional feature carrying the first reconstruction features corresponding to the first 16 frames.

[0138] The second reconstruction features corresponding to the first 16 frames are generated through the first generation model and the second generation model.

[0139] The first generation model is trained based on the first sample features corresponding to the first 16 frames, the first reconstruction features corresponding to the first 16 frames, and the second reconstruction features corresponding to the first 16 frames.

[0140] Optionally, the training of the first generation model based on the first sample features, the first reconstruction features, and the second reconstruction features includes determining a first loss function based on the first sample features and the first reconstruction features, determining a second loss function based on the first sample features and the second reconstruction features, and training the first generation model based on the first loss function and the second loss function. In this way, the first generation model participates in the generation of the first reconstruction features and the second reconstruction features, and can learn the correlation between the first sample features and the first reconstruction features and the correlation between the first sample features and the second reconstruction features in the training process. Thus, the trained first generation model can generate reconstructed four-dimensional features based on the original four-dimensional features, and the accuracy of any frame of three-dimensional features in the reconstructed four-dimensional features is high, that is, the four-dimensional feature reconstruction accuracy of the first generation model is significantly improved.

[0141] Optionally, the first loss function is determined based on the first sample features and the first reconstruction features, including determining a first difference between the first sample features corresponding to any frame and the first reconstruction features corresponding to any frame, and determining the first loss function based on the first differences corresponding to the frames. Alternatively, the first loss function is determined based on the difference between the first four-dimensional features and the second four-dimensional features.

[0142] Optionally, the second loss function is determined based on the first sample features and the second reconstruction features, including determining a second difference between the first sample features corresponding to any frame and the second reconstruction features corresponding to any frame, and determining the second loss function based on the second differences corresponding to the frames. Alternatively, the second reconstruction features corresponding to the frames are mapped to a four-dimensional space to generate eighth four-dimensional features, and the second loss function is determined based on the difference between the first four-dimensional features and the eighth four-dimensional features.

[0143] Optionally, the first generation model is trained based on the first loss function and the second loss function, including determining a first total loss function based on the first loss function and the second loss function, and training the first generation model based on the first total loss function. For example, the first total loss function can be determined by weighted summation of the first loss function and the second loss function.

[0144] For example, the formula of the first total loss function is as follows:

[0145] L1=L AE(X1,X2)+βL KL (X1,X2)+αL reg (X1,X8)

[0146] Among them, X1 is the first four-dimensional feature, X2 is the second four-dimensional feature, X8 is the eighth four-dimensional feature, L1 is the first total loss function, L AE (X1,X2),L KL (X1,X2) are different first loss functions, L reg (X1,X8) is the second loss function, L AE (·) is the reconstruction loss function, L KL (·) is the KL (Kullback-Leibler) divergence loss function, and α and β are coefficients.

[0147] For example, L AE (·), L reg (·) Includes the reconstruction cross entropy loss function and the Lovász-Softmax loss function. The Lovász-Softmax loss function is a loss function used to optimize deep learning models, especially for segmentation tasks. It aims to improve the performance of image segmentation models by directly optimizing the mean intersection over union (mIoU).

[0148] The first generation model in this disclosure can be applied to generating dynamic scene information such as driving scenarios, XR (Extended Reality) scenarios, robotics scenarios, film and television scenarios, gaming scenarios, surveillance scenarios, and traffic planning. XR scenarios can include AR (Augmented Reality) and VR (Virtual Reality) scenarios.

[0149] In the first case, taking the driving scene as an example, a sample driving video can be determined, and features can be extracted from any frame of the sample driving image in the sample driving video to determine the sample features of the image of any frame in two-dimensional space. Based on the sample features of the image of any frame in two-dimensional space, the first sample features corresponding to any frame can be determined.

[0150] Therefore, in this embodiment, the sample driving video can be used to determine the first sample features, and then the trained first generation model can be used to reconstruct the four-dimensional features of the driving scene, which significantly improves the accuracy of the first generation model in reconstructing the four-dimensional features of the driving scene.

[0151] In the second case, taking the robot scene as an example, a sample robot action video can be determined, feature extraction is performed on any frame sample robot action image in the sample robot action video, sample features of the image of any frame in two-dimensional space are determined, and the first sample features corresponding to any frame are determined based on the sample features of the image of any frame in two-dimensional space.

[0152] Thus, in the embodiment, the first sample features can be determined by using the sample robot action video, and the trained first generation model can be used to reconstruct the four-dimensional features of the robot scene, thereby significantly improving the four-dimensional feature reconstruction accuracy of the first generation model in the robot scene.

[0153] In the third case, taking the XR scene as an example, a sample extended reality video can be determined, feature extraction is performed on any frame sample extended reality image in the sample extended reality video, sample features of the image of any frame in two-dimensional space are determined, and the first sample features corresponding to any frame are determined based on the sample features of the image of any frame in two-dimensional space.

[0154] Thus, in the embodiment, the first sample features can be determined by using the sample extended reality video, and the trained first generation model can be used to reconstruct the four-dimensional features of the XR scene, thereby significantly improving the four-dimensional feature reconstruction accuracy of the first generation model in the XR scene.

[0155] It should be noted that the execution timing of steps S101 to S104 is not limited in the present disclosure, Figure 1 only taking the sequential execution of steps S101 to S104 as an example.

[0156] The training method of the generation model provided by the embodiment of the present disclosure maps the first sample features of M frames of voxels in three-dimensional space to four-dimensional space to generate first four-dimensional features, M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension. The first generation model processes the first four-dimensional features to generate second four-dimensional features in the four-dimensional space, the second four-dimensional features carry the first reconstruction features of M frames of voxels in three-dimensional space, the second generation model is used to reconstruct the features of voxels in three-dimensional space of any frame, and the first generation model is trained based on the first sample features, the first reconstruction features and the second reconstruction features. Thus, the second generation model can be used to supervise the training of the first generation model, and the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model during the training process, so that the trained first generation model and the second generation model maintain consistent three-dimensional features reconstructed for the same frame, thereby significantly improving the four-dimensional feature reconstruction accuracy of the first generation model, i.e., significantly improving the reconstruction accuracy of dynamic scene information of the first generation model.

[0157] Figure 2is a flowchart of a training method of a generation model according to another exemplary embodiment, as shown in Figure 2 The training method of the generation model of the embodiment of the present disclosure comprises the following steps.

[0158] In S201, the first sample features of voxels in M frames in a three-dimensional space are mapped to a four-dimensional space to generate first four-dimensional features, M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension.

[0159] In S202, the first four-dimensional features are processed by a first generation model to generate second four-dimensional features in the four-dimensional space, and the second four-dimensional features carry the first reconstructed features of voxels in M frames in the three-dimensional space.

[0160] The related content of steps S201-S202 can be referred to the above-mentioned embodiments, which will not be repeated here.

[0161] In S203, the first sample features corresponding to any frame are projected to a two-dimensional space by a second encoder of a second generation model to generate third reconstructed features of images of any frame in the two-dimensional space, and the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension.

[0162] In the embodiment, the first generation model comprises a first encoder and a first decoder, and the second generation model comprises a second encoder and a second decoder. The first sample features corresponding to any frame are projected to a two-dimensional space by a second encoder of a second generation model to generate third reconstructed features of images of any frame in the two-dimensional space, which can be realized by any feature projection method in the related art, and will not be limited here.

[0163] The first four-dimensional features are processed by the first generation model to generate second four-dimensional features in the four-dimensional space, which comprises projecting the first four-dimensional features to a two-dimensional space by a first encoder of the first generation model to generate second projection features, the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension, and the second projection features are decoded by a first decoder of the first generation model to generate the second four-dimensional features.

[0164] It should be noted that the related content of generating the second four-dimensional features can be referred to the above-mentioned embodiments, which will not be repeated here.

[0165] In S204, the third reconstructed features corresponding to each frame are fused to generate first projection features.

[0166] It should be noted that the third reconstructed features corresponding to each frame are fused to generate first projection features, which can be realized by any feature fusion method in the related art, and will not be limited here.

[0167] For example, continue to take the two-dimensional space 1 as an example, the two-dimensional space 1 is constructed based on the spatial dimensions x, y, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 1 are (x, y), the third reconstruction features of each frame in the two-dimensional subspace (x, y) can be fused to generate the fusion features of the two-dimensional subspace (x, y), and the first projection features under the two-dimensional space 1 are generated based on the fusion features of each two-dimensional subspace (x, y).

[0168] Continue to take the two-dimensional space 2 as an example, the two-dimensional space 2 is constructed based on the spatial dimensions x, z, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 2 are (x, z), the third reconstruction features of each frame in the two-dimensional subspace (x, z) can be fused to generate the fusion features of the two-dimensional subspace (x, z), and the first projection features under the two-dimensional space 2 are generated based on the fusion features of each two-dimensional subspace (x, z).

[0169] Continue to take the two-dimensional space 3 as an example, the two-dimensional space 2 is constructed based on the spatial dimensions y, z, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 3 are (y, z), the third reconstruction features of each frame in the two-dimensional subspace (y, z) can be fused to generate the fusion features of the two-dimensional subspace (y, z), and the first projection features under the two-dimensional space 3 are generated based on the fusion features of each two-dimensional subspace (y, z).

[0170] Continue to take the two-dimensional space 4 as an example, the two-dimensional space 4 is constructed based on the spatial dimension x and the time dimension t, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 4 are (x, t), the third reconstruction features of the i-th frame in the two-dimensional subspace (x, i) can be taken as the mapping features of the two-dimensional subspace (x, i), and the first projection features under the two-dimensional space 4 are generated based on the mapping features of each two-dimensional subspace (x, i).

[0171] Continue to take the two-dimensional space 5 as an example, the two-dimensional space 5 is constructed based on the spatial dimension y and the time dimension t, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 5 are (y, t), the third reconstruction features of the i-th frame in the two-dimensional subspace (y, i) can be taken as the mapping features of the two-dimensional subspace (y, i), and the first projection features under the two-dimensional space 5 are generated based on the mapping features of each two-dimensional subspace (y, t).

[0172] Continue to take the two-dimensional space 6 as an example, the two-dimensional space 6 is constructed based on the spatial dimension z and the time dimension t, that is, the coordinates of the two-dimensional subspace of the two-dimensional space 6 are (z, t), the third reconstruction features of the i-th frame in the two-dimensional subspace (z, i) can be taken as the mapping features of the two-dimensional subspace (z, i), and the first projection features under the two-dimensional space 6 are generated based on the mapping features of each two-dimensional subspace (z, t).

[0173] S205, decoding the first projected feature through a first decoder of the first generative model to generate a third four-dimensional feature in the four-dimensional space, the third four-dimensional feature carrying the second reconstructed feature corresponding to each frame.

[0174] In this embodiment, the second encoder of the second generative model is configured to project the first sample feature corresponding to a frame to a two-dimensional space to generate a third reconstructed feature corresponding to the frame.

[0175] The second reconstructed feature is generated by the second encoder of the second generative model and the first decoder of the first generative model, that is, the second reconstructed feature is generated by the latent space feature encoding capability of the second encoder and the latent space feature decoding capability of the first decoder, so as to train the first generative model. The second encoder has good latent space feature encoding capability, so that the training of the first encoder of the first generative model can be supervised by the second encoder. In the training process, the first encoder can learn the latent space feature encoding capability of the second encoder, so that the trained first encoder and the second encoder keep consistent for the reconstructed latent space feature of the same frame, which significantly improves the latent space feature encoding precision of the first encoder, and further significantly improves the four-dimensional feature reconstruction precision of the first generative model, that is, the reconstruction precision of the dynamic scene information of the first generative model is significantly improved.

[0176] For example, taking M=16 as an example, the first sample feature corresponding to the first frame is projected to a two-dimensional space 1 by the second encoder of the second generative model to generate a third reconstructed feature of the first frame in the two-dimensional space 1. The first sample feature corresponding to the first frame is projected to a two-dimensional space 2 by the second encoder to generate a third reconstructed feature of the first frame in the two-dimensional space 2. The first sample feature corresponding to the first frame is projected to a two-dimensional space 3 by the second encoder to generate a third reconstructed feature of the first frame in the two-dimensional space 3. The first sample feature corresponding to the first frame is projected to a two-dimensional space 4 by the second encoder to generate a third reconstructed feature of the first frame in the two-dimensional space 4. The first sample feature corresponding to the first frame is projected to a two-dimensional space 5 by the second encoder to generate a third reconstructed feature of the first frame in the two-dimensional space 5. The first sample feature corresponding to the first frame is projected to a two-dimensional space 6 by the second encoder to generate a third reconstructed feature of the first frame in the two-dimensional space 6.

[0177] The related content of the third reconstructed features of the second to sixteenth frames in the two-dimensional spaces 1 to 6 can refer to the related content of the third reconstructed features of the first frame in the two-dimensional spaces 1 to 6, which will not be described here.

[0178] The third reconstruction features of the 1st to 16th frames in the two-dimensional space 1 are fused to generate the first projection features in the two-dimensional space 1. The third reconstruction features of the 1st to 16th frames in the two-dimensional space 2 are fused to generate the first projection features in the two-dimensional space 2. The third reconstruction features of the 1st to 16th frames in the two-dimensional space 3 are fused to generate the first projection features in the two-dimensional space 3. The third reconstruction features of the 1st to 16th frames in the two-dimensional space 4 are fused to generate the first projection features in the two-dimensional space 4. The third reconstruction features of the 1st to 16th frames in the two-dimensional space 5 are fused to generate the first projection features in the two-dimensional space 5. The third reconstruction features of the 1st to 16th frames in the two-dimensional space 6 are fused to generate the first projection features in the two-dimensional space 6.

[0179] The first projection features in the two-dimensional spaces 1 to 6 are decoded by the first decoder of the first generation model to generate third four-dimensional features in a four-dimensional space, the third four-dimensional features carrying the second reconstruction features corresponding to the 1st to 16th frames.

[0180] In S206, the first four-dimensional features are projected to a two-dimensional space by the first encoder of the first generation model to generate second projection features, the two-dimensional space being constructed based on at least one of the following manners: any two spatial dimensions, any spatial dimension and the time dimension.

[0181] The related content of S206 can be referred to the above-mentioned embodiments, which will not be repeated here.

[0182] In S207, the fourth reconstruction features of the images of each frame in the two-dimensional space are extracted from the second projection features.

[0183] In S208, the second reconstruction features corresponding to any frame are generated by decoding the fourth reconstruction features corresponding to the frame by the second decoder of the second generation model.

[0184] In the embodiment, the second projection features carry the fourth reconstruction features of the images of each frame in the two-dimensional space, and the second decoder of the second generation model is used to decode the fourth reconstruction features corresponding to a certain frame to generate the second reconstruction features corresponding to the frame.

[0185] The second reconstruction feature is generated by the first encoder of the first generation model and the second decoder of the second generation model in cooperation, that is, the latent space feature encoding capability of the first encoder and the latent space feature decoding capability of the second decoder are utilized to jointly generate the second reconstruction feature to train the first generation model. The second decoder has good latent space feature decoding capability, so that the training of the first decoder of the first generation model can be supervised by the second decoder of the second generation model. During the training process, the first decoder can learn the latent space feature decoding capability of the second decoder, so that the trained first decoder and the second decoder keep consistent for the three-dimensional features reconstructed for the same frame, significantly improving the latent space feature decoding precision of the first decoder, and further significantly improving the four-dimensional feature reconstruction precision of the first generation model, that is, the reconstruction precision of the dynamic scene information of the first generation model.

[0186] For example, taking M=16 as an example, the first four-dimensional feature is projected into a two-dimensional space 1 by the first encoder of the first generation model to generate a second projection feature under the two-dimensional space 1. The first four-dimensional feature is projected into a two-dimensional space 2 by the first encoder to generate a second projection feature under the two-dimensional space 2. The first four-dimensional feature is projected into a two-dimensional space 3 by the first encoder to generate a second projection feature under the two-dimensional space 3. The first four-dimensional feature is projected into a two-dimensional space 4 by the first encoder to generate a second projection feature under the two-dimensional space 4. The first four-dimensional feature is projected into a two-dimensional space 5 by the first encoder to generate a second projection feature under the two-dimensional space 5. The first four-dimensional feature is projected into a two-dimensional space 6 by the first encoder to generate a second projection feature under the two-dimensional space 6.

[0187] The fourth reconstruction features 1 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 1. The fourth reconstruction features 2 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 2. The fourth reconstruction features 3 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 3. The fourth reconstruction features 4 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 4. The fourth reconstruction features 5 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 5. The fourth reconstruction features 6 corresponding to the 1st to 16th frames are extracted from the second projection features under the two-dimensional space 6.

[0188] The fourth reconstruction features 1 to 6 corresponding to the 1st frame are decoded by the second decoder of the second generation model to generate the second reconstruction feature corresponding to the 1st frame.

[0189] The fourth reconstruction features 1 to 6 corresponding to the 2nd frame are decoded by the second decoder to generate the second reconstruction feature corresponding to the 2nd frame.

[0190] The fourth reconstructed features 1-6 corresponding to the third frame are decoded by the second decoder to generate second reconstructed features corresponding to the third frame.

[0191] The related content of the second reconstructed features corresponding to the fourth-sixteenth frames can refer to the related content of the second reconstructed features corresponding to the first-three frames, which will not be repeated here.

[0192] In S209, the first generation model is trained based on the first sample features, the first reconstructed features, and the second reconstructed features.

[0193] The related content of S209 can refer to the above-mentioned embodiments, which will not be repeated here.

[0194] For example, the formula of the first total loss function is as follows:

[0195]

[0196] Wherein, x 21 is the second reconstructed features generated by performing S203-S205, x 22 is the second reconstructed features generated by performing S206-S208, is a different second loss function, and α1 and α2 are coefficients.

[0197] For example, including reconstruction cross-entropy loss function, Lovász-Softmax loss function, etc.

[0198] It should be noted that the execution timing of steps S201-S209 is not limited in the present disclosure, Figure 2 only taking the sequential execution of steps S201-S209 as an example. For example, steps S201-S205 and S209 can be implemented as independent embodiments, and steps S201-S202 and S206-S209 can be implemented as independent embodiments.

[0199] The training method of the generation model provided by the embodiments of the present disclosure can supervise the training of the first encoder of the first generation model by using the second encoder of the second generation model. During the training process, the first encoder can learn the latent space feature encoding ability of the second encoder, so that the trained first encoder and the second encoder keep consistent for the latent space features reconstructed for the same frame. The latent space feature encoding precision of the first encoder is significantly improved, and the four-dimensional feature reconstruction precision of the first generation model is also significantly improved, that is, the reconstruction precision of the dynamic scene information of the first generation model is significantly improved.

[0200] And / or, the second decoder of the second generative model can be used to supervise the training of the first decoder of the first generative model, and the first decoder can learn the latent space feature encoding capability of the second decoder during the training process, so that the trained first decoder is consistent with the three-dimensional features reconstructed by the second decoder for the same frame, which significantly improves the latent space feature decoding accuracy of the first decoder, and further significantly improves the four-dimensional feature reconstruction accuracy of the first generative model, that is, the reconstruction accuracy of the dynamic scene information of the first generative model is significantly improved.

[0201] On the basis of any of the above embodiments, the method further comprises determining initial parameters of the first generative model based on at least one of the model parameters of the trained second generative model and the model parameters of the trained third generative model, wherein the third generative model is used to reconstruct a fourth four-dimensional feature in a four-dimensional space, and the fourth four-dimensional feature carries fifth reconstructed features of voxels in a three-dimensional space for N frames, N being an integer greater than 1 and less than M.

[0202] It can be understood that the first generative model and the third generative model are both used to reconstruct four-dimensional features, but the number of frames of three-dimensional features carried by the four-dimensional features reconstructed by the first generative model is greater than the number of frames of three-dimensional features carried by the four-dimensional features reconstructed by the third generative model, that is, the number of parallel processing frames of the first generative model is greater than that of the third generative model.

[0203] Therefore, the first generative model can be trained based on the trained second generative model and / or the third generative model with relatively few parallel processing frames, avoiding repeated training of the generative model, which helps to speed up the convergence speed of the first generative model and further improve the training efficiency of the first generative model.

[0204] It should be noted that the third generative model is used to reconstruct the fourth four-dimensional feature and also used to reconstruct the N three-dimensional features.

[0205] Figure 3 is a flowchart of a training method of a generative model according to another exemplary embodiment, as shown in Figure 3 The training method of the generative model of the present disclosure comprises the following steps.

[0206] S301, mapping M frames of first sample features of voxels in a three-dimensional space to a four-dimensional space to generate a first four-dimensional feature, M being an integer greater than 1, and the four-dimensional space being constructed based on three spatial dimensions and a time dimension.

[0207] S302, processing the first four-dimensional feature by a first generative model to generate a second four-dimensional feature in a four-dimensional space, the second four-dimensional feature carrying first reconstructed features of voxels in a three-dimensional space for M frames.

[0208] S303, generating, by the first generative model and the second generative model, a second reconstruction feature of voxels in the three-dimensional space for each frame, the second generative model being configured to reconstruct the feature of voxels in the three-dimensional space for any frame.

[0209] S304, training the first generative model based on the first sample feature, the first reconstruction feature, and the second reconstruction feature.

[0210] The related content of steps S301-S304 can be referred to the above embodiments, which will not be repeated here.

[0211] S305, mapping the second sample feature of voxels in the three-dimensional space for K frames to a four-dimensional space to generate a fifth four-dimensional feature, K being an integer greater than M.

[0212] It should be noted that the related content of step S305 can be referred to the related content of step S101, which will not be repeated here.

[0213] S306, processing the fifth four-dimensional feature by a fourth generative model to generate a sixth four-dimensional feature in the four-dimensional space, the sixth four-dimensional feature carrying a sixth reconstruction feature of voxels in the three-dimensional space for K frames.

[0214] It should be noted that the fourth generative model is configured to reconstruct the sixth four-dimensional feature and is also configured to reconstruct the three-dimensional feature of K frames.

[0215] Optionally, the fourth generative model comprises a fourth encoder and a fourth decoder.

[0216] Optionally, the processing of the fifth four-dimensional feature by the fourth generative model to generate the sixth four-dimensional feature in the four-dimensional space comprises projecting the fifth four-dimensional feature to a two-dimensional space by a fourth encoder of the fourth generative model to generate a fourth projection feature, and decoding the fourth projection feature by a fourth decoder of the fourth generative model to generate the sixth four-dimensional feature.

[0217] Optionally, the method further comprises determining initial parameters of the fourth generative model based on at least one of model parameters of the trained second generative model and model parameters of the trained first generative model.

[0218] It can be understood that the first generative model and the fourth generative model are both configured to reconstruct four-dimensional features, however, the number of frames of three-dimensional features carried by the four-dimensional feature reconstructed by the first generative model is smaller than the number of frames of three-dimensional features carried by the four-dimensional feature reconstructed by the fourth generative model, i.e., the number of parallel processing frames of the first generative model is smaller than the number of parallel processing frames of the fourth generative model.

[0219] Thus, the fourth generation model can be trained based on the trained second generation model and / or the first generation model with a relatively small number of processing frames, avoiding repeated training of the generation model, helping to speed up the convergence speed of the fourth generation model, and further improving the training efficiency of the fourth generation model.

[0220] In S307, a seventh four-dimensional feature in the four-dimensional space is generated by the first generation model and the fourth generation model, and the seventh four-dimensional feature carries a seventh reconstruction feature of the voxel in the three-dimensional space of each frame.

[0221] In this embodiment, the seventh reconstruction feature is generated by the trained first generation model and the fourth generation model, that is, the first generation model and the fourth generation model can be used to generate the seventh reconstruction feature together to train the fourth generation model. The first generation model has good four-dimensional feature reconstruction capability, so the first generation model can be used to supervise the training of the fourth generation model. During the training process, the fourth generation model can learn the four-dimensional feature reconstruction capability of the first generation model, so that the trained fourth generation model is consistent with the three-dimensional feature reconstructed by the first generation model for the same frame, significantly improving the four-dimensional feature reconstruction accuracy of the fourth generation model, that is, significantly improving the reconstruction accuracy of the dynamic scene information of the fourth generation model.

[0222] In addition, using the first generation model to supervise the training of the fourth generation model also helps to speed up the convergence speed of the fourth generation model, and further improves the training efficiency of the fourth generation model.

[0223] In S308, the fourth generation model is trained based on the second sample feature, the sixth reconstruction feature, and the seventh reconstruction feature.

[0224] For example, taking M=16 and K=32 as an example, the second sample features corresponding to the first to 32 frames are mapped to the four-dimensional space to generate the fifth four-dimensional feature. The fourth encoder of the fourth generation model projects the fifth four-dimensional feature to two-dimensional space 1 to generate the fourth projection feature in two-dimensional space 1. The fourth encoder projects the fifth four-dimensional feature to two-dimensional space 2 to generate the fourth projection feature in two-dimensional space 2. The fourth encoder projects the fifth four-dimensional feature to two-dimensional space 3 to generate the fourth projection feature in two-dimensional space 3. The fourth encoder projects the fifth four-dimensional feature to two-dimensional space 4 to generate the fourth projection feature in two-dimensional space 4. The fourth encoder projects the fifth four-dimensional feature to two-dimensional space 5 to generate the fourth projection feature in two-dimensional space 5. The fourth encoder projects the fifth four-dimensional feature to two-dimensional space 6 to generate the fourth projection feature in two-dimensional space 6.

[0225] The fourth projection feature under the two-dimensional space 1-6 is decoded by a fourth decoder of the fourth generation model to generate a sixth four-dimensional feature under the four-dimensional space, and the sixth four-dimensional feature carries the sixth reconstruction feature corresponding to the first 32 frames.

[0226] The seventh four-dimensional feature under the four-dimensional space is generated by the first generation model and the fourth generation model, and the seventh four-dimensional feature carries the seventh reconstruction feature corresponding to the first 16 frames.

[0227] The fourth generation model is trained based on the third sample feature corresponding to the first 32 frames, the sixth reconstruction feature corresponding to the first 32 frames, and the seventh reconstruction feature corresponding to the first 32 frames.

[0228] Optionally, the fourth generation model is trained based on the second sample feature, the sixth reconstruction feature, and the seventh reconstruction feature, including determining a third loss function based on the second sample feature and the sixth reconstruction feature, determining a fourth loss function based on the second sample feature and the seventh reconstruction feature, and training the fourth generation model based on the third loss function and the fourth loss function. In this way, the fourth generation model participates in the generation of the sixth reconstruction feature and the seventh reconstruction feature, and can learn the correlation between the second sample feature and the sixth reconstruction feature and the correlation between the second sample feature and the seventh reconstruction feature in the training process, so that the trained fourth generation model can generate a reconstructed four-dimensional feature based on the original four-dimensional feature, and the accuracy of any frame of three-dimensional feature in the reconstructed four-dimensional feature is higher, that is, the four-dimensional feature reconstruction accuracy of the fourth generation model is significantly improved.

[0229] Optionally, the third loss function is determined based on the second sample feature and the sixth reconstruction feature, including determining a third difference between the second sample feature corresponding to any frame and the sixth reconstruction feature corresponding to any frame, and determining the third loss function based on the third difference corresponding to each frame. Alternatively, the third loss function is determined based on the difference between the fifth four-dimensional feature and the sixth four-dimensional feature.

[0230] Optionally, the fourth loss function is determined based on the second sample feature and the seventh reconstruction feature, including determining a fourth difference between the second sample feature corresponding to any frame and the seventh reconstruction feature corresponding to any frame, and determining the fourth loss function based on the fourth difference corresponding to each frame. Alternatively, the fourth loss function is determined based on the difference between the fifth four-dimensional feature and the seventh four-dimensional feature.

[0231] Optionally, the fourth generation model is trained based on the third loss function and the fourth loss function, including determining a second total loss function based on the third loss function and the fourth loss function, and training the fourth generation model based on the second total loss function. For example, the third loss function and the fourth loss function are weighted and summed to determine the second total loss function.

[0232] For example, the formula of the second total loss function is as follows:

[0233] L2=L AE (X5,X6)+βL KL (X5,X6)+αL reg (X5,X7)

[0234] wherein X5 is a fifth four-dimensional feature, X6 is a sixth four-dimensional feature, X7 is a seventh four-dimensional feature, L2 is the first total loss function, L AE (X5,X6), L KL (X5,X6) is a different third loss function, l reg (X5,X7) is a fourth loss function.

[0235] The fourth generation model in the present disclosure can be applicable to dynamic scene information generation of driving scenes, XR scenes, robot scenes, film and television scenes, game scenes, monitoring scenes, and traffic planning. The related content of the second sample feature can refer to the related content of the first sample feature, which will not be repeated here.

[0236] It should be noted that the present disclosure does not limit the execution timing of steps S301 to S308, Figure 3 only taking the sequential execution of steps S301 to S308 as an example. For example, steps S305-S308 can be implemented as independent embodiments.

[0237] The training method of the generation model provided by the embodiments of the present disclosure maps the second sample feature of the K frames of voxels in the three-dimensional space to the four-dimensional space, generates the fifth four-dimensional feature, K is an integer greater than M, processes the fifth four-dimensional feature through the fourth generation model, generates the sixth four-dimensional feature in the four-dimensional space, the sixth four-dimensional feature carries the sixth reconstructed feature of the K frames of voxels in the three-dimensional space, generates the seventh four-dimensional feature in the four-dimensional space through the first generation model and the fourth generation model, the seventh four-dimensional feature carries the seventh reconstructed feature of the voxels in each frame in the three-dimensional space, and trains the fourth generation model based on the third sample feature, the sixth reconstructed feature and the seventh reconstructed feature. Thus, the first generation model can be used to supervise the training of the fourth generation model, and the fourth generation model can learn the four-dimensional feature reconstruction capability of the first generation model in the training process, so that the trained fourth generation model is consistent with the three-dimensional feature reconstructed by the first generation model for the same frame, and the four-dimensional feature reconstruction accuracy of the fourth generation model is significantly improved, that is, the reconstruction accuracy of the dynamic scene information of the fourth generation model is significantly improved.

[0238] On the basis of any of the above embodiments, the seventh four-dimensional feature in the four-dimensional space is generated by the first generation model and the fourth generation model, including at least one of the following ways:

[0239] Manner 1, the fifth four-dimensional feature is projected to a two-dimensional space by a first encoder of the first generation model to generate a third projected feature, the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension and a time dimension, and the third projected feature is decoded by a fourth decoder of the fourth generation model to generate a seventh four-dimensional feature.

[0240] In this embodiment, the seventh reconstructed feature is generated by the first encoder of the first generation model and the fourth decoder of the fourth generation model, that is, the latent space feature encoding capability of the first encoder and the latent space feature decoding capability of the fourth decoder are utilized to jointly generate the seventh reconstructed feature to train the fourth generation model. The first encoder has good latent space feature encoding capability, so that the training of the fourth encoder of the fourth generation model can be supervised by the first encoder of the first generation model. During the training process, the fourth encoder can learn the latent space feature encoding capability of the first encoder, so that the latent space features reconstructed by the trained first encoder and the fourth encoder for the same frame are consistent, which significantly improves the latent space feature encoding precision of the fourth encoder, and further significantly improves the four-dimensional feature reconstruction precision of the fourth generation model, that is, the reconstruction precision of the dynamic scene information of the fourth generation model is significantly improved.

[0241] For example, taking M=16 and K=32 as an example, the fifth four-dimensional feature is projected to a two-dimensional space 1 by a first encoder of the first generation model to generate a third projected feature under the two-dimensional space 1. The fifth four-dimensional feature is projected to a two-dimensional space 2 by the first encoder to generate a third projected feature under the two-dimensional space 2. The fifth four-dimensional feature is projected to a two-dimensional space 3 by the first encoder to generate a third projected feature under the two-dimensional space 3. The fifth four-dimensional feature is projected to a two-dimensional space 4 by the first encoder to generate a third projected feature under the two-dimensional space 4. The fifth four-dimensional feature is projected to a two-dimensional space 5 by the first encoder to generate a third projected feature under the two-dimensional space 5. The fifth four-dimensional feature is projected to a two-dimensional space 6 by the first encoder to generate a third projected feature under the two-dimensional space 6.

[0242] The third projected features under the two-dimensional spaces 1 to 6 are decoded by a fourth decoder of the fourth generation model to generate a seventh four-dimensional feature, and the seventh four-dimensional feature carries the seventh reconstructed features corresponding to the first to 32 frames.

[0243] Manner 2, the fifth four-dimensional feature is projected to a two-dimensional space by a fourth encoder of the fourth generation model to generate a fourth projected feature, and the fourth projected feature is decoded by a first decoder of the first generation model to generate a seventh four-dimensional feature.

[0244] In this embodiment, the seventh reconstruction feature is generated by the fourth encoder of the fourth generative model and the first decoder of the first generative model, that is, the latent space feature encoding capability of the fourth encoder and the latent space feature decoding capability of the first decoder are utilized to jointly generate the seventh reconstruction feature to train the fourth generative model. The first decoder has good latent space feature decoding capability, so that the training of the fourth decoder of the fourth generative model can be supervised by the first decoder of the first generative model. During the training process, the fourth decoder can learn the latent space feature decoding capability of the first decoder, so that the trained fourth decoder is consistent with the first decoder in reconstructing the three-dimensional feature of the same frame, which significantly improves the latent space feature decoding accuracy of the fourth decoder, and further significantly improves the four-dimensional feature reconstruction accuracy of the fourth generative model, that is, the reconstruction accuracy of the dynamic scene information of the fourth generative model is significantly improved.

[0245] For example, taking M=16 and K=32 as an example, the fifth four-dimensional feature is projected into two-dimensional space 1 by the fourth encoder of the fourth generative model to generate the fourth projection feature under two-dimensional space 1. The fifth four-dimensional feature is projected into two-dimensional space 2 by the fourth encoder to generate the fourth projection feature under two-dimensional space 2. The fifth four-dimensional feature is projected into two-dimensional space 3 by the fourth encoder to generate the fourth projection feature under two-dimensional space 3. The fifth four-dimensional feature is projected into two-dimensional space 4 by the fourth encoder to generate the fourth projection feature under two-dimensional space 4. The fifth four-dimensional feature is projected into two-dimensional space 5 by the fourth encoder to generate the fourth projection feature under two-dimensional space 5. The fifth four-dimensional feature is projected into two-dimensional space 6 by the fourth encoder to generate the fourth projection feature under two-dimensional space 6.

[0246] The fourth projection features under two-dimensional spaces 1 to 6 are decoded by the first decoder of the first generative model to generate the seventh four-dimensional feature, which carries the seventh reconstruction features corresponding to the first to 32 frames.

[0247] On the basis of any of the above embodiments, the method further comprises processing the third sample feature of the voxel under the three-dimensional space of any frame by the second generative model to generate the eighth reconstruction feature of the voxel under the three-dimensional space of any frame, and training the second generative model based on the third sample feature and the eighth reconstruction feature. Thus, during the training process, the second generative model can learn the correlation between the third sample feature and the eighth reconstruction feature, so that the trained second generative model can generate the reconstructed three-dimensional feature of a certain frame based on the original three-dimensional feature of the frame.

[0248] Optionally, the third sample feature of the voxels of any frame in the three-dimensional space is processed by the second generation model to generate an eighth reconstruction feature of the voxels of any frame in the three-dimensional space, including projecting the third sample feature corresponding to any frame to a two-dimensional space by a second encoder of the second generation model to generate a ninth reconstruction feature of an image of any frame in the two-dimensional space, and decoding the ninth reconstruction feature corresponding to any frame by a second decoder of the second generation model to generate the eighth reconstruction feature corresponding to any frame.

[0249] Optionally, the second generation model is trained based on the third sample feature and the eighth reconstruction feature, including determining a fifth loss function based on the third sample feature and the eighth reconstruction feature, and training the second generation model based on the fifth loss function.

[0250] On the basis of any of the above embodiments, the method further includes projecting the first four-dimensional feature to a two-dimensional space by a first encoder of the trained first generation model to generate a sample projection feature, denoising the noisy image based on a sample generation condition of the diffusion model to generate a predicted projection feature of the image in the two-dimensional space, and training the diffusion model based on the sample projection feature and the predicted projection feature. Thus, the trained first encoder has good latent space feature coding capability, so that the first encoder can be used to supervise the training of the diffusion model, and the diffusion model can learn the latent space feature coding capability of the first encoder in the training process, so that the trained diffusion model is consistent with the latent space feature reconstructed by the first encoder for the same frame.

[0251] Figure 4 is a flowchart of an image processing method according to an example embodiment, as shown in Figure 4 The image processing method of the present embodiment includes the following steps.

[0252] S401, based on a preset generation condition, denoising a noisy image to generate a target projection feature of an image in a two-dimensional space, and the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension.

[0253] It should be noted that the execution subject of the image processing method of the present embodiment is an electronic device, such as a vehicle terminal, a vehicle controller, a server, an ISP, etc. The image processing method of the present embodiment can be executed by the image processing device of the present embodiment, and the image processing device of the present embodiment can be configured in any electronic device to execute the image processing method of the present embodiment.

[0254] Optionally, the noise image is denoised based on a preset generation condition to generate the target projection feature of the image in the two-dimensional space, including inputting the generation condition and the noise image into a diffusion model, and denoising the noise image based on the generation condition by the diffusion model to generate the target projection feature of the image in the two-dimensional space.

[0255] It can be understood that the generation condition required in different scenarios can be different. The related content of the generation condition can be referred to the following embodiments, which will not be described here.

[0256] S402, processing the target projection feature by the first generation model to generate a target video.

[0257] The target video is determined based on a target four-dimensional feature in a four-dimensional space, the four-dimensional space is constructed based on three spatial dimensions and a time dimension, the target four-dimensional feature carries a target reconstruction feature of voxels in the three-dimensional space of each frame, and the target four-dimensional feature is obtained by processing the target projection feature by the first generation model.

[0258] The first generation model is trained based on a second reconstruction feature of voxels in the three-dimensional space of each frame, the second reconstruction feature is generated by the first generation model and a second generation model, and the second generation model is used to reconstruct the feature of voxels in the three-dimensional space of any frame.

[0259] In the training process of the first generation model in the present disclosure, the second reconstruction feature is generated by the first generation model and the trained second generation model in cooperation, that is, the four-dimensional feature reconstruction capability of the first generation model and the single-frame three-dimensional feature reconstruction capability of the second generation model are utilized to generate the second reconstruction feature together to train the first generation model. The second generation model has good single-frame three-dimensional feature reconstruction capability, so that the training of the first generation model can be supervised by the second generation model. In the training process, the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model, so that the three-dimensional features reconstructed by the trained first generation model and the second generation model for the same frame are consistent, the four-dimensional feature reconstruction accuracy of the first generation model is significantly improved, and the video reconstruction accuracy of the first generation model is further improved, that is, the reconstruction accuracy of dynamic scene information of the first generation model is significantly improved.

[0260] In addition, the noise image can be denoised based on a preset generation condition to generate the target projection feature of the image in the two-dimensional space, and the target projection feature is processed by the first generation model to generate the target video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the video reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of dynamic scene information of the first generation model is significantly improved.

[0261] Optionally, the target projection feature is processed by the first generation model to generate a target video, including decoding the target projection feature by a first decoder of the first generation model to generate a target four-dimensional feature, processing the target four-dimensional feature to generate the target video.

[0262] The first generation model in the present disclosure can be applicable to dynamic scene information generation of driving scenes, XR scenes, robot scenes, film and television scenes, game scenes, monitoring scenes, and traffic planning.

[0263] In the first case, taking a driving scene as an example, the generation conditions include at least one of a driving trajectory of a target vehicle, a driving instruction for the target vehicle, road information, and obstacle information.

[0264] The driving instruction is not limited too much, such as including left turn, straight ahead, etc.

[0265] The road information is not limited too much, such as including road topology (such as sharp turns, slopes), storage location information, etc.

[0266] The obstacle information is not limited too much, such as including obstacle position, obstacle quantity, pedestrian flow, vehicle flow, etc.

[0267] At least one of the driving trajectory of the target vehicle, the driving instruction for the target vehicle, the road information, and the obstacle information, and the noise image can be input to a diffusion model, and the diffusion model can be used to denoise the noise image based on the above generation conditions to generate the target projection feature of the image in the two-dimensional space.

[0268] The target projection feature is decoded by the first decoder to generate a target four-dimensional feature, and the target four-dimensional feature is processed to generate a target video.

[0269] Therefore, the noise image can be denoised based on the preset generation conditions to generate the target projection feature of the image in the two-dimensional space, and the target projection feature is processed by the first generation model to generate a simulated driving video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the driving video reconstruction accuracy of the first generation model, i.e., significantly improves the dynamic driving scene information reconstruction accuracy of the first generation model.

[0270] After generating the target video, at least one of the following operations is performed:

[0271] Operation 1: The target video is used as a simulated driving video.

[0272] For example, the simulation driving video can be added to the driving video library to expand the driving video library, so that the driving video library can cover various driving scenes, such as driving scenes under various weathers, road conditions, different perspectives, and different time periods, thereby improving the richness and diversity of the driving video library.

[0273] Operation 2: training and / or testing the assisted driving model based on the simulation driving video.

[0274] It should be noted that the assisted driving model is not limited, and can include a perception model, map construction, route planning, behavior decision model, etc.

[0275] Therefore, the simulation driving video generated by the first generation model can be used to train and / or test the assisted driving model, which helps to reduce the collection cost of real driving videos and enhances the adaptability of the assisted driving model to various driving scenes, i.e., improves the robustness and reliability of the assisted driving model.

[0276] Operation 3: sending the simulation driving video to at least one of the terminal device and the target vehicle, and the simulation driving video is used for visual display by at least one of the terminal device and the target vehicle.

[0277] Therefore, the simulation driving video can be sent to at least one of the terminal device and the target vehicle to display the simulation driving video on at least one of the terminal device and the target vehicle, so that the user can intuitively understand the simulation driving situation of the target vehicle, thereby optimizing the user experience.

[0278] Operation 4: training and / or testing the traffic planning model based on the simulation driving video.

[0279] It should be noted that the traffic planning model is not limited, and can include a trip generation model, a trip distribution model, a traffic mode division model, and a traffic assignment model, etc.

[0280] Therefore, the simulation driving video generated by the first generation model can be used to train and / or test the traffic planning model, which helps to reduce the collection cost of real driving videos and enhances the adaptability of the traffic planning model to various driving scenes, i.e., improves the robustness and reliability of the traffic planning model.

[0281] In the second case, taking a robot scene as an example, the generation conditions include at least one of the driving trajectory of the target robot, the control instruction for the target robot, the road information, and the obstacle information.

[0282] The running track of the target robot, the control instruction for the target robot, at least one of road information, obstacle information, and a noise image can be input to the diffusion model, and the diffusion model can be used to denoise the noise image based on the generation condition to generate the target projection feature of the image in the two-dimensional space.

[0283] The target projection feature is decoded by the first decoder to generate a target four-dimensional feature, and the target four-dimensional feature is processed to generate a target video.

[0284] Therefore, the noise image can be denoised based on the preset generation condition to generate the target projection feature of the image in the two-dimensional space, and the target projection feature can be processed by the first generation model to generate the simulated robot action video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the robot action video reconstruction accuracy of the first generation model, i.e., the dynamic robot scene information reconstruction accuracy of the first generation model.

[0285] After the target video is generated, at least one of the following operations is performed:

[0286] Operation 1: The target video is used as a simulated robot action video.

[0287] For example, the simulated robot action video can be added to the robot action video library to expand the robot action video library, so that the robot action video library can cover various robot scenes, such as various weather, road conditions, different perspectives, and different time periods, thereby improving the richness and diversity of the robot action video library.

[0288] Operation 2: The robot control model is trained and / or tested based on the simulated robot action video.

[0289] It should be noted that the robot control model is not limited, such as a perception model, map construction, route planning, behavior decision model, etc.

[0290] Therefore, the simulated robot action video generated by the first generation model can be used to train and / or test the robot control model, which helps to reduce the acquisition cost of real robot action videos and enhances the adaptability of the robot control model to various robot scenes, i.e., improves the robustness and reliability of the robot control model.

[0291] Operation 3: The simulated robot action video is sent to at least one of a terminal device and a target robot, and the simulated robot action video is used for visual display by at least one of the terminal device and the target robot.

[0292] Thus, the simulated robot action video can be sent to at least one of the terminal device and the target robot to display the simulated robot action video on at least one of the terminal device and the target robot, so that the user can intuitively understand the simulated action of the target robot, and the user experience is optimized.

[0293] In the third case, the generation condition includes at least one of simulated environment information of the extended reality scene and object behavior information.

[0294] The simulated environment information is not limited too much, such as including geometric information of a building, lighting, weather, etc.

[0295] The object behavior information is not limited too much, such as including running, walking, etc.

[0296] At least one of the simulated environment information of the extended reality scene and the object behavior information, and the noise image can be input to the diffusion model, and the diffusion model is used to denoise the noise image based on the above generation condition to generate target projection features of the image in the two-dimensional space.

[0297] The target projection features are decoded by the first decoder to generate target four-dimensional features, and the target four-dimensional features are processed to generate the target video.

[0298] Thus, the noise image can be denoised based on the preset generation condition to generate target projection features of the image in the two-dimensional space, and the target projection features are processed by the first generation model to generate new extended reality videos, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the extended reality video reconstruction accuracy of the first generation model, i.e., the dynamic XR scene information reconstruction accuracy of the first generation model is significantly improved.

[0299] After the target video is generated, at least one of the following operations is performed:

[0300] Operation 1: The target video is used as a new extended reality video.

[0301] For example, the new extended reality video can be added to the extended reality video library to expand the extended reality video library, so that the extended reality video library can cover multiple XR scenes, and the richness and diversity of the extended reality video library are improved.

[0302] Operation 2: The new extended reality video is sent to the extended reality device, and the new extended reality video is used for visual display by the extended reality device.

[0303] Thus, the new extended reality video can be sent to the extended reality device to display the new extended reality video on the extended reality device.

[0304] It should be noted that the disclosure does not limit the execution timing of steps S401 to S402. Figure 4 Only the sequential execution of steps S401 to S402 is exemplified.

[0305] The image processing method provided by the embodiment of the disclosure performs denoising processing on the noise image based on the preset generation condition, generates target projection features of the image in the two-dimensional space, the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension and time dimension, the target projection features are processed by the first generation model to generate a target video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the video reconstruction accuracy of the first generation model, i.e. significantly improves the reconstruction accuracy of the dynamic scene information of the first generation model.

[0306] For example, as shown in Figure 5 The first generation model includes a first encoder and a first decoder, and the second generation model includes a second encoder and a second decoder. The second generation model can be pre-trained, and the second encoder of the trained second generation model is used to supervise the training of the first encoder of the first generation model, and the second decoder of the trained second generation model is used to supervise the training of the first decoder of the first generation model.

[0307] The video generation model includes a diffusion model and a first decoder. The trained first encoder of the first generation model can be used to supervise the training of the diffusion model.

[0308] The diffusion model is used to perform denoising processing on the noise image based on the generation condition to generate target projection features of the image in the two-dimensional space, the first decoder of the first generation model is used to decode the target projection features to generate target four-dimensional features, and the target four-dimensional features are processed to generate a target video.

[0309] Figure 6 FIG. 6 is a structural schematic diagram of a training device of a generation model according to an exemplary embodiment.

[0310] Referring to Figure 6 , the training device 600 of the generation model of the embodiment of the disclosure includes a mapping module 601, a first processing module 602, a second processing module 603, and a training module 604.

[0311] The mapping module 601 is configured to map the first sample features of M frames of voxels in a three-dimensional space to a four-dimensional space to generate first four-dimensional features, M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension;

[0312] A first processing module 602 is configured to process the first four-dimensional feature using a first generation model to generate a second four-dimensional feature in the four-dimensional space, where the second four-dimensional feature carries a first reconstructed feature of a voxel of the M frame in the three-dimensional space;

[0313] The second processing module 603 is configured to generate second reconstructed features of voxels in each frame in three-dimensional space using the first generation model and the second generation model, where the second generation model is used to reconstruct features of voxels in any frame in three-dimensional space;

[0314] The training module 604 is configured to train the first generation model based on the first sample feature, the first reconstruction feature and the second reconstruction feature.

[0315] In some possible embodiments, the second processing module 603 is further configured to: project the first sample features corresponding to any frame into a two-dimensional space through the second encoder of the second generation model to generate a third reconstructed feature of the image of any frame in the two-dimensional space, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and the time dimension; fuse the third reconstructed features corresponding to each frame to generate a first projection feature; decode the first projection feature through the first decoder of the first generation model to generate a third four-dimensional feature in the four-dimensional space, where the third four-dimensional feature carries the second reconstructed features corresponding to each frame.

[0316] In some possible embodiments, the second processing module 603 is further configured to: project the first four-dimensional feature into a two-dimensional space through the first encoder of the first generation model to generate a second projection feature, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; extract a fourth reconstruction feature of the image of each frame in the two-dimensional space from the second projection feature; and decode the fourth reconstruction feature corresponding to any frame through the second decoder of the second generation model to generate a second reconstruction feature corresponding to any frame.

[0317] In some possible embodiments, the training module 604 is further configured to determine the initial parameters of the first generation model based on at least one of the model parameters of the trained second generation model and the model parameters of the trained third generation model; wherein the third generation model is used to reconstruct the fourth four-dimensional feature in the four-dimensional space, and the fourth four-dimensional feature carries the fifth reconstructed feature of the voxel of N frames in the three-dimensional space, where N is an integer greater than 1 and less than M.

[0318] In some possible implementation, the training module 604 is further configured to: map the second sample feature of the K frames of voxels in the three-dimensional space to the four-dimensional space to generate a fifth four-dimensional feature, K being an integer greater than M; process the fifth four-dimensional feature by a fourth generative model to generate a sixth four-dimensional feature in the four-dimensional space, the sixth four-dimensional feature carrying a sixth reconstructed feature of the K frames of voxels in the three-dimensional space; generate a seventh four-dimensional feature in the four-dimensional space by the first generative model and the fourth generative model, the seventh four-dimensional feature carrying a seventh reconstructed feature of voxels in each frame in the three-dimensional space; and train the fourth generative model based on the second sample feature, the sixth reconstructed feature and the seventh reconstructed feature.

[0319] In some possible implementation, the training module 604 is further configured to: project the fifth four-dimensional feature to a two-dimensional space by a first encoder of the first generative model to generate a third projected feature, the two-dimensional space being constructed in at least one of the following manners: based on any two spatial dimensions, or based on any spatial dimension and the time dimension; and decode the third projected feature by a fourth decoder of the fourth generative model to generate the seventh four-dimensional feature.

[0320] Or, project the fifth four-dimensional feature to a two-dimensional space by a fourth encoder of the fourth generative model to generate a fourth projected feature; and decode the fourth projected feature by a first decoder of the first generative model to generate the seventh four-dimensional feature.

[0321] As to the apparatus in the above-mentioned embodiments, the specific manners in which various modules perform operations have been described in detail in the embodiments of the method, and thus will not be described herein.

[0322] The training device for generating a model provided by the embodiment of the present disclosure maps the first sample features of voxels of M frames in a three-dimensional space to a four-dimensional space to generate first four-dimensional features, M is an integer greater than 1, the four-dimensional space is constructed based on three spatial dimensions and a time dimension, the first four-dimensional features are processed by a first generation model to generate second four-dimensional features in the four-dimensional space, the second four-dimensional features carry the first reconstructed features of voxels of M frames in the three-dimensional space, the second reconstructed features of voxels of each frame in the three-dimensional space are generated by the first generation model and a second generation model, the second generation model is used to reconstruct the features of voxels of any frame in the three-dimensional space, and the first generation model is trained based on the first sample features, the first reconstructed features and the second reconstructed features. Therefore, the training of the first generation model can be supervised by the second generation model, the first generation model can learn the single-frame three-dimensional feature reconstruction capability of the second generation model in the training process, the three-dimensional features reconstructed by the trained first generation model and the second generation model for the same frame are consistent, and the four-dimensional feature reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of dynamic scene information of the first generation model, is significantly improved.

[0323] Figure 7 FIG. 1 is a structural schematic diagram of an image processing device according to an example embodiment.

[0324] With reference to Figure 7 The image processing device 700 of the embodiment of the present disclosure includes a first processing module 701 and a second processing module 702.

[0325] The first processing module 701 is configured to perform denoising processing on a noisy image based on a preset generation condition to generate target projection features of images in a two-dimensional space, the two-dimensional space being constructed in at least one of the following manners: any two spatial dimensions, any one spatial dimension and a time dimension.

[0326] The second processing module 702 is configured to process the target projection features by a first generation model to generate a target video.

[0327] The target video is determined based on target four-dimensional features in a four-dimensional space, the four-dimensional space being constructed based on three spatial dimensions and a time dimension, the target four-dimensional features carrying target reconstructed features of voxels of each frame in a three-dimensional space, and the target four-dimensional features being obtained by processing the target projection features by the first generation model.

[0328] The first generation model is trained based on second reconstructed features of voxels of each frame in the three-dimensional space, the second reconstructed features being generated by the first generation model and a second generation model, and the second generation model being used to reconstruct the features of voxels of any frame in the three-dimensional space.

[0329] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target vehicle, a driving instruction for the target vehicle, road information, and obstacle information.

[0330] After the target video is generated, the second processing module 702 is further configured to perform at least one of the following operations:

[0331] The target video is used as a simulation driving video.

[0332] A driving assistance model is trained and / or tested based on the simulation driving video.

[0333] The simulation driving video is sent to at least one of a terminal device and the target vehicle, and the simulation driving video is used for visual display by at least one of the terminal device and the target vehicle.

[0334] A traffic planning model is trained and / or tested based on the simulation driving video.

[0335] In some possible implementation manners, the generation condition comprises at least one of a driving track of the target vehicle, a driving instruction for the target vehicle, road information, and obstacle information.

[0336] After the target video is generated, the second processing module 702 is further configured to perform at least one of the following operations:

[0337] The target video is used as a simulation robot action video.

[0338] A robot control model is trained and / or tested based on the simulation robot action video.

[0339] The simulation robot action video is sent to at least one of a terminal device and the target robot, and the simulation robot action video is used for visual display by at least one of the terminal device and the target robot.

[0340] In some possible implementation manners, the generation condition comprises at least one of simulation environment information of an extended reality scene and object behavior information.

[0341] After the target video is generated, the second processing module 702 is further configured to perform at least one of the following operations:

[0342] The target video is used as an extended reality video.

[0343] The extended reality video is sent to an extended reality device, and the extended reality video is used for visual display by the extended reality device.

[0344] As to the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be described in detail here.

[0345] The image processing apparatus provided by the embodiments of the present disclosure performs denoising processing on a noise image based on a preset generation condition, generates a target projection feature of an image in a two-dimensional space, the two-dimensional space is constructed in at least one of the following manners: any two spatial dimensions, any spatial dimension, and a time dimension, and the target projection feature is processed by a first generation model to generate a target video, which significantly improves the four-dimensional feature reconstruction accuracy of the first generation model, and further significantly improves the video reconstruction accuracy of the first generation model, that is, the reconstruction accuracy of dynamic scene information of the first generation model is significantly improved.

[0346] In order to implement the above-mentioned embodiments, the present disclosure further proposes an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, when the processor executes the program, the steps of the training method of the generation model provided by the present disclosure are implemented, and / or the steps of the image processing method provided by the present disclosure are implemented.

[0347] Optionally, the electronic device includes a vehicle-mounted terminal, a vehicle-mounted controller, a server, an ISP, etc.

[0348] Optionally, the electronic device can be a server device, or a separate processing platform.

[0349] Figure 8 FIG. 8 is a structural schematic diagram of an electronic device according to an example embodiment. For example, the electronic device 800 can be a vehicle, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0350] Referring to Figure 8 The electronic device 800 can include one or more of the following components: a processing component 802, a memory 804, a power component 806, a multimedia component 808, an audio component 810, an input / output (I / O) interface 812, a sensor component 814, and a communication component 816.

[0351] The processing component 802 generally controls the overall operation of the electronic device 800 such as the operation associated with displaying, making phone calls, data communications, camera operations, and recording operations. The processing component 802 can include one or more processors 820 to execute instructions to complete all or a subset of the steps described in the above-mentioned method of training a model and / or the above-mentioned method of image processing. Furthermore, the processing component 802 can include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 can include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0352] The memory 804 is configured to store various types of data to support operations of the electronic device 800. Examples of these data include instructions to operate any applications or methods on the electronic device 800, contact data, phonebook data, messages, pictures, videos, and the like. The memory 804 can be implemented by any type of volatile or non-volatile storage devices or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic or optical disks.

[0353] The power component 806 provides power to the various components of the electronic device 800. The power component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 800.

[0354] The multimedia component 808 includes a screen providing an output interface between the electronic device 800 and a user. In some embodiments, the screen can include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive an input signal from a user. The touch panel includes one or more touch sensors to sense a touch, a slide and a gesture on the touch panel. The touch sensor can not only sense a boundary of a touching or a sliding action, but also detect duration and pressure related to the touching or sliding action. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.

[0355] The audio component 810 is configured to output and / or input an audio signal. For example, the audio component 810 includes a microphone (MIC) configured to receive an external audio signal when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting an audio signal.

[0356] The I / O interface 812 provides an interface between the processing component 802 and a peripheral interface module, which can be a keypad, a click wheel, buttons, and the like. The buttons can include, but are not limited to, a home button, a volume button, a start button, and a lock button.

[0357] The sensor assembly 814 includes one or more sensors for providing various aspects of the status assessment of the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and changes in the temperature of the electronic device 800. The sensor assembly 814 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 can also include an optical sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 can also include an accelerometer, a gyroscope, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0358] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0359] In an exemplary embodiment, the electronic device 800 can be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the steps of the above-mentioned training method for generating a model, and / or execute the steps of the above-mentioned image processing method.

[0360] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is further provided, such as a memory 804 including instructions, wherein the instructions can be executed by the processor 820 of the electronic device 800 to complete the steps of the training method for generating the model and / or the steps of the image processing method. For example, the non-transitory computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, etc.

[0361] In order to implement the above embodiments, the present disclosure also proposes a computer-readable storage medium on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the training method for generating a model provided by the present disclosure are implemented, and / or the steps of the image processing method provided by the present disclosure are implemented.

[0362] Alternatively, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, or the like.

[0363] In order to implement the above embodiments, the present disclosure also proposes a chip, which includes an interface circuit and a processing circuit coupled to each other, the interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the training method of the generative model provided by the present disclosure, and / or, implement the steps of the image processing method provided by the present disclosure.

[0364] Figure 9FIG. 1 is a schematic diagram showing the structure of a chip according to an exemplary embodiment. Figure 9 The structure of the chip 900 is shown, but is not limited thereto.

[0365] Chip 900 includes a processing circuit 901, which is configured to execute the steps of any of the above-mentioned training methods for generating a model, and / or execute the steps of any of the above-mentioned image processing methods.

[0366] In some embodiments, chip 900 further includes one or more interface circuits 902. Optionally, interface circuit 902 is connected to memory 903. Interface circuit 902 can be used to receive signals from memory 903 or other devices, and can be used to send signals to memory 903 or other devices. For example, interface circuit 902 can read instructions stored in memory 903 and send the instructions to processing circuit 901.

[0367] In some embodiments, the interface circuit 902 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 901 performs the other steps.

[0368] In some embodiments, terms such as interface circuit, interface, transceiver pin, and transceiver may be used interchangeably.

[0369] In some embodiments, the chip 900 further includes one or more memories 903 for storing instructions. Alternatively, all or part of the memories 903 may be located outside the chip 900 .

[0370] In order to implement the above embodiments, the present disclosure also proposes a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the training method for generating a model provided by the present disclosure, and / or implements the steps of the image processing method provided by the present disclosure.

[0371] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.

[0372] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A training method for a generative model, characterized in that: include: Mapping the first sample features of the voxels of the M frames in the three-dimensional space to the four-dimensional space to generate a first four-dimensional feature, where M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension; Processing the first four-dimensional feature through a first generation model to generate a second four-dimensional feature in the four-dimensional space, where the second four-dimensional feature carries a first reconstructed feature of a voxel of the M frames in the three-dimensional space; generating second reconstruction features of voxels in three-dimensional space for each frame using the first generation model and the second generation model, wherein the second generation model is used to reconstruct features of voxels in three-dimensional space for any frame; The first generation model is trained based on the first sample feature, the first reconstruction feature, and the second reconstruction feature.

2. The method according to claim 1, characterized in that Generating second reconstruction features of voxels of each frame in three-dimensional space by using the first generation model and the second generation model includes: Projecting the first sample features corresponding to any frame into a two-dimensional space through a second encoder of the second generative model to generate third reconstructed features of the image of any frame in the two-dimensional space, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; Performing fusion processing on the third reconstructed features corresponding to each frame to generate a first projection feature; The first projection feature is decoded by a first decoder of the first generation model to generate a third four-dimensional feature in the four-dimensional space, where the third four-dimensional feature carries the second reconstructed feature corresponding to each frame.

3. The method according to claim 1, characterized in that Generating second reconstruction features of voxels of each frame in three-dimensional space by using the first generation model and the second generation model includes: Projecting the first four-dimensional feature into a two-dimensional space by a first encoder of the first generative model to generate a second projected feature, wherein the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; Extracting a fourth reconstruction feature of the image of each frame in a two-dimensional space from the second projection feature; The fourth reconstructed feature corresponding to any frame is decoded by the second decoder of the second generation model to generate the second reconstructed feature corresponding to the any frame.

4. The method according to claim 1, wherein The method further comprises: Determining initial parameters of the first generative model based on at least one of the trained model parameters of the second generative model and the trained model parameters of the third generative model; The third generation model is used to reconstruct a fourth four-dimensional feature in the four-dimensional space, and the fourth four-dimensional feature carries a fifth reconstructed feature of voxels of N frames in the three-dimensional space, where N is an integer greater than 1 and less than M.

5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Mapping the second sample features of the voxels of the K frames in the three-dimensional space to the four-dimensional space to generate a fifth four-dimensional feature, where K is an integer greater than M; Processing the fifth four-dimensional feature through a fourth generation model to generate a sixth four-dimensional feature in the four-dimensional space, the sixth four-dimensional feature carrying a sixth reconstructed feature of the voxel of the K frame in the three-dimensional space; generating a seventh four-dimensional feature in the four-dimensional space by using the first generation model and the fourth generation model, wherein the seventh four-dimensional feature carries a seventh reconstructed feature of a voxel of each frame in the three-dimensional space; The fourth generation model is trained based on the second sample feature, the sixth reconstruction feature, and the seventh reconstruction feature.

6. The method according to claim 5, characterized in that Generating the seventh four-dimensional feature in the four-dimensional space by using the first generation model and the fourth generation model includes: Projecting the fifth four-dimensional feature into a two-dimensional space by a first encoder of the first generative model to generate a third projected feature, wherein the two-dimensional space is constructed based on at least one of any two spatial dimensions, any one spatial dimension, and a time dimension; Decoding the third projection feature by a fourth decoder of the fourth generation model to generate the seventh four-dimensional feature; or, Projecting the fifth four-dimensional feature into a two-dimensional space through a fourth encoder of the fourth generative model to generate a fourth projected feature; The fourth projection feature is decoded by a first decoder of the first generation model to generate the seventh four-dimensional feature.

7. An image processing method, characterized in that: include: Denoising the noisy image based on a preset generation condition to generate target projection features of the image in a two-dimensional space, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any one spatial dimension, and a time dimension; Processing the target projection features through a first generation model to generate a target video; The target video is determined based on a target four-dimensional feature in a four-dimensional space, the four-dimensional space is constructed based on three spatial dimensions and a time dimension, the target four-dimensional feature carries a target reconstruction feature of a voxel in each frame in the three-dimensional space, and the target four-dimensional feature is obtained by processing the target projection feature by the first generation model; The first generative model is trained based on the second reconstruction features of voxels in each frame in three-dimensional space, the second reconstruction features are generated by the first generative model and the second generative model, and the second generative model is used to reconstruct the features of voxels in any frame in three-dimensional space.

8. The method according to claim 7, characterized in that The generation condition includes at least one of a target vehicle's driving trajectory, a driving instruction for the target vehicle, road information, and obstacle information; After generating the target video, performing at least one of the following operations: Using the target video as a simulated driving video; Training and / or testing an assisted driving model based on the simulated driving video; The simulated driving video is sent to a terminal device and at least one device among the target vehicle, and the simulated driving video is used for visual display by the terminal device and at least one device among the target vehicle. Based on the simulated driving video, a traffic planning model is trained and / or tested.

9. The method according to claim 7, characterized in that The generation condition includes at least one of the following information: a target robot's driving trajectory, a control instruction for the target robot, road information, and obstacle information; After generating the target video, performing at least one of the following operations: Using the target video as a simulated robot action video; Training and / or testing a robot control model based on the simulated robot action video; The simulated robot action video is sent to at least one of a terminal device and the target robot, and the simulated robot action video is used for visual display by at least one of the terminal device and the target robot.

10. The method according to claim 7, characterized in that The generation condition includes at least one of the following information: simulation environment information of the augmented reality scene and object behavior information; After generating the target video, performing at least one of the following operations: Using the target video as an extended reality video; The extended reality video is sent to an extended reality device, where the extended reality video is used for visual display by the extended reality device.

11. A training device for generating a model, characterized in that: include: a mapping module configured to map first sample features of voxels of the M frames in the three-dimensional space to a four-dimensional space to generate a first four-dimensional feature, where M is an integer greater than 1, and the four-dimensional space is constructed based on three spatial dimensions and a time dimension; A first processing module is configured to process the first four-dimensional feature through a first generation model to generate a second four-dimensional feature in the four-dimensional space, where the second four-dimensional feature carries a first reconstructed feature of a voxel of the M frame in the three-dimensional space; A second processing module is configured to generate second reconstructed features of voxels in three-dimensional space of each frame using the first generation model and the second generation model, wherein the second generation model is used to reconstruct features of voxels in three-dimensional space of any frame; A training module is configured to train the first generation model based on the first sample feature, the first reconstruction feature and the second reconstruction feature.

12. The device according to claim 11, characterized in that The second processing module is further configured to: Projecting the first sample features corresponding to any frame into a two-dimensional space through a second encoder of the second generative model to generate third reconstructed features of the image of any frame in the two-dimensional space, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; Performing fusion processing on the third reconstructed features corresponding to each frame to generate a first projection feature; The first projection feature is decoded by a first decoder of the first generation model to generate a third four-dimensional feature in the four-dimensional space, where the third four-dimensional feature carries the second reconstructed feature corresponding to each frame.

13. The device according to claim 11 or 12, characterized in that The second processing module is further configured to: Projecting the first four-dimensional feature into a two-dimensional space by a first encoder of the first generative model to generate a second projected feature, wherein the two-dimensional space is constructed based on at least one of any two spatial dimensions, any spatial dimension, and a time dimension; Extracting a fourth reconstruction feature of the image of each frame in a two-dimensional space from the second projection feature; The fourth reconstructed feature corresponding to any frame is decoded by the second decoder of the second generation model to generate the second reconstructed feature corresponding to the any frame.

14. An image processing device, characterized in that: include: a first processing module configured to perform denoising on the noisy image based on a preset generation condition to generate target projection features of the image in a two-dimensional space, where the two-dimensional space is constructed based on at least one of any two spatial dimensions, any one spatial dimension, and a time dimension; A second processing module is configured to process the target projection features through the first generation model to generate a target video; The target video is determined based on a target four-dimensional feature in a four-dimensional space, the four-dimensional space is constructed based on three spatial dimensions and a time dimension, the target four-dimensional feature carries a target reconstruction feature of a voxel in each frame in the three-dimensional space, and the target four-dimensional feature is obtained by processing the target projection feature by the first generation model; The first generative model is trained based on the second reconstruction features of voxels in each frame in three-dimensional space, the second reconstruction features are generated by the first generative model and the second generative model, and the second generative model is used to reconstruct the features of voxels in any frame in three-dimensional space.

15. The device according to claim 14, characterized in that The generation condition includes at least one of a target vehicle's driving trajectory, a driving instruction for the target vehicle, road information, and obstacle information; After generating the target video, the second processing module is further configured to perform at least one of the following operations: Using the target video as a simulated driving video; Training and / or testing an assisted driving model based on the simulated driving video; The simulated driving video is sent to a terminal device and at least one device among the target vehicle, and the simulated driving video is used for visual display by the terminal device and at least one device among the target vehicle. Based on the simulated driving video, a traffic planning model is trained and / or tested.

16. The device according to claim 14, characterized in that The generation condition includes at least one of the following information: a target robot's driving trajectory, a control instruction for the target robot, road information, and obstacle information; After generating the target video, the second processing module is further configured to perform at least one of the following operations: Using the target video as a simulated robot action video; Training and / or testing a robot control model based on the simulated robot action video; The simulated robot action video is sent to at least one of a terminal device and the target robot, and the simulated robot action video is used for visual display by at least one of the terminal device and the target robot.

17. An electronic device, characterized in that: The invention comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the model training method described in any one of claims 1 to 6 are implemented, and / or the steps of the image processing method described in any one of claims 7 to 10 are implemented.

18. A computer program product, characterized in that It includes a computer program, which, when executed by a processor, implements the steps of the model training method described in any one of claims 1 to 6, and / or implements the steps of the image processing method described in any one of claims 7 to 10.