Method, apparatus, and storage medium for processing video
By adjusting the appearance style of the target object image in the original video, videos with different appearance styles are generated, thereby building a three-dimensional Gaussian model with different appearance styles, solving the problem of the three-dimensional Gaussian model being too single, and achieving a richer and more beautiful three-dimensional Gaussian model.
Patent Information
- Application Number
- CN202411991745.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2044-12-31
AI Technical Summary
The three-dimensional Gaussian model reconstructed the appearance is too single and lacks diversity and aesthetics.
By adjusting the appearance style of the target object image in the original video, videos with different appearance styles are generated, thereby building a three-dimensional Gaussian model with different appearance styles. Specific methods include using the target filter template to adjust lighting parameters or modifying the image style by modifying the image style.
It realizes rich appearance and style of the three-dimensional Gaussian model, improving the aesthetics and diversity of the model.
Smart Images

Figure CN119399341B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing, and particularly to a method, apparatus, and storage medium for processing videos. Background Art
[0002] The three-dimensional Gaussian technique is a technique for three-dimensional scene representation and rendering. Through the three-dimensional Gaussian technique, a 1:1 reconstruction of real-world objects can be performed to obtain a three-dimensional Gaussian model with the same shape as the real-world objects.
[0003] In the related art, the three-dimensional Gaussian model reconstructed using the three-dimensional Gaussian technique is the same as the real-world object, resulting in a too single appearance of the reconstructed three-dimensional Gaussian model. Summary of the Invention
[0004] This application provides a method, apparatus, and storage medium for processing videos to enrich the appearance of the three-dimensional Gaussian model. The technical solution is as follows:
[0005] In a first aspect, this application provides a method for processing a video. In this method, a first video is obtained based on an original video. The original video includes multiple frames of pictures taken of a target object from multiple perspectives. The first video includes multiple frames of first object images, and the multiple frames of first object images are images of the target object in the multiple frames of pictures. The appearance style of the multiple frames of first object images included in the first video is adjusted to obtain at least one second video corresponding to at least one appearance style. Each second video includes multiple frames of second object images, and the appearance style of the multiple frames of second object images included in each second video is the appearance style corresponding to each second video. Based on the original video and at least one second video, at least one first three-dimensional Gaussian model is obtained, and the appearance style of the at least one first three-dimensional Gaussian model is the same as the at least one appearance style.
[0006] Since the appearance style of the multiple frames of first object images included in the first video is adjusted to obtain at least one second video corresponding to at least one appearance style, at least one first three-dimensional Gaussian model with different appearance styles is obtained based on the original video and at least one second video. That is to say, one or more first three-dimensional Gaussian models with different appearance styles can be constructed, thereby enriching the appearance and styles of the three-dimensional Gaussian model.
[0007] In a possible implementation, at least one appearance style includes a target appearance style. Based on the target filter template, the parameter values of at least one lighting parameter of each frame of the first object image included in the first video are adjusted to obtain a second video corresponding to the target appearance style. The target filter template is used to store the standard parameter values of at least one lighting parameter, and the standard parameter values are used to indicate the target appearance style. Alternatively, based on an image style modification model and the first video, the appearance styles of multiple frames of the first object images included in the first video are modified to the target appearance style to obtain a second video corresponding to the target appearance style. In this way, second videos with different appearance styles are obtained through filter templates or image style modification models with different appearance styles. Based on the second videos with different appearance styles, three-dimensional Gaussian models with different appearance styles can be constructed.
[0008] In another possible implementation, at least one appearance style includes a target appearance style, and multiple frames of second object images included in the second video corresponding to the target appearance style correspond to the multiple frames of pictures. The camera poses and camera parameters corresponding to each frame of the pictures included in the original video are obtained. Based on the camera poses, camera parameters, and second object images corresponding to each frame of the pictures, a first three-dimensional Gaussian model of the target appearance style is constructed. Since the second object images corresponding to each frame of the pictures are introduced when constructing the three-dimensional Gaussian model, and the appearance styles of the second object images corresponding to each frame of the pictures are the target appearance style, a three-dimensional Gaussian model of the target appearance style can be constructed.
[0009] In another possible implementation, at least one appearance style includes a target appearance style; a second three-dimensional Gaussian model of the target object is obtained based on the original video. The parameter values of at least one lighting parameter of the second three-dimensional Gaussian model are adjusted to obtain a third three-dimensional Gaussian model. When the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style satisfies the difference condition, the second three-dimensional Gaussian model and the third three-dimensional Gaussian model are superimposed to obtain a first three-dimensional Gaussian model of the target appearance style. In this way, the method for obtaining the three-dimensional Gaussian model of the target appearance style is enriched.
[0010] In another possible implementation, when the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not satisfy the difference condition, the parameter values of at least one lighting parameter of the third three-dimensional Gaussian model are adjusted to obtain a fourth three-dimensional Gaussian model. When the difference between the fourth three-dimensional Gaussian model and the second video corresponding to the target appearance style satisfies the difference condition, the second three-dimensional Gaussian model and the fourth three-dimensional Gaussian model are superimposed to obtain a first three-dimensional Gaussian model of the target appearance style. In this way, the three-dimensional Gaussian model of the target appearance style can be accurately obtained.
[0011] In another possible implementation, at least one lighting parameter of the second three-dimensional Gaussian model includes one or more of the following: color or transparency.
[0012] In another possible implementation, the correspondence between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model is saved. The at least one third three-dimensional Gaussian model corresponds to at least one second video, and the difference between each third three-dimensional Gaussian model and the second video corresponding to each third three-dimensional Gaussian model satisfies the difference condition. This can reduce the storage resources required for saving.
[0013] In a second aspect, the present application provides a device for processing videos, the device including: an acquisition unit and an adjustment unit.
[0014] The acquisition unit is configured to acquire a first video based on an original video, where the original video includes multiple frames of pictures taken of a target object from multiple perspectives, and the first video includes multiple frames of first object images, and the multiple frames of first object images are images of the target object in the multiple frames of pictures;
[0015] The adjustment unit is configured to adjust the appearance style of the multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style. Each second video includes multiple frames of second object images, and the appearance style of the multiple frames of second object images included in each second video is respectively the appearance style corresponding to each second video;
[0016] The acquisition unit is further configured to acquire at least one first three-dimensional Gaussian model based on the original video and at least one second video, and the appearance style of the at least one first three-dimensional Gaussian model is the same as the at least one appearance style.
[0017] Since the adjustment unit adjusts the appearance style of the multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, the acquisition unit acquires at least one first three-dimensional Gaussian model with different appearance styles based on the original video and at least one second video. That is to say, one or more first three-dimensional Gaussian models with different appearance styles can be constructed, thus enriching the appearance and style of the three-dimensional Gaussian model.
[0018] In a possible implementation, the at least one appearance style includes a target appearance style. The adjustment unit is configured to:
[0019] Based on a target filter template, adjust the parameter values of at least one lighting parameter of each frame of the first object images included in the first video to obtain a second video corresponding to the target appearance style. The target filter template is used to save the standard parameter values of at least one lighting parameter, and the standard parameter values are used to indicate the target appearance style. Or,
[0020] Based on the image style modification model and the first video, modify the appearance style of multiple frames of first object images included in the first video to the target appearance style, and obtain a second video corresponding to the target appearance style.
[0021] In this way, by using filter templates or image style modification models with different appearance styles, second videos with different appearance styles are obtained. Based on the second videos with different appearance styles, three-dimensional Gaussian models with different appearance styles can be constructed.
[0022] In another possible implementation manner, at least one appearance style includes the target appearance style, and multiple frames of second object images included in the second video corresponding to the target appearance style correspond to the multiple frames of pictures. The obtaining unit is used for:
[0023] Obtain the camera pose and camera parameters corresponding to each frame of picture included in the original video;
[0024] Based on the camera pose, camera parameters, and second object images corresponding to each frame of picture, construct a first three-dimensional Gaussian model of the target appearance style.
[0025] Since the second object images corresponding to each frame of picture are introduced when constructing the three-dimensional Gaussian model, and the appearance style of the second object images corresponding to each frame of picture is the target appearance style, a three-dimensional Gaussian model of the target appearance style can be constructed.
[0026] In another possible implementation manner, at least one appearance style includes the target appearance style;
[0027] The obtaining unit is further used to obtain a second three-dimensional Gaussian model of the target object based on the original video;
[0028] The adjustment unit is further used to adjust the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model to obtain a third three-dimensional Gaussian model;
[0029] The obtaining unit is used to superimpose the second three-dimensional Gaussian model and the third three-dimensional Gaussian model to obtain a first three-dimensional Gaussian model of the target appearance style when the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition. In this way, the method for obtaining the three-dimensional Gaussian model of the target appearance style is enriched.
[0030] In another possible implementation manner, the adjustment unit is further used to adjust the parameter values of at least one lighting parameter of the third three-dimensional Gaussian model to obtain a fourth three-dimensional Gaussian model when the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not meet the difference condition.
[0031] The obtaining unit is further configured to, when the difference between the fourth three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition, superimpose the second three-dimensional Gaussian model and the fourth three-dimensional Gaussian model to obtain a first three-dimensional Gaussian model of the target appearance style. In this way, the three-dimensional Gaussian model of the target appearance style can be accurately obtained.
[0032] In another possible implementation manner, at least one illumination parameter of the second three-dimensional Gaussian model includes one or more of the following: color or transparency.
[0033] In another possible implementation manner, the device further includes a saving unit; the saving unit is configured to save the correspondence between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model, at least one third three-dimensional Gaussian model corresponds to at least one second video, and the difference between each third three-dimensional Gaussian model and the second video corresponding to each third three-dimensional Gaussian model meets the difference condition. In this way, the storage resources required for saving can be reduced.
[0034] In a third aspect, the present application provides a cluster of computing devices, the cluster of computing devices includes at least one computing device, and each computing device includes a processor and a memory;
[0035] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the cluster of computing devices executes the method in the first aspect or any possible implementation manner of the first aspect.
[0036] In a fourth aspect, the present application provides a computer program product including instructions, when the instructions are executed by a cluster of computing devices, the cluster of computing devices executes the method in the first aspect or any possible implementation manner of the first aspect.
[0037] In a fifth aspect, the present application provides a computer-readable storage medium, including computer program instructions, when the computer program instructions are executed by a cluster of computing devices, the cluster of computing devices executes the method in the first aspect or any possible implementation manner of the first aspect.
[0038] In a sixth aspect, the present application provides a chip, including a memory and a processor, the memory is configured to store computer instructions, and the processor is configured to call and run the computer instructions from the memory to execute the method in the first aspect or any possible implementation manner of the first aspect. Description of the Drawings
[0039] Figure 1 is a schematic diagram of a target object provided by an embodiment of the present application;
[0040] Figure 2It is a schematic diagram of an original video provided by an embodiment of the present application;
[0041] Figure 3 It is a schematic diagram of a first video provided by an embodiment of the present application;
[0042] Figure 4 It is a schematic diagram of a three - dimensional Gaussian model provided by an embodiment of the present application;
[0043] Figure 5 It is a schematic diagram of the structure of a computing device provided by an embodiment of the present application;
[0044] Figure 6 It is a flowchart of a method for processing a video provided by an embodiment of the present application;
[0045] Figure 7 It is another flowchart of a method for processing a video provided by an embodiment of the present application;
[0046] Figure 8 It is a schematic diagram of the structure of a device for processing a video provided by an embodiment of the present application;
[0047] Figure 9 It is a schematic diagram of the structure of another computing device provided by an embodiment of the present application;
[0048] Figure 10 It is a schematic diagram of the structure of a computing device cluster provided by an embodiment of the present application;
[0049] Figure 11 It is a schematic diagram of the structure of another computing device cluster provided by an embodiment of the present application. Detailed implementation manners
[0050] A camera can be used to capture an original video of a target object from multiple different perspectives. The original video includes multiple frames of pictures, and the perspective of each frame of picture is different. Each frame of the original picture includes an image of the target object and a background image of the environment where the target object is located. Since the perspective of each frame of picture is different, each frame of picture includes images of different faces of the target object.
[0051] For example, referring to Figure 1 the target object shown, an original video is captured from five different perspectives of the target object. The five perspectives are the front - face perspective, the right - side perspective, the back - face perspective, the left - side perspective, and the top - down perspective. The original video includes five frames of pictures.
[0052] Referring to Figure 2 , use a camera to capture a picture a of the front of the target object from the front - face perspective. Picture a includes the front - face image of the target object and the background image, that is Figure 2Picture a in it. Use a camera to take a picture of the right side of the target object from the right side view to obtain Picture b. Picture b includes the right side image of the target object, that is Figure 2 Picture b in it. Use a camera to take a picture of the back of the target object from the rear view to obtain Picture c. Picture c includes the back image of the target object and the background image, that is Figure 2 Picture c in it. Use a camera to take a picture of the left side of the target object from the left side view to obtain Picture d. Picture d includes the left side image of the target object, that is Figure 2 Picture d in it. Use a camera to take a picture of the top surface of the target object from the top view to obtain Picture e. Picture e includes the top surface image of the target object and the background image, that is Figure 2 Picture e in it. In this way, the five frames of pictures included in the original video are obtained, that is, the original video includes Picture a, Picture b, Picture c, Picture d, and Picture e.
[0053] Obtain the first video based on the original video. The first video includes multiple frames of object images, and the multiple frames of object images are the images of the target object in the multiple frames of pictures. The multiple frames of object images correspond one-to-one with the multiple frames of pictures. Obtain the camera pose and camera parameters corresponding to each frame of picture, and construct a three-dimensional Gaussian model whose shape, color, and texture are the same as those of the target object based on the camera pose, camera parameters, and target object image corresponding to each frame of picture.
[0054] For example, based on Figure 2 the five frames of pictures included in the original video shown, obtain the first video, and the first video includes five frames of object images. Optionally, in implementation, for Figure 2 Picture a included in the original video shown, obtain the front image of the target object from Figure 2 Picture a in it. The front image of the target object is as shown in Figure 3 Figure (1) in it. For Figure 2 Picture b included in the original video shown, obtain the right side image of the target object from Figure 2 Picture b in it. The right side image is as shown in Figure 3 Figure (2) in it. For Figure 2 Picture c included in the original video shown, obtain the back image of the target object from Figure 2 Picture c in it. The back image is as shown in Figure 3 Figure (3) in it. For Figure 2 Picture d included in the original video shown, obtain the left side image of the target object from Figure 2 Picture d in it. The left side image is as shown in Figure 3 Figure (4) in it. For Figure 2 Picture e included in the original video shown, obtain it from Figure 2Obtain the top surface image of the target object from the picture e in []. As shown in Figure 3 (5) in the figure. Therefore, the first video obtained includes five frames of object images, which are the front image, right side image, back image, left side image, and top surface image of the target object respectively. The original pictures including picture a, picture b, picture c, picture d, and picture e respectively correspond one-to-one to the front image, right side image, back image, left side image, and top surface image of the target object included in the first video.
[0055] Obtain the camera poses and camera parameters corresponding to picture a included in the original video, the camera poses and camera parameters corresponding to picture b, the camera poses and camera parameters corresponding to picture c, the camera poses and camera parameters corresponding to picture d, and the camera poses and camera parameters corresponding to picture e. Based on the camera pose, camera parameters, and front image corresponding to picture a, the camera pose, camera parameters, and right side image corresponding to picture b, the camera pose, camera parameters, and back image corresponding to picture c, the camera pose, camera parameters, and left side image corresponding to picture d, and the camera pose, camera parameters, and top surface image corresponding to picture e, construct a three-dimensional Gaussian model whose shape, color, and texture are all the same as those of the target object. The three-dimensional Gaussian model is as shown in Figure 4 shown.
[0056] Since the constructed three-dimensional Gaussian model is the same as the target object in the real world, the appearance of the three-dimensional Gaussian model is too single and not beautiful enough. In order to enrich the appearance of the three-dimensional Gaussian model and improve the aesthetics of the three-dimensional Gaussian model, any of the following embodiments can be used to process the video to obtain at least one three-dimensional Gaussian model with different appearance styles.
[0057] Refer to Figure 5 , an embodiment of the present application provides a computing device 500, and the computing device 500 includes a material processor 501, a material equalizer 502, and a model construction module 503.
[0058] The material processor 501 is used to obtain the original video. The original video includes multiple frames of pictures taken of the target object from multiple perspectives. Based on the original video, obtain the first video. The first video includes multiple frames of first object images, and the multiple frames of first object images are the images of the target object in the multiple frames of pictures.
[0059] The material equalizer 502 is used to adjust the appearance styles of the multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style. Each second video includes multiple frames of second object images. The appearance styles of the multiple frames of second object images included in each second video are respectively the appearance styles corresponding to each second video, and the appearance styles of the multiple frames of second object images included in each second video are different.
[0060] The model construction module 503 is configured to obtain at least one first three-dimensional Gaussian model of the target object based on the original video and at least one second video. The appearance styles of the at least one first three-dimensional Gaussian model are the same as at least one appearance style, and the appearance styles of each first three-dimensional Gaussian model are different.
[0061] In some embodiments, referring to Figure 5 , the computing device 500 further includes a lighting parser 504;
[0062] The lighting parser 504 is configured to parse the lighting condition of each frame of picture included in the original video to obtain the parameter values of at least one lighting parameter of each frame of picture included in the original video;
[0063] The material equalizer 502 is configured to adjust the appearance styles of multiple first object images included in the first video based on the parameter values of at least one lighting parameter of each frame of picture included in the original video, so as to obtain at least one second video corresponding to at least one appearance style.
[0064] In some embodiments, referring to Figure 5 , the computing device 500 further includes a camera estimation module 505; at least one appearance style includes a target appearance style, and multiple second object images included in the second video corresponding to the target appearance style correspond one-to-one with multiple frames of pictures included in the original video.
[0065] The camera estimation module 505 is configured to obtain the camera pose and camera parameters corresponding to each frame of picture included in the original video.
[0066] The model construction module 503 is configured to construct a first three-dimensional Gaussian model of the target appearance style based on the camera pose, camera parameters, and second object images corresponding to each frame of picture.
[0067] In some embodiments, the method for the computing device 500 to process the original video to obtain at least one first three-dimensional Gaussian model of the target object can be applied to one or more of the following scenarios:
[0068] 1. Any media application scenario;
[0069] 2. Game scenario;
[0070] 3. Advertising, promotion, and film and television video production scenarios;
[0071] 4. Any display scenario such as culture and tourism, e-commerce, and enterprise exhibition halls;
[0072] 5. Extended Reality (XR) glasses.
[0073] In the embodiment of the present application, since the material equalizer 502 adjusts the appearance styles of multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, and the appearance styles of each second video are different, the model construction module constructs at least one first three-dimensional Gaussian model of at least one appearance style based on the original video and at least one second video. That is, one or more first three-dimensional Gaussian models with different appearance styles can be constructed, thus enriching the styles of the three-dimensional Gaussian models and improving the aesthetics of the three-dimensional Gaussian models.
[0074] See Figure 6 , the embodiment of the present application provides a method 600 for processing videos. The method 600 can be applied to Figure 5 the computing device shown in the figure. The method 600 obtains a first video based on the original video, adjusts the appearance styles of multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, and multiple frames of second object images included in each second video correspond one-to-one to multiple frames of pictures included in the original video. At least one first three-dimensional object model of at least one appearance style is constructed based on the camera pose, parameters, and second object images corresponding to each frame of picture. The method 600 includes the following processes.
[0075] Step 601: Obtain a first video based on the original video. The original video includes multiple frames of pictures taken of the target object from multiple perspectives. The first video includes multiple frames of first object images, and these multiple frames of first object images are images of the target object in these multiple frames of pictures.
[0076] The camera can be used to take pictures of different faces of the target object to obtain multiple frames of pictures taken of the target object from multiple perspectives, that is, to obtain the original video including these multiple frames of pictures.
[0077] For each frame of picture in the original video, extract the first object image of the target object from this picture. Extract the first object image of the target object from each frame of picture in the same way as above to obtain multiple frames of first object images corresponding one-to-one to these multiple frames of pictures, and thus obtain the first video including these multiple frames of first object images.
[0078] In some embodiments, the computing device includes a material processor, and the material processor extracts the first object images in each frame of picture included in the original video to obtain a second video.
[0079] Step 602: Adjust the appearance styles of multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style. Each second video includes multiple frames of second object images, and the appearance styles of multiple frames of second object images included in each second video are the appearance styles corresponding to each second video respectively.
[0080] At least one appearance style includes a target appearance style. For each frame of the first object image in the first video, adjust the appearance style of this frame of the first object image to the target appearance style to obtain a frame of the second object image with the target appearance style. Adjust the appearance style of each frame of the first object image to the target appearance style in the same manner as above to obtain multiple frames of the second object image with the target appearance style, and thus obtain a second video including these multiple frames of the second object image, that is, obtain the second video corresponding to the target appearance style. At least one appearance style also includes other appearance styles, and second videos corresponding to other appearance styles can be obtained by the same operation as above.
[0081] In some embodiments, the computing device includes at least one filter template, and the at least one filter template corresponds one-to-one with at least one appearance style, and each filter template is respectively used to indicate the appearance style corresponding to each filter template.
[0082] For each filter template included in the at least one filter template, the filter template includes standard parameter values of at least one lighting parameter, and the standard parameter values of the at least one lighting parameter are used to indicate the appearance style corresponding to the filter template.
[0083] The standard parameter values of the lighting parameters included in the filter template corresponding to each appearance style may be different, so that each filter template is respectively used to indicate the appearance style corresponding to each filter template.
[0084] For example, at least one appearance style includes one or more of the following: Ghibli style, sparkling water style, or midsummer style, etc. The standard parameter values of at least one lighting parameter included in the filter template corresponding to the Ghibli style, the standard parameter values of at least one lighting parameter included in the filter template corresponding to the sparkling water style, and the standard parameter values of at least one lighting parameter included in the filter template corresponding to the midsummer style are different.
[0085] The at least one lighting parameter includes one or more of the following: lighting matrix, color temperature, brightness, sharpness, saturation, or white balance, etc.
[0086] For each filter template included in the at least one filter template, for the sake of convenience of description, this filter template is called the target filter template, and the appearance style corresponding to the target filter template is called the target appearance style. In step 602, based on the target filter template, adjust the parameter values of at least one lighting parameter of each frame of the first object image included in the first video to obtain the second video corresponding to the target appearance style.
[0087] Optionally, when implemented, the second video corresponding to the target appearance style can be obtained according to the following process 6021-6023.
[0088] 6021: Obtain the parameter values of at least one lighting parameter for each frame of the picture included in the original video, and based on the parameter values of at least one lighting parameter for each frame of the picture, obtain the average parameter value of the at least one lighting parameter.
[0089] The computing device includes a lighting parser and a material equalizer. The original video can be input into the lighting parser, and the lighting parser can parse the lighting conditions of each frame of the picture included in the original video to obtain the parameter values of at least one lighting parameter for each frame of the picture included in the original video. Based on the parameter values of at least one lighting parameter for each frame of the picture, obtain the average parameter value of the at least one lighting parameter. Next, through the material equalizer, adjust the appearance style of the first video according to the following process to obtain the second video corresponding to the target appearance style.
[0090] 6022: Based on the standard parameter value of at least one lighting parameter included in the target filter template and the average parameter value of the at least one lighting parameter, obtain the offset value of the at least one lighting parameter.
[0091] For each lighting parameter among the at least one lighting parameter, based on the standard parameter value of the lighting parameter included in the target filter template and the average parameter value of the lighting parameter, obtain the offset value of the lighting parameter.
[0092] 6023: Based on the offset value of the at least one lighting parameter, adjust the first parameter value of at least one lighting parameter of each frame of the first object image included in the first video to obtain the second parameter value of at least one lighting parameter of each frame of the first object image. For two adjacent frames of the first object image, adjust the second parameter value of at least one lighting parameter of the two adjacent frames of the first object image to reduce the difference between the second parameter values of at least one lighting parameter of the two adjacent frames of the first object image, and obtain two adjacent frames of the second object image.
[0093] For example, for the first frame of the first object image and the second frame of the first object image, adjust the second parameter value of at least one lighting parameter of the first frame of the first object image and the second parameter value of at least one lighting parameter of the second frame of the first object image to reduce the difference between the second parameter value of at least one lighting parameter of the first frame of the first object image and the second parameter value of at least one lighting parameter of the second frame of the first object image, and obtain the first frame of the second object image and the second frame of the second object image.
[0094] In some embodiments, for two adjacent frames of first object images, the first parameter values of at least one illumination parameter of the two adjacent frames of first object images may also be adjusted to reduce the difference between the first parameter values of at least one illumination parameter of the two adjacent frames of first object images, thereby obtaining two adjacent frames of third object images. Based on the offset value of the at least one illumination parameter, the first parameter values of at least one illumination parameter of each frame of the third object images are adjusted to obtain two adjacent frames of second object images.
[0095] In some embodiments, based on an image style modification model and a first video, the appearance style of multiple frames of first object images included in the first video is modified to a target appearance style, thereby obtaining a second video corresponding to the target appearance style.
[0096] Optionally, the first video and the target appearance style may be input into the image style modification model. The image style modification model modifies multiple frames of first object images included in the first video based on the target appearance style, obtains multiple frames of second object images with the target appearance style, and outputs multiple frames of second object images with the target appearance style. In this way, a second video corresponding to the target appearance style is obtained, and the second video corresponding to the target appearance style includes multiple frames of second object images with the target appearance style.
[0097] Optionally, the image style modification model may be any model for stylizing a video.
[0098] Step 603: Based on the original video and at least one second video, obtain at least one first three-dimensional Gaussian model, and the appearance style of the at least one first three-dimensional Gaussian model is the same as at least one appearance style.
[0099] The at least one first three-dimensional Gaussian model is a three-dimensional Gaussian model of the target object.
[0100] The at least one second video includes a second video with the target appearance style. Based on the original video and the second video corresponding to the target appearance style, a first three-dimensional Gaussian model with the target appearance style is obtained. In implementation, the operations 6031-6032 may be performed as follows.
[0101] 6031: Obtain the camera pose and camera parameters corresponding to each frame of picture included in the original video.
[0102] Optionally, the camera parameters include the internal parameters and / or external parameters of the camera, etc.
[0103] The computing device includes a camera estimation module. For each frame of picture included in the original video, the camera estimation module is used to process the frame of picture to obtain the camera pose and camera parameters when the camera captures the frame of picture, that is, the camera pose and camera parameters corresponding to the frame of picture are obtained.
[0104] 6032: Construct a first three-dimensional Gaussian model of the target appearance style based on the camera pose, camera parameters, and second object image corresponding to each frame of the picture.
[0105] The computing device includes a model construction module that constructs a first three-dimensional Gaussian model of the target appearance style based on the camera pose, camera parameters, and second object image corresponding to each frame of the picture through the model construction module.
[0106] In the embodiments of the present application, at least one second video corresponding to at least one appearance style is obtained, and the appearance styles of each second video are different. Based on the original video and at least one second video, at least one first three-dimensional Gaussian model of at least one appearance style is obtained. That is to say, one or more first three-dimensional Gaussian models with different appearance styles can be obtained, thus enriching the appearance of the three-dimensional Gaussian model.
[0107] See Figure 7 , the embodiments of the present application provide a method 700 for processing video. The method 700 can be applied to Figure 5 the computing device shown. The method 700 obtains a first video based on the original video, adjusts the appearance style of multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style. Each second video includes multiple frames of second object images that correspond one-to-one with the multiple frames of pictures included in the original video. A second three-dimensional Gaussian model of the target object is constructed based on the original video, and the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model are adjusted to obtain a third three-dimensional Gaussian model. When the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style satisfies the difference condition, the second three-dimensional Gaussian model and the third three-dimensional Gaussian model are superimposed to obtain a first three-dimensional Gaussian model of the target appearance style. The method 700 includes the following processes.
[0108] Steps 701-702: Are the same as steps 601-602 of the method 600 shown in Figure 6 respectively, and will not be described in detail here.
[0109] Step 703: Obtain a second three-dimensional Gaussian model of the target object based on the original video.
[0110] In step 701, a first video is obtained based on the original video. The original video includes multiple frames of pictures taken of the target object from multiple perspectives. The first video includes multiple frames of first object images, and these multiple frames of first object images are images of the target object in these multiple frames of pictures. That is to say, the multiple frames of first object images included in the first video correspond one-to-one with the multiple frames of pictures included in the original video.
[0111] In step 703, obtain the camera poses and camera parameters corresponding to each frame of the original video. Based on the camera poses, camera parameters, and the first object image corresponding to each frame of the image, construct a second three-dimensional Gaussian model.
[0112] The appearance style of the second three-dimensional Gaussian model is the same as the appearance style of the target object. That is, the second three-dimensional Gaussian model is a Gaussian model with the same texture, color, and shape as the target object's texture, color, and shape respectively.
[0113] In step 702, at least one second video corresponding to at least one appearance style is obtained. For each appearance style among the at least one appearance style, for the sake of convenience of explanation, this appearance style is called the target appearance style, and the first three-dimensional Gaussian model of the target appearance style is obtained according to the following process.
[0114] Step 704: Adjust the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model to obtain a third three-dimensional Gaussian model.
[0115] Adjusting the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model can change the appearance style of the second three-dimensional Gaussian model, that is, it can change the shape, texture, and / or color, etc. of the second three-dimensional Gaussian model, and obtain a third three-dimensional Gaussian model with an appearance style different from that of the second three-dimensional Gaussian model.
[0116] Step 705: Compare whether the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition. If the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition, execute step 706. If the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not meet the difference condition, execute step 707.
[0117] The second video includes second object images of different faces of the target object, and the third three-dimensional Gaussian model is composed of third object images of different faces of the target object. The multiple frames of second object images included in the second video correspond one-to-one with the multiple frames of third object images of the third three-dimensional Gaussian model.
[0118] For each frame of the second object image and the third object image corresponding to the second object image, obtain the peak signal-to-noise ratio (PSNR) between the second object image and the third object image. According to the above process, multiple PSNRs between multiple frames of second object images and multiple frames of third object images can be obtained, and the average PSNR is obtained based on the multiple PSNRs.
[0119] Among them, if the average PSNR is smaller, the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style is smaller; if the average PSNR is larger, the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style is larger.
[0120] If the average PSNR is less than or equal to the PSNR threshold, comparing the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition; if the average PSNR is greater than the PSNR threshold, comparing the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not meet the difference condition.
[0121] Step 706: Superimpose the second three-dimensional Gaussian model and the third three-dimensional Gaussian model to obtain the first three-dimensional Gaussian model of the target appearance style.
[0122] The second three-dimensional Gaussian model includes a plurality of first Gaussian spheres, and the third three-dimensional Gaussian model includes a plurality of second Gaussian spheres, and the plurality of first Gaussian spheres correspond to the plurality of second Gaussian spheres one by one.
[0123] In step 706, obtain the position, shape, color information, and transparency of the plurality of first Gaussian spheres included in the second three-dimensional Gaussian model, and obtain the position, shape, color information, and transparency of the plurality of second Gaussian spheres included in the third three-dimensional Gaussian model. Based on the positions and shapes of the plurality of first Gaussian spheres and the positions and shapes of the plurality of second Gaussian spheres, determine a plurality of Gaussian sphere pairs, each Gaussian sphere pair includes a first Gaussian sphere and a second Gaussian sphere, and the position and shape of the first Gaussian sphere are respectively the same as the position and shape of the second Gaussian sphere. Superimpose the color information of the first Gaussian sphere included in each Gaussian sphere pair and the color information of the second Gaussian sphere, and superimpose the transparency of the first Gaussian sphere included in each Gaussian sphere pair and the transparency of the second Gaussian sphere, to obtain the color information and transparency of the plurality of Gaussian spheres included in the first three-dimensional Gaussian model of the target appearance style, that is, obtain the first three-dimensional Gaussian model of the target appearance style.
[0124] For other appearance styles except the target appearance style among at least one appearance style, return to start executing the above step 704 to obtain the first three-dimensional Gaussian model of other appearance styles.
[0125] Step 707: Update the second three-dimensional Gaussian model to the third-dimensional Gaussian model, and return to execute step 704.
[0126] That is to say, continue to adjust the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model until the difference between the adjusted third three-dimensional Gaussian model and the second video corresponding to the target appearance is less than the threshold, and then superimpose the second three-dimensional Gaussian model and the adjusted third three-dimensional Gaussian model to obtain the first three-dimensional Gaussian model of the target appearance style.
[0127] At least one third three-dimensional Gaussian model corresponding to at least one second video can be obtained, and the difference between each third three-dimensional Gaussian model and the second video corresponding to each third three-dimensional Gaussian model satisfies the difference condition.
[0128] Save the corresponding relationship between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model. In this way, when at least one first three-dimensional Gaussian model of at least one appearance style is needed, each third three-dimensional Gaussian model in at least one third three-dimensional Gaussian model is superimposed on the second three-dimensional Gaussian model to obtain at least one first three-dimensional Gaussian model of at least one appearance style.
[0129] The total data volume of the corresponding relationship between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model is less than the total data of at least one first three-dimensional Gaussian model of at least one appearance style. Therefore, saving the corresponding relationship between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model can reduce the storage resources required for saving.
[0130] In the embodiment of the present application, at least one second video corresponding to at least one appearance style is obtained, and the appearance styles of each second video are different. A second three-dimensional Gaussian model is obtained based on the original video. The parameter values of at least one lighting parameter of the second three-dimensional Gaussian model are adjusted to obtain a third three-dimensional Gaussian model. When the difference between the third three-dimensional Gaussian model and the second video corresponding to a certain appearance style satisfies the difference condition, the second three-dimensional Gaussian model and the third three-dimensional Gaussian model are superimposed to obtain the first three-dimensional Gaussian model of this appearance style. In this way, at least one first three-dimensional Gaussian model corresponding to at least one appearance style can be obtained, thus enriching the appearance of the three-dimensional Gaussian model.
[0131] See Figure 8 , an apparatus 800 for processing video is provided in the embodiment of the present application. The apparatus 800 can be deployed on Figure 5 the computing device shown in Figure 6 or the computing device of the method 600 shown in Figure 7 or the computing device of the method 700 shown in
[0132] The obtaining unit 801 is configured to obtain a first video based on the original video. The original video includes multiple frames of pictures taken of the target object from multiple perspectives. The first video includes multiple frames of first object images, and the multiple frames of first object images are images of the target object in the multiple frames of pictures;
[0133] An adjustment unit 802, configured to adjust the appearance style of multiple frames of first object images included in the first video, to obtain at least one second video corresponding to at least one appearance style, each second video includes multiple frames of second object images, and the appearance styles of the multiple frames of second object images included in each second video are respectively the appearance styles corresponding to each second video;
[0134] An obtaining unit 801 is further configured to obtain at least one first three-dimensional Gaussian model based on the original video and at least one second video, and the appearance styles of the at least one first three-dimensional Gaussian model are the same as the at least one appearance style.
[0135] Optionally, for the detailed implementation process of the obtaining unit 801 to obtain the first video based on the original video, refer to Figure 6 step 601 of the method 600 shown in Figure 7 the relevant content of step 701 of the method 700 shown in, which will not be elaborated here.
[0136] Optionally, for the detailed implementation process of the adjustment unit 802 to adjust the appearance style of multiple frames of first object images included in the first video, refer to Figure 6 step 602 of the method 600 shown in Figure 7 the relevant content of step 702 of the method 700 shown in, which will not be elaborated here.
[0137] Optionally, for the detailed implementation process of the obtaining unit 801 to obtain at least one first three-dimensional Gaussian model based on the original video and at least one second video, refer to Figure 6 step 603 of the method 600 shown in Figure 7 the relevant content of steps 703-707 of the method 700 shown in, which will not be elaborated here.
[0138] Optionally, the at least one appearance style includes a target appearance style. The adjustment unit 802 is configured to:
[0139] Based on a target filter template, adjust the parameter values of at least one lighting parameter of each frame of first object image included in the first video, to obtain a second video corresponding to the target appearance style, where the target filter template is used to store standard parameter values of at least one lighting parameter, and the standard parameter values are used to indicate the target appearance style. Or,
[0140] Based on an image style modification model and the first video, modify the appearance style of multiple frames of first object images included in the first video to the target appearance style, to obtain a second video corresponding to the target appearance style.
[0141] Optionally, for the detailed implementation process of the adjustment unit 802 to adjust the parameter values of at least one lighting parameter of each frame of first object image included in the first video based on the target filter template, refer to Figure 6Step 602 of the method 600 shown or Figure 7 the relevant content of step 702 of the method 700 shown will not be elaborated here.
[0142] Optionally, for the detailed implementation process of the adjustment unit 802 to modify the appearance style of multiple frames of first object images included in the first video to the target appearance style based on the image style modification model and the first video, refer to Figure 6 Step 602 of the method 600 shown or Figure 7 the relevant content of step 702 of the method 700 shown will not be elaborated here.
[0143] Optionally, at least one appearance style includes the target appearance style, and multiple frames of second object images included in the second video corresponding to the target appearance style correspond to the multiple frames of pictures. The acquisition unit 801 is configured to:
[0144] acquire the camera pose and camera parameters corresponding to each frame of picture included in the original video;
[0145] Based on the camera pose, camera parameters, and second object images corresponding to each frame of picture, construct a first three-dimensional Gaussian model of the target appearance style.
[0146] Optionally, for the detailed implementation process of the acquisition unit 801 to acquire the camera pose and camera parameters corresponding to each frame of picture included in the original video, refer to Figure 6 the relevant content of step 603 of the method 600 shown will not be elaborated here.
[0147] Optionally, for the detailed implementation process of the acquisition unit 801 to construct a first three-dimensional Gaussian model of the target appearance style based on the camera pose, camera parameters, and second object images corresponding to each frame of picture, refer to Figure 6 the relevant content of step 603 of the method 600 shown will not be elaborated here.
[0148] Optionally, at least one appearance style includes the target appearance style;
[0149] The acquisition unit 801 is further configured to acquire a second three-dimensional Gaussian model of the target object based on the original video;
[0150] The adjustment unit 802 is further configured to adjust the parameter values of at least one lighting parameter of the second three-dimensional Gaussian model to obtain a third three-dimensional Gaussian model;
[0151] The acquisition unit 801 is configured to, when the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition, superimpose the second three-dimensional Gaussian model and the third three-dimensional Gaussian model to obtain a first three-dimensional Gaussian model of the target appearance style.
[0152] Optionally, for the detailed implementation process of the obtaining unit 801 to obtain the second three-dimensional Gaussian model of the target object based on the original video, refer to Figure 7 the relevant content of step 703 of the method 700 shown in
[0153] Optionally, for the detailed implementation process of the adjustment unit 802 to adjust the parameter value of at least one illumination parameter of the second three-dimensional Gaussian model to obtain the third three-dimensional Gaussian model, refer to Figure 7 the relevant content of step 704 of the method 700 shown in
[0154] Optionally, for the detailed implementation process of the obtaining unit 801 to superimpose the second three-dimensional Gaussian model and the third three-dimensional Gaussian model to obtain the first three-dimensional Gaussian model of the target appearance style, refer to Figure 7 the relevant content of step 706 of the method 700 shown in
[0155] Optionally, the adjustment unit 802 is further configured to, when the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not meet the difference condition, adjust the parameter value of at least one illumination parameter of the third three-dimensional Gaussian model to obtain the fourth three-dimensional Gaussian model.
[0156] The obtaining unit 801 is further configured to, when the difference between the fourth three-dimensional Gaussian model and the second video corresponding to the target appearance style meets the difference condition, superimpose the second three-dimensional Gaussian model and the fourth three-dimensional Gaussian model to obtain the first three-dimensional Gaussian model of the target appearance style.
[0157] Optionally, for the detailed implementation process of the adjustment unit 802 to adjust the parameter value of at least one illumination parameter of the third three-dimensional Gaussian model to obtain the fourth three-dimensional Gaussian model, refer to Figure 7 the relevant content of step 704 of the method 700 shown in
[0158] Optionally, for the detailed implementation process of the obtaining unit 801 to superimpose the second three-dimensional Gaussian model and the fourth three-dimensional Gaussian model to obtain the first three-dimensional Gaussian model of the target appearance style, refer to Figure 7 the relevant content of step 706 of the method 700 shown in
[0159] Optionally, at least one illumination parameter of the second three-dimensional Gaussian model includes one or more of the following: color or transparency.
[0160] Optionally, the device 800 further includes a storage unit 803;
[0161] A storage unit 803, configured to store the correspondence between at least one third three-dimensional Gaussian model and the second three-dimensional Gaussian model, where the at least one third three-dimensional Gaussian model corresponds to at least one second video, and the difference between each third three-dimensional Gaussian model and the second video corresponding to each third three-dimensional Gaussian model satisfies the difference condition.
[0162] Among them, the obtaining unit 801, the adjusting unit 802, and the storage unit 803 can all be implemented by software or by hardware. Exemplarily, next, taking the obtaining unit 801 as an example, the implementation manner of the obtaining unit 801 will be introduced. Similarly, the implementation manners of the adjusting unit 802 and the storage unit 803 can refer to the implementation manner of the obtaining unit 801.
[0163] As an example of a software functional unit, the obtaining unit 801 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above computing instance may be one or more. For example, the obtaining unit 801 may include code running on multiple hosts / virtual machines / containers. The multiple hosts / virtual machines / containers for running this code may be distributed in the same region, or may be distributed in different regions. Further, the multiple hosts / virtual machines / containers for running this code may be distributed in the same availability zone (AZ), or may be distributed in different AZs, and each AZ includes one data center or multiple geographically proximate data centers. Usually, one region may include multiple AZs.
[0164] Similarly, the multiple hosts / virtual machines / containers for running this code may be distributed in the same virtual private cloud (VPC), or may be distributed in multiple VPCs. Usually, one VPC is set within one region. For cross-region communication between two VPCs within the same region and between VPCs in different regions, a communication gateway needs to be set in each VPC, and the interconnection between VPCs is realized through the communication gateway.
[0165] As an example of a hardware functional unit, the obtaining unit 801 may include at least one computing device, such as a server, etc. Alternatively, the obtaining unit 801 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). Among them, the above PLD may be implemented by a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0166] The multiple computing devices included in the obtaining unit 801 may be distributed in the same region or in different regions. The multiple computing devices included in the obtaining unit 801 may be distributed in the same availability zone (AZ) or in different AZs. Similarly, the multiple computing devices included in the obtaining unit 801 may be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Among them, the multiple computing devices may be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0167] It should be noted that in other embodiments, the obtaining unit 801 may be used to execute any step in the method for processing video, the adjusting unit 802 may be used to execute any step in the method for processing video, and the saving unit 803 may be used to execute any step in the method for processing video. The steps to be implemented by the obtaining unit 801, the adjusting unit 802, and the saving unit 803 can be specified as needed, and the entire function of the video processing apparatus 800 is realized by respectively implementing different steps in the method for processing video through the obtaining unit 801, the adjusting unit 802, and the saving unit 803.
[0168] In the embodiment of the present application, since the adjusting unit adjusts the appearance style of multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, the obtaining unit obtains at least one first three-dimensional Gaussian model with different appearance styles based on the original video and the at least one second video. That is, one or more first three-dimensional Gaussian models with different appearance styles can be constructed, thus enriching the appearance and style of the three-dimensional Gaussian model.
[0169] See Figure 9 , the embodiment of the present application provides a computing device 900. For example, the computing device 900 may beFigure 5 the computing device shown, or Figure 6 the computing device of the method 600 shown, or Figure 7 the computing device of the method 700 shown.
[0170] As Figure 9 shown, the computing device 900 includes: a bus 902, a processor 904, a memory 906, and a communication interface 908. The processor 904, the memory 906, and the communication interface 908 communicate with each other via the bus 902. The computing device 900 can be a server or a terminal device. It should be understood that the present application does not limit the number of processors and memories in the computing device 900.
[0171] The bus 902 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 9 only one line is used in the figure, but it does not mean that there is only one bus or one type of bus. The bus 902 can include a path for transmitting information between various components of the computing device 900 (for example, the processor 904, the memory 906, the communication interface 908).
[0172] The processor 904 can include any one or more of processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).
[0173] The memory 906 can include a volatile memory, such as a random access memory (RAM). The memory 906 can also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).
[0174] Referring to Figure 9 , executable program codes are stored in the memory 906, and the processor 904 executes the executable program codes to respectively implementFigure 8 The functions of the acquisition unit 801, the adjustment unit 802, and the storage unit 803 in the device 800 shown are used to implement the method provided in any of the above embodiments. That is, instructions for executing the method provided in any of the above embodiments are stored in the memory 906. Or,
[0175] The communication interface 908 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 900 and other devices or communication networks.
[0176] The embodiment of the present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be a server, such as a central server, an edge server, or a local server in a local data center. In some embodiments, the computing device can also be a terminal device such as a desktop computer, a laptop computer, or a smart phone.
[0177] As Figure 10 shown, the computing device cluster includes at least one computing device 900. Instructions for executing the method for processing video provided in any of the above embodiments can be stored in the memory 906 of one or more computing devices 900 in the computing device cluster.
[0178] In some possible implementation manners, partial instructions for executing the method for processing video can also be stored separately in the memory 906 of one or more computing devices 900 in the computing device cluster. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the method provided in any of the above embodiments.
[0179] The memory 906 in different computing devices 900 in the computing device cluster can store different instructions, respectively used to execute partial functions of the device 800 for processing video as Figure 8 shown. That is, the instructions stored in the memory 906 of different computing devices 900 can implement the functions of one or more of the acquisition unit 801, the adjustment unit 802, and the storage unit 803.
[0180] In some possible implementation manners, one or more computing devices in the computing device cluster can be connected through a network. Wherein, the network can be a wide area network or a local area network, etc. Figure 11 A possible implementation manner is shown. As Figure 11 shown, two computing devices 900A and 900B are connected through a network. Specifically, they are connected to the network through the communication interfaces in each computing device.
[0181] In this type of possible implementation, instructions for implementing the functions of the acquisition unit 801 and the adjustment unit 802 in the embodiment shown in Figure 8 are stored in the memory 906 of the computing device 900A. At the same time, instructions for implementing the function of the storage unit 803 in the embodiment shown in Figure 8 are stored in the memory 906 of the computing device 900B.
[0182] Figure 11 The connection manner between the computing device clusters shown in
[0183] should be understood that Figure 11 the functions of the computing device 900A shown in
[0184] can also be completed by multiple computing devices 900. Similarly, the functions of the computing device 900B can also be completed by multiple computing devices 900. Figure 11 Another computing device cluster is provided in an embodiment of the present application. The connection relationship between the computing devices in the computing device cluster can be similarly referred to
[0185] the connection manner of the described computing device cluster. The difference is that instructions for executing the method for processing video provided in any of the above embodiments can be stored in the memory 906 of one or more computing devices 900 in the computing device cluster.
[0186] In some possible implementation manners, partial instructions for executing the method provided in any of the above embodiments can also be stored separately in the memory 906 of one or more computing devices 900 in the computing device cluster. In other words, a combination of one or more computing devices 900 can jointly execute the instructions for executing the method provided in any of the above embodiments.
[0187] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center containing one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive), etc. The computer-readable storage medium includes instructions that direct the computing device to execute the method provided in any of the above embodiments.
[0188] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware or by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.
[0189] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present application shall be included within the protection scope of the present application.
Claims
1. A method for processing a video, characterized in that: The method comprises: Acquire a first video based on an original video, wherein the original video includes multiple frames of pictures obtained by shooting a target object from multiple perspectives, and the first video includes multiple frames of first object images, and the multiple frames of first object images are images of the target object in the multiple frames of pictures; Adjusting the appearance styles of the multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, each second video includes multiple frames of second object images, and the appearance styles of the multiple frames of second object images included in each second video are respectively the appearance styles corresponding to each second video; Based on the original video and the at least one second video, acquiring at least one first three-dimensional Gaussian model, wherein the appearance style of the at least one first three-dimensional Gaussian model is the same as the at least one appearance style; wherein the at least one appearance style comprises a target appearance style; The acquiring at least one first three-dimensional Gaussian model based on the original video and the at least one second video includes: Acquire a second three-dimensional Gaussian model of the target object based on the original video; Adjusting a parameter value of at least one illumination parameter of the second three-dimensional Gaussian model to obtain a third three-dimensional Gaussian model; When a difference between the third 3D Gaussian model and the second video corresponding to the target appearance style satisfies a difference condition, the second 3D Gaussian model and the third 3D Gaussian model are superimposed to obtain a first 3D Gaussian model of the target appearance style.
2. The method according to claim 1, characterized in that The at least one appearance style comprises a target appearance style; The step of adjusting the appearance styles of the plurality of frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style includes: Based on a target filter template, adjusting a parameter value of at least one illumination parameter of each frame of the first object image included in the first video to obtain a second video corresponding to the target appearance style, wherein the target filter template is used to store a standard parameter value of the at least one illumination parameter, and the standard parameter value is used to indicate the target appearance style; or Based on the image style modification model and the first video, the appearance styles of the multiple frames of first object images included in the first video are modified to the target appearance style, so as to obtain a second video corresponding to the target appearance style.
3. The method according to claim 1 or 2, characterized in that The method further comprises: When the difference between the third three-dimensional Gaussian model and the second video corresponding to the target appearance style does not satisfy the difference condition, adjusting a parameter value of at least one illumination parameter of the third three-dimensional Gaussian model to obtain a fourth three-dimensional Gaussian model; When the difference between the fourth 3D Gaussian model and the second video corresponding to the target appearance style satisfies the difference condition, the second 3D Gaussian model and the fourth 3D Gaussian model are superimposed to obtain a first 3D Gaussian model of the target appearance style.
4. The method according to claim 3, characterized in that The at least one illumination parameter of the second three-dimensional Gaussian model includes one or more of the following: color or transparency.
5. A device for processing video, characterized in that: The device comprises: an acquisition unit, configured to acquire a first video based on an original video, wherein the original video includes a plurality of frames of pictures obtained by photographing a target object from a plurality of perspectives, and the first video includes a plurality of frames of first object images, wherein the plurality of frames of first object images are images of the target object in the plurality of frames of pictures; an adjusting unit, configured to adjust the appearance styles of the multiple frames of first object images included in the first video to obtain at least one second video corresponding to at least one appearance style, each second video including multiple frames of second object images, and the appearance styles of the multiple frames of second object images included in each second video are respectively the appearance styles corresponding to each second video; The acquisition unit is further used to acquire at least one first three-dimensional Gaussian model based on the original video and the at least one second video, wherein the appearance style of the at least one first three-dimensional Gaussian model is the same as the at least one appearance style; wherein the at least one appearance style comprises a target appearance style; The acquiring at least one first three-dimensional Gaussian model based on the original video and the at least one second video includes: Acquire a second three-dimensional Gaussian model of the target object based on the original video; Adjusting a parameter value of at least one illumination parameter of the second three-dimensional Gaussian model to obtain a third three-dimensional Gaussian model; When a difference between the third 3D Gaussian model and the second video corresponding to the target appearance style satisfies a difference condition, the second 3D Gaussian model and the third 3D Gaussian model are superimposed to obtain a first 3D Gaussian model of the target appearance style.
6. A computing device cluster, characterized in that: comprising at least one computing device, each computing device comprising a processor and a memory; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the method according to any one of claims 1 to 4.
7. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the method according to any one of claims 1 to 4.
8. A computer program product comprising instructions, characterized in that When the instruction is executed by the computing device cluster, the computing device cluster executes the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
3D Gaussian scene style migration method based on 2D prior, computer equipment and computer program product
CN118967915A