Free viewpoint video generation method, device, electronic device and readable storage medium
By identifying the orientation of the subject in the video and performing frame selection processing, and building a neural radiation field network for training, the problem of missing subject parts in the free viewpoint video of a monocular camera is solved, and the viewing effect of the video is improved.
Patent Information
- Application Number
- CN202311811288.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2043-12-25
AI Technical Summary
When shooting free view videos with a monocular camera, the video quality cannot be guaranteed, resulting in the missing part of the subject in the generated free view video, affecting the viewing effect.
By identifying the orientation of the subject in the video, selecting frames on the video, ensuring that the number of frames in each orientation category is the same or similar, building a neural radiation field network, and training it according to the preset loss function to generate a free viewpoint video.
It effectively avoids the missing part of the main body in the generated free viewing video, and improves the viewing effect of the video.
Smart Images

Figure CN118509617B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminals, and in particular, to a method, apparatus, electronic device, and readable storage medium for generating free-viewpoint video. Background Art
[0002] Free-Viewpoint Video (FVV) is a video technology that allows a scene to be viewed from any angle, enabling users to freely change the viewing perspective and select different perspectives to view objects or people in the scene. Generally, free-viewpoint video is obtained by synchronously recording videos from multiple cameras set at different angles and then analyzing and reconstructing the videos from each angle.
[0003] In order to expand the application scenarios of free-viewpoint video and lower the shooting threshold of free-viewpoint video, there is currently a method for shooting and generating monocular free-viewpoint video. Monocular free-viewpoint video refers to capturing a scene using only one camera and generating videos from arbitrary perspectives through computer vision technology.
[0004] However, when only one camera is used to capture a scene, since the quality of the captured video cannot be guaranteed, the generated free-viewpoint video may have missing main parts, affecting the viewing effect. Summary of the Invention
[0005] This application provides a method, apparatus, electronic device, and readable storage medium for generating free-viewpoint video. By identifying the orientation of the main body in the first video and performing frame selection processing on the first video according to the orientation of the main body, a second video with the same or similar number of frames in each orientation category is obtained. The neural radiance field network is trained according to the second video, the orientation of the main body in each frame of the second video, and a preset loss function. The free-viewpoint video corresponding to the first video is output according to the trained model, thereby improving the problem that the generated free-viewpoint video may have missing main parts and affecting the viewing effect.
[0006] To achieve the above object, this application adopts the following technical solutions:
[0007] In a first aspect, a method for generating free-viewpoint video is provided, including: obtaining a first video, where the first video includes at least one main body; obtaining the orientation of the main body in each frame of the first video; obtaining a second video according to the orientation of the main body, where the number of frames in each orientation in the second video is the same or similar; constructing a neural radiance field network, training the neural radiance field network according to the second video, the orientation of the main body in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network; and inputting a preset viewpoint into the trained neural radiance field network to output a free-viewpoint video corresponding to the first video.
[0008] In an embodiment of the present application, the free-viewpoint video generation method can be applied to electronic devices, including mobile phones, tablet computers, servers, cloud servers, wearable devices, augmented reality / virtual reality devices, laptops, ultra-mobile personal computers, netbooks, personal digital assistants, and the like.
[0009] In a first aspect, by identifying the orientation of the subject in the first video and performing frame selection processing on the first video according to the subject orientation, a second video with the same or similar number of frames in each orientation category is obtained. Then, a neural radiance field network is constructed and trained based on the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network. Finally, a preset viewpoint is input into the trained neural radiance field network to output a free-viewpoint video corresponding to the first video. By identifying and ensuring that the number of frames in each orientation category in the second video is the same or similar, it is possible to avoid the situation of missing parts of the subject in the generated free-viewpoint video and improve the viewing effect of the generated free-viewpoint video.
[0010] In some possible implementation manners, the orientation of the subject includes multiple orientation categories, and each orientation category corresponds to an angular interval of the subject orientation.
[0011] Performing frame selection processing on the first video according to the orientation of the subject to obtain a second video includes: screening the frames in each orientation category according to the orientation category and the orientation of the subject in each frame to obtain a second video.
[0012] In some possible implementation manners, after performing frame selection processing on the first video according to the orientation of the subject to obtain a second video, the method further includes: performing frame selection processing on the second video to remove the blurred frames with blurred subjects in the second video to obtain a third video, and the number of frames in each orientation category in the third video is the same or similar.
[0013] Correspondingly, constructing a neural radiance field network and training the neural radiance field network based on the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network includes: constructing a neural radiance field network and training the neural radiance field network based on the third video, the orientation of the subject in each frame of the third video, and a preset loss function to obtain a trained neural radiance field network.
[0014] In some possible implementation manners, removing the blurred frames with blurred subjects in the second video to obtain a third video includes: obtaining the subject region in each frame of the second video through a preset subject segmentation algorithm. Obtaining the blurred region in the subject region through a preset blurred region segmentation algorithm. Determining the frames with the proportion of the blurred region in the subject region greater than a first preset threshold as blurred frames. Removing the blurred frames in the second video to obtain a third video.
[0015] In some possible implementation manners, when the number of frames of the third video is less than a second preset threshold, a neural radiance field network is constructed, and the neural radiance field network is trained according to the third video and a preset loss function to obtain a trained neural radiance field network, including: obtaining M blurred frames according to the orientation category, where M is the difference between the number of frames of the third video and the second preset threshold. The neural radiance field network is trained according to the third video, the M blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in each frame of the blurred frames, and a preset loss function to obtain a trained neural radiance field network, where the blurred regions in the blurred frames are ignored during training.
[0016] In some possible implementation manners, obtaining M blurred frames according to the orientation category includes: sequentially obtaining the corresponding blurred frames according to the orientation category until the number of obtained blurred frames is greater than or equal to M.
[0017] In some possible implementation manners, sequentially obtaining the corresponding blurred frames according to the orientation category includes: in each orientation category, preferentially obtaining the frame with the smallest proportion of the blurred region.
[0018] In some possible implementation manners, ignoring the blurred regions in the blurred frames during training includes: when training according to a preset loss function, if the ray of the neural radiance field passes through a blurred region, ignoring the loss of the ray.
[0019] In a second aspect, a free viewpoint video generation device is provided, including: an acquisition module, configured to acquire a first video, where the first video includes at least one subject. The acquisition module is further configured to acquire the orientation of the subject in each frame of the first video. A frame selection module, configured to perform frame selection processing on the first video according to the orientation of the subject to obtain a second video, where the number of frames in each orientation in the second video is the same or similar. A training module, configured to construct a neural radiance field network and train the neural radiance field network according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network. A generation module, configured to input a preset viewpoint into the trained neural radiance field network and output a free viewpoint video corresponding to the first video.
[0020] In some possible implementation manners, the orientation of the subject includes multiple orientation categories, and each orientation category corresponds to an angular interval of the subject orientation.
[0021] The frame selection module is specifically configured to screen the frames in each orientation category according to the orientation category and the orientation of the subject in each frame to obtain a second video.
[0022] In some possible embodiments, the frame selection module is further configured to perform frame selection processing on the second video, remove the blurred frames with blurred subjects in the second video, and obtain a third video, where the number of frames in each orientation category in the third video is the same or similar.
[0023] Correspondingly, the training module is specifically configured to construct a neural radiance field network, and train the neural radiance field network according to the third video, the orientation of the subject in each frame of the third video, and a preset loss function, so as to obtain a trained neural radiance field network.
[0024] In some possible embodiments, the frame selection module is specifically configured to obtain the subject area in each frame of the second video through a preset subject segmentation algorithm. Through a preset blurred area segmentation algorithm, obtain the blurred area in the subject area. Determine that the frame with the proportion of the blurred area in the subject area greater than a first preset threshold is a blurred frame. Remove the blurred frames in the second video to obtain a third video.
[0025] In some possible embodiments, when the number of frames of the third video is less than a second preset threshold, the training module is specifically configured to obtain M blurred frames according to the orientation category, where M is the difference between the number of frames of the third video and the second preset threshold. Train the neural radiance field network according to the third video, the M blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in each frame of the blurred frames, and a preset loss function, so as to obtain a trained neural radiance field network, where the blurred area in the blurred frames is ignored during training.
[0026] In some possible embodiments, the frame selection module is further configured to sequentially obtain the corresponding blurred frames according to the orientation category until the number of obtained blurred frames is greater than or equal to M.
[0027] In some possible embodiments, the frame selection module is further configured to preferentially obtain the frame with the smallest proportion of the blurred area in each orientation category.
[0028] In some possible embodiments, the training module is specifically configured to ignore the loss of the ray if the ray of the neural radiance field passes through the blurred area when training according to the preset loss function.
[0029] In a third aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor executes the computer program to perform the steps of the processing in the first aspect or any one of the methods in the first aspect.
[0030] In a fourth aspect, a chip is provided, including: a processor, configured to call and run a computer program from a memory, so that a device installed with the chip performs the steps of the processing in the first aspect or any one of the methods in the first aspect.
[0031] In a fifth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. When the program instructions are executed by a processor, the processor is caused to execute the steps of the processing in the first aspect or any one of the methods in the first aspect.
[0032] In a sixth aspect, a computer program product is provided. The computer program product includes computer program code. When the computer program code is run on an electronic device, the electronic device is caused to execute the steps of the processing in the first aspect or any one of the methods in the first aspect.
[0033] Among them, the beneficial effects of the second aspect to the sixth aspect can be referred to the first aspect, and will not be elaborated here. Description of the Drawings
[0034] Figure 1 is a schematic diagram of an application scenario of a free viewpoint video generation method provided by an embodiment of the present application;
[0035] Figure 2 is a schematic diagram of an application scenario of another free viewpoint video generation method provided by an embodiment of the present application;
[0036] Figure 3 is a hardware structure block diagram of an electronic device provided by an embodiment of the present application;
[0037] Figure 4 is a system structure block diagram of the electronic device provided by an embodiment of the present application;
[0038] Figure 5 is a schematic flowchart of a free viewpoint video generation method provided by an embodiment of the present application;
[0039] Figure 6 is a schematic flowchart of a free viewpoint video generation method provided by another embodiment of the present application;
[0040] Figure 7 is a schematic flowchart of implementing S605 in the free viewpoint video generation method provided by an embodiment of the present application;
[0041] Figure 8 is a schematic diagram of a blurred image - blurred region truth value pair in the free viewpoint video generation method provided by an embodiment of the present application;
[0042] Figure 9 is a structure block diagram of a free viewpoint video generation device provided by an embodiment of the present application;
[0043] Figure 10 is a schematic diagram of the structure of a chip provided by an embodiment of the present application. Detailed Embodiments
[0044] The technical solutions in the present application will be described below in conjunction with the accompanying drawings.
[0045] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B. The "and / or" herein is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.
[0046] Hereinafter, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of this embodiment, unless otherwise specified, the meaning of "a plurality of" is two or more.
[0047] Free-Viewpoint Video (FVV) is a video technology that allows a scene to be viewed from any angle, enabling users to freely change the viewing perspective and select different perspectives to view the objects or people in the scene. Generally, free-viewpoint video is obtained by synchronously recording videos from multiple cameras set at different angles and then analyzing and reconstructing the videos from each angle.
[0048] In order to expand the application scenarios of free-viewpoint video and lower the shooting threshold of free-viewpoint video, there is currently a method for shooting and generating monocular free-viewpoint video. Monocular free-viewpoint video refers to capturing a scene using only one camera and generating videos from arbitrary perspectives through computer vision technology.
[0049] For a terminal device, when shooting monocular free-viewpoint video, it can collect video through the camera of the terminal device and then process the video locally or in the cloud to obtain free-viewpoint video.
[0050] However, when the terminal device is held for shooting, due to environmental light, movement of the subject to be photographed, etc., the subject in the collected video may be blurred. In this case, the generated free-viewpoint video will also be blurred, affecting the viewing effect.
[0051] In view of this, the present application provides a free viewpoint video generation method, including: obtaining a first video, where the first video includes at least one subject; performing frame selection processing on the first video to remove blurred frames with blurred subjects in the first video, obtaining a second video; constructing a neural radiance field network, and training the neural radiance field network according to the second video and a preset loss function to obtain a trained neural radiance field network; inputting a preset viewpoint into the trained neural radiance field network, and outputting a free viewpoint video of the subject corresponding to the preset viewpoint.
[0052] In the present application, by identifying and removing the blurred frames in the first video, a second video is obtained. Then, a neural radiance field network is constructed and trained according to the second video and a preset loss function to obtain a trained neural radiance field network. Finally, a preset viewpoint is input into the trained neural radiance field network, and a free viewpoint video corresponding to the preset viewpoint is output. In the present application, by identifying and removing the blurred frames in the first video, it is possible to avoid blurring or missing parts of the subject in the generated free viewpoint video, and improve the viewing effect of the generated free viewpoint video.
[0053] Figure 1 It is a schematic diagram of an application scenario of a free viewpoint video generation method provided by an embodiment of the present application.
[0054] Reference Figure 1 is made to briefly describe the application scenario of the embodiment of the present application first.
[0055] Figure 1 An electronic device 100 is shown in. The electronic device 100 may have a camera. For example, it may be a mobile phone, a tablet computer, a wearable device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. In this scenario, the electronic device needs to have sufficient computing power, or needs to have a neural network processing unit (NPU) to provide a hardware environment for the free viewpoint video generation method.
[0056] Reference Figure 1 is made to where the user holds the electronic device 100 to take a picture, and the subject to be photographed may be as Figure 1 shown, is a person. Alternatively, the subject to be photographed may also be an animal, an object, etc.
[0057] As an example, the user can hold the electronic device 100 and take a full circle around the subject to be photographed, so as to obtain a video including various orientations of the subject to be photographed.
[0058] Then, the electronic device 100 can process the video locally, and then obtain the free viewpoint video of the subject to be photographed.
[0059] Figure 2 FIG. is a schematic diagram of an application scenario of another free viewpoint video generation method provided by an embodiment of the present application.
[0060] Figure 2 The electronic device 100 and the photographing device 101 are shown.
[0061] Among them, the electronic device 100 can be an electronic computer, a server, a cloud server, etc. with sufficient computing power or an NPU. The photographing device 101 can be a user terminal with a photographing function, such as a mobile phone, a tablet computer, a wearable device, an augmented reality / virtual reality device, a notebook computer, a super mobile personal computer, a netbook, a personal digital assistant, etc. Or, the photographing device 101 can also be a digital camera, a drone, an action camera, a camera, etc.
[0062] As an example, the user can hold the photographing device 101 or fix the photographing device 101 on a tripod, and the subject to be photographed rotates one full circle by itself, so that the photographing device 101 can obtain a video including various orientations of the subject to be photographed.
[0063] Or, as shown in Figure 1 the user can hold the photographing device 101 or control the photographing device 101 to take a full circle around the subject to be photographed, so as to obtain a video including various orientations of the subject to be photographed.
[0064] Then, the photographing device 101 sends the photographed video to the electronic device 100, and the electronic device 100 processes the received video, and then obtains the free viewpoint video of the subject to be photographed.
[0065] Figure 3 FIG. is a hardware structure block diagram of an electronic device provided by an embodiment of the present application.
[0066] In the present application, refer to Figure 1 and Figure 2, the electronic device 100 may include a mobile phone, a tablet computer, a handheld game console, a wearable device, an augmented reality (AR) / virtual reality (VR) device, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), etc. The electronic device in this application needs to have sufficient computing power, and its specific form is not limited in this application.
[0067] Reference Figure 3 , the electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0068] It can be understood that the structure illustrated in the embodiments of this application does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0069] As an example, when the electronic device 100 is a mobile phone or a tablet computer, it may include all the components shown in the figure, or only include some of the components shown in the figure. When the electronic device 100 is a server, it does not include the camera 193, the sensor module 180, etc.
[0070] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0071] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can save the instructions or data that the processor 110 has just used or recycled. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.
[0072] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0073] It can be understood that the interface connection relationships among the modules illustrated in the embodiments of the present invention are only illustrative descriptions and do not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0074] The charging management module 140 is used to receive a charging input from a charger.
[0075] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110.
[0076] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modulation and demodulation processor, and baseband processor, etc.
[0077] The electronic device 100 realizes the display function through the GPU, the display screen 194, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0078] The display screen 194 is used to display images, videos, etc.
[0079] The electronic device 100 can realize the shooting function through the ISP, the camera 193, the video codec, the GPU, the display screen 194, and the application processor, etc.
[0080] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also perform algorithm optimization on the noise, brightness, and skin color of the image. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP may be provided in the camera 193.
[0081] The camera 193 is used to capture still images or videos. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.
[0082] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of the electronic device 100 can be realized, such as image recognition, face recognition, speech recognition, text understanding, etc.
[0083] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100.
[0084] The internal memory 121 can be used to store computer-executable program code, and the executable program code includes instructions. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121 and / or the instructions stored in the memory provided in the processor.
[0085] The electronic device 100 can implement audio functions through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor, etc. Such as music playback, recording, etc.
[0086] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.
[0087] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. The microphone 170C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal. The headphone jack 170D is used to connect a wired headphone.
[0088] The pressure sensor 180A is used to sense a pressure signal and can convert the pressure signal into an electrical signal.
[0089] The gyroscope sensor 180B can be used to determine the motion posture of the electronic device 100. The barometric pressure sensor 180C is used to measure the barometric pressure. The magnetic sensor 180D includes a Hall sensor. The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). The distance sensor 180F is used to measure the distance.
[0090] The proximity light sensor 180G may include, for example, a light-emitting diode (LED) and a photodetector. The electronic device 100 emits infrared light outward through the light-emitting diode. The electronic device 100 uses a photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device 100. When insufficient reflected light is detected, the electronic device 100 can determine that there is no object near the electronic device 100.
[0091] The ambient light sensor 180L is used to sense the ambient light brightness. The fingerprint sensor 180H is used to collect fingerprints.
[0092] The temperature sensor 180J is used to detect the temperature. The touch sensor 180K, also known as a "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as a "touch control screen". The keys 190 include a power-on key, a volume key, etc. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light, which can be used to indicate the charging state, the change in battery power, and can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect the SIM card.
[0093] For the scenarios in the above examples, the operating system of the electronic device 100 may include, but is not limited to, operating systems such as Symbian, Android, Windows, MacOS, iOS, Blackberry, HarmonyOS, Linux, or Unix.
[0094] Figure 4 It is a system structure block diagram of the electronic device provided by the embodiment of the present application.
[0095] As an example, when the free viewpoint video generation method provided by the present application runs on the electronic device 100, the operating system of the electronic device 100 may be Android, and its system structure can be referred to Figure 4 .
[0096] Among them, the layered architecture divides software into several layers, and each layer has clear roles and divisions of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer.
[0097] The application layer may include a series of application packages.
[0098] As Figure 4 shown, the application packages may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0099] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0100] As Figure 4 shown, the application framework layer may include a window manager, a content provider, a view system, a telephone manager, a resource manager, a notification manager, etc.
[0101] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0102] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include videos, images, audio, dialed and answered calls, browsing history and bookmarks, phone books, etc.
[0103] The view system includes visible controls, such as controls for displaying characters, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a short message notification icon may include a view for displaying characters and a view for displaying pictures.
[0104] The telephone manager is used to provide the communication function of the electronic device. For example, the management of call states (including answering, hanging up, etc.).
[0105] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, etc.
[0106] The notification manager enables an application to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that a download is complete, a message reminder, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or a scroll bar text, such as a notification of a background-running application, or a notification that appears on the screen in the form of a dialogue window. For example, it can prompt text information in the status bar, emit a prompt sound, vibrate the electronic device, blink the indicator light, etc.
[0107] The Android Runtime includes core libraries and a virtual machine. The Android runtime is responsible for the scheduling and management of the Android system.
[0108] The core libraries consist of two parts: one part is the functional functions that the Java language needs to call, and the other part is the core libraries of Android.
[0109] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as the management of object life cycles, stack management, thread management, security and exception management, and garbage collection.
[0110] The system libraries can include multiple functional modules. For example: the surface manager, the media libraries, the 3D graphics processing library (e.g., OpenGL ES), the 2D graphics engine (e.g., SGL), etc.
[0111] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0112] The media libraries support the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media libraries can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0113] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0114] The 2D graphics engine is the drawing engine for 2D drawing.
[0115] The kernel layer is the layer between the hardware and the software. The kernel layer includes at least a display driver, a camera driver, an audio driver, and a sensor driver.
[0116] Figure 5 It is a schematic flow chart of the free viewpoint video generation method provided by the embodiments of this application.Figure 5 Take the scenario shown in Figure 2 as an example for illustration. That is, the electronic device is a server, and the shooting device is a user terminal, such as a mobile phone, a tablet computer, etc.
[0117] Referring to Figure 5 , the free viewpoint video generation method includes:
[0118] S501. The user terminal shoots a first video and sends it to the server.
[0119] S502. The server receives the first video.
[0120] In some possible implementation manners, when shooting the first video by the user terminal, there is at least one subject in the first video. The subject can be a human body, an animal, a plant or an object. When shooting, the user terminal can be fixed for shooting, and the subject rotates or moves to obtain the first video including various orientations of the subject. Or, the user terminal can also be held and rotated around the subject for one week to obtain the first video including various orientations of the subject.
[0121] In some possible implementation manners, the user terminal can be communicatively connected to the server through a network, and then send the first video to the server. Or, the user terminal can also upload the first video to a specified storage space through a network, such as a cloud disk, a database, etc. Then, the server reads the first video from the specified storage space.
[0122] Wherein, the network includes but is not limited to a wireless local area network, a cellular communication network, a wired network, etc.
[0123] S503. The server obtains the orientation of the subject in each frame of the first video.
[0124] In some possible implementation manners, the subject can be a human body, an animal or an object. Taking the identification of the human body orientation as an example, to obtain the orientation of the subject in each frame of the first video, methods such as pose estimation method based on key points, estimation method based on 3D reconstruction, appearance-based classifier, head orientation-based estimation method, motion-based analysis method, etc. can be used.
[0125] As an example, in the present application, a pose estimation method based on key points can be used to identify the orientation of the subject in each frame of a video.
[0126] In the pose estimation method based on key points, a deep learning model (such as OpenPose, AlphaPose, etc.) can be used to detect the key points of the human body (such as the head, shoulders, hips, etc.). Then, according to the relative positions and angles between these key points, the orientation of the human body is inferred. For example, if the key points of the head and shoulders are both facing the camera, it can be inferred that the person is facing forward.
[0127] For example, for a complete human body, it may include 23 key points and a root node. The root node can be the center point of the human body or a pre-set reference point for determining the positions of other key points. The key points can be the main skeletal joints of the human body and additional feature points, such as: the head (including eyes, nose, ears, etc.), shoulders, elbows, wrists, hips, knees, ankles, etc. The relative positions and angles between these key points can be used to deduce the shape, orientation, etc. of the human body.
[0128] Among them, the orientation can include the vertical orientation and horizontal orientation of the human body, as well as the relative position parameters between the camera and the human body.
[0129] S504. The server performs frame selection processing on the first video according to the orientation of the subject to obtain a second video.
[0130] In some possible implementation manners, the number of frames in each orientation in the second video is the same or similar. The orientation of the subject can be determined according to the horizontal orientation shown in S503. When the number of video frames in each orientation is the same or similar, the training effect on the neural radiance field network is better, and the probability of the subject missing in the free-viewpoint video output by the model is lower. Therefore, the first video can be screened to obtain a second video in which the number of frames in each orientation is the same or similar.
[0131] In some possible implementation manners, the orientation of the subject is divided into multiple orientation categories, and each orientation category corresponds to an angular interval of the subject's orientation.
[0132] As an example, the frames in each orientation category can be screened according to the orientation category and the orientation of the subject in each frame to obtain a second video.
[0133] For example, the subject's orientation can be divided according to intervals. Among them, the horizontal orientation is a total of 360°, and it can be divided into one orientation every 45°, that is, a total of 8 orientation categories are divided (front, back, left, right, front left, front right, back left, back right).
[0134] Among them, the number of frames in each orientation category in the second video is similar, including that the difference between the number of frames in each orientation in the second video and the average value of the number of frames in each orientation is less than a preset difference or a preset ratio.
[0135] For example, if the subject in the second video includes 8 orientations, the average value of the number of frames in each orientation category is 100, and the preset difference is 5. That is, when the number of frames in each orientation category is between 95 and 105, it is determined that the number of frames in each orientation category in the second video is similar.
[0136] If the subjects in the second video include 8 orientation categories, the average number of frames in each orientation category is 200, and the preset ratio is 5%. That is, when the number of frames in each orientation category is between 190 and 210, it is determined that the number of frames of each orientation category in the second video is similar.
[0137] In this embodiment, by screening the first video, a second video with the same or similar number of frames in each orientation is obtained. Since when training a neural radiance network, the more balanced the number of frames in each orientation, the better the output effect of the trained model, and the lower the probability of subject loss in the free-viewpoint video output by the model. Therefore, the free-viewpoint video output by the neural radiance network trained with the second video has a better effect and a better viewing experience.
[0138] S505. Construct a neural radiance field network, and train the neural radiance field network according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network.
[0139] Among them, the neural radiance fields (NeRF) network is a deep learning method for 3D scene reconstruction and rendering. NeRF generates high-quality synthetic images by learning the continuous volume density and color distribution of the scene. It can capture the complex geometric structure and lighting details of the scene from a set of sparse 2D images (such as a video).
[0140] NeRF uses a fully connected neural network to predict the volume density value and color value at a given 3D spatial position (coordinates) and viewing direction (represented by azimuth and elevation angles). For example, NeRF can be implemented using a multi-layer perceptron (MLP). The input of NeRF is the spatial coordinates and the direction vector (viewpoint) of the viewing angle, and the output is the color (such as RGB value) and the volume density.
[0141] During the training process, NeRF obtains the frames in the second video and the relative position parameters between the camera and the subject in each frame. By constructing rays emitted from the camera position and passing through each pixel, and then sampling multiple points along the rays through a grid, and integrating the predicted colors and densities of these points to obtain the final pixel color.
[0142] The training objective can be to minimize the difference between the color predicted by the network and the true image pixel value through a preset loss function.
[0143] As an example, the prediction function of NeRF can be the following formula:
[0144]
[0145] Among them, x is the coordinate of each point in the given 3D space, and p i is the pose of the i-th frame, and F c (T(x, p i )) represents the color value and volume density value of all points in the given 3D space. e i is the preset viewing point, that is, the viewing angle. Through the volume rendering equation (T), NeRF will predict a color value and volume density value for each x, and integrate the points on each ray to obtain the value of each pixel (one ray corresponds to one pixel). After obtaining all the pixels, the image I i at the viewing point can be obtained. i .
[0146] As an example, the preset loss function can be the mean absolute error loss function (L i loss function) calculated pixel by pixel according to I 1 .
[0147] S506. The server inputs the preset viewing point into the trained neural radiance field network and outputs a free-viewpoint video corresponding to the first video.
[0148] In some possible implementation manners, the free-viewpoint video corresponding to the first video includes the main body in the first video.
[0149] Among them, the viewing point refers to the position and angle of viewing the main body. The viewing point can include the spatial coordinates and the direction vector of the viewing angle. The spatial coordinates are used to indicate the position of the viewing main body, and the direction vector of the viewing angle is used to indicate the angle of the viewing main body. Inputting the viewing point into the trained NeRF, the trained NeRF can output a free-viewpoint video with this viewing point as the viewing angle.
[0150] Reference Figure 2 and Figure 5 , the preset viewing point can be sent by the user terminal in response to the user operation to the server, or can be the default viewing point set in the server.
[0151] For example, when a free-viewpoint video application is running on the user terminal, before sending the first video to the server, the application can prompt the user to select the viewing point of the generated free-viewpoint video, such as directly in front (facing the front of the main body), directly behind (facing the back of the main body), diagonally in front (45° to the left or right in front of the main body), diagonally behind (45° to the left or right behind the main body), etc. Then, the free-viewpoint video application responds to the user operation and sends the first video and the spatial coordinates and the direction vector of the viewing angle of the user-selected viewing point to the server.
[0152] In some possible implementation manners, there may be multiple preset viewpoints, and the multiple preset viewpoints may form a camera trajectory. The viewing angle in the output free viewpoint video may move along the camera trajectory to obtain a free viewpoint video with a moving viewing angle.
[0153] In some possible implementation manners, the preset body pose may also be input into the trained neural radiance field network, so that the body in the free viewpoint video output by the trained neural radiance field network will be presented in the preset body pose.
[0154] S507. The server sends the free viewpoint video to the user terminal.
[0155] S508. The user terminal plays the free viewpoint video.
[0156] In some possible implementation manners, referring to S506, the server may send the generated free viewpoint video to the free viewpoint video application in the user terminal, and the free viewpoint video application displays the free viewpoint video on the screen after receiving it.
[0157] It should be noted that in this embodiment, S503, S504, and S506 are executed on the server side. However, in some other possible implementation manners, S503, S504, and S506 may also be executed on the user terminal.
[0158] For example, the user terminal may execute S503 and S504 after shooting the first video to obtain the second video. Then, the second video is sent to the server, and the server executes S505 according to the second video. Then, the trained model parameters are sent to the user terminal, and the user terminal executes S506.
[0159] Figure 6 It is a schematic flowchart of a free viewpoint video generation method provided by another embodiment of the present application. Figure 6 In Figure 2 The illustrated scenario is used as an example for description. That is, the electronic device is a server, and the shooting device is a user terminal, such as a mobile phone, a tablet computer, etc.
[0160] Referring to Figure 6 , the free viewpoint video generation method includes:
[0161] S601. The user terminal shoots the first video and sends it to the server.
[0162] S602. The server receives the first video.
[0163] S603. The server obtains the orientation of the body in each frame of the first video.
[0164] S604. The server performs frame selection processing on the first video according to the orientation of the body to obtain the second video.
[0165] In this embodiment, the implementation manners of S601 to S604 are the same as those of S501 to S504, and will not be elaborated here.
[0166] S605. The server performs frame selection processing on the second video, removes the blurred frames with blurred subjects in the second video, and obtains a third video.
[0167] In some possible implementation manners, there may be blurred frames with blurred subjects in the second video. Using the second video to train NeRF will result in blurred subjects or missing subjects in the free viewpoint video output by the trained NeRF.
[0168] Figure 7 It is a schematic flowchart for implementing S605 in the free viewpoint video generation method provided by the embodiments of the present application.
[0169] In some possible implementation manners, with reference to Figure 7 , the steps for implementing S605 include:
[0170] S6051. Through a preset subject segmentation algorithm, obtain the subject region in each frame of the second video.
[0171] In some possible implementation manners, the subject segmentation algorithm is also called foreground segmentation or object segmentation, which can separate the subject (foreground) in an image or video from the background. Common subject segmentation algorithms include the GraphCut algorithm, Random Forest, deep learning models, the Level Set Method, the Active Contour Model, etc.
[0172] As an example, when the subject is a human body and the subject segmentation algorithm is a deep learning model, subject segmentation can be performed by first training the deep learning model. For example, training samples can be obtained first. The training samples include images containing human bodies, and these images are labeled. The training samples can also be preprocessed, including scaling, normalization, and data augmentation, etc.
[0173] Then, determine a deep learning model and train it according to the training samples to obtain a trained model with subject segmentation ability.
[0174] Finally, input each frame in the second video into the trained model in turn to obtain the subject region in each frame of the second video.
[0175] S6052. Through a preset blurred region segmentation algorithm, obtain the blurred regions in the subject region.
[0176] In some possible implementation manners, the preset fuzzy region segmentation algorithm may be based on the body segmentation algorithm in S6051, and a fuzzy image-fuzzy region truth value pair is added for training to obtain a model with the ability to segment fuzzy regions.
[0177] Figure 8 It is a schematic diagram of a fuzzy image-fuzzy region truth value pair in the free viewpoint video generation method provided by an embodiment of the present application.
[0178] As an example, referring to Figure 8 a in Figure 8 a in shows a blurred image 81 and a blurred region 82. Referring to Figure 8 a in , the blurred image 81 may be a human body image including a blurred region, and the blurred region 82 may be marked by a frame.
[0179] In some possible implementation manners, when generating a fuzzy image-fuzzy region truth value pair, a clear image may be obtained first, and then a mask with a random shape is generated on the clear image. The image corresponding to the region in the mask is subjected to motion blur or focus blur, and then a blurred image is synthesized.
[0180] For example, the blurred image 81 may be generated first according to a clear human body image, and Figure 8 b in shows the mask 83, and then the image corresponding to the region shown in the mask 83 is blurred to obtain the blurred image 81. Among them, the mask 83 may be used as the truth value of the blurred region of the blurred image 81.
[0181] S6053. Determine that a frame in which the proportion of the blurred region in the body region is greater than a first preset threshold is a blurred frame.
[0182] S6054. Remove the blurred frames in the second video to obtain a third video.
[0183] In some possible implementation manners, to calculate the proportion of the blurred region in the body region, the number of body pixels in the body region may be obtained. Then, according to the predicted blurred region, the number of blurred pixels including the body is obtained. The ratio of the number of blurred pixels to the number of body pixels is used as the proportion of the blurred region in the body region.
[0184] As an example, to obtain the number of blurred pixels including the body in the blurred region, the mask of the predicted blurred region may be compared with the body region to obtain the number of pixels that appear in both the body region and the blurred region.
[0185] In some possible implementation manners, the first preset threshold can be set in intervals according to the number of frames of the second video. For example, when the number of frames of the second video is greater than 600 frames and less than 800 frames, the first preset threshold can be set to 20%, that is, only the frames with a blurred area greater than 20% will be determined as blurred frames. When the number of frames of the second video is greater than 800 frames and less than 1000 frames, the first preset threshold can be set to 10%, that is, the frames with a blurred area greater than 20% will be determined as blurred frames. When the number of frames of the second video is greater than 1000 frames, the first preset threshold can be set to 5%, that is, the frames with a blurred area greater than 5% will be determined as blurred frames.
[0186] In this embodiment, after determining the blurred frames in the second video, the blurred frames can be removed, and only the clear frames are retained. The clear frames are used to train the neural radiance field network, so that the free-viewpoint video output by the trained neural radiance field network has fewer cases of blurred or missing subjects.
[0187] In this embodiment, dynamically setting the first preset threshold can further reduce the probability of blurring or missing subjects in the generated free-viewpoint video when there are enough redundant frames, and improve the viewing effect of the generated free-viewpoint video.
[0188] S606. The server determines whether the number of frames of the third video is less than the second preset threshold. If it is less, S608 is executed; if it is greater than or equal to, S607 is executed.
[0189] In some possible implementation manners, the second preset threshold can be determined according to the minimum number of frames required for training the neural radiance field network. For example, if at least 600 frames are required for training the neural radiance field network, the second preset threshold can be an integer greater than or equal to 600, such as 600, 650, or 700, etc.
[0190] S607. The server trains the neural radiance field network according to the third video, the orientation of the subject in each frame of the third video, and a preset loss function, and obtains the trained neural radiance field network.
[0191] In this embodiment, the implementation manner of S607 is the same as that of S505, and will not be elaborated here.
[0192] S608. The server obtains M blurred frames according to the orientation category.
[0193] In some possible implementation manners, M is the difference between the number of frames of the third video and the second preset threshold. When obtaining the blurred frames, the corresponding blurred frames can be obtained in sequence according to the orientation category until the number of obtained blurred frames is greater than or equal to M.
[0194] Among them, in each orientation category, the frame with the smallest proportion of the blurred area can be preferentially obtained.
[0195] For example, in the third video, the subject includes 8 orientation categories and M is 50. Then, starting from the one with the fewest frames among the 8 orientation categories, the blurred frames corresponding to each orientation category can be obtained in sequence. Among them, 2 orientation categories obtain 7 frames, and 6 orientation categories obtain 6 frames.
[0196] In each orientation category, the blurred frames belonging to this orientation category can be sorted in ascending order according to the proportion of the blurred area, and the frame with the smallest proportion of the blurred area can be preferentially obtained.
[0197] S609. The server trains the neural radiance field network according to the third video, M blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in the blurred frames, and a preset loss function, and obtains the trained neural radiance field network. Among them, during training, the blurred area in the blurred frames is ignored.
[0198] In some possible implementation manners, the method for training the neural radiance field network in S609 is similar to that in S505, and will not be elaborated here.
[0199] It should be noted that during training, ignoring the blurred area in the blurred frames means that when the ray of the constructed neural radiance field passes through the blurred area, the loss caused by this ray is ignored.
[0200] For example, referring to the formula in S505, the obtained blurred frames with i blurred angles can be filtered through the mask in S6052. Then, the filtered images without the blurred area are used for calculating the loss function. In this way, the interference of the blurred area on training can be excluded.
[0201] By ignoring the blurred area in the blurred frames during training, the impact caused by the added blurred frames to ensure the minimum number of frames for training can be reduced. It can not only train the neural radiance field network normally, but also reduce the impact of the blurred area on training, improve the training effect, and make the viewing effect of the free viewpoint video output by the trained model better.
[0202] S610. The server inputs the preset viewpoints into the trained neural radiance field network and outputs the free viewpoint video corresponding to the first video.
[0203] S611. The server sends the free viewpoint video to the user terminal.
[0204] S612. The user terminal plays the free viewpoint video.
[0205] In this embodiment, the implementation manners of S610 and S612 are the same as those of S506 to S508, and thus will not be elaborated herein.
[0206] It should be noted that in this embodiment, S603 - S606, S608, and S610 are executed on the server side. However, in some other possible implementation manners, S603 - S606, S608, and S610 may also be executed on the user terminal.
[0207] For example, the user terminal may execute S603 - S605 after shooting the first video to obtain the third video.
[0208] If the number of frames of the third video is greater than the second preset threshold, the user terminal sends the third video to the server, and the server executes S607. If the number of frames of the third video is less than the second preset threshold, then after the user terminal executes S608, it sends the video to the server, and the server executes S607.
[0209] The server executes S607 or S609 according to the received video, and then sends the trained model parameters to the user terminal, and the user terminal executes S610.
[0210] It should be understood that the above examples are for helping those skilled in the art to understand the embodiments of the present application, rather than limiting the embodiments of the present application to the specific numerical values or specific scenarios illustrated.
[0211] Obviously, those skilled in the art can make various equivalent modifications or changes according to the above examples, and such modifications or changes also fall within the scope of the embodiments of the present application.
[0212] Corresponding to the free viewpoint video generation method provided in the above embodiments, Figure 9 is a structural block diagram of a free viewpoint video generation device provided by an embodiment of the present application. For the sake of convenience of description, only the parts related to the embodiments of the present application are shown.
[0213] Referring to Figure 9 , a free viewpoint video generation device is provided, including:
[0214] An acquisition module 91, configured to acquire a first video, where the first video includes at least one subject.
[0215] The acquisition module 91 is further configured to acquire the orientation of the subject in each frame of the first video.
[0216] A frame selection module 92, configured to perform frame selection processing on the first video according to the orientation of the subject to obtain a second video, where the number of frames in each orientation of the second video is the same or close.
[0217] A training module 93 for constructing a neural radiance field network, training the neural radiance field network according to a second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network.
[0218] A generation module 94 for inputting a preset viewpoint into the trained neural radiance field network and outputting a free viewpoint video corresponding to the first video.
[0219] In some possible implementation manners, the orientation of the subject includes multiple orientation categories, and each orientation category corresponds to an angular interval of the subject orientation.
[0220] A frame selection module 92 specifically for screening the frames in each orientation category according to the orientation category and the orientation of the subject in each frame to obtain the second video.
[0221] In some possible implementation manners, the frame selection module 92 is further configured to perform frame selection processing on the second video, remove the blurred frames with blurred subjects in the second video to obtain a third video, and the number of frames in each orientation category in the third video is the same or close.
[0222] Correspondingly, the training module 93 is specifically for constructing a neural radiance field network, training the neural radiance field network according to the third video, the orientation of the subject in each frame of the third video, and a preset loss function to obtain a trained neural radiance field network.
[0223] In some possible implementation manners, the frame selection module 92 is specifically for obtaining the subject region in each frame of the second video through a preset subject segmentation algorithm. Obtaining the blurred region in the subject region through a preset blurred region segmentation algorithm. Determining the frames with the proportion of the blurred region in the subject region greater than a first preset threshold as blurred frames. Removing the blurred frames in the second video to obtain the third video.
[0224] In some possible implementation manners, when the number of frames of the third video is less than a second preset threshold, the training module 93 is specifically for obtaining M blurred frames according to the orientation category, where M is the difference between the number of frames of the third video and the second preset threshold. Training the neural radiance field network according to the third video, the M blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in each frame of the blurred frames, and a preset loss function to obtain a trained neural radiance field network, where the blurred regions in the blurred frames are ignored during training.
[0225] In some possible implementation manners, the frame selection module 92 is further for sequentially obtaining the corresponding blurred frames according to the orientation category until the number of obtained blurred frames is greater than or equal to M.
[0226] In some possible implementation manners, the frame selection module 92 is further configured to preferentially obtain a frame with the smallest proportion of the blurred area in each orientation category.
[0227] In some possible implementation manners, the training module 93 is specifically configured to, when training according to a preset loss function, if a ray of the neural radiance field passes through a blurred area, ignore the loss of the ray.
[0228] Those skilled in the art can clearly understand that, for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.
[0229] Each functional unit and module in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. For example, a "module" can be a software program, a hardware circuit, or a combination of both to implement the above functions. The hardware circuit may include an application specific integrated circuit (ASIC), an electronic circuit, a processor (such as a shared processor, a dedicated processor, or a group of processors, etc.) for executing one or more software or firmware programs, a memory, a merged logic circuit, and / or other suitable components to support the described functions.
[0230] Therefore, the modules in the examples described in the embodiments of the present application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0231] In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of the present application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0232] It should be understood that the hardware system and the chip in the embodiments of the present application can execute various methods for entering the long standby mode in the foregoing embodiments of the present application, that is, the specific working processes of the following various products can refer to the corresponding processes in the foregoing method embodiments.
[0233] The embodiment of the present application also provides another electronic device, including a processor and a memory.
[0234] The memory is used to store a computer program that can run on the processor.
[0235] The processor is used to execute the steps of the method for entering the long standby mode as described above.
[0236] The embodiment of the present application also provides a computer-readable storage medium, in which computer instructions are stored; when the computer-readable storage medium runs on an electronic device, the electronic device is enabled to execute the method as described above.
[0237] The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.).
[0238] The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated media.
[0239] The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid state disk (SSD), etc.).
[0240] The embodiment of the present application also provides a computer program product containing computer instructions, which enables an electronic device to execute the foregoing technical solution when running on the electronic device.
[0241] Figure 10 It is a schematic structural diagram of a chip provided by the embodiment of the present application. Figure 10 The shown chip can be a general-purpose processor or a dedicated processor. The chip includes a processor 1001. Among them, the processor 1001 is used to support the electronic device to execute the foregoing technical solution.
[0242] Optionally, the chip further includes a transceiver 1002, and the transceiver 1002 is used to accept the control of the processor 1001 and support the communication device to execute the foregoing technical solution.
[0243] Optionally, Figure 10The chip shown may also include: a storage medium 1003.
[0244] It should be noted that Figure 10 The chip shown can be implemented using the following circuits or devices: one or more field programmable gate arrays (FPGAs), programmable logic devices (PLDs), controllers, state machines, gate logic, discrete hardware components, any other suitable circuits, or any combination of circuits capable of performing the various functions described throughout this application.
[0245] The electronic device, computer storage medium, computer program product, and chip provided in the embodiments of the present application above are all used to execute the method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the method provided above, and will not be elaborated here.
[0246] It should be understood that the above is only to help those skilled in the art better understand the embodiments of the present application, rather than to limit the scope of the embodiments of the present application. Those skilled in the art can obviously make various equivalent modifications or changes according to the above examples given.
[0247] For example, in some embodiments of the above method, some steps may not be necessary, or some steps may be newly added, etc. Or any combination of any two or any multiple of the above embodiments. Such modified, changed, or combined solutions also fall within the scope of the embodiments of the present application.
[0248] It should also be understood that the above description of the embodiments of the present application focuses on emphasizing the differences between the various embodiments. The same or similar parts not mentioned can be referred to each other. For the sake of brevity, they will not be elaborated here.
[0249] It should also be understood that the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0250] It should also be understood that in the embodiments of the present application, "pre-set" and "pre-defined" can be implemented by pre-saving corresponding codes, tables, or other means that can be used to indicate relevant information in a device (for example, including an electronic device). The present application does not limit its specific implementation manner.
[0251] It should also be understood that the division of the manners, situations, categories, and embodiments in the embodiments of the present application is only for the convenience of description and should not constitute a special limitation. The features in various manners, categories, situations, and embodiments can be combined without conflict.
[0252] It should also be understood that in various embodiments of the present application, if there is no special indication and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be mutually referred to, and the technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships.
[0253] Finally, it should be noted that the above are only specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A free viewpoint video generation method, characterized in that, the method comprises: obtaining a first video, where the first video includes at least one subject; obtaining the orientation of the subject in each frame of the first video; performing frame selection processing on the first video according to the orientation of the subject to obtain a second video, where the number of frames in each orientation in the second video is the same or similar; constructing a neural radiance field network, and training the neural radiance field network according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network; inputting a preset viewpoint into the trained neural radiance field network and outputting a free viewpoint video corresponding to the first video; wherein, after performing frame selection processing on the first video according to the orientation of the subject to obtain a second video, the method further comprises: performing frame selection processing on the second video to remove the blurred frames with blurred subjects in the second video to obtain a third video, where the number of frames in each orientation category in the third video is the same or similar; wherein, the performing frame selection processing on the second video to remove the blurred frames with blurred subjects in the second video to obtain a third video includes: obtaining the subject region in each frame of the second video through a preset subject segmentation algorithm; obtaining the blurred region in the subject region through a preset blurred region segmentation algorithm; determining the frames with the proportion of the blurred region in the subject region greater than a first preset threshold as the blurred frames; removing the blurred frames in the second video to obtain the third video; wherein, the constructing a neural radiance field network, and training the neural radiance field network according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network includes: when the number of frames of the third video is less than a second preset threshold, obtaining M blurred frames according to the orientation category, where M is the difference between the number of frames of the third video and the second preset threshold; training the neural radiance field network according to the third video, M blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in each frame of the blurred frames, and a preset loss function to obtain the trained neural radiance field network, where the blurred regions in the blurred frames are ignored during training.
2. The method according to claim 1, characterized in that, the orientation of the subject includes multiple orientation categories, and each orientation category corresponds to an angular interval of the subject orientation; the performing frame selection processing on the first video according to the orientation of the subject to obtain a second video includes: screening the frames in each orientation category according to the orientation category and the orientation of the subject in each frame to obtain the second video.
3. The method according to claim 1 or 2, characterized in that, The constructed neural radiance field network is trained according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network, and further includes: When the number of frames of the third video is greater than or equal to the second preset threshold, a neural radiance field network is constructed, and the neural radiance field network is trained according to the third video, the orientation of the subject in each frame of the third video, and a preset loss function to obtain a trained neural radiance field network.
4. The method according to claim 1 or 2, characterized in that the obtaining M blurred frames according to the orientation category includes: sequentially obtaining the corresponding blurred frames according to the orientation category until the number of obtained blurred frames is greater than or equal to M.
5. The method according to claim 4, characterized in that the sequentially obtaining the corresponding blurred frames according to the orientation category includes: in each orientation category, preferentially obtaining the frame with the smallest proportion of the blurred area.
6. The method according to claim 1 or 2, characterized in that the ignoring of the blurred area in the blurred frames during training includes: when training according to the preset loss function, if the ray of the neural radiance field passes through the blurred area, the loss of the ray is ignored.
7. A free viewpoint video generation device, characterized in that the device includes: an acquisition module for acquiring a first video, where the first video includes at least one subject; the acquisition module is further configured to acquire the orientation of the subject in each frame of the first video; a frame selection module for performing frame selection processing on the first video according to the orientation of the subject to obtain a second video, where the number of frames in each orientation in the second video is the same or similar; a training module for constructing a neural radiance field network and training the neural radiance field network according to the second video, the orientation of the subject in each frame of the second video, and a preset loss function to obtain a trained neural radiance field network; a generation module for inputting a preset viewpoint into the trained neural radiance field network and outputting a free viewpoint video corresponding to the first video; wherein the frame selection module is further configured to perform frame selection processing on the second video to remove the blurred frames with subject blur in the second video to obtain a third video, where the number of frames in each orientation category in the third video is the same or similar; wherein the frame selection module is specifically configured to: obtain the subject area in each frame of the second video through a preset subject segmentation algorithm; obtain the blurred area in the subject area through a preset blurred area segmentation algorithm; determine the frame with the proportion of the blurred area in the subject area greater than the first preset threshold as the blurred frame; remove the blurred frames in the second video to obtain the third video; wherein the training module is specifically configured to: When the number of frames of the third video is less than a second preset threshold, obtain M frames of the blurred frames according to the orientation category, where M is the difference between the number of frames of the third video and the second preset threshold; Train the neural radiance field network according to the third video, M frames of the blurred frames, the orientation of the subject in each frame of the third video, the orientation of the subject in each frame of the blurred frames, and a preset loss function to obtain the trained neural radiance field network, where the blurred areas in the blurred frames are ignored during training.
8. An electronic device, characterized in that, it includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and when the processor executes the computer program, the method described in any one of claims 1-6 is implemented.
9. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method described in any one of claims 1-6.
Citation Information
Patent Citations
Dynamic human body free viewpoint video generation method based on neural radiation field and device thereof
CN113099208A