Generating 3D video using 2d images and audio having a background related to metadata
Through machine learning models, analyzing 2D images and video information, generating 3D videos and controlling roles and environments using input devices, solving the interaction problem of 2D to 3D videos, and achieving dynamic updates and interaction effects.
Patent Information
- Application Number
- CN202380085902.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-26
- Filing Date
- 2023-11-02
- Publication Date
- 2025-07-29
AI Technical Summary
The prior art is difficult to convert 2D images and videos into 3D videos with interactive features, and it is impossible to effectively interact with the characters and environments in the 3D videos.
Through machine learning models, analyze information in 2D images and videos, generate 3D videos, and use input devices to control 3D characters and environments to realize dynamic updates and interactions of 3D videos.
It realizes the interaction between users and characters and environments in 3D videos, and can present dynamically updated 3D content on a 3D display, enhancing the user experience.
Smart Images

Figure CN120390943A_ABST
Abstract
Description
Technical Field
[0001] This application generally relates to 3D spatial memories, and more particularly, to generating a 3D video with a timeline based on end-user 2D videos. Background Art
[0002] As understood herein, people ("end-users") generate and store 2D images and videos of the same subject (such as a child) over a long period of time (such as several years). Summary of the Invention
[0003] This principle also understands that the storage reproduction of viewing personal audio-visual files can be enhanced by automatically converting content that is essentially a partial history of the subject into 3D with interactive features. End-users can watch and interact with copies of their family and friends in a 3D video presented, for example, on a holographic display device. The system generates 3D interactive content based on existing videos and photos. The main purpose of the 3D content is to use it for 3D viewing devices such as 3D spatial displays, holographic displays, and VR headsets. Input devices such as gamepads or mice and keyboards can be used to control the created characters and environments.
[0004] The system analyzes the situation / story based on various information contained in the video / photo, such as human movement, facial expressions, environment type, location, and time. Based on this information, objects and features (such as humans, animations, environmental maps, lighting, sound effects, etc.) in the prepared 3D content sources are identified for automatically and transparently generating the final 3D video for the user. A part of the textures and data is collected from the 2D video, including face textures, wall and floor textures, audio such as human voices in the 2D video, and environmental sounds.
[0005] The 3D memory content is continuously updated. Characters (such as parents or children) grow with the real user in the 3D video. When the user wants to view past memories, the user can control the time in the virtual world being viewed.
[0006] Accordingly, a device includes at least one processor configured to access 2D image information in a storage device and use the 2D image information to at least partially generate a 3D interactive video by matching at least one feature of at least one object in the 2D image information with at least one 3D model.
[0007] In some examples, the processor is configured to receive at least one interaction signal from at least one input device and at least partially change the presentation of the 3D interactive video based on the interaction signal.
[0008] The input device can be, for example, at least one touch element and / or at least one microphone on at least one computer simulation controller.
[0009] In some implementations, the processor can be programmed to change the presentation of the 3D interactive video in response to an incoming call.
[0010] The features of the 2D image information on which the 3D video is generated can include one or more of human image texture, human image motion, human image facial expression, and environment type.
[0011] In some examples, the processor can be programmed to change the presentation of the 3D interactive video in response to an input from a time slider input element.
[0012] In another aspect, a device includes at least one computer storage device that is not a transient signal and further includes instructions executable by at least one processor to identify at least one texture of at least one human image in a 2D image. The instructions are executable to generate a 3D character using the texture. The instructions are further executable to identify at least one environmental background type in the 2D image and, at least in part based on the environmental background type in the 2D image, select an environmental background type model for generating a 3D environmental background. The instructions are executable to merge the 3D environmental background with the 3D character to generate a 3D video.
[0013] In another aspect, a method includes accessing a 2D image in a 3D space memory and automatically creating a 3D video at least in part using an environmental background model that is selected based on the environmental background in the 2D image classified by a machine learning (ML) model and the texture of the face image in the 2D image.
[0014] Details regarding both the structure and operation of the present application can be best understood with reference to the accompanying drawings, in which like reference numerals refer to like parts and in which: BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 is a block diagram of an example system according to the present principle;
[0016] Figure 2 shows an example system architecture consistent with the present principle;
[0017] Figure 3 shows an example of an example holographic 3D display;
[0018] Figure 4 shows example general logic in an example flowchart format;
[0019] Figure 5An example 3D image generation logic is shown in an example flowchart format;
[0020] Figure 6 and Figure 7 An example screenshot is shown, which shows how the environmental background in 3D can be different from the 2D image that generates the 3D image;
[0021] Figure 8 An example 3D environment generation logic is shown in an example flowchart format;
[0022] Figure 9 An example 3D environment modification logic is shown in an example flowchart format;
[0023] Figure 10 An example 3D image generation logic based on a user ID is shown in an example flowchart format;
[0024] Figure 11 An example detailed use case logic is shown in an example flowchart format;
[0025] Figure 12 and Figure 13 An example screenshot is shown, which shows watching a 3D video based on time;
[0026] Figure 14 A 3D video in response to a phone call is shown; and
[0027] Figure 15 Further description of the details of an example implementation is provided. DETAILED DESCRIPTION
[0028] The present disclosure generally relates to computer ecosystems, which include aspects of a network of consumer electronics (CE) devices such as, but not limited to, a computer gaming network. The systems herein may include server and client components that may be connected via a network such that data may be exchanged between the client and server components. The client components may include one or more computing devices, the one or more computing devices including a game console (such as a Sony PlayStation ®or a game console made by Microsoft or Nintendo or other manufacturers), extended reality (XR) headsets (such as virtual reality (VR) headsets, augmented reality (AR) headsets), portable TVs (e.g., smart TVs, Internet-enabled TVs), portable computers (such as laptop computers and tablet computers), and other mobile devices (including smartphones and additional examples discussed below). These client devices can operate with a variety of operating environments. For example, some of the client computers can employ (by way of example) the Linux operating system, an operating system from Microsoft or Unix operating system, or an operating system produced by Apple or Google, or the Berkeley Software Distribution or Berkeley Standard Distribution (BSD) OS (including derivatives of BSD). These operating environments can be used to execute one or more browser programs, such as a browser made by Microsoft or Google or Mozilla, or other browser programs that can access websites hosted by the Internet servers discussed below. Additionally, the operating environments according to the present principles can be used to execute one or more computer game programs.
[0029] A server and / or gateway can be used, and the server and / or gateway can include one or more processors that execute instructions that configure the server to receive and send data over a network such as the Internet. Alternatively, the client and the server can be connected via a local intranet or a virtual private network. The server or the controller can be instantiated by a game console (such as Sony PlayStation ® ), a personal computer, etc.
[0030] Information can be exchanged between the client and the server over the network. For this purpose and for security reasons, the server and / or the client can include a firewall, a load balancer, a temporary storage device, and a proxy, as well as other network infrastructure for achieving reliability and security. One or more servers can form a device that implements a method of providing a secure community (such as an online social networking site or a gaming network) to network members.
[0031] The processor can be a single-chip or multi-chip processor that can perform logic by means of various lines (such as address lines, data lines, and control lines), as well as registers and shift registers. A processor that includes a digital signal processor (DSP) can be an implementation of a circuit system.
[0032] Components included in one embodiment can be used in any suitable combination in other embodiments. For example, any of the various components described herein and / or depicted in the figures can be combined, interchanged, or excluded from other embodiments.
[0033] A "system having at least one of A, B, and C" (similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, C") includes the following systems: having only A; having only B; having only C; having both A and B; having both A and C; having both B and C; and / or having A, B, and C simultaneously.
[0034] Now referring to Figure 1 , an example system 10 is shown, and the example system may include one or more of the example devices mentioned above and further described below according to this principle. The first of the example devices included in system 10 is a consumer electronics (CE) device, such as an audio-video device (AVD) 12, and the AVD may be, for example but not limited to, a projector-based theater display system or an Internet-enabled TV having a TV tuner (equivalently, a set-top box for controlling the TV). Alternatively, the AVD 12 may also be a computerized Internet-enabled ("smart") phone, a tablet computer, a laptop computer, a head-mounted device (HMD), and / or a head-mounted earphone such as smart glasses or a VR head-mounted earphone, another wearable computerized device, a computerized Internet-enabled music player, a computerized Internet-enabled earphone, a computerized Internet-enabled implantable device (such as an implantable skin device), etc. In any case, it should be understood that the AVD 12 is configured to adopt this principle (for example, communicate with other CE devices to adopt this principle, execute the logic described herein, and perform any other functions and / or operations described herein).
[0035] Therefore, in order to adopt such principles, the AVD 12 can be established by some or all of the components shown. For example, the AVD 12 may include one or more touch-enabled displays 14, and the one or more touch-enabled displays may be implemented by a high-definition or ultra-high-definition "4K" or higher-resolution flat screen. Consistent with this principle, the touch-enabled display 14 may include, for example, a capacitive or resistive touch-sensing layer for touch sensing having an electrode grid.
[0036] According to this principle, the AVD 12 may also include one or more speakers 16 for outputting audio, and at least one additional input device 18 (such as an audio receiver / microphone) for entering audible commands into the AVD 12 to control the AVD 12. Other example input devices include a gamepad or a mouse or a keyboard.
[0037] The exemplary AVD 12 may also include one or more network interfaces 20 for communicating under the control of one or more processors 24 via at least one network 22 (such as the Internet, WAN, LAN, etc.). Thus, the interface 20 can be, but is not limited to, a Wi-Fi transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It should be understood that the processor 24 controls the AVD 12 to employ the present principles, including other elements of the AVD 12 described herein, such as controlling the display 14 to present an image on the display and receiving input from the display. Additionally, it should be noted that the network interface 20 can be a wired or wireless modem or router or other suitable interface, such as a wireless phone transceiver, or a Wi-Fi transceiver as mentioned above, etc.
[0038] In addition to the foregoing, the AVD 12 may also include one or more input and / or output ports 26, such as a high-definition multimedia interface (HDMI) port or a universal serial bus (USB) port for physically connecting to another CE device and / or a headphone port for connecting headphones to the AVD 12 to present audio from the AVD 12 to the user via the headphones. For example, the input port 26 may be connected to a cable or satellite source 26a of audio-visual content in a wired or wireless manner. Thus, the source 26a can be a separate or integrated set-top box or satellite receiver. Alternatively, the source 26a can be a game console or disc player containing content. When the source 26a is implemented as a game console, it may include some or all of the components described below with respect to the CE device 48.
[0039] The AVD 12 may also include one or more computer memories / computer-readable storage media 28 that are not transient signals, such as disk-based storage devices or solid-state storage devices, which in some cases are embodied as separate devices within the housing of the AVD, or as a personal video recording device (PVR) or video disc player for playing back AV programs inside or outside the housing of the AVD, or as a removable memory medium or a server described below. Additionally, in some embodiments, the AVD 12 may include a positioning or location receiver 30 such as, but not limited to, a cell phone receiver, a GPS receiver, and / or an altimeter, which is configured to receive geolocation information from satellites or cell phone base stations and provide the information to the processor 24 and / or determine, in conjunction with the processor 24, the altitude at which the AVD 12 is set.
[0040] Continuing the description of the AVD 12, in some embodiments, in accordance with this principle, the AVD 12 may include one or more cameras 32, which may be thermal imaging cameras, digital cameras (such as webcams), IR sensors, event-based sensors, and / or cameras integrated into the AVD 12 and controllable by the processor 24 to gather pictures / images and / or videos. Also included on the AVD 12 may be a Bluetooth ® transceiver 34 and other near field communication (NFC) elements 36 for communicating with other devices using Bluetooth and / or NFC technologies, respectively. Example NFC elements may be radio frequency identification (RFID) elements.
[0041] In addition, the AVD 12 may include one or more auxiliary sensors 38 that provide input to the processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors that form a layer of the touch-enabled display 14 itself, and may be, but are not limited to, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Other sensor examples include pressure sensors, motion sensors (such as accelerometers, gyroscopes, odometers, or magnetic sensors), infrared (IR) sensors, optical sensors, speed and / or rhythm sensors, event-based sensors, (e.g., gesture sensors for sensing gesture commands). The sensors 38 may thus be implemented by one or more motion sensors (such as a single accelerometer, gyroscope, and magnetometer, and / or an inertial measurement unit (IMU) that typically includes a combination of an accelerometer, gyroscope, and magnetometer to determine the position and orientation of the AVD 12 in three dimensions), or by an event-based sensor (such as an event detection sensor (EDS)). The EDS consistent with this disclosure provides an output that indicates a change in the light intensity sensed by at least one pixel of a light sensing array. For example, if the light sensed by a pixel is decreasing, the output of the EDS may be -1; if the light is increasing, the output of the EDS may be +1. An output binary signal 0 may indicate no change in light intensity below a certain threshold.
[0042] The AVD 12 may also include an over-the-air (OTA) TV broadcast port 40 for receiving OTA TV broadcasts that provide input to the processor 24. In addition to the foregoing, it should be noted that the AVD 12 may also include an infrared (IR) transmitter and / or IR receiver and / or IR transceiver 42, such as an IR data association (IRDA) device. A battery (not shown) may be provided for powering the AVD 12, and a kinetic energy harvester may also be provided, which can convert kinetic energy into electricity to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field programmable gate array 46 may also be included. One or more haptic / vibration generators 47 may be provided for generating haptic signals that can be sensed by a person holding or touching the device. The haptic generator 47 may thus use an electric motor to vibrate all or a part of the AVD 12, where the electric motor is connected to an eccentric and / or unbalanced weight by a rotatable shaft of the motor, such that the shaft can be rotated under the control of the motor (which in turn can be controlled by a processor such as the processor 24) to generate vibrations of various frequencies and / or amplitudes and force simulations in various directions.
[0043] A light source, such as a projector, such as an infrared (IR) projector, may also be included.
[0044] In addition to the AVD 12, the system 10 may include one or more other types of CE devices. In one example, the first CE device 48 may be a computer game console, which can be used to transmit computer game audio and video to the AVD 12 either by commands directly transmitted to the AVD 12 and / or through a server described below, while the second CE device 50 may include components similar to those of the first CE device 48. In the example shown, the second CE device 50 may be configured as a computer game controller manipulated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent or non-transparent display for presenting AR / MR content or VR content (more generally, extended reality (XR) content) respectively. The HMD may be configured as a glasses-type display or a large VR-type display sold by a computer game equipment manufacturer.
[0045] In the example shown, only two CE devices are shown, but it should be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for the AVD 12. Any of the components shown in the figures below may incorporate some or all of the components shown in the case of the AVD 12.
[0046] Now referring to the at least one server 52 described above, in accordance with this principle, the at least one server includes at least one server processor 54; at least one tangible computer-readable storage medium 56 (such as a disk-based storage device or a solid-state storage device); and at least one network interface 58, which under the control of the server processor 54 allows communication with other illustrated devices via the network 22 and can in fact facilitate communication between the server and the client device. Note that the network interface 58 can be, for example, a wired or wireless modem or router, a Wi-Fi transceiver, or other suitable interface (such as, for example, a wireless telephone transceiver).
[0047] Thus, in some embodiments, the server 52 can be an Internet server or an entire server "farm" and can include and execute "cloud" functions such that the devices of the system 10 can access a "cloud" environment via the server 52 in an example embodiment of, for example, an online gaming application. Or the server 52 can be implemented by one or more game consoles or other computers in the same room or nearby as the other illustrated devices.
[0048] The components shown in the figures below can include some or all of the components shown herein. Any user interface (UI) described herein can be combined and / or extended, and UI elements can be mixed and matched between UIs.
[0049] This principle can employ various machine learning models, including deep learning models. Machine learning models consistent with this principle can use various algorithms trained in ways that include supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms that can be implemented by computer circuitry include one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and an RNN type called long short-term memory (LSTM) networks. Support vector machines (SVMs) and Bayesian networks can also be considered examples of machine learning models. In addition to the network types described above, the models herein can also be implemented by classifiers.
[0050] As understood herein, performing machine learning can thus involve accessing training data and then training a model based on the training data so that the model can process additional data to make inferences. An artificial neural network / artificial intelligence model trained by machine learning can thus include an input layer, an output layer, and a plurality of hidden layers therebetween, the plurality of hidden layers being configured and weighted to make inferences regarding an appropriate output.
[0051] Now referring to Figure 2, understand the example description of family / friend interactive 3D video generation for displaying the content in the 3D space memory on a holographic display device placed at home. The 3D space memory may include a memory 200 for storing 2D photos, a memory 202 for storing 2D videos, and a memory 204 for storing 2D audio. The 3D space memory may be located, for example, on a local server or a cloud server. If needed, the memories may be combined into a single memory. The memories may be referred to as part of the 3D space memory because they are accessed by one or more 2D-to-3D models 206 to generate 3D videos for presentation on one or more 3D displays 208 (such as 3D space displays, holographic displays, and VR headsets). The model 206 may be a machine learning (ML) model trained on a 2D image training set to generate 3D images.
[0052] When generating the 3D video, the model 206 may access multiple sub-models 210, such as environmental models. Each environmental model may represent a corresponding environmental type, such as a beach, a lake, a mountain, a city, etc. As further discussed herein, the information in the 2D storage device is used to select which environmental type model to generate a background similar but not necessarily identical to the background in the 2D image for the 3D object, in order to improve the 3D generation speed.
[0053] Figure 3 An example 3D display 208 implemented as a holographic display is shown, which presents a 3D character 300 generated from the images in 2D photos and videos, and an environmental background 302 generated based on the information from the 2D data.
[0054] Figure 4 An example overall logic for enabling the storage reproduction of personal audio-video files accessed at block 400 is shown, and the storage reproduction is enhanced (block 404) by automatically converting (block 402) the content that is essentially a partial history of the subject into 3D with interactive features. The end user can watch and interact with copies of their family and friends in the 3D video presented on, for example, a holographic display device. The system generates 3D interactive content based on existing videos and photos. The 3D content is used in 3D viewing devices, such as 3D space displays, holographic displays, and VR headsets. Input devices such as game pads or mice and keyboards can be used to control the created characters and environments.
[0055] Figure 5Additional details are shown. At block 500, the logic analyzes the situation / story of a scene based on various information such as human motion, facial expressions, environment type, location, and time included in the 2D video / photo. Based on this information, at block 502, objects and features (e.g., humans, animations, environment maps, lighting, sound effects, etc.) in the prepared 3D content source are identified for automatically and transparently generating the final 3D video for the user. As indicated at block 504, a portion of the texture and data is collected from the 2D video, including face texture, wall and floor texture, audio such as human voices in the 2D video, and environmental sounds, to update the 3D video at block 506 when the user generates additional 2D content. The content can be timestamped for later disclosure.
[0056] Thus, as indicated at block 506, the 3D memory content is continuously updated. Characters (e.g., parents or children) "grow" with the real user in the 3D video. When the user wants to view past memories, the user can control the time in the virtual world being viewed.
[0057] Figure 6 and Figure 7 Additional principles are shown. Through analysis, for example, using an ML model trained to identify the environment type in an image, 2D information including Figure 6 person 600 and environment background 602 in Figure 6 can be accessed to access an appropriate 3D model to generate Figure 7 3D background image 700 of environment background 602 in
[0058] However, for the person in the 2D image converted to 3D, both the type of the person (selecting a "close-up" model to start) and the 2D texture can be used to render Figure 6 of 2D image 600 in Figure 7The 3D character 702 in []. If, for example, the ML model determines that the image 600 is a little girl, then a little girl 3D model can be selected, and then the texture of the 2D image 600 and (if necessary) the facial features of the image 600 are applied to render the 3D character 702. This shows that for a person, a user viewing a 3D image may notice a significant difference between the 2D image 600 that the user may take of a friend or family member and the 3D character 702 representing that friend or family member.
[0059] Figures 8 to 10 Further shown. At Figure 8 In the box 800 in [], an analysis engine such as an ML model identifies a type of background environment in the 2D image. The ML model can be trained to do this using a training set of images of various environment types that have ground truth labels as to what those types are. Continuing to box 802, a 3D model type is selected based on the output of box 800 to generate a 3D environment background that is similar (although perhaps not identical) to the actual 2D background image.
[0060] Figure 9 Shows that when the person imaged at 600 in the original 2D video moves, the 2D background can change, and by changing the 3D background environment at box 902 according to the changes in the 2D video, the changes are reflected in the 3D video. Figure 6
[0061] In some cases, Figure 7 Figure 6 the 3D character 702 shown in [] can present an audio dialog box generated based on the ID of the user viewing the 3D video. The ID of the viewing user is received at box 1000. Based on the type of person, i.e., the viewer, being, for example, "father", a 3D character 702 is generated at box 1002 and provided with a basic dialog to speak on one or more audio speakers. For example, the 3D character 702 can be animated to have mouth movements and use the original audio of the person who was originally the subject of the 2D video ( the image 600 in []) to say "Hello, Dad" or other similar dialog. The viewing listener can choose to say a simple sentence such as "How are you?", and using speech recognition, an ML model trained on simple dialogs can make the 3D character respond with "Fine" or a similar simple response.
[0062]
[0063] Figure 11 Shows the details discussed herein. Starting from block 1100, any digital camera device can capture 2D pictures and videos as normal. Moving to block 1102, the picture / video is saved to the 3D spatial memory service.
[0064] Proceeding to block 1104, the ML model accessing the 3D spatial memory automatically creates a 3D environment. As discussed above, the model type selection for minimizing the 3D environment background is based on 2D information (e.g., the actual environmental background in the 2D image classified by the ML model).
[0065] Additionally or alternatively, the environmental background model can be selected based on the geographical region associated with the 2D image, e.g., as indicated in the metadata associated with the 2D video. Thus, for example, if a video of a person is captured in 2D and the metadata associated with the video indicates "Paris", the generated 3D background can include an image of the Eiffel Tower even if the Eiffel Tower does not appear in the underlying 2D video.
[0066] The actions in the 3D video are animated to reflect the actions in the 2D video. If needed, the action type (running, swimming, walking, etc.) in the 2D video can be mapped to a matching motion animation motion model in the model database to select a model that matches the action in 2D. If needed, the character in 3D can be animated to move using an appropriate motion model. Language can also be mapped to the character movement. For example, if the person in the 2D video says "Watch me run!", word recognition can be applied to extract "run" and based on this, a running motion model can be selected. As described above, such motion can be used to generate appropriate SFX in the 3D video.
[0067] Once the 3D video is ready, the user can be notified at block 1106 using the user interface on the electronic device. Then, the user can choose to view the 3D video presented accordingly on the 3D display at block 1108.
[0068] If needed, an image of the viewing user can be generated at block 1110 using any of the cameras described herein and provided to the 3D generation model to create a simple dialogue for the 3D character to say. For example, if the user is identified as a father and the 3D character is based on an image of his daughter, the 3D model can generate a dialogue that is played with the voice of the child recorded in the original 2D video, such as "Good morning, Dad". The lips of the 3D character can be animated in sync with the dialogue.
[0069] The box 1112 indicates that the 3D character can respond to subsequent user responses. In this way, the viewing user can have a simple conversation interaction with the virtual character through voice input. For example, the user may say "Good morning", which can be detected by any microphone described in this article and recognized using speech recognition processing. This enables the 3D generation model to play audio in a child's voice, saying "Um... Good morning, Dad. I'm still sleepy." The 3D character can also respond through body postures such as facial expressions and gestures.
[0070] Moving to box 1114, if needed, the viewing user can use an input device such as a gamepad on a computer simulation controller to control the 3D character and interact with other characters in the 3D content world. In this way, the viewing user can make the 3D character walk on the map, make the 3D character talk, eat food, etc.
[0071] Box 1116 indicates that, as described above, when the analyzed 2D data generated by the user and stored in the 3D memory section is updated, the virtual 3D content is automatically updated to keep showing the latest memories of people.
[0072] Due to these features, box 1118 indicates that the viewing user can choose to watch a 3D video depicting a specific period in the life of the subject of the 3D video. Figure 2 and Figure 13 shown. The viewing user can move the slider 1200 to select which year out of six years the user wishes to watch the 3D video of. In Figure 12 it, the second period (such as one year) out of the six periods covered by the 3D video is selected, so the 3D video depicts an image based on 2D photos and videos taken during the second period. In Figure 13 it, the slider 1200 has been moved to the sixth period, so the 3D video depicts an image based on 2D photos and videos taken during the sixth period. Therefore, the user can control the time slider 1200 and display any past moment in the virtual 3D world.
[0073] Figure 11 and Figure 14 The box 1120 in Figure 14 shows that an incoming call on the communication device 1400 can be detected and input into the 3D model generation system to make the 3D character 1402 presented on the 3D display 1404 make an auditory response, as indicated by the word 1406 in
[0074] Figure 15Further illustration of the principles described herein is provided. 2D videos and photos 1500 in a 3D memory space are provided for analysis 1502 to extract features 1504, including the personal characteristics of people in the 2D images, such as facial texture, type of facial expression, age, gender, type and color of clothing, voice files, type of body movement (in the video), and personal identity. The features may also include environmental features, such as type of location, time and date of the day, and background environmental sounds in the 2D video.
[0075] The features 1504 extracted from the 2D data 1500 are provided to a matching model 1506 to select various 3D models 1508 for creating a related 3D video. These models may include character body models, facial animation models, character voice files, character body animation models, name tags, and background environmental models and background sound files.
[0076] The 3D models 1508 are exported to a 3D model set 1510, which combines (1512) the 3D characters, environments, and sounds output by the selected 3D models 1508 with interaction functions 1514 (such as input device control, feedback functions, and phone call functions for 3D characters as described above). The combined 3D information is presented on a 3D display 1516, such as a holographic display with an audio speaker.
[0077] While specific embodiments have been shown and described in detail herein, it should be understood that the subject matter covered by the present invention is limited only by the claims.
Claims
1. An apparatus, comprising: at least one processor configured to: access 2D image information in a storage device; use the 2D image information to at least partially generate a 3D interactive video by matching at least one feature of at least one object in the 2D image information with at least one 3D model.
2. The apparatus according to claim 1, wherein the processor is configured to: receive at least one interaction signal from at least one input device; and at least partially change the presentation of the 3D interactive video based on the interaction signal.
3. The apparatus according to claim 2, wherein the input device includes at least one touch element on at least one computer simulation controller.
4. The apparatus according to claim 2, wherein the input device includes at least one microphone.
5. The apparatus according to claim 1, wherein the processor is programmed to: change the presentation of the 3D interactive video in response to an incoming call.
6. The apparatus according to claim 1, wherein the at least one feature of the 2D image information includes a person image texture.
7. The apparatus according to claim 1, wherein the at least one feature of the 2D image information includes a person image motion.
8. The apparatus according to claim 1, wherein the at least one feature of the 2D image information includes a person image facial expression.
9. The apparatus according to claim 1, wherein the at least one feature of the 2D image information includes an environment type.
10. The apparatus according to claim 1, wherein the processor is programmed to: change the presentation of the 3D interactive video in response to an input from a time slider input element.
11. A device, comprising: at least one computer storage device, which is not a transient signal and includes instructions that can be executed by at least one processor to: identify at least one texture of at least one person image in a 2D image; use the texture to generate a 3D character; identify at least one environmental background type in the 2D image; select an environmental background type model at least partially based on the environmental background type in the 2D image; use the environmental background type model to generate a 3D environmental background; and merge the 3D environmental background with the 3D character to generate a 3D video.
12. The device according to claim 11, wherein the instructions can be executed to: present the 3D video on at least one 3D display.
13. The device according to claim 11, wherein the instructions can be executed to: receive at least one interaction signal from at least one input device; and at least partially change the presentation of the 3D video based on the interaction signal.
14. The device according to claim 11, wherein the instructions can be executed to: change the presentation of the 3D interactive video in response to an incoming call.
15. The device according to claim 11, which includes the at least one processor.
16. A method, comprising: access a 2D image in a 3D space memory; Automatically create a 3D video at least in part using an environmental background model that is selected based on the environmental background in the 2D image as classified by a machine learning (ML) model and the texture of a face image in the 2D image.
17. The method of claim 16, comprising: Animating at least one character in the 3D video based on an action in the 2D image.
18. The method of claim 16, comprising: Playing an audible dialogue in the 3D video at least in part based on an image of a user viewing the 3D video.
19. The method of claim 16, comprising: Animating at least one character in the 3D video based on an input from an input device.
20. The method of claim 16, comprising: Automatically updating the 3D video when a 2D image is added to the 3D spatial memory.