Generating 3D videos using images and audio in the background keyed to metadata from 2D images

The system converts 2D images and videos into interactive 3D content using machine learning, addressing the lack of immersive 3D experiences by generating realistic characters and environments, enabling time-based viewing on 3D display devices.

JP2025539384AActive Publication Date: 2025-12-05SONY INTERACTIVE ENTERTAINMENT LLC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2025530568
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-26
Filing Date
2023-11-02
Publication Date
2025-12-05
Estimated Expiration
2043-11-02

AI Technical Summary

Technical Problem

Existing 2D image and video collections lack the ability to be enhanced into interactive 3D content with realistic characters and environments, limiting the immersive experience on 3D display devices.

Method used

A system that analyzes 2D images and videos to identify objects and features, generating 3D interactive content using machine learning models, allowing users to interact with characters and environments through input devices, and updating the 3D memory as the user grows.

Benefits of technology

Enables the creation of immersive 3D interactive videos with realistic characters and environments, enhancing user engagement and allowing time-based viewing of personal memories on 3D display devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025539384000001_ABST
    Figure 2025539384000001_ABST
Patent Text Reader

Abstract

The 3D video is generated (402) based on the end user's 2D photos and videos stored in a section of the memory (200, 202) designated as 3D memory. Characters in the 3D video are based on the textures of people in the 2D images (504), and environmental backgrounds in the 3D video are selected from models that most closely resemble the images in the 2D video. The 3D video is updated (506) over time to allow the user to select which period of a person's life to view in 3D.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This application relates generally to 3D spatial memory, and more particularly to generating 3D video using a timeline based on an end-user's 2D video. [Background technology]

[0002] As understood herein, people ("end users") generate and store 2D images and videos of the same subject, such as a child, over an extended period of time, such as several years. Summary of the Invention

[0003] The present principles also recognize that viewing stored copies of personal audio-video files can be enhanced by automatically converting what is essentially a partial history of the subject into 3D with interactive features. End users can, for example, view and interact with copies of family and friends in 3D videos presented on a holographic display device. The system generates 3D interactive content from existing videos and photos. The primary purpose of the 3D content is for use with 3D display devices such as 3D spatial displays, holographic displays, and VR headsets. The created characters and environments can be controlled using input devices such as a gamepad or a mouse and keyboard.

[0004] The system analyzes the situation / story from various information contained in the video / photo, such as human body movements, facial expressions, environment type, location, time, etc. Based on this information, objects and features in the prepared 3D content source (e.g., human bodies, animations, environment maps, lighting, sound effects, etc.) are identified, and the final 3D video for the user is automatically and transparently generated. Some textures and data are collected from the 2D video, including human face textures, wall and floor textures, audio such as human voices in the 2D video, and environmental sounds.

[0005] The 3D memory content keeps updating. The character (e.g., parent or child) grows in the 3D video together with the real user. If the user wants to view past memories, the user can control the time of the virtual world being viewed.

[0006] Accordingly, the apparatus includes at least one processor configured to access the 2D image information from the storage and generate a 3D interactive video using the 2D image information, at least in part, by matching at least one feature of the at least one object in the 2D image information to the at least one 3D model.

[0007] In some examples, the processor is configured to receive at least one interaction signal from the at least one input device and modify the presentation of the 3D interactive video based at least in part on the interaction signal.

[0008] The input device may be, for example, at least one touch element on at least one computer-simulated controller and / or at least one microphone.

[0009] In some implementations, the processor can be programmed to modify the presentation of the 3D interactive video in response to an incoming telephone call.

[0010] The characteristics of the 2D image information on which the 3D video is generated may include one or more of human image texture, human image movement, human image facial expression, and environment type.

[0011] In some examples, the processor may be programmed to modify the presentation of the 3D interactive video in response to input from a time slider input element.

[0012] In another aspect, the device includes at least one computer storage device, which is not a transitory signal, and which includes instructions executable by at least one processor for sequentially identifying at least one texture of at least one person image in a 2D image. The instructions are executable to generate a 3D character using the texture. The instructions are further executable to identify at least one environmental background type in the 2D image and select an environmental background type model to be used to generate a 3D environmental background based at least in part on the environmental background type in the 2D image. The instructions are executable to blend the 3D environmental background with the 3D character to generate a 3D video.

[0013] In another aspect, a method includes accessing a 2D image from a 3D spatial memory and automatically creating a 3D video based on the 2D image and a texture of an image of a human face in the 2D image using, at least in part, an environmental background model selected based on an environmental background of the 2D image classified by a machine learning (ML) model.

[0014] The details of the present application, both as to its structure and operation, can best be understood in reference to the accompanying drawings, in which like reference numerals refer to like parts, and in which: [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a block diagram of an exemplary system in accordance with present principles; [Figure 2] 1 illustrates an exemplary system architecture consistent with the present principles. [Figure 3] 1 illustrates an exemplary holographic 3D display. [Figure 4] The exemplary overall logic is shown in exemplary flow chart form. [Figure 5] 1 illustrates exemplary 3D image generation logic in exemplary flow chart form. [Figure 6]10 shows exemplary screenshots illustrating how a 3D environment background can differ from a 2D image from where the 3D image was generated. [Figure 7] 10 shows exemplary screenshots illustrating how a 3D environment background can differ from a 2D image from where the 3D image was generated. [Figure 8] 1 illustrates exemplary 3D environment generation logic in exemplary flow chart form. [Figure 9] 1 illustrates exemplary 3D environment modification logic in exemplary flowchart form. [Figure 10] 1 illustrates, in exemplary flow chart form, exemplary 3D image generation logic based on user ID. [Figure 11] The logic of an exemplary detailed use case is presented in exemplary flowchart form. [Figure 12] 10 shows exemplary screenshots illustrating time-based viewing of 3D video. [Figure 13] 10 shows exemplary screenshots illustrating time-based viewing of 3D video. [Figure 14] Shows a 3D video of answering a phone call. [Figure 15-1] Further illustrations of details of exemplary embodiments are provided. [Figure 15-2] Further illustrations of details of exemplary embodiments are provided. DETAILED DESCRIPTION OF THE INVENTION

[0016] The present disclosure generally relates to computer ecosystems, including, but not limited to, aspects of consumer electronics (CE) device networks, such as computer gaming networks. The systems herein may include server and client components that may be connected via a network such that data may be exchanged between the client and server components. The client components may include one or more computing devices, including game consoles such as the Sony PlayStation®, or game consoles made by Microsoft®, Nintendo®, or other manufacturers; extended reality (XR) headsets, such as virtual reality (VR) headsets and augmented reality (AR) headsets; portable televisions (e.g., smart TVs, Internet-enabled televisions); portable computers, such as laptop computers and tablet computers; and smartphones and other mobile devices, including additional examples described below. These client devices may operate in a variety of operating environments. For example, some of the client computers may use, as examples, the Linux® operating system, an operating system manufactured by Microsoft®, or a Unix® operating system, or an operating system manufactured by Apple, Inc.®, or Google®, or the Berkeley Software Distribution or Berkeley Standard Distribution (BSD) OS (including derivatives of BSD). These operating environments may be used to run one or more browsing programs, such as browsers made by Microsoft®, Google®, or Mozilla®, or other browser programs capable of accessing websites hosted by Internet servers as described below. An operating environment according to present principles may also be used to run one or more computer game programs.

[0017] Servers and / or gateways may be used, which may include one or more processors that execute instructions that configure the server to receive and transmit data over a network such as the Internet. Alternatively, clients and servers may be connected via a local intranet or virtual private network. The server or controller may be instantiated by a game console such as a Sony PlayStation®, a personal computer, or the like.

[0018] Information may be exchanged between the clients and the servers over a network. To this end, and for security, the servers and / or clients may include firewalls, load balancers, temporary storage, and proxies, as well as other network infrastructure for reliability and security. One or more servers may form an apparatus that implements a method for providing a secure community, such as an online social website or gamer network, to network members.

[0019] The processor may be a single-chip processor or a multi-chip processor capable of performing logic through various lines, such as address lines, data lines, and control lines, as well as registers and shift registers. A processor, including a digital signal processor (DSP), may be an embodiment of a circuit.

[0020] Components included in one embodiment may be used in other embodiments in any suitable combination. For example, any of the various components described herein and / or depicted in the figures may be combined, substituted, or excluded from other embodiments.

[0021] "A system having at least one of A, B, and C" (and similarly, "a system having at least one of A, B, or C" and "a system having at least one of A, B, and C") includes a system having only A, a system having only B, a system having only C, a system having both A and B, a system having both A and C, a system having both B and C, and / or a system having both A, B, and C.

[0022] Referring now to FIG. 1 , an exemplary system 10 is shown. System 10 can include one or more of the exemplary devices in accordance with the present principles referenced above and further described below. A first of the exemplary devices included in system 10 is a consumer electronics (CE) device such as an audio-video device (AVD) 12, such as, but not limited to, a theater display system that may be projector-based, or an Internet-enabled television with a TV tuner (or, equivalently, a set-top box that controls a TV). Alternatively, AVD 12 can also be a computerized Internet-enabled (“smart”) phone, a tablet computer, a notebook computer, a head-mounted device (HMD) and / or headset (such as smart glasses or a VR headset), another computerized wearable device, a computerized Internet-enabled music player, computerized Internet-enabled headphones, a computerized Internet-enabled implantable device (such as an implantable skin device), or the like. In any event, it will be understood that AVD12 is configured to implement the present principles (e.g., to communicate with other CE devices to implement the present principles, to execute the logic described herein, and to perform any other functions and / or operations described herein).

[0023] Thus, to implement such principles, AVD 12 may be established by some or all of the components shown. For example, AVD 12 may include one or more touch-enabled displays 14, which may be implemented by high-definition or ultra-high-definition "4K" or higher flat screens. Touch-enabled display(s) 14 may include, for example, a capacitive or resistive touch-sensing layer with a grid of electrodes for touch sensing consistent with the present principles.

[0024] AVD 12 may also include one or more speakers 16 for outputting audio in accordance with the present principles, and at least one additional input device 18, such as an audio receiver / microphone, for inputting audible commands into AVD 12 to control AVD 12. Other exemplary input devices include a gamepad or a mouse or keyboard.

[0025] The example AVD 12 may also include one or more network interfaces 20 for communicating over at least one network 22, such as the Internet, a WAN, or a LAN, under the control of one or more processors 24. Thus, the interface 20 may be a Wi-Fi® transceiver, which is an example of a wireless computer network interface, such as, but not limited to, a mesh network transceiver. It will be understood that the processor 24 controls the AVD 12 to implement the present principles, including controlling the display 14 to present images thereon, receiving input from the display 14, and other elements of the AVD 12 described herein. Furthermore, it should be noted that the network interface 20 may be a wired or wireless modem or router, or a wireless telephony transceiver or other suitable interface, such as the Wi-Fi® transceiver described above.

[0026] In addition to the foregoing, AVD 12 may also include one or more input and / or output ports 26, such as a High-Definition Multimedia Interface (HDMI®) port or a Universal Serial Bus (USB) port for physically connecting to another CE device, and / or a headphone port for connecting headphones to AVD 12 to present audio from AVD 12 to a user via headphones. For example, input port 26 may be wired or wirelessly connected to a cable or satellite source 26a of audio-video content. Thus, source 26a may be a separate or integrated set-top box or satellite receiver. Alternatively, source 26a may be a game console or disc player containing content. If implemented as a game console, source 26a may include some or all of the components described below in connection with CE device 48.

[0027] AVD 12 may further include one or more computer memory / computer-readable storage media 28, such as disk-based or solid-state storage, which may be embodied within the AVD's chassis as a standalone device, or as a personal video recording device (PVR) or video disc player either internal or external to the AVD's chassis for playing AV programs, or as a removable storage medium or a server as described below. In some embodiments, AVD 12 may also include a location or position receiver, such as, but not limited to, a cellular telephone receiver, a GPS receiver, and / or an altimeter 30 configured to receive geographic location information from a satellite or cellular tower and provide that information to processor 24 and / or configured in cooperation with processor 24 to determine the altitude at which AVD 12 is located.

[0028] Continuing with the description of AVD 12, in some embodiments, AVD 12 may include one or more cameras 32, which may be a thermal imaging camera, a digital camera such as a webcam, an IR sensor, an event-based sensor, and / or a camera integrated into AVD 12 and controllable by processor 24 to collect pictures / images and / or video in accordance with the present principles. AVD 12 may also include a Bluetooth transceiver 34 and other NFC elements 36 for communicating with other devices using Bluetooth and / or near field communication (NFC) technology, respectively. An exemplary NFC element may be a radio frequency identification (RFID) element.

[0029] Furthermore, AVD 12 may include one or more auxiliary sensors 38 that provide input to processor 24. For example, one or more of the auxiliary sensors 38 may include one or more pressure sensors that form a layer of touch-enabled display 14 itself, and may be, without limitation, piezoelectric pressure sensors, capacitive pressure sensors, piezoresistive strain gauges, optical pressure sensors, electromagnetic pressure sensors, etc. Examples of other sensors include pressure sensors, motion sensors such as accelerometers, gyroscopes, cyclometers, or magnetic sensors, infrared (IR) sensors, optical sensors, speed and / or cadence sensors, event-based sensors, and gesture sensors (e.g., sensors for sensing gesture commands). Thus, sensors 38 may be implemented by an inertial measurement unit (IMU), which typically includes one or more motion sensors such as individual accelerometers, gyroscopes, and magnetometers, and / or a combination of accelerometers, gyroscopes, and magnetometers, or by an event-based sensor such as an event detection sensor (EDS), to determine the position and orientation of AVD 12 in three dimensions. An EDS consistent with the present disclosure provides an output indicative of a change in light intensity sensed by at least one pixel of the light-sensing array. For example, if the light sensed by the pixel is decreasing, the output of the EDS may be −1. If it is increasing, the output of the EDS may be +1. No change in light intensity below a certain threshold may be indicated by an output binary signal of 0.

[0030] The AVD 12 may also include a wireless TV broadcast port 40 for receiving over-the-air TV broadcasts, which provides input to the processor 24. Note that, in addition to the foregoing, the AVD 12 may also include an infrared (IR) transmitter and / or receiver and / or transceiver 42, such as an Infrared Data Association (IRDA) device. The AVD 12 may be provided with a battery (not shown) for powering itself, which may be a kinetic energy harvester capable of converting kinetic energy into electrical power to charge the battery and / or power the AVD 12. A graphics processing unit (GPU) 44 and a field-programmable gate array 46 may also be included. One or more haptic / vibration generators 47 may be provided for generating haptic signals that can be sensed by a person holding or in contact with the device. Thus, the haptic generator 47 may vibrate all or part of the AVD 12 using an electric motor connected to an off-center and / or unbalanced weight via the motor's rotating shaft, the shaft rotating under the control of the motor (which may be controlled by a processor such as processor 24) to create simulations of vibrations of various frequencies and / or amplitudes, and forces in various directions.

[0031] A light source such as a projector, such as an infrared (IR) projector, may also be included.

[0032] In addition to the AVD 12, the system 10 may include one or more other CE device types. In one example, the first CE device 48 may be a computer game console that can be used to transmit computer game audio and video to the AVD 12 via commands sent directly to the AVD 12 and / or via a server, as described below, while the second CE device 50 may include similar components to the first CE device 48. In the illustrated example, the second CE device 50 may be configured as a computer game controller operated by a player or a head-mounted display (HMD) worn by a player. The HMD may include a head-up transparent display or a head-up opaque display that presents AR / MR content or VR content (more generally, extended reality (XR) content), respectively. The HMD may be configured as a glasses-type display or as a large VR-type display sold by a computer game console manufacturer.

[0033] In the illustrated example, only two CE devices are shown, and it will be understood that fewer or more devices may be used. The devices herein may implement some or all of the components shown for AVD 12. Any of the components shown in the figures below may incorporate some or all of the components shown in the AVD 12 example.

[0034] Referring now to the aforementioned at least one server 52, it includes at least one server processor 54, at least one tangible computer-readable storage medium 56, such as disk-based or solid-state storage, and at least one network interface 58 that, under the control of the server processor 54, enables communication with the other illustrated devices over the network 22 and may, indeed, facilitate communication between the server and client devices in accordance with the present principles. It should be noted that the network interface 58 may be, for example, a wired or wireless modem or router, a Wi-Fi® transceiver, or other suitable interface, such as, for example, a wireless telephony transceiver.

[0035] Thus, in some embodiments, server 52 may be an entire Internet server or server "farm" that includes and is capable of performing "cloud" functionality, allowing devices of system 10 to access a "cloud" environment via server 52, for example, in embodiments for network gaming applications, or server 52 may be implemented by one or more game consoles or other computers in the same room or nearby as the other devices shown.

[0036] The components shown in the following figures may include some or all of the components shown herein. Any user interfaces (UIs) described herein may be integrated and / or extended, and UI elements may be mixed and matched between UIs.

[0037] The present principles may use a variety of machine learning models, including deep learning models. Machine learning models consistent with the present principles may use a variety of algorithms trained using methods including supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, feature learning, self-learning, and other forms of learning. Examples of such algorithms may be implemented by computer circuitry, including one or more neural networks, such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and a type of RNN known as a long-short-term memory (LSTM) network. Support vector machines (SVMs) and Bayesian networks may also be considered examples of machine learning models. In addition to the types of networks listed above, the models herein may also be implemented by classifiers.

[0038] As understood herein, performing machine learning may include accessing and training a model on training data to enable the model to process additional data and make inferences. As a result, an artificial neural network / artificial intelligence model trained through machine learning may include a weighted input layer, an output layer, and multiple hidden layers in between, configured to make inferences about a suitable output.

[0039] Reference is now made to FIG. 2 for an exemplary illustration of family / friend interactive 3D video generation for displaying the contents of a 3D spatial memory on a holographic display device located in a home. The 3D spatial memory may include a memory 200 for storing 2D photos, a memory 202 for storing 2D video, and a memory 204 for storing 2D audio. The 3D spatial memory may reside, for example, on a local server or a cloud server. If desired, the memories may be combined into a single memory. The memory may be referred to as part of a 3D spatial memory because it is accessed by one or more 2D-to-3D models 206 to generate 3D video for presentation on one or more 3D displays 208, such as a 3D spatial display, a holographic display, and a VR headset. The models 206 may be machine learning (ML) models trained on a training set of 2D images to generate 3D images.

[0040] When generating 3D video, model 206 may access multiple sub-models 210, such as environment models. Each environment model may represent a different environment type, such as a beach, a lake, a mountain, a city, etc. As discussed further herein, information in the 2D storage is used to select which model of the environment type to select for generating backgrounds for 3D objects that are similar to, but not necessarily exactly identical to, the backgrounds of the 2D images in order to increase the speed of 3D generation.

[0041] FIG. 3 shows an exemplary 3D display 208 implemented as a holographic display presenting a 3D character 300 generated from 2D photo and video images along with an environmental background 302 generated based on information from the 2D data.

[0042] FIG. 4 illustrates exemplary overall logic, accessed by block 400, that allows viewing of stored copies of personal audio-video files, augmented by automatically converting (block 402) what is essentially a partial history of the subject into 3D with interactive features (block 404). End users can view and interact with copies of family and friends in 3D video presented on a holographic display device, for example. The system generates 3D interactive content from existing videos and photos. The 3D content is used in 3D display devices such as 3D spatial displays, holographic displays, and VR headsets. The created characters and environments can be controlled using input devices such as a gamepad or a mouse and keyboard.

[0043] Figure 5 shows additional details. In block 500, logic analyzes the situation / story of the scene from various information contained in the 2D video / photo, such as human body movements, facial expressions, environment type, location, and time. Based on this information, objects and features of the prepared 3D content source (e.g., human bodies, animations, environment maps, lighting, sound effects, etc.) are identified in block 502, and the user's final 3D video is automatically and transparently generated. As shown in block 504, some textures and data, including human face textures, wall and floor textures, and audio, such as human voices and environmental sounds in the 2D video, are collected from the 2D video, and the 3D video is updated in block 506 as additional 2D content is generated by the user. The content may be time-stamped for purposes of brief disclosure.

[0044] Thus, the contents of the 3D memory continue to update, as shown in block 506. The character (e.g., parent or child) "grows" in the 3D video along with the real user. If the user wants to view past memories, the user can control the time of the virtual world being viewed.

[0045] 6 and 7 illustrate additional principles. For example, by analyzing 2D information including the person 600 and environmental background 602 of FIG. 6 using an ML model trained to recognize the environmental type of an image, an appropriate 3D model can be accessed to generate a 3D background image 700 of FIG. 7 of the environmental background 602 of FIG. 6 . While both the 2D and 3D environmental backgrounds include trees and grass, the 3D rendering, in some embodiments, is based solely on the type of the environmental background 602 to render a similar environmental background; however, the 3D environmental background 700 need not be an exact 3D rendering of the 2D environmental background 602. This recognizes that as long as the environmental background type remains the same between the 2D and 3D versions of the image, a user viewing the image is unlikely to notice or care whether the environmental backgrounds match exactly. If sand is present in the original 2D background, typically any sandy environmental background may be used in the 3D rendering, which can reduce processing time.

[0046] However, for people in 2D images converted to 3D, both the person type (first selecting a "close" model) and the 2D texture can be used to render the 3D character 702 in Figure 7 for the 2D image 600 in Figure 6. If the image 600 is determined by the ML model to be of, for example, a little girl, a 3D model of the little girl may be selected, and then the texture of the 2D image 600, and optionally the facial features of the image 600, may be applied to render the 3D character 702. This recognizes that with regard to people, a user viewing a 3D image will likely notice substantial differences between the 2D image 600, which the user may have taken of a friend or family member, and the 3D character 702 representing that friend or family member.

[0047] 8-10 are further illustrated. In block 800 of FIG. 8, an analysis engine, such as an ML model, recognizes the type of background environment in a 2D image. To do this, the ML model may be trained using a training set of images of various environment types and ground truth labels of what those types are. Proceeding to block 802, a 3D model type is selected based on the output of block 800 to create a 3D environment background that is similar, though perhaps not exactly the same, as the actual 2D background image.

[0048] FIG. 9 shows that as the person captured at 600 in FIG. 6 moves in the original 2D video, the 2D background may change, and that change is reflected in the 3D video by changing the 3D background environment at block 902 according to the change in the 2D video.

[0049] In some examples, the 3D character 702 shown in FIG. 7 may be presented with audio dialogue generated based on the identity of the user viewing the 3D video. In block 1000, the identity of the viewing user is received. Based on the viewer's character type, e.g., "father," the 3D character 702 is generated in block 1002 and given basic dialogue to speak through one or more audio speakers. For example, the 3D character 702 may be animated with mouth movements and using the original audio of the character originally the subject of the 2D video (image 600 in FIG. 6 ) to say "Hello, Dad" or other similar dialogue. The viewing listener may choose to speak a simple sentence such as "Are you okay?" Using speech recognition, an ML model trained on simple dialogue can cause the 3D character to respond with "I'm okay" or a similar simple response.

[0050] Also, sound effects (SFX) may be added to the 3D video according to the actions of the subject in the video, as shown in block 1004. For example, if the subject is running on a side street, footstep sounds may be added to the 3D presentation whenever the feet hit the ground.

[0051] 11 shows the details discussed herein. Starting at block 1100, 2D images and videos are taken as usual with any digital camera device. Proceeding to block 1102, the pictures / videos are saved to a 3D spatial memory service.

[0052] Proceeding to block 1104, the ML model accessing the 3D spatial memory automatically creates the 3D environment. As described above, the selection of the model type of the 3D environment background is minimized according to 2D information, e.g., the actual environment background of the 2D image(s) classified by the ML model.

[0053] Additionally or alternatively, the environmental background model may be selected based on a geographic region associated with the 2D image, for example, as may be indicated in metadata associated with the 2D video. Thus, for example, if a video of a person is shot in 2D and the metadata associated with the video indicates "Paris," the generated 3D background may include an image of the Eiffel Tower, even though the Eiffel Tower is not evident in the underlying 2D video.

[0054] Actions in the 3D video are animated to mirror the actions in the 2D video. Optionally, the type of motion in the 2D video (running, swimming, walking, etc.) can be mapped to a motion model for matching motion animations in a database of models to select a model that matches the action in 2D. Characters in 3D are animated to move using the appropriate motion model as needed. Language can also be mapped to character movements. For example, if a person in the 2D video says, "Watch me run!", word recognition can be applied to extract "running" and select a running motion model based on that. As described above, such motions can be used to generate appropriate SFX in the 3D video.

[0055] When the 3D video is ready, the user may be notified using the user interface of the electronic device in block 1106. The user may then select to view the 3D video presented correspondingly on the 3D display in block 1108.

[0056] Optionally, using any of the cameras described herein, an image of the viewing user may be generated in block 1110 and fed to a 3D generative model to create simple dialogue for the 3D character to speak. For example, if the user is recognized as a father and the 3D character is based on an image of his daughter, the 3D model may generate dialogue such as "Good morning, Dad," played in the child's voice recorded in the original 2D video. The 3D character's lips may be animated in synchronization with the dialogue.

[0057] Block 1112 shows that the 3D character can be made responsive to subsequent user reactions. In this manner, a viewing user can interact with the virtual character through simple conversations via voice input. For example, if the user says, "Good morning," this can be detected by any microphone described herein and recognized using voice recognition processing. This can cause the 3D generative model to play audio in a child's voice saying, "Hmm... Good morning, Dad. I'm still sleepy." The 3D character can also respond through facial expressions and body gestures, such as hand gestures.

[0058] Moving to block 1114, the viewing user can optionally use an input device, such as a gamepad on a computer-simulated controller, to control the 3D character and interact with other characters in the 3D content world. In this manner, the viewing user can make the 3D character walk around a map, make the 3D character talk, eat, etc.

[0059] Block 1116 indicates that as the analytical 2D data generated by the user and stored in the 3D memory section is updated, as described above, the virtual 3D content is automatically updated to continue to show the most recent memories of the people.

[0060] Because of these features, block 1118 indicates that the viewing user can select to view a 3D video depicting a specific period in the life of the 3D video subject. This is illustrated in FIGS. 2 and 13. A slider 1200 can be moved by the viewing user to select the six-year period for which the user wishes to view the 3D video. In FIG. 12, a second period (e.g., one year) of the six-period time span covered by the 3D video has been selected, so that the 3D video depicts images based on 2D photographs and videos taken during the second period. In FIG. 13, the slider 1200 has been moved to the sixth period, so that the 3D video depicts images based on 2D photographs and videos taken during the sixth period. Thus, the time slider 1200 can be used to control and display any past moment in the virtual 3D world.

[0061] Block 1120 of Figures 11 and 14 indicates that an incoming phone call can be detected at communication device 1400 and input into the 3D model generation system so that a 3D character 1402 presented on 3D display 1404 can respond audibly, as indicated by word 1406 in Figure 14. When a call is received through the system, a 3D model of the registered person can be animated and speak in sync with the voice audio.

[0062] Figure 15 provides a further illustration of the principles described herein. 2D videos and photos 1500 in 3D memory space are subjected to analysis 1502 to extract features 1504, including human characteristics of people in the 2D images, such as facial texture, type of facial expression, age, gender, type and color of clothing, audio files, type of body motion (in the video), personal identification, etc. Features may also include environmental features such as type of location, time and date, and background ambient sounds of the 2D video.

[0063] The features 1504 extracted from the 2D data 1500 are provided to a matching model 1506 to select various 3D models 1508 to use in creating the associated 3D video. These models may include a character body model, a facial animation model, a character voice file, a character body animation model, a name tag, a background environment model, and a background sound file.

[0064] The 3D models 1508 are exported to a 3D model set 1510, which, as described above, blends 1512 the 3D people, environments, and sounds output by the selected 3D models 1508 with interaction features 1514, such as control of the 3D character's input devices, feedback features, and phone call features. The blended 3D information is presented on a 3D display 1516, such as a holographic display with audio speakers. While particular embodiments are shown and described in detail herein, it should be understood that the subject matter encompassed by the present invention is limited only by the scope of the claims.

Claims

1. 1. An apparatus including at least one processor, the at least one processor comprising: Accessing 2D image information in storage, The device is configured to use the 2D image information to generate a 3D interactive video by at least partially matching at least one feature of at least one object in the 2D image information to at least one 3D model.

2. The processor: receiving at least one interaction signal from at least one input device; The device of claim 1 , configured to modify a presentation of the 3D interactive video based at least in part on the interaction signal.

3. The apparatus of claim 2 , wherein the input device comprises at least one touch element on at least one computer-simulated controller.

4. The apparatus of claim 2 , wherein the input device includes at least one microphone.

5. The processor:

10. The device of claim 1, programmed to modify the presentation of the 3D interactive video in response to an incoming telephone call.

6. The apparatus of claim 1 , wherein the at least one feature of the 2D image information includes texture in an image of a human.

7. The apparatus of claim 1 , wherein the at least one feature of the 2D image information includes motion in an image of a person.

8. The apparatus of claim 1 , wherein the at least one feature of the 2D image information includes facial features in images of humans.

9. The apparatus of claim 1 , wherein the at least one characteristic of the 2D image information includes an environment type.

10. The processor:

10. The apparatus of claim 1, programmed to modify the presentation of the 3D interactive video in response to input from a time slider input element.

11. A device, at least one computer storage device containing instructions executable by at least one processor that are not transitory signals, said instructions comprising: identifying at least one texture of the at least one human image in the 2D image; generating a 3D character using said texture; identifying at least one environmental background type of the 2D image; selecting a model of an environmental background type based at least in part on the environmental background type in the 2D image; generating a 3D environment using the environment type model; and and blending the 3D environment background with the 3D character to generate a 3D video.

12. The instruction: The device of claim 11 , wherein the device is operable to present the 3D video on at least one 3D display.

13. The instruction: receiving at least one interaction signal from at least one input device; and The device of claim 11 , wherein the device is operable to modify a presentation of the 3D video based at least in part on the interaction signal.

14. The instruction:

12. The device of claim 11, wherein the device is operable to modify the presentation of the 3D interactive video in response to an incoming telephone call.

15. The device of claim 11 comprising the at least one processor.

16. 1. A method comprising: accessing a 2D image in a 3D spatial memory; automatically creating a 3D video based on the 2D images and textures of images of human faces in the 2D images using, at least in part, an environmental background model selected based on an environmental background of the 2D images classified by a machine learning (ML) model.

17. 17. The method of claim 16, comprising animating at least one character in the 3D video based on an action in the 2D image.

18. 17. The method of claim 16, comprising playing audible dialogue in the 3D video based at least in part on an image of a user watching the 3D video.

19. 17. The method of claim 16, comprising animating at least one character in the 3D video based on a device from an input device.

20. 17. The method of claim 16, including automatically updating 3D video as 2D images are added to the 3D spatial memory.

Citation Information

Patent Citations

  • Game device, game control method and program

    JP2012213453A

  • Program, information processing apparatus, communication system, and communication method

    JP2016018313A

  • Virtual reality-based apparatus and method to generate three-dimensional (3D) human face model using image and depth data

    JP2018170005A

  • Virtual character generation from image or video data

    US20210275925A1

  • Methods and systems for dynamic summary queue generation and provision

    WO2022146754A1