Method and apparatus for providing creation and distribution service of lightweight immersive content
The method addresses limitations in immersive content production and distribution by converting and streaming lightweight 360-degree and 3D data to various devices, enhancing realism and accessibility on existing hardware.
Patent Information
- Application Number
- PCT/KR2024/009982
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2024-07-12
- Publication Date
- 2025-10-30
AI Technical Summary
Existing methods for producing and distributing immersive content, such as 360-degree images and 3D data, are limited by the need for expensive dedicated hardware, difficulty in use, and inefficiencies in processing and distribution, leading to a lack of accessibility and realism on devices like smartphones, PCs, and tablets.
A method and system for producing and distributing lightweight immersive content by performing resource lightweighting on 360-degree image data and 3D data, including conversion to low, medium, and high-resolution formats, and combining these with gaze and location information to create immersive content that can be streamed to user terminals.
Enables immersive content with higher realism and accessibility on existing hardware, supporting various industrial fields through real-time streaming and reducing the weight of high-quality image and metadata, facilitating easy use on smartphones, PCs, tablets, and HMDs.
Smart Images

Figure KR2024009982_30102025_PF_FP_ABST
Abstract
Description
Method for providing production and distribution services for lightweight immersive content and device therefor
[0001] The present invention relates to a method for providing a service for producing and distributing lightweight immersive content and a device therefor, and more particularly, to a method for producing lightweight immersive content by performing resource lightweighting on image data extracted from 360-degree image data or 3D data and then performing a mashup to recombinate the same, and to distributing the produced immersive content to a user terminal and a device therefor.
[0002] With the development and spread of HMD (Head-mounted Display), mobile terminals, 360-degree cameras, etc., technologies for Virtual Reality (VR) and Augmented Reality (AR) have begun to emerge, and recently, even eXtended Reality (XR) technology, which encompasses Virtual Reality (VR) and Augmented Reality (AR), has emerged.
[0003] Technologies like the above are beginning to attract increasing market attention, and recently, the Metaverse environment, which allows each user to have an independent, separate self-identity and engage in diverse activities in various fields such as society, economy, education, culture, and science and technology, is also beginning to gain attention and continues to attract attention.
[0004] However, environments like the metaverse, including the aforementioned ones, while sometimes blurring the lines, are ultimately a kind of virtual world distinct from the real world. Therefore, it's essential to provide users with a high level of realism (or a sense of realism) that mirrors reality. Furthermore, it's clear that only with this level of realism can these environments garner greater attention and further development.
[0005] To this end, various dedicated hardware (e.g., VR and AR devices) have been released to assist with this, but most are either excessively expensive or difficult to use or access, leading to their shunning by users, except for a select few enthusiasts in the relevant market. Consequently, metaverse environments (e.g., Roblox, Zepeto, etc.) have gradually excluded these devices and have begun to be provided indirectly, utilizing only readily available and relatively inexpensive smartphones, PCs, and tablets. As a result, the healthy development of virtual worlds and related environments, including the metaverse, has also stagnated.
[0006] Furthermore, if the metaverse environment provided to users is provided only through smartphones, PCs, and tablets that can be easily found around them, excluding dedicated hardware that can increase their immersion, it can be said that it is nothing more than an indirect virtual world that can only be indirectly experienced through the display, and furthermore, it can be seen as completely different from a direct virtual world that can directly feel realism from a first-person perspective.
[0007] Therefore, in order to provide users with an extended reality (XR) or metaverse environment that allows them to experience a virtual world that feels real, it can be seen that immersive technology must be supported to deeply immerse users in the virtual world.
[0008] Immersive technology can be seen as a digital experience that can replicate the offline experience. In other words, it allows users to experience digital content as if it were a real-world experience. This allows virtual, augmented, and mixed reality technologies to support natural interaction between online and offline environments.
[0009] These immersive technologies, or immersive content technologies, are particularly effective in knowledge transfer, and the creation of "realistic digital experiences" is becoming increasingly important. The digital transformation driven by immersive content is impacting industries across the board, and we are currently evolving beyond the mobile computing era into the era of spatial computing.
[0010] Even from the perspective of content consumption, platforms are evolving alongside digital content, which has evolved from text to images and finally to videos. In the era of text content, platforms such as blogs and communities emerged, centered around internet portals. The advancement of mobile and communications technologies has improved the accessibility of images and videos, ushering in the era of social media and video platforms. If we consider text as one-dimensional, images are two-dimensional, and videos are three-dimensional, combining image and time. Future content is expected to add a new dimension to video: immersive content that adds experience.
[0011] So, going back to the beginning, you could say that dedicated hardware is the only way to increase immersion, but from a different perspective, you could also consider improving the content itself provided to users. Specifically, if you can provide users with more realistic immersive content, this too can significantly improve their immersion.
[0012] Reflecting these practical needs, authoring software has been developed to build virtual spaces and enhance the immersive experience of related content. Specifically, examples include editing tools for building virtual spaces using VR technology and 360-degree images, and authoring software like Unity and Unreal for adding immersive experiences to 3D video games, 3D animation, and 3D architectural visualization.
[0013] However, the conventional method above, that is, the method of constructing a virtual space using 360-degree images, has limitations in terms of space because it requires an actual space for shooting, and therefore the space that can be implemented is also limited. In addition, there is the difficulty of having to configure a large number of editing tasks and a separate server infrastructure in the process of distributing it to the web, etc. using a separate editing tool.
[0014] Furthermore, even if all 360-degree images were created using computer graphics to overcome this issue, the burden of creating 360-degree images for each desired viewpoint would be significant, making detailed editing difficult once the image was created. Furthermore, conventional authoring software relies on creating virtual reality content on a local PC and then uploading it to the web, inevitably requiring excessive network traffic and inherently hindering accessibility.
[0015] For the reasons mentioned above, the production and distribution of immersive content is still limited to mobile and PCs through specific mobile applications or PC installable programs, but this method is not easily accessible from the user's perspective.
[0016] Therefore, it can be seen that a technology is required to produce immersive content that can be easily used on existing hardware such as smartphones, PCs, tablets, and HMDs, but that offers a higher level of realism. Furthermore, it can be seen that a technology that combines the produced immersive content with unit technologies that can meet the needs of various industrial fields and distributes it is also required for the healthy development and activation of the relevant industry.
[0017] Therefore, there is a need for specific methods to solve the above problems.
[0018] In order to solve the above-mentioned conventional problems, one task of the present invention is to provide a service for producing and distributing lightweight immersive content.
[0019] Another object of the present invention is to provide a technology for packaging 3D data for distribution of immersive content.
[0020] Another object of the present invention is to provide a service for distributing immersive content including virtual space through real-time streaming.
[0021] Another task of the present invention is to provide immersive content that can be easily used on existing hardware such as smartphones, PCs, tablets, and HMDs, but provides a higher level of realism.
[0022] Another task of the present invention is to distribute immersive content including virtual space by combining it with unit technologies that can meet the needs of various industrial fields.
[0023] Another object of the present invention is to provide a service for extracting high-quality 360-degree image data and metadata from 3D data produced using various 3D tools.
[0024] According to the present invention, a service is provided for reducing the weight of resources such as high-quality 360-degree image data and metadata extracted from 3D data.
[0025] According to the present invention, a service is provided for creating immersive content by mashing up lightweight, high-quality 360-degree image data and metadata, and distributing the same by streaming it to a user terminal.
[0026] The technical problems to be achieved in the present invention are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.
[0027] In order to solve the above-described problem, a method for providing a service of producing and distributing lightweight immersive content by a server according to an embodiment of the present invention may include the steps of: receiving content data including at least one of 360-degree image data or 3D data for 3D content from a first user terminal; performing resource lightweighting on the received 360-degree image data or image data extracted from the 3D data; and performing a mashup that recombines the image data on which the resource lightweighting was performed according to a predetermined standard when the server receives a request for transmission of immersive content from a second user terminal to produce lightweight immersive content, and then distributing the produced immersive content to the second user terminal.
[0028] According to one embodiment of the present invention, the 360-degree image data in the content data received from the first user terminal may include at least one image taken in a 360-degree × 360-degree full-field of an actual space using a camera or at least one image in a 360-degree × 360-degree full-field of a virtual space created using computer graphics.
[0029] According to one embodiment of the present invention, 3D data for 3D content in the content data received from the first user terminal may include 3D data for a virtual space created through a 3D tool including a game engine or a 3D modeling program.
[0030] According to one embodiment of the present invention, the image data extracted from the 3D data may include at least one image for a 360 degree × 360 degree omnidirectional direction of the virtual space extracted from the 3D data for the virtual space.
[0031] According to one embodiment of the present invention, the step of performing the resource weight reduction may include converting and storing the image data extracted from the 360-degree image data or the 3D data into a low-resolution image of a first predetermined capacity consisting of one side, a medium-resolution image of a second predetermined capacity divided into 6 sides, and a high-resolution image of a third predetermined capacity divided into 24 sides, respectively.
[0032] According to one embodiment of the present invention, the step of performing the resource lightweighting includes performing resource lightweighting on object data of 3D models and objects in the 3D data when the content data received from the first user terminal is 3D data for the 3D content, and the resource lightweighting on the object data may include performing at least one of conversion to a predetermined file format, vertex grouping to reduce the number of vertices for the 3D model and object, vertex merging to merge vertices for the 3D model and object, and resolution step subdivision to store the 3D model and object at a resolution divided into a plurality of steps of a predetermined size.
[0033] According to one embodiment of the present invention, the lightweight immersive content may be produced by performing a mashup that recombines image data on which resource lightweighting has been performed based on location and gaze information of a user of the second user terminal within the immersive content.
[0034] According to one embodiment of the present invention, when the content data received from the first user terminal is 3D data for the 3D content, the lightweight immersive content may be produced by performing a mashup that recombines the image data on which the resource lightweighting has been performed and the object data on which the resource lightweighting has been performed.
[0035] According to one embodiment of the present invention, the distribution may be performed by streaming.
[0036] In order to solve the above-described problem, an apparatus for providing a service of producing and distributing lightweight immersive content according to an embodiment of the present invention comprises a memory unit and a control unit, wherein the control unit receives content data including at least one of 360-degree image data or 3D data for 3D content from a first user terminal, performs resource lightweighting on the received 360-degree image data or image data extracted from the 3D data, and when a request for transmission of immersive content is received from a second user terminal, performs a mashup that recombines the image data on which the resource lightweighting has been performed according to a predetermined standard to produce lightweight immersive content, and then controls distribution of the produced immersive content to the second user terminal.
[0037] According to the present invention as described above, the effects described below can be achieved. However, the effects achieved through the present invention are not limited thereto.
[0038] According to the present invention, there is an effect of being able to provide a service for producing and distributing lightweight immersive content.
[0039] According to the present invention, there is an effect of being able to provide a packaging technology for 3D data for distribution of immersive content.
[0040] According to the present invention, there is an effect of being able to provide a service that distributes immersive content including virtual space through real-time streaming.
[0041] According to the present invention, it is possible to provide immersive content with a higher sense of realism while being easily usable on existing hardware such as smartphones, PCs, tablets, and HMDs.
[0042] According to the present invention, there is an effect that immersive content including virtual space can be distributed by combining it with unit technologies that can meet the needs of various industrial fields.
[0043] According to the present invention, there is an effect of being able to extract high-quality 360-degree image data and metadata from 3D data produced through various 3D tools.
[0044] According to the present invention, there is an effect of reducing the weight of resources such as high-quality 360-degree image data and metadata extracted from 3D data.
[0045] According to the present invention, there is an effect of being able to create immersive content by mashing up lightweight, high-quality 360-degree image data and metadata, and to distribute the same by streaming it to a user terminal.
[0046] FIG. 1 is a schematic diagram illustrating a system (1) for providing a production and distribution service of lightweight immersive content according to one embodiment of the present invention.
[0047] FIG. 2 is a drawing for explaining a function of extracting visual data from 3D data according to one embodiment of the present invention.
[0048] FIG. 3 is a diagram for explaining a weight-based interpolation algorithm according to one embodiment of the present invention.
[0049] FIG. 4 is a drawing for comparing the degree of staircase phenomenon according to a conventional signal removal method and a signal removal method according to an embodiment of the present invention.
[0050] Figure 5 is a diagram for explaining metadata extracted from 3D data.
[0051] FIG. 6 is a drawing for explaining a real-time 360-degree image extracted in a form in which light of a predetermined size is irradiated according to one embodiment of the present invention.
[0052] FIG. 7 is a drawing for explaining a method of dividing a real-time 360-degree image into three stages representing different resolutions and storing them by tiling, according to one embodiment of the present invention.
[0053] FIG. 8 is a drawing for explaining a vertex merging step according to one embodiment of the present invention.
[0054] FIG. 9 is a drawing for explaining a LOD method for performing mesh simplification of a 3D model into multiple level units representing various resolutions according to one embodiment of the present invention.
[0055] FIG. 10 is a drawing for explaining a cubic rendering method for mashing up real-time 360-degree images according to one embodiment of the present invention.
[0056] FIG. 11 is a diagram comparing the loading time and data consumption of a real-time 360-degree image according to one embodiment of the present invention and a 360-degree image according to the prior art.
[0057] FIG. 12 is a drawing for explaining a method of moving space using a mouse pointer according to one embodiment of the present invention.
[0058] FIG. 13 is a drawing for explaining a mouse pointer expressing material characteristics of a 3D model according to one embodiment of the present invention.
[0059] FIG. 14 is another drawing for explaining a mouse pointer expressing material characteristics of a 3D model according to one embodiment of the present invention.
[0060] FIG. 15 is a drawing for explaining a block diagram of a server according to one embodiment of the present invention.
[0061] Figure 16 is a detailed block diagram of a server according to one embodiment of the present invention.
[0062] Figure 17 is a detailed block diagram of a processor according to one embodiment of the present invention.
[0063] FIG. 18 is a drawing for explaining a view provided in relation to an editing function for immersive content according to one embodiment of the present invention.
[0064] Hereinafter, preferred embodiments of the present invention will be described in detail with reference to the accompanying drawings. The detailed description set forth below, together with the accompanying drawings, is intended to illustrate exemplary embodiments of the present invention and is not intended to represent the only embodiments in which the present invention may be practiced.
[0065] These examples are provided solely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined solely by the scope of the claims.
[0066] In some cases, to avoid obscuring the concepts of the present invention, well-known structures and devices may be omitted or illustrated in block diagram form focusing on the core functions of each structure and device. Furthermore, the same components are described using the same reference numerals throughout this specification.
[0067] Throughout the specification, when a part is said to "comprising" or "including" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise stated.
[0068] Additionally, the term "unit" described in the specification means a unit that processes at least one function or operation, which may be implemented by hardware, software, or a combination of hardware and software. Furthermore, the terms "a" or "an," "one," and similar related words may be used in the context of describing the present invention to encompass both singular and plural meanings, unless otherwise indicated in the specification or clearly contradicted by the context.
[0069] Additionally, specific terms used in the embodiments of the present invention are provided to aid in understanding the present invention. Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by those of ordinary skill in the art to which the present invention pertains. The use of these specific terms may be modified in other forms without departing from the technical spirit of the present invention.
[0070] Hereinafter, various embodiments related to a method for providing 3D data packaging and streaming services for distribution of immersive content according to the present invention and a system therefor are provided.
[0071] In particular, this specification discloses a server that streamlines 360-degree image data or 3D data while enabling distribution to user terminals, and provides a system including the server. The server may include various programs, software, and hardware for the present invention.
[0072] Hereinafter, "content" in this specification includes all forms of content that exist or will be developed in the future. For convenience, 3D (3-dimensional) content is used as an example, but is not necessarily limited thereto. In addition, such 3D content may be a concept that includes virtual space itself according to virtual reality (VR), augmented reality (AR), mixed reality (XR), and metaverse environments, or various immersive contents related thereto.
[0073] Meanwhile, terms used in this specification that are not specifically explained and technologies related to such terms may be interpreted according to or refer to known technologies, and it is hereby made clear in advance that only the necessary content is described.
[0074] In addition, the present invention may be combined with ICT (Information and Communication Technology) technology, such as artificial intelligence (AI) technology or blockchain network-based technology, as needed, and if a person with ordinary knowledge in the relevant technical field combines the present invention with the above technology, the combined technology falls within the technical scope of the present invention.
[0075] Referring to the attached drawings, a method for providing 3D data packaging and streaming services for distribution of immersive content according to the present invention is described as follows.
[0076] FIG. 1 is a schematic diagram illustrating a system (1) for providing a production and distribution service of lightweight immersive content according to one embodiment of the present invention.
[0077] Referring to FIG. 1, a system (1) for providing a production and distribution service of lightweight immersive content is illustrated, and the system (1) can be implemented by including a user terminal 1 (110), a user terminal 2 (120), and a server (130).
[0078] However, the configuration of the above system (1) is not necessarily limited to this, and the above system (1) may be implemented by including one or more other components depending on the embodiment. In addition, for convenience of explanation, only two terminals are illustrated, but this is not limited thereto.
[0079] User terminal 1 (110) may represent a terminal that directly creates 3D content using a 3D tool and possesses (or stores) the created 3D content data. In some embodiments, user terminal 1 (110) may represent a device that has not directly created the 3D content, but currently possesses or is connected to a medium that possesses the created 3D content data and is capable of receiving the created 3D content data.
[0080] The "3D tools" described herein may include game engines (e.g., Unreal Engine, etc.) and modeling programs such as Autodesk, 3DS Max, Unity, SketchUP, and Blender.
[0081] Meanwhile, user terminal 1 (110) may represent a terminal that produces 360-degree image data and holds (or stores) the produced 360-degree image data. According to an embodiment, the 360-degree image data may be at least one image (or video) taken in all directions of 360 degrees × 360 degrees of an actual space using a camera, or alternatively, may include at least one image for all directions of 360 degrees × 360 degrees of a virtual space produced using computer graphics.
[0082] Meanwhile, for the production and distribution service of immersive content according to the present invention, user terminal 1 (110) can download and install a plug-in provided by the system (1) or server (130) of the present invention.
[0083] Here, a plug-in refers to a program added to enhance the function of a specific program, and the "plug-in" described in this specification refers to an object that performs functions such as various preprocessing tasks for importing at least one 3D data or 360-degree image data created by various 3D tools to the web and converting it into 3D content data.
[0084] In the present invention, a plug-in of a corresponding type or version can be provided depending on the type of 3D tool and OS used in the terminal. Through this, the server (130) can provide plug-ins of various types or versions and enable the terminal to download and install one or more of them. Of course, only one plug-in can be provided, and this one plug-in can be made to correspond to all types of 3D tools and OS used in the terminal.
[0085] However, the present invention also includes performing the program by providing it in the form of an independent program rather than a plug-in, and for this purpose, the user terminal 1 (110) can download and install a program related to the present invention provided by the system (1) or server (130) of the present invention.
[0086] In addition, unlike this, the present invention can enable the user terminal 1 (110) to utilize 3D data packaging for distribution of lightweight immersive content described below on the web through a web service or platform controlled by the system (1) or server (130) even without installing the plug-in or independent program.
[0087] Meanwhile, if the user terminal 1 (110) is a 3D data production terminal, the user terminal 2 (120) may represent a user-owned terminal that can view the 3D content data and participate in interactions, etc. when the 3D content data is distributed through a streaming method or the like via a server (130).
[0088] That is, in this specification, in order to help understand the present invention and for convenience of explanation, the user terminal 1 (110) and the user terminal 2 (120) are described as being divided into a 3D data production terminal and a 3D content viewing terminal, respectively, but this is not necessarily limited to this, and their roles may be changed. In addition, unlike what is illustrated in FIG. 1, there may be multiple user terminals 1 (3D data production terminals) and user terminals 2 (3D content viewing terminals), and the system (1) of the present invention can provide various supports so that a smooth service can be provided when there is simultaneous access or use of them.
[0089] User terminal 1 (110) and user terminal 2 (120) may each be in the form of a fixed terminal such as a PC or TV, or a mobile terminal such as a smartphone, laptop, tablet, HMD, etc.
[0090] According to an embodiment, user terminal 1 (110) and user terminal 2 (120) may each be in the form of a lightweight immersive content production and distribution service providing system (1) according to the present invention or a dedicated terminal device for the service.
[0091] In another embodiment, user terminal 1 (110) and user terminal 2 (120) may each be a wearable device or a device that operates in conjunction with another terminal. However, user terminal 1 (110) and user terminal 2 (120) according to the present invention are not necessarily limited to the examples described above, and any device including software and hardware necessary for using the system (1) or service is sufficient.
[0092] Meanwhile, it is not necessary for multiple user terminals 1 (110) and 2 (120) to have the same form or performance. Furthermore, it is not necessary for multiple user terminals 1 (110) to use the same 3D tool to create the same type of 3D content, and furthermore, it is not necessary to download and install the same plug-in or the same program from the server (130).
[0093] User terminal 1 (110) can subscribe to and / or log in to a web service related to the present invention operated / controlled by the server (130), package or edit 3D data received from another device or created by the user using a 3D tool into 3D content for distribution to user terminal 2 (120), etc., through the web service related to the present invention, and then distribute the same to user terminal 2 (120), etc., by streaming.
[0094] When the production of 3D data by 360 image data or a 3D tool is completed on user terminal 1 (110), the produced 360 image data or 3D data can be transmitted to the server (130), specifically, to the cloud platform of the server. At this time, the transmission can be performed according to settings such as real-time, a predefined specific time, or the time of user terminal 1 (110) logging in to the web service, or alternatively, it can be performed automatically, or it can be performed through a pre-installed plug-in or an independent program.
[0095] User terminal 2 (120) can subscribe to and / or log in to the web service related to the present invention operated by the server (130), and watch or participate in interaction with immersive content distributed by user terminal 1 (110) through streaming, etc. Depending on the embodiment, user terminal 2 (120) can separately download and install a player provided by the server (130) or the corresponding web page, and then participate in or view immersive content distributed through the player. Alternatively, user terminal 2 (120) can access the URL address of the immersive content, log in, or receive an environment for participating in or viewing the immersive content without logging in.
[0096] In order to use the 3D data streaming service for distributing the immersive content of the present invention through the player, the user terminal 2 (120) may additionally be equipped with or support an interface such as a pointer, stylus pen, or mouse for appropriate content control, in addition to a display and speaker.
[0097] Meanwhile, as described above, the server (130) may provide the user terminal 1 (110) with a plug-in or an independent program corresponding to the 3D tool used by the user terminal 1 (110) when creating 3D data, and if the user terminal 1 (110) has 3D data, the server may allow the user terminal 1 (110) to import (or upload) the 3D data as is, not in cut units, through the plug-in or independent program, and may allow the user terminal 1 (110) to package (or edit) the 3D data on the web and then distribute it to the user terminal 2 (120) or the like by streaming. To this end, the server (130) may include the software and hardware necessary to perform the corresponding function.
[0098] The server (130) can build a cloud platform or database (DB) so that 3D data of user terminal 1 (110) can be imported onto the web through the plug-in or independent program for the 3D data packaging service for distribution of the immersive content of the present invention.
[0099] Meanwhile, the user terminal 1 (110) may store a token issued by communicating with the server (130) for logging in to the cloud platform and / or importing 3D data, and the server (130) may also store the token to determine whether the user has authority and for login management.
[0100] The server (130) may support a related API (Application Program Interface) to support data communication between user terminal 1 (110) that produces 3D data and distributes 3D content and / or user terminal 2 (120) that views the distributed 3D content.
[0101] In addition, communication or data communication between components belonging to the lightweight immersive content production and distribution service provision system (1) of the present invention can be performed through various wired / wireless communication networks developed to date or to be developed in the future, and it is not necessary for all components to use the same communication network or communication protocol.
[0102] Meanwhile, when user terminal 1 (110) creates 3D content using a 3D tool (e.g., game engine, e.g., Unreal engine, Autodesk, 3DS Max, Unity, SketchUP, Blender) used in creating 3D content, the user terminal 1 (110) can perform a preprocessing operation to import 3D data for the 3D content (including virtual space, etc.) to the web through the installed plug-in or independent program, and then convert it into 3D content for external distribution. To this end, the present invention provides a 3D data packaging technology.
[0103] Therefore, below, a 3D data packaging technology according to an embodiment of the present invention will be described in more detail.
[0104] The 3D data packaging technology of the present invention is a technology for easily and quickly building and distributing immersive content including virtual space that can be used in environments such as virtual reality (VR), augmented reality (AR), mixed reality (XR), or metaverse, and can simplify and automate the process of building and distributing immersive content including virtual space, and can solve problems of the existing conventional technology, so that users can more easily and quickly convert 3D data to build and distribute immersive content including virtual space.
[0105] Previously, the production of realistic / immersive content required repetitive processes such as modeling, rendering, and composition based on 3D data, and because multiple processes were required, it took a lot of time. In addition, the final results could not be provided through streaming methods such as the web, and could only be experienced as images, videos, or installations.
[0106] However, unlike existing web builders that receive and process 3D data in cut units, the 3D data packaging technology according to one embodiment of the present invention provides a plug-in or independent program that can be installed on a user terminal. In addition, since this allows the user terminal to easily import pre-produced 3D data onto the web, it is independent of the type of 3D tool used to produce the 3D data, and thus has excellent versatility and convenience.
[0107] Additionally, the terminal can utilize a 3D data packaging service to convert 3D data imported from the web via the plug-in or independent program into 3D content representing its own virtual space. Once the conversion into 3D content representing its own virtual space is completed via the 3D data packaging service, the 3D content can be transmitted to other terminals via streaming, thereby inducing interaction that enhances participation and immersion in the virtual space.
[0108] In other words, the present invention standardizes the process of building and distributing immersive content, including virtual spaces, using a standardized pipeline, thereby dramatically reducing the iterative process time. Furthermore, as will be described later, the present invention enables the operation of 3D content not only in various formats, such as conventional 360-degree images and videos, but also in various device environments, such as web, mobile, and installed platforms, thereby providing greater accessibility and versatility.
[0109] In addition, it provides the functions of building, managing and distributing immersive content including virtual space in the form of a plug-in or independent program, and since each function is provided in a modular unit as described below, it also has the effect of significantly lowering the physical / logical costs and entry barriers required for building, managing and distributing immersive content including virtual space.
[0110] Below, various technologies related to 3D data packaging are described in more detail.
[0111] 1. Extracting visual data from 3D data
[0112] In order to provide users with high-quality immersive content including virtual spaces, the quality of images sampled by shooting 360-degree virtual spaces is a crucial factor. However, due to the characteristics of virtual spaces that express the user's gaze through a camera and the characteristics of immersive content that serves as the source of real-time 360-degree images, the rendering effects applied to the camera vary in real time depending on the position or direction of the camera. Therefore, it is impossible to obtain high-quality 360-degree images required to provide immersive content services including virtual spaces in web or mobile environments using conventional image capture and image merging methods.
[0113] Accordingly, the present invention provides a function for extracting visual data (image data, etc.) from 3D data as illustrated in FIG. 2 to solve such a problem.
[0114] That is, when the terminal imports pre-produced 3D content data onto the web through the plug-in or independent program described above, the present invention performs image extraction, i.e., image sampling, for the 3D content data by rotating the camera in all directions (360 degrees x 360 degrees or 180 degrees x 180 degrees) at a predetermined specific location within the 3D content, based on a predetermined tile.
[0115] After that, the sampled image tiles are overlapped at predetermined locations and interpolated for each pixel value to obtain an interpolated pixel value. Based on the interpolated pixel values, the sampled image tiles are projected onto an equirectangular surface, which is a preset dedicated image format, to extract a high-quality 360-degree image for the 3D content data.
[0116] In the present invention, to distinguish it from prior art, the high-quality 360-degree image extracted as described above is referred to as a real-time 360-degree image (or 360-degree image data). Furthermore, in the above description, interpolation refers to a method of estimating the value of an unconfirmed point using confirmed data, and equirectangular may refer to a method of capturing a 360-degree sphere screen on a plane with a 2:1 ratio.
[0117] However, when the camera moves in all directions and samples images in tile units within the above 3D content, i.e., mixed reality (XR) content, it is important to specify the appropriate position, direction, and movement of the camera.
[0118] Therefore, for this purpose, the present invention uses mathematical models such as 3D matrix transformation, quaternion, and vector calculation, and the above models are not only used to determine the position of the camera, but also to express the direction of the camera, and can be used to calculate movement to an appropriate sampling point for a desired tile.
[0119] In addition, in the present invention, by applying Bezier Curves and Splines collision detection and avoidance algorithms to effectively implement camera movement within 3D content, the camera can be controlled to avoid collision with objects existing within the 3D content, i.e., mixed reality (XR) content.
[0120] For example, a Bezier curve is a mathematically created curve to express a curve of arbitrary shape. This curve passes through all control points, but has the property of having a control line that simultaneously indicates the direction and length, called a tangent vector, for each control point. In other words, the direction and length (vector) of this tangent line that passes through the curve at a specific control point represent the range affected by the curvature direction and the tangent vector of the spline curve, respectively. In addition, a characteristic of a Bezier curve is that the first and last points must be inside the curve, and the two control points always smooth the curve. Therefore, using a Bezier curve like this, you can easily express a curve and has the advantage of always drawing a smooth and accurate curve.
[0121] Accordingly, the present invention provides a curved camera movement by applying the principle of the Bezier curve to select points where the camera does not collide with other objects within the 3D content, i.e., mixed reality (XR) content, as respective control points, and moving the camera while passing through the corresponding control points. As a result, the camera within the 3D content may not collide with other objects within the 3D content, i.e., mixed reality (XR) content, and there is an effect that the camera may move in all directions within the mixed reality (XR) content and perform image sampling in units of tiles.
[0122] Meanwhile, after acquiring multiple images within the 3D content, i.e., mixed reality (XR) content, a process is performed to reassemble them into a single real-time 360-degree image (or high-quality 360-degree image) in a panoramic format. To this end, the present invention interpolates the pixel values as described above and projects the sampled image tiles onto an equirectangular surface, which is a preset dedicated image format, based on these values.
[0123] More specifically, first, multiple images extracted in tile units from 3D content can be combined in the form of a cube map, which is a cube shape corresponding to a hexahedron of the same size or different sizes. The cube map can divide the space within the 3D content into six areas of the same size or different sizes, and can store a 360-degree space or surrounding environment within the 3D content as a six-sided 2D image.
[0124] At this point, image discontinuity can occur at the boundaries of each image tile. Simply put, when two images are stitched together, differences in pixels or colors at their boundaries can significantly reduce the naturalness of the image. Specifically, if some images are enlarged or compressed, distortion or quality degradation can occur.
[0125] Therefore, if the sampled image tiles are simply combined or projected onto an equirectangular surface to be recombined into a 360-degree image without resolving these problems, the visual quality of the 360-degree image itself is bound to be quite low, as the stair-stepping or patterning phenomenon occurs as described below.
[0126] Therefore, to solve the above problem, it is possible to consider overlapping sampled image tiles at predetermined locations and performing linear interpolation on each pixel value. This linear interpolation is intended to minimize image discontinuities that may occur at the boundaries of each image tile, and is intended for so-called smooth tile image combination.
[0127] To do this, we need to obtain data related to the image pair to be combined. First, we need to select two arbitrary points from the image pair to be combined, the first point (e.g., and the second point (e.g., ) is obtained by specifying the value, and then the value is estimated based on the straight line between the first point and the second point, and specifically, the estimation is performed using the following mathematical formula 1. At this time, y is the value to be estimated, and x can be viewed as the location where the value of y is to be estimated.
[0128]
[0129] However, performing linear interpolation alone does not always provide optimal results for edge artifacts that can occur at the boundaries of image pairs to be combined. Edge artifacts here refer to visual defects in the image at the boundary, and can include, among other things, stair-stepping artifacts, where the boundary appears jagged.
[0130] As previously explained, the purpose of the present invention is to provide a technology capable of producing / distributing immersive content that can be easily used on existing hardware such as smartphones, PCs, tablets, and HMDs, while providing a higher sense of realism. Therefore, to achieve this, the visual quality of the immersive content itself needs to be further improved compared to that produced using conventional technologies.
[0131] Therefore, in order to more clearly solve this problem, the present invention has been improved to apply a weight based on the color profile (or color information) of each pixel, thereby enabling more precise adjustment for each pixel value.
[0132] A weight based on the color profile (or color information) of each pixel according to one embodiment of the present invention can be expressed by the following mathematical expression 2.
[0133]
[0134] At this time, represents the color difference between surrounding pixels, is a parameter that controls the strength of the effect of the difference on the weight. In addition, represents an exponential function, and indicates that the greater the color difference, the smaller the weight, and thus the less interpolation is applied. Therefore, an interpolation algorithm applying the above weight according to an embodiment of the present invention can be expressed by the following mathematical expression 3.
[0135]
[0136] By utilizing the weight-based interpolation algorithm according to one embodiment of the present invention, the influence of interpolation at points with large color changes can be reduced, thereby mitigating abrupt color changes that may occur at the boundaries of an image and significantly improving the natural gradation of the overall image. Furthermore, by adjusting the weights in consideration of the color profile of each pixel, there is an effect of significantly reducing the sense of incongruity that may occur at the horizontal and vertical boundaries when each face of a cube map is pasted.
[0137] The above weight-based interpolation algorithm can be applied in both horizontal and vertical directions, as illustrated in Fig. 3. Specifically, the four closest pixels to the boundary of each image in the original images to be combined are selected in a 2×2 grid format, and these pixels can be used as a reference for interpolation.
[0138] Next, horizontal interpolation is performed using the top two pixels and the bottom two pixels of the four selected pixels, respectively, in the process of calculating two middle pixel values for each row, and the calculation can be performed by applying weights to the x-coordinate of the target pixel.
[0139] Afterwards, the final pixel value can be obtained by performing vertical interpolation on the two intermediate pixel values obtained through horizontal interpolation. This can be done by applying a weight to the y-coordinate of the target pixel.
[0140] Regarding the application of the above weight-based interpolation algorithm, it was explained that horizontal interpolation is performed first and then vertical interpolation is performed. However, depending on the settings, vertical interpolation may be performed first and then horizontal interpolation may be performed later.
[0141] Meanwhile, sampled image tiles are overlapped at predetermined positions, and interpolated pixel values are obtained for each pixel value through the weight-based interpolation algorithm, and then seamless stitching is performed to seamlessly connect the sampled image tiles based on the obtained interpolated pixel values to complete a cube map in the shape of a hexahedron, and the completed cube map is projected onto an equirectangular surface to be converted into a 2D image, thereby extracting high-quality 360-degree image data (or real-time 360-degree image) for 3D content data. In the conversion process, a 3D direction vector corresponding to each pixel of the cube map can be calculated and then converted into a polar coordinate system, and the latitude and longitude values at this time can be used as the x, y coordinates in the equirectangular image.
[0142] Meanwhile, since the above method, i.e., equirectangular surface projection, is a process of converting a 3D image into 2D, it is easy for problems with imperfect image quality to occur, such as distortion or quality deterioration, such as part of the image being enlarged or compressed during the process, or a pattern phenomenon, which is an unseen visual pattern that can occur when two grids or lines overlap, or a staircase phenomenon, in which the outline is not smooth but rather uneven and has a staircase shape.
[0143] Therefore, according to one embodiment of the present invention, anti-aliasing can be performed to solve the above problem.
[0144] The above-mentioned signal removal is intended to minimize the staircase and patterning phenomenon that occurs when sampling high resolution to low resolution in an image or graphic. Conventional signal removal methods are performed by simply expressing the boundary of pixels smoothly or increasing transparency.
[0145] However, unlike the conventional method for removing false signals, in the present invention, first, multiple subsamples (smaller pixel units) are generated around each pixel, and then capture is performed for the multiple subsamples, thereby collecting color information of the original image based on the color values of the subsamples in pixel units smaller than the original pixels, and the color values of the collected subsamples are averaged and used as the color value of the original pixel.
[0146] According to the above method, the color transition between pixels (i.e., between the original pixel and the pixel of another original) can be continued more smoothly, and specifically, as shown in Fig. 4, it can be confirmed that the staircase phenomenon and the pattern phenomenon are significantly improved in (a) using the false signal removal method according to the present invention compared to (b) using the conventional false signal removal method, and through this, it is possible to extract high-quality 360-degree image data (or real-time 360-degree image).
[0147] Furthermore, according to one embodiment of the present invention, each of the above steps / functions / methods can be implemented in a modularized manner, thereby enabling the entire real-time 360-degree image extraction process to be modularized and / or automated. Furthermore, immersive content based on the extracted 360-degree image data (or real-time 360-degree images) can be distributed.
[0148]
[0149] 2. Metadata extraction
[0150] In order to distribute real-time 360-degree image data extracted from 3D data by reorganizing it in various environments (web / mobile / HMD, etc.), metadata about the 3D coordinates and directions of resources such as 3D models, markers (including icons, moving points, or objects that users can interact with), etc., including real-time 360-degree images, must also be transmitted.
[0151] Accordingly, the present invention can provide a function for protocolizing metadata required for 3D content distribution, extracting the metadata, and then loading / managing the extracted metadata for each immersive content such as virtual space.
[0152] The metadata extracted in the present invention (or metadata of 3D models, markers (objects, etc.)) may include, as illustrated in FIG. 5, location / direction information of 3D models, markers (objects, etc.), distance information between moving points, and user visible area information, but is not limited thereto, and should be interpreted as a concept that includes everything related to the extracted 3D models, markers (objects, etc.).
[0153] Specifically, as for the position / direction information, the spatial coordinates (position) and direction (pivot position, scale, rotation degree) occupied by a specific image within a 3D model are extracted, and as for the distance information between moving points, the coordinates of moving points that the user can move within immersive content such as virtual space and the actual distance between each moving point are extracted (pivot position, distance, connection link), and the user visible area information is information about the user's viewpoint and the range of content that can be viewed from that viewpoint within immersive content such as virtual space, and may include calculating and extracting the visible area (pivot position, scale, rotation degree) based on the user's position and gaze direction.
[0154] Meanwhile, to efficiently store diverse 3D data information and facilitate easy access and parsing later, it is necessary to define the metadata data structure in a predefined, specific format. To this end, the present invention defines the metadata data structure in the JSON (JavaScript Object Notation) data format, but this is not limited thereto and, depending on the embodiment, it may be defined in other data formats.
[0155] As previously explained, using JSON as a data structure for metadata can have the advantage of allowing each element of the metadata to be clearly identified, making the data easy to parse or share with other systems, and allowing for high extensibility, allowing for easy modification of the data structure when new information or elements are added.
[0156] Meanwhile, in order to accurately position and independently move 3D models and markers (objects, etc.) in immersive content such as virtual spaces, the coordinate system must be aligned and the rotation center axis (pivot) must be corrected.
[0157] Since the data of the above extracted 3D model and marker (object, etc.) records the location, direction, size, etc. based on the global coordinate system, if it is not converted to the local coordinate system, inconsistency in operation may occur in the new global coordinate system rather than the original environment, and the problem of significantly reducing work efficiency may occur.
[0158] Therefore, according to one embodiment of the present invention, a transformation matrix is calculated and applied to compensate for the difference between the global coordinate system and the local coordinate system of the 3D model through an affine transformation matrix operation to match the two coordinate systems, and the position vector is stored separately to accurately adjust the position, rotation, and scale of the model later.
[0159] Therefore, in the present invention, the position, scale, and rotation values on the global coordinate system are matched by performing an affine transformation matrix operation with each vertex of a 3D model and a marker (object, etc.) while the position vector is stored separately so that the size, position, and rotation center axis can be corrected, and the 3D model and marker (object, etc.) extracted through this have the effect of being able to be independently controlled in other environments (web / mobile / HMD, etc.).
[0160] Meanwhile, real-time 360-degree images that can be extracted using the above method are high-quality but static, so there may be limitations in the user's ability to control changes to the material or structure of 3D models, markers, etc. Therefore, in the present invention, extraction can be performed in the form of light of a predetermined size being irradiated from a predetermined direction on the surface of a 3D model or marker, as illustrated in FIG. 6, and depending on the embodiment, this may be based on WebGL (Web Graphic Library), but is not limited thereto.
[0161]
[0162] 3. Resource lightweighting
[0163] In order to smoothly provide immersive content including virtual spaces in environments such as mobile and PC, it is necessary to reduce the weight of extracted real-time 360-degree images, 3D models, and markers.
[0164] At this time, the above resource lightweighting is not only considered in terms of technology and service implementation, but is also closely related to capacity reduction and cost reduction. Therefore, when providing services to a large number of users, it can have the advantage of increasing user experience and minimizing costs by reducing network traffic.
[0165] Accordingly, the present invention includes functions such as image optimization and model optimization as described below, thereby enabling resource lightening and thereby minimizing costs.
[0166] Meanwhile, if real-time 360-degree images extracted from 3D data are loaded as they are to provide immersive content that includes virtual space, users will have to wait a long time for loading, and especially when moving locations within immersive content that includes virtual space, the entire image must be loaded each time the next image is loaded, which may cause the inconvenience of experiencing a white screen or having to endure a very long waiting time during loading.
[0167] Accordingly, in order to prevent and solve such problems, the present invention provides a resource lightweight function by dividing real-time 360-degree images into three levels representing different resolutions (e.g., low resolution, medium resolution, high resolution) and storing them by tiling (multi-level tiling), thereby solving the above problems.
[0168] That is, as illustrated in FIG. 7, a real-time 360-degree image is divided into three stages (e.g., low resolution, medium resolution, and high resolution) representing different resolutions, and specifically, a low-resolution preview image of several tens of kilobytes (KB) representing a real-time 360-degree image on one side, a medium-resolution image of several hundred kilobytes (KB) representing a real-time 360-degree image divided into six sides, and a high-resolution image of several MB (Megabyte) representing a total of 24 sides by further dividing each of the six sides into four, and these can be divided into three multi-stages in total and stored by tiling.
[0169] Meanwhile, according to one embodiment of the present invention, a sparse texture tile technique for efficient data management can be provided.
[0170] In the above tiling process, when texture data of a specific area exhibits low variability, it can be made lightweight using the sparse texture tile technique according to an embodiment of the present invention.
[0171] To this end, we first analyze the color histogram of each tile and quantitatively evaluate the variability within the texture.
[0172] The above color histogram represents the frequency of each color value within the tile image. It can be calculated as shown in the mathematical expression 4 below, and the color histogram specific color Indicates the frequency with which the image appears within the corresponding image tile.
[0173]
[0174] Also, the variability of the texture The distribution of color values within a tile can be quantified by the standard deviation, entropy, and range of the histogram. For example, if the variability of a texture is defined by the standard deviation of the histogram, it can be calculated as in the mathematical expression 5 below. Here, is the number of pixels in a tile, means the histogram average of the entire color.
[0175]
[0176] Volatility If the variability is below a preset threshold, the tile is considered to have low information density (or frequency) and can be compressed to have a lower resolution than the preset criterion. Conversely, if the variability is below a preset threshold, the tile is considered to have low information density (or frequency) and can be compressed to have a lower resolution than the preset criterion. If the tile appears above a preset threshold, the tile is considered to have high information density (or frequency) and can be compressed to have a higher resolution than the preset criterion.
[0177] The above approach can be mainly applied to areas with low variability and high homogeneity, such as the sky or a monochromatic background where the same color is continuously present.
[0178] Utilizing this method will significantly reduce data transfer and processing load when rendering immersive content on user devices in the future, ultimately significantly improving the rendering efficiency of the user device. This method is particularly crucial for optimizing memory usage and performance when processing large-scale immersive content (or large-scale XR spaces), and can significantly contribute to providing users with a smooth and fast graphics experience.
[0179] Meanwhile, in environments where the use of GPUs (Graphics Processing Units) is limited, such as the web, when displaying immersive content including virtual spaces, such as 3D models and markers, if the 3D model or marker itself has a large number of vertices, the rendering time may become long, causing users to experience screen tearing or slowdowns.
[0180] Therefore, in the present invention, as another resource lightweight function to prevent and solve such problems, a file format conversion step for vertex reduction and a vertex merging / grouping step (mesh simplification) can be performed.
[0181] First, the 3D file format is checked. If the 3D file already meets the predefined file format criteria, the file format conversion step may be omitted. However, if the 3D file does not meet the predefined file format criteria, file format conversion may be performed.
[0182] For example, assuming that the predetermined file format standard is OBJ, GLTF, if the format of the confirmed 3D file is FBX file format, this does not correspond to the predetermined file format standard (OBJ, GLTF), and therefore file format conversion can be performed as described above. At this time, the predetermined file format standard may be a file format that allows a 3D model or marker to be expressed with a smaller number of vertices than the predetermined number, and is not limited to the OBJ and GLTF described as examples above.
[0183] 3D data originating from Autodesk's graphic production software (such as 3Ds MAX) as well as game engines (such as Unreal and Unity) are usually saved in the FBX file format.
[0184] However, FBX contains all the 3D scene data required for rendering, including cameras, lighting, geometry information, and animations, so it has good compatibility with other digital content creation tools, but it has the problem of taking up too much space.
[0185] Therefore, in an environment where immersive content is distributed via the web, the OBJ or GLTF file format may be a more appropriate choice than FBX in terms of lightness and performance, and the extracted 3D data and metadata are converted thereto.
[0186] Additionally, when distributing immersive content over the web, you can reduce file size by removing unnecessary data or duplicate information that is not expressed, and by structuring and serializing it.
[0187] A file format conversion process according to one embodiment of the present invention may consist of the following four steps.
[0188] The first stage is the data extraction and parsing stage, which efficiently extracts the necessary 3D data and metadata from the original FBX file and analyzes and refines the data required in this process, such as vectors, vertices, texture coordinates, normal vectors, and UV mapping.
[0189] The second stage is the mesh compression stage, which analyzes the structure of the mesh to remove unimportant vertices and adjusts the positions of neighboring vertices to optimize the file size while preserving the quality of the mesh.
[0190] The third stage is the texture optimization stage, where the resolution and bit rate of the texture are dynamically adjusted to find the optimal balance between quality and file size.
[0191] The fourth step is the serialization and format conversion step, where the extracted data is serialized into OBJ or GLTF format to maximize the efficiency of the data structure and reduce the file size by removing redundant or unnecessary information.
[0192] Depending on the embodiment, the following steps may be further added, specifically, a cloud storage upload step, in which the optimized file is uploaded to cloud storage (e.g., Amazon S3, etc.) so that it can be immediately accessed and rendered in immersive content including virtual space.
[0193] Meanwhile, the vertex merging / grouping step according to one embodiment of the present invention may mean minimizing visual quality loss while reducing the number of vertices and triangles (polygons) of vertices of a 3D model by performing mesh simplification.
[0194] First, vertex clustering is performed through the process of grading vertices to assign weights according to high or low visual importance, triangulating vertices, clustering to group vertices based on geometric accessibility, integrating vertices to derive representative vertices, removing triangles (polygons) and vertices of duplicate vertices, and adjusting normals to form triangles by connecting the remaining vertices. As shown in Fig. 7, through this, it can be confirmed that 18 vertices are reduced to 12 vertices through the grouping.
[0195] Vertex grouping reduces the number of vertices within a 3D model using cluster-based geometric simplification methods. This can be done by clustering spatially adjacent vertices and selecting the vertex closest to the spatial center within each cluster as the representative.
[0196] Accordingly, by utilizing the above technical features of the present invention, the number of vertices of a 3D model or marker itself can be reduced, so that even when displaying a 3D model, marker, etc. in a virtual space in an environment where GPU use is limited, such as the web, the phenomenon of screen tearing or slowing down can be significantly reduced.
[0197] Vertex merging is a method to further reduce model complexity by merging adjacent vertices with similar properties. Vertex merging can be performed with a focus on removing unnecessary details while preserving the model's geometric form. Comparing the before (a) and after (b) vertex merging images, as shown in Figure 8, demonstrates the results of vertex merging, which removes unnecessary details while preserving the model's geometric form.
[0198] When reducing the number of vertices and triangles (polygons) of vertices, if you simplify indiscriminately without distinguishing between points (features) that contain important characteristics of the object, you may feel that the characteristics of the original are rapidly lost and the quality is greatly degraded.
[0199] Therefore, in the present invention, a quadric error measurement method (QEM, Quadric Error Metrics) is used to quantify vertices and edges that affect visual quality, and an algorithm is used to dynamically merge or remove them based on the quantification.
[0200] The above quadric error measurement method (hereinafter, QEM) defines the error between the original form of the model and the simplified form as a quadratic function and simplifies the mesh in the direction of minimizing this error. Here, the quadric (Q) can be expressed as the following matrix mathematical formula 6 using four coefficients (a, b, c, d) when the equation of the plane is ax + by + cz + d = 0.
[0201]
[0202] A quadric is connected to each vertex, and a new error quadric is calculated by combining the quadrics of two vertices as in the following mathematical expression 7, and the pair of vertices that minimizes it are merged.
[0203]
[0204] Error quadrics are used as a criterion for selecting or removing vertices during the triangle (polygon) reduction process of vertices, and play an important role in optimizing memory usage and performance when processing and rendering complex 3D models in real time in the mashup stage, ultimately contributing to providing users with a high-quality graphic experience.
[0205] Meanwhile, the present invention can provide a resolution step subdivision (LOD, Level Of Detail) method that performs mesh simplification of a 3D model into multiple level units representing various resolutions.
[0206] Resolution level-of-detail (LOD) methods can be used to dynamically render 3D model meshes at an optimized level of detail based on a user's viewpoint and interaction, while maintaining a balance of high performance and quality when the 3D model is later used as immersive content on the user's device.
[0207] To this end, first, a mesh simplification and classification step of the 3D model can be performed, and by applying a simplification algorithm based on the QEM, the 3D model can be subdivided into multiple stages of various resolutions and stored. For example, as illustrated in Fig. 9, the 3D model can be subdivided into 5 stages (the mesh of the 3D model is simplified from LOD 5 to LOD 1). The mesh simplification and classification step as described above can be performed with a focus on effectively reducing the size of the data while maintaining the visual integrity of the 3D model.
[0208] Afterwards, a user interaction-based LOD model determination and data-ization step can be performed. This step can be said to be a step of pre-setting data on which LOD model to display to the user based on the user's interaction, particularly the user's screen zoom ratio. This means data-izing the logic for selecting and rendering a mesh with an optimized level of detail when the user approaches or moves away from a specific 3D model within the immersive content or performs screen zoom while using the immersive content.
[0209] To explain more simply, when a user uses immersive content, depending on the user's location or the location where the user's gaze is focused, a 3D model with minimal mesh simplification (e.g., LOD 5 in FIG. 9) is selected and rendered for 3D models that fall within the user's gaze range and require detailed expression, while a 3D model with maximum mesh simplification (e.g., LOD 1 to 4 in FIG. 9) is selected and rendered for 3D models that fall outside the user's gaze range and do not require detailed expression. This can be said to be a method of setting data in advance on which LOD model to show based on the user's existing or current interaction.
[0210] Using the above method can significantly reduce the rendering load, resulting in an optimized user experience and efficient resource use.
[0211]
[0212] 4. Mashup
[0213] Mashup refers to the process of recombining (or repackaging) extracted and lightweight content for the purpose of distributing (or publishing) immersive content, and the purpose is to recombine various 3D models and metadata, etc. to enable users to interact with, explore, and manipulate immersive content in real time in a web environment.
[0214] This process combines image tiling techniques with web-based 3D rendering technology, ultimately enabling users to experience immersive 3D environments. Therefore, mashup technology is optimized for web services, making it suitable for large-scale users and ensuring smooth service delivery even in low-spec environments.
[0215] Therefore, for the distribution of immersive content, the resources and metadata sampled in the real-time 360-degree image extraction process and loaded on the server must be recombined and reproduced according to predetermined criteria. To this end, the present invention may apply the cubic rendering method as illustrated in FIG. 10 according to one embodiment.
[0216] The above-described cubic rendering method is a method of recombining real-time 360-degree image tiles and resources / metadata that have been stored by tiling as described above, as illustrated in FIG. 7, and is a method of rendering only the necessary images according to the user's viewpoint (i.e., the location where the user's gaze is focused) and interaction, while also performing step-by-step loading over time. Therefore, according to one embodiment of the present invention, immersive content including virtual space can be provided at a fast speed, and there is also an effect of minimizing network traffic consumption.
[0217] Specifically, when a user quickly moves to a specific location in an immersive content provision service that includes a virtual space, only some of the tiles visible in the foreground need to be loaded at the location where the user briefly stops while moving, so only about 1 / 6 of the data that originally needed to be loaded needs to be loaded, which can significantly reduce the burden on the terminal and the amount of data consumed.
[0218] Furthermore, predicting the user's movement path and preloading and rendering tiles can result in much smoother screen transitions and minimize network traffic consumption. Since only a portion of the tiles are loaded based on the user's field of view, representing only one-sixth of the entire 360-degree image, the computational and communication requirements on the user's device can be reduced accordingly.
[0219] As illustrated in FIG. 11, in the case of a general 360-degree image, the initial screen loading time (time from user entry to first screen rendering) took 1650 ms, but when utilizing a real-time 360-degree image and a method of tiling and storing the same according to an embodiment of the present invention, the initial screen loading time takes 320 ms, showing an effect of about 80% reduction, and the data consumption when using an immersive content service including a virtual space can also be reduced by about 40% compared to a general 360-degree image. Therefore, according to the present invention, when using an immersive content service including a virtual space, there is an effect of minimizing network traffic consumption, terminal burden, and data consumption.
[0220] Meanwhile, according to one embodiment of the present invention, progressive image loading, adaptive image delivery, and local caching functions may be provided.
[0221] First, the progressive image loading feature is a feature that minimizes delay time in the user interface based on the LOD described above, and has the following characteristics.
[0222] (1) Multi-level resolution extraction and initial rendering: 360-degree images or 3D models that are saved in various resolutions (low, medium, high, or LOD 1 to LOD 5, etc.) during the extraction stage can be rendered using the image or 3D model with the lowest resolution when the user first accesses the site. Initial rendering in this way can prevent a white screen from appearing on the user's terminal until the high-resolution image is fully loaded, and can further prevent the user from experiencing long waiting times.
[0223] (2) Progressive Resolution Upgrade: As a user navigates an immersive content space (e.g., an XR space) that includes a virtual space, low-resolution images loaded initially can be sequentially upgraded to medium-resolution and then high-resolution images for rendering and display. This process can be achieved by continuously monitoring the user's viewpoint (i.e., where their gaze is focused) and interactions.
[0224] (3) Application of dynamic loading algorithm: This can be done by collecting interaction data such as the user's viewpoint, movement speed, and direction, as well as environmental data such as network status and device performance in real time, and analyzing the collected data to determine and transmit the images to be transmitted next in the server area.
[0225] Next, regarding the adaptive image transmission function, when a user uses immersive content, selecting an appropriate quality image to load on the user's terminal is affected not only by the user's field of view, but also by the target device's specifications, screen resolution, and network speed (3G, 4G, Wi-Fi, etc.). Therefore, it can be utilized as an image transmission technology that enhances the user experience by detecting this in real time, converting it into an image with a specific resolution and compression ratio according to predetermined criteria based on this information, and ultimately transmitting it to the user's device.
[0226] The adaptive image transmission function can be performed in the following order:
[0227] (1) Real-time environment detection stage: In this stage, the specifications of the user terminal, screen resolution, and network status (e.g., 3G, 4G, Wi-Fi, etc.) can be continuously monitored.
[0228] (2) Resolution adjustment engine stage: In this stage, the image resolution can be dynamically adjusted according to the screen resolution and GPU rendering capability of the user terminal to ensure that the image is rendered at the optimal resolution while preserving the overall quality of the image.
[0229] (3) Network Adaptive Compression Stage: This stage compresses images at an optimal compression rate based on the current network speed and conditions. For example, on slow 3G connections, a high compression rate can be used to minimize transmission time, while on fast Wi-Fi connections, a low compression rate can be used to improve image quality.
[0230] (4) Server-side image conversion stage: This stage performs image conversion in real time, allowing users to always receive optimized images at high speed.
[0231] This maximizes user experience by providing images that are most appropriate for the user's device and network conditions in real time, and promotes efficient use of resources.
[0232] Next, regarding the local caching feature, image tiles can be stored locally on the user's device. When the user revisits immersive content containing virtual spaces, these cached tiles can be directly retrieved from the user's device's local storage, enabling rapid rendering without additional network requests. This reduces data usage on the user's device and significantly improves the response time of immersive content containing virtual spaces.
[0233] During this process, the user's movement patterns, frequently visited areas, and recent tile access history are analyzed, and based on this information, the tiles that will be requested in the future with the highest probability can be predicted and cached preferentially (i.e., smart caching). At this time, the user's location within the immersive content and the zoom / reduction information of the screen (or viewpoint) can be identified in real time through the graphics library and 3D space recognition engine. Based on this information, not only the real-time changes in the user's location and screen (or viewpoint) within the immersive content, but also the behavioral patterns within the immersive content can be analyzed to store frequently used image tiles in the local cache in advance.
[0234] In addition, the local storage capacity and available resources of the user terminal can be monitored in real time, and the cache size can be dynamically adjusted based on this, and old tiles or tiles that are not used often can be automatically deleted to maintain the efficiency of the cache (i.e., cache optimization), the resolution of image tiles can be adjusted in real time according to the zoom level of the user's screen (or viewpoint), and the cache size can be adjusted by detecting the current network status and determining the optimal resolution for streaming.
[0235] Meanwhile, in the present invention, in relation to the spatial movement method used by a user using immersive content including a virtual space to move a location or gaze, etc., to a specific location within the immersive content, there is a problem that the conventional movement points (or conventional spatial movement methods) indicated by a plurality of arrows as shown in FIG. 12(a) can only provide limited and restricted movements. In order to solve this problem, a spatial movement method using a mouse pointer as shown in FIG. 12(b) can be provided.
[0236] More specifically, according to one embodiment of the present invention, when a user selects and clicks any area of the entire virtual space in order to move to a specific location on an immersive content service including a virtual space, the 2D coordinates (x, y) for the clicked area are checked, and then it is determined whether the clicked area is the floor within the 3D content through the checked 2D coordinates (x, y). If the clicked area is determined to be the floor, it can be determined that the user wishes to move, and in this case, among the predetermined mouse pointer areas that the user can move to, the mouse pointer area closest to the clicked area is searched, and then the user's location can be moved to the searched mouse pointer area.
[0237] At this time, the search for the mouse pointer area closest to the clicked area can be performed through Euclidean distance calculation, and specifically, after specifying at least one mouse pointer area adjacent to the clicked area, calculating the distance between the specified at least one mouse pointer area and the clicked area, and then specifying the mouse pointer area with the shortest calculated distance value as the mouse pointer area to which the user's position movement should be performed, and then rendering a tile image related to the mouse pointer area, the user's position movement can be performed to the mouse pointer area.
[0238] Meanwhile, the system (1) according to the present invention can identify the actual 3D spatial position and direction of each object or wall surface by utilizing metadata and precise x, y, z coordinate information extracted from a 3D model, and can define a scene movement point at an appropriate position and angle on a 2D image accordingly.
[0239] In this process, the height and surface angle of the texture are adjusted through Bump Mapping and Normal Mapping technologies, and further, when combined with texture mapping on a 2D image, the details and curves of the surface of an object or wall can be precisely expressed.
[0240] In addition, by using environment mapping technology to realistically express the reflection and texture of the material of an object on a 2D image, such as the physical characteristics and texture of the extracted 3D model, it is possible to provide a high level of immersion to the user.
[0241] Meanwhile, in the case of real-time 360-degree images, since the images are composed of 2D due to the nature of the images, even if a user positions a scene movement point on a 3D object such as an object or wall placed within immersive content that includes a virtual space, it is generally seen that the curvature of the 3D model or marker cannot be reflected and expressed in the scene movement point itself.
[0242] However, in order to provide a better user experience and immersion, the system (1) of the present invention enables the reflection material of a 3D model to be expressed even in a 2D image.
[0243] More specifically, when a user moves the mouse to a specific location on the screen, the 2D screen coordinates (e.g., x, y, z, etc.) of the specific location are recognized, and then the 2D screen coordinates are converted into geometry coordinates of 3D content (i.e., immersive content).
[0244] Afterwards, based on the extracted 3D coordinates, a mouse pointer of a specific shape (a circle is basically exemplified in FIG. 13(b) but is not limited thereto) is displayed within the immersive content, and for each pixel of the mouse pointer of the specific shape, a property value related to a reflection material for a 3D object on the 3D coordinates is output, so that, as shown in FIG. 13(b) and FIG. 14, the mouse pointer of the specific shape can express material features such as the curvature of a 3D model or marker (i.e., a 3D interactive pointing function) even in a real-time 360-degree image.
[0245] Meanwhile, the system (1) according to the present invention may further include a real-time rendering-based preset provision function for 3D data creation, a module unit marker provision function for interaction, a cloud login function, etc.
[0246] According to one embodiment of the present invention, the system (1) according to the present invention may further include a function of providing a real-time rendering-based preset (a set of predefined settings) for 3D data authoring in the process of providing a 3D data packaging and streaming service for distribution of immersive content.
[0247] The above presets are intended to provide a uniform approach for maintaining and utilizing immersive content including constructed virtual spaces, and can automatically apply the necessary setting values after pre-optimization. In the present invention, the types of presets include environment setting presets for constructing realistic spaces, material setting presets for setting various materials, and lighting presets that define actual light values, but these are merely examples and are not limiting.
[0248] That is, each set value or piece of information is individually embedded in the form of one or more presets, allowing the virtual space manager (including the creator) to conveniently apply them at any desired time. Therefore, the present invention facilitates the quality of virtual spaces by providing various quantified and automated presets, thereby reducing the time and effort required for production and lowering the skill requirements for creators.
[0249] Meanwhile, according to one embodiment of the present invention, all of the presets can be provided in a modularized form, such as an environment setting preset, a material setting preset, a lighting preset, etc., through which a virtual space can be constructed and maintained / utilized in module units.
[0250] Specifically, the present invention can provide a service that allows a virtual space manager to build and manage a virtual space in the form of a module-based structure and module-based functions through the functions described above and modularized presets, etc., when building a virtual space, and this also has the effect of reducing the time and cost burden consumed in building and managing a virtual space.
[0251] Meanwhile, the system (1) according to the present invention may further include a function of providing a module unit marker for interaction in the process of providing a 3D data packaging and streaming service for distribution of immersive content.
[0252] For virtual space services, there's a need to go beyond simply displaying a virtual space to users and allowing them to explore it, offering additional interactive activities. In other words, it's necessary to provide various forms of interactive objects or markers, such as icons, screens, and 3D models, within the virtual space to facilitate diverse user interactions while ensuring scalability.
[0253] According to one embodiment of the present invention, the object or marker operates based on web-specific functions, so it can be freely edited on a web-based converting tool separately from a virtual space, and can also be linked with various types of external services.
[0254] According to one embodiment of the present invention, the object or marker for the interactive activity can be applied during the mashup (or 3D content packaging process), and specifically, can be applied to be inserted into a specific area of a virtual space, or to express specific information or perform a predefined action through a viewer or the like when a specific button is clicked. For example, it can be applied to give a lecture in a conference hall in a virtual space through a video marker, or to display detailed product information through a viewer, and can be applied to send / receive data to / from a chat server using a Web Socket, or to conduct a group video chat or webinar in the viewer in the form of an iframe by linking an external API.
[0255] According to one embodiment of the present invention, objects or markers may include, but are not limited to, pop-up markers that display images, videos, PDFs, etc. in a modal form as pop-ups in a virtual space to enable acquisition of additional information; viewer markers used to play images or videos on a specific wall or object in a virtual space; text markers for conveying information by expressing text in a virtual space; link markers that connect external websites as in / out links when clicked; 3D model markers that can be used to mark models themselves having various shapes, such as human bodies, placed in a virtual space and grant them the functions of basic markers or custom functions when interacting with them; external interaction markers that enable data communication with the outside in the form of post messages when virtual space content is included in an external web service; special markers that place AI humans created using TTS (Text-To-Speech) and face reconstruction technology in a virtual space; and other markers that enable linking with external links or services (surveys, video chats, webinars, etc.).
[0256] Meanwhile, in the process of providing a streaming service for distributing immersive content according to the present invention, a user terminal can generate a dynamic URL and transmit an HTTP GET request to retrieve a specific tile image from the server. The dynamic URL generation prevents unauthorized access by generating a personalized URL based on the user's session information and request history. Furthermore, the URL automatically expires after a validity period, thereby providing secure data request / transmission.
[0257] If abnormal requests are detected more frequently than a predetermined number of times, the requests are temporarily restricted to prevent denial-of-service attacks (DDoS), etc., and all communications between the user's terminal browser and the server are transmitted safely through advanced encryption. The key used for encryption is changed at predetermined intervals to minimize the risk of key exposure.
[0258] In addition, the data request method and speed can be dynamically adjusted according to the network speed and stability of the user terminal, and if the data request fails, it can be automatically requested and processed without service interruption through a retry mechanism.
[0259] Meanwhile, in the process of providing a streaming service for distribution of immersive content according to the present invention, the user terminal receives the necessary tile image in real time from the server, and the received tile image is stored in a temporary buffer and can be managed for data transmission to the GPU.
[0260] By utilizing multi-threading technology to decode multiple image tiles simultaneously, it can quickly render them to the user terminal, and the decoded images can be converted into a form optimized for the GPU, so that only the minimum data conversion work required for rendering is performed.
[0261] Additionally, based on the user's current viewport and activity within the immersive content, the priority of required image tiles is dynamically adjusted, ensuring that important data is transmitted, received, and processed quickly. Furthermore, when the temporary buffer runs out of space, inactive or less important data is intelligently removed, enabling efficient buffer management.
[0262] Additionally, it can improve transfer speed by minimizing data copy operations between temporary buffers and the GPU, and provide seamless image tile rendering performance through synchronization between the GPU rendering pipeline and buffer management system.
[0263] Meanwhile, based on the user's movement and screen (or gaze) zoom patterns, the priority of which image tiles should be loaded first can be dynamically determined. This means that when the user moves or changes the screen (or gaze), the user terminal continuously communicates with the server to request and receive tile image information in real time, allowing the user to experience uninterrupted streaming.
[0264] Specifically, it is possible to learn users' screen (or gaze) zoom in / out and movement patterns in advance and predict the user's next action based on this, while dynamically adjusting the priority of image tiles to be loaded later.
[0265] In addition, by minimizing the connection setup time with the server, data transmission can be started quickly, and through continuous communication between the server and the user terminal, the necessary image tile information can be requested and received in real time when the user moves or the screen (or gaze) changes, and unnecessary network traffic and buffer usage can be minimized by recognizing and processing duplicate image tile requests.
[0266] Fig. 15 is a block diagram of a server (130). First, Fig. 15 is a block diagram of a server (130). According to an embodiment of the present invention, a server (130) may be implemented by including a memory (241) and a processor (or control unit) (240).
[0267] The above memory (1501) may be one or more databases (DBs), and may correspond to the various databases described above, and the processor (240) may perform various functions described in FIGS. 1 to 14, and specifically, may receive content data including at least one of 360-degree image data or 3D data for 3D content from a first user terminal (110), perform resource lightening on the received 360-degree image data or image data extracted from the 3D data, and when a request for transmission of immersive content is received from a second user terminal (120), perform a mashup that recombines the image data on which the resource lightening has been performed according to a predetermined standard to produce lightened immersive content, and then control to distribute the produced immersive content to the second user terminal (120).
[0268] Meanwhile, FIG. 16 is a block diagram of a server (130) for 3D data packaging and streaming service for distribution of immersive content according to an embodiment of the present invention, and FIG. 17 is a detailed block diagram of a processor (240) according to an embodiment of the present invention.
[0269] First, referring to FIG. 16, a server (130) for 3D data packaging and streaming service for distribution of immersive content can be implemented by including an account (210), an Identity and Access Management (IAM) (220), an editor (230), a processor (240), a player (250), and a Contents Delivery Network (CDN) (260). However, the present invention is not limited thereto, and may be implemented by including one or more other components, or vice versa.
[0270] Meanwhile, even if it is expressed as a single component, it may be in the form of a module containing multiple components, or vice versa. For example, the processor (240) of FIG. 17 may be expressed as a single module with a configuration similar to that of FIG. 15. In FIG. 16, the editor (230) and the processor (240) can be viewed as components related to processing immersive content before distribution, while the player (250) and CDN (260) can be viewed as components related to distribution and processing of immersive content.
[0271] The account (210) can be responsible for registering the user terminal (110), i.e., the user account. The IAM (220) is for registering and managing the user account and may be token-based. In addition, the IAM (220) can also manage the access level and permissions of the user account. In addition, the IAM (220) can also manage the user data of the user account.
[0272] In the present invention, the editor (230) verifies the user through the account (210) and IAM (220) when the user logs in, and can perform operations according to the user's authority. For example, if the user terminal (110) has authority, the user terminal (110) can access the editor of the server (130). The editor (230) is a type of web editor, and can provide a preview function of 3D content representing a virtual space, and can handle objects within immersive content including a virtual space based on coordinates, and can provide a function for distributing them. The editor (230) can receive and provide data related to the aforementioned function from a cloud platform.
[0273] If the aforementioned editor (230) is related to basic settings, etc., the processor (240) can be seen as related to processing actual data according to basic settings, etc. The processor (240) can download 3D content data from a plug-in via an API and receive and process metadata, 3D model files, and / or various 360 images via the editor (230). Meanwhile, depending on the embodiment, the editor (230) may mean including the function of the aforementioned processor (240).
[0274] Meanwhile, the data communication of plug-ins or independent programs, i.e. protocols, is as follows.
[0275] In relation to login, for example, a login request may be sent to the Authzip (IAM) server (220) via a RESTful API (Application Programming Interface) to obtain a token, and an authentication procedure may be performed in subsequent communication with the server (130) via the token. Meanwhile, in relation to data download / upload, for example, tour information managed by Backstage (310) may be read via a RESTful API, and the constructed data may be processed to create or update. Resources such as 3D models and 360-degree images may also be uploaded via the Backstage API.
[0276] Referring to FIG. 17, the processor (240) may include a backstage (310), a relational database (RDB) (320), a Lambda (330), and a simple storage service (S3) (340).
[0277] The backstage (310) can receive and actually process immersive content data including virtual spaces received through a plug-in or an independent program, and metadata, 3D model files, and 360 images received through the editor (230). The backstage (310) can perform tour schema management, spatial metadata management, 3D model and 360 image resource management, editing permission management, version management, concurrent access control, etc.
[0278] Backstage (310) can transmit a resource exceed check to a payment module for pay management, product plan management, recurring payment support, etc.
[0279] The relational database (320) can store metadata processed in the backstage (310). The relational database (320) can also store data related to transaction guarantee.
[0280] Lambda (330) can receive an image from backstage (310) and perform image processing such as cropping and compression. Lambda (330) can also receive a model from backstage (310) and perform model processing.
[0281] The Simple Storage Service (S3, 340) can store data processed in Backstage (310) and Lambda (330). The Simple Storage Service (S3, 340) can store metadata files distributed through Backstage (310) as well as 3D models and 360-degree images. At this time, the metadata files may be of JSON type, but are not limited thereto.
[0282] The player (250) can support playback of immersive content distributed at the request of a user terminal (120). The player (250) is a web-based service that supports playback of immersive content distributed through extraction, lightweighting, and mesh-up processes from 3D data according to the present invention, thereby enabling a terminal to visit or participate in immersive content including a virtual space, and can provide functions such as a function of loading and rendering resource images (such as video, images, and 3D markers) based on metadata of the distributed immersive content, a function of moving between scenes and changing themes, a function of providing additional information through various markers, and a responsive web service.
[0283] The player (250), as described above, can be provided to allow general users to experience and participate in immersive content without a separate login process. However, when provided in an embedded form on an external homepage, access can be restricted or bypassed based on the HTTP protocol's referrer header so that the immersive content can be accessed only through a designated URL. The referrer is a header value that records the domain before moving to another page on the web. When entering the player (250) from a specific website, a type of firewall can be implemented that compares the domain in the whitelist, allows access if it matches, and restricts access if it does not. Users load metadata and resources distributed on AWS S3 (Simple storage service) and AWS CloudFront (CDN) via HTTP.
[0284] The player (250) allows the user to utilize immersive content including virtual space, and can respond to various actions (requests) of the user terminal (120) regarding the immersive content being played. At this time, the various actions may include, for example, providing a view according to the user's selection, switching screens, moving between floors, etc., and further include various functions that can be provided through markers, etc. described in the present invention. Similarly, the player (250) may support query parameter options of the user terminal (120), and the query parameter options may include embed, start scene, etc. described above. Meanwhile, the player (250) may also provide a responsive web service.
[0285] Meanwhile, to ensure that anyone can easily and quickly experience immersive content in real time, establishing a standardized and streamlined infrastructure operating environment is essential. An ideal standard infrastructure environment begins with ensuring consistency and reducing complexity, achieved through consistent components, interfaces, and processes.
[0286] A streamlined infrastructure is easier to manage and operate, facilitating provisioning, expansion, troubleshooting, and disaster recovery. Using a well-known and simplified infrastructure allows for efficient infrastructure operation with a smaller workforce based on standardized operating procedures and processes. Therefore, efficient and stable infrastructure operation ultimately enables the provision of high-quality, immersive content services to users in real time. Therefore, the present invention utilizes a cloud-based infrastructure rather than an on-premise one, comprehensively considering the following: ensuring integrity and high availability under high traffic conditions, either temporarily or temporarily; preventing waste of computing resources; providing flexibility to support changing workloads; and ensuring security to safely store established virtual space assets.
[0287] In the past, expanding infrastructure required adding servers, which was inconvenient because it took a considerable amount of time for the purchased servers to physically arrive. However, in contrast, the cloud can be used immediately after payment, so if traffic surges, immediate infrastructure expansion is possible, allowing for a quick and easy response.
[0288] Similarly, by using the cloud's CDN (AWS CloudFront) instead of a web server, not only can it provide HTTPS endpoints, but it can also reduce latency and deliver content quickly, eliminating the need to secure idle servers to prepare for user surges, resulting in cost efficiency and the stability to deliver immersive content in real time from anywhere in the world.
[0289] Additionally, if the scale of the infrastructure required to build a service is not as anticipated, and in the past, corresponding costs were incurred when infrastructure was insufficient or surplus, using the cloud allows for real-time increases and decreases, but you only pay for what you use, which allows for cost savings and efficient cost management.
[0290] Additionally, immersive content can be safely stored with high security, and global cloud service providers have data centers on major continents around the world, allowing users to experience immersive content at high speeds anywhere.
[0291] CDN (260) supports caching with AWS Cloud Front according to the request of the player (250) and can forward the request of the player (250) to the simple storage service (S3, 340).
[0292] The server (130) (including an editor and processor) can provide a web editor service for highly realistic immersive content utilizing real-time 360-degree images and 3D models. Furthermore, the server (130) can provide management functions, tour main functions, import functions, editing functions, and distribution functions, as well as statistical functions, to provide highly realistic immersive content.
[0293] Specifically, statistical functions can analyze customer behavior data from users who access immersive content. These data can be stored and provided as custom events, including the number of users accessing the content, duration of stay, inflow paths, preferred pages, click events, users by country, and retention rates. While Google Analytics can be utilized, this is merely an example and can be provided in various other formats. Furthermore, in addition to statistical information, it may also be possible to collect unstructured user log data in real time and provide it in the form of a graphical dashboard.
[0294] The above management function can handle not only the creation and management of tours, but also the sharing of tours when multiple user accounts work as a team. The above management function can also provide guides.
[0295] The above tour main function can provide, for example, basic settings such as title / description / tag / logo / representative image for the tour, management of guidance text such as start message, as well as CMS project linkage functions (e.g., chat, customer inquiry, statistics).
[0296] The import function can provide functions such as creating / editing / deleting images required for the tour, filtering and search functions, and creating / editing / deleting floors and themes. Meanwhile, the import function can also support multiple updates.
[0297] The editing function may provide a keymap / preview function, 3D modeling function support, movement and various action (hotspot) editing function, undo / redo function, auto-save function, simultaneous work limitation function, and other editing-related functions. For example, Fig. 18 (a) illustrates a TOP view related to the editing function, and Fig. 18 (b) illustrates an ISO view. In Figs. 18 (a) and 18 (b), circular items in the images represent hotspots related to camera points, and the number of such hotspots can be arbitrarily set. In addition, when one of multiple hotspots in the image is selected, the camera point set to the selected hotspot, for example, the camera angle, may also be displayed.
[0298] In addition, with regard to the editing function, if multiple hotspots among the aforementioned hotspots are sequentially selected at a given time, the order of the hotspots is determined according to the selection order, and the right screen (a kind of preview function) can be set to be automatically played as in (a) and (b) of FIG. 18 arbitrarily according to the above order. Alternatively, if one hotspot among multiple hotspots is selected, the next recommended hotspot can be displayed to help understand the preset or selected hotspot, thereby serving as a tour guide.
[0299] Although not illustrated, each hotspot in the left images of Fig. 18 (a) and Fig. 18 (b) is provided with a numbering, and if the first numbering and the last numbering or the second numbering (each, any numbering) from the first numbering are selected within a preset time, they may be automatically played according to the numbering order. The numbering may be set by the user terminal (110) or the server (130), or may be automatically set based on the number of views frequently viewed by the user terminal (120).
[0300] In the above, selection can also be made by dragging.
[0301] In addition, unlike the above, when one hotspot is selected in the left images of Figs. 18(a) and 18(b) or the right images of Figs. 18(a) and 18(b) are provided according to the default hotspot, it is possible to provide movement to the next hotspot according to the change in the camera angle even if another hotspot on the left is not selected. For example, in this case, the left images of Figs. 18(a) and 18(b) may provide information on selectable hotspots according to the change in the camera angle by highlighting them in a similar manner to the selected hotspot, and the right image may provide a guide on whether to select movement to another hotspot.
[0302] Regarding distribution features, a final review screen (preview) and distribution functions, along with scene-specific access URLs, can be provided. These distribution functions can include methods such as hosting and embedding.
[0303] Meanwhile, the editor's data communication, or protocol, is as follows.
[0304] In terms of login, you can process integrated login by redirecting to the account domain and receive a token through a callback to process login.
[0305] In relation to tour creation / change / deletion, for example, loading / saving can be performed step by step in the backstage (310) via RESTful API. The case of uploading resource files (image, model, etc.) from the user terminal (110) is also as described above.
[0306] With regard to limiting concurrent connections, for example, you can manage sessions currently connected to the backstage (310) using Websocket.
[0307] The above player (250) is a web service that can play a tour created with a web builder, and can provide functions such as loading and rendering resources (images, videos, 3D models, etc.) based on distributed tour metadata, moving between scenes, changing themes, providing various other functions in the form of hotspot buttons or attached to content, providing additional functions in the form of query parameters, and supporting responsive web. In the above, the various functions may include an image slider, a video viewer, a 3D model viewer, a custom HTML viewer, etc.
[0308] Meanwhile, data communication related to the player (250), i.e., the protocol, is as follows. Unlike the user terminal (110), which is a 3D content maker, it can have a structure that can be viewed by general users without a login process. Regarding data loading, metadata and resources distributed on the simple storage service and cloud front can be loaded via HTTP. Regarding chat, data can be sent and received to the chat server using, for example, a web socket. Regarding the inquiry form, data can be transmitted to the Zapier service using, for example, a RESTful API. Regarding statistics, as described above, for example, customer behavior data can be stored in the form of a custom event on an analysis server (e.g., Google Analytics). The customer behavior data can include scene visit / stay time (the default time window can be counted as a preset arbitrary time (e.g., 15 seconds)) and can also include hotspot actions, etc.
[0309] Meanwhile, the above-described method can be written as a program that can be executed on a computer, and can be implemented on a general-purpose digital computer that executes the program using a computer-readable medium. In addition, the structure of the data used in the above-described method can be recorded on a computer-readable medium through various means. Computer-readable media that store executable computer code for performing various methods of the present invention include storage media such as magnetic storage media (e.g., ROM, floppy disks, hard disks, etc.) and optical reading media (e.g., CDs, DVDs, etc.).
[0310] Those skilled in the art will appreciate that the embodiments of the present invention can be implemented in modified forms without departing from the essential characteristics of the above description. Therefore, the disclosed methods should be considered illustrative rather than restrictive. The scope of the present invention is determined by the claims, not the detailed description, and all differences within the scope equivalent thereto should be construed as being included within the scope of the present invention.
[0311] The method and device for providing a service for producing and distributing lightweight immersive content according to the present invention can be used in methods and devices for providing a service for producing and distributing various immersive contents.
Claims
1. A method for providing a service for producing and distributing lightweight immersive content by a server, A step in which the server receives content data including at least one of 360-degree image data or 3D data for 3D content from a first user terminal; A step of performing resource lightening on the received 360-degree image data or image data extracted from the 3D data; and When the server receives a request for transmission of immersive content from a second user terminal, a step of performing a mashup that recombines the image data on which the resource has been lightened according to a predetermined standard to produce light-weighted immersive content, and then distributing the produced immersive content to the second user terminal is included. Method for providing production and distribution services for lightweight immersive content.
2. In paragraph 1, The 360-degree image data in the content data received from the first user terminal includes at least one image of a 360-degree × 360-degree full-field image of a real space taken using a camera or at least one image of a 360-degree × 360-degree full-field image of a virtual space created using computer graphics. Method for providing production and distribution services for lightweight immersive content.
3. In paragraph 1, The 3D data for the 3D content in the content data received from the first user terminal includes 3D data for a virtual space created through a 3D tool including a game engine or a 3D modeling program. Method for providing production and distribution services for lightweight immersive content.
4. In paragraph 3, Image data extracted from the above 3D data is, Containing at least one image for a 360 degree X 360 degree omnidirectional direction of the virtual space extracted from 3D data for the virtual space, Method for providing production and distribution services for lightweight immersive content.
5. In paragraph 1, The steps for performing the above resource lightweighting are: Including converting and storing the image data extracted from the 360-degree image data or the 3D data into a low-resolution image of a first predetermined capacity consisting of one side, a medium-resolution image of a second predetermined capacity divided into 6 sides, and a high-resolution image of a third predetermined capacity divided into 24 sides, respectively. Method for providing production and distribution services for lightweight immersive content.
6. In paragraph 1, The steps for performing the above resource lightweighting are: In the case where the content data received from the first user terminal is 3D data for the 3D content, it includes performing resource lightening for object data of 3D models and objects in the 3D data, The resource lightening for the above object data includes performing at least one of conversion to a predetermined file format, vertex grouping to reduce the number of vertices for 3D models and objects, vertex merging to merge vertices for 3D models and objects, and resolution step subdivision to save 3D models and objects at a resolution divided into multiple steps of a predetermined size. Method for providing production and distribution services for lightweight immersive content.
7. In paragraph 1, The above lightweight immersive content is, It is produced by performing a mashup that recombines the image data on which the resource has been lightened based on the location and gaze information of the user of the second user terminal in the immersive content. Method for providing production and distribution services for lightweight immersive content.
8. In paragraph 6, If the content data received from the first user terminal is 3D data for the 3D content, the lightweight immersive content is produced by performing a mashup that recombines the image data for which the resource lightweighting has been performed and the object data for which the resource lightweighting has been performed. Method for providing production and distribution services for lightweight immersive content.
9. In paragraph 1, The above distribution is done by streaming method. Method for providing production and distribution services for lightweight immersive content.
10. For a device that provides production and distribution services for lightweight immersive content, memory section; and It is composed of a control unit; The control unit receives content data including at least one of 360-degree image data or 3D data for 3D content from a first user terminal, Perform resource lightening on the received 360-degree image data or image data extracted from the 3D data, When a request for transmission of immersive content is received from a second user terminal, a mashup is performed to reassemble the image data on which the resource has been lightened according to a predetermined standard to produce light-weighted immersive content, and then the produced immersive content is controlled to be distributed to the second user terminal. A device that provides production and distribution services for lightweight immersive content.
Citation Information
Patent Citations
Resin composition
KR1020230149742A
Separator and electrochemical conversion cell comprising same
KR1020230150210A
Mobility apparatus
KR1020250150840A
Main and immersive video coordination system and method
US20150289032A1
Rendering Content in a 3D Environment
US20200279429A1