Image provision method, and server for performing same
The described method and server system address the passive nature of existing content delivery in virtual environments by generating personalized images based on user viewpoints, thereby enhancing user engagement and immersion.
Patent Information
- Application Number
- PCT/KR2023/020725
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2023-12-15
- Publication Date
- 2025-06-12
AI Technical Summary
Existing technologies for providing content in virtual environments do not allow viewers to actively engage with the content, as they are limited to passively watching pre-captured video feeds.
A method and server system that generate a point cloud from camera information, identify and classify objects as dynamic or static, and create a virtual background. The system then determines the user's target viewpoint and generates a personalized image by combining the virtual background with the dynamic object's relevant areas, allowing for active viewing experiences.
Enables users to actively view content by providing personalized images based on their viewpoint, enhancing engagement and immersion in virtual environments.
Smart Images

Figure KR2023020725_12062025_PF_FP_ABST
Abstract
Description
Image providing method and server performing the method
[0001] The embodiments below relate to a technology for providing a content image to a user, and more specifically, to a technology for providing a content image corresponding to a user's viewpoint in a virtual environment.
[0002] With the advancement of two-way communication technology and internet media, services that allow viewers to watch performances in real time through video devices and performers to monitor audience reactions in real time are becoming widespread. Regardless of their individual viewing time, viewers all view the same video captured by the camera. In other words, viewers passively watch the performance video provided by the performance organizer.
[0003] One embodiment may provide a method for providing content that allows viewers to actively view it.
[0004] One embodiment may provide a server that provides content that viewers can actively view.
[0005] However, technical challenges are not limited to the technical challenges described above, and other technical challenges may exist.
[0006] According to one embodiment, a method for providing an image may include: generating a point cloud for a scene based on information received from one or more cameras; identifying one or more objects included in the scene based on the point cloud; determining each of the one or more objects as a dynamic object or a static object; generating a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object; acquiring a first target viewpoint of a first user through a first user terminal; determining an area corresponding to the first target viewpoint among a plurality of areas of the point cloud constituting the dynamic object as a first main area and determining an area not corresponding to the first target viewpoint as a first auxiliary area; generating a first target image based on the first main area, the first auxiliary area, and the virtual background; and transmitting the first target image to the first user terminal.
[0007] The above first target image can be output by the first user terminal.
[0008] The image providing method may further include an operation of acquiring a second target viewpoint of the first user through the first user terminal, an operation of determining an area corresponding to the second target viewpoint among the plurality of areas of the point cloud constituting the dynamic object as a second main area, and an operation of determining an area not corresponding to the second target viewpoint as a second auxiliary area, an operation of generating a second target image based on the second main area, the second auxiliary area, and the virtual background, and an operation of transmitting the second target image to the first user terminal.
[0009] Each point in the above point cloud may include color information.
[0010] The action of determining each of the one or more objects as a dynamic object or a static object may include an action of determining the object as a dynamic object if the type of the identified object is a person or an animal.
[0011] The operation of generating a virtual background by rendering the background of the scene based on at least a portion of the point cloud corresponding to the static object may include the operation of determining background points corresponding to background objects determined as static objects among the one or more objects, the operation of generating a three-dimensional background structure by performing three-dimensional modeling on the background points, and the operation of generating the virtual background by performing texturing on the three-dimensional background structure.
[0012] The operation of determining an area corresponding to the first target viewpoint among the plurality of areas of the point cloud constituting the dynamic object as a first main area and determining an area not corresponding to the first target viewpoint as a first auxiliary area may include an operation of determining an area directly corresponding to the first target viewpoint as the first target area, and an operation of determining the first target area and areas adjacent to the first target area as the first main area.
[0013] The above image providing method may further include an operation of acquiring a third target viewpoint of a second user through a second user terminal.
[0014] The operation of determining an area corresponding to the first target viewpoint among the plurality of areas of the point cloud constituting the dynamic object as a first main area and determining an area not corresponding to the first target viewpoint as a first auxiliary area may include an operation of determining an area corresponding to the first target viewpoint or the third target viewpoint among the plurality of areas of the point cloud constituting the dynamic object as a first main area and determining an area not corresponding to the first target viewpoint and the third target viewpoint as the first auxiliary area.
[0015] The operation of generating a first target image based on the first main region, the first auxiliary region, and the virtual background may include an operation of adjusting the number of points in the first auxiliary region, an operation of generating a three-dimensional foreground structure by performing three-dimensional modeling on the dynamic object based on points in the first main region and points in the first auxiliary region, an operation of generating a virtual foreground by performing texturing on the three-dimensional foreground structure, and an operation of generating the first target image based on the virtual foreground and the virtual background.
[0016] According to one embodiment, a server includes a communication module, a processor, and a memory, and the memory may store instructions that, when executed by the processor, cause the server to generate a point cloud for a scene based on information received from one or more cameras, identify one or more objects included in the scene based on the point cloud, determine each of the one or more objects as a dynamic object or a static object, generate a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object, acquire a first target viewpoint of a first user through a first user terminal, determine an area corresponding to the first target viewpoint among a plurality of areas of the point cloud constituting the dynamic object as a first main area, determine an area not corresponding to the first target viewpoint as a first auxiliary area, generate a first target image based on the first main area, the first auxiliary area, and the virtual background, and transmit the first target image to the first user terminal.
[0017] According to one embodiment, a method for providing an image, performed by a system server including a main server and an auxiliary server, includes an operation in which the main server identifies one or more objects included in the scene based on the point cloud, an operation in which the main server determines each of the one or more objects as a dynamic object or a static object, an operation in which the main server generates a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object, an operation in which the main server transmits points corresponding to the dynamic object among points of the point cloud and the virtual background to the auxiliary server, an operation in which the auxiliary server acquires a first target view point of a first user through a first user terminal, an operation in which the auxiliary server determines an area corresponding to the first target view point among a plurality of areas of the point cloud constituting the dynamic object as a first main area and determines an area not corresponding to the first target view point as a first auxiliary area, an operation in which the auxiliary server generates a first target image based on the first main area, the first auxiliary area, and the virtual background, and an operation in which the auxiliary server transmits the first target image to the first user terminal, wherein the first The target image can be output by the first user terminal.
[0018] The above auxiliary server may be a ME`C (mobile edge computing) server.
[0019] Figure 1 is a configuration diagram of a content provision system according to one embodiment.
[0020] Figure 2 is a configuration diagram of a server according to one embodiment.
[0021] Figure 3 is a flowchart of an image providing method according to one embodiment.
[0022] Figure 4 is a flowchart of a method for generating a virtual background according to an example.
[0023] Figure 5 is a flowchart of a method for determining a first main area according to an example.
[0024] FIG. 6 is a flowchart of a method for generating a first target image based on a virtual foreground and a virtual background, according to an example.
[0025] FIG. 7 illustrates, according to an example, a first main region corresponding to a first target viewpoint among multiple regions of a point cloud and a first auxiliary region not corresponding to the first target viewpoint.
[0026] FIG. 8 is a flowchart of a method for transmitting a second target image generated based on a second target viewpoint of a first user to a first user terminal according to an example.
[0027] FIG. 9 is a flowchart of a method for determining a plurality of regions of a point cloud, a first main region and a first auxiliary region, based on a third target viewpoint of a second user, according to an example.
[0028] Figure 10 is a configuration diagram of a system server including a main server and an auxiliary server according to an example.
[0029] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Therefore, the actual implementation is not limited to the specific embodiments disclosed, and the scope of this specification includes modifications, equivalents, or alternatives within the technical concepts described in the embodiments.
[0030] Although terms such as "first" or "second" may be used to describe various components, these terms should be interpreted solely to distinguish one component from another. For example, a first component may be referred to as a second component, and similarly, a second component may also be referred to as a first component.
[0031] When it is said that a component is "connected" to another component, it should be understood that it may be directly connected or connected to that other component, but there may also be other components in between.
[0032] Singular expressions include plural expressions unless the context clearly dictates otherwise. In this specification, the terms "comprises" or "has" should be understood to indicate the presence of a described feature, number, step, operation, component, part, or combination thereof, but not to exclude the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof.
[0033] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art. Terms defined in commonly used dictionaries should be interpreted to have a meaning consistent with their meaning in the context of the relevant technology, and will not be interpreted in an idealized or overly formal sense unless explicitly defined herein.
[0034] Hereinafter, embodiments will be described in detail with reference to the attached drawings. In the description with reference to the attached drawings, identical components are assigned the same reference numerals regardless of the drawing numbers, and redundant descriptions thereof will be omitted.
[0035]
[0036] Figure 1 is a configuration diagram of a content provision system according to one embodiment.
[0037] According to one embodiment, the content providing system may include one or more cameras (112, 114), a server (120), and a user terminal (130). For example, each of the one or more cameras (112, 114) may be a color camera (e.g., an RGB camera), a depth camera (e.g., a D (depth) camera), or an RGB-D camera. A scene (10) may be captured through the one or more cameras (112, 114). For example, the scene (10) may be a performance scene. The scene (10) may include a background and a foreground.
[0038] According to one embodiment, the server (120) can acquire points for a scene (10) using one or more cameras (112, 114). The server (120) can generate a point cloud for the scene (10) based on the acquired points. For example, the server (120) can generate a point cloud for the scene (10) by aligning points acquired through one or more cameras (112, 114).
[0039] The server (120) can generate a virtual scene for the scene (10) based on the point cloud. The server (120) can provide content to the user terminal (130) based on the virtual scene. For example, the server (120) can obtain a time point at which the user views the virtual scene from the user terminal (130), generate an image corresponding to the obtained time point, and transmit the generated image to the user terminal (130) to provide content to the user. Since the image corresponding to the user's viewing time point is provided to the user by the server (120) and the user terminal (130), the user can actively view the content.
[0040] According to one embodiment, the user terminal (130) may include a head-mounted display (HMD). For example, the user terminal (130) may determine the point of view of the user's eyes through a sensor of the HMD or a sensor positioned adjacent to the HMD. The user terminal (130) may receive an image of content from the server (120) and output the received image through the HMD.
[0041] A method of providing content to a user is described in detail with reference to FIGS. 2 to 10 below.
[0042]
[0043] Figure 2 is a configuration diagram of a server according to one embodiment.
[0044] According to one embodiment, a server (200) includes a communication unit (210), a processor (220), and a memory (230).
[0045] The communication unit (210) is connected to a processor (220), a memory (230), and a camera (e.g., the cameras (112, 114) of FIG. 1) to transmit and receive data. The communication unit (210) may be connected to other external devices to transmit and receive data. Hereinafter, the expression "transmitting and receiving "A" may refer to transmitting and receiving "information or data representing A."
[0046] The communication unit (210) may be implemented as a circuitry within the server (200). For example, the communication unit (210) may include an internal bus and an external bus. As another example, the communication unit (210) may be an element that connects the server (200) to an external device. The communication unit (210) may be an interface. The communication unit (210) may receive data from an external device and transmit the data to the processor (220) and the memory (230).
[0047] The processor (220) processes data received by the communication unit (210) and data stored in the memory (230). A "processor" may be a data processing device implemented as hardware having a circuit with a physical structure for executing desired operations. For example, the desired operations may include code or instructions included in a program. For example, a data processing device implemented as hardware may include a microprocessor, a central processing unit, a processor core, a multi-core processor, a multiprocessor, an ASIC (Application-Specific Integrated Circuit), and an FPGA (Field Programmable Gate Array).
[0048] The processor (220) executes computer-readable code (e.g., software) stored in memory (e.g., memory (230)) and instructions generated by the processor (220).
[0049] The memory (230) stores data received by the communication unit (210) and data processed by the processor (220). For example, the memory (230) may store a program (or application, software). The stored program may be a set of syntaxes that are coded to provide content and can be executed by the processor (220).
[0050] According to one aspect, the memory (230) may include one or more volatile memory, non-volatile memory, and random access memory (RAM), flash memory, a hard disk drive, and an optical disk drive.
[0051] The memory (230) stores a set of instructions (e.g., software) that operate the server (200). The set of instructions that operate the server (200) is executed by the processor (220).
[0052] The camera can generate information used to create a point cloud of the scene by capturing the scene. The camera can be an RGB-D camera, but is not limited to the described embodiments.
[0053] The communication unit (210), processor (220), and memory (230) are described in detail below with reference to FIGS. 3 to 10.
[0054]
[0055] Figure 3 is a flowchart of an image providing method according to one embodiment.
[0056] The following operations 310 to 380 can be performed by the server (200) described above with reference to FIG. 2.
[0057] In operation 310, the processor (220) of the server (200) may generate a point cloud for a scene (e.g., scene (10) of FIG. 1) based on information received from one or more cameras (e.g., one or more cameras (112, 114) of FIG. 1). For example, the server (200) may generate points for the scene based on information received from one or more cameras, and may generate a point cloud for the scene by aligning the generated points.
[0058] According to one embodiment, each point of the point cloud may include color information, texture information, and / or volume information.
[0059] In operation 320, the processor (220) of the server (200) may identify one or more objects included in the scene based on the point cloud. For example, the identified objects may include a plurality of regions. Each of the plurality of regions may be a three-dimensional tile including at least a portion of points of the point cloud. Each of the plurality of regions may be a three-dimensional region.
[0060] In operation 330, the processor (220) of the server (200) may determine each of one or more objects as a dynamic object or a static object. For example, an object may be determined as a dynamic object or a static object based on the type of the object. A dynamic object may be an object that moves within a scene. For example, if the type of the identified object is a person or an animal, the object may be determined as a dynamic object. A static object may be an object that does not move within the scene.
[0061] In operation 340, the processor (220) of the server (200) can generate a virtual background by rendering the background of the scene based on at least a portion of the point cloud corresponding to the static object.
[0062] In one embodiment, since objects corresponding to a virtual background in a scene captured by one or more cameras do not change over time, a virtual background may be generated once to reduce the amount of data to be processed later. For example, the identified one or more objects may include a first object, a second object, and a third object, and the first and second objects may be determined as static objects, and the third object may be determined as a dynamic object. The virtual background may be generated by rendering the background of the scene using first points corresponding to the first object, second points corresponding to the second object, and third points corresponding to the third object. Since the third object is a dynamic object, the shape or pose of the third object may change over time. Since the user is observing the changing third object, data processing for the third object may need to be performed in real time.
[0063] In operation 350, the communication unit (210) of the server (200) can obtain a first target viewpoint of the first user through the first user terminal (e.g., the user terminal (130) of FIG. 1). For example, the first user terminal can determine a part or area that the user is gazing at within a virtual scene using a sensor. For example, when the first user terminal includes an HMD, the first user terminal can obtain the posture of the user's head using an IMU (inertial measurement unit) of the HMD, and can obtain the first target viewpoint based on the posture of the user's head. For example, the first user terminal can capture the user's pupils using a camera of the first user terminal (or HMD), and determine which area of the display the user is gazing at based on the position of the user's pupils. The first target viewpoint can be obtained based on the area of the display that is being gazed at. The first target viewpoint can include a first left target viewpoint corresponding to the first user's left eye and a first right target viewpoint corresponding to the first user's right eye.
[0064] In one embodiment, the first user's position within a virtual scene may be set by the first user. For example, if the virtual scene is a performance, the first user may set a desired location among virtual areas designated as audience seats within the virtual scene as the first user's position. For example, the first target viewpoint may be defined as the viewpoint from which the first user looks at a specific area of the virtual scene from the first user's position.
[0065] According to one embodiment, the communication unit (210) of the server (200) may receive, from the first user terminal, at least one of information about the posture of the first user's head and information about the area of the display that the user is gazing at. The processor (220) of the server (200) may determine the first target viewpoint of the first user based on at least one of information about the posture of the first user's head and information about the area of the display that the first user is gazing at.
[0066] In operation 360, the processor (220) of the server (200) may determine an area corresponding to a first target viewpoint among multiple areas of a point cloud constituting a dynamic object as a first main area, and may determine an area not corresponding to the first target viewpoint as a first auxiliary area. For example, the main area may be a region of interest (ROI).
[0067] The method of determining the first major area and the second major area is described in detail below with reference to Fig. 5.
[0068] In operation 370, the processor (220) of the server (200) can generate a first target image based on the first main area, the first auxiliary area, and the virtual background.
[0069] According to one embodiment, the processor (220) of the server (200) can adjust the number of points in the first auxiliary area. For example, if the number of points to be processed is reduced, the amount of data transmitted and received between devices and the amount of data processed can be reduced.
[0070] According to one embodiment, the processor (220) of the server (200) can generate a virtual foreground based on the first main region and the first auxiliary region. The virtual foreground may be a model or image corresponding to a dynamic object gazed by the user. For example, the virtual foreground may be a three-dimensional model. For example, the virtual foreground may be a two-dimensional image generated by projecting the three-dimensional model from the first target viewpoint.
[0071] According to one embodiment, the processor (220) of the server (200) may generate a first target image based on a virtual foreground and a virtual background. For example, the first target image may be a two-dimensional image generated by projecting the virtual foreground and the virtual background from a first target viewpoint.
[0072] In one embodiment, when the first target viewpoint includes a first left target viewpoint and a first right target viewpoint, a first left target image corresponding to the first left target viewpoint and a first right target image corresponding to the first right target viewpoint may be generated. The first left target image and the first right target image may be stereoscopic images.
[0073] Below, a method for generating a first target image is described in detail with reference to FIGS. 6 and 7.
[0074] In operation 380, the communication unit (220) of the server (200) may transmit a first target image to a first user terminal. The first user terminal, which receives the first target image from the server (200), may output the first target image through a display. For example, the display may be an HMD. The user may actively view content provided by a content provider through the first target image, rather than unilaterally viewing the content.
[0075] In one embodiment, if the first target images are stereoscopic images, the user can view a three-dimensional virtual scene through the first target images.
[0076] According to one embodiment, operations 350 to 380 with reference to FIG. 3 may be performed by a streaming server of the server (200).
[0077] According to one embodiment, although operations 350 to 380 are described as being performed by the server (200) with reference to FIG. 3, after operation 340 is performed according to the embodiment, the server (200) may transmit points (e.g., foreground points) of the point cloud corresponding to the generated virtual background and dynamic object to the first user terminal. The first user terminal may obtain the first target viewpoint of the first user. The first user terminal may determine a first main region among a plurality of regions of the point cloud constituting the dynamic object and may determine the first main region as the first auxiliary region. The first user terminal may generate a first target image based on the first main region, the first auxiliary region, and the virtual background. The first user terminal may output the generated first target image.
[0078]
[0079] Figure 4 is a flowchart of a method for generating a virtual background according to an example.
[0080] According to one embodiment, operations 410 to 430 below may be associated with operation 340 described above with reference to FIG. 3. For example, operation 340 may include operations 410 to 430. Operations 410 to 430 may be performed by the server (200) described above with reference to FIG. 2.
[0081] In operation 410, the processor (220) of the server (200) may determine background points corresponding to background objects determined as static objects among one or more objects.
[0082] In operation 420, the processor (220) of the server (200) may generate a three-dimensional background structure by performing three-dimensional modeling on background points. For example, the three-dimensional background structure may be composed of polygon meshes generated based on background points.
[0083] In operation 430, the processor (220) of the server (200) may generate a virtual background by performing texturing on a three-dimensional background structure. For example, the processor (220) may perform texturing on the three-dimensional background structure based on texture information contained in each of the background points.
[0084]
[0085] Figure 5 is a flowchart of a method for determining a first main area according to an example.
[0086] According to one embodiment, operations 510 and 520 below may be associated with operation 360 described above with reference to FIG. 3. For example, operation 360 may include operations 510 and 520. Operations 510 and 520 may be performed by the server (200) described above with reference to FIG. 2.
[0087] In operation 510, the processor (220) of the server (200) may determine an area directly corresponding to the first target time point as the first target area.
[0088] In operation 520, the processor (220) of the server (200) may determine a first target area and areas adjacent to the first target area as first main areas. The areas adjacent to the first target area may be areas indirectly corresponding to the first target viewpoint. For example, areas located within a preset distance from the first target area may be determined as the first main area. The first main area may correspond to the first user's field of view or ROI based on the first target viewpoint.
[0089] According to one embodiment, among the plurality of regions of the point cloud constituting the dynamic object, the remaining regions that are not determined as the first main region may be determined as the first auxiliary region.
[0090]
[0091] FIG. 6 is a flowchart of a method for generating a first target image based on a virtual foreground and a virtual background, according to an example.
[0092] According to one embodiment, operations 610 to 640 below may be associated with operation 370 described above with reference to FIG. 3. For example, operation 370 may include operations 610 to 640. Operations 610 to 640 may be performed by the server (200) described above with reference to FIG. 2.
[0093] In operation 610, the processor (220) of the server (200) may adjust the number of points within the first auxiliary area. For example, the processor (220) may reduce the number of points within the first auxiliary area.
[0094] According to one embodiment, the processor (220) may adjust the number of points within the first auxiliary area based on the distance between the boundary of the first main area and the first auxiliary area. For example, as the distance between the boundary of the first main area and the first auxiliary area increases, the number of points within the first auxiliary area may be significantly reduced.
[0095] In one embodiment, the number of points within the first auxiliary area may be adjusted so that no points are included in the first main area.
[0096] In operation 620, the processor (220) of the server (200) may generate a 3D foreground structure by performing 3D modeling on a dynamic object based on points of a first main region and points of a first auxiliary region. Each of the points of the first main region and the points of the first auxiliary region may be referred to as a foreground point. For example, the 3D foreground structure may be composed of polygon meshes generated based on the foreground points. As the number of points of the first auxiliary region decreases, the amount of data processing required to generate the 3D foreground structure may decrease.
[0097] In operation 630, the processor (220) of the server (200) may generate a virtual foreground by texturing a three-dimensional foreground structure. For example, the processor (220) may perform texturing on the three-dimensional foreground structure based on texture information contained in each foreground point.
[0098] In operation 640, the processor (220) of the server (200) may generate a first target image based on the virtual foreground and virtual background. For example, the processor (220) may generate the first target image as a two-dimensional image by rendering the virtual foreground and virtual background. The first target image may be a two-dimensional image generated by projecting the virtual foreground and virtual background from a first target viewpoint.
[0099]
[0100] FIG. 7 illustrates, according to an example, a first main region corresponding to a first target viewpoint among multiple regions of a point cloud and a first auxiliary region not corresponding to the first target viewpoint.
[0101] According to one embodiment, a target object, which is a dynamic object, may be composed of a plurality of regions (701, 702, 703, 704, 705, 706, 711, 712, 713, 714, 715, 716) that are part of a point cloud (700). The server (200) may determine the regions (701, 702, 703, 704) that the user is looking at as a first main region (730), and determine the regions (705, 706, 711, 712, 713, 714, 715, 716) that the user is not looking at as a first auxiliary region.
[0102] According to one embodiment, since the areas (705, 706, 711, 712, 713, 714, 715, 716) determined as the first auxiliary areas do not correspond to the user's field of view, the points of the areas (705, 706, 711, 712, 713, 714, 715, 716) may be excluded in the process of generating a three-dimensional foreground structure. The server (200) may generate a three-dimensional foreground structure (740) by performing three-dimensional modeling on the target object based on the points of the first main area (730), excluding the points of the first auxiliary area. The processor (220) of the server (200) may generate a virtual foreground by performing texturing on the three-dimensional foreground structure (740). The generated virtual foreground may be placed within a virtual background.
[0103]
[0104] FIG. 8 is a flowchart of a method for transmitting a second target image generated based on a second target viewpoint of a first user to a first user terminal according to an example.
[0105] According to one embodiment, operations 810 to 840 below may be performed after operation 370 described above with reference to FIG. 3 is performed. Operations 810 to 840 may be performed by the server (200) described above with reference to FIG. 2.
[0106] In operation 810, the communication unit (210) of the server (200) may acquire a second target time point of the first user through the first user terminal (e.g., the user terminal (130) of FIG. 1). The second target time point may be a different time point from the first target time point. The description of operation 810 may be replaced with the description of operation 350 described above with reference to FIG. 3.
[0107] In operation 820, the processor (220) of the server (200) may determine an area corresponding to a second target viewpoint among multiple areas of a point cloud constituting a dynamic object as a second main area, and may determine an area not corresponding to the second target viewpoint as a second auxiliary area. The description of operation 820 may be replaced with the description of operation 820 described above with reference to FIG. 3.
[0108] In operation 830, the processor (220) of the server (200) may generate a second target image based on the second main area, the second auxiliary area, and the virtual background. The description of operation 830 may be replaced with the description of operation 370 described above with reference to FIG. 3.
[0109] In operation 840, the communication unit (220) of the server (200) may transmit the second target image to the first user terminal. The description of operation 840 may be replaced with the description of operation 380 described above with reference to FIG. 3.
[0110]
[0111] FIG. 9 is a flowchart of a method for determining a plurality of regions of a point cloud, a first main region and a first auxiliary region, based on a third target viewpoint of a second user, according to an example.
[0112] According to one embodiment, operations 910 and 920 below may be performed after operation 350 described above with reference to FIG. 3 is performed. Operations 910 and 920 may be performed in parallel with and independently of operations 350 and 360. Operations 910 and 920 may be performed by the server (200) described above with reference to FIG. 2.
[0113] In operation 910, the communication unit (210) of the server (200) can acquire the third target viewpoint of the second user through the second user terminal. For example, the first user terminal and the second user terminal can simultaneously access the server (200), and the first user of the first user terminal and the second user of the second user terminal can simultaneously view content provided by the server (200). The description of operation 910 can be replaced with the description of operation 350 described above with reference to FIG. 3.
[0114] In operation 920, the processor (220) of the server (200) may determine an area corresponding to a first target viewpoint or a third target viewpoint among multiple areas of a point cloud constituting a dynamic object as a first main area, and may determine an area not corresponding to the first target viewpoint or the third target viewpoint as a first auxiliary area. For example, even if an area does not correspond to the first user's field of view, if it corresponds to the third user's field of view, it may be determined as the first main area.
[0115] After operation 920 is performed, operation 370 described above with reference to FIG. 3 may be performed.
[0116]
[0117] Figure 10 is a configuration diagram of a system server including a main server and an auxiliary server according to an example.
[0118] According to one embodiment, the system server (1000) may include a main server (1010) and an auxiliary server (1020). For example, the auxiliary server (1020) may be a mobile edge computing (MEC) server.
[0119] The main server (1010) can generate a point cloud for the scene (10) based on information received from one or more cameras (112, 114).
[0120] The main server (1010) can identify one or more objects included in the scene (10) based on the point cloud.
[0121] The main server (1010) can determine each of one or more objects as a dynamic object or a static object.
[0122] The main server (1010) can generate a virtual background by rendering the background of the scene (10) based on at least a portion of a point cloud corresponding to a static object.
[0123] The main server (1010) can transmit points corresponding to dynamic objects and a virtual background among the points of the point cloud to the auxiliary server (1020).
[0124] The auxiliary server (1020) can obtain the first target time point of the first user through the first user terminal (130).
[0125] The auxiliary server (1020) may determine an area corresponding to a first target point time among multiple areas of a point cloud constituting a dynamic object as a first main area, and may determine an area not corresponding to the first target point time as a first auxiliary area.
[0126] The auxiliary server (1020) can generate a first target image based on the first main area, the first auxiliary area, and the virtual background.
[0127] The auxiliary server (1020) can transmit the first target image to the first user terminal (130). The first target image can be output by the first user terminal (130).
[0128]
[0129] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing device may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0130] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.
[0131] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0132] The hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0133] Although the embodiments described above have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0134] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims described below.
[0135]
[0136] [National Research and Development Project Supporting This Invention]
[0137] [Project ID] 1425172083
[0138] [Assignment Number] S3276711
[0139] [Ministry Name] Ministry of SMEs and Startups
[0140] [Name of Project Management (Specialist) Agency] Small and Medium Business Technology Information Promotion Agency
[0141] [Research Project Name] Small and Medium Enterprise Technology Innovation Development Project 'Export-Oriented Project'
[0142] [Research Project Title] Development of a Platform for Creating and Sharing Ultra-Realistic Content Based on AI Motion Capture
[0143] [Contribution rate] 50 / 100
[0144] [Name of the project performing organization] 3RII Co., Ltd.
[0145] [Research Period] January 1, 2023 - December 31, 2023
[0146] [National Research and Development Project Supporting This Invention]
[0147] [Project ID] 1711195689
[0148] [Project Number] RS-2023-00229330
[0149] [Ministry Name] Ministry of Science and ICT
[0150] [Name of Project Management (Specialist) Agency] Information and Communications Technology Planning and Evaluation Institute
[0151] [Research Project Name] Development of Core Technologies for Realistic Content - Development of Core Technologies for Metaverse Media
[0152] [Research Project Name] Streaming 3D Digital Media Service Technology
[0153] [Contribution rate] 50 / 100
[0154] [Name of the project performing organization] Seoul National University of Science and Technology Industry-Academic Cooperation Foundation
[0155] [Research Period] January 1, 2023 - December 31, 2023
Claims
1. An operation of generating a point cloud for a scene based on information received from one or more cameras; An operation of identifying one or more objects included in the scene based on the point cloud; An action of determining each of the one or more objects as a dynamic object or a static object; An operation of generating a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object; An action of acquiring a first target view point of a first user through a first user terminal; An operation of determining an area corresponding to the first target time point among a plurality of areas of the point cloud constituting the dynamic object as a first main area, and determining an area not corresponding to the first target time point as a first auxiliary area; An operation of generating a first target image based on the first main area, the first auxiliary area and the virtual background; and An operation of transmitting the first target image to the first user terminal Including, The above first target image is output by the first user terminal, How to provide images.
2. In paragraph 1, An operation of acquiring a second target time point of the first user through the first user terminal; An operation of determining an area corresponding to the second target time point among the plurality of areas of the point cloud constituting the dynamic object as a second main area, and determining an area not corresponding to the second target time point as a second auxiliary area; An operation of generating a second target image based on the second main area, the second auxiliary area and the virtual background; and An operation of transmitting the second target image to the first user terminal Including more, How to provide images.
3. In paragraph 1, Each point in the above point cloud contains color information. How to provide images.
4. In paragraph 1, The action of determining each of the above one or more objects as a dynamic object or a static object is: If the type of the identified object is a person or an animal, an action to determine the object as a dynamic object Including, How to provide images.
5. In paragraph 1, An operation of generating a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object, An operation of determining background points corresponding to background objects determined as static objects among one or more of the above objects; An operation of generating a three-dimensional background structure by performing three-dimensional modeling on the above background points; and An operation of generating the virtual background by performing texturing on the above three-dimensional background structure. Including, How to provide images.
6. In paragraph 1, An operation of determining an area corresponding to the first target point time among a plurality of areas of the point cloud constituting the dynamic object as a first main area, and determining an area not corresponding to the first target point time as a first auxiliary area, An operation of determining an area directly corresponding to the first target time point as the first target area; and An operation of determining the first target area and areas adjacent to the first target area as the first main area. Including, How to provide images.
7. In paragraph 6, An action to obtain a third target point of view of a second user through a second user terminal Including more, An operation of determining an area corresponding to the first target point time among a plurality of areas of the point cloud constituting the dynamic object as a first main area, and determining an area not corresponding to the first target point time as a first auxiliary area, An operation of determining an area corresponding to the first target time point or the third target time point among the plurality of areas of the point cloud constituting the dynamic object as the first main area, and determining an area not corresponding to the first target time point and the third target time point as the first auxiliary area. Including, How to provide images.
8. In paragraph 1, The operation of generating a first target image based on the first main area, the first auxiliary area, and the virtual background is: An operation for adjusting the number of points within the first auxiliary area; An operation of generating a three-dimensional foreground structure by performing three-dimensional modeling on the dynamic object based on points of the first main area and points of the first auxiliary area; An operation of generating a virtual foreground by performing texturing on the above three-dimensional foreground structure; and An operation of generating the first target image based on the virtual foreground and the virtual background. Including, How to provide images.
9. A computer program stored on a computer-readable recording medium to execute the method of claim 1 by being combined with hardware.
10. The server, Communication module; processor; and Contains memory, The above memory, when executed by the processor, causes the server to: Generate a point cloud for a scene based on information received from one or more cameras, Identifying one or more objects included in the scene based on the point cloud, Determine each of the above one or more objects as a dynamic object or a static object, Generating a virtual background by rendering the background of the scene based on at least a portion of the point cloud corresponding to the static object, Obtain the first target view point of the first user through the first user terminal, Among the multiple regions of the point cloud constituting the dynamic object, a region corresponding to the first target time point is determined as a first main region, and a region not corresponding to the first target time point is determined as a first auxiliary region. Generate a first target image based on the first main area, the first auxiliary area, and the virtual background, Transmitting the above first target image to the above first user terminal Stores instructions that are set to be performed, Server.
11. An image providing method performed by a system server including a main server and a secondary server, An operation of the main server to identify one or more objects included in the scene based on the point cloud; An action by which the main server determines each of the one or more objects as a dynamic object or a static object; An operation in which the main server generates a virtual background by rendering a background of the scene based on at least a portion of the point cloud corresponding to the static object; An operation in which the main server transmits points corresponding to the dynamic object among the points of the point cloud and the virtual background to the auxiliary server; An operation in which the auxiliary server acquires the first target view point of the first user through the first user terminal; An operation in which the auxiliary server determines an area corresponding to the first target time point among a plurality of areas of the point cloud constituting the dynamic object as a first main area, and determines an area not corresponding to the first target time point as a first auxiliary area; The operation of the auxiliary server generating a first target image based on the first main area, the first auxiliary area, and the virtual background; and An operation in which the auxiliary server transmits the first target image to the first user terminal. Including, The above first target image is output by the first user terminal, How to provide images.
12. In paragraph 11, The above auxiliary server is a MEC (mobile edge computing) server. How to provide images.
Citation Information
Patent Citations
Image forming device, image forming method, and program
JP2023047882A
Depth image-based representation method for 3D objects, modeling method and apparatus using it, and rendering method and apparatus using the same
KR1020060107899A
A facemask in which the air may enter or leave through the earlobe straps
KR1020220115888A
Manufacturing Method for a face cleansing gloves
KR102030370B1
KR20190064050A