Method for generating a digital twin of a physical environment and method and system for surfing said digital twin

EP4747855A1Pending Publication Date: 2026-05-27OVER HOLDING SRL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
EP · EP
Patent Type
Applications
Current Assignee / Owner
OVER HOLDING SRL
Filing Date
2024-07-18
Publication Date
2026-05-27

AI Technical Summary

Technical Problem

Current methods for generating surfable digital twins of physical environments require high computing power, making them inaccessible to devices with lesser capabilities, and are also cumbersome due to the use of expensive 3D scanners and large file sizes.

Method used

A method that allows user devices with less computing power to generate and surf within a digital twin of a physical environment by acquiring and processing three-dimensional digital images, using a graphic processing unit to calculate and select points for a new three-dimensional mesh, and applying rendering algorithms to produce a surfable digital twin.

Benefits of technology

Enables the generation and interaction with digital twins on a wider range of user devices, including those with lower computing power, while reducing the complexity and cost associated with existing systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2024056954_23012025_PF_FP_ABST
    Figure IB2024056954_23012025_PF_FP_ABST
Patent Text Reader

Abstract

Method (1) for generating with a user device (2) a three-dimensional digital image of a physical environment, with respect to a desired viewing position and according to a desired viewing direction and a desired viewing angle in said physical environment, the method (1) including the following operational steps: A. through a data processing unit (3) of said user device (2), acquiring the desired viewing position, viewing direction and viewing angle and receiving one set of three-dimensional digital images of said physical environment, wherein each of said three-dimensional digital images is defined by at least: one vision point, one vision direction and one vision angle; color data and depth data; and wherein two three-dimensional digital images of said set, having adjacent vision points, represent at least one common part of said physical environment; and the vision points of said three-dimensional digital images are comprised within the neighborhood of said viewing position; B. through said processing unit (3) of said user device (2), transmitting said three-dimensional digital images to one graphic processing unit (4) of said user device (2) and calculating with said graphic processing unit (4) one three-dimensional mesh of points for each three-dimensional digital image of said received set of three-dimensional digital images; C. through said graphic processing unit (4) of said user device (2), selecting the portions of each mesh and of each three-dimensional digital image of said set of three-dimensional digital images, consistent with the desired viewing position, viewing direction and viewing angle, thereby obtaining the points of a new three-dimensional mesh of points, representing a new three-dimensional digital image; and D. through said graphic processing unit (4) of said user device (2) and starting from the new three-dimensional mesh of points, obtaining a new three-dimensional digital image at said desired viewing position and according to the desired viewing direction and viewing angle.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHOD FOR GENERATING A DIGITAL TWIN OF A PHYSICAL ENVIRONMENT AND METHOD AND SYSTEM FOR SURFING SAID DIGITAL TWIN

[0002] * * *

[0003] The present invention relates to a method for generating a digital twin of a physical environment, as well as a method and system for surfing within said digital twin, optionally in real time and interactively. The present invention also relates to a corresponding computer program as well as a medium for such a computer program.

[0004] Nowadays, with the advent of augmented reality and virtual reality systems and the spread of the metaverse, the need has arisen to digitally reproduce a copy of a physical environment, also known as a digital twin, so that one can virtually surf and interact within it, optionally in real time. Prior art document WO 2021 / 092455 Al teaches, for example, a method of generating an arbitrary two- dimensional view starting from a set of two-dimensional images stored in a database of an arbitrary image generation system of a scene. Document US 2012 / 127169 Al teaches a method for the guided navigation of a virtual camera in an interactive three-dimensional environment, the article by Moreland K. et al. entitled "Remote rendering for ultrascale data" teaches a rendering algorithm, and US application US 2019 / 026956 Al teaches the use of machine learning algorithms to derive 3D data from 2D images, according to the prior art.

[0005] In order to obtain a surfable digital twin of a physical environment, 3D modelling systems are currently available, which implement very time-consuming and computationally intensive operations. In fact, 3D scanners can be used, which are very expensive and generate very large files, which are complex in terms of their structure / architecture and take a long time to be downloaded. Not only that, the fruition of the digital twin of a physical environment generated by these systems, especially if it has a very complex structure, requires considerable computing power to allow a user device to surf through it, and this greatly limits the type of devices with which this activity can be carried out.

[0006] There is, therefore, a need to improve the state of the art in the field of the generation of surfable digital twins of physical environments and, more particularly, the main object of the present invention is to enable the generation of a digital twin of a physical environment by means of user devices with less computing power than currently used systems.

[0007] Another object of the present invention is to enable surfing within a digital twin of a physical environment by means of user devices with less computing power than currently used systems.

[0008] It is a specific object of the present invention a method for generating with a user device a three- dimensional digital image of a physical environment, with respect to a desired viewing position and according to a desired viewing direction and a desired viewing angle in said physical environment, said method including the following operational steps: A. through a data processing unit of said user device, acquiring the desired viewing position, viewing direction and viewing angle and receiving one set of three-dimensional digital images of said physical environment, wherein each of said three-dimensional digital images is defined by at least:

[0009] - one vision point, one vision direction and one vision angle;

[0010] - color data and depth data; and wherein

[0011] - two three-dimensional digital images of said set, having adjacent vision points, resulting in an overlapping representation of at least one common part of said physical environment; and

[0012] - the vision points of said three-dimensional digital images are comprised in the neighborhood of said viewing position;

[0013] B. through said processing unit of said user device, transmitting said three-dimensional digital images to one graphic processing unit of said user device and calculating with said graphic processing unit one three-dimensional mesh of points for each three-dimensional digital image of said received set of three-dimensional digital images;

[0014] C. through said graphic processing unit of said user device, automatically selecting the points of each three-dimensional mesh of points calculated for each three-dimensional digital image of said set of three-dimensional digital images, consistent with the desired viewing position, viewing direction and viewing angle, and calculating based on said points thus selected the points of a new three-dimensional mesh of points, representing a new three-dimensional digital image; and

[0015] D. through said graphic processing unit of said user device and starting from the new three- dimensional mesh of points, obtaining a new three-dimensional digital image at said desired viewing position and according to the desired viewing direction and viewing angle.

[0016] According to another aspect of the invention, each three-dimensional mesh of points can be processed, for each digital image of said three-dimensional digital images, starting from the respective depth data and is centered on the respective vision point.

[0017] According to a further aspect of the invention, step C can comprise the application of one automatic algorithm that determines the points of the new three-dimensional mesh of points associated to the new three-dimensional digital image and the corresponding color data, at the desired viewing position and according to the desired viewing direction and viewing angle, among the points of the three- dimensional meshes that have been calculated at step B and the corresponding color data and based on the distance between the points of each three-dimensional mesh of points, with respect to the viewing position, discarding the points of each mesh that are more distant, with respect to the desired viewing point, the desired viewing direction and viewing angle, than other points of at least another mesh, which instead are closer to the desired viewing point, according to the desired viewing direction and viewing angle. According to an additional aspect of the invention, step D can comprise applying at least one rendering algorithm, optionally a texturing algorithm, to the new three-dimensional mesh of points of the new three-dimensional digital image and to corresponding color data.

[0018] According to another aspect of the invention, if the user device is in said physical environment, the desired viewing position, the desired viewing direction and viewing angle can be directly and automatically acquired by the user device and depend on the position and orientation of the user device with respect to said physical environment, optionally provided to the data processing unit through position sensors of said user device, operatively connected thereto.

[0019] According to a further aspect of the invention, if the user device is located remotely with respect to said physical environment, the desired viewing position, desired viewing direction and viewing angle can be interactively provided in input to the user device, by a user of the same, optionally through I / O means of said user device operatively connected to the data processing unit.

[0020] It is also an object of the invention a method for generating a surfable digital twin of a physical environment with a user device, the method comprising the following operational steps:

[0021] I. transmitting a viewing position, a viewing direction and a viewing angle to a remote device, when they are acquired by a data processing unit of said user device, automatically through position sensors, or interactively through I / O means of said user device operatively connected thereto;

[0022] II. carrying out, through the user device, the method for generating a three-dimensional image of a physical environment as described above, based on the acquired viewing position, viewing direction and viewing angle, and a set of three-dimensional digital images, received in reply from the remote device, by the data processing unit, thereby obtaining a new three-dimensional digital image, at the acquired viewing position and according to the viewing direction and viewing angle;

[0023] III. representing, through a display of the user device that is operatively connected to the graphic processing unit, the new three-dimensional digital image thereby obtained, and going back to step I.

[0024] According to an aspect of the invention, said method can comprise, between step I and step II, through said remote device: a. receiving the viewing position, viewing direction and viewing angle sent by the user device; b. selecting, through a processing unit of the remote device, among a plurality of pre-computed three-dimensional digital images, said set of digital images having vision points comprised within the neighborhood of said viewing position; and c. transmitting said set of pre-computed three-dimensional digital images to said user device.

[0025] According to a further object of the invention, said method can comprise, before step b, through said user device or another user device: i. acquiring, through an image acquisition unit of the user device, operatively connected to a data processing unit, a plurality of bidimensional digital images of the physical environment, ii. calculating, for each bidimensional digital image of the plurality of bidimensional digital images, starting from data provided by position sensors of said user device, respective capture points and respective capture directions, and ill. transmitting the plurality of bidimensional images, the capture points and the respective capture directions to the remote device, for their processing into three-dimensional digital images.

[0026] According to an additional aspect of the invention, said step ii can comprise applying a calibrated slam model, optionally based on arkit and arcore technology, to data provided by said position sensors of said user device, obtaining said capture points and said respective capture directions.

[0027] According to another aspect of the invention, the method can comprise, after step iii, by said remote device: iv. receiving the plurality of bidimensional digital images of the physical environment, with the respective capture points and capture directions, v. pre-computing, using neural rendering algorithms, based on the received bidimensional digital images and the respective capture points and the respective capture directions, a corresponding plurality of three-dimensional digital images; and vi. storing the plurality of three-dimensional digital images in the remote device.

[0028] According to a further object of the invention, the number N of three-dimensional digital images of the plurality of three-dimensional digital images pre-computed at step v can vary based on the computing power of the remote device, the number of bidimensional digital images of the plurality of bidimensional digital images of the physical environment provided by the user device, the complexity of the physical environment to be digitally represented, the computing power of the user device that will have to download said three-dimensional digital images for implementing the method.

[0029] It is further an object of the invention a system for implementing a method as described above, wherein the user device comprises at least one data processing unit, one graphic processing unit, position sensor devices, at least one display, I / O means and one image acquisition unit operatively connected with each other and, said user device is a device comprised between: a smartphone, a tablet, a laptop, a desktop PC, a hand-held device; and wherein the remote device comprises at least one graphic processing unit and one processing unit.

[0030] According to an aspect of the invention, the user device can be configured for carrying out steps i-iii and / or steps l-lll of the method for generating a surfable digital twin of a physical environment, as well as method for generating a three-dimensional digital image, and wherein the remote device can be configured for carrying out steps iv-vi and / or steps a-c of the method for generating a surfable digital twin of a physical environment.

[0031] It is also an object of the invention a computer program, comprising instructions that cause the execution of the method for generating a surfable digital twin of a physical environment and / or the method or generating a three-dimensional digital image by at least said user device, and / or the execution of the method for generating a surfable digital twin of a physical environment, by the remote device.

[0032] It is also an object of the invention a computer readable medium, having stored therein the computer program above.

[0033] The present invention offers numerous advantages.

[0034] Indeed, with it, it is possible to obtain a digital twin of a physical environment, which can optionally be surfed in real time by devices with less computing power than conventional devices. This greatly expands the range of user devices capable of generating a digital twin of a physical environment and surfing within it.

[0035] The present invention will now be described, by way of illustration but not limitation, according to its preferred embodiment, with particular reference to the accompanying Figures of the drawings, wherein:

[0036] Figure 1 shows a flow chart of the main steps of the method for generating a digital image of a physical environment, according to the invention;

[0037] Figure 2a illustrates a user device suitable for implementing that method;

[0038] Figure 2b is a block diagram representation of the main components of the user device of Figure 2a;

[0039] Figure 3 is a representation of a preferred embodiment of three-dimensional digital images that can be provided as input to the method of Figure 1;

[0040] Figure 4 shows a three-dimensional mesh of points obtained through the method of Figure 1;

[0041] Figure 5 shows the three-dimensional mesh of Figure 3 to which color data has been associated;

[0042] Figure 6 is a corresponding digital image of a physical environment, generated through the method of Figure 1, obtained at a desired viewing position and with respect to a desired viewing direction and desired viewing angle, in a reproduced digital twin of said physical environment;

[0043] Figure 7 illustrates a flowchart of the method for generating a digital twin of a physical environment and surfing within it, according to the present invention; and

[0044] Figure 8 represents a system suitable for implementing the above methods of the invention.

[0045] Before going into the merits of the present invention, it should be noted that by "physical environment" in the present description and in the following claims it is construed a real three- dimensional space, such as a room inside a building, a town square, etc.

[0046] The term "digital twin of a physical environment" in the present description and in the following claims means a digital object (e.g. comprising one or more digital files or a data stream) which, when reproduced on a display screen of a user device, visually represents a copy of the physical environment in question or a part thereof. The term " surfable " or "surfed" referred to said digital twin, in the present description and in the following claims, refers to the fact that the visual representation of the digital twin, reproduced on the display screen of the user device varies according to a viewing position and / or a viewing direction and / or a viewing angle received as input by said user device, and the result for the user using the user device is to observe the digital twin first hand.

[0047] The term "viewing position", in the present description and in the following claims, is construed as a point in the three-dimensional space of the physical environment at which it is desired that the digital twin of that physical environment is displayed.

[0048] The term "viewing direction", in the present description and in the following claims, is construed as a direction along which, at the viewing position, the digital twin of the physical environment is desired to be displayed.

[0049] The term "viewing angle", in the present description and in the following claims, is construed as an angular amplitude at which the digital twin of the physical environment can be viewed, at the viewing position and along the viewing direction. Typically, this is an angle between 0° and 360°.

[0050] With reference, moreover, to a three-dimensional digital image of a physical environment, which will be further discussed below, it is clarified that in the present description and the following claims:

[0051] - the term "vision point" is construed as a point in three-dimensional space of said physical environment, at which a corresponding image acquisition device would acquire that digital image;

[0052] - the term "vision direction" is construed as a direction along which, at the vision point, the abovesaid image acquisition device would be oriented to acquire that digital image; and

[0053] - the term "vision angle" is construed as the magnitude of the maximum angle at which that image acquisition device would acquire that digital image, placed at the vision point and oriented along the "vision direction".

[0054] The method of the present invention, configured to generate a three-dimensional digital image of a physical environment with respect to a viewing position in said physical environment and according to a desired viewing direction and desired viewing angle, is depicted in Figure 1 and referred to therein by reference numeral 1 and can be advantageously implemented by a user device 2, depicted in Figures 2a and 2b and further discussed below.

[0055] Method 1 comprises a first operational step A of receiving as input, via a data processing unit 3 of the user device 2, a set of three-dimensional digital images of the physical environment in question, wherein each one of the three-dimensional digital images is defined by at least:

[0056] - a vision point, a vision direction and a vision angle, as well as by

[0057] - color data, representing, the colors of the objects present in the portion of said physical environment represented in the three-dimensional digital image, and - depth data, representing the distance of such objects from the vision point of such three-dimensional digital image.

[0058] Specifically, vision point, vision direction, vision angle and color and depth data can be represented in any suitable way. For example, the vision point may be represented by spatial coordinates expressed in the form of a numerical set of dimension 3, wherein the numbers in the set represent the cartesian coordinates (referring to a predefined reference point in that physical environment - optionally the origin of a cartesian coordinate system defined for that physical environment) of the point in the physical environment at which a corresponding imaging device would acquire that digital image. The vision direction can be represented by a vector, originating at the vision point, and the vision angle can be represented by a solid angle between 0° and 360°. In addition, the color and depth data may be represented in the form of two numerical matrices, wherein the values of the fields of said numerical matrices represent, in a first matrix, the colors of the objects populating the portion of that physical environment represented in that three-dimensional digital image, and in a second matrix, represent the distance of said objects, with respect to the vision point of said three-dimensional digital image. The person skilled in the art, however, will have no difficulty in understanding how vision points, vision direction, vision angle, color and depth data can be represented in any other suitable way, without thereby falling outside the scope of protection of the present invention.

[0059] In the abovesaid set of three-dimensional digital images, according to a particularly advantageous aspect of the invention, two three-dimensional digital images of the set of three- dimensional digital images, having adjacent vision points (i.e., between which another vision point of another three-dimensional digital image of the set is not included), represent at least a common part of the physical environment, optionally with an overlap of not less than 5%, optionally not less than 15%, more optionally not less than 60%, whereby they have at least some common features in terms of color data. To clarify, purely by way of illustration and not limitation, given two adjacent three-dimensional digital images and the amplitudes of the areas of their respective projections on the ground, with respect to the physical environment they represent, there would be an overlap of no less than 60% between these three-dimensional digital images, if at least 60% of the projection area of one of the aforementioned three-dimensional digital images were to overlap with the projection area of the other three-dimensional digital image. In any case, the person skilled in the art will have no difficulty in understanding how other methodologies can also be implemented to establish the degree of overlap between the aforementioned adjacent three-dimensional digital images, depending also on the amplitude of the respective vision angles, without thereby falling outside the scope of protection of the present invention.

[0060] Returning to invention method 1, again according to a particularly advantageous aspect, the vision points of each three-dimensional digital image of the set received as input at step A of invention method 1 are comprised within a neighborhood of the viewing position, i.e. within a predefined distance d therefrom. Thus, for example, according to a preferred embodiment of the invention, given the cartesian coordinate system of the physical environment described above, the set of three-dimensional digital images received at step A may comprise: least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is above the viewing position, at least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is below the viewing position, at least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is in front of the viewing position, at least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is behind the viewing position, at least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is on the right of the viewing position, and, finally, at least one three-dimensional digital image of the physical environment whose vision point is a point in the space of the physical environment that is on the left of the display position. All images of the set of three-dimensional digital images, as anticipated above, are within a predefined distance d from the viewing position, and this predefined distance d may depend on various factors, including, for example, the computing power and memory capacity of the user device 2 implementing the invention method 1, as well as the quality (e.g., in terms of data download speed) of a connection network to which that user device 2, as will also be further discussed below, may connect to acquire those three-dimensional digital images.

[0061] Purely by way of example and not limitation, Figure 3 shows the color data of two digital images according to a preferred embodiment of the invention, wherein the viewing angle is 360°, which can be received as input by the user device 2, during step A. The color and depth data, as mentioned above, can be represented for each three-dimensional digital image in the form of two numerical matrices, the fields of which are populated with values representing, in a first matrix Ml, the colors of the objects that populate the portion of the physical environment represented in the 360° digital image, and in a second matrix M2, the distance of these objects with respect to the vision point of that 360° digital image. Specifically, the first row of images in Figure 3 represents the color and depth data (Ml,l and M2,l) of a 360° digital image at a first vision point and the second row of images in Figure 3 represents the color and depth data (Ml, 2 and M2, 2) of a 360° digital image at a second vision point in the physical environment.

[0062] The method 1 of the invention thus provides, in a subsequent step B, forwarding the received three-dimensional digital images, to a graphic processing unit 4 of the user device 2 operatively connected to the data processing unit 3, and the calculation by said graphic processing unit 4 of a three- dimensional mesh of points for each three-dimensional digital image of the abovesaid received set of three-dimensional digital images. Each three-dimensional mesh of points, according to a preferred embodiment of the invention, is processed based on the respective depth data which represent, as mentioned above, the distance of the objects of the physical environment represented in said three- dimensional digital image, relative to the vision point. Thus, in the example above, each three- dimensional mesh of points is obtained, based on matrices M2,l and M2, 2 associated with the two 360° digital images provided as input to the user device 2.

[0063] In this case, also, each three-dimensional mesh of points can be represented in any suitable way. By way of non-limiting example, it can be represented as a data file in .stl or .gib format, or even in .fbx do .obj format.

[0064] The invention method 1 advantageously provides a subsequent step C of automatic selection of the portions of each three-dimensional mesh of points and each corresponding three-dimensional digital image, consistent with said desired viewing position, viewing direction and viewing angle, based on:

[0065] - the vision direction and vision angle of each one of the images of the three-dimensional digital image set,

[0066] - the abovesaid viewing direction and the abovesaid desired viewing angle, and

[0067] - the distance of the points of each three-dimensional mesh of points, with respect to the viewing position.

[0068] In this step, in practice, by applying automatic (i.e. without user intervention) linear algorithms, by way of example but not limitation for example shader algorithms, two or more three-dimensional meshes of points and the corresponding three-dimensional digital images are "merged" by the graphic processing unit 4 of the user device 2, discarding the points of these that would be in the background (further away), with respect to the viewing point and according to the desired viewing direction and viewing angle, because they would be "screened" by other points of the other mesh or of another mesh, which would instead be in the foreground (closer) to the viewing point and according to the desired viewing direction and viewing angle.

[0069] The result of the selection made at step C of method 1 is a new three-dimensional mesh of points (e.g. as represented in Figure 4), representative of a new three-dimensional digital image, with the corresponding color data, e.g. as represented in Figure 5, at the desired viewing position and viewing direction and viewing angle.

[0070] Purely by way of example and not limitation, and according to an embodiment of the invention, the new three-dimensional mesh of points can be obtained at step C through Voxelisation and Averaging, as follows. The physical environment to be digitally represented is subdivided by the graphic processing unit 4 of the user device 2 into a grid of cubes, called voxels, of a predefined size that can be set a priori. Each voxel contains a number of points from the three-dimensional meshes of points calculated at step B. For each voxel, at step C the graphic processing unit 4 of the user device performs an averaging of the coordinates of the points included in the voxel and of the color data. According to an advantageous embodiment of the invention, instead of calculating a simple arithmetic mean, which could be affected by noise or spurious values, the median is calculated, which is a more robust value that better represents the data. In this respect, it should be noted that for a set of n ordered coordinate or color values, if n is odd, the median is the central value. If n is even, the median is the mean of the two central values. Formally, given xltx2. ■ ■ xnordered in ascending order, the median:

[0071] This formula is then applied both to the spatial (x, y, z) co-ordinate values of the points and to the color (rgb) values of the points included in each voxel, with the clarification that if, in a given voxel, the points taken are only the points belonging to the same three-dimensional mesh of points, then these points are not averaged. This makes it possible to preserve the original structure of the three-dimensional mesh of points in those areas, avoiding distortion or unwanted merging of details that do not strictly overlap.

[0072] Then, based on the points of the new three-dimensional mesh of points of the new three- dimensional digital image and the corresponding color data, the corresponding new three-dimensional digital image is automatically processed by the graphic processing unit 4 of the user device 2 at the next step D of the invention method 1 (Figure 6).

[0073] This processing is carried out by applying to the new three-dimensional mesh of points of the new three-dimensional digital image and the corresponding color data a rendering algorithm, for example a suitable algorithm selected among texturing algorithms. According to the non-limiting example given above, the graphic processing unit 4 of user device 2 automatically manages the display of the new three-dimensional mesh of points that are in the foreground with respect to the desired viewing position, viewing direction and viewing angle, ignoring points that are behind. This can be implemented by means of depth buffering and occlusion culling techniques, which allow to only those parts of the new three-dimensional mesh of points point that are visible from the desired viewing position, viewing direction and viewing angle to be displayed correctly.

[0074] In view of the above, it is quite apparent that, with the invention method 1, it is possible to obtain a new three-dimensional digital image of a physical environment at any desired viewing position and according to a desired viewing direction and desired viewing angle of said physical environment, starting from a reduced set of previously acquired three-dimensional digital images, optionally from six (6) previously acquired three-dimensional digital images located in the neighborhood of that viewing position. It is thus highlighted the clear difference in approach between the present method and the method for generating an arbitrary view taught in WO 2021 / 092455 Al, wherein the arbitrary view is a two-dimensional - and not three-dimensional - image obtained (according to method 300) from a set of two-dimensional - and not three-dimensional - images stored on a database, the two-dimensional images stored on the database being themselves optionally previously obtained (according to a method 400) from a three-dimensional mesh model. In the present invention, instead, the new digital image obtained at step D is a three-dimensional image, obtained from a three-dimensional mesh of points, in turn obtained from a set of three-dimensional meshes of points of corresponding three-dimensional digital images, previously stored on a database.

[0075] The desired viewing position, viewing direction and viewing angle according to which, with the invention method, the desired new three-dimensional digital image of the physical environment is generated, are provided to the user device 2 in the form of input data, according to two different options.

[0076] According to a first option of the present invention, in fact, if the user device 2 is physically present in the physical environment whose twin is to be digitally reproduced, the desired viewing position, viewing direction and viewing angle are directly obtained by the user device 2 and depend on the position and orientation of the user device 2 in the physical environment, optionally on the position within the physical space of the position sensors 5 thereof, optionally an accelerometer and / or motion sensors or other sensors suitable for the purpose, mounted on the 2 user device. By way of example and not by way of limitation, this is the case of a user who is, for example, inside an archaeological site and wants to observe the corresponding digital twin of the archaeological site via a display 6 of their user device 2, while they physically moves within the archaeological site itself.

[0077] According to a second option of the present invention, however, if the user device 2 is located at a distance from the physical environment of which the digital twin is to be reproduced, the desired viewing position, viewing direction and viewing angle are supplied manually as input to the user device 2, in any suitable way, by a user of the user device 2 itself, for example via suitable I / O means 7, such as a keyboard and / or a mouse and / or a touch pad and / or or a touch screen and / or a joystick, etc. of the user device 2 or operatively connected thereto. By way of example and not by way of limitation, this is the case of a user who, using their user device 2 for example from their home, wants to observe, for example, the digital twin of the interior of a museum or the streets of a city of their choice, reproduced on their user device 2.

[0078] Advantageously, the aforementioned method 1 for generating a digital image is included in a method 10 for generating a surfable digital twin of a physical environment, represented in Figure 7, which also forms the subject of the present invention. That method 10 can advantageously be implemented by a system 100 represented in Figure 8, comprising at least one user device 2 as described above and at least one remote device 9, which will be better described below, and is executed every time a data processing unit 3 of the user device 2 acquires (step I) the aforementioned desired viewing position, the desired viewing direction and the desired viewing angle in the ways described above and that is, for example, transmitted by position sensors 5 and / or by I / O means 7 of the user device 2 operatively connected thereto. When this happens, the desired viewing position, viewing direction and viewing angle are transmitted to the remote device 9 (see Figures 7 and 8), and the invention method 10 therefore provides, in one subsequent step II, the execution by the user device 2 of the method 1 described above, based on the aforementioned viewing position, the desired viewing direction and desired viewing angle, as well as a set of three-dimensional digital images as described above, transmitted in reply by the remote device 9, at the end of which the graphic processing unit 4 of the user device 2 generates the new three- dimensional digital image at the viewing position and with respect to the desired viewing direction and desired viewing angle.

[0079] Therefore, the method 10 for generating the digital twin of the physical environment according to the invention provides in a subsequent step III to reproduce, via the display 6 of the user device 2, operatively connected to the graphic processing unit 4, the new digital image thus obtained (at the viewing position and with respect to the desired viewing direction and viewing angle), optionally a portion thereof in two-dimensional format (see for example Figure 6) and, method 10 then resumes from step I, when a new viewing position and / or a new viewing direction and / or viewing angle are acquired by the data processing unit 3 of the user device 2.

[0080] By way of example but not by way of limitation, the repetition of steps l-lll of the abovesaid method can occur when, for example, according to the first option of the present invention described above, the user moves within the archaeological site so that their user device 2 automatically detects from the position sensors 5 thereof a change in their position in the physical space and / or a change in their orientation. According to the second option of the present invention, the repetition of the method steps l-lll described above can occur when the new viewing position and / or the new viewing direction and / or the new viewing angle are actively set by the user by sending a corresponding signal to the processing unit 3 of the user device 2, via the I / O means 7 (for example by touching the touch screen of the user device 2).

[0081] Returning to step I of method 10 and the transmission of the viewing position, viewing direction and viewing angle to the remote device 9, method 10 provides that the remote device 9, having received in a first step (a) that viewing position, viewing direction and viewing angle transmitted, selects in reply, in a second step (b), by means of its own processing unit (not represented in the drawings) and between a plurality of pre-calculated three-dimensional digital images, the set of three-dimensional digital images having spatial coordinates in the neighborhood of the received viewing position. The invention method 10 therefore provides, at subsequent step (c), that the selected set of pre-calculated three-dimensional digital images is transmitted back to the user device 2 for the execution of step II of the method 10, for example via any suitable network, wired or non-wired.

[0082] The plurality of three-dimensional digital images precalculated according to the invention method 10 can have been previously received by the remote device 9 and stored thereon or, according to a preferred variant of the present invention (represented in Figure 7), it can be generated directly and stored in the remote device 9, at least preliminarily to step (b) above, starting from a plurality of two- dimensional digital images of the physical environment, taken at respective capture points and along respective capture directions and received by the remote device 9 together with the respective capture points and respective capture directions, in a step iv of the aforementioned method 10. In this case, the invention method 10, following phase iv, comprises a step v of pre-calculation of the corresponding plurality of three-dimensional digital images, for example by means of neural rendering or NeRF algorithms, based on the two-dimensional digital images received, the respective capture points and the respective capture directions, after which, at the subsequent step vi, that plurality of pre-calculated three-dimensional digital images is stored in the remote device 9.

[0083] In turn, the plurality of two-dimensional digital images of the physical environment, provided at step iv, can be received by the remote device 9 in a known manner, after a user device 2 equal to or different from the user device 2 configured to perform steps I -III of invention method 10, has acquired (in a step i) by means of its own image acquisition unit 8 operatively connected to the data processing unit 3, such two-dimensional digital images of the physical environment, has then calculated ( at step ii), for each one of these two-dimensional digital images the respective capture points and the respective capture directions, starting from data provided by its own position sensors 5, and has therefore transmitted (in phase iii) the two-dimensional digital images, with the respective capture points and the respective capture directions to remote device 9, for their processing into three-dimensional digital images.

[0084] The capture points and capture directions of the two-dimensional digital images captured with the user device 2 can be derived according to a variant of the present invention, by applying a calibrated slam model, optionally based on arkit and arcore technology, to signals outputted by the position sensors 5 of the user device 2.

[0085] Going back to the plurality of three-dimensional digital images pre-calculated and stored at step vi of method 10, this is obtained starting from the plurality of two-dimensional digital images, transmitted with the respective capture points and the respective capture directions, by applying standard photogrammetry algorithms thereto (for example, the Colmap algorithm) and standard neural networks optimized for calculating the capture positions and capture directions of the two-dimensional digital images themselves in terms of translation and rotation, which algorithms use standard training datasets, without the need for a dedicated training. Then, based on the respective capture points and respective capture directions calculated in terms of translation and rotation and applying corresponding neural rendering algorithms of any suitable type, for example NeRF artificial intelligence algorithms, the plurality of pre-calculated three-dimensional digital images is obtained.

[0086] According to a particularly advantageous aspect of the present invention, the images of the plurality of precalculated three-dimensional digital images are precalculated in a predetermined number N of viewing positions, greater than the number of capture positions of the plurality of starting two- dimensional digital images. This number N can vary, for example, depending on the computing power of the remote device 9, the number of digital images of the plurality of two-dimensional digital images of the physical environment supplied as input to the remote device 9 by the user device 2, the complexity of the physical environment to be digitally represented, the computing power of the user device 2 which will then have to download the aforementioned three-dimensional digital images, via the above network, for the implementation of method 1 described above.

[0087] In this way, therefore, a sort of grid of N three-dimensional digital images pre-calculated on the remote device 9 according to the method 10 of the present invention is obtained, a certain number of them included in the vicinity of the desired viewing position is downloaded onto the user device 2 and with them, by implementing the method 1 described above, a new "approximate" three-dimensional digital image of the physical environment is generated, which is reproduced on the display 6 of the user device 2, at each desired viewing position, according to the respective viewing direction and the respective viewing angle, received from the user device 2. This occurs without compromising the calculation capacity of the user device 2, since the method 1 of the invention described above allows the new three-dimensional digital image to be "approximated" starting from at least one pair of three- dimensional digital images in its surroundings, implementing linear algorithms implemented directly and in an optimized manner by the graphic processing unit 4 which, being specialized for this type of implementation, allows performances that are several orders of magnitude higher than the ones that can be obtained by executing the same algorithms via the central processing unit 3 data of the same user device 2.

[0088] The method 10 for generating a surfable digital twin of a physical environment and the method 1 for generating a digital image of this physical environment are implemented, as anticipated above, by a system 100 comprising at least one user device 2 and a remote device 9.

[0089] The remote device 9 is, for example, a server device, which, having to implement the steps of the method iv-vi above, includes at least a GPU and a processing unit having high computing power. The user device 2, instead, can be a smartphone, a tablet, a laptop PC, a desktop PC, a PDA, etc. provided that it is equipped with a data processing unit 3, a graphic processing unit 4, position sensor devices 5, a display 6, I / O means 7 of the type described above and an image acquisition unit 8, operatively connected to each other as described above and configured to implement methods 1 and 10 of the invention. In this regard, it is specified that the user device 2 which carries out steps i-iii of method 10 can be the same user device 2 which subsequently carries out steps l-lll of method 10. Alternatively, steps i-iii and steps I- III of method 10 can be performed by 2 different user devices. The user device 2 that carries out steps i- iii must necessarily be in the physical environment of which one wants to generate the digital twin, while the user device 2 that generates the digital image of the physical environment for each viewing position and according to the viewing direction and viewing angle acquired by method 1 described above can be even placed at a distance. In any case, both user devices can surf (in the sense described above) through the digital twin of the physical environment thus obtained.

[0090] The aforementioned devices are configured to have installed thereon a computer program 20, which is also the object of the present invention, comprising instructions that cause the execution of the method 1 by at least one user device 2 and / or of the method 10, by that at least one user device 2 or remote device 9, as described above.

[0091] Finally, it is an object of the present invention a computer-readable medium 200, having stored therein the above-mentioned computer program 20.

[0092] In view of the above, it is clear that the methods, the system, the computer program and computer readable medium described above solve the problems set out in the introduction.

[0093] The method 1 described above for generating a three-dimensional digital image, not requiring high computing power and being performed by the graphic processing unit 4 of the user device 2, can be implemented by a wide range of user devices 2 which up to now have not been used for this type of application, such as smartphones. Method 10 and the corresponding system 100 allow one to generate a digital twin of a physical environment and surf within it, quickly, optionally in real time, offering an adequate user experience.

[0094] In the foregoing, the preferred embodiments have been described and variations of the present invention have been suggested, but it is to be understood that those skilled in the art will be able to make modifications and changes without thereby departing from the relevant scope of protection, as defined by the appended claims.

Claims

CLAIMS1. Method (1) for generating with a user device (2) a three-dimensional digital image of a physical environment, with respect to a desired viewing position and according to a desired viewing direction and a viewing angle in said physical environment, said method (1) including the following operational steps:A. through a data processing unit (3) of said user device (2), acquiring the desired viewing position, viewing direction and viewing angle and receiving one set of three-dimensional digital images of said physical environment, wherein each of said three-dimensional digital images is defined by at least:- one vision point, one vision direction and one vision angle;- color data and depth data; and wherein- two three-dimensional digital images of said set, having adjacent vision points, resulting in an overlapping representation of at least one common part of said physical environment; and- the vision points of said three-dimensional digital images are comprised in the neighborhood of said viewing position;B. through said processing unit (3) of said user device (2), transmitting said three-dimensional digital images to one graphic processing unit (4) of said user device (2) and calculating with said graphic processing unit (4) one three-dimensional mesh of points for each three-dimensional digital image of said received set of three-dimensional digital images;C. through said graphic processing unit (4) of said user device (2), automatically selecting the points of each three-dimensional mesh of points calculated for each three-dimensional digital image of said set of three-dimensional digital images, consistent with the desired viewing position, viewing direction and viewing angle, and calculating based on said points thus selected the points of a new three-dimensional mesh of points, representing a new three-dimensional digital image; andD. through said graphic processing unit (4) of said user device (2) and starting from the new three- dimensional mesh of points, obtaining a new three-dimensional digital image at said desired viewing position and according to the desired viewing direction and viewing angle.

2. Method (1) according to claim 1, wherein each three-dimensional mesh of points is processed, for each digital image of said three-dimensional digital images, starting from the respective depth data and is centered on the respective vision point.

3. Method (1) according to claim 1 or 2, wherein step C comprises the application of one automatic algorithm that determines the points of the new three-dimensional mesh of points associated to the new three-dimensional digital image and the corresponding color data, at the desired viewing position and according to the desired viewing direction and viewing angle, among the points of the three-dimensional meshes that have been calculated at step B and the corresponding color data and based on the distancebetween the points of each three-dimensional mesh of points, with respect to the viewing position, discarding the points of each mesh that are more distant, with respect to the desired viewing point, the desired viewing direction and viewing angle, than other points of at least another mesh, which instead are closer to the desired viewing point, according to the desired viewing direction and viewing angle.

4. Method (1) according to any previous claim, wherein step D comprises applying at least one rendering algorithm, optionally a texturing algorithm, to the new three-dimensional mesh of points of the new three-dimensional digital image and to corresponding color data.

5. Method (1) according to any previous claim, wherein if the user device (2) is in said physical environment, the desired viewing position, the desired viewing direction and viewing angle are directly and automatically acquired by the user device (2) and depend on the position and orientation of the user device (2) with respect to said physical environment, optionally provided to the data processing unit (3) through position sensors (5) of said user device (2), operatively connected thereto.

6. Method (1) according to any claim 1 to 4, wherein if the user device (2) is located remotely with respect to said physical environment, the desired viewing position, desired viewing direction and viewing angle are interactively provided in input to the user device (2), by a user of the same, optionally through I / O means (7) of said user device (2) operatively connected to the data processing unit (3).

7. Method (10) for generating a surfable digital twin of a physical environment with a user device (2), the method comprising the following operational steps:I. transmitting a viewing position, a viewing direction and a viewing angle to a remote device (9), when they are acquired by a data processing unit (3) of said user device (2), in an automatic way through position sensors (5), or interactively through I / O means (7) of said user device (2) operatively connected thereto;II. carrying out, through the user device (2), the method (1) according to any previous claim, based on the acquired viewing position, viewing direction and viewing angle, and a set of three-dimensional digital images received, in reply from the remote device (9), by the data processing unit (3), obtaining a new three-dimensional digital image, at the acquired viewing position and according to the viewing direction and viewing angle;III. representing, through a display (6) of the user device (2) that is operatively connected to the graphic processing unit (4), the new three-dimensional digital image thereby obtained, and going back to step I.

8. Method (10) according to claim 7, comprising, between step I and step II, through said remote device (9): a. receiving the viewing position, viewing direction and viewing angle sent by the user device (2);b. selecting, through a processing unit of the remote device (9), among a plurality of pre-computed three-dimensional digital images, said set of digital images having vision points comprised in the neighborhood of said viewing position; and c. transmitting said set of pre-computed three-dimensional digital images to said user device (2).

9. Method (10) according to claim 8, comprising, before step b, through said user device (2) or another user device (2): i. acquiring, through an image acquisition unit (8) of the user device (2), operatively connected to a data processing unit (3), a plurality of bidimensional digital images of the physical environment, ii. calculating, for each bidimensional digital image of the plurality of bidimensional digital images, starting from data provided by position sensors (5) of said user device (2), respective capture points and respective capture directions, and ill. transmitting the plurality of bidimensional images, the capture points and the respective capture directions to the remote device (9), for their processing into three-dimensional digital images.

10. Method (10) according to claim 9, wherein said step ii comprises applying a calibrated slam model, optionally based on arkit and arcore technology, to data provided by said position sensors (5) of said user device (2), obtaining said capture points and said respective capture directions.

11. Method (10) according to claim 9 or 10, comprising, after step iii, by said remote device (9): iv. receiving the plurality of bidimensional digital images of the physical environment, with the respective capture points and capture directions, v. pre-computing, using neural rendering algorithms, based on the received bidimensional digital images and the respective capture points and the respective capture directions, a corresponding plurality of three-dimensional digital images; and vi. storing the plurality of three-dimensional digital images in the remote device (9).

12. Method (10) according to claim 11, wherein the number N of three-dimensional digital images of the plurality of three-dimensional digital images pre-computed at step v varies based on the computing power of the remote device (9), the number of bidimensional digital images of the plurality of bidimensional digital images of the physical environment provided by the user device (2), the complexity of the physical environment to be digitally represented, the computing power of the user device (2) that will have to download said three-dimensional digital images for implementing the method (1).

13. System (100) for implementing a method (1,10) according to any claim 7 to 12, wherein the user device (2) comprises at least one data processing unit (3), one graphic processing unit (4), position sensor devices (5), at least one display (6), I / O means (7) and one image acquisition unit (8) operatively connected with each other and, said user device is a device comprised between: a smartphone, a tablet, a laptop, a desktop PC, a hand-held device; and wherein the remote device (9) comprises at least one graphic processing unit and one processing unit.

14. System (100) according to claim 13, wherein the user device (2) is configured for carrying out steps i-iii and / or steps l-lll of the method (10) according to claims 7 , 9 and 10, as well as method (1) according to any claim 1 to 6, and wherein the remote device (9) is configured for carrying out steps iv-vi and / or steps a-c of the method 10 according to claims 8, 11 and 12.

15. Computer program (20), comprising instructions that cause the execution of the method (1) according to any claim 1 to 7 and / or the method (10) according to claim 9 and 10, by at least said user device (2), and / or the execution of the method (10) according to any claim 8, 11 and 12, by the remote device (9).

16. Computer readable medium (200), having stored therein the computer program (20) according to claim 15.