Method for generating a three-dimensional digital twin of a physical environment

A smartphone-based computer vision method using SfM and NeRF simplifies and accelerates the creation of high-precision digital twins, addressing the limitations of existing technologies by providing accessible and efficient 3D model generation.

FR3166733A1Pending Publication Date: 2026-03-27SN SNCF
0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
FR · FR
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-26
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing methods for creating three-dimensional digital twins, such as CAD, photogrammetry, and laser scanning, are complex, expensive, and require specialized skills, making them inaccessible and impractical for rapid generation in applications like the railway sector.

Method used

A method using standard imaging devices like smartphones and computer vision techniques, including Structure from Motion (SfM) and Neural Radiance Fields (NeRF), to generate high-precision digital twins by acquiring, processing, and reconstructing images to create detailed and realistic 3D models.

Benefits of technology

Enables rapid, cost-effective, and accessible creation of high-precision digital twins, suitable for real-time updates and diverse applications, without the need for expensive equipment or specialized technical skills.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader

Abstract

METHOD FOR GENERATION OF A THREE-DIMENSIONAL DIGITAL TWIN OF A PHYSICAL ENVIRONMENT. The invention relates to a method for generating a three-dimensional digital twin 3.4 of a physical environment 1 comprising the following steps: a step E0 of acquiring, by at least one camera, a plurality of images 3 of said physical environment 1; a step E1 of receiving said plurality of images 3; a step E2 of estimating poses from each camera 2 and extracting depth data from the received images; a step E3 of reconstructing a digital scene 3.3 of said physical environment 1 from said plurality of received images 3, said estimated camera poses, and said extracted depth data; a step E4 of generating at least one three-dimensional representation of said physical environment 1 from said reconstructed digital scene 3.3; a step E5 of generating a digital twin 3.4 of said physical environment 1 from said three-dimensional representation representing said physical environment 1. Figure for the abbreviation: figure 1.
Need to check novelty before this filing date? Find Prior Art

Description

Title of the invention: Method for generating a three-dimensional digital twin of a physical environment. Technical field of the invention

[0001] The present invention relates to the technical field of three-dimensional modeling of objects and physical environments, and more particularly, the technical field of computer vision for the generation of high-precision three-dimensional digital twins of physical environments from images acquired by optical sensors. Technological background

[0002] Existing techniques for creating three-dimensional digital twins are mainly based on methods such as Computer-Aided Design (CAD), photogrammetry, and laser scanning (LiDAR). Each of these approaches has advantages, but also limitations that make their use complex, expensive, and not readily accessible for certain applications, particularly in the railway sector.

[0003] Computer-Aided Design is widely used to create detailed and accurate 3D models of technical structures. However, this method requires specialized modeling skills, and the creation processes are often lengthy and laborious. Furthermore, CAD models, although extremely detailed, lack visual realism because they are based on theoretical plans rather than real-world data.

[0004] Photogrammetry, which involves creating 3D models from a series of photographs taken from different angles, is another common technique. Although this method produces realistic results, it requires a large number of images, precise camera calibration, and in-depth knowledge of photogrammetry software to obtain accurate reconstructions. Furthermore, data processing time is often lengthy, making this method less practical for applications requiring the rapid creation of 3D models.

[0005] Laser scanning, on the other hand, is a technology that allows for the precise capture of the shape and structure of a scene or object using laser beams. This method is particularly effective for capturing complex environments with high accuracy. However, it requires specialized and expensive equipment, as well as technical expertise to interpret the captured data. Furthermore, the process of acquiring and processing LiDAR data is generally slow, which can be a disadvantage for applications requiring rapid updates of digital models.

[0006] Despite these advanced technologies, there is a pressing need for a simpler, faster, and more accessible method for creating three-dimensional digital twins, particularly in contexts where time and resources are limited, such as in the railway sector. Current approaches are often inaccessible to non-specialist users and do not allow for the generation of 3D models in real time or dynamically, thus limiting their practical application in many fields.

[0007] The present invention aims to solve these problems by introducing a process that simplifies and accelerates the creation of 3D digital twins while maintaining a high level of accuracy and realism, without requiring expensive equipment or specialized technical skills. Objectives of the invention

[0008] The invention aims to provide a method for generating high-precision three-dimensional digital twins of physical environments from visual data acquired by standard imaging devices, such as conventional cameras, including smartphones. This method leverages substantial advances in computer vision and deep learning to facilitate and accelerate the creation of three-dimensional models.

[0009] More specifically, one embodiment of the invention aims to provide a method enabling the efficient and rapid creation of three-dimensional representations, thus minimizing the time and financial constraints inherent in conventional methodologies such as three-dimensional laser scanning and photogrammetry. Description of the invention

[0010] To this end, the invention relates to a method for generating a three-dimensional digital twin of a physical environment comprising the following steps:

[0011] an acquisition step by at least one camera of a plurality of images of said physical environment.

[0012] a step of receiving said plurality of acquired images;

[0013] a step of estimating a plurality of poses for each camera, called poses camera, and extraction of estimated depth data from said plurality of received images;

[0014] a step of reconstructing a digital scene of said physical environment from said plurality of received images, said estimated camera poses and said extracted depth data;

[0015] a step of generating at least one three-dimensional representation representing said physical environment from said reconstructed digital scene;

[0016] a step of generating a digital twin of said physical environment from said three-dimensional representation representing said physical environment.

[0017] By proceeding in this way, the invention proposes to simplify, accelerate and improve the quality of the rendering when creating a three-dimensional digital twin of a physical environment.

[0018] This solution makes it possible in particular to acquire images from inexpensive equipment such as a smartphone or a conventional digital camera.

[0019] Advantageously, the step of acquiring a plurality of images of said physical environment makes it possible to capture a large quantity of images in a short time. Thus, the plurality of images obtained provides a rich visual data base, essential for an accurate reconstruction of the physical environment.

[0020] The step of receiving a plurality of images including the plurality of images of said physical environment acquired allows the creation of the digital twin to be carried out on the basis of said acquired images.

[0021] The step of estimating a plurality of camera poses and extracting depth data from said plurality of received images makes it possible to understand the position and orientation of each camera at each instant, as well as the distance of objects from each camera in the environment in order to properly align the images and create an accurate three-dimensional representation of the physical environment.

[0022] Advantageously, an algorithm called "SfM", from the English "Structure from Motion", can be used to extract depth information from the images composing said plurality of images.

[0023] The step of reconstructing a digital scene of said physical environment from said plurality of received images, said estimated camera poses and said extracted depth data makes it possible to create a faithful virtual representation of the physical environment by transforming the raw data into a coherent digital scene, facilitating subsequent modeling and analysis steps.

[0024] Advantageously, the digital scene reconstruction step is executed by an algorithm based on neural architectures of the “Neural Radiance Fields” (NeRF) type; in particular, this step can be executed using models such as “fast-NeRF”, “TensorRF” or “instant-ngp”.

[0025] The step of generating at least one three-dimensional representation of said physical environment from said reconstructed digital scene makes it possible to structure the digital scene into a geometric shape usable for various applications, such as simulation or analysis. The three-dimensional representation provides a detailed and manipulable representation of the environment, fundamental for the creation of a high-resolution digital twin.

[0026] Advantageously, the step of generating a digital twin of said physical environment from said three-dimensional representation of said physical environment concludes the process by creating a high-fidelity virtual copy of the physical environment, ready for use in various applications. The digital twin thus created can be used for simulation, planning, predictive maintenance, and other applications.

[0027] If desired, the generation of said three-dimensional representation representing the physical environment can be optimized by using three-dimensional representation simplification techniques to reduce the complexity of the model while retaining the desired details.

[0028] If desired, the application of high-resolution textures to said three-dimensional representation can be carried out in order to improve the visual realism of the digital twin.

[0029] In the context of the present invention, the term "physical environment" means any set of real objects, structures, or scenes capable of being captured by imaging devices such as cameras or smartphones, and subsequently modeled in three dimensions as digital twins. Such a physical environment may include indoor or outdoor spaces, industrial infrastructure, rolling stock, and any other physical setting or element.

[0030] Advantageously, and according to the invention, the environment may include, in particular, the interior of a railway vehicle, specifically all the components of a train car, the exterior of said vehicle, as well as the transition zone between the exterior and interior of said railway vehicle, or even an area of ​​a station or any other object. This specificity makes it possible to capture precise and relevant details for the confined and complex environments of trains, where every element, from the seating arrangement to the safety equipment, plays a crucial role.

[0031] According to this aspect of the invention, the railway environment can include the interior of a railway vehicle. This feature makes it possible to accurately capture and model interior spaces, focusing specifically on the interior of a railway vehicle, thus facilitating predictive maintenance, space optimization, and an improved passenger experience. By modeling the interior, Operators can identify areas requiring repairs or improvements and schedule interventions without interrupting service. The process addresses the specific needs of this sector, such as predictive maintenance, space optimization, and improving the passenger experience.

[0032] The step of acquiring a plurality of images of the physical environment makes it possible to collect comprehensive and continuous visual data of the interior of the railway vehicle. This continuous collection ensures complete and detailed coverage of the environment, capturing all the variations and nuances necessary for accurate modeling. The plurality of images guarantees that even hard-to-reach areas or complex angles are well documented, which is essential for a faithful representation of the train's interior.

[0033] Advantageously, receiving a plurality of images acquired by the cameras allows for the centralization and standardization of visual data. By using cameras to capture the plurality of images, the process ensures uniform image quality and perfect data synchronization. This integration also facilitates subsequent image processing, as all the data originates from a single, consistent source, thus simplifying the processing and reconstruction steps.

[0034] Estimating a plurality of camera poses and extracting depth data from the plurality of received images makes it possible to understand the position and orientation of each camera at each instant, as well as the distance of objects from the camera. This dual estimation is crucial for creating an accurate three-dimensional representation of the physical environment. The depth data adds an additional dimension to the captured images, allowing for the faithful reconstruction of volumes and distances in the physical environment.

[0035] Reconstructing a digital scene of the physical environment from the plurality of received images, estimated camera poses, and extracted depth data makes it possible to create a detailed digital representation. This digital scene serves as the basis for all subsequent modeling and analysis steps. The digital reconstruction allows for immersive visualization and interaction with the environment, thus facilitating the planning, maintenance, and optimization of spaces.

[0036] Advantageously and according to the invention, said plurality of images is acquired by at least one camera, in particular at least one camera of a smartphone.

[0037] According to this aspect of the invention, the plurality of images is acquired by a camera, in particular a smartphone camera. The use of a smartphone camera democratizes access to three-dimensional digital twin generation technology. Since smartphones are widespread and equipped with high-quality cameras, this technical feature It facilitates the capture of images of the physical environment without requiring expensive specialized equipment. Furthermore, smartphone cameras often incorporate advanced technologies such as image stabilization and high-resolution capture, further improving the quality of the data acquired for digital scene reconstruction.

[0038] Furthermore, the use of a smartphone camera simplifies the image acquisition process, making the technology accessible to a wider range of users, including those without advanced technical skills. Users can easily capture multiple images of their physical environment using devices they already own and are familiar with. This lowers barriers to entry and enables broader adoption of digital twin technology.

[0039] Furthermore, smartphones offer optimal portability and flexibility, allowing users to capture images in various environments and conditions. This mobility is particularly useful for applications requiring outdoor capture or capture in hard-to-reach areas. The ability to capture high-quality images on the go improves the accuracy and fidelity of the generated digital twins.

[0040] Finally, the integration of video capture via smartphones allows for real-time updating of digital twins. Users can easily record new image multiples to reflect changes in the physical environment, thus ensuring that the digital twin remains up-to-date and relevant. This technical feature is essential for applications requiring continuous monitoring or frequent updates of the modeled environment.

[0041] Alternatively, the plurality of images is acquired by a camera mounted on a drone, thus enabling the capture of images of the physical environment from different angles and altitudes. In addition, a fixed camera installed within the physical environment is used to provide a continuous plurality of images, ensuring real-time monitoring and constant updating of the digital twin. Another possibility is the use of several synchronized cameras, positioned at various strategic locations, to cover the entire physical environment and obtain a more comprehensive and detailed overview. These cameras are also existing security cameras, integrated into the system to minimize costs and maximize the use of available resources.

[0042] Advantageously and according to the invention, the method of generating a three-dimensional digital twin of a physical environment includes a step of adding noise to at least a part of the images of said plurality of images.

[0043] According to this aspect of the invention, the noise addition step allows controlled variations to be introduced into the images, thereby improving the robustness of the image processing and reconstruction algorithms. Indeed, noise addition facilitates and improves the robustness of matching common points in the images from different viewing angles.

[0044] The estimation of a plurality of camera poses and the extraction of depth data from said plurality of received images are optimized through noise reduction. The pose and depth calculation algorithms become more robust to visual disturbances, which improves the accuracy of the estimations and the extracted data. This results in a better reconstruction of the digital scene.

[0045] Advantageously and according to the invention, the type of noise employed is a Gaussian noise.

[0046] According to this aspect of the invention, the method uses Gaussian noise to improve the quality of the acquired data. Gaussian noise, characterized by its normal distribution, makes it possible to model and filter the random variations present in the images captured by the camera. By applying this type of noise, the system manages to reduce the artifacts and measurement errors that can occur during image acquisition, leading to a more accurate estimation of the camera's poses.

[0047] This results in better quality depth data, which is essential for the accurate reconstruction of the digital scene of the physical environment.

[0048] Advantageously and according to the invention, the method of generating a three-dimensional digital twin of a physical environment comprises an acquisition step, from a plurality of viewpoints of said physical environment, and in that the step of creating a digital scene is carried out from said plurality of viewpoints and said estimated camera poses.

[0049] Advantageously and according to the invention, the step of acquiring a plurality of points of said physical environment is carried out by means of a LiDAR sensor.

[0050] According to this aspect of the invention, the step of acquiring a plurality of points of said physical environment makes it possible to capture high-precision data on distance and depth.

[0051] This point capture method offers superior resolution and accuracy compared to traditional image-based methods, which improves the quality of the depth data used for digital scene reconstruction.

[0052] The step of creating a digital scene from said plurality of viewpoints and said estimated camera poses makes it possible to combine the precise depth information obtained by the LiDAR sensor with the position data and camera orientation. This combination enriches the digital scene by providing a more detailed and faithful representation of the physical environment, which facilitates the generation of a more realistic and accurate three-dimensional model.

[0053] The combined use of points acquired by the LiDAR sensor and estimated camera poses optimizes the digital scene reconstruction process. By integrating these two data sources, the process reduces errors and uncertainties that can occur during reconstruction, leading to better consistency and greater fidelity of the digital model compared to the real physical environment.

[0054] Generating a digital twin from this digital scene enriched by LiDAR data and multiple images makes it possible to create an extremely precise and detailed three-dimensional representation of the physical environment. This increased precision is essential for various applications, such as simulation, planning, predictive maintenance, and analysis, where an accurate understanding of the environment is crucial for informed and effective decision-making. LiDAR acquisition also serves as a reference for evaluating and improving tools based solely on multiple images.

[0055] Advantageously and according to the invention, the process includes a step of generating at least one mesh representing the physical environment which can be optimized using local mesh refinement techniques to increase the overall quality of the mesh without a significant increase in the number of elements.

[0056] According to this aspect of the invention, generating at least one mesh representing the physical environment from the reconstructed digital scene makes it possible to create a detailed geometric structure of the environment. The mesh serves as a skeleton for the digital twin, capturing the shapes and contours of the environment with high precision. This step is essential to ensure that the digital twin is a faithful and usable representation of the environment, enabling accurate analyses and simulations.

[0057] Generating a digital twin of the physical environment from the mesh representing the physical environment makes it possible to create an exact virtual copy of the physical environment. This digital twin can be used for various applications, such as scenario simulation, staff training, maintenance planning, and improving the passenger experience. By having an accurate and detailed digital twin, railway operators can make informed decisions and optimize the use and maintenance of their vehicles.

[0058] Advantageously and according to the invention, the step of generating at least one initial three-dimensional model representing said physical environment includes a sub-step of structural evaluation of said generated initial three-dimensional model.

[0059] According to this aspect of the invention, the step of generating at least one initial three-dimensional model representing the physical environment includes a substep of structural evaluation of said generated initial three-dimensional model. This substep makes it possible to verify the accuracy and consistency of the initial three-dimensional model by identifying and correcting structural anomalies. By proceeding in this way, the initial three-dimensional model obtained faithfully reflects the geometric characteristics of the physical environment, thereby improving the quality and reliability of the final digital twin.

[0060] Structural evaluation of the initial generated three-dimensional model also makes it possible to detect potential errors that could occur during the reconstruction of the digital scene. By identifying these errors at an early stage, it becomes possible to correct them before finalizing the digital twin, thus reducing the risk of failures or inaccuracies in subsequent applications of the digital twin.

[0061] By incorporating this structural evaluation substep, the process ensures that the initial generated three-dimensional model is optimized for high performance in various applications, such as simulation, analysis, and visualization. This guarantees that the digital twin can be used efficiently and accurately in diverse contexts, ranging from urban planning to predictive maintenance.

[0062] Finally, this structural evaluation substep contributes to the robustness of the process by enabling continuous validation of the initial three-dimensional model throughout the digital twin generation process. This ensures better data management and smoother integration of the different process steps, resulting in a high-quality and reliable digital twin.

[0063] Advantageously and according to the invention, the structural evaluation substep is carried out by a similarity measure based on an asymmetric Chamfer distance.

[0064] According to this aspect of the invention, the structural evaluation substep is performed by a similarity measurement based on an asymmetric Chamfer distance. This method makes it possible to efficiently compare the shapes and structures between the digital model and the physical environment. The asymmetric Chamfer distance, by calculating the average distance between the points of two sets of points, offers increased accuracy in evaluating structural correspondences. This allows for the highly accurate detection of differences and similarities between the two representations, which is crucial for the fidelity of the digital twin.

[0065] By using the asymmetric Chamfer distance, the method ensures robust reconstruction through point-to-point comparison of the data, even in the presence of Noise or irregularities in the results from the generation of the digital twin from the images. The ability of the asymmetric Chamfer distance is to evaluate the quality of the reconstruction against a reference and a tolerance threshold.

[0066] The asymmetric Chamfer distance-based similarity measure also offers optimization of computational resources. By focusing on the average distances between points, this method reduces the complexity of the calculations required to assess the correspondence between structures. This results in improved efficiency of the overall process, enabling faster and less expensive generation of digital twins, while maintaining a high level of accuracy.

[0067] Finally, the use of asymmetric Chamfer distance in structural evaluation facilitates the integration of this method into various systems and applications. Its simplicity and efficiency make it a method adaptable to different types of data and camera configurations, rendering the method versatile and applicable to a wide range of use cases. This flexibility is essential to meet the diverse needs of users and industries that rely on the accuracy of digital twins for their operations and analyses.

[0068] Advantageously and according to the invention, for two subsets A and B of the three-dimensional space, said asymmetric Chamfer distance is given by:

[0069] [Math.l] beB œA ~

[0070] According to this aspect of the invention, the distance makes it possible to quantify the dissimilarity between two sets of points, which is crucial for evaluating the accuracy of the digital twin reconstruction relative to the real physical environment. By using this measurement, the process ensures a more faithful match between the digital model and reality, thereby improving the quality and reliability of the generated digital twin.

[0071] Advantageously, the asymmetric Chamfer distance, as a measure of dissimilarity, also facilitates the optimization of reconstruction algorithms. By minimizing this distance, the algorithms can adjust the parameters more efficiently, leading to a more accurate and faster reconstruction of the physical environment. This results in a reduction in the computation time and resources required to generate the digital twin, making the process more efficient.

[0072] Advantageously, the asymmetric Chamfer distance allows us to calculate the offset of set A with respect to set B and also that of set B by relation to set A, which notably allows us to detect the case where set B is a subset of A and vice versa.

[0073] Furthermore, the use of the asymmetric Chamfer distance makes it possible to detect and correct reconstruction errors. When the asymmetric Chamfer distance exceeds a certain threshold, this indicates significant discrepancies between the digital model and the physical environment, which triggers automatic correction mechanisms. These mechanisms adjust the parameters used to estimate camera poses and depth data to improve the accuracy of the reconstructed digital scene.

[0074] Advantageously and according to the invention, the step of generating at least one initial three-dimensional model representing said physical environment includes a sub-step of visually evaluating the texture of said generated initial three-dimensional model.

[0075] According to this aspect of the invention, the step of generating at least one initial three-dimensional model representing the physical environment includes a substep of visually evaluating said generated initial three-dimensional model. This substep makes it possible to verify the quality and accuracy of the initial three-dimensional model, particularly in real time, which ensures that the digital model faithfully reflects the physical environment. By incorporating a visual evaluation, it becomes possible to quickly detect and correct errors or anomalies in the initial three-dimensional model, thus guaranteeing a more accurate and reliable representation of the environment.

[0076] Advantageously and according to the invention, the visual evaluation sub-step is carried out by two measurements of the visual quality of the initial three-dimensional model, based on a PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index Measure) metric.

[0077] The Peak Signal-to-Noise Ratio (PSNR) metric quantifies the quality of the reconstruction by comparing the appearance of the initial generated three-dimensional model with the original images, thus ensuring high visual fidelity. Using this metric, reconstruction errors can be detected and minimized, thereby guaranteeing that the digital twin faithfully reflects the physical environment.

[0078] The SSIM (Structural Similarity Index Measure) metric focuses on the human perception of visual quality, evaluating changes in luminance, contrast, and structure between the appearance of the initial three-dimensional model and the original images. This approach ensures that the digital twin not only matches the raw data but is also visually consistent and pleasing to the human eye. By integrating these two metrics, the process optimizes the accuracy and visual quality of the digital twin, which is crucial for applications requiring high visual fidelity, such as augmented reality or immersive simulation.

[0079] Advantageously and according to the invention, the step of generating an initial three-dimensional model of said physical environment includes a sub-step of cleaning said initial three-dimensional model generated.

[0080] According to this aspect of the invention, the step of generating an initial three-dimensional model of said physical environment includes a substep of cleaning up said generated initial three-dimensional model. This cleaning substep improves the quality and accuracy of the initial three-dimensional model by eliminating artifacts and errors that may occur during the reconstruction of the digital scene. By cleaning the data and correcting anomalies, the resulting initial three-dimensional model more accurately reflects the real physical environment, which is crucial for applications requiring high accuracy, such as simulation, planning, and predictive maintenance.

[0081] The remediation of the initial three-dimensional model includes processes such as surface smoothing, texture correction, and the removal of unwanted noise. These processes ensure that the three-dimensional model is not only visually consistent but also geometrically accurate. This allows the initial three-dimensional model to be used in contexts where data quality is paramount, such as in the fields of architecture, engineering, and construction, where critical decisions depend on the reliability of digital models.

[0082] Advantageously and according to the invention, the step of generating an initial three-dimensional model of said physical environment includes a sub-step of rescaling said initial three-dimensional model generated.

[0083] According to this aspect of the invention, the step of generating an initial three-dimensional model of said physical environment includes a substep of rescaling said generated initial three-dimensional model. This substep ensures that the initial three-dimensional model accurately reflects the actual dimensions of the physical environment. By adjusting the proportions and dimensions of the digital model, a precise match with the physical space is ensured, which is crucial for applications requiring high accuracy, such as architectural planning or scenario simulation.

[0084] Rescaling the initial three-dimensional model also corrects distortions or measurement errors that may occur during data acquisition. By recalibrating the digital model, a more accurate and reliable representation of the physical environment is obtained, thereby improving the quality and usefulness of the initial three-dimensional model for various analyses and decision-making processes.

[0085] Furthermore, this substep facilitates the integration of the initial three-dimensional model with other systems or databases that use specific units of measurement. By standardizing the dimensions of the digital model, the processes of comparison, data merging, and interoperability between different tools and platforms are simplified, thereby optimizing the efficiency of workflows and collaborative projects.

[0086] Advantageously and according to the invention, the step of generating an initial three-dimensional model of said physical environment includes a sub-step of aligning reference frames.

[0087] According to this aspect of the invention, the step of generating an initial three-dimensional model of said physical environment includes a substep of coordinate system alignment. This substep ensures that the various acquired and reconstructed data are consistent with each other. By aligning the coordinate systems, a precise correspondence is ensured between the objects in the physical world and their digital representation. This avoids positioning and dimensioning errors that could occur if the coordinate systems were not correctly aligned. Furthermore, this consistency facilitates the integration of new data or updates into the initial three-dimensional model, since the new information can be directly mapped onto the existing coordinate system without requiring recalibration or complex transformation.Consequently, the updated initial three-dimensional model obtained is more accurate and reliable, which is crucial for applications requiring high fidelity of the digital representation relative to the real physical environment.

[0088] Alternatively, the step of generating an initial three-dimensional model of said physical environment includes a substep of coordinate system alignment, which ensures a precise correspondence between the coordinates of the digital model and those of the real physical environment. This substep may include the use of reference markers placed in the physical environment, which are detected and used to adjust the camera pose data and depth data. Alternatively, algorithms for recognizing specific shapes or features of the environment may be used to perform this alignment. Furthermore, registration techniques based on known control points or pre-existing models of the environment may be integrated to improve the accuracy of the alignment.

[0089] Advantageously and according to the invention, the step of generating an initial three-dimensional model of said physical environment includes a sub-step of aligning points by an iterative closest point method, known in English as "Iterative Closest Points" (ICP).

[0090] According to this aspect of the invention, the step of generating an initial three-dimensional model of said physical environment comprises a sub-step of point alignment by an iterative near-point (ICP) method. This substep refines the accuracy of the numerical model by aligning the points more precisely. Using an iterative method, the process progressively adjusts the points until they correspond optimally, thus improving the fidelity of the initial three-dimensional model to the real physical environment.

[0091] Advantageously and according to the invention, the step of generating an initial three-dimensional model of said physical environment includes a sub-step of verifying the consistency of the viewpoints from a plurality of camera poses.

[0092] According to this aspect of the invention, the step of generating an initial three-dimensional model of said physical environment includes a substep of verifying the consistency of viewpoints from a plurality of camera poses. By proceeding in this way, it is possible to ensure that the different images captured by the camera are correctly aligned with each other, thus guaranteeing a faithful and accurate representation of the physical environment. By verifying the consistency of viewpoints, the method minimizes reconstruction errors and improves the quality of the initial three-dimensional model.

[0093] Verifying the consistency of viewpoints from camera poses also makes it possible to detect and correct any inconsistencies or anomalies in the extracted depth data. This ensures that the depth information used for reconstructing the digital scene is reliable and accurate, which is crucial for obtaining an accurate initial three-dimensional model of the physical environment. Consequently, the generated initial three-dimensional model faithfully reflects the dimensions and characteristics of the real environment.

[0094] By incorporating this substep, the method optimizes the reconstruction process by reducing the need for manual recalculations or adjustments. This makes the process more efficient and less prone to human error, while also enabling robust automation. Verifying the consistency of viewpoints thus contributes to a faster and more accurate generation of the initial three-dimensional model, facilitating its use in various applications such as simulation, planning, and predictive maintenance.

[0095] Finally, this verification substep improves the robustness of the method to variations in capture conditions, such as changes in brightness or unexpected camera movements. By ensuring that the camera poses are consistent, the method maintains high reconstruction quality even in dynamic or complex environments.

[0096] The invention also relates to a system for generating a three-dimensional digital twin of a physical environment comprising:

[0097] - a camera configured to acquire a plurality of images comprising a plurality of images of the physical environment;

[0098] - a computing unit, comprising a processor, configured to receive the plurality from camera images, estimate a plurality of camera poses and extract depth data from the plurality of images and the plurality of camera poses, create a digital scene of the physical environment from the plurality of received images, the estimated poses of each camera and the extracted depth data, generate at least one initial three-dimensional model representing the physical environment from the digital scene;

[0099] - a storage memory connected to the computing unit, configured to store the plurality of images, depth data, camera poses, the reconstructed digital scene, and the initial three-dimensional model;

[0100] - a digital twin generation module integrated into the computing unit, configured to generate a digital twin of the physical environment from the plurality of acquired images and the initial three-dimensional model.

[0101] According to this aspect of the invention, the camera, configured to acquire a plurality of images comprising a plurality of images of the physical environment, makes it possible to capture detailed and continuous visual data of the environment. This real-time acquisition ensures an accurate and dynamic representation of the physical space, thus facilitating the creation of a faithful digital model.

[0102] The computing unit, comprising a processor, receives the plurality of images from the camera, estimates a plurality of camera poses, and extracts depth data from the plurality of images. This computing unit processes the visual information to determine the camera's position and orientation at different times, which is essential for an accurate three-dimensional reconstruction of the environment. The extraction of depth data enriches the visual information with distance measurements, enabling more detailed and realistic modeling.

[0103] The storage memory connected to the computing unit stores the plurality of images, depth data, camera poses, the reconstructed digital scene, and the initial three-dimensional model. This storage capacity ensures the preservation and accessibility of the data required at each stage of the digital twin generation process. By centralizing this information, the storage memory facilitates data management and manipulation for subsequent analyses or updates to the digital model.

[0104] The digital twin generation module integrated into the computing unit generates a digital twin of the physical environment from the initial three-dimensional model. This module transforms the initial three-dimensional model into a representation A complete and interactive digital model of the physical environment. This transformation allows for the creation of a virtual model that can be used for various applications, such as simulation, planning, and visualization, thus providing a deeper understanding and improved interaction with the physical environment.

[0105] Alternatively, the system for generating a three-dimensional digital twin of a physical environment may include alternative embodiments for some of its features. For example, the camera configured to acquire a plurality of images may be replaced by an array of cameras arranged at different angles to capture varied perspectives of the physical environment, thereby improving the accuracy of the depth data and camera poses. The computing unit, comprising a processor, may be supplemented by a graphics processing unit (GPU) to accelerate the calculations required for estimating camera poses and extracting depth data. Furthermore, the storage memory connected to the computing unit may be extended to include cloud storage solutions, thus enabling greater storage capacity and increased data accessibility.The digital twin generation module integrated into the computing unit can also be configured to use Artificial Intelligence algorithms to improve the accuracy and fidelity of the generated digital twin.

[0106] A module may, for example, consist of a computing device such as a computer, a set of computing devices, an electronic component or a set of electronic components, or, for example, a computer program, a set of computer programs, a computer program library or a computer program function executed by a computing device such as a computer, a set of computing devices, an electronic component or a set of electronic components. List of figures

[0107] Other objects, features and advantages of the invention will become apparent from the following description, given by way of non-limiting example only, and with reference to the accompanying figures, in which:

[0108] [Fig.1] is a schematic view of a block diagram of a method for generating a three-dimensional digital twin of a physical environment according to an embodiment of the invention.

[0109] [Fig.2] is a schematic view of a physical environment and a trajectory of a camera used for acquiring images of said physical environment according to an embodiment of the invention.

[0110] [Fig.3] is a schematic view of an image acquired by a camera of a smartphone of a physical environment according to an embodiment of the invention.

[0111] [Fig.4] is a schematic view of a depth image of a physical environment according to an embodiment of the invention.

[0112] [Fig.5] is a schematic view of a three-dimensional digital scene reconstructed according to an embodiment of the invention.

[0113] [Fig.6] is a schematic view of a three-dimensional digital twin obtained by the process of [Fig.1].

[0114] [Fig.7] is a schematic view of a system for generating a three-dimensional digital twin of a physical environment according to an embodiment of the invention.

[0115] A detailed description of an embodiment of the invention

[0116] In the figures, scales and proportions are not strictly respected, for the purposes of illustration and clarity.

[0117] In addition, identical, similar or analogous elements are designated by the same references in all figures.

[0118] A block diagram of a method for generating three-dimensional digital twins 3.4 of a physical environment 1 is shown in [Fig. 1]. The method begins with a step E0 of acquiring a plurality of images 3 of the physical environment 1 using a camera 2 moving along a spatial trajectory 2.2 so as to obtain different views of the physical environment 1 from different perspectives. This acquisition step aims to capture the entirety of the environment 1 and to record the images for subsequent processing to generate the digital twin.

[0119] Following step E0 of acquiring a plurality of images 3, the process continues with a step PI of adding noise to at least a portion of the images 3.2 of the acquired plurality of images 3. The noise addition method is Gaussian and is used to improve the robustness of subsequent image processing.

[0120] In parallel with step PI, the process includes a step El for receiving the plurality of acquired images 3. This step centralizes the visual data of the environment, acquired in the form of images 3.2, thus allowing said images to be prepared for the subsequent processing necessary for the creation of the digital twin of the physical environment 1.

[0121] Following step E1, the process continues with a step E2 of estimating a plurality of camera 2 poses 2.1 from said plurality of images and of extracting depth data from the plurality of images and said plurality of received poses 3. To do this, the depth data extraction is achieved by applying an algorithm from the family of algorithms known as "Structure from Motion", in particular the COLMAP tool.

[0122] In parallel with step E2, the method advantageously includes a step L1 for acquiring a plurality of points of the physical environment 1 by means of a LiDAR sensor. These points, combined with the estimated camera poses 2.2 and the depth information determined in step E2, make it possible to improve the accuracy and fidelity of the digital twin 3.4 of the physical environment 1.

[0123] Following step E2, the process continues with a step E3 of reconstruction of a digital scene of the physical environment 1. This reconstruction is carried out from said plurality of received images 3, the estimated camera poses and the depth data extracted in step E2. An algorithm, based on Neural Radiance Fields (NeRF), is used to reconstruct the scene.

[0124] Concurrent with step E3, the process performs a step M1 of generating a plurality of meshes representing at least a part of the physical environment 1, from the reconstructed digital scene.

[0125] Following step E3, the process continues with step E4, which involves generating and optimizing an initial three-dimensional model of the physical environment from the generated mesh. This step comprises five sub-steps:

[0126] a remediation substep E4.1 of the initial three-dimensional model to correct imperfections;

[0127] a substep E4.2 of rescaling the initial three-dimensional model to ensure dimensional accuracy;

[0128] a substep E4.3 of alignment of the reference frames to ensure spatial consistency. The determination of the alignment proceeds by calculating the barycenter of each reconstructed model, then a transformation matrix, resulting from the composition of a spatial rotation and a translation, is determined so that the first reference frame is transformed into the second reference frame by application of said transformation matrix;

[0129] a substep E4.4 of aligning points by an iterative method of nearby points, strengthening the spatial correspondence. This substep begins by identifying a correspondence set between a reference point cloud and a reconstructed point cloud using a matrix, called the transformation matrix. Subsequently, this transformation matrix is ​​updated from the solutions of a cost function optimization problem established from said correspondence set; and finally

[0130] a sub-step E4.5 of verifying the consistency of viewpoints, carried out on the basis of camera poses, thus ensuring a precise and consistent digital representation of the physical environment.

[0131] Following step E4, the process continues with a step SI for evaluating the quality of the initial three-dimensional model. This step comprises two substeps that can be executed in parallel. Step SI thus comprises:

[0132] a substep Sl.l of structural evaluation of the generated mesh, thus verifying the geometric conformity of the mesh, in particular through checks of the Jacobian ratio, the aspect ratio, the co4rbure tensors, as well as other structural quality metrics, and

[0133] a substep SI.2 of visual evaluation of the mesh carried out by a measurement of the visual quality of the mesh based on a PSNR (Peak Signal-to-Noise Ratio) and SSIM (Structural Similarity Index) metric, which allows validation of the visual appearance and the accuracy of the reconstruction.

[0134] The results of these structural and visual evaluation steps allow for the determination of a score which, in turn, serves as a quality test of the initial three-dimensional model for its use as the basis for creating the digital twin. When the result of the T-test is below a predetermined threshold, called the acceptance threshold, the initial three-dimensional model is returned to step E4 for further evaluation of substeps E4.1 to E4.5, until the T-criterion is above said acceptance threshold.

[0135] The process concludes with a step E5 of generating a digital twin from said initial three-dimensional model.

[0136] Thus, the process results in the creation of a faithful and detailed three-dimensional digital twin, optimally representing the initially captured physical environment.

[0137] A schematic view of a physical environment 1 and the path followed by a camera 2 of a smartphone used for acquiring images of said physical environment, corresponding to step E0 of the process in [Fig.1], is shown in [Fig.2].

[0138] A schematic view of an image acquired by a camera 2 of a smartphone of a physical environment 1, resulting from step El of the process of [Fig.1], is shown in [Fig.3].

[0139] A schematic view of a depth image of a physical environment, obtained during step E2 of the process of [Fig.1], is shown in [Fig.4]. In this Figure, the color scale encodes the depth of the objects in the image scene.

[0140] A schematic view of a reconstructed three-dimensional digital scene, generated during step E3 of the process in [Fig.1], is shown in [Fig.5].

[0141] A schematic view of a three-dimensional digital twin obtained by the process described from [Fig.1] is shown in [Fig.6], corresponding to the final step E5.

[0142] A system for generating three-dimensional digital twins of a physical environment is shown in [Fig. 7]. This generation system includes a camera 2 configured to acquire a plurality of images 3 consisting of a plurality of images of the physical environment 1. This environment is illustrated as the interior of a railway vehicle. The camera is intended to capture the visual data necessary for creating the digital twin.

[0143] The acquired plurality of images 3 is then transmitted to a computing unit 4, which includes a processor. This computing unit is shown in the drawing and is configured to receive the plurality of images, estimate a plurality of camera poses, extract depth data from the plurality of images 3 and the plurality of camera poses, and create a digital scene 3.3 of the physical environment 1 from the received plurality of images, the estimated camera poses, and the extracted depth data. From this digital scene, the computing unit generates an initial three-dimensional model and a plurality of meshes representing the physical environment.

[0144] A storage memory 5 is also shown in the drawing, and is connected to the computing unit 4. This memory is configured to store the various critical elements of the process, including the plurality of images, depth data, camera poses, the reconstructed digital scene, and the generated mesh.

[0145] Finally, a digital twin generation module is integrated into the computing unit 4, also shown in the drawing. This module is responsible for generating the three-dimensional digital twin 3.4 of the physical environment 1, which is then displayed on a device screen 6. The digital twin thus created is a faithful and accurate representation of the physical environment initially captured by the camera 2.

Claims

Demands

1. A method for generating a three-dimensional digital twin (3.4) of a physical environment (1) characterized in that it comprises the following steps: a step (EO) of acquiring by at least one camera a plurality of images (3) of said physical environment (1); a step (E1) of receiving said plurality of images (3) acquired; a step (E2) of estimating a plurality of poses from each camera (2), called camera poses, from said plurality of images and of extracting depth data from said plurality of images (3) and said plurality of poses received; a step (E3) of reconstructing a digital scene (3.3) said physical environment (1) from said plurality of images (3) received, said estimated camera poses and said extracted depth data; a step (E4) of generating at least one three-dimensional model, said initial three-dimensional model, representing said physical environment (1) from said reconstructed digital scene (3.3); a step (E5) of generating a digital twin (3.4) of said physical environment (1) from said three-dimensional representation representing said physical environment (1).

2. A method according to claim 1, characterized in that the physical environment (1) is an interior environment of a railway vehicle.

3. A method according to any one of claims 1 to 2, characterized in that it comprises a step (PI) of adding noise to at least a part of the images of said plurality of images (3).

4. Method according to claim 3, characterized in that the type of noise used is a Gaussian noise.

5. A method according to any one of claims 1 to 4, characterized in that it comprises a step (L1) of acquiring, by means of a LiDAR sensor, a plurality of points of said physical environment (1).

6. A method according to any one of claims 1 to 5, characterized in that the step (E3) of reconstructing a digital scene is executed by an algorithm based on neural architectures of the “Neural Radiance Fields” (NeRF) type.

7. A method according to any one of claims 1 to 6, characterized in that the step (E4) of generating and optimizing an initial three-dimensional model of said physical environment (1) comprises the following substeps: a substep (E4.1) of cleaning said initial generated three-dimensional model; a substep (E4.2) of rescaling said initial generated three-dimensional model; a substep (E4.3) of aligning reference frames; a substep (E4.4) of aligning points by an iterative method of nearby points, and a substep (E4.5) of verifying the consistency of viewpoints from a plurality of camera poses (2).

8. A method according to any one of claims 1 to 7, characterized in that it comprises a step (SI) of evaluating the quality of said initial three-dimensional model.

9. Method according to claim 8, characterized in that the step (SI) of evaluating the quality of said initial three-dimensional model comprises, on the one hand, a substep (S 1.1) of structural evaluation of said generated initial three-dimensional model and, on the other hand, a substep (S 1.2) of visual evaluation of said initial three-dimensional model.

10. A system for generating a three-dimensional digital twin (3.4) of a physical environment (1), characterized in that it comprises: at least one camera (2) configured to acquire a plurality of images (3) comprising a plurality of images (3.1) of the physical environment (1); a computing unit (4), comprising a processor, configured to receive the plurality of images (3) from each camera (2), estimate a plurality of poses from each camera and extract depth data from the plurality of images (3) and the plurality of camera poses, create a digital scene (3.3) of the physical environment (1) from the plurality of received images, the estimated poses from each camera and the extracted depth data, generate at least one initial three-dimensional model representing the physical environment (1) from the digital scene (3.3) ; a storage memory (5) connected to the computing unit (4), configured to store the plurality of images, the data of. depth, camera poses, the reconstructed digital scene, and the initial three-dimensional model; a digital twin generation module integrated into the computing unit (4), configured to generate a digital twin (3.4) of the physical environment (1) from the initial three-dimensional model.