A system and method for efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment.

JP2026529652APending Publication Date: 2026-09-01SIEMENS AG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026509182
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-08-14
Filing Date
2024-08-12
Publication Date
2026-09-01

Smart Images

  • Figure 2026529652000001_ABST
    Figure 2026529652000001_ABST
Patent Text Reader

Abstract

A system, apparatus, and method for the efficient rendering of a scene for a user interacting in a computer simulation environment is disclosed. The method includes the steps of: receiving visual data from a data acquisition device configured to collect data about entities interacting within a facility using a processing unit; identifying a scene to be rendered in a computer simulation environment for a specific user and entities within it; generating one or more clusters of each identified entity from the scene to be rendered; assigning each cluster in the scene to one computing device from a plurality of computing devices installed in the facility based on a device assignment model; and synchronizing each cluster of the scene received from each of the plurality of computing devices in order to render the scene for a user interacting in a computer simulation environment.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates generally to computer simulation environments, and more specifically to methods and systems for efficient rendering of one or more scenes to one or more users interacting in a computer simulation environment using distributed computing devices.

[0002] Background of the Invention An industrial environment includes a plurality of machines or assets within an automated factory, or interacting IoT devices. Accordingly, industrial environments often include a plurality of interconnected components that communicate signals with each other, either directly or via a network. One emerging concept complementing rapid industrial development is the "industrial metaverse". The industrial metaverse is a next-generation fully immersive three-dimensional collaborative space that integrates multiple technical directions, including digital twins, Internet of Things, Industrial Internet, augmented reality (AR), virtual reality (VR), mixed reality (MR), and the like. A metaverse is a virtual universe with a shared 3D virtual space where virtual assets can be owned, placed, and interacted with. Different users can also interact with each other in a collaborative environment. These virtual assets can be as simple as chairs and tables, or as complex as industrial machinery.

[0003] For this purpose, a typical IIoT (Industrial Internet of Things) solution in the metaverse would involve capturing real-world data and rendering it photorealistically within the same industrial environment to provide users with an immersive experience. Within industrial environments such as manufacturing plants and factory floors, there are numerous machines and corresponding standard operating procedures, which operators / workers must follow for optimal functioning of the industrial environment. In such cases, metaverse simulation can be used to schedule and train workers / operators using photorealistic rendering of scenarios, and to immerse those workers / operators in a virtual world to create a real-world experience. For example, a manufacturing plant could host a digital workflow for repairing machinery in a VR space. Employees log into the virtual space via VR, meet and communicate in a shared space via 3D avatars, and repair machinery together.

[0004] However, reproducing real-world objects and their movements in a virtual environment in real time requires sufficient computing power to process and infer information along with a wealth of resource data. To avoid missing critical decision-making opportunities, changes in the real environment should be easily visible in the virtual environment in near real-time. In current scenarios, the solution to the above problem is to use high-speed networks (5G / 6G) and high-performance GPU servers. However, it should be understood that such high-speed networks and high-performance GPU servers consume resources and energy intensively, thereby increasing the carbon dioxide emissions of such virtual environments.

[0005] An efficient, seamless, and near real-time metaverse would allow for the recreation of real-world events within AR / VR spaces. Processing these events on a single device with limited computing power (such as standalone VR hardware or IoT devices available in a factory setting) is difficult because the processes of object recognition, tracking, and analysis are not scalable. Overcoming this requires dedicated computing devices (GPU servers) stations. Furthermore, such applications would eventually exceed the computing and storage (main memory) capacity of the device, especially as the number of entities in the scene increases.

[0006] The deployment of centralized management systems presents challenges such as single points of failure and performance degradation, which cause lag, delays, and inconsistencies in the animation of the environment and its objects. When the continuous movement of objects is disrupted, the animation begins to appear fragmented. This significantly degrades the quality of the experience and confuses the observer. While using cloud computing is a promising solution to this problem, cloud computing itself has a set of challenges.

[0007] Furthermore, even with such state-of-the-art solutions, lag, delays, and inconsistencies are observed when rendering objects within the metaverse. Moreover, such delays and lags become even more pronounced when multiple participants are involved in the scene to be rendered. It should be noted that this problem becomes more significant if network interruptions occur during transmission. Another challenge arises when multiple collaborators are present in the same virtual environment. Therefore, if the continuous movement of objects is disrupted, the animation begins to appear fragmented, which significantly degrades the user experience quality in the virtual environment.

[0008] Further prior art references include the paper “VR immersive social user interaction under metaverse” by Zhu, Yixiao, Xin Wang, and XiangWei Qi et al., Second International Conference on Electronic Information Engineering, Big Data, and Computer Technology (EIBDCT 2023). Vol. 12642. SPIE, 2023.

[0009] Further prior art can be found in the document “Design and Implementation of Distributed Rendering System.” 2022 IEEE Smartworld, Ubiquitous Intelligence & Computing, Scalable Computing & Communications, Digital Twin, Privacy Computing, Metaverse, Autonomous & Trusted Vehicles (SmartWorld / UIC / ScalCom / DigitalTwin / PriComp / Meta). IEEE, 2022. by Liu, Dan et al.

[0010] Further prior art can be found in "A Layered Architecture Enabling Metaverse Applications in Smart Manufacturing Environments" by Bujari, Armir et al., 2023 IEEE International Conference on Metaverse Computing, Networking and Applications (MetaCom). IEEE, 2023.

[0011] From the above perspective, there is a need to provide a system and method for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment using distributed computing devices.

[0012] Summary of the Invention The object of the present invention is to provide a system and method for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment.

[0013] As used herein, the term “facility” refers to any industrial or non-industrial environment that includes one or more assets interacting within it. The term “facility” may refer to any building that may house various entities, such as factories, manufacturing facilities of any kind, or office buildings. In this specification, the term “factory” or “industrial environment” refers to any machine or part thereof that cooperates to enable any production process of any kind to be carried out. Examples of factories include any industrial facility with multiple assets, such as power plants, wind farms, power grids, manufacturing facilities, and processing plants. While in general, the term “facility” in this disclosure is described as some form of factory site, it should be broadly understood that as used in preferred embodiments and claims of this disclosure, the term “facility” includes not only factory sites but also other indoor facilities and buildings such as hospitals, office spaces, apartments, complexes, schools, and training centers. The term “facility” may also include other outdoor facilities such as parking IoTs and traffic intersections, as well as vehicles such as airplanes, ships, and trucks.

[0014] The term “one or more entities” refers to specific objects, elements, components, or subjects within a facility. For example, if the facility is an industrial environment, one or more entities may be assets, devices, machinery, robots, equipment, assembly lines, conveyors, motors, pumps, compressors, or other mechanical, electrical, or electronic equipment, workers, operators, supervisors, etc. For example, if the facility is an office building, one or more entities may be desks, chairs, tables, cabinets, shelves, partitions, workstations, conference room furniture, lighting fixtures, or any other furniture or equipment typically found in an office environment. Furthermore, one or more entities may be equipment or systems used in office operations (including computers, printers, scanners, photocopiers, telephones, projectors, audiovisual systems, network devices), or any other technical or electronic equipment commonly used in an office setting. In another example, the facility is an aircraft under maintenance, and one or more entities may be the physical framework, airframe, or fuselage of the aircraft (including wings, tail section, landing gear, engine nacelles, cockpit, passenger cabin, doors, windows, and any other structural components that contribute to the overall shape and integrity of the aircraft), jet engines, turboprops, propellers, fuel systems, exhaust systems, thrust reversers, or any other elements related to the generation and control of the aircraft's thrust, electronic systems and instruments used for the navigation, communication, monitoring, and control of the aircraft (including flight control systems, flight management systems, autopilot systems, navigation systems, communication systems, radar systems, or any other electronic devices or subsystems installed on the aircraft), seating arrangements, overhead storage, lavatories, galley equipment, lighting systems, entertainment systems, safety equipment, passengers, pilots, crew, etc.

[0015] Throughout this disclosure, the term “computer simulation environment” as used herein refers to a three-dimensional (3D) representation of the real or physical world. It can be understood as a virtual world. The computer simulation environment is accessible to users, i.e., accessible from the real / physical world. This includes data exchange between the computer simulation environment and the real / physical world. In particular, the computer simulation environment can be understood as a “metaverse.” It is also possible to interact with the computer simulation environment, i.e., to influence or use processes, components, and / or functions within the computer simulation environment. Thus, processes within the computer simulation environment can directly influence processes in the real / physical world, for example, through the virtual modeling of control processes.

[0016] For example, a user may be able to access a computer simulation environment through an interface such as a virtual reality (VR) interface or an augmented reality (AR) interface. The corresponding part of the computer simulation environment does not necessarily have to exist in reality, but may be, for example, a 3D model. Also, physical forces and phenomena, such as gravity, may be represented in a different way in the computer simulation environment than in the real world (e.g., gravitational acceleration). For the purposes of this invention, the metaverse consists of one or more animated scenes rendered corresponding to multiple entities interacting in an industrial environment.

[0017] A metaverse can contain multiple computer simulation components. A computer simulation component can be understood, for example, as a representation of a real-world or physical component, particularly a 3D representation. A component may be, for example, a room, a building, an item, or an object. A computer simulation component may have different functions / features, such as an access interface. A computer simulation component may further contain specific data of the component, such as sensor data of a virtual sensor that can be collected via the access interface. Access to a computer simulation component may include, for example, use, modification, or connection to other computer simulation components. A computer simulation component can interact with a computer simulation environment. For the purposes of the present invention, a computer simulation component may be one or more entities rendered in a computer simulation co-environment or metaverse.

[0018] The metaverse can be implemented through a hosting environment. This hosting environment can be implemented, for example, as a cloud environment, an edge cloud environment, and / or on a specific device (e.g., a mobile device).

[0019] Throughout this disclosure, the term “one or more data acquisition devices” refers to any electronic devices configured to capture, collect, and transmit data from an industrial environment to a device. In one embodiment, the data acquisition device may be at least one of a camera, an image sensor, a scanner, a digital camera, a surveillance camera, a 3D imaging device, or a combination thereof. In particular, the primary function of the data acquisition device is to collect visual data from an industrial environment, such as a factory floor, and to transmit the collected data to a device for further processing. In one embodiment, the visual data includes at least one of still image data and video data.

[0020] The object of the present invention is achieved by a method for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment. The method includes a step of a processing unit receiving visual data from one or more data acquisition devices configured to collect data about one or more entities interacting within a facility, where the visual data is collected from the viewpoint of one or more users interacting in the computer simulation environment. The method includes a step of the processing unit identifying one or more scenes to be rendered in the computer simulation environment for a particular user, and one or more entities within them. The method includes a step of the processing unit generating one or more clusters for each of the one or more entities identified from the one or more scenes to be rendered, where each cluster includes a set of entities relating to one another. The method includes a step of the processing unit assigning each cluster in the one or more scenes to one computing device from a plurality of computing devices installed within the facility, based on a device assignment model, where each of the plurality of computing devices is configured to process one or more clusters using one or more machine learning models. This method includes the step of having a processing unit synchronize each cluster of one or more scenes received from each of several computing devices in order to render one or more scenes for one or more users interacting in a computer simulation environment.

[0021] In one embodiment, the method further includes the step of generating an animation from synchronized visual data for a scene to be rendered for a specific user. Furthermore, the method includes the step of rendering the generated animation for the specific user in a computer simulation environment.

[0022] In one embodiment, a method for synchronizing a cluster of one or more scenes to be rendered includes a processing unit determining the relationship between each scene of a cluster transmitted from each of a plurality of computing devices based on the timestamp of each scene. Furthermore, the method for synchronizing a cluster of one or more scenes to be rendered includes a processing unit positioning each cluster in its respective scene based on the coordinates of one or more entities in each cluster and the determined scene relationship. Furthermore, the method for synchronizing a cluster of one or more scenes to be rendered includes a processing unit synchronizing one or more scenes to be rendered based on the timestamp of each scene, where one or more scenes include one or more associated cluster arrangements.

[0023] In one embodiment, a method for synchronizing clusters of one or more scenes received from each of several computing devices in order to render one or more scenes for one or more users interacting in a computer simulation environment includes a processing unit generating metadata for each cluster of one or more scenes processed by the computing devices, wherein the metadata includes one or more parameters that define visual data relating to one or more entities within one or more clusters. Furthermore, the method includes the processing unit synchronizing the metadata received from each computing device based on the arrival time and service rate of each metadata received from the computing device.

[0024] In one embodiment, a method for generating clusters for each of one or more entities identified from a scene to be rendered includes a processing unit generating bounding boxes for a set of entities from one or more entities, where the sets of entities interact with each other. Furthermore, the method includes the processing unit determining the superposition score of the bounding boxes generated over multiple frames, where the superposition score is determined based on a comparison of the superposition area between one or more bounding boxes and the joint area between one or more bounding boxes. Furthermore, the method includes the processing unit determining the frame relating to the set of entities with the highest superposition score. Furthermore, the method includes the processing unit generating a first set of association graphs for the set of entities with the highest superposition score. Furthermore, the method includes the processing unit generating one or more clusters based on the generated first set of association graphs.

[0025] In one embodiment, a method for generating clusters for each of one or more entities identified from a scene to be rendered includes a processing unit generating bounding boxes for a set of entities from one or more entities, where the sets of entities do not interact with each other. Furthermore, the method includes the processing unit calculating relative distance scores between each of the one or more bounding boxes over a number of frames, where the relative distance score is the distance between the sets of entities in each bounding box. Furthermore, the method includes the processing unit determining the frame relating to the set of entities having the minimum relative distance score. Furthermore, the method includes the processing unit generating a second set of association graphs for the set of entities having the minimum relative distance score. Furthermore, the method includes the processing unit generating one or more clusters based on the generated second set of association graphs.

[0026] In one embodiment, a method of allocating each cluster to one computing device from a plurality of computing devices based on a device allocation model comprises identifying, by a processing unit, a plurality of computing devices installed in a facility. Further, the method comprises determining, by the processing unit, a configuration of each of the identified computing devices. Further, the method comprises listing, by the processing unit, each of one or more clusters to be rendered based on a rendering order over a predetermined period. Further, the method comprises determining, by the processing unit, a complexity score for each cluster to be rendered, wherein the complexity score is an indicator of a processing complexity of the one or more clusters. The method comprises determining, by the processing unit, a priority score for each cluster to be rendered, wherein the priority score is determined based on the complexity score and a priority of the rendering order. The method comprises allocating, by the processing unit, the cluster to be rendered to a computing device having the highest computing capability based on the highest priority score of the scene.

[0027] In one embodiment, a method of allocating each cluster to one computing device from a plurality of computing devices based on a device allocation model comprises allocating each set of clusters in a first association graph and a second association graph to computing devices based on a proximity between a user and the computing devices.

[0028] In one embodiment, a method of allocating each cluster to one computing device from a plurality of computing devices based on a device allocation model comprises determining, by a processing unit, one or more new entities in a previously rendered scene in a computer simulation environment for a specific user when the user initiates interaction with the one or more new entities in the computer simulation environment. The method comprises allocating, by the processing unit, the one or more new entities to a computing device allocated for a previously rendered scene for the user.

[0029] In one embodiment, a method of assigning each cluster to one computing device from among a plurality of computing devices based on an assignment model comprises determining, by a processing unit, one or more overlapping entities from respective viewpoints of a plurality of collaborating users in a computer simulation environment. The method comprises assigning, by the processing unit, the determined overlapping entities to a specific computing device for processing visual data for all of the plurality of users.

[0030] In one embodiment, a method of assigning each cluster to one computing device from among a plurality of computing devices based on an assignment model comprises detecting, by a processing unit, a failure of a computing device during processing of a cluster assigned for processing. The method comprises determining, by the processing unit, a progress status of processing completion of the cluster assigned for processing. The method comprises assigning, by the processing unit, another computing device that is capable of processing the cluster assigned for processing and is available within a facility.

[0031] The object of the present invention is also achieved by an apparatus for efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment. The apparatus comprises one or more processing units and a memory communicatively coupled to the one or more processing units. The memory comprises a module stored in a form of machine-readable instructions executable by the one or more processing units. The module is configured to perform the aforementioned method steps.

[0032] The object of the present invention is also achieved by a system for the efficient rendering of one or more scenes for one or more users in a computer simulation environment. The system includes a computer simulation co-environment that renders one or more scenes corresponding to real-world entities within a facility. The system includes a plurality of computing devices that are communicatively coupled to the computer simulation environment, where the plurality of computing devices are installed within the facility. The device is communicatively coupled to the plurality of computing devices and the computer simulation co-environment. The device is configured to efficiently render one or more scenes for one or more users interacting in the computer simulation environment according to the method steps described above.

[0033] The object of the present invention can also be achieved by a computer program product that includes machine-readable instructions causing one or more processing units to perform the method steps described above, when executed by one or more processing units.

[0034] The object of the present invention is further addressed by a computer-readable medium on which a program code section of a computer program is stored, the program code section being loadable into the system and / or executable within the system so as to cause the system to perform the method steps described above when the program code section is executed within the system. This outline is provided to introduce in a simplified form the concepts which will be described in more detail in the following description. It is not intended to identify any features or essential characteristics of the subject matter described in the claims. Furthermore, the subject matter described in the claims is not limited to embodiments that solve some or all of the defects described in any part of the present invention.

[0035] The present invention will be further described below with reference to embodiments shown and depicted in the accompanying drawings. [Brief explanation of the drawing]

[0036] [Figure 1A] This is a block diagram of a system for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention. [Figure 1B] This is an exemplary depiction of an avatar viewing one or more scenes in a computer simulation environment, according to one embodiment of the present invention. [Figure 2] This is a block diagram of an exemplary apparatus for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention. [Figure 3] This flowchart shows the steps of a method for efficiently rendering one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention. [Figure 4A] This flowchart shows the steps of a method for generating one or more clusters for one or more entities that interact with each other, according to one embodiment of the present invention. [Figure 4B] This is an illustrative depiction of one or more clusters of one or more entities interacting with each other, according to one embodiment of the present invention. [Figure 5A] This flowchart shows the steps of a method for generating one or more clusters of one or more entities that do not interact with each other, according to one embodiment of the present invention / another embodiment of the present invention. [Figure 5B] This is an illustrative depiction of one or more clusters of one or more entities that do not interact with each other, according to one embodiment of the present invention. [Figure 6] This flowchart shows the steps of a method for assigning each cluster to a computing device for processing, according to one embodiment of the present invention. [Figure 7]This flowchart shows a workflow for a method of synchronizing one or more clusters received from a computing device, according to one embodiment of the present invention. [Figure 8] This is a flowchart illustrating a workflow for a method of efficiently rendering one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention.

[0037] Embodiments for carrying out the present invention are described in detail below. Various embodiments are described with reference to the drawings, in which similar reference numerals are used to refer to similar elements. In the following description, numerous specific details are shown for illustrative purposes in order to provide a complete understanding of one or more embodiments. It will be apparent that such embodiments may be practiced without those specific details.

[0038] Detailed description of the embodiment Figure 1A is a block diagram of system 100A for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment. System 100A includes a computer simulation environment 102, one or more entities 104-1 to 104-N rendered within the computer simulation environment 102, one or more data acquisition devices 105, multiple computing devices 108-1 to 108-N, and a device 110 that communicates via a communication network 106. The computer simulation environment 102 is a three-dimensional (3D) representation of the real world or physical world. It can be understood as a virtual world representing one or more entities 104-1 to 104-N from real-world entities within a facility. As used herein, the term “facility” refers to any industrial or non-industrial environment that includes one or more assets interacting therein. The term “facility” can refer to any building that may house various entities, such as a factory or any kind of manufacturing facility or office building. As used herein, the term “factory” or “industrial environment” refers to all or part of machinery that cooperate to enable any kind of production process to be carried out. Examples of factories include any industrial facility with multiple assets, such as power plants, wind farms, power grids, manufacturing facilities, and processing plants. In this disclosure, the term "facility" is generally used to describe any form of factory site, but it should be understood more broadly that the term “facility” as used in the preferred embodiments and claims of this disclosure includes not only factory sites but also other indoor facilities and buildings such as hospitals, office spaces, apartments, complexes, schools, and training centers. The term “facility” may also include other outdoor facilities such as parking IoT systems and traffic intersections, as well as vehicles such as airplanes, ships, and trucks.

[0039] The term “one or more entities” 104-1 to 104-N refers to specific objects, elements, components, or subjects within a facility. For example, if the facility is an industrial environment, one or more entities 104-1 to 104-N may be assets, devices, machinery, robots, equipment, assembly lines, conveyors, motors, pumps, compressors, or other mechanical, electrical, or electronic equipment, workers, operators, supervisors, etc. For example, if the facility is an office building, one or more entities may be 104-1 to 104-N, desks, chairs, tables, cabinets, shelves, partitions, workstations, conference room furniture, lighting fixtures, or any other furniture or equipment typically found in an office environment. Furthermore, one or more entities may be equipment or systems used in office operations (including computers, printers, scanners, photocopiers, telephones, projectors, audiovisual systems, network devices), or any other technical or electronic equipment commonly used in an office setting. In another example, the facility is an aircraft under maintenance, and one or more entities 104-1 to 104-N may be the physical framework, airframe, or fuselage of the aircraft (including wings, tail sections, landing gear, engine nacelles, cockpit, passenger cabin, doors, windows, and any other structural components that contribute to the overall shape and integrity of the aircraft), jet engines, turboprops, propellers, fuel systems, exhaust systems, thrust reversers, or any other elements related to the generation and control of the aircraft's thrust, electronic systems and instruments used for the navigation, communication, monitoring, and control of the aircraft (including flight control systems, flight management systems, autopilot systems, navigation systems, communication systems, radar systems, or any other electronic devices or subsystems installed on the aircraft), seating arrangements, overhead storage spaces, lavatories, galley equipment, lighting systems, entertainment systems, safety equipment, passengers, pilots, crew, etc.

[0040] One or more entities 104-1 to 104-N are captured using one or more data acquisition devices 106 and then rendered on a computer simulation environment 102. The one or more data acquisition devices 106 may be any electronic devices configured to capture and collect data from an industrial environment and transmit it to the apparatus 110. In one embodiment, the data acquisition device may be at least one of a camera, image sensor, scanner, digital camera, surveillance camera, 3D imaging device, or a combination thereof. In particular, the main function of the data acquisition device is to collect visual data from an industrial environment such as a factory floor and to transmit the collected data to the apparatus 110 for further processing. In one embodiment, the visual data includes at least one of still image data and video data.

[0041] The computer simulation environment 102 is accessible to the user, i.e., accessible from the real world / physical world. In particular, the computer simulation environment 102 can be understood as a "metaverse." It is also possible to interact with the computer simulation environment 102, i.e., to influence or use the processes, components, and / or functions within the computer simulation environment 102. The user or avatar can interact with entities rendered within the metaverse.

[0042] For example, a user may access the computer simulation environment 102 via an interface such as a virtual reality (VR) interface or an augmented reality (AR) interface. For the purposes of the present invention, the metaverse consists of one or more animated scenes rendered in relation to multiple entities interacting within the facility. The metaverse may include multiple computer simulation components. Computer simulation components can be understood, for example, as representations of real-world or physical components, particularly 3D representations. Components may be, for example, rooms, buildings, items, or objects. Computer simulation components may have different functions / features, such as access interfaces. The metaverse can be implemented by a hosting environment. This hosting environment can be implemented, for example, as a cloud environment, an edge cloud environment, and / or on a specific device (e.g., a mobile device).

[0043] In one embodiment, the device 110 is deployed in a cloud computing environment. As used herein, “cloud computing environment” refers to a processing environment that includes configurable physical and logical computing resources such as networks, servers, storage, applications, and services, and data distributed over a network 106, such as the Internet. This cloud computing environment provides on-demand network access to a shared pool of configurable physical and logical computing resources. The device 110 may include modules for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment.

[0044] In particular, System 100 includes a cloud computing device configured to efficiently render one or more scenes for one or more users interacting in a computer simulation environment. This cloud computing device includes a cloud communication interface, cloud computing hardware and OS, and a cloud computing platform. The cloud computing hardware and OS may include one or more servers with an operating system (OS) installed and containing one or more processing units, one or more storage devices for storing data, and other peripherals necessary to provide cloud computing functionality. The cloud computing platform is a platform that implements functionalities such as data storage, data analysis, data visualization, and data communication on the cloud hardware and OS via APIs and algorithms, and provides the aforementioned cloud services using cloud-based applications.

[0045] Throughout this disclosure, the term “computation device” 108-1 to 108-N (more specifically, the first computing device 108-1, the second computing device 108-2, and the Nth computing device 108-N) refers, in each case, to a computer (system), client, smartphone, device, or server located within a facility that can perform one or more functions, such as processing visual data using one or more machine learning models. In certain embodiments, a computing device may be a physical device or a virtual device. In many embodiments, a computing device may be any device capable of performing operations, such as a dedicated processor, a part of a processor, a virtual processor, a part of a virtual processor, a part of a virtual device, or a virtual device. In some embodiments, a processor may be a physical processor or a virtual processor. In some embodiments, a virtual processor may correspond to one or more parts of one or more physical processors. In some embodiments, instructions / logic may be distributed and executed across one or more virtual or physical processors for the execution of such instructions / logic. Computation devices may be located within a facility that is rendered in the metaverse. Advantageously, arranging computing devices in a distributed manner ensures that the processing of visual data is allocated among the computing devices depending on the availability and computing power of the edge devices. In a preferred embodiment, these computing devices are configured to run one or more machine learning models for object detection, activity recognition, pose tracking, object tracking, face recognition, and other computer vision techniques.

[0046] Figure 1B is an exemplary depiction 100B of a user viewing one or more scenes in a computer simulation environment according to one embodiment of the present invention. As shown in the figure, user 112 is viewing one or more scenes S1, S2, ... S in the computer simulation environment 102. N These are time points t0, t1 and t respectively. NThe wearable device 114 is configured to visualize the user. In one example, the wearable device 114 includes a display module that renders one or more entities in a realistic and immersive format and presents visual information to the user. This display module may include a high-resolution screen, a holographic display, augmented reality (AR) glasses, or any other suitable technology for visually presenting virtual content to the user 112. Additionally, the wearable device 114 incorporates a tracking system that captures the user's movements and gestures, enabling real-time interaction and movement within the metaverse. The tracking system may utilize sensors, cameras, motion trackers, or any other suitable means for capturing and interpreting the user's movements. Furthermore, the wearable device 114 includes connectivity features that facilitate communication and data exchange with the metaverse infrastructure. These features may include wireless communication capabilities such as Wi-Fi, Bluetooth, and cellular connectivity, enabling the wearable device to connect to the metaverse platform, collect asset data, and transmit user behavior or preferences. The wearable device 114 can also incorporate input mechanisms such as touch-sensitive surfaces, buttons, voice recognition, and motion sensors, enabling the user to provide commands, make selections, or manipulate virtual assets within the metaverse environment. Designed for visualizing assets within the metaverse, the wearable device 114 enhances the user's experience and immersion in the virtual world. It allows the user to seamlessly perceive, interact with, and navigate across virtual assets, objects, and environments, thereby providing a novel and immersive way to explore and visualize digital content within the metaverse.

[0047] Figure 2 is a block diagram of an exemplary apparatus 110 for efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention. In the exemplary embodiment, the apparatus 110 is communicably coupled to a computer simulation environment 102 for rendering one or more entities 104-1 to 104-N, one or more data acquisition devices 105, and one or more computing devices 108-1 to 108-N.

[0048] The device 110 may be a personal computer, laptop computer, tablet, server, virtual machine, etc. The device 110 includes a processing unit 202, memory 204 including module 206, storage unit 217 including database 222, input unit 224, output unit 226, and bus 228.

[0049] As used herein, processing units 202 mean any type of computing circuit, including but not limited to microprocessors, microcontrollers, complex instruction set arithmetic microprocessors, reduced instruction set arithmetic microprocessors, extra-long instruction word type microprocessors, explicit parallel instruction arithmetic microprocessors, graphics processors, digital signal processors, or any other type of processing circuit. Processing units 202 may also include embedded controllers such as general-purpose or programmable logic devices or arrays, application-specific integrated circuits, and single-chip computers.

[0050] Memory 204 may be non-temporary volatile memory and / or non-volatile memory. Memory 204 may be coupled to a computer-readable storage medium, such as a computer-readable storage medium, for communication with the processing unit 202. The processing unit 202 can execute instructions and / or code stored in memory 204. Various computer-readable instructions may be stored in and accessed from memory 204. Memory 204 may include any suitable elements for storing data and machine-readable instructions, such as read-only memory, random-access memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, a hard drive, or a removable media drive for handling compact disks, digital video disks, floppy disks, magnetic tape cartridges, memory cards, etc.

[0051] In this embodiment, the memory 204 includes a module 206 that is stored in the form of machine-readable instructions in any of the above-mentioned storage media, communicates with the processing unit 202, and is executed by the processing unit 202. When the machine-readable instructions are executed by the processing unit 202, the module 206 causes the processing unit 202 to efficiently render one or more scenes for one or more users interacting in a computer simulation environment.

[0052] Module 206 further includes a data acquisition module 208, a scene identification module 210, a cluster generation module 212, a device assignment module 216, a synchronization module 218, and a rendering module 220.

[0053] The data acquisition module 208 is configured to collect visual data from an industrial environment. This data acquisition module 208 is configured to collect data about one or more entities 104-1 to 104-N interacting in the industrial environment from the perspective of a specific user 112. The data acquisition module 208 is configured to collect data in one or more formats that may include at least one of still image data, video data, signal data, and acoustic data. The data acquisition module 208 is configured to preprocess data received from one or more data acquisition devices. The data acquisition module 208 is configured to collect visual data from data acquisition devices such as image sensors, camera modules, digital cameras, surveillance cameras, aerial imaging devices, 3D imaging devices, thermal imaging devices, and wearable cameras.

[0054] The scene identification module 210 is configured to identify one or more real-world scenes from the industrial environment in real time. The scene identification module 210 is configured to identify scenes from visual data received from the data acquisition module 208. The scene identification module 210 is configured to identify one or more scenes from the viewpoint of a particular avatar. It should be understood that the scene identification module is configured to determine one or more scenes over several frames to be rendered from the user's viewpoint at a given instantaneous point in time, rather than the entire industrial environment being processed and rendered simultaneously.

[0055] The cluster generation module 212 is configured to generate one or more clusters for each of one or more entities identified from one or more scenes to be rendered. Each cluster contains a set of entities that relate to one another. The cluster generation module 212 is configured to generate clusters for one or more entities that interact with one another. The cluster generation module 212 is configured to generate clusters for one or more entities that do not interact with one another. The cluster generation module 212 is configured to generate clusters for one or more entities based on an association graph generated based on the relationships between one or more entities.

[0056] The device assignment module 216 is configured to assign each cluster in one or more scenes to one computing device from among multiple computing devices installed within the facility, based on a device assignment model. The device assignment module 216 is configured to identify the configuration of each computing device within the facility. Furthermore, the device assignment module 216 is configured to determine the complexity score of one or more scenes based on the time required for scene processing and the complexity of the scenes. Furthermore, the device assignment module 216 is configured to assign one or more clusters to the computing device based on the configuration of the computing device and the complexity of one or more scenes to be rendered.

[0057] The synchronization module 218 is configured to synchronize each cluster of one or more scenes received from each of multiple computing devices in order to render one or more scenes for one or more scenes interacting in a computer simulation environment. The synchronization module 218 is configured to determine the relationship between each cluster transmitted from the computing device and a particular scene. The synchronization module 218 is configured to position each cluster in its respective scene based on the coordinates of one or more entities in each cluster. Furthermore, the synchronization module 218 is configured to synchronize one or more scenes to be rendered based on the timestamp of each scene.

[0058] The rendering module 220 is configured to generate animations from synchronized visual data relating to a scene to be rendered for a specific user. Furthermore, the rendering module 220 is configured to render the generated animations for that specific user in a computer simulation environment.

[0059] Processing unit 202 is configured to perform all the functionalities of module 206. Processing unit 302 is configured to receive visual data from one or more data acquisition devices configured to collect data about one or more entities interacting within the facility. The visual data is collected from the viewpoint of one or more avatars interacting in the computer simulation environment. Processing unit 304 is configured to identify one or more scenes to be rendered and one or more entities within them, from the viewpoint of a specific user in the computer simulation environment. Processing unit 302 is configured to generate one or more clusters for each of the one or more entities identified from the one or more scenes to be rendered. Each cluster contains a set of entities relating to one another. Processing unit 302 is configured to assign each cluster in one or more scenes to one of several computing devices installed in the facility, based on a device assignment model. Each of the multiple computing devices is configured to process the visual data using one or more machine learning models. Processing unit 302 is configured to synchronize each cluster of one or more scenes received from each of the multiple computing devices in order to render one or more scenes for one or more users interacting in the computer simulation environment.

[0060] The storage unit 318 includes a database 320 for storing a first knowledge graph and a second knowledge graph. The storage unit 318 and / or the database 320 may be provided using various types of storage technologies such as solid-state drives, hard disk drives, and flash memory, and may be stored in various formats such as relational databases, non-relational databases, flat files, spreadsheets, and extended markup files.

[0061] The input unit 322 can provide ports to receive input from input devices such as a keypad, a touch-sensitive display, and a camera (such as a camera that receives gesture-based input), and can receive requirements sets for generating a scene of a specific area within a factory floor, details of a factory floor model to be stored in a first knowledge graph, and details of a worker behavior model to be stored in a second knowledge graph. The display unit 324 can provide ports to output data via an output device having a graphical user interface for displaying one or more scenes in a computer simulation virtual environment. The bus 326 acts as an interconnection between the processing unit 302, memory 304, storage unit 318, input unit 322, and display unit 324.

[0062] Those skilled in the art will recognize that the hardware shown in Figure 3 may vary depending on the specific implementation. For example, other peripheral devices such as optical disc drives, local area network (LAN) / wide area network (WAN) / wireless (e.g., Wi-Fi) adapters, graphics adapters, disk controllers, and input / output (I / O) adapters may be used in addition to or instead of the illustrated hardware. The illustrated examples are provided for illustrative purposes only and do not constitute an implicit architectural limitation of this disclosure.

[0063] Figure 3 is a flowchart showing the steps of Method 300 for efficient rendering of one or more scenes in a computer simulation environment, according to one embodiment of the present invention.

[0064] In step 302, visual data is received from one or more data acquisition devices configured to collect data about one or more entities interacting within the facility. This visual data is collected from the viewpoint of one or more users interacting in a computer simulation environment.

[0065] The data acquisition device 105 is an electronic device configured to capture and collect data from an industrial environment and transmit it to the apparatus 110. In one embodiment, the data acquisition device may be at least one of a camera, an image sensor, a scanner, a digital camera, a surveillance camera, a 3D imaging device, or a combination thereof. In particular, the main function of the data acquisition device is to collect visual data from an industrial environment such as a factory floor and to transmit the collected data to the apparatus 110 for further processing. In one embodiment, the visual data includes at least one of still image data and video data.

[0066] In exemplary embodiments, the data acquisition device 105 is a camera configured to capture still image and / or video data of one or more entities within a facility. The term “camera” is understood to mean anything from a single camera to a group of cameras arranged in an appropriate configuration and outputting one or more video feeds of corresponding coverage areas of an industrial environment. In one example, the data acquisition device may be localized within a facility, for example, cameras may be placed throughout the industrial environment and installed in various locations therein. In another example, the data acquisition device may be integrated into an aerial mobile device, for example, cameras may be mounted on the airframe of a drone and pointed in various directions to cover various areas of the industrial environment, or pointed at specific entities based on a set of requirements.

[0067] In another exemplary embodiment, the data acquisition device 105 is a surveillance camera used for security and monitoring purposes. These surveillance cameras are designed to capture and record images or videos of a specific area or premises. The surveillance cameras may include features such as night vision, motion detection, remote monitoring, and video analytics for intelligent video processing.

[0068] In another exemplary embodiment, the data acquisition device 105 is a 3D imaging device that captures depth information in addition to visual images, enabling the creation of a three-dimensional representation of an object or scene. Such a device may use techniques such as structured light, time-of-flight method, stereo vision, or LiDAR (Light Detection and Ranging) to acquire 3D data.

[0069] In another exemplary embodiment, the data acquisition device 105 is a thermal imaging device that captures infrared radiation emitted from assets in a factory setting and generates images based on temperature differences. In particular, this device would provide additional information for generating high-quality scenes during nighttime or in low-light conditions.

[0070] In another exemplary embodiment, the data acquisition device 105 is a wearable camera that may be placed on the worker / operator's body to enable the system to capture visual data of assets, machinery, or activities being performed from the worker / operator's field of vision on the factory floor. Advantageously, the system can switch between different views to provide a more immersive experience for avatars in the metaverse when rendering one or more scenes.

[0071] It should be understood that the data acquisition device 105 is configured to collect visual data from the viewpoint of a user viewing a specific scene in the metaverse. For example, when a user or avatar in the metaverse is walking down a corridor in the metaverse, the user's viewpoint may be a part / area of ​​a factory site, and when this is detected in the metaverse, the corresponding part / area is collected from the real-world factory site in accordance with the user's movement and changes in the user's viewpoint, and rendered in the metaverse.

[0072] In step 304, for a specific user 112, one or more scenes to be rendered and one or more entities 104-1 to 104-N within them are identified in the computer simulation environment 102. The real-world scenes are identified from the facility based on the viewpoint of the specific user 112. It should be understood that one or more scenes are identified from the viewpoint of user 112 who is viewing specific scenes (S1 to SN) in the metaverse. For example, when user 112 or their avatar is walking down a corridor in the metaverse, the user's viewpoint may be a part / area of ​​the factory site, and when this is detected in the metaverse, one or more scenes from the corresponding part / area are identified and rendered.

[0073] Real-world scenes may be identified in real time from the facility. Real-world scenes are identified using image processing techniques on the received visual data. The collected visual data, such as images and video data, may undergo preprocessing steps to improve image quality or remove noise and artifacts. Preprocessing techniques may include noise reduction, image filtering, color correction, or image enhancement algorithms. Furthermore, this method includes extracting relevant features from the collected images. Features can be various visual descriptors such as color, texture, shape, edges, and keypoints. These features are derived to represent features specific to objects, scenes, or patterns in the image. The extracted features are then used to analyze the scene and identify one or more entities within the scene. Furthermore, entities such as machines, conveyor belts, motors, levers, boxes, carts, robots, sensors, actuators, operators, and workers can be identified using object detection algorithms in the scene. This can be achieved using object recognition algorithms, machine learning techniques, or pattern matching approaches.

[0074] Furthermore, visual data are meaningful regions or segments based on their visual characteristics. This may include classifying each pixel or region into different categories, such as asset areas, fire evacuation areas, or conveyor belt areas. The method further includes using one or more machine learning algorithms for scene classification of the visual data. Labels or categories are assigned to the entire scene based on its overall content. The method includes training a machine learning model to recognize and classify scenes based on predefined categories, such as machine areas, storage areas, packaging areas, or electrical areas. Once one or more entities are identified and analyzed, it should be understood that the method includes recognizing or classifying the scene based on the one or more identified entities within the scene. This may include comparing extracted features to a database of known scenes, or using machine learning algorithms to match scene characteristics to predefined scene categories. As an example, the method includes performing object detection on the scene using a computer vision-based object detection model. This model identifies and locations different entities (objects) present in the scene. For this purpose, object detection algorithms such as YOLO (You Only Look Once), SSD (Single Shot Multibox Detector), or Faster R-CNN (Region-based Convolutional Neural Networks) can be used.

[0075] Step 306 generates one or more clusters for each of the one or more entities 10⁴-1 to 10⁴-N identified from one or more scenes to be rendered. Here, each cluster contains a set of entities 10⁴-1 to 10⁴-N that relate to each other. One or more clusters can be understood as a group of entities identified from one or more scenes. When an object is detected (as described in Step 304), the coordinates or bounding box of each entity in the scene are obtained. These bounding boxes define the position and size of the detected entity within the image or frame. Furthermore, the method includes analyzing the overlapping regions between the bounding boxes of the detected entities. Bounding boxes that significantly overlap or interact with each other are likely to be part of the same cluster or group. To create clusters of entities, a clustering algorithm is applied to group the overlapping bounding boxes. Various clustering algorithms can be used, such as hierarchical clustering, k-means clustering, and DBSCAN (Density-Based Spatial Clustering of Applications with Noise). The clustering algorithm assigns each bounding box to a specific cluster based on its overlap pattern with other bounding boxes. Bounding boxes that do not significantly overlap with any other bounding boxes may be treated as separate entities or may form smaller clusters. After the initial clustering, post-processing steps may be applied to refine the clusters. For example, overlapping clusters may be merged, or noise and outliers may be removed, to ensure accurate and consistent cluster formation. In one embodiment, one or more clusters are generated based on association graphs created for one or more entities identified in one or more scenes. Subsequently, one or more clusters are generated based on the association graphs. Such methods are described in more detail in Figures 4A-4B and Figures 5A-5B.

[0076] In step 308, each cluster within one or more scenes is assigned to one computing device from among several computing devices installed within the facility, based on a device assignment model. Each of the multiple computing devices 108-1 to 108-N is configured to process visual data using one or more machine learning models. Generally, in connection with embodiments of the present invention, “computing device” should be understood to mean a computer (system), client, smartphone, device, or server, etc., that is capable of performing one or more functions, such as processing visual data using one or more machine learning models in each case located within the facility.

[0077] In a given embodiment, computing devices 108-1 to 108-N may be physical or virtual devices. In many embodiments, computing devices may be any device capable of performing calculations, such as a dedicated processor, a part of a processor, a virtual processor, a part of a virtual processor, a part of a virtual device, or a virtual device. In some embodiments, a processor may be a physical or virtual processor. In some embodiments, a virtual processor may correspond to one or more parts of one or more physical processors. In some embodiments, instructions / logic may be distributed and executed across one or more virtual or physical processors to execute the instructions / logic.

[0078] In preferred embodiments, one or more computing devices 108-1 to 108-N are edge devices deployed within the facility. Therefore, in the following specification, the terms “edge device” and “computing device” are used interchangeably. Edge devices are characterized by their ability to perform local data processing and analysis, reducing the need for large-scale data transmission to centralized cloud servers or data centers. Operating at the edge of the network, these devices efficiently process data near the source or consumption of data, minimizing latency, bandwidth utilization, and reliance on remote computing resources. Edge devices may include a variety of electronic devices, such as sensors, gateways, microcontrollers, embedded systems, mobile devices, smart appliances, or any other computing devices deployed within the facility. The invention further encompasses communication modules and protocols that enable seamless integration and interaction with other components of the network infrastructure. These communication capabilities facilitate data exchange, control signals, or other forms of data interaction between edge devices and central systems or other edge devices. Advantageously, the local processing capabilities of edge devices enable real-time or near-real-time processing of visual data, allowing for efficient rendering and improved overall system efficiency. Furthermore, the present invention encompasses the ability of edge devices to adapt to dynamic network conditions, scale resources as needed, and optimize their operation to meet computing requirements.

[0079] In the context of the present invention, this method includes assigning one or more clusters to one computing device from a plurality of computing devices 108-1 to 108-N installed within a facility, based on a device assignment model. It should be understood that the assignment model can assign computing devices 108-1 to 108-N based on multiple factors, such as the computing power of the computing device, the availability of the computing device, and the proximity of the computing device to the user. The functionality of the assignment model is explained in more detail in Figure 6.

[0080] Figure 6 is a flowchart illustrating the steps of a method 600 according to one embodiment of the present disclosure, which assigns each cluster to computing devices 108-1 to 108-N for processing. The method includes assigning one or more clusters to one computing device from a plurality of computing devices 108-1 to 108-N based on an assignment model. The assignment model is executed on a processing unit 202. In one embodiment, the assignment model is based on a set of rules. In another embodiment, the assignment model is based on a machine learning model trained to accurately assign computing devices to process one or more clusters of visual data to be rendered. Step 602 identifies a plurality of computing devices 108-1 to 108-N installed within a facility. The facility may be scanned to identify computing devices 108-1 to 108-N capable of performing the aforementioned processing of visual data. In an exemplary embodiment, the computing devices are edge devices identified within the facility after scanning. Step 604 determines the configuration of each of the identified computing devices 108-1 to 108-N. In one example, since the computing devices are edge devices, the configuration of each of the edge devices is determined. Determining the configuration of edge devices involves identifying the network operating parameters, settings, and functionality of individual edge devices or interconnected edge devices. The configuration provides details about how the edge devices are connected, which software and hardware components they utilize, and how they interact with the network and other devices. First, information about each device's type, model, and unique identifier (e.g., MAC address or serial number) is determined. Furthermore, device specifications such as processing power, memory capacity, storage, operating system, and any dedicated hardware components or accelerators are determined. Finally, software configurations are identified, including the operating system version, firmware, drivers, and any specific applications or services installed on each edge device.Furthermore, network settings for edge devices are determined, including IP addresses, subnet masks, gateway addresses, DNS configurations, and whether they use wired or wireless connections. Additionally, data storage and management policies are determined, such as local storage capacity, data retention policies, data synchronization, and backup strategies.

[0081] In step 606, each of the one or more clusters to be rendered is listed based on the rendering order over a given period. In step 608, a complexity score is determined for each cluster to be rendered. Here, the complexity score is an indicator of the complexity of processing one or more clusters. The complexity score of a cluster to be rendered may be based on the number of entities in the cluster, the type of entities in the cluster, the geometry of the entities, the texture of the entities, the lighting and shadows of the entities, the transparency and reflections of the entities to be rendered, etc. In step 610, a priority score is determined for each of the clusters to be rendered. Here, the priority score is determined based on the complexity score and the priority of the rendering order. In step 612, the cluster to be rendered is assigned to a computing device with the best configuration based on the highest priority score of the scene. If the best computing device is not available, the cluster is assigned to the next best computing device.

[0082] In one embodiment, a method for assigning each cluster to one computing device from a plurality of computing devices 108-1 to 108-N based on a device assignment model includes assigning each cluster to a computing device based on the proximity between the user and the computing device. For example, if the user is near a conveyor belt and is visually observing the conveyor belt's operation, the edge device physically closest to the conveyor belt is selected to process one or more clusters in the scene to be rendered.

[0083] In another embodiment, a method for assigning each cluster to one computing device from a plurality of computing devices 108-1 to 108-N based on a device assignment model includes determining one or more new entities in a previously rendered scene in the computer simulation environment 102 for a particular user 112 when the user begins interacting with the environment. Furthermore, the method includes assigning one or more new entities to the computing device assigned for the previously rendered scene for user 112. For example, a scene of conveyor belt operation including a conveyor belt, workers, boxes moving on the conveyor belt is rendered to user 112, and when user 112 begins interacting with one or more boxes, image data regarding the position of the boxes in accordance with the user's actions is provided to the same edge device that rendered the original scene.

[0084] In another embodiment, a method for assigning each cluster to one computing device from a plurality of computing devices 108-1 to 108-N based on a device assignment model includes determining one or more superimposed entities 104-1 to 104-N from the respective viewpoints of multiple users collaborating in a computer simulation environment 102. Furthermore, the method includes assigning the determined superimposed entities to a specific computing device for processing visual data for all of the multiple users. For example, if multiple users are collaborating in the environment 102 and can view one or more superimposed entities or scenes, the method includes automatically identifying superimposed entities in multiple scenes that should be rendered for multiple users with superimposed viewpoints. These superimposed entities are grouped into clusters and then assigned to a single computing device for processing. Once the superimposed entities are processed, rendering is performed for the multiple users as needed. Advantageously, the method eliminates the need for repeated calculations each time a scene common to multiple users is rendered.

[0085] In one embodiment, a method for assigning each cluster to one computing device from a plurality of computing devices 108-1 to 108-N based on an allocation model includes detecting a failure of the computing device during processing of the cluster assigned for processing. For example, if the assigned edge device fails while processing of the cluster is in progress, the anomaly may be detected and reported to the processing unit. Furthermore, the method includes determining the progress of the cluster's processing completion. Furthermore, the method includes assigning another computing device available within the facility that can resume the processing completion of the cluster assigned for processing. Advantageously, if an edge device within the facility fails, rendering of the scene in the metaverse should not be interrupted, and the system reassigns the cluster to the next available edge device, thereby enabling seamless rendering of the scene in the metaverse.

[0086] In step 310, each of the one or more clusters of scenes received from each of the multiple computing devices is synchronized to render one or more scenes for one or more users 112 interacting in the computer simulation environment 102. It should be understood that the one or more clusters assigned to different computing devices are processed individually and need to be synchronized to render scenes that are meaningful to the users viewing the scenes. In one embodiment, a method for synchronizing one or more clusters of scenes includes determining the scene association of each cluster transmitted from each of the multiple computing devices based on the timestamp of each scene. In particular, each cluster is associated with a specific scene identified as described above when it is generated. For example, a first cluster may be associated with timestamp t1, a second cluster with timestamp t2, and a third cluster with timestamp t3. Also, a first scene may be associated with timestamp t1, a second scene with timestamp t2, and a third scene with timestamp t3. In such a scenario, based on timestamps t1, t2, and t3, the first cluster is associated with the first scene, the second cluster with the second scene, and the third cluster with the third scene. Furthermore, the method includes positioning each cluster in its respective scene based on the coordinates of one or more entities in each cluster and the determined scene associations. Once the association of each cluster is determined based on the timestamps, the next step is to determine the coordinates of each entity within the cluster. For each cluster to be positioned in a particular scene, the cluster positioning is done according to the real-world coordinate arrangement. Furthermore, the method includes synchronizing one or more scenes to be rendered based on the timestamps of each scene, where one or more scenes include one or more associated cluster positions. Advantageously, the one or more scenes are synchronized to render a scene meaningful to the user.

[0087] In one embodiment, a method for synchronizing each cluster of one or more scenes received from each of a plurality of computing devices 108-1 to 108-N in order to render one or more scenes for one or more users interacting in a computer simulation environment includes generating metadata for each cluster of one or more scenes processed by the computing devices, wherein the metadata includes one or more parameters that define the visual data relating to one or more entities within one or more clusters. As used herein, the term “metadata” refers to additional information or descriptive data associated with the processed data of one or more clusters. The metadata provides context, properties, or details relating to the visual data of the clusters, which are then used to render the scenes in the metaverse. The metadata may be stored in an image file or attachment, or it may be stored in a database.For example, metadata may include, but is not limited to, Exchangeable Image File Format (Exif) data containing information about camera settings and capture conditions for a specific cluster, such as camera manufacturer and model, exposure settings (shutter speed, aperture, ISO), date and time of capture, GPS coordinates, and focal length; visual data property metadata containing information about cluster size (width and height in pixels or inches), color space (e.g., RGB, CMYK), bit depth (8-bit, 16-bit, etc.), file format (JPEG, PNG, TIFF, etc.); and descriptive keywords assigned to clusters that make it easy to classify, organize, and search visual data based on specific attributes or content. The metadata may include tags, a cluster editing or processing history including the order of modifications, applied filters, and adjustments made (which helps in understanding the cluster's development), camera calibration data providing information about camera lens distortion, intrinsic parameters, and extrinsic parameters, cluster processing parameters (if the cluster has performed specific image processing operations (e.g., contrast enhancement, noise reduction, sharpening), the metadata can store the parameters used for those operations), image annotations such as cluster annotations, comments, and captions associated with the cluster, providing additional context or information about specific areas or features within the cluster, and geospatial information such as latitude, longitude, altitude, and possibly direction data. Furthermore, the method includes synchronizing metadata received from each computing device based on the arrival time and service rate of each metadata received from the computing device. Advantageously, instead of synchronizing the processed visual data from each cluster, the processed visual data from each cluster is converted into metadata and then transmitted to a processing unit for synchronization and possibly rendering of one or more scenes in the metaverse. Here, the metadata plays a crucial role in efficiently managing, organizing, analyzing, and rendering the vast image collection in the metaverse.

[0088] In an exemplary embodiment, a method for synchronizing metadata about a cluster of one or more scenes received from each of several edge devices in order to render one or more scenes in a metaverse includes determining the configuration of each assigned edge device that transmits the metadata. Furthermore, the method includes calculating the service time of the metadata received in the queue based on the determined configuration of each edge device, where service time is the time it takes for the edge device to process and handle the cluster processing request. In one example, service time includes the time required for processing, computation, data transmission, and any other tasks associated with the supply of metadata. Furthermore, the method includes calculating the arrival time of the metadata received in the queue based on the determined configuration of each edge device, where arrival time is the time it is entered into the queuing system to request service. Arrival time represents the moment when the metadata or request arrives at the synchronization device and is added to the waiting queue. Furthermore, the method includes calculating the wait time for each of the metadata arranged in the queue based on the calculated service time and calculated arrival time. Furthermore, the method includes retrieving the metadata from the queue for synchronization according to the calculated wait time for each of the metadata in the queue. Furthermore, the method includes synchronizing received metadata in order to render a scene for a specific user in a computer simulation environment. Advantageously, the method uses the arrival and service times of the metadata to optimize performance by analyzing queues, predicting behavior, and determining the most optimal service capacity, personnel levels, and processing strategies, thereby enabling one or more scenes to be efficiently rendered within the metaverse. An example of synchronizing metadata received from a computing device is illustrated in detail in Figure 7. The method further includes generating animations from one or more synchronized clusters of scenes to be rendered for a specific user. Furthermore, the method includes rendering the generated animations for the specific user in a computer simulation environment.

[0089] Referring to Figure 7, a flowchart is shown illustrating the workflow of a method 700 for synchronizing metadata received from computing devices according to one embodiment of the present disclosure. As can be seen from the figure, the device assignment model 216 provides clusters to each of the edge devices, such as “Edge Device 1” 702-1, “Edge Device 2” 702-2, “Edge Device 3” 702-3, and “Edge Device N” 702-N. Each of these edge devices 702-1 to 702-N then processes each of the clusters using one or more machine learning models and then converts the processed visual data into metadata. The metadata from each of the edge devices 702-1 to 702-N is then transmitted to a synchronization module 218, which is configured to arrange the metadata into multiple queues based on the service time and arrival time of the edge devices 702-1 to 702-N. In one example, metadata from edge device 1, specifically 702-1, is placed in queue 1, specifically 704-1; metadata from edge device 2, specifically 702-2, is placed in queue 2, specifically 704-2; metadata from edge device 3, specifically 702-3, is placed in queue 3, specifically 704-3; and metadata from edge device N, specifically 702-N, is placed in queue N, specifically 704-N. Furthermore, the synchronization module 218 synchronizes all metadata in a meaningful manner. Once the metadata is synchronized, an animation is generated from the metadata, and then the synchronized metadata is transmitted to the rendering module 220 to render the same thing to one or more users in the metaverse.

[0090] Referring to Figure 4A, a flowchart is shown illustrating the steps of a method 400A for generating one or more clusters for one or more entities that interact with each other, according to one embodiment of the present invention. In step 402, bounding boxes are generated from one or more entities for a set of entities, where one or more entities are interacting with each other. Interaction means that one or more entities are in close proximity to each other and actions are being performed on the entities. In one example, bounding boxes should be generated around a person interacting with a machine on a workbench. In step 404, the superposition score of the generated bounding boxes is determined over multiple frames, where the superposition score is determined based on a comparison of the superposition area between one or more bounding boxes and the combined area between one or more bounding boxes. In one example, the superposition score is calculated using the Intersection Factor (IoU) technique, where IoU measures the degree of superposition between the predicted bounding boxes and the ground truth bounding boxes of one or more objects, determining the accuracy of locating entities in one or more scenes. IoU is calculated by dividing the intersection area between the predicted bounding boxes and the ground truth bounding boxes by their combined area. Based on this evaluation, an IoU score is calculated for each bounding box. The resulting value, ranging from 0 to 1, represents the degree of overlap between the two bounding boxes. A higher IoU score indicates better alignment between the predicted bounding box and the ground truth bounding box, meaning improved object detection accuracy. In step 406, a frame is determined for the set of entities with the highest IoU score. In particular, if the IoU score is calculated over multiple frames for one or more bounding boxes, the frame with the highest IoU score is selected. Advantageously, calculating the IoU score over multiple frames ensures high reliability of the bounding box accuracy. In step 408, a first set of association graphs is generated for the set of entities with the highest IoU score. These first sets of association graphs are the bounding boxes of the selected frame with the highest IoU score. In step 410, one or more clusters are generated based on the generated first set of association graphs.Advantageously, since the first association graph is generated for entities that interact with each other, the one or more clusters generated consist of one or more entities that interact with each other.

[0091] Referring to Figure 4B, an exemplary depiction of one or more clusters of one or more entities interacting with each other within facility 400B according to one embodiment of the present invention. As can be seen from the figure, three clusters are depicted in the scene, in detail a first cluster "412", a second cluster "414", and a third cluster "416". The first cluster 412 is formed from a set of entities 418A, 418B, and 418C that interact with each other. The second cluster 414 is formed from a set of entities 420A, 420B, and 420C that interact with each other. The third cluster 418 is formed from a set of entities 422A, 422B, and 422C that interact with each other.

[0092] Referring to Figure 5A, a flowchart shows the steps of a method 500A for generating one or more clusters for one or more entities that do not interact with each other, according to one embodiment of the present invention / another embodiment of the present invention. In step 502, bounding boxes are generated for a set of entities from one or more entities, where the set of entities does not interact with each other. Entities that are separated from other entities and do not interact with each other are then enclosed in bounding boxes for subsequent processing. In one example, if a worker is in an exclusive area that does not interact with other entities, a bounding box is generated around a single worker. In step 504, relative distance scores are calculated between each of the one or more bounding boxes over a series of frames, where the relative distance score is the distance between sets of entities in each bounding box. Relative distance scores are calculated between bounding boxes that have a single entity that does not interact with any other entities. Furthermore, relative distance scores are calculated between bounding boxes to determine whether there are two distinct, non-interacting entities that can be clustered together over a series of frames. In step 506, a frame is determined for the set of entities that has the minimum relative distance score. It should be understood that these entities may move closer to or further away from each other across multiple frames. Therefore, across multiple frames, it is observed which particular entities are close enough to be clustered together. The frames with the minimum relative distance are then selected for clustering. In step 508, a second set of association graphs is generated for the set of entities with the minimum relative distance score. It should be understood that an association graph is generated for single entities that have the minimum relative distance, as an association may be established between such entities. In step 510, one or more clusters are generated based on the generated second set of association graphs. It should be understood that multiple association graphs may be generated based on the relative distance calculation of multiple single entities. Furthermore, these clusters are determined based on the second set of association graphs.

[0093] Figure 5B is an exemplary depiction of one or more clusters of one or more entities that do not interact with each other within facility 500B, according to one embodiment of the present invention. As can be seen from the figure, entities 512, 514, and 516 are surrounded by different bounding boxes 518, 520, and 522 because they do not interact with any other entities. Furthermore, a distance "d1" is calculated between bounding boxes 518 and 520 over multiple frames, and a distance "d2" is calculated between bounding boxes 518 and 522 over multiple frames. In this scenario, as can be seen from the figure, cluster 524 is generated with the combination of bounding boxes 518 and 520 because the relative distance "d2" is shorter than the relative distance "d1".

[0094] Figure 8 is a flowchart of the workflow of Method 800 for efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment, according to one embodiment of the present invention. As can be seen from the figure, in step 802, the actual scene is captured via a camera. In step 804, scene analysis is performed and one or more entities to be rendered are identified. In step 806, one or more clusters are generated based on the association graphs stored in the association graph storage location 808. In step 810, the facility is scanned for available edge devices and the configuration of each edge device is determined. In step 812, one or more clusters are assigned to one or more edge devices based on the device assignment model. In step 814, the association between one or more clusters and their corresponding entities is received by each edge device. In step 816, one or more clusters are processed, entities are identified and tracked using one or more machine learning models to understand the scene. In step 818, metadata is generated for each of the processed one or more clusters. In step 820, the generated metadata is aggregated from each edge device and thereby aggregated as metadata from other edge devices received in step 822. In step 824, metadata is synchronized to generate a meaningful scene. In step 826, the synchronized metadata is converted into an animation. In step 828, the animation is rendered for one or more users sharing the metaverse.

[0095] This invention provides a system and method for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment. Advantageously, the invention provides a method for rendering real-world scenes to multiple users in a metaverse in an energy-efficient manner. The invention supports energy-efficient, seamless, and near real-time metaverse rendering. The invention eliminates the need for a centralized processing system to process the visual data to be rendered in real time, thereby eliminating lag, delay, inconsistency, single-point failures, and performance deficiencies in metaverse rendering. Furthermore, the invention aims to enhance the user experience for users in the metaverse by enabling the efficient rendering of scenes in the metaverse. In particular, because the invention relies on a distributed computing system within a facility, it ensures that scene rendering is not interrupted even in the event of a network failure.

[0096] While the present invention has been described in detail with reference to specific embodiments, it should be recognized that this disclosure is not limited to these embodiments. The examples described herein are provided for illustrative purposes only and should not be construed as limiting the present invention as disclosed herein. While the present invention has been described with reference to various embodiments, it should be understood that the terms used herein are for illustrative and illustrative purposes only, and not limiting. Furthermore, while the present invention has been described with reference to specific means, materials, and embodiments, the present invention is not intended to be limited to what is disclosed herein, but rather extends to all functionally equivalent structures, methods, and uses as included in the appended claims. A person skilled in the art who has benefited from the teachings herein will be able to make numerous modifications and changes to the present invention without departing from the scope of the invention in various aspects of the present invention. [Explanation of Symbols]

[0097] 100A System 102 Computer Simulation Environment 10⁴-1 to 10⁴-N: One or more entities 105 Data Acquisition Devices 106 Communication Networks 110 Equipment 100B Drawing of a user viewing one or more scenes 112 users 114 Wearable Devices 202 One or more processing units 204 memory units 206 modules 208 Data Acquisition Modules 210 Scene Identification Module 212 Cluster Generation Module 216 Device Assignment Module 218 Synchronization Module 220 rendering modules 217 Storage Units 222 Databases 224 Input Units 226 Output Unit 228 Bus 300 Flowchart showing steps for efficient rendering of one or more scenes in a computer simulation environment 400A A flowchart illustrating the steps of a method for generating one or more clusters for one or more entities that interact with each other. 400B facility 412 First cluster 414 Second cluster 416 Third cluster 418A, 418B, 418C entities 420A, 420B, 420C physical units 422A, 422B, 422C entities 500A A flowchart illustrating the steps of a method for generating one or more clusters of one or more entities that do not interact with each other. 500B facility 512, 514, 516 entities 518,522 Bounding Box 524 clusters 600 Method for assigning each cluster to a computing device for processing visual data. 700 How to synchronize metadata received from computing devices 702-1~702-N Edge Devices 704-1~704-N Queue 800 A method for efficiently rendering one or more scenes for one or more users interacting in a computer simulation environment.

Claims

1. A computer implementation method for efficiently rendering one or more scenes for one or more users interacting in a computer simulation environment (102) that refers to a three-dimensional (3D) representation of the physical world, The aforementioned method, The processing unit (202) receives visual data from one or more data acquisition devices (105) configured to collect data about one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) interacting within the facilities (400B, 500B), wherein the visual data is collected from the viewpoints of one or more users interacting in the computer simulation environment (102). The processing unit (202) includes the steps of identifying one or more scenes to be rendered in the computer simulation environment (102) for a specific user, and one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) within those scenes, The processing unit (202) generates one or more clusters (412, 414, 416, 524) for each of the one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) identified from the one or more scenes to be rendered, wherein each cluster (412, 414, 416, 524) includes a set of entities that are related to each other. The processing unit (202) performs the steps of assigning each of the clusters (412, 414, 416, 524) in one or more scenes to one computing device from a plurality of computing devices installed in the facility (400B, 500B) based on a device assignment model, wherein each of the plurality of computing devices is configured to process the one or more clusters (412, 414, 416, 524) using one or more machine learning models; The processing unit (202) performs the steps of synchronizing each cluster (412, 414, 416, 524) of the one or more scenes received from each of the plurality of computing devices in order to render the one or more scenes for the one or more users who interact in the computer simulation environment (102), Computer implementation methods, including those mentioned above.

2. The above method further, The steps include generating animation from synchronized visual data regarding a scene to be rendered for a specific user, and The method according to claim 1, further comprising the step of rendering the generated animation to a specific user in the computer simulation environment (102).

3. The step of synchronizing the clusters (412, 414, 416, 524) of one or more scenes to be rendered further includes: The processing unit (202) performs the steps of determining the scene relationships of each of the clusters (412, 414, 416, 524) transmitted from each of the plurality of computing devices based on the timestamp of each scene, The processing unit (202) performs the steps of placing each of the clusters (412, 414, 416, 524) into the respective scenes based on the coordinates of one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) in each of the clusters (412, 414, 416, 524) and the determined scene relationships, The method according to claims 1 and 2, comprising the step of synchronizing one or more scenes to be rendered based on the timestamp of each scene using the processing unit (202), wherein the one or more scenes include one or more associated cluster arrangements.

4. The step of synchronizing each cluster (412, 414, 416, 524) of the one or more scenes received from each of the plurality of computing devices in order to render the one or more scenes for the one or more users interacting in the computer simulation environment (102) is: The processing unit (202) generates metadata for each of the clusters (412, 414, 416, 524) of the one or more scenes processed by the computing device, wherein the metadata includes one or more parameters that define visual data relating to one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) within the one or more clusters (412, 414, 416, 524); The method according to any one of claims 1 to 3, further comprising the step of synchronizing metadata received from each of the computing devices based on the arrival time and service rate of each metadata received from the computing devices using the processing unit (202).

5. The step of generating clusters (412, 414, 416, 524) for each of the one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) identified from the scene to be rendered, The processing unit (202) generates a bounding box for a set of entities from one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516), wherein the set of entities interacts with each other. A step of determining the superposition score of the bounding boxes (518, 522) generated across multiple frames by the processing unit (202), wherein the superposition score is determined based on a comparison of the superposition area between one or more bounding boxes (518, 522) and the combined area between the one or more bounding boxes (518, 522). The processing unit (202) performs the steps of determining a frame relating to the set of entities having the highest value of the superposition score, The processing unit (202) performs the steps of generating a first set of association graphs for the set of entities having the highest superposition score, The method according to any one of claims 1 to 4, further comprising the step of generating one or more clusters (412, 414, 416, 524) based on the set of first association graphs generated by the processing unit (202).

6. The step of generating clusters (412, 414, 416, 524) for each of the one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) identified from the scene to be rendered, The processing unit (202) generates a bounding box (518, 522) for a set of entities from one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516), wherein the set of entities does not interact with each other. The processing unit (202) calculates relative distance scores between each of the one or more bounding boxes (518, 522) over a plurality of frames, wherein the relative distance score is the distance between the set of entities in each of the bounding boxes (518, 522). The processing unit (202) determines a frame relating to the set of entities having the minimum value relative distance score, The processing unit (202) performs the steps of generating a second set of association graphs for the set of entities having the minimum relative distance score, The method according to any one of claims 1 to 5, further comprising the step of generating one or more clusters (412, 414, 416, 524) based on the set of second association graphs generated by the processing unit (202).

7. The step of assigning each cluster (412, 414, 416, 524) from multiple computing devices to one computing device based on the device allocation model is: The processing unit (202) identifies the plurality of computing devices installed within the facility (400B, 500B), The processing unit (202) determines the configuration of each of the identified computing devices, The processing unit (202) performs the steps of listing each of one or more clusters (412, 414, 416, 524) to be rendered based on the rendering order over a predetermined period of time, A step of determining a complexity score for each cluster (412, 414, 416, 524) to be rendered by the processing unit (202), wherein the complexity score is an indicator of the complexity of processing for one or more clusters (412, 414, 416, 524). The processing unit (202) determines a priority score for each of the clusters (412, 414, 416, 524) to be rendered, wherein the priority score is determined based on the complexity score and the priority of the rendering order. The method according to any one of claims 1 to 6, comprising the step of assigning the clusters (412, 414, 416, 524) to be rendered by the processing unit (202) to a computing device having an optimal configuration based on the highest priority score of the clusters (412, 414, 416, 524).

8. The step of assigning each cluster (412, 414, 416, 524) from multiple computing devices to one computing device based on the device allocation model is: The method according to any one of claims 1 to 7, comprising the step of assigning each set of clusters (412, 414, 416, 524) in a first association graph and in a second association graph to a computing device based on the proximity between the user and the computing device.

9. The step of assigning each cluster (412, 414, 416, 524) from multiple computing devices to one computing device based on the device allocation model is: The processing unit (202) performs the steps of determining, for a particular user, one or more new entities in a previously rendered scene in the computer simulation environment (102) when the user begins interacting with one or more new entities in the computer simulation environment (102); The method according to any one of claims 1 to 8, further comprising the step of assigning one or more new entities to a computing device allocated for a scene previously rendered for the user by the processing unit (202).

10. The step of assigning each cluster (412, 414, 416, 524) from multiple computing devices to one computing device based on the allocation model is: The processing unit (202) determines one or more entities (418A, 418B, 418C, 420A, 420B, 420C, 422A, 422B, 422C, 512, 514, 516) that are superimposed from the perspectives of multiple users who are working together in the computer simulation environment (102), The method according to any one of claims 1 to 9, further comprising the step of assigning the superimposed entities determined by the processing unit (202) to a specific computing device for processing visual data for all of the multiple users.

11. The step of assigning each cluster (412, 414, 416, 524) from multiple computing devices to one computing device based on the allocation model is: The processing unit (202) includes the step of detecting a failure in the computing device during processing of the clusters (412, 414, 416, 524) that have been allocated for processing, The processing unit (202) determines the progress of the processing completion of the clusters (412, 414, 416, 524) that have been assigned for processing, The method according to any one of claims 1 to 10, further comprising the step of assigning another computing device to the processing unit (202) that is capable of processing the clusters (412, 414, 416, 524) allocated for processing and is available within the facility (400B, 500B).

12. A device (110) for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment (102), The aforementioned device (110) is as follows: One or more processing units (202), Includes a memory (204) that is communicably coupled to one or more processing units (202), The memory (204) includes a module in which machine-readable instructions executable by one or more processing units (202) are stored, and the apparatus (110) is configured to perform the method steps according to any one of claims 1 to 11.

13. A system (100A) for the efficient rendering of one or more scenes for one or more users interacting in a computer simulation environment (102), The aforementioned system (100A) is as follows: A computer simulation collaborative environment (102) that renders one or more scenes corresponding to real-world entities within the facility (400B, 500B), Multiple computing devices are connected to the aforementioned computer simulation shared environment (102) in a communicative manner, wherein the multiple computing devices are installed within the aforementioned facilities (400B, 500B), A system (100A) comprising the apparatus (110) according to claim 12, which is communicably coupled to the plurality of computing devices and the computer simulation cooperative environment (102), wherein the apparatus (110) is configured to efficiently render one or more scenes for one or more users interacting in the computer simulation environment (102) according to the method according to any one of claims 1 to 11.

14. A computer program product comprising a computer-readable instruction stored internally which, when executed by a processing unit (202), causes the processing unit (202) to perform the method step described in any one of claims 1 to 11.

15. A computer-readable medium on which a program code section of a computer program is stored, wherein the program code section is loadable into and / or executable within a system (110A) so as to cause the system (100A) to perform the method step described in any one of claims 1 to 11 when the program code section is executed within the system (100A).