Streaming visual for configuring 3D object
Patent Information
- Application Number
- AE202602569
- Authority / Receiving Office
- AE · AE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2025-01-29
Smart Images

Figure FULLTEXT_1 
Figure FULLTEXT_2 
Figure FULLTEXT_3
Abstract
Description
STREAMING VISUAL FOR CONFIGURING 3D OBJECTTECHNICAL FIELDThe present disclosure relates to methods for streaming visuals for configuring three-dimensional (3D) objects. The present disclosure also relates to systems for streaming visuals for configuring 3D objects.BACKGROUNDAdvancements in technology have revolutionized visuals and / or visual experiences, making them increasingly engaging, interactive, detail-rich, and lifelike. These visual experiences are not limited to entertainment but extend to education, training and simulation, product showcasing, gaming, environment designing, healthcare and wellness, personal care, and more. Despite these advancements, existing technologies face significant limitations that hinder the delivery of high-quality visual experiences.Three-dimensional (3D) visualizers or configurators, which allow users to visualize objects in 3D, often suffer from extremely high loading times. These delays frustrate users and result in high bounce rates, sometimes exceeding 70 percent. Once loaded, these visualizers can overwhelm users with complex navigation, leading to confusion and a suboptimal user experience.Additionally, product showcasing experiences, such as digital solutions for vehicles, furniture, and real estate, often lack performance, realism, device compatibility, and responsiveness. Custom development is typically required to support extended reality (XR) formats. Furthermore, visual experiences that provide product comparison functionality usually rely on textual information, which limits visual engagement and fails to offer an immersive experience. Comparative visual experiences in 3D are not yet widely available, and there is no scalable, no-code solution for creating them.Therefore, in the light of foregoing discussion, there exists a need to overcome the aforementioned drawbacks associated with the existing technology for providing visual comparison experiences.SUMMARYThe present disclosure seeks to provide a method for streaming a visual for configuring a three-dimensional (3D) object. The present disclosure also seeks to provide a system for streaming a visual for configuring a 3D object. An aim of the present disclosure is to provide a solution that overcomes at least partially the problems encountered in prior art.In a first aspect, an embodiment of the present disclosure provides a method for streaming a visual for configuring a three-dimensional (3D) object, the method comprising:- defining at least one first variable that pertains to a complexity of the visual;- defining at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual;- generating an adaptive render streaming output by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual;- rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and- streaming the adaptive render streaming output from a server to one or more user devices, via the communication network, according to the corresponding bit rates.Optionally, in the method, at least one first variable comprises at least one of: a rendering quality for deep learning super sampling (DLSS), an upscaling quality for DLSS, quantization for pixel values, a foveated rendering scheme, and a resolution.Optionally, the step of rendering the set of image frames representing the visual, at each time instant, comprises:- applying the foveated rendering scheme for adjusting the resolution of the set of image frames.Optionally, the step of rendering the set of image frames representing the visual, at each time instant, comprises:- rendering the set of image frames according to a given value of the rendering quality for DLSS;- employing DLSS for enhancing the resolution of the set of image frames, from the given value of the rendering quality for DLSS to a target value of the upscaling quality for DLSS; and- applying the quantization for pixel values of pixels in the image frames of the set, according to a value of the quantization.Optionally, in the method, at least one second variable comprises at least one of: a bandwidth, a latency, a data transfer rate, a compression ratio, and a data packet loss ratio.Optionally, in the method, a value of at least one first variable is obtained from at least one of: a rendering engine, a graphics processing unit configured for rendering image frames, an input received from the one or more user devices, and a predefined configuration that is pre-stored at a memory that is communicably coupled with the server.Optionally, in the method, a value of at least one second variable is obtained from at least one of: network monitoring and / or diagnostic data received from a network device in the communication network, and communication metrics received from the one or more user devices.Optionally, the method further comprises adjusting a value of at least one first variable, based on a value of at least one second variable, prior to implementing the step of rendering the set of image frames representing the visual, at a given time instant.Optionally, the method further comprises integrating an extended-reality (XR) plugin into the visual, thereby enabling the visual to be provided as an XR visual.Optionally, in the method, the 3D object is a vehicle. Optionally, in the method, the one or more time instants comprise at least one of: a start time instant of streaming the visual, a time instant of receiving an input from the one or more user devices, a time instant of detecting a change in the network parameters of the communication network, and a time instant of detecting a change of the communication network.Optionally, in the method, the visual for configuring the 3D object is a part of a product suite comprising a plurality of visuals for interactively viewing, configuring, and comparing three-dimensional (3D) objectsIn a second aspect, an embodiment of the present disclosure provides a system for streaming a visual for configuring a three-dimensional (3D) object, the system comprising a server configured to:- define at least one first variable that pertains to a complexity of the visual;- define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual;- generate an adaptive render streaming output by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual;- rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and- stream the adaptive render streaming output to one or more user devices, via the communication network, according to the corresponding bit rates.In a third aspect, an embodiment of the present disclosure provides a non-transitory computer-readable medium storing a set of instructions for streaming a visual for configuring a three-dimensional (3D) object, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device, cause the device to:- define at least one first variable that pertains to a complexity of the visual;- define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual;- generate an adaptive render streaming output by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual;- rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and- stream the adaptive render streaming output from a server to one or more user devices, via the communication network, according to the corresponding bit rates.Embodiments of the present disclosure substantially eliminate or at least partially address the aforementioned problems in the prior art, and enable a more efficient, responsive, and user-friendly object configuration, allowing users to quickly load, navigate, and interact / configure 3D object in real time or near-real time.Additional aspects, advantages, features and objects of the present disclosure would be made apparent from the drawings and the detailed description of the illustrative embodiments construed in conjunction with the appended claims that follow.It will be appreciated that features of the present disclosure are susceptible to being combined in various combinations without departing from the scope of the present disclosure as defined by the appended claims.BRIEF DESCRIPTION OF THE DRAWINGSThe summary above, as well as the following detailed description of illustrative embodiments, is better understood when read in conjunction with the appended drawings. For the purpose of illustrating the present disclosure, exemplary constructions of the disclosure are shown in the drawings. However, the present disclosure is not limited to specific methods and instrumentalities disclosed herein. Moreover, those skilled in the art will understand that the drawings are not to scale. Wherever possible, like elements have been indicated by identical numbers.Embodiments of the present disclosure will now be described, by way of example only, with reference to the following diagrams wherein:FIG. 1A illustrates steps of a method for streaming a visual for configuring a three-dimensional (3D) object, FIG. 1B illustrates additional optional steps of the method as described in FIG. 1A, in accordance with various embodiments of the present disclosure; FIG. 2 illustrates a block diagram of an architecture of a system for streaming a visual for configuring a 3D object, in accordance with an embodiment of the present disclosure;FIG. 3 illustrates an exemplary schematic process flow for generating a three-dimensional (3D) scene, in accordance with an embodiment of the present disclosureFIG. 4 illustrates a flowchart of method steps involved in minimizing a reduction in a number of frames per second (FPS) when generating a visual for visual comparison of a plurality of objects, in accordance with an embodiment of the present disclosure; andFIG. 5 illustrates a schematic flowchart showing minimization of a reduction in a number of frames per second (FPS) when generating the visual for visual comparison of a plurality of objects, in accordance with an embodiment of the present disclosure.In the accompanying drawings, an underlined number is employed to represent an item over which the underlined number is positioned or an item to which the underlined number is adjacent. A non-underlined number relates to an item identified by a line linking the non-underlined number to the item. When a number is non-underlined and accompanied by an associated arrow, the non-underlined number is used to identify a general item at which the arrow is pointing.DETAILED DESCRIPTIONThe following detailed description illustrates embodiments of the present disclosure and ways in which they can be implemented. Although some modes of carrying out the present disclosure have been disclosed, those skilled in the art would recognize that other embodiments for carrying out or practising the present disclosure are also possible.The present disclosure provides a method and a system for streaming visuals and / or visual experience for configuring a three-dimensional (3D) object. Herein, rendering streaming outputs of the 3D configurator is done based on the evolving complexity of the visual content and the changing conditions of the communication network. By adapting the visual experience to the available resources, users are likely to experience smoother and more consistent streaming. This contributes to an improved user experience, minimizing disruptions or buffering issues. Additionally, the adaptive nature of the streaming output makes the visual content more accessible across a range of devices and network conditions and thereby, users with varying capabilities and network speeds still enjoy the content with an optimized streaming experience. The adaptive nature of the streaming method allows for scalability, enables to handle different levels of content complexity and network conditions, making it suitable for different scenarios. By dynamically adjusting to network parameters, the streaming method can help reduce latency, providing a more responsive and interactive experience for users. Referring to FIGs. 1A and 1B, FIG. 1A illustrates steps of a method for streaming a visual for configuring a 3D object, FIG. 1B illustrates additional optional steps of the method as described in FIG. 1A, in accordance with various embodiments of the present disclosure. The steps illustrated in the aforesaid FIGs. 1A and 1B will now be discussed hereinbelow in detail.With reference to FIG. 1A, at step 102, at least one first variable is defined that pertains to a complexity of the visual. At step 104, at least one second variable is defined that pertains to network parameters of a communication network that is to be employed for streaming the visual. At step 106, an adaptive render streaming output is generated. The step 106 comprises steps 108a, 108b, and 108c. At step 108a, values of at least one first variable and at least one second variable are obtained, at one or more time instants of streaming the visual. At step 108b, a set of image frames representing the visual is rendered, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant. At step 108c, a corresponding bit rate to be employed is determined for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant, wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically. At step 110, the adaptive render streaming output is streamed from a server to one or more user devices, via the communication network, according to the corresponding bit rates.With reference to FIG. 1B, the step 108b of rendering the set of image frames representing the visual, at each time instant, comprises steps 112, 114, 116, and 118. Additionally optionally, at step 112, the foveated rendering scheme is applied for adjusting the resolution of the set of image frames. Additionally optionally, at step 114, the set of image frames is rendered according to a given value of the rendering quality for DLSS. At step 116, DLSS is employed for enhancing the resolution of the set of image frames, from the given value of the rendering quality for DLSS to a target value of the upscaling quality for DLSS. At step 118, the quantization for pixel values of pixels is applied in the image frames of the set, according to a value of the quantization. The aforementioned steps are only illustrative and other alternatives can also be provided where one or more steps are added, one or more steps are removed, or one or more steps are provided in a different sequence without departing from the scope of the claims herein. Each of these steps is described later in more detail.Referring to FIG. 2, illustrated is a block diagram of an architecture of a system 200 for streaming a visual for configuring a 3D object, in accordance with an embodiment of the present disclosure. The system 200 comprises at least one server (for example, depicted as a server 202). The server 202 is shown to be communicably coupled to a user device 204. Notably, at least one server 202 is configured to:- define at least one first variable that pertains to a complexity of the visual;- define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual;- generate an adaptive render streaming output by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual;- rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and- stream the adaptive render streaming output to one or more user devices 204, via the communication network, according to the corresponding bit rates.It may be understood by a person skilled in the art that the FIG. 2 includes a simplified architecture of the system 200, for sake of clarity, which should not unduly limit the scope of the claims herein. It is to be understood that the specific implementation of the system 200 is not to be construed as limiting it to specific numbers or types of server and / or processors and user devices. The person skilled in the art will recognize many variations, alternatives, and modifications of embodiments of the present disclosure.In one embodiment, the 3D object is to be shown as a visual in a 3D scene along with other visuals or information. Further, the 3D object may be configured such that a visual effect or a layer of another visual is added on the visual of the 3D object which collectively creates a new visual in the 3D scene. Thus, the new visual collectively created by the visual of the 3D object and other visuals along with the addition of layers or visual effects may lead to different visual experiences.Throughout the present disclosure, the term "visual experience" refers to an experience that enables in real-time modification, customization, and visualization of a 3D object. Such experience is facilitated via the interactive user interface of the user device. The visual experience may involve manipulation of and / or interaction with a plurality of objects, thereby allowing for a more intuitive and informative visual experience.Optionally, at least one first variable that pertains to a complexity of the visual experience is defined. At least one first variable comprises at least one of: a rendering quality for deep learning super sampling (DLSS), an upscaling quality for DLSS, quantization for pixel values, a foveated rendering scheme, and a resolution. DLSS is a technology that uses artificial intelligence (AI) to upscale lower-resolution images to higher resolutions. The upscaling results in sharper and more detailed visuals, improved temporal stability, and reduced ghosting. DLSS helps improve performance by reducing the computational load on the server 202, allowing for higher frame rates without sacrificing quality. Quantization for pixel values may refer to the process of mapping a large range of continuous pixel intensity values to a finite range of discrete levels. Each pixel's intensity value is rounded to the nearest available level within a defined bit depth (e.g., 8-bit, 16-bit), effectively reducing the number of unique values that each pixel can represent. This process helps to compress image data and make it more manageable for storage and processing, although it can introduce some loss of detail and subtle variations in color and brightness. Quantization is an important step in digitizing images and plays a significant role in various image compression techniques. The foveated rendering scheme is a technique used in computer graphics, particularly in virtual reality (VR) and augmented reality (AR) systems, to optimize the rendering process by reducing the workload. The foveated rendering scheme is based on the way human vision works: our eyes have a small area called the fovea, where we perceive high-resolution detail, while the peripheral vision has much lower resolution. In foveated rendering, the system renders the high-resolution detail only in the region of the screen where the user is directly looking (the foveal region). The rest of the screen, which falls into the user's peripheral vision, is rendered at a lower resolution. This selective rendering significantly reduces the amount of computation required, leading to better performance and higher frame rates without noticeable loss in visual quality. Resolution refers to the amount of detail an image or display can render and is commonly measured in pixels. In digital imaging, resolution indicates the number of pixels that make up the width and height of an image (e.g., 1920x1080 pixels for Full HD). Higher resolution means more pixels and finer detail, resulting in a clearer and sharper image. For displays, resolution impacts the quality and clarity of what is viewed on the screen, with higher resolutions providing better visual experiences. Optionally, a value of at least one first variable is obtained from at least one of: a rendering engine, a graphics processing unit configured for rendering image frames, an input received from the one or more user devices 204, and a predefined configuration that is pre-stored at a memory that is communicably coupled with the server 202.Optionally, at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual experience is defined. at least one second variable comprises at least one of: a bandwidth, a latency, a data transfer rate, a compression ratio, and a data packet loss ratio. Bandwidth is a crucial network parameter that refers to the maximum amount of data that can be transmitted over a network connection in a given amount of time, typically measured in bits per second (bps). Essentially, bandwidth indicates the capacity of a network link to handle data traffic. High bandwidth means more data can be transferred quickly, resulting in faster internet speeds and more efficient communication between devices. Conversely, low bandwidth can lead to slower data transfer rates and potential bottlenecks, affecting the performance of network-dependent applications. Latency is a network parameter that measures the delay between a data packet being sent and received across a network. Latency is usually measured in milliseconds (ms) and is often referred to as "ping." High latency results in noticeable delays, affecting real-time applications. Conversely, low latency ensures smoother and more responsive communication. Factors contributing to latency include the physical distance between devices, network congestion, and the quality of the connection. Managing and minimizing latency is essential for maintaining optimal network performance and user experience.Data transfer rate may refer to the speed at which data is transmitted between devices over a network. Data transfer rate is usually measured in bits per second (bps), kilobits per second (Kbps), megabits per second (Mbps), or gigabits per second (Gbps). A higher data transfer rate means more data can be sent or received in a given amount of time, which is crucial for high-speed internet and efficient data communication. Compression Ratio is the ratio of the original data size to the compressed data size. Compression ratio is a measure of how much data can be reduced in size by a compression algorithm. For example, a compression ratio of 4:1 means the compressed data is one-fourth the size of the original. Compression is important for reducing storage requirements and speeding up data transmission. Data packet loss ratio measures the percentage of data packets that are lost during transmission over a network. A high packet loss ratio can lead to poor network performance, as lost packets need to be retransmitted, causing delays and potentially disrupting applications like streaming. Ideally, networks aim to minimize packet loss to ensure smooth and reliable data transfer.Optionally, a value of at least one second variable is obtained from at least one of: network monitoring and / or diagnostic data received from a network device in the communication network, and communication metrics received from the one or more user devices 204.Optionally, an adaptive render streaming output is generated by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual;- rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically.The one or more time instants comprise at least one of: a start time instant of streaming the visual, a time instant of receiving an input from the one or more user devices 204, a time instant of detecting a change in the network parameters of the communication network, and a time instant of detecting a change of the communication network.Bit rate is often measured in bits per second (bps) and refers to the amount of data processed or transmitted per unit of time in the communication network. Bit rate indicates the data transfer speed. Higher bit rates generally result in better quality because more data is allocated to capture detail, but they also require more bandwidth and storage. Conversely, lower bit rates save on bandwidth and storage but may compromise quality. Bit rate is an important factor for optimizing performance and quality in applications such as streaming.Optionally, a value of at least one first variable is adjusted, based on a value of at least one second variable, prior to implementing the step of rendering the set of image frames representing the visual, at a given time instant.Optionally, the step of rendering the set of image frames representing the visual, at each time instant, comprises:- applying the foveated rendering scheme for adjusting the resolution of the set of image frames.Optionally, the step of rendering the set of image frames representing the visual, at each time instant, comprises:- rendering the set of image frames according to a given value of the rendering quality for DLSS;- employing DLSS for enhancing the resolution of the set of image frames, from the given value of the rendering quality for DLSS to a target value of the upscaling quality for DLSS; and- applying the quantization for pixel values of pixels in the image frames of the set, according to a value of the quantization.Optionally, the adaptive render streaming output is rendered from the server 202 to one or more user devices 204, via the communication network, according to the corresponding bit rates.Optionally, the method further comprises integrating an extended-reality (XR) plugin into the visual, thereby enabling the visual to be provided as an XR visual.Optionally, the visual for configuring the 3D object is a part of a product suite comprising a plurality of visuals for interactively viewing, configuring, and comparing three-dimensional (3D) objects.Optionally, the 3D object is a vehicle. In this regard, the streaming of the visual experience may be indicative of the vehicle that is to be configured. Examples of the vehicle include, but are not limited to, a car, a truck, a motorcycle, a bus, and a bicycle. A technical benefit of the 3D object being the vehicle is that the interactive user interface allows the user to configure the vehicle in real-time, meaning that the user can see immediate changes as he / she modifies different aspects of the vehicle. For example, the user might change the vehicle's exterior color, add or remove accessories like roof racks or spoilers, or adjust interior features such as seat materials and dashboard designs. Such real-time feedback creates an immersive experience where users can visualize their customizations instantly, helping them better understand how their choices affect an overall appearance and functionality of the vehicle. Additionally, said real-time configurability facilitates in a streamline decision-making, i.e., rather than having to imagine or mentally visualize how different changes look, users can experiment freely with different vehicle configurations, and quickly arrive at decisions. This reduces hesitation and indecision, making an overall purchasing or customization process smoother and efficient. Displaying a 3D representation of the vehicle and updating it in real-time based on interactions of the user, it can be ensured that changes are accurately reflected in the vehicle, enabling the user to see a realistic preview of customizations.The server 202 may include suitable logic, circuitry, interfaces, and / or codes, executable by the circuitry, that may be configured to perform the one or more operations for streaming visuals for configuring the 3D object. For example, the server 202 may be configured to control and manage various functionalities and operations such as defining variables, generating adaptive render streaming output, and streaming the adaptive render streaming output. The various functionalities and operations may be controlled and managed by means of one or more internal components of the server 202, such as an variable definition module 206, an output generation module 208, and a streaming module 210.The variable definition module 206 may include suitable logic, circuitry, interfaces, and / or codes, executable by the circuitry, that may be configured to perform the one or more operations for defining variables. The variable definition module 206 may define at least one first variable that pertains to a complexity of the visual and define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual. The output generation module 208 may include suitable logic, circuitry, interfaces, and / or codes, executable by the circuitry, that may be configured to perform the one or more operations for generating the adaptive render streaming output. The output generation module 208 may generate an adaptive render streaming output by: obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual; rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant. The adaptive render streaming output may comprise one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically. The streaming module 210 may include suitable logic, circuitry, interfaces, and / or codes, executable by the circuitry, that may be configured to perform the one or more operations for streaming the adaptive render streaming output. The streaming module 210 may stream the adaptive render streaming output to one or more user devices 204, via the communication network, according to the corresponding bit ratesIt will be appreciated that a first video is displayed on the interactive user interface of the user device while the object configurator is being loaded in a background. The first video can be understood to an initial audiovisual or visual presentation that showcases features and perspectives of the 3D object, serving as an introductory mode to capture user interest, while the object configurator is being loaded in a background. Optionally, the first video comprises a plurality of first image frames representing the 3D object. Optionally, the first representation of the 3D object is shown in the plurality of first image frames. Optionally, the plurality of first image frames are pre-generated and pre-stored in at least one data repository that is communicably coupled to at least one processor. At least one data repository could be implemented, for example, as a memory of at least one server 202, a memory of a computing device, a memory of the user device 204, a removable memory, a cloud-based database, a digital library, or similar. Examples of the computing device include, but are not limited to, a laptop, a tablet, a phablet, and a smartphone.Throughout the present disclosure, the term "object configurator" refers to a software module that enables the user to configure (namely, customize or modify) attributes, features, and settings of the 3D object. It will be appreciated that when the first video is displayed, visual representation of the 3D object is being displayed to the user, and simultaneously, at least one server 202 initializes the loading of the object configurator in the background. This could involve fetching or preparing necessary software tools, assets, and resources required for configuring the 3D object. The loading process may include loading 3D models, textures, configuration options, and other user interface elements that will be necessary for the user to interactively configure the 3D object.It will be appreciated that an initial display of the first video keeps the user engaged until the object configurator is fully loaded and ready. A technical benefit of such an approach is a reduction in a perceived load time on the interactive user interface. Traditional 3D configurators often have high load times, frustrating users and causing disengagement. By displaying the first video (for example, of a vehicle such as a car) while the object configurator loads in the background, at least one processor enables in masking a loading process, ensuring that users are engaged with visual content without waiting idly for the object configurator to load. Such a seamless switching from the first video to the 3D scene of the object configurator (when the predefined criterion is satisfied) makes the user feel that the loading process is fast and smooth, enhancing overall user satisfaction and reducing frustration. It will be appreciated that preparing the object configurator in advance, the method enables the user to interact with and configure the 3D object as soon as the object configurator is ready, without any delays. This leads to a more responsive user interface, contributing to higher satisfaction in scenarios where real-time customization is essential.Throughout the present disclosure, the term "predefined criterion" refers to a condition that is to be fulfilled (namely, satisfied) to trigger a transition from the first video to the object configurator. It will be appreciated that once the predefined criterion is satisfied, at least one server 202 initiates the switching from the first video to the object configurator, and displays the 3D scene of the object configurator on the interactive user interface. The transition may occur automatically or upon confirmation by the user, depending on the predefined criterion.Optionally, the predefined criterion is detected to be satisfied upon one of: receiving a second input for switching from the first video to the object configurator, while the first video is being displayed; elapsing of a total time duration of the first video.In this regard, the term "second input" refers to a user-initiated action that triggers the switching from displaying the first video of the 3D object to the object configurator. For example, in some implementations, the user may expressly provide the second input (via the user device) for switching from the first video to the object configurator (for example, the user may select the "enter configurator mode" option displayed alongside / over the first video of the object, or sends a voice command for the same, or similar). Such an approach ensures a highly responsive user experience by allowing the user to control a timing of the switching according to his / her preference. Additionally, such a user-initiated transition enhances engagement and satisfaction by providing a more interactive and intuitive interface. However, in other implementations, at least one processor tracks a playback time of the first video, and the switching from the first video to the object configurator can happen when an entirety of the first video has been played. In such implementations, a total duration of the first video is to be equal to or greater than a loading time required by the object configurator. For example, when it is known that the object configurator loads fully in a range of 7 to 10 seconds, the total time duration of the first video could be 10 seconds or more. During the playback of the first video, at least one server 202 continuously monitors the elapsed time. Once the total time duration of the first video has played, at least one server 202 confirms that the object configurator has sufficient time to load. Upon completion of the first video, at least one processor enables in switching from the first video to the object configurator, ensuring that the object configurator is ready for object interaction / object configuration purposes. It will be appreciated that by setting the duration of the first video to match or exceed the loading time of the object configurator, the method ensures that the object configurator is fully operational by the time the first video ends. This approach provides a seamless user experience by synchronizing video playback with the loading process of the object configurator, eliminating the need for additional user input to switch modes. Additionally, this approach eliminates the risk of the object configurator being unready or incomplete when the user transitions from the first video, enhancing reliability and reducing potential frustration. A technical effect of the aforementioned feature is that it provides flexibility and reliability in managing transitions from presentation mode to configurator mode. By incorporating multiple criteria either user-initiated input or the completion of the total time duration of the first video, the method ensures that the transition occurs seamlessly regardless of user interaction.Optionally, the method further comprising providing a first switching element on the interactive user interface while the first video is being displayed, wherein the second input is received upon activation of the first switching element. In this regard, the term "switching element" refers to an interactive component or interface control element within the interactive user interface that enables the user to initiate or trigger a transition, change, action and similar. For example, the switching element can be a button, an icon, a toggle, or any other interactive widget that the user can engage with to perform a specific function, such as the switching between different modes or screens, initiating a process, modifying states, or similar. The method involves displaying the first switching element on the interactive user interface while the first video plays. When the user activates the first switching element by clicking it, tapping it, or using another form of input, at least one server 202 detects this activation as the second input. Upon receiving the second input, at least one server 202 initiates the transition from the first video to the object configurator. This transition process typically involves stopping the video playback and loading or displaying the object configurator.In some cases, when a duration of the first video is greater than a loading time duration of the object configurator, the switching element could be provided upon elapsing of the loading time since a start time of displaying the first video. For example, when duration of the first video may be 15 seconds and the loading time duration of the object configurator may be 8 seconds, the switching element can appear on the interactive user interface after 8 seconds have passed since a start time of the first video. In other cases, when a duration of the first video is greater than a loading time duration of the object configurator, the first switching element can be provided at a start time of displaying the first video, but its activation could be disabled until the loading time has elapsed. It will be appreciated that displaying the first switching element only after the object configurator has sufficiently loaded, or by ensuring that the first switching element is visible but inactive until the object configurator is ready, the method may facilitate in minimising a risk of premature user interactions that could lead to incomplete or erroneous configuration experiences. Such an approach ensures that the user can only initiate transitions when the object configurator is fully prepared, thereby reducing frustration and improving overall user experience. Additionally, allowing the first switching element to appear after the loading time or to be initially visible but inactive provides clear and immediate feedback to the user about an availability of switching options. Such a clarity in a design of the interactive user interface prevents confusion and sets appropriate expectations regarding a timing of available actions.Throughout the present disclosure, the term "three-dimensional scene" refers to a 3D visual scene representing at least a visual representation of a given 3D object and at least one configuration element. Herein, the term "visual representation" may encompass colour information associated with the given 3D object, and additionally optionally at least one of: depth information, transparency information, brightness information, contrast information, and the like, associated with the given 3D object. It will be appreciated that the 3D scene may not necessarily represent only the given 3D object, but may also represent one or more 2D objects (for example, such as 2D virtual objects), in addition to the given 3D object. Herein, the term "2D object" refers to a computer-generated object that could be superimposed in the 3D scene, for example, to create a visual effect. The 2D object could, for example, be a symbol, a coloured icon, a pointer, or similar. It is to be understood that the 3D scene of the object configurator represents the first representation of the 3D object. The term "configuration element" refers to an interactive component within the interactive user interface designed to modify or adjust attributes of the 3D object. Examples of the configuration element may include, but are not limited to, buttons, sliders, drop-down menus, icons, and the like.The user interacts with the plurality of configuration elements to adjust and personalize the 3D object according to their preferences. It will be appreciated that displaying the 3D scene of the object configurator with the plurality of configuration elements provides a dynamic and immersive user experience, allowing the user to intuitively visualize and personalize the 3D object. In a first example, in a virtual car customization application, where a user may complete watching the first video about a new car model. Once the first video has ended, or a specific interaction point is reached, at least one server 202 displays the 3D scene on the interactive user interface featuring a detailed visual of the car that is the first representation along with the plurality of configuration elements such as paint colour options, wheel designs, interior materials, accessory choices, and similar. Such a setup enables the user to interactively modify appearance of the car and features by selecting different colours, trims, or additional accessories directly within the 3D scene.Optionally, the method further comprising:- receiving a third input indicative of a user’s interaction with at least one configuration element from amongst the plurality of configuration elements;- generating an updated 3D scene comprising a second representation of the 3D object, wherein the second representation is generated by applying at least one visual effect to the first representation, and wherein at least one visual effect corresponds to the user’s interaction with at least one configuration element; and- displaying the updated 3D scene on the interactive user interface.In this regard, the term "third input" refers to a user-generated action that indicates interaction with at least one configuration element within the interactive user interface. The second representation is an updated visual depiction of the 3D object within the updated 3D scene that reflects changes or modifications applied as a result of the user interaction with at least one configuration element. The user provides the third input, for example, by clicking on an interactive play button at the interactive user interface, selecting an option from a dropdown menu displayed via the interactive user interface, tapping at a specific area of a screen of the user device 204, or similar. Then, at least one processor 202 processes the third input to determine which configuration element has been interacted with and interprets a result of the user's interaction. At least one visual effect is applied to the first representation of the 3D object. The updated 3D scene, which includes at least one visual effect, is then rendered and presented on the interactive user interface. This updated 3D scene enables the user to assess and view the effects of their interactions in real-time. It will be appreciated that incorporating the third input to update the 3D scene provides an interactive and responsive user experience. By reflecting the user interactions in real-time through the second representation, the user can instantly visualize the impact of their modifications, facilitating a more intuitive and engaging design process. This approach not only enhances the user satisfaction by offering immediate feedback but also allows for efficient and precise adjustments, improving the overall effectiveness of customization tool.Optionally, at least one visual effect comprises at least one of: changing a color of at least a portion of the three-dimensional object, changing a part of the three-dimensional object, adding a new part to the three-dimensional object, removing a part of the three-dimensional object, changing a size of a part of the three-dimensional object, changing a finish of a part of the three-dimensional object, resizing the three-dimensional object, changing a variant of the three-dimensional object.In this regard, by customizing the 3D object by way of applying different visual effects, the method enables the user to achieve a high level of personalization and visual detail while configuring the 3D object, which is essential for applications where precise and varied adjustments are necessary, such as in product design, virtual simulations, interactive visualizations, and similar. In an example, at least one configuration element may enable the user to click on a color palette slider to change a colour of the 3D object (for example, such as a car). Moreover, at least one configuration element (such as sliders or buttons) may enable the user to modify attributes like a wheel design of the car or interior features of the car. At least one configuration element may enable the user to modify specific parts of the 3D object, for example, such as replacing a headlight of the car or adjusting a grille design of the car. Such an adjustment alters a visual appearance of the car by updating the selected part to reflect a new design of the car. Furthermore, at least one configuration element may enable the user to enhance the 3D object by adding the new part, for example, equipping the car with additional accessories like a roof rack or spoilers. The user can interact with at least one configuration element to remove existing components from the 3D object, for example, such as detaching side mirrors or bumpers from the car. Such an action updates a visual representation of the car by eliminating the selected part, which helps in refining or simplifying the design of the car. At least one configuration element may enable the user to adjust a size of the part of the 3D object, for example, such as resizing wheels of the car or adjusting dimensions of windows of the car. Such a customization modifies a scale of individual components to better fit user preferences or functional requirements. At least one configuration element may enable the user to alter a finish of a part of the 3D object, for example, such as changing a texture or a material of a roof the car. This adjustment updates an appearance by applying different finishes, such as matte or glossy, to enhance visual appeal or match design themes. At least one configuration element may enable the user to resize the 3D object, for example, such as increasing or decreasing the overall dimensions of the car. This resizing action adjusts the scale of the entire object, enabling it to fit within different contexts or design specifications. A configuration element may enable the user to select different variants of the 3D object, for example, such as switching between different models or versions of the car with distinct features or configurations. This change updates the visual representation to reflect the chosen variant, providing the user with various options to explore and customize. Thus, different interaction elements correspond to different visual effects, and the 3D scene is continuously updated to reflect latest customizations, ensuring that the user receives immediate and accurate visual feedback. A technical effect of applying at least one visual effect is that it provides the user with a high degree of flexibility and precision in customizing the 3D object, enabling a comprehensive and adjusted configuration experience.Optionally, the updated 3D scene further comprises a second switching element, and wherein the method further comprises:- receiving a fourth input indicative of activation of the second switching element;- generating a second video representing the second representation of the 3D object, in real-time upon receiving the fourth input; and- displaying the second video on the interactive user interface.In this regard, the term "fourth input" refers to an input that triggers an activation of the second switching element within the interactive user interface. The second video is a sequence of image frames representing the second representation of the 3D object. Optionally, the second video comprises a plurality of second image frames. Upon generating the updated 3D scene, when the user interacts with the second switching element in the updated 3D scene, the fourth input is received indicating that the user wants to switch to a presentation mode from a configurator mode. Thus, the second video representing the second representation of the 3D object with user-customized configurations applied to the 3D object, is generated by at least one processor and is displayed on the interactive user interface in real time or near-real time. The second representation reflects customizations made by the user in the configurator mode, such as modified colors, textures, and shapes. The second video is generated in real-time and consists of the plurality of image frames that visually present the 3D object, reflecting customizations made by the user from different perspectives or angles. It will be appreciated that generating the second video in real-time ensures that users can instantly visualize and review customizations they applied to the 3D object, offering immediate feedback and enhancing user satisfaction. Additionally, an ability to seamlessly switch from the configurator mode to the presentation mode provides an intuitive user experience, allowing the user to engage with the 3D object dynamically. By displaying the second video, an immersive and continuous viewing experience is provided to the users, showcasing the customized 3D object from various angles, which aids in better decision-making and overall interaction quality. Continuing with the first example, the user may start by interacting with the 3D model of the car in the configurator mode. Using the plurality of configuration elements, such as sliders and dropdown menus, the user may customize the car by changing its colour, applying a matte texture, and selecting different wheel designs. Once done with the modifications, the user may notice a "Present My Design" button, which serves as the second switching element displayed in the updated 3D scene. Upon clicking said button, the second video is generated in real-time. The second video captures the aforesaid customized car from various angles, showcasing the applied changes such as the new color, the matte texture, and new wheel designs.Optionally, a duration of the second video lies in a range of a first time duration to a total time duration of the first video, wherein the first time duration is equal to a difference between the total time duration of the first video and an elapsed time duration of the first video. In this regard, when the object configurator is displayed after an entirety of the first video was played back, the duration of the second video is equal to the total time duration of the first video. However, when the object configurator is displayed upon receiving the second input during the playback of the first video, the second video would resume from a time point at which the second input was received. Thus, a time duration of the second video should be a remaining time of the un-played part of the first video. By ensuring that the duration of the second video lies within the range of the first time duration to the total time duration of the first video allows for an uninterrupted and cohesive user experience, as the second video can seamlessly integrate with or follow the first video’s playback, whether it is played in full or partially. It will be appreciated that the said method optimizes the viewing experience by making efficient use of the available video content, reducing redundancy and keeping focus of the user on the configured 3D object within an allocated time frame. Continuing with the first example, the first video may be a promotional video showcasing different features of the car may be displayed to the user, wherein the duration of the first video is 5 minutes. During the playback of the first video, the user may choose to customize an appearance of the car using the object configurator. The elapsed time duration of the first video may be 2 minutes, when the user begins configuring the car. In this regard, a remaining time duration of the first video would be 3 minutes. Upon using the object configurator, the duration of the second video (which represents the customized car) should have a duration that fits within this remaining time. Therefore, the duration of the second video is adjusted to a range from the remaining 3 minutes to a complete duration of 5 minutes, depending on when the second video is displayed on the interactive user interface. When the object configurator is displayed immediately after an entirety of the first video has been played, the second video would have a duration of 5 minutes. However, when the object configurator is displayed during a playback of the first video (for example, after 2 minutes of playing the first video), the second video would dynamically adjust its timing / playback to reflect a remaining 3 minutes, ensuring that the visual representation of the customized car complements a remaining time of the first video. A technical effect of setting the duration of the second video to lie in the aforesaid range is that it facilitates in providing a seamless and coherent viewing experience, as the duration of the second video is aligned with a state of the playback of the first video.Optionally, the step of generating the second video comprises employing a generative artificial intelligence model for generating a plurality of second image frames, based on the second representation of the 3D object and one or more first image frames amongst a plurality of first image frames of the first video. In this regard, the term "generative artificial intelligence model" refers to a machine learning (ML) model that is designed to autonomously generate new content (for example, such as images, text, videos, audio, and similar) by learning patterns and features from training data provided to the ML model.Optionally, an input of the generative artificial intelligence (AI) model comprises the second representation of the 3D object and the one or more first image frames from the first video, whereas an output of the generative AI model comprises the second video. In this regard, the second representation of the 3D object represents an updated, customized configuration of the 3D object based on the user's interaction, and the one or more first image frames from the first video serve as base frames, for training the generative AI model to generate second image frames of the second video. The generative AI model may adjust the second representation of the 3D object to match lighting, shadows, and / or spatial details present in the first video. As a result, the second video is generated either by replacing the 3D object in relevant first image frames or by generating entirely new second image frames. It will be appreciated that the generative AI model enables real-time generation of the second video, ensuring that customized changes to the 3D object are immediately reflected without perceptible delays, thereby enhancing the user experience. It will also be appreciated that the generative AI model approach reduces computational complexity by leveraging only necessary one or more first image frames, rather than requiring the entire video to be re-rendered from scratch.In an example, the first video may comprise 500 first image frames. In one case, when the third input for switching from the first video to the object configurator is received while displaying the 420th first image frame of the first video, a second video comprising 80 second image frames may be generated using remaining 80 first image frames by replacing a representation of the 3D object in said remaining 80 first image frames with the second representation of the 3D object. Additionally, the second video may further comprise 420 second image frames which are generated using previously-displayed 420 first image frames and the second representation of the 3D object. Such 420 second image frames can be played back after previously-generated 80 second image frames have been displayed. In another case, when the total time duration of the first video has elapsed before the step of displaying the 3D scene of the object configurator, the second video may comprise 500 second image frames, which are generated by replacing a representation of the 3D object in said 500 first image frames with the second representation of the 3D object.Optionally, the method further comprising optimizing at least one of: the first video, the second video, for streaming to the user device 204, by employing at least one of: a video compression technique, a bandwidth-based resolution adjustment technique, a variable bitrate encoding technique, an adaptive streaming technology, a frame rate setting, a graphics optimization technique, a content delivery network. In this regard, the aforesaid step of optimizing at least one of: the first video, the second video, is only performed when at least one processor is external to the user device (i.e., at least one processor is not a part of the user device). The technical benefit of optimizing at least one of: the first video, the second video, for streaming to the user device 204 is that it ensures a quick loading and a smooth playback across various user devices. This is because by employing at least one of the aforesaid techniques, a bandwidth usage and a latency during transmission can be minimised. This potentially improves both a streaming efficiency and a device performance, resulting in a smoother user experience with lower latency and faster transitions between visual content.The "video compression technique" is a technique that is employed to reduce a size of a given video by encoding data, thereby allowing for a fast transmission and storage while maintaining an acceptable level of visual quality. The video compression techniques could, for example, be High Efficiency video Coding (HEVC), H.264, VP9, or similar. The "bandwidth-based resolution adjustment technique" is a technique used to dynamically alter a resolution of a video based on an available network bandwidth. The "variable bitrate encoding technique" is a technique for video encoding that adjusts a bitrate dynamically, based on a complexity of content of video and a visual quality level. The "adaptive streaming technology" is a method of delivering video content over the internet that automatically adjusts a quality of the video stream in real-time based on an available bandwidth and device capabilities. The "frame rate setting" refers to a specification of a number of individual frames or images displayed per second in a video. Typically, the frame rate setting is measured in frames per second (fps). By adjusting a frame rate of the video based on available bandwidth and capabilities of the user device 204, the frame rate setting can ensure efficient video delivery while minimizing buffering and lag. Additionally, the frame rate setting reduces unnecessary data transmission during less dynamic moments, leading to lower bandwidth usage and improved streaming efficiency. The "graphics optimization technique" is a technique that is designed to improve a visual performance and a rendering efficiency of graphics content in digital media. The graphics optimization technique may include methods but are not limited to, reducing polygon counts in 3D models, applying texture compression, optimizing shaders, and implementing Level of Detail (LOD) techniques. The "content delivery network" is a distributed network of servers strategically located in various geographical regions, designed to efficiently deliver digital content such as videos, images, web pages, and similar to different user devices. For example, when a user sends a request for streaming at least one of: the first video, the second video, the content delivery network (CDN) directs the request to a nearest server that hosts at least one of: the first video, the second video, reducing a physical distance that data must travel. This minimizes latency and decreases loading times, resulting in faster video playback.With reference to FIG. 3, illustrated is an exemplary schematic process flow for generating a three-dimensional (3D) scene, in accordance with an embodiment of the present disclosure. is shown, with consistent frame rates even when the 3D models of the plurality of objects is loaded. Herein, the 3D model of the first object 302a is loaded into the 3D scene (not shown for sake of simplicity). Herein, a sequence of image frames representing the 3D scene is displayed at the user device 310 at a frame rate of, for example, 120 frames per second (FPS). At step 312, the 3D model of the second object 302a is also loaded into the (same) 3D scene. Herein, a sequence of other image frames representing the 3D scene (comprising both the 3D models) is displayed at the user device 310 at a frame rate of, for example, 108 FPS. Notably, during said loading, matching model data is shared between the 3D models. Beneficially, due to this, a frame rate loss may only be 10 percent of an original frame rate (namely, from 120 FPS to 108 FPS), as compared to the prior art where a frame rate loss is almost 50 percent of the original frame rate (for example, from 120 FPS to 60 FPS). In an example, when wheel rim designs of both the cars are identified to be same, during said loading of the 3D model of the first object 302a and the 3D model of the second object 302a, a wheel rim design is instantiated once instead of instantiating it individually for both the aforesaid 3D models.Referring to FIG. 4, illustrated is a flowchart of method steps involved in minimizing a reduction in a number of frames per second (FPS) when generating the visual experience for visual comparison of the plurality of objects, in accordance with an embodiment of the present disclosure. At step 402, a first three-dimensional model of a first object is loaded in the visual experience. At step 404, a particular FPS is achieved as per a size and a complexity of the first three-dimensional model. At step 406, loading of a second three-dimensional model of a second object in the visual experience, alongside the first three-dimensional model of the first object, is initialized. At step 408, matching model data between the first three-dimensional model and the second three-dimensional model, is identified. The matching model data between the first three-dimensional model and the second three-dimensional model may be identified by comparing individual model data of the first three-dimensional model and the second three-dimensional model (which are pre-generated three-dimensional models). At step 410, the second three-dimensional model of the second object is loaded alongside the first three-dimensional model of the first object, with a minimal FPS loss (i.e., with a minimal reduction of FPS from the particular FPS achieved at the step 404. It will be appreciated that the aforementioned method may include additional steps, may exclude one or more of the aforementioned steps, or may have a different order of the aforementioned steps, or similar modifications, for the purpose of minimizing a reduction in a number of frames per second (FPS) when generating the visual experience for visual comparison of the plurality of objects. The aforementioned method may be implemented by a system comprising at least one processor that is configured to execute the steps of the aforementioned method.A technical effect of the aforementioned method is that processing resources are shared amongst pre-generated 3D models of a plurality of objects, for loading similar portions of the pre-generated 3D models, for displaying said 3D models adjacently in a manner that a resultant performance is comparable to a performance that is achieved when a single 3D model is displayed. In other words, a technical problem of proportional performance loss in case of multi-3D experience is effectively mitigated by the aforementioned method. The method can beneficially be employed for streamed visual experiences, thereby effectively controlling streaming costs.It will also be appreciated that the aforementioned method is not limited to generation of visual experiences for visual comparison of only two objects, as is given for example in FIG. 8. The aforementioned method can be used when generating visual experiences for visual comparison of more than two objects too, for example, for visual comparison of 3 objects, 4 objects, 5 objects, 6 objects, and so on.The aforementioned method may further comprise one or more steps of: defining a layout of the visual experience; loading the pre-generated three-dimensional models into their corresponding positions in the layout such that distinct portions of the individual model data are used for loading distinct portions of the pre-generated three-dimensional models, whereas the matching model data is shared therebetween for the loading of the pre-generated three-dimensional models.Referring to FIG. 5, illustrated is a schematic flowchart showing minimization of a reduction in a number of frames per second (FPS) when generating the visual experience for visual comparison of a plurality of objects, in accordance with an embodiment of the present disclosure. As shown in view 502, when the first three-dimensional model of the first object (shown for example as a first car 504) is loaded in the visual experience, 120 FPS is achieved as per size and complexity of the first three-dimensional model. Then, the second three-dimensional model of the second object (shown for example as a second car 506 in view 510) is added to the visual experience at step 508. For example, when the second three-dimensional model of the second object is added, it may be detected that tyres, leather textures, window materials, and similar features of the first car 504 and the second car 506 are similar, and thus their rendering data can be shared (and not be re-calculated for the second car 506). This, when loading the first three-dimensional model and the second three-dimensional model, matching model data (i.e., overlapping model data or similar model data) is shared therebetween. Thereafter, as shown in view 510, the first three-dimensional model and the second three-dimensional model are loaded alongside each other in the visual experience, and 110 FPS performance is achieved. By using the aforementioned method, the reduction in FPS is minimized to be around 10 FPS (for example). Without employing the method, the reduction in FPS would have been considerably larger (for example, such as of the order of 60 FPS. The visual experience shown in the view 510 is displayed on a user device (shown for example, as a smartphone).It will be appreciated that in FIG. 5, the plurality of objects are shown for example as cars, but the plurality of objects could optionally include any other objects, such as furniture, scientific models, sculptures, buildings, toys, sports equipment, clothing, and the like.The present disclosure also relates to the system as described above. Various embodiments and variants disclosed above, with respect to the aforementioned method, apply mutatis mutandis to the system.Optionally, at least one server is further configured to:- define at least one first variable that pertains to a complexity of the visual experience;- define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual experience;- generate an adaptive render streaming output by:- obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual experience;- rendering a set of image frames representing the visual experience, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and- determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant,wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and- stream the adaptive render streaming output to one or more user devices, via the communication network, according to the corresponding bit rates.The present disclosure also provides an interactive object configuration experience, a product suite comprising a plurality of visual experiences for interactively viewing, configuring, and comparing three-dimensional (3D) objects, and a method for creating such a product suite. Various embodiments and variants disclosed above, with respect to the aforementioned method and the aforementioned system, apply mutatis mutandis to the interactive object configuration experience, the product suite, and the method for creating such a product suite.Notably, the present disclosure also provides the interactive object configuration experience, wherein a 3D object that is to be configured is presented in a first video representing the 3D object, whilst loading of an object configurator for configuring the 3D object is simultaneously initialized,and wherein a 3D scene of the object configurator is displayed upon detection that a predefined criterion for switching from the first video to the object configurator is satisfied, the 3D scene comprising a first representation of the 3D object and a plurality of configuration elements for configuring the 3D object.It will be appreciated that the interactive object configuration experience is provided in a seamless and efficient manner, by way of displaying the first video that keeps a user engaged until the object configurator is fully loaded and ready. This facilitates in a reduction of a perceived load time on the interactive user interface, thereby enhancing user experience, and also provides a fluid and uninterrupted customization of the 3D object.Notably, the present disclosure also provides the product suite comprising the plurality of visual experiences for interactively viewing, configuring, and comparing the three-dimensional (3D) objects, wherein the plurality of visual experiences comprise:a first visual experience for viewing a 3D object in a presentation mode of the product suite,a second visual experience for configuring the 3D object in a configurator mode of the product suite, wherein an interactive object configuration experience comprises the first visual experience and the second visual experience, and a third visual experience for interactively comparing the 3D object with at least one another 3D object in a 3D comparison mode of the product suite, wherein each visual experience comprises at least one switching element for enabling seamless switching from said visual experience to at least one other visual experience.In this regard, the third visual experience is the interactive visual comparison experience of the aforementioned method and the aforementioned system. The first visual experience enables in viewing the 3D object. In this regard, the presentation mode could, for example, has engaging visuals, swift loading and an all-inclusive showcase of features of the 3D object. The second visual experience enables in configuring the 3D object. In this regard, the configurator mode (namely, a 3D visualiser mode) could, for example, at least showcases life-like 3D views of detailed 3D configurations of the 3D object in real time, and allows for tailored customisation, interior views, and exterior views of the 3D object. It will be appreciated that the aforementioned product suite may be implemented by way of a software product. Such a software product is device agnostic, has a user-friendly-interface, and requires nominal computational resources for loading purposes and for running any of the aforesaid experiences. In an example implementation, the software product may optionally allow a user using the configurator mode to contact a manufacturer of the 3D object, may optionally also provide a lead to a manufacturer of the 3D object, as to user's interest in the 3D object. In another example implementation, the software product may minimise performance losses (for example, minimise reduction in frames per second) even when multiple 3D models of the plurality of objects are presented alongside each other in the third visual experience. It will be appreciated that at least one switching element in the product suite enables in seamless switching (i.e., toggling) between the viewing mode, the configurator mode, and the 3D comparison mode (for example, upon receiving user input).Notably, the present disclosure also provides the method for creating the product suite, the method comprising:creating the first visual experience, the second visual experience, and the third visual experience; andintegrating the first visual experience, the second visual experience, and the third visual experience into a software product, wherein when a user interacts with the software product, said experiences are provided to the user.The aforesaid method for creating the product suite is simple, and can be implemented with ease. Creating and integrating the first visual experience, the second visual experience, and the third visual experience into the software product enables in addressing specific specialised functionalities to provide a comprehensive solution for users (namely, a holistic interactive visual experience to the users).Various methods described could also be incorporated into computer implemented methods. A special purpose computer may be used to implement. In another embodiment a general purpose computer may be configured based on the teachings in the disclosures herein to implement disclosed functionalities. Modifications to embodiments of the present disclosure described in the foregoing are possible without departing from the scope of the present disclosure as defined by the accompanying claims. Expressions such as "including", "comprising", "incorporating", "have", "is" used to describe and claim the present disclosure are intended to be construed in a non-exclusive manner, namely allowing for items, components or elements not explicitly described also to be present. Reference to the singular is also to be construed to relate to the plural.ABSTRACTDisclosed is a method for streaming visual for configuring a three-dimensional (3D) object, method comprising: defining a first variable that pertains to a complexity of the visual; defining a second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual; generating an adaptive render streaming output by: obtaining values of the first variable and the second variable, at one or more time instants of streaming the visual; rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants; and determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant; and streaming the adaptive render streaming output from a server to one or more user devices, via the communication network, according to the corresponding bit rates.FIG. 1A
Claims
I / We Claim:
1. A method for streaming a visual for configuring a three-dimensional (3D) object, the method comprising: defining at least one first variable that pertains to a complexity of the visual; defining at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual; generating an adaptive render streaming output by: obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual; rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant, wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and streaming the adaptive render streaming output from a server to one or more user devices, via the communication network, according to the corresponding bit rates.
2. The method as claimed in claim 1, wherein at least one first variable comprises at least one of: a rendering quality for deep learning super sampling (DLSS), an upscaling quality for DLSS, quantization for pixel values, a foveated rendering scheme, and a resolution.
3. The method as claimed in claim 2, wherein the step of rendering the set of image frames representing the visual, at each time instant, comprises:applying the foveated rendering scheme for adjusting the resolution of the set of image frames.
4. The method as claimed in claim 2, wherein the step of rendering the set of image frames representing the visual, at each time instant, comprises: rendering the set of image frames according to a given value of the rendering quality for DLSS; employing DLSS for enhancing the resolution of the set of image frames, from the given value of the rendering quality for DLSS to a target value of the upscaling quality for DLSS; and applying the quantization for pixel values of pixels in the image frames of the set, according to a value of the quantization.
5. The method as claimed in claim 1, wherein at least one second variable comprises at least one of: a bandwidth, a latency, a data transfer rate, a compression ratio, and a data packet loss ratio.
6. The method as claimed in claim 1, wherein a value of at least one first variable is obtained from at least one of: a rendering engine, a graphics processing unit configured for rendering image frames, an input received from the one or more user devices, and a predefined configuration that is pre-stored at a memory that is communicably coupled with the server.
7. The method as claimed in claim 1, wherein a value of at least one second variable is obtained from at least one of: network monitoring and / or diagnostic data received from a network device in the communication network and communication metrics received from the one or more user devices.
8. The method as claimed in claim 1, further comprising adjusting a value of at least one first variable, based on a value of at least one second variable, prior to implementing the step of rendering the set of image frames representing the visual, at a given time instant.
9. The method as claimed in claim 1, further comprising integrating an extended -reality (XR) plugin into the visual, thereby enabling the visual to be provided as an XR visual.
10. The method as claimed in claim 1, wherein the 3D object is a vehicle.
11. The method as claimed in claims 1, wherein the one or more time instants comprise at least one of: a start time instant of streaming the visual, a time instant of receiving an input from the one or more user devices, a time instant of detecting a change in the network parameters of the communication network, and a time instant of detecting a change of the communication network.
12. The method as claimed in claim 1, wherein the visual for configuring the 3D object is a part of a product suite comprising a plurality of visuals for interactively viewing, configuring, and comparing three-dimensional (3D) objects.
13. A system for streaming a visual for configuring a three-dimensional (3D) object, the system comprising a server configured to: define at least one first variable that pertains to a complexity of the visual; define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual; generate an adaptive render streaming output by: obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual; rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant, wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and stream the adaptive render streaming output to one or more user devices, via the communication network, according to the corresponding bit rates.
14. A non-transitory computer-readable medium storing a set of instructions for streaming a visual for configuring a three-dimensional (3D) object, the set of instructions comprising: one or more instructions that, when executed by one or more processors of a device, cause the device to: define at least one first variable that pertains to a complexity of the visual; define at least one second variable that pertains to network parameters of a communication network that is to be employed for streaming the visual; generate an adaptive render streaming output by: obtaining values of at least one first variable and at least one second variable, at one or more time instants of streaming the visual; rendering a set of image frames representing the visual, at each time instant amongst the one or more time instants, based on the value of at least one first variable at said time instant; and determining a corresponding bit rate to be employed for streaming the set of image frames rendered at each time instant, based on the values of at least one first variable and at least one second variable at said time instant, wherein the adaptive render streaming output comprises one or more sets of image frames dynamically rendered at the one or more time instants along with their corresponding bit rates determined dynamically; and stream the adaptive render streaming output from a server to one or more user devices, via the communication network, according to the corresponding bit rates.