Computer-based methods, computer systems, computer programs (generative AI-based questionnaires and transcript generation for virtual reality environments)
Patent Information
- Application Number
- JP2026015227
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-02-28
- Filing Date
- 2026-02-02
- Publication Date
- 2026-09-09
Smart Images

Figure 2026144993000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates generally to the field of computing, and more specifically to systems for generative AI-based questionnaire and transcript generation for virtual reality (VR) environments. Summary of the Invention Problem to be Solved by the Invention
[0002] Consumer satisfaction surveys are important for keeping consumers engaged by providing valuable insights into consumer perceptions of a company. Consumer satisfaction surveys help evaluate successes and failures across the company. There are several reasons why companies implement consumer satisfaction surveys, including but not limited to: building good relationships with consumers by showing them that their opinions matter; correcting mistakes; identifying which offerings are effective; seeking out new opportunities; building target profiles; tracking progress over time; and making well-informed decisions. Means for Solving the Problem
[0003] According to one embodiment, a method, computer system, and computer program product are provided for generating generative AI-based questionnaires and transcripts for a virtual reality (VR) environment. The method, computer system, and computer program product may include a step of receiving real-time and historical data from one or more sources within the VR environment. The method, computer system, and computer program product may also include a step of generating a personalized questionnaire for a user about the VR experience based on the real-time and historical data using a first generative artificial intelligence (AI) model. The method, computer system, and computer program product may further include a step of generating transcripts for one or more virtual avatars for interacting with the user within the VR environment based on the personalized questionnaire using a second generative AI model. The method, computer system, and computer program product may also include a step of generating video of a virtual dialogue session between the one or more virtual avatars and the user within the VR environment based on the personalized questionnaire and the transcript using a generative adversarial network (GAN). The method, computer system, and computer program product may further include a step of identifying the user's first one or more interactions with the one or more virtual avatars. The method, computer system, and computer program product may also include a step of using the GAN to fit the video of the virtual dialogue session, based on the step of determining, based on the first one or more dialogues, whether it is possible to derive answers exceeding a threshold confidence level to one or more questions presented by one or more virtual avatars according to the transcript. [Brief explanation of the drawing]
[0004] These and other objects, features and advantages of the present invention will become apparent from the following detailed description of exemplary embodiments of the invention, which will be read in conjunction with the accompanying drawings. The examples are for clarity to facilitate understanding of the invention by those skilled in the art, in conjunction with the detailed description; therefore, various features in the drawings are not to exact scale. In these drawings:
[0005] [Figure 1] An exemplary computing environment is shown in at least one embodiment.
[0006] [Figure 2A] This document provides an operational flowchart for generating AI-based questionnaires and transcripts for a virtual reality (VR) environment in a VR questionnaire and transcript generation process according to at least one embodiment. [Figure 2B] This document provides an operational flowchart for generating AI-based questionnaires and transcripts for a virtual reality (VR) environment in a VR questionnaire and transcript generation process according to at least one embodiment.
[0007] [Figure 3] This is an exemplary diagram showing a user interacting with a product in a VR environment according to at least one embodiment.
[0008] [Figure 4] This is an exemplary diagram showing a user providing responses to a questionnaire through interaction in at least one embodiment. [Modes for carrying out the invention]
[0009] Detailed embodiments of the claimed structure and method are disclosed herein; however, it should be understood that the disclosed embodiments are merely illustrative of the claimed structure and method, which may be embodied in various forms. Nevertheless, the present invention may be embodied in many different forms and should not be construed as being limited to the exemplary embodiments described herein. In the description, well-known features and technical details may be omitted to avoid unnecessarily obscuring the presented embodiments.
[0010] Unless otherwise explicitly indicated by the context, the singular forms "a," "an," and "the" should be understood to include multiple referents. Therefore, unless otherwise explicitly indicated by the context, for example, a reference to "component surfaces" includes a reference to one or more such surfaces.
[0011] Embodiments of the present invention relate to the field of computing, and more specifically to a system for generating AI-based questionnaires and transcripts for virtual reality (VR) environments. Exemplary embodiments described below provide, among other things, a system, method, and program product for generating video of virtual dialogue sessions between a virtual avatar and a user in a VR environment based on personalized questionnaires and transcripts using a generative adversarial network (GAN), and accordingly for adapting the generated video of the virtual dialogue sessions based on one or more interactions between the user and the virtual avatar using the GAN. Thus, these embodiments have the ability to enhance VR technology by adjusting the appearance of the VR interface, thereby facilitating a more flexible VR experience by minimizing the need for verbal conversation between the user and the avatar. Additionally, these embodiments have the ability to improve the graphical user interface (GUI) by dynamically adjusting product offerings to the user in the VR environment, thereby reducing the need for the user to drill down through many layers to obtain desired products.
[0012] As mentioned earlier, consumer satisfaction surveys are crucial for keeping consumers engaged by providing valuable insights into how consumers perceive a company. They help evaluate the successes and failures of the company. There are several reasons why companies implement consumer satisfaction surveys, including, but are not limited to, building good relationships with consumers by demonstrating that their opinions matter, correcting mistakes, recognizing which offerings are effective, exploring new opportunities, building targeted profiles, tracking progress over time, and making informed decisions. Traditional methods of obtaining consumer feedback in VR environments are often impersonal and inefficient. This problem is typically addressed by providing consumers with standardized, generic questionnaires after a transaction is complete. However, standardized, generic questionnaires fail to elicit specific and meaningful responses from users.
[0013] Therefore, it may become essential to implement a system that generates personalized questionnaires to collect information in real time during the VR experience.
[0014] According to at least one embodiment, a computer-based method, computer system, and computer program product are provided for generating generative AI-based questionnaires and transcripts for a VR environment. The method comprises the steps of: receiving real-time and historical data from one or more sources in a VR environment; generating a personalized questionnaire for a user about a VR experience based on the real-time and historical data using a first generative AI model; generating a transcript for one or more virtual avatars to interact with the user in the VR environment based on the personalized questionnaire using a second generative AI model; generating video of a virtual dialogue session between the one or more virtual avatars and the user in the VR environment based on the personalized questionnaire and the transcript using a GAN; identifying the first one or more interactions of the user with the one or more virtual avatars; determining, based on the first one or more interactions, whether answers exceeding a threshold confidence level can be derived for one or more questions presented by the one or more virtual avatars according to the transcript; and, based on the determination that answers can be derived, fitting the video of the virtual dialogue session using the GAN. This embodiment has the advantage of enhancing the VR experience by personalizing questionnaires for individual users.
[0015] According to at least one embodiment, the method may further comprise the steps of: iterating until a response can be derived, based on a determination that the response cannot be derived; using the second generative AI model to generate a modified transcript for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; using the GAN to generate video of a modified virtual dialogue session between the one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the modified transcript; and identifying one or more subsequent interactions of the user with the one or more virtual avatars. This embodiment has the advantage of reducing the possibility of misinterpreting the user's response.
[0016] According to at least one embodiment, the method may further comprise the steps of identifying one or more additional interactions between the user and the one or more virtual avatars in the adapted video, and obtaining one or more responses to the personalized questionnaire based on the additional one or more interactions. This embodiment has the advantage of facilitating a seamless VR experience by minimizing the need for verbal conversation between the one or more virtual avatars and the user.
[0017] According to at least one embodiment, the step of generating the personalized questionnaire for the user may further include the steps of identifying one or more items in which the user is currently engaged within the VR environment and one or more of the user's preferences regarding the VR experience, and tailoring the personalized questionnaire to the user based on the identified one or more items and one or more preferences. This embodiment has the advantage of engaging the user in real time.
[0018] According to at least one embodiment, the step of generating the transcript for the one or more virtual avatars may further include associating a first virtual avatar of the one or more virtual avatars with a first segment of the VR environment, and a second virtual avatar of the one or more virtual avatars with a second segment of the VR environment, wherein at least one of the one or more questions presented by the first virtual avatar corresponds to the first segment of the VR environment, and at least one of the one or more questions presented by the second virtual avatar corresponds to the second segment of the VR environment. This embodiment has the advantage of presenting contextually relevant content to the user to avoid redundancy.
[0019] According to at least one embodiment, the step of adapting the image of the virtual dialogue session may further include a step of adjusting the placement of one or more items in the VR environment based on the user's first one or more interactions using the GAN. This embodiment has the advantage of improving the graphical user interface (GUI) of the VR device by allowing the user to engage with their preferred items without having to perform several steps to display the preferred items.
[0020] According to at least one embodiment, the first one or more interactions may take the form of verbal feedback. Verbal feedback has the advantage of ensuring that the user's preferences are aligned with the VR environment. According to at least one embodiment, the first one or more interactions may take the form of facial expressions. Facial expressions have the advantage of facilitating a seamless and intuitive VR experience. According to at least one embodiment, the first one or more interactions may take the form of body gestures. Body gestures have the advantage of facilitating a seamless and intuitive VR experience.
[0021] Various aspects of this disclosure are described by explanatory text, flowcharts, block diagrams of computer systems and / or block diagrams of machine logic included in computer program product (CPP) embodiments. With respect to any flowchart, depending on the technology involved, operations may be performed in a different order than those shown in a given flowchart. For example, again depending on the technology involved, two operations shown in consecutive flowchart blocks may be performed in reverse order, as a single integrated stage, simultaneously, or with at least partial time overlap.
[0022] Computer program product embodiment ("CPP embodiment" or "CPP") is a term used in this disclosure to describe any set of one or more storage media (also called "mediums") that are collectively comprised of a set of one or more storage devices that collectively contain machine-readable code corresponding to instructions and / or data for performing computer operations specified in a given CPP claim. A "storage device" is any tangible device capable of holding and storing instructions for use by a computer processor. Computer-readable storage media may be, but are not limited to, electronic storage media, magnetic storage media, optical storage media, electromagnetic storage media, semiconductor storage media, mechanical storage media, or any suitable combination thereof. Some known types of storage devices, including these media, include diskettes, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), compact disc read-only memory (CD-ROM), digital versatile disk (DVD), memory sticks, floppy disks, mechanically encoded devices (such as punch cards or pits / lands formed on the main surface of a disk), or any suitable combination of the foregoing. Computer-readable storage media should not be interpreted as storage in the form of transient signals themselves, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through waveguides, optical pulses passing through optical fiber cables, electrical signals communicated through wires, and / or other transmission media, as used in this disclosure.As will be understood by those skilled in the art, data is typically moved at some irregular points during normal operation of a storage device, such as during access, defragmentation or garbage collection, but this does not render the storage device temporary because the data is not transitory while it is stored.
[0023] The exemplary embodiments described below provide systems, methods and program products for generating, by means of a GAN, video of a virtual interaction session between a virtual avatar and a user in a VR environment based on personalized questionnaires and transcripts, and accordingly adapting, by means of a GAN, the generated video of the virtual interaction session based on one or more interactions of the user with the virtual avatar.
[0024] Referring to Figure 1, an exemplary computing environment 100 according to at least one embodiment is shown. The computing environment 100 includes an example of an environment for executing at least some of the computer code involved in performing the method of the present invention, such as a questionnaire generation program 150. In addition to block 150, the computing environment 100 includes, for example, a computer 101, a wide area network (WAN) 102, an end user device (EUD) 103, a remote server 104, a public cloud 105, and a private cloud 106. In this embodiment, the computer 101 includes a processor set 110 (including processing circuits 120 and a cache 121), a communication fabric 111, volatile memory 112, persistent storage 113 (including the operating system 122 and block 150 identified above), a peripheral device set 114 (including a user interface (UI) device set 123, storage 124, and an Internet of Things (IoT) sensor set 125), and a network module 115. The remote server 104 includes a remote database 130. The public cloud 105 includes a gateway 140, a cloud orchestration module 141, a host physical machine set 142, a virtual machine set 143, and a container set 144.
[0025] The computer 101 may be in the form of a desktop computer, a laptop computer, a tablet computer, a smartphone, a smart watch or other wearable computer, a mainframe computer, a quantum computer, or any other form of computer or mobile device that is currently known or will be developed in the future and is capable of executing programs, accessing a network or querying a database such as the remote database 130, for example. As is well understood in the field of computer technology, and depending on the technology, the execution of the computer-implemented method may be distributed among a plurality of computers and / or between a plurality of locations. On the other hand, in this presentation of the computing environment 100, to keep the presentation as simple as possible, the detailed discussion focuses on a single computer, specifically the computer 101. Although the computer 101 is not shown in the cloud in Fig. 1, it may be located in the cloud. On the other hand, the computer 101 is not required to be in the cloud, except to any extent that may be explicitly indicated.
[0026] The processor set 110 includes one or more computer processors of any type that are currently known or will be developed in the future. The processing circuit 120 may be distributed across a plurality of packages, for example, a plurality of coordinated integrated circuit chips. The processing circuit 120 may implement a plurality of processor threads and / or a plurality of processor cores. The cache 121 is a memory located within the processor chip package, and is typically used for data or code that should be available for high-speed access by a thread or core executing on the processor set 110. Cache memory is typically organized into a plurality of levels according to the relative proximity to the processing circuit. Alternatively, some or all of the caches for the processor set may be located "off-chip". In some computing environments, the processor set 110 may be designed to operate using qubits and perform quantum computing.
[0027] Computer-readable program instructions are typically loaded onto computer 101 to cause the processor set 110 of computer 101 to execute a series of operational steps, thereby executing a computer implementation method, and as a result, the instructions thus executed instantiate the method specified in the flowcharts and / or descriptions of the computer implementation method contained herein (collectively referred to as the "Method of the Invention"). These computer-readable program instructions are stored in various types of computer-readable storage media, such as cache 121 and other storage media discussed below. The program instructions and associated data are accessed by the processor set 110 to control and direct the execution of the Method of the Invention. In the computing environment 100, at least some of the instructions for executing the Method of the Invention may be stored in blocks 150 in persistent storage 113.
[0028] The communication fabric 111 is a signal conduction path that enables various components of the computer 101 to communicate with one another. Typically, this fabric is made up of switches and conductive paths, such as buses, bridges, physical input / output ports, and similar components. Other types of signal communication paths, such as optical fiber communication paths and / or wireless communication paths, may be used.
[0029] The volatile memory 112 is any type of volatile memory currently known or to be developed in the future. Examples include dynamic type random access memory (RAM) or static type RAM. Typically, the volatile memory 112 is characterized by random access, but this is not required unless explicitly stated. In computer 101, the volatile memory 112 is located in a single package and is internal to computer 101, but alternatively or additionally, the volatile memory 112 may be distributed across multiple packages and / or located externally to computer 101.
[0030] The persistent storage 113 is any form of non-volatile storage for a computer, currently known or to be developed in the future. The non-volatility of this storage means that the stored data is maintained regardless of whether power is supplied to the computer 101 and / or directly to the persistent storage 113. The persistent storage 113 may be read-only memory (ROM), but typically, at least a portion of the persistent storage 113 allows for writing, deleting, and rewriting of data. Some well-known forms of persistent storage 113 include magnetic disks and solid-state storage devices. The operating system 122 can take several forms, such as various known proprietary operating systems or open-source portable operating system interface type operating systems using a kernel. The code contained in block 150 typically includes at least a portion of computer code involved in performing the method of the present invention.
[0031] The peripheral device set 114 includes a set of peripheral devices for the computer 101. Data communication connections between the computer 101's peripheral device set 114 and other components can be implemented in various ways, including Bluetooth® connections, near-field communication (NFC) connections, connections made by cables (such as Universal Serial Bus (USB) type cables), insert-type connections (e.g., Secure Digital (SD) cards), connections made through local area communication networks, and even connections made through wide area networks such as the Internet. In various embodiments, the UI device set 123 may include components such as a display screen, speakers, microphones, wearable devices (such as goggles and smartwatches), keyboards, mice, printers, touchpads, game controllers, and haptic devices. Storage 124 is external storage, such as an external hard drive, or insertable storage, such as an SD card. Storage 124 may be persistent and / or volatile. In some embodiments, storage 124 may take the form of a quantum computing storage device for storing data in the form of qubits. In embodiments where computer 101 is required to have a large amount of storage (for example, computer 101 locally stores and manages a large database), this storage may be provided by peripheral storage devices designed to store very large amounts of data, such as a storage area network (SAN) shared by multiple geographically distributed computers. The IoT sensor set 125 consists of sensors that can be used in Internet of Things applications. For example, one sensor may be a thermometer, and another may be a motion detector. The peripheral device set 114 may also include, but is not limited to, a VR headset and / or sensors, cameras, and microphones communicatively coupled to the VR headset.
[0032] The network module 115 is a collection of computer software, hardware, and firmware that enables computer 101 to communicate with other computers via the WAN 102. The network module 115 may include hardware such as a modem or Wi-Fi® signal transceiver, software for packetizing and / or depackaging data for communication network transmission, and / or web browser software for communicating data over the Internet. In some embodiments, the network control and network forwarding functions of the network module 115 are performed on the same physical hardware device. In other embodiments (e.g., embodiments utilizing Software-Defined Networking (SDN)), the control and forwarding functions of the network module 115 are performed on physically separate devices, such that the control function manages several different network hardware devices. Computer-readable program instructions for performing the method of the present invention can typically be downloaded to computer 101 from an external computer or external storage device via a network adapter card or network interface included in the network module 115.
[0033] WAN102 is any wide area network (e.g., the Internet) that can communicate computer data over non-local distances using any currently known or future-developed technology for communicating computer data. In some embodiments, the WAN may be replaced and / or complemented by a local area network (LAN), such as a Wi-Fi® network, designed to communicate data between devices located within a local area. WAN102 and / or LAN typically include computer hardware such as copper transmission cables, optical transmission fibers, wireless transmissions, routers, firewalls, switches, gateway computers, and edge servers.
[0034] An end-user device (EUD) 103 is any computer system used and controlled by an end-user (e.g., a customer of the company operating computer 101) and can take any of the forms discussed above in relation to computer 101. Typically, EUD 103 receives useful and valuable data from the operation of computer 101. For example, in a hypothetical case where computer 101 is designed to provide recommendations to an end-user, these recommendations would typically be communicated from the network module 115 of computer 101 to EUD 103 via WAN 102. Thus, EUD 103 can display or otherwise present recommendations to the end-user. In some embodiments, EUD 103 may be a client device, such as a thin client, heavy client, mainframe computer, desktop computer, and similar.
[0035] The remote server 104 is any computer system that provides at least some data and / or functionality to computer 101. The remote server 104 may be controlled and used by the same entity that operates computer 101. The remote server 104 represents a machine that collects and stores useful and valuable data for use by other computers, such as computer 101. For example, in a hypothetical case where computer 101 is designed and programmed to provide recommendations based on historical data, this historical data may be provided to computer 101 from a remote database 130 of the remote server 104.
[0036] The public cloud 105 is any computer system available for use by multiple entities, providing on-demand availability of computer system resources and / or other computer functions, particularly data storage (cloud storage) and computing power, without direct active management by the user. Cloud computing typically leverages resource sharing to achieve coherence and economies of scale. Direct and active management of the computing resources of the public cloud 105 is performed by the computer hardware and / or software of the cloud orchestration module 141. The computing resources provided by the public cloud 105 are typically implemented by virtual computing environments running on various computers that make up the computers of the host physical machine set 142, which is a collection (universe) of physical computers within and / or available to the public cloud 105. The virtual computing environment (VCE) typically takes the form of virtual machines from the virtual machine set 143 and / or containers from the container set 144. These VCEs may be stored as images and are understood to be transferable either as images or after instantiation of the VCEs between various physical machine hosts. The cloud orchestration module 141 manages the transfer and storage of images, deploys new VCE instances, and manages the active instantiation of VCE deployments. The gateway 140 is a collection of computer software, hardware, and firmware that enables the public cloud 105 to communicate over the WAN 102.
[0037] Here, some further explanation of virtualized computing environments (VCEs) is provided. A VCE can be stored as an "image." A new active instance of a VCE can be instantiated from an image. Two well-known types of VCEs are virtual machines and containers. A container is a VCE that uses operating system-level virtualization. This refers to an operating system feature in which the kernel enables the existence of multiple isolated user-space instances called containers. These isolated user-space instances typically behave as real computers in terms of the programs running in them. Computer programs running on a normal operating system can utilize all the resources of that computer, including connected devices, files and folders, network shares, CPU power, and quantifiable hardware capabilities. However, programs running inside a container can only use the contents of the container and the devices allocated to the container; this feature is known as containerization.
[0038] The private cloud 106 is similar to the public cloud 105, except that its computing resources are available only for use by a single enterprise. Although the private cloud 106 is shown communicating with the WAN 102, in other embodiments, the private cloud 106 may be completely isolated from the internet and accessible only through a local / private network. A hybrid cloud is a combination of multiple clouds of different types (e.g., private, community, or public cloud types), often implemented by different vendors. Each of the multiple clouds remains a distinct and isolated entity, but a larger hybrid cloud architecture is bound together by standardized or proprietary technologies that enable orchestration, management, and / or data / application portability between the multiple configured clouds. In this embodiment, both the public cloud 105 and the private cloud 106 are part of a larger hybrid cloud.
[0039] According to this embodiment, the questionnaire generation program 150 may be a program that receives real-time and historical data from one or more sources within the VR environment, generates video of a virtual dialogue session between a virtual avatar and a user within the VR environment based on a personalized questionnaire and transcript using a GAN, and adapts the generated video of the virtual dialogue session based on one or more interactions between the user and the virtual avatar using a GAN. Furthermore, despite its depiction within the computer 101, the questionnaire generation program 150 may be stored and / or executed individually or in any combination on an end-user device 103, a remote server 104, a public cloud 105, and a private cloud 106. The questionnaire generation method will be described in more detail below with respect to Figures 2A, 2B, 3, and 4. It should be understood that the examples described below are not intended to be limiting, and that the parameters used in these examples may differ in embodiments of the present invention.
[0040] Referring here to Figures 2A and 2B, an operational flowchart for generating AI-based questionnaires and transcripts for a VR environment in a VR questionnaire and transcript generation process 200 according to at least one embodiment is shown according to at least one embodiment. In 202, the questionnaire generation program 150 receives real-time and historical data from one or more sources within the VR environment. The VR environment may be a VR shopping environment where a user can purchase products and / or services. Real-time data may include social network data. Examples of social network data may include, but are not limited to, comments, posts, likes, dislikes, emojis, profiles, shared content and / or media (e.g., music and video). Real-time data may also include a user's interaction with one or more virtual avatars within the VR environment, which will be described in more detail below with respect to step 210. As shown in Figure 3, real-time data may further include the item the user is interacting with and the items in the user's virtual cart. For example, if a user is reading the label on a bottle of source, the user may be interacting with this item. Continuing with this example, a box of pasta can be identified in the user's virtual shopping cart.
[0041] The historical data may include, but is not limited to, purchase history, return history, and / or purchase patterns (e.g., repetition of certain items being purchased). The historical data may be stored in a customer database, such as a remote database 130. For example, a purchase pattern may be that a user purchases flour, butter, cocoa powder, and cake mix 7 out of 10 times in the VR environment.
[0042] Next, in step 204, the survey generation program 150 generates a personalized survey for the user regarding the VR experience using a first generative AI model. The personalized survey is generated based on real-time and historical data. The personalized survey regarding the VR experience may include one or more questions about products and / or services offered within the VR environment. It should be understood that "product" and "item" are used synonymously in this specification. The first generative AI model may be a pre-trained generative transformer model. The real-time and historical data described above with respect to step 202 may be included in a Large-Scale Customer Model (LCM). The survey generation program 150 may derive the user's preferences regarding the VR experience based on purchase history, return history and / or purchase patterns. For example, if a user has purchased flour, butter, cocoa powder and cake mix 7 out of 10 times within the VR environment, the user's preferences may include baking. In another example, if a user has returned bottles of ketchup and salad dressing, the user's preferences may not include spices.
[0043] A generative transformer model can be pre-trained with a diverse set of user-based questionnaires to understand patterns and preferences in user responses. The Lifecycle Management Model (LCM) can be fed into the pre-trained generative transformer model. The input data can then be pre-processed to extract relevant information, such as the type and quantity of products purchased and previous purchase patterns. The pre-trained generative transformer model can then generate relevant and engaging questions tailored to the user, taking into account various factors, such as the extracted content within the LCM.
[0044] According to at least one embodiment, the step of generating a personalized questionnaire may include identifying one or more items that the user in the VR environment is currently engaged with, and one or more of the user's preferences regarding the VR experience. For example, the user may be looking at butter and flour on a virtual shelf in the VR environment. Continuing with this example, baking may be one of the user's preferences. The personalized questionnaire may be tailored to the user based on the identified one or more items and one or more preferences.
[0045] For example, a personalized survey could include questions such as: "What type of oil or butter do you typically prefer for cooking or baking? Olive oil, rapeseed oil, butter, margarine or something else (please specify)"; "You currently buy butter, are you considering any supplementary items or ingredients for your baking projects? Vanilla extract, leavening agents, baking soda, eggs, milk, nuts or something else (please specify)"; and "Are you interested in looking for any particular brand or type of flour or sugar for your baking recipes? All-purpose flour, whole wheat flour, granulated sugar, brown sugar, organic varieties, gluten-free options or something else (please specify)."
[0046] Personalized questionnaires can be stored in a customer database along with user profile information. Over time, a pre-trained generative transformer model can refine the question generation process and better adapt to user preferences by continuously learning from user interactions and feedback.
[0047] Next, in 206, the questionnaire generation program 150 generates transcripts for one or more virtual avatars for interacting with the user in the VR environment using a second generative AI model. These transcripts are generated based on a personalized questionnaire. The second generative AI model may be a Bi-LSTM (Bi-LSTM) model. The personalized questionnaire may be supplied to the Bi-LSTM model along with the user's preferences. The Bi-LSTM model may use an attention mechanism to focus on relevant parts of the personalized questionnaire and dynamically weight the importance of different parts of the personalized questionnaire. The Bi-LSTM may include an encoder that processes the personalized questionnaire into a fixed-length representation and a decoder that generates responses for one or more virtual avatars. Contextual information about the user's preferences and purchase history may be embedded in the Bi-LSTM model to provide additional information for generating transcripts, thereby ensuring that the generated responses are relevant and coherent based on the user's profile. The Bi-LSTM model may be trained on a dataset of transcripts of everyday conversations between avatars and users. During training, the Bi-LSTM model can optimize its parameters to minimize discrepancies between the generated transcript and ground truth conversation (e.g., loss function) by learning to predict the next word in a conversation based on personalized questionnaires and contextual information, guiding the Bi-LSTM toward more accurate predictions. The Bi-LSTM model can be evaluated on a separate validation dataset. To further improve the accuracy and generalization ability of the model, fine-tuning techniques such as gradient descent, optimization, and normalization may be applied.
[0048] According to at least one embodiment, the step of generating transcripts for one or more virtual avatars may include associating a first virtual avatar of one or more virtual avatars with a first segment of the VR environment, and a second virtual avatar of one or more virtual avatars with a second segment of the VR environment. The segments of the VR environment may be different sections of the VR environment. For example, similar to how products are arranged in the aisles of a grocery store or supermarket, the VR environment may be divided into segments having different types of products. Continuing with this example, the first segment of the VR environment may include an aisle of baking products, and the second segment of the VR environment may include an aisle of hardware products (e.g., kitchen appliances). At least one of the one or more questions presented by the first virtual avatar may correspond to the first segment of the VR environment, and at least one of the one or more questions presented by the second virtual avatar may correspond to the second segment of the VR environment. For example, at least one question corresponding to the first segment may be related to baking (e.g., What items will you be buying for baking?), and at least one question corresponding to the second segment may be related to hardware (e.g., What size and color temperature of light bulb do you prefer?).
[0049] The output of the Bi-LSTM model may be a transcript for one or more avatars to interact with the user.
[0050] For example, the spoken transcript might include questions such as, "Welcome to our VR shopping experience! We understand that everyone has their own preferences when it comes to cooking or baking. Could you share what type of oil or butter you usually prefer? We have options like olive oil, rapeseed oil, butter, and margarine, but please let us know if there's anything else you have in mind"; "Hello! Baking can be such a fun experience, especially when you're using whole flour or sugar. Is there any particular brand or variety you have in mind that you're looking for? We have a wide range of options, including all-purpose flour, whole wheat flour, granulated sugar, brown sugar, and even organic varieties. Let us know which products are your favorites!"; and, "Oh, butter is such a versatile ingredient for baking! While you're here, would you be interested in looking for any other items or ingredients that are perfect for your baking project? We have a variety of options, such as vanilla extract, leavening agents, and eggs. Please let us know if there's anything else you're looking for to enrich your baking experience!"
[0051] Next, in 208, the questionnaire generation program 150 uses a GAN to generate video of a virtual dialogue session between one or more virtual avatars and users in a VR environment. The video is generated based on a personalized questionnaire and transcript. The VR environment, personalized questionnaire, and transcript may be supplied to a GAN, specifically a GAN generator. The GAN may include a GAN generator that synthesizes virtual dialogue video based on the input data. The output of the GAN generator may be supplied as input to a GAN discriminator. In any GAN, the goal of the GAN generator is to deceive the GAN discriminator into classifying artificially generated (i.e., fake) video as real. In addition to supplying the output from the GAN generator to the GAN discriminator, the GAN discriminator is also supplied with unaltered (e.g., real) virtual dialogue video as part of a training session. Next, the GAN discriminator may output a number between 0 and 1, where 0 instructs the GAN discriminator to classify the image as fake, and 1 instructs the GAN discriminator to classify the image as real. For example, if the GAN discriminator receives a virtual dialogue video, it may output a number to classify the virtual dialogue video as real or fake. The GAN discriminator may also provide feedback to the GAN generator to improve the performance of the GAN generator. Additionally, the GAN training process may include optimizing an adversarial loss function that measures the difference between the distributions of real and fake videos. Auxiliary losses, such as coherence, relevance, and visual fidelity, may also be incorporated to favor specific attributes in the generated videos.
[0052] Once trained, a GAN can generate a composite video of a virtual dialogue session between one or more virtual avatars and a user. One or more virtual avatars may be included in the virtual dialogue video generated by the GAN generator. The video of the virtual dialogue session may be displayed to the user in a VR environment. For example, one or more virtual avatars may be standing adjacent to the user in the VR environment. One or more virtual avatars may present one or more questions to the user when the user is interacting with one or more products in the VR environment. One or more questions may be spoken aloud by one or more virtual avatars, or one or more questions may be written in front of the user's field of view in the VR environment. One or more questions may be presented to the user by one or more virtual avatars according to a transcript.
[0053] For example, the transcript might say, "Welcome to our VR shopping experience! We understand that everyone has their own preferences when it comes to cooking or baking. Could you share what type of oil or butter you usually prefer? We have options like olive oil, rapeseed oil, butter, and margarine, but please let us know if there's anything else you have in mind."; "Hello! Baking can be a very enjoyable experience, especially when you're using whole flour or sugar. Is there any particular brand or variety you have in mind that you're looking for? We have..." These questions may be presented to the user if they include questions such as, "We offer a wide range of options, including all-purpose flour, whole wheat flour, granulated sugar, brown sugar, and even organic varieties. Let us know which product is your favorite!" and, "Oh, butter is truly a versatile ingredient for baking! While you're here, would you be interested in looking for any other items or ingredients that are perfect for your baking projects? We have a variety of options available, such as vanilla extract, leavening agents, and eggs. Let us know if there's anything you're looking for to enrich your baking experience!"
[0054] Next, in 210, the questionnaire generation program 150 identifies the user's first one or more interactions with one or more virtual avatars. As used herein, “first one or more interactions” refers to the user's first one or more interactions with one or more virtual avatars. There may be several examples in which transcripts are modified and / or video is adapted, as will be described in more detail below. Examples of the first one or more interactions may include, but are not limited to, verbal feedback from the user, the user’s facial expressions, and / or bodily expressions made by the user.
[0055] Sensors, microphones, and / or cameras may be attached to the user's VR device to capture the first or subsequent interactions. The user's first or subsequent interactions may consist of responses to one or more questions presented by one or more virtual avatars.
[0056] For example, if the avatar asks, "Could you tell us what type of oil or butter you usually prefer? We have options like olive oil, rapeseed oil, butter, and margarine, or if there's anything else you have in mind," the user might respond verbally, "I like olive oil and margarine." If the avatar asks, "Is there any particular brand or variety you have in mind that you're looking for? We have a wide range of options, including all-purpose flour, whole wheat flour, granulated sugar, brown sugar, and even organic varieties," the user might respond with a "thumbs up" gesture (e.g., a physical expression), as shown in Figure 4. If the avatar asks, "While you're here, are you interested in looking for any other items or ingredients that would be perfect for your baking project? We have a variety of options, such as vanilla extract, leavening agents, and eggs," the user might respond with a smile (e.g., a facial expression), as shown in Figure 4.
[0057] Next, in step 212, the questionnaire generation program 150 determines whether a response exceeding a threshold confidence level can be derived for one or more questions presented by one or more avatars according to the transcript. This determination is based on the first one or more interactions. As described above with respect to step 210, one or more virtual avatars may present one or more questions to the user according to the transcript, and the user may respond to one or more questions by interacting with one or more virtual avatars. There may be certain cases in which a response exceeding a threshold confidence level (e.g., 50%) cannot be derived from the user's response.
[0058] For example, if the avatar asks, "Could you tell us what type of oil or butter you usually prefer? We have options like olive oil, rapeseed oil, butter, and margarine, or if there's anything else you have in mind," the user might respond with the words, "I like olive oil and margarine." In this example, the user explicitly states their preference, so the answer can be derived with 100% confidence. If the avatar asks, "Is there any particular brand or variety you have in mind that you're looking for? We have a wide range of options including all-purpose flour, whole wheat flour, granulated sugar, brown sugar, and even organic varieties," the user might respond with a "thumbs up" gesture. In this example, it's not clear which specific product the user is referring to with the "thumbs up" gesture, so the answer can be derived with 30% confidence. If the avatar asks, "While you're here, would you be interested in looking for any other items or ingredients perfect for your baking project? We have a variety of options available, such as vanilla extract, baking powder, and eggs," the user might respond with a smile. In this example, since it's not clear which specific product the user is smiling at, the answer can be derived with 40% confidence.
[0059] In response to a determination that a response exceeding the threshold confidence level can be derived (step 212, branch to "yes"), the VR questionnaire and transcript generation process 200 proceeds to step 214 to adapt the video of the virtual dialogue session. In response to a determination that a response exceeding the threshold confidence level cannot be derived (step 212, branch to "no"), the VR questionnaire and transcript generation process 200 returns to step 206 to generate a revised transcript for one or more virtual avatars to interact with the user in the VR environment.
[0060] In embodiments where a response exceeding a threshold confidence level cannot be derived, it may be understood that steps 206, 208, and 210 may be repeated until a response can be derived. A second generative AI model may generate a modified transcript for one or more virtual avatars to interact with the user in a VR environment, based on a personalized questionnaire. The modified transcript may include paraphrasing of one or more questions in the original transcript. For example, the original question, "Is there any particular brand or variety you have in mind that you are looking for? We offer a wide range of options, including all-purpose flour, whole wheat flour, granulated sugar, brown sugar, and even organic varieties," may be paraphrased as, "Is there any particular brand or variety you have in mind that you are looking for? In particular, do you like all-purpose flour? Do you like whole wheat flour? Do you like granulated sugar? Do you like brown sugar? Or do you have any preferred organic varieties?" Continuing with this example, if, after being asked "Do you like whole wheat?", the user smiles or makes a "thumbs up" gesture, then the smile or the "thumbs up" gesture may be associated with whole wheat, and the answer may be derived with 80% confidence.
[0061] A GAN can generate video of a modified virtual dialogue session between one or more avatars and a user in a VR environment, based on a personalized questionnaire and a modified transcript. For example, the video of the modified virtual dialogue session may include one or more virtual avatars presenting the user with the paraphrased questions in the example described above. Then, one or more subsequent interactions of the user with one or more virtual avatars can be identified. As used herein, “one or more subsequent interactions” refers to one or more interactions of the user with one or more virtual avatars in which one or more virtual avatars present the user with the paraphrased questions. For example, if there was a smile in the response to the original question and a “thumbs up” gesture in the response to the paraphrased question, the smile may be part of the first or more interactions, and the thumbs-up gesture may be part of the subsequent or more interactions.
[0062] Next, in 214, the questionnaire generation program 150 uses a GAN to fit the video of the virtual dialogue session. The user's first one or more interactions may be supplied to the GAN, which may provide output based on this updated data. This output may be a fitted video. The fitted video may include some or all of the content of the generated original video, as well as one or more adjustments to the appearance of the generated original video. The fitted video of the virtual dialogue session may be displayed to the user in a VR environment. The GAN may dynamically adjust the color and layout of the VR environment in the fitted video. Additionally, the GAN may adjust the placement of one or more items in the VR environment based on the user's first one or more interactions.
[0063] For example, blue may be the user's favorite color, determined by the user's social network data. The adapted video may include making the background of the VR environment blue and / or presenting products in blue. In another example, in the generated original video, all-purpose flour, whole wheat flour, granulated sugar, and brown sugar may be presented to the user on a shelf in the VR environment. If it is determined, based on verbal, facial, and / or bodily feedback, that the user prefers whole wheat flour, the whole wheat flour may be moved to the user's eye level on the shelf and made available for purchase. Additionally, or alternatively, removing all-purpose flour, granulated sugar, and brown sugar from the shelf may simplify the content displayed to the user, allowing the user to focus only on their preferred products.
[0064] According to at least one embodiment, if no answer can be derived and a video of a modified virtual dialogue session is generated, the questionnaire generation program 150 may adapt the video of the modified virtual dialogue session as described above, the difference being that the adapted video version of the modified dialogue session may present the user with rephrased questions.
[0065] Next, in 216, the questionnaire generation program 150 identifies one or more additional interactions of the user with one or more virtual avatars in the adapted video. As used herein, “one or more additional interactions” refers to one or more interactions of the user with one or more virtual avatars after the video has been adapted. For example, if whole wheat flour is placed on the shelf at eye level and other products are removed, the user may nod their head up and down. In this example, the nodding of the head may be part of one or more additional interactions. In another example, if all-purpose flour, granulated sugar, and brown sugar are removed from the shelf, the user may indicate dissatisfaction with the removal of the items by making a “thumbs down” gesture. According to at least one embodiment, if the user is dissatisfied with the removal of an item, the item may reappear on the shelf and become available for purchase.
[0066] Next, in step 218, the survey generation program 150 obtains one or more responses to the personalized survey. One or more responses are obtained based on one or more additional interactions. Once preferred items are confirmed based on one or more additional interactions, one or more responses may be obtained. The obtained one or more responses may be mapped to the personalized survey and stored in a customer database linked to the user profile. For example, a user response obtained to the question, "Do you particularly like all-purpose flour? Do you like whole wheat flour? Do you like granulated sugar? Do you like brown sugar? Or do you have any preferred organic varieties?" might be "I like all-purpose flour." In this example, the response "I like all-purpose flour" may be mapped to the survey and stored in the customer database.
[0067] Referring here to Figure 3, an exemplary diagram 300 is shown illustrating a user 302 engaging with a product in a VR environment according to at least one embodiment. The user 302 may be wearing a VR headset 304. Through the VR headset 304, the user 302 may be able to see a plurality of floating icons 306. The plurality of floating icons 306 may include, but are not limited to, a cart icon, a home icon, a mobile phone icon, and / or a search icon, and through these icons, the user 302 can navigate to different segments of the VR environment. The user 302 may be interacting with one or more items 308 in the VR environment. The user 302 may select one or more items 308 for purchase by placing one or more items into a virtual shopping cart 310. In embodiments of the present invention, one or more items 308 in which the user is interacting, and other items already in the virtual shopping cart 310, may be used to tailor a personalized questionnaire to the user 302's preferences.
[0068] Referring here to Figure 4, an exemplary diagram 400 is shown illustrating a user 302 (Figure 3) providing answers to a questionnaire through interaction according to at least one embodiment. In diagram 400, a camera 402 and a microphone 404 may be used to capture one or more interactions 406, 408, 410 of user 302 (Figure 3). The camera 402 may capture body gestures 408 and / or facial expressions 410. The microphone 404 may capture verbal feedback 406 of user 302 (Figure 3). One or more interactions 406, 408, 410 may be used to derive answers to one or more questions in a personalized questionnaire.
[0069] It should be understood that Figures 2A, 2B, 3, and 4 provide only one example of an implementation and do not imply any limitation on how different embodiments may be implemented. Many modifications to the shown environment may be made based on design and implementation requirements.
[0070] The descriptions of various embodiments of the present invention are presented for illustrative purposes only and are not intended to be exhaustive or limitful to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein has been selected to best describe the principles, practical applications, or technological improvements over commercially available technologies of the embodiments, or to enable other those skilled in the art to understand the embodiments disclosed herein.
Claims
1. A computer-based method for generating AI-based questionnaires and transcripts for virtual reality (VR) environments, The stage of receiving real-time and historical data from one or more sources within a VR environment; A step in which a first generative artificial intelligence (AI) model generates a personalized questionnaire for the user regarding the VR experience based on the real-time and historical data; A second generative AI model generates transcripts for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; A step of generating video of a virtual dialogue session between one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the transcript, using a Generative Adversarial Network (GAN); A step of identifying the first one or more interactions of the user with the one or more virtual avatars; A step of determining, based on the first one or more interactions, whether an answer exceeding a threshold confidence level can be derived for one or more questions presented by the one or more virtual avatars in accordance with the transcript; and Based on the determination that the above answer can be derived, the GAN is used to adapt the video of the virtual dialogue session. A computer-based method comprising the following features.
2. A step of repeating the process until the aforementioned answer can be derived, based on the determination that the aforementioned answer cannot be derived; The second generation AI model generates a modified transcript for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; The GAN generates video of a modified virtual dialogue session between the one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the modified transcript; and Step of identifying one or more subsequent interactions of the user with the one or more virtual avatars. The computer-based method according to claim 1, further comprising:
3. The steps include identifying one or more additional interactions of the user with the one or more virtual avatars in the adapted video; and The step of obtaining one or more responses to the personalized questionnaire based on the aforementioned additional one or more interactions. The computer-based method according to claim 1, further comprising:
4. The step of generating the personalized questionnaire for the user is: A step of identifying one or more items in which the user is currently engaged within the VR environment, and one or more of the user's preferences regarding the VR experience; and The step of tailoring the personalized questionnaire to the user based on the identified one or more items and one or more preferences. It further possesses, The computer-based method according to claim 1.
5. The step of generating the transcript for one or more virtual avatars is: The step of associating the first virtual avatar of the one or more virtual avatars with the first segment of the VR environment, and the second virtual avatar of the one or more virtual avatars with the second segment of the VR environment, wherein at least one of the one or more questions presented by the first virtual avatar corresponds to the first segment of the VR environment, and at least one of the one or more questions presented by the second virtual avatar corresponds to the second segment of the VR environment. It further possesses, The computer-based method according to claim 1.
6. The step of adapting the video of the virtual dialogue session is: The GAN adjusts the placement of one or more items in the VR environment based on the user's initial one or more interactions. It further possesses, The computer-based method according to claim 1.
7. The computer-based method according to any one of claims 1 to 6, wherein the initial one or more interactions are selected from the group consisting of verbal feedback, facial expressions, and body gestures.
8. One or more processors, one or more computer-readable memories, one or more computer-readable tangible storage media, and program instructions stored in at least one of the one or more computer-readable tangible storage media for execution by at least one of the one or more processors via at least one of the one or more computer-readable memories. Equipped with, The stage of receiving real-time and historical data from one or more sources within a VR environment; A step in which a first generative artificial intelligence (AI) model generates a personalized questionnaire for the user regarding the VR experience based on the real-time and historical data; A second generative AI model generates transcripts for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; A step of generating video of a virtual dialogue session between one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the transcript, using a Generative Adversarial Network (GAN); A step of identifying the first one or more interactions of the user with the one or more virtual avatars; A step of determining, based on the first one or more interactions, whether an answer exceeding a threshold confidence level can be derived for one or more questions presented by the one or more virtual avatars in accordance with the transcript; and Based on the determination that the above answer can be derived, the GAN is used to adapt the video of the virtual dialogue session. A computer system capable of performing methods including those mentioned above.
9. The aforementioned method, A step of repeating the process until the aforementioned answer can be derived, based on the determination that the aforementioned answer cannot be derived; The second generation AI model generates a modified transcript for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; The GAN generates video of a modified virtual dialogue session between the one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the modified transcript; and Step of identifying one or more subsequent interactions of the user with the one or more virtual avatars. The computer system according to claim 8, further comprising:
10. The aforementioned method, The steps include identifying one or more additional interactions of the user with the one or more virtual avatars in the adapted video; and The step of obtaining one or more responses to the personalized questionnaire based on the aforementioned additional one or more interactions. The computer system according to claim 8, further comprising:
11. The step of generating the personalized questionnaire for the user is: A step of identifying one or more items in which the user is currently engaged within the VR environment, and one or more of the user's preferences regarding the VR experience; and The step of tailoring the personalized questionnaire to the user based on the identified one or more items and one or more preferences. Further including, The computer system according to claim 8.
12. The step of generating the transcript for one or more virtual avatars is: The step of associating the first virtual avatar of the one or more virtual avatars with the first segment of the VR environment, and the second virtual avatar of the one or more virtual avatars with the second segment of the VR environment, wherein at least one of the one or more questions presented by the first virtual avatar corresponds to the first segment of the VR environment, and at least one of the one or more questions presented by the second virtual avatar corresponds to the second segment of the VR environment. Further including, The computer system according to claim 8.
13. The step of adapting the video of the virtual dialogue session is: The GAN adjusts the placement of one or more items in the VR environment based on the user's initial one or more interactions. Further including, The computer system according to claim 8.
14. The computer system according to any one of claims 8 to 13, wherein the initial one or more interactions are selected from the group consisting of verbal feedback, facial expressions, and body gestures.
15. Equipped with program instructions, The aforementioned program instruction is, The stage of receiving real-time and historical data from one or more sources within a VR environment; A step in which a first generative artificial intelligence (AI) model generates a personalized questionnaire for the user regarding the VR experience based on the real-time and historical data; A second generative AI model generates transcripts for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; A step of generating video of a virtual dialogue session between one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the transcript, using a Generative Adversarial Network (GAN); A step of identifying the first one or more interactions of the user with the one or more virtual avatars; A step of determining, based on the first one or more interactions, whether an answer exceeding a threshold confidence level can be derived for one or more questions presented by the one or more virtual avatars in accordance with the transcript; and Based on the determination that the above answer can be derived, the GAN is used to adapt the video of the virtual dialogue session. Executable by a processor capable of performing a method including, Computer program.
16. The aforementioned method, A step of repeating the process until the aforementioned answer can be derived, based on the determination that the aforementioned answer cannot be derived; The second generation AI model generates a modified transcript for one or more virtual avatars to interact with the user in the VR environment, based on the personalized questionnaire; The GAN generates video of a modified virtual dialogue session between the one or more virtual avatars and the user in the VR environment, based on the personalized questionnaire and the modified transcript; and Step of identifying one or more subsequent interactions of the user with the one or more virtual avatars. The computer program according to claim 15, further comprising:
17. The aforementioned method, The steps include identifying one or more additional interactions of the user with the one or more virtual avatars in the adapted video; and The step of obtaining one or more responses to the personalized questionnaire based on the aforementioned additional one or more interactions. The computer program according to claim 15, further comprising:
18. The step of generating the personalized questionnaire for the user is: A step of identifying one or more items in which the user is currently engaged within the VR environment, and one or more of the user's preferences regarding the VR experience; and The step of tailoring the personalized questionnaire to the user based on the identified one or more items and one or more preferences. Further including, The computer program according to claim 15.
19. The step of generating the transcript for one or more virtual avatars is: The step of associating the first virtual avatar of the one or more virtual avatars with the first segment of the VR environment, and the second virtual avatar of the one or more virtual avatars with the second segment of the VR environment, wherein at least one of the one or more questions presented by the first virtual avatar corresponds to the first segment of the VR environment, and at least one of the one or more questions presented by the second virtual avatar corresponds to the second segment of the VR environment. Further including, The computer program according to claim 15.
20. The step of adapting the video of the virtual dialogue session is: The GAN adjusts the placement of one or more items in the VR environment based on the user's initial one or more interactions. Further including, A computer program according to any one of claims 15 to 19.