Automatic generation of a virtual environment
Generative artificial intelligence automates the creation and editing of virtual environments, addressing the need for specialized skills and time by enabling non-specialist users to generate immersive experiences efficiently.
Patent Information
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- ORANGE SA
- Filing Date
- 2024-10-18
- Publication Date
- 2026-04-24
AI Technical Summary
Existing systems for creating virtual environments require specialized skills and manual effort, limiting accessibility to non-specialist users and being time-consuming.
A method utilizing generative artificial intelligence to analyze content, extract contextual and narrative elements, and generate representations, which are then aggregated to create or edit virtual environments, allowing user interaction and customization.
Enables non-specialist users to generate high-quality, immersive virtual environments quickly and intuitively, reducing manual work and enhancing user engagement through interactive and personalized experiences.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Automatic generation of a virtual environment. Technical field.
[0001] This disclosure falls within the domain of virtual environments. More specifically, it relates to a method for automatically generating a virtual environment, a method for editing a virtual environment, a method for using immersive reality, a computer program, an automatic generator of a virtual environment, and an immersive reality system. Previous technique
[0002] The state of the art includes systems for creating virtual environments, particularly for video game or virtual reality applications. These systems require specialized skills to imagine, design, and manually produce the various elements that make up a 3D world, such as objects, background images, soundscapes, and, where applicable, characters and their narration. Artificial intelligence software, such as language models (LLMs) and image generators, allows for the partial automation of these tasks, but still requires setup time and advanced technical knowledge to achieve a satisfactory result.
[0003] In this context, there is a continuing need to simplify and automate the entire process of creating a 3D universe in order to allow non-specialist users to generate such environments without special technical skills and while having control over the appearance and atmosphere of the universe. Summary
[0004] This disclosure improves the situation.
[0005] A method for automatically generating a virtual environment is proposed, the method comprising: an extraction, in the form of a query intended for a generative artificial intelligence, of elements of a content using an analysis by a first artificial intelligence, the elements including at least contextual elements, and a generation by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
[0006] The use of a first artificial intelligence to analyze content and extract contextual elements can make it possible to process complex information with increased accuracy, reducing human error and facilitating the automated management of large volumes of data
[0007] The automatic generation of contextual representations by a second artificial intelligence can provide a gain in terms of speed and flexibility in the creation of rich environments, by automating a process that was formerly manual.
[0008] The aggregation of representations can allow a fluid composition of the elements extracted in the virtual environment, and can make it possible to offer a high-quality immersive experience.
[0009] According to another aspect, a method for editing a virtual environment comprising: is proposed. an extraction, in the form of a query intended for a generative artificial intelligence, of elements of a content using an analysis by a first artificial intelligence, the elements including at least contextual elements, and a generation by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
[0010] The editing process allows for the adaptation of an already generated virtual environment, offering the possibility of modifying elements without having to generate a complete environment from the beginning, which reduces processing and adjustment time.
[0011] Interacting with artificial intelligence to extract contextual elements and generate new representations can facilitate rapid modification cycles.
[0012] According to another aspect, a method of using an immersive reality previously obtained by at least one of the aforementioned generation method and the aforementioned editing method is proposed, the method of use comprising a reproduction of the virtual environment obtained.
[0013] Reproducing a virtual environment allows a user to be immersed in an interactive space. This improves user engagement and makes interactions more natural and intuitive.
[0014] According to another aspect, a computer program is proposed that includes instructions for implementing all or part of a process as defined herein when this program is executed by a processor. According to another aspect, a non-transient, computer-readable recording medium is proposed on which such a program is recorded.
[0015] According to another aspect, an automatic generator of a virtual environment is proposed, the generator comprising: an extractor, in the form of a query intended for a generative artificial intelligence, of elements of a content using an analysis by a first artificial intelligence, the elements including at least contextual elements, and a generator by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
[0016] According to another aspect, an immersive reality system is proposed comprising the aforementioned automatic generator, a device for reproducing the generated virtual environment, and an adapter of the virtual environment generated based on user interaction with the reproduced virtual environment
[0017] The automation of tasks by artificial intelligence modules and the automatic aggregation of generated elements can allow a user to generate a complete virtual environment without needing knowledge of modeling, interactive storytelling, integration, and / or programming. This makes creating an environment more accessible, even to a non-specialist user.
[0018] Even when the user has specialized technical knowledge, obtaining contextual and / or narrative elements through a first artificial intelligence module, combined with the generation of corresponding representations by a second module, can significantly reduce the amount of manual work required to create the environment.
[0019] User input can be used to generate a personalized environment, notably by taking into account at least one user preference or instruction. One possible benefit is to allow the user some control over aspects such as the aesthetics and atmosphere of the virtual world.
[0020] The features described in the following paragraphs may optionally be implemented independently of each other or in combination with each other.
[0021] In one example, the extraction further extracts at least one narrative element from the content, and the generation of representations takes into account at least one narrative element. In another example, the extraction further extracts at least one narrative element from the content, and the aggregation of the generated representations takes into account at least one narrative element.
[0022] Extracting narrative elements, in addition to contextual elements, improves the coherence and richness of generated environments. This can allow for the construction of more immersive scenarios by aligning representations and / or aggregating them. elements with the progression of a story, thus increasing narrative depth.
[0023] In one example, The extraction provides a descriptive graph of the virtual environment to be generated, including the extracted elements, and The generated representations are aggregated according to the descriptive graph.
[0024] Providing a descriptive graph can improve the organization of extracted elements and their relationships within the virtual environment. This can enable better management of interactions, for example between characters, places, and objects, and can facilitate the adaptation of the environment to a story or to a user's actions.
[0025] In one example, the automatic generation process and / or the editing process includes obtaining user input, the user input enabling the content to be obtained.
[0026] Obtaining user input allows for greater customization of the generated or edited virtual environment. The user can specify preferences, such as visual style or elements to be highlighted, making it possible to generate environments tailored to specific needs or customized immersive experiences.
[0027] In an example, the user input includes at least one of the following: the content; part of the content; a link to the content; a link to a portion of the content.
[0028] Providing content in user input can facilitate the direct generation of a virtual environment based on the provided elements, reducing or avoiding intermediate content retrieval. This allows for faster adaptation of the environment from a defined and immediately accessible source.
[0029] Providing part of the content in the user input, for example an extract or a specific section of text, can make it possible to focus the extraction on subsets of the environment, facilitating the creation of specific scenes or elements without processing the entire content.
[0030] Providing in the user input, rather than the content or part of the content, a link to that content or part of the content in the user input may allow the first artificial intelligence to asynchronously access external resources, such as files stored online, reducing or avoiding, at the level of an entity interacting with the first artificial intelligence, local storage and the handling of large files.
[0031] In one example, the content used by the extraction includes at least one of the following: user input; text from a work; audio content; written content; graphic content.
[0032] Each type or nature of content used by extraction has its own advantages; for example, contextual elements reflecting a designer's style can be extracted from graphic content, and such contextual elements can facilitate the creation of a visual environment faithful to that graphic style. The aesthetic consistency of the generated virtual environment can thus be preserved, respecting the unique visual characteristics of the original content.
[0033] When the content used by the extraction includes several types of content of different kinds (text, audio, graphics), various dimensions of a virtual environment can be more easily combined. For example, narrative descriptions in a text can be combined with visual elements of graphic content to facilitate the generation of a richer and more immersive environment.
[0034] When the content used by the extraction includes user input, the user can directly influence the creation of the virtual environment, for example by specifying preferences or guidelines. For instance, a user can specify a theme, style, or specific elements they wish to see reflected in the environment. Including user input in this framework can therefore personalize the experience of generating or editing the virtual environment, allowing the user to play an active role in its creation and to generate or modify an environment to match their specific expectations or needs.
[0035] In one example, the automatic generation process and / or the editing process includes an aggregation of representations to create the virtual environment.
[0036] In one example, the aggregation of representations is carried out in a rendering engine.
[0037] The use of a rendering engine can allow for a smoother and more realistic visualization of the aggregated elements in the virtual environment. This can improve the graphic and / or sound quality of the reproduced environment.
[0038] In one example, the extraction of elements includes sending successive requests to the first artificial intelligence, in which the sending of successive requests includes at least one sending of a request dependent on a previous response from the first artificial intelligence module.
[0039] The chaining of queries can allow for a progressive refinement of the extraction, by increasing the granularity of the extracted elements. This can contribute to enriching the generated or edited environment by taking into account more complete and contextualized information.
[0040] In one example, the representations of the elements include visual and / or sound representations.
[0041] The inclusion of visual representations can create a more immersive environment, with graphic details that help to materialize the extracted elements (places, objects, characters) within the virtual environment. The addition of sound elements can enhance immersion by adding audio effects that correspond to the visual environment and the story. This can, for example, allow for the synchronization of soundscapes (e.g., forest or city sounds) or specific interactions (such as dialogues or object noises), thus improving the feeling of being immersed in the environment.
[0042] In one example, the virtual environment is three-dimensional.
[0043] Generating a three-dimensional virtual environment allows the user to navigate and interact within a realistic simulated space, offering multiple perspectives and more natural interactions. This enhances depth perception, enables better representation of distances, movements, and spatial interactions between objects, thus creating a more convincing and interactive immersive experience.
[0044] In one example, the method of use includes navigation in the virtual environment.
[0045] Navigation within the virtual environment can provide personalized control, tailored to the user's specific preferences and actions. Brief description of the drawings
[0046] Other features, details and advantages will become apparent from reading the detailed description below and from analyzing the accompanying drawings, in which: Fig. 1
[0047] [Fig.1] illustrates an immersive reality system in an example of implementation. Fig. 2
[0048] [Fig.2] illustrates an immersive reality system in an example of implementation. Fig. 3
[0049] [Fig.3] illustrates an element extractor from a content in an example embodiment. Fig. 4
[0050] [Fig.4] illustrates an interaction between a content analyzer and an extractor of content elements in an example implementation. Fig. 5
[0051] [Fig.5] illustrates a descriptive graph of a content in an example of an embodiment. Description of the implementation methods
[0052] The drawings and description can not only serve to better explain the proposed technique, but also contribute to its definition, where appropriate.
[0053] The proposed technique relates to the generation and editing of virtual environments, as well as the use of these environments in immersive experiences.
[0054] The general principle of the proposed technique is based on: extracting information from content (such as text or multimedia), by formulating queries intended for a first artificial intelligence to obtain contextual and / or narrative elements, and the processing of the extracted elements by a second artificial intelligence, which generates representations of these elements.
[0055] The representations can then be aggregated to create a virtual environment.
[0056] The proposed technique also covers a method of editing a pre-existing virtual environment, allowing elements to be modified, added or removed in an already generated environment.
[0057] In addition, the proposed technique also covers a method of using an immersive reality, where a user can interact with the generated or edited virtual environment, either by navigating freely or by following a specific narrative.
[0058] Any suitable hardware and / or software may be used for the practical implementation of the proposed technique. In general, although aspects of the proposed technique may be described in this document as a process, device, system, method, or method, it should be noted that the proposed technique may also cover computer memory that can be connected to a processor possibly connected to a communication interface, the memory storing instructions which, when executed by such a processor, enable the implementation of the processes, devices, systems, procedures, or methods described in this document.
[0059] The proposed technique can be implemented within a cloud computing environment, where the described processes and methods can be distributed across multiple servers. The use of cloud services can, for example, facilitate the use of artificial intelligence for element extraction and representation generation, remote access to the virtual environment and editing tools, and collaboration between multiple connected users.
[0060] Some specific terms are now clarified for a better understanding of the proposed technique.
[0061] A virtual environment refers to a digital representation, which may be graphical, for example three-dimensional or two-dimensional, simulating a space in which a user can interact, either as an observer (for example, in an interactive video) or as an actor (for example, in a video game). The environment may contain objects, characters, and visual and sound elements that simulate a realistic or imaginary experience. For example, a virtual environment may represent a city, a forest, or a fantasy world. A digital twin refers to a virtual environment that is a digital copy of a real or imaginary space.
[0062] Immersive reality is a digital environment in which the user is fully immersed, interacting with visual, auditory, and / or physical elements that create the illusion of presence within that environment. Immersive reality can be simulated through devices such as virtual reality headsets, motion tracking systems, or panoramic screens. In immersive reality, the user can manipulate objects, move through space, and participate in narrative or interactive experiences, either autonomously or by following a defined storyline.
[0063] Virtual environment generation refers to the process by which a virtual environment is created from input data (such as textual descriptions or pre-existing models). This process can be automated, using software, and can involve several steps, such as the creation of landscapes, objects, and characters, as well as the application of textures and sounds. For example, a virtual environment can be generated from a literary work to recreate the places and events described in the story.
[0064] Editing a virtual environment involves modifying or adjusting the elements of a previously created virtual environment. This can include adding or removing objects, changing the arrangement of elements, or altering visual characteristics such as color or brightness. For example, a user can edit a virtual environment by changing the scenery or adjusting the soundscape.
[0065] The reproduction of a virtual environment refers to the display or execution of the virtual environment so that the user can explore or interact with it. This can be done on various devices, such as a computer screen, a virtual reality headset, or a mobile device. For example, after being generated, a virtual environment can be reproduced to allow the user to navigate in a 3D world. A rendering engine is a software program or module. Used to transform digital models into images or visual sequences that the user can view. A rendering engine allows, in particular, the control of the display of objects, textures, and lighting effects by a display device within the context of reproducing a virtual environment.
[0066] Navigation in a virtual environment refers to the movement or exploration of that environment by the user, either using controls (such as a keyboard or a gamepad), or in virtual reality via motion-tracking devices. The user can move freely within the environment, such as walking in a virtual city or flying over a landscape.
[0067] The term “content” in this document refers to all data or information used to generate or modify a virtual environment. This may include text, graphics, sounds, or any other descriptive element. For example, narrative text, reference images, or audio files may constitute content. Content may include popular works such as films, comics, manga, etc., documentary sources such as timelines, geographical maps, architectural plans, photographic archives, or a combination of several types of content from different sources or in distinct formats.All or part of the content may be provided by the user in the form of a user input or as a link contained within a user input, or it may be a popular work accessible through a simple interpretation of the user input, for example from a text or a summary description.
[0068] Content elements are the individual components that make up the overall content. These elements can include visual descriptions, dialogue, characters, or specific locations within a story. For example, character descriptions in a book or scenes in a literary work can be content elements. Contextual content elements are information that describes the environment, atmosphere, or setting in which the action takes place. They can include descriptions of locations, historical periods, or weather conditions. Narrative content elements are information that relates to the plot, the characters' actions, and the progression of events.Content extraction refers to the process by which specific elements, such as descriptions of places or characters, are identified and extracted from a text or other medium for use in generating a virtual environment.
[0069] A descriptive graph is a data structure that represents the relationships between the different elements extracted from a content. Such a data structure can be used to guide the generation and organization of representations in a virtual environment. Structured elements can be stored and organized in different formats, such as graph-oriented databases (e.g., Neo4j), relational databases (RDS), or RDF / Sparql models.
[0070] Artificial intelligence (AI) is a technology that enables computer systems to perform tasks that normally require human intelligence, such as image recognition, natural language understanding, or decision-making. Generative artificial intelligence is a subcategory of AI that creates new data or content from given instructions or examples. Generative AIs can be used to create different types of content in the virtual environment. For example, specialized AIs can generate 3D objects, images, sounds, dynamic landscapes or "skyboxes," audio descriptions for visually impaired users, or subtitles in different languages for dialogues. These AIs make it possible to generate diverse elements that can, for example, be adapted to the user's preferences and needs.
[0071] In the context of the proposed technique, generative artificial intelligence is used to analyze content and to generate representations of elements extracted from such content. Analysis by generative artificial intelligence is a process by which this AI examines given content (such as text) to identify and extract relevant information, and then uses this information to generate new elements. The generation of representations by generative artificial intelligence refers to the process by which this AI creates representations of the content elements, whether, for example, in the form of images, sounds, or 3D objects.
[0072] A request to a generative artificial intelligence is an instruction or question submitted with the aim of receiving a response or generating data. For example, a request might ask the generative artificial intelligence to generate a representation of a character in the form of an image based on a contextual element such as a textual description of the character. Sending a request to a generative artificial intelligence is the act of transmitting this request (a request for information or creation) to the AI through a means of communication, such as an application programming interface (API). Receiving a response from a generative artificial intelligence refers to receiving the content or information generated by the AI after processing the sent request.
[0073] Representation aggregation is the process by which the different generated representations are combined to form a coherent virtual environment. For example, representations of a character, its environment, and ambient sounds can be aggregated to create an immersive scene.
[0074] Reference is now made to [Fig. 1], which represents an immersive reality system, according to a possible embodiment of the proposed technique. [Fig. 1] shows different logic modules, each defined by a specific function: an interface 11 for obtaining user input, a content analyzer 12, an element extractor 13 from the content, a generator 14 of representations of the extracted elements, an aggregator 15 of the representations, a virtual environment generator 16, a module 17 for reproducing the virtual environment, and an interface 18 for navigating the virtual environment.
[0075] The user input retrieval interface 11 allows for the collection of one or more initial pieces of information provided by a user. The user input may include all or part of a piece of content such as a descriptive text, a literary work, and / or multimedia files. The user input may include a link to all or part of the content or instructions for accessing all or part of the content. The user input may include instructions regarding a desired immersive experience, for example, a desired visual or sound style, a desired type of immersive experience (narrative or exploratory), and / or a preference for focusing on certain elements of the work (specific characters, locations, etc.). The interface 11 may be implemented in various forms, such as an online form or a voice command. Depending on the implementation, the interface may transmit the collected information to subsequent modules for processing.It is possible that the user input interface will only provide partial information (for example, simple text), and that any preferences regarding the immersive experience will only be determined at a later stage by other modules. It is also conceivable that interface 11 will be replaced by an automated module that retrieves content without user intervention.
[0076] The content analyzer 12 analyzes the content to identify relevant information. If the content is text, the analyzer can use a natural language processing (NLP) technique to identify entities such as locations, characters, actions, and atmospheres. If the content is multimedia, the analyzer can process metadata and / or directly analyze visual and / or audio elements to extract relevant information. The analysis can include structuring the identified elements to facilitate their subsequent use. In one embodiment, the analyzer 12 can be limited to analyzing certain aspects of the content (for example, only the characters). In one embodiment, the analyzer 12. Content can be omitted or integrated into the extractor 13. Content elements. In such an implementation, the content is directly transmitted to the extraction module without prior analysis.
[0077] The content element extractor 13 extracts at least contextual elements and optionally one or more narrative elements from the analyzed content. The extractor 13 can operate iteratively, providing increasingly detailed information about each element, for example, a location and its components (buildings, rooms, objects, etc.). The extractor can be configured to explore relationships between nested elements, for example, spatial relationships between rooms in a building or between a room and objects placed within it, or logical relationships between objects and their function in the narrative. The extraction result can be represented as a descriptive graph, facilitating the subsequent generation of representations. In some cases, the extractor can be merged with the analyzer, forming a single module that performs both analysis and extraction.It is also possible that the extraction will be limited to a small number of elements or to a specific category (such as objects).
[0078] The extraction result can be formulated as a query for a generative AI, or as a portion of such a query. Such a query can be a set of instructions or content elements suitable for use by a generative artificial intelligence to generate visual, auditory, or other representations. This query can include text (e.g., descriptions of characters or environments), images (visual references, plans, sketches), audio files (voice instructions, sounds to be used in the environment), or a combination of several of these elements. For example, a query might contain a detailed textual description of a location, accompanied by a reference image and an audio file describing the soundscape to be generated. The formulation of this query depends on the nature of the element to be generated.For example, to generate a 3D representation of a location, the request might include textual information about the arrangement of objects, visual references, and metadata about scale and dimensions. For dialogues, the request might include both text and audio files of specific voices, or even narrative elements to guide the interaction between characters. Specific languages exist for structuring and sending these requests, such as graph databases using Cypher, or extraction-oriented APIs that allow retrieving structured and relevant information for generating an environment.
[0079] The element representation generator 14 receives the structured data from the extractor 13 and calls upon one or more generative artificial intelligences to create visual, sound, or textual representations of the extracted elements. It can use generative artificial intelligences specialized, for example, in The creation of 3D objects, characters, or landscapes. For example, for a location described in a text, the generator can produce a 3D representation of that location, including textures and precise dimensions. It can also generate avatars for characters, interactive objects, or landscapes. The generation process can be influenced by parameters defined in the user input, such as the visual style (realistic, manga, etc.). The generator can leverage several AIs specialized in specific tasks, such as generating 3D objects or dialogue.
[0080] The interaction mechanism between the generator 14 and the generative artificial intelligence(s) may include several types of requests or commands intended for this or these generative artificial intelligence(s), depending on the information available and the objectives of the generation.
[0081] The request may, for example, consist solely of the extracted contextual element in the form of a request for generative artificial intelligence. Generally speaking, a "contextual element in the form of a request for generative artificial intelligence" refers to information extracted from content (text, image, audio, video, etc.) that is formatted and sent as a request to a generative AI to create a visual, auditory, or other representation of that element in the virtual environment. Such a request may be simple or complex, and may contain isolated pieces of information or combinations of contextual and narrative elements.
[0082] The request may include one or more of the following: a single contextual element, for example a specific place or object, part of a contextual element, for example a detail or feature of a place, several contextual elements combined to give a richer description (for example, the combination of a place and the weather), additional information, such as narrative elements or user instructions (visual preferences, themes, etc.).
[0083] For example, when an extracted contextual element is a description of a place such as a forest, an example of a request sent to an image-generating artificial intelligence could be: "A dense forest with large trees and a humid climate" or, equivalently, "Generate a visual representation of a dense forest with large trees and a humid climate".
[0084] For example, where an extracted contextual element is a description of a rainforest, this description mentioning the presence of a river as part of the contextual element, an example of a query might be: "Generate a visual representation of a river surrounded by trees in a rainforest."
[0085] For example, when a descriptive graph includes a first extracted contextual element entitled "a shopping street in a city" associated with a second element In a contextual extract titled "rainy weather," an example query might be: "Generate a soundscape of a city with pedestrians walking in the rain." Alternatively, a query might include multiple elements, such as a location and objects within that location. For example: "Generate a visual representation of a bedroom with a four-poster bed and dim lighting."
[0086] A request can also include a narrative element, such as an action or event, in addition to the contextual element. For example: "Generate a visual scene of a character sitting on a park bench at sunset, reading a book." In this case, generative artificial intelligence can take into account both the environment and the character's action to create a more dynamic and contextual representation.
[0087] Requests to the second AI are not limited to textual instructions. They can be more generic and include images, text, or any other instruction format, or a combination of instructions in distinct formats, such as text combined with one or more images. A request may contain a textual description and a reference image, the textual description being, for example: "Generate a visual scene of a busy street with architectural elements similar to the attached image." A request may, for example, include an ambient soundtrack to which the generative artificial intelligence is tasked with associating one or more visuals, as well as a textual instruction comprising a contextual element and a reference to the soundtrack, such as: "Generate a visual scene of a rainforest corresponding to the sounds of birds and wind in the attached recording."
[0088] One challenge encountered when generating visual representations is that some may not meet the user's expectations in terms of resemblance to the desired element. To overcome this, several versions of the same contextual element can be generated, and the user can vote for the representation that best suits them. This iterative process, or "fine-tuning," can continue by generating new variants inspired by the chosen version, thus refining the final representation until it satisfies the user.
[0089] The element representation aggregator 15 can combine the different representations generated by module 14 to form a coherent environment. The aggregator can position objects in virtual space and adjust their spatial relationships (for example, the relationship between an object and its environment). It can also manage sound aspects, such as adding dynamic sound effects or soundtracks, depending on the scenes or actions planned. In some implementations, the extractor can also adjust the positions of element representations based on user interaction. The aggregator may not be necessary if The representation generator directly generates a pre-structured environment. In some implementations, the aggregator can be replaced by a simpler module that merely arranges objects in a predefined way.
[0090] The virtual environment generator 16 can finalize the creation of the virtual environment by aggregating elements. It can compile the aggregated representations into an executable application that allows the virtual environment to be reproduced on different media (virtual reality headsets, computer screens, etc.). The generator can configure the interactive aspects of the environment, such as the ability to follow a narrative, change perspective, or move objects. The generator can adjust the interactive aspects to the technical constraints of the target platforms and can manage, for example, graphics performance or memory resources. The generator can be omitted in scenarios where the environment is not intended to be run on an interactive platform (for example, in the case of simple visualization).In some implementations, separate elements are generated (for example, visual elements and, separately, sound elements) and are only assembled when reproducing the virtual environment.
[0091] The virtual environment reproduction module 17 can control the display and execution of the virtual environment on various devices (computer, virtual reality headset, tablet, etc.) to reproduce graphics and / or sounds. Module 17 can also manage multi-user environments, allowing several people to interact in the same virtual world. Module 17 can rely on a rendering engine, which calculates visual perspectives, lighting, shadows, and textures in real time, allowing the user to see and interact with the virtual environment. Module 17 can be omitted in embodiments where the environment is not reproduced in real time, for example, if it is simply pre-recorded. In some embodiments, an external rendering engine can be used for reproduction.
[0092] The virtual environment navigation interface 18 allows the user to interact with the virtual environment, either by navigating freely or by following a predefined narrative. It may include movement controls via a keyboard, gamepad, navigation buttons, or virtual reality motion tracking devices. The navigation interface also allows the user to change their perspective (for example, switching from a first-person to a third-person view) or to interact with objects in the environment (opening doors, picking up objects, etc.). In the case of interactive storytelling, this module can also adjust the possible actions accordingly. of the unfolding of the story. The navigation interface can be omitted in implementations where the experience is purely observational.
[0093] It is possible to consider several different sequences for the execution of the modules, while keeping all modules active. According to one possible sequence, the overall operation of the algorithm relies on a pipeline where each logical module can act based on the data processed by the preceding modules.
[0094] According to one embodiment, obtaining user input can trigger content analysis, thereby providing the information necessary for extracting elements from the content. The extracted elements are then used to generate representations, which are subsequently aggregated. A virtual environment can be generated from the aggregated representations, reproduced, and the user can then interact with the environment via the navigation interface.
[0095] According to one embodiment, content analysis by analyzer 12 and element extraction by extractor 13 precede obtaining user input through interface 11. In such an example, specific pre-existing content is imposed, and the user input contains, for example, a display or interaction preference.
[0096] According to an example embodiment, a draft of a virtual environment is generated by the generator 16 starting from a previously defined structure, before all or part of the representations of the elements are created by the generator 14 and aggregated by the aggregator 15.
[0097] According to one embodiment, the interface 11 can obtain user input containing one or more additional pieces of information, relating for example to reproduction and / or navigation preferences, for example after the extraction of relevant elements and before the generation of representations, or after the generation of representations.
[0098] It is also possible to provide several modules having the same function. For example, several interfaces 11 can be configured to receive user input in different formats. One interface 11 might be designed to receive user input in the form of text, a second to accept PDF or multimedia files (images, videos, audio files), a third to receive voice commands, and a fourth to accept other inputs. In one variation, the three aforementioned interfaces can each be configured to receive a portion of user input, with the input then being reconstructed from these portions.
[0099] Similarly, several 12 content analyzers may be provided, each specialized in processing specific types of content. For example, a first analyzer may process texts exclusively, while a second analyzer is dedicated to the analysis of images or videos, and a third analyzer specializes in the analysis of audio files. It may also be possible to have several analyzers capable of analyzing the same content format and working in parallel to accelerate the process of analyzing large volumes of data.
[0100] Several extractors 13 can also be considered to process different elements of the content. For example, a first extractor can be tasked with focusing on extracting contextual elements and a second extractor on narrative elements. These extractors can be configured to operate simultaneously.
[0101] It is also possible to provide several representation generators 14, each specialized in generating specific types of elements. A first generator can be responsible for creating 3D visual representations, a second generator can be responsible for creating sound representations, and a third generator can be responsible for creating textual representations such as subtitles.
[0102] Several aggregators 15 can also be used, for example in a modular fashion to combine the different types of representations created by the generators. For example, a first aggregator can handle the merging of visual representations of contextual elements, a second aggregator can be dedicated to integrating the sound representations of contextual elements into the environment, and a third aggregator can be responsible for synchronizing the appearance of representations of specific contextual elements taking into account one or more narrative elements.
[0103] It is also possible to provide several virtual environment generators 16 and several reproduction modules 17, each of which can be configured to generate and reproduce the environment on different media (e.g., computer screens, virtual reality headsets, mobile devices). These generators and reproduction modules can be coordinated to offer experiences adapted to the capabilities of different devices or to manage multi-user sessions on several types of devices simultaneously.
[0104] Other modules offering additional functionalities may be provided.
[0105] For example, a sharing module may be provided to share a generated virtual environment with other users who can explore it themselves.
[0106] For example, an automatic storytelling module may be provided to interact with interface 11 in order to offer the user predefined or dynamic stories or scenarios. The automatic storytelling module may also be linked to representation generator 14 to generate specific dialogues or actions within the environment.
[0107] For example, a virtual environment adjustment module may be provided to interact with interface 11 or with navigation module 18 to offer narrative elements tailored to a user profile or past user behavior.
[0108] For example, a machine translation and / or cultural adaptation module can be linked to the analyzer 12 to translate the analyzed content into a different language or to adapt or ignore certain cultural elements, taking into account user preferences. The machine translation and / or cultural adaptation module can also be linked to the representation generator 14, to adapt or ignore representations of visual and sound elements taking into account local preferences or specific cultural expectations.
[0109] For example, a real object integration module may be provided to scan real physical objects using a capture device and to provide representations of these objects to the aggregator 15 for their integration into the virtual environment.
[0110] For example, a module for managing variations in atmosphere based on changes in context or narrative may be provided. For example, such a management module could adjust the lighting, weather, or background music according to the development of the story or the actions of the characters. The module for managing variations in atmosphere may be linked to the representation generator 14 in order to generate several possible representations of the same element. The module for managing variations in atmosphere may be linked to the virtual environment generator 16 in order, for example, to select a representation from among several previously generated representations of the same element, or to modify a representation of an element to adapt it to a desired atmosphere.
[0111] For example, an avatar creation, modification or personalization module can be linked to interface 11, virtual environment generator or navigation interface 18, in order to allow the user to create, modify or personalize their avatar in the virtual environment.
[0112] Reference is now made to [Fig. 2], which represents an immersive reality system according to a possible implementation of the proposed technique. [Fig. 2] is similar to [Fig. 1] except for the following two additions: a module 20 for obtaining a virtual environment, and a virtual environment editor 26.
[0113] Module 20 for obtaining a virtual environment is configured to obtain a pre-existing virtual environment. The virtual environment is already generated and can be obtained from a database, a stored file, or another source. It can be an environment previously created via the generation process in [Fig. 1] or another existing environment. It can be a shared environment.
[0114] The virtual environment editor 26 replaces the generator 16 of [Fig. 1]. Whereas the generator 16 created a new environment from scratch, the editor 26 is responsible for modifying an existing environment obtained via the module 20. It allows the user to add, delete, or modify representations of interactive elements in the environment, according to user instructions. The editor 26 can also adjust global parameters such as the atmosphere, lighting, or spatial arrangement of objects.
[0115] The other modules retain their functions as described in relation to [Fig.1].
[0116] As part of the editing process, user input may include specific instructions on the changes to be made to the existing environment. For example, the user may indicate which objects or characters to modify, which visual or sound elements to adjust, or specify areas of the environment to be modified.
[0117] The element representation generator 14 is configured to generate new representations of elements to be added or modified in the environment. For example, if the user wants to replace a character or object, generator 14 creates the new version of the element to be integrated. It can also generate texture or lighting adjustments based on various criteria, for example, based on user instructions.
[0118] The element representation aggregator 15 manages added representations and modifications made to existing representations. For example, if an object is modified in terms of texture or position, the aggregator integrates this modification into the virtual environment. It can also merge modified elements with elements that have remained unchanged.
[0119] An interaction loop may be provided, involving one or more of the following modules: the interface 11, the generator 14, the aggregator 15, and the editor 26, to allow adjustment of the representations of the virtual elements composing the virtual environment. This adjustment may include modifications to textures, colors, object layout, or soundscapes. The interaction loop may allow the user to interact with these elements, test different configurations, and observe the results to adjust the appearance and interactions of the environment to their specific preferences. These adjustments may be made manually or via a semi-automated process assisted by artificial intelligence, also known as "fine-tuning."
[0120] Reference is now made to [Fig. 3], which represents a possible operation of a content element extractor 13. The extractor 13 is designed to analyze the supplied content and extract information relevant for the generation or the modification of a virtual environment. It can be structured into several sub-modules, each with a specific function in the extraction and organization of data from the content.
[0121] The extractor 13 may include a contextual element extractor 31. Such an extractor 31 is designed to extract, specifically, elements intended to define a setting and / or atmosphere of the virtual environment. It can extract information such as locations, weather conditions, a time period (past, present, or future), descriptions of environments, for example, natural or urban, as well as objects that are part of a setting. For example, if the content is a literary text describing a scene in a forest, the extractor 31 can identify contextual elements such as trees, climate, lighting, and the spatial layout of the scene.
[0122] The extractor 13 may include a narrative element extractor 32. Such an extractor 32 is designed to extract, specifically, elements that contribute to the progression of a story or the interaction of characters within the environment. It can extract information such as characters, their actions, their dialogue, and events that influence the unfolding of a story. For example, in a film or novel, the extractor 32 can extract interactions between characters, key dialogue, decisions made by the protagonists, and narrative transitions from one scene to another. Such narrative elements can be used to generate interactive sequences within the virtual environment, following the logic of the extracted narrative elements.
[0123] Extractor 13 may include a module 33 for providing a descriptive graph. Once the contextual elements (and optionally one or more narrative elements) have been extracted, module 33 structures this information into a descriptive graph. This graph represents the relationships between the different extracted elements, for example, between characters and their environment, or between events and objects in the scene. The descriptive graph helps guide the generation of representations in the virtual environment, thus facilitating the aggregation of the different elements. For example, a character can be associated with specific objects, or a location can be linked to a key plot point.
[0124] It is possible to consider different configurations for the extractor 13 and its modules in order to adapt the extraction process to the specific needs of a virtual environment. For example, several contextual element extractors 31 can be used simultaneously to process specific aspects of the content. A first extractor 31 can be responsible for extracting geographical information, while a second extractor 31 can be responsible for extracting cultural and historical elements specific to a given region.
[0125] Similarly, several narrative element extractors 32 can work in parallel to extract different types of elements. A first extractor 32 can extract interactions between characters, a second extractor 32 can be responsible for extracting plot events, while a third extractor 32 can be dedicated to extracting dialogues.
[0126] Module 33, which provides a descriptive graph, can be configured to first generate several descriptive graphs to represent different layers of the virtual environment: for example, a first graph for the spatial relationships between objects and places, a second graph for the interactions between characters, and a third graph for the progression of events. These graphs can then be merged into a single global graph or used independently.
[0127] The operation of the modules can be adapted according to the order of intervention. For example, in some implementations, the extraction of narrative elements by extractor 32 can take place before the extraction of contextual elements by extractor 31. In another configuration, the descriptive graph can be provided upstream, before all the elements are extracted, to guide how the different extractors 31 and 32 interact with the content. This can allow the extraction to be adapted as the graph becomes more complex with new elements.
[0128] It is also possible to omit certain modules in certain configurations. For example, module 33, which provides a descriptive graph, can be omitted if the extraction of contextual and narrative elements is used directly to generate representations without prior structuring. Similarly, extractor 32, which extracts narrative elements, can be omitted if the virtual environment is purely exploratory and does not require narrative progression.
[0129] Extractor 13 may include, or collaborate with, one or more additional modules offering various functionalities.
[0130] For example, a spatio-temporal relationship analysis module may be provided to analyze spatial and / or temporal relationships between extracted events. For example, such a module may establish a chronology of actions in a plot, and facilitate the management of a succession of events involving characters, objects, places, etc., in the virtual environment.
[0131] For example, a contextual enrichment module may be provided to supplement extracted contextual elements by enriching them with additional data from external sources, such as historical databases or cultural repositories. For example, a description of a geographical location may thus be supplemented by taking into account information on local fauna or flora.
[0132] A consolidation module for extracted elements may be provided to validate and merge elements extracted by different extractors. Such a consolidation module may be configured to eliminate duplicates or correct inconsistencies, and thus consolidate the extracted elements before their integration into the descriptive graph and / or their use for generating representations of these elements.
[0133] Reference is now made to [Fig.4] which represents a possible process of interaction between an extractor 13 of elements from a content and an analyzer 12, in this case an artificial intelligence, within the framework of a process of extracting information from a given content, in an example of an embodiment.
[0134] Unlike [Fig.3] which focuses on the internal logical modules of the extractor or content analyzer, [Fig.4] focuses on the interaction between the different stages of processing a request intended for an external AI, for example, and describes a cascading extraction scheme of contextual and narrative elements.
[0135] The process begins with the sending 41 of a request, or "prompt," to the parser. The request is a specific question about the content, asking the AI to extract relevant information from the content. A request might ask the parser to generate a list of contextual elements, for example, the main locations in a story corresponding to the analyzed content. Following the sending 41 of the request, the process continues with obtaining 42 a response from the parser. The response is based on the parser's analysis of the content. In the given example, this response might include a list of the main locations in the story, as described in the content (for example, a city, a forest, a house). Finally, once the list of locations is obtained, an extraction 43 of the listed elements can be performed.
[0136] Each element thus listed can also serve as a basis for formulating additional queries. Continuing the previous example, for each location identified in the response, a new query can be prepared asking the parser to provide a list of the parts or sub-spaces of that location (for example, for a house, this could be the different rooms). Then, one or more additional queries can be generated for each part of the location or sub-space of the location, requesting, for example, a list of the objects present in each room, or information on the dimensions of the locations. This process can continue down to a very fine level of detail. For example, the chaining of queries can make it possible to obtain, for each part of the location, the information necessary to predict a query intended for an artificial intelligence tasked with generating a visual or spatial representation of the identified locations.
[0137] A similar mechanism can also be applied to narrative elements. For example, an initial query might request a list of the main characters in the story corresponding to the analyzed content. Once this list is obtained, further queries can be sent for each character, requesting a list of their main actions or interactions in the story. A query can also be formulated to identify the locations corresponding to each action, in order to establish links between the characters and the environments where the events take place.
[0138] The set of information extracted via these query chains can be structured, for example, by the contextual element extractor 31 and the narrative element extractor 32. That is to say, the extractor 31 can be responsible for preparing queries to extract the contextual elements and for interpreting the parser's responses to these queries in order to extract the contextual elements. Similarly, the extractor 32 can be responsible for preparing queries to extract the narrative elements and for interpreting the parser's responses to these queries in order to extract the narrative elements.
[0139] Module 33, which provides a descriptive graph, can be used to interpret the parser's responses to queries in order to identify relationships between the extracted elements and structure them in the descriptive graph. For example, the descriptive graph can specify that a particular character is in a certain location at a certain point in the plot, or that a given object is associated with a given scene in a given room.
[0140] The extraction of narrative elements can take place before or after the extraction of contextual elements, or these extractions can occur simultaneously or in parallel. The descriptive graph can be generated after all elements have been extracted. In some embodiments, the descriptive graph can be generated as elements are extracted, and the extractor 13 can take into account the descriptive graph being generated to guide the preparation of queries aimed at extracting new elements.
[0141] Reference is now made to [Fig. 5], which represents a possible example of a descriptive graph 50 in an example embodiment. The descriptive graph can be multidimensional and can link several types of elements to the virtual environment.
[0142] For example, the descriptive graph can list 51 places, which are environments identified in the analyzed content. These places can represent physical spaces (e.g., a city, a house, a forest). Each place is an entity in the graph that can be linked to other elements to structure the space.
[0143] The descriptive graph may also identify sub-places 52, which represent subdivisions of places 51. For example, if place 51 is a house, the sub-places 52 may correspond to the rooms of that house (living room, bedroom, kitchen). These Sub-places allow for a finer level of detail and can be used to structure the internal space of complex environments. Sub-places can also include sub-sections of natural spaces, such as clearings in a forest or specific areas within a city.
[0144] Objects 53 constitute another type of element that can be included in the descriptive graph. These objects can be interactive or decorative elements present in locations or sub-locations. For example, in a room of a house, furniture, works of art, or particular objects can be identified and integrated into the graph. Objects can be associated with specific actions or events that take place within the narrative context.
[0145] The descriptive graph can include 54 events, which represent points or periods in time where actions or changes occur in the virtual environment. These events help structure the chronology of actions taking place in the different locations and sub-locations. For example, an event might correspond to a moment when a particular character arrives at a given location or performs a specific action. This type of temporal relationship can help guide a narrative or interactive sequences.
[0146] The 55 characters are entities that can be enumerated in the graph. The characters represent individuals interacting in the virtual environment. Each character can be linked to places, sub-places, objects, or events. For example, a character can be associated with a specific place where it resides or with objects it uses.
[0147] The descriptive graph may include 56 actions, which represent the interactions or movements performed by the characters in the virtual environment. These actions may be linked to specific locations, objects, or events. For example, an action may be moving an object from one room to another, or interacting with a character in a scene.
[0148] In summary, the descriptive graph 50 can link the contextual and narrative elements extracted from a content, for example the places, sub-places, objects, events, characters and actions of the content in a coherent manner to create a structure which guides the generation of representations of the extracted contextual elements and their interaction in the virtual environment.
Claims
Demands
1. Method for automatically generating a virtual environment, the method comprising: an extraction, in the form of a query intended for a generative artificial intelligence, of elements of a content using an analysis by a first artificial intelligence, the elements comprising at least contextual elements, and a generation by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
2. A method according to the preceding claim, the extraction further extracting at least one narrative element from the content, and the generation of representations and / or the aggregation of the generated representations taking into account at least one narrative element.
3. A method according to any one of the preceding claims, the extraction providing a descriptive graph of the virtual environment to be generated comprising the extracted elements, the generated representations being aggregated according to the descriptive graph.
4. A method according to any one of the preceding claims, the method comprising: obtaining a user input, the user input enabling the content to be obtained.
5. Method according to the preceding claim, the user input comprising at least one of the following: the content; a part of the content; a link to the content; a link to a part of the content.
6. A method according to any one of the preceding claims, wherein the content used by the extraction comprises at least one of the following: user input; text from a work; audio content; written content; graphic content.
7. A method according to any one of the preceding claims, comprising: an aggregation of representations to create the virtual environment.
8. A method according to any one of the preceding claims, wherein the extraction of elements includes sending successive requests to the first artificial intelligence, wherein the sending of successive requests includes at least one sending of a request dependent on a prior response from the first artificial intelligence module.
9. A method according to any one of the preceding claims, wherein the representations of the elements include visual and / or sound representations.
10. A method according to any one of the preceding claims, wherein the aggregation of representations is carried out in a rendering engine.
11. A method according to any one of the preceding claims, wherein the virtual environment is three-dimensional.
12. A method for editing a virtual environment comprising: an extraction, in the form of a query intended for a generative artificial intelligence, of elements of a content using an analysis by a first artificial intelligence, the elements including at least contextual elements, and a generation by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
13. Method of using an immersive reality previously obtained by at least one of the following methods: the generation method according to the method of one of claims 1 to 11; the editing method according to claim 12; the method of use comprising: a reproduction of the virtual environment obtained.
14. Method according to the preceding claim, comprising: navigation in the virtual environment.
15. Computer program comprising instructions for carrying out the method according to any one of the preceding claims when this program is executed by a microprocessor.
16. Automatic generator of a virtual environment, the generator comprising: an extractor, in the form of a query intended for a generative artificial intelligence, of elements of content using an analysis by a first artificial intelligence, the elements including at least contextual elements, and a generator by a second artificial intelligence of representations of the elements, the generated representations being aggregated in the generated virtual environment.
17. Immersive reality system comprising: an automatic generator of a virtual environment according to the preceding claim, a device for reproducing the generated virtual environment, and an adapter of the generated virtual environment based on user interaction with the reproduced virtual environment.