Data processing apparatus and method

The data processing apparatus addresses the limitations of manual environmental adjustments by using AI to generate scene schemas from audio input, enabling automatic and immersive adjustments of lighting and sound in simulated experiences.

WO2025162809A1PCT designated stage Publication Date: 2025-08-07SONY GROUP CORP +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051649
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-31
Filing Date
2025-01-23
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Current solutions for adjusting environmental elements like lighting and sound in simulated experiences, such as role-playing games or museum exhibits, rely on manual adjustments from a limited number of settings, leading to predictable atmospheres and disrupting user engagement.

Method used

A data processing apparatus that uses a processor to analyze audio input, generate a scene schema, and control environmental output devices like lights and speakers automatically, based on a virtual environment, utilizing generative AI and semantic databases to create immersive and context-dependent simulations.

Benefits of technology

Enables richer, seamless, and immersive simulated experiences by automatically adjusting lighting and sound in response to audio input, reducing the need for manual user intervention and maintaining engagement with the simulated environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051649_07082025_PF_FP_ABST
    Figure EP2025051649_07082025_PF_FP_ABST
Patent Text Reader

Abstract

A data processing apparatus comprising circuitry configured to: receive information derived from digital content; generate a virtual environment using one or more digital assets corresponding to the received information; extract information from the virtual environment; and output a control signal to one or more output devices of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] DATA PROCESSING APPARATUS AND METHOD BACKGROUND Field of the Disclosure The present disclosure relates to a data processing apparatus and method. Description of the Related Art The “background” description provided is for the purpose of generally presenting the context of the disclosure. Work of the presently named inventors, to the extent it is described in the background section, as well as aspects of the description which may not otherwise qualify as prior art at the time of filing, are neither expressly or impliedly admitted as prior art against the present disclosure. There are many scenarios where it would be beneficial for environmental elements such as lighting and sound to be automatically adjusted to help provide a particular atmosphere. For example, in a table top role playing game, depending on the environmental setting narrated by the game master, it may be desirable for sound and lighting to be adjusted to correspond with the environment (e.g. darker torch-like lighting and ominous sounds for a cave or brighter day-like lighting and nature sounds for a forest). This helps improve player immersion in the game. Other examples which might benefit from such adjustable lighting and sound (as well as other potential environmental adjustments, such as temperature, smell or the like) could include museum exhibits, escape games, karaoke, storytelling (e.g. bedtime stories for children or audiobooks), spa and massage settings, medical intervention (e.g. calming sound and lighting for a patient in an medical scanner), meditation, stage play, travel (e.g. aeroplane cabins) and commercial (e.g. luxury brand events) settings. Although the use of environmental adjustments is known, current solutions tend to rely on manual adjustments and / or selection from a relatively small number of settings. This can limit the effectiveness of atmospheres created. For instance, the environmental settings become predictable and, for manual adjustment, the action of having to manually adjust, say, the lighting and sound makes it difficult for the user doing the manual adjustments to stay engaged with the created environment (in other words, manual adjustment breaks the continuity of the simulated experience). There is thus a desire for a technical solution to allow simulated experiences of this kind to be richer, less predictable and implementable in a seamless, context dependent and response / automatic manner. SUMMARY The present disclosure is defined by the claims. BRIEF DESCRIPTION OF THE DRAWINGS Non-limiting embodiments and advantages of the present disclosure are explained with reference to the following detailed description taken in conjunction with the accompanying drawings, wherein: Fig.1 shows an example system; Fig.2 shows some example functions executed by a processor; Fig.3 shows some example steps carried out by an extractor function; Fig.4 shows some example steps carried out by an asset builder function; Fig.5 shows some example outputs of the extractor and asset builder functions; Figs. 6A and 6B shows an example simulated environment and corresponding physical environment; Fig. 7 shows example ways in which how light interacts with objects in the simulated environment can be specified; and Fig.8 shows an example method. Like reference numerals designate identical or corresponding parts throughout the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS Fig.1 shows an example system. The system comprises a data processing apparatus / device 100, environmental output devices 110 and a microphone 109. The data processing apparatus 100 comprises a processor 101 for executing electronic instructions, a memory 102 for storing the electronic instructions to be executed and electronic input and output information associated with the electronic instructions, a storage medium 103 (e.g. a hard disk drive or solid state drive) for long term storage of information, a communication interface 104 for sending electronic information to and / or receiving electronic information from one or more other apparatuses and a user interface 105 (e.g. a touch screen, a non-touch screen, buttons, a keyboard and / or a mouse) for receiving commands from and / or outputting information to a user. Each of the processor 101, memory 102, storage medium 103, communication interface 104 and user interface 105 are implemented using appropriate circuitry, for example. The processor 101 controls the operation of each of the memory 102, storage medium 103, communication interface 104 and user interface 105. The environmental output devices 110 are devices configured to adjust a characteristic of the environment under control of the data processing apparatus 100. They are each connected (e.g. via a suitable wired or wireless connection, such as via the Zigbee ® protocol) to the communication interface 104 of the data processing apparatus. In this example, the devices 110 include one or more lights 106 (e.g. smart lights comprising a plurality of LEDs to enable light of any colour in a predetermined colour space, such as the YCbCr colour space, to be reproduced), one or more loudspeakers (speakers) 107 and, optionally, one or more other devices 108. The one or more other devices could include, for example, a heating and / or cooling apparatus to change the temperature of the environment, an air-blowing apparatus (e.g. fan) to create an airflow in the environment (e.g. simulating wind), a scent diffuser (to diffuse a safe chemical compound, e.g. certain essential oils, into the environment with a particular smell) and / or an electronic display (e.g. liquid crystal display, LCD, or organic light emitting diode, OLED, display) for displaying images. In the following examples, the lights 106 and loudspeakers 107 are discussed in detail. However, it will be appreciated the present technology may be adapted accordingly to allow control of the other devices 108 as appropriate. The microphone 109 is configured to capture audio in the environment (in particular, speech uttered by one or multiple users in the environment) and provide it, via the communication interface 104, to the data processing apparatus 100. As will be explained, speech captured by the microphone 109 is processed to determine how to control the environmental output devices 110. More generally, there may be one or more microphones 109. Although the data processing apparatus 100 is shown as a single apparatus, the functionality of the data processing apparatus 100 may instead be spread between multiple data processing apparatuses (each having the structure of the data processing apparatus 100) connected to each other (e.g. via their respective communication interfaces 104 over a local area network (LAN) and / or the internet). The data processing apparatus 100 (or apparatuses) may also be local (e.g. connected via a local network) or remote (e.g. connected via the internet) to the environmental output devices 100 and / or microphone 109. In one example, this allows the microphone 109 to be placed in a first location (e.g. to capture speech in the room a role playing game is being placed) and the environmental output devices 100 to be placed in a second, different, location (e.g. in a room with someone watching the role playing game being played via a video link or the like). This helps improve the sense of immersion of users watching the game remotely. Fig. 2 shows some example functions executed by the processor 101 of the data processing apparatus 100. The functions include an extractor 201 for extracting scene schema from audio, an asset builder 205 for gathering or generating digital assets I(that is, digital content such as colour information (e.g. lighting luminance and / or chrominance values and / or dynamics), images, videos and / or audio files) for the scene and extracting information from them, a simulator 204 for taking the scene schema and generated assets and / or asset information and modelling them in a game engine (the simulator 204 including a composer component that controls the dynamics and lighting composition of the scene to provide the realism of the output), a location editor 203 for specifying the dimensions of the simulation (e.g. to correspond with the size of the room in which the environmental output devices 110 are placed) and where each of the environmental output devices 110 (so-called “smart devices” in this example) are located and a streamer 202 for streaming values extracted from the simulation and transmitting them to the smart devices 110. Fig.2 also shows the microphone 109 (which provides captured audio data to the extractor 201), a generative artificial intelligence (AI) model 206 and semantic and asset databases 207 and 208 (stored in storage medium 103, for example). The generative AI model 206 comprises, for example, any suitable known text-to-image model for generating images from natural language prompts. An example of such a model is DALL-E from OpenAI ®. The asset database 208 stores assets generated by the asset builder 205 for use by the simulator 204. The semantic database 207 stores assets to be retrieved by finding a closest vector embedding in the semantic database 207 to a vector embedding representing input text (the text being determined based on speech detected in audio picked up by the microphone 109). In an example, to build the semantic database 207, suitable known semantic embedding model(s) (e.g. CLIP from Open AI for image assets and / or CLever Audio Plugin, CLAP for audio assets) are used to associate each asset in the semantic database 207 with a respective vector embedding. When a new text prompt is provided, it is converted into a vector and efficiently checked against the vectors in the semantic database 207. The asset associated with the vector most closely matched to the vector representing the new text prompt is then retrieved. The database 207 uses Facebook ® AI Similarity Search (FAISS), which is an efficient vector database, to find the closest vector, for example. The asset builder 205 thus uses the generative model 206 and / or the semantic database 207 to obtain assets from input speech / text. The generative model 206 allows new assets to be generated whereas the semantic database 207 allows existing assets to be retrieved quickly and efficiently. Generated and / or retrieved assets are added to the asset database 208 for use by the simulator 204. The semantic database 207 thus acts as catalogue of assets while the asset database 208 consists of the assets usable by the simulator 204 to realise the current simulation. This allows fast retrieval of the assets by the simulator 204 based on relevant information in the scene schema (the scene schema is described below). Assets added to the asset database 208 may also be processed before being made available to the simulator 204 to make them more appropriate for use by the simulator. For example, image or video assets may be subsampled to obtain a lower spatial and / or temporal resolution to improve the processing speed of the simulator 204. The scene extractor 201 starts from a raw audio narration (picked up by microphone 109) and outputs a scene specification (scene schema). The scene schema is a machine- readable file (e.g. a json or yaml file) specifying all of the relevant information to define a scene. The scene schema is used by the asset builder 205 and simulator 204 to construct a simulated environment (simulated scene). The simulated environment / scene may also be referred to as a virtual environment / scene. The streamer 202 then extracts values from the simulated environment which are used to control the environmental output devices 110. The scene extractor 201 implements multiple steps. These are illustrated in Fig.3. A first step 301 is to take the audio stream input to the microphone 109 and convert speech in the audio stream to text. Any suitable known speech-to-text library (e.g. Whisper from OpenAI ®) may be used for this step. The next step 302 is a scene monitor step. The scene monitor step acts as a gateway to the scene extractor 201. In particular, it serves to determine the context of the audio description and, based on this context, decide when to transmit information (in particular, the obtained text information) to the next step of the scene extractor (along with information to contextualise what is happening, such as a scene change). In one example, the scene monitor transmits every sentence as it is completed (e.g. when words are uttered but then there is a pause in further words being uttered for a predetermined period of time, it is determined the sentence has ended and the text of the sentence is transmitted). In this case, the scene monitor does not know anything about its environment or the task it is being used for. Rather, it buffers streamed transcripts (converted text) and passes them to the extractor in batches (e.g. after every sentence). In another example, conversation is actively monitored and irrelevant content is not transmitted. In this case, the scene monitor needs to be provided with information about the rules and protocols of the game so as to be able to determine when the received text is relevant (and should thus be further processed) and when it is not (and should thus be discarded). The monitor may also detect when the scene changes (e.g. from a forest scene to a cave scene) and provide this information to the next step of the scene extractor 201. This may be automatic based on a suitable classification of the latest received input text or based on a set phrase or the like (e.g. “New Scene” or “Scene Change”) uttered by one of the users. The scene extracting step 303 generates the scene schema. Based on the text passed on by the scene monitor at step 302, the scene schema is generated by identifying the setting, expanding on the setting, identifying virtual light and sound sources (and, if appropriate, any other virtual sources such as wind, heat or smell sources), placing the virtual sources in a virtual environment and identifying any dynamics for the sources. Virtual sources (which may be referred to simply as “sources”) are sources to be included in a simulated environment which generate, virtually, a detectable sensation. For example, light sources (e.g. the sun, a fire, lightning) generate light in the simulated environment and audio / sound sources (e.g. thunder, a fire crackling, a gun firing) generate sound in the simulated environment. The generation of the scene schema uses an instruction-trained large language model (LLM) such as ChatGPT from OpenAI ®. In particular, a sequence of instructions is provided to the LLM based on the text received from the scene monitor step 302. The sequence of instructions include instructions to: ^ Understand, based on the text received from the scene monitor step 302, the semantic context of the scene. ^ Expand on the scene using general knowledge available to the LLM of what is typically present in a scene. ^ Focus the context on sources in the scene such as audio and light sources (and, potentially, other sources such as wind, heat or smell sources). ^ Arrange the location of the sources according to an explicit instruction (e.g. the fireplace is on the right) and / or based on general knowledge (e.g. fireplaces are not on the ceiling but can be in a wall). ^ Output the resulting definition of the scene as a scene schema (e.g. as a json or yaml file) in a predetermined format. Providing a sequence of instructions in this way allows the LLM to incrementally generate information about the scene in context and consider all relevant information when generating the output scene schema. An example sequence of instructions provided to the LLM to generate the scene schema is shown in the Appendix (which forms part of this description). In an example, information about the dynamics of the sources (e.g. the light and sound sources) is determined. For example, output of a source may occur a single time or may occur continuously. For instance, there is a difference between a background noise of gunshots in the distance and a single gunshot nearby. The former is an ambient noise and the latter is more likely to correspond with an event in the story. In an example, the task of composing source dynamics into a coherent set of events is carried out by a composer part of the simulator 204. It does this based on general information about the type of dynamics that are occurring for a particular source (e.g. “continuous”, “periodic”, “single”, etc.) defined in the scene schema, for example. Dynamics other than those defined by this general information may be directly captured in the relevant asset itself. So, for example, a description “the ambient sound of the sea” will map to a first asset representing a long continuous sound of the sea while a description “a single wave splashing” will map to a second asset representing a single short sound of a wave. In an example, a more complex scene may be simulated across multiple rooms simultaneously. This may include, for example, different sources in different virtual rooms map to corresponding environmental output devices 110 in different physical rooms. For instance, if the virtual rooms are those of a castle, one room may be barracks where knights are training, another room may a food court where people are eating and music is playing and another room may be a dungeon which is dark and lit by flaming torches. These may be mapped to respective physical rooms in a house, museum or the like so environmental output devices 110 in each of these rooms output lighting and sound (and any other output(s)) corresponding to the relevant virtual room. This allows a more complex and immersive experience and may be suitable for museums, parties, escape rooms or the like. The asset builder 205 takes the sources defined in the scene schema and obtains corresponding assets (the assets being generated by the model 206 and / or retrieved from the semantic database 207) that can be included in a simulated scene created by the scene simulator 204. The obtained assets are stored in the asset database 208 for use by the simulator 204. The obtained assets may include audio sources and / or light sources (as previously mentioned) but also images and / or videos (and, depending on the availability of specialist environmental output devices 110, heat sources, wind sources, smell sources, etc.). In an example, images and / or videos may be used in an indirect way to extract lighting information (e.g. chrominance information) to be applied to a light source in a simulated scene. For instance, pixel information of an image or video asset may be analysed to extract light information algorithmically. For example, threshold(s) may be determined to exclude pixels with certain luminance (Y) value(s) (e.g. those with a luminance below a predetermined threshold to exclude shadows) and an average of the pixel values (Y, Cb, Cr) of the remaining pixels may be determined as the dominant light colour either a single time (for an image) or multiple times (for different frames of a video). This is only an example and any suitable technique may be used to obtain light information for a light source from an image or video. Images may also be used directly as textures in the simulated scene (as described later). The asset builder 205 carries out a plurality of functions to generate and / or obtain assets based on the scene schema received from the scene extractor 201. These are illustrated in Fig.4. An audio function 401 retrieves or generates an audio asset that matches a textual description in the scene schema. As previously described, there may also be additional information (dynamics information) provided in the scene schema indicating the dynamics of the audio (e.g. indicating whether the sound is an ambient, continuous sound or whether it is associated with a single event). This information can be included in the text description itself (e.g. “an ambient sound of the sea”), meaning the dynamics information is used directly in the selection or generation of a sound asset. Alternatively, for retrieval of existing sound assets from the semantic database 207, for example, it can be separate to the text description (e.g. as a separate value of the scene schema) and used to direct queries against databases tables that are partitioned according to their type (e.g. “ambient” sound assets or “single” sound assets). As previously mentioned, to enable retrieval of a suitable audio asset from the semantic database 207, the audio assets of the semantic database 207 are represented as a vector embedding (e.g. using CLAP). The semantic database 207 may be a FAISS database, for example. During retrieval, the textual description of the audio from the scene schema is embedded as a vector (e.g. again using CLAP) and this semantic embedding is queried against the semantic database 207. The closest matching asset is then extracted. Alternatively, instead of retrieving an existing audio asset from the semantic database 207, the vector embedding of the textual description may be input to any suitable generative sound neural network (e.g. implemented as at least part of the generative model 206). A first light function 402 (static light function) takes, as input, a textual description of a type of static light included in the scene schema and outputs a colour representing this type of light. For example, for the text “the sun over the horizon”, a colour representing this type of light is output. Variations in the context might change the output light (e.g. “setting sun” vs “midday” sun). As previously mentioned, the colour of light of a virtual light source may be determined by obtaining a suitable image, using predetermined threshold(s) to exclude pixels in the image unlikely to be related to lighting (e.g. those with a luminance value less than a predetermined threshold value to exclude shadows) and averaging the pixel values of the remaining pixels. The luminance of the virtual light source may be fixed or adjustable (while keeping the chrominance values the same) to provide appropriate illumination of the simulated environment. A fixed luminance value for a virtual light source may be indicated in the scene schema, for example. An adjustable luminance value for a virtual light source may be adjustable based on an instruction received from a user via user interface 105, for example. When the luminance is adjustable, the chrominance values (e.g. Cb and Cr) of the virtual light source may remain fixed (e.g. to those derived from the original image used for generating the virtual light source). As also previously mentioned, to enable retrieval of a suitable image from the semantic database 207, the images of the semantic database 207 are represented as a vector embedding (e.g. using CLIP). The semantic database 207 may be a FAISS database, for example. During retrieval, the textual description of the light source is embedded as a vector (e.g. again using CLIP) and this semantic embedding is queried against the semantic database 207. The closest matching asset is then extracted. Alternatively, instead of retrieving an existing image in the semantic database 207, the vector embedding of the textual description may be input to any suitable generative image neural network (e.g. implemented as at least part of the generative model 206). A second light function 403 (dynamic light function) takes, as input, a textual description of a type of dynamic light included in the scene schema and outputs a list of colours to be output in sequence (e.g. periodically in a round robin fashion or in a (pseudo) random order) to represent this type of light. For example, text inputs such as “fire”, “fireworks”, “thunder” or the like may be associated with such a colour list which characterises the dynamics of the light source. In an example, this works in a similar way to the generation of a single colour for a static light source, except that multiple frames of a video associated with the textual description are used to obtain multiple respective colours to be added to the list. A dynamic mask may be used to focus on parts of the video that change most over time. The pixels of such part(s) of a given frame (e.g. those with a change in luminance value which exceeds a threshold amount with respect to the previous frame) are thus included when generating the output colour for that frame (while other pixels are excluded). In an example, information indicative of single colours (for static light sources) and colour lists (for dynamic light sources) are themselves stored as assets in the asset database 208 once generated by the asset builder 205. A location function 404 takes into account that a scene is not just a collection of light and sound sources but has other features which help characterise the scene. For instance, light sources reflect off surfaces which absorb some of the spectrum of the light, changing whatever light then returns to the eye. As such, it is desirable to capture characteristic(s) of surface(s) in the scene illuminated by the light from the light source(s). Thus, in an example, one or more image(s) corresponding to a textual description in the scene schema of what is surrounding the user are retrieved or generated (such image(s) being retrieved or generated in the same way as described for the first light function 402, for example). The extracted image(s) are used to create a complete simulated environment (scene). These can then be used to determine a colour profile that can be captured by a virtual light sensor (for example, a camera) in the simulated scene and used to stream corresponding colour information to the physical light(s) 106. Thus, for example, if the location is a desert during daylight (with sand at the bottom of the picture and the sky at the top of the picture), then the virtual sensor will capture more sand-like colours from nearer the bottom of the simulated environment and more sky-like colours from nearer the top of the simulated environment. Fig.5 shows example outputs of the extractor 201 and asset builder 205. They include an audio file 701 of speech detected by the microphone 109, a speech-to-text transcript 702 of the audio file 701, the scene schema 703, image assets 704 generated and / or retrieved based on the scene schema 704 (which may be used as textures and / or used to determine chrominance and / or luminance values of a virtual light source) and audio assets 705 (in this case, audio files incorporating sounds of ambient conversation and sounds of a crackling fire) to be used for virtual sound sources. An example of a simulated environment including sounds, colours and images generated by functions 401 to 404 is shown in Fig.6A. A corresponding physical environment 503 (room) with three lights 106A-C and a speaker 107 to which values generated from virtual sensors in the simulated environment are streamed is shown in Fig.6B. The simulated (virtual) environment 505 is generated by the simulator 204, which takes the scene schema and retrieved and / or generated assets and uses them to construct the simulated environment 505 using, for example, a suitable 3D game engine. The scene schema may not completely capture all information required by the simulator 204 to create the simulated environment 505. The simulator 204 thus comprises a composer function that fills in any gaps and makes the overall simulation more coherent. This involves, for example, setting up and running the dynamic light and sound elements as well as ensuring the overall lighting conditions work well together. The composer is described later. In an example, functionalities of the simulator 204 are fully supported by any suitable existing game engine. Such functionalities include, for example: ● Placing the light source(s) within the scene. ● A lighting engine to simulate the behaviour of light from the light source(s). ● An ability to create and place object(s) that can be used to construct a scene. ● Applying textures to object(s). ● Placing a first virtual camera within the scene. ● Placing the sound source(s) within the scene. ● Placing a virtual sound sensor in the scene to detect sound from the virtual sound source(s) relative to its position in the scene. The virtual sound sensor represents a virtual user (virtual listener). ● Placing a second virtual camera in the scene which captures visual information from the perspective of the virtual user. This may be different to the first virtual camera which, for example, captures an image of the scene including the position of the virtual user from a wide angle perspective. The simulated environment 505 of Fig.6A (shown from the perspective of the first virtual camera) is generated using these functionalities. The scene is a forest with a dynamic light source representing lightning and an audio source representing thunder. These are located at the same location 501 in the simulated environment 505, in this example. The simulated environment 505 also includes, as objects, virtual walls with textures on them corresponding to the forest scene. The walls include three perpendicular vertical walls 506A, a base wall 506B and an upper wall 506C. A tree texture is applied to the vertical walls 506A, a ground texture is applied to base wall 506B and a sky texture is applied to upper wall 506C. These textures are obtained by the location function 404 of asset builder 205, for example. The textures are illuminated by the dynamic light source (lightning) based on the simulated behaviour of light emitted by the dynamic light source. The virtual user is positioned at virtual user position 502. Audio output of the speaker 107 is controlled based on the virtual user position 502. For example, audio output is controlled to be louder when the virtual listener position 502 is closer to the source location 501 and quieter when the virtual listener position 502 is farther from the source location 501. In this sense, the output of sound from the source location is localized. Similarly, light output (including chrominance and / or luminance output) of the lights 106A-C is controlled based on the virtual user position 502. In particular, one or more virtual cameras are placed at the virtual user position 502 and a portion of the scene captured from the perspective of each of these virtual cameras is used to determine the output of the correspondingly positioned lights 106A-C in the physical room 503. To ease illustration, such portions of the scene are illustrated in Fig.6A as light patches 504 distinct from the rest of the scene. In reality, they are portions of the scene itself (e.g. rectangular portions of one or more of the walls 506A-C positioned depending on the orientation of the relevant virtual camera) and are images used to generate a light colour which the lights 106A-C are then controlled to output. In particular, image 504A (representing a portion of the scene captured from the perspective of one virtual camera located at user position 502) is used to generate an output of light 106A, image 504B (representing a portion of the scene captured from the perspective of another virtual camera located at user position 502) is used to generate an output of light 106B and image 504C (representing a portion of the scene captured from the perspective of yet another virtual camera located at user position 502) is used to generate an output of light 106C. In this way, a mapping is established between positions in the virtual environment 505 and positions in the room 503. This provides a real life user 507 in the room 503 with lighting which corresponds with that of the simulated scene 505. The generation of a light colour from each of the images 504A-C may be determined in a similar way to that of the colour of virtual light sources from images (as previously described). For example, threshold(s) may again be determined to exclude pixels with certain luminance (Y) value(s) (e.g. those with a luminance below a predetermined threshold to exclude shadows) and an average of the (Y, Cb, Cr) values of the remaining pixels may be determined as the light colour. Information representing the light colour (Y, Cb, Cr) of each image 504A-C is then provided (by streamer 202) to its respective light 106A-C. Note that, in this description, Y relates to luminance (where Y’, luma, can also be used), Cb is the blue-difference chroma component and Cr is the red-difference chroma component. The asset builder 205 thus provides the scene schema (indicating the position 501 of the virtual sound and light sources and the positions of the walls 506A, 506B and 506C in the simulated environment 505), a sound asset (representing the thunder), a colour list (representing the lightening, which is a dynamic light source in this case – if the light source is static, a single colour is provided instead of the colour list) and the texture (which is an image) for each wall to the simulator 205 to build the simulated environment 505. The schema may also indicate, for example, a brightness of the light source (e.g. as a luminance value Y, where the colours of the light source, represented by chrominance values Cb and Cr, have already been determined by the dynamic light source function 403 of the asset builder 205 in the way previously described) and a volume of the audio asset. Other information such as spatial sound information (e.g. a level of echo and / or reverb) can also be provided (e.g. as part of the scene schema or through manual configuration of the simulator 204 via user interface 105) applied as effects to the audio asset. In particular, if there are multiple speakers 107 distributed in the physical environment 503, correspondingly placed virtual sound sensors in the simulated environment 505 can extract the appropriate localized sound(s) (from the virtual sound source(s)) and apply appropriate spatial sound information. For example, if the virtual environment is a cave, spatial sound information can be applied to help make the audio sound like it would in a cave. The composer function of the simulator 204 takes the scene schema and assets and, using e.g. the 3D game engine, sets various material and / or lighting properties to create coherent lighting conditions in the simulated scene. In one example, the textured walls 506A-C are associated with predefined (e.g. fixed) values for the material properties of the wall surfaces. However, to create more realistic lighting information (to be provided to the lights 106A-C) and thus a more immersive experience, different material properties may also be defined for each wall surface to capture the way in which light from the virtual light source interacts with surfaces. This takes into account, for example, that wet surfaces and metal surfaces reflect light differently than dry matt surfaces. Example ways in which how light interacts with objects can be specified to emulate different materials is shown in Fig.7. These are applicable to objects (e.g. textured walls 506A-C) in the simulated environment 505 by the game engine using any suitable known technique(s). Example 601 shows an ambient light interaction with an object surface. Example 602 shows a diffuse light interaction with an object surface. Example 603 shows a specular light interaction with an object surface. Example 604 shows a combined (e.g. diffuse combined with specular) light interaction with an object surface. The composer function is also configured to transform the scene schema and assets provided by the asset builder 204 into a convincing dynamic simulation (e.g. one in which dynamic light and / or sound sources change in a convincing way over time). For example, the composer function may be configured to construct a timeline or definition of the dynamics of when certain events are triggered. For example, for audio assets, the composer function may apply certain rules depending on a characteristic of the sound (e.g. whether the sound asset is classed as “ambient” or “single”). For instance, “ambient” sounds (e.g. thunder, background conversation or the like) may be continuously looped and “single” sounds are provided either once (e.g. a single gunshot) or multiple times separated by an interval (e.g. a bird call). The interval may be fixed or may change (e.g. it may be determined probabilistically). In an example, a graphical interface of the location editor 203 may be provided (e.g. via user interface 105) for constructing the dimensions of a real room (which can be input as boundaries to the simulator 204). Information representing the relationship between different connected physical rooms (each being associated with a different respective simulated scene) can also be provided by this graphical interface. This allows for a simulation to simultaneously occur across multiple rooms or for users to physically change rooms as they change scenes in a story, for example. The graphical interface may also be used to adjust the virtual user position 502, the number and position of virtual sound sensor(s) (in particular, if there is more than one virtual sound sensor which, by default, may be located at the virtual user position 502) and the number and viewing perspectives of virtual cameras located at the virtual user position 502 (which generate the images 504A-C). The graphical user interface may also be used to associate the location of each portion of the scene from which each image 504A-C is generated with its corresponding light 106A-C. For example, the graphical user interface may show the simulated environment 505 from the perspective of the above-mentioned first virtual camera (as exemplified in Fig. 6A) and allow the user to indicate (e.g. via a touch screen or mouse) the position of each light 106A-C in the physical room at a corresponding position in the virtual environment 505. The position of each light in the virtual environment 505 then becomes the position of the portion of the scene captured by the virtual camera used to generate the image used to generate the light output for that light (e.g. image 504A for light 106A, image 504B for light 106B and image 504C for light 106C). Once the simulator 204 is running the dynamic simulated scene 505, the streamer 202 is used to stream values from the simulator to the environmental output devices 110. For example, using user interface 105, the user indicates the environmental output devices 110 they have, how they correspond to virtual devices and / or where those virtual devices are placed in the simulated (e.g. as exemplified above for the lights 106A-C). The streamer 204 places virtual measuring devices in the simulator (e.g. a virtual microphone / sound sensor and virtual cameras / light sensors at user position 502, takes appropriate measurements and outputs information based on these measurements to the appropriate environmental output devices 110. For example, light values (Y, Cb, Cr) are output to each of the lights 106A-C in real time and an audio stream (e.g. based on a looped thunder sound asset) is output to the speaker 107. Other information (e.g. simulated heat, wind, smell or the like) may also be output to appropriate other environmental devices (e.g. heater, fan, scene diffuser or the like). Although only a single speaker 107 is shown, it will be appreciated that multiple speakers 107 could be used. Such speakers may be configured with, for example, localisation or panning effects to allow sound in the simulated environment 505 to appear to be come from a corresponding location in the physical environment 503. In one example, each speaker in the physical environment 503 is controlled to output sound as detected by a respectively positioned virtual sound sensor in the simulated environment 505. In an example, to take into account any latency of smart light control (e.g.1 / 10 second), only significant changes to the light output (e.g. changes in Y, Cb and / or Cr which exceed a predetermined threshold) are provided to the light(s) 106. This helps reduce the perceived effects of smart light latency. In particular, dynamic changes in the lighting of the simulated environment 505 are captured by the virtual sensor(s). These changes run at a particular frame rate of the game engine (for example, 60 frames per second, corresponding to 60 potential changes per second). The physical light(s) 106 need to be controlled according to these changes. However, the physical lights may run at a lower change rate than the frame rate of the game engine. For example, it may be that the lights only be changed every 1 / 10 of a second. Instead of indicating all changes to the light (e.g.60 changes in a second), it is thus desirable to only indicate changes at the refresh rate of the light (e.g.10 changes in a second). Thus, for example, if, each second, 10 light changes are to be indicated to a physical light from 60 possible changes generated by the game engine, those 10 light changes need to be chosen. In an example, the nthlight change (where n = 1 to 10) indicated to the physical light is chosen from the nthset of 6 possible changes generated by the game engine (so the 1stlight change of the physical light is chosen from changes 1 to 6 of the game engine, the 2ndlight change of the physical light is chosen from changes 7 to 12 of the game engine, and so on). Over a given time period (in this case, 1 / 10thsecond), each change of the physical light is thus determined from the 6 corresponding lighting changes of the game engine. For example, the physical light output may be controlled to be the average of the 6 game engine light outputs or to be the one of the 6 game engine light outputs with the maximum brightness. Moreover, in an example, if the light output of the game engine does not change (or does not change by more than a threshold brightness, for example) over the given time period, no update is transmitted to the physical light (so the physical light maintains is current light colour and brightness). This helps reduce the communication overhead. In the above examples, the input to the extractor 201 is audio comprising speech captured by the microphone 109. However, the present technique is not limited to this and other inputs may be used (instead of or in addition to captured speech) by the extractor 201 to generate words and, using the generated words (e.g. in an LLM), generate the scene schema. For example, the input to the extractor 201 may be an image comprise text (e.g. a captured image of a page of a book) and the extractor 201 may be configured to perform optical character recognition (OCR) on the image to obtain the text. This text may then be used to generate the scene schema in the way described. Text may also be provided to the extractor 201 directly via user interface 105 (e.g. by typing). In another example, the input to the extractor may be an audio or video file or stream (e.g. received over a network via communication interface 104). In this case, the extractor 201 may detect and perform text-to-speech on speech on audio. It may also be configured to classify other sounds (e.g. using any suitable known machine learning sound classification technique) and output a textual classification of such sounds for use in generating the scene schema. For example, sounds in an audio or video file or stream such as “gunshot”, “water flowing”, “dog barking” or the like may be detectable and classifiable by the extractor 201. Furthermore, for an image or video file or stream, object and / or image classification may be used by the extractor 201 (e.g. using any suitable known machine learning object detection and / or image classification techniques) to classify images or video frames (or detected objects in images or video frames) and output a suitable textual classification. For example, an image of a parasol on a beach at sunset may be classified as “sunset” or and / or detected objects in the image may be classified as “sand”, “sea” and “parasol”. Thus more generally, inputs to the extractor 201 for generation of the scene schema may include audio (as picked up by microphone 109 or as included in an audio or video file or stream), images, video and / or text. In an example, the same virtual environment 505 may be generated for users in a plurality of geographical locations. In this case, the virtual object(s), light source(s) and sound sources(s) are the same for all users but the position(s) of the virtual light and sound sensor(s) may be adjusted according to the light and sound device(s) 106 and 107 of each user. For instance, a first user in a first geographical location may only have a single light device 106 and a single speaker 107 located in the centre of a room and a second user in a second geographical location may have four light devices 106 and four speakers 107 located in each corner of a room. For the first user, a single virtual camera capturing a central portion of the virtual scene is provided and a single virtual sound sensor is placed in the centre of the virtual scene. On the other hand, for the second user, four virtual cameras each capturing respective corner portion of the virtual scene is provided and four virtual sound sensors are each placed in a respective corner of the virtual scene. This allows each user to have a bespoke immersive experience. To implement this, each user may have an instance of apparatus 100. One of the apparatuses 100 generates the virtual scene (e.g. from speech uttered by its corresponding user) and transmits data indicative of the virtual scene (e.g. over a network such as the internet via communication interface 104) to the apparatus 100 of each of the other users. Each user then provides information on the environmental output devices 110 available to them (including the number and location of light and sound device(s) 106 and 107) via, for example, the user interface 105. Each apparatus 100 then outputs control signals to its connected environmental output devices 110 accordingly. In an example, once a scene schema has been generated, it may be updated over time as more information is received (e.g. speech input via the microphone 109) to reflect changes in the scene. For instance, the scene is a scene in a forest at night but a narrating user says “Dawn begins to break”, the light source in the scene schema (and therefore the virtual scene 505) may be updated from the moon to a rising sun while the rest of the scene schema (and therefore the virtual scene 505) remains the same. This provides an efficient way of updating the characteristics of a single virtual scene over time. On the other hand, when the scene changes (e.g. going from a forest scene to inside a tavern, as detected by the scene monitor step 302), a completely new scene schema may be generated. The present technique thus allows environments to be simulated dynamically based on detected relevant speech of users using lights, speakers and the like. For instance, simply based on the speech of a game master (“Our heroes walk into a tavern, they sit beside a fireplace and listen to the conversation of the clients around them” or “Walking along the beach at night with the sea to their right and fire to their left, our heroes hear distant birds”), the output lighting and sounds are automatically adjusted to provide a corresponding ambience in the room which matches the described scene. This helps provide new and immersive simulated experiences for users. Furthermore, since it is based on speech already being uttered by the user (e.g. as part of a role playing game), there is a reduced need for the user to take extra steps (e.g. to manually configure settings of the simulated environment), therefore reducing the burden on the user. Fig. 8 shows an example method according to the present technique. The method is implemented by the processor 101, for example. The method starts at step 801. At step 802, information derived from digital content is received. The digital content is, for example, audio (e.g. audio picked up by microphone 109 or included in an audio or video file or stream), an image, a video and / or text. In some of the example(s) above, the digital content is a digital representation of speech uttered by users playing a table top role playing game and picked up by microphone 109. The information derived from the digital content is then the scene schema generated by providing a speech-to-text transcript of the speech and appropriate sequential instructions to an LLM (as previously described). The information may also be the speech-to-text transcript itself (or, indeed, an audio recording of the speech) and the scene schema may then be generated from this information. At step 803, a virtual environment (e.g. virtual environment 505) is generated using one or more digital assets corresponding to the received information. For example, when the virtual environment is generated based on a scene schema, the scene schema may define: ^ A position of one or more virtual light sources in the virtual environment (e.g. position 501 for the lightning in virtual environment 505) and information for determining a brightness and / or colour of the virtual light source. The brightness and / or colour information (e.g. values in the YCbCr colour space) is determined from a digital image or video asset (e.g. a video asset of lightning in virtual environment 505). ^ A position and form of one or more textured objects in the virtual environment (e.g. walls 506A-C in virtual environment 505) and a texture image (e.g. tree texture, ground texture or sky texture) to be applied to a surface of each textured object as a texture. ^ A position of one or more virtual sound sources in the virtual environment (e.g. position 501 for the thunder in virtual environment 505) and audio content (e.g. an audio asset of thunder) to be output by the virtual sound source(s). At step 804, information from the virtual environment is extracted. For example, a portion of the virtual environment may be captured as an image using a virtual camera (e.g. images 504A-C are captured from respective virtual cameras positioned at virtual user position 502 in the virtual environment 505) and audio content output by the virtual sound source may be captured using a virtual sound sensor (also located at virtual user position 502). At step 805, a control signal to one or more output devices (e.g. light(s) 106 and speaker(s) 107) of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment. For example, a brightness and / or colour of each of the lights 106A-C is controlled based on the respective virtual camera images 504A-C (as previously explained) and the speaker 107 is controlled to output the audio content (thunder) as captured by the virtual sound sensor at user position 502. The method ends at step 806. Example(s) of the present disclosure are defined by the following numbered clauses: 1. A data processing apparatus comprising circuitry configured to: receive information derived from digital content; generate a virtual environment using one or more digital assets corresponding to the received information; extract information from the virtual environment; and output a control signal to one or more output devices of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment. 2. A data processing apparatus according to clause 1, wherein: the one or more digital assets include a virtual light source image or video; the virtual environment comprises a virtual light source; and a brightness and / or colour of the virtual light source is determined using the virtual light source image or video. 3. A data processing apparatus according to clause 2, wherein: the one or more digital assets include a texture image; the virtual environment comprises a textured object, the textured object comprising a surface to which the texture image is applied as a texture; and the surface of the textured object is illuminated in the virtual environment by the virtual light source. 4. A data processing apparatus according to any preceding clause, wherein: the one or more output devices of the physical environment comprise one or more light devices; and the circuitry is configured to: capture a one or more portions of the virtual environment using one or more virtual light sensors; and control a brightness and / or colour of the one or more light devices based on the one or more captured portions of the virtual environment. 5. A data processing apparatus according to any preceding clause, wherein there is a mapping between positions in the virtual environment and positions in the physical environment. 6. A data processing apparatus according to any preceding clause, wherein: the one or more digital assets include audio content; and the virtual environment comprises one or more virtual sound sources which output the audio content. 7. A data processing apparatus according to clause 6, wherein: the one or more output devices of the physical environment comprise one or more speaker devices; and the circuitry is configured to: capture the audio content output of the one or more virtual sound source(s) using one or more virtual sound sensor; and control the one or more speaker devices to output the audio content captured by the one or more virtual sound sensors. 8. A data processing apparatus according to clause 6 or 7, wherein, depending on a characteristic of each virtual sound source, the virtual sound source outputs the audio content a single time, multiple times or on a continuous loop. 9. A data processing apparatus according to any preceding clause, wherein: the received information derived from the digital content comprises a scene schema, the scene schema defining the virtual environment; and the circuitry is configured to generate the scene schema by: providing words derived from the digital content to a large language model (LLM); providing sequential instructions to the LLM to build the scene schema based on the words. 10. A data processing apparatus according to clause 9, wherein the scene schema defines one or more of: a position of a virtual light source in the virtual environment and information for determining a brightness and / or colour of the virtual light source; a position and form of one or more textured objects in the virtual environment and a texture image to be applied to a surface of each textured object as a texture; and a position of a virtual sound source in the virtual environment and audio content to be output by the virtual sound source. 11. A data processing apparatus according to any preceding clause, wherein the digital content comprises one or more of audio, images, video and text. 12. A data processing apparatus according to clause 11, wherein the audio comprises speech uttered by one or more users. 13. A data processing apparatus according to any preceding clause, wherein: the circuitry is configured to transmit information representing the virtual environment to a second data processing apparatus; and a position of virtual light and / or sound sources is adjustable depending on respective physical environments of the data processing apparatus and second data processing apparatus. 14. A data processing apparatus according to any preceding clause, wherein the one or more digital assets are generated by a generative artificial intelligence, AI, model and / or retrieved from a database. 15. A computer-implemented data processing method comprising: receiving information derived from digital content; generating a virtual environment using one or more digital assets corresponding to the received information; extracting information from the virtual environment; and outputting a control signal to one or more output devices of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment. 16. A program for controlling a computer to perform a method according to clause 15. 17. A computer-readable storage medium storing a program according to clause 16. Numerous modifications and variations of the present disclosure are possible in light of the above teachings. It is therefore to be understood that, within the scope of the claims, the disclosure may be practiced otherwise than as specifically described herein. In so far as embodiments of the disclosure have been described as being implemented, at least in part, by one or more software-controlled information processing apparatuses, it will be appreciated that a machine-readable medium (in particular, a non-transitory machine-readable medium) carrying such software, such as an optical disk, a magnetic disk, semiconductor memory or the like, is also considered to represent an embodiment of the present disclosure. In particular, the present disclosure should be understood to include a non-transitory storage medium comprising code components which cause a computer to perform any of the disclosed method(s). It will be appreciated that the above description for clarity has described embodiments with reference to different functional units, circuitry and / or processors. However, it will be apparent that any suitable distribution of functionality between different functional units, circuitry and / or processors may be used without detracting from the embodiments. Described embodiments may be implemented in any suitable form including hardware, software, firmware or any combination of these. Described embodiments may optionally be implemented at least partly as computer software running on one or more computer processors (e.g. data processors and / or digital signal processors). The elements and components of any embodiment may be physically, functionally and logically implemented in any suitable way. Indeed, the functionality may be implemented in a single unit, in a plurality of units or as part of other functional units. As such, the disclosed embodiments may be implemented in a single unit or may be physically and functionally distributed between different units, circuitry and / or processors. Although the present disclosure has been described in connection with some embodiments, it is not intended to be limited to these embodiments. Additionally, although a feature may appear to be described in connection with particular embodiments, one skilled in the art would recognize that various features of the described embodiments may be combined in any manner suitable to implement the present disclosure.

[0002] Appendix Example chain-of-thought reasoning: INPUT: "Walking along the beach at night with the sea to far to my right and fire to my left. I hear distant birds every now and again." SETTING: Instruction: "The following is a narrative piece of text. Extract the setting and describe the surroundings in very simple language: " Output: “The setting is a beach at night. The surroundings are simple and peaceful. On one side, there is the vast expanse of the sea, stretching far into the distance. On the other side, there is a fire, emitting a warm glow that illuminates the immediate area. The sound of birds can be heard from a distance, intermittently breaking the quiet atmosphere.” LOCATION: Instruction: “Provide an image caption for what can be seen in the setting from different directions using the following schema: View Front: description | View Right: description | View Left: description | View Ground: description” Output: 'View Front: A serene beach stretches out before you, with gentle waves lapping onto the shore. The soft sand glows in the moonlight. View Right: The sea extends far into the horizon, its dark waters glistening under the starry night sky. It seems endless and mysterious. View Left: A crackling fire emits a warm glow, casting dancing shadows on the sand. Its flickering flames create a cozy ambiance. View Ground: The sand beneath your feet is cool and inviting, with scattered seashells and small pebbles adding texture to the beachscape.’ EXPANDED SETTING: Instruction: "List the sources of sound and light present in this setting using simple language. Include where they are located and how far away they are. " Output: 'Sources of sound and light in this setting are as follows:\n\nSound Sources:\n1. Distant Birds: The sound of birds can be heard sporadically, coming from a distance.\n2. Gentle Waves: The waves of the sea create a soothing sound as they break onto the shore.\ Light Sources: 1. Moonlight: The soft glow of the moon illuminates the entire beach and surrounding area.2. Fire: A crackling fire is located to the left, casting a warm and flickering light nearby.' SOURCES: Instruction: “ Light sources and can either be ambient or directed. A directed light comes from an obvious position in the setting. A directed light is typically brighter than an ambient light source. An ambient light doesn't have an obvious source but provide general illumination across the setting. Usually, the further away a sound source is, the quieter it is and the closer it is, the louder it is. Light and sound sources have particular vertical and horizontal positions. The vertical positions are: below, middle, above, far below, far above. Where 'far' here means a greater distance above or below. Horizontal positions are: left, centre, right, far left, far right. Where 'far' here means a greater distance right or left. The distances: must be either near, medium or far. Extract a list of light sources and sound sources." Output: 'Light Sources:\n1. Directed Light: The fire to the left provides a directed light source that is brighter than the ambient light.\n2. Ambient Light: The moonlight provides an ambient light source that illuminates the entire setting.\n\nSound Sources:\n1. Distant Sound: The birds can be heard from a distance, indicating a sound source that is far away.\n2. Close Sound: The gentle waves breaking onto the shore create a sound source that is closer to the listener.', SCHEMA: Instruction: “Provide a list of light sources and sound sources in json according to the following schema. Do not include any explanations, only provide a RFC8259 compliant JSON response following this format without deviation. Output: '[{\n "modality": "light",\n "source_type": "directed",\n "description": "fire",\n "volume": null,\n "intensity": "high",\n "frequency": "continuous",\n "horizontal_position": "left",\n "vertical_position": "middle",\n "action": "static"\n}, {\n "modality": "light",\n "source_type": "ambient",\n "description": "moonlight",\n "volume": null,\n "intensity": "medium",\n "frequency": "continuous",\n "horizontal_position": null,\n "vertical_position": null,\n "action": "static"\n}, {\n "modality": "sound",\n "source_type": null,\n "description": "birds",\n "volume": "quiet",\n "intensity": null,\n "frequency": "occasional",\n "horizontal_position": null,\n "vertical_position": null,\n "action": "static"\n}, {\n "modality": "sound",\n "source_type": null,\n "description": "gentle waves",\n "volume": "moderate",\n "intensity": null,\n "frequency": "continuous",\n "horizontal_position": null,\n "vertical_position": null,\n "action": "static"\n}

Claims

CLAIMS 1. A data processing apparatus comprising circuitry configured to: receive information derived from digital content; generate a virtual environment using one or more digital assets corresponding to the received information; extract information from the virtual environment; and output a control signal to one or more output devices of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment.

2. A data processing apparatus according to claim 1, wherein: the one or more digital assets include a virtual light source image or video; the virtual environment comprises a virtual light source; and a brightness and / or colour of the virtual light source is determined using the virtual light source image or video.

3. A data processing apparatus according to claim 2, wherein: the one or more digital assets include a texture image; the virtual environment comprises a textured object, the textured object comprising a surface to which the texture image is applied as a texture; and the surface of the textured object is illuminated in the virtual environment by the virtual light source.

4. A data processing apparatus according to claim 1, wherein: the one or more output devices of the physical environment comprise one or more light devices; and the circuitry is configured to: capture a one or more portions of the virtual environment using one or more virtual light sensors; and control a brightness and / or colour of the one or more light devices based on the one or more captured portions of the virtual environment.

5. A data processing apparatus according to claim 1, wherein there is a mapping between positions in the virtual environment and positions in the physical environment.

6. A data processing apparatus according to claim 1, wherein: the one or more digital assets include audio content; and the virtual environment comprises one or more virtual sound sources which output the audio content.

7. A data processing apparatus according to claim 6, wherein: the one or more output devices of the physical environment comprise one or more speaker devices; andthe circuitry is configured to: capture the audio content output of the one or more virtual sound source(s) using one or more virtual sound sensor; and control the one or more speaker devices to output the audio content captured by the one or more virtual sound sensors.

8. A data processing apparatus according to claim 6, wherein, depending on a characteristic of each virtual sound source, the virtual sound source outputs the audio content a single time, multiple times or on a continuous loop.

9. A data processing apparatus according to claim 1, wherein: the received information derived from the digital content comprises a scene schema, the scene schema defining the virtual environment; and the circuitry is configured to generate the scene schema by: providing words derived from the digital content to a large language model (LLM); providing sequential instructions to the LLM to build the scene schema based on the words.

10. A data processing apparatus according to claim 9, wherein the scene schema defines one or more of: a position of a virtual light source in the virtual environment and information for determining a brightness and / or colour of the virtual light source; a position and form of one or more textured objects in the virtual environment and a texture image to be applied to a surface of each textured object as a texture; and a position of a virtual sound source in the virtual environment and audio content to be output by the virtual sound source.

11. A data processing apparatus according to claim 1, wherein the digital content comprises one or more of audio, images, video and text.

12. A data processing apparatus according to claim 11, wherein the audio comprises speech uttered by one or more users.

13. A data processing apparatus according to claim 1, wherein: the circuitry is configured to transmit information representing the virtual environment to a second data processing apparatus; and a position of virtual light and / or sound sources is adjustable depending on respective physical environments of the data processing apparatus and second data processing apparatus.

14. A data processing apparatus according to claim 1, wherein the one or more digital assets are generated by a generative artificial intelligence, AI, model and / or retrieved from a database.

15. A computer-implemented data processing method comprising: receiving information derived from digital content; generating a virtual environment using one or more digital assets corresponding to the received information; extracting information from the virtual environment; and outputting a control signal to one or more output devices of a physical environment to control the one or more output devices to generate an output corresponding to the information extracted from the virtual environment.

16. A program for controlling a computer to perform a method according to claim 15.

17. A computer-readable storage medium storing a program according to claim 16.

Citation Information

Patent Citations

  • Real world acoustic and lighting modeling for improved immersion in virtual reality and augmented reality environments

    US20140132628A1

  • Ambient Light Control and Calibration via Console

    US20170368459A1

  • Generating simulation scenarios and digital media using digital twins

    US20230334773A1

  • Methods and systems for unified rendering of light and sound content for a simulated 3D environment

    US20230401789A1