Apparatus and method for immersion support, head-mounted device and method for a head-mounted device, information processing system for immersion support and method for an information processing system

The apparatus and method enhance literary immersion by synchronizing sensory stimulation with literature content, addressing the disconnect between narrative and surroundings, thereby enriching the reading experience.

WO2026008749A1PCT designated stage Publication Date: 2026-01-08SONY GROUP CORP +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/068917
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-03
Filing Date
2025-07-03
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing reading experiences in literature fail to provide immersive engagement due to the disconnect between the reader's physical surroundings and the narrative environment, hindering full immersion and enjoyment.

Method used

An apparatus and method that utilize processing circuitry to determine output content and format, integrating sensory stimulation through visual, audio, thermal, and vibrational outputs synchronized with literature content, enhancing immersion via augmented or virtual reality environments.

Benefits of technology

Seamlessly integrates textual narratives with immersive environments, providing a captivating and enriching literary experience by bridging the gap between the reader's surroundings and the narrative, increasing engagement and enjoyment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025068917_08012026_PF_FP_ABST
    Figure EP2025068917_08012026_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an apparatus for immersion support. The apparatus includes processing circuitry configured to receive input data representing literature content that a user is consuming. The processing circuitry is further configured to: determine, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content. The output content is determined for outputting it, via a user device, to the user in accordance with the output content format.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] APPARATUS AND METHOD FOR IMMERSION SUPPORT, HEAD-MOUNTED DEVICE AND METHOD FOR A HEAD-MOUNTED DEVICE, INFORMATION PROCESSING SYSTEM FOR IMMERSION SUPPORT AND METHOD FOR AN INFORMATION PROCESSING SYSTEM

[0002] Field

[0003] The present disclosure relates to an apparatus and a method for immersion support, a headmounted device and a method for a head-mounted device, an information processing system for immersion support and a method for an information processing system. The present disclosure may be applied in the field of augmented or virtual reality.

[0004] Background

[0005] Head-mounted devices are generally known. Such devices may be used for providing virtual reality (VR) or augmented reality (AR). A user of such devices may experience immersion when watching a movie or playing a video game and may thus, disregard the real world when being immersed into the VR / AR scene. Different technologies may be applied for generating an AR / VR experience and different devices may be used, such as a display, speakers, and the like.

[0006] Although ARA / R technology is generally known, it may be desirable to provide concepts to provide user immersion.

[0007] Summary

[0008] This demand is met by an apparatus for immersion support, a head-mounted device, an information processing system for immersion support, a method for immersion support, a method for a head-mounted device, and a method for an information processing system in accordance with the independent claims. Further embodiments are set forth in the dependent claims.

[0009] According to a first aspect, the present disclosure provides an apparatus for immersion support. The apparatus comprises processing circuitry configured to receive input data representing literature content that a user is consuming. The processing circuitry is further configured to determine, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content. The output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

[0010] According to a second aspect, the present disclosure provides a head-mounted device. The head-mounted device comprises processing circuitry configured to receive output content and an output content format. The output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content. The processing circuitry is further configured to control output of the output content in accordance with the output content format via at least one output device of the head-mounted device.

[0011] According to a third aspect, the present disclosure provides an information processing system for immersion support. The information processing system comprises an apparatus for immersion support according to the first aspect. The information processing system further comprises a user device configured to receive output content and an output content format from the apparatus for immersion support. The output content is for output to a user of the information processing system in accordance with the output content format, thereby supporting immersion of the user while consuming the literature content. The user device is further configured to control output of the output content in accordance with the output content format via at least one output device of the information processing system.

[0012] According to a fourth aspect, the present disclosure provides a method for immersion support. The method further comprises receiving input data representing literature content that a user is consuming. The method further comprises determining, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content. The output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

[0013] According to a fifth aspect, the present disclosure provides a method for a head-mounted device. The method comprises receiving output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content. The method further comprises controlling output of the output content in accordance with the output content format via at least one output device of the head-mounted device. According to a sixth aspect, the present disclosure provides a method for an information processing system. The method comprises receiving, by an apparatus for immersion support, input data representing literature content that a user is consuming. The method further comprises determining, by the apparatus for immersion support based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content. The output content is determined for outputting it, via a user device, to the user in accordance with the output content format. The method further comprises receiving, by the user device and from the apparatus for immersion support, output content and an output content format. The output content is for output to a user of the user device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content. The method further comprises controlling, by the user device, output of the output content in accordance with the output content format via at least one output device.

[0014] Brief description of the Figures

[0015] Some examples of apparatuses and / or methods will be described in the following by way of example only, and with reference to the accompanying figures, in which:

[0016] Fig. 1 depicts an apparatus for immersion support according to the present disclosure;

[0017] Fig. 2 depicts a head-mounted device according to the present disclosure;

[0018] Fig. 3 depicts an information processing system according to the present disclosure;

[0019] Fig. 4 depicts a method for immersion support according to the present disclosure;

[0020] Fig. 5 depicts a method for a head-mounted device according to the present disclosure;

[0021] Fig. 6 depicts a method for an information processing system according to the present disclosure; and

[0022] Fig. 7 depicts a method for supporting immersion according to the present disclosure.

[0023] Detailed Description Some examples are now described in more detail with reference to the enclosed figures. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of the features as well as equivalents and alternatives to the features. Furthermore, the terminology used herein to describe certain examples should not be restrictive of further possible examples.

[0024] Throughout the description of the figures same or similar reference numerals refer to same or similar elements and / or features, which may be identical or implemented in a modified form while providing the same or a similar function. The thickness of lines, layers and / or areas in the figures may also be exaggerated for clarification.

[0025] When two elements A and B are combined using an “or”, this is to be understood as disclosing all possible combinations, i.e. only A, only B as well as A and B, unless expressly defined otherwise in the individual case. As an alternative wording for the same combinations, "at least one of A and B" or "A and / or B" may be used. This applies equivalently to combinations of more than two elements.

[0026] If a singular form, such as “a”, “an” and “the” is used and the use of only a single element is not defined as mandatory either explicitly or implicitly, further examples may also use several elements to implement the same function. If a function is described below as implemented using multiple elements, further examples may implement the same function using a single element or a single processing entity. It is further understood that the terms "include", "including", "comprise" and / or "comprising", when used, describe the presence of the specified features, integers, steps, operations, processes, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components and / or a group thereof.

[0027] Fig. 1 schematically illustrates an apparatus 1 for immersion support according to the present disclosure.

[0028] The apparatus includes processing circuitry 2. For example, the processing circuitry 2 may be a single dedicated processor, a single shared processor, or a plurality of individual processors, some of which or all of which may be shared, a digital signal processor (DSP) hardware, an application specific integrated circuit (ASIC), a system-on-a-chip (SoC) a neuromor- phic processor or a field programmable gate array (FPGA). The processing circuitry 2 may optionally be coupled to, e.g., memory such as read only memory (ROM) for storing software, random access memory (RAM) and / or non-volatile memory. For example, the apparatus 1 may include memory configured to store instructions, which when executed by the processing circuitry 2, cause the processing circuitry 2 to perform the steps and methods described herein.

[0029] The processing circuitry 2 is configured to receive input data representing literature content that a user is consuming. The input data may be provided from various sources. For example, the input data may be received from a server, a computer, an AR / VR headset (or headmounted device, HMD), an electronic reader, or the like. The apparatus 1 may optionally include a receiver or a transceiver (not illustrated in Fig. 1) configured for (e.g., wireless or wired) reception of the input data from the source. However, it is to be noted that the present application is not limited to the aforementioned examples. The input data may have an appropriate data format for the apparatus 1 (or the processing circuitry 2) to process. Depending on the input data format (e.g., audio, video, etc.), the processing circuitry may be configured to transform the received data into an appropriate format to be processed.

[0030] Literature content according to the present disclosure may include any visual or audio / speech piece of text (or text data) to be consumed by a user (visual: when the user reads the text; audio: when the text is read to the user), e.g., for entertainment, information, work, or the like. If the literature content is visual, it may be presented to a user in digital form, e.g., on a display (of an e-reader, of an HMD, of a computer screen, a television, or the like). On the other hand, the user may hold a piece of text (e.g., a book, a magazine, a newspaper, or the like). If the literature content is based on audio or speech (or both), the literature content may be output via speakers (e.g., of an HMD, a sound system, a television, a computer, or the like). Moreover, text may be output on a display and at the same time, the text may be read to the user, in some embodiments. It should be noted that the literature content may also include or be based on pictures. For example, the literature content may include comics, mangas, storybooks, picture-books, and the like.

[0031] The apparatus 1 (or the processing circuitry 2) is further configured determine, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content. The output content is determined for outputting it, via a user device, to the user in accordance with the output content format. There may be various ways how, based on the input data, an output content and an output content format are determined. For example, a trained algorithm may be used which is configured to process the input data and generate or determine an appropriate output content format and an output content, which will be discussed below.

[0032] Immersion support may relate to a sensory stimulation (e.g., sound, images, smell, skin sensation, or the like) of a user and thus, supporting immersion may relate to or be based on a generation of such sensory stimulation. For example, if the user is reading a book and thus, consuming literature content, the user may already experience immersion, but there may be occasions in which the user gets distracted, e.g., by events in the real world. Thus, the sensory stimulation that is to be generated is supposed to enhance or support the immersion by recognizing the literature content and based on the literature content, find an output content format and an output content that fits to the literature content. For example, if the literature content is about the desert, the sensory stimulation may include outputting an image of a desert and possibly to output heat.

[0033] Output content may relate to data that is output in order to support immersion for the user, i.e., to generate the sensory stimulation. The output content is output in the output content format. The output content format may indicate what type of output device is to output the output content. For example, the output content format may include at least one of a visual output, an audio output, a thermal output, and a vibrational output. Thus, if the output content format is visual, a display may be used to output the output content. Further explications regarding the output content format are made further below.

[0034] The output content and the output content format are (supposed) to support immersion while the user is consuming the literature content. That means that the output content may be synchronized with the literature content or in other words, that the output content may fit to the literature content while the user is consuming the literature content. Hence, in some embodiments, the literature content it may be recognized, as will be discussed further below.

[0035] The output content is determined for outputting it via a user device. The user device may have various forms and may include, for example, a headset (e.g., an AR or VR head- set / HMD), a television, a laptop, a computer, a mobile phone, a (smart) monitor / screen, or the like. On the other hand, the user device may relate to processing circuitry (e.g., as discussed above) that is configured to control any of the above mentioned device to output the output content in the output content format. On the other hand, the user device may be provided externally, e.g., as a remote device (e.g., server) that is configured to control any of the device above to output the output content in the output content format. Thereby, processing power may be externalized and saved in the particular device. The user device may be provided in or may be equal to an information processing system as discussed herein.

[0036] According to the present disclosure, in contrast to “traditional reading”, a reader’s (or user’s) immersion into the content is supported and the reader does not need to imagine the content completely by themselves and thus, the reader does not need to concentrate that much on the content. Thereby, immersion is made easier and a user’s reading duration can be increased.

[0037] The present disclosure is based on the finding that traditional reading experiences may fail to provide an immersive engagement as in other media forms, such as movies or video games. Readers may struggle to vividly imagine scenes described in the text they are reading, especially when a real-world environment contradicts the settings depicted in the narrative / litera- ture content. For instance, when a reader is sitting in a bustling coffee shop and reading a book that intends to transport them into the vast expanses of space, the jarring contrast between the ambient noise of the coffee shop and the serene silence of the cosmos may significantly diminish the reading experience. This discrepancy between the reader's physical surroundings and the imagined world of the text may create a disconnect, hindering their ability to fully immerse themselves into the story. Without visual or sensory cues to complement the textual descriptions, readers may find it challenging to evoke the intended emotions and atmosphere portrayed by the author / narrator. As a result, the reading experience may become fragmented, limiting the reader's engagement and enjoyment of the material. Thus, it has been recognized that there may be a need for a solution that can bridge this gap and seamlessly integrate textual narratives with immersive environments, providing readers with a captivating and enriching literary experience.

[0038] This need is fulfilled at least by the aspects described above. In the following, various examples will be described in greater detail to highlight the aspects of the present disclosure.

[0039] For example, the apparatus 1 for immersion support (or its processing circuitry 2) is further configured to control the user device to output the output content in accordance with the output content format, as discussed herein. It should be noted that several implementations are possible how the output content is output and what the user device entails, as discussed above. For example, the apparatus 1 (or its processing circuitry 2) is further configured to: determine the input data based on the literature content, as discussed herein. For example, trained models may be used for recognizing language / literature content, e.g., based on acquired images, based on text data (e.g., if the literature content is stored in digital form), based on audio data, and the like. Moreover, it may be determined which text passage the user reads, e.g., based on a recognition of the literature content, e.g., of a page, or by text recognition. For example, a gaze or a line of sight of the user may be determined, e.g., based on eye tracking circuitry, in order to determine the literature content. If the literature content is displayed on a display, metadata may be used that indicate the literature content that is displayed. A similar approach may be taken, if the literature content is audio content.

[0040] For example, the apparatus 1 (or its processing circuitry 2) is further configured to receive literature content data from the user device. The apparatus 1 (or its processing circuitry 2; and determine the literature content based on the literature content information received from the user device.

[0041] In such examples, the user device (or corresponding processing circuitry of the user device) may produce / generate the input data and receive the output content and the output content format. Hence, the user device may pre-process the literature content into the input data for the apparatus 1 (or the processing circuitry 2) to determine the output data (output content and output content format). For example, as already discussed above, the user device may be included in an HMD (which may include a sensor, such as a camera) and may thus process images taken from the literature content (in accordance with the line of sight of the user) and transform the images into input data for the apparatus 1 according to the present disclosure. The user device may then receive the output content and the output content format and translate this into the output content in the output content format, such as images to be shown on a display.

[0042] For example, the input data includes at least one latent vector representing the literature content. In such examples, the apparatus 1 (or its processing circuitry 2) may be further configured to determine at least one of the output content and the output content format based on the at least one latent vector. For example, if an encoder is used to generate the latent vector based on the input data, a corresponding decoder may translate the latent vector into an output that is close to the input data, as will be discussed below.

[0043] However, the present disclosure is not limited to the use of latent vectors. Depending on the machine-learning model that is used, other types of input data may be used, such as tokens.

[0044] A machine-learning (or machine-learned) model may refer to a data structure and / or set of rules representing a statistical model that the processing circuitry 2 uses to determine the output content and the output content format. The data structure and / or set of rules represents learned knowledge (e.g. based on training performed by a machine-learning algorithm as described below). In machine-learning, instead of a rule-based transformation of data, a transformation of data may be used, that is inferred from an analysis of training data.

[0045] The machine-learning model may be trained based on a machine-learning algorithm. The term "machine-learning algorithm" denotes a set of instructions that are used to create, train or use a machine-learning model. For the machine-learning model to determine the output content and the output content format, the machine-learning model may be trained using training data, as commonly known. By training the machine-learning model with a large set of training data and associated training content information, the machine-learning model "learns" to determine the output content and the output content format. By training the machine-learning model using training information on output content and output content formats (based on input data), the machine-learning model "learns" a transformation between training input data and appropriate output content and an output content format.

[0046] The machine-learning model may be trained using training input data (e.g. literature content that is associated with an output content and an output content format). For example, the machine-learning model may be trained using a training method called "supervised learning". In supervised learning, the machine-learning model is trained using a plurality of training samples, wherein each sample may include a plurality of input data values, and a plurality of desired output values, i.e., each training sample is associated with a desired output value. By specifying both training samples and desired output values, the machine-learning model "learns" which output value to provide based on an input sample that is similar to the samples provided during the training. For example, a training sample may include input data that is based on literature content and desired output content in an output content format. Also, the transformation of literature content into appropriate input data may be trained in a similar fashion r with a different approach.

[0047] Apart from supervised learning, semi-supervised learning may be used. In semi-supervised learning, some of the training samples lack a corresponding desired output value. Supervised learning may be based on a supervised learning algorithm (e.g. a classification algorithm or a similarity learning algorithm). Classification algorithms may be used as the desired outputs of the trained machine-learning model are restricted to a limited set of values (categorical variables), i.e. , the input is classified to one of the limited set of values. Similarity learning algorithms are similar to classification algorithms but are based on learning from examples using a similarity function that measures how similar or two related objects are.

[0048] Apart from supervised or semi-supervised learning, unsupervised learning may be used to train the machine-learning model. In unsupervised learning, (only) input data are supplied and an unsupervised learning algorithm is used to find structure in the input data (e.g. by grouping or clustering the input data, finding commonalities in the data). Clustering is the assignment of input data including a plurality of input values into subsets (clusters) so that input values within the same cluster are similar according to one or more (pre-defined) similarity criteria, while being dissimilar to input values that are included in other clusters.

[0049] Reinforcement learning is a third group of machine-learning algorithms. In other words, reinforcement learning may be used to train the machine-learning model. In reinforcement learning, one or more software actors (called "software agents") are trained to take actions in an environment. Based on the taken actions, a reward is calculated. Reinforcement learning is based on training the one or more software agents to choose the actions such that the cumulative reward is increased, leading to software agents that become better at the task they are given (as evidenced by increasing rewards).

[0050] Furthermore, additional techniques may be applied to some of the machine-learning algorithms. For example, feature learning may be used. In other words, the machine-learning model may at least partially be trained using feature learning, and / or the machine-learning algorithm may include a feature learning component. Feature learning algorithms, which may be called representation learning algorithms, may preserve the information in their input but also transform it in a way that makes it useful, often as a pre-processing step before performing classification or predictions. Feature learning may be based on principal components analysis or cluster analysis, for example. For example, the machine-learning model may be an Artificial Neural Network (ANN). AN Ns are systems that are inspired by biological neural networks, such as can be found in a retina or a brain. ANNs include a plurality of interconnected nodes and a plurality of connections, so-called edges, between the nodes. There are usually three types of nodes, input nodes that receive input values (e.g., based on literature content), hidden nodes that are (only) connected to other nodes, and output nodes that provide output values (e.g., output content and output content format). Each node may represent an artificial neuron. Each edge may transmit information from one node to another. The output of a node may be defined as a (non-linear) function of its inputs (e.g. of the sum of its inputs). The inputs of a node may be used in the function based on a "weight" of the edge or of the node that provides the input. The weight of nodes and / or of edges may be adjusted in the learning process. In other words, the training of an ANN may include adjusting the weights of the nodes and / or edges of the ANN, i.e. , to achieve a desired output for a given input.

[0051] Alternatively, the machine-learning model may include a different structure and, e.g., be a support vector machine, a random forest model or a gradient boosting model. Alternatively, the machine-learning model may be based on a genetic algorithm, which is a search algorithm and heuristic technique that mimics the process of natural selection.

[0052] In examples, the machine-learning model may be a combination of any of above examples.

[0053] In examples, the apparatus 1 (or its processing circuitry 2) is further configured to generate further output data to control the user device to embed the output content in accordance with the output content format into a scene generated by the user device.

[0054] The further output data may have the same, similar or different data structure as the output data mentioned above. The further output data may be provided as meta-data such that compatibility with other devices may not be compromised.

[0055] Embedding may refer to a process or method in which the output content is aligned with other content (such as the literature content), such that the literature content may be consumed without obstructing or without covering at least a part of the literature content. For example, if the output content is visual content, the output content should not hinder the user in consuming the literature content. For example, if the output content is audio content, a volume or a placement of certain sounds (e.g., in 3D sound) should not distract the user from consuming the literature content.

[0056] For example, if the user is reading a book and at the same time wear an HMD according to the present disclosure, it may be desirable not to obstruct the user’s field of view by the content and the scene. If the scene is generated according to the literature content (e.g., a desert scene as a narrative framework), and further describes camels, these camels might be embedded into the scene. However, the display may be segmented out in accordance with the line of sight of the user, such that the book can still be read. Similarly, if the literature content is displayed on the display, the scene and the embedded output content should be positioned such that the user can still read the content.

[0057] In examples, the scene (as well as the (embedded) output content) is an augmented reality scene or a virtual reality scene, as discussed herein.

[0058] As indicated above, in some examples, the output content format includes at least one of a visual output, an audio output, a thermal output, and a vibrational output.

[0059] Visual output may include at least one of a still image and a moving image (such as a video). Also, a mixture may be envisaged, such as a still image on one part of the display and a moving image on another part of the display. Audio output may be based on sound and may further include a volume of the output sound. A thermal output may be based on a temperature that is generated by the HMD. It may also be envisaged that heat is output only at a predetermined part of the HMD (such as, on the right, because the scene describes fire on the right side of a character). A vibrational output may include a certain vibration pattern and intensity in accordance with the literature content. It should be noted that the present disclosure is not limited to any output content format, and it may depend on the model that it used for determining the output content format.

[0060] Hence, in some examples, at least one of the output content and the output content format is determined based on a trained machine-learning model (e.g., such as a machine-learning model discussed above).

[0061] In some examples the trained machine-learning model is a transformer based multimodality model. The transformer based multimodality model (also referred to as meta-transformer) may be trained based on unified multimodal learning. A meta-transformer may utilize the same backbone to encode natural language (e.g., text), image (e.g., based on 2D vision or artificially generated images), point cloud (e.g., based on 3D vision), audio (e.g., sound, speech, music), video (e.g., spatio-temporal data), infrared (e.g., thermal information), hyper- spectral (e.g., remote sensing data), x-ray (e.g., based on medical applications), time-series (e.g., for forecasting), data table (e.g., for tabular analysis), inertial measurement unit (IMU; e.g., based on inertia measurements), graph data (e.g., molecular structure), and the like. In other words, in multimodal learning, models may be used (or built) that can process and relate information from multiple modalities. An exemplary multimodal learning approach can be found in the publication “Meta-Transformer: A Unified Framework for Multimodal Learning” by Zhan et al., that can be retrieved under arxiv.org / pdf / 2307.10802. The aforementioned multimodal learning approach may be used for training the multimodality model. However, the present disclosure is not limited thereto. Other training approaches (e.g., one of the above- mentioned) may be used as well.

[0062] However, the present disclosure is not limited to the use of such a meta-transformer, since any model which is configured to use input data according to the present disclosure and is able to determine different output content formats (and corresponding output content) may be used. According to the present disclosure, a seamless integration of sensory stimulation into the reading experience may be provided based on such a model. A (transformative) model according to the present disclosure may use processing input of various formats, such as words, sounds, images, and the like, and interpret a context of the input. Thereby, the model may be able to translate / transform the input, such as a narrative, into immersive scenes, such as audiovisual scenes within AR / VR environments. Additionally, the integration of sensory stimulation may further enrich the user reading experience, transcending traditional boundaries to engage readers on a multisensory level.

[0063] A model according to the present disclosure may use, for example, text and images as an input (as provided by a book) and generate other closely related outputs (e.g., audio or other sensory information). To achieve this, the model may process an input by an encoder that creates, for example, latent vectors representing the input. These latent vectors may then be processed by a decoder to create other forms of output. The output may then be virtually embedded into the AR / VR scene for instance as 3D objects, audio output or other correlating feedback.

[0064] The model may enable the apparatus 1 (or a user device, such as an HMD or an information processing system according to the present disclosure, as discussed below) to dynamically generate vivid scenes that are not only visually and auditorily depicting the narrative but also stimulate other senses such as smell, touch, and temperature. For instance, when encountering a description of a serene beach, the system can evoke the scent of the ocean, the warmth of the sun, and the sound of waves crashing against the shore. Ambient lighting within the AR / VR environment could mimic the soft glow of the sun at sunset, enhancing the visual immersion. Additionally, the integration of audio outputs simulating the soothing sounds of sea birds and the gentle rustle of palm leaves complements the sensory experience. This comprehensive approach to immersion ensures that readers are fully enveloped in the narrative world, engaging their senses of sight, sound, and smell to create a deeply memorable experience.

[0065] By combining (multimodal) inputs and outputs with sensory stimulation, the present disclosure may offer a reading (or listening) experience that captivates and enthralls readers in ways previously unattainable. This integration of technologies may provide immersive storytelling within ARA / R environments, thereby supporting literary engagement providing interactive reading experiences.

[0066] Fig. 2 depicts a head-mounted device (HMD) 11 according to the present disclosure. The HMD 11 is arranged such that it can be mounted on a head of a user, such that a display of the HMD 11 is in a field of view of the user. The HMD 11 may be configured as smart glasses, for example. In some embodiments, the HMD 11 is configured as a see-through HMD, i.e., including a display through which the user can at least partially see. On the other hand, the HMD 11 is configured, in some embodiments, as a non-see-through HMD, which completely obstructs the field-of-view. Depending on the specific configuration, the HMD 11 may be capable of providing virtual reality content (in the case of non-see-through or partially see- through) or augmented reality content (in the case of (partially) see-through or in the case of non-see-through when cameras are provided which obtain images of the real world).

[0067] The HMD 11 includes a display and camera system 12 including a display and at least one camera, as is commonly known. The camera and corresponding circuitry of the system 12 is configured to obtain images of the real world and to capture information regarding literature content that a user of the HMD 11 is consuming. However, the present disclosure is not limited in that regard. In some embodiments, the literature content is displayed on the display of the HMD 11 such that there is no need to capture it with a camera. Also, in some embodiments, literature content relates to an audio content (e.g., an audio book, a podcast, audio drama, or the like) such that it may be output with speakers of the HMD 11 (not depicted), as discussed above.

[0068] Generally, the HMD 11 includes circuitry for obtaining literature content in a certain content format, such as visual (e.g., an image of a book), audio (e.g., an audio book), as text (e.g., an e-book), and the like, wherein the system 12 is depicted for illustrational purposes and should not limit the present disclosure.

[0069] The HMD 11 further includes an apparatus for immersion support 13, as discussed herein, e.g., under reference of Fig. 1.

[0070] The HMD 11 constitutes a user device according to the present disclosure and thus, includes further processing circuitry 14 which is configured to control an output device of the HMD 11 , such as the display 12, to output the output content. For selecting an output device, the output content format may be used. For example, if the output content format is a visual output, such as an image or a video, the circuitry 14 may be configured to select the display to output the output content. Such information (i.e. , a matching between the content format and a corresponding device) may be stored in a database or may be retrieved from a trained algorithm which is trained to do such matching between output content and output content format. In some instances, an output content format may indicate more than one output device, such as visual and audio. For example, if the literature content describes a classical concert, the audio output content may include a classical piece of music to be played via headphones of the HMD 11 , and the visual output may include a concert hall being depicted in a part of the display of the HMD 11. Hence, the present disclosure is not limited to a single output content format being determined. It may depend on the literature content and ultimately on the input data, which and how many output content formats are determined.

[0071] In the following, various examples will be described in greater detail for the HMD 11 .

[0072] As indicated above, the head-mounted device 11 (or its processing circuitry 14) may be further configured to determine an eye gaze of the user, as discussed above, for example based on eye-tracking technology. The eye-tracking technology may be based on infrared light, time-of-flight technology, reflectivity measurements, or the like. The HMD 11 may be further configured to determine the literature content based on the eye gaze (or line of sight), as discussed herein. Thereby, the output content may be brought into accordance with the literature content that the user is consuming.

[0073] As discussed above, there may be several ways how the user consumes the literature content. One way includes the user holding the literature content, e.g., when reading a book, a magazine, a newspaper, or the like. For such cases, the HMD 11 includes at least one sensor configured to detect the literature content that the user is holding. The sensor may include a camera (e.g., RGB, infrared, multispectral, time-of-flight, or the like)

[0074] In such examples, the HMD 11 (or its processing circuitry 14) is further configured to acquire literature content data based on sensor data indicating the detected literature content. For example, a camera may take an image of the literature content that the user is holding, and the processing circuitry may apply a text-recognition method / technology in order to identify the literature content information.

[0075] There may be various sensors for carrying out such a detection. As already discussed above, a camera may be used for that purpose. Additionally or alternatively, if the user reads the literature content on an electronic device, such as a smartphone or an electronic reader, the HMD 11 (or the processing circuitry 14) may be configured to communicate with the electronic device and thereby obtain the literature content. Hence, the corresponding sensor may include a communication interface configured to communicate via a communication protocol, such as Wi-Fi, a standardized 3GPP communication protocol, Bluetooth, or the like.

[0076] Literature content data may refer to text that is recognized in such a way, without limiting the present disclosure in that regard. Based on the literature content data, the input data may be obtained. For example, the literature content data already corresponds to the input data, or alternatively, the literature content data may be fed into a machine-learning model that is configured to transform them into the input data, such as one or more of the machine-learning models described above.

[0077] As indicated above, the head-mounted device 11 may further include a display configured to display the literature content, as discussed herein.

[0078] As also indicated above, the head-mounted device 11 (or its processing circuitry) may be further configured to segment a display of the head-mounted device, if the output content format includes a visual output. Thereby, the visual output is displayed on a portion of the display that is different from the portion of the display the user uses to consume the literature content, as discussed herein. In other words, the user’s sight is not obstructed by the output content and the user can still consume the literature content.

[0079] A user device according to the present disclosure (such as the HMD discussed under reference with Fig. 2 or at least a part of the information processing system as discussed under reference with Fig. 3) may have at least one output device which may include a force feedback device for sensory feedback. Such a device may exemplarily be configured to simulate raindrops by applying a slight force feedback on the user’s head together with a specific raindrop sound. For temperature sensation, an HMD or information processing system according to the present disclosure may additionally or alternatively include an integrated resistance heating device and for olfactory sensation it may additionally or alternatively include an integrated digital scent device. These two devices may enable feeling a presence of a fire or fireplace by applying a warm feeling and the smell of burned wood. Combined with a 2D or 3D representation of a fireplace in the AR / VR scene and hearing the crackling sounds of a typical fire it may enable an unprecedented virtual experience.

[0080] Fig. 3 depicts an information processing system for immersion support 20. The information processing system 20 includes an apparatus for immersion support 21 , as discussed herein, such as under reference of Fig. 1. In this embodiment, the apparatus for immersion support

[0081] 21 is included in an electronic reader and thus, the literature content is available in digital form. Hence, the literature content is electronically available for the apparatus 21 .

[0082] The apparatus 21 includes processing circuitry configured to generate (and thereby, receive) input data based on the electronic literature content (e.g., based on content of a page that is displayed) and to generate output content and an output content format, as discussed herein.

[0083] The output content and the output content format are transmitted to a user device - in this example, a smart TV 22 - that includes corresponding output devices (such as speakers and a display) to support the immersion of the user, as discussed herein. Hence, the user device

[0084] 22 is configured to receive the output content and the output content format from the apparatus 21. The output is output via at least one output device of the user device 22, as discussed herein. It should be noted that the information processing system is not limited to such a case. It may also be envisaged to provide an information processing system according to the present disclosure in a type of library or user immersion room, such that, when consuming the literature, the room may have corresponding output means (such as a plurality of displays, speakers, heating devices, odor generators, and so on). Thereby, the whole room may contribute to a (full) immersion of the user into the literature content. Moreover, it should be noted that also an HMD (or any other type of headset), such as the HMD 11 shall be understood as an information processing system. Also, a laptop, any type of headset, or the like may constitute or be part of an information processing system according to the present disclosure.

[0085] Fig. 4 depicts a block diagram of a method for immersion support 30 according to the present disclosure.

[0086] The method 30 includes, at 31 , receiving input data that represent literature content that a user is consuming, as discussed herein.

[0087] The method 30 further includes, at 32, determining, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format, as discussed herein.

[0088] The method 30 may be carried out by an apparatus for immersion support, as discussed herein, such as the apparatus 1 , 13 or 21.

[0089] Fig. 5 depicts a block diagram of a method for a head-mounted device 40.

[0090] The method 40 includes, at 41 , receiving output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content, as discussed herein.

[0091] The method 40 further includes, at 42, controlling output of the output content in accordance with the output content format via at least one output device of the head-mounted device, as discussed herein. The method 40 may be carried out by an HMD according to the present disclosure, such as the HMD 11 or the HMD 63.

[0092] Fig. 6 depicts a block diagram of a method for an information processing system 50.

[0093] The method 50 includes, at 51 receiving, by an apparatus for immersion support, input data representing literature content that a user is consuming, as discussed herein.

[0094] The method 50 further includes, at 52, determining, by the apparatus for immersion support based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format, as discussed herein.

[0095] The method 50 further includes, at 53, receiving, by the user device and from the apparatus for immersion support, output content and an output content format, wherein the output content is for output to a user of the user device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content, as discussed herein.

[0096] The method 50 further includes, at 54, controlling, by the user device, output of the output content in accordance with the output content format via at least one output device.

[0097] The method 50 may be carried out by an information processing system according to the present disclosure, such as the information processing system 20, or in an HMD, such as the HMD 11 or the HMD 63.

[0098] Fig. 7 depicts a method 70 for supporting immersion according to the present disclosure which is exemplarily carried out for / in an HMD 63, but the present disclosure is not limited in that regard since it may generally be carried out for immersion support regardless of the device that is used and, as discussed above, several functions may be distributed among different devices.

[0099] In the method 60, a user 61 holds and reads a book 62, representing literature content according to the present disclosure. The user 61 wears the HMD 63, which includes an apparatus for immersion support according to the present disclosure. The HMD 63 further includes sensors, a display, speakers, and the like, as generally known to the person skilled in the art. Based on the literature content, input data is generated for a multimodality model 64, as described herein. The multimodality model 64 is configured to determined, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, as discussed herein.

[0100] One aspect of the output content is depicted with reference numeral 65 including visual output (represented as trees) and audio output (represented as soundwaves). Moreover, the multimodality model 64 determines sensory stimulation 66 (represented as wind), wherein the HMD 63 has corresponding devices to output such sensory stimulation (e.g., a fan to generate wind, a heating to generate heat, and so on) to the user 61.

[0101] Thereby, immersion into the literature content is supported for the user 61 .

[0102] Analogously to what is described above under reference of Figs. 1 to 3, the methods discussed herein may achieve the same or similar effects as the apparatus for immersion support, the HMD, and the information processing system according to the present disclosure.

[0103] More details and aspects of the methods described herein are explained in connection with the proposed technique or one or more examples described above (e.g., Figs. 1 to 3). The methods may include one or more additional optional features corresponding to one or more aspects of the proposed technique or one or more examples described above.

[0104] The following examples pertain to further embodiments of the present disclosure:

[0105] (1) An apparatus for immersion support, the apparatus including processing circuitry configured to: receive input data representing literature content that a user is consuming; and determine, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

[0106] (2) The apparatus of (1), further configured to: control the user device to output the output content in accordance with the output content format.

[0107] (3) The apparatus of (1) or (2), further configured to: determine the input data based on the literature content.

[0108] (4) The apparatus of (3), further configured to: receive literature content information from the user device; and determine the literature content based on the literature content information received from the user device.

[0109] (5) The apparatus of anyone of (1) to (4), wherein the input data includes at least one latent vector representing the literature content the apparatus being further configured to: determine at least one of the output content and the output content format based on the at least one latent vector.

[0110] (6) The apparatus of anyone of (1) to (5), further configured to: generate further output data to control the user device to embed the output content in accordance with the output content format into a scene generated by the user device.

[0111] (7) The apparatus of (6), wherein the scene is an augmented reality scene or a virtual reality scene.

[0112] (8) The apparatus of anyone of (1) to (7), wherein the output content format includes at least one of a visual output, an audio output, a thermal output, and a vibrational output.

[0113] (9) The apparatus of anyone of (1) to (8), wherein at least one of the output content and the output content format is determined based on a trained machine-learning model.

[0114] (10) The apparatus of anyone of (1) to (9), wherein the trained machine-learning model is a transformer based multimodality model.

[0115] (11) A head-mounted device including processing circuitry configured to: receive output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and control output of the output content in accordance with the output content format via at least one output device of the head-mounted device.

[0116] (12) The head-mounted device of (11), further including an apparatus for immersion support according to anyone of (1) to (10).

[0117] (13) The head-mounted device of (11) or (12) further configured to: determine an eye gaze of the user; and determining the literature content based on the eye gaze.

[0118] (14) The head-mounted device of anyone of (11) to (13), wherein the user holds the literature content, and wherein the head-mounted device further includes at least one sensor configured to: detect the literature content that the user is holding; wherein the head-mounted device is further configured to: acquire literature content information based on sensor data indicating the detected literature content.

[0119] (15) The head-mounted device of anyone of (11) to (14), further including a display configured to display the literature content.

[0120] (16) The head-mounted device of anyone of (11) to (15), further configured to: if the output content format includes a visual output, segment a display of the headmounted device, such that the visual output is displayed on a portion of the display that is different from the portion of the display the user uses to consume the literature content.

[0121] (17) An information processing system for immersion support, the information processing system including: an apparatus for immersion support according to anyone of (1) to (10); and a user device configured to: receive output content and an output content format from the apparatus for immersion support, wherein the output content is for output to a user of the information processing system in accordance with the output content format, thereby supporting immersion of the user while consuming the literature content; and control output of the output content in accordance with the output content format via at least one output device of the information processing system.

[0122] (18) A method for immersion support, the method including: receiving input data representing literature content that a user is consuming; and determining, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

[0123] (19) A method for a head-mounted device, the method including: receiving output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and controlling output of the output content in accordance with the output content format via at least one output device of the head-mounted device.

[0124] (20) A method for an information processing system, the method including: receiving, by an apparatus for immersion support, input data representing literature content that a user is consuming; determining, by the apparatus for immersion support based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for out- putting it, via a user device, to the user in accordance with the output content format; receiving, by the user device and from the apparatus for immersion support, output content and an output content format, wherein the output content is for output to a user of the user device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and controlling, by the user device, output of the output content in accordance with the output content format via at least one output device. The aspects and features described in relation to a particular one of the previous examples may also be combined with one or more of the further examples to replace an identical or similar feature of that further example or to additionally introduce the features into the further example.

[0125] Examples may further be or relate to a (computer) program including a program code to execute one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Thus, steps, operations or processes of different ones of the methods described above may also be executed by programmed computers, processors or other programmable hardware components. Examples may also cover program storage devices, such as digital data storage media, which are machine-, processor- or computer-readable and encode and / or contain machine-executable, processorexecutable or computer-executable programs and instructions. Program storage devices may include or be digital storage devices, magnetic storage media such as magnetic disks and magnetic tapes, hard disk drives, or optically readable digital data storage media, for example. Other examples may also include computers, processors, control units, (field) programmable logic arrays ((F)PLAs), (field) programmable gate arrays ((F)PGAs), graphics processor units (GPU), application-specific integrated circuits (ASICs), integrated circuits (ICs) or system-on-a-chip (SoCs) systems programmed to execute the steps of the methods described above.

[0126] It is further understood that the disclosure of several steps, processes, operations or functions disclosed in the description or claims shall not be construed to imply that these operations are necessarily dependent on the order described, unless explicitly stated in the individual case or necessary for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a certain order. Furthermore, in further examples, a single step, function, process or operation may include and / or be broken up into several sub-steps, -functions, -processes or -operations.

[0127] If some aspects have been described in relation to a device or system, these aspects should also be understood as a description of the corresponding method. For example, a block, device or functional aspect of the device or system may correspond to a feature, such as a method step, of the corresponding method. Accordingly, aspects described in relation to a method shall also be understood as a description of a corresponding block, a corresponding element, a property or a functional feature of a corresponding device or a corresponding system.

[0128] The following claims are hereby incorporated in the detailed description, wherein each claim may stand on its own as a separate example. It should also be noted that although in the claims a dependent claim refers to a particular combination with one or more other claims, other examples may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are hereby explicitly proposed, unless it is stated in the individual case that a particular combination is not intended. Furthermore, features of a claim should also be included for any other independent claim, even if that claim is not directly defined as dependent on that other independent claim.

Claims

ClaimsWhat is claimed is:

1. An apparatus for immersion support, the apparatus comprising processing circuitry configured to: receive input data representing literature content that a user is consuming; and determine, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

2. The apparatus of claim 1 , further configured to: control the user device to output the output content in accordance with the output content format.

3. The apparatus of claim 1 , further configured to: determine the input data based on the literature content.

4. The apparatus of claim 3, further configured to: receive literature content data from the user device; and determine the literature content based on the literature content data received from the user device.

5. The apparatus of claim 1 , wherein the input data includes at least one latent vector representing the literature content the apparatus being further configured to: determine at least one of the output content and the output content format based on the at least one latent vector.

6. The apparatus of claim 1 , further configured to: generate further output data to control the user device to embed the output content in accordance with the output content format into a scene generated by the user device.

7. The apparatus of claim 6, wherein the scene is an augmented reality scene or a virtual reality scene.

8. The apparatus of claim 1 , wherein the output content format includes at least one of a visual output, an audio output, a thermal output, and a vibrational output.

9. The apparatus of claim 1 , wherein at least one of the output content and the output content format is determined based on a trained machine-learning model.

10. The apparatus of claim 1 , wherein the trained machine-learning model is a transformer based multimodality model.

11. A head-mounted device comprising processing circuitry configured to: receive output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and control output of the output content in accordance with the output content format via at least one output device of the head-mounted device.

12. The head-mounted device of claim 11, further comprising an apparatus for immersion support according to claim 1 .

13. The head-mounted device of claim 11 , further configured to: determine an eye gaze of the user; and determining the literature content based on the eye gaze.

14. The head-mounted device of claim 11 , wherein the user holds the literature content, and wherein the head-mounted device further comprises at least one sensor configured to: detect the literature content that the user is holding; wherein the head-mounted device is further configured to: acquire literature content data based on sensor data indicating the detected literature content.

15. The head-mounted device of claim 11 , further comprising a display configured to display the literature content.

16. The head-mounted device of claim 11 , further configured to: if the output content format includes a visual output, segment a display of the headmounted device, such that the visual output is displayed on a portion of the display that is different from the portion of the display the user uses to consume the literature content.

17. An information processing system for immersion support, the information processing system comprising: an apparatus for immersion support according to claim 1 ; and a user device configured to: receive output content and an output content format from the apparatus for immersion support, wherein the output content is for output to a user of the information processing system in accordance with the output content format, thereby supporting immersion of the user while consuming the literature content; and control output of the output content in accordance with the output content format via at least one output device of the information processing system.

18. A method for immersion support, the method comprising: receiving input data representing literature content that a user is consuming; and determining, based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for outputting it, via a user device, to the user in accordance with the output content format.

19. A method for a head-mounted device, the method comprising: receiving output content and an output content format, wherein the output content is for output to a user of the head-mounted device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and controlling output of the output content in accordance with the output content format via at least one output device of the head-mounted device.

20. A method for an information processing system, the method comprising:receiving, by an apparatus for immersion support, input data representing literature content that a user is consuming; determining, by the apparatus for immersion support based on the input data, an output content and an output content format for the output content to support immersion of the user while consuming the literature content, wherein the output content is determined for out- putting it, via a user device, to the user in accordance with the output content format; receiving, by the user device and from the apparatus for immersion support, output content and an output content format, wherein the output content is for output to a user of the user device in accordance with the output content format, thereby supporting immersion of the user while consuming literature content; and controlling, by the user device, output of the output content in accordance with the output content format via at least one output device.

Citation Information

Patent Citations

  • Entertaining device for Reading and the driving method thereof

    KR1020180130422A

  • Animated document using an integrated projector

    US20150070264A1

  • System and method for augmented or virtual reality entertainment experience

    US20150302651A1

  • System and Methods for Enhancing a Printed Graphic Novel, Book, and e-Book

    US20200285665A1

  • Immersive reading experience using eye tracking

    WO2006100645A2