Calibration system in video space

The in-spatial video calibration system addresses XR calibration challenges by using a viewing device and recognition technology to synchronize user location and time information, ensuring an immersive and shared experience across multiple users.

WO2026083555A1PCT designated stage Publication Date: 2026-04-23HASHILUS INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HASHILUS INC
Filing Date
2024-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing cross-reality (XR) technologies face challenges in accurately calibrating positional and temporal discrepancies among multiple users, leading to incorrect rendering and inability to provide an immersive shared experience.

Method used

An in-spatial video calibration system that includes a viewing device and a current information recognition device, utilizing controller holders and docks for calibration processing, along with input means to detect user hand placement and external information, to synchronize location and time information across users.

Benefits of technology

Enables users to experience an immersive and highly realistic XR space by accurately calibrating positional and temporal information, allowing multiple users to share an emotional experience synchronously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024037043_23042026_PF_FP_ABST
    Figure JP2024037043_23042026_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To provide a calibration system in a video space for performing calibration processing related to mutual positional relationships and video timing so that no positional or temporal deviation occurs among users viewing video related to XR via viewing devices. [Solution] This calibration system in a video space for calibrating position information of users in an XR space, and position information and time information of which content is displayed, comprises a viewing device that is worn by a user to be able to view the video, and a current information recognition device that recognizes current defined information of one or a plurality of users each wearing a viewing device, wherein the viewing device is provided with: a display means for displaying video on the basis of position information; a calibration means for calibrating position information and time information in a space configured by video of each user on the basis of information recognized by the current information recognition device; and an information holding means for holding, inter alia, calibrated position information of one or a plurality of users in the space configured by video, position information and time information of which content is displayed, and video.
Need to check novelty before this filing date? Find Prior Art

Description

In-spatial calibration system for video

[0001] The present invention relates to a system for providing a video space viewable by users through XR (Cross Reality), and more particularly to an in-video space calibration system that performs calibration processing on positional relationships and video timing so that positional and temporal discrepancies do not occur between one or more users viewing XR video via a viewing device, and performs the calibration processing in a simple and low-cost configuration, functionally, quickly, and in a manner that matches the user's experience, thereby enabling users to easily and intuitively experience an immersive and highly real-world XR space, and enabling multiple users to share an emotional experience by synchronously participating in the same XR space.

[0002] Over time, numerous cross-reality (XR) technologies have been developed and are being utilized in various situations, enabling users to gain various simulated experiences by viewing videos. Cross-reality technologies are being used in all fields, and many XR technologies have been developed and are in use that allow users to obtain all kinds of information from superimposed images while viewing the real world, enjoy such images, and experience being in a virtual space while remaining indoors.

[0003] Extended Reality (XR) is a technology that creates and provides new experiences to users by fusing the real world with a digitally composed virtual world, and includes VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality). Various business tools and amusement devices using this technology have been developed, providing users with convenience and entertainment, and it has the potential for further use and application in the future.

[0004] As a technology related to a system using cross-reality, for example, there is Japanese Patent Application Laid-Open No. 2023-524446. Here, it is a technology related to a cross-reality system that processes an image obtained using a portable device and restricts the result of position identification based on the estimated direction of gravity of a persistent map and a coordinate frame in which data within a position identification request is pose-attached, thereby enabling the position identification of the portable device with respect to the persistent map quickly and accurately. It is disclosed that with this configuration, it becomes possible to access the stored map and render virtual content defined in relation to those maps efficiently and accurately.

[0005] According to this technology, it is certainly considered that it becomes possible to realistically render so that the user can interact with virtual objects. However, in virtual content in which multiple users participate, the mutual positional relationship and the timing of video become problems. If these are not accurately adjusted for each user, correct rendering cannot be performed, or there is a deviation in the video and information viewed by the users, and there is a problem that it is impossible to provide XR video that can share an immersive experience.

[0006] Also, in Japanese Patent Application Laid-Open No. 2023-501952, a technology related to a cross-reality system that is shareable among a plurality of user devices is disclosed. Here, a technology for rendering virtual content and providing an immersive user experience by providing quality information about a shared map is disclosed.

[0007] According to this technology, it is certainly considered that it becomes possible to render virtual content that can be shared by multiple users in virtual content in which multiple users participate. However, also in virtual content in which multiple users participate, it is necessary to calibrate the mutual positional relationship and the timing of video for each user, and there is a possibility that there is a problem that correct rendering cannot be performed, or there is a deviation in the video and information viewed by the users, and it is impossible to provide XR video that can share an immersive experience.

[0008] Therefore, there was a need for the development of an in-video space calibration system that would enable users to easily and intuitively experience an immersive and highly immersive XR space by performing calibration processing on positional relationships and video timing to prevent positional and temporal discrepancies among one or more users viewing XR-related video via a viewing device, and by performing the calibration processing in a simple and low-cost configuration, and in a functional, rapid, and intuitively matching manner, thereby enabling multiple users to synchronously participate in the same XR-related space and share the emotional experience.

[0009] Special table 2023-524446 publication Special table 2023-501952 publication

[0010] The present invention relates to a system for providing a video space viewable by users through XR (Cross Reality), and more particularly, aims to provide an in-video space calibration system that enables users to easily and intuitively experience an immersive and highly real-world XR space by performing calibration processing on positional relationships and video timing so that one or more users viewing XR-related video via a viewing device do not experience positional and temporal discrepancies, and by performing the calibration processing in a simple and low-cost configuration, as well as in a functional, rapid, and sensory-responsive manner, and by enabling multiple users to synchronously participate in the same XR-related space and share an emotionally impactful experience.

[0011] To achieve the above objective, the present invention provides an in-spatial video calibration system for calibrating the location information of each user, the location information of content displayed, and the time information in a space composed of images provided by XR (cross reality), comprising: a viewing device that a user can wear to view the images; a current information recognition device that recognizes the current definition information of one or more users wearing the viewing device; a display means that displays the images based on the definition information; and a calibration means that calibrates the location information and / or time information of each user in the space composed of the images based on the information recognized by the current information recognition device.

[0012] Furthermore, the calibration means recognizes the definition information of each user recognized via the current information recognition device, and calibrates the positional and temporal information in the space composed of the video based on this definition information, after which the display means displays the video.

[0013] Furthermore, the current information recognition device comprises one or more controller holders, each having a different shape, and one or more controller docks, each having a shape corresponding to the controller holders. Each user attaches and fixes the controller holder to the controller dock, thereby recognizing the user's defined information and transmitting it to the calibration means of the viewing device. The calibration means then performs calibration processing of positional information and / or time information within the space composed of the video. The current information recognition device also comprises a docking station in which the controller docks are installed at predetermined intervals.

[0014] Furthermore, the video consists of multiple images and includes at least 3DCG content, the entirety of live-action footage, a part of the video, and assets that constitute the video. The viewing device includes a video selection means that selects and displays one or more of the multiple videos. The video selection means is configured to automatically select and display a viewable video and / or the video itself based on the positional relationship in which the video is displayed, according to the fixing conditions when the controller holder is attached and fixed to the controller dock. The video selection means is also configured to automatically select and display a viewable video and / or the video itself based on the positional relationship in which the video is displayed, according to one or more combinations of the fixed combination information when the controller holder is attached and fixed to the controller dock, the angle when the controller holder is attached and fixed to the controller dock, and the on / off status of one or more of the buttons on the controller holder.

[0015] Furthermore, the video space calibration system according to the present invention is a video space calibration system for calibrating the location information of each user, the location information and / or time information on which content is displayed in a space composed of images provided by XR (cross reality), and comprises a viewing device that a user can wear to view the images and recognize the current definition information of the user wearing the device, the viewing device comprising a display means for displaying the images based on the definition information, and a calibration means for calibrating the location information and / or time information of each user in the space composed of the images based on the definition information. Furthermore, the calibration means recognizes the current definition information of each user and calibrates it to the location information and / or time information in the space composed of the images, after which the display means displays the images related to the space.

[0016] Furthermore, the viewing device is equipped with an input means for acquiring external information, and the viewing device is configured to detect the user's hand based on the external information acquired by the input means, and to detect when the user's hand is placed, thereby causing the calibration means to perform calibration processing of each user's position information within the space composed of the video. Furthermore, the viewing device is configured to detect when the user's hand is placed in a predetermined position, thereby causing the calibration means to perform calibration processing of each user's position information within the space composed of the video. Furthermore, the viewing device is configured to detect when the user's hand is placed in a predetermined position corresponding to the position of the hand when the hand is placed on the sheet.

[0017] Furthermore, the viewing device includes an angle detection means for detecting the angle of the user's hand, and an angle adjustment means for adjusting the angle of the image displayed on the display means. The angle adjustment means is configured to adjust the angle of the image displayed on the display means according to the angle of the user's hand detected by the angle detection means during or after the calibration process of each user's position information by the calibration means. In addition, the viewing device is configured to select and display one or more of the following from 3DCG content, live-action footage, the entire video, a part of the video, or assets constituting the video when the input means detects information in the external information acquired that matches a defined information.

[0018] Furthermore, the viewing device is configured to perform calibration processing of location information on which content is displayed when it detects information in the external information acquired by the input means that matches the definition information designated as applicable. Furthermore, the viewing device is configured to perform calibration processing of time information when it detects information in the external information acquired by the input means that matches the definition information designated as applicable.

[0019] Furthermore, the external information consists of a two-dimensional code or information such as a string obtained therefrom, external video and / or audio, or other matching target information, and the definition information consists of a two-dimensional code or information such as a string obtained therefrom corresponding to the external information, external video and / or audio, or other matching target information. Furthermore, the viewing device is configured to select and display one or more of the following based on the information when it detects information existing on the web that can be accessed based on the URL information obtained from the external information: 3DCG content, the whole of live-action video, a part of the video, or assets that constitute the video.

[0020] Furthermore, the viewing device is configured to perform calibration processing of location information where content is displayed based on the information detected when it detects information existing on the web that can be accessed based on the URL information obtained from the external information. Furthermore, the viewing device is configured to perform calibration processing of time information based on the information detected when it detects information existing on the web that is accessed based on the URL information obtained from the external information. Furthermore, the viewing device is configured to detect the user's hand based on the external information acquired by the input means, and then detect when the user's hand is placed in a predetermined position corresponding to the position of the hand when the hand is placed on the sheet.

[0021] Furthermore, the viewing device is configured to superimpose and display image elements on the detected image of the user's hand to initiate one or more of the following processes: time information calibration processing, location information calibration processing, and content display processing. Additionally, the viewing device is configured to superimpose and display image elements on the image of one of the detected user's hands, and to initiate one or more of the following processes: time information calibration processing, location information calibration processing, and content display processing when the user's other hand is detected at the location where the image elements are placed.

[0022] Furthermore, the viewing device includes tracking means for tracking specific elements included in the external information acquired by the input means, and the tracking means is configured to track the user's hand, which is an element included in the external information acquired and detected by the input means. Furthermore, the video consists of multiple videos and includes at least 3DCG content, live-action footage, the whole video, a part of the video, and assets that constitute the video, and the viewing device includes video selection means for selecting and displaying one or more of the videos consisting of multiple videos, the input means acquires the image attached to the sheet, and the video selection means recognizes the acquired image and then compares it with definition information based on the information read from the image or accesses information on the web via a URL.

[0023] Furthermore, the viewing device is configured to perform time information calibration when it receives trigger information. It is also configured to start displaying content when it receives trigger information. Moreover, the viewing device is configured to either hold time-related information or acquire time-related information from an external source, and is configured to perform time information calibration processing at the start of the video display process and / or during the video display process.

[0024] As the present invention has the configuration described in detail above, it has the following effects: 1. Because the current configuration includes an information recognition device, it is possible to acquire reference definition information, and the user's location information, the location information and time information on which the content is displayed can be calibrated and unified, enabling users to easily experience an immersive and highly realistic XR space, and enabling multiple users to synchronously participate in the same XR space and share the emotional experience. Furthermore, because the viewing device is configured to include a calibration means, it is possible to easily, functionally, and quickly calibrate each user's location information, the location information and time information on which the content is displayed. 2. Because the calibration means is configured to calibrate to the location information and time information in the XR space, accurate calibration enables users to experience an immersive and highly realistic XR space, and enables them to share the emotional experience by synchronously participating in the same XR space.

[0025] 3. The current information recognition device has a configuration in which the controller holder is attached and fixed to the controller dock, making it possible to perform calibration processing of position information and time information functionally and reliably with a simple structure. 4. The current information recognition device has a configuration in which a docking station is provided, making it possible to install multiple controller docks at predetermined intervals.

[0026] 5. Since the viewing device is configured to include a video selection means, it becomes possible to select and display any video according to the user's selection. Furthermore, since the device is configured to select viewable videos and / or the videos themselves from the positional relationship in which the videos are displayed according to the combination information of the controller holder fixed to the controller dock, it becomes possible to select viewable videos and / or the videos themselves from the positional relationship in which the videos are displayed according to the number of controller docks. 6. Since the video selection means is further configured to select videos in combination with the angle of the controller holder fixed to the controller dock and the on / off status of one or more of the controller buttons, it becomes possible to increase the number of viewable videos and / or the videos themselves from the positional relationship in which the selectable options are displayed.

[0027] 7. The calibration means is configured to calibrate each user's location and time information within the XR space based on definition information, enabling users to enjoy the XR space across platforms, even without using a controller and regardless of the type of viewing device or equipment. Furthermore, functional and accurate calibration provides users with an immersive and highly realistic XR space, allowing each user to simultaneously share an emotional experience. 8. The calibration means is configured to calibrate the definition information to location and time information within the XR space, enabling simple, functional, rapid, and accurate calibration processing.

[0028] 9. The viewing device is configured to initiate calibration processing when it detects the user's hand, making it possible to perform calibration processing intuitively and simply by detecting the placement of the user's hand, thus enabling the provision of an XR space where calibration processing is performed in a low-cost, intuitive way that matches the user's experience. 10. The viewing device is configured to initiate calibration processing when it detects that the user's hand is placed in a predetermined position, making it possible to perform calibration processing linked to that predetermined position and display the image, thus enabling the provision of an XR space where calibration processing is performed in a low-cost, functional, and intuitive way that matches the user's experience. 11. The viewing device is configured to detect when the user's hand is placed in a predetermined position corresponding to the position of the hand when placed on a seat, thus enabling the provision of an XR space where calibration processing is performed in a simple, low-cost, functional, and intuitive way that matches the user's experience.

[0029] 12. Because the viewing device is configured to include angle detection means and angle adjustment means, it is possible to accurately perform calibration processing of each user's position information and angle adjustment of the image displayed on the display means using the hand angle information detected by the user rotating their hand, thereby providing an XR space that matches the user's experience. 13. Because the viewing device is configured to perform image display processing when certain external information is detected, it is possible to perform image display processing simply and reliably.

[0030] 14. The viewing device is configured to perform location information calibration processing when it detects certain external information, making it possible to start the calibration process in a simpler way. This allows users to easily share and enjoy an immersive and highly immersive XR space that has undergone location information calibration processing. 15. The viewing device is configured to perform time information calibration processing when it detects certain external information, making it possible to start the calibration process in a simpler way. This allows users to easily share and enjoy an immersive and highly immersive XR space that has undergone time calibration processing.

[0031] 16. The system is configured to acquire external information such as two-dimensional codes, strings of characters obtained from two-dimensional codes, external video, audio, or other matching target information. By determining whether this matches the corresponding definition information, it becomes possible to start each process, and it is also possible to store all kinds of information as definition information for starting processing. 17. The system is configured to perform video display processing based on certain information when the viewing device detects that information on the accessed web. This makes it possible to perform video display processing simply, functionally, and reliably.

[0032] 18. The viewing device is configured to perform calibration processing on the location information where the content is displayed based on certain information present on the accessed web, so that users can easily share and enjoy an immersive and highly immersive XR space that has undergone calibration processing. 19. The viewing device is configured to perform calibration processing on the time information where the content is displayed based on certain information present on the accessed web, so that users can easily share and enjoy an immersive and highly immersive XR space that has undergone calibration processing on time information, which is inherently prone to discrepancies.

[0033] 20. The viewing device is configured to detect the user's hand and then to detect that the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet. This makes it easier to detect the user's hand in its proper position than to detect a hand arbitrarily placed in the air without the sheet's guidance, enabling far more reliable and accurate calibration of the user's position information. 21. The viewing device is configured to place image elements on the image of the detected user's hand. This makes it possible to start content display processing based on the user's action on the image elements.

[0034] 22. The viewing device is configured to place image elements on the video of one of the user's hands that it has detected, and then detect whether the other hand is superimposed on it, making it possible to start content display processing in response to the intuitive movement of the user's hand. 23. The viewing device is configured to include tracking means, making it possible to perform processing to track the user's hand in the video, and making it easy to trigger various processes even when the user's hand moves.

[0035] 24. The image selection means is configured to match information read from an image with definition information or to access information on the web via a URL, making it possible to perform any processing, such as content display processing, simply by appropriately changing the image in which the information can be recognized. 25. The system is configured to perform time information calibration when trigger information is received, making it possible to start the calibration process by manipulating image elements superimposed on the user's hand, or by operations performed by staff other than the user.

[0036] 26. The system is configured to start displaying content when trigger information is received, making it possible to initiate calibration processing through operations such as those performed by the user on the superimposed image elements on their hand, or by operations performed by staff other than the user. 27. The viewing device is configured to either hold or acquire time-related information from an external source, making it possible to perform time information calibration processing with accurate time.

[0037] The in-spatial calibration system according to the present invention will be described in detail below based on the embodiments shown in the drawings. Figure 1 is a schematic diagram of the in-spatial calibration system according to the present invention, and Figure 2 is a diagram showing the controller holder and controller dock. Figure 3 is a schematic diagram showing an embodiment of the in-spatial calibration system, and Figure 4 is a schematic diagram showing another embodiment of the in-spatial calibration system. Figure 5 is a plan view of the sheet.

[0038] The video space calibration system 1 according to the present invention, as shown in Figure 1, consists of a viewing device 100 and a current information recognition device 200, and is a video space calibration system that enables the synchronous provision of an immersive and highly immersive XR space and the sharing of emotional experiences by calibrating the location information of each user, the location information where content is displayed, and the time information in any space composed of video 10 provided in XR (cross reality). The video space calibration system 2 according to the present invention can also be configured to consist of multiple viewing devices.

[0039] In this invention, Extended Reality (XR) refers to a technology that creates and provides new experiences to users by constructing a virtual world and / or fusing it with reality. This concept includes VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and SR (Substitutional Reality).

[0040] The viewing device 100 is a device consisting of one or more components worn by the user, which plays and displays XR-related images 10, enabling the user to view them. In this embodiment, the viewing device 100 mainly consists of a goggle-type projector, but is not limited to this. Any projector that can be worn by the user and that can display the real world and the image superimposed on it can be appropriately selected and used. For example, a retinal projection system that directly forms and projects an image onto the retina can be used.

[0041] In this embodiment, the viewing device 100 is equipped with a storage means 170 consisting of a computing device (not shown) and an arbitrary storage medium, and the video 10 and the like are stored in the storage means 170. The video 10 displayed on the viewing device 100 is an XR-related video, and may be configured to display a video in a virtual space, or to overlay the video onto the real world in front of the user. The video 10 may include 3DCG content, the whole of a video such as live-action footage, a part of a video, or assets that make up a video.

[0042] As shown in FIG. 1, the viewing device 100 can be configured to be connected to a computer such as a server or the cloud via the Internet. In this embodiment, although the video 10 is configured to be managed, calculated, and stored by the viewing device 100, a storage means (not shown) composed of a calculation device (not shown) and an arbitrary storage medium can be installed on a computer or the cloud, the video 10 is stored in the storage means, and each viewing device 100 can be configured to acquire the video 10 via a network. Also, a computer such as a server or the cloud can be configured to be used for holding and calculating position information and the like.

[0043] The current information recognition device 200 is a device that recognizes the current defined information 40 of each user. In this embodiment, as shown in FIG. 1, the user wears the viewing device 100, and the current information recognition device 200 shown in FIG. 2 recognizes the defined information 40 of each user. The recognized information is configured to be subjected to calibration processing by a calibration means 130 described later. Note that, in this embodiment, the term "recognition" is used to mean temporarily acquiring and holding the defined information 40 of each user, and indicates the system state in which the current information recognition device 200 grasps the information. Also, the calibrated position information refers to, in this embodiment, the origin coordinates and orientation of the user, but is not limited thereto. Also, the coordinates of the calibrated position information can be appropriately selected and used, such as the position of the user's hand, the position of the head, the position of the head-mounted display, and the like.

[0044] In this embodiment, the current information recognition device 200 is composed of a member independent of the viewing device 100 and is configured to be installed at an arbitrary location, regardless of indoors or outdoors. One or more users wearing the viewing device 100 are configured to transmit the defined information 40 via the current information recognition device 200. The defined information 40 means the fixing conditions when the controller holder 210 of the current information recognition device 200 is fixed to the controller dock 220. This fixing condition will be described later.

[0045] As shown in Figure 1, the viewing device 100 in this embodiment is configured to include an information holding means 110, a display means 120, and a calibration means 130. The current user definition information 40 acquired by the information recognition device 200 is calibrated by the calibration means 130, which will be described later. Each piece of content, consisting of the calibrated location information, time information, and video 10, is held by the information holding means 110.

[0046] Furthermore, the space formed by the video 10 may contain either one user or multiple users. The information holding means 110 can be configured to hold the calibrated location and time information of other users for each user. This makes it possible to hold the location information of each user, as well as the location and time information where the content is displayed.

[0047] In this embodiment, the information holding means 110 consists of software that is read and processed by a computing device incorporated in the viewing device 100 from the storage means 170. However, it is not limited to this configuration. For example, it is also possible to set up an external computer or cloud, and have the information holding means 110 equipped on the computer or on the cloud perform calculations and store information related to each content, including the location information of each user after calibration, the location information where the content is displayed, time information, and the video 10. It is also possible to migrate only a part of the functions to the cloud or a web application and then perform the processing there.

[0048] The display means 120 is a means for displaying the video 10 so that the user can visually recognize it based on the definition information 40. In this embodiment, the video 10 stored or temporarily acquired by the storage means 170 or the like is output and displayed in the video display area of the viewing device 100 (HMD: Head-Mounted Display) having a goggle shape. This video display area can be configured to simultaneously display the video of the external real world acquired by the input means 150 described later, and is a configuration in which a video composed of CG or the like is superimposed on the real world. The video 10 is configured such that the display means 120 displays it in the video display area of the HMD after being arithmetic processed to be in an optimal position based on the calibrated position information.

[0049] The display means 120 controls the display of the video 10 based on the position information and time information of the user calibrated by the calibration means 130 based on the definition information 40. When the video 10 is a complete virtual space, it is displayed after calculating the standing position and orientation of each user, etc. When the video 10 is a partial video that is superimposed on the real world, it calculates and displays the place, positional relationship, etc. where the video is superimposed. At this time, it includes, but is not limited to, a stencil test, which is a process of determining whether to draw an object on a pixel using data referred to and edited to achieve occlusion, and as a result, it is possible to display a more realistic and depth-perceivable video 10 that correctly reflects the drawing order in front of and behind the overlapping of objects by causing occlusion, that is, the occlusion of CG by real objects.

[0050] In this embodiment, the display means 120 is composed of software that is read from the storage means 170 and processed by the arithmetic device incorporated in the viewing device 100, but it is not limited to this. For example, by providing an external computer or cloud, or using a web application, it is also possible to configure the display means 120 equipped on the computer or on the cloud or web application to perform arithmetic processing and output the video 10. Also, it is possible to configure to perform each process after migrating only a part of the functions to the cloud or web application.

[0051] The calibration means 130 is a means for calibrating location information and time information. In this embodiment, based on the definition information 40 recognized by the information recognition device 200, the location information of each user, the location information and time information where content is displayed in the space composed of the video 10 are calibrated, and the information holding means 110 holds the calibrated location information and time information. This configuration makes it possible to calibrate and synchronize the location information of each user, the location information and time information where content is displayed, and the location information of users participating in the XR space, as well as the location information and time information where content is displayed, in a simple, functional, rapid, and accurate manner. In other words, it becomes possible to define location in the XR space, which is different from the real world (defining the XR world), and to synchronize time, allowing users to share an immersive and highly real-world XR space that matches their intuition.

[0052] The calibration process by the calibration means 130 will now be explained. First, the calibration means 130 recognizes the current definition information 40 of each user via the current information recognition device 200.

[0053] Next, the definition information 40 is converted into positional and temporal information in the space composed of the video 10 and calibrated. After that, the information holding means 110 holds the calibrated positional information of each user in the space composed of the video 10, as well as the positional and temporal information where the content is displayed. This enables the display means 120 to display the video 10 synchronously and accurately, and through synchronous and accurate calibration, it becomes possible to provide an immersive and highly realistic XR space that matches the user's experience, allowing users to enjoy XR videos while sharing an emotional experience.

[0054] In this embodiment, the calibration means 130 consists of software that is read from the storage means 170 and processed by a computing device incorporated in the viewing device 100. However, it is not limited to this configuration. For example, it is also possible to configure the calibration means 130 to perform calculations and calibration processing by providing an external computer or cloud, or by using a web application. It is also possible to migrate only a part of the functions to the cloud or web application and then perform each processing there.

[0055] Next, the current information recognition device 200 will be described. In this embodiment, the current information recognition device 200 consists of a controller holder 210 and a controller dock 220, as shown in Figure 2. The controller holder 210 is a component attached to the controller and has the function of holding the controller, and is a component for fitting into the controller dock 220, which will be described later, when using the video space calibration system 1. One or more controller holders 210 are provided, and each has a different shape, and in this embodiment, they are shaped like uppercase letters of the alphabet, but are not limited to this shape.

[0056] The controller dock 220 is a component for fitting and fixing the controller holder 210, and the configuration includes one or more controller docks 220 (a number corresponding to the controller holder 210) that have a shape corresponding to the controller holder 210.

[0057] When a user uses the video space calibration system 1, they first attach and secure the controller holder 210 to the controller dock 220. This action triggers the current information recognition device 200 to recognize the user's current definition information 40. Based on the recognized information, the calibration means 130 performs calibration processing on each user's position information, the position information where content is displayed, and the time information within the space composed of the video 10. This configuration makes it possible to start the calibration processing of position information and time information quickly, functionally, and reliably with a simple structure.

[0058] Currently, the information recognition device 200 can be configured to include a docking station 230, as shown in Figure 2. The docking station 230 is a component that allows controller docks 220 to be installed at predetermined intervals, is portable and easy to handle, allows controller docks 220 to be installed together in one place, enables appropriate placement, and is easily visible to the user, making it easy for the user to find the location of the docks. As a result, the controller holders 210 can be installed at predetermined intervals, making it possible to provide a current information recognition device 200 that can perform calibration processing simply, quickly, and functionally.

[0059] The video 10 viewable by the display means 120 of the viewing device 100 is not limited to a single video but consists of multiple videos of any kind, and is configured to allow selection and display of a wide variety of videos 10. In this embodiment, the video 10 includes at least 3DCG content, the entirety of live-action footage, a part of a video, and assets that constitute a video, and multiple videos 10 selected from among these can be viewed on the viewing device 100.

[0060] As shown in Figure 1, the viewing device 100 is configured to include a video selection means 140. The video selection means 140 selects and displays one or more videos from a plurality of videos 10. In this embodiment, after selecting a video 10, the video selection means 140 performs a process to adjust the display of the video 10 according to the positional relationship as seen from the user's position. The viewing device 100 is capable of viewing the real world, and as shown in Figure 3, it is possible to superimpose the various videos selected by the video selection means 140 onto the real world in front of the user's eyes. At this time, the video 10 may be a video relating to a complete virtual reality world, or it may be configured to superimpose the video 10 onto the entire real world. Alternatively, it may be configured to display the selected video 10 on a part of the real world, or it may be configured to superimpose various video assets onto a large number of objects visible in the real world.

[0061] The display means 120 of the viewing device 100 displays various contents composed of images 10 held by the information holding means 110. In this embodiment, the content consists of, for example, 3DCG content, the whole of live-action footage, a part of an image, or assets that make up an image. Specifically, possible configurations include, for example, a configuration in which a space appears on the floor and outer space or a space probe is displayed, a configuration in which a box becomes transparent and dancers sing and dance, a configuration in which tourist attractions appear in the real world, a configuration in which various advertisements are superimposed on the real world, a configuration in which lively illuminations are superimposed and displayed, a configuration in which objects are displayed on an empty table and shared by each user, and other configurations that perform image processing such as breaking down a wall displayed as the real world, or configurations that perform effects such as making the floor or box displayed as the real world transparent.

[0062] For example, in a configuration where an object is displayed as an image 10 on a table, the positional relationship in which the object is seen will differ because each user is in a different position. Furthermore, if all users cannot view the object at the same time, a problem arises in that a synchronized experience cannot be achieved in the XR space. Since the object display process is performed by each viewing device 100, this problem occurs if the positional and temporal information is not calibrated. By calibrating the positional and temporal information using the calibration means 130, all users participating in the XR space can view the same object synchronously from any positional relationship, making it possible to share and enjoy the excitement in an immersive and highly real-world XR space that matches the physical sensations.

[0063] In this embodiment, the video selection means 140 is configured to automatically select a video 10 according to the fixing conditions when the controller holder 210 is attached and fixed to the controller dock 220. In this embodiment, it automatically selects and displays a viewable video or the video 10 itself based on the positional relationship in which the video 10 is displayed. For example, the controller dock 220 is provided in five types, labeled "A," "B," "C," "D," and "E," and the user can select one of them. The video selection means 140 selects a video 10 from information relating to the combination of the selected controller holder 210 and controller dock 220. With this configuration, it becomes possible to select a positional relationship in which videos are displayed according to the number of controller holders 210 and controller docks 220, making it possible for the user to select videos in an easy-to-understand and simple way.

[0064] Furthermore, the video selection means 140 is configured to automatically select the video 10 itself or a viewable video based on the positional relationship in which the video 10 is displayed, according to the combination of the angle at which the controller holder 210 is attached and fixed to the controller dock 220 and the on / off status of one or more buttons provided on the controller holder 210. For example, the controller dock 220 can be set to allow selection of three installation angles for the controller holder 210: vertical, diagonal, and planar, with three options for 0 degrees, 90 degrees left, and 90 degrees right. Adding the on / off status of the buttons (or more if there are multiple buttons) makes it possible to select a total of 18 different videos. In other words, it is possible to increase the number of videos that can be selected in a way that is easy for the user to understand.

[0065] In this embodiment, the video selection means 140 consists of software that is read and processed by a computing device incorporated in the viewing device 100 from the storage means 170. However, it is not limited to this configuration. For example, it is also possible to configure the video selection means 140 to be installed on an external computer or cloud, or to be located on a web application, and to perform calculation processing to select a video. It is also possible to migrate only a part of the functions to a cloud or web application and then perform each processing there.

[0066] Next, another embodiment of the video space calibration system according to the present invention will be described. The video space calibration system 2 is a system for calibrating the location information of one or more users, the location information and time information of content displayed in a space composed of video provided by XR (cross reality), and as shown in Figure 4, it can be configured to consist of one or more viewing devices 100.

[0067] The viewing device 100 is a component that allows a user to wear it and view the video 10. In this embodiment, the viewing device 100 is configured to recognize the current definition information 40 of the user wearing it.

[0068] As shown in Figure 4, the viewing device 100 is configured to include information holding means 110 that each holds information about the location of one or more users in the space composed of the video 10, location information where content is displayed, time information, and each piece of content composed of the video 10, and display means 120 that displays the video 10 based on definition information 40.

[0069] Furthermore, as shown in Figure 4, the viewing device 100 is configured to include a calibration means 130. The calibration means 130 is a means for calibrating the user's location information, the location information and time information where content is displayed, within the space composed of the video 10. In this embodiment, based on the definition information 40, the calibration means 130 performs calculation processing to calibrate the location information of each user within the space composed of the video 10, and the location information and time information where content is displayed, and the information holding means 110 holds this information as calibrated location and time information. This configuration makes it possible to perform calibration processing without using a separate current information recognition device, and by calibrating each user's location information, the location information and time information where content is displayed simply, quickly, and accurately, it becomes possible to share an emotional experience in an immersive and highly realistic XR space that is more in line with the user's experience.

[0070] The calibration process by the calibration means 130 will now be explained. First, the calibration means 130 recognizes the current definition information 40 of each user.

[0071] Next, each definition information 40 is converted into positional and temporal information in the space composed of the video 10 and calibrated. After that, the information holding means 110 holds the calibrated positional information of each user in the space composed of the video 10, as well as the positional and temporal information where the content is displayed. This enables the display means 120 to display the video 10 synchronously and accurately, and through synchronous and accurate calibration, it becomes possible to provide an immersive and highly realistic XR space that matches the user's experience, allowing users to enjoy XR videos while sharing an emotional experience.

[0072] In this embodiment, the calibration means 130 consists of software that is read from the storage means 170 and processed by a computing device incorporated in the viewing device 100. However, it is not limited to this configuration. For example, it is also possible to configure the calibration means 130 to perform calculations and calibration processing by providing an external computer or cloud, or by using a web application. It is also possible to migrate only a part of the functions to the cloud or web application and then perform each processing there.

[0073] In this embodiment, the viewing device 100 is configured to include input means 150, as shown in Figures 1 and 4. The input means is a means for acquiring external information 20, such as video, audio, and various data, and includes, but is not limited to, cameras, microphones, various sensors, communication functions for acquiring data wirelessly or via wired connections, and ports.

[0074] External information 20 is any information that the input means 150 can acquire, and in this embodiment it consists of, but is not limited to, information such as a two-dimensional code or a string obtained therefrom, external video, audio, or other matching target information. Furthermore, the audio as external information 20 may be the audio data itself, or it may be acquired in a form in which the audio data has been converted into a string and then converted into text.

[0075] The viewing device 100 is configured to detect the user's hand based on external information 20 acquired by the input means 150. After detecting the hand from the external information 20 acquired by the input means 150, the viewing device 100 detects that the user's hand has been placed. Based on the detected information, the calibration means 130 performs calibration processing on the positional and temporal information of each user within the space composed of the video 10. This makes it possible to perform calibration processing intuitively and simply by detecting the placement of the user's hand, and makes it possible to provide an XR space that has been calibrated in a low-cost and tactile way.

[0076] In this embodiment, the viewing device 100 can be configured such that the calibration means 130 performs calibration processing of each user's position information within the space composed of the video 10 when it detects that the user's hand is placed in a predetermined location. For example, when the user places their hand on a predetermined location where an illustration visible to the user is displayed, the calibration means 130 starts the calibration process and the video 10 is displayed. This configuration makes it possible to display the video after performing calibration processing linked to the predetermined location, and makes it possible to provide an XR space in which calibration processing has been performed in a low-cost, functional, and sensory-responsive manner.

[0077] In this embodiment, as shown in Figures 4 and 5, it is possible to provide a sheet S for use in initiating the calibration process by the calibration means 130. The sheet S may be configured to have markings for placing the user's hand, such as a handprint, or it may be configured to have a star mark to encourage the user to place their hand accurately. Alternatively, if the user's hand can be positioned similarly, it is also possible to detect the user's hand without using the guide sheet S and make corrections as needed.

[0078] The viewing device 100 detects a hand from external information 20 acquired by the input means 150, and then detects that the user's hand is placed on a predetermined position corresponding to the position of the hand when placed on the sheet S. Based on the detected information, the calibration means 130 performs calibration processing on the positional and temporal information of each user within the space composed of the video 10. This makes it possible to provide an XR space that is functionally and intuitively calibrated with simple equipment.

[0079] In this embodiment, the sheet S is made of paper or the like, and for example, has a design such as a handprint or a star printed on it. However, it is not limited to this, and it may also have a three-dimensional shape such as a recess for placing a hand, or a configuration in which the sheet S is displayed on a screen and a design such as a handprint or a star is displayed on the sheet S, or other configurations can be selected and used. This configuration makes it possible to provide the in-video space calibration system 1 at a low cost.

[0080] Furthermore, as shown in Figure 4, the viewing device 100 can be configured to include an angle detection means 180 and an angle adjustment means 190. The angle detection means 180 is a means for detecting the angle of the user's hand when the input means 150 detects the user's hand. The user's hand may be detected at any location, or it may be configured to detect whether it is placed in a predetermined position. For example, in this embodiment, after detecting that the user's hand is placed on a marker for placing the user's hand displayed on a sheet S or the like, if the user rotates their hand left or right to change the angle of their hand, the angle detection means 180 detects in which direction and by how many degrees the hand has rotated.

[0081] The angle adjustment means 190 is a means for adjusting the display angle of each content, which consists of the video 10 displayed on the display means 120. The video 10 displayed on the display means 120 is configured to be viewed in an optimal state from the user's position through calibration processing by the calibration means 130, but fine adjustment of the display angle may be necessary. The angle adjustment means 190 makes it possible to perform more optimal video display processing to meet such needs.

[0082] In this embodiment, the angle adjustment means 190 adjusts the angle of the image 10 displayed on the display means 120 according to the angle of the user's hand detected by the angle detection means 180 during or after the calibration process of each user's position information by the calibration means 130. The rotation angle of the hand and the angle of the image 10 and their respective rotation axes may be matched, or they may be different angles and rotation axes. Alternatively, the angle and rotation distance obtained by multiplying the angle and rotation distance of the user's hand by a coefficient may be used as the changed angle and rotation distance of the image 10, ensuring that the optimal angle adjustment does not deviate from the user's perception. This configuration makes it possible to perform optimal image display processing that matches the user's perception.

[0083] The viewing device 100 is configured to perform video display processing when it detects information in the external information 20 that matches the definition information 30. The definition information 30 is information that has been defined in advance and stored in the storage means 170 of the viewing device 100, and consists of information such as a two-dimensional code or a string obtained therefrom, external video, audio, or other matching target information. The definition information 30 corresponds to the external information 20, and the device is configured to start some kind of processing when it detects that the acquired external information 20 matches the definition information 30.

[0084] The viewing device 100 is configured to perform display processing of one or more videos from the entire 3DCG content, live-action footage, or any part of the video, or from the assets that constitute the video, when the external information 20 acquired by the input means 150 contains information that matches the definition information 30 which has been determined to be applicable as described above.

[0085] In this embodiment, the viewing device 100 adjusts to display the image 10 according to the positional relationship as seen from the user's position. It is also possible to adjust the display so that the image is superimposed on the displayed image of the real world. Furthermore, it is possible to configure the device to display a preview of the image 10 in a certain area of ​​the display screen. Examples of image 10 displays include, for example, an effect in which a wall breaks down and a dinosaur appears near the location where the display target is projected, an effect in which the area below the floor becomes outer space and a falcon appears, an effect in which a wooden box becomes transparent and a girl dancing and singing appears, and an effect in which one boards a vehicle synchronized with an existing object. Also, for example, in the case of displaying Kaminarimon, it is possible to configure the device so that when the user passes through Kaminarimon, the image of Kaminarimon moves relative to the user and is displayed in a different positional relationship depending on the user's location.

[0086] For example, if the acquired external information 20 is a video in the form of a two-dimensional code, the viewing device 100 reads the two-dimensional code, decodes it, and obtains the content indicated by the two-dimensional code. The two-dimensional code can be converted into text information, for example, if a specific string is predefined and stored as definition information 30, the viewing device 100 compares the converted content with the string defined as definition information 30. If the two match, the aforementioned process is executed.

[0087] External information 20 may also be strings, images, audio, or audio data converted to text, or a combination of these. The system can be configured to execute processing when this information (or a combination thereof) matches the definition information 30. It is also possible to define a two-dimensional code as the definition information 30 and directly compare it with the external information 20 consisting of the two-dimensional code.

[0088] In another embodiment, the viewing device 100 can be configured to perform calibration processing on the location information (e.g., origin coordinates and orientation) on which the content is displayed when the input means 150 detects information in the external information 20 acquired by the user that matches the definition information 30 that has been determined to be applicable, in order to display content corresponding to the user's current definition information 40.

[0089] For example, if a specific string is predefined and stored as definition information 30, the viewing device 100 reads the two-dimensional code, obtains the content indicated by the two-dimensional code, and then compares the content of the two-dimensional code with the string defined as definition information 30. If the two match, the calibration process for the location information (origin coordinates and orientation) where the content is displayed is initiated. This configuration makes it possible to start the calibration process in a simpler, faster, and more functional way, allowing users to easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration process has been completed.

[0090] In another embodiment, the viewing device 100 can be configured to perform a time information calibration process when it detects information in the external information 20 acquired by the input means 150 that matches the defined information 30 that has been determined to be applicable.

[0091] For example, if a specific string is predefined and stored as definition information 30, the viewing device 100 reads the two-dimensional code, obtains the content indicated by the two-dimensional code, and then compares the content of the two-dimensional code with the string defined as definition information 30. If the two match, the time information calibration process is started. This configuration makes it possible to start the calibration process in a simpler way, allowing users to easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration process has been completed.

[0092] In another embodiment, the viewing device 100 can be configured to access the web based on URL information obtained from external information 20. In this configuration, certain information exists on the web accessed based on the URL information. Based on the information present on the accessed web, the device is configured to display one or more of the following: 3DCG content, live-action footage, the entirety of the video, a part of the video, or assets that constitute the video. The external information 20 used to obtain the URL information in this example consists of a two-dimensional code or text information of the URL, but is not limited to this configuration.

[0093] For example, when the viewing device 100 detects a two-dimensional code as external information 20, it decodes the two-dimensional code. After obtaining the URL text information as the content of the two-dimensional code, it accesses the web based on that information and displays the obtained content in a certain area of ​​the display means 120 of the viewing device 100. As for the display method, in this embodiment, for example, a configuration such as preview display or a configuration that uses the functions of a browser can be considered, but it is not limited to this configuration, and it is of course possible to select and use an appropriate display method, such as a configuration that uses the functions of other web applications.

[0094] Alternatively, the system may be configured to obtain the URL text information as the content of the two-dimensional code, and then launch a web application from that URL string. Another configuration may be used to obtain content consisting of 3DCG content, live-action footage, or other video content from the web that can be displayed on the viewing device 100, or content consisting of a part of a video or assets that constitute a video.

[0095] This configuration makes it possible to display a variety of images 10 based on content information obtained by accessing a web superimposed on the real world. For example, a configuration in which a space appears on the floor and outer space or a space probe is displayed, a configuration in which a box becomes transparent and a dancer sings and dances, a configuration in which tourist attractions appear in the real world, a configuration in which various advertisements are superimposed on the real world, a configuration in which lively illuminations are superimposed and displayed, a configuration in which an object is displayed on an empty table and shared by each user, and other configurations that perform image processing such as breaking down a wall displayed as the real world, or configurations that perform effects such as making the floor or box displayed as the real world transparent.

[0096] Furthermore, in another embodiment, the viewing device 100 can be configured to perform calibration processing on the location information (e.g., origin coordinates and orientation) where the content is displayed, based on the information detected on the web accessed based on URL information obtained from external information 20. For example, the viewing device 100 accesses the web based on the acquired URL information and displays it. Certain information exists on the web, and the calibration processing is initiated based on this information. With this configuration, users can easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration processing.

[0097] Furthermore, in another embodiment, the viewing device 100 can be configured to perform time information calibration processing when it detects information existing on the web accessed based on URL information obtained from external information 20. For example, the viewing device 100 accesses and displays the web based on the acquired URL information. In this configuration, certain information exists on the web, and based on this, the time information calibration processing is initiated. With this configuration, users can easily share and enjoy an immersive and highly realistic XR space that matches the sense of presence after the calibration processing has been performed.

[0098] The viewing device 100 detects the user's hand based on external information 20 acquired by the input means 150. Subsequently, it detects when the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet S. The input means 150 acquires video as external information 20. When the user's hand is detected in the video, it detects that the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet S. This makes it possible to detect the user's hand in its proper position, and the placement of the hand on the sheet S can be used as a trigger to start the calibration process for accurate user position information.

[0099] Furthermore, the viewing device 100 can be configured to superimpose and display image elements on the detected video of the user's hand to prompt the system to perform some kind of processing. In this embodiment, image elements for starting time information calibration processing, position information calibration processing, and content display processing are superimposed and displayed. Each of these processes may be configured to execute one of them, or two or more processes may be selected and executed. This configuration makes it possible to initiate each configuration process or content display process based on the user's action on the image elements.

[0100] Furthermore, the viewing device 100 is configured to start time information calibration processing, position information calibration processing, and content display processing when it detects the other hand at the position where the image elements are placed, while the image elements are superimposed on the video of one of the detected user's hands. Each of these processes may be executed individually, or two or more processes may be selected and executed. This configuration makes it possible to start each configuration process and content display process in response to the intuitive movement of the user's hand. Alternatively, the device may be configured to start time information calibration processing, position information calibration processing, and content display processing when it detects the other hand at the position where one or both of the image elements are placed, while the image elements are superimposed on the video of both of the user's hands.

[0101] Furthermore, the viewing device 100 can be configured to include a tracking means 160, as shown in Figures 1 and 4. The tracking means 160 is a means for tracking a specific element included in the external information 20 acquired by the input means 150, and is configured to track the user's hand, which is an element included in the external information 20 acquired and detected by the input means 150. This configuration makes it possible to easily track the user's hand.

[0102] In this embodiment, the in-video space calibration system 2 is configured such that, as shown in Figure 5, the input means 150 acquires an image I attached to the sheet S. The image I is a pattern that the input means 150 can recognize and acquire information from. In this embodiment, it consists of a two-dimensional code, but is not limited to this; any pattern that the input means 150 can acquire information from can be appropriately selected and used. The video selection means 140 recognizes the acquired image I and, based on the information read from this image I, compares it with definition information or accesses information on the web via a URL. This configuration simplifies and speeds up the calibration of location information and time information, as well as the display of content, making it easy to share and enjoy an XR space consisting of images that match the user's perception.

[0103] The viewing device 100 can be configured to perform time information calibration when it receives trigger information. Trigger information is information other than the external information 20 actively acquired by the input means 150. For example, it may be information generated by the user themselves operating image elements superimposed on their hand, or information generated primarily by system administrators or staff operating the system. When operating from a network device, for example, it may be information generated using an application on a smartphone. This makes it possible to generate trigger information and operate the system easily and functionally, but it is not limited to these. Accuracy of time information is required for the operation of the system, but discrepancies often occur inevitably during execution. With the configuration of the present invention, it becomes possible for users or staff operating the system other than users to initiate the calibration process, thereby ensuring the accuracy of time information.

[0104] Furthermore, the viewing device 100 can be configured to start displaying content when it receives trigger information. This configuration makes it possible for the user or staff other than the user to initiate the calibration process.

[0105] Furthermore, the viewing device 100 can be configured to hold time-related information. The viewing device 100 can also be configured to acquire time-related information from an external source. Moreover, in this embodiment, the device is configured to perform time information calibration processing at the start of the video 10 display process and / or during the video 10 display process. An example of an embodiment in which calibration processing is performed at the start of the video 10 display process or during the display process is the case of live streaming of a competitive race in an XR space.

[0106] Time information can be obtained from time information acquired from an NTP server or from reference time information on the web. This configuration allows for the acquisition and storage of accurate time for each device, enabling time information calibration based on precise time.

[0107] Schematic diagram of the in-video calibration system according to the present invention; diagram showing the controller holder and controller dock; schematic diagram showing an embodiment of the in-video calibration system; schematic diagram showing another embodiment of the in-video calibration system; plan view of the sheet.

[0108] 1.2 In-video space calibration system 10 Video 20 External information 30 Definition information 40 Definition information 100 Viewing device 110 Information holding means 120 Display means 130 Calibration means 140 Video selection means 150 Input means 160 Tracking means 170 Storage means 180 Angle detection means 190 Angle adjustment means 200 Current information transmission means 210 Controller holder 220 Controller dock 230 Docking station S Sheet I Image

Claims

1. An in-video space calibration system (1) for calibrating the location and time information of each user in a space composed of images (10) provided by XR (cross reality), characterized in that it comprises: a viewing device (100) that a user can wear to view the images; a current information recognition device (200) that recognizes the current definition information (40) of one or more users wearing the viewing device (100); a display means (120) that displays the images based on the definition information (40); and a calibration means (130) that calibrates the location and / or time information of each user in the space composed of the images (10) based on the information recognized by the current information recognition device (200).

2. The in-video space calibration system according to claim 1, characterized in that the calibration means (130) recognizes each user's defined information (40) recognized via the current information recognition device (200), and calibrates the positional information and time information in the space composed of the video (10) based on the defined information (40), and then the display means (120) displays the video (10).

3. The current information recognition device (200) comprises one or more controller holders (210) having different shapes, and one or more controller docks (220) having shapes corresponding to the controller holders, wherein each user attaches and fixes the controller holder (210) to the controller dock (220), thereby recognizing the user's defined information (40) and transmitting it to the calibration means (130) of the viewing device (100), and the calibration means (130) then performs calibration processing of each user's position information and / or time information in the space composed of the video (10), as described in 2.

4. The video spatial calibration system according to claim 3, characterized in that the current information recognition device (200) comprises a docking station (230) in which the controller docks (220) are installed at predetermined intervals.

5. The video (10) consists of a plurality of videos, and includes at least 3DCG content, the whole of a video such as live-action footage, a part of a video, and assets that constitute a video; the viewing device (100) is provided with a video selection means (140) that selects and displays one or more of the plurality of videos (10); and the video selection means (140) automatically selects and displays videos that can be viewed from the positional relationship in which the video (10) is displayed, and / or the video (10) itself, according to the video spatial calibration system according to claim 3 or 4, depending on the fixing conditions when the controller holder (210) is attached and fixed to the controller dock (220).

6. The video spatial calibration system according to claim 5, characterized in that the video selection means (140) automatically selects and displays a video that can be viewed from the positional relationship in which the video (10) is displayed, and / or the video (10) itself, according to one or more combinations of the fixed combination information when the controller holder (210) is attached and fixed to the controller dock (220), the angle when the controller holder (210) is attached and fixed to the controller dock (220), and the on / off status of one or more of the buttons on the controller holder (210).

7. An in-video space calibration system (2) for calibrating the location and / or time information of each user in a space composed of images provided by XR (cross reality), comprising a viewing device (100) that a user can wear to view the images (10) and recognize the current definition information (40) of the user wearing the device, wherein the viewing device (100) comprises a display means (120) that displays the images (10) based on the definition information (40), and a calibration means (130) that calibrates the location and / or time information of each user in the space composed of the images based on the definition information (40), characterized in that it comprises a display means (120) that displays the images (10) based on the definition information (40), and a calibration means (130) that calibrates the location and / or time information of each user in the space composed of the images.

8. The in-video space calibration system according to claim 7, characterized in that the calibration means (130) recognizes the current definition information (40) of each user and calibrates it to the position information and / or time information in the space composed of the video (10), and then the display means (120) displays the video (10) relating to the space.

9. The video space calibration system according to 8, wherein the viewing device (100) is equipped with an input means (150) for acquiring external information (20), and the viewing device (100) detects the user's hand based on the external information (20) acquired by the input means (150), and detects that the user's hand has been placed, causing the calibration means (130) to perform calibration processing of the position information of each user in the space composed of the video (10).

10. The video space calibration system according to 98, characterized in that the viewing device (100) detects when a user's hand is placed in a predetermined position, and the calibration means (130) performs calibration processing of the position information of each user in the space composed of the video (10).

11. The video space calibration system according to claim 10, characterized in that the viewing device (100) is configured to detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when the hand is placed on the sheet (S).

12. The viewing device (100) comprises an angle detection means (180) for detecting the angle of the user's hand and an angle adjustment means (190) for adjusting the angle of the image (10) displayed on the display means (120), wherein the angle adjustment means (190) adjusts the angle of the image (10) displayed on the display means (120) according to the angle of the user's hand detected by the angle detection means (180) during or after the calibration process of each user's position information by the calibration means (130), characterized in that the in-video space calibration system according to any one of claims 9 to 11.

13. The video space calibration system according to any one of claims 9 to 11, characterized in that when the viewing device (100) detects information in the external information (20) acquired by the input means (150) that matches the definition information (30) that has been determined to be applicable, it selects and displays one or more of the following: 3DCG content, live-action footage, the whole of the video, a part of the video, or assets that constitute the video.

14. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) performs calibration processing of location information on which content is displayed when it detects information in the external information (20) acquired by the input means (150) that matches the definition information (30) that has been determined to be applicable.

15. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) performs a time information calibration process when it detects information in the external information (20) acquired by the input means (150) that matches the defined information (30) that has been determined to be applicable.

16. The video space calibration system according to any one of claims 9 to 15, characterized in that the external information (20) consists of a two-dimensional code or information such as a string obtained therefrom, external video and / or audio, and the definition information (30) consists of a two-dimensional code or information such as a string obtained therefrom corresponding to the external information (20), external video and / or audio.

17. The video space calibration system according to any one of claims 9 to 11, characterized in that when the viewing device (100) detects information existing on the web that can be accessed based on URL information obtained from the external information (20), it selects and displays one or more of the following based on the information: 3DCG content, the whole of live-action footage, a part of the footage, or assets that constitute the footage.

18. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) detects information existing on the web that can be accessed based on URL information obtained from the external information (20), and performs calibration processing of location information on which content is displayed based on said information.

19. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) detects information existing on the web that is accessed based on URL information obtained from the external information (20), and performs a calibration process for time information based on said information.

20. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) detects the user's hand based on the external information (20) acquired by the input means (150), and then detects that the user's hand is placed in a predetermined position corresponding to the position of the hand when the hand is placed on the sheet.

21. The video space calibration system according to 18, characterized in that the viewing device (100) superimposes and displays image elements for initiating one or more of the following processes on the detected video of the user's hand: time information calibration processing, position information calibration processing, and content display processing.

22. The video space calibration system according to 20, characterized in that the viewing device (100) superimposes and displays image elements on the video of one of the user's hands that it has detected, and when it detects the user's other hand at the position where the image elements are placed, it starts one or more of the following processes: time information calibration process, position information calibration process, and content display process.

23. The viewing device (100) is further equipped with tracking means (160) for tracking specific elements included in the external information (20) acquired by the input means (150), and the tracking means (160) tracks the user's hand, which is an element included in the external information (20) acquired and detected by the input means (150), as described in any one of claims 20 to 22.

24. The video (10) consists of a plurality of videos, and includes at least 3DCG content, the whole of a video such as live-action footage, a part of a video, and assets that constitute a video; the viewing device (100) is provided with video selection means (140) that selects and displays one or more of the plurality of videos (10); the input means (150) acquires the image (I) attached to the sheet (S); and the video selection means (140) recognizes the acquired image (I) and then compares it with definition information based on the information read from the image (I) or accesses information on the web via a URL, characterized in that the video spatial calibration system according to any one of claims 9 to 23.

25. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) performs time information calibration when it receives trigger information.

26. The video space calibration system according to any one of claims 9 to 11, characterized in that the viewing device (100) starts displaying content when it receives trigger information.

27. The viewing device (100) is configured to hold time-related information or to acquire time-related information from an external source, and the video spatial calibration system according to any one of claims 9 to 11 is characterized in that it performs time information calibration processing at the start of the video (10) display processing and / or during the video display processing.

Citation Information

Patent Citations

  • Information processing device, information processing program, and information processing method

    JP2023092003A

  • Cross-reality systems for large-scale environments

    JP2023524446A