Intra-video space calibration system

The in-video space calibration system addresses XR calibration challenges by synchronizing positional and temporal information, ensuring an immersive and realistic shared XR experience for multiple users.

WO2026084046A1PCT designated stage Publication Date: 2026-04-23HASHILUS INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HASHILUS INC
Filing Date
2025-10-17
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing cross-reality (XR) technologies face challenges in accurately calibrating positional and temporal discrepancies among multiple users, leading to incorrect rendering and inability to provide an immersive shared experience.

Method used

An in-video space calibration system that performs calibration processing on positional relationships and video timing using a viewing device, including a current information acquisition device, display means, and calibration means, to ensure synchronized and immersive XR experiences for multiple users.

Benefits of technology

Enables users to easily and intuitively experience an immersive and highly realistic XR space by accurately calibrating positional and temporal information, allowing multiple users to share an emotional experience without discrepancies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025036587_23042026_PF_FP_ABST
    Figure JP2025036587_23042026_PF_FP_ABST
Patent Text Reader

Abstract

[Problem] To provide: an intra-video space calibration system that performs calibration processing relating to a mutual positional relationship and video timing so that deviations in position and time do not occur between users viewing video relating to XR via viewing devices; a user assistance system in which AI provides related information on the basis of acquired information; and a display function in which AI superimposes and arranges content related to real-world objects. [Solution] The present invention comprises: a viewing device that is worn by a user and enables the user to view the video; and a current information acquisition device that recognizes current defining information of each user wearing the viewing device. The viewing device is configured to include: a display means that displays video; a calibration means that calibrates position information and time information in a space constituted by video of each user; and an information holding means that holds each of the calibrated user position information in the space constituted by video, the position information and the time information at which content is displayed, video, and the like.
Need to check novelty before this filing date? Find Prior Art

Description

In-spatial calibration system for video

[0001] The present invention relates to a system for providing a video space viewable by users through XR (Cross Reality), and more particularly to an in-video space calibration system that performs calibration processing on positional relationships and video timing so that positional and temporal discrepancies do not occur between one or more users viewing XR video via a viewing device, and performs the calibration processing in a simple and low-cost configuration, as well as in a functional, rapid, and sensory-responsive manner, thereby enabling users to easily and intuitively experience an immersive and highly real-world XR space, and enabling multiple users to share an emotional experience by synchronously participating in the same XR space.

[0002] Over time, numerous cross-reality (XR) technologies have been developed and are being utilized in various situations, enabling users to gain various simulated experiences by watching videos. Cross-reality technologies are being used in all fields, and many XR technologies have been developed and are in use that allow users to obtain all kinds of information from superimposed images while viewing the real world, enjoy such images, and experience being in a virtual space while remaining indoors.

[0003] Extended Reality (XR) is a technology that creates and provides new experiences to users by fusing the real world with a digitally composed virtual world, and includes VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality). Various business tools and amusement devices using this technology have been developed, providing users with convenience and entertainment, and it has the potential for further use and application in the future.

[0004] As a technology related to a system using cross-reality, for example, there is Japanese Patent Application Laid-Open No. 2023-524446. Here, it is a technology related to a cross-reality system that processes an image obtained using a portable device and restricts the result of position identification based on the estimated direction of gravity of a persistent map and a coordinate frame in which data within a position identification request is pose-aware, thereby enabling the position identification of the portable device with respect to the persistent map quickly and accurately. It is disclosed that this configuration enables access to the stored map and efficient and accurate rendering of virtual content defined in relation to those maps.

[0005] According to this technology, it is certainly considered that it becomes possible for a user to interact with a virtual object realistically. However, in virtual content in which multiple users participate, the mutual positional relationship and the timing of video become problems. If these are not accurately adjusted for each user, correct rendering cannot be achieved, or there is a deviation in the video and information viewed by users, and there is a problem that it is impossible to provide XR video that can share an immersive experience.

[0006] Also, in Japanese Patent Application Laid-Open No. 2023-501952, a technology related to a cross-reality system that is shareable among a plurality of user devices is disclosed. Here, a technology for rendering virtual content and providing an immersive user experience by providing quality information about a shared map is disclosed.

[0007] According to this technology, it is certainly considered that it becomes possible to render virtual content that can be shared by multiple users in virtual content in which multiple users participate. However, also in virtual content in which multiple users participate, it is necessary to calibrate the mutual positional relationship and the timing of video for each user, and there is a possibility that there is a problem that correct rendering cannot be achieved, or there is a deviation in the video and information viewed by users, and it is impossible to provide XR video that can share an immersive experience.

[0008] Therefore, there was a need for the development of an in-video space calibration system that would enable users to easily and intuitively experience an immersive and highly immersive XR space by performing calibration processing on positional relationships and video timing to prevent positional and temporal discrepancies among one or more users viewing XR-related video via a viewing device, and by performing the calibration processing in a simple and low-cost configuration, and in a functional, rapid, and intuitively matching manner, thereby enabling multiple users to synchronously participate in the same XR-related space and share the emotional experience.

[0009] Special table 2023-524446 publication Special table 2023-501952 publication

[0010] The present invention relates to a system for providing a video space viewable by users through XR (Cross Reality), and more particularly to providing an in-video space calibration system that enables users to easily and intuitively experience an immersive and highly realistic XR space by performing calibration processing on positional relationships and video timing so that there is no positional and temporal discrepancy for one or more users viewing XR video via a viewing device, and by performing the calibration processing in a simple and low-cost configuration, as well as in a functional, rapid, and sensory-responsive manner, and by enabling multiple users to synchronously participate in the same XR space and share an emotional experience. Furthermore, in connection with the in-video space calibration system, the invention provides a user assistance system in which AI provides relevant information based on acquired information, and an information content display function in which AI superimposes and arranges content related to real-world objects.

[0011] To achieve the above objective, the present invention provides an in-spatial video calibration system for calibrating the positional and temporal information of each user in a space composed of images provided by XR (cross reality), comprising: a viewing device worn by a user to enable viewing of the images; a current information acquisition device that acquires definitional information for one or more users wearing the viewing device; a display means that displays the images based on the definitional information; and a calibration means that calibrates the positional and / or temporal information of each user in the space composed of the images based on the information acquired by the current information acquisition device, wherein the calibration means acquires the definitional information of each user via the current information acquisition device The display means is configured to display the video after acquiring the definition information and, based on the definition information, calibrating the position information and time information in the space composed of the video. The current information acquisition device consists of one or more controller holders, each with a different shape, and one or more controller docks, each with a shape corresponding to the controller holder. Each user attaches and fixes the controller holder to the controller dock, thereby acquiring the user's definition information and transmitting it to the calibration means of the viewing device. The calibration means then performs calibration processing on the position information and / or time information of each user in the space composed of the video.

[0012] Furthermore, the current information acquisition device is configured to include a docking station in which the controller docks are installed at predetermined intervals. The video consists of multiple videos and includes at least 3DCG content, live-action footage, the entirety of a video, a part of a video, and assets that constitute a video. The viewing device includes a video selection means that selects and displays one or more of the videos consisting of multiple videos. The video selection means is configured to automatically select and display a video that can be viewed from the positional relationship in which the video is displayed, and / or the video itself, depending on the fixing conditions when the controller holder is attached and fixed to the controller dock, and the time of installation.

[0013] Furthermore, the video selection means is configured to automatically select and display videos that can be viewed from the positional relationship in which the video is displayed, and / or the video itself, according to one or more combinations of the fixed combination information when the controller holder is attached and fixed to the controller dock, the angle when the controller holder is attached and fixed to the controller dock, and the on / off status of one or more of the buttons on the controller holder.

[0014] Furthermore, the system is configured to include an accessibility function that provides the user with information in a conversational format, among other methods, by having the AI ​​select relevant information based on acquired location information, information obtained through dialogue, and information obtained from an external camera, and by generating and / or selecting various types of information and content.

[0015] Furthermore, the system is configured to include an information content display function in which the AI ​​generates and / or selects and displays content related to real-world objects based on the location information and information acquired from an external camera. Additionally, when the AI ​​generates and / or selects content related to real-world objects based on the location information and information acquired from an external camera, the system is configured to include an information content display function in which feature points and frames on which the content should be displayed are generated and displayed, and these are superimposed and positioned on the real-world object.

[0016] Furthermore, the video space calibration system according to the present invention is a video space calibration system for calibrating the location information and / or time information of each user in a space composed of images provided by XR (cross reality), and comprises a viewing device that a user can wear to view the images and acquire the current definition information of the user wearing the device, the viewing device comprising a display means for displaying the images based on the definition information and a calibration means for calibrating the location information and / or time information of each user in the space composed of the images based on the definition information, the calibration means acquires the current definition information of each user and calibrates it to the location information and / or time information in the space composed of the images, and then the display means displays the images related to the space.

[0017] Furthermore, the viewing device includes a current location acquisition means for acquiring the user's current location, and the calibration means is configured to acquire the location information acquired by the current location acquisition means as the current definition information for each user, and then calibrate it to the location information and / or time information in the space composed of the video. Furthermore, the current location acquisition means is configured to acquire the user's current location using one or more of the following: GNSS, VPS, beacons, location markers, and image markers.

[0018] Furthermore, the current location acquisition means is configured to acquire directional information that allows the user to infer their viewing direction using a compass (magnetic sensor). The current location acquisition means is also configured to acquire acceleration and angular velocity information that allows the user to infer their viewing direction using a MEMS gyroscope. In addition, the system is configured to have a function in which the AI ​​selects relevant information based on the acquired location information, information acquired based on dialogue, and information acquired from an external camera, generates and / or selects various information and content, and provides this information to the user in a presentation method that includes a dialogue format.

[0019] Furthermore, the system is configured to include an information content display function in which the AI ​​generates and / or selects and displays information and content related to real-world objects based on the location information and information acquired from an external camera. In addition, when the AI ​​generates and / or selects information and content related to real-world objects based on the location information and information acquired from an external camera, the system is configured to also include a function to generate and display feature points and frames on which the content should be displayed, and to superimpose and position these on the real-world object.

[0020] Furthermore, the viewing device is equipped with an input means for acquiring external information, and the viewing device detects the user's hand based on the external information acquired by the input means, and also detects when the user's hand is placed in a predetermined position, thereby acquiring this as definition information for each user, and the calibration means performs calibration processing of the position information of each user in the space composed of the video. In addition, the viewing device is equipped with a means for detecting when the user's hand is placed in a predetermined position corresponding to the position of the hand when the hand is placed on the sheet.

[0021] Furthermore, the viewing device is configured to detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when placed on a real-world object. The viewing device is also configured to detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when placed on a plate-shaped object drawn in a virtual space.

[0022] Furthermore, the viewing device is configured to detect when a user's hand is placed on it, with the prerequisite being that the hand first clenches and then opens. The viewing device is also configured to guide the user to a sheet or a real-world object by displaying one or more images, videos, or other content within the virtual space, illustrating the procedure for placing a hand on the sheet or a real-world object.

[0023] Furthermore, the viewing device is configured such that when the input means detects the user's hand on an object that actually exists, it displays a flat object to be drawn in the virtual space. The viewing device is also configured such that the current position acquisition means detects the plane of the real world visible to the user in a predetermined manner, and then raycasts in a predetermined manner to project and display a flat object.

[0024] Furthermore, the viewing device includes an angle detection means for detecting the angle of the user's hand, and an angle adjustment means for adjusting the angle of the image displayed on the display means. The angle adjustment means is configured to adjust the angle of the image displayed on the display means according to the angle of the user's hand detected by the angle detection means, either during or after the calibration process of each user's position information by the calibration means.

[0025] Furthermore, the viewing device includes finger position detection means for detecting the position of each fingertip of the user's hand. The finger position detection means detects the position of each fingertip of the user's hand, and then the plane detection means estimates a plane from the position information of each finger using a plane regression method, thereby obtaining a plane that matches the user's hand.

[0026] Furthermore, the viewing device is configured to select and display one or more of the following from among 3DCG content, live-action footage, the entire video, a part of the video, or assets that constitute the video, when it detects information in the external information acquired by the input means that matches the definition information designated as applicable.

[0027] Furthermore, the viewing device is configured to perform calibration processing of location information on which content is displayed when it detects information in the external information acquired by the input means that matches the definition information designated as applicable.

[0028] Furthermore, the viewing device is configured to perform a time information calibration process when it detects information in the external information acquired by the input means that matches the defined information that has been designated as applicable.

[0029] Furthermore, the external information consists of a two-dimensional code or information such as a string obtained therefrom, external video and / or audio, and the definition information consists of a two-dimensional code or information such as a string obtained therefrom corresponding to the external information, external video and / or audio.

[0030] Furthermore, the viewing device is configured to select and display one or more of the following based on URL information obtained from the external information: 3DCG content, the entirety of live-action footage, a part of the footage, or assets that constitute the footage, when it detects information existing on the web that can be accessed based on the external information.

[0031] Furthermore, the viewing device is configured to perform calibration processing of location information where content is displayed based on the URL information obtained from the external information when it detects information that exists on the web that can be accessed.

[0032] Furthermore, the viewing device is configured to perform time information calibration processing based on information that is accessed on the web based on URL information obtained from the external information.

[0033] Furthermore, the viewing device is configured to detect the user's hand based on the external information acquired by the input means, and then to detect that the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet / object.

[0034] The viewing device is configured to superimpose and display image elements on the detected video of the user's hand, in order to initiate one or more of the following processes: time information calibration processing, location information calibration processing, and content display processing.

[0035] Furthermore, the viewing device is configured to superimpose and display image elements on the video of one of the user's hands that has been detected, and to initiate one or more of the following processes when the user's other hand is detected at the position where the image elements are placed: time information calibration processing, position information calibration processing, and content display processing.

[0036] The viewing device includes tracking means for tracking specific elements included in the external information acquired by the input means, and the tracking means is configured to track the user's hand, which is an element included in the external information acquired and detected by the input means.

[0037] Furthermore, the video consists of multiple images and includes at least 3DCG content, the entirety of live-action footage, a part of the video, and assets that constitute the video. The viewing device is equipped with video selection means for selecting and displaying one or more of the multiple videos. The input means acquires the image attached to the sheet / object, and the video selection means recognizes the acquired image and then compares it with definition information based on the information read from the image or accesses information on the web via a URL.

[0038] Furthermore, the viewing device is configured to perform time information calibration when it receives trigger information. Additionally, the viewing device is configured to start displaying content when it receives trigger information.

[0039] Furthermore, the viewing device is configured to either hold time-related information or acquire time-related information from an external source, and is configured to perform time information calibration processing at the start of the video display process and / or during the video display process.

[0040] As the present invention has the configuration described in detail above, it has the following effects: 1. Because the current information acquisition device is provided, it is possible to acquire reference definition information, and the user's location information, the location information and time information on which the content is displayed can be calibrated and unified, enabling users to easily experience an immersive and highly realistic XR space, and enabling multiple users to synchronously participate in the same XR space and share the emotional experience. Furthermore, because the viewing device is configured to include a calibration means, it is possible to easily, functionally, and quickly calibrate each user's location information, the location information and time information on which the content is displayed. Also, because the calibration means is configured to calibrate to the location information and time information in the XR space, accurate calibration enables users to experience an immersive and highly realistic XR space, and enables them to share the emotional experience by synchronously participating in the same XR space. Furthermore, because the current information acquisition device is configured to attach and fix the controller holder to the controller dock, it is possible to perform calibration processing of location information and time information with a simple structure, functionally and reliably. 2. Currently, the information acquisition device is configured to include a docking station, making it possible to install multiple controller docks at predetermined intervals.

[0041] 3. Since the viewing device is configured to include a video selection means, it becomes possible to select and display any video according to the user's choice. Furthermore, since the device is configured to select viewable videos and / or the videos themselves based on the positional relationship in which the videos are displayed according to the combination information of the controller holder fixed to the controller dock, it becomes possible to select viewable videos and / or the videos themselves based on the positional relationship in which the videos are displayed according to the number of controller docks. 4. Since the video selection means is further configured to select videos based on a combination of the angle of the controller holder fixed to the controller dock and the on / off status of one or more of the controller buttons, it becomes possible to increase the number of viewable videos and / or the videos themselves based on the positional relationship in which the selectable options are displayed.

[0042] 5. To assist users, the system is configured to obtain information from the user in a conversational format, the AI ​​selects relevant information, and provides various information and content (including, but not limited to, relevant language information, map information, and / or video) in a conversational format. This enables verbal interaction with the system and allows the system to provide information and content that aligns with the user's needs and intentions (for example, if the content is about the Asakusa area, it can provide information related to the history, tourism, and safety of Asakusa, and display maps, videos, and other content). 6. In addition to information obtained through conversation, the system is configured to generate and / or select various information and content (including, but not limited to, relevant language information, map information, and / or video) in association with location information acquired by the AI ​​and information acquired from an external camera, making it possible to provide the user with the most optimal information and content.

[0043] 7. Further, based on the position information and the information obtained from the external camera, the AI is configured to have a function of generating, and / or selecting, and displaying information and content related to objects in the real world. Therefore, the user can not only visually recognize the objects actually visible in the real world, but also visually recognize, appreciate, and experience the extended information, explanations, images, and videos thereof in a superimposed manner (for example, it becomes possible to superimpose the explanations and historical videos of the Five-Story Pagoda in Asakusa). 8. Also, when the AI generates, and / or selects, information and content related to objects in the real world based on the position information and the information obtained from the external camera, characteristic points and frames for displaying the information and content are generated and displayed (including, but not limited to, the case where it is generated and displayed in a wireframe), and the AI or the user can easily detect the deviation from the real world by using the characteristic points and frames (for example, wireframes), and by automatically adjusting the AI or the user adjusting to a comfortable position by themselves, it becomes possible to adjust and display accurate content at each detail of the object (for example, the exact positions such as the tip, the third stage, and the first stage of the Five-Story Pagoda in Asakusa).

[0044] 9. Since the calibration means is configured to calibrate the position information and time information of each user in the XR space based on the definition information, the user can enjoy the XR space across platforms without using a controller, regardless of the type of viewing device or device. Also, by functional and accurate calibration, it is possible to provide the user with an immersive and highly realistic XR space, and each user can share a moving experience at the same time. Furthermore, since the calibration means is configured to calibrate the definition information to the position information and time information in the XR space, it is possible to perform the calibration process simply, functionally, quickly, and accurately. 10. Since it is configured to include a current position acquisition means, it is possible to calibrate the position information and time information related to the video to be displayed to an optimal one based on the definition information of the user.

[0045] 11. Because the current location acquisition means is configured to use GNSS, VPS, beacons, location markers, image markers, etc., it is possible to acquire the user's current location (approximate absolute position based on XYZ coordinates) by any means. 12. Because the current location acquisition means is configured to use a compass (magnetic sensor), it is possible to acquire the direction and orientation that allows the user to infer their viewing direction.

[0046] 13. Because the current position acquisition means uses a MEMS gyroscope, it is possible to acquire acceleration and angular velocity information that allows the user to infer their viewing direction. 14. For user assistance, the system is configured to acquire information from the user in a conversational format, the AI ​​selects relevant information, and provides various information and content (including, but not limited to, relevant language information, map information, and / or video) in a conversational format. This enables verbal interaction with the system and allows the user to be provided with information and content that aligns with their needs and intentions.

[0047] 15. In addition to information obtained through dialogue, the AI ​​is configured to generate and / or select various information and content (including, but not limited to, related language information, map information, and / or video) in association with acquired location information and information obtained from external cameras, making it possible to provide the user with optimal information and content. Furthermore, the AI ​​is configured to generate and / or select and display information and content related to real-world objects based on the aforementioned location information and information obtained from external cameras, so that the user can not only see objects that are actually visible in the real world, but also superimpose and view, appreciate, and experience their extended information, explanations, images, and videos (for example, it is possible to superimpose explanations and historical videos of the five-story pagoda in Asakusa). 16. Furthermore, based on the aforementioned location information and information acquired from an external camera, the AI ​​generates and / or selects information and content related to a real-world object. The system is configured to generate and display feature points and frames where such content should be displayed (including, but not limited to, generating and displaying them as wireframes), and to superimpose and position these on the real-world object. This means that information and content are not simply placed near the object (for example, the Asakusa Five-Storied Pagoda), but rather, by using feature points and frames (for example, wireframes), the AI ​​or the user can easily detect discrepancies with the real world. Additionally, the AI ​​can automatically adjust the display, or the user can adjust it to a comfortable position, enabling the precise adjustment and display of content at each detail of the object (for example, the tip, third tier, first tier, etc., of the Asakusa Five-Storied Pagoda).

[0048] 17. Since the configuration is such that the viewing device starts the calibration process by detecting the user's hand, it becomes possible to perform the calibration process intuitively and simply just by detecting the placement of the user's hand, and it becomes possible to provide an XR space in which the calibration process is performed in a low-cost and intuitive manner that matches the physical sensation. Also, since the viewing device is configured to start the calibration process by detecting that the user's hand is placed at a predetermined position, it becomes possible to perform the calibration process associated with the predetermined position and display the video, and it becomes possible to provide an XR space in which the calibration process is performed in a low-cost, functional, and intuitive manner that matches the physical sensation.

[0049] 18. Since the configuration is such that the viewing device detects that the user's hand is placed at a predetermined position corresponding to the position of the hand when placed on the sheet, it becomes possible to provide an XR space in which the calibration process is performed in a simple, low-cost, functional, and intuitive manner that matches the physical sensation. 19. Since the configuration is such that it detects that a hand is placed at a predetermined position of an object, it becomes possible to provide an XR space in which the calibration process is performed in a simple, low-cost, functional, and intuitive manner that matches the physical sensation.

[0050] 20. Since the configuration is such that it detects that a hand is placed at a predetermined position of a plate-shaped object drawn in the virtual space, it is possible to start the calibration process and the video display process in a non-contact state, and it becomes possible to provide an XR space in which the calibration process is performed in a simple, low-cost, functional, and intuitive manner without causing hygiene problems or psychological resistance to the user. 21. Since the configuration is such that grasping the hand prior to detection and detecting the opening of the hand subsequent to detection are set as preconditions for detection, it becomes possible to prevent misdetection.

[0051] 22. The system is designed to guide users through the process of placing their hands on sheets and objects by placing markers and movement paths in images and videos within the virtual space. This prevents difficulty in finding the object and operational errors, providing a comfortable XR experience that prevents users from overlooking the object and minimizes delays in operation. 23. The system is designed to display a board-shaped object in the virtual space when the input means detects the user's hand on an object. This allows users to perform calibration and image display processing without directly touching paper or objects, enabling hygienic content delivery without causing psychological resistance to the user.

[0052] 24. The current position acquisition means is configured to detect the plane of the real world visible to the user in a defined manner, and to raycast in a defined manner and project and display a plate-shaped object. This makes it possible to provide accuracy and precision through the arrangement of plate-shaped objects, and to provide an XR space that has been calibrated in a more functional and intuitive way. The defined method here includes, but is not limited to, distance measuring sensors such as LiDAR and image processing technologies for both plane detection and raycasting. The term "plate-shaped" here includes, but is not limited to, a triangular plate or an arrow-shaped plate, and also includes, but is not limited to, a flat plane or a three-dimensional shape with thickness. 25. The viewing device is configured to include an angle detection means and an angle adjustment means. This makes it possible to accurately calibrate the position information of each user and adjust the angle of the image displayed on the display means using the angle information of the hand detected by the user rotating their hand, and to provide an XR space that matches the user's experience.

[0053] 26. Because the viewing device is configured to include a finger position detection means and a plane calculation means using regression analysis, it becomes possible to acquire a plane that matches the user's hand, and it becomes possible to adjust the video display processing and calibration processing based on the orientation and angle of the plane that matches the hand, and to accurately acquire whether or not the hand is placed on an object and the situation therein. 27. Because the viewing device is configured to perform video display processing when certain external information is detected, it becomes possible to perform video display processing simply and reliably.

[0054] 28. The viewing device is configured to perform location information calibration processing when it detects certain external information, making it possible to start the calibration process in a simpler way. This allows users to easily share and enjoy an immersive and highly immersive XR space that has undergone location information calibration processing. 29. The viewing device is configured to perform time information calibration processing when it detects certain external information, making it possible to start the calibration process in a simpler way. This allows users to easily share and enjoy an immersive and highly immersive XR space that has undergone time calibration processing.

[0055] 30. The system is configured to acquire external information such as two-dimensional codes, strings of characters obtained from two-dimensional codes, external video, audio, or other matching target information. By determining whether this matches the corresponding definition information, it becomes possible to start each process, and it is also possible to store all kinds of information as definition information for starting processing. 31. The system is configured to perform video display processing based on certain information when the viewing device detects such information on the accessed web. This makes it possible to perform video display processing simply, functionally, and reliably.

[0056] 32. The viewing device is configured to perform calibration processing on the location information where the content is displayed based on certain information that is detected on the accessed web. As a result, users can easily share and enjoy an immersive and highly immersive XR space that has undergone calibration processing. 33. The viewing device is configured to perform calibration processing on the time information where the content is displayed based on certain information that is detected on the accessed web. As a result, users can easily share and enjoy an immersive and highly immersive XR space that has undergone calibration processing for time information, which is inherently prone to discrepancies.

[0057] 34. The viewing device is configured to detect the user's hand and then to detect that the user's hand is placed in a predetermined position corresponding to the position of the hand when placed on a sheet or object. This makes it easier to detect the user's hand in its proper position than detecting a hand arbitrarily placed in mid-air without the guidance of a sheet or object, enabling far more reliable and accurate calibration of the user's position information. 35. The viewing device is configured to place image elements on the image of the detected user's hand. This makes it possible to start content display processing based on the user's action on the image elements.

[0058] 36. The viewing device is configured to place image elements on the video of one of the user's hands that it has detected, and then detect whether the other hand is superimposed on it, making it possible to start content display processing in response to the intuitive movement of the user's hand. 37. The viewing device is configured to include tracking means, making it possible to perform processing to track the user's hand in the video, and making it easy to trigger various processes even when the user's hand moves.

[0059] 38. The image selection means is configured to match information read from an image with definition information or to access information on the web via a URL, so that any processing, such as content display processing, can be performed simply by appropriately changing the image in which the information can be recognized. 39. The system is configured to perform time information calibration when trigger information is received, so that calibration processing can be started by the user's own manipulation of image elements superimposed on their hand, or by operations performed by staff other than the user.

[0060] 40. The system is configured to start displaying content when trigger information is received, making it possible to initiate calibration processing through operations such as those performed by the user on the superimposed image elements on their hand, or by operations performed by staff other than the user. 41. The viewing device is configured to either hold or acquire time-related information from an external source, making it possible to perform time information calibration processing with accurate time.

[0061] The in-spatial calibration system according to the present invention will be described in detail below based on the embodiments shown in the drawings. Figure 1 is a schematic diagram of the in-spatial calibration system according to the present invention, and Figure 2 is a diagram showing the controller holder and controller dock. Figure 3 is a schematic diagram showing an embodiment of the in-spatial calibration system, and Figure 4 is a schematic diagram showing another embodiment of the in-spatial calibration system. Figure 5 is a plan view of the sheet, and Figure 6 is a diagram showing the generation and display of feature points and frames of an image, etc.

[0062] The video space calibration system 1 according to the present invention, as shown in Figure 1, consists of a viewing device 100 and a current information acquisition device 200, and is a video space calibration system that enables the synchronous provision of an immersive and highly immersive XR space and the sharing of emotional experiences by calibrating the location information of each user, the location information where content is displayed, and the time information in any space composed of video 10 provided in XR (cross reality). The video space calibration system 2 according to the present invention can also be configured to consist of multiple viewing devices.

[0063] In this invention, Extended Reality (XR) refers to a technology that creates and provides new experiences to users by constructing a virtual world and / or fusing it with reality. This concept includes VR (Virtual Reality), AR (Augmented Reality), MR (Mixed Reality), and SR (Substitutional Reality).

[0064] The viewing device 100 is a device consisting of one or more components worn by the user, which plays and displays XR-related images 10, enabling the user to view them. In this embodiment, the viewing device 100 mainly consists of a goggle-type projector, but is not limited to this. Any projector that can be worn by the user and that can display the real world and the image superimposed on it can be appropriately selected and used. For example, a retinal projection system that directly forms and projects an image onto the retina can be used.

[0065] In this embodiment, the viewing device 100 is equipped with a storage means 170 consisting of a computing device (not shown) and an arbitrary storage medium, and the video 10 and the like are stored in the storage means 170. The video 10 displayed on the viewing device 100 is an XR-related video, and may be configured to display a video in a virtual space, or to overlay the video onto the real world in front of the user. The video 10 may include 3DCG content, the whole of a video such as live-action footage, a part of a video, or assets that make up a video.

[0066] As shown in Figure 1, the viewing device 100 can be configured to connect to a computer such as a server or to the cloud via the internet. In this embodiment, the video 10 is managed, calculated, and stored by the viewing device 100, but it is also possible to configure the system so that a computer or cloud is equipped with a computing device (not shown) and storage means (not shown) consisting of any storage medium, the video 10 is stored in the storage means, and each viewing device 100 acquires the video 10 via the network. Furthermore, it is also possible to configure the system so that a computer such as a server or to the cloud is used for storing and calculating location information, etc.

[0067] The current information acquisition device 200 is a device that acquires user-defined information 40 for each user. In this embodiment, as shown in Figure 1, the user is wearing a viewing device 100, and the current information acquisition device 200 shown in Figure 2 acquires each user's defined information 40. This acquired information is then subjected to calibration processing by the calibration means 130, which will be described later. In this embodiment, "acquisition" is used to mean temporarily holding each user's defined information 40, indicating a system state in which the current information acquisition device 200 has grasped that a certain amount of information exists. Furthermore, the calibrated position information in this embodiment refers to the user's origin coordinates and orientation, but is not limited to these. In addition, the coordinates of the calibrated position information can be appropriately selected and used, such as the position of the user's hands, head, or head-mounted display.

[0068] In this embodiment, the current information acquisition device 200 is a component independent of the viewing device 100 and is configured to be installed in any location, whether indoors or outdoors. One or more users wearing the viewing device 100 transmit the definition information 40 via the current information acquisition device 200.

[0069] Here, the defining information 40 is information used when selecting whether or not to perform calibration processing, the timing of such processing, and the video to be displayed. In this case, it is determined by the fixing conditions such as the installation position and angle when the controller holder 210 of the information acquisition device 200 is fixed to the controller dock 220, and the time of installation. According to this information, the calibration means 130 performs calibration processing of the user's location information and time information in the virtual space, and the display means 120 changes or selects video content according to the information. These fixing conditions will be described later.

[0070] As shown in Figure 1, the viewing device 100 in this embodiment is configured to include an information holding means 110, a display means 120, and a calibration means 130. The current user definition information 40 acquired by the current information acquisition device 200 is calibrated by the calibration means 130, which will be described later. Each piece of content, consisting of the calibrated location information, time information, and video 10, is held by the information holding means 110.

[0071] Furthermore, the space formed by the video 10 may contain either one user or multiple users. The information holding means 110 can be configured to hold the calibrated location and time information of other users for each user. This makes it possible to hold the location information of each user, as well as the location and time information where the content is displayed.

[0072] In this embodiment, the information holding means 110 consists of software that is read and processed by a computing device incorporated in the viewing device 100 from the storage means 170. However, it is not limited to this configuration. For example, it is also possible to set up an external computer or cloud, and have the information holding means 110 equipped on the computer or on the cloud perform calculations and store information related to each content, including the location information of each user after calibration, the location information where the content is displayed, time information, and the video 10. It is also possible to migrate only a part of the functions to the cloud or a web application and then perform the processing there.

[0073] The display means 120 is a means for displaying the video 10 in a way that is visible to the user, based on the definition information 40. In this embodiment, the video 10 stored in or temporarily acquired by the storage means 170, etc. is output and displayed on the video display area of ​​the goggle-shaped viewing device 100 (HMD: head-mounted display). This video display area can be configured to simultaneously display external real-world images acquired by the input means 150, which will be described later, so that images consisting of CG, etc., are superimposed on the real world. The video 10 is processed to be in the optimal position based on the calibrated position information, and the display means 120 displays it on the video display area of ​​the HMD.

[0074] The display means 120 controls the display of the video 10 based on the user's position information and time information calibrated by the calibration means 130 based on the definition information 40. If the video 10 is a completely virtual space, it calculates and displays the standing position and orientation of each user, and if the video 10 is a partial video superimposed on the real world, it calculates and displays the location and positional relationship where the video will be superimposed. This process includes, but is not limited to, a stencil test, which is a process that determines whether or not to draw objects on pixels using data referenced and edited to achieve occlusion. As a result, by causing occlusion, that is, the occlusion of CG by real objects, it is possible to display a more realistic video 10 that conveys a sense of depth, while correctly reflecting the drawing order of objects in front and behind.

[0075] In this embodiment, the display means 120 consists of software that is read and processed by a computing device incorporated in the viewing device 100 from the storage means 170. However, it is not limited to this configuration. For example, it is also possible to configure the display means 120, which is equipped on an external computer or cloud, or on a web application, to perform calculations and output the video 10 by using an external computer or cloud or web application. It is also possible to configure the display means 120 to perform calculations and output the video 10 by migrating only a part of the functions to the cloud or web application and then performing each processing there.

[0076] The calibration means 130 is a means for calibrating location information and time information. In this embodiment, based on the definition information 40 acquired by the current information acquisition device 200, the location information of each user, the location information and time information where content is displayed in the space composed of the video 10 are calibrated, and the information holding means 110 holds the calibrated location information and time information. This configuration makes it possible to calibrate and synchronize the location information of each user, the location information and time information where content is displayed, and the location information of users participating in the XR space, as well as the location information and time information where content is displayed, in a simple, functional, rapid, and accurate manner. In other words, it becomes possible to define location in the XR space, which is different from the real world (defining the XR world), and to synchronize time, allowing users to share an immersive and highly real-world XR space that matches their intuition.

[0077] In this embodiment, the position information calibration process refers to the process of aligning the user's real-world location with their location in the virtual space, and includes initial setup processing at the initial stage of starting content, as well as one or more calibration processes during operation. Position information refers to position information that has been accurately calibrated as a relative position starting from the origin by the in-video space calibration system (controller dock or palm calibration) according to the present invention. However, if GNSS or the like can also be used, it refers to position information that has been accurately calibrated as a relative position starting from the origin by the in-video space calibration system (controller dock or palm calibration) according to the present invention, based on absolute position information (and viewing direction inferred by a compass (magnetic sensor) and MEMS gyro) roughly acquired by GNSS or the like. The in-video space calibration system (controller dock or palm calibration) according to the present invention sets an origin in the virtual space and performs processing to appropriately display the user's relative position from now on. Furthermore, the time information calibration process performed in this embodiment refers to the process of aligning the trigger time for each user with the progression of time within the virtual space, and is primarily a process to ensure the playback and progression timing of video related to the virtual space is consistent. This time information calibration process includes initial setup processing at the initial stage of starting content, and one or more calibration processes during operation. It also performs processes such as displaying video that corresponds to the user's actual time.

[0078] The calibration process by the calibration means 130 will now be explained. First, the calibration means 130 acquires the current definition information 40 of each user via the current information acquisition device 200. Next, based on the definition information 40, it performs a calibration process for positional and temporal information in the space composed of the video 10. That is, based on the information specified by the definition information 40 (fixing conditions such as the installation position and installation angle when the controller holder 210 is fixed to the controller dock 220, and the installation time), it performs a calibration process for the user's positional and temporal information in the virtual space according to the fixing conditions. After that, based on the calibrated positional and temporal information, the display means 120 changes and selects video content according to the fixing conditions and installation time, and then selects the video to be viewed based on the positional relationship with the user in the video 10 and displays it in the video display area of ​​the viewing device 100.

[0079] Subsequently, the information holding means 110 holds the position information of each user after calibration within the space composed of the video 10, as well as the position and time information where the content is displayed. This enables the display means 120 to display the video 10 synchronously and accurately, and through synchronous and accurate calibration, it becomes possible to provide an immersive and highly realistic XR space that matches the user's experience, allowing users to enjoy XR videos while sharing an emotional experience.

[0080] In this embodiment, the calibration means 130 consists of software that is read from the storage means 170 and processed by a computing device incorporated in the viewing device 100. However, it is not limited to this configuration. For example, it is also possible to configure the calibration means 130 to perform calculations and calibration processing by providing an external computer or cloud, or by using a web application. It is also possible to migrate only a part of the functions to the cloud or web application and then perform each processing there.

[0081] Next, the current information acquisition device 200 will be described. In this embodiment, the current information acquisition device 200 consists of a controller holder 210 and a controller dock 220, as shown in Figure 2. The controller holder 210 is a component attached to the controller and has the function of holding the controller, and is a component for fitting into the controller dock 220, which will be described later, when using the video space calibration system 1. One or more controller holders 210 are provided, and each has a different shape, and in this embodiment, they are shaped like uppercase letters of the alphabet, but are not limited to this shape.

[0082] The controller dock 220 is a component for fitting and fixing the controller holder 210, and the configuration includes one or more controller docks 220 (a number corresponding to the controller holder 210) that have a shape corresponding to the controller holder 210.

[0083] When a user uses the in-video space calibration system 1, they first attach and secure the controller holder 210 to the controller dock 220. This action triggers the current information acquisition device 200 to acquire the user's current definition information 40. Based on the acquired information, the calibration means 130 performs calibration processing on the position information of each user, the position information and time information on which content is displayed within the space composed of the video 10, as described above. In this embodiment, the visible position, angle, and timing of movement of objects and other elements in the video are calibrated to be optimal so that there is no discrepancy between the video 10 that the user can view through the viewing device 100 and the real video. This configuration makes it possible to start the calibration processing of position information and time information quickly, functionally, and reliably with a simple structure.

[0084] Currently, the information acquisition device 200 can be configured to include a docking station 230, as shown in Figure 2. The docking station 230 is a component that allows the controller docks 220 to be installed at predetermined intervals, is portable and easy to handle, allows the controller docks 220 to be installed together in one place, enables appropriate placement, and is easily visible to the user, making it easy for the user to find the location of the docks. As a result, the controller holders 210 can be installed at predetermined intervals, making it possible to provide an information acquisition device 200 that can perform calibration processing simply, quickly, and functionally.

[0085] The video 10 viewable by the display means 120 of the viewing device 100 is not limited to a single video but consists of multiple videos of any kind, and is configured to allow selection and display of a wide variety of videos 10. In this embodiment, the video 10 includes at least 3DCG content, the entirety of live-action footage, a part of a video, and assets that constitute a video, and multiple videos 10 selected from among these can be viewed on the viewing device 100.

[0086] As shown in Figure 1, the viewing device 100 is configured to include a video selection means 140. The video selection means 140 selects and displays one or more videos from a plurality of videos 10. In this embodiment, after selecting a video 10, the video selection means 140 performs a process to adjust the display of the video 10 according to the positional relationship as seen from the user's position. The viewing device 100 is capable of viewing the real world, and as shown in Figure 3, it is possible to superimpose the various videos selected by the video selection means 140 onto the real world in front of the user's eyes. At this time, the video 10 may be a video relating to a complete virtual reality world, or it may be configured to superimpose the video 10 onto the entire real world. Alternatively, it may be configured to display the selected video 10 on a part of the real world, or it may be configured to superimpose various video assets onto a large number of objects visible in the real world.

[0087] The display means 120 of the viewing device 100 displays each piece of content, which is composed of the video 10 held by the information holding means 110, in the video display area of ​​the viewing device 100. In this embodiment, the content consists of, for example, 3DCG content, the whole of a video such as live-action footage, a part of a video, or assets that make up a video. Specifically, possible configurations include, for example, a configuration in which a space appears on the floor and outer space or a space probe is displayed, a configuration in which a box becomes transparent and dancers sing and dance, a configuration in which tourist attractions appear in the real world, a configuration in which various advertisements are superimposed on the real world, a configuration in which lively illuminations are superimposed and displayed, a configuration in which objects are displayed on an empty table and shared by each user, and other configurations such as image processing that breaks down walls displayed as the real world, and configurations that perform effects such as making the floor or box displayed as the real world transparent.

[0088] For example, in a configuration where an object is displayed as an image 10 on a table, the positional relationship in which the object is seen will differ because each user is in a different position. Furthermore, if all users cannot view the object at the same time, a problem arises in that a synchronized experience cannot be achieved in the XR space. Since the object display process is performed by each viewing device 100, this problem occurs if the positional and temporal information is not calibrated. By calibrating the positional and temporal information using the calibration means 130, all users participating in the XR space can view the same object synchronously from any positional relationship, making it possible to share and enjoy the excitement in an immersive and highly real-world XR space that matches the physical sensations.

[0089] In this embodiment, the video selection means 140 is configured to automatically select a video 10 according to the fixing conditions when the controller holder 210 is attached and fixed to the controller dock 220, and the time of installation. In this embodiment, it automatically selects and displays a viewable video or the video 10 itself based on the positional relationship in which the video 10 is displayed. For example, the controller dock 220 is provided in five types, labeled "A," "B," "C," "D," and "E," and the user can select one of them. The video selection means 140 selects a video 10 from information relating to the combination of the selected controller holder 210 and controller dock 220. With this configuration, it becomes possible to select a positional relationship in which videos are displayed according to the number of controller holders 210 and controller docks 220, making it possible for the user to select a video in an easy-to-understand and simple way.

[0090] Furthermore, the video selection means 140 is configured to automatically select the video 10 itself or a viewable video based on the positional relationship in which the video 10 is displayed, according to the combination of the angle at which the controller holder 210 is attached and fixed to the controller dock 220 and the on / off status of one or more buttons provided on the controller holder 210. For example, the controller dock 220 can be set to allow selection of three installation angles for the controller holder 210: vertical, diagonal, and planar, with three options for 0 degrees, 90 degrees left, and 90 degrees right. Adding the on / off status of the buttons (or more if there are multiple buttons) makes it possible to select a total of 18 different videos. In other words, it is possible to increase the number of videos that can be selected in a way that is easy for the user to understand.

[0091] In this embodiment, the video selection means 140 consists of software that is read and processed by a computing device incorporated in the viewing device 100 from the storage means 170. However, it is not limited to this configuration. For example, it is also possible to configure the video selection means 140 to be installed on an external computer or cloud, or to be located on a web application, and to perform calculation processing to select a video. It is also possible to migrate only a part of the functions to a cloud or web application and then perform each processing there.

[0092] The video space calibration system 1 according to the present invention can be configured to incorporate artificial intelligence (AI). The video space calibration system 1 according to the present invention can be configured to have a user assistance function that provides information to the user in a method including dialogue, by having the installed AI select relevant information based on location information, information acquired based on dialogue, and information acquired from an external camera, and by generating and / or selecting various information and content (various information and content include, but are not limited to, relevant language information, map information, and / or video). With this configuration, for user assistance, information is acquired from the user in a dialogue format, the AI ​​selects relevant information, and various information and content (including, but not limited to, relevant language information, map information, and / or video) are provided in a dialogue format, enabling verbal interaction with the system and making it possible to provide the user with information and content that aligns with the user's needs and intentions (for example, if the content is about the Asakusa area, information related to the history, tourism, and safety of Asakusa is provided, and maps, videos, and other content are displayed).

[0093] Furthermore, by combining information obtained through dialogue with location information acquired by the AI ​​and information acquired from external cameras, the system is configured to generate, and / or select, and display various types of information and content (including, but not limited to, related language information, map information, and / or video), making it possible to provide users with optimal information and content.

[0094] Furthermore, the system is configured to include an information content display function in which the AI ​​generates and / or selects and displays information and content related to real-world objects based on location information and information acquired from external cameras. As a result, users can not only visually perceive objects that are actually visible in the real world, but also superimpose and view, appreciate, and experience their extended information, explanations, images, and videos (for example, it becomes possible to overlay explanations and historical videos of the five-story pagoda in Asakusa).

[0095] Furthermore, based on location information and information acquired from external cameras, when the AI ​​generates and / or selects information and content related to real-world objects, it generates and displays feature points and frames where such content should be displayed (including, but not limited to, generating and displaying them as wireframes), as shown in Figure 6. This configuration includes a function to superimpose and position these on real-world objects. As a result, information and content are not simply placed near the object (e.g., the Asakusa Five-Storied Pagoda), but the feature points and frames (e.g., wireframes) allow the AI ​​or the user to easily detect discrepancies with the real world. Additionally, the AI ​​automatically adjusts the display, or the user adjusts it to a comfortable position, making it possible to precisely adjust and display content on each detail of the object (e.g., the tip, third tier, first tier, etc., of the Asakusa Five-Storied Pagoda).

[0096] Next, another embodiment of the video space calibration system according to the present invention will be described. The video space calibration system 2 is a system for calibrating the location information of one or more users, the location information and time information of content displayed in a space composed of video provided by XR (cross reality), and as shown in Figure 4, it can be configured to consist of one or more viewing devices 100.

[0097] The viewing device 100 is a component that allows a user to wear it and view the video 10. In this embodiment, the viewing device 100 is configured to acquire the current definition information 50 of the user wearing it.

[0098] Here, the definition information 50 is information used when determining whether or not to perform calibration processing, the timing of such processing, and selecting the video to display. In this case, it is information determined according to the actions performed when the user places their hand over the sheet S or objects O and P, which will be described later. Based on this information, the calibration means 130 performs calibration processing of the user's location information and time information in the virtual space, and the display means 120 changes or selects the video content according to that information. Furthermore, the definition information 50 is not limited to this, and may be fixed information in advance. In this case, at any arbitrary timing, calibration processing of the user's location information and time information in the virtual space is performed according to that fixed information, and the video content is selected.

[0099] As shown in Figure 4, the viewing device 100 is configured to include information holding means 110 that each holds information about the location of one or more users in the space composed of the video 10, location information where content is displayed, time information, and each piece of content composed of the video 10, and display means 120 that displays the video 10 based on definition information 50.

[0100] Furthermore, as shown in Figure 4, the viewing device 100 is configured to include a calibration means 130. The calibration means 130 is a means for calibrating the user's location information, the location information and time information where content is displayed, within the space composed of the video 10. In this embodiment, based on the definition information 50, the calibration means 130 performs calculation processing to calibrate the location information of each user within the space composed of the video 10, and the location information and time information where content is displayed, and the information holding means 110 holds this information as calibrated location information and time information. This configuration makes it possible to perform calibration processing without using a separate current information acquisition device, and by calibrating each user's location information, the location information and time information where content is displayed simply, quickly, and accurately, it becomes possible to share an emotional experience in an immersive and highly realistic XR space that is more in line with the user's experience.

[0101] In this embodiment, the position information calibration process refers to the process of aligning the position, and includes the initial setup process at the initial stage of starting the content, as well as one or more calibration processes during operation. Position information refers to position information that has been accurately calibrated as a relative position starting from the origin by the in-video space calibration system (controller dock or palm calibration) according to the present invention. However, if GNSS or the like can also be used, it refers to position information that has been accurately calibrated as a relative position starting from the origin by the in-video space calibration system (controller dock or palm calibration) according to the present invention, based on absolute position information roughly acquired by GNSS or the like (and inferred by a compass (magnetic sensor) and MEMS gyro). In the in-video space calibration system (controller dock or palm calibration) according to the present invention, an origin is set in the virtual space, and processes are performed to select and display images corresponding to the relative position of each user and the orientation of the user. Furthermore, the time information calibration process performed in this embodiment refers to the process of aligning trigger information and the progression of time within each user's virtual space, and is primarily a process to ensure the synchronization of the playback and progression timing of video related to the virtual space. This time information calibration process includes initial setup processing at the initial stage of starting content, and one or more calibration processes during operation. It also performs processes such as displaying video that corresponds to the user's actual time.

[0102] The calibration process by the calibration means 130 will now be explained. First, the calibration means 130 acquires the current definition information 50 for each user. Next, based on each definition information 50, it performs a process to calibrate the position information and time information of each user in the space composed of the video 10. That is, it performs a calibration process for the user's position information and time information in the virtual space based on the information identified by the definition information 50 (information determined according to actions performed by the user, etc.). After that, based on the calibrated position information and time information, the display means 120 changes and selects video content, and adjusts and displays the video to be viewed based on the positional relationship with the user in the video 10. In this embodiment, the configuration is such that a process is performed to calibrate the visible position, angle, and timing of movement of objects O and P in the video to the optimal settings so that there is no discrepancy between the video 10 that the user can view through the viewing device 100 and the real video.

[0103] Subsequently, the information holding means 110 holds the position information of each user after calibration within the space composed of the video 10, as well as the position and time information where the content is displayed. This enables the display means 120 to display the video 10 synchronously and accurately, and through synchronous and accurate calibration, it becomes possible to provide an immersive and highly realistic XR space that matches the user's experience, allowing users to enjoy XR videos while sharing an emotional experience.

[0104] In this embodiment, the calibration means 130 consists of software that is read from the storage means 170 and processed by a computing device incorporated in the viewing device 100. However, it is not limited to this configuration. For example, it is also possible to configure the calibration means 130 to perform calculations and calibration processing by providing an external computer or cloud, or by using a web application. It is also possible to migrate only a part of the functions to the cloud or web application and then perform each processing there.

[0105] In this embodiment, the viewing device 100 can be configured to include a current position acquisition means 155, as shown in Figure 4. The current position acquisition means 155 is a component for acquiring the user's current position, and the calibration means 130 acquires the position information acquired by the current position acquisition means 155 as the current definition information 50 for each user, and then calibrates it to position information and time information in the space composed of the video 10.

[0106] The system obtains the current location of the user based on the location information acquired by the current location acquisition means 155. This makes it possible to select a video 10 that is specific to the acquired location. As a result, it becomes possible to obtain the current definition information 50 for each user simply by acquiring the user's location information, and it becomes possible to calibrate the location and time information of each user related to the video 10, as well as to select a timely video 10 as the user moves.

[0107] In this embodiment, the current location acquisition means 155 can be configured to acquire the user's current location using one or more of the following: GNSS, VPS, beacons, location markers, and image markers. This configuration makes it possible to acquire the user's current location (an approximate absolute location based on XYZ coordinates) by any means.

[0108] Furthermore, in this embodiment, the current position acquisition means 155 can be configured to acquire directional information related to the user's current position using a compass (magnetic sensor). This configuration makes it possible to acquire directional information that allows the user to infer their viewing direction.

[0109] Furthermore, in this embodiment, the current position acquisition means 155 can be configured to acquire acceleration and angular velocity information related to the user's current position using a MEMS gyroscope. This configuration makes it possible to acquire acceleration and angular velocity information that allows the user to infer their viewing direction.

[0110] The video space calibration system 2 according to the present invention can be configured to incorporate artificial intelligence (AI). The video space calibration system 2 is configured to have a function to provide information to the user in a way that includes dialogue, by having the installed AI select relevant information based on acquired location information, information acquired based on dialogue, and information acquired from an external camera, and by generating and / or selecting various information and content (including, but not limited to, related language information, map information, and / or video). This configuration enables verbal interaction with the system and makes it possible to provide the user with information and content that aligns with the user's needs and intentions.

[0111] Furthermore, the video spatial calibration system 2 according to the present invention can be configured to include a user assistance function that provides information to the user in a method including dialogue, by having the AI ​​select relevant information based on acquired location information, information acquired based on dialogue, and information acquired from an external camera, and by generating and / or selecting various information and content (including, but not limited to, relevant language information, map information, and / or video). With this configuration, for user assistance, information is acquired from the user in a dialogue format, the AI ​​selects relevant information, and various information and content (including, but not limited to, relevant language information, map information, and / or video) are provided in a dialogue format, enabling verbal interaction with the system and making it possible to provide the user with information and content that aligns with the user's needs and intentions (for example, if the content is about the Asakusa area, information related to the history, tourism, and safety of Asakusa is provided, and maps, videos, and other content are displayed).

[0112] Furthermore, the system is configured to include an information content display function that generates and / or selects and displays various information and content (including, but not limited to, related language information, map information, and / or video) in association with location information acquired by the AI ​​and information acquired from external cameras, in addition to information obtained through dialogue. This makes it possible to provide users with optimal information and content. Moreover, the system is configured to include an information content display function that generates and / or selects and displays content related to real-world objects based on the aforementioned location information and information acquired from external cameras. This allows users not only to visually perceive objects that are actually visible in the real world, but also to superimpose and view, appreciate, and experience their extended information, explanations, images, and videos (for example, it becomes possible to superimpose explanations and historical videos of the five-story pagoda in Asakusa).

[0113] Furthermore, based on the location information and information acquired from an external camera, when the AI ​​generates and / or selects content related to a real-world object, it generates and displays feature points and frames where the content should be displayed (including, but not limited to, generating and displaying them as wireframes), as shown in Figure 6. This configuration includes an information content display function that superimposes and places these on the real-world object. As a result, information and content are not simply placed near the object (for example, the Asakusa Five-Storied Pagoda), but feature points and frames (for example, wireframes) are used to allow the AI ​​or the user to easily detect discrepancies with the real world. Additionally, the AI ​​automatically adjusts the display, or the user adjusts it to a comfortable position, making it possible to precisely adjust and display content on each detail of the object (for example, the tip, third tier, first tier, etc., of the Asakusa Five-Storied Pagoda).

[0114] In this embodiment, the viewing device 100 is configured to include an input means 150, as shown in Figures 1 and 4. The input means 150 is a means for acquiring external information 20, such as video, audio, and various data, and includes, but is not limited to, cameras, microphones, various sensors, communication functions for acquiring data wirelessly or via wired connections, and ports.

[0115] External information 20 is any information that the input means 150 can acquire, and in this embodiment it consists of, but is not limited to, information such as a two-dimensional code or a string obtained therefrom, external video, audio, or other matching target information. Furthermore, the audio as external information 20 may be the audio data itself, or it may be acquired in a form in which the audio data has been converted into a string and then converted into text.

[0116] The viewing device 100 is configured to detect the user's hand based on external information 20 acquired by the input means 150. After detecting the hand from the external information 20 acquired by the input means 150, the viewing device 100 detects that the user's hand has been placed. Based on this, the viewing device 100 acquires the user's definition information 50, and the calibration means 130 performs calibration processing on the position and time information of each user in the space composed of the video 10 based on the detected information. This makes it possible to perform calibration processing intuitively and simply by detecting the placement of the user's hand, and makes it possible to provide an XR space that has been calibrated in a low cost and in a way that matches the user's experience.

[0117] In this embodiment, the viewing device 100 can be configured to detect when a user's hand is placed in a predetermined location, acquire this information as user definition information 50, and then have the calibration means 130 perform calibration processing on the user's position information within the space composed of the video 10. For example, when a user places their hand in a predetermined location where an illustration visible to the user is displayed, the calibration means 130 starts the calibration process and the video 10 is displayed. This configuration makes it possible to display the video after performing calibration processing linked to the predetermined location, making it possible to provide an XR space in a low-cost, functional, and sensory-responsive manner.

[0118] In this embodiment, as shown in Figures 4 and 5, it is possible to provide a sheet S for use in initiating the calibration process by the calibration means 130. The sheet S may be configured to have markings for placing the user's hand, such as a handprint, or it may be configured to have a star mark to encourage the user to place their hand accurately. Alternatively, if the user's hand can be positioned similarly, it is also possible to detect the user's hand without using the guide sheet S and make corrections as needed.

[0119] The viewing device 100 detects a hand from external information 20 acquired by the input means 150, and then detects that the user's hand is placed on a predetermined position corresponding to the position of the hand when placed on the sheet S. Based on the detected information, the calibration means 130 performs calibration processing on the positional and temporal information of each user within the space composed of the video 10. This makes it possible to provide an XR space that is functionally and intuitively calibrated with simple equipment.

[0120] In this embodiment, the sheet S is made of paper or the like, and for example, has a design such as a handprint or a star printed on it. However, it is not limited to this, and it may also have a three-dimensional shape such as a recess for placing a hand, or a configuration in which the sheet S is displayed on a screen and a design such as a handprint or a star is displayed on the sheet S, or other configurations can be selected and used. This configuration makes it possible to provide the in-video space calibration system 1 at a low cost.

[0121] The viewing device 100 can be configured to detect a hand from external information 20 acquired by the input means 150, and then detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when placed on an actual object O. Based on the detected information, the calibration means 130 performs calibration processing of the position information and time information of each user in the space composed of the video 10. Subsequently, the display means 120 changes and selects video content based on the calibrated position information and time information, and adjusts and displays the video to be viewed based on the positional relationship with the user in the video 10.

[0122] In this embodiment, object O consists of a real-world object whose position and shape are predetermined. For example, it could be a pillar or bronze statue erected at a certain location, or a piece of paper placed arbitrarily. It may be an object that already exists, or it may be an object erected at an arbitrary location. This configuration makes it possible to provide an XR space that has undergone calibration processing in a simple, low-cost, functional, and intuitive way that matches the user's experience.

[0123] Furthermore, the viewing device 100 can be configured to detect a hand from external information 20 acquired by the input means 150, and then detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when placed on a plate-shaped object P drawn in the virtual space. Based on the detected information, the calibration means 130 performs calibration processing of the position information and time information of each user in the space composed of the video 10. Subsequently, the display means 120 changes and selects video content based on the calibrated position information and time information, and adjusts and displays the video to be viewed based on the positional relationship with the user in the video 10.

[0124] In this embodiment, object P is an image or video superimposed on a realistic image such as scenery, patterns, or objects viewed by the viewing device 100. For example, a triangular board or an arrow board (a board includes, but is not limited to, a flat surface with no thickness or a three-dimensional shape with thickness) can be displayed as if it were fixed to a specific object. With this configuration, it is not necessary to pre-specify the real objects to be used for calibration processing, and it becomes possible to place an object that enables the start of calibration processing at any location. This allows calibration processing to be started in a non-contact state, making it possible to provide an XR space in which calibration processing is performed in a simple, low-cost, functional, and intuitive way that matches the user's experience, without causing psychological resistance to the user.

[0125] In this embodiment, when the viewing device 100 detects that the user's hand has been placed on the sheet S or object O / P, it is possible to make it a prerequisite for detection that the user first clenches their hand and then opens it. This hand movement of the user is acquired as external information 20 by the input means 150 and analyzed. This configuration makes it possible to prevent false detection of movements for calibration processing, and to provide a more accurate and comfortable XR space.

[0126] Furthermore, the viewing device 100 can be configured to display the procedure for placing a hand on the sheet S or a real-world object O. This procedure may involve placing images that serve as markers or movement paths within the virtual space, or placing content such as videos, and can be configured to show one or more points. This configuration makes it possible to guide the user to the sheet S or object O, making it easier for the user to find it, and providing a comfortable XR space with reduced delays in operation.

[0127] Furthermore, the viewing device 100 can be configured to display a board-shaped object P to be drawn in the virtual space when the input means 150 detects the user's hand on a real-world object O. This configuration makes it possible for the user to perform calibration and image display processing without directly touching paper or objects, thus enabling the provision of content hygienically and without causing the user any psychological resistance.

[0128] Furthermore, the viewing device 100 can be configured such that the current position acquisition means 155 detects the plane of the real world visible to the user in a defined manner, and then raycasts in a defined manner to display a plate-shaped object P. The defined manner here includes, but is not limited to, distance measuring sensors such as LiDAR and image processing technologies for both plane detection and raycasting. LiDAR is a technology that measures the time and wavelength of laser light emitted from an object until it hits an object, reflects back, and returns, in order to calculate the distance to the object or detect the shape of the object. Here, "plate-shaped" includes, but is not limited to, triangular plates and arrow-shaped plates, and also includes, but is not limited to, flat surfaces or three-dimensional shapes with thickness. By using distance measuring sensors such as LiDAR and image processing technologies, it becomes possible to provide accuracy and precision in the arrangement of plate-shaped objects P, and to provide an XR space that has undergone calibration processing in a more functional and intuitive way.

[0129] Furthermore, as shown in Figure 4, the viewing device 100 can be configured to include an angle detection means 180 and an angle adjustment means 190. The angle detection means 180 is a means for detecting the angle of the user's hand when the input means 150 detects the user's hand. The user's hand may be detected at any location, or it may be configured to detect whether it is placed in a predetermined position. For example, in this embodiment, after detecting that the user's hand is placed on a marker for placing the user's hand displayed on a sheet S or object O / P, if the user rotates their hand left or right to change the angle of their hand, the angle detection means 180 detects in which direction and by how many degrees the hand has rotated.

[0130] The angle adjustment means 190 is a means for adjusting the display angle of each content, which consists of the video 10 displayed by the display means 120. The video 10 displayed by the display means 120 is configured to be viewed in an optimal state from the user's position through calibration processing by the calibration means 130, but fine adjustment of the display angle may be necessary. The angle adjustment means 190 makes it possible to perform more optimal video display processing to meet such needs.

[0131] In this embodiment, the angle adjustment means 190 adjusts the angle of the image 10 displayed by the display means 120 according to the angle of the user's hand detected by the angle detection means 180 during or after the calibration process of each user's position information by the calibration means 130. The rotation angle of the hand and the angle of the image 10 and their respective rotation axes may be matched, or they may be different angles and rotation axes. Alternatively, the angle and rotation distance obtained by multiplying the angle and rotation distance of the user's hand by a coefficient may be used as the angle and rotation distance of the image 10 after the change, so that the optimal angle adjustment does not deviate from the user's perception. With this configuration, it is possible to perform optimal image display processing that matches the user's perception.

[0132] Furthermore, the viewing device 100 can be configured to include a finger position detection means 185 and a plane detection means 186, as shown in Figure 4. The finger position detection means 185 is a means for detecting the position of each fingertip of the user's hand, and the plane calculation means 186 by regression analysis is a means for estimating a plane from the position information of each of the user's fingers detected by the finger position detection means 185.

[0133] In this embodiment, the finger position detection means 185 detects the position of each fingertip of the user's hand, and then the plane detection means 186 estimates a plane from the position information of each finger using a plane regression method. This makes it possible to obtain a plane that matches the user's hand. With this configuration, it becomes possible to obtain a plane that matches the user's hand, and to perform video display processing and calibration processing based on the orientation and angle of the plane that matches the hand, as well as to accurately determine whether or not the hand is placed on an object.

[0134] The viewing device 100 is configured to perform video display processing when it detects information in the external information 20 that matches the definition information 30. The definition information 30 is information that has been defined in advance and stored in the storage means 170 of the viewing device 100, and consists of information such as a two-dimensional code or a string obtained therefrom, external video, audio, or other matching target information. The definition information 30 corresponds to the external information 20, and the device is configured to start some kind of processing when it detects that the acquired external information 20 matches the definition information 30.

[0135] The viewing device 100 is configured to perform display processing of one or more videos from the entire 3DCG content, live-action footage, or any part of the video, or from the assets that constitute the video, when the external information 20 acquired by the input means 150 contains information that matches the definition information 30 which has been determined to be applicable as described above.

[0136] In this embodiment, the viewing device 100 adjusts to display the image 10 according to the positional relationship as seen from the user's position. It is also possible to adjust the display so that the image is superimposed on the displayed image of the real world. Furthermore, it is possible to configure the device to display a preview of the image 10 in a certain area of ​​the display screen. Examples of image 10 displays include, for example, an effect in which a wall breaks down and a dinosaur appears near the location where the display target is projected, an effect in which the area below the floor becomes outer space and a falcon appears, an effect in which a wooden box becomes transparent and a girl dancing and singing appears, and an effect in which one boards a vehicle synchronized with an existing object. Also, for example, in the case of displaying Kaminarimon, it is possible to configure the device so that when the user passes through Kaminarimon, the image of Kaminarimon moves relative to the user and is displayed in a different positional relationship depending on the user's location.

[0137] For example, if the acquired external information 20 is a video in the form of a two-dimensional code, the viewing device 100 reads the two-dimensional code, decodes it, and obtains the content indicated by the two-dimensional code. The two-dimensional code can be converted into text information, for example, if a specific string is predefined and stored as definition information 30, the viewing device 100 compares the converted content with the string defined as definition information 30. If the two match, the aforementioned process is executed.

[0138] External information 20 may also be strings, images, audio, or audio data converted to text, or a combination of these. The system can be configured to execute processing when this information (or a combination thereof) matches the definition information 30. It is also possible to define a two-dimensional code as the definition information 30 and directly compare it with the external information 20 consisting of the two-dimensional code.

[0139] In another embodiment, the viewing device 100 can be configured to perform calibration processing on the location information (e.g., origin coordinates and orientation) where the content is displayed when the input means 150 detects information in the external information 20 acquired by the user that matches the definition information 30 that has been determined to be applicable, in order to display content corresponding to the user's current definition information 50.

[0140] For example, if a specific string is predefined and stored as definition information 30, the viewing device 100 reads the two-dimensional code, obtains the content indicated by the two-dimensional code, and then compares the content of the two-dimensional code with the string defined as definition information 30. If the two match, the calibration process for the location information (origin coordinates and orientation) where the content is displayed is initiated. This configuration makes it possible to start the calibration process in a simpler, faster, and more functional way, allowing users to easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration process has been completed.

[0141] In another embodiment, the viewing device 100 can be configured to perform a time information calibration process when it detects information in the external information 20 acquired by the input means 150 that matches the defined information 30 that has been determined to be applicable.

[0142] For example, if a specific string is predefined and stored as definition information 30, the viewing device 100 reads the two-dimensional code, obtains the content indicated by the two-dimensional code, and then compares the content of the two-dimensional code with the string defined as definition information 30. If the two match, the time information calibration process is started. This configuration makes it possible to start the calibration process in a simpler way, allowing users to easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration process has been completed.

[0143] In another embodiment, the viewing device 100 can be configured to access the web based on URL information obtained from external information 20. In this configuration, certain information exists on the web accessed based on the URL information. Based on the information present on the accessed web, the device is configured to display one or more of the following: 3DCG content, live-action footage, the entirety of the video, a part of the video, or assets that constitute the video. The external information 20 used to obtain the URL information in this example consists of a two-dimensional code or text information of the URL, but is not limited to this configuration.

[0144] For example, when the viewing device 100 detects a two-dimensional code as external information 20, it decodes the two-dimensional code. After obtaining the URL text information as the content of the two-dimensional code, it accesses the web based on that information and displays the obtained content in a certain area of ​​the display means 120 of the viewing device 100. As for the display method, in this embodiment, for example, a configuration such as preview display or a configuration that uses the functions of a browser can be considered, but it is not limited to this configuration, and it is of course possible to select and use an appropriate display method, such as a configuration that uses the functions of other web applications.

[0145] Alternatively, the system may be configured to obtain the text information of a URL as the content of a two-dimensional code, and then launch a web application from that URL string. Another configuration may be used to obtain content consisting of 3DCG content, live-action footage, or other video content from the web that can be displayed on the viewing device 100, or content consisting of a part of a video or assets that constitute a video.

[0146] This configuration makes it possible to display a variety of images 10 based on content information obtained by accessing a web superimposed on the real world. For example, a configuration in which a space appears on the floor and outer space or a space probe is displayed, a configuration in which a box becomes transparent and a dancer sings and dances, a configuration in which tourist attractions appear in the real world, a configuration in which various advertisements are superimposed on the real world, a configuration in which lively illuminations are superimposed and displayed, a configuration in which an object is displayed on an empty table and shared by each user, and other configurations that perform image processing such as breaking down a wall displayed as the real world, or configurations that perform effects such as making the floor or box displayed as the real world transparent.

[0147] Furthermore, in another embodiment, the viewing device 100 can be configured to perform calibration processing on the location information (e.g., origin coordinates and orientation) where the content is displayed, based on the information detected on the web accessed based on URL information obtained from external information 20. For example, the viewing device 100 accesses the web based on the acquired URL information and displays it. Certain information exists on the web, and the calibration processing is initiated based on this information. With this configuration, users can easily share and enjoy an immersive and highly realistic XR space that matches the experience after the calibration processing.

[0148] Furthermore, in another embodiment, the viewing device 100 can be configured to perform time information calibration processing when it detects information existing on the web accessed based on URL information obtained from external information 20. For example, the viewing device 100 accesses and displays the web based on the acquired URL information. In this configuration, certain information exists on the web, and based on this, the time information calibration processing is initiated. With this configuration, users can easily share and enjoy an immersive and highly realistic XR space that matches the sense of presence after the calibration processing has been performed.

[0149] The viewing device 100 detects the user's hand based on external information 20 acquired by the input means 150. Subsequently, it detects when the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet S or object O / P. The input means 150 acquires video as external information 20. When the user's hand is detected in the video, it detects that the user's hand is placed in a predetermined position corresponding to the position of the hand when it is placed on the sheet S or object O / P. This makes it possible to detect the user's hand in its proper position, and when the hand is placed on the sheet S or object O / P, it becomes possible to start a calibration process for accurate user position information as a trigger.

[0150] Furthermore, the viewing device 100 can be configured to superimpose and display image elements on the detected video of the user's hand to prompt the system to perform some kind of processing. In this embodiment, image elements for starting time information calibration processing, position information calibration processing, and content display processing are superimposed and displayed. Each of these processes may be configured to execute one of them, or two or more processes may be selected and executed. This configuration makes it possible to initiate each configuration process or content display process based on the user's action on the image elements.

[0151] Furthermore, the viewing device 100 is configured to start time information calibration processing, position information calibration processing, and content display processing when it detects the other hand at the position where the image elements are placed, while the image elements are superimposed on the video of one of the detected user's hands. Each of these processes may be executed individually, or two or more processes may be selected and executed. This configuration makes it possible to start each configuration process and content display process in response to the intuitive movement of the user's hand. Alternatively, the device may be configured to start time information calibration processing, position information calibration processing, and content display processing when it detects the other hand at the position where one or both of the image elements are placed, while the image elements are superimposed on the video of both of the user's hands.

[0152] Furthermore, the viewing device 100 can be configured to include a tracking means 160, as shown in Figures 1 and 4. The tracking means 160 is a means for tracking a specific element included in the external information 20 acquired by the input means 150, and is configured to track the user's hand, which is an element included in the external information 20 acquired and detected by the input means 150. This configuration makes it possible to easily track the user's hand.

[0153] The video 10 viewable by the display means 120 of the viewing device 100 is not limited to a single video but consists of multiple videos of any kind, and is configured to allow selection and display of a wide variety of videos 10. In this embodiment, the video 10 includes at least 3DCG content, the entirety of live-action footage, a part of a video, and assets that constitute a video, and multiple videos 10 selected from among these can be viewed on the viewing device 100.

[0154] As shown in Figure 4, the viewing device 100 is configured to include a video selection means 140. The video selection means 140 selects and displays one or more videos from a plurality of videos 10. In this embodiment, after selecting a video 10, the video selection means 140 performs a process to adjust the display of the video 10 according to the positional relationship as seen from the user's position. The viewing device 100 is capable of viewing the real world, and as shown in Figure 3, it is possible to superimpose the various videos selected by the video selection means 140 onto the real world in front of the user's eyes. At this time, the video 10 may be a video relating to a complete virtual reality world, or it may be configured to superimpose the video 10 onto the entire real world. Alternatively, it may be configured to display the selected video 10 on a part of the real world, or it may be configured to superimpose various video assets onto a large number of objects visible in the real world.

[0155] The display means 120 of the viewing device 100 displays each piece of content, which is composed of the video 10 held by the information holding means 110, in the video display area of ​​the viewing device 100. In this embodiment, the content consists of, for example, 3DCG content, the whole of a video such as live-action footage, a part of a video, or assets that make up a video. Specifically, possible configurations include, for example, a configuration in which a space appears on the floor and outer space or a space probe is displayed, a configuration in which a box becomes transparent and dancers sing and dance, a configuration in which tourist attractions appear in the real world, a configuration in which various advertisements are superimposed on the real world, a configuration in which lively illuminations are superimposed and displayed, a configuration in which objects are displayed on an empty table and shared by each user, and other configurations such as image processing that breaks down walls displayed as the real world, and configurations that perform effects such as making the floor or box displayed as the real world transparent.

[0156] For example, in a configuration where an object is displayed as an image 10 on a table, the positional relationship in which the object is seen will differ because each user is in a different position. Furthermore, if all users cannot view the object at the same time, a problem arises in that a synchronized experience cannot be achieved in the XR space. Since the object display process is performed by each viewing device 100, this problem occurs if the positional and temporal information is not calibrated. By calibrating the positional and temporal information using the calibration means 130, all users participating in the XR space can view the same object synchronously from any positional relationship, making it possible to share and enjoy the excitement in an immersive and highly real-world XR space that matches the physical sensations.

[0157] In this embodiment, the in-video space calibration system 2 is configured such that, as shown in Figure 5, the input means 150 acquires an image I attached to a sheet S or object O. The image I is a pattern that the input means 150 can recognize and acquire information from. In this embodiment, it consists of a two-dimensional code, but is not limited to this; any pattern that the input means 150 can acquire information from can be appropriately selected and used. The video selection means 140 recognizes the acquired image I and, based on the information read from this image I, compares it with definition information or accesses information on the web via a URL. This configuration simplifies and speeds up the calibration of location information and time information, as well as the display of content, making it easy to share and enjoy an XR space consisting of images that match the user's perception.

[0158] The viewing device 100 can be configured to perform time information calibration when it receives trigger information. Trigger information is information other than the external information 20 actively acquired by the input means 150. For example, it may be information generated by the user themselves operating image elements superimposed on their hand, or information generated primarily by system administrators or staff operating the system. When operating from a network device, for example, it may be information generated using an application on a smartphone. This makes it possible to generate trigger information and operate the system easily and functionally, but it is not limited to these. Accuracy of time information is required for the operation of the system, but discrepancies often occur inevitably during execution. With the configuration of the present invention, it becomes possible for users or staff operating the system other than users to initiate the calibration process, thereby ensuring the accuracy of time information.

[0159] Furthermore, the viewing device 100 can be configured to start displaying content when it receives trigger information. This configuration makes it possible for the user or staff other than the user to initiate the calibration process.

[0160] Furthermore, the viewing device 100 can be configured to hold time-related information. The viewing device 100 can also be configured to acquire time-related information from an external source. Moreover, in this embodiment, the device is configured to perform time information calibration processing at the start of the video 10 display process and / or during the video 10 display process. An example of an embodiment in which calibration processing is performed at the start of the video 10 display process or during the display process is the case of live streaming of a competitive race in an XR space.

[0161] Time information can be obtained from time information acquired from an NTP server or from reference time information on the web. This configuration allows for the acquisition and storage of accurate time for each device, enabling time information calibration based on precise time.

[0162] Schematic diagram of the in-video space calibration system according to the present invention; Diagram showing the controller holder and controller dock; Schematic diagram showing an embodiment of the in-video space calibration system; Schematic diagram showing another embodiment of the in-video space calibration system; Plan view of the sheet; Diagram showing the generation and display of feature points and frames of images, etc.

[0163] 1.2 In-video space calibration system 10 Video 20 External information 30 Definition information 40 Definition information 50 Definition information 100 Viewing device 110 Information holding means 120 Display means 130 Calibration means 140 Video etc. selection means 150 Input means 155 Current position acquisition means 160 Tracking means 170 Storage means 180 Angle detection means 185 Finger position detection means 186 Plane detection means 190 Angle adjustment means 200 Current information transmission means 210 Controller holder 220 Controller dock 230 Docking station S Sheet I Image O Object P Object

Claims

1. An in-video space calibration system (1) for calibrating the location and time information of each user in a space composed of images (10) provided by XR (Cross Reality) comprises: a viewing device (100) that a user can wear to view the images; a current information acquisition device (200) that acquires definition information (40) for one or more users wearing the viewing device (100); a display means (120) that displays the images based on the definition information (40); and a calibration means (130) that calibrates the location and / or time information of each user in the space composed of the images (10) based on the information acquired by the current information acquisition device (200). The calibration means (130) acquires each user's definition information (40) acquired via the current information acquisition device (200), and calibrates the position information and time information in the space composed of the video (10) based on the definition information (40), and the display means (120) displays the video (10). The current information acquisition device (200) consists of one or more controller holders (210) of different shapes and one or more controller docks (220) of shapes corresponding to the controller holders. A video space calibration system characterized in that each user attaches and fixes the controller holder (210) to the controller dock (220), thereby acquiring the user's definition information (40) and transmitting it to the calibration means (130) of the viewing device (100), and the calibration means (130) then performs calibration processing of each user's position information and / or time information within the space composed of the video (10).

2. The video spatial calibration system according to claim 1, characterized in that the current information acquisition device (200) comprises a docking station (230) in which the controller dock (220) is installed at predetermined intervals.

3. The video (10) consists of a plurality of videos, and includes at least 3DCG content, the whole of a video such as live-action footage, a part of a video, and assets that constitute a video; the viewing device (100) is provided with a video selection means (140) that selects and displays one or more of the plurality of videos (10); and the video selection means (140) automatically selects and displays videos that can be viewed from the positional relationship in which the video (10) is displayed, and / or the video (10) itself, according to claim 1 or 2, characterized in that the video spatial calibration system is described in claim 1 or 2, depending on the fixing conditions when the controller holder (210) is attached and fixed to the controller dock (220) and the time of installation.

4. The video spatial calibration system according to claim 3, characterized in that the video selection means (140) automatically selects and displays a video that can be viewed from the positional relationship in which the video (10) is displayed, and / or the video (10) itself, according to one or more combinations of the fixed combination information when the controller holder (210) is attached and fixed to the controller dock (220), the angle when the controller holder (210) is attached and fixed to the controller dock (220), and the on / off status of one or more of the buttons on the controller holder (210).

5. The video space calibration system according to claim 1 or 2, characterized in that the AI ​​selects relevant information based on acquired location information, information acquired based on dialogue, and information acquired from an external camera, generates and / or selects various types of information, and provides the information to the user in a presentation method including a dialogue format.

6. The video space calibration system according to claim 5, characterized in that the AI ​​has an information content display function that generates and / or selects and displays information and content related to real-world objects based on the location information and information acquired from an external camera.

7. The video space calibration system according to claim 5, characterized in that the AI ​​generates and / or selects and displays information and content related to real-world objects based on the location information and information acquired from an external camera.

8. The video space calibration system according to claim 5, characterized in that, when the AI ​​generates and / or selects information and content related to real-world objects based on the location information and information acquired from an external camera, it generates and displays feature points and frames on which such content should be displayed, and superimposes and places them on real-world objects.

9. An in-video space calibration system (2) for calibrating the location information and / or time information of each user in a space composed of images provided by XR (cross reality) comprises a viewing device (100) that a user can wear to view the images (10) and acquire the current definition information (50) of the user wearing the device, the viewing device (100) comprises a display means (120) that displays the images (10) based on the definition information (50), and a calibration means (130) that calibrates the location information and / or time information of each user in the space composed of the images based on the definition information (50), the calibration means (130) acquires the current definition information (50) of each user, calibrates it to the location information and / or time information in the space composed of the images (10), and then the display means (120) displays the images (10) related to the space.

10. The video space calibration system according to claim 9, wherein the viewing device (100) includes a current location acquisition means (155) for acquiring the user's current location, and the calibration means (130) acquires the location information acquired by the current location acquisition means (155) as the current definition information (50) for each user, and then calibrates it to the location information and / or time information in the space composed of the video (10).

11. The video space calibration system according to 10, characterized in that the current location acquisition means (155) acquires the user's current location using one or more of GNSS, VPS, beacons, location markers, and image markers.

12. The video space calibration system according to claim 10, characterized in that the current position acquisition means (155) acquires directional information that allows the user to infer the direction of view using a compass (magnetic sensor).

13. The video space calibration system according to claim 10, characterized in that the current position acquisition means (155) acquires acceleration and angular velocity information that allows the user to infer the direction of view using a MEMS gyroscope.

14. The video space calibration system according to claim 9 or 10, characterized in that the AI ​​selects relevant information based on acquired location information, information acquired based on dialogue, and information acquired from an external camera, generates and / or selects various information and content, and provides the information to the user in a presentation method including a dialogue format.

15. The video space calibration system according to 14, characterized in that the AI ​​has an information content display function that generates and / or selects and displays information and content related to real-world objects based on the location information and information acquired from an external camera.

16. The video space calibration system according to 14, characterized in that, when the AI ​​generates and / or selects information and content related to a real-world object based on the location information and information acquired from an external camera, it generates and displays feature points and frames on which the content should be displayed, and superimposes and positions these on the real-world object.

17. The video space calibration system according to any one of claims 9 to 11, wherein the viewing device (100) includes an input means (150) for acquiring external information (20), and the viewing device (100) detects the user's hand based on the external information (20) acquired by the input means (150), and detects that the user's hand has been placed in a predetermined position, thereby acquiring this as definition information (50) for each user, and the calibration means (130) performs calibration processing of the position information of each user in the space composed of the video (10).

18. The video space calibration system according to claim 9, characterized in that the viewing device (100) is configured to detect when the user's hand is placed on a predetermined position corresponding to the position of the hand when the hand is placed on the sheet (S).

19. The video space calibration system according to claim 9, characterized in that the viewing device (100) is configured to detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when it is placed on an object (O) that actually exists.

20. The video space calibration system according to claim 9, characterized in that the viewing device (100) is configured to detect when the user's hand is placed at a predetermined position corresponding to the position of the hand when the hand is placed on a plate-shaped object (P) drawn in a virtual space.

21. The video space calibration system according to claim 9, characterized in that the viewing device (100) has a configuration that makes it a prerequisite for detection that the user's hand is placed on it, by first gripping the hand and then opening it.

22. The video space calibration system according to claim 9, characterized in that the viewing device (100) is configured to guide the user to a sheet (S) or an object (O) that actually exists by displaying one or more images, videos, or other content in a virtual space as a procedure for placing a hand on a sheet (S) or an object (O) that actually exists.

23. The video space calibration system according to claim 9, characterized in that the viewing device displays a plate-shaped object (P) to be drawn in a virtual space when the input means (150) detects a user's hand on an object (O) that actually exists.

24. The video space calibration system according to claim 9, characterized in that the viewing device has a current position acquisition means (155) that detects the plane of the real world visible to the user in a predetermined manner, and raycasts in a predetermined manner to project and display a plate-shaped object (P).

25. The viewing device (100) comprises an angle detection means (180) for detecting the angle of the user's hand, and an angle adjustment means (190) for adjusting the angle of the image (10) displayed on the display means (120), wherein the angle adjustment means (190) adjusts the angle of the image (10) displayed on the display means (120) according to the angle of the user's hand detected by the angle detection means (180) during or after the calibration process of each user's position information by the calibration means (130), characterized in that the in-video space calibration system according to any one of claims 9 to 24.

26. The video space calibration system according to 25, wherein the viewing device (100) includes a finger position detection means (185) for detecting the position of each fingertip of the user's hand, and the finger position detection means (185) detects the position of each fingertip of the user's hand, and then the plane detection means (186) estimates a plane from the position information of each finger using a plane regression method to obtain a plane that matches the user's hand.

27. The video space calibration system according to any one of claims 9 to 24, characterized in that when the viewing device (100) detects information in the external information (20) acquired by the input means (150) that matches the definition information (30) that has been determined to be applicable, it selects and displays one or more of the following: 3DCG content, live-action footage, or any part of the video, or assets that constitute the video.

28. The video space calibration system according to any one of claims 9 to 24, characterized in that the viewing device (100) performs calibration processing of location information on which content is displayed when it detects information in the external information (20) acquired by the input means (150) that matches the definition information (30) that has been determined to be applicable.

29. The video space calibration system according to any one of claims 9 to 24, characterized in that the viewing device (100) performs a time information calibration process when it detects information in the external information (20) acquired by the input means (150) that matches the defined information (30) that has been determined to be applicable.

30. The video space calibration system according to any one of claims 9 to 29, characterized in that the external information (20) consists of a two-dimensional code or information such as a string obtained therefrom, external video and / or audio, and the definition information (30) consists of a two-dimensional code or information such as a string obtained therefrom corresponding to the external information (20), external video and / or audio.

31. The video space calibration system according to any one of claims 9 to 29, characterized in that when the viewing device (100) detects information existing on the web that can be accessed based on URL information obtained from the external information (20), it selects and displays one or more of the following based on the information: 3DCG content, the whole of live-action footage, a part of the footage, or assets that constitute the footage.

32. The video space calibration system according to any one of claims 9 to 29, characterized in that the viewing device (100) detects information existing on the web that can be accessed based on URL information obtained from the external information (20), and performs calibration processing of location information on which content is displayed based on said information.

33. The video space calibration system according to any one of claims 9 to 29, characterized in that the viewing device (100) detects information existing on the web that is accessed based on URL information obtained from the external information (20), and performs calibration processing of time information based on said information.

34. The video space calibration system according to any one of claims 18 to 24, characterized in that the viewing device (100) detects the user's hand based on the external information (20) acquired by the input means (150), and then detects that the user's hand is placed in a predetermined position corresponding to the position of the hand when the hand is placed on the sheet / object.

35. The video space calibration system according to 32, characterized in that the viewing device (100) superimposes and displays image elements for initiating one or more of the following processes on the detected video of the user's hand: time information calibration processing, position information calibration processing, and content display processing.

36. The video space calibration system according to 34, characterized in that the viewing device (100) superimposes and displays image elements on the video of one of the user's hands that has been detected, and when the other hand of the user is detected at the position where the image elements are placed, it starts one or more of the following processes: time information calibration process, position information calibration process, and content display process.

37. The viewing device (100) is further equipped with tracking means (160) for tracking specific elements included in the external information (20) acquired by the input means (150), and the tracking means (160) is characterized in that it tracks the user's hand, which is an element included in the external information (20) acquired and detected by the input means (150), as described in any one of claims 34 to 36.

38. The video (10) consists of a plurality of images and includes at least 3DCG content, the whole of live-action footage, a part of an image, and assets constituting an image; the viewing device (100) is provided with a video selection means (140) that selects and displays one or more of the plurality of images (10); the input means (150) acquires the image (I) attached to the sheet (S) / object (O); and the video selection means (140) recognizes the acquired image (I) and then compares it with definition information based on the information read from the image (I) or accesses information on the web via a URL, characterized in that the video spatial calibration system according to any one of claims 9 to 37.

39. The video space calibration system according to any one of claims 9 to 24, characterized in that the viewing device (100) performs time information calibration when it receives trigger information.

40. The video space calibration system according to any one of claims 9 to 24, characterized in that the viewing device (100) starts displaying content when it receives trigger information.

41. The viewing device (100) is configured to hold time-related information or to acquire time-related information from an external source, and the video spatial calibration system according to any one of claims 9 to 24 is characterized in that it performs time information calibration processing at the start of the video (10) display processing and / or during the video display processing.

Citation Information

Patent Citations

  • Eye Tracking Calibration Technique

    JP2022173312A

  • Information processing device, information processing program, and information processing method

    JP2023092003A

  • Cross-reality systems for large-scale environments

    JP2023524446A

  • VR video space generation system

    JP7547501B2

  • Program, information processing device, and information processing method

    WO2024209802A1