Display system, first terminal, head-mounted display, and display method

WO2026177102A1PCT designated stage Publication Date: 2026-08-27PANASONIC INTELLECTUAL PROPERTY CORP OF AMERICA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2026/005562
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-18
Filing Date
2026-02-16
Publication Date
2026-08-27

Smart Images

  • Figure JP2026005562_27082026_PF_FP_ABST
    Figure JP2026005562_27082026_PF_FP_ABST
Patent Text Reader

Abstract

This display system has a first head-mounted display (10) having a first processor, and a second head-mounted display (20) having a second processor. The first processor is configured to: display a first image; acquire identification information for identifying an object in the first image on the basis of a user operation; acquire change information for the object on the basis of a user operation; display the first image changed on the basis of the change information; and transmit the identification information and the change information to the second head-mounted display. The second processor is configured to: display a second image; receive the identification information and the change information; identify an object in the second image on the basis of the identification information; and display the second image changed on the basis of the change information.
Need to check novelty before this filing date? Find Prior Art

Description

Display system, first terminal, head-mounted display, and display method

[0001] This disclosure relates to display systems, and in particular to display systems including multiple head-mounted displays.

[0002] In recent years, Augmented Reality (AR) technology has attracted attention as a technology that overlays virtual information onto the real world environment. This technology extends the user's visual experience of the real world and is being applied in various fields. For example, in entertainment, education, medicine, and manufacturing, displaying virtual objects and information in the real world environment in real time brings new conveniences and effects compared to conventional operations or experiences.

[0003] In particular, development of AR goggles (head-mounted displays) is progressing as one of the AR technologies. AR goggles are devices that, when worn by the user, display information that combines images of the real world with virtual objects within the user's field of vision. This allows users to intuitively recognize, manipulate, and observe virtual objects placed on their physical environment while viewing it.

[0004] Patent No. 6654625

[0005] This disclosure aims to provide display systems and other technologies that further expand the convenience and possibilities of AR technology.

[0006] A display system according to one aspect of the present disclosure is a display system comprising: a first head-mounted display having a first camera, a first display, and a first processor; and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor is configured to display a first image captured by the first camera on the first display, acquire identification information to identify an object in the first image based on user operation, acquire change information for the object based on user operation, display the first image modified based on the change information on the first display, and transmit the identification information and the change information to the second head-mounted display; and the second processor is configured to display a second image captured by the second camera on the second display, receive the identification information and the change information, identify the object in the second image based on the identification information, and display the second image modified based on the change information on the second display.

[0007] A first terminal according to one aspect of the present disclosure has a processor and memory, the processor being configured to use the memory to acquire identification information and modification information for objects in an image taken in front of the user, and to transmit the identification information and modification information to another head-mounted display.

[0008] A head-mounted display according to one aspect of the present disclosure comprises a camera, a display, and a processor, wherein the processor is configured to display an image captured by the camera on the display, receive identification information and modification information about an object from another head-mounted display, identify the object in the image based on the identification information, and display the modified image on the display based on the modification information.

[0009] A display method according to one aspect of the present disclosure is a display method performed by a first head-mounted display having a first camera, a first display, and a first processor, and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor displays a first image captured by the first camera on the first display, acquires identification information to identify an object in the first image based on user operation, acquires change information for the object based on user operation, displays the first image modified based on the change information on the first display, transmits the identification information and the change information to the second head-mounted display, and the second processor displays a second image captured by the second camera on the second display, receives the identification information and the change information, identifies the object in the second image based on the identification information, and displays the second image modified based on the change information on the second display.

[0010] A display method according to one aspect of the present disclosure is a display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, acquires identification information to identify an object in the image based on user operation, acquires change information to the object based on user operation, displays the modified image based on the change information on the display, and transmits the identification information and the change information to another head-mounted display.

[0011] A display method according to one aspect of the present disclosure is a display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, receives identification information and modification information about an object from another head-mounted display, identifies the object in the image based on the identification information, and displays the modified image on the display based on the modification information.

[0012] Herein, in this disclosure, “object” is a concept that includes all tangible things, such as machines, devices and tools, as well as parts or all of people, animals, plants, etc. “Modification” is a concept that includes virtual modifications to an object that exists as a substance, such as virtual changes to the color of part or all of an object, virtual changes to the size of part or all of an object, addition of virtual objects (parts, etc.), virtual removal of part or all of an object (parts, etc.) that exists as a substance, or changes in position and orientation.

[0013] In this disclosure, "identifying information for an object" may be information used to handle the same object or the same target area at both the transmitting and receiving terminals. For example, identifying information may include an identifier associated with the object or target area (such as an ID number or UUID), feature quantities of the object (such as shape, color, or size), image information of the object, a two-dimensional or three-dimensional model of the object, the position and orientation of the object or target area, or a combination thereof. Furthermore, if the identifying information includes position and orientation, such position and orientation may be defined in the local coordinate system of the transmitting terminal, the local coordinate system of the receiving terminal, or the world coordinate system.

[0014] According to the display system of this disclosure, a user of the first head-mounted display can specify specific information and modification information for a particular object in real space, and the first and second images are modified based on that specific information and modification information. This allows the user to share the virtual changes they have made to the object with other users, further expanding the convenience and possibilities of AR technology.

[0015] Figure 1A is a system configuration diagram of a display system according to the first embodiment of the present disclosure. Figure 1B is a system configuration diagram of another display system according to the first embodiment of the present disclosure. Figure 2 is a diagram showing the usage state of the display system according to the first embodiment of the present disclosure. Figure 3 is a diagram showing the usage state of the display system according to the first embodiment of the present disclosure. Figure 4 is a hardware configuration diagram of the first HMD and second HMD according to the first embodiment of the present disclosure. Figure 5A is another hardware configuration diagram of the first HMD according to the first embodiment of the present disclosure. Figure 5B is another hardware configuration diagram of the second HMD according to the first embodiment of the present disclosure. Figure 6A is a functional configuration diagram of the first HMD according to the first embodiment of the present disclosure. Figure 6B is a functional configuration diagram of the second HMD according to the first embodiment of the present disclosure. Figure 7 is a flowchart of the first HMD and second HMD according to the first embodiment of the present disclosure. Figure 8 is a diagram showing the software architecture according to the first embodiment of the present disclosure. Figure 9 is a block diagram showing the functional configuration of the display system according to the first embodiment of the present disclosure. Figure 10 is a sequence diagram according to the first embodiment of this disclosure. Figure 11 is a diagram illustrating the procedure for 3D modeling and identifying an object according to the second embodiment of this disclosure. Figure 12 is a diagram illustrating the procedure for 3D modeling and identifying an object according to the second embodiment of this disclosure. Figure 13 is a diagram illustrating the procedure for 3D modeling and identifying an object according to the second embodiment of this disclosure. Figure 14 is a diagram illustrating the procedure for 3D modeling and identifying an object according to the second embodiment of this disclosure. Figure 15 is a block diagram illustrating the functional configuration of a display system according to the second embodiment of this disclosure. Figure 16 is a sequence diagram according to the second embodiment of this disclosure. Figure 17 is a block diagram illustrating the functional configuration of a display system according to the fourth embodiment of this disclosure. Figure 18 is a sequence diagram according to the fourth embodiment of this disclosure. Figure 19 is a block diagram illustrating the functional configuration of a display system according to a modified example of the fourth embodiment of this disclosure. Figure 20 is an example of an ID list according to the fourth embodiment of this disclosure. Figure 21 is an example of processing using the ID list according to the fourth embodiment of this disclosure. Figure 22 is a diagram illustrating a coordinate transformation according to the fifth embodiment of this disclosure. Figure 23 is a diagram showing a coordinate transformation according to the fifth embodiment of this disclosure.Figure 24 is a diagram showing a coordinate transformation according to the fifth embodiment of the present disclosure. Figure 25 is a block diagram showing the functional configuration of a display system according to the fifth embodiment of the present disclosure. Figure 26 is a sequence diagram according to the fifth embodiment of the present disclosure.

[0016] The embodiments of this disclosure will be described below with reference to the drawings. The embodiments described below are all specific examples of this disclosure. Therefore, the components, their arrangement positions, and connection configurations shown in the following embodiments are examples and are not intended to limit this disclosure. Furthermore, components in the following embodiments that are not described in an independent claim will be described as optional components.

[0017] Furthermore, each figure is a schematic diagram and not necessarily a strictly accurate representation. Note that in each figure, substantially identical components are denoted by the same reference numerals, and redundant explanations may be omitted or simplified.

[0018] Furthermore, in this specification, ordinal numbers such as "first," "second," etc., do not mean the number or order of components unless otherwise specified, but are used to avoid confusion and to distinguish similar components.

[0019] (Knowledge forming the basis of the invention) Currently, development of head-mounted displays (HMDs) is progressing. However, the challenges of using multiple HMDs remain unresolved. For example, when multiple people view the same image or communicate with each other, synchronization or alignment between HMDs is difficult. Also, if the specifications of each HMD, such as rendering performance or field of view, differ, it is not possible to provide a consistent video experience. Furthermore, if the HMD is difficult to put on or operate, the convenience and comfort of AR technology are reduced. For these reasons, conventional technology has the problem that usability and quality are not sufficiently ensured when using multiple HMDs.

[0020] Based on the above, it is necessary to further expand the convenience and potential of AR technology.

[0021] Therefore, a display system according to a first aspect of the present disclosure is a display system having a first head-mounted display having a first camera, a first display, and a first processor, and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor is configured to display a first image captured by the first camera on the first display, acquire identification information to identify an object in the first image based on user operation, acquire change information for the object based on user operation, display the first image modified based on the change information on the first display, and transmit the identification information and the change information to the second head-mounted display, and the second processor is configured to display a second image captured by the second camera on the second display, receive the identification information and the change information, identify the object in the second image based on the identification information, and display the second image modified based on the change information on the second display.

[0022] This allows the user of the first head-mounted display to specify specific information and modification information for a particular object in real space, and the first and second images are modified based on that specific information and modification information. Therefore, the virtual changes made by the user to the object can be shared with the user of the second head-mounted display, resulting in a display system that further expands the convenience and possibilities of AR technology.

[0023] For example, in the display system according to the second aspect of this disclosure, the specific information includes information of a 3D model of the object in the display system according to the first aspect.

[0024] This enables accurate object mapping even with differences in viewpoint and partial occlusion, resulting in a display system capable of highly accurate mapping.

[0025] For example, in the display system according to the third aspect of this disclosure, in the display system according to the first aspect, the identifying information includes the ID of the identified object in the list of objects present in space.

[0026] As a result, when a list including IDs is pre-registered, a display system with a high processing speed is realized.

[0027] Further, for example, in the display system according to the fourth aspect of the present disclosure, in the display system according to the first aspect, the specific information includes the position coordinates of the object.

[0028] As a result, a display system in which the second head-mounted display can receive the position coordinates of the object and thus can accurately associate the object is realized.

[0029] The first terminal according to the fifth aspect of the present disclosure has a processor and a memory. The processor uses the memory to obtain specific information for identifying an object in an image captured in front of the user and change information for the object, and is configured to transmit the specific information and the change information to another head-mounted display.

[0030] As a result, the user of the first terminal designates specific information and change information for a specific object existing in the real space, and the specific information and the change information are transmitted to another head-mounted display. When an image is changed based on the specific information and the change information, the virtual change made by the user to the object can be shared with the users of other head-mounted displays, so that a first terminal that expands the further convenience and possibilities of AR technology is realized.

[0031] [[ID=十六]] Further, for example, in the first terminal according to the sixth aspect of the present disclosure, in the first terminal according to the fifth aspect, the specific information and the change information are generated based on a user operation.

[0032] As a result, since information is explicitly generated using a user operation as a trigger, mistransmission of specific information and change information can be suppressed.

[0033] Further, for example, the first terminal according to the seventh aspect of the present disclosure further has a display in the first terminal according to the fifth aspect, and the processor displays the image changed based on the change information on the display.

[0034] This allows the user of the first terminal to confirm that the image is changed based on specific information and change information.

[0035] For example, in the first terminal according to the eighth aspect of this disclosure, the first terminal according to the fifth aspect is a first head-mounted display, and the specific information includes position coordinate information of the object in a first local coordinate system with the transmitting first head-mounted display as the origin, and position coordinate information of the receiving second head-mounted display in the first local coordinate system.

[0036] As a result, the first terminal (first head-mounted display) sends specific information including the position coordinates of the second head-mounted display, simplifying the coordinate transformation process of the receiving second head-mounted display and enabling a more accurate understanding of the positional relationship.

[0037] For example, in the first terminal according to the ninth aspect of this disclosure, the position coordinate information of the second head-mounted display is determined based on the respective image positions of objects that are displayed in common in the image acquired by the first head-mounted display and the image acquired by the second head-mounted display.

[0038] This allows for the calculation of relative coordinates using objects common to both the first and second head-mounted displays, thus enabling a more accurate understanding of positional relationships.

[0039] For example, in the first terminal according to the tenth aspect of this disclosure, the first terminal according to the fifth aspect is a first head-mounted display, and the specific information includes position coordinate information of the object in a world coordinate system fixed in space.

[0040] This allows multiple devices to operate while referencing the same world coordinate system by using a common world coordinate system, enabling a more accurate understanding of their relative positions.

[0041] For example, in the first terminal according to the eleventh aspect of the present disclosure, the first terminal according to the fifth aspect is a first head-mounted display, and the specific information includes position coordinate information of the object in a first local coordinate system with the transmitting first head-mounted display as the origin.

[0042] This allows local coordinates to be used based on a single device, the first head-mounted display, making it easier to configure the system for determining positional relationships.

[0043] For example, in the first terminal according to the twelfth aspect of this disclosure, the position coordinate information of the other head-mounted display in the first local coordinate system is predetermined based on the respective image positions of objects that are displayed in common in the image acquired by the first head-mounted display and the image acquired by the other head-mounted display.

[0044] This allows for the calculation of relative coordinates using objects common to both the first head-mounted display and the other head-mounted display, thus enabling a more accurate understanding of positional relationships.

[0045] A head-mounted display according to a thirteenth aspect of the present disclosure comprises a camera, a display, and a processor, wherein the processor is configured to display an image captured by the camera on the display, receive identification information and modification information about an object from another head-mounted display, identify the object in the image based on the identification information, and display the modified image on the display based on the modification information.

[0046] As a result, images are modified and displayed based on specific and modified information received from other head-mounted displays. This allows virtual changes made to objects in the head-mounted display and other head-mounted displays to be shared, thus realizing a head-mounted display that further expands the convenience and possibilities of AR technology.

[0047] For example, in a head-mounted display according to a 14th aspect of the present disclosure, the received identification information in a head-mounted display according to a 13th aspect is first position coordinate information of the object in a first local coordinate system with the other head-mounted display on the transmitting side as the origin, and position coordinate information of the head-mounted display on the receiving side in the first local coordinate system. The processor converts the first local coordinate system to a second local coordinate system with the position of the head-mounted display as the origin, based on the position coordinate information of the head-mounted display on the receiving side in the first local coordinate system, converts the first position coordinate information to second position coordinate information in the second local coordinate system, and identifies the object in the image based on the second position coordinate information.

[0048] This transforms the first local coordinate system into the second local coordinate system, allowing the head-mounted display to process data in its own local coordinate system (the second local coordinate system), thus enabling a more accurate understanding of spatial relationships.

[0049] A display method according to a 15th aspect of the present disclosure is a display method performed by a first head-mounted display having a first camera, a first display, and a first processor, and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor displays a first image captured by the first camera on the first display, acquires identification information to identify an object in the first image based on user operation, acquires change information for the object based on user operation, displays the first image modified based on the change information on the first display, transmits the identification information and the change information to the second head-mounted display, and the second processor displays a second image captured by the second camera on the second display, receives the identification information and the change information, identifies the object in the second image based on the identification information, and displays the second image modified based on the change information on the second display.

[0050] This allows the user of the first head-mounted display to specify specific information and modification information for a particular object in real space, and the first and second images are modified based on that specific information and modification information. Therefore, the virtual changes made by the user to the object can be shared with the user of the second head-mounted display, resulting in a display system that further expands the convenience and possibilities of AR technology.

[0051] A display method according to a sixteenth aspect of the present disclosure is a display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, acquires identification information to identify an object in the image based on user operation, acquires change information for the object based on user operation, displays the modified image based on the change information on the display, and transmits the identification information and the change information to another head-mounted display.

[0052] This allows a user of a head-mounted display to specify specific information and modification information for a particular object in the real world, and this information is then transmitted to other head-mounted displays. If the image on the other head-mounted display is modified based on the specified information and modification information, the virtual changes made by the user to the object can be shared with the user of the other head-mounted display, thus realizing a display method that further expands the convenience and possibilities of AR technology.

[0053] A display method according to a 17th aspect of this disclosure is a display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, receives identification information and modification information about an object from another head-mounted display, identifies the object in the image based on the identification information, and displays the modified image on the display based on the modification information.

[0054] As a result, images are modified and displayed based on specific information and change information received from other head-mounted displays. This allows virtual changes made to objects in the head-mounted display and other head-mounted displays to be shared, thus realizing a display method that expands the convenience and possibilities of AR technology.

[0055] These comprehensive or specific embodiments may be implemented as a system, method, integrated circuit, computer program, or recording medium such as a computer-readable CD-ROM, or as any combination of a system, method, integrated circuit, computer program, or recording medium.

[0056] (First Embodiment) This disclosure relates to a system in which a first user 1 wears a first head-mounted display (first HMD) 10 and a second user 2 wears a second head-mounted display (second HMD) 20, and displays an object (model 3) present in the same space as they each view it on the first HMD 10 and the second HMD 20. This disclosure relates in particular to a display system 200 in which, when the first user 1 performs a virtual modification operation on a specific object in the first image 17 that he is viewing, a virtual modification is also made to the same specific object in the second image 27 that the second user 2 is viewing, and both the first user 1 and the second user 2 can see the image of the virtually modified object.

[0057] A head-mounted display (HMD) is a device that displays images in the user's field of vision. For example, it is used to experience XR (Extended Reality), which includes VR (Virtual Reality), AR (Augmented Reality), and MR (Mixed Reality). HMDs are worn so that the display is positioned in front of the eyes, and are often shaped like goggles or glasses, but are not limited to these forms.

[0058] Figure 1A is a system configuration diagram of a display system 200 according to the first embodiment of the present disclosure. Figure 1A is also a diagram showing an example of the display system 200 according to the first embodiment of the present disclosure. The display system 200 includes a first terminal (first HMD 10) used by a first user 1 and a second terminal (second HMD 20) used by a second user 2. In the display system 200, the first terminal and the second terminal communicate with each other, and the second terminal receives data requested by the second terminal from the first terminal. The first terminal can also be described as the transmitting terminal, i.e., the transmitting first HMD 10, and the second terminal as the receiving terminal, i.e., the receiving second HMD 20. Conversely, the first terminal receives data requested by the first terminal from the second terminal. Figure 1A shows two terminals communicating, but there may be three or more terminals.

[0059] Figure 1B is a system configuration diagram of another display system 200a according to the first embodiment of the present disclosure. The display system 200a includes a first terminal (first HMD 10) used by a first user 1, a second terminal (second HMD 20) used by a second user 2, and a server 100. The server 100 communicates with the first terminal and the second terminal and holds data from the first terminal and the second terminal. The server 100 transmits data from the first terminal requested by the second terminal to the second terminal, and transmits data from the second terminal requested by the first terminal to the first terminal.

[0060] Figures 2 and 3 are diagrams showing the usage state of the display system 200 according to the first embodiment of the present disclosure. Figures 2 and 3 also show how the display system 200 according to the first embodiment of the present disclosure is being used. The first user 1 is wearing the first HMD 10 and looking at a model of a car 3 (object). The first HMD 10 is equipped with a first camera 11 that takes pictures of the front, and the first image 17 of the model 3 taken by the first camera 11 is displayed on the first display 12 of the first HMD 10. That is, the first user 1 is looking at the first image 17 of the model 31 displayed on the first display 12 of the first HMD 10.

[0061] The second user 2 is wearing the second HMD 20 and looking at the car model 3 (object). The second HMD 20 is equipped with a second camera 21 that takes pictures of the front, and the second image 27 of the model 3 taken by the second camera 21 is displayed on the second display 22 of the second HMD 20. In other words, the second user 2 is looking at the second image 27 of the model 32 displayed on the second display 22 of the second HMD 20.

[0062] The first user 1 performs an operation on the first HMD 10. The first processor 15 of the first HMD 10 acquires identification information to identify an object in the first image 17 based on the operation of the first user 1, and acquires change information for the object based on the operation of the first user 1. More specifically, it is as follows:

[0063] The first user 1 performs a modification operation on the model 31 included in the first image 17. The modification operation is performed, for example, by the first user 1 saying, "Change the color of the car roof from white to black." When the first sensor 13 (microphone) of the first HMD 10 acquires voice information, the first processor 15 of the first HMD 10 uses voice recognition technology to acquire identification information (car roof 33) that identifies the target object from the voice information and modification information regarding the content of the modification (change the color from white to black). As a modification operation, known interface technologies such as the first sensor 13 acquiring gestures or eye movements of the first user 1 may be used in addition to voice commands. Furthermore, if there is a controller attached to the first HMD 10, the modification operation may be performed using that controller. Hereinafter, information including identification information and / or modification information may be referred to as control information.

[0064] The specific information may include, for example, an identifier (ID number or UUID, etc.) associated with an object or a part of an object, the position coordinates of the object in the image, region information indicating the area containing the object (e.g., rectangle, contour, polygon, or mask), the position and orientation of the object, or the feature quantities of the object (size, shape, color, texture, feature points, or 2D / 3D model, etc.). More specifically, the specific information may be any information that identifies an object in the first image 17 (e.g., the roof 33 of the model 31 (car)). As described above, the specific information may be information that identifies a part of an object in the first image 17, such as "the roof of the car." Alternatively, for example, the specific information may be information that identifies the entire object in the first image 17, such as "car."

[0065] The identifying information may be information that identifies one or more objects among the multiple objects shown in the first image 17, and may particularly identify two or more objects. Furthermore, even if only one object is shown in the first image 17, the identifying information may be information that identifies multiple parts (for example, four tires) of that single object.

[0066] Change information may be any information relating to the nature of the changes made to an object. Change information is information that instructs a change in the display manner of the object, more specifically, a change in the display manner of the object identified by the specific information. If a part of the object is identified by the specific information, the change information will be information that instructs a change in the display manner of that part. Change information is not limited to changes in the display manner, and may also include, for example, changes in the position, orientation, or size of the object, deformation or replacement of the shape of the object, addition, deletion, or duplication of the object, and information regarding the timing or duration of applying the changes. Change information may also include area information indicating the area to be changed and information indicating the processing content to be applied to that area.

[0067] Furthermore, the change information is used to instruct a change in the display mode of objects in the first image 17 displayed on the first display 12, and to instruct a change in the display mode of objects in the second image 27 displayed on the second display 22.

[0068] As mentioned above, changing the display manner of an object may involve changing the display color of the object, such as "changing the color from white to black," but is not limited to this. For example, examples of changing the display manner include the object flashing in the image, or a mark such as a "star mark" or "arrow mark" being displayed on the object. Also, as mentioned above, examples of changing the display color of an object may include changing an opaque color to transparent, or changing it to be displayed with a pattern such as dots or stripes.

[0069] Furthermore, the first processor 15 displays the first image 17, which has been modified based on the change information, on the first display 12. More specifically, it is as follows:

[0070] The first processor 15 of the first HMD 10 identifies the roof 33 of the model 31 (automobile) from the first image 17 displayed on the first display 12 using image recognition technology based on specific information (the roof 33 of the automobile), and changes the color of the roof 35 of the model 31 (automobile) from white to black based on the change information and displays it (Figure 3). The color change may be performed by changing the display color of the first image 17, or an image of the black roof 33 may be generated and superimposed on the roof 33 of the model 31 (automobile). The first processor 15 then transmits the specific information and the change information to the second HMD 20 via the first communication device 14.

[0071] The second processor 25 of the second HMD 20 receives identification information and modification information transmitted from the first HMD 10. Based on the identification information, the second processor 25 identifies an object in the second image 27 and displays the modified second image 27 on the second display 22 based on the modification information. More specifically, it is as follows:

[0072] The second processor 25 identifies the roof 34 of the model 32 (automobile) from the second image 27 displayed on the second display 22 using image recognition technology based on the received specific information (automobile roof 33), and changes the color of the roof 36 of the model 32 (automobile) from white to black and displays it based on the received change information (Figure 3). The color change may be performed by changing the display color of the second image 27, or an image of the black roof 34 may be generated and superimposed on the roof 34 of the model 32 (automobile).

[0073] In addition to the modifications described above, the modifications may also include various operations such as changing the color of specific areas of an object, adding or removing components, modifying components, changing sizes, and masking (concealing). The modified image may be displayed by superimposing the image generated based on the modifications onto the original image.

[0074] The specific information transmitted from the first HMD 10 to the second HMD 20 may include text information such as "car roof," or it may include an image of an object recognized by the first processor 15 in the first image 17, and the second processor 25 may identify the object in the second image 27 using image recognition based on the received image. Furthermore, the specific information may be voice information uttered by the first user 1, or it may include spatial features, the position coordinates of an object, feature points of an object, a 3D model of an object, or predetermined object identification information. It is not limited to any format as long as it allows the first HMD 10 and the second HMD 20 to identify the object. In other words, the specific information may include information on the 3D model of an object, or it may include the position coordinates of an object.

[0075] It goes without saying that the images described above as specific information may be not only still images but also videos. For example, it may be the video file itself, or it may be information extracted from the video. For example, the first processor 15 may select some still images (frames) or regions from the video captured by the first HMD 10, and send the selected frames or regions to the second HMD 20 as specific information. Alternatively, the first processor 15 may extract features (edges, textures, colors, etc.) within the frames from the video captured by the first HMD 10, and send that feature information to the second HMD 20 as specific information. These selection or extraction processes may be performed by the second processor 25.

[0076] The specific information transmitted from the first HMD 10 to the second HMD 20 may be the output result when the aforementioned video or spatial features are input to a machine learning model that has already been trained. In other words, it may be information that identifies the type and shape or location of an object, output from a machine learning model. The second processor 25 may take the video received by the second HMD 20 as specific information as input and perform machine learning object identification processing to obtain information that identifies the type and shape or location of an object. Alternatively, the first processor 15 may send the video to the server 100, the server 100 may perform machine learning object identification processing and send the output result to the second HMD 20 as specific information.

[0077] As a typical example, we have described an example in which the object to be changed in display mode is identified by a change operation (e.g., voice command) performed by the first user 1 wearing the first HMD 10. However, the identification of an object does not necessarily have to be based on an operation by the first user 1. For example, the first processor 15 may identify a characteristic object included in the captured image as the object to be changed. Alternatively, if multiple characteristic objects are detected in the captured image, multiple candidates may be presented to the first HMD 10 or the second HMD 20.

[0078] Figure 4 is a hardware configuration diagram of the first HMD 10 and the second HMD 20 according to the first embodiment of the present disclosure. The first terminal and / or the second terminal have a processor 71, a memory 72, and a communication interface 73 interconnected via a bus 74.

[0079] The processor 71 may be implemented by one or more CPUs (Central Processing Units), GPUs (Graphics Processing Units), processing circuits, etc., which may consist of one or more processor cores, and executes various functions and processes of the first terminal and / or the second terminal according to data such as programs, instructions, and parameters necessary to execute said programs or instructions stored in the memory 72.

[0080] The memory 72 is composed of, for example, RAM (Random Access Memory) or ROM (Read-Only Memory). Alternatively, the memory 72 may be internal memory integrated into the CPU or GPU. Furthermore, the memory may store a program or various data that performs the display processing described herein.

[0081] The communication interface (communication IF) 73 is implemented by various communication circuits that communicate with other terminals and server devices. The communication interface 73 is, for example, a communication module compatible with communication methods such as Bluetooth® or WiGig®. The first HMD 10 and the second HMD 20 communicate with each other, for example, via the communication interface 73, and send and receive control information, etc. The received control information, etc. is stored in, for example, the memory 72.

[0082] The communication interface 73 is composed of, for example, a signal processing circuit and an antenna corresponding to the communication method. The communication method is not limited to Bluetooth® and WiGig®, but may also be LTE (Long-Term Evolution), NR (New Radio), or Wi-Fi®, etc.

[0083] However, the hardware configuration described above is merely an example, and the first HMD 10 and the second HMD 20 according to this disclosure may be implemented by any other suitable hardware configuration.

[0084] The first HMD 10 includes a first display 12, and the first processor 15 that performs various processes and the first memory 16 connected to the first processor 15 may be from different terminals or servers 100. The first HMD 10 does not need to have any other configurations. The same applies to the second HMD 20. The first terminal and / or the second terminal may be mobile terminals such as smartphones or tablet devices.

[0085] Figure 5A is a diagram of another hardware configuration of the first HMD 10 according to the first embodiment of the present disclosure. Figure 5B is a diagram of another hardware configuration of the second HMD 20 according to the first embodiment of the present disclosure.

[0086] The first HMD 10 includes a first camera 11, a first display 12, a first sensor 13, a first communication device 14, a first processor 15, and a first memory 16. The second HMD 20 includes a second camera 21, a second display 22, a second sensor 23, a second communication device 24, a second processor 25, and a second memory 26.

[0087] The first HMD 10 includes a first camera 11 located outside the first HMD 10 that captures the direction of the first user 1's gaze (forward), a first display 12 that is directed towards the first user 1 and displays a first image 17 captured by the first camera 11, a first sensor 13 that detects operation information of the first user 1 regarding the first image 17, such as the first user 1's gestures, gaze, and voice, a first communication device 14 that can communicate with the second communication device 24 of the second HMD 20, a first processor 15 that controls the first camera 11, the first display 12, the first sensor 13, and the first communication device 14, etc., and a first memory 16 that stores programs executed by the first processor 15 and various information.

[0088] The first camera 11 is comprised of a shooting device that captures the direction of the first user 1's line of sight (forward).

[0089] The first display 12 is the part that projects images onto the first user 1. The first display 12 may use a panel such as liquid crystal or organic EL, or it may project images directly onto the retina using a light source such as a laser. The first display 12 is often mounted so that it is positioned in front of the first user 1's eyes.

[0090] The first sensor 13 is used to track the gaze, posture, movements (gestures), and / or voice of the first user 1. The first sensor 13 may be, for example, an accelerometer, gyroscope, infrared sensor, depth sensor, microphone, camera, etc. The first sensor 13 may detect head movement or position.

[0091] The first communication device 14 is implemented by various communication circuits that communicate with other devices and server devices.

[0092] The first processor 15 is realized by one or more CPUs, GPUs, processing circuits, etc., which may consist of one or more processor cores.

[0093] The second HMD 20 includes a second camera 21 located outside the second HMD 20 that captures the direction of the second user 2's gaze (forward), a second display 22 facing the second user 2 that displays a second image 27 captured by the second camera 21, a second sensor 23 that detects the second user 2's operation information such as gestures, gaze, and voice, a second communication device 24 that can communicate with the first communication device 14 of the first HMD 10, a second processor 25 that controls the second camera 21, the second display 22, the second sensor 23, and the second communication device 24, and a second memory 26 that stores programs executed by the second processor 25 and various information. The first HMD 10 and the second HMD 20 can use HMDs with the same structure.

[0094] The second camera 21 is comprised of a shooting device that captures the direction of the second user 2's line of sight (forward).

[0095] The second display 22 is the part that projects images onto the second user 2. The second display 22 may use a panel such as liquid crystal or organic EL, or it may project images directly onto the retina using a light source such as a laser. The second display 22 is often mounted so that it is positioned in front of the second user 2's eyes.

[0096] The second sensor 23 is used to track the gaze, posture, movements (gestures), and / or voice of the second user 2. The second sensor 23 may be, for example, an accelerometer, gyroscope, infrared sensor, depth sensor, microphone, camera, etc. The second sensor 23 may also detect head movement or position.

[0097] The second communication device 24 is implemented by various communication circuits that communicate with other devices and server devices.

[0098] The second processor 25 is implemented by one or more CPUs, GPUs, processing circuits, etc., which may consist of one or more processor cores.

[0099] Figure 6A is a functional configuration diagram of the first HMD 10 according to the first embodiment of the present disclosure. Figure 6B is a functional configuration diagram of the second HMD 20 according to the first embodiment of the present disclosure. Each functional unit named "○○ unit" corresponds to the functions performed by the first processor 15 and the second processor 25 of the first HMD 10 and the second HMD 20 when they execute a display program.

[0100] The first HMD 10 has the following functional configuration: the first HMD 10 includes a first image acquisition unit 41, a first image display unit 42, a specific information acquisition unit 43, a change information acquisition unit 44, a first object identification unit 45, a first change image display unit 46, and a transmission unit 47. The second HMD 20 has the following functional configuration: the second HMD 20 includes a second image acquisition unit 51, a second image display unit 52, a second object identification unit 53, a second change image display unit 54, and a receiving unit 55.

[0101] The first image acquisition unit 41 of the first HMD 10 acquires a first image 17 of the area in front of the first user 1, which is captured by the first camera 11. The first image display unit 42 displays the first image 17 acquired by the first image acquisition unit 41 on the first display 12. The identification information acquisition unit 43 acquires information for identifying an object shown in the first image 17 based on the operation information of the first user 1 detected by the first sensor 13. The change information acquisition unit 44 acquires change information for the identified object based on the operation information of the first user 1 detected by the first sensor 13. The first object identification unit 45 identifies a predetermined object from among the objects shown in the first image 17 based on the identification information. The first changed image display unit 46 changes the image of the identified object based on the change information and displays it on the first display 12. The transmission unit 47 transmits the identification information and the change information to the receiving unit 55 of the second HMD 20.

[0102] The second image acquisition unit 51 of the second HMD 20 acquires a second image 27 of the area in front of the second user 2, which is captured by the second camera 21. The second image display unit 52 displays the second image 27 acquired by the second image acquisition unit 51 on the second display 22. The receiving unit 55 receives identification information and modification information from the first HMD 10. The second object identification unit 53 identifies a predetermined object from among the objects shown in the second image 27 based on the identification information. The second modified image display unit 54 modifies the image of the identified object based on the modification information and displays it on the second display 22.

[0103] In the first embodiment, the first HMD 10 is configured to transmit specific information and modified information based on the operation of the first user 1 to the second HMD 20. However, the first HMD 10 and the second HMD 20 may each be equipped with all the functions of both HMDs, so that specific information and modified information can be sent from either the first HMD 10 or the second HMD 20 to the other HMD. That is, in addition to the first image acquisition unit 41, first image display unit 42, specific information acquisition unit 43, modified information acquisition unit 44, first object identification unit 45, first modified image display unit 46, and transmission unit 47 functions of the first HMD 10, the first HMD 10 may also be equipped with the second image acquisition unit 51, second image display unit 52, second object identification unit 53, second modified image display unit 54, and receiving unit 55 functions. Furthermore, in addition to the second image acquisition unit 51, second image display unit 52, second object identification unit 53, second modified image display unit 54, and receiving unit 55, the second HMD 20 may also be equipped with the following functional units: first image acquisition unit 41, first image display unit 42, identification information acquisition unit 43, modified information acquisition unit 44, first object identification unit 45, first modified image display unit 46, and transmission unit 47.

[0104] Figure 7 is a flowchart of the first HMD 10 and the second HMD 20 according to the first embodiment of the present disclosure. Figure 7 is also a flowchart of the processing performed by the processors (first processor 15 and second processor 25) of the display system 200 of the present disclosure according to the display program.

[0105] The first processor 15 of the first HMD 10 acquires the first image 17 captured by the first camera 11 (step S11). The first processor 15 then displays the first image 17 on the first display 12 (step S12), allowing the first user 1 to view the image showing what is in front of them.

[0106] Next, the first user 1 looks at the first display 12 and selects an object in the image displayed there, and decides on a modification operation to that object. Then, by inputting identification information for identifying the object and modification information for modifying the object using the first sensor 13, the first processor 15 acquires the identification information and modification information (steps S13, S14). In other words, the first processor 15 acquires identification information and modification information based on the operation of the first user 1.

[0107] Next, the first processor 15 identifies an object from the image displayed on the first display 12 (more specifically, within the first image 17) using image recognition technology based on the identified information (step S15). Then, the first image 17, in which a modification operation has been performed on the identified object based on the modification information, is displayed on the first display 12 (step S16). The modification operation refers to an operation that makes a change to the identified object based on the modification information, for example, an operation that changes the color of the object from white to black and displays it. The modification operation is not limited to changing the display color, and may also include, for example, changing the position, orientation, or size of the object, deforming or replacing the shape of the object, adding or deleting an object, or highlighting an object (flashing, marking, etc.). Furthermore, the modification information may include area information indicating the area to be modified and information indicating the processing content to be applied to that area. In addition, the determination of whether or not to perform the modification processing may be made based on whether or not the target object is traceable, or whether or not the application of the modification processing is permitted, in addition to the mode setting.

[0108] Next, the first processor 15 transmits the specific information and the change information to the second HMD 20 (step S17).

[0109] The second processor 25 of the second HMD 20 acquires the second image 27 captured by the second camera 21 (step S21). The second processor 25 then displays the second image 27 on the second display 22 (step S22), allowing the second user 2 to view the image showing what is in front of them.

[0110] Next, the second processor 25 determines whether or not to perform modification processing on the second image 27 based on the specific information and modification information transmitted from the first HMD 10 (step S23). The determination of whether or not to perform modification processing may be made by, for example, the second user 2 setting a mode in advance for whether or not to perform modification processing, and obtaining that setting. If modification processing is not performed (step S23 is No), the second processor 25 does not perform modification processing, and the second image 27 is displayed as is on the second display 22. Alternatively, if modification processing is not performed, the second processor 25 may notify the first HMD 10 that modification processing will not be performed.

[0111] If a modification process is to be performed (step S23 is Yes), the second processor 25 receives and acquires specific information and modification information from the first HMD 10 (step S24). Then, based on the specific information, the second processor 25 uses image recognition technology to determine whether or not an object (target object) exists in the second image 27 displayed on the second display 22 (step S25). If it is determined that an object exists in the second image 27 (step S25 is Yes), the second processor 25 displays the modified second image 27 on the second display 22, which has been modified based on the object identified in the modification information (step S26).

[0112] On the other hand, if there is no object within the field of view of the second camera 21, or if there is an object within the field of view but it is hidden by another object, or if it cannot be determined to be the same object depending on the viewing angle, etc., and it cannot be determined that an object exists in the second image 27 (step S25 is No), the following processing is performed. That is, in this case, the determination process is repeated until it can be determined that an object exists in the second image 27, for example, by changing the orientation of the second HMD 20. At this time, the second processor 25 may notify the first HMD 10 that it cannot determine that an object exists in the second image 27, and in response to this notification, the first processor 15 may acquire identification information and change information again and send them to the second HMD 20. Alternatively, the second processor 25 may prompt the second user 2 to change the orientation of the second HMD 20, for example by displaying on the second display 22 that it cannot identify the object and to change the orientation of the second HMD 20. Furthermore, if there are multiple candidate objects in the second image 27, the second processor 25 may notify the second user 2 by marking the multiple objects in the second image 27, and the second user 2 may select and identify an object from the multiple objects.

[0113] In the process shown in Figure 7, the display images on either the first HMD 10 or the second HMD 20, or both, may be switched between the original image and the modified image at the discretion of either the first user 1 or the second user 2. Alternatively, control rights for image display may be set so that only one of the first user 1 or the second user 2 can switch the images displayed on either one or both HMDs.

[0114] Figure 8 shows a software architecture according to the first embodiment of this disclosure. As shown in Figure 8, each HMD has multiple layers, and Figure 8 shows how communication and control are performed via API (Application Programming Interface: an interface (function) that connects software, programs, or web services). The first HMD 10 and the second HMD 20 each have a configuration that includes three logical layers: application / UI layers 101 and 201, software layers 102 and 202, and hardware layers 103 and 203, respectively. Communication or control between these layers and between each HMD is achieved via API.

[0115] The application UI layers 101 and 201 are layers that provide an interface with the user, accepting user input and presenting output information visually or in other formats. These layers transmit operation commands to the software layers 102 and 202 via APIs.

[0116] Software layers 102 and 202 function as intermediate layers that link the application / UI layers 101 and 201 with the hardware layers 103 and 203. These layers receive commands from the application / UI layers 101 and 201 via APIs and control the hardware layers 103 and 203 as needed. They also process data acquired from the hardware layers 103 and 203 and transmit it to the application / UI layers 101 and 201.

[0117] Hardware layers 103 and 203 are layers that include the physical structure and electronic components of the HMD, including sensors, displays, communication devices, etc. These layers receive control commands from software layers 102 and 202 via APIs and execute corresponding operations.

[0118] In the sequence diagram described below, "device" corresponds to hardware layers 103 and 203, and "software" corresponds to application / UI layers 101 and 201 and / or software layers 102 and 202. The operation of "software" may be an operation within a single layer or an operation between different layers. Furthermore, the information exchanged between "software" and "device" may be information exchanged with a third layer such as the application layer or middleware layer.

[0119] Figure 9 is a block diagram showing the functional configuration of the display system 200 according to the first embodiment of the present disclosure. Figure 9 is also a block diagram showing the functional configurations of the first HMD 10 and the second HMD 20 of the display system 200 according to the present disclosure.

[0120] The display system 200 includes a first HMD 10 and a second HMD 20.

[0121] The first HMD 10 includes a sensor 111, an input receiving unit 112, a control information generation unit 113, a region identification unit 114, an image generation unit 115, a transmission unit 116, and a display unit 117. The second HMD 20 includes a receiving unit 211, a sensor 212, a region identification unit 213, an image generation unit 214, and a display unit 215.

[0122] Figure 10 is a sequence diagram according to the first embodiment of the present disclosure. Figure 10 is also a sequence diagram showing the operation of the software executed by the respective devices (hardware) and processors (first processor 15 and second processor 25) of the first HMD 10 and the second HMD 20.

[0123] The sensor 111 (first camera 11) of the first HMD 10 captures images in front of the first HMD 10, and the captured data acquired by the sensor 111 is obtained. In addition, the sensor 212 (second camera 21) of the second HMD 20 captures images in front of the second HMD 20, and the captured data acquired by the sensor 212 is obtained.

[0124] The input receiving unit 112 (first sensor 13) of the first HMD 10 acquires a change instruction from the first user 1. The change instruction from the first user 1 corresponds to the operation of the first user 1 described above. Based on the captured data and the change instruction, the control information generation unit 113 interprets the content of the change instruction to identify the area to be changed and to identify the content of the change process (311). The control information generation unit 113 then generates control information indicating the target area and the content of the process (311), and the control information is sent to the transmission unit 116. In other words, here, the control information (specific information and change information) is generated based on the operation of the first user 1. The target area indicated by the control information indicates the object specified by the specific information (for example, the roof 33 of the model 31 (car)), and the content of the process indicated by the control information is the content of the change indicated by the change information (for example, "change the color from white to black").

[0125] Furthermore, when the region identification unit 114 identifies a target region in the image, the image generation unit 115 generates an image (AR display image) to be displayed on the presentation unit 117 (first display 12) of the first HMD 10 based on the processing content (312), and sends the generated image to the transmission unit 116 and the presentation unit 117. The presentation unit 117 displays the AR display image superimposed on the captured data captured by the sensor 111.

[0126] The transmitting unit 116 transmits control information and an AR display image to the second HMD 20, and the receiving unit 211 (second communication device 24) of the second HMD 20 receives the transmitted control information and the AR display image. Based on the control information and the AR display image received by the receiving unit 211, the area identification unit 213 interprets the control information to identify the area to be changed and to identify the content of the change process (313). Then, the image generation unit 214 generates an image (AR display image) to be displayed on the presentation unit 215 (second display 22) of the second HMD 20 based on the target area and the content of the process (313), and sends it to the presentation unit 215. The presentation unit 215 displays the AR display image superimposed on the captured data captured by the sensor 212.

[0127] As described above, when the first user 1 wearing the first HMD 10 performs an operation to modify a specific object in the first image 17, the first image 17 displayed on the first display 12 of the first HMD 10 is modified, and the second image 27 displayed on the second display 22 of the second HMD 20 is also modified. As a result, the first user 1 and the second user 2 can simultaneously view and share the modified images.

[0128] (Second Embodiment) The method for identifying and sharing the same object between the first HMD 10 and the second HMD 20 is not limited to image recognition technology on each display, as in the first embodiment. In the second embodiment, the first HMD 10 generates a 3D model of the object, transmits the generated 3D model as identification information to the second HMD 20, and the second processor 25 identifies the object corresponding to the 3D model in the second image 27 as the object to be modified.

[0129] Figures 11 to 14 are diagrams illustrating the procedure for 3D modeling and identifying an object according to the second embodiment of this disclosure.

[0130] In Figure 11, when the first user 1 identifies an object in the first image 17, the first processor 15 generates a 3D model 61 of that object with its position information in the virtual space 6. The 3D model 61 may also be generated by scanning the object with the first camera 11. When the first processor 15 transmits the information of the 3D model 61 to the second HMD 20, the second processor 25 displays the 3D model 62 superimposed on the second image 27 on the second display 22 (Figure 12). When the second user 2 manipulates the position of the 3D model 62 on the second image 27 to superimpose it on the corresponding object in the second image 27 (Figure 13), the second processor 25 identifies that object as the object identified by the first user 1. The second processor 25 then changes the color of the object (the roof 36 of the model 32 (car)) from white to black according to the change information (Figure 14).

[0131] Figure 15 is a block diagram showing the functional configuration of the display system 200b according to the second embodiment of this disclosure. Figure 15 is also a block diagram showing the functional configurations of the first HMD 10 and the second HMD 20 of the display system 200b according to the second embodiment.

[0132] The display system 200b includes a first HMD 10 and a second HMD 20 according to the second embodiment.

[0133] The first HMD 10 according to the second embodiment includes a sensor 121, an input receiving unit 122, a control information generation unit 123, a region identification unit 124, an image generation unit 125, a transmission unit 126, a display unit 127, and a 3D model generation unit 128. The second HMD 20 according to the second embodiment includes a receiving unit 221, a sensor 222, a region identification unit 223, an image generation unit 224, and a display unit 225.

[0134] Figure 16 is a sequence diagram according to a second embodiment of the present disclosure. Figure 16 is also a sequence diagram showing the operation of the software executed by the respective devices (hardware) and processors (first processor 15 and second processor 25) of the first HMD 10 and the second HMD 20.

[0135] The sensor 121 (first camera 11) of the first HMD 10 captures images in front of the first HMD 10, and the captured data acquired by the sensor 121 is obtained. In addition, the first user 1 inputs an instruction to generate a 3D model of the target object (hereinafter sometimes referred to as a 3D model generation instruction) from the sensor 121 (first sensor 13) via voice or gesture. In addition, the sensor 222 (second camera 21) of the second HMD 20 captures images in front of the second HMD 20, and the captured data acquired by the sensor 222 is obtained.

[0136] The input receiving unit 122 (first sensor 13) of the first HMD 10 acquires a change instruction from the first user 1. Based on the acquired shooting data, 3D model generation instruction, and change instruction, the 3D model generation unit 128 generates a 3D model (321), and the control information generation unit 123 interprets the content of the change instruction to identify the area to be changed and to identify the content of the change process (321). The control information generation unit 123 then generates control information indicating the target area and the content of the process (321), and the control information is sent to the transmission unit 126 along with the 3D model.

[0137] Furthermore, when the region identification unit 124 identifies a target region in the image, the image generation unit 125 generates an image (AR display image) to be displayed on the presentation unit 127 (first display 12) of the first HMD 10 based on the processing content (322), and sends the generated image to the transmission unit 126 and the presentation unit 127. The presentation unit 127 displays the AR display image superimposed on the captured data captured by the sensor 121.

[0138] The transmitting unit 126 transmits control information, a 3D model, and an AR display image to the second HMD 20, and the receiving unit 221 (second communication device 24) of the second HMD 20 receives the transmitted control information, 3D model, and AR display image. Based on the control information, 3D model, and AR display image received by the receiving unit 221, the area identification unit 223 interprets the control information to identify the area to be modified and to identify the content of the modification process (323). Then, the image generation unit 224 generates an image (AR display image) to be displayed on the presentation unit 225 (second display 22) of the second HMD 20 based on the target area and the content of the process (323), and sends it to the presentation unit 225. The presentation unit 225 displays the AR display image superimposed on the captured data captured by the sensor 222.

[0139] In this way, the second processor 25 can identify the object through the manipulation of the 3D model by the second user 2, and therefore can accurately identify the object to be modified.

[0140] In the example above, a 3D model of an object was used as the identifying information. However, in addition to or instead of a 3D model, the target object may be captured from, for example, the first image 17 displayed on the first display 12 of the first HMD 10, and this captured object image and its annotations may be transmitted to the second HMD 20 as identifying information. When the second processor 25 displays the object image and its annotations superimposed on the second image 27 of the second display 22, the second user 2 can manipulate the position of the object image on the second image 27 to superimpose it onto the corresponding object in the second image 27, thereby allowing the second processor 25 to identify that object as the object identified by the first user 1.

[0141] (Third Embodiment) In the first embodiment, the first processor 15 and the second processor 25 each identify objects using image recognition technology, so depending on the position and / or direction of gaze of the first user 1 and / or the second user 2, it may not be possible to identify an object. Also, depending on the position, direction of gaze, and / or manner of modification of the first user 1 and / or the second user 2, it may not be possible to accurately modify the display of the object in the first image 17 (roof 33 of model 31) and the object in the second image 27 (roof 34 of model 32).

[0142] In the third embodiment, the first processor 15 and the second processor 25 share 3D models of objects that have been pre-generated or newly generated in space and display them on the first display 12 and the second display 22, respectively. When the first user 1 performs a modification operation on the 3D model while viewing the first image 17, the first processor 15 processes the modification on the 3D model and displays it on the first display 12, and also transmits the modification information to the 3D model to the second HMD 20. When the second HMD 20 receives the modification information, the second processor 25 displays the modified 3D model on the second display 22 according to the modification information.

[0143] With this configuration, by sharing the 3D model of an object and displaying it on the first display 12 and the second display 22, information about the identification or modification of an object can be easily shared, and the display of changes to the object can be processed accurately even when the first user 1 and / or the second user 2 move and their line of sight changes, or when there are complex changes.

[0144] (Fourth Embodiment) In the first embodiment, since the object in the first image 17 is identified based on the voice or gestures of the first user 1, the object may not be identified accurately depending on the accuracy of the voice recognition, the accuracy of the motion recognition, the position of the first user 1 and / or the direction of the first user 1's gaze, etc.

[0145] In the fourth embodiment, a list of objects present in the space (for example, a table summarizing the ID, name, and image (image information) of each object as shown in Figure 20) is created in advance and stored in the first memory 16 and the second memory 26, respectively. The first processor 15 recognizes whether or not there are any objects listed in the first image 17 displayed on the first display 12, and if so, displays that fact (for example, a bounding box surrounding the object and its name) on the first display 12. The first user 1 selects an object in the first image 17 while looking at the first display 12, and the first processor 15 transmits identification information (for example, the object's ID, name, and image) to the second HMD 20 to identify the selected object. That is, in the fourth embodiment, the identification information may include the ID of the identified object in the list of objects present in the space. The space in which the objects exist is the space in which the first user 1 and the second user 2 exist.

[0146] The second processor 25 recognizes whether or not an object listed in the list is present in the second image 27 displayed on the second display 22, and if so, displays a message to that effect (for example, a bounding box surrounding the object and its name) on the second display 22. The second processor 25 then searches the second image 27 for an object identified by specific information based on the image included in the specific information, and if so, displays a message to that effect (for example, by changing the display color of the bounding box and its name) on the second display 22. The second processor 25 may also modify the display image for the identified object based on the modification information.

[0147] Figure 17 is a block diagram showing the functional configuration of the display system 200c according to the fourth embodiment of this disclosure. Figure 17 is also a block diagram showing the functional configurations of the first HMD 10 and the second HMD 20 of the display system 200c according to the fourth embodiment.

[0148] The display system 200c includes a first HMD 10 and a second HMD 20 according to the fourth embodiment.

[0149] The first HMD 10 according to the fourth embodiment includes a sensor 131, an input receiving unit 132, a control information generation unit 133, a region identification unit 134, an image generation unit 135, a transmission unit 136, a display unit 137, and an ID list generation unit 138. The second HMD 20 according to the fourth embodiment includes a receiving unit 231, a sensor 232, a region identification unit 233, an image generation unit 234, and a display unit 235.

[0150] Figure 18 is a sequence diagram according to a fourth embodiment of the present disclosure. Figure 18 is also a sequence diagram showing the operation of the software executed by the respective devices (hardware) and processors (first processor 15 and second processor 25) of the first HMD 10 and the second HMD 20.

[0151] The sensor 131 (first camera 11) of the first HMD 10 captures images in front of the first HMD 10, and the captured data acquired by the sensor 131 is obtained. In addition, the first user 1 inputs an instruction to generate an ID list indicating the IDs of the target objects (hereinafter sometimes referred to as an ID list generation instruction) from the sensor 131 (first sensor 13) via voice or gesture. In addition, the sensor 232 (second camera 21) of the second HMD 20 captures images in front of the second HMD 20, and the captured data acquired by the sensor 232 is obtained.

[0152] The input receiving unit 132 (first sensor 13) of the first HMD 10 acquires a change instruction from the first user 1. Based on the acquired shooting data, ID list generation instruction, and change instruction, the ID list generation unit 138 generates an ID list (331), and the control information generation unit 133 interprets the contents of the change instruction to identify the area to be changed and to identify the contents of the change process (331). The control information generation unit 133 then generates control information indicating the target area and the contents of the process (331), and the control information is sent to the transmission unit 136 along with the ID list. The specific information included in this control information includes the ID of the identified object.

[0153] Furthermore, when the region identification unit 134 identifies a target region in the image, the image generation unit 135 generates an image (AR display image) to be displayed on the presentation unit 137 (first display 12) of the first HMD 10 based on the processing content (332), and sends the generated image to the transmission unit 136 and the presentation unit 137. The presentation unit 137 displays the AR display image superimposed on the captured data captured by the sensor 131.

[0154] The transmitting unit 136 transmits control information, an ID list, and an AR display image to the second HMD 20, and the receiving unit 231 (second communication device 24) of the second HMD 20 receives the transmitted control information, ID list, and AR display image. Based on the control information, ID list, and AR display image received by the receiving unit 231, the area identification unit 233 interprets the control information to identify the area to be changed and to identify the content of the change process (333). At this time, since the identification information contained in the control information includes the ID of the object, the second processor 25 can easily identify the object. Then, the image generation unit 234 generates an image (AR display image) to be displayed on the presentation unit 235 (second display 22) of the second HMD 20 based on the target area and the content of the process (333), and sends it to the presentation unit 235. The presentation unit 235 displays the AR display image superimposed on the captured data captured by the sensor 232. Once the ID list has been sent from the first HMD10 to the second HMD20, it does not need to be sent again.

[0155] With this configuration, a list of objects present in the space is created in advance, and the first processor 15 displays the objects it recognizes in the first image 17 on the first display 12, so the first user 1 can easily identify the objects. The second processor 25 then identifies the objects in the second image 27 based on the object's ID, name, image, etc., so it can identify the objects with high accuracy.

[0156] In the example above, the first HMD 10 generated the ID list, but this is not the only example. Figure 19 is a block diagram showing the functional configuration of a display system 200d according to a modified example of the fourth embodiment of this disclosure.

[0157] The display system 200d comprises a first HMD 10 and a second HMD 20 according to a modification of the fourth embodiment, and a server 100.

[0158] The first HMD 10 according to a modification of the fourth embodiment includes a sensor 131, an input receiving unit 132, a control information generation unit 133, a region identification unit 134, an image generation unit 135, a transmission unit 136, a display unit 137, and a receiving unit 140. The second HMD 20 according to the fourth embodiment includes a receiving unit 231, a sensor 232, a region identification unit 233, an image generation unit 234, a display unit 235, and a conversion unit 236. The server 100 according to a modification of the fourth embodiment includes an ID list generation unit 139.

[0159] In the configuration shown in Figure 19, the conversion unit 236 of the second HMD 20 performs conversion processing to align the ID list received by the receiving unit 231, and / or specific information, change information, control information, etc. transmitted from the first HMD 10, with an internal representation usable by the second HMD 20. For example, the conversion unit 236 may perform coordinate system conversion to convert coordinate information (position, orientation, vertices or boundaries of a region, etc.) included in the received information to the camera coordinate system, display coordinate system, or world coordinate system referenced by the second HMD 20. Furthermore, considering that identifiers assigned to each terminal may differ even for the same object, the conversion unit 236 may perform ID conversion to generate or update the correspondence between identifiers included in the ID list transmitted from the server 100 (e.g., global ID) and identifiers managed internally by the second HMD 20 (e.g., local ID).

[0160] Furthermore, the conversion unit 236 may acquire or estimate the position and orientation corresponding to the identifier contained in the received specific information in a predetermined coordinate system, so that the region identification unit 233 can re-identify the target region in the second image 27 based on said position and orientation. Here, the predetermined coordinate system is the local coordinate system H of the first HMD 10. 1 , local coordinate system H of the second HMD20 2 It may be either the world coordinate system W, and the transformation unit 236 may, if necessary, switch to the local coordinate system H 1 The position and orientation defined in the local coordinate system H 2Alternatively, convert to the world coordinate system W, or convert the position and orientation defined in the local coordinate system W to the local coordinate system H. 2 It may also be converted to [another format]. In addition, the conversion unit 236 may convert the received representation of the area to be modified (for example, information about a two-dimensional mask, rectangle, polygon, three-dimensional area, or three-dimensional shape) into another representation that is easy for the area identification unit 233 or the image generation unit 234 to handle. If the information arrives out of order due to communication delay or processing delay, the conversion unit 236 may rearrange, interpolate, or determine applicability of the information based on a timestamp or the like. Depending on the implementation, the processing of the conversion unit 236 may be integrated into the area identification unit 233 or the receiving unit 231, etc., or the conversion unit 236 may not be provided as an independent component.

[0161] As shown in Figure 19, an ID list generation unit 139 of an external server 100 connected to the first HMD 10 and the second HMD 20 may generate an ID list and transmit this ID list to the first HMD 10 and the second HMD 20, respectively. Note that communication between terminals or between terminals and the server 100 is not limited to internet communication, and short-range wireless communication may be used. For example, after the other terminal is detected and a connection is established via short-range wireless communication, the specific information, change information, control information, or ID list may be transmitted and received.

[0162] Figure 20 is an example of an ID list according to the fourth embodiment of this disclosure. Figure 21 is an example of processing using the ID list according to the fourth embodiment of this disclosure. Figure 20 shows an example of an ID list, and Figure 21 shows an overview of an example of processing a change on an object using the ID list. More specifically, Figure 21(a) is a diagram showing multiple objects existing in real space, Figure 21(b) is a diagram showing an example in which bounding boxes and labels are displayed, Figure 21(c) is a diagram showing a concealment image 405 in the first HMD 10, and Figure 21(d) is a diagram showing a concealment image 406 in the second HMD 20.

[0163] The ID list shown in Figure 20 assigns ID numbers 1 through 4 to the objects "heart" (401), "triangle" (402), "thunder" (403), and "moon" (404) that exist in real space (Figure 21(a)). In addition, the ID list shown in Figure 20 registers the name of each object as a label, and also registers an image (image information) of each object.

[0164] The first HMD 10 recognizes objects listed in the ID list from the image displayed on its first display 12 using image recognition technology, and displays a bounding box and a label for each object (Figure 21(b)). Based on this image, the first user 1 says, for example, "Hide the heart." When the first sensor 13 detects the first user 1's voice, the first processor 15 uses voice recognition technology to superimpose a concealment image 405 onto the area where the "heart" is displayed, thereby hiding the "heart" (Figure 21(c)).

[0165] Then, control information is sent to the second HMD 20 indicating that a concealing image should be superimposed on the "heart". The second HMD 20 retrieves the image of the "heart" from the ID list and recognizes the "heart" from the image on the second display 22. By superimposing the concealing image 406 onto the recognized "heart", the "heart" is concealed in the image on the second display 22 (Figure 21(d)). As a result, when the first user 1 performs a modification operation on the "heart", the image on the second display 22 is changed.

[0166] In the example above, the ID list contains image information of an object, but this image information may be multiple images of the object taken from multiple viewpoints, or a 2D or 3D model of the object generated based on these multiple images. Furthermore, in addition to or instead of image information, a barcode attached to the object may also be included. This improves the speed or accuracy of the second processor 25 in searching for an object.

[0167] Alternatively, object tracking technology may be applied instead of an ID list to identify objects. When the first user 1 identifies an object in the first image 17, the first processor 15 acquires the object's position coordinates or features (size, shape, color, 3D model, etc.) and transmits them to the second HMD 20 as identification information. The second processor 25 identifies the object in the second image 27 based on the received position coordinates or features. In other words, the second HMD 20 tracks the object recognized by the first HMD 10. This allows the first HMD 10 and the second HMD 20 to independently identify and share objects in real time.

[0168] Furthermore, the positional coordinate feature may include an identifier (such as an ID number or UUID) associated with the object or target area. Also, if the positional coordinate feature includes position and orientation, such position and orientation may be defined in the local coordinate system of the transmitting terminal, the local coordinate system of the receiving terminal, or the world coordinate system, and the receiving terminal may identify the object or target area in its own local coordinate system based on the received identifier and / or position and orientation.

[0169] (Fifth Embodiment) In the fifth embodiment, the local coordinate systems of multiple terminals are aligned in order to identify the object to be modified. That is, the object to be modified can be identified in the second HMD 20 from the position information of the object relative to the first HMD 10 and the relative position and orientation information of the first HMD 10 and the second HMD 20 based on the alignment result.

[0170] Normally, multiple devices operate in a local coordinate system with the origin being the location at the time of each device's startup, making it difficult to identify and share the same real-world object. By performing alignment, each device can accurately determine its position in real space based on the position information of the object 80 that is the target of modification, which is included in the identification information. In other words, each device can accurately identify the same position in real space.

[0171] Furthermore, in order to perform alignment, information used for alignment may be exchanged between multiple terminals, in addition to the specific information described above. The information used for alignment (alignment information) may be transmitted as part of the specific information or control information.

[0172] Below, we will explain three alignment methods, (1) to (3), but these are not the only methods available.

[0173] (1) Determine the relative position and orientation of the HMDs. Figure 22 is a diagram showing the coordinate transformation according to the fifth embodiment of this disclosure.

[0174] The first method uses the relative positional relationship between multiple terminals, i.e., relative coordinates, to determine the location of the same object in real space across multiple terminals. As an example, we will explain the process by which the second HMD 20 identifies the object specified by the first HMD 10 on the image displayed on the second HMD 20, based on the alignment information transmitted by the first HMD 10.

[0175] First, the first HMD 10 requests video data being captured by the second camera 21 of the second HMD 20 from the second HMD 20. In response to the request from the first HMD 10, the second HMD 20 transmits the video being captured by its own terminal (second HMD 20) to the first HMD 10. At this time, the video data may be transmitted directly to the first HMD 10, or it may be transmitted via the server 100 using a network API. In addition, the video data may be encoded in an appropriate format and converted into a format suitable for streaming before transmission.

[0176] Next, the first HMD 10 identifies the object 80 that is visible in both the first camera 11 and the second camera 21 of the second HMD 20. Specifically, the first HMD 10 uses image processing techniques to extract objects 80 with common features or shapes from within each video frame. The first HMD 10 then checks where the identified object 80 is located within the field of view of each HMD and determines its local coordinate system (local coordinate system H 1), information on the position of the first HMD 10, and the local coordinate system of the first HMD 10 (local coordinate system H 1 ), acquire information on the position of the second HMD 20. Thus, the position coordinate information (position information) of the second HMD 20 is obtained based on the position of each object (object 80) that is commonly projected on the image acquired by the first HMD 10 and the image acquired by the second HMD 20, and is acquired. In other words, the position coordinate information of another HMD (for example, the second HMD 20) in the first local coordinate system (local coordinate system H 1 ) is obtained and acquired in advance based on the position of each object (object 8) that is commonly projected on the image acquired by the first HMD 10 and the image acquired by another HMD.

[0177] Then, the first HMD 10 uses the acquired information on the position of the object 80 in its own local coordinate system H 1 to calculate the relative position and orientation between the first HMD 10 and the second HMD 20. For example, the first HMD 10 performs this calculation using an 8-point algorithm. This calculation clarifies the spatial relationship between the HMDS and enables adjustment of the local coordinate system.

[0178] Next, the first HMD 10 transmits to the second HMD 20 the information on the position of the object 80 that is the modification target identified in its own local coordinate system H 1 and the information on the position of the second HMD 20 on the local coordinate system H 1 (distance and / or direction from the first HMD 10). The information on the position of the object 80 that is the modification target identified in the local coordinate system H 1 and the information on the position of the second HMD 20 on the local coordinate system H 1 are an example of alignment information. As described above, the alignment information is part of the control information. Therefore, the specific information included in the control information is the first local coordinate system (local coordinate system H 1) includes position coordinate information (position information) of the object (object 80) in ). Furthermore, the specific information included in the control information is a first local coordinate system (local coordinate system H) with the first HMD 10 of the transmitting side as the origin. 1 This includes the position coordinate information (position information) of the object (object 80) in the first local coordinate system, and the position coordinate information (position information) of the receiving side's second HMD 20 in the first local coordinate system.

[0179] In this case, for example, if the object 80 to be modified is a wall in the room where the first user 1 is located, the position of the wall may be determined by obtaining the shape of the room using multiple sensors provided on the first HMD 10.

[0180] Finally, the second HMD20 receives the transmitted alignment information and, based on the received alignment information, sets its own local coordinate system H 2 The position of object 80 is identified above.

[0181] Local coordinate system H fixed to the first HMD10 1 The position coordinates of the object (object 80) can be determined, for example, from parallax information based on the position of the object (object 80) on the first display 12 and the position of the object (object 80) on the second display 22.

[0182] According to this method, each HMD independently uses its own local coordinate system (for example, local coordinate system H 1 or H 2 Because it can operate in a shared world coordinate system, it has the advantage of high system configuration flexibility, as it does not require the establishment of a shared world coordinate system or the prior generation of a common map. In addition, high-precision alignment is possible in environments with abundant feature points.

[0183] Furthermore, it can be said that the following processing is performed in the alignment method of (1). That is, the specific information is placed in a first local coordinate system (local coordinate system H) with the first HMD 10 of the transmitting side as the origin. 1The first local coordinate system is the position coordinate information (position information) of the object (object 80) in the first local coordinate system, and the position coordinate information of the receiving side's second HMD 20 in the first local coordinate system. The second processor 25 then converts the first local coordinate system to a second local coordinate system (local coordinate system H) with the position of the second HMD 20 as the origin. 2 The system converts the first position coordinate information to a second position coordinate information in a second local coordinate system, and identifies objects in the image based on the second position coordinate information.

[0184] (2) Each HMD aligns its local coordinate system with the existing world coordinate system. Figure 23 shows a coordinate transformation according to the fifth embodiment of this disclosure.

[0185] The second method involves multiple terminals identifying the location of the same object in real space using information about the position and orientation of each terminal in the world coordinate system of each terminal, and the position information of the target object in the world coordinate system. As an example, the process flow of identifying the object specified by the first HMD 10 on the image displayed on the second HMD 20, based on the position coordinates of the object in the world coordinate system W transmitted by the first HMD 10, will be explained.

[0186] The first HMD 10 identifies an object 80 that serves as the reference object for the world coordinate system W on its own image. Then, based on the identified position, the first HMD 10 calculates its own position and orientation in the world coordinate system W using methods such as solving a PnP problem.

[0187] Similarly, the second HMD 20 identifies object 80, which serves as the reference object for the world coordinate system W, on its own image. Then, based on the identified position, the second HMD 20 calculates its own position and orientation in the world coordinate system W using methods such as solving a PnP problem.

[0188] Next, based on the operation of the first user 1, the first HMD 10 moves the object 80, which is the object to be modified on its own image, to the local coordinate system H fixed to the first HMD 10. 1The location is then identified. The position and orientation information of the first HMD10 in the world coordinate system W and the local coordinate system H fixed to the first HMD10 are then used. 1 Based on the position information of the object 80 to be modified, the first HMD 10 calculates the position information of the object 80 to be modified in the world coordinate system W.

[0189] Next, the first HMD 10 transmits information about the position of the object 80 to be modified in the world coordinate system W to the second HMD 20. Local coordinate system H 1 The position information of object 80 identified is an example of alignment information. As described above, since alignment information is part of control information, the specific information included in the control information includes the position coordinate information (position information) of the object (object 80) in the world coordinate system W fixed in space.

[0190] The second HMD 20 uses the position information of the object 80 to be modified in the world coordinate system W and the position and orientation information of the second HMD 20 in the world coordinate system W to determine the local coordinate system H fixed to the second HMD 20. 2 The second HMD 20 calculates the position information of the object 80 that is to be modified. Based on this position information, the second HMD 20 identifies the object 80 that is to be modified on the image of the second HMD 20.

[0191] With this method, once alignment is complete, all terminals can operate in the same coordinate system, making it easy to share alignment information. Furthermore, calculations can be performed quickly, allowing for real-time information sharing.

[0192] (3) Align the world coordinate system and the local coordinate system of the other HMD with the local coordinate system of one HMD. Figure 24 is a diagram showing the coordinate transformation according to the fifth embodiment of this disclosure.

[0193] The third method involves multiple terminals identifying the location of the same object in real space based on the world coordinate system in the local coordinate system of one terminal and the position and orientation information of other terminals. As an example, the process flow for the second HMD 20 to identify the object 80 to be modified on the image displayed on the second HMD 20 will be explained, based on the position coordinates of the object 80 to be modified in the local coordinate system transmitted by the first HMD 10.

[0194] The first HMD 10 identifies an object 80 (more specifically, the position of object 80) that serves as a reference for the world coordinate system W on its own image. Then, based on the identified position, the first HMD 10 calculates its own position and orientation in the world coordinate system W using methods such as solving a PnP problem. Note that the estimation of the terminal's position and orientation is not limited to solving a PnP problem, but may also be performed by self-position estimation based on sensor information (for example, estimation combining inertial measurement and image measurement), or by using estimation results provided by the terminal's execution environment.

[0195] Similarly, the second HMD20 identifies object 80 (more specifically, the position of object 80) on its own image, which serves as the reference for the world coordinate system W. Then, based on the identified position, the second HMD20 calculates its own position and orientation in the world coordinate system W using methods such as solving a PnP problem.

[0196] Next, the second HMD20 uses the position and orientation information of the first HMD10 in the world coordinate system W and the position and orientation information of the second HMD20 in the world coordinate system W to determine the local coordinate system H of the second HMD20 fixed to the first HMD10. 1 Calculate the position and orientation at [location].

[0197] Next, based on the operation of the first user 1, the first HMD 10 fixes the object 80, which is the object to be modified on its own image, to the first HMD 10 in a local coordinate system H 1 Identify it based on that.

[0198] Then, the first HMD 10 is fixed to the first HMD 10 of the object 80 that is to be modified, and the local coordinate system H1 The position information in the local coordinate system H is transmitted to the second HMD20. 1 The position information of object 80, which was identified as the object to be modified, is an example of alignment information.

[0199] The second HMD20 is fixed to the local coordinate system H of the first HMD10 of the object 80 that is to be modified. 1 Information on the position in and the local coordinate system H fixed to the first HMD10 of the second HMD20. 1 Based on the position and orientation information in the second HMD20, the local coordinate system H is fixed to the second HMD20. 2 The position of the object 80 to be modified is calculated. The second HMD 20 then identifies the object 80 to be modified on the image of the second HMD 20 based on the calculated position information of the object 80.

[0200] This method simplifies setup because alignment is performed using a single device as a reference. Furthermore, it facilitates alignment between devices, enabling highly accurate positioning.

[0201] Figure 25 is a block diagram showing the functional configuration of the display system 200f according to the fifth embodiment of the present disclosure. Figure 25 is also a block diagram showing the functional configurations of the first HMD 10 and the second HMD 20 of the display system 200f according to the present disclosure.

[0202] The display system 200f includes a first HMD 10 and a second HMD 20 according to the fifth embodiment.

[0203] The first HMD 10 according to the fifth embodiment includes a sensor 141, an input receiving unit 142, a control information generation unit 143, a region identification unit 144, an image generation unit 145, a transmission unit 146, and a display unit 147. The second HMD 20 according to the fifth embodiment includes a receiving unit 241, a sensor 242, a region identification unit 243, an image generation unit 244, a display unit 245, and a conversion unit 246.

[0204] The conversion unit 246 of the second HMD 20 shown in Figure 25 adjusts the region information contained in the control information (including the target region and processing content) received by the receiving unit 241 to an internal representation usable by the second HMD 20. For example, the conversion unit 246 adjusts the received region information (location, boundary, or shape of the region, etc.) to the local coordinate system H of the second HMD 20. 2 Alternatively, the region information may be converted to a display coordinate system so that the region identification unit 243 can identify the target region within the second image. Furthermore, if the representation of the region information is unsuitable for processing, the conversion unit 246 may convert the region information to another representation (rectangle, polygon, or mask, etc.). Depending on the implementation, the processing of the conversion unit 246 may be integrated into the region identification unit 243 or the receiving unit 241, etc.

[0205] Figure 26 is a sequence diagram according to a fifth embodiment of the present disclosure. Figure 26 is also a sequence diagram showing the operation of the software executed by the respective devices (hardware) and processors (first processor 15 and second processor 25) of the first HMD 10 and the second HMD 20.

[0206] The sensor 141 (first camera 11) of the first HMD 10 captures images in front of the first HMD 10, and the captured data acquired by the sensor 141 is recorded. In addition, the sensor 242 (second camera 21) of the second HMD 20 captures images in front of the second HMD 20, and the captured data acquired by the sensor 242 is recorded.

[0207] The input receiving unit 142 (first sensor 13) of the first HMD 10 acquires a change instruction from the first user 1. Based on the acquired image data and the change instruction, the control information generation unit 143 interprets the content of the change instruction to identify the area to be changed and to identify the content of the change process (341). The control information generation unit 143 then generates control information indicating the target area and the content of the process (341), and the control information is sent to the transmission unit 146. At this time, the control information includes location information of the target area.

[0208] Furthermore, the control information may include, in addition to the processing details related to the change instruction, information used to identify the target (such as an identifier or coordinate system) or information about the status of the target, for example, a tracking status (such as trackable / untrackable), a sharing status (such as shareable / currently shared), or a persistence status (such as saved, deleted, or expiration date). This allows the receiving terminal to re-identify the target area, apply / suppress the change processing, or perform resynchronization processing according to the status.

[0209] Furthermore, when the region identification unit 144 identifies a target region in the image, the image generation unit 145 generates an image (AR display image) to be displayed on the presentation unit 147 (first display 12) of the first HMD 10 based on the processing content (342), and sends the generated image to the transmission unit 146 and the presentation unit 147. The presentation unit 147 displays the AR display image superimposed on the captured data captured by the sensor 141.

[0210] The transmitting unit 146 transmits control information and an AR display image to the second HMD 20, and the receiving unit 241 (second communication device 24) of the second HMD 20 receives the transmitted control information and the AR display image. Based on the control information and the AR display image received by the receiving unit 241, the area identification unit 243 interprets the control information to identify the area to be modified and to identify the content of the modification process (343). At this time, the image generation unit 244 converts the position information of the object included in the control information into a local coordinate system fixed to the second HMD 20 based on the methods (1) to (3) described above to identify the target area (343). Then, based on the target area and the content of the process, it generates an image (AR display image) to be displayed on the presentation unit 245 (second display 22) of the second HMD 20 (343) and sends it to the presentation unit 245. The presentation unit 245 displays the AR display image superimposed on the captured data captured by the sensor 242.

[0211] With this configuration, objects present in space are identified based on their positional information, allowing for highly accurate object identification.

[0212] In this embodiment, if the target object (object 80) is not within the image range of the second display 22, the second processor 25 does not perform any modification processing on the displayed image. Alternatively, the second processor 25 may inform the second user 2 of the location of the target object by displaying on the second display 22, for example, "The target object is on the right." Then, when the second user 2 changes the direction of their face or otherwise brings the target object into the image range of the second display 22, the second processor 25 may perform modification processing on the target object. This allows the object to be modified to be quickly shared between the first user 1 and the second user 2.

[0213] Alternatively, instead of performing a three-dimensional coordinate system transformation as described in (1) to (3) above, the second processor 25 may project the position coordinates of the target object onto the second display 22 and identify the target object on the second image 27.

[0214] (Sixth Embodiment) In the sixth embodiment, a scenario is assumed in which two HMDs (Terminal A and Terminal B) exist in the same real space, and each image generation unit modifies and displays a portion of the real space it has captured. The user of Terminal A specifies an object (e.g., a part) in their field of view and instructs Terminal A to change its color. Terminal A identifies the area of ​​the specified object and generates and displays a virtual object (e.g., an image of the same shape but a different color) superimposed on that area. Terminal A also generates control information by marking the area of ​​the specified object on an image of its field of view and sends it to Terminal B. Terminal B transforms the received control information to match its own coordinate system or viewpoint and identifies the area of ​​the same object. Terminal B then generates and displays a virtual object in the same way. As a result, the user of Terminal A and the user of Terminal B can confirm that the color of the same object has changed.

[0215] Furthermore, consider the case where terminal C, located in a different space, presents images via a web conferencing system as if it were in the same space as terminal A. Terminal C receives control information transmitted from terminal A and transforms it to match its own coordinate system or viewpoint. Then, terminal C identifies the region of objects that are the same as those in terminal A's field of view and generates and displays virtual objects superimposed on that region. Terminal C also generates control information by capturing an image of its own field of view and marking the region of objects that are the same as those in terminal A's field of view, and transmits this generated control information to terminal A. Terminal A transforms the received control information to match its own coordinate system or viewpoint and identifies the region of objects that are the same as those in terminal C's field of view. Then, it generates and displays virtual objects superimposed on that region. As a result, the user of terminal A and the user of terminal C can get the feeling that they are in the same space through the web conference.

[0216] The above explanation focused on cases where VR or AR is provided using an HMD, but these are not the only functions that can be provided by an HMD. For example, an HMD may be used to present content such as 2D images. For instance, by displaying a large-screen image in front of the user's eyes using an HMD, a highly immersive viewing experience can be provided. Furthermore, the viewpoint or field of view of the image may be changed according to the movement or position of the head, or the direction of the gaze, using an HMD. This allows the user to freely look around, zoom in, and zoom out of the image. HMDs can be used for entertainment such as movies and games, or for learning such as education and training.

[0217] (Seventh Embodiment) In the seventh embodiment, multiple terminals located in the same space use short-range wireless communication to discover other terminals or nearby groups of participants, and share a group identifier (e.g., GUID or UUID) that identifies the discovered group of terminals. The first terminal registers information associated with a predetermined location in real space (hereinafter referred to as spatial fixed information) in association with the group identifier in a shareable manner, and notifies the second terminal of the identifier that identifies the spatial fixed information as specific information.

[0218] The second terminal acquires or estimates the position and orientation corresponding to the spatially fixed information based on the group identifier and the identifier, and determines the coordinate system in which the position and orientation are defined (for example, the world coordinate system W or the local coordinate system H of the second terminal). 2 In the second image, an object or target area is identified. The first terminal transmits, in addition to the identification information, change information or control information (processing content) for the object, and the second terminal performs the modification process of the second image based on this information.

[0219] Furthermore, the sharing of spatially fixed information, as well as the acquisition or estimation of position and orientation, may be performed asynchronously, and information regarding the sharing status (shareable / sharing in progress, etc.) or validity period of spatially fixed information may be included in change information or control information.

[0220] (Eighth Embodiment) In the eighth embodiment, an object (hereinafter referred to as a spatial element object) provided by the terminal's execution environment (runtime) that is associated with an element in real space is used. The spatial element object may include, for example, a plane, a marker, or an object that corresponds to information fixed at a predetermined position in real space. The first terminal acquires the identifier of the spatial element object as identifying information that identifies the object and transmits it to the second terminal along with change information or control information. The second terminal acquires or estimates the position and orientation corresponding to the same spatial element object based on the identifier and identifies the target area or object in the second image in the coordinate system in which the position and orientation are defined.

[0221] Furthermore, if management such as saving or deleting is performed with respect to the spatial element object, information regarding saving, deleting, or management operations may be sent and received as change information or control information.

[0222] In this specification, "spatial fixed information" refers to information that is associated with a predetermined location in real space and can be used in common by multiple terminals, and may include, for example, an identifier associated with the predetermined location, the location and orientation corresponding to the identifier, and information about the coordinate system in which the location and orientation are defined. Furthermore, "spatial element object" refers to an object associated with an element in real space managed by the terminal's execution environment, and may include, for example, a plane, a marker, or an object corresponding to the spatial fixed information. The process by which the receiving terminal makes the location and orientation available based on the identifier, etc., is called "acquisition or estimation."

[0223] (Other Embodiments) The following are examples of modifications of each of the above embodiments. In the above embodiments, the second HMD 20 modified the image based on specific information and modification information sent from the first HMD 10. Instead, the first HMD 10 may send the modified image of the object as modification information to the second HMD 20, and the second HMD 20 may superimpose the image received from the first HMD 10 onto the position of the object in the second image 27.

[0224] In the above embodiment, a display system having two head-mounted displays, a first HMD and a second HMD, was described, but the number of HMDs may be three or more. In that case, specific information and modified information generated by one of the HMDs may be transmitted to some or all of the other HMDs, and the image may be modified and displayed on some or all of the other HMDs.

[0225] The above embodiments or modifications can be implemented in any combination.

[0226] (Note) The following technologies are disclosed by the above embodiments and modifications.

[0227] (Technical Configuration Example 0) A first terminal having a processor and memory, wherein the processor is configured to use the memory to acquire identification information and modification information for objects in an image taken in front of the user, and to transmit the identification information and modification information to another head-mounted display.

[0228] (Technical Configuration Example 1) The first terminal described in Technical Configuration Example 0 is configured to receive list information generated by an external server as information used for generating or interpreting the specific information and the change information, and to transmit the said list information to other head-mounted displays.

[0229] (Technical Configuration Example 2) The first terminal described in Technical Configuration Example 0, wherein the change information includes information indicating the area to be changed in the image and information indicating the content of the processing to be performed on the area.

[0230] (Technical Configuration Example 3) A first terminal as described in Technical Configuration Example 0, wherein the specific information and the modified information are transmitted on the premise that they will be interpreted after undergoing a transformation process to align them with a predetermined internal representation including a coordinate system in another head-mounted display.

[0231] (Technical Configuration Example 4) The first terminal described in Technical Configuration Example 0, wherein the identification information includes, in addition to the position coordinates of the object, feature quantities of the object, and another head-mounted display tracks and identifies the object based on the feature quantities.

[0232] (Technical Configuration Example 5) A first terminal as described in Technical Configuration Example 0, wherein the degree of freedom in the type of information to be prepared prior to the transmission of the specific information and the change information is high because each head-mounted display operates independently in its respective local coordinate system.

[0233] (Technical Configuration Example 6) The first terminal described in Technical Configuration Example 0, wherein the change information further includes information regarding the processing content to be applied to the target area.

[0234] (Technical Configuration Example 7) A first terminal as described in Technical Configuration Example 0 or 6, wherein the change information includes status information regarding whether or not the object is traceable.

[0235] (Technical Configuration Example 8) A first terminal according to Technical Configuration Example 0, 6, or 7, wherein the first terminal communicates with the other head-mounted display by short-range wireless communication and transmits the specific information and the change information.

[0236] (Technical Configuration Example 9) The first terminal described in Technical Configuration Example 8, wherein the first terminal acquires or generates a group identifier corresponding to a group of participating terminals including the other head-mounted displays, and transmits the specific information and the change information in association with the group identifier.

[0237] (Technical Configuration Example 10) A first terminal described in any one of Technical Configuration Examples 0 and 6 to 9, wherein the first terminal asynchronously transmits at least a portion of the specified information and the modified information, and controls the transmission order based on a timestamp in response to communication delays or retransmissions.

[0238] As a display system that provides unprecedented AR technology, it can be suitably used in a wide range of fields, including business, science, engineering, design, fashion, food and beverage, leisure, tourism, sports, and games.

[0239] 1 First User 2 Second User 3, 31, 32 Model 6 Virtual Space 10 First HMD 11 First Camera 12 First Display 13 First Sensor 14 First Communication Device 15 First Processor 16 First Memory 17 First Image 20 Second HMD 21 Second Camera 22 Second Display 23 Second Sensor 24 Second Communication Device 25 Second Processor 26 Second Memory 27 Second Image 33, 34, 35, 36 Roof 41 First Image Acquisition Unit 42 First Image Display Unit 43 Identification Information Acquisition Unit 44 Change Information Acquisition Unit 45 First Object Identification Unit 46 First Changed Image Display Unit 47 Transmission Unit 51 Second Image Acquisition Unit 52 Second Image Display Unit 53 Second Object Identification Unit 54 Second Changed Image Display Unit 55 Receiving Unit 61, 62 3D Model 71 Processor 72 Memory 73 Communication Interface 74 Bus 80 Object 100 Server 101, 201 Application / UI Layer 102, 202 Software Layer 103, 203 Hardware Layer 111, 121, 131, 141 Sensor 112, 122, 132, 142 Input Reception Unit 113, 123, 133, 143 Control Information Generation Unit 114, 124, 134, 144 Area Identification Unit 115, 125, 135, 145 Image Generation Unit 116, 126, 136, 146 Transmission Unit 117, 127, 137, 147 Presentation Unit 128 D-Model Generation Unit 138, 139 ID List Generation Unit 140 Receiving Unit 200, 200a, 200b, 200c, 200d, 200f Display system 211, 221, 231, 241 Receiving unit 212, 222, 232, 242 Sensor 213, 223, 233, 243 Area identification unit 214, 224, 234, 244 Image generation unit 215, 225, 235, 245 Presentation unit 236, 246 Conversion unit

Claims

1. A display system comprising: a first head-mounted display having a first camera, a first display, and a first processor; and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor is configured to display a first image captured by the first camera on the first display, acquire identification information to identify an object in the first image based on user operation, acquire change information for the object based on user operation, display the first image modified based on the change information on the first display, and transmit the identification information and the change information to the second head-mounted display; and the second processor is configured to display a second image captured by the second camera on the second display, receive the identification information and the change information, identify the object in the second image based on the identification information, and display the second image modified based on the change information on the second display.

2. The display system according to claim 1, wherein the specific information includes information on a 3D model of the object.

3. The display system according to claim 1, wherein the identifying information includes the ID of the identified object in a list of objects present in space.

4. The display system according to claim 1, wherein the specific information includes the position coordinates of the object.

5. A first terminal having a processor and memory, wherein the processor is configured to use the memory to acquire identification information and modification information for objects in an image taken in front of the user, and to transmit the identification information and modification information to another head-mounted display.

6. The first terminal according to claim 5, wherein the specific information and the modified information are generated based on user operations.

7. The first terminal according to claim 5, further comprising a display, wherein the processor displays the modified image on the display based on the modification information.

8. The first terminal according to claim 5, wherein the first terminal is a first head-mounted display, and the specific information includes position coordinate information of the object in a first local coordinate system with the transmitting first head-mounted display as the origin, and position coordinate information of the receiving second head-mounted display in the first local coordinate system.

9. The first terminal according to claim 8, wherein the position coordinate information of the second head-mounted display is determined based on the respective image positions of objects that are displayed in common in the image acquired by the first head-mounted display and the image acquired by the second head-mounted display.

10. The first terminal according to claim 5, wherein the first terminal is a first head-mounted display, and the specific information includes position coordinate information of the object in a world coordinate system fixed in space.

11. The first terminal according to claim 5, wherein the first terminal is a first head-mounted display, and the specific information includes position coordinate information of the object in a first local coordinate system with the transmitting first head-mounted display as the origin.

12. The first terminal according to claim 11, wherein the position coordinate information of the other head-mounted display in the first local coordinate system is predetermined based on the respective image positions of objects that are displayed in common in the image acquired by the first head-mounted display and the image acquired by the other head-mounted display.

13. A head-mounted display comprising a camera, a display, and a processor, wherein the processor is configured to display an image captured by the camera on the display, receive identification information and modification information about an object from another head-mounted display, identify the object in the image based on the identification information, and display the modified image on the display based on the modification information.

14. The head-mounted display according to claim 13, wherein the received identification information comprises first position coordinate information of the object in a first local coordinate system with the other head-mounted display of the transmitting side as the origin, and position coordinate information of the head-mounted display of the receiving side in the first local coordinate system, wherein the processor converts the first local coordinate system to a second local coordinate system with the position of the head-mounted display as the origin, based on the position coordinate information of the head-mounted display of the receiving side in the first local coordinate system, converts the first position coordinate information to second position coordinate information in the second local coordinate system, and identifies the object in the image based on the second position coordinate information.

15. A display method performed by a first head-mounted display having a first camera, a first display, and a first processor, and a second head-mounted display having a second camera, a second display, and a second processor, wherein the first processor displays a first image captured by the first camera on the first display, acquires identification information to identify an object in the first image based on user operation, acquires change information for the object based on user operation, displays the first image modified based on the change information on the first display, transmits the identification information and the change information to the second head-mounted display, and the second processor displays a second image captured by the second camera on the second display, receives the identification information and the change information, identifies the object in the second image based on the identification information, and displays the second image modified based on the change information on the second display.

16. A display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, acquires identification information to identify an object in the image based on user operation, acquires change information for the object based on user operation, displays the modified image on the display based on the change information, and transmits the identification information and the change information to another head-mounted display.

17. A display method performed by a head-mounted display having a camera, a display, and a processor, wherein the processor displays an image captured by the camera on the display, receives identification information and modification information about an object from another head-mounted display, identifies the object in the image based on the identification information, and displays the modified image on the display based on the modification information.