Video processing device, video processing system, control method for video processing device, and program

The video processing apparatus addresses the limitations of existing XR training technologies by generating composite videos that align real and virtual body parts, enhancing training convenience and accuracy for XR training.

JP2025096918APending Publication Date: 2025-06-30CANON KK
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2023212912
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-18
Publication Date
2025-06-30

AI Technical Summary

Technical Problem

Existing XR training technologies are limited in usability, as they can only train based on pre-stored movement information and are restricted to training finger movements, failing to effectively teach movements of other body parts.

Method used

A video processing apparatus that acquires and transmits real video from a first user, receives operation information from a second user, generates virtual body parts based on this information, and aligns these with real video to create a composite video, enhancing XR training convenience.

Benefits of technology

The solution provides enhanced usability and convenience for XR training by allowing real-time operation training of various body parts, regardless of geographical distance, thereby improving training accuracy and effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025096918000001_ABST
    Figure 2025096918000001_ABST
Patent Text Reader

Abstract

To provide a video processing device, a video processing system, a control method for the video processing device, and a program having excellent convenience which can be used for training etc. by means of XR, for example.SOLUTION: An image processing device 101A comprises: video acquisition means (first user visual field video acquisition means 201) which acquires a reality video viewed by a first user US1; video transmission means (first user visual field video transmission means 202) which transmits the reality video to an image processing device 101B of a second user US2; information reception means (second user tracking information reception means 203) which receives action information of the second user US2; virtual object generation means (second user virtual body portion generation means 206) which generates a virtual object of the second user US2; and video generation means (second user virtual body portion display means 208) which generates a composite video of the reality video and the virtual object. The video generation means aligns the reality video and the virtual object.SELECTED DRAWING: Figure 1B
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image processing apparatus, an image processing system, a control method for an image processing apparatus, and a program.

Background Art

[0002] As a technology for making a user perceive a different sensation from reality by performing image processing on an image (captured image) of a subject existing in the real space, there is an XR (Extended Reality) technology. Further, as a typical device for displaying an image to which the XR technology is applied, an HMD (Head Mounted Display) used by being mounted on a person's head is known. Further, as one of the utilization examples of the XR technology, there is XR training. In XR training, there may be a worker who receives training and a supporter who supports the training for the worker. The worker receives training while wearing an HMD. On the HMD, devices, tools, etc. as virtual objects are displayed in CG. The worker can learn how to use devices, tools, etc. by operating virtual objects on the HMD with the support of the supporter. Note that in XR training, it is not necessary for the worker and the supporter to be in the same real space, that is, the distance in the real space is close. For example, the worker may be in a foreign country and the supporter may be in the country. For example, Patent Document 1 discloses a technology in which the user performs XR training following the movement of a partner by reproducing the pre-stored movement information of the partner on an HMD. Further, Patent Document 2 discloses a technology in which the fingers of a supporter are generated as virtual objects on an HMD and superimposed on the hand of a worker to teach the movement of the fingers of the supporter to the worker.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

[0004] However, with the technology of Patent Document 1, training for an operator can only be performed based on the movement information of the other party stored in advance. Further, with the technology of Patent Document 2, only the movement of the fingers can be trained, and the movement of other body parts such as the arm cannot be taught. Therefore, both the technologies of Patent Document 1 and Patent Document 2 had a problem of poor usability, that is, a difficult-to-use part.

[0005] The present invention has been made in view of the above problems. An object of the present invention is to provide a video processing apparatus, a video processing system, a control method of a video processing apparatus, and a program, which are excellent in convenience and can be used for training by, for example, XR (Extended Reality). MEANS FOR SOLVING THE PROBLEMS

[0006] In order to achieve the above object, a video processing apparatus of the present invention is a video processing apparatus used by a first user to process video, and includes: a video acquisition unit that acquires a real video of a real space viewed by the first user; a video transmission unit that transmits the real video acquired by the video acquisition unit to another video processing apparatus used by a second user different from the first user; an information reception unit that receives operation information regarding the operation of the second user from the other image processing apparatus; a virtual object generation unit that generates a body part of the second user as a virtual object that can be displayed in the video based on the operation information received by the information reception unit; and a video generation unit that generates a composite video by fusing the real video and the virtual object, wherein the video generation unit aligns the real video and the virtual object when generating the composite video to generate the composite video. EFFECTS OF THE INVENTION

[0007] According to the present invention, it is excellent in convenience and can be used, for example, for training by XR (Extended Reality).

Brief Description of the Drawings

[0008]

Figure 1A

Figure 1B

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Modes for Carrying Out the Invention

[0009] Hereinafter, each embodiment of the present invention will be described in detail with reference to the drawings. However, the configurations described in the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in each embodiment. For example, each part constituting the present invention can be replaced with any configuration that can exhibit the same function. Also, any component may be added. Further, any two or more configurations (features) among the embodiments can be combined.

[0010] <First Embodiment> Hereinafter, the first embodiment will be described with reference to FIGS. 1A to 9. FIG. 1A is a block diagram showing an example of the hardware configuration of an image processing apparatus included in the video processing system according to the first embodiment. FIG. 1B is a diagram showing an example of the usage state of the video processing system shown in FIG. 1A. As shown in FIG. 1A, the video processing system 1000 includes an image processing apparatus 101 as a video processing apparatus. As shown in FIG. 1B, the image processing apparatus 101 includes an image processing apparatus 101A as a first video processing apparatus used by a first user US1, and an image processing apparatus 101B as a second video processing apparatus (another video processing apparatus) used by a second user US2. In the present embodiment, the image processing apparatus 101A is a head-mounted display (HMD) that can be detachably attached to the head of the first user US1. The image processing apparatus 101B is composed of a desktop or notebook personal computer and a digital camera communicably connected to the personal computer. Note that the image processing apparatus 101A and the image processing apparatus 101B may each be, for example, a personal computer having a built-in web camera.

[0011] As shown in FIG. 1A, the image processing apparatus 101 includes a CPU 102, a ROM 103, a RAM 104, a sensing unit 105, a photographing unit 106, a display unit (display means) 107, an operation unit 108, and a communication unit 109, which are communicably connected to each other via a bus 110. The CPU 102 is an arithmetic processing unit (computer) that comprehensively controls the image processing apparatus 101. The CPU 102 executes various programs stored in the ROM 103 and the like to perform various processes. The ROM 103 is a read-only non-volatile memory device that stores programs and parameters such as initial data. The programs include, for example, an image processing program for causing the CPU 102 to execute each part and each means of the video processing apparatus (control method of the video processing apparatus). The RAM 104 temporarily stores input information, calculation results in image processing, and the like. The RAM 104 also functions as a memory device that provides a work area for the CPU 102. The sensing unit 105 is a device such as a sensor.

[0012] The sensing unit 105 can perform, for example, motion tracking and eye tracking. Thereby, it is possible to acquire movement information and gaze information of the body of the user who uses the image processing apparatus 101. The photographing unit 106 is a photographing device that performs photographing processing. The display unit 107 is composed of a liquid crystal display or the like. On the display unit 107, for example, a captured image obtained by the photographing unit 106 and a virtual object described later are displayed. In addition, characters, symbols, figures, items (for example, icons and contents) and the like are also displayed on the display unit 107. The operation unit 108 is an operation unit including various buttons such as a power button and operation members such as a dial. The communication unit 109 transmits and receives data to and from an external device by wired communication or wireless communication (for example, wireless LAN or local 5G, etc.). The communication unit 109 is a device compliant with a communication standard such as Ethernet or IEEE802.11.

[0013] As shown in FIG. 1B, in this embodiment, the video processing system 1000 is used for on-site training at the work site. The first user US1 is, for example, a training for receiving on-site training in operating the operation panel 2000 at the work site of a manufacturing factory. The second user US2 is a trainer, different from the first user US1. The second user US2 remotely teaches or supports the first user US1 in operating the operation panel 2000. Note that the use of the video processing system 1000 is not limited to use in on-site training at the work site.

[0014] FIG. 2 is a block diagram showing an example of the software configuration (functional configuration) of the first video processing apparatus. As shown in FIG. 2, the image processing apparatus 101A, which is the first video processing apparatus, includes a first user's field-of-view video acquisition means (video acquisition means) 201, a first user's field-of-view video transmission means (video transmission means) 202, and a second user tracking information reception means (information reception means) 203. Further, the image processing apparatus 101A includes an alignment reference position setting means (position setting means) 204 and an alignment means 205. Furthermore, the image processing apparatus 101A includes a second user's virtual body part generation means (virtual object generation means) 206, a second user's virtual body part control setting means (permission setting means) 207, and a second user's virtual body part display means (video generation means) 208. The processing performed by these software configurations can be realized by the CPU 102 of the image processing apparatus 101A executing a program. Therefore, it can be said that the CPU 102 of the image processing apparatus 101A has these software configurations.

[0015] Note that the CPU 102 is not limited to being able to implement all the processes of these software configurations. For example, the image processing apparatus 101A may have a dedicated processing circuit for implementing one or more processes. Note that each software configuration may transfer information to other software configurations by any method. For example, each software configuration may save the acquired information or the generated information in the RAM 104 so that other software configurations can acquire these information. In this way, in the image processing apparatus 101A, information transfer between software configurations may be performed via the RAM 104. Also, each software configuration may output the acquired information or the generated information to other software configurations without saving it in the RAM 104.

[0016] The first user's field-of-view video acquisition means 201 acquires a real video of the real space that is captured by the imaging unit 106 and visible to the first user US1, that is, a field-of-view video (video acquisition step).

[0017] The first user's field-of-view video transmission means 202 transmits the real video acquired by the first user's field-of-view video acquisition means 201 to the image processing apparatus 101B via the communication unit 109 (video transmission step).

[0018] The second user tracking information reception means 203 receives operation information regarding the operation of the second user US2 (hereinafter referred to as "second user tracking information") from the image processing apparatus 101B via the communication unit 109 (information reception step). The second user tracking information is not particularly limited, and examples include information regarding the positions of body parts such as the arms and hands of the second user US2, and information regarding the line of sight of the second user US2. Then, this second user tracking information is stored in the RAM 104 of the image processing apparatus 101A.

[0019] As will be described later, the real image acquired by the first user's field of view image acquisition means 201 and the virtual object that becomes the body part of the virtual second user US2 are fused to generate a composite image that can be displayed on the display unit 107 of the image processing apparatus 101A. The alignment reference position setting means 204 sets a reference position for aligning the real image and the virtual object when generating the composite image. The reference position can be an arbitrary predetermined position and serves as a common reference position for aligning the real image and the virtual object. The three-dimensional coordinates of this reference position are stored in the RAM 104. FIG. 3 is a diagram showing a relationship table between the reference position and the method for estimating the three-dimensional coordinates corresponding to the reference position. As shown in FIG. 3, as the reference position, for example, it is possible to set a body part such as the shoulder, waist, elbow, etc. of the first user US1, or an object of interest (object of attention) that is located at the tip of the line of sight of the second user US2 and that the second user US2 is paying attention to. When the reference position is a body part of the first user US1, the position coordinates of the body part are estimated by the CPU 102 by acquiring the position coordinate information of the image processing apparatus 101A from the sensing unit 105 of the image processing apparatus 101A worn on the head of the first user US1. Also, the position coordinates of the body part may be estimated based on the detection result of a sensor device (not shown) worn on the body of the first user US1. The position coordinates of the object of interest are acquired by calculation by the CPU 102. Note that the reference position can be changed from the body part of the first user US1 to the object of interest, and vice versa, according to, for example, the usage state of the video processing system 1000. Also, the object of interest is not particularly limited, and examples include the operation button 2001 (see FIG. 1B) of the operation panel 2000 displayed on the display unit 107 of the image processing apparatus 101B. Also, the object of interest may be an object that exists as a physical entity in the real space or a virtual object.

[0020] The alignment means 205 aligns the real image and the virtual object based on the reference position determined by the alignment reference position setting means 204. For example, when the reference position is determined to be the head of the first user US1, the three-dimensional coordinates of the head position of the first user US1 are overlaid (corresponded) with the three-dimensional coordinates of the head position of the second user US2 stored in the second user tracking information storage means 304 described later. Also, when the reference position is determined to be the object of interest, the three-dimensional coordinates of the object of interest are overlaid with the three-dimensional coordinates of the hand stored in the second user tracking information storage means 304.

[0021] As described above, the second user tracking information receiving means 203 receives the second user tracking information. The second user virtual body part generating means 206 generates the body part of the second user US2 as a virtual object that can be displayed in the video based on this second user tracking information (virtual object generation step). The generation of the virtual object is performed using the joint position coordinate information of the second user US2 stored in the second user tracking information storage means 304 described later. Then, the data of the virtual object is stored in the RAM 104. In the present embodiment, the second user virtual body part generating means 206 generates the arm of the second user US2 as a virtual object among the body parts of the second user US2.

[0022] The second user virtual body part control setting means 207 sets whether to permit an operation on the virtual object of the second user US2 on the composite video. Thereby, it is possible to guarantee the security in which the degree of freedom of the operation on the virtual object is appropriately set, that is, without excess or deficiency. Note that the operation on the virtual object of the second user US2 on the composite video can be performed by various means (operation means) such as means for recognizing gestures.

[0023] The second user virtual body part display means 208 displays the virtual object of the second user US2 on the display unit 107 of the image processing apparatus 101A. At this time, the second user virtual body part display means 208 fuses the virtual object aligned by the alignment means 205 with the real video to generate a composite video, and displays the composite video (video generation step). Note that, in the present embodiment, the second user virtual body part display means 208 and the alignment means 205 have separate configurations, but are not limited thereto, and the alignment means 205 may be included as a part of the second user virtual body part display means 208.

[0024] FIG. 4 is a block diagram showing an example of the software configuration (functional configuration) of the second video processing apparatus. As shown in FIG. 4, the image processing apparatus 101B, which is the second video processing apparatus, includes a first user view video receiving means (video receiving means) 301 and a video display means (display means) 302. The image processing apparatus 101B also includes a second user tracking information transmitting means (information transmitting means) 303, a second user tracking information storage means 304, and a second user tracking information obtaining means (information obtaining means) 305. The processing performed by these software configurations can be realized by the CPU 102 of the image processing apparatus 101B executing a program. Therefore, it can be said that the CPU 102 of the image processing apparatus 101B has these software configurations.

[0025] The first user view video receiving means 301 receives the real video from the first user view video transmitting means 202 via the communication unit 109.

[0026] The video display means (display means) 302 displays the real video received by the first user view video receiving means 301 on the display unit 107 of the image processing apparatus 101B. Note that the real video preferably includes a virtual object.

[0027] The second user tracking information acquisition means 305 acquires second user tracking information. The second user tracking information includes, for example, the position coordinate information of the joints of the second user US2, the line-of-sight information of the second user US2, and the like. Then, the second user tracking information is stored in the second user tracking information storage means 304. FIG. 5 is a diagram showing an example of the joint position coordinate information stored in the second user tracking information storage means. As shown in FIG. 5, the joint position coordinate information includes the position coordinate information of the joints such as the shoulder 501, elbow 502, hand 503, and knee 504 of the second user US2. The second user tracking information is obtained by extracting and analyzing the joint position coordinate information from the video obtained by the imaging unit 106 of the image processing apparatus 101B capturing the second user US2. Note that the second user tracking information may be acquired from, for example, a tracking device for motion capture attached to the second user US2. Also, the line-of-sight information of the second user US2 is acquired by the sensing unit 105 measuring which area of the display unit 107 the second user US2 is paying attention to.

[0028] The second user tracking information transmission means 303 transmits the second user tracking information stored in the second user tracking information storage means 304 to the image processing apparatus 101A via the communication unit 109.

[0029] FIG. 6 is a flowchart showing the processing executed by the first video processing apparatus. The program based on the flowchart shown in FIG. 6 operates when the CPU 102 of the image processing apparatus 101A executes it. Here, as an example, it is assumed that this program operates in a state where the video processing system 1000 is used in on-site training at a work site (see FIG. 1B). As shown in FIG. 9, in step S601, the first user visual field video acquisition means 201 acquires the real video of the real space that is captured by the imaging unit 106 and is visible to the first user US1 who is receiving operation training on the operation panel 2000, that is, the visual field video (visual field information). This visual field video is stored in the RAM 104.

[0030] In step S602, the first user's field of view video transmission means 202 transmits the field of view video of the first user US1 stored in the RAM 104 in step S601 to the image processing device 101B via the communication unit 109. As a result, the field of view video of the first user US1 is displayed on the display unit 107 of the image processing device 101B. As shown in the upper right figure of Fig. 1B, the second user US2 can visually recognize the operation panel 2000 included in the field of view video of the first user US1. In particular, the second user US2 is paying attention to the operation button 2001 of the operation panel 2000.

[0031] In step S603, the second user tracking information receiving means 203 receives the second user tracking information from the second user tracking information storage means 304 of the image processing device 101B via the communication unit 109. The second user tracking information is gesture information as if the second user US2 operates the operation button 2001 on the display unit 107 of the image processing device 101B. And this second user tracking information is stored in the RAM 104.

[0032] In step S604, the second user virtual body part generation means 206 generates the body part of the second user US2 as a virtual object based on the second user tracking information stored in the RAM 104 in step S603. Note that the body part of the second user US2 is not particularly limited, but here it is the arm.

[0033] In step S605, the alignment reference position setting means 204 sets a reference position for displaying the virtual object generated in step S604 on the display unit 107 of the image processing device 101A. As described above, the reference position is a three-dimensional coordinate for aligning the real video and the virtual object during composite video generation. And this reference position is stored in the RAM 104.

[0034] In step S606, the alignment means 205 aligns the real video and the virtual object based on the reference position stored in the RAM 104 in step S605.

[0035] In step S607, the second user virtual body part display means 208 generates a composite image of the real video and the virtual object that have been aligned in step S606, and displays it on the display unit 107 of the image processing apparatus 101A. Also, it is assumed that the avatar 800 of the first user US1 is included in the composite image. In this embodiment, this avatar 800 (see FIGS. 7 and 8) is generated, for example, by the second user virtual body part display means 208.

[0036] FIGS. 7 and 8 are diagrams each showing an example of a composite image of a real video and a virtual object. FIG. 7 is a diagram showing a composite image when the body part of the first user US1 is set at the reference position. In the composite image shown in FIG. 7, the root of the arm 802 of the avatar 800, that is, the position information of the shoulder (the body part of the first user US1), is made to correspond to the root of the arm of the second user US2 included in the second user tracking information, that is, the position information of the shoulder. As a result, the virtual object 803, which is the arm of the second user US2, extends from the shoulder of the avatar 800 of the first user US1. The first user US1 can view the composite image shown in FIG. 7. Thereby, as shown in the lower left figure of FIG. 1B, the first user US1 can feel as if the virtual object 803 extending from himself / herself is operating the operation panel 2000.

[0037] FIG. 8 is a diagram showing a composite image when the object of interest is set at the reference position. In the composite image shown in FIG. 8, the position information of the operation button 2001 of the operation panel 2000, which is the object of interest, is made to correspond to the tip of the arm of the second user US2 included in the second user tracking information, that is, the position information of the hand. As a result, the virtual object 803, which is the arm of the second user US2, extends from the shoulder of the avatar 800 of the first user US1 and is in a state of operating the operation button 2001. The first user US1 can view the composite image shown in FIG. 8. Thereby, as shown in the lower right figure of FIG. 1B, the first user US1 can feel as if the virtual object 803 is operating the operation button 2001.

[0038] In addition, when the body part of the first user US1 is included in the composite image, the second user virtual body part display means 208 may temporarily suppress the display of the body part. Thereby, the virtual object 803 is emphasized, and the first user US1 can easily visually recognize the virtual object 803. In this way, the second user virtual body part display means 208 may have a function as a suppression means for temporarily suppressing the display of the avatar 800, that is, it may be configured to enable diminished reality. Further, when the body part of the first user US1 is included in the composite image, the second user virtual body part display means 208 may eliminate the physical difference between the first user US1 and the second user US2. Specifically, the physical difference between the first user US1 and the second user US2 in the body part of the first user US1 and the body part of the second user US2 (virtual object 803) being displayed may be eliminated. Thereby, for example, the length of the arm 802 of the avatar 800 and the length of the virtual object 803 can be made the same, and the first user US1 can feel as if the virtual object 803 is operating the operation button 2001 without a sense of discomfort. In this way, the second user virtual body part display means 208 may have a function as an elimination means for eliminating the physical difference between the first user US1 and the second user US2. In this case, the length of the arm of the first user US1 is acquired in advance.

[0039] As shown in FIG. 6, in step S608, the second user virtual body part control setting means 207 determines whether to permit the operation of the virtual object for the second user US2. As a result of the determination in step S608, if it is determined that the operation of the virtual object is permitted, the process proceeds to step S609. On the other hand, as a result of the determination in step S608, if it is determined that the operation of the virtual object is not permitted, the process ends.

[0040] In step S609, when the second user virtual body part display means 208 detects that the second user US2 has operated the virtual object, it generates a composite image reflecting the operation result and displays it on the display unit 107. Note that information such as vibrations generated by the second user US2 operating the virtual object can be fed back and shared with the first user US1.

[0041] FIG. 9 is a flowchart showing the processing executed by the second video processing apparatus. The program based on the flowchart shown in FIG. 6 operates when the CPU 102 of the image processing apparatus 101B executes it.

[0042] In step S701, the first user view video receiving means 301 receives the view video of the first user US1 transmitted in step S602 via the communication unit 109.

[0043] In step S702, the video display means 302 displays the first user view video received in step S701 on the display unit 107.

[0044] In step S703, the second user tracking information acquisition means 305 acquires the second user tracking information. This second user tracking information is stored in the RAM 104. When acquiring the second user tracking information, while the second user US2 visually recognizes the view video of the first user US1 displayed on the display unit 107 of the image processing apparatus 101B, the second user US2 moves their own arm to operate the operation button 2001 included in the view video.

[0045] In step S704, the second user tracking information transmission means 303 transmits the second user tracking information stored in step S703 to the image processing apparatus 101A via the communication unit 109.

[0046] In the video processing system 1000 configured as described above, the first user US1 can receive operation training of the operation panel 2000 from the second user US2 based on the motion information regardless of the body part where the second user US2 gestures. As a result, the first user US1 can accurately operate the operation panel 2000. Thus, the video processing system 1000 is a system with excellent convenience that can be used, for example, in training using XR or the like.

[0047] <Second Embodiment> Hereinafter, with reference to FIGS. 10 to 12, the second embodiment will be described. The description will focus on the differences from the above-described embodiments, and the description of the same matters will be omitted. This embodiment is the same as the first embodiment except that an operation on the avatar of the first user is possible based on the second user tracking information. FIG. 10 is a block diagram showing an example of the software configuration (functional configuration) of the first video processing apparatus according to the second embodiment. As shown in FIG. 10, the image processing apparatus 101A includes a first user view video acquisition unit 201, a first user view video transmission unit 202, and a second user tracking information reception unit 203. Further, the image processing apparatus 101A includes a first user avatar operation permission unit 1001, a first user avatar generation unit (avatar generation unit) 1002, a first user avatar storage unit 1003, and a first user avatar control unit (operation control unit) 1004. The first user avatar operation permission unit 1001 sets whether to permit an operation on the avatar 1200 (see FIG. 12) of the first user US1 for the second user US2. This setting information is stored in the RAM 104. The first user avatar generation unit 1002 generates an avatar 1200 that can be operated in the view video. The data of the avatar 1200 is stored in the first user avatar storage unit 1003. The first user avatar control unit 1004 controls the operation of the avatar 1200 according to the operation when an operation on the avatar 1200 based on the second user tracking information received from the image processing apparatus 101B is performed.

[0048] FIG. 11 is a flowchart showing the processing executed by the first video processing apparatus. As shown in FIG. 11, in step S1101, the first user avatar generation means 1002 generates an avatar 1200. The information of this avatar 1200 is stored in the first user avatar storage means 1003. Also, the avatar 1200 is displayed on the display unit 107 of the image processing apparatus 101A. Thereby, the first user US1 can operate the avatar 1200 on the visual field video. After step S1101 is executed, the processing proceeds to steps S601 to S603 in order.

[0049] In step S1102 after step S603 is executed, the first user avatar operation permission means 1001 determines whether to permit the second user US2 to operate the avatar 1200 on the visual field video. This determination is made based on the setting information stored in the RAM 104 as to whether to permit the operation of the avatar 1200. If, as a result of the determination in step S1102, it is determined to permit the operation of the avatar 1200, the processing proceeds to step S1103. On the other hand, if, as a result of the determination in step S1102, it is determined not to permit the operation of the avatar 1200, the processing ends.

[0050] In step S1103, when an operation on the avatar 1200 based on the second user tracking information received from the image processing device 101B is performed, the first user avatar control means 1004 controls the operation of the avatar 1200 according to the operation. FIG. 12 is a diagram showing an example of a composite image of a real image and a virtual object. Here, the "operation on the avatar 1200 based on the second user tracking information received from the image processing device 101B" will be described. For example, when the second user US2 moves his own arm and presses the operation button 2001 of the operation panel 2000, arm movement information is acquired as the second user tracking information. Then, according to this movement of the arm, the arm 1201 of the avatar 1200 will press the operation button 2001. In the composite image shown in FIG. 12, the arm 1201 of the avatar 1200 is in a state of moving based on the second user tracking information (arm movement information). Thus, the first user US1 can accurately operate the operation panel 2000 in actuality by referring to the composite image shown in FIG. 12.

[0051] As described above, the preferred embodiments of the present invention have been described. However, the present invention is not limited to the above-described embodiments, and various modifications and changes are possible within the scope of the gist thereof. The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors of a computer of the system or device read and execute the program. Further, the present invention can also be realized by a circuit (for example, ASIC) that realizes one or more functions. Also, in the above embodiment, the image processing device 101A is a head-mounted display having the CPU 102 to the communication unit 109, but it is not limited thereto. For example, in the image processing device 101A, the sensing unit 105, the photographing unit 106, and the display unit 107 may be omitted, and these may constitute a head-mounted display communicably connected to the image processing device 101. In this case, the image processing device 101 and the head-mounted display may be connected by wire or wirelessly.

[0052] In the video processing system 1000, the image processing device 101A can also be used as a terminal device, and the image processing device 101B can be used as a server that is communicably connected to a plurality of terminal devices. In the video processing system 1000 in this case, for example, even when the server is located outside Japan and the terminal device is located within Japan, each file and data can be transmitted from the server to the terminal device, and the terminal device can receive each file and data. Even when the server is located outside Japan in this way, the transmission and reception (transmission and reception) of files and data in this system are integrated, that is, without going through a separate operation by the user of the terminal device. And since the system functions when the terminal device in Japan receives each file and data, the transmission and reception can be regarded as having been performed within the country. In this system, for example, even when the server is located outside Japan and the terminal device is located within Japan, the terminal device can perform the main functions of this system, and the effects of the functions can be manifested within Japan. For example, even if the server is located outside Japan, if the terminal device constituting this system is located within Japan, it is possible to use this system within the country using the terminal device. And the use of this system can, for example, have an impact on the economic interests of the patent holder.

[0053] The disclosure of each embodiment includes the following configurations, methods, and programs. (Configuration 1) A video processing device used by a first user to process video, video acquisition means for acquiring a real video of the real space visually recognized by the first user, video transmission means for transmitting the real video acquired by the video acquisition means to another video processing device used by a second user different from the first user, information reception means for receiving operation information regarding the operation of the second user from the other image processing device, virtual object generation means for generating the body part of the second user as a virtual object that can be displayed in the video based on the operation information received by the information reception means video generation means for generating a composite video by fusing the real video and the virtual object, The video generation means aligns the real video and the virtual object when generating the composite video, and generates the composite video. A video processing apparatus characterized by that. (Configuration 2) When generating the composite video, the video generation means aligns the real video and the virtual object with a predetermined position as a common reference position, and generates the composite video. The video processing apparatus according to Configuration 1, characterized by that. (Configuration 3) The video processing apparatus according to Configuration 2, further comprising position setting means for setting the reference position. (Configuration 4) The position setting means sets, as the reference position, a body part of the first user or a target object located at the tip of the line of sight of the second user and attracting the attention of the second user. The video processing apparatus according to Configuration 3. (Configuration 5) When the body part of the first user is set at the reference position, the video generation means, as alignment of the real video and the virtual object, corresponds the position information of the body part of the first user with the position information of the body part of the second user included in the motion information. The video processing apparatus according to Configuration 4, characterized by that. (Configuration 6) When the target object is set at the reference position, the video generation means, as alignment of the real video and the virtual object, corresponds the position information of the target object with the position information of the body part of the second user included in the motion information. The video processing apparatus according to Configuration 4, characterized by that. (Configuration 7) The video processing apparatus according to any one of Configurations 3 to 6, wherein the position setting means can change the reference position. (Configuration 8) The information acquisition means receives, as the motion information, information on the position of the body part of the second user or information on the line of sight of the second user. The video processing apparatus according to any one of Configurations 1 to 7, characterized by that. (Configuration 9) The virtual object generation means generates the arm of the second user as the virtual object, and the video processing apparatus according to any one of Configurations 1 to 8, characterized in that. (Configuration 10) An operation means capable of operating on the virtual object on the composite video, And permission setting means for setting whether to permit the operation by the second user, and the video processing apparatus according to any one of Configurations 1 to 9, characterized in that. (Configuration 11) Avatar generation means for generating an avatar of the first user, And operation control means for controlling the operation of the avatar of the first user based on the operation information, and the video processing apparatus according to any one of Configurations 1 to 10, characterized in that. (Configuration 12) A video processing apparatus according to any one of Configurations 1 to 11, characterized by comprising display means for displaying the composite video generated by the video generation means. (Configuration 13) A video processing apparatus according to Configuration 12, characterized by comprising suppression means for suppressing the display of the body part of the first user when the body part of the first user is included in the composite video displayed on the display means. (Configuration 14) A video processing apparatus according to Configuration 12 or 13, characterized by comprising elimination means for eliminating the physical difference between the first user and the second user in the body part of the first user and the body part of the second user displayed by the virtual object when the body part of the first user is included in the composite video displayed on the display means. (Configuration 15) A video processing apparatus according to any one of Configurations 12 to 14, characterized in that it is a head-mounted display. (Configuration 16) A video processing system comprising a first video processing apparatus used by a first user to process video and a second video processing apparatus used by a second user different from the first user to process video, The first video processing apparatus, Video acquisition means for acquiring a real video of the real space viewed by the first user, Video transmission means for transmitting the real video obtained by the video acquisition means to another video processing device used by a second user different from the first user; Information reception means for receiving operation information regarding the operation of the second user from the other image processing device; Virtual object generation means for generating a body part of the second user as a virtual object that can be displayed in the video based on the operation information received by the information reception means; Video generation means for generating a composite video by fusing the real video and the virtual object, comprising: When generating the composite video, the video generation means aligns the real video and the virtual object to generate the composite video. The second video processing device: Video reception means for receiving the real video from the video transmission means; Display means for displaying the real video received by the video reception means; Information acquisition means for acquiring operation information regarding the operation of the second user; Information transmission means for transmitting the operation information acquired by the information acquisition means to the first video processing device. A video processing system characterized by comprising: (Method 1) A method for controlling a video processing device used by a first user to process video, comprising: A video acquisition step of acquiring a real video of a real space visually recognized by the first user; A video transmission step of transmitting the real video acquired in the video acquisition step to another video processing device used by a second user different from the first user; An information reception step of receiving operation information regarding the operation of the second user from the other image processing device; A virtual object generation step of generating a body part of the second user as a virtual object that can be displayed in the video based on the operation information received in the information reception step; A video generation step of generating a composite video by fusing the real video and the virtual object. In the video generation step, when generating the composite video, the control method of the video processing apparatus is characterized by aligning the real video and the virtual object to generate the composite video. (Program 1) A program for causing a computer to execute each means of the video processing apparatus according to any one of Configurations 1 to 15.

Explanation of Reference Numerals

[0054] 101A Image processing apparatus 101B Image processing apparatus 201 First user's field of view video acquisition means 202 First user's field of view video transmission means 203 Second user tracking information reception means 204 Alignment reference position setting means 206 Second user's virtual body part generation means 208 Second user's virtual body part display means 1000 Video processing system

Claims

1. A video processing apparatus used by a first user to process video, comprising: video acquisition means for acquiring a real video of the real space visually recognized by the first user; video transmission means for transmitting the real video acquired by the video acquisition means to another video processing apparatus used by a second user different from the first user; information reception means for receiving operation information regarding the operation of the second user from the other image processing apparatus; virtual object generation means for generating, based on the operation information received by the information reception means, a body part of the second user as a virtual object that can be displayed in the video; video generation means for generating a composite video by fusing the real video and the virtual object, wherein the video generation means performs alignment of the real video and the virtual object when generating the composite video to generate the composite video.

2. The video processing apparatus according to claim 1, wherein the video generation means performs alignment of the real video and the virtual object with a predetermined position as a common reference position when generating the composite video to generate the composite video.

3. The video processing apparatus according to claim 2, further comprising position setting means for setting the reference position.

4. The video processing apparatus according to claim 3, wherein the position setting means sets, as the reference position, a body part of the first user or an object of attention located at the tip of the line of sight of the second user and attracting the attention of the second user.

5. The video processing apparatus according to claim 4, wherein when a body part of the first user is set at the reference position, the video generation means corresponds the position information of the body part of the first user and the position information of the body part of the second user included in the operation information as alignment of the real video and the virtual object.

6. The video processing apparatus according to claim 4, wherein when the object of attention is set at the reference position, the video generation means corresponds the position information of the object of attention and the position information of the body part of the second user included in the operation information as alignment of the real video and the virtual object.

7. The video processing apparatus according to claim 3, wherein the position setting means can change the reference position.

8. The video processing apparatus according to claim 1, wherein the information receiving means receives, as the operation information, information regarding the position of a body part of the second user or information regarding the line of sight of the second user.

9. The video processing apparatus according to claim 1, wherein the virtual object generation means generates the arm of the second user as the virtual object.

10. operation means capable of operating on the virtual object in the composite video; permission setting means for setting whether or not to permit the operation by the second user, the video processing apparatus according to claim 1, characterized in that it comprises.

11. avatar generation means for generating an avatar of the first user; operation control means for controlling the operation of the avatar of the first user based on the operation information, the video processing apparatus according to claim 1, characterized in that it comprises.

12. The video processing apparatus according to claim 1, further comprising display means for displaying the composite video generated by the video generation means.

13. The video processing apparatus according to claim 12, further comprising suppression means for suppressing the display of the body part of the first user when the body part of the first user is included in the composite video displayed on the display means.

14. When the body part of the first user is included in the composite video displayed on the display means, the video processing apparatus according to claim 12, further comprising elimination means for eliminating the physical difference between the first user and the second user in the body part of the first user and the body part of the second user displayed as the virtual object.

15. The video processing apparatus according to claim 12, characterized in that it is a head-mounted display.

16. A video processing system comprising a first video processing apparatus used by a first user to process video and a second video processing apparatus used by a second user different from the first user to process video, wherein the first video processing apparatus video acquisition means for acquiring a real video of the real space viewed by the first user; video transmission means for transmitting the real video acquired by the video acquisition means to another video processing apparatus used by a second user different from the first user; information receiving means for receiving operation information regarding the operation of the second user from the other image processing apparatus; Virtual object generation means for generating, based on the operation information received by the information receiving means, a body part of the second user as a virtual object that can be displayed in the video; Video generation means for generating a composite video by fusing the real video and the virtual object, comprising: When generating the composite video, the video generation means aligns the real video and the virtual object to generate the composite video. The second video processing device is: Video receiving means for receiving the real video from the video transmitting means; Display means for displaying the real video received by the video receiving means; Information acquisition means for acquiring operation information regarding the operation of the second user; An information transmission means for transmitting the operation information acquired by the information receiving means to the first video processing device, and a video processing system characterized by comprising:

17. A method for controlling a video processing device used by a first user to process video, comprising: A video acquisition step of acquiring a real video of a real space visually recognized by the first user; A video transmission step of transmitting the real video acquired in the video acquisition step to another video processing device used by a second user different from the first user; An information reception step of receiving operation information regarding the operation of the second user from the other image processing device; A virtual object generation step of generating, based on the operation information received in the information reception step, a body part of the second user as a virtual object that can be displayed in the video; A video generation step of generating a composite video by fusing the real video and the virtual object, and In the video generation step, when generating the composite video, the real video and the virtual object are aligned to generate the composite video, and a method for controlling a video processing device characterized by this.

18. A program for causing a computer to execute each means of the video processing device according to claim 1.

Citation Information

Patent Citations

  • Physical activity supporting system, method and program

    JP2020195551A

  • Work support system and program

    JP2021039567A