Information processing device, head-mounted display, and program

The information processing device addresses the challenge of missynchronization in virtual space technologies by synchronizing virtual and real videos based on audio signals, providing a seamless and enhanced video experience.

JP2025079565AActive Publication Date: 2025-05-22TSUBURAYA FIELDS HLDG CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2023192325
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-05-22
Estimated Expiration
2043-11-10

AI Technical Summary

Technical Problem

Existing virtual space technologies, such as Augmented Reality (AR) and Mixed Reality (MR), struggle to provide a seamless and synchronized video experience by combining virtual and real videos, often leading to discomfort due to missynchronization.

Method used

An information processing device that captures real videos and audio using a head-mounted display, extracts audio signals from the real audio, and synchronizes virtual videos with real videos based on these audio signals to create a composite video that is frame-synchronized with the video content.

Benefits of technology

The solution provides a new video experience by seamlessly integrating virtual and real videos, reducing user discomfort caused by missynchronization and enhancing the video content with additional information and effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025079565000001_ABST
    Figure 2025079565000001_ABST
Patent Text Reader

Abstract

To provide a user with a new visual experience utilizing a virtual space technique.SOLUTION: An information processing device comprises: a storage unit which stores a virtual moving image; a video input unit which inputs a real-world video captured in an angle of view including a display in a real-world space by a camera mounted on a head-mounted display; an audio input unit which inputs real-world audio in the real-world space detected by a microphone mounted on the head-mounted display; an extraction unit which extracts an audio signal embedded in the audio of a moving image content displayed on the display from the real-world audio; a composite video generation unit which combines the virtual moving image with the real-world video such that the virtual moving image is frame-synchronized with the moving image content on the basis of the audio signal and generates a composite video; and a transmission unit which transmits the composite video data to the head-mounted display.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] An embodiment of the present invention relates to an information processing device and a program. [Background technology]

[0002] Virtual space technologies such as Augmented Reality (AR) and Mixed Reality (MR) are known that superimpose and display virtual objects and virtual images on images of real space captured by a camera mounted on a head mounted display (HMD) (for example, Patent Document 1). Because these technologies can use images of actual real space, they are expected to be applied to various fields such as surgery simulations, teaching workers assembly procedures, explaining artworks in museums, and introducing furniture in newly built homes. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2017-084323 A Summary of the Invention [Problem to be solved by the invention]

[0004] The aim is to provide users with a new video experience using virtual space technology. [Means for solving the problem]

[0005] The information processing device of the present disclosure comprises a memory unit that stores a virtual video, a video input unit that inputs a real video captured by a camera equipped on a head-mounted display with an angle of view that includes the display in real space, an audio input unit that inputs real audio of the real space detected by a microphone equipped on the head-mounted display, an extraction unit that extracts an audio signal embedded in the audio of the video content displayed on the display from the real audio, a composite video generation unit that combines the virtual video with the real video based on the audio signal so that the virtual video is frame-synchronized with the video content, and generates a composite video, and a transmission unit that transmits data of the composite video to the head-mounted display. [Brief description of the drawings]

[0006] [Figure 1] FIG. 1 is a diagram showing the configuration of a video distribution system including an information processing device according to the first embodiment. [Diagram 2] FIG. 2 is a hardware configuration diagram of the information processing device in FIG. [Diagram 3] FIG. 3 is a functional configuration diagram of the information processing device in FIG. [Figure 4] FIG. 4 is a diagram showing an example of a signal waveform of a real sound processed by the sound signal processing unit of FIG. [Diagram 5] FIG. 5 is a supplementary diagram for explaining the details of the processing by the display frame identification unit, the preprocessing unit, and the combining unit in FIG. [Figure 6] FIG. 6 is a flowchart illustrating an example of a procedure of preparation processing for composite video distribution processing by the information processing device according to the first embodiment. [Figure 7] FIG. 7 is a flowchart showing an example of a composite video creation process of the composite video distribution process by the information processing device according to the first embodiment. [Figure 8] FIG. 8 is a diagram showing a configuration of a video distribution system including an information processing device according to a modification of the first embodiment. [Figure 9] FIG. 9 is a diagram showing an example of a signal waveform of real sound processed by the sound signal processing unit included in the information processing device of FIG. [Figure 10] FIG. 10 is a supplementary diagram for explaining the details of the display frame specifying unit, the preprocessing unit, and the combining unit included in the information processing device in FIG. [Figure 11] FIG. 11 is a functional configuration diagram of an information processing device according to the second embodiment. [Figure 12] FIG. 12 is a functional configuration diagram of an information processing device according to the third embodiment. [Figure 13] FIG. 13 is a supplementary diagram for explaining details of a display frame specifying unit and a preprocessing unit included in the information processing device according to the third embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0007] Hereinafter, an information processing device according to an embodiment of the present invention will be described with reference to the drawings. In the following description, components having substantially the same functions and configurations are denoted by the same reference numerals, and repeated description will be given only when necessary.

[0008] (First embodiment) The information processing device according to the first embodiment is a computer device having a function of creating a composite image by combining a virtual video with a real image captured by a camera equipped on a head mounted display at an angle of view including a display on which video content is displayed, and a function of transmitting the created composite image to the head mounted display. This makes it possible to provide a new video experience that extends the video content to a user wearing the head mounted display. In addition, when creating the composite image, the information processing device according to the first embodiment uses real sound detected by a microphone equipped on the head mounted display to frame-synchronize the video content and the virtual video. This makes it possible to avoid a situation in which a user wearing a head mounted display feels uncomfortable due to a missynchronization between the video content in the composite image and the virtual video.

[0009] Typically, the information processing device according to the first embodiment is configured as follows. In the first embodiment, the terms are defined as follows. User: A person wearing a head-mounted display. Real image: An image of real space captured by a camera mounted on a head mounted display (HMD). In particular, in the first embodiment, the real image is an image captured with an angle of view that includes a display of a television, smartphone, or the like arranged in the real space. Video content: Videos being displayed on a display arranged in real space. For example, in the first embodiment, video of movies, dramas, concerts, variety shows, sports, etc., corresponds to video content. Virtual video: A video that is combined with real video. Virtual videos are created in advance to match the video content. For example, virtual videos include videos that aim to expand the range of images shown in the video content, videos that aim to decorate the video content, videos that aim to increase the dramatic effect of the video content, and videos that aim to provide users with new information about the video content. A video that aims to expand the range of images shown in the video content has a pattern that is continuous with the pattern shown in the video content. A video that aims to provide users with information about the video content is a video that shows viewers' comments on the video content or explains the contents of the video content. Composite image: An image that combines a real image with a virtual moving image. Real sound: Sound related to the real space detected by a microphone equipped on the head mounted display. In particular, in the first embodiment, the real sound includes audio content. Of course, the real sound may include not only audio content but also sounds of daily life and the user's voice. Audio content: Audio corresponding to video content. For example, audio content is output from a speaker connected to a display. For example, in the first embodiment, background music of the video content, lines of characters appearing in the video content, sound effects of the video content, etc., correspond to audio content.

[0010] 1, an information processing device 1 according to a first embodiment (hereinafter simply referred to as the information processing device 1) configures a video distribution system together with a server device 7 and a head mounted display 5. The information processing device 1 is a computer device such as a PC, a tablet, or a smartphone.

[0011] The information processing device 1 and the server device 7 are connected to be able to communicate with each other via a network 9 such as the Internet. A plurality of virtual moving image files are stored in the server device 7. Each virtual moving image file is associated with an identification code for identifying moving image content. When the server device 7 receives a request to acquire a virtual moving image file from the information processing device 1, it reads out the virtual moving image file associated with the identification code included in the acquisition request from the storage device and transmits the virtual moving image file to the information processing device 1.

[0012] The information processing device 1 and the head mounted display 5 are connected to each other so as to be able to communicate with each other by short-distance wireless communication such as Bluetooth. The head mounted display 5 has a display device 55 arranged at a position blocking the user's field of vision, a camera 51 that captures an image in the direction of the user's line of sight, a microphone 53 that detects real sound, and a speaker (not shown) that outputs the sound detected by the microphone 53 to the user. A real image captured by the camera 51 is displayed on the display device 55. Since the real image is an image corresponding to the user's line of sight, the user can see the real image as if he or she were directly looking at the real space with his or her eyes. Video data representing the real image captured by the camera 51 and audio data representing the real sound detected by the microphone 53 are transmitted to the information processing device 1. The information processing device 1 creates a composite image by combining the real image captured by the camera 51 and the virtual video acquired from the server device 7, and transmits the created composite image to the head mounted display 5. As a result, the composite image created by the information processing device 1 is displayed on the display device 55 of the head mounted display 5.

[0013] As shown in Fig. 2, the information processing device 1 has a processor 11. A RAM 12, a ROM 13, a storage device 14, and a communication device 15 are connected to the processor 11 via a data / control bus 10. The processor 11 has a CPU and a GPU. The RAM 12 functions as the main memory, work area, etc. of the processor 11. The ROM 13 or the storage device 14 stores a BIOS (Basic Input Output System), an OS (operation system), and a composite video distribution program executed by the processor 11. The communication device 15 transmits and receives data between the server device 7 and the head mounted display 5 under the control of the processor 11.

[0014] As shown in Figure 3, when the distribution program stored in the storage device 14 is executed by the processor 11, the information processing device 1 functions as a video input unit 21, an audio input unit 22, a video transmission unit 23, a virtual moving image acquisition unit 24, a memory unit 25, an audio signal processing unit 26, a display frame identification unit 27, and a composite video generation unit 28.

[0015] The video input unit 21 is realized by the function of the communication device 15 shown in Fig. 2. The video input unit 21 inputs video data representing a real image captured by a camera 51 mounted on the head mounted display 5.

[0016] The voice input unit 22 is realized by the function of the communication device 15 shown in Fig. 2. The voice input unit 22 inputs voice data representing real voice detected by a microphone 53 equipped on the head mounted display 5.

[0017] The video transmission unit 23 is realized by the function of the communication device 15 shown in Fig. 2. The video transmission unit 23 transmits image data of the composite video generated by the composite video generation unit 28 to the head mounted display 5.

[0018] The virtual moving image acquisition unit 24 acquires from the server device 7 a virtual moving image file corresponding to the identification code extracted by the audio signal processing unit 26. The virtual moving image file acquired by the virtual moving image acquisition unit 24 is stored in the storage unit 25.

[0019] The storage unit 25 is realized by the function of the storage device 14 shown in FIG. 2. The storage unit 25 stores various information required for processing related to the distribution program (composite video distribution processing). Specifically, the storage unit 25 stores a virtual video file acquired from the server device 7. The virtual video file includes data on a plurality of frames (referred to as virtual frames) constituting the virtual video, data on the waiting time from the extraction of the start code to the start of the composite video creation processing, and the like. An area (embedding area) into which the video content is embedded is set in each virtual frame. The embedding area is set to a transparency of 100% so that the video content appears when the virtual frame is superimposed on a frame (referred to as a real frame) of the real video. In addition, the aspect ratio of the embedding area is set to 16:9 or 4:3, which is the same as the aspect ratio of a general video. In addition, the storage unit 25 stores a code list. The code list is a list of a plurality of codes embedded in the audio content. By referring to the code list, the type of the code extracted from the real audio can be identified.

[0020] The audio signal processing unit 26 extracts a code (sound signal) included in the real audio by executing a predetermined process on the audio data representing the real audio. A known method can be adopted as a method for extracting a code from audio data. For example, the audio signal processing unit 26 extracts a code from audio content included in the real audio by combining existing signal processes such as AD conversion processing, filtering processing, and decoding processing. The audio signal processing unit 26 refers to a code list stored in the storage unit 25 to identify the type of the extracted code.

[0021] The display frame identification unit 27 identifies the size and position of the display frame in which the video content is displayed in the real image. A known method can be adopted as a method for identifying the display frame in the real image. For example, the display frame identification unit 27 can identify the display frame in the real image by pattern matching using a pattern having a preset aspect ratio. In addition, when the display frame identification unit 27 inputs the real image, it can identify the display frame by using a trained model that has been machine-learned to output the display frame in the real image.

[0022] The composite image generating unit 28 generates a composite image by combining the virtual moving image with the real image so that the virtual moving image is frame-synchronized with the moving image content in the real image. The composite image generating unit 28 includes a selecting unit 29, a preprocessing unit 30, and a combining unit 31.

[0023] The selection unit 29 selects a virtual frame to be combined with a real frame input by the video input unit 21 from among the multiple virtual frames. Specifically, the selection unit 29 selects a virtual frame to be combined with a real frame from among the multiple virtual frames in accordance with the frame order. The selection unit 29 also determines whether or not the virtual video is synchronized with video content when a time code is extracted from the real audio. When the virtual video is synchronized with the video content, the selection unit 29 selects a virtual frame to be combined with a real frame from among the multiple virtual frames in accordance with the frame order. When the virtual video is not synchronized with the video content, the selection unit 29 selects a virtual frame corresponding to the time code from among the multiple virtual frames as a virtual frame to be combined with a real frame.

[0024] The pre-processing unit 30 performs pre-processing on the virtual frame so that the video content can be fitted into the fitting area set in the virtual frame. The pre-processing unit 30 will be described in detail later. The combining unit 31 combines the virtual frames preprocessed by the preprocessing unit 30 with the real frames to generate image data of a composite image frame.

[0025] Hereinafter, with reference to FIG. 4, the sound signal extracted from the real sound by the sound signal processing unit 26 will be described. The real sound includes sound content. As shown in FIG. 4, sound signals such as an identification code, a start code, a time code, and an end code are embedded in the sound content by digital watermarking technology. The identification code is identification information for identifying the video content, and is embedded in the audio content at a predetermined time point after the start of the main part of the video content. The extraction of the identification code triggers the start of the composite video distribution process. The start code is information that triggers the start of the composite video creation process, and is embedded in the audio content 10 seconds before the start of the distribution of the composite video. The time code is information used to correct the deviation of synchronization of the virtual video with the video content, and is information that can identify the current frame of the video content being displayed on the display. For example, the time code is information that indicates the time on the video content, information that indicates the elapsed time from a predetermined point such as the distribution point of the composite video, and the like. The time code may be embedded in the audio content at equal time intervals or at unequal time intervals. The end code is information that triggers the end of the video production process, and is embedded in the audio content just before the end of the composite video distribution. If the identification code, start code, time code, and end code are considered as one code set, only one code set may be embedded in the audio content, or multiple code sets may be embedded.

[0026] Hereinafter, the display frame specifying unit 27, the preprocessing unit 30, and the combining unit 31 will be described in detail with reference to FIG. 5, the display frame specification unit 27 specifies the size and position of a display frame D1 of the video content displayed on the television 100 in the real frame Gf1. Note that the video content may be displayed over the entire area of ​​the display of the television 100, the smartphone 110, or the like. Therefore, the display frame D1 of the video content may be the display frame.

[0027] The preprocessing unit 30 resizes the virtual frame Kf1 based on the size of the display frame D1. For example, the preprocessing unit 30 enlarges or reduces the virtual frame Kf1 while maintaining the aspect ratio so that the size of the fitting area Kr1 in the virtual frame Kf1 matches the size of the display frame D1. The resized virtual frame is denoted as Kf2, and the fitting area in the virtual frame Kf2 is denoted as Kr2. The preprocessing unit 30 trims the virtual frame Kf2 based on the position of the display frame D1. Specifically, the preprocessing unit 30 trims a range Tr1 that fits within the real frame Gf1 while aligning the fitting area Kr2 of the virtual frame Kf2 with the display frame D1. The trimmed virtual frame is denoted as Kf3, and the fitting area in the virtual frame Kf3 is denoted as Kr3. The preprocessing unit 30 sets the joining position of the virtual frame Kf3 with respect to the real frame Gf1 based on the position of the display frame D1. Specifically, the preprocessing unit 30 sets the joining position of the virtual frame Kf3 with respect to the real frame Gf1 such that the fitting area Kr3 of the virtual frame Kf3 coincides with the display frame D1 of the real frame Gf1.

[0028] The combining unit 31 combines the virtual frame Kf3 after pre-processing (resizing and trimming) is performed on the virtual frame Kf1 with the real frame Gf1. The combined frame is denoted as Mf1. Frame Mf1 is one frame that constitutes the composite video. In frame Mf1 of the composite video, a dotted line represents the boundary between the video content and the virtual video. By accurately aligning the position of the virtual frame Kf3 with the real frame Gf1, the image of the video content and the image of the virtual video are continuous at the boundary, and a composite video that does not feel strange due to discontinuity of the images can be provided to the user.

[0029] Hereinafter, the composite video distribution process by the information processing device 1 according to the first embodiment will be described with reference to FIG. 6 and FIG. 7. FIG. 6 is a flowchart showing the procedure of preparation process for composite video distribution process. As shown in FIG. 6, a real video captured by a camera 51 mounted on the head mounted display 5 and a real sound detected by a microphone 53 mounted on the head mounted display 5 are input to the information processing device 1 (S11, S12). The information processing device 1 executes audio signal processing on the audio data representing the real sound (S13). The process of step S13 is repeatedly executed until an identification code of the video content is extracted from the real sound (S14; No). Upon extracting the identification code of the video content from the real sound (S14; Yes), the information processing device 1 acquires a virtual video file corresponding to the identification code from the server device 7 and stores it (S15). Then, the information processing device 1 waits until a start code is extracted from the audio data representing the real sound (S16; No). When the information processing device 1 extracts a start code from the audio data representing the real audio (S16; Yes), it waits until a predetermined waiting time has elapsed since the extraction (S17; No), and when the predetermined waiting time has elapsed (S17; Yes), it starts the composite image creation process shown in Figure 7.

[0030] Hereinafter, the procedure of the composite video creation process will be described with reference to FIG. 7. FIG. 7 is a flowchart showing the procedure of the composite video creation process in the composite video distribution process. As shown in FIG. 7, the information processing device 1 initializes the frame number (m) of the virtual video, and sets the initial value (1) to the frame number (m) (S18). Next, the information processing device 1 specifies the size and position of the display frame in the real frame (S19), resizes the (m)th virtual frame of the virtual video based on the size of the display frame (S20), trims the resized (m)th virtual frame based on the position of the display frame (S21), and sets the combining position. Then, the information processing device 1 creates a composite video frame by combining the (m)th virtual frame of the virtual video subjected to the processes of steps S20 and S21 with the real frame based on the position of the display frame (S22), and transmits image data of the created composite video frame to the head mounted display 5 (S23). The information processing device 1 repeatedly executes the processes of steps S19 to S23 while incrementing the frame number (m) (S28) until an end code or a time code is extracted from the audio data representing the real audio (S24; No, S25; No). As a result, a composite video combining the real video and the virtual video is displayed on the display device 55 of the head mounted display 5. The composite video creation process is executed upon extraction of the start code embedded in the audio content, so in principle, the video content and the virtual video should be frame-synchronized in the created composite video. However, the processing time in the head mounted display 5, the processing time in the information processing device 1, the data download time, and the like are accumulated, and there is a possibility that the virtual video will not be synchronized with the video content and will be out of sync. The information processing device 1 uses the time code embedded in the audio content to correct the time lag of the virtual video relative to the video content.

[0031] Specifically, when the information processing device 1 extracts a time code from audio data representing real audio (S25; Yes), it compares the frame number (n) of the video content represented by the time code with the frame number (m) of the virtual video. If the frame number (m) of the virtual video does not match the frame number (n) of the video content represented by the time code (S26; No), the information processing device 1 determines that the video content and the virtual video are not frame-synchronized, aligns the frame number (m) of the virtual video to the frame number (n) of the video content (S27), and transitions to processing of step S28. On the other hand, if the frame number (m) of the virtual video matches the frame number (n) of the video content represented by the time code (S26; Yes), the information processing device 1 determines that the video content and the virtual video are frame-synchronized, transitions to processing of step S28 without correcting the frame number (m), and the frame number (m) is incremented. The information processing device 1 repeatedly executes the processes of steps S19 to S28 until an end code is extracted from the audio data representing real audio (S24; Yes). As a result, even if the virtual video is out of sync in time with the video content in the composite video, the time code embedded in the audio content can correct the lag, and a composite video in which the video content and the virtual video are frame-synchronized can be displayed on the display device 55 of the head mounted display 5.

[0032] According to the information processing device 1 according to the first embodiment described above, a composite image in which video content included in a real image and a virtual video are integrated can be displayed on the display device 55 of the head mounted display 5. The virtual video can widen the video range of the video content, decorate the video content, increase the dramatic effect of the video content, and supplement content that cannot be expressed by the video content alone. Therefore, the composite image can be said to be video content with a higher level of interest than the video content alone, and can provide a new video experience to the user wearing the head mounted display 5 that cannot be obtained when viewing the video content alone.

[0033] Furthermore, even if the virtual video becomes out of sync with the video content for some reason during the creation of the composite video, the information processing device 1 can correct the frame number of the virtual video to be combined with the real video at any time based on the time code embedded in the audio content, and can frame-synchronize the video content and the virtual video in the composite video. This makes it possible to reduce the discomfort and strangeness felt by the user due to the missynchronization of the virtual video with the video content in the composite video.

[0034] When a sound for a virtual moving image (referred to as a virtual sound) is associated with the virtual moving image, the information processing device 1 may have a function of generating a composite sound by mixing the virtual sound with the sound content. When creating the composite sound, the information processing device 1 corrects the timing of mixing the virtual sound with the sound content based on the time code embedded in the sound content, in the same way as when creating a composite video. This allows the virtual sound to be linked to the virtual moving image.

[0035] In the first embodiment, the video content displayed on the display of the television 100, the smartphone 110 owned by the user, etc. is expanded. However, the device on which the video content is displayed is not limited to these. Hereinafter, with reference to FIG. 8 to FIG. 10, an information processing device 1 according to a modified example of the first embodiment will be described. The hardware configuration and the functional configuration of the information processing device 1 according to the modified example are the same as those of the first embodiment. Therefore, detailed description of these will be omitted. As shown in FIG. 8, in the modified example, the video content is displayed on a display 210 equipped on a pinball game machine (pachinko machine) 200. Of course, the video content may be displayed on a display equipped on other types of game machines such as a slot machine (pachislot machine), or may be displayed on a large screen in an amusement park or a movie theater, or on a display installed in a game center.

[0036] In the modified example, the real image is an image captured at an angle of view including the display 210 mounted on the pachinko machine 200. The video content is a presentation video displayed on the display 210 mounted on the pachinko machine 200. For example, a presentation video for notifying a jackpot, a presentation video for notifying a loss, a presentation video in a probability bonus state, and the like correspond to the video content. The real sound includes audio content corresponding to the video content being displayed on the display 210 of the pachinko machine 200. For example, the audio content is background music, music, dialogue, sound effects, and the like linked to the video content output from a speaker (not shown) equipped on the pachinko machine 200.

[0037] With reference to FIG. 9, a sound signal extracted from real sound by the sound signal processor 26 will be described. Real sound includes sound content. As shown in FIG. 9, an identification code, a start code, a time code, and an end code are embedded in the sound content by digital watermarking technology. The identification code is embedded in the sound content after it is determined that the big win has occurred and the big win performance video is determined. The start code is embedded, for example, at the start point of display of the big win performance video. The time code may be embedded in the sound content at equal time intervals or at unequal time intervals. The end code is embedded in the sound content immediately before the end of the big win performance video. In addition, when the identification code, the start code, the time code, and the end code are one code set, only one code set may be embedded in the sound content, or multiple code sets may be embedded.

[0038] Hereinafter, the display frame specifying unit 27, the preprocessing unit 30, and the combining unit 31 will be described in detail with reference to Fig. 10. As shown in Fig. 10, the display frame specifying unit 27 specifies the size and position of the display frame D2 of the video content displayed on the display 210 in the real frame Gf2.

[0039] The preprocessing unit 30 resizes the virtual frame Kf5 based on the size of the display frame D2. For example, the preprocessing unit 30 enlarges or reduces the virtual frame Kf5 while maintaining the aspect ratio so that the size of the fitting area Kr5 in the virtual frame Kf5 matches the size of the display frame D2. The resized virtual frame is denoted as Kf6, and the fitting area in the virtual frame Kf6 is denoted as Kr6. The preprocessing unit 30 trims the virtual frame Kf6 based on the position of the display frame D2. Specifically, the preprocessing unit 30 trims a range Tr2 that fits within the real frame Gf2 while aligning the fitting area Kr6 of the virtual frame Kf6 with the display frame D2. The trimmed virtual frame is denoted as Kf7, and the fitting area in the virtual frame Kf7 is denoted as Kr7. The preprocessing unit 30 sets the joining position of the virtual frame Kf7 with respect to the real frame Gf2 based on the position of the display frame D2. Specifically, the preprocessing unit 30 sets the joining position of the virtual frame Kf7 with respect to the real frame Gf2 so that the fitting area Kr7 of the virtual frame Kf7 coincides with the display frame D2 of the real frame Gf2.

[0040] The combining unit 31 combines the virtual frame Kf7 obtained after pre-processing (resizing and trimming) has been performed on the virtual frame Kf5 with the real frame Gf2. The combined frame is denoted as Mf2. Frame Mf2 is one frame that constitutes the composite video. In frame Mf2 of the composite video, the dotted line represents the boundary between the video content and the virtual video. By accurately aligning the position of the virtual frame Kf7 with the real frame Gf2, the image of the video content and the image of the virtual video are continuous at the boundary, and a composite video that does not feel strange due to discontinuity of the images can be provided to the user.

[0041] Second embodiment In the first embodiment, a composite image obtained by combining a real image captured by a camera 51 mounted on the head mounted display 5 with a virtual moving image is displayed on the head mounted display 5. That is, the first embodiment is based on the assumption of a fully immersive head mounted display. However, by applying the first embodiment, it is possible to provide a new video experience to a user wearing a see-through head mounted display.

[0042] Hereinafter, the information processing device 2 according to the second embodiment will be described with reference to Fig. 11. The information processing device 2 according to the second embodiment corresponds to a case where a see-through type head mounted display is used. Fig. 11 is a functional configuration diagram of an information processing device 2 according to the second embodiment. As shown in Fig. 11, the information processing device 2 according to the second embodiment has a video input unit 21, an audio input unit 22, a video transmission unit 23, a virtual moving image acquisition unit 24, a storage unit 25, an audio signal processing unit 26, a display frame specification unit 27, and a playback unit 40. In the information processing device 2 according to the second embodiment, the functions related to the audio input unit 22, the video transmission unit 23, the virtual moving image acquisition unit 24, the storage unit 25, the audio signal processing unit 26, and the display frame specification unit 27 are similar to those in the first embodiment, and therefore detailed description thereof will be omitted.

[0043] The playback unit 40 plays back the virtual moving image so that the virtual moving image is synchronized with the moving image content that the user is directly viewing through the display device 55. Specifically, the playback unit 40 has a selection unit 41, a preprocessing unit 42, and an image data generation unit 43.

[0044] The selection unit 41 selects a virtual frame to be played from the plurality of virtual frames. The virtual frame to be played is a virtual frame displayed on the display device 55 of the head mounted display 5, and is a virtual frame to be processed by the preprocessing unit 42 and the image data generating unit 43. Specifically, the selection unit 41 selects a virtual frame to be played from the plurality of virtual frames in accordance with the frame order. In addition, the selection unit 41 determines whether or not the virtual video is synchronized with the video content when a time code is extracted from the real sound. When the virtual video is synchronized with the video content, the selection unit 41 selects a virtual frame to be played from the plurality of virtual frames in accordance with the frame order. When the virtual video is not synchronized with the video content, the selection unit 41 selects a virtual frame corresponding to the time code as a virtual frame to be played from the plurality of virtual frames.

[0045] The pre-processing unit 42 performs pre-processing on the virtual frame so that the video content that the user is viewing through the display device 55 of the head mounted display 5 is fitted into the fitting area set in the virtual frame. Details of the pre-processing are the same as those in the first embodiment.

[0046] The image data generating unit 43 generates image data corresponding to the virtual frame preprocessed by the preprocessing unit 42. The image data of the virtual frame generated by the image data generating unit 43 is transmitted to the head mounted display 5 by the video transmitting unit 23. As a result, the virtual video is displayed on the display device 55 of the head mounted display 5. The user can simultaneously view the virtual video displayed on the display device 55 of the head mounted display 5 together with the video content appearing through the display device 55 of the head mounted display 5, thereby viewing a new image in which the video content and the virtual video are fused together. Furthermore, even if the virtual video becomes out of sync with the video content due to some factor when playing back the virtual video, the information processing device 2 can correct the synchronization gap by correcting the frame number of the virtual video to be played back as needed based on the time code embedded in the audio content corresponding to the video content. As a result, it is possible to suppress the discomfort felt by the user due to the synchronization gap between the virtual video and the video content.

[0047] Third embodiment In the first embodiment, the composite image is an image in which a virtual moving image is combined with a real image. However, if a new video experience can be provided to the user, the target to be combined with the real image is not limited to a virtual moving image. For example, the composite image may be an image in which a two-dimensional model is combined with a real image. Hereinafter, an information processing device according to the third embodiment will be described with reference to Figs. 12 and 13.

[0048] FIG. 12 is a functional configuration diagram of the information processing device 3 according to the third embodiment. As shown in FIG. 12, the information processing device 3 according to the third embodiment has a video input unit 21, an audio input unit 22, a video transmission unit 23, a three-dimensional model acquisition unit 61, a storage unit 62, an audio signal processing unit 26, a display frame specification unit 63, a two-dimensional model generation unit 64, and a composite video generation unit 65. In the information processing device 3 according to the third embodiment, the functions related to the video input unit 21, the audio input unit 22, and the audio signal processing unit 26 are similar to those in the first embodiment, so detailed explanations are omitted. Note that the server device 7 stores a plurality of three-dimensional model files. Each of the three-dimensional model files is associated with an identification code for identifying video content. When the server device 7 receives a request to acquire a three-dimensional model file from the information processing device 3, it reads out the three-dimensional model file associated with the identification code included in the acquisition request from the storage device and transmits it to the information processing device 3.

[0049] The three-dimensional model acquisition unit 61 acquires, from the server device 7, a three-dimensional model file corresponding to the identification code extracted by the audio signal processing unit 26. The three-dimensional model file acquired by the three-dimensional model acquisition unit 61 is stored in the storage unit 62.

[0050] The storage unit 62 stores a three-dimensional model file acquired from the server device 7. The three-dimensional model file includes a waiting time from when the start code is extracted until the composite image creation process starts, and a plurality of ordered three-dimensional models. The ordering indicates the display order, and the display interval is set in advance. The plurality of three-dimensional models are displayed according to the ordering. As a result, the three-dimensional models are dynamically displayed. The three-dimensional models are models intended to decorate the video content, models intended to increase the production effect of the video content, models linked with the video content, and the like. The three-dimensional model file includes information on the relative size of the two-dimensional model with respect to the size of the display frame, and information on the relative position of the two-dimensional model with respect to the position of the display frame. For example, the information on the relative position of the two-dimensional model represents the position of the feature point of the two-dimensional model with respect to the center position of the display frame, and is given as display coordinates with the center position of the display frame as the origin. The information on the relative size and relative position of the two-dimensional model with respect to the display frame is used for preprocessing of the two-dimensional model by the preprocessing unit 67.

[0051] The display frame specification unit 63 specifies the orientation, size, and position of the display frame of the video content displayed on the display in the real image. For example, the orientation of the display frame can be specified by a known method such as pattern matching.

[0052] The two-dimensional model generating unit 64 generates a plurality of two-dimensional models corresponding to the plurality of three-dimensional models, respectively, based on the orientation of the display frame specified by the display frame specifying unit 63.

[0053] The composite image generating unit 65 generates a composite image by combining a real image with a two-dimensional model so that the movement of the two-dimensional model is synchronized with the video content. The composite image generating unit 65 includes a selecting unit 66, a preprocessing unit 67, and a combining unit 68. The selection unit 66 selects a 2D model to be combined with the real frame from the multiple 2D models. Specifically, the selection unit 66 selects a 2D model to be combined with the real frame from the multiple 2D models according to the ordering. In addition, the selection unit 66 determines whether or not the movement of the 2D model is synchronized with the video content when a time code is extracted from the real audio. When the 2D model is synchronized with the video content, the selection unit 66 selects a 2D model to be combined with the real frame from the multiple 2D models according to the ordering. When the 2D model is not synchronized with the video content, the selection unit 66 selects a 2D model corresponding to the time code from the multiple 2D models as a 2D model to be combined with the real frame. The preprocessing unit 67 performs preprocessing on the two-dimensional model so that the size and position of the two-dimensional model relative to the size and position of the display frame become predetermined relative size and position. The combining unit 68 combines the two-dimensional model preprocessed by the preprocessing unit 67 with the real frame to generate image data of a composite image frame.

[0054] Hereinafter, a process for generating a composite image in which a two-dimensional model is combined with a real image will be described with reference to Fig. 13. Fig. 13 shows an example in which a two-dimensional model that resembles a firework is combined with a real image. A plurality of three-dimensional models that resemble fireworks are stored in the storage unit 62. The plurality of three-dimensional models correspond to a plurality of times from when the fireworks are shot up from the ground into the sky, when they burst into flames in the sky, and when they disappear.

[0055] 13, the display frame specification unit 63 specifies the orientation, size, and position of the display frame D3 of the video content displayed on the television 300 in the real frame Gf3. The orientation of the display frame D3 is expressed by the up-down angle and the left-right angle of the user's line of sight (i.e., the imaging direction of the camera) with respect to the front direction, assuming that the direction perpendicular to the display surface of the television 300 is the front direction. For example, when the orientation of the display frame D3 is up-down (-5 degrees) and left-right (+12 degrees), this means that the line of sight of the user watching the television 300 is tilted downward by 5 degrees with respect to the front direction and tilted left by 12 degrees with respect to the front direction.

[0056] The two-dimensional model generation unit 64 performs rendering processing on the three-dimensional model Gm1 based on the orientation of the display frame D3 (the user's line of sight) to create a two-dimensional model Gm2. The two-dimensional model Gm2 is a planar model when the three-dimensional model Gm1 is projected from the orientation of the display frame D3 (the user's line of sight).

[0057] The pre-processing unit 67 resizes the two-dimensional model Gm2 based on the size of the display frame D3. For example, the pre-processing unit 67 enlarges or reduces the two-dimensional model Gm2 so that the size of the two-dimensional model Gm2 becomes a preset relative size with respect to the size of the display frame D3. The resized two-dimensional model is denoted as Gm3. The pre-processing unit 67 sets the joining position of the two-dimensional model Gm3 based on the position of the display frame D3. Specifically, the joining position of the two-dimensional model Gm3 is set so that the position of the two-dimensional model Gm3 becomes a preset relative position with respect to the position of the display frame D3.

[0058] A combining unit 68 combines the preprocessed two-dimensional model Gm3 with the real frame Gf3. The combined frame Mf3 is one frame that constitutes the compound image.

[0059] The information processing device 3 according to the third embodiment described above provides the same effects as those of the first embodiment. That is, a composite image in which the video content included in the real image and the two-dimensional model are integrated can be displayed on the display device 55 of the head mounted display 5. The two-dimensional model decorates the video content, enhances the dramatic effect of the video content, and can supplement the content that cannot be expressed by the video content alone. Therefore, the composite image can be said to be a more interesting video content than the video content alone, and can provide a new video experience to the user wearing the head mounted display 5 that cannot be obtained when watching the video content alone.

[0060] Furthermore, even if the display order of the 2D models is delayed relative to the video content for some reason when creating the composite video, the information processing device 3 can synchronize the video content and the movements of the 2D models in the composite video by correcting the order of the 2D models to be combined with the real video as needed based on the time code embedded in the audio content corresponding to the video content. This makes it possible to suppress the discomfort caused by the missynchronization of the movements of the 2D models relative to the video content in the composite video.

[0061] The information processing devices 1, 2, and 3 according to the first, second, and third embodiments may have the functions of the server device 7. That is, the information processing devices may store a virtual video file and a three-dimensional model file.

[0062] The functions of the information processing devices 1, 2, and 3 according to the first, second, and third embodiments and the functions of the server device 7 may all be mounted on the head mounted display 5, and the head mounted display 5 may function as a stand-alone device. For example, the head mounted display 5 having the functions of the information processing device 1 according to the first embodiment and the functions of the server device 7 has the following configuration. That is, the head mounted display 5 has a camera 51 that captures a real image captured at an angle of view including the display in the real space, a microphone 53 that detects real sound in the real space, a display device 55 that displays a composite image, and a storage device 14 that stores a virtual video, and has a function of extracting a sound signal embedded in the sound of the video content displayed on the display from the real sound, and a function of combining the virtual video with the real image so that the virtual video is frame-synchronized with the video content based on the sound signal, and generating a composite video. Similarly, the head mounted display 5 can be configured to have all the functions according to the second embodiment, and can also be configured to have all the functions according to the third embodiment.

[0063] Although some embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included in the scope of the invention and its equivalents described in the claims, as well as in the scope and spirit of the invention. [Explanation of symbols]

[0064] 1...information processing device, 5...head mounted display, 7...server device, 9...network, 10...data / control bus, 11...processor, 12...RAM, 13...ROM, 14...storage device, 15...communication device, 21...video input unit, 22...audio input unit, 23...video transmission unit, 24...virtual moving image acquisition unit, 25...memory unit, 26...audio signal processing unit, 27...display frame identification unit, 28...composite image generation unit.

Claims

1. A storage unit that stores the virtual video; an image input unit that inputs a real image captured by a camera equipped in the head mounted display with an angle of view including the display in real space; a voice input unit that inputs a real voice of the real space detected by a microphone equipped in the head mounted display; an extracting unit that extracts, from the real sound, a sound signal embedded in the sound of the video content displayed on the display; a composite image generating unit that combines the virtual moving image with the real image based on the sound signal so that the virtual moving image is frame-synchronized with the moving image content, thereby generating a composite image; a transmitting unit that transmits data of the composite image to the head mounted display; An information processing device comprising:

2. the sound signal includes a time code representing a time on the video content, the composite image generating unit selects a frame corresponding to the time code from a plurality of frames constituting the virtual moving image, and combines the selected frame with a frame of the real image; 2. The information processing device according to claim 1.

3. a display frame specifying unit that specifies a size and a position of a display frame of the video content from the real image, the composite image generating unit sets a frame size of the virtual moving image based on a size of the display frame, and sets a combining position of the frame of the virtual moving image with respect to the frame of the real image based on a position of the display frame.

2. The information processing device according to claim 1.

4. the composite image generating unit trims frames of the virtual video so that the frames of the virtual video fit within the frames of the real image based on a combination position of the frames of the virtual video with respect to the frames of the real image; 4. The information processing device according to claim 3.

5. A storage unit that stores a plurality of ordered three-dimensional models; an image input unit that inputs a real image captured by a camera equipped in the head mounted display with an angle of view including the display in real space; a display frame specification unit that specifies an orientation of a display frame of video content displayed on the display in the real image; a two-dimensional model generating unit that generates a plurality of two-dimensional models corresponding to the plurality of three-dimensional models, based on the identified orientations; a voice input unit that inputs a real voice of the real space detected by a microphone equipped in the head mounted display; an extraction unit that extracts a sound signal embedded in the audio of the video content from the real audio; a composite image generating unit that generates a composite image by combining the two-dimensional model with the real image based on the sound signal so that the two-dimensional model is synchronized with the video content; a transmitting unit that transmits the composite image data to the head mounted display; An information processing device comprising:

6. the sound signal includes a time code representing a time on the video content, the composite image generating unit selects a two-dimensional model corresponding to the time code from the plurality of two-dimensional models, and combines the selected two-dimensional model with a frame of the real image.

6. The information processing device according to claim 5.

7. The display frame specifying unit specifies a size and a position of the display frame as well as an orientation of the display frame; the composite image generating unit sets a size of the two-dimensional model based on a size of the display frame, and sets a position of the two-dimensional model with respect to a frame of the real image based on a position of the display frame.

6. The information processing device according to claim 5.

8. A storage unit that stores the virtual video; an image input unit that inputs a real image captured by a camera equipped in the head mounted display with an angle of view including the display in real space; a voice input unit that inputs a real voice of the real space detected by a microphone equipped in the head mounted display; an extracting unit that extracts, from the real sound, a sound signal embedded in the sound of the video content displayed on the display; a playback unit that plays the virtual moving image based on the sound signal so that the virtual moving image is frame-synchronized with the moving image content; a transmission unit that transmits the reproduced virtual video to the head mounted display; An information processing device comprising:

9. the sound signal includes a time code representing a time on the video content, the reproduction unit selects a frame corresponding to the time code from a plurality of frames constituting the virtual moving image, and reproduces the selected frame.

9. The information processing device according to claim 8.

10. a display frame specifying unit that specifies a size and a position of a display frame of the video content from the real image, the composite image generating unit sets a frame size of the virtual moving image based on a size of the display frame, and sets a display position of the frame of the virtual moving image relative to the frame of the real image based on a position of the display frame.

9. The information processing device according to claim 8.

11. the composite image generating unit trims the frames of the virtual video based on display positions of the frames of the virtual video relative to the frames of the real image so that the frames of the virtual video fit within a display range of a display device of the head mounted display. The information processing device according to claim 10.

12. A camera that captures a real image captured at an angle of view that includes a display in real space; a microphone for detecting real sound in the real space; A storage unit that stores the virtual video; an extracting unit that extracts, from the real sound, a sound signal embedded in the sound of the video content displayed on the display; a composite image generating unit that combines the virtual moving image with the real image based on the sound signal so that the virtual moving image is frame-synchronized with the moving image content, thereby generating a composite image; a display unit for displaying the composite image; A head mounted display comprising:

13. A computer that stores virtual videos A means for inputting a real image captured by a camera mounted on the head mounted display with an angle of view including the display in real space; a means for inputting a real sound of the real space detected by a microphone equipped in the head mounted display; means for extracting, from the real sound, a sound signal embedded in the sound of the video content displayed on the display; means for combining the virtual animation with the real image based on the sound signal such that the virtual animation is frame synchronized with the animation content to generate a composite image; means for transmitting the composite image data to the head mounted display; A program to achieve this.

Citation Information

Patent Citations

  • Wearable smart glasses

    EP3096517A1

  • Apparatus and method for manipulating instruments in the anatomy

    JP2007526788A

  • Method for providing mobile device with second screen information

    JP2015061112A

  • Sound conversion adapter

    JP2015076695A

  • Computer program

    JP2018081410A