Information Processing Apparatus, Head-Mounted Display, and Program

The information processing apparatus addresses the challenge of synchronizing virtual and real-world content in head-mounted displays by extracting identification codes from audio data to retrieve and synchronize virtual video with real-world video, resulting in a seamless and engaging composite video experience.

JP7689173B2Active Publication Date: 2025-06-05TSUBURAYA FIELDS HLDG CO LTD
View PDF 12 Cites 0 Cited by

Patent Information

Application Number
JP2023192325
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-11-10
Publication Date
2025-06-05
Estimated Expiration
2043-11-10

AI Technical Summary

Technical Problem

Existing virtual space technologies, such as augmented and mixed reality, struggle to provide users with a seamless and synchronized video experience when combining real-world videos with virtual content in head-mounted displays.

Method used

An information processing apparatus that includes a head-mounted display with a camera and microphone, connected via short-range wireless communication to an external server. This apparatus processes real-world video and audio data to extract identification codes, which are used to retrieve and synchronize virtual video content with the real-world video, generating a composite video for display.

Benefits of technology

The solution provides a new video experience by seamlessly integrating virtual and real-world content, ensuring synchronization between video content and virtual video, thus enhancing user engagement and reducing discomfort caused by synchronization deviations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007689173000001
    Figure 0007689173000001
  • Figure 0007689173000002
    Figure 0007689173000002
  • Figure 0007689173000003
    Figure 0007689173000003
Patent Text Reader

Abstract

To provide a user with a new visual experience utilizing a virtual space technique.SOLUTION: An information processing device comprises: a storage unit which stores a virtual moving image; a video input unit which inputs a real-world video captured in an angle of view including a display in a real-world space by a camera mounted on a head-mounted display; an audio input unit which inputs real-world audio in the real-world space detected by a microphone mounted on the head-mounted display; an extraction unit which extracts an audio signal embedded in the audio of a moving image content displayed on the display from the real-world audio; a composite video generation unit which combines the virtual moving image with the real-world video such that the virtual moving image is frame-synchronized with the moving image content on the basis of the audio signal and generates a composite video; and a transmission unit which transmits the composite video data to the head-mounted display.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to an information processing apparatus , Head-Mounted Display and a program.

Background Art

[0002] Virtual space technologies such as augmented reality (AR) and mixed reality (MR), which superimpose virtual objects, virtual images, etc. on a video of the real space captured by a camera mounted on a head-mounted display (HMD), are known (for example, Patent Document 1). Since videos of the actual real space can be used, these technologies are expected to be applied to various fields such as surgical simulations, teaching assembly procedures to workers, explaining artworks in art museums, and introducing furniture in newly built houses.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] To provide a user with a new video experience using virtual space technology.

Means for Solving the Problems

[0005] An information processing apparatus according to the present disclosure includes A head-mounted display having a camera for imaging a display, a microphone for detecting sound output from a speaker connected to the display, and a display unit is connected under a short-range wireless communication standard and connected to an external server device for virtual video distribution via an Internet line. The information processing device includes a receiving unit that receives, from the head-mounted display, data of an image captured by the camera and data related to sound detected by the microphone, and extracts a code by executing predetermined processing on the data related to sound, and determines whether the code is an identification code by referring to a code list. When an identification code is extracted from the data related to sound, a request for acquiring data of a virtual video identified by the identification code is transmitted to the external server device, and an acquisition unit that receives the data of the virtual video transmitted from the external server device in response to the acquisition request, combines the received virtual video with the image captured by the camera, a composite video generation unit that generates a composite video, and for display on the display unit, transmits the data of the composite video to a head-mounted display Composite Video Transmission Unit and includes the above.

Brief Description of the Drawings

[0006]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

BEST MODE FOR CARRYING OUT THE INVENTION

[0007] Hereinafter, with reference to the drawings, an information processing apparatus according to an implementation mode of the present invention will be described. In the following description, components having substantially the same functions and configurations are denoted by the same reference numerals, and redundant descriptions are made only when necessary.

[0008] (First Embodiment) The information processing apparatus according to the first embodiment is a computer device having a function of creating a composite video in which a virtual video is combined with a real video captured at an angle of view including a display on which video content is displayed by a camera mounted on a head-mounted display, and a function of transmitting the created composite video to the head-mounted display. Thereby, a new video experience in which video content is extended can be provided to a user wearing a head-mounted display. Further, in creating the composite video, the information processing apparatus according to the first embodiment synchronizes the video content and the virtual video in frame using the real sound detected by a microphone mounted on the head-mounted display. Thereby, it is possible to avoid a situation in which a user wearing a head-mounted display feels discomfort due to a synchronization deviation between the video content and the virtual video in the composite video.

[0009] Typically, the information processing apparatus according to the first embodiment is configured as follows. In the first embodiment, terms are defined as follows. User: A wearer wearing a head-mounted display. Real video: A video related to the real space captured by a camera mounted on a head-mounted display (HMD). In particular, in the first embodiment, the real video is a video captured at an angle of view including a display such as a television or a smartphone arranged in the real space. Video content: A video being displayed on a display arranged in the real space. For example, in the first embodiment, videos such as movies, dramas, concerts, variety shows, and sports correspond to video content. Virtual video: A video that is combined with real - world video. The virtual video is created in advance according to the video content. For example, the virtual video is a video aimed at expanding the video range represented by the video content, a video aimed at decorating the video content, a video aimed at enhancing the production effect of the video content, a video aimed at providing users with new information about the video content, etc. A video aimed at expanding the video range represented by the video content is a video having a pattern continuous with the pattern represented by the video content. A video aimed at providing users with information about the video content is a video of comments from viewers on the video content or a video for explaining the content of the video content. Composite video: A video obtained by combining a virtual video with real - world video. Real - world audio: Audio related to the real - world space detected by a microphone equipped on the head - mounted display. In particular, in the first embodiment, the real - world audio includes audio content. Of course, the real - world audio may include not only audio content but also ambient sounds and the user's voice. Audio content: Audio corresponding to video content. For example, the audio content is output from a speaker connected to the display. For example, in the first embodiment, the BGM of the video content, the lines of characters appearing in the video content, the sound effects of the video content, etc. correspond to the audio content.

[0010] As shown in FIG. 1, the information processing apparatus 1 according to the first embodiment (hereinafter simply referred to as the information processing apparatus 1) constitutes a video distribution system together with the server apparatus 7 and the head - mounted display 5. The information processing apparatus 1 is a computer device such as a PC, a tablet, or a smartphone.

[0011] The information processing device 1 and the server device 7 are communicably connected to each other via a network 9 such as the Internet. A plurality of virtual video files are stored in the server device 7. An identification code for identifying video content is associated with each virtual video file. When the server device 7 receives a request to acquire a virtual video file from the information processing device 1, the server device 7 reads out from the storage device the virtual video file associated with the identification code included in the acquisition request and transmits it to the information processing device 1.

[0012] The information processing device 1 and the head-mounted display 5 are communicably connected to each other by short-range wireless communication such as Bluetooth. The head-mounted display 5 includes a display device 55 disposed at a position that blocks the user's field of view, a camera 51 that images the user's line-of-sight direction, a microphone 53 that detects ambient sound, and a speaker (not shown) that outputs the sound detected by the microphone 53 to the user. The display device 55 displays the real-world video imaged by the camera 51. Since the real-world video corresponds to the user's line-of-sight direction, the user can view the real-world video as if directly looking at the real space with their eyes. The video data representing the real-world video imaged by the camera 51 and the audio data representing the ambient sound detected by the microphone 53 are transmitted to the information processing device 1. The information processing device 1 creates a composite video by combining the real-world video imaged by the camera 51 and the virtual video obtained from the server device 7, and transmits the created composite video to the head-mounted display 5. Thereby, the composite video created by the information processing device 1 is displayed on the display device 55 of the head-mounted display 5.

[0013] As shown in FIG. 2, the information processing apparatus 1 has a processor 11. A RAM 12, a ROM 13, a storage device 14, and a communication device 15 are connected to the processor 11 via a data / control bus 10. The processor 11 has a CPU and a GPU. The RAM 12 functions as a main memory, a work area, etc. of the processor 11. The ROM 13 or the storage device 14 stores a BIOS (Basic Input Output System), an OS (operation system), and a composite video distribution program to be executed by the processor 11. The communication device 15 transmits and receives data to and from the server device 7 and the head-mounted display 5 according to the control of the processor 11.

[0014] As shown in FIG. 3, when the distribution program stored in the storage device 14 is executed by the processor 11, the information processing apparatus 1 functions as a video input unit 21, an audio input unit 22, a video transmission unit 23, a virtual video acquisition unit 24, a storage unit 25, an audio signal processing unit 26, a display frame specifying unit 27, and a composite video generation unit 28.

[0015] The video input unit 21 is realized by the function of the communication device 15 shown in FIG. 2. The video input unit 21 inputs video data representing real-world video captured by a camera 51 equipped on the head-mounted display 5.

[0016] The audio input unit 22 is realized by the function of the communication device 15 shown in FIG. 2. The audio input unit 22 inputs audio data representing real-world audio detected by a microphone 53 equipped on the head-mounted display 5.

[0017] The video transmission unit 23 is realized by the function of the communication device 15 shown in FIG. 2. The video transmission unit 23 transmits the image data of the composite video generated by the composite video generation unit 28 to the head-mounted display 5.

[0018] The virtual video acquisition unit 24 acquires a virtual video file corresponding to the identification code extracted by the audio signal processing unit 26 from the server device 7. The virtual video file acquired by the virtual video acquisition unit 24 is stored in the storage unit 25.

[0019] The storage unit 25 is realized by the function of the storage device 14 shown in FIG. 2. The storage unit 25 stores various types of information necessary for the processing related to the distribution program (composite video distribution processing). Specifically, the storage unit 25 stores the virtual video file acquired from the server device 7. The virtual video file includes data of a plurality of frames (referred to as virtual frames) constituting the virtual video, data related to the waiting time from extracting the start code until starting the composite video creation process, and the like. For each virtual frame, an area (embedding area) where video content is to be embedded is set. The embedding area is set with a transparency of 100% so that the video content appears when the virtual frame is overlaid on the frame of the real video (referred to as the real frame). Also, the aspect ratio of the embedding area is set to 16:9 or 4:3, which is the same as the aspect ratio of a general video. Further, the storage unit 25 stores a code list. The code list is a list of a plurality of codes embedded in the audio content. By referring to the code list, the type of the code extracted from the real audio can be specified.

[0020] The audio signal processing unit 26 extracts a code (audio signal) included in the real audio by performing a predetermined process on the audio data representing the real audio. As a method for extracting the code from the audio data, a known method can be adopted. For example, the audio signal processing unit 26 extracts the code from the audio content included in the real audio by combining existing signal processes such as AD conversion processing, filtering processing, and decoding processing. The audio signal processing unit 26 refers to the code list stored in the storage unit 25 to specify the type of the extracted code.

[0021] The display frame specifying unit 27 specifies the size and position of the display frame in which the video content in the real video is displayed. As a method for specifying the display frame in the real video, a known method can be adopted. For example, the display frame specifying unit 27 can specify the display frame in the real video by pattern matching using a pattern having a preset aspect ratio. Also, when the real video is input, the display frame specifying unit 27 can specify the display frame by using a learned model that has been machine-learned to output the display frame in the real video.

[0022] The composite video generation unit 28 generates a composite video in which the virtual video is combined with the real video so that the virtual video is frame-synchronized with the video content in the real video. The composite video generation unit 28 includes a selection unit 29, a preprocessing unit 30, and a combining unit 31.

[0023] The selection unit 29 selects a virtual frame to be combined with the real frame input by the video input unit 21 from a plurality of virtual frames. Specifically, the selection unit 29 selects, in frame order, a virtual frame to be combined with the real frame from the plurality of virtual frames. Also, when the time code is extracted from the real audio, the selection unit 29 determines whether the virtual video is synchronized with the video content. When the virtual video is synchronized with the video content, the selection unit 29 selects, in frame order, a virtual frame to be combined with the real frame from the plurality of virtual frames. When the virtual video is not synchronized with the video content, the selection unit 29 selects, as the virtual frame to be combined with the real frame from the plurality of virtual frames, the virtual frame corresponding to the time code.

[0024] The preprocessing unit 30 performs preprocessing on the virtual frame so that the video content can be embedded in the embedding area set in the virtual frame. Details of the preprocessing unit 30 will be described later. The combining unit 31 generates the image data of the frame of the composite video in which the virtual frame preprocessed by the preprocessing unit 30 is combined with the real frame.

[0025] Hereinafter, with reference to FIG. 4, the sound signal extracted from the actual sound by the audio signal processing unit 26 will be described. The actual sound includes audio content. As shown in FIG. 4, in the audio content, sound signals such as an identification code, a start code, a time code, and an end code are embedded by an electronic watermarking technique. The identification code is identification information for identifying video content. For example, it is embedded in the audio content at a time point after a predetermined time has elapsed since the start of the main part of the video content. The composite video distribution process is started upon extraction of the identification code. The start code is information that triggers the start of the composite video creation process. For example, it is embedded in the audio content 10 seconds before the start of the distribution of the composite video. The time code is information used to correct the synchronization deviation of the virtual video with respect to the video content, and is information capable of specifying the current frame of the video content being displayed on the display. For example, the time code is information representing the time on the video content, information representing the elapsed time from a predetermined time point such as the distribution time of the composite video, and the like. The time code may be embedded at equal time intervals or unequal time intervals with respect to the audio content. The end code is information that triggers the end of the video creation process, and is embedded in the audio content immediately before the end of the distribution of the composite video. Note that when the identification code, the start code, the time code, and the end code are regarded as one code set, only one code set may be embedded in the audio content, or a plurality of code sets may be embedded.

[0026] Hereinafter, with reference to FIG. 5, the details of the display frame specifying unit 27, the preprocessing unit 30, and the combining unit 31 will be described. As shown in FIG. 5, the display frame specifying unit 27 specifies the size and position of the display frame D1 of the video content displayed on the television 100 within the actual frame Gf1. Note that the video content may be displayed over the entire area of a display such as the television 100 or the smartphone 110. Therefore, the display frame D1 of the video content may be the display frame.

[0027] The preprocessing unit 30 resizes the virtual frame Kf1 based on the size of the display frame D1. For example, the preprocessing unit 30 enlarges or reduces the virtual frame Kf1 while maintaining the aspect ratio so that the size of the embedding area Kr1 within the virtual frame Kf1 matches the size of the display frame D1. The resized virtual frame is denoted as Kf2, and the embedding area within the virtual frame Kf2 is denoted as Kr2. The preprocessing unit 30 trims the virtual frame Kf2 based on the position of the display frame D1. Specifically, the preprocessing unit 30 trims the embedding area Kr2 of the virtual frame Kf2 to the range Tr1 that fits within the real frame Gf1 with the embedding area Kr2 of the virtual frame Kf2 aligned with the display frame D1. The trimmed virtual frame is denoted as Kf3, and the embedding area within the virtual frame Kf3 is denoted as Kr3. The preprocessing unit 30 sets the combination position of the virtual frame Kf3 with respect to the real frame Gf1 based on the position of the display frame D1. Specifically, the preprocessing unit 30 sets the combination position of the virtual frame Kf3 with respect to the real frame Gf1 so that the embedding area Kr3 of the virtual frame Kf3 matches the display frame D1 of the real frame Gf1.

[0028] The combining unit 31 combines the virtual frame Kf3 after the preprocessing is performed on the virtual frame Kf1 (after the resizing process and the trimming process) with the real frame Gf1. The combined frame is denoted as Mf1. The frame Mf1 is one frame that constitutes the composite video. In the frame Mf1 of the composite video, the dotted line represents the boundary line between the video content and the virtual video. By accurately aligning the virtual frame Kf3 with respect to the real frame Gf1, it is possible to provide the user with a composite video in which the pattern of the video content and the pattern of the virtual video are continuous at the boundary line and there is no sense of discomfort due to the discontinuity of the pattern.

[0029] Next, with reference to FIGS. 6 and 7, the composite video distribution process by the information processing apparatus 1 according to the first embodiment will be described. FIG. 6 is a flowchart showing the procedure of the preparation process for the composite video distribution process. As shown in FIG. 6, the real video captured by the camera 51 equipped in the head-mounted display 5 and the real audio detected by the microphone 53 equipped in the head-mounted display 5 are input to the information processing apparatus 1 (S11, S12). The information processing apparatus 1 executes audio signal processing on the audio data representing the real audio (S13). The process of step S13 is repeatedly executed until the identification code of the video content is extracted from the real audio (S14; No). When the information processing apparatus 1 extracts the identification code of the video content from the real audio (S14; Yes), it acquires and stores the virtual video file corresponding to the identification code from the server apparatus 7 (S15). Then, the information processing apparatus 1 waits until the start code is extracted from the audio data representing the real audio (S16; No). When the information processing apparatus 1 extracts the start code from the audio data representing the real audio (S16; Yes), it waits until a predetermined standby time elapses after the extraction (S17; No), and when the predetermined standby time elapses (S17; Yes), it starts the composite video creation process shown in FIG. 7.

[0030] The following describes the procedure of the composite video creation process with reference to FIG. 7. FIG. 7 is a flowchart showing the procedure of the composite video creation process in the composite video distribution process. As shown in FIG. 7, the information processing apparatus 1 initializes the frame number (m) of the virtual video and sets the initial value (1) to the frame number (m) (S18). Next, the information processing apparatus 1 specifies the size and position of the display frame within the real frame (S19), resizes the m-th virtual frame of the virtual video based on the size of the display frame (S20), trims the resized m-th virtual frame based on the position of the display frame (S21), and sets the combination position. Then, the information processing apparatus 1 creates a frame of the composite video in which the m-th virtual frame of the virtual video, for which the processes of steps S20 and S21 have been performed, is combined with the real frame based on the position of the display frame (S22), and transmits the image data of the created composite video frame to the head-mounted display 5 (S23). The information processing apparatus 1 repeatedly executes the processes of steps S19 to S23 while incrementing the frame number (m) until it extracts the start code and time code embedded in the audio content (S24; No, S25; No). As a result, a composite video in which the real video and the virtual video are combined is displayed on the display device 55 of the head-mounted display 5. Since the composite video creation process is executed triggered by the extraction of the start code embedded in the audio content, in principle, the video content and the virtual video should be frame-synchronized in the created composite video. However, the processing time in the head-mounted display 5, the processing time in the information processing apparatus 1, the data download time, etc. accumulate, and there is a possibility that the virtual video may be out of sync with the video content and deviate. The information processing apparatus 1 uses the time code embedded in the audio content to correct the temporal deviation of the virtual video with respect to the video content.

[0031] Specifically, when the information processing apparatus 1 extracts a time code from the audio data representing the real voice (S25; Yes), it compares the frame number (m) of the virtual video with the frame number (n) of the video content represented by the time code. If the frame number (m) of the virtual video does not match the frame number (n) of the video content represented by the time code (S26; No), the information processing apparatus 1 determines that the video content and the virtual video are not frame-synchronized, aligns the frame number (m) of the virtual video with the frame number (n) of the video content (S27), and transitions to the process of step S28. On the other hand, if the frame number (m) of the virtual video matches the frame number (n) of the video content represented by the time code (S26; Yes), the information processing apparatus 1 determines that the video content and the virtual video are frame-synchronized, and without correcting the frame number (m), transitions to the process of step S28, and the frame number (m) is incremented. The information processing apparatus 1 repeatedly executes the processes of steps S19 to S28 until it extracts an end code from the audio data representing the real voice (S24; Yes). Thereby, even if the virtual video is temporally shifted with respect to the video content in the composite video, the shift can be corrected by the time code embedded in the audio content, and a composite video in which the video content and the virtual video are frame-synchronized can be displayed on the display device 55 of the head-mounted display 5.

[0032] According to the information processing apparatus 1 according to the first embodiment described above, a composite video in which the video content included in the real video and the virtual video are integrated can be displayed on the display device 55 of the head-mounted display 5. The virtual video can expand the video range of the video content, decorate the video content, increase the production effect of the video content, and supplement content that cannot be expressed only by the video content. Therefore, it can be said that the composite video is a video content with higher interest than the video content alone, and a new video experience that cannot be obtained when the user views the video content alone can be provided to the user wearing the head-mounted display 5.

[0033] Furthermore, when creating the composite video, even if the synchronization of the virtual video with respect to the video content is shifted due to some factor, the information processing apparatus 1 can correct the frame number of the virtual video to be combined with the real video at any time based on the time code embedded in the audio content, and can perform frame synchronization between the video content and the virtual video in the composite video. Thereby, it is possible to suppress the discomfort and sense of incongruity of the user due to the synchronization deviation of the virtual video with respect to the video content in the composite video.

[0034] In addition, when virtual audio (referred to as virtual audio) for the virtual video is associated with the virtual video, the information processing apparatus 1 may have a function of generating a composite audio in which the virtual audio is mixed with the audio content. When creating the composite audio, the information processing apparatus 1 corrects the timing of mixing the virtual audio with respect to the audio content based on the time code embedded in the audio content, in the same manner as when creating the composite video. Thereby, the virtual audio can be linked to the virtual video.

[0035] In the first embodiment, it was for expanding the video content displayed on a display such as the television 100 or the smartphone 110 held by the user. However, the device on which the video content is displayed is not limited to these. Hereinafter, with reference to FIGS. 8 to 10, the information processing apparatus 1 according to a modification of the first embodiment will be described. The hardware configuration and the functional configuration of the information processing apparatus 1 according to the modification are the same as those of the first embodiment. Therefore, the details of these descriptions will be omitted. As shown in FIG. 8, in the modification, the video content is displayed on the display 210 provided in the pachinko machine 200. Of course, the video content may be displayed on a display provided in another type of gaming machine such as a spinning reel gaming machine (pachislot machine), or may be displayed on a large screen in an amusement park or a movie theater, or a display arranged in a game center.

[0036] In a modified example, the real video is a video captured at an angle of view including the display 210 mounted on the pachinko machine 200. The video content is a production video displayed on the display 210 mounted on the pachinko machine 200. For example, production videos for jackpot notifications, losing notifications, production videos in the probability-variable state, etc. correspond to the video content. The real audio includes audio content corresponding to the video content being displayed on the display 210 of the pachinko machine 200. For example, the audio content is BGM, music, dialogue, sound effects, etc. linked to the video content output from a speaker (not shown) equipped on the pachinko machine 200.

[0037] Referring to FIG. 9, the sound signal extracted from the real audio by the audio signal processing unit 26 will be described. The real audio includes audio content. As shown in FIG. 9, in the audio content, an identification code, a start code, a time code, and an end code are embedded by an electronic watermarking technique. The identification code is determined to be a jackpot, and is embedded in the audio content after the production video for the jackpot is determined. The start code is embedded, for example, at the start time of display of the production video for the jackpot. The time code may be embedded at equal time intervals or unequal time intervals with respect to the audio content. The end code is embedded in the audio content at a time point just before the end of the production video for the jackpot. Note that when the identification code, the start code, the time code, and the end code are regarded as one code set, only one code set may be embedded in the audio content, or a plurality of code sets may be embedded.

[0038] Hereinafter, referring to FIG. 10, the details of the display frame specifying unit 27, the preprocessing unit 30, and the combining unit 31 will be described. As shown in FIG. 10, the display frame specifying unit 27 specifies the size and position of the display frame D2 of the video content displayed on the display 210 within the real frame Gf2.

[0039] The preprocessing unit 30 resizes the virtual frame Kf5 based on the size of the display frame D2. For example, the preprocessing unit 30 enlarges or reduces the virtual frame Kf5 while maintaining the aspect ratio so that the size of the embedding area Kr5 within the virtual frame Kf5 matches the size of the display frame D2. The resized virtual frame is denoted as Kf6, and the embedding area within the virtual frame Kf6 is denoted as Kr6. The preprocessing unit 30 trims the virtual frame Kf6 based on the position of the display frame D2. Specifically, the preprocessing unit 30 trims the embedding area Kr6 of the virtual frame Kf6 to the range Tr2 that fits within the real frame Gf2 with the embedding area Kr6 aligned with the display frame D2. The trimmed virtual frame is denoted as Kf7, and the embedding area within the virtual frame Kf7 is denoted as Kr7. The preprocessing unit 30 sets the combination position of the virtual frame Kf7 with respect to the real frame Gf2 based on the position of the display frame D2. Specifically, the preprocessing unit 30 sets the combination position of the virtual frame Kf7 with respect to the real frame Gf2 such that the embedding area Kr7 of the virtual frame Kf7 matches the display frame D2 of the real frame Gf2.

[0040] The combining unit 31 combines the virtual frame Kf7 after preprocessing is performed on the virtual frame Kf5 (after the resizing process and the trimming process) with the real frame Gf2. The combined frame is denoted as Mf2. The frame Mf2 is one frame constituting the composite video. In the frame Mf2 of the composite video, the dotted line represents the boundary line between the video content and the virtual video. By accurately aligning the virtual frame Kf7 with respect to the real frame Gf2, it is possible to provide the user with a composite video in which the patterns of the video content and the virtual video are continuous at the boundary line and there is no sense of discomfort due to the discontinuity of the patterns.

[0041] (Second Embodiment) In the first embodiment, a composite video obtained by combining a virtual video with a real video captured by a camera 51 mounted on a head-mounted display 5 is displayed on the head-mounted display 5. That is, the first embodiment assumes a fully immersive head-mounted display. However, by applying the first embodiment, a new video experience can be provided to a user wearing a see-through type head-mounted display as well.

[0042] Hereinafter, the information processing apparatus 2 according to the second embodiment will be described with reference to FIG. 11. The information processing apparatus 2 according to the second embodiment corresponds to the case where a see-through type head-mounted display is used. FIG. 11 is a functional configuration diagram of the information processing apparatus 2 according to the second embodiment. As shown in FIG. 11, the information processing apparatus 2 according to the second embodiment includes a video input unit 21, an audio input unit 22, a video transmission unit 23, a virtual video acquisition unit 24, a storage unit 25, an audio signal processing unit 26, a display frame specifying unit 27, and a playback unit 40. In the information processing apparatus 2 according to the second embodiment, the functions of the audio input unit 22, the video transmission unit 23, the virtual video acquisition unit 24, the storage unit 25, the audio signal processing unit 26, and the display frame specifying unit 27 are the same as those in the first embodiment, and thus detailed descriptions thereof are omitted.

[0043] The playback unit 40 plays back a virtual video so that the virtual video is synchronized with video content that the user is directly viewing through the display device 55. Specifically, the playback unit 40 includes a selection unit 41, a preprocessing unit 42, and an image data generation unit 43.

[0044] The selection unit 41 selects a virtual frame to be played back from a plurality of virtual frames. The virtual frame to be played back is a virtual frame displayed on the display device 55 of the head-mounted display 5 and is a virtual frame to be processed by the preprocessing unit 42 and the image data generation unit 43. Specifically, the selection unit 41 selects a virtual frame to be played back from a plurality of virtual frames according to the frame order. Further, when the time code is extracted from the real audio, the selection unit 41 determines whether the virtual video is synchronized with the video content. When the virtual video is synchronized with the video content, the selection unit 41 selects a virtual frame to be played back from a plurality of virtual frames according to the frame order. When the virtual video is not synchronized with the video content, the selection unit 41 selects a virtual frame corresponding to the time code as the virtual frame to be played back from a plurality of virtual frames.

[0045] The preprocessing unit 42 performs preprocessing on the virtual frame so that the video content being viewed by the user through the display device 55 of the head-mounted display 5 is fitted into the embedding area set in the virtual frame. The details of the preprocessing are the same as those in the first embodiment.

[0046] The image data generation unit 43 generates image data corresponding to the virtual frame preprocessed by the preprocessing unit 42. The image data of the virtual frame generated by the image data generation unit 43 is transmitted to the head-mounted display 5 by the video transmission unit 23. Thereby, a virtual video is displayed on the display device 55 of the head-mounted display 5. The user can view a new video in which the video content and the virtual video are fused by simultaneously viewing the virtual video displayed on the display device 55 of the head-mounted display 5 together with the video content appearing through the display device 55 of the head-mounted display 5. Also, when playing the virtual video, even if the synchronization between the virtual video and the video content is shifted for some reason, the information processing apparatus 2 can correct the synchronization shift by correcting the frame number of the virtual video to be played at any time based on the time code embedded in the audio content corresponding to the video content. Thereby, it is possible to suppress the discomfort felt by the user due to the synchronization shift between the virtual video and the video content.

[0047] (Third Embodiment) In the first embodiment, the composite video was a video in which a virtual video was combined with the real video. However, as long as a new video experience can be provided to the user, the object to be combined with the real video is not limited to the virtual video. For example, the composite video may be a video in which a 2D model is combined with the real video. Hereinafter, with reference to FIGS. 12 and 13, the information processing apparatus according to the third embodiment will be described.

[0048] FIG. 12 is a functional configuration diagram of the information processing apparatus 3 according to the third embodiment. As shown in FIG. 12, the information processing apparatus 3 according to the third embodiment includes a video input unit 21, an audio input unit 22, a video transmission unit 23, a three-dimensional model acquisition unit 61, a storage unit 62, an audio signal processing unit 26, a display frame specifying unit 63, a two-dimensional model generation unit 64, and a composite video generation unit 65. In the information processing apparatus 3 according to the third embodiment, the functions of the video input unit 21, the audio input unit 22, and the audio signal processing unit 26 are the same as those in the first embodiment, and thus detailed descriptions thereof are omitted. Note that a plurality of three-dimensional model files are stored in the server device 7. An identification code for identifying video content is associated with each of the three-dimensional model files. When the server device 7 receives a request for acquiring a three-dimensional model file from the information processing apparatus 3, the server device 7 reads out the three-dimensional model file associated with the identification code included in the acquisition request from the storage device and transmits it to the information processing apparatus 3.

[0049] The three-dimensional model acquisition unit 61 acquires, from the server device 7, a three-dimensional model file corresponding to the identification code extracted by the audio signal processing unit 26. The three-dimensional model file acquired by the three-dimensional model acquisition unit 61 is stored in the storage unit 62.

[0050] The storage unit 62 stores the 3D model file acquired from the server device 7. The 3D model file includes the waiting time from when the start code is extracted until the start of the composite video creation process, and a plurality of 3D models arranged in order. The order represents the display order, and the display interval is set in advance. The plurality of 3D models are displayed according to the order. As a result, the 3D models are dynamically displayed. The 3D models are models for the purpose of decorating video content, models for the purpose of enhancing the production effect of video content, models linked to video content, etc. The 3D model file includes information regarding the relative size of the 2D model with respect to the size of the display frame, and information regarding the relative position of the 2D model with respect to the position of the display frame. For example, the information regarding the relative position of the 2D model represents the position of the feature points of the 2D model with respect to the center position of the display frame, and is given as display coordinates with the center position of the display frame as the origin. The information regarding the relative size and relative position of the 2D model with respect to the display frame is used for preprocessing the 2D model by the preprocessing unit 67.

[0051] The display frame specifying unit 63 specifies the orientation, size, and position of the display frame of the video content displayed on the display within the real video. For example, the orientation of the display frame can be specified by a known method such as pattern matching.

[0052] The 2D model generation unit 64 generates a plurality of 2D models corresponding to the plurality of 3D models respectively based on the orientation of the display frame specified by the display frame specifying unit 63.

[0053] The composite video generation unit 65 generates a composite video in which the 2D model is combined with the real video so that the movement of the 2D model is synchronized with the video content. The composite video generation unit 65 includes a selection unit 66, a preprocessing unit 67, and a combining unit 68. The selection unit 66 selects a 2D model to be combined with the real frame from a plurality of 2D models. Specifically, the selection unit 66 selects, in accordance with the ordering, a 2D model to be combined with the real frame from the plurality of 2D models. Further, when a time code is extracted from the real audio, the selection unit 66 determines whether the movement of the 2D model is synchronized with the video content. When the 2D model is synchronized with the video content, the selection unit 66 selects, in accordance with the ordering, a 2D model to be combined with the real frame from the plurality of 2D models. When the 2D model is not synchronized with the video content, the selection unit 66 selects, as the 2D model to be combined with the real frame, the 2D model corresponding to the time code from the plurality of 2D models. The preprocessing unit 67 performs preprocessing on the 2D model so that the size and position of the 2D model relative to the size and position of the display frame become predetermined relative sizes and relative positions. The combining unit 68 generates image data of a frame of a composite video obtained by combining the 2D model preprocessed by the preprocessing unit 67 with the real frame.

[0054] Hereinafter, with reference to FIG. 13, the generation process of a composite video obtained by combining a 2D model with a real video will be described. FIG. 13 shows an example in which a 2D model simulating a launched fireworks is combined with a real video. A plurality of 3D models simulating launched fireworks are stored in the storage unit 62. The plurality of 3D models respectively correspond to a plurality of times from when the launched fireworks rise from the ground into the sky, open in the sky, and disappear.

[0055] As shown in FIG. 13, the display frame specifying unit 63 specifies the orientation, size, and position of the display frame D3 of the video content displayed on the television 300 within the real frame Gf3. The orientation of the display frame D3 is represented by the vertical and horizontal angles of the user's line-of-sight direction (i.e., the imaging direction of the camera) with respect to the front direction, where the direction orthogonal to the display surface of the television 300 is taken as the front direction. For example, when the orientation of the display frame D3 is (-5 degrees) in the vertical direction and (+12 degrees) in the horizontal direction, it means that the line-of-sight direction of the user viewing the television 300 is a direction inclined 5 degrees downward and 12 degrees to the left with respect to the front direction.

[0056] Based on the orientation of the display frame D3 (the user's line-of-sight direction), the two-dimensional model generation unit 64 performs rendering processing on the three-dimensional model Gm1 to create a two-dimensional model Gm2. The two-dimensional model Gm2 is a planar model when the three-dimensional model Gm1 is projected from the orientation of the display frame D3 (the user's line-of-sight direction).

[0057] The preprocessing unit 67 resizes the two-dimensional model Gm2 based on the size of the display frame D3. For example, the preprocessing unit 67 enlarges or reduces the two-dimensional model Gm2 so that the size of the two-dimensional model Gm2 becomes a preset relative size with respect to the size of the display frame D3. The resized two-dimensional model is denoted as Gm3. The preprocessing unit 67 sets the combination position of the two-dimensional model Gm3 based on the position of the display frame D3. Specifically, the combination position of the two-dimensional model Gm3 is set so that the position of the two-dimensional model Gm3 becomes a preset relative position with respect to the position of the display frame D3.

[0058] The combining unit 68 combines the preprocessed two-dimensional model Gm3 with the real frame Gf3. The combined frame Mf3 is one frame constituting the composite video.

[0059] According to the information processing apparatus 3 according to the third embodiment described above, the same effects as those of the first embodiment can be achieved. That is, the display device 55 of the head-mounted display 5 can display a composite video in which the video content included in the real video and the 2D model are integrated. The 2D model can decorate the video content, increase the production effect of the video content, and supplement the content that cannot be expressed only by the video content. Therefore, it can be said that the composite video is a video content with higher interest than the video content alone, and a new video experience that cannot be obtained when the user views the video content alone can be provided to the user wearing the head-mounted display 5.

[0060] Furthermore, when creating the composite video, even if the display order of the 2D model lags behind the video content for some reason, the information processing apparatus 3 can, based on the time code embedded in the audio content corresponding to the video content, correct the order of the 2D models combined with the real video at any time, thereby synchronizing the movement of the video content and the 2D model in the composite video. Thereby, the discomfort caused by the synchronization deviation of the movement of the 2D model with respect to the video content in the composite video can be suppressed.

[0061] The information processing apparatuses 1, 2, and 3 according to the first, second, and third embodiments may have the functions of the server device 7. That is, virtual video files and 3D model files may be stored in the information processing apparatus.

[0062] All of the functions of the information processing apparatuses 1, 2, and 3 and the functions of the server apparatus 7 according to the first, second, and third embodiments may be mounted on the head-mounted display 5 so that the head-mounted display 5 functions as a stand-alone device. For example, the head-mounted display 5 having the functions of the information processing apparatus 1 and the functions of the server apparatus 7 according to the first embodiment has the following configuration. That is, the head-mounted display 5 includes a camera 51 that captures a real video imaged at an angle of view including a display in the real space, a microphone 53 that detects real sound in the real space, a display device 55 that displays a composite video, and a storage device 14 that stores virtual videos. The head-mounted display 5 has a function of extracting a sound signal embedded in the sound of video content displayed on the display from the real sound, and a function of combining the virtual video with the real video so that the virtual video is frame-synchronized with the video content based on the sound signal, and generating a composite video. Similarly, the head-mounted display 5 can be configured to have all the functions according to the second embodiment, and can also be configured to have all the functions according to the third embodiment.

[0063] Although some embodiments of the present invention have been described, these embodiments are presented by way of example and are not intended to limit the scope of the invention. These embodiments can be implemented in various other forms, and various omissions, replacements, and changes can be made without departing from the gist of the invention. These embodiments and their modifications are included in the scope and gist of the invention, and are also included in the invention described in the claims and the equivalent scope thereof.

Explanation of Reference Numerals

[0064] 1... Information processing apparatus, 5... Head-mounted display, 7... Server apparatus, 9... Network, 10... Data / control bus, 11... Processor, 12... RAM, 13... ROM, 14... Storage device, 15... Communication device, 21... Video input unit, 22... Audio input unit, 23... Video transmission unit, 24... Virtual video acquisition unit, 25... Storage unit, 26... Audio signal processing unit, 27... Display frame specification unit, 28... Composite video generation unit.

Claims

1. An information processing apparatus that is connected under a short-range wireless communication standard to a head-mounted display having a camera for imaging a display, a microphone for detecting sound output from a speaker connected to the display, and a display unit, and is connected via an Internet line to an external server device for virtual video distribution, a receiving unit that receives, from the head-mounted display, data of an image captured by the camera and data related to sound detected by the microphone; a voice signal processing unit that extracts a code by performing predetermined processing on the data related to the sound, and determines whether the code is an identification code or a start code by referring to a code list; an acquisition unit that, when the identification code is extracted from the data related to the sound, transmits a request for acquisition of data of a virtual video identified by the identification code to the external server device, and receives the data of the virtual video transmitted from the external server device in response to the acquisition request; a composite video generation unit that combines the received virtual video with the video captured by the camera to generate a composite video; a composite video transmission unit that transmits the data of the composite video to the head-mounted display for display on the display unit, and comprises: the composite video generation unit waits for generation of the composite video from when the data of the virtual video is received until the start code is extracted from the data related to the sound, and after a predetermined waiting time has elapsed since the start code is extracted from the data related to the sound, starts a process of combining the received virtual video with the video captured by the camera to generate the composite video. Information processing apparatus.

2. The voice signal processing unit extracts a code by performing predetermined processing on the data related to the sound, and determines whether the code is a time code by referring to a code list, and when the time code is extracted from the data related to the sound, the composite video generation unit selects a frame corresponding to the time code from a plurality of frames constituting the virtual video, and combines the selected frame with a frame of the video captured by the camera. The information processing apparatus according to claim 1.

3. The apparatus further comprises a display frame specifying unit that specifies the size and position of the display frame from the video imaged by the camera. The composite video generation unit sets the frame size of the virtual video based on the size of the display frame, and sets the combination position of the frame of the virtual video with respect to the frame of the video imaged by the camera based on the position of the display frame. The information processing apparatus according to claim 1.

4. The composite video generation unit trims the frame of the virtual video so that the frame of the virtual video fits within the frame of the video imaged by the camera, based on the combination position of the frame of the virtual video with respect to the frame of the video imaged by the camera. The information processing apparatus according to claim 3.

5. An information processing apparatus connected under a short-range wireless communication standard to a head-mounted display having a camera for imaging a display, a microphone for detecting sound output from a speaker connected to the display, and a display unit, and connected via an Internet line to an external server device for 3D model distribution, a receiving unit that receives, from the head-mounted display, data of the video imaged by the camera and data related to the sound detected by the microphone; a voice signal processing unit that extracts a code by performing predetermined processing on the data related to the sound, and determines whether the code is an identification code or a start code by referring to a code list; when the identification code is extracted from the data related to the sound, a acquisition unit that transmits a request for acquisition of data of a plurality of 3D models identified by the identification code to the external server device, and receives the data of the plurality of 3D models transmitted from the external server device in response to the acquisition request; a 2D model generation unit that generates a plurality of 2D models respectively corresponding to the plurality of 3D models; a composite video generation unit that combines the 2D model with the video imaged by the camera to generate a composite video; and a composite video transmission unit that transmits the data of the composite video to the head-mounted display for display on the display unit. The composite video generation unit waits for the generation of the composite video from when the data of the 3D model is received until the start code is extracted from the data related to the audio, and after the start code is extracted from the data related to the audio and a predetermined waiting time has elapsed, it combines the 2D model with the video captured by the camera and starts the process of generating the composite video. An information processing apparatus.

6. The audio signal processing unit extracts a code by performing a predetermined process on the data related to the audio, and determines whether the code is a time code by referring to a code list. When the time code is extracted from the data related to the audio, the composite video generation unit selects a 2D model corresponding to the time code from the plurality of 2D models, and combines the selected 2D model with the frame of the video captured by the camera. The information processing apparatus according to claim 5.

7. The apparatus further includes a display frame specifying unit that specifies the orientation, size, and position of the display frame of the display from the video captured by the camera. Based on the specified orientation, the 2D model generation unit generates a plurality of 2D models respectively corresponding to the plurality of 3D models. The composite video generation unit sets the size of the 2D model based on the size of the display frame, and sets the position of the 2D model with respect to the frame of the video captured by the camera based on the position of the display frame. The information processing apparatus according to claim 5.

8. A camera for imaging a display, A microphone for detecting audio output from a speaker connected to the display, A storage unit that stores data of a plurality of virtual videos, An audio signal processing unit that extracts a code by performing a predetermined process on the data related to the audio detected by the microphone, and determines whether the code is an identification code or a start code by referring to a code list. A selection unit that selects one virtual video identified by the identification code from the plurality of virtual videos when the identification code is extracted from the data related to the audio. A composite video generation unit that combines the one virtual video with the video captured by the camera to generate a composite video, And a display unit that displays the composite video. The composite video generation unit waits for the generation of the composite video from when the identification code is extracted until the start code is extracted from the data related to the audio, and after a predetermined waiting time has elapsed since the start code is extracted from the data related to the audio, it combines the one virtual video with the video captured by the camera and starts the process of generating the composite video. Head-mounted display.

9. A computer that is connected under a short-range wireless communication standard to a head-mounted display having a camera for imaging the display, a microphone for detecting audio output from a speaker connected to the display, and a display unit, and is connected to an external server device for virtual video distribution via the Internet line. Means for receiving, from the head-mounted display, data of the video captured by the camera and data related to the audio detected by the microphone. Means for extracting a code by performing a predetermined process on the data related to the audio, and determining whether the code is an identification code or a start code by referring to a code list. When the identification code is extracted from the data related to the audio, means for transmitting a request for acquisition of data of the virtual video identified by the identification code to the external server device, and receiving the data of the virtual video transmitted from the external server device in response to the acquisition request. Means for combining the received virtual video with the video captured by the camera to generate a composite video. Means for transmitting the data of the composite video to the head-mounted display for display on the display unit. The means for generating the composite video waits for the generation of the composite video from when the data of the virtual video is received until the start code is extracted from the data related to the audio, and after a predetermined waiting time has elapsed since the start code is extracted from the data related to the audio, it combines the received virtual video with the video captured by the camera and starts the process of generating the composite video. Program.

Citation Information

Patent Citations

  • Wearable smart glasses

    EP3096517A1

  • Apparatus and method for manipulating instruments in the anatomy

    JP2007526788A

  • Method for providing mobile device with second screen information

    JP2015061112A

  • Sound conversion adapter

    JP2015076695A

  • Information processing device, method, and program

    JP2017084323A