Information processing apparatus, control method for information processing apparatus, and program
By using acoustic input and sound recognition technology to generate CG virtual paper in XR environment, the problem of high accuracy and speed requirements in virtual keyboard gesture operations is solved, and comfortable and efficient text input and information organization is achieved.
Patent Information
- Application Number
- JP2023183241
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-25
- Publication Date
- 2025-05-12
AI Technical Summary
In XR environment, when using a virtual keyboard for gesture operation, extremely high gesture recognition accuracy and speed are required, otherwise it will lead to error pressing adjacent keys or slow operation speed, affecting the comfort and efficiency of text input.
The sound input device and sound recognition technology are used to convert voice into text and generate images of CG virtual paper. Combined with real-time video in the virtual space, the generation and synthesis of virtual paper is realized.
It realizes text input in a comfortable and efficient way in the XR environment, and is suitable for meetings, seminars and other scenarios, improving the convenience of information organization and presentation.
Smart Images

Figure 2025072846000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, a control method for an information processing device, and a program. [Background technology]
[0002] In recent years, XR (cross reality) has been developing remarkably. XR is a technology that realizes the fusion of the real world and the virtual world, and is a general term for VR (virtual reality), AR (augmented reality), and MR (mixed reality). XR has a technology for inputting characters using a virtual keyboard. Furthermore, in recent XR, a technology for inputting characters using voice has come to be used. For example, Patent Document 1 discloses a technology for arranging virtual objects of characters that are generated by converting the user's voice in a three-dimensional space. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent Publication No. 2021-5157 Summary of the Invention [Problem to be solved by the invention]
[0004] In the technology of inputting characters by a virtual keyboard, when the virtual keyboard is operated by hand gestures, the recognition accuracy of the hand gestures for the joint points of the hand must be very high. This is because, if the recognition accuracy of the hand gestures is low, the user may mistakenly press an adjacent key when operating the virtual keyboard. Also, the recognition speed of the hand gestures must be very high. This is because, if the recognition speed of the hand gestures is slow, the user can only tell whether or not a key on the virtual keyboard has been pressed at a later time from the time of pressing, and therefore the user can only perform the pressing operation slowly.
[0005] Since it is technically difficult to realize such extremely high recognition accuracy and extremely high speed hand gesture recognition, it is difficult to input characters comfortably and quickly on a virtual keyboard in XR, as if operating a real keyboard with real hands. Therefore, there is a problem that it is difficult for a user to write text such as main points and questions on a virtual object in real time during a review, seminar, or briefing session conducted in XR. In addition, when the technology disclosed in Patent Document 1 is applied, text such as main points and questions expressed by the user's voice becomes text on a virtual object and is scattered in a three-dimensional space. Therefore, there is a problem that the technology disclosed in Patent Document 1 is not suitable for organizing information during a review, seminar, or briefing session conducted in XR.
[0006] The present invention has been made in consideration of the above problems. It is an object of the present invention to provide an information processing device, a control method for the information processing device, and a program that enable easy and fast character input for organizing information in reviews, seminars, and briefing sessions conducted in XR. [Means for solving the problem]
[0007] In order to achieve the above-mentioned object, the information processing device of the present invention is characterized by comprising: a voice input means for inputting voice; a voice recognition means for recognizing the voice input to the voice input means; a first input means for inputting character data representing the content of the voice recognized by the voice recognition means into a virtual paper; a generation means for generating a CG image of the virtual paper onto which the character data has been input by the first input means; and a synthesis means for synthesizing the CG image of the virtual paper generated by the generation means into an XR space in which real space and virtual space are merged. Effect of the Invention
[0008] According to the present invention, information can be organized with comfortable and high-speed character input during reviews, seminars, and information sessions conducted in XR. [Brief description of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a configuration of a head mounted display (hereinafter, referred to as "HMD") according to an embodiment of the present invention. [Diagram 2] FIG. 2 is a block diagram showing the configuration of the HMD shown in FIG. 1 specialized and modified to the first embodiment. [Diagram 3] FIG. 1 is an image diagram for explaining generation of a sticky note of a virtual object (hereinafter referred to as a "virtual sticky note") onto which the contents of a voice (hereinafter referred to as "voice contents") are written by voice input. [Figure 4] FIG. 13 is an image diagram for explaining attaching a virtual sticky note in response to a voice input instruction. [Diagram 5] 10 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging in the first embodiment. [Figure 6] FIG. 13 is a block diagram showing a configuration of the HMD shown in FIG. 1 specialized and modified according to a second embodiment. [Figure 7] FIG. 13 is an image diagram for explaining attaching of a virtual sticky note by line-of-sight calculation. [Figure 8] 13 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging in the second embodiment. [Figure 9] FIG. 13 is a block diagram showing a configuration of the HMD shown in FIG. 1 specialized and modified according to a third embodiment. [Figure 10] 11 is an image diagram illustrating the generation of a virtual sticky note on which voice content is written by voice input, and the changing of the color of the virtual sticky note by person recognition based on the voice. [Figure 11] FIG. 13 is an image diagram for explaining generation of a virtual sticky note on which voice content is written by voice input, and writing of a name on the virtual sticky note by person recognition based on the voice. [Figure 12] 13 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging in the third embodiment. [Figure 13] FIG. 13 is a block diagram showing a configuration of the HMD shown in FIG. 1 specialized and modified to a fourth embodiment. [Figure 14]FIG. 13 is an image diagram illustrating a change in color of a virtual sticky note through emotion recognition based on voice. [Figure 15] 13 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging in the fourth embodiment. [Figure 16] FIG. 13 is a block diagram showing a configuration of the HMD shown in FIG. 1 specialized and modified to a fifth embodiment. [Figure 17] This is an image diagram to explain the generation of virtual sticky notes on which document contents are written by voice input, and the association, recording, and switching of video recording files. [Figure 18] FIG. 13 is an image diagram for explaining playback of a recorded file. [Figure 19] 13 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and compositing, and from the start to the end of recording, in the fifth embodiment. [Figure 20] 13 is a flowchart showing the flow of processing for playing a recorded file in the fifth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0010] Each embodiment of the present invention will be described in detail below with reference to the drawings. However, the configurations described in each of the following embodiments are merely examples, and the scope of the present invention is not limited by the configurations described in each embodiment. For example, each part constituting the present invention can be replaced with any configuration that can perform the same function. In addition, any configuration may be added. In addition, any two or more configurations (features) of each embodiment can be combined. In addition, in each embodiment, an HMD will be described as an application example of the present invention. Note that the present invention is not limited to an HMD, and may be applied to devices such as smart glasses, for example.
[0011] <Configuration of the HMD according to this embodiment> FIG. 1 is a block diagram showing the configuration of an HMD according to this embodiment. The HMD shown in FIG. 1 is mainly for displaying a three-dimensional space (hereinafter referred to as "XR space") in which real space and virtual space are fused on the same display device. The HMD shown in FIG. 1 can be applied to various image processing technologies (VR, AR, MR) included in XR. For example, the HMD shown in FIG. 1 may display only the virtual space on the same display device, unlike FIG. 1, by being applied to VR. The HMD is composed of a lens device, an HMD main body, and a computer graphics (hereinafter referred to as "CG") device. First, the lens device will be described. The optical system 101 has a lens unit for forming an image of a real object on the imaging unit 102.
[0012] Next, the HMD main body and the CG device will be described. The imaging unit 102 has an imaging sensor, and performs photoelectric conversion on the light image formed through the optical system 101 to output an electrical signal. The imaging signal processing unit 103 converts the electrical signal output from the imaging unit 102 into a video signal. The video signal processing unit 104 processes the video signal output from the imaging signal processing unit 103 according to the application. The image synthesis unit 105, the display unit 106, the imaging control unit 107, and the HMD control unit 108 will be described later. The power supply unit 109 supplies power to the entire HMD according to the application. The HMD operation unit 110 is an operation unit used by the user to operate the HMD, and outputs an operation signal to the HMD control unit 108.
[0013] The HMD shake detection unit 111 detects the amount of shake applied to the HMD and outputs a detection signal to the HMD control unit 108. The HMD control unit 108 has a CPU, ROM, and RAM, and controls the entire HMD. The HMD control unit 108 communicates with a CG communication control unit 113 of the CG device via an HMD communication control unit 112. That is, mutual communication is performed by the HMD communication control unit 112 and the CG communication control unit 113. The microphone control unit 114 encodes audio input from the microphone unit 115. The gaze camera control unit 116 acquires a video signal output from the gaze camera unit 117. The CG communication control unit 113 receives each piece of control information of the HMD body described above (hereinafter referred to as "HMD control information").
[0014] The CG communication control unit 113 passes the voice information and video information of the received HMD control information to the voice recognition unit 118, the person recognition unit 119, the emotion recognition unit 120, the gaze calculation unit 121, and the object recognition unit 122. The voice recognition unit 118 extracts the content of the received voice information in text form. The voice recognition unit 118 passes the extracted text to the CG generation unit 123. The CG generation unit 123 generates a CG image of a virtual sticky note to which the received text is input. The CG image of the virtual sticky note thus generated is combined with the real image output from the video signal processing unit 104 by the image synthesis unit 105 together with the three-dimensional CG space generated by the CG generation unit 123, that is, the virtual space, and is displayed on the display unit 106. As a result, the image displayed on the display unit 106 enters the user's eyes as a synthesized image via the eyepiece optical system 124.
[0015] The person recognition unit 119 recognizes a person from the received voice information by referring to the person registered in the recording unit 125. The person recognition unit 119 passes the recognized person information to the CG generation unit 123. The CG generation unit 123 changes the color of the virtual sticky note for which the CG image has been generated to a color corresponding to the received person information. The CG image of the virtual sticky note whose color has been changed in this way is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105 together with the three-dimensional CG space generated by the CG generation unit 123, and is displayed on the display unit 106. As a result, the image displayed on the display unit 106 enters the user's eyes as a synthesized image via the eyepiece optical system 124. Alternatively, the CG generation unit 123 inputs character data of a name corresponding to the received person information into the virtual sticky note for which the CG image has been generated. The CG image of the virtual sticky note into which the character data of the person's name has been input in this way is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105, together with the three-dimensional CG space generated by the CG generation unit 123, and displayed on the display unit 106. As a result, the image displayed on the display unit 106 enters the user's eye as a synthesized image via the eyepiece optical system 124.
[0016] The emotion recognition unit 120 recognizes an emotion from the received voice information. The emotion recognition unit 120 passes the recognized emotion to the CG generation unit 123. The CG generation unit 123 changes the color of the virtual sticky note that generated the CG image to a color corresponding to the received emotion. The CG image of the virtual sticky note whose color has been changed in this way is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105 together with the three-dimensional CG space generated by the CG generation unit 123, and is displayed on the display unit 106. As a result, the image displayed on the display unit 106 enters the user's eye as a synthesized image via the eyepiece optical system 124. Alternatively, the CG generation unit 123 inputs the received emotion information to the virtual sticky note that generated the CG image so that it can be used for business judgment. The CG image of the virtual sticky note into which the emotion information for business judgment has been input in this way is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105 together with the three-dimensional CG space generated by the CG generation unit 123, and is displayed on the display unit 106. As a result, the image displayed on the display unit 106 passes through the eyepiece optical system 124 and reaches the user's eye as a composite image.
[0017] In addition, in the HMD shown in FIG. 1, it is possible to paste a virtual sticky note as follows. In the object recognition unit 122, the name of the object is extracted from the received voice information, and a real object with the extracted name is image-recognized in a real image output from the video signal processing unit 104, that is, in a real space. Alternatively, in the object recognition unit 122, a virtual object with the extracted name is image-recognized in a three-dimensional CG space generated by the CG generation unit 123. The object recognition unit 122 acquires the spatial coordinates of the image-recognized real object or virtual object, and passes the acquired spatial coordinates to the CG generation unit 123. The CG generation unit 123 associates the generated CG image of the virtual sticky note with the received spatial coordinates and pastes it. The CG image of the virtual sticky note pasted in this way is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105 together with the three-dimensional CG space generated by the CG generation unit 123, and is displayed on the display unit 106. As a result, the image displayed on the display unit 106 passes through the eyepiece optical system 124 and reaches the user's eye as a composite image.
[0018] 1, it is also possible to attach a virtual sticky note as follows. The line of sight calculation unit 121 receives video information from the line of sight camera unit 117 and calculates the direction of the user's line of sight from the pupil. The line of sight calculation unit 121 passes the calculated line of sight direction information to the object recognition unit 122. The object recognition unit 122 performs image recognition of a real object in the line of sight direction of the received information in a real image output from the video signal processing unit 104. Alternatively, the object recognition unit 122 performs image recognition of a virtual object in the line of sight direction of the received information in a three-dimensional CG space generated by the CG generation unit 123.
[0019] The object recognition unit 122 acquires spatial coordinates of the image-recognized real object or virtual object, and passes the acquired spatial coordinates to the CG generation unit 123. The CG generation unit 123 associates the generated CG image of the virtual sticky note with the received spatial coordinates and pastes it. The CG image of the virtual sticky note pasted in this way, together with the three-dimensional CG space generated by the CG generation unit 123, is synthesized with the real image output from the video signal processing unit 104 by the image synthesis unit 105, and displayed on the display unit 106. As a result, the image displayed on the display unit 106 enters the user's eye as a synthesized image via the eyepiece optical system 124. The imaging control unit 107 executes exposure control, distance measurement control, and the like for the imaging unit 102. Various data such as audio information and video information are recorded in the recording unit 125.
[0020] First Embodiment Hereinafter, the first embodiment will be described with reference to FIG. 1 to FIG. 5. In the first embodiment, a virtual sticky note (virtual paper) on which voice content is written by voice input is generated in an XR space. In addition, a real object or a virtual object indicated by the name of the object input by voice is recognized. Furthermore, the generated virtual sticky note is attached to the recognized real object or virtual object. FIG. 2 is a block diagram when the configuration of the HMD shown in FIG. 1 is specialized and modified to the first embodiment. In an HMD 200 (information processing device) according to the first embodiment, an image of a real space is formed as a real image by an imaging unit 202 through an optical system 201. The formed real image is image-processed into a digital image by an image processing unit 203.
[0021] The microphone control unit 204 (voice input means) passes voice information of the voice inputted by the microphone unit 205 to the voice recognition unit 206 and the object recognition unit 207. The voice recognition unit 206 (voice recognition means) extracts the voice content from the received voice information and converts the extracted voice content into text. The voice recognition unit 206 passes the converted text to the CG generation unit 208. The CG generation unit 208 (first input means) (generation means) generates a CG image of a virtual sticky note to which the received text is inputted. The object recognition unit 207 extracts the name of the object from the received voice information. Note that the extraction of the name of the object may be performed by the voice recognition unit 206 that received the voice information. However, in this case, the voice recognition unit 206 passes information related to the extracted name of the object to the object recognition unit 207.
[0022] The object recognition unit 207 (object recognition means) acquires the spatial coordinates of a real object indicated by the object name through image recognition in the real image input from the image processing unit 203. Alternatively, the object recognition unit 207 (object recognition means) acquires the spatial coordinates of a virtual object indicated by the object name through image recognition in the three-dimensional CG space input from the CG generation unit 208, that is, the virtual space. A virtual sticky note associating unit 209 (pasting means) associates a CG image of a virtual sticky note with the acquired spatial coordinates of the real object or virtual object, thereby pasting the virtual sticky note.
[0023] The CG image of the virtual sticky note pasted in this manner, together with the three-dimensional CG space generated by the CG generation unit 208, is synthesized by the image synthesis unit 210 (synthesizing means) with the real image converted into digital image by the image processing unit 203, and displayed on the display unit 211. As a result, the image displayed on the display unit 211 is seen by the user as a synthesized image. The camera control unit 212 is part of the HMD control unit 108 in FIG. 1. The virtual sticky note association unit 209 is part of the CG generation unit 123 in FIG. 1. The image processing unit 203 corresponds to the imaging signal processing unit 103 and the video signal processing unit 104 in FIG. 1.
[0024] Fig. 3 is an image diagram for explaining the generation of a virtual sticky note on which voice content is written by voice input. In the first embodiment, for example, in a review, seminar, or briefing session held in XR, as shown in Fig. 3, when a user utters a voice 301 with the voice content of "aiueo...", a CG image of a virtual sticky note 302 on which the voice content of "aiueo..." is written is generated. This allows for hassle-free real-time character input unlike a case where a virtual keyboard is operated by hand gestures, and therefore makes it possible to provide an excellent user experience when inputting to a virtual sticky note in XR.
[0025] Fig. 4 is an image diagram for explaining pasting of a virtual sticky note by a voice input instruction. In the first embodiment, further, for example, as shown in Fig. 4, when a user utters a voice 401 of "paste it on the first button", a CG image of the virtual sticky note 302 is associated with spatial coordinates acquired by image recognition of the first button 402 of a real object or a virtual object. This allows the CG image of the virtual sticky note 302 to be pasted on the first button 402 of a real object or a virtual object without the user having to perform a hand gesture, making it possible to provide an excellent user experience with pasting the CG image of a virtual sticky note in XR.
[0026] FIG. 5 is a flowchart showing the flow of processing from generation of virtual sticky notes to pasting and merging. The processing (method of controlling an information processing device) shown in the flowchart of FIG. 5 is realized by a CPU (computer) in the HMD control unit 108 expanding a program stored in a ROM into a RAM and executing it. The processing shown in the flowchart of FIG. 5 is started in response to a predetermined operation being performed on the HMD operation unit 110. However, the processing shown in the flowchart of FIG. 5 may also be started in response to the voice recognition unit 118 recognizing, for example, voice content such as "Please start sticky note mode" input to the microphone unit 115. Note that these points are also the same for the processing shown in the flowcharts of FIG. 8, FIG. 12, FIG. 15, FIG. 19, and FIG. 20 described later.
[0027] When the process shown in the flowchart of FIG. 5 starts, first, in step S501, the microphone control unit 204 passes voice information of the voice input to the microphone unit 205 to the voice recognition unit 206 (voice input process). In step S502, the voice recognition unit 206 performs voice recognition to extract the voice content from the received voice information (voice recognition process). Furthermore, the voice recognition unit 206 converts the extracted voice content into text and passes the converted text to the CG generation unit 208. In step S503, the CG generation unit 208 inputs the received text into a virtual sticky note (input process). In step S504, the CG generation unit 208 generates a CG image of the virtual sticky note into which the text was input in step S503 (generation process).
[0028] In step S505, the microphone control unit 204 determines whether or not there is voice input to the microphone unit 205. If the microphone control unit 204 determines that there is voice input to the microphone unit 205, it passes voice information of the voice input to the microphone unit 205 to the object recognition unit 207. Thereafter, the process proceeds to step S506. On the other hand, if the microphone control unit 204 determines that there is no voice input to the microphone unit 205, the process proceeds to step S509.
[0029] In step S506, the object recognition unit 207 performs voice recognition to extract the name of the object from the received voice information. The voice recognition in step S506 may be performed by the voice recognition unit 206. In this case, however, in step S505, when the microphone control unit 204 determines that voice has been input to the microphone unit 205, it passes the voice information of the voice input to the microphone unit 205 to the voice recognition unit 206. Then, the voice recognition unit 206 passes information on the name of the object extracted by performing the voice recognition in step S506 to the object recognition unit 207.
[0030] In step S507, the object recognition unit 207 acquires the spatial coordinates of the real object indicated by the object name by image recognition in the real image input from the image processing unit 203. Alternatively, the object recognition unit 207 acquires the spatial coordinates of the virtual object indicated by the object name by image recognition in the three-dimensional CG space input from the CG generation unit 208. In step S508, the virtual sticky note association unit 209 associates the CG image of the virtual sticky note created in step S504 with the spatial coordinates of the real object or virtual object acquired in step S507, thereby pasting it. In step S509, the image synthesis unit 210 synthesizes the CG image of the virtual sticky note and the three-dimensional CG space generated by the CG generation unit 208 with the real image output from the image processing unit 203, and displays it on the display unit 211 (synthesis step). Thereafter, the process shown in the flowchart of FIG. 5 ends.
[0031] As described above, in the first embodiment, when a user speaks during a review, seminar, or information session conducted in XR, a CG image of a virtual sticky note with the contents of the voice input is generated, and the generated CG image of the virtual sticky note is composited with a real image together with a three-dimensional CG space. In this way, the HMD 200 according to the first embodiment can realize information organization with comfortable and high-speed character input during a review, seminar, or information session conducted in XR.
[0032] <Second embodiment> The second embodiment will be described below with reference to Figs. 1, 3, and 6 to 8. In the second embodiment, a virtual sticky note (virtual paper) on which voice content is written by voice input is generated in an XR space. Furthermore, the generated virtual sticky note is attached to a real object or a virtual object in the direction of the line of sight calculated from the pupil of the image. Fig. 6 is a block diagram when the configuration of the HMD shown in Fig. 1 is specialized and modified to the second embodiment. In an HMD 600 (information processing device) according to the second embodiment, an image of the real space is imaged as a real image by an imaging unit 602 through an optical system 601. The imaged real image is image-processed into a digital image by an image processing unit 603.
[0033] The microphone control unit 604 (voice input means) passes voice information of the voice inputted by the microphone unit 605 to the voice recognition unit 606. The voice recognition unit 606 (voice recognition means) extracts the voice content from the received voice information and converts the extracted voice content into text. The voice recognition unit 606 passes the converted text to the CG generation unit 607. The CG generation unit 607 (first input means) (generation means) generates a CG image of a virtual sticky note with the received text inputted. Meanwhile, the video information outputted from the gaze camera unit 608 is passed to the gaze calculation unit 609. The gaze calculation unit 609 (gaze calculation means) receives the video information from the gaze camera unit 608 and calculates the direction of the user's gaze from the pupil. The gaze calculation unit 609 passes the calculated gaze direction information to the object recognition unit 610.
[0034] The object recognition unit 610 (object recognition means) acquires the spatial coordinates of a real object in the direction of the line of sight by image recognition in the real image input from the image processing unit 603. Alternatively, the object recognition unit 610 (object recognition means) acquires the spatial coordinates of a virtual object in the direction of the line of sight by image recognition in the three-dimensional CG space input from the CG generation unit 607, that is, in the virtual space. A virtual sticky note associating unit 611 (pasting means) associates a CG image of a virtual sticky note with the acquired spatial coordinates of the real object or virtual object, thereby pasting the virtual sticky note.
[0035] The CG image of the virtual sticky note pasted in this manner, together with the three-dimensional CG space generated by the CG generation unit 607, is synthesized by the image synthesis unit 612 (synthesizing means) with the real image converted into digital image by the image processing unit 603, and displayed on the display unit 613. As a result, the image displayed on the display unit 613 is seen by the user as a synthesized image. The camera control unit 614 is part of the HMD control unit 108 in FIG. 1. The virtual sticky note association unit 611 is part of the CG generation unit 123 in FIG. 1. The image processing unit 603 corresponds to the imaging signal processing unit 103 and the video signal processing unit 104 in FIG. 1.
[0036] In the second embodiment, when a user utters a voice 301 with the voice content of "aiue..." in a review, seminar, or briefing session held in XR, as shown in the above-mentioned Fig. 3, a CG image of a virtual sticky note 302 with the voice content of "aiue..." written on it is generated. This allows for hassle-free real-time character input unlike a case where a virtual keyboard is operated by hand gestures, and therefore makes it possible to provide an excellent user experience when inputting to a virtual sticky note in XR.
[0037] Fig. 7 is an image diagram for explaining pasting of a virtual sticky note by line of sight calculation. In the second embodiment, for example, as shown in Fig. 7, a CG image of a virtual sticky note 302 is associated with spatial coordinates acquired by image recognition of a real object or a first button 402 of a virtual object in the direction of a user's line of sight 701. This allows the CG image of the virtual sticky note 302 to be pasted on the real object or the first button 402 of a virtual object without the user having to perform a hand gesture, making it possible to provide an excellent user experience with pasting of a CG image of a virtual sticky note in XR.
[0038] FIG. 8 is a flowchart showing the process flow from generating a virtual sticky note to pasting and merging. When the process shown in the flowchart in FIG. 8 starts, first, in step S801, the microphone control unit 604 passes voice information of the voice input to the microphone unit 605 to the voice recognition unit 606. In step S802, the voice recognition unit 606 performs voice recognition to extract the voice content from the received voice information. Furthermore, the voice recognition unit 606 converts the extracted voice content into text and passes the converted text to the CG generation unit 607. In step S803, the CG generation unit 607 inputs the received text into the virtual sticky note. In step S804, the CG generation unit 607 generates a CG image of the virtual sticky note into which the text was input in step S803.
[0039] In step S805, the gaze calculation unit 609 receives image information from the gaze camera unit 608, calculates the direction of the user's gaze from the pupil, and passes the calculated gaze direction information to the object recognition unit 610. In step S806, the object recognition unit 610 acquires the spatial coordinates of a real object in the gaze direction by image recognition in the real image input from the image processing unit 603. Alternatively, the object recognition unit 610 acquires the spatial coordinates of a virtual object in the gaze direction by image recognition in the three-dimensional CG space input from the CG generation unit 607. In step S807, the virtual sticky note association unit 611 attaches the virtual sticky note by associating the CG image of the virtual sticky note created in step S804 with the spatial coordinates of the real object or virtual object acquired in step S806. In step S808, the image synthesis unit 612 synthesizes the CG image of the virtual sticky note and the three-dimensional CG space generated by the CG generation unit 607 with the real image output from the image processing unit 603, and displays the result on the display unit 613. After that, the process shown in the flowchart in FIG. 8 ends.
[0040] As described above, in the second embodiment, when a user speaks during a review, seminar, or information session conducted in XR, a CG image of a virtual sticky note with the contents of the voice input is generated, and the generated CG image of the virtual sticky note is composited with a real image together with a three-dimensional CG space. In this way, the HMD 600 according to the second embodiment can realize information organization with comfortable and high-speed character input during a review, seminar, or information session conducted in XR.
[0041] <Third embodiment> Hereinafter, the third embodiment will be described with reference to FIG. 1 and FIG. 9 to FIG. 12. In the third embodiment, a virtual sticky note (virtual paper) on which voice content is written by voice input is generated in an XR space. A person is also recognized from the input voice. Furthermore, the color of the virtual sticky note is changed to a color corresponding to the recognized person, or a name corresponding to the recognized person is written on the virtual sticky note. Furthermore, a real object or virtual object indicated by the name of the object input by voice is also recognized. Furthermore, the generated virtual sticky note is attached to the recognized real object or virtual object. FIG. 9 is a block diagram when the configuration of the HMD shown in FIG. 1 is specialized and modified to the third embodiment.
[0042] In an HMD 900 (information processing device) according to the third embodiment, an image of real space is imaged as a real image by an imaging unit 902 through an optical system 901. The imaged real image is image-processed as a digital image by an image processing unit 903. A microphone control unit 904 (audio input means) passes audio information of a voice inputted by a microphone unit 905 to a voice recognition unit 906, a person recognition unit 907, and an object recognition unit 908. The voice recognition unit 906 (voice recognition means) extracts the voice content from the received audio information and converts the extracted voice content into text. The voice recognition unit 906 passes the converted text to a CG generation unit 909. The CG generation unit 909 (first input means) (generation means) generates a CG image of a virtual sticky note with the received text inputted therein.
[0043] The person recognition unit 907 (person recognition means) recognizes a person from the received voice information, and passes information on the recognized person to the color and name setting unit 910. The color and name setting unit 910 passes color or name information corresponding to the person in the received information to the CG generation unit 909. The CG generation unit 909 (color change means) (second input means) changes the color of the virtual sticky note that generated the CG image to the color of the received information, or inputs character data of the name of the received information into the virtual sticky note that generated the CG image. The object recognition unit 908 extracts the name of the object from the received voice information. Note that the extraction of the name of the object may be performed by the voice recognition unit 906 that received the voice information. However, in this case, the voice recognition unit 906 passes information on the extracted name of the object to the object recognition unit 908.
[0044] The object recognition unit 908 (object recognition means) acquires the spatial coordinates of a real object indicated by the object name through image recognition in the real image input from the image processing unit 903. Alternatively, the object recognition unit 908 (object recognition means) acquires the spatial coordinates of a virtual object indicated by the object name through image recognition in the three-dimensional CG space input from the CG generation unit 909, that is, the virtual space. A virtual sticky note association unit 911 (pasting means) associates a CG image of a virtual sticky note with the acquired spatial coordinates of the real object or virtual object, thereby pasting the virtual sticky note.
[0045] The CG image of the virtual sticky note pasted in this manner, together with the three-dimensional CG space generated by the CG generation unit 909, is synthesized by the image synthesis unit 912 (synthesizing means) with the real image converted into digital image by the image processing unit 903, and displayed on the display unit 913. As a result, the image displayed on the display unit 913 is seen by the user as a synthesized image. The camera control unit 914 is a part of the HMD control unit 108 in FIG. 1. The color / name setting unit 910 and the virtual sticky note association unit 911 are parts of the CG generation unit 123 in FIG. 1. The image processing unit 903 corresponds to the imaging signal processing unit 103 and the video signal processing unit 104 in FIG. 1.
[0046] FIG. 10 is an image diagram for explaining generation of a virtual sticky note with voice content written by voice input and change of color of the virtual sticky note by person recognition based on the voice. In the third embodiment, for example, as shown in FIG. 10, in a review, a seminar, or an explanatory meeting conducted in XR, when person A utters a voice 301 with the voice content of "Aiue...", a CG image of a virtual sticky note 302 with the voice content of "Aiue..." written is generated. At that time, when person A is recognized from the voice 301, the color of the CG image of the virtual sticky note 302 is changed to a color corresponding to person A, that is, the color of person A. Similarly, when person B utters a voice 301 with the voice content of "Aiue...", a CG image of a virtual sticky note 302 with the voice content of "Aiue..." is generated. At that time, when person B is recognized from the voice 301, the color of the CG image of the virtual sticky note 302 is changed to a color corresponding to person B, that is, the color of person B.
[0047] FIG. 11 is an image diagram for explaining the generation of a virtual sticky note with voice content written by voice input, and the writing of a name on the virtual sticky note by person recognition based on the voice. In the third embodiment, for example, as shown in FIG. 11, in a review, a seminar, or an explanatory meeting conducted in XR, when a person A utters a voice 301 with the voice content of "Aiue...", a CG image of a virtual sticky note 302 with the voice content of "Aiue..." written is generated. At that time, when the person A is recognized from the voice 301, a name according to the person A, that is, the name of the person A, is written in the CG image of the virtual sticky note 302. Similarly, when a person B utters a voice 301 with the voice content of "Aiue...", a CG image of a virtual sticky note 302 with the voice content of "Aiue..." is generated. At that time, when a person B is recognized from the voice 301, a name according to the person B, that is, the name of the person B, is written in the CG image of the virtual sticky note 302.
[0048] As a result, unlike the case of operating a virtual keyboard with hand gestures, text can be input in real time without hassle, and an excellent user experience can be provided for inputting to a virtual sticky note in XR. Furthermore, the color of the CG image of the virtual sticky note is changed according to the person who uttered the voice, and the name according to the person who uttered the voice is written in the CG image of the virtual sticky note, so that an excellent user experience can be provided for the value-added function of voice input in XR. In both cases of FIG. 10 and FIG. 11, the subsequent pasting of the virtual sticky note 302 is the same as in the first embodiment. However, the pasting of the virtual sticky note 302 may be performed in the same manner as in the second embodiment.
[0049] FIG. 12 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging. When the processing shown in the flowchart in FIG. 12 is started, first, in step S1201, the microphone control unit 904 passes voice information of a voice input to the microphone unit 905 to the voice recognition unit 906, the person recognition unit 907, and the object recognition unit 908. In step S1202, the voice recognition unit 906 performs voice recognition to extract voice content from the received voice information. Furthermore, the voice recognition unit 906 converts the extracted voice content into text, and passes the converted text to the CG generation unit 909. In step S1203, the CG generation unit 909 inputs the received text into the virtual sticky note. In step S1204, the CG generation unit 909 generates a CG image of the virtual sticky note into which the text was input in step S1203.
[0050] In step S1205, the person recognition unit 907 recognizes a person from the received voice information, and passes information on the recognized person to the color and name setting unit 910. In step S1206, the color and name setting unit 910 determines whether to change the color of the virtual sticky note. Note that this determination is made, for example, based on a pre-setting by the user using the HMD operation unit 110. If the color and name setting unit 910 determines to change the color of the virtual sticky note, it passes color information corresponding to the person in the information received from the person recognition unit 907 to the CG generation unit 909. Thereafter, the process proceeds to step S1207. On the other hand, if the color and name setting unit 910 determines not to change the color of the virtual sticky note, the process proceeds to step S1210. In step S1207, the CG generation unit 909 changes the color of the CG image of the virtual sticky note to the color of the information received from the color and name setting unit 910.
[0051] In step S1208, the color / name setting unit 910 determines whether or not to input a name to the virtual sticky note. Note that this determination is made based on, for example, a pre-setting by the user using the HMD operation unit 110. If the color / name setting unit 910 determines that a name is to be input to the virtual sticky note, it passes name information corresponding to the person of the information received from the person recognition unit 907 to the CG generation unit 909. Thereafter, the process proceeds to step S1209. On the other hand, if the color / name setting unit 910 determines that a name is not to be input to the virtual sticky note, the process proceeds to step S1210. In step S1209, the CG generation unit 909 inputs character data of the name corresponding to the person of the information received from the color / name setting unit 910 to the CG image of the virtual sticky note. In step S1210, the process from pasting to compositing of the virtual sticky note is performed. Note that this process is similar to each process of steps S505 to S509 in the first embodiment described above, and therefore a detailed description of step S1210 is omitted. After that, the process shown in the flowchart of FIG. 12 ends.
[0052] As described above, in the third embodiment, when a person speaks during a review, seminar, or information session held in XR, a CG image of a virtual sticky note with the contents of that voice input is generated, and the generated CG image of the virtual sticky note is composited with a real image together with a three-dimensional CG space. In this way, the HMD 900 according to the third embodiment can realize information organization with comfortable and high-speed character input during a review, seminar, or information session held in XR.
[0053] <Fourth embodiment> The fourth embodiment will be described below with reference to Figs. 1, 3, and 13 to 15. In the fourth embodiment, a virtual sticky note (virtual paper) on which voice content is written by voice input is generated in an XR space. Furthermore, emotions are recognized from the input voice. Furthermore, the color of the virtual sticky note is changed to a color corresponding to the recognized emotion. Furthermore, a real object or virtual object indicated by the name of the object input by voice is recognized. Furthermore, the generated virtual sticky note is affixed to the recognized real object or virtual object. Fig. 13 is a block diagram when the configuration of the HMD shown in Fig. 1 is specialized and modified to the fourth embodiment.
[0054] In the HMD 1300 (information processing device) according to the fourth embodiment, an image of a real space is imaged as a real image by an imaging unit 1302 through an optical system 1301. The imaged real image is image-processed as a digital image by an image processing unit 1303. A microphone control unit 1304 (audio input means) passes voice information of a voice inputted by a microphone unit 1305 to a voice recognition unit 1306, an emotion recognition unit 1307, and an object recognition unit 1308. The voice recognition unit 1306 (voice recognition means) extracts voice contents from the received voice information and converts the extracted voice contents into text. The voice recognition unit 1306 passes the converted text to a CG generation unit 1309. The CG generation unit 1309 (first input means) (generation means) generates a CG image of a virtual sticky note with the received text inputted therein.
[0055] The emotion recognition unit 1307 (emotion recognition means) recognizes emotions from the received voice information, and passes information on the recognized emotions to the color setting unit 1310. The color setting unit 1310 passes color information corresponding to the emotion of the received information to the CG generation unit 1309. The CG generation unit 1309 (color change means) changes the color of the virtual sticky note that generated the CG image to the color of the received information. The object recognition unit 1308 extracts the name of the object from the received voice information. Note that the extraction of the object name may be performed by the voice recognition unit 1306 that received the voice information. In this case, however, the voice recognition unit 1306 passes information on the extracted object name to the object recognition unit 1308.
[0056] The object recognition unit 1308 (object recognition means) acquires the spatial coordinates of a real object indicated by the object name through image recognition in the real image input from the image processing unit 1303. Alternatively, the object recognition unit 1308 (object recognition means) acquires the spatial coordinates of a virtual object indicated by the object name through image recognition in the three-dimensional CG space input from the CG generation unit 1309, that is, the virtual space. A virtual sticky note associating unit 1311 (pasting means) associates a CG image of a virtual sticky note with the acquired spatial coordinates of the real object or virtual object, thereby pasting the virtual sticky note.
[0057] The CG image of the virtual sticky note pasted in this manner, together with the three-dimensional CG space generated by the CG generation unit 1309, is synthesized by the image synthesis unit 1312 with the real image converted into digital image by the image processing unit 1303, and displayed on the display unit 1313. The image displayed on the display unit 1313 by the image synthesis unit 1312 (synthesizing means) is seen by the user as a synthesized image. The camera control unit 1314 is a part of the HMD control unit 108 in FIG. 1. The color setting unit 1310 and the virtual sticky note association unit 1311 are a part of the CG generation unit 123 in FIG. 1. The image processing unit 1303 corresponds to the imaging signal processing unit 103 and the video signal processing unit 104 in FIG. 1.
[0058] In the fourth embodiment, in a review, seminar, or briefing session held in XR, for example, as shown in FIG. 3 above, when a person utters a voice 301 with the voice content of "aiue...", a CG image of a virtual sticky note 302 with the voice content of "aiue..." written on it is generated. This allows for hassle-free real-time character input unlike a case where a virtual keyboard is operated by hand gestures, and therefore provides an excellent user experience for inputting to a virtual sticky note in XR. Furthermore, in the fourth embodiment, a person's emotions are recognized from the voice 301.
[0059] FIG. 14 is an image diagram for explaining the change in color of the virtual sticky note by emotion recognition based on voice. In the fourth embodiment, the emotion of a person is recognized from the voice 301. Furthermore, for example, as shown in FIG. 14, when the recognized emotion is a first emotion, the color of the CG image of the virtual sticky note 302 is changed to a first color corresponding to the first emotion. When the recognized emotion is a second emotion, the color of the CG image of the virtual sticky note 302 is changed to a second color corresponding to the second emotion. When the recognized emotion is a third emotion, the color of the CG image of the virtual sticky note 302 is changed to a third color corresponding to the third emotion. When the recognized emotion is a fourth emotion, the color of the CG image of the virtual sticky note 302 is changed to a fourth color corresponding to the fourth emotion. When the recognized emotion is a fifth emotion, the color of the CG image of the virtual sticky note 302 is changed to a fifth color corresponding to the fifth emotion. This allows the color of the CG image of the virtual sticky note to change depending on the emotion of the person who uttered the voice, making it possible to provide an excellent user experience with the added-value function of voice input in XR. Note that the subsequent attachment of the virtual sticky note 302 is similar to that of the first embodiment. However, the attachment of the virtual sticky note 302 may also be performed in the same manner as in the second embodiment.
[0060] FIG. 15 is a flowchart showing the flow of processing from generation of a virtual sticky note to pasting and merging. When the processing shown in the flowchart of FIG. 15 is started, first, in step S1501, the microphone control unit 1304 passes voice information of the voice input to the microphone unit 1305 to the voice recognition unit 1306, the emotion recognition unit 1307, and the object recognition unit 1308. In step S1502, the voice recognition unit 1306 performs voice recognition to extract voice content from the received voice information. Furthermore, the voice recognition unit 1306 converts the extracted voice content into text, and passes the converted text to the CG generation unit 1309. In step S1503, the CG generation unit 1309 inputs the received text into the virtual sticky note. In step S1504, the CG generation unit 1309 generates a CG image of the virtual sticky note into which the text was input in step S1503.
[0061] In step S1505, the emotion recognition unit 1307 recognizes emotion from the received voice information, and passes information on the recognized emotion to the color setting unit 1310. In step S1506, the color setting unit 1310 determines whether to change the color of the virtual sticky note. Note that this determination is made, for example, based on a setting in advance by the user using the HMD operation unit 110. If the color setting unit 1310 determines to change the color of the virtual sticky note, it passes color information corresponding to the emotion of the information received from the emotion recognition unit 1307 to the CG generation unit 1309. Then, the process proceeds to step S1507.
[0062] On the other hand, if the color setting unit 1310 determines not to change the color of the virtual sticky note, the process proceeds to step S1508. In step S1507, the CG generation unit 1309 changes the color of the CG image of the virtual sticky note to the color of the information received from the color setting unit 1310. In step S1508, the process from pasting the virtual sticky note to compositing is performed. Note that this process is similar to the processes in steps S505 to S509 in the first embodiment described above, and therefore a detailed description of step S1508 will be omitted. Thereafter, the process shown in the flowchart in FIG. 15 ends.
[0063] As described above, in the fourth embodiment, when a person speaks during a review, seminar, or information session held in XR, a CG image of a virtual sticky note with the content of the voice input is generated, and the generated CG image of the virtual sticky note is composited with a real image together with a three-dimensional CG space. In this way, the HMD1300 according to the fourth embodiment can realize information organization with comfortable and high-speed character input during a review, seminar, or information session held in XR.
[0064] <Fifth embodiment> Hereinafter, the fifth embodiment will be described with reference to Figs. 16 to 20. In the fifth embodiment, a virtual sticky note (virtual paper) on which voice content is written by voice input is generated in the XR space. Furthermore, a real object or a virtual object indicated by the name of the object input by voice is recognized. Furthermore, the generated virtual sticky note is attached to the recognized real object or virtual object. Furthermore, in the fifth embodiment, recording of the XR space is started at the timing of voice input. The recording file in which recording is started is recorded in association with the generated virtual sticky note. Furthermore, at the timing of another voice input, another virtual sticky note is generated, and another recording file is recorded, which is switched from the recording file being recorded. The switched recording file is recorded in association with another virtual sticky note. Furthermore, when a virtual sticky note is selected in the XR space, the recording file associated with the selected virtual sticky note is played.
[0065] Fig. 16 is a block diagram of the HMD configuration shown in Fig. 1 specialized and modified to the fifth embodiment. In an HMD 1600 (information processing device) according to the fifth embodiment, an image of real space is formed as a real image by an imaging unit 1602 through an optical system 1601. The formed real image is image-processed as a digital image by an image processing unit 1603. A microphone control unit 1604 (audio input means) passes audio information of a voice inputted by a microphone unit 1605 to a voice recognition unit 1606, an object recognition unit 1607, and a recording control unit 1608. The voice recognition unit 1606 (voice recognition means) extracts the voice content from the received voice information and converts the extracted voice content into text.
[0066] The voice recognition unit 1606 passes the converted text to the CG generation unit 1609. The CG generation unit 1609 (first input means) (generation means) generates a CG image of a virtual sticky note with the received text input thereto. The CG generation unit 1609 passes information about the virtual sticky note from which the CG image has been generated to the recording control unit 1608. The object recognition unit 1607 extracts the name of the object from the received voice information. Note that the extraction of the object name may be performed by the voice recognition unit 1606 that received the voice information. In this case, however, the voice recognition unit 1606 passes information about the extracted object name to the object recognition unit 1607.
[0067] The object recognition unit 1607 (object recognition means) acquires spatial coordinates of a real object indicated by the name of the object through image recognition in the real image input from the image processing unit 1603. Alternatively, the object recognition unit 1607 (object recognition means) acquires spatial coordinates of a virtual object indicated by the name of the object through image recognition in the three-dimensional CG space input from the CG generation unit 1609, that is, the virtual space. The virtual sticky note association unit 1610 (pasting means) associates the CG image of the virtual sticky note with the acquired spatial coordinates of the real object or virtual object, thereby pasting the CG image. The CG image of the virtual sticky note pasted in this manner is composited by the image synthesis unit 1611 with the real image converted into a digital image by the image processing unit 1603, together with the three-dimensional CG space generated by the CG generation unit 1609, and is displayed on the display unit 1612. The image displayed on the display unit 1612 by the image synthesis unit 1611 (synthesis means) enters the user's eyes as a composite image.
[0068] The recording control unit 1608 uses the audio information received from the microphone control unit 1604 for timing control of the start of recording, switching of recording files, and end of recording. The recording control unit 1608 records the recording file in the recording unit 1613 in association with the virtual sticky note of the information received from the CG generation unit 1609. Such association of the virtual sticky note with the recording file is performed, for example, when recording starts and when recording files are switched. In addition, the digital image output from the image processing unit 1603 is input to a hand gesture selection unit 1614 that detects the user's hands. This allows the user to select spatial coordinates by hand gestures in the composite image seen through the display unit 1612.
[0069] A sticky note selection determination unit 1615 determines whether the selected spatial coordinates and the virtual sticky note overlap. If the selected spatial coordinates and the virtual sticky note overlap, a recording file associated with the virtual sticky note is played by a recording file playback unit 1616. The played recording file is displayed on a display unit 1612 and is seen by a user. The camera control unit 1617 is a part of the HMD control unit 108 in FIG. 1. The virtual sticky note association unit 1610 is a part of the CG generation unit 123 in FIG. 1. The image processing unit 1303 corresponds to the imaging signal processing unit 103 and the video signal processing unit 104 in FIG. 1. The recording control unit 1608, the hand gesture selection unit 1614, the sticky note selection determination unit 1615, and the recording file playback unit 1616 are added to the configuration of the HMD in FIG. 1 in order to realize the fifth embodiment.
[0070] FIG. 17 is an image diagram for explaining the generation of a virtual sticky note in which voice content is written by voice input, and the association, recording, and switching of a recorded file. In the fifth embodiment, for example, as shown in FIG. 17, when a person utters a voice 1701 with the voice content of "start work" in an XR space where a work support briefing is being held, a CG image of a virtual sticky note 1702 in which the voice content of "start work" is written is generated. Furthermore, recording of the work support briefing being held in the XR space is started. The recorded file recorded by this is recorded in association with the virtual sticky note 1702. After that, when a person utters a voice 1703 with the voice content of "press the button here", a CG image of a virtual sticky note 1704 in which the voice content of "press the button here" is written is generated. Furthermore, the recorded file being recorded is switched to another recorded file. The switched other recorded file is recorded in association with the virtual sticky note 1704. Note that the subsequent attachment of the virtual sticky note 1702 and the virtual sticky note 1704 is the same as in the first embodiment. However, the virtual sticky note 1702 and the virtual sticky note 1704 may be pasted in the same manner as in the second embodiment.
[0071] Fig. 18 is an image diagram for explaining playback of a recorded file. In the fifth embodiment, for example, as described above in Fig. 17, virtual sticky notes 1702 and virtual sticky notes 1704 are attached in the XR space. In the fifth embodiment, further, for example, as shown in Fig. 18, when the user clicks and selects the virtual sticky note 1704 with a hand gesture or the like, the recorded file associated with the virtual sticky note 1704 is played. Note that, unlike Fig. 18, when the user clicks and selects the virtual sticky note 1702 with a hand gesture or the like, the recorded file associated with the virtual sticky note 1702 is played.
[0072] This allows for hassle-free real-time character input, unlike when operating a virtual keyboard with hand gestures, and therefore provides an excellent user experience for inputting to a virtual sticky note in XR. Furthermore, by associating a virtual sticky note with a recording file each time voice input is made, it is possible to play back the recording file associated with the selected virtual sticky note, thereby providing an excellent user experience for the value-added function of voice input in XR.
[0073] FIG. 19 is a flowchart showing the process flow from the generation of a virtual sticky note to pasting and synthesis, and from the start to the end of recording. When the process shown in the flowchart in FIG. 19 starts, first, in step S1901, the microphone control unit 1604 passes voice information of the voice input to the microphone unit 1605 to the voice recognition unit 1606 and the recording control unit 1608. In step S1902, the voice recognition unit 1606 performs voice recognition to extract the voice content from the received voice information. Furthermore, the voice recognition unit 1606 converts the extracted voice content into text, and passes the converted text to the CG generation unit 1609. In step S1903, the CG generation unit 1609 inputs the received text into the virtual sticky note. In step S1904, the CG generation unit 1609 generates a CG image of the virtual sticky note into which the text was input in step S1903.
[0074] In step S1905, the recording control unit 1608 (first determination means) determines whether recording has started. If the recording control unit 1608 determines that recording has not started, the process proceeds to step S1906. On the other hand, if the recording control unit 1608 determines that recording has started, the process proceeds to step S1907. In step S1906, the recording control unit 1608 (first recording means) starts recording. Furthermore, the recording control unit 1608 continues to record the recording file for which recording has started in the recording unit 1613, in association with the virtual sticky note of the information received from the CG generation unit 1609, that is, the virtual sticky note for which the CG image was generated in step S1904.
[0075] In step S1907, the recording control unit 1608 (second recording means) switches the recording file being recorded to another recording file. Furthermore, the recording control unit 1608 continues to record the switched-to other recording file in the recording unit 1613 in association with the virtual sticky note of the information received from the CG generation unit 1609, that is, the virtual sticky note whose CG image was generated in step S1904. In step S1908, the process from attaching the virtual sticky note to compositing is performed. Note that this process is similar to the processes in steps S505 to S509 in the first embodiment described above, and therefore a detailed description of step S1908 will be omitted.
[0076] In step S1909, microphone control unit 1604 determines whether or not there is voice input to microphone unit 1605. If microphone control unit 1604 determines that there is voice input to microphone unit 1605, it passes voice information of the voice input to microphone unit 1605 to voice recognition unit 1606 and recording control unit 1608. Then, the process proceeds to step S1910. On the other hand, if microphone control unit 1604 determines that there is no voice input to microphone unit 1605, the process returns to step S1909.
[0077] In step S1910, the voice recognition unit 1606 determines whether the voice input to the microphone unit 1605 instructs the end of recording. This determination is made based on the voice information received from the microphone control unit 1604. If the voice recognition unit 1606 determines that the voice input to the microphone unit 1605 instructs the end of recording, the process proceeds to step S1911. On the other hand, if the voice recognition unit 1606 determines that the voice input to the microphone unit 1605 does not instruct the end of recording, the process returns to step S1902. In step S1911, the recording control unit 1608 ends recording. Thereafter, the process shown in the flowchart of FIG. 19 ends.
[0078] FIG. 20 is a flowchart showing the flow of processing for playing a recorded file in the fifth embodiment. When the processing shown in the flowchart of FIG. 20 starts, first, in step S2001, the sticky note selection determination unit 1615 (second determination means) determines whether a virtual sticky note has been clicked. When the spatial coordinates selected by a hand gesture and the virtual sticky note overlap in the XR space, the sticky note selection determination unit 1615 determines that the virtual sticky note has been clicked. Then, the processing proceeds to step S2002. On the other hand, when the spatial coordinates selected by a hand gesture and the virtual sticky note do not overlap in the XR space, the sticky note selection determination unit 1615 determines that the virtual sticky note has not been clicked. Then, the processing returns to step S2001. In step S2002, the recorded file playback unit 1616 (playback means) plays the recorded file associated with the virtual sticky note determined to have been clicked in step S2002. Then, the processing shown in the flowchart of FIG. 20 ends.
[0079] As described above, in the fifth embodiment, when a person speaks during a work support briefing held in XR, a CG image of a virtual sticky note with the contents of the voice input is generated, and the generated CG image of the virtual sticky note is synthesized with a real image together with a three-dimensional CG space. In this way, the HMD 1600 according to the fifth embodiment can realize information organization with comfortable and high-speed character input during a work support briefing held in XR. This also applies to reviews, training sessions, and other briefings held in XR.
[0080] <Other> Although the preferred embodiments of the present invention have been described above, the present invention is not limited to the above-mentioned embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. For example, in the fourth embodiment, in addition to changing the color of the CG image of the virtual sticky note according to the emotion of the person who uttered the voice, a name according to the person who uttered the voice may be written in the CG image of the virtual sticky note. Also, in each embodiment, the virtual sticky note is attached to a real object or a virtual object, but it may also be attached to a position where there are no real objects or virtual objects, that is, in space.
[0081] In each embodiment, the virtual sticky note may be attached by the user pinching the virtual sticky note with hand tracking. In each embodiment, the virtual sticky note may be synthesized at a predetermined position in the real image when the CG image is generated. In each embodiment, the virtual sticky note may be another virtual object onto which the audio content is written (for example, a virtual object such as a memo paper, a piece of paper, a label, or a tag). In each embodiment, the HMD control information may be shared with other HMDs.
[0082] In addition, although a video see-through type HMD is used in each embodiment, an optical see-through type HMD may be used. In this case, the image synthesis unit of the HMD synthesizes a CG image of the virtual sticky note with the real space, not with the real image, so the display unit of the HMD does not display the real image.
[0083] The present invention can also be realized by a process in which a program for realizing one or more functions of each of the above-mentioned embodiments is supplied to a system or device via a network or a storage medium, and one or more processors of a computer in the system or device read and execute the program. The present invention can also be realized by a circuit (e.g., ASIC) for realizing one or more functions.
[0084] The disclosure of each embodiment includes the following configurations, methods, and programs. (Configuration 1) A voice input means for inputting voice; a voice recognition means for recognizing a voice input to the voice input means; a first input means for inputting character data representing the contents of the voice recognized by the voice recognition means into a virtual paper; a generating means for generating a CG image of the virtual paper to which character data has been input by the first input means; and a synthesis means for synthesizing the CG image of the virtual paper generated by the generation means into an XR space in which real space and virtual space are integrated. (Configuration 2) A person recognition means for recognizing a person from a voice input to the voice input means; The information processing device according to configuration 1, further comprising: a color change means for changing a color of the virtual paper on which the CG image is generated by the generation means to a color corresponding to the person recognized by the person recognition means. (Configuration 3) The information processing device according to configuration 2, further comprising: a second input means for inputting character data of a name corresponding to a person recognized by the person recognition means into the virtual paper on which a CG image has been generated by the generation means. (Configuration 4) An emotion recognition means for recognizing an emotion from a voice input to the voice input means; The information processing device according to configuration 1, further comprising: a color change means for changing a color of the virtual paper on which the CG image is generated by the generation means to a color corresponding to the emotion recognized by the emotion recognition means. (Configuration 5) A person recognition means for recognizing a person from a voice input to the voice input means; and a second input means for inputting character data of a name corresponding to a person recognized by the person recognition means into the virtual paper on which a CG image has been generated by the generation means. (Configuration 6) An object recognition means for acquiring spatial coordinates of an object indicated by a name of the object recognized from the voice input to the voice input means by image recognition of a real object in the real space or a virtual object in the virtual space; 6. The information processing device according to any one of configurations 1 to 5, further comprising: a pasting means for pasting a CG image of a virtual paper generated by the generation means onto the spatial coordinates of the object acquired by the object recognition means. (Configuration 7) The information processing device according to configuration 6, wherein the object recognition means recognizes the name of an object from the voice input to the voice input means. (Configuration 8) The information processing device according to configuration 6, wherein the voice recognition means recognizes a name of an object from the voice input to the voice input means, and passes information relating to the recognized name of the object to the object recognition means. (Configuration 9) A line of sight calculation means for calculating a line of sight direction; an object recognition means for acquiring spatial coordinates of an object in the line of sight direction calculated by the line of sight calculation means by image recognition of a real object in the real space or a virtual object in the virtual space; 6. The information processing device according to any one of configurations 1 to 5, further comprising: a pasting means for pasting a CG image of a virtual paper generated by the generation means onto the spatial coordinates of the object acquired by the object recognition means. (Configuration 10) A first determination means for determining whether recording has already started when the CG image of the virtual paper is generated by the generation means; a first recording means for associating a recording file, which is to be recorded by starting recording when it is determined by the first determination means that recording has not started, with the virtual paper on which the CG image has been generated by the generation means; and a second recording means for recording a recording file to be switched from the recording file being recorded when the first determination means determines that recording has already started, in association with the virtual paper on which the CG image has been generated by the generation means. (Configuration 11) An object recognition means for acquiring spatial coordinates of an object indicated by a name of the object recognized from the voice input to the voice input means by image recognition of a real object in the real space or a virtual object in the virtual space; 11. The information processing apparatus according to configuration 10, further comprising: pasting means for pasting a CG image of a virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means. (Configuration 12) The information processing device according to configuration 11, wherein the object recognition means recognizes the name of an object from the voice input to the voice input means. (Configuration 13) The information processing device according to configuration 11, wherein the voice recognition means recognizes a name of an object from the voice input to the voice input means, and passes information regarding the recognized name of the object to the object recognition means. (Configuration 14) A line of sight calculation means for calculating a line of sight direction; an object recognition means for acquiring spatial coordinates of an object in the line of sight direction calculated by the line of sight calculation means by image recognition of a real object in the real space or a virtual object in the virtual space; 11. The information processing apparatus according to configuration 10, further comprising: pasting means for pasting a CG image of a virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means. (Configuration 15) A second determination means for determining whether or not a virtual paper is selected in the XR space; 15. The information processing apparatus according to any one of configurations 10 to 14, further comprising: a playback unit that plays back a recorded file associated with the virtual paper determined to have been selected by the second determination unit. (Configuration 16) The information processing device according to any one of configurations 1 to 15, wherein the virtual paper is a virtual sticky note. (Method 1) a voice input step for inputting voice; a voice recognition step for recognizing the voice inputted in the voice input step; an input step of inputting character data representing the content of the voice recognized by the voice recognition step into a virtual paper; a generating step of generating a CG image of the virtual paper to which the character data has been inputted by the input step; A control method for an information processing device, comprising: a synthesis step of synthesizing the CG image of the virtual paper generated by the generation step into an XR space in which real space and virtual space are integrated. (Program 1) A program for causing a computer to execute each means of the information processing device according to any one of configurations 1 to 16. [Explanation of symbols]
[0085] 200 HMD (information processing device) 204 Microphone control unit (voice input means) 206 Voice recognition unit (voice recognition means) 208 CG generation unit (first input means) (generation means) 210 Image synthesis unit (synthesis means)
Claims
1. A voice input means for inputting voice; a voice recognition means for recognizing a voice input to the voice input means; a first input means for inputting character data representing the contents of the voice recognized by the voice recognition means into a virtual paper; a generating means for generating a CG image of the virtual paper to which character data is input by the first input means; a synthesis means for synthesizing the CG image of the virtual paper generated by the generation means into an XR space in which real space and virtual space are integrated.
2. a person recognition means for recognizing a person from the voice input to the voice input means; 2. The information processing apparatus according to claim 1, further comprising a color change unit that changes a color of the virtual paper on which the CG image is generated by the generation unit to a color corresponding to the person recognized by the person recognition unit.
3. an emotion recognition means for recognizing emotions from the voice inputted to the voice input means; 2. The information processing apparatus according to claim 1, further comprising a color change unit that changes a color of the virtual paper on which the CG image is generated by the generation unit to a color corresponding to the emotion recognized by the emotion recognition unit.
4. a person recognition means for recognizing a person from the voice input to the voice input means; 2. The information processing apparatus according to claim 1, further comprising: a second input means for inputting character data of a name corresponding to a person recognized by said person recognition means into the virtual paper on which the CG image has been generated by said generation means.
5. an object recognition means for acquiring spatial coordinates of an object indicated by a name of the object recognized from the voice input to the voice input means by image recognition of a real object in the real space or a virtual object in the virtual space; 2. The information processing apparatus according to claim 1, further comprising: pasting means for pasting a CG image of the virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means.
6. 6. The information processing apparatus according to claim 5, wherein the object recognition means recognizes the name of an object from the voice inputted to the voice input means.
7. 6. The information processing apparatus according to claim 5, wherein the voice recognition means recognizes a name of an object from the voice inputted to the voice input means, and passes information relating to the recognized name of the object to the object recognition means.
8. A line of sight calculation means for calculating a line of sight direction; an object recognition means for acquiring spatial coordinates of an object in the line of sight direction calculated by the line of sight calculation means by image recognition of a real object in the real space or a virtual object in the virtual space; 2. The information processing apparatus according to claim 1, further comprising: pasting means for pasting a CG image of the virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means.
9. a first determination means for determining whether recording has already started when the CG image of the virtual paper is generated by the generation means; a first recording means for associating a recording file, which is to be recorded by starting recording when it is determined by the first determination means that recording has not started, with the virtual paper for which a CG image has been generated by the generation means; The information processing device according to claim 1, further comprising: a second recording means for recording a recording file to be switched from the recording file being recorded when the first determination means determines that recording has already started, in association with the virtual paper on which the CG image has been generated by the generation means.
10. an object recognition means for acquiring spatial coordinates of an object indicated by a name of the object recognized from the voice input to the voice input means by image recognition of a real object in the real space or a virtual object in the virtual space; 10. The information processing apparatus according to claim 9, further comprising: pasting means for pasting the CG image of the virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means.
11. 11. The information processing apparatus according to claim 10, wherein the object recognition means recognizes the name of an object from the voice inputted to the voice input means.
12. 11. The information processing apparatus according to claim 10, wherein the voice recognition means recognizes a name of an object from the voice inputted to the voice input means, and passes information relating to the recognized name of the object to the object recognition means.
13. A line of sight calculation means for calculating a line of sight direction; an object recognition means for acquiring spatial coordinates of an object in the line of sight direction calculated by the line of sight calculation means by image recognition of a real object in the real space or a virtual object in the virtual space; 10. The information processing apparatus according to claim 9, further comprising: pasting means for pasting the CG image of the virtual paper generated by said generating means onto the spatial coordinates of the object acquired by said object recognizing means.
14. A second determination means for determining whether or not a virtual paper sheet has been selected in the XR space; 10. The information processing apparatus according to claim 9, further comprising a playback unit that plays back a recorded file associated with the virtual paper determined to have been selected by the second determination unit.
15. 2. The information processing apparatus according to claim 1, wherein the virtual paper is a virtual sticky note.
16. a voice input step of inputting voice; a voice recognition step for recognizing the voice inputted in the voice input step; an input step of inputting character data representing the content of the voice recognized by the voice recognition step into a virtual paper; a generating step of generating a CG image of the virtual paper to which the character data has been inputted by the input step; A control method for an information processing device, comprising: a synthesis step of synthesizing the CG image of the virtual paper generated by the generation step into an XR space in which real space and virtual space are integrated.
17. 2. A program for causing a computer to execute each of the means of the information processing apparatus according to claim 1.
Citation Information
Patent Citations
Image processing apparatus and image processing method
JP2021005157A