Head-mounted display device, image display method, and image display program

The HMD dynamically adjusts display content based on approaching individuals, facilitating communication with relevant persons by prioritizing their images, thus improving user interaction and productivity.

JP2026112008APending Publication Date: 2026-07-06JVC KENWOOD CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
JVC KENWOOD CORP
Filing Date
2024-12-24
Publication Date
2026-07-06

AI Technical Summary

Technical Problem

Existing head-mounted display (HMD) technologies do not adapt the display content based on the presence of individuals approaching the user, making it difficult for users to communicate with relevant persons during tasks.

Method used

A head-mounted display device equipped with an image acquisition unit, person information acquisition unit, and display control unit that determines the relationship between approaching individuals and the displayed content, adjusting the display accordingly based on gestures, facial recognition, and voice analysis.

Benefits of technology

Enables dynamic adjustment of the display to prioritize communication with relevant individuals by displaying their images prominently while maintaining task-related information, enhancing user interaction and productivity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026112008000001_ABST
    Figure 2026112008000001_ABST
Patent Text Reader

Abstract

The present invention provides a head-mounted display device, an image display method, and an image display program that can change the displayed image depending on the person approaching the user. [Solution] The system includes a display unit 19 that displays real-world images and PC images presented during the user's work, an image acquisition unit 11 that acquires images of the user's surroundings, a person information acquisition unit 121 that acquires person information of people included in the surrounding images, and a display control unit 18 that determines whether or not there is a relationship between the person and the PC image displayed on the display unit 19 based on the person information, and controls the image displayed on the display unit 19 according to whether or not there is a relationship.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a head-mounted display device, an image display method, and an image display program.

Background Art

[0002] For the purpose of improving the convenience and security of work using an electronic terminal such as a personal computer, virtual desktop technology has been put into practical use. In a virtual desktop, by displaying an image of the real space and a PC image of the personal computer operated by the user on a head-mounted display (hereinafter abbreviated as "HMD") worn by the user, the user can perform work using the electronic terminal without choosing a work location. In addition, since the user can recognize an image of the real space during work, the user can grasp the surrounding situation even during work.

[0003] Patent Document 1 proposes an HMD for a game machine that changes display content according to the user's surrounding environment. When the user is wearing the HMD and operating the game machine, for example, when there is an incoming call on a smartphone, an image of this smartphone is displayed on the HMD. The user can perform operations on the smartphone while continuing the game operation.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, Patent Document 1 does not describe changing the HMD screen display in consideration of a person approaching the user. For example, when a user is wearing an HMD and performing a task, if a person related to this task, such as the person who sent the user an email or a supervisor instructing the user, speaks to the user, it would be desirable to switch the HMD display image from the PC image to the person's image. Patent Document 1 does not mention changing the HMD display content when a person related to the user's task approaches. As a result, there was a problem in that it was difficult for the user to communicate with people related to the user's task.

[0006] The present invention was made to solve these conventional problems, and its objective is to provide a head-mounted display device, an image display method, and an image display program that can change the displayed image according to a person approaching the user. [Means for solving the problem]

[0007] A head-mounted display device according to one or more embodiments for achieving the above objective is a head-mounted display device worn by a user, comprising: a display unit that displays a real-space image and a presentation image to be presented to the user; an image acquisition unit that acquires an image of the user's surroundings; a person information acquisition unit that acquires person information of a person included in the surrounding image; and a display control unit that determines whether or not there is a relationship between the person and the presentation image displayed on the display unit based on the person information, and controls the image to be displayed on the display unit according to whether or not there is a relationship.

[0008] Furthermore, an image display method according to one or more embodiments is an image display method for displaying an image on the display unit of a head-mounted display device worn by a user, wherein an image acquisition unit acquires an image of the user's surroundings, a person information acquisition unit acquires person information of a person included in the surrounding image, and a display control unit determines whether or not there is a relationship between the person and the presented image displayed on the display unit based on the person information, and changes the image displayed on the display unit according to whether or not there is a relationship.

[0009] Furthermore, an image display program according to one or more embodiments causes a computer to execute the image display method described above. [Effects of the Invention]

[0010] According to the present invention, it becomes possible to change the displayed image depending on the person approaching the user. [Brief explanation of the drawing]

[0011] [Figure 1] Figure 1 is a block diagram showing the configuration of a head-mounted display device according to an embodiment. [Figure 2] Figure 2 shows the background image displayed on the HMD. [Figure 3] Figure 3 is an explanatory diagram showing how a PC image is superimposed on the background image displayed in display area D1. [Figure 4] Figure 4 shows an example of what is displayed when a person related to the user's work approaches a user who is wearing an HMD and working. [Figure 5] Figure 5 shows a first example of the display when a user wearing an HMD and working is approached by a user whose work is not related to the user's work. [Figure 6] Figure 6 shows a second example of the display when a user wearing an HMD and working is approached by a user whose work is not related to the user's task. [Figure 7] Figure 7 shows a third example of the display when a user unrelated to the user's work approaches a user who is wearing an HMD and working. [Figure 8] Figure 8 is a flowchart showing the processing procedure for displaying an image on the display unit of the HMD according to the embodiment. [Figure 9] Figure 9 is a flowchart showing the process of displaying an image on the display unit using voice recognition. [Modes for carrying out the invention]

[0012] Hereinafter, a head-mounted display device 1 (hereinafter abbreviated as "HMD1") according to an embodiment of the present invention will be described with reference to the drawings. Figure 1 is a block diagram showing the configuration of HMD1 according to the embodiment. As shown in Figure 1, HMD1 includes an image acquisition unit 11, a identification unit 12, a database 13, a voice acquisition unit 14, a voice determination unit 15, an acceleration sensor 16, an action determination unit 17, a display control unit 18, and a display unit 19. HMD1 is also connected to a terminal device 2.

[0013] The specific unit 12, database 13, voice determination unit 15, operation determination unit 17, and display control unit 18 shown in Figure 1 can be implemented using a microcomputer equipped with a CPU (Central Processing Unit), memory, and input / output unit. Computer programs for each of the above information processing units (specific unit 12, database 13, voice determination unit 15, operation determination unit 17, and display control unit 18) are installed on the microcomputer and executed. As a result, the microcomputer functions as each information processing unit. While this example shows each information processing unit implemented using software, it is also possible to configure each information processing unit by preparing dedicated hardware. Furthermore, the information processing performed by each information processing unit may be configured using separate hardware.

[0014] The HMD1 is used, for example, in a virtual desktop. The HMD1 is worn on the user's head, displays an image output from the terminal device 2, and executes an operation using the terminal device 2. The terminal device 2 includes a personal computer, a game console, and the like. In the present embodiment, an example will be described in which the user operates a personal computer to execute work and displays an image related to the work (hereinafter referred to as a "PC image") on the display unit 19 of the HMD1.

[0015] The image acquisition unit 11 shown in FIG. 1 is, for example, a plurality of cameras, and acquires a peripheral image of the user when the user is wearing the HMD1. The peripheral image is an image around the user and includes an image of a person approaching the user. The image acquisition unit 11 captures an image of a person (hereinafter referred to as a "surrounding person") existing around the real space where the user exists.

[0016] The specifying unit 12 includes a person information acquisition unit 121 and a relevance determination unit 122.

[0017] The person information acquisition unit 121 analyzes the peripheral image acquired by the image acquisition unit 11 and extracts an image of a surrounding person included in this image. Specifically, using an image analysis algorithm, the outline and feature points of the person in the image are detected, and the person is separated from the background. The person information acquisition unit 121 acquires person information of the person included in the peripheral image. The person information here includes data related to persons such as the user's family, friends, colleagues, superiors, etc., and is further identification information such as a face image and a name. The person information acquisition unit 121 determines whether or not the surrounding person is performing a specific gesture based on the extracted image of the surrounding person. The specific gesture can be, for example, a gesture indicating an action when talking to the user, such as waving a hand facing the user's direction or beckoning. For the determination of the gesture, a motion pattern recognition technique is used.

[0018] As will be described later, the database 13 stores templates of various gestures including specific gestures. The person information acquisition unit 121 compares the gesture executed by the surrounding person with the gesture template stored in the database 13, and calculates the degree of coincidence between the gesture executed by the surrounding person and the specific gesture.

[0019] Based on, for example, the image acquired by the image acquisition unit 11, the person information acquisition unit 121 calculates the correlation coefficient between the gesture executed by the surrounding person and the template using well-known technologies such as AI (Artificial Intelligence) and image recognition technology, and calculates the degree of coincidence between the two based on the correlation coefficient within the range of 0 to 100%. When the above-mentioned degree of coincidence exceeds a predetermined threshold value (for example, 70%), the person information acquisition unit 121 determines that a specific gesture is being executed.

[0020] The person information acquisition unit 121 performs pattern matching for the entire surrounding person or for each predetermined area of the surrounding person to extract the feature points of the surrounding person. For example, the person information acquisition unit 121 may extract the feature points of the surrounding person using motion vectors. For example, based on the captured video captured by the camera, information indicating the motion of the surrounding person is generated.

[0021] The specifying unit 12, for example, uses the captured video in which the entire surrounding person is shown to detect the motion vector of the surrounding person, and generates the motion information of the surrounding person based on the detected motion vector. The specifying unit 12 may detect the motion vector, for example, using the current frame and the previous frame of the video of the surrounding person. The specifying unit 12 performs pattern matching from the detected motion vector to extract the feature points of the user. Further, the specifying unit 12 may analyze the motion of the face in time series to extract the feature points.

[0022] Note that the specifying unit 12 is not limited to generating the motion information of the surrounding person using the captured image captured by the camera as long as it can detect the natural motion of the user. The specifying unit 12 performs pattern User features can be extracted using machine learning methods other than matching.

[0023] The identification unit 12 extracts characteristic points of the surrounding person's movements and calculates the positional relationship of these points. Based on the positional relationship of the feature points, the person information acquisition unit 121 detects the movements of the surrounding person by identifying the movements of the surrounding person captured in the video from a pre-prepared set of person movement patterns as movement patterns with the same or similar positional relationship of the feature points. For example, after the identification unit 12 extracts characteristic points of the surrounding person's movements, it records the positional relationship of the feature points as coordinate positions. For example, the surrounding person 12 refers to a database in which pre-prepared gestures and their associated characteristic points are stored, and if there is data for a movement that corresponds to the same or similar positional relationship of the user's movement's characteristic points, it identifies the user's movement as a movement pattern such as "waving."

[0024] The person information acquisition unit 121 identifies a person in the vicinity when it detects the presence of a person performing a specific gesture. For example, it compares the facial images of multiple people stored in the database 13 with the image of the person acquired by the image acquisition unit 11 to determine whether or not this person is related to the user.

[0025] The person information acquisition unit 121 uses face recognition technology to extract facial feature points and identify individuals by matching them with face images in a database. Specifically, it detects face regions from images acquired by the image acquisition unit 11 using Haar-like features and HOG (Histogram of Oriented Gradients), and extracts feature points such as eyes, nose, and mouth from the detected face regions using landmark detectors such as Dlib or OpenCV. Next, it generates a facial feature vector based on the extracted feature points, and a deep learning model (e.g., FaceNet or DeepFace) can be used for this. The generated feature vector is compared with the feature vectors of face images stored in the database 13, and the degree of agreement is calculated by calculating cosine similarity and Euclidean distance. If the degree of agreement exceeds a predetermined threshold (e.g., 90%), it is determined that the surrounding person is related to the user. In this way, the person information acquisition unit 121 can determine whether or not a surrounding person performing a specific gesture is related to the user.

[0026] The person information acquisition unit 121 acquires gestures performed by a person and, if the gesture is a specific gesture, determines whether there is a relationship between the person and the displayed image shown on the display unit 19. The displayed image is, for example, a PC image of a PC operated by a user.

[0027] Specifically, the person information acquisition unit 121 analyzes the facial images of people in the surrounding area using facial recognition technology and compares them with data of people's facial images that have been registered in the database 13 in advance. This allows the system to determine whether the person in the surrounding area who performed the gesture is related to the user. As will be described later, the database 13 contains data about people such as the user's family, friends, colleagues, and superiors, and also registers identification information such as facial images and names. The person information acquisition unit 121 uses this information to perform the above comparison. For example, the database 13 stores a table that associates the relationship between a face and a person, and the degree of that relationship (numerical value), and the person information acquisition unit 121 may use this table to derive the relationship. For example, a high numerical value may be assigned to the user's family, and a medium numerical value may be assigned to a colleague or friend. In this way, the system can determine how related the people in the surrounding area are to the user.

[0028] The relationship determination unit 122 evaluates the relationship between the surrounding people identified by the person information acquisition unit 121 and the content of the PC image resulting from the work being performed by the user. The content of the PC image refers to information contained in the PC image, such as emails, chats, and documents.

[0029] The relevance determination unit 122 determines the relevance between the surrounding people and the content of the PC image, for example, from 0 to 100%. The correlation is quantified. The correlation determination unit 122 determines that a person in the vicinity is an important person related to the user if the correlation is above a predetermined threshold (for example, 50%). For example, if a user is performing a difficult task and their supervisor speaks to the user using a specific gesture, the unit evaluates this supervisor as a person in the vicinity with high correlation to the user. The correlation is set high because the supervisor is likely to want to check on the user's work status.

[0030] Furthermore, for example, if a user sends an email to person A through the operation of a computer, and the person information acquisition unit 121 determines that the person in the vicinity is person A, then person A is given a high degree of relevance to the task being performed by the user. The relevance determination unit 122 outputs a command to the display control unit 18 to change the image displayed on the display unit 19 according to the relevance.

[0031] Database 13 stores gesture template data used in the aforementioned person information acquisition unit 121, data for recognizing a person's face image, and voice templates indicating voice instruction information used in the voice determination unit 15, which will be described later. The data for face image recognition includes identification information such as the face images and names of the user's family, friends, colleagues, and superiors. Database 13 may be installed inside the HMD1 or on the cloud.

[0032] The voice acquisition unit 14 is, for example, a microphone, which acquires ambient sounds from the user wearing the HMD1. Multiple voice acquisition units 14 are preferably installed around the HMD1. The voice acquisition unit 14 acquires, for example, the voices of people in the user's vicinity who speak.

[0033] The voice determination unit 15 compares the voice acquired by the voice acquisition unit 14 with the voice templates stored in the database 13 to determine whether the voice is a voice addressing the user. A "voice addressing the user" is a voice spoken by people in the vicinity when addressing the user, and includes, for example, "Mr. / Ms. [User's Last Name]" or "Hello" as a greeting to the user. The database 13 stores voice templates such as "Mr. / Ms. [User's Last Name]" and "Hello". In other words, the voice determination unit 15 determines, based on the surrounding voice, whether the voice spoken by a person in the vicinity is a voice addressing the user.

[0034] Specifically, the voice determination unit 15 converts the voice acquired using speech recognition technology into text data and compares the converted text data with the text data of the voice template stored in the database 13. For the comparison, it calculates the similarity of the voices using dynamic time stretching (DTW) or deep learning models (e.g., RNN or Transformer), and determines that it is a call to action if the similarity exceeds a predetermined threshold (e.g., 80%). In this way, the voice determination unit 15 can determine whether the voice spoken by a person in the vicinity is a call to action directed at the user.

[0035] The acceleration sensor 16 detects acceleration generated in the HMD1. The acceleration sensor 16 detects acceleration in three dimensions or two dimensions. For example, the acceleration sensor 16 detects acceleration generated when a user wearing the HMD1 turns their head. The acceleration sensor 16 detects whether the user turned their head up and down or from side to side. The acceleration sensor 16 outputs the detection result to the motion determination unit 17.

[0036] The motion determination unit 17 determines whether or not the user has performed a head-turning motion based on the acceleration data detected by the acceleration sensor 16. In particular, the motion determination unit 17 determines whether or not the user has performed a vertical head-turning motion. When the voice determination unit 15 detects a voice call from a person in the vicinity, and the acceleration sensor 16 further detects a vertical head-turning motion by the user, the motion determination unit 17 outputs a command to the display control unit 18 to change the displayed image.

[0037] Specifically, the motion determination unit 17 analyzes the temporal changes in acceleration data and uses FFT (Fast Fourier Transform) and wavelet transform to detect characteristic patterns of head-shaking motion. It also uses machine learning algorithms (e.g., SVM or LSTM) to classify head-shaking motion from acceleration data. As a result, the motion determination unit 17 can understand the user's intent and perform appropriate display control.

[0038] Specifically, when the motion determination unit 17 detects that a person in the vicinity has called out to it, and further detects that the user has nodded, it outputs a command to the display control unit 18 to change the displayed image.

[0039] The display unit 19 is equipped with a display that shows a PC image (presented image), an image of surrounding people (real-world image), and a background image to be presented to the user. The display unit 19 includes main image areas D2 and D3 (see Figures 3 and 4 described later), which are set in part of the overall display area D1, and a sub-image area D4 which is narrower than the main image areas D2 and D3. The "background image" refers to an image that is displayed as a background over the entire display area D1.

[0040] The display control unit 18 controls the image displayed on the display unit 19. Based on the display image change command output from the relationship determination unit 122 and the operation determination unit 17, the display control unit 18 changes the image displayed on the display unit 19. Examples of image display on the display unit 19 will be described below with reference to Figures 2 to 7.

[0041] Figure 2 shows a background image displayed on the display unit 19 of the HMD1, depicting an office interior. The background image shown in Figure 2 is displayed across the entire display area D1 of the display unit 19. In other words, the background image in this case refers to the scenery that the user sees when they remove the HMD1. To improve user workability, the background image displayed in the display area D1 should be displayed with increased transparency. Specifically, if the transparency range is 0-100%, it should be set to around 50%. Note that 0% transparency refers to normal image display, and 100% transparency refers to the image not being displayed.

[0042] Figure 3 is an explanatory diagram showing how a PC image is superimposed on a background image displayed in display area D1, when a user operates a terminal device 2 such as a personal computer to perform a task. As shown in Figure 3, a main image area D2 is set in the center of display area D1, occupying a portion of the display area D1. The PC image is displayed in the main image area D2. Therefore, the user can perform tasks using the PC while viewing the PC image displayed in the main image area D2. In addition, since the background image has a certain degree of transparency, the user can concentrate on performing tasks on the PC image without being distracted by the background image.

[0043] This embodiment describes an example in which a PC image is superimposed on the surrounding scenery visible to the user in real space as a background image. In other words, assuming the user is not wearing a head-mounted display, the PC screen is superimposed on the scenery the user would see. In this case, for example, the surrounding scenery, which is an image of the surrounding scenery captured by a camera, may be displayed on the display unit 19, thereby providing the user with the surrounding scenery as a background image through the display unit 19. Alternatively, a transparent head-mounted display may be used, for example, to transmit ambient light (visible light in the surroundings) through the display unit 19, thereby providing the user with the scenery around them. That is, the user can recognize the image of the actual scenery as a background image through the display unit 19.

[0044] Figure 4 shows an example of how the display looks when a person approaches and waves at a user wearing the HMD1 while they are working, and this person is related to the user's work. As shown in Figure 4, the main image area D3, which is set near the center of the display unit 19, displays the person in the surrounding area. In addition, a reduced image of the PC image is displayed in the sub-image area D4, which is located in the lower right corner of the main image area D3. That is, if a person performing a specific gesture approaches a user who is working while looking at the PC image shown in Figure 3, and this person is related to the user's work, the display will switch from the display shown in Figure 3 to the display shown in Figure 4. The sub-image area D4 is narrower than the main image area D3.

[0045] Figure 5 shows a first display example where a person approaches and waves to a user wearing the HMD1 while they are working, but this person is not related to the user's work. As shown in Figure 5, the main image area D3, which is set near the center of the display unit 19, displays the person in the vicinity, as in Figure 4. Also, a reduced image of the PC image is not displayed in the lower right corner of the main image area D3. In other words, if a person approaches a user who is working while looking at the PC image shown in Figure 3 and performs a specific gesture, and this person is not related to the user's work, the display can be switched from the display shown in Figure 3 to the display shown in Figure 5.

[0046] Figure 6 shows a second display example where a person approaches and waves at a user wearing the HMD1 while they are working, but this person is not related to the user's work. As shown in Figure 6, the main image area D3, which is set in the center of the display unit 19, displays the person in the surrounding area, similar to Figure 4 described above. In addition, the transparency of the background image displayed in the display area D1 is set to 0%, and a reduced image of the PC image is displayed with high transparency in the lower right corner of the main image area D3.

[0047] Figure 7 shows a third display example where a person approaches and waves at a user wearing the HMD1 while they are working, but this person is not related to the user's work. As shown in Figure 7, the main image area D3, which is set in the center of the display unit 19, displays the person in the surrounding area, similar to Figure 4 described above. In addition, a reduced image of the PC image is displayed with high transparency in the lower right corner of the main image area D3.

[0048] Specifically, the display control unit 18 determines whether there is a relationship between a person and a PC image based on the person information, and controls the image to be displayed on the display unit 19 according to whether or not there is a relationship. Normally, the display control unit 18 displays a predetermined background image in the display area D1 and displays the PC image (presented image) in the main image area D3, and controls the display to show the person's image in the main image area D3 when a person is included in the surrounding image. In addition, the display control unit 18 changes the image displayed on the display unit 19 depending on whether or not the surrounding sound is a voice calling out to the user.

[0049] [Description of operation of this embodiment] Next, the process of setting the image to be displayed on the display unit 19 of the HMD1 when a person nearby approaches a user wearing the HMD1 while working will be explained with reference to the flowchart shown in Figure 8.

[0050] First, in step S11 of Figure 8, the display control unit 18 displays a background image in the overall display area D1 of the display unit 19. The background image can be, for example, an image of the office interior. Alternatively, the background image can be an image of the user's surroundings acquired in advance by the image acquisition unit 11. In this case, the transparency of the background image is set to a predetermined level. For example, the transparency should be around 50%. The display control unit 18 then displays a PC image related to the user's work, output from the terminal device 2, in the main image area D2.

[0051] As a result, as shown in Figure 3, the background image is displayed with a predetermined transparency, and the PC image is displayed in the main image area D2 set in the center of the display unit 19. The user can perform tasks while viewing the PC image displayed in the main image area D2. The background image displayed in area D1 is shown with a predetermined level of transparency, allowing the user to concentrate on the PC image and perform their work.

[0052] In step S12, the image acquisition unit 11 acquires images of the user's surroundings. As a result, images of people in the surroundings who approach and speak to the user are captured.

[0053] In step S13, the person information acquisition unit 121 determines whether or not there is a person performing a specific gesture in the image acquired by the image acquisition unit 11. As mentioned above, a specific gesture refers to gestures such as waving or beckoning. Furthermore, the person information acquisition unit 121 identifies the individual of the person performing the specific gesture based on the facial image of the person and a template registered in the database 13. As a result, an individual such as "Mr. / Ms. XX" is identified.

[0054] In step S14, the person information acquisition unit 121 determines whether the surrounding person identified in step S13 (the surrounding person who performed the specific gesture) is related to the content displayed on the PC screen. For example, if the surrounding person is a supervisor or instructor who gave instructions to the user during work, it is determined that this surrounding person is related to the content displayed on the PC screen. If they are related (S14; YES), the process proceeds to step S16; otherwise (S14; NO), the process proceeds to step S17.

[0055] In step S15, the relevance determination unit 122 quantifies the relevance between the PC image the user is working on and the people around them, and determines whether the relevance is above a predetermined threshold. For example, the relevance is quantified in the range of 0 to 100%, and it is determined whether the relevance is 50% or higher. For example, if the user is performing a difficult task and the user's supervisor speaks to the user using a specific gesture, the supervisor is evaluated as a person in the vicinity with a high degree of relevance to the user. Since the supervisor is likely to want to check on the status of the user's work, the relevance is set high. If the relevance is above the threshold (S15; YES), the process proceeds to step S16; otherwise (S15; NO), the process proceeds to step S17. As an example of a specific method for determining difficulty, if the user is spending more time than usual on a particular task, it is determined that the task is difficult. Also, if the user is frequently making errors while working, it can be determined that the task is difficult. Furthermore, if the user is frequently requesting support or help, or if overall productivity is decreasing, these are also indicators of difficulty. By combining these indicators and making a comprehensive judgment, it may be possible to more accurately determine whether or not the user is having difficulty.

[0056] In step S16, as shown in Figure 4, the display control unit 18 displays an image of the user's surroundings (an image of the person who performed a specific gesture) in the main image area D3 and displays the PC image in the sub-image area D4. That is, if a person highly relevant to the PC image (the work the user is performing) displayed on the display unit 19 approaches, the image of this person is displayed in the main image area D3, and a reduced version of the PC image is left in a sub-image area D4, which is part of the display area D1. This allows the user to communicate with people around them while confirming important information. Also, since the PC image is displayed in a reduced size in the sub-image area D4, it is possible to avoid problems such as forgetting to ask the supervisor questions about the current work. After that, this process ends.

[0057] In step S17, the display control unit 18 displays the user's surrounding image in the main image area D3 and hides the PC image, as shown in Figure 5. That is, since the reduced PC image is not displayed in the sub-image area D4 as shown in Figure 4 above, the user can concentrate on communicating with people around them without being distracted by the content of the PC image. The image displayed in 9 may be the display configuration shown in Figure 6 or Figure 7, instead of the display example shown in Figure 5.

[0058] As shown in Figures 6 and 7, the reduced PC image is displayed with a predetermined level of transparency, making it possible to immediately switch to work after communication with people around you has ended.

[0059] To summarize the above process, when a user wearing HMD1 is performing work using terminal device 2, a background image with a predetermined transparency is displayed in display area D1, as shown in Figure 3, and the PC image is displayed in the main image area D2. Therefore, the user can concentrate on the PC image and perform their work.

[0060] If a person approaches the user and performs a specific gesture, and there is a high degree of correlation between this person and the PC image (presented image) displayed on the display unit 19, then, as shown in Figure 4, the main image area D3 displays the person performing the gesture, and the sub-image area D4 displays a reduced version of the PC image. This allows the user to communicate with the person while checking the PC image.

[0061] Furthermore, if a person approaching the user performs a specific gesture, and there is little correlation between this person and the PC image displayed on the display unit 19, the main image area D3 will display the person performing the gesture, as shown in Figure 5, and the PC image will be hidden. This allows the user to communicate with the person around them without being affected by the PC image.

[0062] Next, referring to the flowchart shown in Figure 9, the processing procedure for changing the image displayed on the display unit 19 using voice recognition will be explained.

[0063] First, in step S31 of Figure 9, the display control unit 18 displays the background image in the overall display area D1 of the display unit 19. At this time, the transparency of the background image is set to a predetermined transparency. For example, the transparency should be about 50%. The display control unit 18 displays the PC image related to the user's work output from the terminal device 2 in the main image area D2.

[0064] As a result, as shown in Figure 3, the background image is displayed with high transparency, and the PC image is displayed in the main image area D2 set in the center of the display unit 19. The user can perform tasks while viewing the PC image displayed in the main image area D2. In addition, since the background image is displayed with a predetermined level of transparency, the user can concentrate on the PC image while performing tasks.

[0065] In step S32, the image acquisition unit 11 acquires images of the user's surroundings. As a result, for example, images of people in the surroundings who approach and call out to the user are captured.

[0066] In step S33, the voice acquisition unit 14 acquires sounds occurring around the user. Specifically, it acquires the voices spoken by people in the vicinity who are calling out to the user.

[0067] In step S34, the voice determination unit 15 determines whether the voice acquired by the voice acquisition unit 14 is a voice that addresses the user (a greeting voice). As mentioned above, a greeting voice is, for example, "Mr. / Ms. XX" or "Excuse me." These greeting voices are stored as templates in the database 13. The voice determination unit 15 compares the voice acquired by the voice acquisition unit 14 with the templates registered in the database 13 to determine whether it is a greeting voice. If it is a greeting voice (S34; YES), the process proceeds to step S35; otherwise (S34; NO), this process ends.

[0068] In step S35, the motion determination unit 17 determines whether the user has nodded or raised their head based on the acceleration data detected by the acceleration sensor 16. If it is detected that the user has nodded or raised their head (S35; YES), the process proceeds to step S36; otherwise (S35; NO), the process ends. In other words, if the user does not nod or raise their head, it is presumed that they do not intend to respond to people around them, so the display of the PC image is maintained in the main image area D2 shown in Figure 3.

[0069] In step S36, the display control unit 18 displays the PC image being operated by the user in the sub-image area D4, as shown in Figure 4, and displays the surrounding image (image of the person calling out to the user) in the main image area D3. After that, the process ends. In this way, when the sound of a person calling out to the user is detected, the PC image that the user is working on is displayed in the main image area D3 as shown in Figures 4 and 5, so that the user can communicate with the person calling out to them while looking at this image.

[0070] Thus, in the HMD1 according to this embodiment, when a user wears the HMD1 and performs work using a personal computer with the HMD1, the PC image is displayed in the main image area D2 (see Figure 3) set on the display unit 19 of the HMD1. Therefore, the user can perform work while viewing the PC image.

[0071] In this embodiment, if a person approaches the user and performs a specific gesture such as waving, an image of that person is displayed in the main image area D3 (see Figures 4 to 7). That is, the image switches from the PC image displayed in the main image area D2 shown in Figure 3 to the image of the surrounding person displayed in the main image area D3 shown in Figures 4 to 7. Therefore, the user can instantly recognize the surrounding person approaching them.

[0072] In this embodiment, the background image displayed in display area D1 has a predetermined transparency, while the surrounding people displayed in the main image area D3 are normal images (0% transparency). As a result, the user can instantly recognize these surrounding people and smoothly transition into communication.

[0073] In this embodiment, if a person approaching the user is related to the task the user is performing, the PC image of the user's PC is displayed in the sub-image area D4 (see Figure 4). Therefore, for example, if a supervisor who has instructed the user to perform a task approaches the user, the supervisor's image is displayed in the main image area D2 and the PC image is displayed in the sub-image area D4. This allows the user to communicate with their supervisor while confirming the content of the task they are performing using the PC image displayed in the sub-image area D4.

[0074] In this embodiment, if a person visiting the user is not related to the task the user is performing, the PC image is either not displayed or displayed with increased transparency in the sub-image area D4. This allows the user to communicate with the person around them without being affected by the PC image.

[0075] In this embodiment, if the voice acquired by the voice acquisition unit 14 is a voice calling out to the user, and the user then nods, the image of the person in the surrounding area acquired by the image acquisition unit 11 at that time is displayed in the main image area D2. Therefore, when a person who has business with the user approaches the user and makes a voice call, the image of the person who made the call is displayed in the main image area D2, enabling a smooth response.

[0076] The embodiments of the present invention have been described above, but the descriptions and drawings that constitute part of this disclosure are not part of this invention. This disclosure should not be understood as limiting. Various alternative embodiments, examples, and operational techniques will become apparent to those skilled in the art from this disclosure. [Explanation of Symbols]

[0077] 1. Head-mounted display device (HMD) 2 Terminal devices 11 Image acquisition unit 12 Specific part 13 Databases 14. Voice acquisition unit 15. Voice Judgment Unit 16. Accelerometer 17 Operation judgment section 18 Display Control Unit 19 Display section 121 Personal Information Acquisition Department 122 Relationship Determination Unit D1 display area D2, D3 Main Image Area D4 Sub-image region

Claims

1. A head-mounted display device worn by the user, A display unit that displays a real-space image and a presentation image to be presented to the user, An image acquisition unit that acquires images of the user's surroundings, A person information acquisition unit that acquires person information of a person included in the surrounding image, A display control unit determines whether or not there is a relationship between the person and the image displayed on the display unit based on the person information, and controls the image displayed on the display unit according to whether or not there is a relationship. A head-mounted display device equipped with the following features.

2. The person information acquisition unit acquires the gestures performed by the person and determines whether or not there is a correlation when the gesture is a specific gesture. The head-mounted display device according to claim 1.

3. A head-mounted display device worn by the user, A display unit that displays a real-space image and a presentation image to be presented to the user, A sound acquisition unit that acquires ambient sounds, A voice determination unit that determines whether the ambient sound is a voice calling out to the user, based on the ambient sound, A display control unit that changes the image displayed on the display unit depending on whether the ambient sound is a voice calling out to the user, A head-mounted display device equipped with the following features.

4. An image display method for displaying an image on the display unit of a head-mounted display device worn by a user, The image acquisition unit acquires images of the user's surroundings, The person information acquisition unit acquires person information of the person included in the surrounding image, The display control unit determines, based on the person information, whether there is a relationship between the person and the image displayed on the display unit, and changes the image displayed on the display unit according to whether or not there is a relationship. Image display method.

5. An image display program that causes a computer to execute the image display method described in claim 4.

Citation Information

Patent Citations

  • Program and display control device

    JP2022120553A