Information processing apparatus, and method for controlling information processing apparatus

The information processing device enhances user identification accuracy and secure information presentation by employing dual identification methods and object recognition, addressing the inflexibility of existing machine learning models.

JP2025119443APending Publication Date: 2025-08-14CANON KK
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024014333
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-01
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Machine learning models used for user identification and information presentation lack flexibility in adjusting computational speed and inference accuracy based on situational requirements, leading to inappropriate information display due to inaccurate user identification.

Method used

An information processing device with dual user identification capabilities: low-accuracy identification performed locally and high-accuracy identification via a server, combined with object recognition and information acquisition based on user identification results, ensuring appropriate information display.

Benefits of technology

Enables accurate and secure presentation of personalized information by adjusting identification accuracy and protecting personal data privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025119443000001_ABST
    Figure 2025119443000001_ABST
Patent Text Reader

Abstract

To acquire appropriate information on an object on the basis of a result of identification of a user.SOLUTION: An information processing apparatus has: first acquisition means that can acquire a first result of identification of a user from an image of the user with first accuracy; second acquisition means that can acquire a second result of identification of the user from the image with second accuracy higher than the first accuracy; and information acquisition means that acquires first information on an object located around the user on the basis of the first result and acquires second information on the object on the basis of the second result.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing device and a control method for the information processing device. [Background technology]

[0002] Conventionally, when performing inference processing on a device using a machine learning model, a method has been proposed in which the inference processing is performed using a model selected from multiple machine learning models. For example, Patent Document 1 discloses a method of providing a device that performs inference processing using a machine learning model with a model extracted from the machine learning model based on device information such as specifications. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-43772 [Non-patent literature]

[0004] [Non-Patent Document 1] S. Haykin, “Neural Networks A Comprehensive Foundation 2nd Edition”, Prentice Hall, pp.156-255, July 1998 Summary of the Invention [Problem to be solved by the invention]

[0005] However, the computational speed and inference accuracy required of a machine learning model change depending on the situation. For example, when a machine learning model is used to identify a user and information is presented to the user based on the identification result, the required identification accuracy differs depending on the content of the information to be displayed. If a machine learning model is provided in advance based on device information such as specifications, the user will not be identified using a machine learning model with accuracy appropriate to the situation, making it difficult to display appropriate information to the user.

[0006] Therefore, an object of the present invention is to provide an information processing device that acquires appropriate information about an object based on a user identification result. [Means for solving the problem]

[0007] The information processing device of the present invention is characterized by having a first acquisition means capable of acquiring a first result that identifies a user with a first accuracy from an image of the user, a second acquisition means capable of acquiring a second result that identifies the user with a second accuracy higher than the first accuracy from the image, and an information acquisition means that acquires first information about objects in the vicinity of the user based on the first result, and acquires second information about the objects based on the second result. [Effects of the Invention]

[0008] According to the present invention, appropriate information about the object can be obtained based on the user identification result. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a schematic diagram illustrating the configuration of an HMD according to a first embodiment. [Figure 2] 1 is a schematic diagram of a configuration for performing gaze detection of an HMD according to a first embodiment. [Figure 3] 10 is a flowchart illustrating a related information acquisition process. [Figure 4] FIG. 10 is a diagram illustrating detection of an object. [Figure 5] FIG. 10 is a diagram illustrating a threshold value of similarity according to the identification accuracy of a user. [Figure 6] FIG. 10 is a diagram showing an example of display of related information of an object. [Figure 7] FIG. 1 is a diagram for explaining the principle of a gaze detection method. [Figure 8] FIG. 10 is a diagram showing an eye image. [Figure 9] 10 is a flowchart illustrating an example of a gaze detection process. [Figure 10] FIG. 10 is a diagram illustrating an example of the structure of a neural network that estimates a viewing position. [Figure 11] FIG. 1 is a diagram illustrating a feature detection process and a feature integration process of a CNN. [Figure 12] FIG. 10 is a schematic diagram illustrating the configuration of an imaging device according to a second embodiment. [Figure 13] FIG. 10 is a block diagram of an imaging device according to a second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0010] <Embodiment 1> 1 and 2, a configuration of a head-mounted display (HMD) 100 as an example of a display device having an information processing device according to the present disclosure will be described.

[0011] FIG. 1 is a schematic diagram of a configuration of a head-mounted display (HMD) 100 as an example of a display device having an information processing device according to embodiment 1. The HMD 100 is a display device that is worn on the head of a user. On the left side of FIG. 1, the configuration of the HMD 100 as seen from the top of the head of a user wearing the HMD 100 is shown, and on the right side, a block diagram showing the functional configuration of the information processing device is included. Note that the HMD 100 can also be considered as an example of an information processing device having a display unit.

[0012] A user wears a housing 103 of the HMD 100 on their head and can observe real space with their left eye 101 and right eye 102 through a transmissive left display 104 and right display 105, respectively. In the example of FIG. 1, the HMD 100 is an optical see-through HMD. The optical see-through HMD 100 can superimpose virtual objects, information about virtual objects, information about real objects, and the like onto the real world that the user sees through the left display 104 and right display 105.

[0013] The HMD 100 may be a video see-through HMD having a non-transmissive display. The video see-through HMD 100 has eyepieces placed between each of the left and right eyes and the non-transmissive display. The video see-through HMD 100 also has a left camera or a right camera. In the non-transmissive mode, the non-transmissive display displays video stored inside (such as a captured video or game video). In the transparent mode, the non-transmissive display displays an image captured by the left-eye camera 106 or the right-eye camera 107 so that the real space appears transparent. The non-transmissive display may display an image that combines the internal video and the captured image.

[0014] The HMD 100 includes a CPU 128 and a memory 129. The HMD 100 also includes a motion detection unit 111, an image processing unit 112, a display control unit 113, and a gaze detection unit 114 as functional units.

[0015] The CPU 128 controls the entire HMD 100 (each functional unit). The memory 129 has a function of storing image signals from the image sensor 125 and the eye image sensor 123 shown in FIG. 2, and records images (image information) to be displayed on the displays (left display 104 and right display 105). The memory 129 also stores programs to be executed by the CPU 128. The CPU 128 can implement the functions of the motion detection unit 111, image processing unit 112, display control unit 113, and gaze detection unit 114 by executing the programs stored in the memory 129.

[0016] The motion detection unit 111 detects the head motion of the user wearing the HMD 100. The motion detection unit 111 detects the head motion based on the output value (output result) of the motion sensor 110. The motion sensor 110 includes, for example, an acceleration sensor and a gyro sensor, and can detect the acceleration of the HMD 100 in each of the X, Y, and Z directions, and the angular velocity around the X, Y, and Z axes.

[0017] The motion detection unit 111 may detect the motion of the head based on a change in the image captured by the background camera 130. The motion detection unit 111 can calculate a motion vector within the image by calculating a difference image in the time direction for the image captured by the background camera 130, for example.

[0018] Image processing unit 112 acquires images captured by background camera 130 and performs various image processing. Image processing unit 112 can detect objects that appear in the captured images. Background camera 130 can capture images of the real space in the front direction of the user.

[0019] The display control unit 113 displays images of virtual objects, information about virtual objects, information about real objects, etc. on the left display 104 and the right display 105. In this way, the display control unit 113 can superimpose images of virtual objects and various pieces of information onto the real world seen by the user through the left display 104 and the right display 105.

[0020] The gaze detection unit 114 detects the gaze position of the user on the display surfaces of the left display 104 and the right display 105, using the eye images captured by the left eye imaging unit 108 and the right eye imaging unit 109. For example, the gaze detection unit 114 can estimate the area the user is viewing on the left display 104, or the area the user is viewing on the right display 105.

[0021] 2 is a cross-sectional view of the HMD 100 cut along the YZ plane defined by the Y-axis and Z-axis shown in FIG. 1, and shows a schematic diagram of the configuration for performing gaze detection. FIG. 2 is a cross-sectional view seen from the left eye side of a user wearing the HMD 100, and shows a detailed configuration for detecting the gaze. Below, the mechanism for detecting the gaze of the left eye 101 will be described, but the HMD 100 also has a similar mechanism on the right eye side, and can detect the gaze of the right eye 102.

[0022] 2, the HMD 100 has a left display 104, an image sensor 125, an aperture mechanism 126, a focus mechanism 127, a light source 120, a light receiving lens 122, and an eye image sensor 123 arranged in a housing 103. The background camera 130 includes the image sensor 125, the aperture mechanism 126, and the focus mechanism 127, and is capable of capturing an image of the background. The memory 129 has a function of storing an image signal (captured image) from the image sensor 125.

[0023] The left eye imaging unit 108 includes a light source 120, a light receiving lens 122, and an eye imaging element 123. The light source 120 is a light source that illuminates the left eye 101 for gaze detection. For example, the light source 120 is an infrared light emitting diode that emits infrared light that is insensitive to the user. The light source 120 may include multiple infrared light emitting diodes.

[0024] An optical image of the illuminated left eye 101 (eye image; an image formed by light emitted from the light source 120 and reflected by the left eye 101) is formed on an eye image sensor 123, which has a two-dimensional array of photoelectric elements such as CMOS, by a light receiving lens 122. The eye image includes a reflection image (corneal reflection image; Purkinje image) formed by corneal reflection of light from the light source 120.

[0025] The light receiving lens 122 positions the pupil of the left eye 101 and the ocular imaging element 123 in a conjugate imaging relationship. The line of sight of the left eye 101 is detected from the positional relationship between the pupil (pupil image) and the corneal reflection image in the eyeball image formed on the ocular imaging element 123, using a predetermined algorithm described in Figures 7, 8(A), 8(B), and 9. The memory 129 has a function of storing imaging signals from the ocular imaging element 123.

[0026] Referring to FIG. 3, a process for acquiring and displaying information about objects in the user's vicinity (hereinafter referred to as related information) will be described. Objects in the user's vicinity are, for example, objects detected from the background image by the image processing unit 112. The HMD 100 can present appropriate information to the user by superimposing related information associated in advance with objects included in the background image. Note that objects in the user's vicinity are not limited to objects included in the background image, but may also include objects in the area where the user is located (for example, landmarks such as the nearest station). FIG. 3 is a flowchart illustrating the related information acquisition process.

[0027] The processing from step S301 to step S303 will be described with reference to Figures 4(A) and 4(B). Figures 4(A) and 4(B) are diagrams for explaining the detection of an object. In step S301, the image processing unit 112 captures an image of the surroundings of the user wearing the HMD 100 (the direction the user is looking) using the background camera 130, and acquires a background image 400 as shown in Figure 4(A).

[0028] In step S302, the image processing unit 112 detects objects from the background image 400. As shown in FIG. 4(B), the image processing unit 112 performs object detection processing on the background image 400 to detect objects 402, 404, 406, and 408. The object detection processing is a process of detecting objects using, for example, an existing AI model (machine learning model). Through the object detection processing, the image processing unit 112 can recognize the detected objects 402 and 404 as stores, and the detected objects 406 and 408 as trees.

[0029] Note that objects in the user's vicinity are not limited to objects detected from the background image, but may also be objects that exist within the area where the user is located (for example, within a predetermined distance from the user's location) based on the user's location information.

[0030] In step S303, CPU 128 acquires related information (third information) that is not linked to personal information from the related information of the object. Related information that is not linked to personal information is information that can be acquired without being related to a specific individual, such as the name of a store, the type of tree, or the address of the current location. Related information that is not linked to personal information can be acquired without using the user identification result. Whether the related information of the object is information that is not linked to personal information is defined in advance.

[0031] In the example of Fig. 4(B), CPU 128 acquires information on the type of tree for objects 406 and 408 that have been recognized as trees. CPU 128 can acquire the type of tree for objects 406 and 408 using an AI model that recognizes the type of tree from an image of a tree. For example, the recognition result that object 406 is a ginkgo tree and object 408 is a maple tree can be obtained.

[0032] The CPU 128 may transmit images of the objects 406, 408 to an external server via a network and receive the results of recognition by the AI model on the server side. The CPU 128 is not limited to sending images of objects detected from a background image to the server, but may also transmit user location information to the server. The CPU 128 can obtain the user's location using a location information service such as a GPS provided in the HMD 100. The server identifies objects in the vicinity of the user's location and stores related information of the identified objects. and transmit it to the HMD 100. The CPU 128 can acquire, for example, the direction to the nearest station and the time required to get there based on the user's location information, as related information not linked to personal information. In this way, the HMD 100 can acquire related information of an object based on the user's location information.

[0033] In step S304, CPU 128 determines whether user identification with low accuracy (first accuracy) has already been performed on the user. CPU 128 can identify the user using an AI model configured from a trained neural network model, with the user's eye image captured by eye image sensor 123 as input. In the following description, users who can be identified by the AI model are assumed to be users whose eye images have already been collected when training the AI model. The configuration and training method of the AI model for identifying users will be described later with reference to FIGS. 10 and 11.

[0034] User identification does not necessarily have to be performed by the CPU 128 of the HMD 100, but may also be performed by transmitting eye images to an external server and having the processing performed on the server. Devices such as the HMD 100 that are intended to be carried by the user have size and power constraints that limit the size of the AI model used to identify users. Therefore, the HMD 100's user identification performance is also limited.

[0035] On the other hand, the server is larger than the HMD100 and can use a more accurate AI model. Therefore, the server can improve the user identification performance. When identifying users on the server, the time required to send eye images to the server and receive the identification results increases, but the time required to identify users can be shortened depending on the performance of the server.

[0036] The AI model for identifying users is generated by training so that it can recognize users even if their eye images differ from the images used for training. Furthermore, for more stable and accurate user recognition, it is preferable that the eye images input to the AI model be captured in the same environment as the images used for training. For example, if eye images captured with the user's gaze primarily facing forward are used for training, it is preferable that the HMD 100 identify the user with the user facing forward.

[0037] Taking into account the characteristics of user identification, the HMD 100 according to the first embodiment identifies users with two different levels of accuracy depending on the situation. The first identification process is a user identification process executed by the CPU 128 of the HMD 100 regardless of the direction of the user's gaze, and is referred to as low-accuracy identification. The HMD 100 may execute low-accuracy identification by instructing the user, or may execute low-accuracy identification using captured eye images without notifying the user and obtain an identification result. The second identification process is a user identification process executed by the server after instructing the user to gaze forward and sending captured eye images to the server. This is referred to as high-accuracy identification. Note that high-accuracy identification only needs to identify the user with higher accuracy than low-accuracy identification, and high-accuracy identification and low-accuracy identification are not limited to the above examples.

[0038] For example, the HMD 100 can obtain a classification result based on low-accuracy classification and a classification result based on high-accuracy classification using different machine learning models. The first AI model used to obtain a classification result based on low-accuracy classification may be generated so as to have lower classification accuracy than the second AI model used to obtain a classification result based on high-accuracy classification. Furthermore, the first AI model may be generated so as to satisfy at least one of the following: a smaller data size than the second AI model (including a smaller number of intermediate hierarchical layers and a smaller file size) and a shorter processing time than the second AI model. The HMD 100 performs low-accuracy classification using the first AI model and high-accuracy classification using the second AI model. Identification can be performed.

[0039] Furthermore, the low-accuracy classification is not limited to being performed by the HMD 100 and the high-accuracy classification is not limited to being performed by the server. For example, the HMD 100 may not perform user identification by a server, and may instead use a first AI model for low-accuracy classification and a second AI model for high-accuracy classification that is different from the first AI model. The first AI model is a model with lower classification accuracy and shorter processing time than the second AI model. Conversely, the second AI model is a model with higher classification accuracy and longer processing time than the first AI model.

[0040] The HMD 100 may also use the same AI model to obtain both low-accuracy and high-accuracy classification results. The HMD 100 can achieve low-accuracy and high-accuracy classification using the same AI model by applying different thresholds to the similarity used to identify the user. The HMD 100 simply sets a first threshold of similarity used to identify the user with low accuracy lower than a second threshold of similarity used to identify the user with high accuracy.

[0041] With reference to Figure 5, we will explain a specific example of similarity thresholds according to the accuracy of user identification. The False Acceptance Rate (FAR) and False Rejection Rate (FRR) are used as indicators of the accuracy of user identification. FAR is the probability that a different person will be mistakenly accepted as the user. FRR is the probability that the user will be mistakenly rejected (not accepted as the user). The vertical axis of the graph in Figure 5 is the probability of the FAR and FRR of the AI model that identifies the user. The horizontal axis is similarity. The closer the similarity value is to 1, the higher the possibility that the user is the correct person.

[0042] The similarity threshold is set based on the FAR and FRR. When the similarity threshold is set to the threshold 502 (first threshold), the FAR becomes higher than the FRR. In this case, a certain degree of false detection is tolerated, and the actual person is less likely to be detected as a different person. In other words, the HMD 100 can perform low-accuracy user identification by setting the first threshold so that the FAR is higher than the FRR.

[0043] When the threshold for the similarity is set to the threshold 504 (second threshold), the FAR becomes lower than the FRR. In this case, if the similarity is not higher than the threshold 504, the user will not be identified as the person. That is, the HMD 100 can perform highly accurate user identification by setting the second threshold so that the FAR is lower than the FRR.

[0044] In step S304 of Fig. 3, CPU 128 determines whether low-accuracy classification has been performed on the user. CPU 128 can determine that low-accuracy classification has been performed, for example, when a result of identifying the user with low accuracy is obtained from an image of the user's eyes. If low-accuracy classification has been performed (S304: YES), the process proceeds to step S307. If low-accuracy classification has not been performed (S304: NO), the process proceeds to step S305.

[0045] In step S305, CPU 128 determines whether low-accuracy identification can be performed. For example, CPU 128 notifies the user that low-accuracy identification will be performed, and if the user allows it, CPU 128 can determine that low-accuracy identification can be performed. If low-accuracy identification can be performed (S305: YES), the process proceeds to step S306. If low-accuracy identification cannot be performed (S305: NO), the process proceeds to step S313.

[0046] It should be noted that the CPU 128 may omit the process of step S305. In this case, the CPU 128 performs low-accuracy classification without notifying the user and obtains the classification result.

[0047] In step S306, the CPU 128 performs low-accuracy user identification. As described above, the CPU 128 can acquire an identification result in which the user is identified with low accuracy from the user's eye image. The CPU 128 may perform low-accuracy identification in the HMD 100 to acquire the identification result, or may transmit the user's eye image to an external server and acquire the low-accuracy identification result from the server.

[0048] In step S307, CPU 128 acquires, from the related information of the object detected in step S302 based on the low-accuracy user identification result, publicly available related information (first information) linked to the personal information of the identified user. Publicly available related information linked to personal information is, for example, information based on the user's preferences, and is general information that can be made public to users other than the user. Whether the related information of the object is publicly available information linked to personal information is defined in advance.

[0049] For example, if the store of the object 402 in Fig. 4(B) is a bookstore, the CPU 128 acquires information about new books in the user's favorite genre that has been identified with low accuracy. The user's favorite information may be stored in advance in the memory 129 of the HMD 100, or may be acquired by estimating it from related information that the user directed their gaze at among related information that has been displayed in the past.

[0050] In step S308, CPU 128 determines whether or not non-public related information (second information) linked to personal information is associated with the object detected in step S302. Non-public related information linked to personal information is information that is not disclosed to other users and can identify an individual, such as a user's purchase history or store reservation information by the user. Whether or not the related information of an object is non-public information linked to personal information is defined in advance.

[0051] For example, CPU 128 queries the server of the store of object 404 in FIG. 4(B) or the server of a reservation site registered with the store to acquire purchase history or reservation information for object 404. However, because high-precision identification is not performed in step S308, CPU 128 acquires information on whether purchase history or reservation information is associated with the object, but does not acquire detailed information on the respective information. Details of the purchase history or reservation information can be acquired after high-precision identification is performed. If private related information linked to personal information is associated with the object in step S308, processing proceeds to step S309. If private related information linked to personal information is not associated with the object, processing proceeds to step S313.

[0052] In step S309, the CPU 128 prompts the user to perform high accuracy identification. For example, the CPU 128 instructs the user to face forward in order to perform high accuracy identification.

[0053] In step S310, CPU 128 determines whether high accuracy identification can be performed. For example, CPU 128 prompts the user to perform high accuracy identification, and determines that high accuracy identification can be performed when the user faces forward in accordance with the instructions of CPU 128. Note that if the eye image for performing high accuracy identification can be acquired, such as when the user faces forward, CPU 128 can omit the processing of steps S309 and S310. If high accuracy identification can be performed (S310: YES), the processing proceeds to step S311. If high accuracy identification cannot be performed (S310: NO), the processing returns to step S311. Proceed to step S313.

[0054] In step S311, CPU 128 performs high-accuracy user identification. As described above, CPU 128 can obtain an identification result that identifies the user with high accuracy from the user's eye image. CPU 128 may transmit the user's eye image to an external server and obtain the high-accuracy identification result from the server.

[0055] In step S312, based on the highly accurate user identification result, CPU 128 acquires non-public related information (second information) linked to personal information of the identified user from the related information of the object detected in step S302. CPU 128 acquires, for example, reservation details such as reservation date and time, number of people, and reserved course for the store of object 404 in FIG. 4(B).

[0056] 4B are stored on different servers, the CPU 128 may prompt the user to perform high-accuracy identification before acquiring the associated information. That is, the CPU 128 may acquire a result of identifying the user with high accuracy before acquiring the associated information linked to personal information for each of the multiple objects.

[0057] Furthermore, from a security perspective, CPU 128 may set a validity period for the identification result of high-accuracy identification. That is, CPU 128 invalidates the result of high-accuracy identification after a predetermined time has passed since the user was identified with high accuracy. If the user wishes to obtain non-public related information linked to personal information again, CPU 128 may prompt the user to perform high-accuracy identification again.

[0058] In step S313, the CPU 128 causes the display control unit 113 to display the object-related information acquired in steps S303, S307, and S312 on the left display 104 and the right display 105. A display example of object-related information will be described with reference to Fig. 6 .

[0059] 6(A) shows an example of the display of related information when it is determined in step S305 that low-accuracy classification is not possible. In FIG. 6(A), the display control unit 113 displays the related information acquired in step S303, i.e., related information not linked to personal information, to notify the user. In the example of FIG. 6(A), the display control unit 113 displays, as related information 606 for object 406, that the type of tree is "ginkgo," and as related information 608 for object 408, that the type of tree is "maple."

[0060] FIG. 6(B) shows an example of the display of related information when it is determined in step S308 that there is no private related information linked to personal information, or when it is determined in step S310 that high-accuracy identification is not possible. In FIG. 6(B), the display control unit 113 notifies the user by displaying the related information acquired in step S307, i.e., the publicly available related information linked to personal information, in addition to the related information acquired in step S303. In the example of FIG. 6(B), the display control unit 113 displays the release information of a new book at a bookstore, "New Comic Y Release Available Tomorrow," as related information 610 for the object 402, in addition to the related information displayed in FIG. 6(A). The related information 610 is information acquired based on the user's preferred genre. Note that the display control unit 113 may display either related information not linked to personal information or publicly available related information linked to personal information.

[0061] 6C shows an example of display of related information when high accuracy identification is performed in step S311. In FIG. 6C, the display control unit 113 displays the related information acquired in step S312, that is, the private information linked to personal information, in addition to the related information acquired in steps S303 and S307. 6(C), the display control unit 113 displays the store reservation information "Reservation for 2 people from 12 noon on June 1st" as related information 612 for the object 404, in addition to the related information displayed in FIG. 6(B). The related information 612 is information acquired based on the user's personal information from a server of the store of the object 404, etc. Note that the display control unit 113 may select and display at least any of related information not linked to personal information, publicly available related information linked to personal information, and privately available related information linked to personal information.

[0062] 6(A) to 6(C), the display control unit 113 may display information related to an object that the user is looking at. The object that the user is looking at is an object that exists at the user's line of sight. When acquiring information related to the object in steps S303, S307, and S312, the CPU 128 may acquire information related to the object that exists at the user's line of sight.

[0063] Furthermore, CPU 128 may display the related information for a predetermined time and then hide it. The time for displaying the related information can be determined based on, for example, the amount of information (such as the number of characters) of the related information to be displayed.

[0064] In the above-described first embodiment, the HMD 100 acquires and displays information (related information) about objects around the user according to the user's identification accuracy. This allows the HMD 100 to protect the security of personal information and provide appropriate information to the user. Therefore, the usability of the HMD 100 is improved for the user wearing the HMD 100.

[0065] (Explanation of gaze detection method) Here, a gaze detection method for acquiring and displaying related information about an object present at the user's gaze position will be described in detail. In step S303, the HMD 100 transmits images of the objects 406 and 408 detected in FIG. 4B to a server and acquires the related information about each object. However, the HMD 100 may also acquire the related information by transmitting an image of an object that the user wearing the HMD 100 is looking at to the server. For example, if the user is looking at object 406, the HMD 100 may transmit only the image of object 406 to the server to acquire the related information. By detecting the user's gaze, the HMD 100 can efficiently acquire the related information about the object that the user is looking at.

[0066] The gaze detection method (gaze detection algorithm) will be described using Figures 7, 8(A), 8(B), and 9. Figure 7 is a diagram for explaining the principle of the gaze detection method, and is a schematic diagram of an optical system for performing gaze detection. As shown in Figure 7, light sources 713a and 713b are arranged approximately symmetrically with respect to the optical axis of a light receiving lens 716, and illuminate the user's eyeball 14. A portion of the illumination light emitted from light sources 713a and 713b and reflected by the eyeball 714 is collected by the light receiving lens 716 onto an eye imaging element 717.

[0067] FIG. 8(A) is a schematic diagram of an eye image captured by the eye image sensor 717 (eyeball image projected onto the eye image sensor 717), and FIG. 8(B) is a diagram showing the output intensity of the CCD in the eye image sensor 717.

[0068] FIG. 9 is a flowchart illustrating the gaze detection process. When the gaze detection process starts, in step S901, the light sources 713a and 713b emit infrared light toward the user's eyeball 714. An image of the user's eyeball illuminated by the infrared light is formed on the eye image sensor 717 through the light receiving lens 716 and is photoelectrically converted by the eye image sensor 717. As a result, an electrical signal of the eye image that can be processed is obtained. In step S902, the CPU 128 The eye image (image data, image signal) obtained from the imaging element 717 is acquired.

[0069] In step S903, the CPU 128 obtains the corneal reflection images Pd and Pe of the light sources 713a and 713b and the coordinates of a point corresponding to the pupil center c from the eye image obtained in step S902. Infrared light emitted from the light sources 713a and 713b illuminates the cornea 742 of the user's eyeball 714. At this time, the corneal reflection images Pd and Pe formed by a portion of the infrared light reflected from the surface of the cornea 742 are condensed by the light receiving lens 716 and formed on the ocular imaging element 717 as corneal reflection images Pd' and Pe' in the eye image. Similarly, light beams from the edges a and b of the pupil 741 are also formed on the ocular imaging element 717 as pupil edge images a' and b' in the eye image.

[0070] Figure 8(B) shows the luminance information (luminance distribution) of region α in the eye image of Figure 8(A). In Figure 8(B), the horizontal direction of the eye image is the X-axis direction, and the vertical direction is the Y-axis direction. Figure 8(B) shows the luminance distribution in the X-axis direction. The coordinates of the corneal reflection images Pd' and Pe' in the X-axis direction (horizontal direction) are Xd and Xe, and the coordinates of the pupil edge images a' and b' in the X-axis direction are Xa and Xb.

[0071] As shown in FIG. 8(B), an extremely high level of luminance is obtained at the coordinates Xd and Xe of the corneal reflection images Pd' and Pe'. In the region from coordinate Xb to coordinate Xa, which corresponds to the region of the pupil 741 (the region of the pupil image obtained when the light beam from the pupil 741 is focused on the ocular imaging element 717), an extremely low level of luminance is obtained except for coordinates Xd and Xe. In the region of the iris 743 outside the pupil 741 (the region of the iris image outside the pupil image obtained when the light beam from the iris 743 is focused), a luminance intermediate between the above two types of luminance is obtained. Specifically, a luminance intermediate between the above two types of luminance is obtained in the region where the X coordinate (coordinate in the X-axis direction) is smaller than coordinate Xa and the region where the X coordinate is larger than coordinate Xb.

[0072] 8(B), the CPU 128 can obtain the X-coordinates Xd and Xe of the corneal reflection images Pd' and Pe' and the X-coordinates Xa and Xb of the pupil edge images a' and b'. Specifically, the CPU 128 can obtain the coordinates where the luminance is extremely high as the coordinates of the corneal reflection images Pd' and Pe', and the coordinates where the luminance is extremely low as the coordinates of the pupil edge images a' and b'.

[0073] Furthermore, when the rotation angle θx of the optical axis of the eyeball 714 relative to the optical axis of the light receiving lens 716 is sufficiently small, the coordinate Xc of the pupil-centered image c' (center of the pupil image) obtained when the light beam from the pupil center c is focused on the ocular imaging element 17 can be expressed as Xc ≒ (Xa + Xb) / 2. In other words, the CPU 128 can calculate the coordinate Xc of the pupil-centered image c' from the X-coordinates Xa and Xb of the pupil edge images a' and b'. In this way, the CPU 128 can estimate the coordinates of the corneal reflection images Pd' and Pe' and the coordinate of the pupil-centered image c'.

[0074] In step S904, CPU 128 calculates the imaging magnification β of the eyeball image. The imaging magnification β is determined by the position of eyeball 714 relative to light receiving lens 716, and can be calculated using a function of the distance (Xd-Xe) between corneal reflection images Pd' and Pe'.

[0075] In step S905, CPU 128 calculates the rotation angle of the optical axis of eyeball 714 relative to the optical axis of light receiving lens 716. The X coordinate of the midpoint between corneal reflection image Pd and corneal reflection image Pe approximately coincides with the X coordinate of the center of curvature O of cornea 742. Therefore, if the standard distance from the center of curvature O of cornea 742 to the center c of pupil 741 is Oc, then the rotation angle θx of the optical axis of eyeball 714 in the ZX plane (plane perpendicular to the Y axis) can be calculated using the following equation 1. The rotation angle θy of eyeball 714 in the ZY plane (plane perpendicular to the X axis) can also be calculated using a method similar to that for calculating rotation angle θx. β×Oc×SINθx≒{(Xd+Xe) / 2}-Xc (Formula 1)

[0076] In step S906, CPU 128 uses the rotation angles θx, θy calculated in step S905 to determine (acquire) the user's viewpoint (gaze position; the position at which the user is looking) on the screen of left display 104. If the coordinates (Hx, Hy) of the viewpoint are coordinates corresponding to the pupil center c, the coordinates (Hx, Hy) of the viewpoint can be calculated using the following equations 2 and 3. Hx=m×(Ax×θx+Bx) (Formula 2) Hy=m×(Ay×θy+By) (Formula 3)

[0077] Parameter m in equations 2 and 3 is a constant that represents the relationship between the user's eyeball angle and the position of the viewpoint on the left display 104, and is a conversion coefficient that converts the rotation angles θx and θy into coordinates corresponding to the pupil center c on the left display 104. Parameter m is determined in advance and stored in memory 129. Parameters Ax, Bx, Ay, and By are gaze correction parameters that correct for individual differences in gaze, and can be obtained through calibration. Parameters Ax, Bx, Ay, and By are stored in memory 129 before the gaze detection process starts.

[0078] In step S907, CPU 128 stores the coordinates (Hx, Hy) of the viewpoint (pupil center c) calculated in step S906 in memory 129, and ends the gaze detection process. Note that, although Fig. 9 shows an example in which the corneal reflection images of light sources 713a and 713b are used to obtain the rotation angle of the eyeball and to obtain the coordinates of the viewpoint on left display 104, the present invention is not limited to this. A method for obtaining the rotation angle of the eyeball from the eyeball image may be, for example, a method of measuring the gaze from the pupil center position.

[0079] By performing the above-described gaze detection process, the HMD 100 can acquire related information about the object the user is looking at among multiple objects around the user, thereby enabling efficient acquisition of related information in a short time. Furthermore, the HMD 100 can limit the information superimposed on the background image 900 shown in FIG. 4 to information related to the object the user is looking at. The HMD 100 can improve usability by acquiring related information about the object the user is looking at and superimposing it on the background image 900.

[0080] (Neural network explanation) 10 and 11, we will explain the neural network, which is an AI model used for low-accuracy and high-accuracy user identification described in step S304 of Fig. 3. Using a CNN (Convolutional Neural Network) as an example, we will explain the basic configuration of a neural network inference model. Fig. 10 shows the basic configuration of a CNN that extracts features from an input image, which is two-dimensional image data.

[0081] CNN includes multiple layers, each of which includes two layers called the feature detection layer (S layer) and the feature integration layer (C layer). In the example of Figure 10, the input image input to the CNN is processed in order from the first layer to the Xth layer.

[0082] In CNN, first, in layer S, features of the input image are detected based on the features detected in the previous layer. Next, the features detected in layer S are integrated in layer C and input to the next layer as the detection result of the current layer.

[0083] The S layer includes multiple feature-detecting cell planes, each of which detects a different feature. The C layer includes multiple feature integration cell planes and pools the detection results from the feature detection cell planes in the S layer. In the example of Figure 10, the final layer, the output layer (Xth layer), is composed of the S layer without using the C layer. The feature detection cell planes and feature integration cell planes are collectively referred to as cell planes.

[0084] The feature detection process at the feature detection cell plane and the feature integration process at the feature integration cell plane will be described in detail with reference to Figure 11. In Figure 11, rectangles indicate cell planes. The S layer at the Lth level includes multiple feature detection cell planes, and the C layer at the L-1th level and the C layer at the Lth level each include multiple feature integration cell planes.

[0085] The feature detection cell surface is composed of multiple feature detection neurons, which are connected to the C layer of the previous layer in a predetermined structure. The feature integration cell surface is composed of multiple feature integration neurons, which are connected to the S layer of the same layer in a predetermined structure.

[0086] In the m-th cell plane of the S layer of the Lth hierarchy, the output value of the feature detection neuron at position (ξ,ζ) is y m LS (ξ,ζ), the feature of the position (ξ,ζ) in the mth cell plane of the C layer of the Lth layer The output value of the integration neuron is y m LC (ξ,ζ). The connection coefficients of each neuron The number w m LS (n,u,v), w m LC Assuming (u,v), each output value can be expressed as follows:

number

[0087] In Equation 4, f is an activation function, which may be a sigmoid function such as a logistic function or a hyperbolic tangent function, and is realized, for example, by a tanh function. m LS (ξ,ζ) is the L-level The output value of the feature detection neuron shown in Equation 4 is the internal state u of the feature detection neuron at position (ξ,ζ) on the m-th cell plane of the layer S of the eye. m LS It is calculated by transforming (ξ,ζ) with the activation function f.

[0088] The output value of the feature integration neuron shown in Equation 5 is calculated by the coupling coefficient w without using an activation function. m LC It is calculated by a simple linear sum of (u,v) and the mth output value of the Sth layer of the Lth hierarchy. Without the use of a quantization function, the internal state u of the feature integration neuron m LC (ξ,ζ) and output value y m LC (ξ,ζ) is equal to y in Eq. n L-1C (ξ+u,ζ+v), y in Eq. m LS (ξ+u, ζ+v) are called the output values of the feature detection neuron and the feature integration neuron, respectively.

[0089] The following explains ξ, ζ, u, v, and n in Equations 4 and 5. The position (ξ, ζ) corresponds to the position coordinates in the input image. y m LS (ξ,ζ) is the output of the feature detection neuron at other positions. If the force value is higher than the force value, it means that the feature detected in the m-th cell plane of the S layer of the Lth hierarchy is likely to exist at the pixel position (ξ, ζ) of the input image.

[0090] The n in Equation 4 means the n-th cell plane in the C layer of the L-1th layer, and is called the target feature number. Basically, a product-sum operation is performed for each cell plane in the C layer of the L-1th layer. In other words, the internal state u of the feature detection neuron m LS (ξ,ζ) is the number of layers in the C layer of the L-1th layer. For the cell surface, the coupling coefficient w m LS (n,u,v) and the output value of the feature detection neuron y n L-1C It is calculated by multiplying and adding (ξ+u,ζ+v). (u,v) is the coupling coefficient It is a relative position coordinate, and the product-sum operation is performed within a finite range (u,v) depending on the size of the feature to be detected. The finite range (u,v) is called the receptive field. The size of the receptive field is determined by the size of the connected It is expressed as the number of horizontal pixels x the number of vertical pixels of the range, and is hereinafter referred to as the receptive field size.

[0091] In Equation 4, L=1, that is, the first S layer, y n L-1C (ξ+u,ζ+v) is the input Force Image in_image (ξ+u,ζ+v) or input position map y in_posi_map (ξ+u, The distribution of neurons and pixels is discrete, and the connection feature numbers are also discrete. Therefore, ξ, ζ, u, v, and n are not continuous variables, but take discrete values. Here, ξ and ζ are non-negative integers, n is a natural number, and u and v are integers, all of which have a finite range.

[0092] w in Equation 4 m LS (n,u,v) is the distribution of coupling coefficients for detecting a given feature. By adjusting the connection coefficients to appropriate values, it becomes possible to detect specific features. In the construction (learning) of CNN, various test patterns are presented and y m LS (ξ,ζ) is The coupling coefficients are adjusted by repeatedly modifying them gradually to obtain appropriate output values.

[0093] w in Equation 5 m LC (u,v) is expressed as Equation 6 using a two-dimensional Gaussian function. do.

number

[0094] (u,v) is a finite range, and as in the explanation of feature detection neurons, the finite range (u,v) is called the receptive field, and the size of the receptive field range is called the receptive field size. The receptive field size can be set to an appropriate value depending on the size of the feature detected on the mth cell plane of the S layer of the Lth hierarchy. In Equation 6, σ is the feature size factor, and is set to an appropriate constant depending on the receptive field size. Specifically, the feature size factor σ is preferably set to a value that allows the value of the coupling coefficient at the outermost edge of the receptive field to be considered nearly 0.

[0095] By performing the above calculations at each layer of the neural network, the features extracted at the final layer, S, are treated as features of the input image.

[0096] We will explain the specific learning method of neural networks. The connection coefficients are adjusted by supervised learning. In supervised learning, test patterns are given to actually obtain the output values of neurons, and the connection coefficients w m LS (n,u,v) is the actual neuron output value and the training signal. The coupling coefficients can be corrected in relation to the signal (the desired output value that the neuron should output). For example, the coupling coefficients can be corrected using the least squares method in the feature detection layer at the final layer, and using the backpropagation method in the feature detection layers at the intermediate layers. The coupling coefficients can be corrected using the least squares method, the backpropagation method, or other known methods such as those disclosed in Non-Patent Document 1.

[0097] When training a neural network in advance, many test patterns are prepared for training, including specific patterns to be detected and patterns that should not be detected. When the activation function is set to a tanh function and a specific pattern to be detected is presented, a teacher signal is given to neurons in the area of the feature detection cell plane at the final layer where the specific pattern exists so that the output value becomes 1. Conversely, when a pattern that should not be detected is presented, a teacher signal is given to neurons in the area where the pattern that should not be detected exists so that the output value becomes -1.

[0098] By the above method, a neural network that can identify users from input images is constructed. The actual detection (identification) process is performed using the connection coefficients w m LS (n,u , v). If the neuron output on the feature detection cell plane of the final layer is equal to or greater than a predetermined value, the input image is determined to be an image of a specific user registered during learning. In the first embodiment, the AI model for identifying a user is a model trained on images of the user's eyes, and the user can be identified by inputting images of the eyes of a user wearing the HMD 100. Note that the AI model is not limited to eye images, and may be a model that identifies a user by training on images of the user's face or part of their face.

[0099] <Embodiment 2> Similar to the HMD 100 described in the first embodiment, the second embodiment is an embodiment of an imaging device having an information processing device according to the present disclosure. The configuration of an imaging device 1200 according to the second embodiment will be described with reference to Fig. 12. Fig. 12 is a cross-sectional view of the imaging device 1200, and is a schematic configuration diagram of the imaging device 1200. Note that the imaging device 1200 can also be regarded as an example of an information processing device having a display unit (display device 1210).

[0100] The photographing lens unit 1270 of the interchangeable lens camera includes two lenses 1201 and 1202, an aperture 1261, and an aperture drive unit 1262. The photographing lens unit 1270 also includes a lens drive motor 1263, a lens drive member 1264, a photocoupler 1265, a pulse plate 1266, a mount contact 1267, a focus adjustment circuit 1268, etc. The lens drive member 1264 is made up of a drive gear, etc. The photocoupler 1265 detects the rotation of the pulse plate 1266 that is linked to the lens drive member 1264 and transmits this to the focus adjustment circuit 1268. The focus adjustment circuit 1268 drives the lens drive motor 1263 based on information from the photocoupler 1265 and information from the camera housing 1271 (information on the lens drive amount) to move the lens 1201 and change the focus position. Mount contact 1267 is an interface between taking lens unit 1270 and camera body 1271. For simplicity, two lenses 1201 and 1202 are shown, but in reality, taking lens unit 1270 contains more than two lenses.

[0101] The camera housing 1271 contains an image sensor 1272, a CPU 1273, a memory 1274, a display device 1210, a display device drive circuit 1211, etc. The image sensor 1272 is disposed at a planned imaging plane of the photographing lens unit 1270. The CPU 1273 is a central processing unit of a microcomputer, and controls the entire imaging device 1200. The memory 1274 records images captured by the image sensor 1272, etc. The display device 1210 is composed of a liquid crystal display or the like, and displays the captured images, etc. on the screen (display surface) of the display device 1210. The display device drive circuit 1211 drives the display device 1210.

[0102] The camera housing 1271 also contains light sources 1213a and 1213b, a light splitter 1215, a light receiving lens 1216, an eye imaging element 1217, etc. The light sources 1213a and 1213b are used to detect the gaze direction from the relationship between the pupil and a reflection image of light reflected by the cornea (corneal reflection image), and are light sources for illuminating the user's eyeball 1214. Specifically, the light sources 1213a and 1213b are infrared light emitting diodes or the like that emit infrared light that is insensitive to the user, and are arranged around the eyepiece 1212. An optical image of the illuminated eyeball 1214 (eye image; an image formed by light emitted from the light sources 1213a and 1213b and reflected by the eyeball 1214) passes through the eyepiece 1212 and is reflected by the light splitter 1215. The eyeball image is then formed by a light receiving lens 1216 on an ocular imaging element 1217, which has a two-dimensional array of photoelectric elements such as a CCD or CMOS. The light receiving lens 1216 positions the pupil of the user's eyeball 1214 and the ocular imaging element 1217 in a conjugate imaging relationship. From the position of the corneal reflection image in the eyeball image formed on the ocular imaging element 1217, the user's line of sight (the viewpoint on the screen of the display device 1210) is detected by the predetermined algorithm described in the first embodiment.

[0103] Operation members 1241 to 1243 that accept various operations from the user are arranged on the back of camera housing 1271. For example, operation member 1241 is a touch panel that accepts touch operations, operation member 1242 is an operation lever that can be pushed down in all directions, and operation member 1243 is a four-way key that can be pressed in each of four directions. Operation member 1241 (touch panel) is equipped with a display panel such as a liquid crystal panel, and has the function of displaying images on the display panel.

[0104] Figure 13 is a block diagram showing the electrical configuration within the imaging device 1200. Components that are the same as those in Figure 12 are assigned the same reference numerals. The CPU 1273 is connected to a gaze detection circuit 1301, a photometry circuit 1302, an autofocus detection circuit 1303, a signal input circuit 1304, a display device drive circuit 1211, and an illumination light source drive circuit 1305. The CPU 1273 also transmits signals via mount contacts 1267 to a focus adjustment circuit 1218 disposed in the photographing lens unit 1270 and an aperture control circuit 1306 included in an aperture drive section 1262 within the photographing lens unit 1270. A memory 1274 associated with the CPU 1273 has a function of storing image capture signals from the image sensor 1272 and the eye image sensor 1217, and a function of storing gaze correction data that corrects for individual differences in the gaze.

[0105] The gaze detection circuit 1301 A / D converts the output (eye image captured by the eye image sensor 1217 (CCD-EYE) when an eyeball image is formed on the eyeball image sensor 1217, and transmits the result to the CPU 1273. The CPU 1273 extracts each feature point of the eyeball image used for gaze detection according to the predetermined algorithm described in the first embodiment, and calculates the user's gaze (the viewpoint on the screen of the display device 1210) from the position of each feature point.

[0106] The photometry circuit 1302 amplifies, logarithmically compresses, and A / D converts the signal obtained from the image sensor 1272, which also functions as a photometry sensor, specifically the luminance signal corresponding to the brightness of the field, and sends the result to the CPU 1273 as field luminance information.

[0107] The autofocus detection circuit 1303 A / D converts signal voltages from multiple detection elements (multiple pixels) used for phase difference detection, which are included in the CCD in the image sensor 1272, and sends the converted signal to the CPU 1273. The CPU 1273 calculates the distance to the subject corresponding to each focus detection point from the signals from the multiple detection elements. This is a well-known technique known as image plane phase difference AF.

[0108] The signal input circuit 1304 is connected to a switch SW1 that is turned on with the first stroke of a release button (not shown) and that starts photometry, distance measurement, line-of-sight detection processing, etc. of the imaging device 1200. The signal input circuit 1304 is also connected to a switch SW2 that is turned on with the second stroke of the release button and that starts a release operation. On signals from the switches SW1 and SW2 are input to the signal input circuit 1304 and transmitted to the CPU 1273. A light source drive circuit 1305 drives the light sources 1213a and 1213b.

[0109] The operation member 1241 (touch panel compatible liquid crystal display), operation member 1242 (lever type operation member), and operation member 1243 (button type four-way key) are configured to transmit their respective operation signals to the CPU 1273. The CPU 1273 controls the movement of the viewpoint that has been detected (acquired) in accordance with the received operation signals.

[0110] The image captured by the image sensor 1272 corresponds to the background image 400 described in FIG. 4(A) of the first embodiment. The eye image captured by the eye image sensor 1217 corresponds to the eye image described in FIG. 8(A) of the first embodiment. The image capturing device 1200 can acquire and display related information of the object using the background image captured by the image sensor 1272 and the eye image captured by the eye image sensor 1217, similar to the related information acquisition process according to the first embodiment described in FIG. 3. do.

[0111] In the second embodiment, the image capturing device 1200, like the HMD 100 according to the first embodiment, acquires and displays information (related information) about objects around the user according to the user's identification accuracy. This allows the image capturing device 1200 to protect the security of personal information and provide appropriate information to the user. This improves the usability of the image capturing device 1200 for the photographer.

[0112] Although the embodiments of the present invention have been described in detail, the present invention is not limited to these specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention. Furthermore, each of the above-described embodiments merely shows one embodiment of the present invention, and each embodiment can be combined as appropriate.

[0113] <Other embodiments> The present invention can also be realized by supplying a program that realizes one or more functions of the above-described embodiments to a system or device via a network or a storage medium, and having one or more processors in the computer of the system or device read and execute the program. It can also be realized by a circuit (e.g., ASIC) that realizes one or more functions.

[0114] The disclosure of this embodiment includes the following configuration, method, program, and medium. (Configuration 1) a first acquisition means capable of acquiring a first result that identifies the user with a first accuracy from an image of the user; a second acquisition means capable of acquiring a second result in which the user is identified from the image with a second accuracy higher than the first accuracy; and an information acquisition means for acquiring first information about an object in the vicinity of the user based on the first result, and acquiring second information about the object based on the second result. (Configuration 2) The information acquisition means acquires third information about the object that can be acquired without using the first result and the second result. 2. The information processing device according to configuration 1, (Configuration 3) The information acquisition means acquires information about the object present at the user's line of sight. 3. The information processing device according to configuration 1 or 2. (Configuration 4) The information acquisition means acquires information about the object based on the location information of the user. 4. The information processing device according to any one of configurations 1 to 3. (Configuration 5) When the second information is associated with the object, the information acquisition means prompts the user to identify the object with the second accuracy. 5. The information processing device according to any one of configurations 1 to 4. (Configuration 6) The second acquisition means invalidates the second result after a predetermined time has elapsed since the user was identified with the second accuracy. 6. The information processing device according to any one of configurations 1 to 5. (Configuration 7) The information acquisition means acquires the second information for each of the plurality of objects. Obtain the second result 7. The information processing device according to any one of configurations 1 to 6. (Configuration 8) The first acquisition means acquires the first result without notifying the user. 8. The information processing device according to any one of configurations 1 to 7. (Configuration 9) the first result and the second result are results of identifying the user using the same machine learning model; A first threshold of similarity used to identify the user with the first accuracy is set lower than a second threshold of similarity used to identify the user with the second accuracy. 9. The information processing device according to any one of configurations 1 to 8. (Configuration 10) the first threshold and the second threshold are set based on a false acceptance rate and a false rejection rate, which are indicators of accuracy in identifying the user; the first threshold is set so that the false acceptance rate is higher than the false rejection rate; The second threshold is set so that the false acceptance rate is lower than the false rejection rate. 10. The information processing device according to configuration 9. (Configuration 11) the first result and the second result are results of identifying the user using different machine learning models; a first machine learning model used to obtain the first result satisfies at least one of the following: a data size is smaller than a second machine learning model used to obtain the second result; and a processing time is shorter than that of the second machine learning model. 9. The information processing device according to any one of configurations 1 to 8. (Configuration 12) The image is an image of the user's eye. 12. The information processing device according to any one of configurations 1 to 11. (Configuration 13) The information acquisition device further includes a display control device for controlling the display of the information acquired by the information acquisition device. 13. The information processing device according to any one of configurations 1 to 12. (method) A method for controlling an information processing device having a first acquisition means capable of acquiring a first result in which a user is identified with a first accuracy from an image of the user, and a second acquisition means capable of acquiring a second result in which the user is identified with a second accuracy higher than the first accuracy, obtaining first information about objects in the user's vicinity based on the first result; obtaining second information about the object based on the second result; 1. A method for controlling an information processing device, comprising: (program) 14. A program for causing a computer to function as each means of the information processing device according to any one of configurations 1 to 13. (medium) 14. A computer-readable storage medium storing a program for causing a computer to function as each means of the information processing device according to any one of configurations 1 to 13. [Explanation of symbols]

[0115] 100: HMD (information processing device), 128: CPU

Claims

1. a first obtaining means for obtaining a first result of identifying the user with a first accuracy from an image of the user; a second acquisition means capable of acquiring a second result in which the user is identified from the image with a second accuracy higher than the first accuracy; and an information acquisition means for acquiring first information about an object in the vicinity of the user based on the first result, and acquiring second information about the object based on the second result.

2. The information acquisition means acquires third information about the object that can be acquired without using the first result and the second result.

2. The information processing apparatus according to claim 1, wherein:

3. The information acquisition means acquires information about the object present at the user's line of sight.

2. The information processing apparatus according to claim 1, wherein:

4. The information acquisition means acquires information about the object based on the location information of the user.

2. The information processing apparatus according to claim 1, wherein:

5. When the second information is associated with the object, the information acquisition means prompts the user to identify the object with the second accuracy.

2. The information processing apparatus according to claim 1, wherein:

6. The second acquisition means invalidates the second result after a predetermined time has elapsed since the user was identified with the second accuracy.

2. The information processing apparatus according to claim 1, wherein:

7. The information acquisition means acquires the second result before acquiring the second information for each of a plurality of objects.

2. The information processing apparatus according to claim 1, wherein:

8. The first acquisition means acquires the first result without notifying the user.

2. The information processing apparatus according to claim 1, wherein:

9. the first result and the second result are results of identifying the user using the same machine learning model; A first threshold of similarity used to identify the user with the first accuracy is set lower than a second threshold of similarity used to identify the user with the second accuracy.

2. The information processing apparatus according to claim 1, wherein:

10. the first threshold and the second threshold are set based on a false acceptance rate and a false rejection rate, which are indicators of accuracy in identifying the user; the first threshold is set so that the false acceptance rate is higher than the false rejection rate, the second threshold is set so that the false acceptance rate is lower than the false rejection rate; 10. The information processing apparatus according to claim 9,

11. the first result and the second result are results of identifying the user using different machine learning models; a first machine learning model used to obtain the first result satisfies at least one of the following: a data size is smaller than that of a second machine learning model used to obtain the second result; and a processing time is shorter than that of the second machine learning model.

2. The information processing apparatus according to claim 1, wherein:

12. The image is an image of the user's eye.

2. The information processing apparatus according to claim 1, wherein:

13. The information acquisition device further includes a display control device for controlling the display of the information acquired by the information acquisition device.

2. The information processing apparatus according to claim 1, wherein:

14. A method for controlling an information processing device having a first acquisition means capable of acquiring a first result in which a user is identified with a first accuracy from an image of the user, and a second acquisition means capable of acquiring a second result in which the user is identified with a second accuracy higher than the first accuracy, obtaining first information about objects in the user's vicinity based on the first result; obtaining second information about the object based on the second result; and 1. A method for controlling an information processing device, comprising:

15. A program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 13.

16. A computer-readable storage medium storing a program for causing a computer to function as each of the means of the information processing device according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Providing apparatus, providing method, and program

    JP2021043772A