Information processing device, information processing method, and program
The information processing device addresses accuracy and robustness issues in identifying target subjects by combining area setting, code detection, and image recognition, ensuring precise and reliable information retrieval.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- CANON KK
- Filing Date
- 2024-11-15
- Publication Date
- 2026-05-27
Smart Images

Figure 2026087258000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a technique for acquiring information on a target subject.
Background Art
[0002] In images displayed on cameras, smartphones, smart glasses, etc., the technology of robustly and highly accurately displaying information on a subject that the user is interested in not only makes the information searchable on the spot, but also increases its importance in systems such as settlement.
[0003] In Patent Document 1, a technique is disclosed in which information on an IC tag embedded in a product is read using a mobile terminal, and product information is downloaded from a server based on the information on the IC tag. In addition, a technique for recognizing a subject on an image using machine learning is generally known.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, in the method described in Patent Document 1, when a plurality of subjects are adjacent, it is necessary to bring the mobile terminal closer to identify the target subject for the user. As another method, a method of recognizing an image to acquire subject information is also known, but there is a possibility that information cannot be correctly acquired depending on the shooting angle and lighting conditions. Also, in the case of such a method of recognizing an image to acquire subject information, the recognition accuracy by learning may decrease for a subject for which it is difficult to obtain learning images.
[0006] This invention has been made in view of the above problems, and aims to provide an information processing device, an information processing method, and a program that can acquire information about a target subject with high accuracy and robustness. [Means for solving the problem]
[0007] To solve the above problems, the information processing apparatus of the present invention is characterized by comprising: area setting means for setting a setting area containing a subject in an image; first detection means for detecting a code assigned to the subject; first information acquisition means for acquiring first information corresponding to the code; second information acquisition means for acquiring second information corresponding to an image within the setting area; information matching means for matching the code with the setting area by comparing the first information with the second information; and output means for outputting at least one of the first information and the second information based on the result of the information matching means. [Effects of the Invention]
[0008] According to the present invention, detailed information about the subject can be acquired with high accuracy and robustness. [Brief explanation of the drawing]
[0009] [Figure 1] This is a block diagram of an information processing device according to the first embodiment. [Figure 2] Figure 1 is a flowchart illustrating the operation of the information processing device. [Figure 3] This is a conceptual diagram of the code. [Figure 4] Figure 1 is an explanatory diagram of the information matching unit's processing. [Figure 5] This is an explanatory diagram of the processing of the information processing device according to the second embodiment. [Figure 6] This is an explanatory diagram of the processing of the information processing device according to the third embodiment. [Figure 7] This is an explanatory diagram of the processing of the information processing device according to the fourth embodiment. [Modes for carrying out the invention]
[0010] Embodiments of the present invention will be described below with reference to the drawings.
[0011] [First Embodiment] Figure 1 is a block diagram showing an example configuration of an information processing device 100 in a first embodiment of the present invention. The information processing device 100 includes an imaging unit 110, a processing unit 120, and a display unit 130. The processing unit 120 also includes an image acquisition unit 101, a subject detection unit 102, a code detection unit 103, a code information acquisition unit 104, an image information acquisition unit 105, an information matching unit 106, a specified position acquisition unit 107, and a user interface 108. The information processing device 100 can acquire information from a server device 140 via the Internet N, which is an example of an external network.
[0012] The imaging unit 110 acquires an image of the target. The imaging unit 110 may be a digital camera, smartphone, smart glasses, etc., but is not limited to these as long as it is a device capable of acquiring images. The processing unit 120 functions as the main part of the information processing device 100. The display unit 130 functions as a display means for displaying information, and may be part of the imaging unit 110, or it may be displayed on a display separate from the imaging unit 110. The information processing device 100 only needs to be able to implement the functions of at least the blocks of the processing unit 120 described later.
[0013] The image acquisition unit 101 acquires the image acquired by the imaging unit 110. The subject detection unit 102 detects the regions containing each subject in the image (hereinafter referred to as subject regions) as rectangular frames based on the image acquired by the image acquisition unit 101. Methods for detecting subject regions include machine learning-based methods. Specifically, convolutional neural networks and Transformers can be used as subject detection means, but the method is not limited to machine learning as long as it can detect subjects. Also, the frames of the subject regions do not have to be rectangular; for example, segmentation may be performed for each subject.
[0014] That is, the subject detection unit 102 functions as a second detection means for detecting a subject area. Further, the subject detection unit 102 functions as an area setting means for setting the subject area as a setting area.
[0015] The code detection unit 103 detects a code physically attached to the subject. Details of the code will be described later. That is, the code detection unit 103 functions as a first detection means for detecting the code attached to the subject.
[0016] The code information acquisition unit 104 acquires information about the subject included in the code itself detected by the code detection unit 103. Alternatively, the code information acquisition unit 104 acquires information about the subject from the server device 140 via the Internet N based on information such as a link destination included in the code. That is, the code information acquisition unit 104 functions as a first information acquisition means for acquiring information (code information) corresponding to the code.
[0017] Based on the subject area acquired by the subject detection unit 102, the image information acquisition unit 105 recognizes an image of the subject area and acquires information about the subject as a recognition result. Alternatively, based on the result of recognizing the image of the subject area, the image information acquisition unit 105 acquires information about the subject from the server device 140 via the Internet N. That is, the image information acquisition unit 105 functions as a second information acquisition means for acquiring information (image information) corresponding to the image within the setting area.
[0018] Based on the result of collating the information acquired by the code information acquisition unit 104 and the image information acquisition unit 105, the information collation unit 106 performs association between the detected code and the subject area. Details of the association will be described later. That is, the information collation unit 106 functions as an information collation means for collating the first information and the second information to associate the code with the setting area.
[0019] The designated position acquisition unit 107 acquires the designated position of the user within the image. In this embodiment, it is assumed that the display unit 130 has a touch panel function, and user input is performed via the touch panel. The designated position acquisition unit 107 acquires information on the position where a finger or the like on the touch panel touches. That is, the designated position acquisition unit 107 functions as designated position acquisition means for acquiring the position within the image designated by the user.
[0020] In the case where the imaging unit 110 is a smart glass or the like, a sensor for detecting and acquiring the line-of-sight position of the user on the image may be provided, and the designated position of the user may be acquired based on the line of sight. That is, the sensor may function as line-of-sight detection means for detecting the line of sight of the user, and the position may be acquired based on the detected line of sight. That is, the method for acquiring the designated position of the user is not limited to the above, as long as it is a method capable of specifying the subject position of interest to the user.
[0021] When the subjects are close to each other, there may be a plurality of codes in the vicinity of the designated position of the user, and it may be difficult to determine which code the user has designated. Therefore, as will be described later, by associating the code with the subject area by the information collation unit 106, the user can visually and easily specify the target subject.
[0022] The user interface 108 specifies the subject area closest to the designated position of the user, and outputs and displays either or both of the code information and the image information associated by the information collation unit 106 to the display unit 130. That is, the user interface 108 functions as output means for outputting at least one of the first information and the second information based on the result of the information collation means.
[0023] When displaying on the display unit 130, there are variations such as enlarged display of the information corresponding to the designated position of the user, display with a changed color, and change in the thickness of the characters. Note that as long as the information of the designated position of the user can be displayed on the display unit 130 in a form that can be differentiated from other information, it is not limited to the above method.
[0024] The above example describes a scenario where the user first specifies a location, but the display unit 130 may first display the subject area and its corresponding information, allowing the user to select the display location. Alternatively, only the subject area (setting area) may be displayed first, followed by the display of information about the subject area selected by the user. In other words, the user interface 108 may display information associated with the subject area corresponding to the user's specified location. Furthermore, if the detected subject area is below a predetermined threshold, only the subject area may be displayed first. This prevents the user from specifying a location that cannot be selected, such as a subject for which no subject area could be detected, by informing the user in advance that the selection range is limited.
[0025] (Processing flow) The processing flow in this embodiment will be explained with reference to Figure 2. Figure 2 is a flowchart showing an example of processing (information processing method).
[0026] First, in step 201, the image acquisition unit 101 acquires an image from the imaging unit 110. Next, in step 202, the subject detection unit 102 detects the subject area from the image acquired by the image acquisition unit 101. Next, in step 203, the code detection unit 103 detects the codes contained in the image. Next, in step 204, the code detection unit 103 acquires the code information contained in the codes detected in step 203 and accesses the server device 140 to download the information based on the acquired code information. Alternatively, the code information itself may be output without accessing the server device 140.
[0027] Next, in step 205, the image information acquisition unit 105 performs image recognition within the subject area detected by the subject detection unit 102, and accesses the server device 140 to download information (image information) based on the recognition result. Alternatively, the image recognition result itself may be output as image information without accessing the server device 140.
[0028] Furthermore, steps 202 through 205 are not limited to being executed in the order described above. Step 202 must be executed before step 205, and step 203 must be executed before step 204, but steps 202 and 203 may be executed in parallel, and steps 204 and 205 may be executed in parallel.
[0029] Next, in step 206, the information matching unit 106 matches the code information acquired by the code information acquisition unit 104 with the image information acquired by the image information acquisition unit 105 to associate the subject area with the code. Next, in step 207, the user interface 108 displays the information for each subject on the display unit 130 based on the results of the information matching unit 106. For codes and subject areas that were not associated in step 206, the information may be displayed based on the code information alone (code information) or the subject area information alone (image information).
[0030] Next, in step 208, the designated position acquisition unit 107 acquires the user's designated position. Then, in step 209, the user interface 108 highlights information and the detection frame of the subject closest to the user's designated position on the display unit 130 for the user to see. In other words, the user interface 108 functions as a display control means that displays at least one of the first information and the second information on the display means based on the position acquired by the designated position acquisition means.
[0031] According to the flowchart in Figure 2, step 202 functions as the region setting step, step 203 as the first detection step, and step 204 as the first information acquisition step. Furthermore, step 205 functions as the second information acquisition step, step 206 as the information matching step, and step 207 as the output step. Alternatively, each step in the flowchart in Figure 2 may be configured as a program executed by the CPU.
[0032] Note that in Figure 2, information about each subject (code information, image information) is displayed before the user makes a selection. However, information about nearby subjects may be displayed after the user has selected a location.
[0033] (code) Next, the code will be explained with reference to Figure 3. Figure 3 shows a conceptual diagram of the code. 301 and 302 are examples of two-dimensional codes, with 301 being an example of a QR code (registered trademark) and 302 being an example of a barcode. 301 and 302 are recognizable (readable) in visible light, but similar patterns recognizable (readable) in invisible light may be embedded in the subject to avoid damaging its appearance. In that case, the imaging unit 110 shall be equipped with an image sensor capable of detecting invisible light. If the code is small, the image may be divided into regions, and the region may be enlarged before detecting the code.
[0034] The code may be a two-dimensional code like 301 or 302, or a three-dimensional pattern. It may also be an object that can be attached to or detached from the object, such as an RF tag (IC tag), and which emits electromagnetic waves itself. Alternatively, it may be in the form of an RFID (Radio Frequency Identification) system that emits electromagnetic waves from the information processing device 100 and operates based on the energy of those electromagnetic waves. Furthermore, two or more of these may be combined. In other words, the code in this embodiment is any medium that is physically attached to the object and contains information that can identify the object, or information that can link to such information.
[0035] The number of codes assigned to a subject is not limited to one; as will be explained later, multiple codes may be assigned to the same subject to account for cases where subjects overlap and hide each other. Furthermore, the code may be embedded not in the subject itself, but in tags attached to the subject, such as the RF tags mentioned above.
[0036] (Information Verification Unit) Next, the operation of the information matching unit 106 will be explained with reference to Figure 4. The advantage of using codes compared to image recognition is that if a code can be detected, information can be reliably obtained. However, when subjects are close together, such as in a street or inside a store, it becomes necessary to associate the code with which subject. For this purpose, the information matching unit 106 associates the code with the subject area.
[0037] Figure 4 schematically illustrates a portion of the situation where subjects are in close proximity. In reality, there are often other subjects besides those shown, but they are omitted for illustrative purposes. In Figure 4, 411 and 412 represent the subject areas corresponding to subjects 401 and 402, respectively. Subjects 401 and 402 are also assigned codes 421 and 422, respectively.
[0038] Since subject area 412 also partially includes code 421, which corresponds to subject 401, it is necessary to accurately associate each subject area with its corresponding code. Table 430 shows subject areas (regions) and codes in a matrix format, with each cell in the table containing image information within the region on the left and code information on the right. For example, the cell where "region 411" and "code 422" intersect indicates that the image information is a "cube" and the code information for code 422 is a "cylinder".
[0039] Table 430 simplifies the process by only listing the shape of the subject, but it is also acceptable to include multiple pieces of information about the subject's characteristics, such as color, material, descriptive text, images, and shape. When matching information, methods such as converting the information into feature vectors and then measuring the distance between feature vectors may be used. The information matching unit 106 matches the image information within the subject area with the code information. In the case of Table 430, the combination of the top-left and bottom-right cells that match is adopted to associate the subject area with the code. That is, the top-left cell combination, subject area 411 and code 421, are associated, and the bottom-right cell combination, subject area 412 and code 422, are associated.
[0040] The above describes examples where it is possible to associate a code with a subject area. However, if, for example, the code within the subject area is hidden and cannot be detected, or if the subject cannot be recognized in the image, the association may not be performed, and the code or the subject area may be treated separately.
[0041] In this way, by using image recognition in conjunction with codes embedded in the subject, information about the subject of interest to the user can be obtained with high accuracy and robustness.
[0042] [Second Embodiment] This embodiment describes a case where code information is used to assist in image-based detection, recognition, and display. Note that the same configuration as in the first embodiment will be omitted from the description, and the following description will focus on the differences from the first embodiment.
[0043] In Figure 4, since part of subject 402 is hidden by subject 401, the accuracy of image detection and recognition may decrease. In this case, if code 422 is detected, the code information acquisition unit 104 can acquire information about subject 402, and the subject detection unit 102 and the image information acquisition unit 105 can use the code information of code 422 to perform image recognition and image information acquisition. For example, if information such as the shape and color of subject 402 is embedded in code 422, it is possible to improve the accuracy of subject detection and recognition by using that information. In other words, the subject detection unit 102 may detect the subject area based on the first information.
[0044] Furthermore, if code 422 includes a link to image information or 3D shape information of the entire subject 402, the information may be acquired based on that link and displayed on the display unit 130. An example of this is explained using Figure 5. In Figure 5(A), 501 is the user's designated position acquired by the designated position acquisition unit 107. 502 is the closest subject detection frame highlighted based on the position of 501. In this way, it is possible to clearly display to the user which subject the user has designated.
[0045] Figure 5(B) shows an example where, as in Figure 5(A), an image (appearance information) of the subject 402 is downloaded from, for example, the server device 140 based on the information of code 422 at a user-specified location and superimposed on the subject 402. 503 is the downloaded image, and even the hidden parts of the subject 402 are visible. Note that even without user specification, the image acquisition unit 101 may compare the subject information on the image acquired by the image acquisition unit 101 with the information downloaded by the image information acquisition unit 105 and perform an occlusion determination. If occlusion is determined, the user interface 108 may automatically display the superimposed image.
[0046] Furthermore, the information downloaded is not limited to images of the subject (external appearance information); it may also include and display the 3D shape information of the subject, and this information may be displayed in a different position rather than superimposed on subject 402. In this way, the user can obtain information about the overall appearance of the hidden subject 402.
[0047] The above describes an example of using code information to assist in image-based detection, recognition, and display. Conversely, the results of image-based subject recognition and detection may also be used as auxiliary information for code detection. For example, to facilitate code detection, the image may be cropped to the size of the subject detection frame, enlarged, and then the code may be detected. In other words, the code detection unit 103 may detect the code based on the second information.
[0048] According to this embodiment, the accuracy of subject detection and recognition can be improved by using code information for subject detection and recognition. Furthermore, by displaying appearance information downloaded based on the code, it becomes possible to display the entire image of a hidden subject.
[0049] [Third Embodiment] This embodiment describes a method for prioritizing information when multiple codes are assigned to the same subject. Note that the same configuration as in the first embodiment will be omitted from the description, and the following description will focus on the differences from the first embodiment.
[0050] In Figure 6, 601 is a tag assigned to subject 401, and the tag 601 has code 602 embedded within it. In other words, in Figure 6, subject 401 is assigned code 602 in addition to code 421 which is embedded within it.
[0051] In this embodiment, the code includes a link to detailed information and a category (information indicating the category) related to the linked information. Table 603 shows examples of categories (information categories) embedded in the code. In Table 603, code 421 shows a case where information representing the characteristics of the subject, such as shape, color, and material, is embedded as the category, and code 602 shows a case where price information of the subject is embedded as the category, but the combination is not limited to the above.
[0052] The code information acquisition unit 104 outputs the code information to the information matching unit 106, including, for example, the information shown in Table 603. The information matching unit 106 assigns priorities based on the content of the acquired code information, and the user interface 108 displays the information with the highest priority obtained from the information matching unit 106 on the display unit 130 first. Alternatively, the user interface 108 may display the information with the highest priority in a way that is easily noticeable to the user.
[0053] In other words, the information matching unit 106 functions as an information matching means that assigns priority to the acquired first information when two or more codes are associated with a specific setting area. The user interface 108 then outputs the first information based on the priority and displays it on the display means.
[0054] In this embodiment, a rule-based method is used to prioritize information, which prioritizes the information of the tag 601 associated with the subject 401 over the information of the subject 401. The reason for this is that the information of the subject 401 is general-purpose and rarely changes, whereas tag information is often attached under specific circumstances such as in a store, and therefore often contains information that should be prioritized.
[0055] Furthermore, when prioritizing code information, the category of the code information or the user's own location information may be used. Location information can be obtained, for example, by having a means of acquiring location information such as a GPS receiver. For example, if it is determined that the user is in the store, the priority of code information belonging to product-related categories such as price is increased. In other words, the information matching unit 106 assigns priority based on location information.
[0056] Alternatively, the system could retain the categories of code information selected by the user and learn their preferred categories. This would allow the system to prioritize the user's preferred categories, thus enabling the presentation of information tailored to individual preferences. Prioritization information could also be included within the code information itself.
[0057] According to this embodiment, even when multiple codes are assigned to the same subject, it becomes possible to display appropriate information depending on the situation.
[0058] [Fourth Embodiment] In this embodiment, we will explain how to improve image recognition accuracy by using code information that could not be associated with the subject area to perform learning. Note that the same configuration as in the first embodiment will be omitted from the explanation, and the following will focus on the differences from the first embodiment.
[0059] Image recognition accuracy may decrease if the environment in which the images were used for training differs from the environment in which the recognition actually takes place. For example, this could occur if the subject was trained using images taken during the day, but then photographed again at dusk. In this case, even if a code is detected, the subject region may not be matched. A specific example is illustrated in Figure 7.
[0060] Figure 7 shows a situation where subject region 412 and code 422 are associated, but subject 401 is not recognized in the image and code 421 is not associated with any subject region. In this case, a certain region containing code 421 is extracted from the image, and the extracted image is used as input data, while the code information obtained from code 421 is used as the ground truth data to train the machine learning model. This makes it possible to retrain the model for subject 401, which failed to be recognized.
[0061] For example, in the configuration shown in Figure 1, the image information acquisition unit 105 has a machine learning-type image recognition function. The image acquisition unit 101 or the subject detection unit 102 extracts a certain region containing codes that have not been matched by the information matching unit 106. The image information acquisition unit 105 then inputs a training dataset consisting of the extracted image data and code information to train a machine learning model.
[0062] Based on the above, according to this embodiment, it is possible to improve the recognition rate by performing retraining using the code information for subjects that failed to be recognized.
[0063] The above embodiments can be modified in various ways. Furthermore, although the first to fourth embodiments were described only when the subject is an object, the methods can also be applied when a code is assigned to a person or animal. For example, in a sports scene, a tag can be embedded in a player's uniform to display the player's information.
[0064] The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.
[0065] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of its gist. Embodiments of this disclosure include the following configurations, methods, and programs.
[0066] [Configuration 1] A region setting means for setting a region that includes a subject within an image, A first detection means for detecting a code assigned to the subject, A first information acquisition means for acquiring first information corresponding to the aforementioned code, A second information acquisition means for acquiring second information corresponding to the image within the aforementioned setting area, Information matching means for matching the code with the setting area by comparing the first information and the second information, Based on the results of the information matching means, an output means outputs at least one of the first information and the second information, An information processing device characterized by having the following features. [Configuration 2] The first information is information contained in the code, or information obtained from an external network based on the code. The second piece of information is an image of the setting area, or information obtained from an external network based on an image within the setting area. The information processing device according to configuration 1, characterized by the above. [Configuration 3] The information processing device according to configuration 1 or 2, characterized in that the code is one or more of the following: two-dimensional information or three-dimensional information that can be read using visible light or invisible light, and a tag that emits electromagnetic waves. [Structure 4] The system further includes a second detection means for detecting a subject region, which is a region containing the subject in the aforementioned image. The information processing apparatus according to any one of configurations 1 to 3, characterized in that the area setting means sets the subject area detected by the second detection means to the setting area. [Composition 5] A specified position acquisition means for acquiring a position within the image specified by the user, It further comprises a means for displaying information, The output means includes a display control means that displays at least one of the first information and the second information on the display means based on the position acquired by the designated position acquisition means. An information processing device according to any one of configurations 1 to 4, characterized by the above. [Composition 6] The display control means displays area information indicating the set area on the display means, The information processing device according to configuration 5, characterized in that it subsequently displays the first information and the second information associated with the setting region indicated by the region information corresponding to the position acquired by the designated position acquisition means. [Composition 7] The information processing device according to configuration 6, characterized in that the display control means displays the area information when the number of set areas is less than or equal to a predetermined threshold. [Structure 8] The information processing apparatus according to configuration 5 or 6, wherein the position detection means includes a gaze detection means for detecting the user's gaze, and the position is acquired based on the detected gaze. [Composition 9] The information processing apparatus according to any one of configurations 1 to 8, characterized in that the first detection means detects the code based on the second information. [Configuration 10] The information processing apparatus according to configuration 4, characterized in that the second detection means detects the subject area based on the first information. [Composition 11] The information processing apparatus according to any one of configurations 1 to 10, characterized in that the first information acquisition means acquires at least one of the appearance information of the subject and the shape information of the subject as the first information. [Composition 12] The information matching means assigns priority to the acquired first information when two or more codes are associated with a specific setting area. The information processing apparatus according to any one of configurations 1 to 11, characterized in that the output means outputs the first information based on the priority order. [Composition 13] The system further includes location information acquisition means for acquiring location information of the information processing device, The information processing apparatus according to configuration 12, characterized in that the information matching means assigns the priority order based on the location information. [Composition 14] The first information includes information indicating the category of information relating to the subject to which the code has been assigned, The information processing apparatus according to configuration 12 or 13, characterized in that the information matching means assigns the priority based on the information indicating the category. [Composition 15] The second information acquisition means includes a machine learning type image recognition means, The image recognition means takes the image and the first information as input to train a machine learning model. An information processing device according to any one of configurations 1 to 14, characterized by the above. [Method 1] An information processing method performed by an information processing device, A region setting step in which a setting area containing the subject is set within the image, A first detection step involves detecting a code assigned to the subject, A first information acquisition step for acquiring first information corresponding to the aforementioned code, A second information acquisition step involves acquiring second information corresponding to the image information within the aforementioned setting area, An information matching step of matching the first information and the second information to associate the code with the setting area, An output step which outputs at least one of the first information and the second information based on the results of the information matching means, An information processing method characterized by having the following features. [program] A program characterized by causing a computer to execute the information processing method described in Method 1. [Explanation of Symbols]
[0067] 100 Information Processing Devices 101 Image acquisition unit 102 Subject detection unit (area setting means, second detection means) 103 Code detection unit (first detection means) 104 Code Information Acquisition Unit (First Information Acquisition Means) 105 Image information acquisition unit (second information acquisition means) 106 Information Verification Unit (Information Verification Means) 107 Specified position acquisition unit (specified position acquisition means) 108 User interface (output means, display control means) 110 Imaging Unit 120 Processing Unit 130 Display section (display means)
Claims
1. A region setting means for setting a region that includes a subject within an image, A first detection means for detecting a code assigned to the subject, A first information acquisition means for acquiring first information corresponding to the aforementioned code, A second information acquisition means for acquiring second information corresponding to the image within the aforementioned setting area, Information matching means for matching the code with the setting area by comparing the first information and the second information, Based on the results of the information matching means, an output means outputs at least one of the first information and the second information, An information processing device characterized by having the following features.
2. The first information is information contained in the code, or information obtained from an external network based on the code. The second piece of information is an image of the setting area, or information obtained from an external network based on an image within the setting area. The information processing apparatus according to feature 1.
3. The information processing apparatus according to claim 1, characterized in that the code is one or more of the following: two-dimensional information or three-dimensional information that can be read using visible light or invisible light, and a tag that emits electromagnetic waves.
4. The system further includes a second detection means for detecting a subject region, which is a region containing the subject in the aforementioned image. The information processing apparatus according to claim 1, characterized in that the area setting means sets the subject area detected by the second detection means to the setting area.
5. A specified position acquisition means for acquiring a position within the image specified by the user, It further comprises a means for displaying information, The output means includes a display control means that displays at least one of the first information and the second information on the display means based on the position acquired by the designated position acquisition means. The information processing apparatus according to feature 1.
6. The information processing apparatus according to claim 5, characterized in that the display control means displays region information indicating the set region on the display means, and then displays the first information and the second information associated with the set region indicated by the region information corresponding to the position acquired by the designated position acquisition means.
7. The information processing apparatus according to claim 6, characterized in that the display control means displays the area information when the number of set areas is less than or equal to a predetermined threshold.
8. The information processing apparatus according to claim 5, wherein the position detection means includes a gaze detection means for detecting the user's gaze, and the position is acquired based on the detected gaze.
9. The information processing apparatus according to claim 1, characterized in that the first detection means detects the code based on the second information.
10. The information processing apparatus according to claim 4, characterized in that the second detection means detects the subject area based on the first information.
11. The information processing apparatus according to claim 1, characterized in that the first information acquisition means acquires at least one of the appearance information of the subject and the shape information of the subject as first information.
12. The information matching means assigns priority to the acquired first information when two or more codes are associated with a specific setting area. The information processing apparatus according to claim 1, characterized in that the output means outputs the first information based on the priority order.
13. The system further includes location information acquisition means for acquiring location information of the information processing device, The information processing apparatus according to claim 12, characterized in that the information matching means assigns the priority order based on the location information.
14. The first information includes information indicating the category of information relating to the subject to which the code has been assigned, The information processing apparatus according to claim 12, characterized in that the information matching means assigns the priority order based on the information indicating the category.
15. The second information acquisition means includes a machine learning type image recognition means, The image recognition means takes the image and the first information as input to train a machine learning model. The information processing apparatus according to feature 1.
16. An information processing method performed by an information processing device, A region setting step in which a setting area containing the subject is set within the image, A first detection step for detecting a code assigned to the subject, A first information acquisition step for acquiring first information corresponding to the aforementioned code, A second information acquisition step involves acquiring second information corresponding to the image information within the aforementioned setting area, An information matching step of matching the first information and the second information to associate the code with the setting area, An output step which outputs at least one of the first information and the second information based on the results of the information matching means, An information processing method characterized by having the following features.
17. A program characterized by causing a computer to execute the information processing method described in claim 16.