Information processing device, information processing method, and program

By displaying feature points of input images alongside correct data and inference results, the mechanism addresses the lack of evaluation in existing technologies, enabling improved accuracy and reliability in skeletal information estimation.

JP7807693B1Active Publication Date: 2026-01-28CANON MARKETING JAPAN INC +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2024230551
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2024-12-26
Publication Date
2026-01-28
Estimated Expiration
2044-12-26

AI Technical Summary

Technical Problem

Existing technologies for estimating skeletal information lack a mechanism for evaluating the performance of skeletal information estimation techniques, as evidenced by Patent Document 1's failure to disclose evaluation details.

Method used

A mechanism is provided for evaluating skeletal information estimation techniques by displaying feature points of an input image using a trained model in association with both correct data and inference results, allowing users to compare and assess accuracy and reliability.

Benefits of technology

Enables accurate evaluation of skeletal information estimation techniques, enhancing the ability to improve performance by identifying discrepancies and optimizing training data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807693000001_ABST
    Figure 0007807693000001_ABST
Patent Text Reader

Abstract

To provide a technology that makes it possible to evaluate technology for estimating skeletal information, etc. [Solution] An input image to be subjected to inference processing by a trained model is input, and control is exercised to display the results of inference processing on an object related to the input image by the trained model that has been made to learn feature points of the object. In a first setting, control is exercised to display feature points of the object related to the input image identified by inference processing by the trained model in association with each other, and in a second setting, control is exercised to display feature points of the object related to the input image identified by inference processing by the trained model in association with feature points related to the ground truth data related to the object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system, a control method thereof, and a program. [Background technology]

[0002] It is known that there are techniques for estimating joint positions and bone structure from image data.

[0003] Patent Document 1 discloses that skeletal information (small skeletal information) of each finger is extracted from an arbitrary hand region image based on its reference points and image features using a finger prediction model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Patent Publication No. 2020-198019 DISCLOSURE OF THE INVENTION [Problem to be solved by the invention]

[0005] To accurately grasp skeletal information, a high-performance estimation technology is required, but as a prerequisite, a mechanism for users to grasp (evaluate) the performance of the estimation technology is necessary. Patent Document 1 discloses outputting extracted skeletal information, but does not disclose details related to evaluation.

[0006] Therefore, an object of the present invention is to provide a mechanism for evaluating techniques for estimating skeletal information and the like. [Means for solving the problem]

[0007] The device comprises an input means for inputting an input image that is the subject of inference processing using a trained model, and a display control means for controlling the display of the results of inference processing for an object related to the input image using the trained model that has learned the feature points of the object, wherein the display control means controls in a first setting to display the feature points of the object related to the input image identified by the inference processing using the trained model in association with each other, and in a second setting to display the feature points of the object related to the input image identified by the inference processing using the trained model in association with the feature points related to the correct data for the object. [Effects of the Invention]

[0008] According to the present invention, it is possible to provide a mechanism for evaluating techniques for estimating skeletal information and the like. [Brief explanation of the drawings]

[0009] [Figure 1] FIG. 1 is a diagram illustrating an example of the configuration of a posture estimation system according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing an example of a hardware configuration of a client terminal 101 according to an embodiment of the present invention. [Figure 3] 10 is a flowchart illustrating an example of a process for switching the display of an inference result in an embodiment of the present invention. [Figure 4] FIG. 10 is a diagram showing an example of a user's operation screen for displaying an inference result in an embodiment of the present invention. [Figure 5] FIG. 10 is a diagram showing an example of a screen displaying image data in the embodiment of the present invention. [Figure 6] FIG. 10 is a diagram illustrating an example of an image of a trained model storage folder in an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0011] FIG. 1 is a diagram illustrating an example of the configuration of a posture estimation system according to an embodiment of the present invention.

[0012] A client terminal 101 and a server 102 are connected via a network 100 so as to be able to communicate with each other.

[0013] The client terminal 101 may be any device that incorporates the functions shown in FIG. 2, and may be, for example, a personal computer (hereinafter referred to as a PC), a mobile terminal such as a smartphone, or a tablet terminal.

[0014] The network 100 can take the form of a wired LAN, a wireless LAN, a USB, or the like, depending on the physical interface that the server 102 has.

[0015] The server 102 can store trained models, setting files for storing setting parameters used during training, dataset information, inference results for inference datasets, etc. The trained models and the like that can be stored in the server 102 described above may be stored in the ROM 202 or external memory 211 of the client terminal 101.

[0016] FIG. 2 is a block diagram showing an example of the hardware configuration of the client terminal 101 according to the embodiment of the present invention.

[0017] 2 is a block diagram showing an example of the hardware configuration of a client terminal 101 (a client terminal is an example of an information processing device) according to an embodiment of the present invention. The file server 102 also has a similar configuration.

[0018] As shown in FIG. 2, each information processing device is connected to a CPU (Central Processing Unit) 201, a ROM (Read Only Memory) 202, a RAM (Random Access Memory) 203, an input controller 205, a video controller 206, a memory controller 207, and a communication I / F controller 208 via a system bus 204.

[0019] The CPU 201 comprehensively controls each device and controller connected to the system bus 204 .

[0020] ROM202 or external memory 211 stores the BIOS (Basic Input / Output System) and OS (Operating System), which are control programs executed by CPU201, computer-readable and executable programs for realizing this information processing method, and various necessary data (including data tables).

[0021] The RAM 203 functions as a main memory, a work area, etc. for the CPU 201. The CPU 201 loads programs and the like required for executing processing from the ROM 202 or the external memory 211 into the RAM 203, and executes the loaded programs to realize various operations.

[0022] The input controller 205 controls input from input devices such as a keyboard 209 and pointing devices such as a mouse, touchpad, etc. (not shown). If the input device is a touch panel, the user can issue various instructions by pressing (touching with a finger, etc.) icons, cursors, or buttons displayed on the touch panel.

[0023] The touch panel may also be a touch panel capable of detecting positions touched by multiple fingers, such as a multi-touch screen.

[0024] The video controller 206 controls the display on an external output device such as a display 210. The display also includes the display of a notebook PC integrated with the main body. Note that the external output device is not limited to a display, and may be, for example, a projector. In addition, for the device capable of receiving the above-mentioned touch operation, an input device is also provided.

[0025] In the explanation using the flowcharts below, the display destination when displaying is assumed to be the display 210 unless otherwise specified.

[0026] The video controller 206 can control a video memory (VRAM) for display control, and can use part of the RAM 203 as a video memory area, or can provide a separate dedicated video memory.

[0027] The memory controller 207 controls access to the external memory 211. The external memory may be an external storage device (hard disk) that stores a boot program, various applications, font data, user files, edited files, and various data, a flexible disk (FD), or a CompactFlash (registered trademark) memory connected to a PCMCIA card slot via an adapter.

[0028] The communication I / F controller 208 connects to and communicates with external devices via a network 214 (for example, the network 101 shown in FIG. 1), and executes communication control processing on the network 214. For example, communication using TCP / IP, telephone lines such as ISDN, and 3G, 4G, and 5G lines of mobile phones are possible.

[0029] The CPU 201 enables display on the display 210 by, for example, executing a process of expanding (rasterizing) an outline font into a display information area in the RAM 203. The CPU 201 also enables user instructions using a mouse cursor (not shown) or the like on the display 210.

[0030] Before explaining the flowchart in Figure 3, an example of an image diagram of a trained model storage folder will be explained using Figure 6.

[0031] Trained model storage files 602, which are divided into files for each job ID, are stored in a trained model storage folder 601. In FIG. 6, trained model storage files for job ID 1 and job ID 2 are stored.

[0032] The trained model storage file 602 stores a trained model 603, a setting file 604, dataset information 605, and an inference result 606.

[0033] The trained model 603 is an AI model trained using the training dataset stored in the dataset information 605.

[0034] The setting file 604 stores the parameters set during learning as a file.

[0035] The dataset information 605 is a file in which a training dataset and an inference dataset are stored, and an image folder 607 and an annotation folder 609 are stored.

[0036] Image data 608 used for learning or inference is stored in the image folder 607. In Fig. 6, the image data is stored in JPG format, but it may be stored in other formats such as PNG format. The file name of the image data 608 is given a name corresponding to the file name of the annotation file 610, and in Fig. 6 it is named "IMG_001".

[0037] The annotation folder 609 stores an annotation file 610. The file name of the annotation file 610 corresponds to the file name of the image data 608, and is named "IMG_001" in FIG.

[0038] In this embodiment, feature points (key points) refer to points that are characteristic in understanding the structure and shape of an object to which the present invention is applied, and in the case of a human body, for example, these refer to the joints in the skeleton, the main points that form the skeleton (such as the head, waist, toes, and fingertips, which are not the positions of specific joints but are necessary for identifying the skeleton), and organs such as the eyes, nose, and ears.

[0039] The annotation process refers to the process of assigning the positions of key points to learning images.

[0040] Therefore, the key points of the correct answer data described in this embodiment are key points added by annotation work when creating learning data, and the key points of the inference result are key points output as a result of the trained model performing inference processing on the input image. The screen viewed by the user (image data display unit 418) displays the key points and an image in which the key points are connected by lines. How the lines are connected will be described later with reference to FIG. 5.

[0041] In this embodiment, an example will be described in which finger joints are inferred as feature points from an image of a hand, but the application is not limited to the hand and may be any part of the body, or even the entire body. Furthermore, the application is not limited to the human body, and may also be to a non-human organism. Furthermore, the present invention may also be applied to a technology for inferring moving parts of tools, robots, etc.

[0042] The annotation file stores key point information of the correct answer data for the corresponding image data 608. That is, it stores predetermined position information, label names, etc. related to the correct answer data annotated to the image data 608. For example, in the case of a joint related to the right wrist, it stores information such as coordinate information related to the correct answer data placed at the joint and a label name indicating that the joint is a joint related to the right wrist.

[0043] The result of inference performed using the inference dataset is stored in the inference result 606. Specifically, information on inference accuracy (information displayed in the inference data accuracy display unit 411), information on the reliability of the inference result (information displayed in the score display unit 415), and key point information of the inference result are stored.

[0044] In the learning function of this system, the user trains the AI ​​model with the necessary information, and the trained AI model (trained model 603) and learning outcomes (configuration file 604, dataset 605) are managed with an identifier called a job ID. The trained model 603 performs inference for each job ID based on the learning outcomes, and generates inference results 606. The inference results 606 are also managed linked to the job ID.

[0045] When inference using the trained model has been completed and the client terminal 101 receives an instruction to start processing, the processing in FIG. 3 starts.

[0046] Next, the processing executed by the client terminal 101 in the embodiment of the present invention will be described with reference to FIGS.

[0047] The process executed by the posture estimation system according to the embodiment of the present invention will be described with reference to the flowchart of FIG.

[0048] 3 is a flowchart showing an example of processing for switching the display of inference results, in which the CPU 201 of the client terminal 101 reads and executes a predetermined control program. The processing of each step is executed by the CPU 201 of each device.

[0049] This posture estimation system is equipped with learning and inference functions using a trained model. Figure 3 is a flowchart showing the processing that is executed when an instruction to start the system is received after AI-based learning and inference have been completed.

[0050] In S301, the CPU 201 accepts a job ID selection from the user. In S302 to S304, the CPU 201 acquires the images, annotations, and inference results 606 corresponding to the job ID for which the selection has been accepted.

[0051] In S302, the CPU 201 acquires image data 608 in the dataset used for inference by the trained model. The image data 608 is stored in the server 102 in association with the job ID selected in S301. The CPU 201 displays the acquired image data on the display 210. In this embodiment, the CPU 201 displays a preview of the acquired image data in the inference data preview display section 410 of the verification data list 4a.

[0052] In S303, the CPU 201 acquires annotations in the dataset used for inference by the trained model. The annotations are stored as annotation files 610 in the annotation folder 609 and are managed in association with the image data used for inference by the trained model. Specifically, the CPU 201 acquires the correct answer data and inference results (inferred keypoints) for each image data acquired in S302.

[0053] In S304, the CPU 201 acquires the inference result in the dataset used for inference by the trained model. The inference result 606 is managed in association with the job ID, and stores information on the inference accuracy (information displayed in the inference data accuracy display unit 411), information on the reliability of the inference result (information displayed in the score display unit 415), and key point information of the inference result. The CPU 201 displays the inference result on the screen in association with the image data acquired in S302.

[0054] In S305, the CPU 201 accepts a selection of image data to be displayed on the image data display unit 418. The CPU 201 displays the image data for which the selection has been accepted. Specifically, the CPU 201 displays, on the image data display unit 418, the image data for which the selection has been accepted from the user, from among the image data preview-displayed on the inference data preview display unit 410.

[0055] In S306, CPU 201 accepts a selection from the user regarding key points to be displayed superimposed on the image data acquired in S305. Specifically, the selection is accepted as to whether to display one of the key points of the supervised data acquired in S303 and the key points of the inference result acquired in S304 (first setting), or whether to display both the key points of the inference result and the supervised data (second setting). If a selection is accepted for both the inference result and the supervised data, the process proceeds to S307, and if a selection is accepted for either the inference result or the supervised data, the process proceeds to S308.

[0056] In S307, the CPU 201 connects the key points of the inference result and the correct answer data with lines. The processing result of S307 will be described later using a specific example in FIG.

[0057] In S308, when the CPU 201 receives a selection of an inference result from the user, it connects the key points of the inference result with a line, and when the CPU 201 receives a selection of supervised data, it connects the key points of the supervised data with a line. The processing result of S308 will be described later using a specific example in FIG. 5.

[0058] In S309, the CPU 201 displays the keypoint information on the screen. Specifically, the result of connecting the keypoints with lines in S307 or S308 is displayed on the image data display unit 418, and the comparison result between the keypoint information related to the inference result acquired in S303 and the keypoint information related to the correct answer data is displayed in the keypoint information (4b).

[0059] The details of the processing in S309 will be explained with reference to FIG.

[0060] Figure 4 is an example of a screen displaying inference results, etc. Note that Figure 4 shows only a portion of the screen for the purpose of explanation, and the verification data list 4a and key point information 4b include data that is not shown in Figure 4.

[0061] The screen is composed of a screen switching section 401, a display switching section 402, a key point size changing section 403, a line color changing section 404, a line display switching section 405, a job ID selection section 406, a job description section 407, a tag 408, a verification data list 4a, key point information 4b, and an image data display section 418.

[0062] First, a screen switching section 401, a display switching section 402, a key point size changing section 403, a line color changing section 404, and a line display switching section 405 on the left side of the screen will be described.

[0063] The screen switching unit 401 is a part that can switch to either a screen for executing the learning function or a screen for executing the analysis (inference) function. When a pull-down menu is pressed, options of "learning" and "analysis" are displayed. When a selection of one of the options is selected, the screen switches according to the selected option. In FIG. 4, "analysis" is selected, and a screen for executing the analysis (inference) function is displayed. When "learning" is selected, a screen for executing the learning function is displayed. In this embodiment, only the case where "analysis" is selected will be described.

[0064] The display switching unit 402 is a part that can switch the display of key points superimposed on the image data display unit 418 (S306). Specifically, it is possible to switch between displaying key points from the correct answer data, displaying key points from the inference result, displaying both key points from the correct answer data and inferred key points, or displaying none of the key points. Depending on the key point display switching, the display of the lines connecting the key points also changes. When a pull-down button is pressed by the user, options (correct answer 402a and inference 402b) are displayed. When an option is pressed, the key points displayed in the image data display unit 418 change depending on the selected option. In FIG. 4, both the correct answer 402a and inference 402b are selected, and both the key points from the correct answer data and the key points from the inference result are displayed in the image data display unit 418. To cancel a selection, press the “x” button on the right end of the option to cancel the selection for each option. When deselecting all options in the display switching section 402, pressing the "x" button on the right end of the display switching section 402 causes all options to be deselected.

[0065] In this way, users can switch the key points displayed depending on their purpose. For example, if they want to know the difference between the correct data and the inference result, they can display the key points of the correct data and the inference result; if they want to check the correct data, they can display only the key points of the correct data; and if they want to check the inference result, they can display only the key points of the inference result.

[0066] Details of the lines connecting the key points displayed in the image data display section 418 will be explained with reference to FIG.

[0067] The keypoint size change unit 403 can change the size of the object indicating the position of the keypoint displayed in the image data display unit 418. The size can be changed from 5 to 50, with smaller values ​​indicating smaller keypoints. The user can intuitively adjust the size of the keypoint by dragging the pointer left or right. In Figure 4, it is set to 25. Large keypoints make it easier to find keypoints within the image data, but they may overlap with the inference target or with each other, making them difficult to see. On the other hand, small keypoints rarely overlap with the inference target or with each other, making them difficult to see, but they may make it more difficult to find the keypoints themselves. The function to change the keypoint size allows the user to adjust the keypoint size to an easy-to-see size depending on the inference target and keypoint distribution.

[0068] A line color change section 404 can change the color of the line connecting the key points.

[0069] When the line color change unit 404 is pressed, available colors are displayed, and the user can select any color. As a selection method, the user may select from colors displayed as a color chart, or may select a color adjusted by setting RGB values. In either case, the user can select an easy-to-see line color depending on the color of the inference target or key points.

[0070] Note that the color of the line connecting the key points may be limited to a color different from the color of the key points. By making the color of the line connecting the key points different from the color of the key points, the user can easily distinguish between the color of the line connecting the key points and the key points.

[0071] The line display switching unit 405 is a section that can switch whether or not to display lines connecting key points in the image data display unit 418. When a pull-down button is pressed by the user, options of "Yes" and "No" are displayed. When an option is pressed, the screen switches according to the option received. In FIG. 4, "Yes" is selected, and lines connecting key points are displayed in the image data display unit 418. The user can switch the display according to their purpose, for example, hiding the lines connecting key points when they want to check only individual key points, or displaying the lines connecting key points when they want to check the distance between key points.

[0072] In this way, the user can intuitively perform operations for switching screens, switching key point display, changing key point size, changing line color, and switching line display.

[0073] Next, the job ID selection section 406, the job explanation section 407, and the tag 408 will be described.

[0074] The job ID selection unit 406 is a section where the user can select a job ID to be inferred (analyzed) (S301). The job ID is an identifier for managing the trained model 603, configuration file 604, dataset information 605, and inference result 606 used when performing inference. When a pull-down menu is pressed, selectable job IDs are displayed. When a job ID selection by the user is accepted, information is displayed in the job description unit 407, tag 408, verification data list 4a, keypoint information 4b, and image data display unit 418 according to the accepted job ID. In FIG. 4, a job ID named "1" is selected, and data corresponding to "1" is displayed in the job description unit 407, tag 408, verification data list 4a, keypoint information 4b, and image data display unit 418.

[0075] The job description section 407 and tag 408 display a description of the job ID selected in the job ID selection section 406. For example, the content learned during job learning is displayed as a description of the job ID and a tag. The description of the job ID is set by the user during learning.

[0076] As shown in FIG. 4, the job description section 407 and tag 408 may be left blank, and are set as necessary.

[0077] Next, the verification data list 4a will be described. The verification data list 4a is a list made up of an inference data name display section 409, an inference data preview display section 410, an inference data accuracy display section 411, and an image data selection section 417.

[0078] The inference data name display section 409 displays the file name of the image data to be inferred, and corresponds to the "Name" column in the verification data list 4a.

[0079] The inference data preview display section 410 is a display section where a preview of the image data to be inferred is displayed, and is the "Preview" column in the verification data list 4a. Specifically, the image data acquired in S302 is displayed as a preview. Data corresponding to the file name of the inference target displayed in the inference data name 409 is displayed. The user can select image data to be displayed in the image data display section 418 by referring to the image data displayed in the inference data preview display section 410. Furthermore, even after selecting image data, the user can easily confirm what image data is displayed in the image data display section 418.

[0080] The image data selection unit 417 is a section for selecting an image to be displayed in the image data display unit 418, and is the check box to the left of the inference data name display unit 409 in FIG. 4. When a blank check box is pressed by the user, the image data corresponding to the pressed row is selected and displayed in the image data display unit 418. The entire selected column is highlighted, making it easy to determine which image data the user has selected. When image data is selected and the selected check box is pressed again, the selection is cancelled and the image displayed in the image data display unit 418 is hidden. When image data is selected and a check box other than the selected check box is pressed, the image data corresponding to the other check box is displayed in the image data display unit 418.

[0081] The inference data accuracy display unit 411 is a section that displays the inference results acquired in S304, and is the "PCK", "AUC", and "EPE" columns in Figure 4. "PCK" and "AUC" are indices that indicate the accuracy of the inference results, with higher values ​​indicating higher accuracy. "EPE" is an index that indicates the error rate of the inference results, with lower values ​​indicating lower error (higher accuracy). "PCK", "AUC", and "EPE" displayed in the inference data accuracy display unit 411 are generally indices used to measure the accuracy of the inference results of a pose estimation AI, but other indices may also be used.

[0082] Next, the keypoint information 4b will be described. The keypoint information 4b includes a keypoint name display section 412, a keypoint display selection section 413, a line length display section 414, a score display section 415, and a number of learning images display section 416, and displays information related to the inference accuracy of each keypoint related to the image selected in S305.

[0083] The content displayed in the key point information 4b is not affected by which key points are displayed (or not displayed) by the display switching unit 402. In other words, the content displayed in the key point information 4b remains unchanged whether only the key points of the correct answer data are displayed, only the key points of the inference result are displayed, or both the key points of the correct answer data and the inference result are displayed. For example, when only the key points of the inference result are displayed, the key points of the correct answer data are not displayed, so nothing is displayed in the line length display unit 414. Note that in this embodiment, an example has been described in which the displayed content remains unchanged, but values ​​may be hidden or changed as necessary.

[0084] The keypoint name display section 412 displays the name of each keypoint (keypoint name) and corresponds to the "Name" column in the keypoint information 4b. The keypoint names are named after the joints or bones of the target of inference, such as "wrist" for the wrist or "thumb_cmc (short for thumb-carpometacarpal)" for the carpometacarpal bone of the thumb. The background color of the keypoint name display section is a different color for each keypoint name, i.e., each joint or bone structure represented by the keypoint. The background color of the keypoint name display section is the same as the color of the keypoints displayed in the image data display section 418, and the same background color is never used for different keypoint names. This allows the user to identify all keypoints in the image data. For example, the background color of the keypoint name "thumb_cmc" is light blue, and the keypoint corresponding to the keypoint name "thumb_cmc" displayed in the image data display section 418 is also light blue. This makes it easy for the user to confirm which keypoint name corresponds to which keypoint in the image data.

[0085] The background color of the key point name display area and the color of the key points may be changed by the user to any color, allowing the user to select a color that is easy to see depending on the image data and the inference target, making it easier to distinguish the inference target and key points.

[0086] The key point display selection section 413 is a section where it is possible to select for each key point whether or not to display the key point in the image data display section 418, and is the "Visible" column in the key point information 4b. When a blank check box is pressed, the key point corresponding to the pressed check box is selected and displayed in the image data display section 418. When a check box that has already been pressed is pressed, the selection is cancelled and the key point displayed in the image data display section 418 is no longer displayed. Figure 4 shows a state where the check boxes corresponding to all key points have been pressed and all key points are displayed.

[0087] The user can adjust the number of key points to be displayed to a number that is easy to see depending on the purpose using the key point display selection unit 413. For example, if displaying all key points makes it difficult to see the inference results, it is possible to display only specific key points, and if the user wants to check the inference results for the entire image, it is possible to display all key points.

[0088] The distance display unit 414 displays the distance between the keypoints in the inference result and the keypoints in the correct data, and corresponds to the "Distance" column in the keypoint information 4b. A value greater than or equal to 0 is displayed, allowing users to check the accuracy of the inference made by the trained model. The lower the value (i.e., the shorter the line), the smaller the difference (hereinafter sometimes indicated as distance) between the correct data and the inference result, indicating higher accuracy. The higher the value (i.e., the longer the line), the larger the difference (hereinafter sometimes indicated as distance) between the correct data and the inference result, indicating lower accuracy. The user can check the difference between the correct data and the inference result from the value displayed in the distance display unit 414 and identify which keypoints have been accurately inferred and which keypoints have not. Based on these results, the user can take measures, such as increasing the amount of training data or revising annotations, for areas where the difference between the correct data and the inference result is large. In other words, displaying the difference between the correct data and the inference result in a way that users can easily recognize contributes to the creation of higher-performance AI models.

[0089] Note that if the joints or skeleton are hidden by the shadow of the inference target, etc., and the key points of the correct data cannot be seen from the image data to be inferred, the value displayed in the line length display unit 414 will be "None." For example, in the image displayed in the image data display unit 418, the metacarpophalangeal joint of the little finger is hidden behind the back of the hand, and the key point "little_finger_mcp" of the correct data corresponding to the metacarpophalangeal joint of the little finger cannot be seen. Therefore, the value displayed in the line length display unit 414 corresponding to little_finger_mcp is None (not shown).

[0090] The score display section 415 displays the reliability of the inference result, and is the "Score" column in the keypoint information 4b. The numerical value displayed in the score display section 415 is a value between 0 and 1, with a higher numerical value indicating a higher reliability of the inference result based on the trained model, and a lower numerical value indicating a lower reliability of the inference result based on the trained model.

[0091] When the training data is normal, there is a correlation between the distance between the correct answer data and the inference result, which is displayed in the line length display section 414, and the reliability of the inference result, which is displayed in the score display section 415. The smaller the value indicating the distance between the correct answer data and the inference result, the higher the accuracy of the inference. Therefore, it is predicted that a correlation will be observed, such that, for example, the smaller the value indicating the distance between the correct answer data and the inference result (the higher the accuracy of the inference), the higher the reliability of the inference result. However, if no correlation is observed, it is predicted that there is some problem with the training data, and this can provide an opportunity for the user to review the training data.

[0092] By checking the correlation, it is possible to check whether there are any defects in the annotations added to the images selected by the image data selection unit 417, whether the amount of training data is sufficient, and so on.

[0093] If there is a correlation, that is, if the value in the line length display section 414 is small and the value in the score display section 415 is large (high accuracy and high reliability of the inference result), it can be confirmed that there is no deviation in the annotation during learning and that the amount of learning data is sufficient. If the value in the line length display section 414 is large and the value in the score display section 415 is small (low accuracy and low reliability of the inference result), although a correlation is observed, it indicates that the accuracy and reliability of the inference result are low for some reason, such as an insufficient amount of learning data, and that improvement is necessary.

[0094] If there is no correlation, the following improvement measures can be considered. For example, if the value in the line length display section 414 is large and the value in the score display section 415 is large (low accuracy and high reliability of the inference result), the annotations added to the selected image during learning may be misaligned. Therefore, by adding annotations to the correct positions and re-learning, it may be possible to improve the accuracy and reliability of the inference result. Also, if the value in the line length display section 414 is small and the value in the score display section 415 is small (high accuracy and low reliability of the inference result), the amount of learning data may be insufficient. Therefore, in order to further increase the reliability, the user needs to increase the number of images annotated with low-reliability keypoints during learning.

[0095] In the example shown in FIG. 4, for the keypoint with the keypoint name "thumb-tip," the value displayed in the line length display section 414 is 0.02, and the value (reliability of the inference result) displayed in the score display section 415 corresponding to "thumb-tip" is 0.79. Because a correlation between the distance and the reliability of the inference result can be seen, the user can understand that the learning has been performed correctly. On the other hand, it can also be seen that the reliability is not as high as expected, given the low value of 0.02. In that case, measures can be taken to increase the reliability, such as by increasing the amount of training data.

[0096] In this way, by comparing the distance between the correct data and the inference result (the accuracy of inference by the trained model) and the reliability of the inference result, it is possible to check for abnormalities in the training data and use this information to add or change the training data as necessary.

[0097] Note that the countermeasures taken after confirming the correlation are merely examples, and the present invention displays the distance between the correct data and the inference result and the reliability of the inference result displayed in the score display unit 415, but the presence or absence of a correlation and the countermeasures to be taken are determined by the user.

[0098] Note that a predetermined notification may be output when no correlation is found. For example, a correlation coefficient is calculated using a known method between the value in the line length display section 414 and the value in the score display section 415. If the calculated correlation coefficient is less than a preset threshold, a notification to review the training data is output. For example, the notification may identify and display rows related to keypoint names for which no correlation is found, and notify the user to increase the training data or review the dataset information 605. The notification method is not limited to this, and any notification method that suggests the user to review the training data may be used.

[0099] In this embodiment, when only the key points of the inference result for the key point name "wrist" are displayed, two key points are marked, one on each of the right and left hands. In this case, the accuracy of the inference by the trained model displayed in the line length display unit 414 and the reliability of the inference result displayed in the score display unit 415 are displayed as average values ​​for the left and right hands. In this way, even when only the inference result or correct answer data is displayed, if two or more key points are displayed in the image data display unit 418, the accuracy of the inference by the trained model displayed in the line length display unit 414 and the reliability of the inference result displayed in the score display unit 415 are displayed as average values ​​for the two or more key points.

[0100] The accuracy of the inference based on the trained model, displayed in the line length display section 414, and the reliability of the inference result, displayed in the score display section 415, may be displayed as separate values ​​for the two or more marks. For example, in the case of Figure 4, the numerical values ​​for the right hand and the left hand are displayed separately. This allows for more accurate data to be confirmed, as the reliability of the inference result and the accuracy of the inference based on the trained model may differ between the right and left hands. In addition, the results for the right and left hands can be compared, allowing for confirmation of which hand has a higher reliability of the inference result and higher accuracy of the inference.

[0101] The learning number display section 416 displays the number of images annotated for each keypoint, and is the AnnTotal column in the keypoint information 4b. An integer greater than or equal to 0 is displayed, with a higher number indicating a greater number of images annotated. Specifically, the number of labels assigned to each keypoint in the annotation folder 609 is counted. For example, for the "wrist" keypoint, depending on the image data 608, the wrist may be cut off and an annotation may not be possible. The number of images in the image folder 607 that have been annotated for the wrist (i.e., the number of images in the annotation folder 609 with the label name "wrist") is counted and displayed in the learning number display section 416.

[0102] This allows users to examine the distribution of the number of images annotated by each keypoint. Users can identify annotations with low values ​​(i.e., a small number of images annotated) and take measures such as increasing the number of images containing the annotation to be trained.

[0103] It is also possible to set a threshold for the number of images to be learned, and if the number is less than the threshold, the learning number display unit 416 including a value less than the threshold may indicate that the number of images to which annotations have been added is insufficient.

[0104] The image data display section 418 displays image data selected from the list of verification data by pressing the image data selection section 417 (S305). The image data display section 418 displays "Image Name: egocentric_0030" in the outer upper left corner, along with the file name of the selected image data. This allows the user to view the image data while confirming which image data they selected. The bottom of the image data display section 418 also displays "Correct answers are displayed as circles, inferences as squares," clearly indicating that keypoints in the correct answer data are displayed as circles, and keypoints in the inference results are displayed as squares. The color of the displayed keypoints is the same as the background color of each keypoint name displayed in the keypoint name display section 412. This makes it easy for the user to confirm which keypoint name corresponds to which keypoint in the image data.

[0105] The content displayed on the image data display section 418 will be described in detail with reference to FIG.

[0106] 5 is a diagram showing an example of a screen displaying image data, in which image data display section 418 in FIG.

[0107] Screen 5a is a screen that displays the keypoints of the correct data and the keypoints of the inference results (S307). The keypoints of the correct data are displayed as circles, and the keypoints of the inference results are displayed as squares. By displaying the keypoints of the correct data and the keypoints of the inference results using different symbols, the user can easily distinguish between the keypoints of the correct data and the keypoints of the inference results. The keypoints of the correct data and the keypoints of the inference results that have the same keypoint name are the same color, making it easier to recognize the correspondence between the keypoints.

[0108] On screen 5a, keypoint 501 in the correct data corresponding to the keypoint name middle_finger_mcp in the right hand, i.e., the metacarpophalangeal joint of the middle finger, is connected by line 502 to keypoint 503 in the inference result corresponding to keypoint name middle_finger_mcp in the right hand. As with the right hand, keypoint 504 in the correct data corresponding to keypoint name middle_finger_mcp in the left hand is connected by line 505 to keypoint 506 in the inference result corresponding to keypoint name middle_finger_mcp in the left hand. In this way, the keypoint in the correct data and the keypoint in the inference result corresponding to that keypoint are displayed connected by a line.

[0109] If the inference target is divided into two or more parts, i.e., if there are four or more correct keypoints and inference keypoints in total, the keypoints for each divided inference target are connected by a line. For example, if the inference target is divided into a right hand and a left hand, the keypoints within the right hand are connected by a line, and the keypoints within the left hand are connected by a line.

[0110] By displaying the data in this way, the user can easily grasp the distance between the inference result and the correct data (the difference between the inference result and the correct data), making it easier to evaluate the accuracy of the inference process. In other words, the longer the line connecting the inference result and the key points in the correct data, the lower the accuracy, and the user can be made aware that measures need to be taken to improve accuracy.

[0111] The present invention is useful not only for comparing keypoints within image data, but also for comparing inference results for multiple pieces of image data. For example, by comparing the inference results for image data A and image data B, it is possible to determine which image has a higher accuracy of inference results, or whether there is a difference in the accuracy of inference results even for the same keypoint names.

[0112] When displaying the keypoints of the supervised data and the keypoints of the inference result on screen 5a, any display form may be used as long as it is clear that the keypoints are associated with each other. In addition to displaying the keypoints by connecting them with lines as in this embodiment, for example, when an instruction is received from the user to display only the keypoints corresponding to a specific keypoint name (e.g., middle_finger_mcp), the keypoints of the inference result and the supervised data corresponding to that keypoint name may be displayed blinking or highlighted.

[0113] Screen 5b is a screen that displays only the keypoints of the inference results (S308). On screen 5b, keypoint 501 of the inference results corresponding to keypoint name middle_finger_mcp is connected by line 508 to keypoint 507 of the inference results corresponding to keypoint name wrist. In this way, on a screen that displays only specific keypoints (when it is not a screen that displays both the correct answer data and the inference results as in Figure 5a), keypoints are displayed connected to each other along the shape of the skeleton. Using this display format makes it easier to understand whether the positions of the keypoints have been identified as appropriate based on the shape of the skeleton.

[0114] Only the key points of the correct answer data may be displayed on screen 5b, allowing the user to confirm where the key points will be displayed if the inference is correct.

[0115] In this way, the present invention has the effect of making the accuracy of the inference results visible by having the CPU 201 associate key points and display them superimposed on the image data, thereby facilitating the user's confirmation work.

[0116] As described above, according to this embodiment, the evaluation results can be easily confirmed. Specifically, when displaying both the correct answer data and the inference results, the differences between the correct answer data and the inference results are displayed in an easily understandable manner by connecting the key points of the corresponding correct answer data with the key points of the inference results with lines. When displaying only the inference results, the inference results can be appropriately evaluated by displaying in an easily understandable manner whether annotations have been made in appropriate positions along the inference target.

[0117] The present invention can be embodied as, for example, a system, an apparatus, a method, a program, a recording medium, etc. Specifically, the present invention may be applied to a system consisting of multiple devices, or may be applied to an apparatus consisting of a single device.

[0118] The various controls described above as being performed by CPU 201 may be performed by a single piece of hardware, or the entire device may be controlled by multiple pieces of hardware (e.g., multiple processors or circuits) sharing the processing.

[0119] Furthermore, the program of the present invention is a program that enables a computer to execute the processing method of the flowchart shown in Fig. 3, and the storage medium of the present invention stores a program that enables a computer to execute the processing method of Fig. 3. Note that the program of the present invention may be a program for each processing method of each device in Fig. 1.

[0120] As described above, it goes without saying that the object of the present invention can also be achieved by supplying a recording medium on which a program that realizes the functions of the above-mentioned embodiments is recorded to a system or device, and having the computer (or CPU or MPU) of that system or device read and execute the program stored on the recording medium.

[0121] In this case, the program itself read from the recording medium will realize the novel functions of the present invention, and the recording medium on which the program is recorded will constitute the present invention.

[0122] Examples of recording media for supplying the program include flexible disks, hard disks, optical disks, magneto-optical disks, CD-ROMs, CD-Rs, DVD-ROMs, magnetic tapes, non-volatile memory cards, ROMs, EEPROMs, and silicon disks.

[0123] Furthermore, it goes without saying that not only are the functions of the above-mentioned embodiments realized by the computer executing a program it has read, but also cases are included in which an OS (operating system) running on the computer performs some or all of the actual processing based on the instructions of the program, and the functions of the above-mentioned embodiments are realized through that processing.

[0124] Furthermore, it goes without saying that this also includes cases where a program read from a recording medium is written into a memory provided on a function expansion board inserted into a computer or a function expansion unit connected to the computer, and then a CPU or the like provided on the function expansion board or function expansion unit performs some or all of the actual processing based on the instructions of the program code, thereby realizing the functions of the above-mentioned embodiments.

[0125] Furthermore, the present invention may be applied to a system consisting of multiple devices, or to a device consisting of a single device. It goes without saying that the present invention can also be applied to a system or device that is achieved by supplying a program to the system or device. In this case, the system or device can enjoy the effects of the present invention by reading a recording medium that stores a program for achieving the present invention into the system or device.

[0126] Furthermore, by downloading and reading a program for achieving the present invention from a server, database, etc. on a network using a communication program, the system or device can enjoy the effects of the present invention. Note that the present invention also includes configurations that combine the above-mentioned embodiments and their modified examples. [Explanation of symbols]

[0127] 101 client terminals 102 Server

Claims

1. In the first setting, control is performed to associate and display feature points of objects related to an input image identified by inference processing using a trained model, In the second setting, a display control means is provided that controls displaying feature points of an object related to the input image identified by inference processing using the trained model in association with feature points related to the correct answer data related to the object. An information processing device characterized by:

2. the first setting is a setting indicating that feature points identified by the inference process are displayed, and feature points related to the correct answer data related to the object are not displayed; The second setting is a setting indicating that the feature points identified by the inference process and the feature points related to the correct answer data related to the object are to be displayed.

2. The information processing device according to claim 1,

3. the display control means controls the display of an object at the position of the feature point; The display control means controls display of the feature points in association with each other by connecting the objects related to the feature points with lines.

2. The information processing device according to claim 1,

4. Controlling to display objects related to feature points at the same location on the object in the same color.

4. The information processing device according to claim 3,

5. The input image is an input image that is the subject of inference processing by a trained model.

2. The information processing device according to claim 1,

6. The display control means controls to display a result of an inference process for an object related to the input image using the trained model that has learned feature points of the object.

2. The information processing device according to claim 1,

7. At least one of the size of the object related to the feature point and the color of the line connecting the objects related to the feature point can be changed.

4. The information processing device according to claim 3,

8. The display control means controls the display of the object related to the feature point and the line connecting the objects related to the feature point in different colors.

4. The information processing device according to claim 3,

9. 2. The information processing apparatus according to claim 1, wherein the feature points are points that indicate the positions of the joints of the object.

10. The display control means of the information processing device comprises: In the first setting, control is performed to associate and display feature points of objects related to the input image identified by inference processing using the trained model, In the second setting, control is performed so that feature points of an object related to the input image identified by the inference process using the trained model are displayed in association with feature points related to the correct answer data related to the object. An information processing method comprising:

11. 10. A program for causing at least one computer to function as each of the means of the information processing device according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Skeleton recognition device, learning method, and learning program

    JP7571796B2

  • Skeletal tracking using previous frames

    US20210248373A1

  • Whole body segmentation

    US20230419497A1

  • Method, device and program for skeleton extraction

    JP2020198019A

  • JPP7571796B