Machine learning model development equipment, annotation equipment for model development.

JP2026131213APending Publication Date: 2026-08-14QUALTEC CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-02-03
Publication Date
2026-08-14

AI Technical Summary

Benefits of technology

【0025】 タスクは個別毎に最適化され、効率的で精度の高いアノテーションを実現する。また、アノテーターの集中力とモチベーションを維持させるために、入力インターフェース、視覚的な改善、ゲーミフィケーションの活用により、単調な作業を楽しい体験に変える。これらにより、高品質なアノテーションが実現し、高精度なAIモデル構築が可能となる。 モデルの判断が倫理的または社会的に影響を及ぼす場合に、人間が最終判断を行うことができるため、倫理的·社会的配慮に対応できる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026131213000001_ABST
    Figure 2026131213000001_ABST
Patent Text Reader

Abstract

High-precision pixel-level annotation is required, but the process is time-consuming and labor-intensive, there is variability in quality among annotators, and maintaining annotator motivation is difficult. [Solution] In AI model development, particularly AI development for image segmentation, personalized analysis of the annotator is performed before and during the work. Based on the annotator's work status, work history, skill level, work tendencies, past work quality, and task completion rate, as measured by cameras, the optimal task is designed. This optimizes each task individually, resulting in efficient and highly accurate annotation. Furthermore, to maintain the annotator's concentration and motivation, the input interface, visual improvements, and gamification are used to transform monotonous work into an enjoyable experience, thereby achieving high-quality annotation and enabling the construction of highly accurate AI models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a system used for learning a machine learning model and a machine learning model development device. It also relates to a learning data generation device and method, a program learning device and method, a classification device and method, a prediction device and method, and an evaluation device and method for a machine learning model.

[0002] The invention relates to a method for managing an annotation operation, a device and a system for assisting the same. It also relates to a device and a method for more efficiently managing the annotation operation. Further, it provides a method for ensuring the accuracy of annotation results, a device and a system for assisting the method.

Background Art

[0003] The annotation operation means an operation of tagging label information for each data in order to generate a learning data set. Annotation is an important process that greatly affects the learning of a machine learning model and the evaluation of a learned model. Since the annotation operation is generally performed by a person, a considerable human cost and time cost increase in order to generate a large amount of learning data sets.

[0004] In the present invention, the collected data is used as learning data for newly learning a machine learning model or for updating or re-learning an operating machine learning model after being annotated. In order to quickly generate a learning data set, automatic annotation is performed on the data using a machine learning model.

[0005] Image segmentation is the process of dividing an image into multiple segments (parts). This segmentation involves grouping pixels within an image based on specific criteria, with each segment corresponding to a particular object or region. The primary purpose of image segmentation is to make the image content easier to understand and to extract information necessary for specific tasks (such as object recognition or region detection).

[0006] Active learning (AL) is a type of machine learning that efficiently builds high-performance models by "actively selecting" the data that the model uses for training. It is particularly useful when labeling is costly. For example, it is effective when there is a large amount of data but insufficient accurate labels (such as those assigned manually).

[0007] In this invention, machine learning involves labeling the entire dataset and using that labeling to train a machine learning model. On the other hand, in active learning, the model itself selects "which data to have labeled," and humans only label that specific data. This reduces the cost of labeling while building a highly accurate model.

[0008] Human-in-the-Loop (HITL) machine learning refers to an approach that incorporates humans as a crucial element of the system in the machine learning process. In this method, humans intervene in the learning process and decision-making of algorithms and models, providing assistance and adjustments. The Human-in-the-Loop approach of this invention not only improves the accuracy and reliability of machine learning models but also enables ethical judgments and adaptation to specific tasks.

[0009] The purpose and role of HITL in this invention is to improve model accuracy by allowing humans to label unlabeled data and provide more accurate training data.

[0010] When a model's predictions are ambiguous or uncertain, human judgment can be used to handle highly uncertain processes. Error detection and correction are possible by having humans review and correct model errors. Even when a model's judgment has ethical or social implications, a human can make the final decision, thus allowing for consideration of ethical and social factors. [Prior art documents] [Patent Documents]

[0011] [Patent Document 1] Japanese Patent Publication No. 2023-021647 [Overview of the project] [Problems that the invention aims to solve]

[0012] Developing machine learning models requires a series of processes, from creating training data to training the model and evaluating its performance. However, each step requires specialized knowledge and skills. Furthermore, there are numerous challenges, including quality control of training data, selection of appropriate algorithms, and evaluation of model performance. Therefore, without sufficient expertise, project success becomes difficult.

[0013] Developing machine learning models, particularly those for image segmentation, requires a large amount of high-quality training data. Creating this training data involves essential annotation work to accurately label the boundaries of each object, but the time-consuming and labor-intensive nature of image segmentation annotation is a significant challenge. Models for image segmentation require pixel-level accuracy, and even slightly inaccurate annotations can affect the model's performance.

[0014] Annotators, who actually perform the annotation work, must work carefully to avoid incorrect labeling and draw accurate boundaries. Therefore, accuracy and time are essential. To reduce variations in labeling among annotators, standardization of detailed annotation guidelines and quality standards is crucial. In particular, in segmentation, where the fine handling of boundaries directly impacts performance, the development of standardized guidelines is indispensable.

[0015] Developing machine learning models, particularly image segmentation models, requires highly accurate pixel-level annotation. However, this process is time-consuming and labor-intensive, and there is variability in quality among annotators. Furthermore, the monotonous nature of the work makes it difficult to maintain annotator motivation. Furthermore, annotation work is monotonous yet requires a high level of concentration, making it extremely difficult for annotators to maintain their motivation. [Means for solving the problem]

[0016] The present invention automates most of the processes involved in developing a machine learning model, such as creating training data, training the model, and evaluating it, providing a system that allows on-site personnel to easily carry out machine learning model development without requiring specialized knowledge.

[0017] The present invention incorporates a human-in-the-loop mechanism, enabling users to continuously improve the performance of machine learning models while participating in the process. This system allows even non-experts to efficiently build high-quality machine learning models and optimize them for practical use.

[0018] To ensure the performance of machine learning models, the quality of annotation data is paramount. By implementing personalized quality control tailored to the personality and work quality of each annotator, flexible data management based on individual characteristics becomes possible.

[0019] The invention of the present application analyzes the work tendencies and past achievements (such as task completion rate and working hours) of each annotator, and realizes the improvement of the quality of annotation data and an efficient work process by providing appropriate feedback and task assignment.

[0020] As a specific embodiment, at the start of work, the work quality of the annotator is evaluated by comparing with the set of true answers (Ground truth answers), and further the working hours and the accuracy for each label are analyzed. Based on this information, by selecting a dataset suitable for the annotator, efficient and accurate annotation can be realized.

[0021] In order to maintain the concentration of the annotator and make their work efficient and attractive, it is important to devise interfaces such as the input screen, input pen, and input mouse. In addition, the working state of the user is monitored by a camera or the like to promote close interaction between the annotator and the system. Also, an introduction of a mechanism that can provide information to the annotator step by step and frequently.

[0022] By the above implementation, not only the work efficiency but also the satisfaction and motivation of the annotator can be improved. In particular, by visual improvement and the use of gamification, monotonous work can be changed into an enjoyable experience. As a result, high-quality annotation data can be obtained and high-precision model construction becomes possible.

[0023] In the invention of the present application, personalized analysis of the annotator is performed before and during work. Based on the working state, work history, skill level, work tendency, past work quality, and achievements such as task completion rate of the annotator by a camera or the like, an optimal task is designed.

[0024] The present invention has a first step of performing a personalized analysis of an annotator before or during work, and a second step of designing an optimal task based on one of the working state, work history, skill level, work tendency, past work quality, and task completion rate of the annotator. The annotation method is characterized in that the task is optimized by the first step and the second step.

Effects of the Invention

[0025] Tasks are optimized individually, realizing efficient and accurate annotation. In addition, in order to maintain the concentration and motivation of the annotator, by utilizing the input interface, visual improvement, and gamification, monotonous work is changed into an enjoyable experience. As a result, high-quality annotation is realized and high-precision AI model construction becomes possible. When the judgment of the model has an ethical or social impact, a human can make the final judgment, so it is possible to respond to ethical and social considerations.

Brief Description of the Drawings

[0026] [Figure 1] It is an explanatory diagram and a block diagram of a machine learning model development device of the present invention and an annotation device (standalone environment) for model development. [Figure 2] It is an explanatory diagram and a block diagram of a machine learning model development device of the present invention and an annotation device (local network environment) for model development. [Figure 3] It is an explanatory diagram and a block diagram of a machine learning model development device of the present invention and an annotation device (cloud / internet environment) for model development. [Figure 4] It is a configuration diagram and an explanatory diagram of a user interface (UI) of the present invention. [Figure 5] It is an explanatory diagram of a workspace screen program of the present invention. [Figure 6] It is an explanatory diagram of an annotation screen program of the present invention. [Figure 7] This is an explanatory diagram of the annotation screen program of the present invention. [Figure 8] This is an explanatory diagram of the annotation screen program of the present invention. [Figure 9] This is an explanatory diagram of the annotation screen program of the present invention. [Figure 10] This is an explanatory diagram and flowchart of the administrator's setup process for the present invention. [Figure 11] This is an explanatory diagram of how to include precautions in the example data of the present invention. [Figure 12] These are explanatory diagrams and flowcharts of the annotator analysis of the present invention. [Figure 13] These are explanatory diagrams and flowcharts of the active learning program of the present invention. [Figure 14] These are explanatory diagrams and flowcharts of the annotator analysis of the present invention. [Figure 15] This is an explanatory diagram of the indicator for annotator analysis according to the present invention. [Figure 16] This is a diagram illustrating the configuration and operation of the input mouse of the present invention. [Figure 17] This is a diagram illustrating the configuration and operation of the input pen of the present invention. [Figure 18] This is an explanatory diagram of the machine learning model method and annotation method for model development according to the present invention. [Figure 19] This is an explanatory diagram illustrating the operation and usage method of the input tablet and input device of the present invention. [Figure 20] This is an explanatory diagram of the machine learning method and data input monitoring method of the present invention. [Figure 21] This is an explanatory diagram of the machine learning method and data input monitoring method of the present invention. [Figure 22] This is an explanatory diagram of the machine learning method and input data selection method of the present invention. [Figure 23] This is an explanatory diagram of a machine learning model development apparatus and an annotation apparatus for model development (standalone environment) in another embodiment of the present invention. [Figure 24] This is an explanatory diagram of a machine learning model development apparatus and an annotation apparatus for model development (standalone environment) in another embodiment of the present invention. [Figure 25] This is an explanatory diagram of a machine learning model development apparatus and an annotation apparatus for model development (standalone environment) in another embodiment of the present invention. [Figure 26] This is an explanatory diagram of a machine learning model development apparatus and an annotation apparatus for model development (standalone environment) in another embodiment of the present invention. [Modes for carrying out the invention]

[0027] The following description will refer to the attached drawings to explain an embodiment of the present invention: a machine learning model development apparatus and an annotation apparatus for model development. Figure 1 is an explanatory diagram of the system configuration of the present invention under a standalone environment.

[0028] The information processing device A104 includes a UI (User Interface) management program 105, a login program 106, a workspace program 107, an annotation program 108, a management program 109, a setup program 110, a user analysis program 111, an active learning program 112, an AI learning program 113, an AI evaluation program 114, an AI deployment program 115, an AI inference program 116, and the like.

[0029] It also includes an image database (DB) 117, an annotation database (DB) 118, a model database (DB) 119, a work database (DB) 120, an AI model base (DB) 121, a deployed AI model base (DB) 122, a personal database (DB) 123, and an administrator settings base (DB) 124.

[0030] The information processing device A104 is connected to a display 100, a pen 101, a mouse 102, a camera (204), a controller 103, a numeric keypad (not shown), a keyboard (not shown), and other devices.

[0031] The display 100 includes those with and without touch functionality. Those with touch functionality are called touch panels 100 and input tablets 100. Furthermore, with touch panel operation, images displayed on the display can be moved by sliding a finger. Images displayed on the display can also be enlarged and reduced using pinch-in and pinch-out gestures. Pinch-in and pinch-out gestures can be modified and configured to suit the preferences of users, workers, and administrators, as well as to accommodate their physical characteristics.

[0032] The input tablet 100 is not limited to devices that use an input pen 101 or an input mouse 102 for input work, but also includes devices that use blinking (eye contact, etc.) or gestures for input. The input pen 101 is used with displays or tablet devices that have touch functionality, and is particularly useful during annotation work.

[0033] Physical disabilities often include "visual impairment," "hearing and balance disorders," "speech, language, and chewing disorders," "limb disabilities," and "internal disabilities due to diseases of internal organs."

[0034] Camera 204 observes the viewpoint position, gaze direction, eye opening and closing, facial expressions, face orientation, and mouth opening and closing of administrators and workers (users), and transmits the observed data to the information processing device. This data is then used for annotation, input work, AI learning, AI inference, deployment processing, and management settings. It also allows for the modification or configuration of menus and displays on the workspace screen, annotation screen, and management screen.

[0035] This system can be operated even without a touch panel and stylus 101, by substituting a non-touch display and input mouse 102. A combination of a non-touch display, input mouse 102, input pen 101, etc., is also acceptable.

[0036] The input mouse 102 and input pen 101 are coordinate input devices, function selection devices, display control devices, and control switching devices, and have the functions of control methods, control methods, modification methods, and setting methods for these devices of the present invention. The same applies to the camera 204.

[0037] As shown in Figure 16, the input mouse 102 is equipped with an operation button 551, a scroll wheel button 552, and a selection button 553. The usage state and operating state can be changed using the operation button 551, the scroll wheel button 552, the selection button 553, or a combination thereof.

[0038] Furthermore, settings can be modified and configured to accommodate the preferences of users, workers, and administrators, as well as their physical characteristics. The usage and operation status can also be changed in combination with the output data from the surveillance camera 204.

[0039] The wheel button 552 and the selection button 553 can be switched on or off. The user can understand the usage status and operation status by observing the on / off status of the wheel button 552 and the selection button 553, and can also change the usage status or settings. Furthermore, the on / off status can be controlled by user operation or the controller 103.

[0040] The direction in which the input mouse 102 is moved is detected by the accelerometer 206. The accelerometers 205 and 206 function not only as acceleration and movement speed sensors but also as direction sensors. By detecting or recognizing the speed, acceleration, and direction using the input mouse 102, the marker state, selection state, and operation state of the target shape 558 are monitored, and annotation work is applied appropriately. This can also be applied to management work and work tasks. The input mouse 102 has the following functions.

[0041] Basic shapes (rectangles, circles, triangles, etc.) can be easily drawn. Simply drag with the mouse input 102 to create the desired shape. For example, holding down button 553 while drawing allows you to draw shapes such as squares and circles while maintaining their proportions.

[0042] It features a grid and snapping function. Pressing button 553 while drawing will display the grid (scale). Furthermore, you can set the drawn shapes to snap to specific positions or lines.

[0043] It also has a straight-line assistance function. When you select button 553 and use input mouse 102 to draw a straight line, guide lines will be displayed. This function allows you to draw straight lines and angles accurately.

[0044] You can intuitively manipulate drawn shapes using the input mouse (102) by selecting them and resizing, rotating, flipping, etc. Functions can be set and switched using buttons (553) and the scroll wheel button (552). Furthermore, you can easily modify shapes using drag-and-drop or handles.

[0045] Additionally, by adjusting settings using button 553 and wheel button 552, you can add color to drawn shapes and fill the inside of shapes with color using the fill function. Color changes can be easily made by simply clicking and selecting with buttons 553 and 555.

[0046] The input mouse 102 allows you to access a context menu when using buttons 553a and 553b to input shapes. It also allows you to toggle the display of annotations on and off.

[0047] By rotating the wheel button 552, you can zoom in and out of the canvas. This is useful for detailed work and checking the overall composition when drawing shapes. In addition, the target shapes 558 illustrated and explained in Figures 9, 11, 18, and 26 can be displayed in sections, and the number of sections can be set and changed.

[0048] Buttons 553 and 555 allow you to instantly execute frequently used tools and operations. You can also switch between high, medium, and low display settings for annotations.

[0049] By pressing button 553 and button 555 on the input pen 101 while operating the mouse, you can restrict certain actions. For example, drawing with the rectangle tool while holding down button 553 will draw a square, and drawing with the circle tool will draw a perfect circle.

[0050] Double-clicking buttons 553 and 555 allows you to edit the properties (size, color, style, etc.) of selected shapes and objects. Additionally, double-clicking a specific tool opens its settings screen.

[0051] The input sensitivity can be adjusted by selecting button 553, button 555 on the input pen 101, etc. Adjusting or setting the sensitivity can improve the smoothness of drawing. For example, setting the sensitivity high will cause the cursor on the screen to move quickly with just a slight movement of the mouse, enabling high-speed operation. For example, the embodiment shown in Figure 18(b) is illustrated. These functions and methods can be modified and configured to suit the preferences of users, workers, and administrators, as well as their physical characteristics.

[0052] The input pen 101 has a pen pressure sensing function (pressure-sensitive axis, pressure-sensitive axis, pressure sensor, pressure sensor, pressure sensor) on its pen shaft 557. In annotation processing, the software can use the pen pressure sensing to adjust the line thickness and transparency.

[0053] The input mouse 102 and input pen 101 of the present invention can be used to assign tool switching and shortcut functions for specific applications in various work scenarios, such as annotation display and annotation methods, using multiple buttons 553 and 555. For example, switching to the "selection tool" or "pen tool" can be done quickly with the buttons. These features, methods, and settings can be modified and configured to suit the preferences of users, workers, and administrators, as well as to accommodate their physical characteristics.

[0054] The input mouse 102 and input pen 101 of this invention have built-in LED lights. The LED lights allow the RGB colors of the input mouse 102 and input pen 101 to be freely changed. The RGB colors of the displayed input mouse 102 and input pen 101 can be freely changed in the login program, workspace screen program, annotation screen program, and management screen program shown in Figure 4. Furthermore, preferred colors and lighting patterns can be set in the annotation display. This adds visual enjoyment, such as the lights changing according to the work state.

[0055] The input mouse 102 and input pen 101 of this invention are compatible with touchpads and touchscreens. They can be operated by swiping with the thumb, and operations can be switched by tapping.

[0056] Furthermore, the input mouse 102 and input pen 101 have a weight adjustment function. By adding or removing internal weights, users can set their preferred weight, improving the precision and comfort of operation. The weight setting can be changed and adjusted to suit the preferences of the user, worker, or administrator, as well as to accommodate their physical characteristics.

[0057] Furthermore, buttons 553 and 55 can be customized by the user in terms of their placement, shape, and weight. Side buttons can also be added. The grip can be replaced to achieve a comfortable operating feel tailored to the user's hand. The operating feel can be changed and configured to suit the preferences of the user, worker, and administrator, as well as their physical characteristics.

[0058] As illustrated and explained in Figure 21 and other figures, the present invention includes an input pen 101 with a pressure-sensing (pen pressure) function. It is equipped with a function to sense pen pressure, allowing for real-time adjustment of line thickness and darkness during handwriting.

[0059] For example, a light touch can draw a thin line, while a firm press can draw a thick line. Furthermore, the system observes, measures, and evaluates user and worker fatigue, skill level, and attention span to set and modify input methods, annotation display, and annotation control methods. Settings can also be modified to accommodate user, worker, and manager preferences, as well as physical characteristics.

[0060] Furthermore, the display and settings of the login program, management screen program, workspace screen program, and annotation screen program, as illustrated in Figure 4, can be switched, configured, and modified. In addition, the control method, configuration method, and display can be switched, configured, and modified based on observations and results of viewpoints, as explained in Figure 20, etc. Furthermore, modifications and settings can be made to accommodate the preferences of users, workers, and administrators, as well as their physical characteristics.

[0061] The input pen 101 of this invention has hundreds of pressure sensitivity levels, enabling extremely delicate operation. This allows for a richer expression of nuances during drawing.

[0062] The input pen 101 has a tilt detection function. This function changes the thickness and direction of the line according to the angle at which the stylus is tilted. Like a pencil or brush, you can draw wide, thick lines by tilting the pen diagonally, enabling more natural and diverse expressions. It allows for 3D modeling and shading. This feature enables the creation of a three-dimensional effect when applying shading.

[0063] Note that input devices are not limited to touch panels, input mice 102, and input pens 101. For example, an air mouse may also be used. An air mouse is a device that controls a cursor by moving the hand in the air, and unlike a regular mouse or touchpad, it operates without physical contact. Voice input devices may also be used to control shapes and designs by voice using a microphone and voice recognition technology. Instead of displaying annotations, voice commands can be used to instruct shape input, create drawings, and select tools. These functions can be modified and configured to accommodate the preferences of users, workers, and administrators, as well as their physical characteristics.

[0064] Motion sensors are also given as an example. Using motion sensors such as Kinect and Leap Motion, hand and body movements are detected and input into a computer. It is possible to move your hands in the air and draw shapes in a digital space.

[0065] Both the input pen 101 and the input mouse 102 can be used simultaneously, one or the other can be selected, or they can be used alternately. The operation buttons 551, selection buttons 553, selection buttons 555, and the display on the display unit 556 of the input pen 101 and the input mouse 102 can be linked. It can be linked to the observation status from camera 204. For example, the user's viewpoint position and viewpoint movement can be monitored, and the display can be linked to changes or changes in the state.

[0066] The settings, modifications, and controls of the apparatus and method of the present invention based on measurements, observations, or output data from the camera 204 can be modified and configured to suit the preferences of the user, operator, or manager, as well as their physical characteristics.

[0067] The input mouse 102 is equipped with an accelerometer 205. The output of the accelerometer 205 detects the speed and acceleration of the mouse movement by the user. It also detects the direction in which the user moves the mouse. By detecting or recognizing the speed, acceleration, and direction, the marker state, selection state, and operation state of the target shape 558 are monitored and appropriately applied to annotation work, input work, management work, etc. The same applies to the input pen 101.

[0068] As shown in Figure 17, the input pen 101 is equipped with a selection button 555 and a display unit 556. The usage state and operation state can be changed using the display unit 556 and the selection button 555. The selection button 555 can also be turned on or off by user operation or control by the controller 103.

[0069] For example, users, administrators, and controllers 103 can understand the usage status and operation status by the illumination or de-illumination of the selection button 555 and the display unit 556, and can also change the usage status or settings using the selection button 555.

[0070] Furthermore, it can be linked to the observation status from camera 204. For example, the user's viewpoint position and viewpoint movement can be monitored, and the display can be linked to changes or changes in the state.

[0071] The input pen 101's accelerometer 205 detects the direction of movement, angle of movement, and speed of movement. By detecting or recognizing the speed, acceleration, and direction, the marker state, selection state, and operation state of the target shape 558 are monitored, and appropriate annotation, input, and management tasks are performed. Furthermore, the camera 204 can also detect the direction, angle, and speed of movement of the input mouse 102 and input pen 101.

[0072] The input pen 101 is equipped with an accelerometer 205. The output of the accelerometer 205 detects the speed and acceleration of the mouse movement by the user. It also detects the direction and speed of the input pen 101 movement by the user. The controller 103 monitors the marker state, selection state, and operation state of the target shape 558 by detecting or recognizing the speed, acceleration, and direction of the input pen 101, and applies them appropriately to annotation and input tasks.

[0073] For example, the movement speed, movement direction, input speed, and input direction related to the coordinate input of the input pen 101 are processed with acceleration, delay, etc. The same applies to the input mouse 102.

[0074] Controller 103 functions as a commander during image editing, a game console controller, etc., enabling easy movement, zooming, and scaling of images. It also appropriately controls annotation, input, and management tasks.

[0075] Furthermore, the controller 103 can, through AI learning, annotation display, or user actions, appropriately select and change / configure the control method, management method, annotation display, input mouse 102, input pen 101, and camera 204 control according to the user's preferences and physical characteristics.

[0076] The UI (User Interface) control program 105 controls the login program 106, workspace program 107, annotation program 108, and management program 109 to enable transitions between UI (User Interface) screens.

[0077] The UI (User Interface) control program 105 allows users authenticated by the login program 106 to change the screen display content of the workspace program 107, the screen display content of the annotation program 108, and the screen display content of the management program 109, and to permit or restrict screen transitions.

[0078] For example, if an authenticated user is an annotator, the screen display will be changed to one previously customized by that user to facilitate annotation work. Furthermore, access to the management program 109 will be restricted to prevent users from becoming aware of specialized and detailed settings.

[0079] The login program 106 consists of authentication processing (sign-in processing, login processing) for users of this system, registration processing (sign-up processing), and screen display processing for these processes. After the login program 106 authenticates the user, the workspace program 107 displays the workspace screen.

[0080] Workspace Program 107 is a program that allows you to create a workspace for each AI development project, task, or project, enabling independent management and operation of each. It allows you to create, edit, and delete workspaces. It also includes screen display processing to show information for each workspace.

[0081] When a user selects a workspace from the workspace list displayed by the workspace program 107, the annotation program 108 is launched, allowing the user to begin annotation work for that workspace. Additionally, only users with administrator privileges can launch the management program 109, which allows them to configure the parameters required for AI development.

[0082] AI development tasks are diverse and include tasks involving continuous values, object detection, semantic segmentation, sequence labeling, language generation, information retrieval, video, and audio data.

[0083] In one embodiment, the task of developing AI for semantic segmentation involves assigning a specific class to each pixel in an image. In other words, it divides the image into small sections and identifies what each region corresponds to. This is suitable for tasks that require detailed, pixel-level analysis, such as in medicine (tumor detection), autonomous driving (road analysis), and agriculture (crop classification).

[0084] The annotation program 108 performs the pixel-level annotation. This pixel-level annotation is carried out using the control and operation methods described for the input mouse 102 in Figure 16 and the input pen 101 in Figure 17.

[0085] You can select which applications can be launched for each workspace. Applications can be configured as a combination of one or more of the following: annotation, training, evaluation, and inference.

[0086] For example, in annotation and learning, the first step is to label the data. This may involve adding bounding boxes to objects in images or sentiment labels to text. It is also preferable to coordinate this with the input mouse 102, the operation buttons 551, wheel button 552, selection button 553, selection button 555, and display unit 556 of the input pan 101.

[0087] A bounding box, in this invention, is a sub-region that encloses an object in an image or video. In object detection, bounding boxes are used to estimate the position and classify objects within an image. The dimensions of the bounding box can be specified in terms of time or other non-spatial quantities.

[0088] Data preprocessing is performed (data is cleaned and normalized to prepare it for AI model learning, such as unifying image sizes, imputing missing data, and tokenizing text). Next, an AI model to be trained is selected or built.

[0089] Next, the model is trained using annotated data. The annotated data is divided into training data (for example, about 80%) and validation data (for example, about 20%). Then, the model's performance is evaluated using the validation data and confirmed using metrics such as accuracy, recall, and F-score. Finally, the model is retrained after hyperparameter tuning and data augmentation to improve its performance.

[0090] The annotation program 108 enables users to add correct labels and supplementary information to the image database (DB) 117 used for training the AI ​​model database (DB) 121. This allows the AI ​​model database (DB) 121 to understand the meaning and structure of the data and to make accurate predictions and classifications even for unknown data.

[0091] For example, in image recognition, objects in an image are labeled, and in natural language processing, the sentiment and intent of text are tagged. It is also preferable to coordinate this with the input mouse 102, the input pan 101's operation buttons 551, the wheel button 552, the selection button 553, the selection button 555, and the display unit 556. The quality of annotation is a crucial process because it directly affects the model's performance. Inaccurate annotations and inconsistent quality compromise the reliability of the training data, significantly impacting the model's accuracy and generalization performance. For example, this can include label mismatches or mislabeling. In semantic segmentation, a pixel representing a car might be mistakenly labeled as a "building."

[0092] For example, there may be inconsistencies in the granularity of annotation. Some data may be labeled finely (e.g., each finger), while other data may be labeled coarsely (e.g., one label for the entire hand).

[0093] For example, there might be an imbalance in the class distribution within the dataset. The "road" class might be heavily labeled, while the "pedestrian" class is hardly annotated at all. For example, annotation errors in complex areas can occur. The boundaries between overlapping objects in an image (e.g., a tree in the foreground and a building in the background) may not be accurately labeled.

[0094] For example, consider a change in the label scheme. Initially, the data distinguished between "cars," "buses," and "trucks," but at some point, they were unified into a single class called "vehicles." For example, there is ambiguity in annotation. When a person is only partially visible on screen, it is unclear whether to label it as "person" or "background."

[0095] For example, differences in environment and context can be a factor. Data collected in urban areas may be accurately annotated, while data from rural areas or at night may contain inaccurate labels. For example, there are limitations to annotation tools. Tools with auto-completion features automatically omit fine details.

[0096] Management program 109 is a program for setting various settings in a series of AI development processes using this system, such as the number of annotations, hyperparameters during training, and thresholds during evaluation.

[0097] Hyperparameters are parameters set to control the model's learning process and structure. Since the model cannot adjust them during training, it is necessary to select appropriate values ​​beforehand.

[0098] Examples include the learning rate, number of epochs, batch size, loss function, optimizer, regularization parameter, momentum, weight initialization, learning rate scheduling, number of hidden layers and units, kernel size (for CNNs), gradient clipping, data augmentation, and mini-batch shuffling. Choosing appropriate hyperparameters has a significant impact on model performance.

[0099] Generally, the administrator who manages the annotator's work sets these hyperparameters, but the annotator can also set them. Sometimes the annotator also acts as the administrator, and vice versa. The created AI model can be downloaded from the administrator screen and used in other applications. Examples of situations where an annotator would handle the setup include cases where the annotator themselves is as proficient in AI learning as an expert.

[0100] Setup program 110 is a program that performs the necessary tasks for a series of AI development projects using this system. It includes programs for image input for annotation, image preprocessing, exploratory data search, diversity sampling, sampling of sub-datasets for annotation tasks (datasets assigned to individual annotation tasks), sampling of example data, annotation and writing of notes to the example data, and various settings using management program 109.

[0101] For example, when using images as input for annotation, preprocessing steps include unifying image sizes, denoising, normalizing and adjusting images, cropping (trimming) unnecessary parts, filtering data, applying automatic region detection, format conversion, grouping image data, data augmentation, background removal, renaming and organizing images, pre-setting ROIs (Regions of Interest), color space conversion, and controlling dynamic ranges (for HDR images).

[0102] Exploratory data analysis is an initial analytical process aimed at gaining a deep understanding of data and discovering features and patterns. It uses statistical methods and visualization to grasp the structure, relationships, missing values, and inconsistencies of a dataset. This process involves reviewing the data overview, checking basic statistics, checking and handling missing values, checking data distribution, analyzing correlations, discovering features using visualization, checking categorical data, detecting outliers, analyzing target variables, and verifying data assumptions.

[0103] Diversity sampling reduces data redundancy and improves learning efficiency and model generalization performance by selecting samples with diverse features from the data, that is, by selecting highly representative samples. In the case of image data, this involves selecting images with different backgrounds, angles, and lighting conditions, or selecting prediction samples that the model is not confident in, or samples with features different from those previously selected, so that the training dataset better covers the entire target distribution.

[0104] The annotation task sub-datasets are randomly sampled or diversity-sampled to ensure representativeness and diversity. Annotation tasks are generated based on predefined rules and conditions, such as Annotation Task 1, Annotation Task 2, Annotation Task 3, ... Annotation Task N, and the design and assignment of each annotation task are adjusted to maintain efficiency and fairness. Annotation tasks are personalized by being adjusted to the user's needs and objectives. Requests and progress management for these annotation tasks are carried out by a dedicated program or system.

[0105] User Analysis Program 111 is designed to provide personalized annotation quality and annotation work management tailored to the individual user's personality and skill level. It includes programs for annotating sample data, measuring the degree of agreement with the ground truth answers and the measurement time, analyzing annotator data, personalizing annotation task 1 to suit the annotator, and requesting annotation task 1.

[0106] The example data provides examples of standard annotations and high-quality sample data, enabling users to properly manage and generate example data by applying specific rules and feedback results. The degree of agreement with this example data is calculated, for example, by comparing the annotator's annotation results with the ground truth answers and using evaluation metrics such as correlation coefficients and F1 scores.

[0107] The active learning program 112 is a program that implements an approach to efficiently build high-performance models by "actively selecting" the data that the AI ​​model database (DB) 121 uses for training. It is particularly effective when labeling is costly, for example, when there is a large amount of data but insufficient accurate labels (such as those assigned manually). The main flow of active learning is: 1. Initial model construction (train the model with a small amount of initial labeled data). 2. Sample selection (select useful data for the model from the unlabeled data and label it). 3. Label the selected data (usually handled by experts or annotators). 4. Retrain the model (retrain the model using newly labeled data).

[0108] 5. This is repeated (repeat steps 2 to 4 as needed). The program includes AI model training, AI model evaluation, AI deployment, AI inference initiation, AI model inference, inference result analysis, uncertainty score calculation, sampling based on the score, personalization of annotation task N, and requesting annotation task N.

[0109] AI learning program 113 is a program that implements a process to improve the ability to learn knowledge and patterns from data and use that knowledge to perform reasoning tasks (such as classification and identification). In this process, by analyzing data and finding regularities and features, it becomes possible to make predictions and judgments about new data.

[0110] Finding regularities and features is a crucial process for models to understand patterns and relationships hidden behind data and to make predictions and classifications. For example, in pattern recognition in image recognition, this involves capturing features of cats (round ears, whiskers, eye shape, etc.) in different images to find features that distinguish cats from other animals (such as dogs), or in convolutional neural networks (CNNs), which learn specific edges, textures, and shapes in images to discover regularities that distinguish cats from other animals.

[0111] Convolutional Neural Networks (CNNs) are a type of deep learning and a machine learning algorithm specifically designed for recognizing images, audio, and time-series data. CNNs mimic neurons, the nerve cells in the human brain, and have a structure built on stacks of layers, such as convolutional and pooling layers. CNNs are particularly effective at detecting patterns within images to recognize objects, classes, and categories.

[0112] The core of AI learning is a technique called "machine learning," which also includes advanced methods such as deep learning. Deep learning is a method that uses multi-layered neural networks to automatically extract and learn features from data, and it is preferable that it can make highly accurate predictions and classifications by utilizing large amounts of data and computing resources.

[0113] The AI ​​Evaluation Program 114 is a program that measures the performance, reliability, and safety of AI systems and models, and implements processes for improving them. This process is crucial for determining how effective an AI model is for a specific task and for verifying whether it functions properly in the real world.

[0114] For example, the criteria for judging effectiveness in an image segmentation task relate to how accurately the model can divide different regions or objects within an image. Evaluation metrics include accuracy, the Jackard coefficient (IoU: Intersection over Union), and the Dice coefficient.

[0115] AI Deployment Program 115 is a program that enables the process of deploying trained AI models for use in real-world environments and integrating them into operations. To utilize AI models in real-world applications and systems, it's not enough to simply develop the models; a mechanism is needed to efficiently execute, monitor, and maintain them.

[0116] Mechanisms for monitoring and maintenance include, for example, model performance monitoring, data drift detection, model retraining, logging and alerting systems, model version control, security and access management, user feedback systems, and continuous integration (CI) / continuous delivery (CD).

[0117] The AI ​​inference program 116 is a program that implements a process of making predictions and decisions based on new (unknown) data using a pre-trained AI model. This process is a major task in the operational phase of the AI ​​system and is a phase different from training (the process of building a model). The AI ​​inference program aims to generate appropriate outputs (results) for unknown input data using the model obtained as a result of training.

[0118] To generate appropriate outputs (results) for unknown input data, a combination of methods such as regularization and cross-validation is necessary to prevent the model from overfitting to the training data. Figure 2 is a system configuration diagram and explanatory diagram of the present invention under a local network environment.

[0119] The main difference from Figure 1 is the presence of an information processing device B200 equipped with a web browser program 201, to which a display 100, pen 101, mouse 102, camera 204, and controller 103 are connected.

[0120] The web browser program 201 renders web pages using HTML, CSS, JavaScript, etc., and displays them visually to the user. Web browsers enable comfortable and safe internet use. Rendering is the process of processing data to generate images, videos, audio, etc.

[0121] The B200 information processing unit is a desktop computer equipped with an operating system (OS), internet connectivity, storage capabilities, input devices (keyboard and mouse), graphics and display functions, peripheral device connectivity, and software execution capabilities, making it a highly versatile device capable of handling a variety of tasks.

[0122] Information processing unit B200 and information processing unit A104 are connected via a local network or the internet. Each program and data within information processing unit A104 can be accessed via the web browser program 201.

[0123] Even with low CPU and memory performance, the information processing device B200 can utilize the processing results of the high-performance information processing device A104 through the web browser program 201. Furthermore, using the web browser program 201 allows multiple information processing devices B200 to share the processing results of information processing device A104. The information processing device B200 may be integrated with a display and camera. For example, a tablet device could be used. Figure 3 is a system configuration diagram and explanatory diagram of the present invention under a cloud internet environment.

[0124] The main difference from Figure 2 is that the network to which the information processing device B200 and the information processing device A104 are connected is not a local network, but rather the cloud 300, the internet 300, and the cloud-internet 300. Each program and data within the information processing device A104 can be accessed or worked on via the web browser program 201 through the cloud-internet 300.

[0125] Cloud Internet 300 is a service that provides distributed resources accessible via the internet for data storage, processing, and application execution. Administrators, acting as system administrators, access Cloud Internet 300 to manage and monitor resources. Users, as general users, access Cloud Internet 300 and utilize the services they need. Billing is often based on the amount or duration of resources used, for example, by the time spent using storage capacity or computing resources.

[0126] Figure 4 is an explanatory diagram of the UI (User Interface) control program. It shows how the UI control program controls the login program 106, workspace program 107, annotation program 108, and management program 109 to realize the transitions between UI screens. The configuration is designed so that when annotators and administrators divide the work, they can perform the necessary tasks in the fewest steps possible from their respective positions.

[0127] The login program is used to authenticate users and control their access to the system. It handles sign-in and sign-up processes, and after successful authentication, transitions the user to the next program according to their permissions. In conjunction with other programs, it sets access permissions for workspace programs and annotation programs.

[0128] The Workspace program allows users to create, edit, and delete workspaces within AI development projects. When a user selects a workspace, they can begin annotation and management tasks associated with that workspace. Workspace information is integrated with other programs (e.g., annotation and management programs) to ensure that tasks and settings appropriate for the workspace are reflected. Application launches and configurations are also based on the workspace.

[0129] An annotation program enables the task of labeling images and data. After the annotator selects a workspace, it provides an interface for performing annotation tasks related to that workspace. The annotation program plays a role in adding correct labels and supplementary information to datasets for AI model training, and customizes the screen to allow the annotator to work efficiently. In conjunction with other programs, the annotation results are used as data for training and evaluation, contributing to the improvement of the overall system accuracy.

[0130] The management program handles the overall system configuration and management. Administrators manage various parameters necessary for AI development, including setting the number of annotations, hyperparameters during training, and evaluation criteria. The settings configured by the administrator affect the workspace and annotation process, ensuring appropriate data flow and work processes.

[0131] The management program manages user permissions and work progress through integration with other programs (login program, workspace program, annotation program). Additionally, created AI models can be downloaded and exported for use in other applications.

[0132] Figure 5 is an explanatory diagram showing the components displayed on the screen by the workspace program. The workspace program includes various panels, toolbars, and menus to provide a user-friendly interface. Key components include the workspace name, workspace creation time display, status of each process, and progress.

[0133] The program operates in real time based on user input, reflecting task updates and progress. Furthermore, it includes integration capabilities with other programs, such as database and storage connections, enabling seamless import and export of various information and files.

[0134] Furthermore, a dashboard is provided to visualize progress. The dashboard uses graphs and charts to visually represent task completion status and important notifications, allowing users to see them at a glance.

[0135] Users can directly monitor their work progress from the dashboard and take necessary actions quickly. The display menu also includes options such as logout, log output, file management, settings changes, user account management, notification settings, and help, each designed for easy access using icons and dropdown menus.

[0136] Figures 6 and 7 are explanatory diagrams illustrating the components displayed on the screen by the annotation program. These components provide a personalized user interface (UI) that can be customized to suit each annotator's work style, in order to maintain the annotator's concentration, motivation, and self-efficacy.

[0137] These customization histories are saved. For example, the display screen shown in Figure 6 shows the toolbox 202 on the right side of the screen and the file display 203 on the left side. The tools are set to the left and the work progress display and file list on the right side to accommodate left-handed users. For right-handed users, for example, the toolbox 202 is displayed on the left side and the file display 203 is displayed on the right side of the screen.

[0138] In Figure 7, the file display 203 is shown on the right side of the screen, and the toolbox 202 is shown on the left side of the screen. The present invention allows the screen display 601 to be changed according to the dominant hand. Furthermore, the display position can be changed or set according to the user's settings. In addition, the presence or absence of buttons, their arrangement, the number of buttons, and the types of buttons can also be customized.

[0139] For example, this includes features such as rearranging the layout of tool buttons and showing / hiding tool buttons as needed. For instance, you can make efficient use of screen space by displaying only frequently used tools (such as the pen tool or selection tool) and grouping other tools in a dropdown menu.

[0140] The presentation also illustrates the ability to rearrange tool buttons, allowing users to create a user-optimized workflow. By placing frequently used tools on the left and less frequently used tools on the far right, access can be streamlined.

[0141] The documentation also provides examples of how to adjust button sizes, allowing users to change button sizes to suit their visibility and ease of use. Larger buttons improve visibility, while smaller buttons can be selected for limited workspace to maximize screen space.

[0142] Furthermore, the system demonstrates the ability to customize shortcut buttons, allowing users to freely customize shortcut buttons to quickly execute specific actions. For example, assigning frequently performed operations during annotation work to buttons can improve work speed.

[0143] For example, grouped tool buttons can be organized by their related functions and arranged as expandable buttons. Tools such as "pen," "eraser," and "line drawing" can be combined into a single button, and menus can be expanded as needed, thus organizing the screen.

[0144] Furthermore, examples include dynamically changing button displays and features where the content and placement of buttons dynamically change according to the task and its progress. By setting buttons to appear only when a specific task is underway, unnecessary buttons can be eliminated, providing an interface that allows users to concentrate on their work.

[0145] Furthermore, the display of context menus will be customized based on the action taken, and the context menus that appear when specific tools or objects are selected will be made customizable. This will allow users to quickly access the functions most relevant to the tool they are using.

[0146] Furthermore, the system will allow users to change button colors and themes, providing a visually conducive environment for focused work. For users working in dark mode, the button colors will be adjusted to be easier on the eyes.

[0147] The work progress display 602 shows the progress of annotation. While the original plan might involve annotating around 4000 images in total, to maintain the annotator's self-efficacy, the amount required for a single annotation (e.g., 10 images) can be arbitrarily set. Alternatively, only the amount of a sub-dataset (e.g., 10 images) can be displayed to indicate the completion status.

[0148] Additionally, a scoring system can be displayed to improve annotators' motivation by awarding points for accurate and efficient annotation work, and by introducing ranking and reward systems.

[0149] Furthermore, a work completion ranking is provided as an example. The ranking is displayed based on the number of annotations completed and the progress of the annotator. The annotator who has completed the most annotations can be ranked higher. This ranking visualizes the speed and efficiency of work and facilitates comparison with other annotators.

[0150] Furthermore, an accuracy ranking is provided as an example. The ranking is displayed based on how accurately annotators performed their annotations. The fewer incorrect annotations, the higher the score, emphasizing quality. This provides an incentive for annotators to perform accurate annotations.

[0151] Furthermore, an efficiency ranking is provided as an example. The ranking is based on annotation speed (number of annotations per hour). It evaluates how efficiently the work was completed within the given time, and displays highly efficient annotators at the top. This serves as an incentive to improve work speed.

[0152] For example, an error rate ranking is presented. This ranking displays data based on the number of errors and corrections that occurred in annotations. The goal is to minimize errors, with lower error rates resulting in higher ratings. This ranking serves as an incentive for quality improvement.

[0153] Furthermore, an improvement ranking is provided as an example. This ranking displays an evaluation of how much improvement has been made since the previous work. Focusing on self-improvement, annotators who have improved the accuracy and efficiency of their annotations compared to the previous work are displayed at the top of the ranking. This is effective in boosting self-efficacy.

[0154] Furthermore, a motivation ranking is provided as an example. The ranking is displayed based on regular work progress and continuous participation. For instance, annotators who complete many tasks within a certain period and work continuously can receive special recognition. This ranking encourages continued active participation.

[0155] Team / group rankings are also provided as an example. When multiple annotators work together, team-based rankings can be implemented. Displaying performance within the team and clearly showing individual contributions can promote collaboration throughout the group.

[0156] Furthermore, daily, weekly, and monthly rankings are provided as examples. Annotators' progress is ranked in short-term increments. By displaying daily, weekly, and monthly rankings, short-term results are made easier to see, and progress towards goals can be visualized. This encourages users to prioritize short-term results.

[0157] Compensation systems can include, for example, monetary rewards, point systems, gifts and in-kind rewards, status and special recognition (rank-based positions and responsibilities), flexible working hours, and special leave, all of which serve as strong incentives for annotators to perform efficient and high-quality work. By combining various types of rewards, it is possible to promote not only monetary rewards but also personal growth and a sense of social accomplishment, thereby maintaining motivation in the long term. Additionally, to enhance the sense of accomplishment, badges or medals can be displayed when a certain amount of corrections or difficult errors are fixed. The above functions, controls, and displays can be set or changed individually. Furthermore, multiple functions, controls, and displays can be set or changed in conjunction with or in combination with each other. File display 203 shows a list of images in this workspace that the user should annotate.

[0158] There is an annotation incomplete display function that, when checked, shows only images that have not yet been annotated. There is also an all-images display function that, when checked, shows all images, both those that have been annotated and those that have not.

[0159] Switching between displaying incomplete annotations and a full image list is done by monitoring the state of a checkbox and detecting whether it is on or off. When the checkbox is on, only incomplete images are displayed; when it is off, all images are displayed. There is also an annotation completion setting function that allows you to check off images that have been annotated.

[0160] Toolbox 202 consists of tools that implement the following functions: "Label Settings" for setting labels, "Add by Enclosing" for adding a specified label to an area enclosed with a stylus, "Delete by Enclosing" for deleting an area enclosed with a stylus, "Add by Painting (Pen)" for adding a specified label to an area painted with a stylus, "Delete by Painting (Eraser)" for deleting an area painted with a stylus, "Set Paint Size" for setting the size of the circles used in "Add by Painting (Pen)" and "Delete by Painting (Eraser)", "Delete All Areas by Batch" for deleting all areas painted with a stylus, and "Save" for temporarily saving the annotation results.

[0161] The "Label Setting" function allows you to assign appropriate labels to images to be annotated. Users can select or add new label types and apply them to specific areas within the image. This function enables clear and consistent classification of objects and areas within an image. Labels can be used for various purposes, such as object recognition, scene classification, and importance assessment, depending on the set objective.

[0162] The "Enclose and Add" feature allows users to use input devices such as a mouse or stylus to enclose an area within an image and add a specified label to that area. By enclosing an area, objects or regions can be clearly identified and labeled, making it particularly effective for annotation tasks such as object detection and region segmentation. Labels are automatically applied to the enclosed area, enabling efficient annotation work.

[0163] The "Select and Delete" function allows users to delete areas they select using an input device such as a mouse or stylus. This is useful when labels have been applied incorrectly or when you want to remove unwanted areas. By undoing the selected area, annotation results can be corrected, leading to more accurate work. This function makes it easy to remove unwanted labels and areas.

[0164] The "Paint and Add (Pen)" function allows users to paint specific areas on an image using input devices such as a mouse or stylus, and then add a designated label to that area. Because the painting action allows for more precise area selection, it is useful when annotating objects or areas with freeform shapes. It is particularly effective for annotating objects with unclear boundaries or complex contours.

[0165] The "Paint and Delete (Eraser)" function allows users to delete painted areas using input devices such as a mouse or stylus. It's used to correct accidentally painted areas or unwanted parts, improving the accuracy of annotations. The ability to intuitively specify the area to delete makes correction work easy.

[0166] The "Fill Size Setting" feature allows users to adjust the size of the circles used for "Fill and Add (Pen)" and "Fill and Delete (Eraser)" operations using input devices such as a mouse or stylus. Users can change the circle size as needed for their work, enabling them to efficiently perform finer corrections or large-area filling tasks. This allows workers to work with the optimal size for each annotation target, improving workflow flexibility.

[0167] The "Batch Delete Area" function allows users to delete all areas they have colored using an input device such as a mouse or stylus in one go. This function is extremely useful when you need to delete multiple areas at once or when you have accidentally colored a large area. It saves you the trouble of deleting each area individually and allows you to correct mistakes in a short amount of time.

[0168] The "Save" function allows you to temporarily save the results of your annotation work in progress. It's useful for interrupting your work or ensuring you keep track of your progress. Since saved data can be resumed later, you can handle interruptions and unexpected work, allowing you to proceed with confidence.

[0169] Split display 604 is a function that divides the target image into multiple sections for display. This allows users and annotators to visually separate the areas to be worked on, enabling them to work more efficiently.

[0170] Annotation result 605 indicates the areas where annotation has been completed. This display allows you to see at a glance which areas have been processed and which areas are still pending, enabling you to manage the progress of your work.

[0171] The "Start Learning" button 607 has the function of starting learning once annotation is complete. The size of the "Start Learning" button 607 is set to be small to make it difficult to press by mistake or by users who do not understand AI learning, and it is designed to blend in with the background to make it difficult to press easily.

[0172] For example, the "Start Learning" button 607 should have the same background color tone, but with only its brightness changed. The size of the "Start Learning" button 607 should be less than half the size of the checkbox.

[0173] The learning start button 607 should not only start learning, but also effectively and usefully include functions for stopping and completing learning. Furthermore, it would be effective and useful to change the button's display, image, character, color, and size depending on the learning status.

[0174] A camera 204 is installed or positioned on the display. The camera 204 monitors the user's gaze, detecting movement of the gaze, direction, direction of movement, angle of movement, speed of movement, and reactions, and can evaluate the user's work state and fatigue level, prompting breaks or adjustments to the work pace at appropriate times.

[0175] Figures 7, 8, and 9 are diagrams illustrating the process of dividing, displaying, and processing an image. The divided display 604 is a function that divides and displays an image to clearly show which area to start working on, without the user getting lost when beginning their work.

[0176] Figures 8 and 9 show examples of dividing the target image into four sections. Dividing the target image reduces the amount of information the worker has to see at once, making it easier to concentrate on the work. Also, for annotation beginners, it eliminates the need to wonder where to start with annotation. Furthermore, the application of the split-screen display function becomes easier, improving the accuracy of annotation and corrections.

[0177] The split-screen display function includes a split-screen display settings function that allows users to configure how images are displayed when they are split. It also includes a split-screen display location change function that allows users to change the location of the split-screen display when an image is split. The split-screen display location change function allows users to freely switch between the displayed split areas and adjust the display order of the areas to improve work efficiency.

[0178] Figure 9 shows the target image divided into four sections: 1, 2, 3, and 4. Figure 9(a) shows section 1 of the target image. Figure 9(b) shows section 2 of the target image. Figure 9(c) shows section 3 of the target image. Figure 9(d) shows section 4 of the target image.

[0179] For example, the user adds annotations to unprocessed areas of a divided image. This ensures that the overall annotation progresses smoothly. The administrator checks the annotation progress of the divided image and provides appropriate feedback to the worker. In this way, the accuracy and speed of the work can be managed. The controller 103 can perform annotation processing on the divided image in a batch. This operation streamlines the process and reduces processing time.

[0180] The embodiment shown in Figure 9 is an example in which a single target figure is divided into multiple parts, but the present invention is not limited to this. For example, as shown in Figure 24(b), it can also be applied when a single object is divided into multiple cross-sections. The images (figures) of the multiple cross-sections are often similar, and the drawing process, annotation display, and annotation operation are also similar. Therefore, it is preferable to perform the drawing process instruction, annotation display, annotation operation process, and AI learning process by considering the cross-sections together.

[0181] It has a display reset function that resets the image scaling and adjusts it appropriately to fit the screen window. It also has an image movement function that moves the image up, down, left, and right.

[0182] The display reset function returns the displayed image to its initial scale and position. This function allows users to revert any accidentally changed scale or position and start over from the beginning.

[0183] Furthermore, the image movement function allows users to move the displayed image to any position on the screen. This function allows users to adjust the display position of the image to check its details and perform annotation work, thereby improving work efficiency.

[0184] The annotation result display / hide function 605 is designed to avoid readability issues during annotation and / or modification by switching the display / hide of AI inference and manually colored elements on a label-by-label basis. Furthermore, the annotation result display / hide function 605 can be changed and configured based on monitoring data from camera 204. Additionally, the annotation result display / hide function 605 can be changed and configured based on the position, movement, and operation status of the input mouse 102 and input pen 101.

[0185] AI inference is an inference process that automatically predicts annotations based on machine learning algorithms and identifies regions based on specified labels. By using this inference, manual annotation work can be complemented, and the work can be carried out more efficiently. Furthermore, by displaying the inference results, the consistency between manual annotations and the inference results can be checked, making it easier to see which parts need to be changed during revisions. In addition, by toggling the display on and off, it is possible to avoid clutter and confusion during revisions.

[0186] Annotation result display / hide 605 allows you to independently switch the display / hide of annotated areas by label. By label, this means setting the display / hide based on the type of object, such as "person," "car," or "animal," or even more detailed classifications.

[0187] This allows users to display only the necessary labels and work more efficiently. Furthermore, annotation results can be displayed while annotation work is in progress and hidden when the work is completed and the user switches to correction mode, providing an appropriate display state based on the progress of the work. This can also be changed and configured based on the position, movement, and operation status of the input mouse 102 and input pen 101.

[0188] As mentioned above, Toolbox 606 consists of functions such as "Label Settings" for setting labels, "Add by Enclosing" for adding a specified label to an area enclosed by the stylus, "Delete by Enclosing" for deleting an area enclosed by the stylus, "Add by Painting (Pen)" for adding a specified label to an area painted with the stylus, "Delete by Painting (Eraser)" for deleting an area painted with stylus 101 (input pen 101), "Setting the size of the circle for "Add by Painting (Pen)" and "Delete by Painting (Eraser)", "Delete All Areas at Once" for deleting all areas painted with stylus 101, and "Save" for temporarily saving the annotation results.

[0189] These tools utilize a touch panel to enable editing and modification through an intuitive interface. Furthermore, settings can be changed and configured based on the position, movement, and operation status of the input mouse 102 and input pen 101. In this case, it is preferable to link these settings with the operation buttons 551, wheel button 552, selection button 553, selection button 555, and display unit 556 of the input mouse 102 and input pan 101.

[0190] By adopting the touch panel 100, users can perform annotation work intuitively using touch, swipe, and pinch-in / pinch-out gestures, without relying on physical buttons or mouse operations. Furthermore, by maximizing the real-time responsiveness and operational flexibility of the annotation work of the present invention, work efficiency is improved and the burden on the user is reduced.

[0191] By linking with camera 204, the input process automatically recognizes the images to be annotated, allowing users to proceed with annotation work intuitively. The output results from camera 204 can be processed, and based on that information, annotation processing and administrator processing can be arbitrarily or automatically changed and configured.

[0192] The learning start button 607 is a function for launching the active learning program 112 based on the annotated data. Even users without specialized knowledge can perform these complex processes with a single click. The present invention is designed to allow users to intuitively start and configure the learning process easily. This enables users to easily proceed with annotation work and obtain results, even without special knowledge or operational skills.

[0193] The learning start button 607 is displayed in a smaller size compared to the other operation buttons. This is intended to encourage careful operation to prevent accidental activation of the learning process. This design reduces accidental activation and minimizes the risk of users starting learning at an unintended time.

[0194] The learning start button 607 changes only the brightness while keeping the background color the same. Alternatively, it changes only the chromaticity while keeping the brightness the same as the background. It can also be set to display symbols, characters, etc. A function is included to notify users of appropriate break times in order to adjust their work pace. To notify users of appropriate break times, the camera observes the user's gaze and automatically determines that a break is necessary if the user is working for a long time or if their eyes are strained. This notification prevents users from forgetting to take breaks due to excessive concentration and helps maintain a healthy work pace. Furthermore, if the user works for a long time continuously, the notification interval is shortened to allow for appropriate rest, thus considering the user's work efficiency and health. Figure 10 is an explanatory diagram and flowchart illustrating the administrator's setup process and setup method. First, the administrator authenticates with the system and logs in as a user with administrator privileges (S00).

[0195] Next, the administrator creates a workspace for each AI development project, task, or case, and inputs the target image sets into the system (S01). An image set is a collection of images to be annotated, and may include, for example, images of various objects or images taken under different conditions. Furthermore, an image set contains data necessary for AI training, and images with the same categories or characteristics are grouped together.

[0196] The image dataset is preprocessed according to the characteristics of the task and the requirements of the system resources. Preprocessing may include, for example, unifying the image resolution or removing unwanted noise. If the characteristics of the task involve identifying a specific object, preprocessing may involve highlighting and filtering areas related to that object. When system resources are high, the image size may be reduced or the data compressed to reduce the processing load. Image data preprocessing is extremely important; proper preprocessing improves model performance and makes training more efficient (S02). The purpose of the pretreatment is exemplified by the following in one embodiment.

[0197] Data consistency is ensured by standardizing image size and format so that the model can learn efficiently. Data augmentation is performed by reducing noise (unnecessary information (noise) in the image and highlighting important features) and increasing the training data to improve the generalization performance of the model. Computational efficiency is improved by removing unnecessary information and saving computational resources. Examples include resizing, normalization, color space conversion, noise reduction, histogram flattening, centering, data augmentation, binarization, background removal, sharpening, and color correction.

[0198] In particular, performing noise reduction preprocessing removes unnecessary information from the image, allowing the model to focus more easily on important features. This improves training speed and enables more accurate predictions. Furthermore, data augmentation increases the diversity of the training data, improving the generalization performance of the model.

[0199] Exploratory data analysis is a process that enables a deep understanding of the characteristics and relationships of data by visualizing, summarizing, and discovering patterns and features in the input data (S03).

[0200] In particular, understanding the data overview (checking basic statistics such as mean, median, and variance, and examining the distribution) is a crucial step in the fundamentals of data analysis. Furthermore, detecting outliers and missing values ​​(identifying inconsistencies, missing values, and outliers in the data) is an important task that improves data quality and leads to the construction of accurate models. In addition, understanding the relationships between variables (analyzing the correlations between features and between features and target variables) is useful for designing predictive models and selecting important features.

[0201] If the data is biased, failing to consider methods to correct for the bias could lead to incorrect conclusions. Furthermore, with large datasets, it's necessary to implement appropriate sampling and distributed processing to enable efficient analysis.

[0202] In particular, conducting exploratory data analysis aimed at detecting outliers and missing values ​​improves the reliability of the dataset and enhances the predictive accuracy of the model. Furthermore, accurately understanding the relationships between variables allows for the selection of important features, significantly improving the efficiency of analysis and model building. Diversity sampling is a method for selecting as many different types of data as possible based on the results of the exploratory data analysis described above (S04). This method aims to improve the efficiency of model training and enhance generalization performance by ensuring sample diversity.

[0203] Diversity sampling includes the following methods: model-based outlier sampling (actively selecting outliers that the model has determined are difficult to learn from), cluster-based sampling (clustering data and selecting representative samples from each cluster), representative point sampling (creating diverse datasets by selecting points that represent the overall data distribution), and real-world diversity sampling (eliminating biases in specific categories or attributes and constructing datasets that reflect actual usage scenarios).

[0204] In particular, when data is biased towards specific categories or classes, diversity sampling allows for more balanced model training and improved prediction accuracy. Furthermore, it enhances adaptability to complex real-world patterns, enabling the construction of more versatile and reliable models.

[0205] The results of diversity sampling are used as a sub-dataset for annotation task 1 (the first annotation task performed by the user in the workspace) (S05). From the sub-dataset 705, one or more images are sampled to be used as example data (S06).

[0206] The administrator will perform annotation work on this example data to serve as a model for annotation, and will also record any notes to be taken during the annotation process (S07). The annotation results and notes are stored in the example database (DB) 119.

[0207] In Annotation Task 2 (the second annotation task performed by the user in the workspace), sampling is performed again depending on the task's progress and any newly added sub-datasets, and a sub-dataset is created based on the results. From this sub-dataset, example data is selected, the administrator performs the annotation work, adds notes, and stores it in the example database. Note that the selection of example data, the annotation work by the administrator, and the inclusion of notes in annotation task 2 can be omitted.

[0208] Furthermore, in annotation task N (the Nth annotation task performed by the user in the workspace), the same process is repeated, involving the selection of example data based on a new sub-dataset, annotation by the administrator, and the writing of notes. This ensures that appropriate example data and notes are accumulated for each task, establishing a system that supports users in performing annotation work efficiently and accurately.

[0209] Note that the selection of example data, the annotation work by the administrator, and the inclusion of notes in annotation task N can be omitted. In the administrator settings, it is possible to configure in detail who will be assigned to annotation task 1 and how the work will be assigned (S08). The administrator can set the workload for annotation task 1 (for example, 10 sheets), thereby limiting the amount of work an annotator has to do and lowering the barrier to entry for annotators. This reduces the workload and makes it possible to proceed with annotation work efficiently.

[0210] Furthermore, it is possible to adjust the difficulty level, assigning easy tasks to annotation beginners and more difficult tasks to experienced annotators. This feature ensures appropriate task assignments according to the annotator's skill level, contributing to an overall improvement in work quality. These settings are stored in the administrator settings database (DB) 124 and can be referenced and updated as needed.

[0211] Similarly, in annotation task 2, the workload and difficulty can be adjusted according to the characteristics of the target data and the annotator's skill level. The same process is applied to annotation task N, and the administrator settings are reflected in all tasks.

[0212] Furthermore, the annotations and notes (S07) are stored in the example database (DB) 119. This example database (DB) 119 is used by annotators to refer to it during their work, allowing them to learn proper annotation techniques and improve the accuracy of their work.

[0213] On the other hand, administrator settings (S08) are stored in the administrator settings database (DB) 124. This administrator settings database (DB) 124 manages the assignment criteria, workload, and difficulty adjustment criteria for each task, and is used for future task design and operational efficiency improvements. This makes it possible to continuously optimize the workflow and ensure the smooth progress of the entire project.

[0214] Figure 11 shows an example screen in the administrator setup process where the administrator performs annotation work on sample data to serve as an annotation example, and also writes notes (S07) regarding the annotation process. The notes (S07) can be any text and can be placed at any position on the screen.

[0215] Notes are automatically displayed when the annotator begins work and serve as a reference. The annotator can also manually redisplay the notes if needed. To hide them, click the "Hide" button on the screen. A setting can be applied to automatically hide the notes after a certain period of time. Furthermore, the display / hide settings can be controlled using the operation buttons 551, wheel button 552, selection button 553, and selection button 555 on the input mouse 102 and input pan 101.

[0216] This allows annotators to efficiently acquire necessary information, which is expected to prevent confusion and errors during the work process. Furthermore, the ability to flexibly customize the way notes are written and displayed makes it adaptable to different tasks and datasets.

[0217] Figure 11 is an explanatory diagram of how to add notes to the example data. In Figure 11, a soldered connecting copper foil 534 is formed or configured on the printed circuit board 535, and lead pins 537 of electronic components, etc., are connected to the connecting copper foil 534 with solder 538.

[0218] Voids 536 and cracks 539 occur in solder 538. Voids 536 and cracks 539 can be detected by imaging with an X-ray CT scanner, etc. However, many of the imaged voids 536 and cracks 539 are unclear, so it is necessary to create training data manually and accumulate data through AI processing.

[0219] During the administrator's setup process, the user or worker will see an example screen where the administrator performs annotation work on sample data to serve as an annotation guide, and also displays notes (532a, 532b) to be noted during the annotation process. These notes can be any text and placed at any position on the screen.

[0220] The notes are automatically displayed when the annotator begins work and serve as a reference for the work. Based on the notes given during annotation, the user or worker should fill in the voids 536a, etc. 533.

[0221] The precautions can be different for each void 536, and the user or worker will perform the work based on the precautions given during annotation. Different precautions will also be displayed for other objects such as cracks 539, and the work will be performed accordingly. The work can also be modified based on information from cameras 204, etc. The objects (target shapes) include not only voids 536 and cracks 539, but also lead pins 537, connecting copper foil 534, printed circuit boards 535, etc.

[0222] Furthermore, annotation display and annotation behavior can be changed and configured based on the object (target shape). The annotator analysis process can also be changed and configured based on the object (target shape). Additionally, personalized analysis can be changed and configured based on the object (target shape).

[0223] The annotator can manually redisplay the notes as needed. To hide them, click the "Hide" button on the screen. A setting can be applied to automatically hide them after a certain period of time. Additionally, the display and hide settings can be controlled using the operation buttons 551, wheel button 552, selection button 553, and selection button 555 on the input mouse 102 and input pan 101.

[0224] Figure 12 is an explanatory diagram and flowchart of the annotator analysis process of the present invention. Figure 12 shows the annotator analysis process after the administrator setup process is completed and before the user who will become the annotator starts the annotation work. User authentication is performed during login (S21). The user authenticates with the system and logs in as a user with user privileges.

[0225] After logging in, the user selects an existing workspace (S22). If personalized analysis has not been performed in the workspace selected by the user, personalized analysis is performed first (S23).

[0226] Personalized analysis is a method of providing optimized work content and settings for individual users based on their work history, skill level, work tendencies (personality and annotation style), and past work quality and task completion rates. This analysis enables efficient and effective task assignment and interface settings, which is expected to reduce user burden and improve work efficiency.

[0227] If the personalization analysis result is YES (implemented), the existing personalization settings will be used as is, allowing the user to smoothly start the assigned annotation task N. This eliminates unnecessary effort and allows the task to progress quickly.

[0228] On the other hand, if the result of the personalization analysis is NO (not implemented), a new personalization analysis will be conducted to generate annotation tasks that are optimal for the user. In this process, adjustments will be made as needed based on the user's current skills and the characteristics of the tasks, improving the work experience in the workspace. These steps provide an optimized environment for each user, resulting in improved work efficiency and highly accurate annotation.

[0229] From the example database (DB) 119 created during the administrator setup process, data corresponding to the selected workspace is extracted and presented to the user. The user then performs annotation (S24).

[0230] The annotator's work accuracy (agreement) is evaluated by comparing the annotated data with the ground truth answers (S25). Work accuracy (agreement) can be measured using metrics such as Intersection over Union (IoU). This is an evaluation metric that shows how much the annotated area overlaps with the area of ​​the true answers; a higher IoU value indicates higher annotation accuracy. For example, label agreement rate can be used.

[0231] This measure the percentage of annotations that match the true labels, indicating how accurately classification is performed. For example, pixel precision can be used. By calculating the percentage of all pixels in an image that are correctly annotated, the precision of the details can be checked.

[0232] The results of accuracy assessments based on these metrics are provided to users as feedback. Furthermore, opportunities for additional guidance and retraining are provided as needed to support the improvement of annotators' skills. This process ensures the overall quality and consistency of annotations. Additionally, work time and accuracy per label are measured to conduct annotator characteristics and data quality analysis (S26). Based on the analysis results, annotation task N is customized for each user (S27).

[0233] For users who have high work accuracy (agreement) but have spent a long time on the task, this is an extremely difficult task. Therefore, customize the task to reduce the number of annotations. On the other hand, for users with low work accuracy (agreement) but short work time, the system will be customized to assign them tasks of lower difficulty.

[0234] Furthermore, the Annotation Task 1 Request (S28) is the process of formally requesting the user who will be the annotator responsible for Annotation Task 1 to perform the specific annotation work for Annotation Task 1. This request includes information about the task content, objectives, notes, and deadline.

[0235] After the annotation task 1 request (S28), a decision is made (S31) as to whether the task can be started. This decision is a step to confirm whether the annotator has accepted the task and is ready to start working.

[0236] If "Task Start OK" is answered with YES, Annotation Task 1 officially begins. The annotator works on the task within the system and proceeds with the work. During this time, the system records the progress and provides support until the task is completed.

[0237] If the task start status is "NO" instead of "OK," the reason why the work cannot be started is recorded. For example, this could be due to unclear task content, excessive workload, or a short deadline. In this case, the manager will implement a process to adjust the task content or review the annotator's working conditions. This creates a system that maximizes the quality and efficiency of annotation work, ensuring that tasks proceed smoothly.

[0238] A personalized annotation task N request is sent to the user, who becomes available to perform the annotation, and the user then executes this annotation task N (S32).

[0239] Annotation review is a function that allows annotators to notify their administrators and request a review in order to check whether their annotation results are correct or incorrect and to consult on how to improve their annotations (S33). This function gives annotators an opportunity to evaluate and improve their work.

[0240] Furthermore, when working on tasks collaboratively with other annotators, the system includes a feature that allows for mutual feedback. This enables users to review annotation results from multiple perspectives and improve quality. In addition, a chat function is available for discussing particularly difficult corrections with other annotators and administrators.

[0241] For example, if the annotation results are ambiguous, you can use the chat function to send a specific question to the administrator, such as "Is the labeling for region A in this image appropriate?" This allows you to receive appropriate advice from the administrator or other annotators.

[0242] Discussion is a process of clearly sharing problems and, when necessary, referring to the example database (DB119) to derive the optimal solution. For example, when discussing "which label is appropriate?" for a particular image area, each user provides justification, and the administrator makes the final decision, resulting in efficient and accurate annotation.

[0243] These review and discussion features improve the accuracy of annotations while simultaneously enhancing the skills of annotators and improving overall team collaboration. After the annotation review, once annotation task N is deemed complete, an analysis of the annotator's corrections, work time, etc., is performed again (S34). If the analysis confirms that the start of learning is approved, the instruction to start learning is added to the system's queue database DB125 (S35).

[0244] Similarly, an annotator receives an N+1 annotation task request and, once it is ready to perform the annotation, executes this N+1 annotation task. This process is repeated.

[0245] Figure 13 is an explanatory diagram and flowchart of the active learning program of the present invention. The aim is to automate the learning and evaluation processes of AI development, making it easy for on-site personnel to perform these tasks without specialized knowledge. By incorporating a human-in-the-loop mechanism, the program realizes a system that allows users to continuously improve the performance of the AI ​​model while participating in the process.

[0246] Based on the data in Queue 1 database (DB) 126, AI model training is performed (S51). Model training is the process of learning patterns from data and optimizing parameters (weights and biases) for making predictions. Learning algorithms used at this stage include neural networks, decision tree analysis (DCA), and support vector machines (SVM). The AI ​​model is trained according to the learning method appropriate to the purpose, such as supervised learning, unsupervised learning, or reinforcement learning.

[0247] Decision tree analysis is a data mining technique used for purposes such as prediction, discrimination, and classification. It involves identifying explanatory variables that influence the dependent variable in customer information, survey results, etc., and creating a tree-like model.

[0248] A Port Vector Machine (SVM) is an algorithm that performs classification and regression by determining a boundary line or hyperplane that divides two sets of data into two classes. SVM uses the concepts of support vectors and margin maximization. Support vectors are the data points closest to the boundary line, and the margin is the distance between the boundary line and the support vectors.

[0249] Once the AI ​​model has finished training, it reaches an optimal state for making predictions based on the training data. This AI model includes the algorithms, weights, parameters, etc., for performing predictions and is generated as a result of machine learning. The created model is stored in the AI ​​model database DB121 and saved for future use.

[0250] Next, the AI ​​model is evaluated (S52). Validation and test data are used for the evaluation. These datasets are new data not used for model training and are used to measure how accurately the model can make predictions in a real-world operating environment. Evaluation metrics such as accuracy, precision, recall, and F1 score are used to confirm how well the model performs as expected.

[0251] After evaluation, the newly created AI model is compared to the existing AI model (S53). This confirms how much the new model's performance has improved compared to the existing model. If the new model is superior to the existing model, the new AI model is "deployed" (S54) and put into use in the actual operating environment. Deployment means integrating the new AI model into actual business operations or services and using it to perform real-time AI inference (S55).

[0252] "Deploy" means "to spread out" or "to place." In the IT field, it refers to a series of operations in the system development process, such as for web applications, where the functions and services of an application are placed and deployed on a server, making them available for use.

[0253] If a new AI model's performance falls short of existing models—that is, if it fails to achieve the expected accuracy or predictive power—the new model will not be deployed. In this case, the AI ​​inference system will continue to use the existing AI model to perform AI inference. Deploying an underperforming model could lead to inaccurate predictions and ultimately a decline in the quality of services and operations; therefore, new models will not be used until sufficient performance is ensured.

[0254] AI model inference is performed using data from the queue 2 database (DB) 127, the unlabeled database (DB) 12, and the AI ​​database for inference (DB) 129 (S56).

[0255] AI model inference refers to using a pre-trained AI model to make predictions or classifications on new (unknown) input data. Specifically, it includes performing inference on unknown data using the prediction results and class labels provided by the AI ​​model. For example, it can provide classification results for unlabeled data and generate predicted values ​​based on specific features.

[0256] AI models output multiple prediction results for input data. In the case of a label classification problem, the confidence level (probability) for each class is calculated and analyzed (S57).

[0257] Confidence level (probability) is an indicator of how confident a model is in the class or predicted value it outputs. Confidence level is usually expressed as the probability of the class predicted by the model. A higher confidence level indicates that the model is more confident in its prediction.

[0258] The uncertainty score is calculated based on these confidence levels (S58). A lower confidence level indicates higher uncertainty, suggesting the model is less confident in its predictions. This uncertainty score is an important indicator for evaluating how reliable the model's predictions are.

[0259] After that, it is determined whether the score is low (for example, less than 60%) (S59). If the score is low (YES), the prediction result is considered to be sufficiently reliable and the process can proceed to the next step as it is. On the other hand, if the score is high (NO), further processing such as re-evaluation or collection of additional data is required because the reliability of the prediction result is low.

[0260] Sampling based on the uncertainty score (S60) includes minimum confidence sampling, confidence margin sampling, confidence ratio sampling, and entropy-based sampling.

[0261] Minimum confidence sampling is a method of selecting data that the model is least confident about. Specifically, in a classification task, data with the lowest probability of the class predicted by the model is extracted. This allows training to focus on data where the model's learning is insufficient.

[0262] Confidence margin sampling is a method of selecting data where the difference (margin) between the highest probability value and the second highest probability value predicted by the model is small. This method aims to extract data where the model is having difficulty distinguishing classes and to promote learning efficiently.

[0263] Confidence ratio sampling is a method of selecting data where the ratio of the probability values of the top two classes predicted by the model is high. This method identifies data where the balance of the degree of confidence between different classes by the model is not good, and learning proceeds based on this.

[0264] Entropy-based sampling is a method of selecting data with high entropy (a measure of uncertainty) of the model's prediction probability distribution. Data with high entropy indicates that the model does not have strong confidence in any class, and by utilizing these data, the prediction accuracy and generalization performance of the model can be improved.

[0265] During active learning, the confidence level distribution for each image is maintained, and by comparing it with the corrected inference result and dynamically changing the confidence level, only the areas where the next inference result is more certain will be colored.

[0266] Personalization of annotation task N (S61) uses data from at least one of the personal database (DB) 123 and the administrator settings database (DB) 124.

[0267] Personalizing annotation tasks N can be achieved, for example, by adjusting the task's difficulty. The difficulty of annotation can be adjusted according to the user's skills and experience. For example, beginners could be assigned easy tasks (e.g., images with clear boundaries), while experienced annotators could be given more complex tasks (e.g., images with ambiguous boundaries).

[0268] One example is task optimization based on past work results. Personalized tasks are suggested based on the results of annotation work performed by the user in the past. For example, annotators who are skilled in a particular type of label or area can be assigned tasks related to that area, enabling more efficient work.

[0269] Another feature is interface customization. The interface can be personalized to improve the annotator's usability. For example, frequently used tools and shortcuts can be individually configured for each annotator, thereby increasing work efficiency.

[0270] Another improvement is the enhanced feedback function. This function provides individual feedback to annotators. By advising annotators on how to proceed with annotation based on their past work, improvements in the accuracy and speed of individual work are expected.

[0271] For example, this could involve managing the progress of a task. The system would suggest the next task to proceed with based on the annotator's work progress. For instance, after a specific step is completed, the system would automatically assign the most suitable next task to that annotator.

[0272] Furthermore, when multiple people are working on a task, if annotator B's annotation results are of low quality, the system may attempt to improve the annotation results by delegating the task to annotator A without setting a completion flag.

[0273] These personalization features aim to provide an optimal annotation environment based on the individual needs and work history of each annotator, supporting efficient and highly accurate work.

[0274] Finally, for the personalized annotation task N, the user who will be the annotator is formally requested to perform the specific annotation work (S62). Figure 14 is an explanatory diagram and flowchart of the predictive annotation method in the present invention.

[0275] Figure 14 shows "predictive annotation," a process that selects whether or not to apply the results of AI inference to the target of annotation task N (where N is 2 or greater).

[0276] Correcting predictions from machine learning models is a simple task for annotators. In "predictive annotation," machine learning can correctly predict data that is easy to annotate, so annotators are left to handle the cumbersome data that machine learning failed to predict well. Correcting misspecified boundaries often takes more time than starting from scratch, which can significantly increase annotator stress.

[0277] To reduce workload, the AI's uncertainty score is checked during the pre-inference phase, and if it falls below a certain threshold, the AI's inference result is rejected. By setting this threshold, annotators are not forced to receive predictions with low confidence, allowing for more efficient work.

[0278] For example, if the uncertainty score of the data predicted by the AI ​​is high (e.g., 60% or higher), the prediction result is not provided to the annotator, and instead, they are instructed to perform new annotations. This allows the annotator to focus on data they are confident in and avoid unnecessary correction work.

[0279] On the other hand, the AI ​​allows for selective adoption of its inference results to help users understand the importance of their work by demonstrating that the annotator's corrections are immediately reflected in the model and that the process influences subsequent predictions. Once these selections are complete, the annotator can perform the annotation task.

[0280] If the answer to "Continue from where you left off?" (S72) is YES, the task will resume from where the previous annotation results were saved and only the changes will be corrected. If the answer is NO, the annotation process will start from the beginning.

[0281] If the answer to "Is there a prior inference result?" (S73) is YES, the AI's prediction is displayed, and the annotator makes corrections based on it. If NO, there is no prior inference result, so the annotator performs the annotation work from scratch. The prior inference result is, for example, the class or value predicted by the AI ​​after analyzing the input data, and if the annotator needs to make corrections, they will review the result and make the necessary changes.

[0282] Figure 15 shows the situation when the inference result of the AI is reflected in real time based on the threshold of the uncertainty score during pre-inference. Multiple candidates proposed by the model are presented, and the annotator can select by clicking or dynamically moving the threshold of the indicator 540 with a slider.

[0283] The 0 of the indicator 540 indicates the lowest confidence level, indicating that the inference result is uncertain. On the other hand, the 1 of the indicator 540 indicates the highest confidence level, indicating that the inference result is very certain. The position of the indicator 540 indicates how confident the model is in its inference result, and the closer the value is to 1, the higher the model evaluates its inference result.

[0284] Conversely, the closer the value of the indicator 540 is to 0, the lower the confidence level of the model, indicating that the prediction result is uncertain. For example, if the value of the indicator 540 is 0.5, it indicates that the confidence level of the model for that inference result is medium.

[0285] The value of the indicator 540 is preferably controlled by the annotator adjusting the threshold. This allows the annotator to make appropriate data selection and modification based on the reliability of the inference result.

[0286] Figure 15(a) shows the state where the indicator 540 shows a threshold of 0.2, indicating that the inference result is very uncertain. Figure 15(b) shows the state where the indicator 540 shows a threshold of 0.5, indicating that the confidence level for the inference result is medium.

[0287] Figure 15(c) shows the state where the indicator 540 shows a threshold of 0.8, indicating that the inference result is very certain. By changing or setting the indicator 540 with a slider, the annotator can select an appropriate threshold according to their own criteria and perform annotation work based on the inference result.

[0288] It is also effective to change or set the threshold and value of indicator 540 based on the user's fatigue level, skill level, etc. Furthermore, it is also effective to change or set the threshold and value of indicator 540 based on the output data from camera 204.

[0289] The system monitors the annotation work log (pen trajectory log) while simultaneously using visual sensors such as camera 204 and gaze sensors to estimate the user's facial expressions, emotions, and gaze, and to measure their level of concentration on the annotation task.

[0290] Accordingly, the pen will light up, vibrate, etc., to alert the user. In addition, the display unit 556 of the input pen 101, the buttons of the input mouse 102, etc. will light up. When we determine that the level of concentration has clearly decreased, we will conduct another user analysis test. Figure 16 is a diagram illustrating the configuration and explanatory diagram of the input mouse 102 of the present invention.

[0291] As shown in Figure 16, the input mouse 102 is equipped with an operation button 551, a scroll wheel button 552, and a selection button 553. The operation state and operating state can be changed using the operation button 551, scroll wheel button 552, and selection button 553.

[0292] The operation button 551 can be enabled or disabled. The wheel button 552 can also be enabled or disabled. The selection button 553 can also be enabled or disabled.

[0293] Enablement or disablement can be set or changed based on user preferences and ease or difficulty of use. It can also be set or changed based on the user's physical characteristics, ease or difficulty of use, administrators can set or change it to be appropriate for the user, and settings or changes based on the type and content of the work, or the duration and level of fatigue of the work.

[0294] The wheel button 552 and the selection button 553 can be illuminated or deactivated. The user can understand the usage status and operation status by observing the illumination or deactivation of the wheel button 552 and the selection button 553, and can also change the usage status or settings. The illumination or deactivation can be controlled by the user, the administrator, or the controller 103.

[0295] Furthermore, this is not limited to simply turning on or off a light; it may also involve generating or suppressing sound or other noises. It may also involve generating or suppressing vibrations.

[0296] The direction in which the input mouse 102 is moved is detected by the accelerometer 206. The accelerometers 205 and 206 function not only as acceleration and movement speed sensors but also as direction sensors. By detecting or recognizing the speed, acceleration, and direction using the input mouse 102, the marker state, selection state, and operation state of the target shape 558 are monitored and applied appropriately to annotation work, etc.

[0297] Both the input pen 101 and the input mouse 102 can be used simultaneously, one or the other can be selected, or they can be used alternately. The operation buttons 551, selection buttons 553, selection buttons 555, and the display on the display unit 556 of the input pen 101 and the input mouse 102 can be linked.

[0298] The input mouse 102 is equipped with an accelerometer 205. The output of the accelerometer 205 detects the speed and acceleration of the mouse movement by the user. It also detects the direction in which the user moves the mouse. By detecting or recognizing the speed, acceleration, and direction, the marker state, selection state, and operation state of the target shape 558 are monitored and appropriately applied to annotation work, input work, management work, etc. The same applies to the input pen 101. As one embodiment, Figure 16(a) is an explanatory diagram showing the state in which both selection buttons 553a and 553b are not illuminated. Figure 16(b) is an explanatory diagram showing the state where selection button 553a is lit and selection button 553b is not lit. Figure 16(c) is an explanatory diagram showing the state where the selection button 553a is not lit and the selection button 553b is lit. Figure 16(d) is an explanatory diagram showing the state in which both selection buttons 553a and 553b are lit.

[0299] Figure 16(e) is an explanatory diagram showing the state when the wheel button 552 is lit. The wheel button 552 is active, and when the wheel button 552 is operated, the display scrolls in the normal vertical direction.

[0300] Figure 16(f) is an explanatory diagram showing the state when the wheel button 552 is not lit. The wheel button 552 is active, but when the wheel button 552 is operated, the normal vertical scrolling becomes slow.

[0301] Figure 16(g) is an explanatory diagram showing the state when the wheel button 552 is not illuminated. The wheel button 552 is active, but when the wheel button 552 is operated, scrolling to the left becomes more effective or the amount of movement to the left increases compared to normal.

[0302] Figure 16(h) is an explanatory diagram showing the state when the wheel button 552 is not lit. The wheel button 552 is active, but when the wheel button 552 is operated, scrolling to the right becomes more effective or the amount of movement to the right increases compared to normal.

[0303] The embodiment shown in Figure 16 is just one example, and it goes without saying that the enable / disable status of the selection button 553 and wheel button 552, as well as the movement speed and direction of movement, can be set and changed according to the user's selection, the content of the work, physical characteristics, administrator settings, and control of the controller 103.

[0304] By illuminating or deactivating the selection button 553, the user can visually recognize whether the selection button 553 is enabled or disabled. Alternatively, the user may set the button to illuminate or deactivate as they prefer. Furthermore, the administrator may set or control the button based on the user's work. Alternatively, the controller 103 may control the button's settings or functionality.

[0305] Figure 18 is an explanatory diagram of an embodiment in which an object or object figure 18 is enclosed by an outline. However, the embodiment in Figure 18 is just one example relating to the machine learning model development apparatus, annotation apparatus for model development, or machine learning model development method, annotation method for model development, etc. of the present invention.

[0306] During the administrator's setup process, the user performs annotation work on sample data to serve as an example for the administrator, and notes regarding the annotation process are displayed.

[0307] Figure 18(a) shows that the contour line drawing process starts at point a, but at point b the contour line is drawn deviating from the predetermined direction and continues towards point c. Furthermore, the drawing from point c is shifted towards point d. This can occur, for example, due to user fatigue. It can also occur depending on the user's strengths and weaknesses in line drawing.

[0308] The AI ​​learning function, controller 103, and annotation program detect the amount of deviation from a predetermined position and control the input mouse 102. Alternatively, they set and change the direction of movement, movement speed, button settings, etc.

[0309] In Figure 18(a), the direction of movement from point c is made more likely to be to the right (Figure 16(h)). Also, the speed of movement to points e and f is made slower compared to point b (Figure 16(f)).

[0310] Figure 18(b) shows that contour line drawing is performed from point a to point b, with good contour line drawing being achieved. Contour line drawing from point a is performed at a relatively slow speed. In this case, it is estimated that good drawing can be achieved even if the drawing speed is increased. Therefore, the drawing speed is increased from point b onwards.

[0311] Notes will be displayed if changes occur or are anticipated. These notes will automatically appear when the annotator begins work, serving as a reference. Users or workers will perform tasks such as drawing outlines based on these annotation notes.

[0312] As described above, by setting or changing the direction and speed of movement during the process, the user can configure the settings appropriately. Furthermore, annotation display can be implemented correctly. Additionally, the AI ​​learning effect can be enhanced.

[0313] It is effective to have users perform test drawing tasks before starting work, to detect and evaluate their strengths and weaknesses, and to select appropriate input devices (input mouse 102, input pen 101). It is also effective to appropriately set or change the accuracy, movement speed, and movement angle of the input devices.

[0314] Although the above embodiments were described using the input mouse 102 as an example, it goes without saying that the input device in the present invention is not limited to the input mouse 102, and may be, for example, an input pen 101 or the like.

[0315] As shown in Figure 17, the input pen 101 is equipped with a selection button 555 and a display unit 556. The display unit 556 and the selection button 555 allow the user to change the usage status, operation status, and the display on the display 100.

[0316] Furthermore, the selection button 555 can be turned on or off by user operation, administrator operation, or control by controller 103. For example, the user, administrator, and controller 103 can understand the usage status and operation status by the lighting or unlighting of the selection button 555 and the display unit 556, and the usage status or settings can be changed using the selection button 555.

[0317] The input pen 101's accelerometer 205 detects the direction of movement. By detecting or recognizing speed, acceleration, and direction, the marker state, selection state, and operation state of the target shape 558 are monitored and appropriately applied to annotation, input, and management tasks.

[0318] The input pen 101 is equipped with an accelerometer 205. The output of the accelerometer 205 detects the speed and acceleration of the mouse movement by the user. It also detects the direction in which the user moves the input pen 101. The controller 103 monitors the marker state, selection state, and operation state of the target shape 558 by detecting or recognizing the speed, acceleration, and direction of the input pen 101, and applies them appropriately to annotation and input tasks. For example, it processes the movement speed, movement direction, input speed, and input direction related to coordinate input of the input pen 101 by accelerating or delaying them. The same applies to the input mouse 102.

[0319] Controller 103 functions as a commander during image editing, a game console controller, etc., enabling easy movement, resizing, and scaling of images. It also allows for configuration of image movement and resizing operations. Furthermore, it appropriately controls annotation, input, and management tasks. As one embodiment, Figure 17(a) is an explanatory diagram showing the state in which both selection buttons 555a and 555b are not illuminated, and the display unit 556 is not illuminated. Figure 17(b) is an explanatory diagram showing the state where the selection button 555a is lit, and the selection button 555b and the display unit 556 are not lit. Figure 17(c) is an explanatory diagram showing the state where the selection button 555b is lit, and the selection button 555a and the display unit 556 are not lit. Figure 17(b) is an explanatory diagram showing the illumination states of the selection button 555a, selection button 555b, and display unit 556.

[0320] The embodiment shown in Figure 17 is just one example, and it goes without saying that the enable / disable status of the selection button 555 and the display unit 556, as well as the movement speed and movement direction, can be set and changed by user selection, by the content of the work, by administrator settings, and by control of the controller 103.

[0321] By illuminating or deactivating the selection button 555, the user can visually recognize whether the button is enabled or disabled. Users may also set the button to illuminate or deactivate as they prefer. Furthermore, administrators may configure or control the button's illumination based on the user's work. Alternatively, the settings or controls may be managed by the controller 103.

[0322] The input pen 101 has a pressure-sensitive axis 557, and the output data from the pressure-sensitive axis 557 changes depending on the pen pressure. The AI ​​learning function, controller 103, and annotation program also control the pen pressure.

[0323] The AI ​​learning function, controller 103, and annotation program detect deviations from a predetermined position (error amount, error amount, etc.) and control the input mouse 102. Alternatively, they set and change the direction of movement, movement speed, button settings, etc.

[0324] The work efficiency and working conditions change depending on the pressure applied to the input pen 101. Therefore, it is effective to have the user perform a test drawing task before starting work, to detect and evaluate their strengths and weaknesses regarding pen pressure, and to select an appropriate cushion sheet 560 and adjust the pen pressure applied to the input pen 101 accordingly.

[0325] The embodiment shown in Figure 19 is an embodiment in which the input tablet 100 and display 100 of the device of the present invention are configured so that a cushion sheet 560 selected from a plurality of cushion sheets 560 can be placed or mounted on the display unit 556.

[0326] The Cushion Sheet 560 (Cushion Sheet 560a, Cushion Sheet 560b, Cushion Sheet 560c, Cushion Sheet 560d) differs in flexibility. By selecting a Cushion Sheet 560, you can adjust the flexibility, and change and set the sensitivity and contact feel to suit the user. If the user has high pen pressure, select a flexible Cushion Sheet 560. If the user has low pen pressure, select a relatively hard Cushion Sheet 560.

[0327] The present invention allows for the use of a cushion sheet by changing or setting the type, material, flexibility, color, brightness, size, etc., to correspond to or adapt to the characteristics, personality, work conditions, age, recognition level, etc., of the worker or user.

[0328] The present invention is characterized by monitoring the user's viewpoint movement using a camera 204 and reflecting this viewpoint movement in the work status, annotation display, annotation operation, and AI learning.

[0329] The system observes the worker using camera 204 and detects or evaluates whether the worker is fatigued. Based on the evaluation, it sets and switches annotation displays, annotation instructions, etc. It also switches, sets, and changes the display or settings of the login program, management screen program, workspace screen program, and annotation screen program, as illustrated in Figure 4, etc. Furthermore, the control method, setting method, and display are switched, set, and changed based on observations of viewpoints, etc., as explained in Figure 20, etc.

[0330] As illustrated and explained in Figure 20, the method of observing, detecting, or evaluating worker fatigue during graphic input work using camera 204 involves analyzing the worker's physical changes, behavioral patterns, and attention span to assess signs of fatigue. Annotation displays and annotation instructions are set and switched based on the level of fatigue.

[0331] Furthermore, the display and settings of the login program, administration screen program, workspace screen program, and annotation screen program, as illustrated in Figure 4, etc., can be switched, configured, or changed.

[0332] Furthermore, the control method, setting method, and display are switched, set, and changed based on observations of viewpoints, etc., as explained in Figure 20, etc. Also, as illustrated in Figure 21, etc., pen pressure, etc., are controlled, changed, and set.

[0333] Using facial recognition technology, the system analyzes the worker's facial expressions and movements in real time. When a worker is tired, signs such as eye strain, decreased concentration, facial tension, blank expression, blinking, mouth movements, and eye movement speed may be observed. It is also possible to monitor eye movements, blinking frequency, and eyebrow elevation. Using facial recognition software (such as OpenCV or DeepFace), the system tracks the worker's facial features in real time and detects expressions and signs related to fatigue.

[0334] For example, eye strain can lead to increased blinking and longer periods of eye closure. Facial muscle tension can cause frown lines and a stiff face. Changes in facial expression can result in a blank or gloomy appearance.

[0335] Cameras monitor changes in posture and body movements to detect signs of fatigue. As workers become tired, their posture deteriorates, the muscles that support their bodies become fatigued, and their posture worsens unconsciously. Posture recognition algorithms (such as OpenPose or MediaPipe) are used to track the joints and posture of the worker's body. If a worker is working in the same posture for a long time, forward bending of the body, rounding of the back, and drooping of the shoulders can be observed, and these are evaluated as signs of fatigue.

[0336] For example, it can detect postural imbalances such as hunching or drooping shoulders. Examples include frequent changes in posture (moving the body frequently as a result of discomfort) and fatigue in the hands and arms (the hands and arms become tired from holding them in the same position for a long time).

[0337] One method involves using cameras to monitor the movements of a worker's mouse or pen, as well as their finger and hand gestures, to predict signs of fatigue. When fatigued, movements may become slower and more awkward. When the hands and arms become tired, the smoothness of movements is lost, and mouse and pen operation may become unstable. The camera tracks the smoothness and speed of hand and arm movements and analyzes the movement patterns. Slow and choppy movements may indicate fatigue.

[0338] For example, it measures and evaluates movement delays (slow reaction time), unnatural movements (shaking arms or hands, awkward movements), and changes in wrist and finger position (muscle fatigue due to prolonged use). Furthermore, it controls and modifies the display, annotations, and other settings of the display 100 in response to the measurement and evaluation results.

[0339] Cameras are used to detect subtle movements of the face and neck, and from these, heart rate and respiratory rate are estimated. For example, there is a method to detect heart rate from the minute movements of the arteries in the face (facial reflection method). When fatigue or stress accumulates, heart rate and breathing tend to increase. Infrared cameras and high-frame-rate cameras are used to estimate heart rate and respiratory rate from facial and neck movements and monitor changes in these values.

[0340] For example, symptoms include an increased heart rate (due to mental and physical stress during work), changes in breathing (shallow breathing or shortness of breath), and redness or paleness of the face (indicating changes in blood flow).

[0341] Using eye-tracking technology, which tracks eye movements, we analyze workers' concentration levels and fatigue. When fatigue sets in, gaze becomes unstable and eyes move frequently. Also, as concentration decreases, workers may spend more time looking away from the screen. Eye tracking devices and cameras are used to monitor eye movement patterns. The speed of eye movements and the duration of fixation are analyzed to assess decreased concentration and fatigue.

[0342] For example, we monitor and evaluate eye movement (frequent eye movements, unfocused gaze), blinking frequency (excessive blinking or frequent eye rubbing), and visual distraction (increased frequency of looking away from the screen).

[0343] The above visual data (facial expressions, posture, movements, eye movements, etc.) is integrated and a machine learning algorithm is used to comprehensively evaluate the worker's fatigue level. Based on the trained model, fatigue patterns can be identified and the worker's fatigue level can be scored in real time.

[0344] A fatigue score is calculated based on features extracted from image and video data. Based on this score, an alert can be issued to prompt workers to take a break. In addition, the display on Display 100, annotation display, and annotation control can be switched in response to the measurement and evaluation results.

[0345] In particular, eye tracking is a technology that tracks eye movements in real time and is extremely useful for evaluating visual attention, cognitive state, concentration, and fatigue levels. Because eye movements reflect changes in a worker's cognitive state and awareness, they can help detect fatigue, stress, and decreased concentration.

[0346] Infrared camera-based eye tracking is an effective method for tracking eye movements. It works by shining infrared light into the eye and reading the reflection patterns of the cornea and iris with a camera. This technology is highly accurate because it is less affected by external light conditions.

[0347] Figure 20 shows the measurement of changes in the user's viewpoint direction during a simple task on the target image displayed on screen 601. The viewpoint, visual direction, change time, change speed, etc., are detected and measured using camera 204. The work state, annotation display, and annotation operation are set and modified taking into account viewpoint movement, etc.

[0348] Figure 20(a) records the horizontal change in the viewpoint direction. The vertical axis represents the horizontal viewpoint angle, and the horizontal axis represents time (seconds). Figure 20(b) records the vertical change in the viewpoint direction. The vertical axis represents the vertical viewpoint angle, and the horizontal axis represents time (seconds).

[0349] In Figure 20(a), the horizontal angle is initially less than c°(DEG.), but for example, from 12 seconds onward in Figure 20, the angle becomes greater than or equal to c°(DEG.) (dotted line A). The horizontal angle is presumed to be caused by user fatigue and changes in concentration.

[0350] In Figure 20(b), the time of change in the vertical direction is initially a constant value d1, but after 10 seconds it becomes d2 (dotted line B). This change is presumed to occur due to changes in user fatigue and concentration.

[0351] The direction and duration of viewpoint shifts occur due to user fatigue and changes in concentration levels, but they also occur depending on each user's state, personality, etc., and the timing and angle of these shifts vary accordingly.

[0352] The present invention modifies or sets the learning data generation method, program learning method, prediction method, and machine learning model evaluation method in response to changes in the operator's or user's viewpoint position, speed, user's state, characteristics, personality, physical characteristics, etc.

[0353] Furthermore, the system provides methods for managing annotation work, and for managing annotation work more efficiently, by setting, changing, adjusting, and varying these methods. It also provides methods for ensuring the accuracy of annotation results, and devices and systems to support these methods.

[0354] It goes without saying that the above points may be implemented or modified based on observations of not only viewpoints, but also the facial expressions, mouth movements, blinking, etc. of users and workers using camera 204.

[0355] This invention monitors marker status, selection status, and operation status by measuring or evaluating the user's viewpoint position, etc., and applies them appropriately to annotation and input tasks. It performs processing such as acceleration and delay on the movement speed, movement direction, input speed, and input direction related to coordinate input of the input pen 101 and input mouse 102.

[0356] User fatigue, concentration level, etc., are reflected in pen pressure. The system detects and measures the user's pen pressure. Based on pen pressure and other factors, the system sets and modifies the work state, annotation display, and annotation operations.

[0357] Figure 21 shows the changes in user pen pressure and touch response status measured during a simple task on the target image displayed on screen 601. Figure 21(b) records the changes in pen pressure. The vertical axis represents pen pressure, and the horizontal axis represents time (seconds). Figure 21(b) shows Touch (on, input) and Left (off, non-contact) by the arrow. As shown by the dotted line in Figure 21(a), the pen pressure increases as the working time progresses. Changes in pen pressure occur due to user fatigue and changes in concentration, but they also occur depending on each user's characteristics, and the timing and pressure at which they occur vary.

[0358] The present invention modifies or sets methods for generating training data, program training, prediction, and evaluating machine learning models in response to changes in pen pressure, speed, user characteristics, state, and personality. It also controls, modifies, and sets methods for managing annotation work and methods for managing annotation work more efficiently. Furthermore, it provides methods for ensuring the accuracy of annotation results, as well as devices and systems that support such methods.

[0359] This invention monitors marker status, selection status, and operation status by measuring or evaluating the user's pen pressure, and appropriately applies and reflects these states in annotation and input tasks. Furthermore, it performs processing such as acceleration and delay on the movement speed, movement direction, input speed, and input direction related to coordinate input for the input devices, input pen 101 and input mouse 102.

[0360] While it is possible to measure only one of the viewpoint position and pen pressure data and incorporate it into the processing and apparatus of the present invention, it is preferable to use both the viewpoint position and pen pressure data and incorporate them into the processing and apparatus of the present invention.

[0361] It goes without saying that the above points may be implemented or modified based on observations of not only viewpoints, but also the facial expressions, mouth movements, blinking, etc. of users and workers using camera 204.

[0362] The difficulty level of a task changes depending on the user's fatigue and concentration level. Furthermore, the appropriate difficulty level for a task varies depending on the user's condition, personality, and characteristics.

[0363] As illustrated in Figure 22, the present invention provides users with work files categorized by difficulty level, taking into account fatigue, concentration level, and individual characteristics. These work files are organized in a database, categorized as either easy or difficult. The user, as the worker, may find that their work efficiency varies depending on the color and shape of the object being worked on.

[0364] Figures 23(a1) and 23(a2) illustrate voids 536 that have formed in the solder 538. The user performs a process to select and detect voids 536 during the operation. In Figure 23(a1), voids 536 are shown in white, and in Figure 23(a2), voids 536 are shown in black.

[0365] The difficulty level of the task shown in Figure 23(a1) or Figure 23(a2) changes depending on the user's fatigue and level of concentration. It may also vary depending on the user's personality and characteristics. The present invention is configured to allow switching between Figure 23(a1) and Figure 23(a2) by setting the display settings in Figure 6 and the input device settings in Figures 16 and 17.

[0366] Furthermore, the difficulty level varies depending on the color tone, detail, brightness, size, and number of the target shapes 558 to be displayed. This invention sets and changes the work state, annotation display, and annotation operation, taking into consideration the color tone, detail, brightness, size, and number of the target shapes 558 to be displayed.

[0367] Figures 23(b1) and 23(b2) illustrate the solder joint 538 on the lead pin 537. The user performs a process to select and detect voids 536 during the operation. In Figure 23(b1), each component is shown in white, while in Figure 23(b2), each component is colored.

[0368] The difficulty level of the task in Figure 23(b1) or Figure 23(b2) changes depending on the user's fatigue and level of concentration. It may also vary depending on the user's personality and characteristics. The difficulty level also changes depending on the color tone, detail, and brightness of the displayed target shape 558. The present invention is configured to allow switching between Figure 23(b1) and Figure 23(b2) by setting the display in Figure 6 and the input devices in Figures 16 and 17. Figure 24 is an explanatory diagram illustrating the method for detecting cracks 539 and voids 536 in each XY cross-sectional image.

[0369] For images showing solder joints (solder 538), the system processes the image to show where the cracks begin and end, and how the cracks are expected to propagate in the future. Similarly, it detects and processes areas (areas) where voids (voids 536) are present in the solder 538, or areas (areas) where voids are predicted to occur in the future. Handa 538 is an X-ray CT scanner, and as shown in Figure 24(b), it takes X-ray images to visualize the three-dimensional distribution of cracks 539 and voids 536. For each slice image, the operator detects cracks (539) and voids (536), and these results are superimposed to visualize the 3D distribution.

[0370] First, slice images are extracted one by one from the X-ray CT data. Next, solder regions and crack regions in the images are detected using the aforementioned segmentation model (U-Net, etc.).

[0371] As shown in Figure 24(b), for N XY cross-sectional images, the operator performs the task of detecting voids 536, etc., using multiple types of AI trained through deep learning. The images are then processed, and the results of each processing are combined to visualize the 3D distribution.

[0372] As shown in Figure 24, even in the results of the deep tillage treatment, cracks 539 and voids 536 projected onto the Z-plane at the joint of solder 538 can be measured and detected.

[0373] The difficulty level varies depending on the color tone, detail, and brightness of the target shape 558 to be displayed. This invention takes into account the color tone, detail, and brightness of the target shape 558 to be displayed and performs drawing processing instructions, annotation display, annotation operation processing, and AI learning processing.

[0374] Although the present invention has been explained using examples such as the detection of cracks 539 in joints, etc., it goes without saying that the present invention can also be applied to non-jointed areas, differences in crystal orientation, etc. Although the above examples mainly describe applications to the processing of solder 538 and the like, it goes without saying that the present invention can be implemented on other materials as well. For example, as shown in Figure 26, this is an embodiment in which the position and status of a vehicle are detected and processed by an operator or other person.

[0375] We will select a tool for extracting the car parts. Representative tools include LabelImg, VGG Image Annotator (VIA), and CVAT, which allow for manual annotation of images. "Car parts" refer to the parts of the car shown in the image (body, tires, windows, bumpers, etc.). It is important to define specifically what to extract. For example, decide whether to group the entire car under one label or to annotate individual parts (tires, doors, etc.) separately. As one embodiment of annotation,

[0376] Bounding Box Annotation: This annotation encloses parts of a car in a rectangle (square) and adds a label. You can enclose the entire car in one rectangle, or enclose individual parts of the car individually.

[0377] Polygon annotation: This method involves drawing polygons to match the shape of the vehicle. This allows for more accurate annotation even when the vehicle has a shape more complex than a rectangle.

[0378] Segmentation Mask: To extract the car portion in more detail, a mask is created for each pixel in the image indicating whether it is part of the car. This helps to accurately capture the contours and fine details of the object.

[0379] When performing annotation, appropriate labels are applied to the car body. For example, labels such as "car_body" and "tire" are used. After annotation, the annotation data for each image is saved in an appropriate format (XML, JSON, COCO format, etc.). This allows it to be used for subsequent data processing and training of machine learning models. Finally, the annotations are reviewed to check for any errors in annotation or missing car parts. Any inaccurate annotations are corrected. Through these processes, the car body can be accurately extracted, and data that can be effectively used in machine learning models and image analysis systems can be generated.

[0380] This invention is not limited to the above. Needless to say, it can non-destructively analyze a wide variety of things, such as air bubbles in adhesive resins, spaces in concrete blocks, inclusions of other metals in iron parts, and fat granules in the internal organs of living organisms.

[0381] Although embodiments of this disclosure have been described in detail above, this disclosure is not limited to the specific embodiments described above, and various modifications and changes are possible within the scope of the gist of this disclosure as described in the claims.

[0382] It should be noted that the present invention is not limited to the embodiments described above, and various modifications are possible without departing from the spirit of the invention. For example, the embodiments described above are described in detail for the purpose of clearly illustrating the present invention, and are not necessarily limited to those having all the configurations described. Furthermore, it is possible to add, delete, or replace some of the configurations of the above embodiments with other configurations.

[0383] While embodiments of this disclosure have been described above with reference to the attached drawings, a person with ordinary skill in the art to which this disclosure belongs will understand that this disclosure can be implemented in other specific forms without altering the technical idea or essential features. Therefore, it should be understood that the above-described embodiment is illustrative in all respects and not limiting.

[0384] The scope of protection of this disclosure should be interpreted in accordance with the claims below, and all technical ideas within the same scope should be interpreted as being included in the scope of rights of the technical ideas defined by this disclosure. [Industrial applicability]

[0385] This disclosure enables annotation and other processes for generating machine learning models and is suitable for use in annotation devices and learning devices. [Explanation of symbols]

[0386] 100 displays, touch panels 101 Input Pen, Stylus Pen 102 Input Mouse 103 Controller 104 Information Processing Device A 105 User Interface (UI) Management Program 106 Login Program 107 Workspace Program 108 Annotation Programs 109 Management Programs 110 Setup Program 111 User Analysis Program 112 Active Learning Programs 113 AI Learning Programs 114 AI Evaluation Program 115 AI Deployment Program 116 AI Inference Programs 117 Image Data Database (DB) 118 Annotation Data Database (DB) 119. Example Data Database (DB) 120. Work Data Database (DB) 121 AI Model Database (DB) 122 Deployed AI Model Database (DB) 123 Personal Data Database (DB) 124 Administrator Settings Database (DB) 125 Queue Database (DB) 126 Queue Database 1 (DB) 127 Queue Database 2 (DB) 128 Unlabeled Databases (DBs) 129. Inference AI Model Database (DB) 200 Information Processing Device B 201 Web Browser Programs 202 Toolbox 203 File Display 204 (Web) Camera 205 Accelerometer, Directional Sensor 205 Accelerometer, Directional Sensor 300 Cloud Internet 532 Annotation Display 533 Fill Display 534 Connecting copper foil 535 Printed circuit board 536 Void 537 Lead pins 538 Handa 539 Crack 540 Indicators 551 Operation Buttons 552 Wheel Button 553 Select button 555 Select button 556 Display section 557 Pressure-sensitive axis 558 Target Shapes 559 Input screen 560 Cushion Seat 601 Display screen 602 Work progress display 604 Split display 605 Annotation result 607 Start Learning button

Claims

[Claim 1] The first step involves conducting a personalized analysis of the annotator before or during the work, The process includes a second step in designing the optimal task based on one of the following conditions: the annotator's work status, work history, skill level, work trends, past work quality, and task completion rate. An annotation method characterized by optimizing the task through the first and second steps described above.

Citation Information

Patent Citations

  • Learning device and method, prediction device and method, program, and evaluation method of machine learning model

    JP2023021647A