Annotation system for surgical procedures

JP2025514094A5Pending Publication Date: 2026-03-30VERB SURGICAL INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-04-20
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Obtaining reliable annotations for training machine learning models in surgical applications is time-consuming and challenging due to the need for expert input, making it difficult to gather sufficient training data.

Method used

An annotation system that configures annotation jobs to specify annotators, labels, and decision targets, allowing for efficient annotation of surgical content items through user interfaces that facilitate swipe gestures or label selection, and updates machine learning model parameters based on annotated data.

Benefits of technology

The system enables efficient collection of labels for surgical content items, improving the training of machine learning models by reducing the time and effort required for annotation and enhancing the reliability of predictions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

The annotation system facilitates the collection of labels for images, videos, or other content items relevant to training a machine learning model associated with a surgical application or other medical application. The annotation system enables an administrator to configure annotation jobs associated with training a machine learning model. The job configuration controls the presentation of content items to various participating annotators via the annotation application, and the collection of labels via the annotation application's user interface. The annotation application enables participating annotators to provide input in a simple and efficient manner, such as by providing gesture-based input or by selecting graphical elements associated with different possible labels.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The described embodiments relate to an annotation system for capturing annotations of surgical content for training a machine learning model. [Background technology]

[0002] Machine learning models enable the automated generation of predictions from videos, images, animations (e.g., gifs), and other medical data. These predictions are useful for assisting physicians in diagnosing patients, performing surgical procedures, and recommending treatments. In a supervised learning process, machine learning models are trained using large datasets that have been annotated to describe their content or characteristics. However, obtaining reliable annotations traditionally involves a significant time investment from physicians or researchers with deep expertise in the specific field involved. As a result, it can be difficult to obtain sufficient training data to meaningfully train or improve these types of machine learning models. Summary of the Invention [Means for solving the problem]

[0003] In a first embodiment, a method facilitates training a machine learning model to predict a characteristic of a content item, including an image or video associated with a surgical application. An annotation job associated with the machine learning model is configured to specify at least one annotator, a predetermined set of selectable labels, and a target number of judgments. A content item is obtained. The machine learning model is applied to generate a prediction associated with the content item and a confidence metric associated with the prediction. The confidence metric is evaluated to determine whether the confidence metric meets a predetermined confidence threshold. In response to the confidence metric not meeting the confidence threshold, the content item is added to an annotation set for receiving an annotation via an annotation application. The annotation application facilitates presentation of the content item to the at least one annotator via a user interface of the annotation application. At least one selected label from a predetermined set of selectable labels is obtained from the annotation application. It is determined whether the target number of judgments is met for the content item. In response to the target number of judgments being met, parameters of the machine learning model are updated based on the at least one selected label and the content item.

[0004] In one embodiment, obtaining the at least one selected label includes identifying a swipe gesture performed on the user interface and selecting between a predetermined set of labels based on a direction of the swipe gesture.

[0005] In one embodiment, obtaining the at least one selected label includes facilitating presentation, via a user interface, of user interface elements each associated with a predetermined set of labels, and selecting among the predetermined set of labels based on a selection of one of the user interface elements.

[0006] In one embodiment, configuring the annotation job further includes obtaining labeling rules that indicate a number of selectable labels that may be selected for a content item by an annotator.

[0007] In one embodiment, obtaining the at least one selected label includes performing a selection of one and only one of the predefined set of labels via a user interface.

[0008] In one embodiment, obtaining the at least one selected label includes enabling selection of any number of the predefined set of labels via a user interface.

[0009] In one embodiment, the annotation system automatically facilitates presentation of another content item via a user interface of an annotation application on a client device in response to obtaining the at least one selected label.

[0010] In one embodiment, the target number of judgments includes the number of unique annotators that provided at least one label for the content item.

[0011] In another embodiment, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors, cause the one or more processors to perform the following steps: A content item associated with the annotation job is presented to an annotator via a user interface of an annotation application. The content item includes a surgical image or video and a predetermined set of selectable labels associated with the surgical image or video. A swipe gesture performed by the annotator via the user interface is identified. At least one selected label from the predetermined set of selectable labels is determined based on the swipe gesture. The selected label is stored in association with the content item. An additional content item associated with the annotation job is presented via the user interface.

[0012] In one embodiment, determining the at least one selected label includes selecting only a single selected label for association with the content item.

[0013] In one embodiment, determining the at least one selected label includes selecting a plurality of selected labels for association with the content item.

[0014] In one embodiment, determining the at least one selected label includes selecting between a first predefined label in response to the swipe gesture being performed in a first direction and a second predefined label in response to the swipe gesture being performed in a second direction opposite the first direction.

[0015] In one embodiment, determining the at least one selected label includes selecting between a first predetermined label in response to a swipe gesture being performed in a first direction, a second predetermined label in response to a swipe gesture being performed in a second direction, a third predetermined label in response to a swipe gesture being performed in a third direction, and a fourth predetermined label in response to a swipe gesture being performed in the third direction.

[0016] In one embodiment, the content item includes at least one of an image, a video, and an animation.

[0017] In one embodiment, the instructions, when executed, further cause the one or more processors to perform steps including presenting a control element for flagging a content item for review by an administrator, and storing the flag in association with the content item in response to selection of the control element.

[0018] In one embodiment, the instructions, when executed, further cause the one or more processors to perform steps including tracking a state of the annotation job recording the progress of annotations received for a set of content items associated with the annotation job, and configuring the annotation job based on the tracked state in response to the annotation application closing the reopen.

[0019] In another embodiment, a method facilitates training a machine learning model to predict a characteristic of a content item including an image or video associated with a surgical application. An annotation job associated with the machine learning model is configured to specify a predetermined set of selectable labels for labeling the content item and a labeling rule indicating a number of selectable labels that may be selected for the content item by an annotator. A content item is retrieved. A presentation of the content item is facilitated via a user interface of an annotation application. At least one selected label from the predetermined set of selectable labels is retrieved from the annotation application. Parameters of the machine learning model are updated based on the at least one selected label and the content item.

[0020] In one embodiment, obtaining the at least one selected label includes identifying a swipe gesture performed on the user interface and selecting between a predetermined set of labels based on a direction of the swipe gesture.

[0021] In one embodiment, obtaining the at least one selected label includes facilitating presentation, via a user interface, of user interface elements each associated with a predetermined set of labels, and selecting among the predetermined set of labels based on a selection of one of the user interface elements.

[0022] In one embodiment, configuring the annotation job further includes identifying a set of annotators associated with the job for presenting the content item. [Brief description of the drawings]

[0023] [Figure 1] is an example embodiment of a computing environment for facilitating collection of labels associated with training a machine learning model for a medical application. [Diagram 2] 1 is an exemplary embodiment of a job management engine for managing annotation jobs. [Diagram 3] 1 is an exemplary sequence of user interface screens associated with an annotation application. [Figure 4] 1 is another example of a user interface screen associated with an annotation application. [Diagram 5] 1 is yet another example of a user interface screen associated with an annotation application. [Figure 6] 1 is an example of a user interface screen for creating an annotation job. [Figure 7] 1 is an example of a user interface screen for assigning an annotation job to a set of annotators. [Figure 8] 1 is an example of a user interface screen for viewing a list of annotation jobs. [Figure 9] 1 is an example embodiment of a process for facilitating training of a machine learning model based on labels collected via an annotation system. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0024] The Figures (FIGS.) and the following description illustrate specific embodiments by way of example only. Those skilled in the art will readily appreciate from the following description that alternative embodiments of the structures and methods illustrated herein may be employed without departing from the principles described herein. Reference will now be made to certain embodiments, examples of which are illustrated in the accompanying drawings. Wherever possible, similar or similar reference numbers may be used in the drawings and these reference numbers may indicate similar or similar functionality.

[0025] The annotation system facilitates collection of labels for images, videos, animations (e.g., gifs), or other content items relevant to training machine learning models associated with surgical applications or other medical applications. The annotation system enables an administrator to configure annotation jobs associated with training machine learning models. The configuration of annotation jobs controls the presentation of content items to various participating annotators via the annotation application and facilitates the collection of labels via the annotation application's user interface. The annotation application enables participating annotators to provide input in a simple and efficient manner, such as by providing gesture-based input or by selecting graphical elements associated with different possible labels.

[0026] 1 illustrates an example embodiment of an annotation system 100 for capturing annotations of medical images, videos, animations (e.g., gifs), or other medical data. The annotation system 100 includes an annotation server 110, one or more client devices 120, one or more management devices 130, and a network 140.

[0027] Annotation server 110 comprises one or more computing devices that interact with client devices 120 and management device 130 over network 140 to perform various functions described herein. Annotation server 110 may comprise a single physical server, a set of distributed physical servers, a cloud server, one or more virtual machines or containers running on one or more physical servers, or a combination thereof. Annotation server 110 includes at least a processor and a non-transitory computer-readable storage medium that stores instructions executed by the processor to perform functions attributed to annotation server 110 described herein.

[0028] In one embodiment, annotation server 110 includes a user interface engine 112, a job management engine 114, a machine learning (ML) engine 126, a content database 116, an annotation database 118, and a machine learning (ML) model database 124. Alternative embodiments may include different or additional components.

[0029] The job management engine 114 controls a set of jobs associated with obtaining annotations to facilitate training of machine learning models. Generally, each job involves identifying a training set of content items stored in the content database 116, obtaining a set of annotations for the training content items for storage in the annotation database 118, and generating or updating a machine learning model (in the ML model database 124) based on the training content items and associated annotations. Jobs may be created and managed based on a job description provided by an administrator via a management application 132 (described in more detail below) of the management device 130 that indicates various parameters associated with the job.

[0030] The job description may identify the content items associated with the job directly (e.g., by associating a unique identifier for the content item with the job) or based on a set of configurable characteristics. For example, the job description may identify the content items for inclusion in the training set by specifying one or more types of the content item (e.g., image, video, animation (e.g., gift), illustration, text, etc.), one or more sources of the content item, a time or location associated with the creation of the content item, one or more characteristics of the content item (e.g., encoding type, quality, resolution, size / length, etc.), one or more tags associated with the content item, or various metadata describing the content item. In this manner, content items that meet the specified criteria may be automatically added to a training set associated with a particular job based on configured parameters in the job description.

[0031] The job description may also specify labeling rules, including various parameters or constraints associated with obtaining the annotations. For example, the job description may specify a set of predefined labels between which the annotator selects, the number of which may vary depending on the labeling rule. For example, some jobs may have labeling rules that implement a binary classification, where the annotator is required to select one label from a set of two predefined labels for each content item. Other jobs may have labeling rules that specify a multi-label classification request, where the annotator is required to select one label from a set of three or more predefined labels. Other jobs may be configured to allow multi-label selection, where the annotator selects any number of labels from a predefined set of labels (e.g., selecting anywhere between zero labels and all labels), or some other predefined quantity or range of quantities of labels (e.g., selecting between zero and two labels from a set of four labels, selecting exactly two labels from a set of five labels, etc.). Other jobs may allow free-form text labels, where the annotator is not necessarily limited to a predefined set of labels.

[0032] The job description may further optionally specify sets of control inputs that correspond to different ones of the given labels. For example, the job description may assign specific touch screen gestures to specific labels (e.g., swipe left, swipe right, swipe up, swipe down, etc.). In other embodiments, the control inputs may be assigned automatically.

[0033] The job description may further specify a configured number of judgments representing the number of labels that, when received from different annotators for a given content item, will cause the content item to be added to a training set for training the machine learning model. For example, if the number of judgments is set to 5, the content item will be marked as unannotated if it has been reviewed and annotated by less than 5 annotators. Once the content item receives 5 labels, the job management engine 114 may mark the content item as annotated, which may be utilized in the training set. In an exemplary embodiment, the job management engine 114 may enforce a requirement that all received labels match before adding the content item to the training set. Alternatively, the job management engine 114 may enforce a predetermined threshold level of match while still allowing some mismatch (e.g., 5 out of 6 labels match). In some embodiments, the job management engine 114 may remove a content item from the content database 116 (or remove its association with the job) if it observes at least a threshold level of mismatch for the labels received from the annotators.

[0034] The job description can further identify a set of annotators to receive content items associated with the job for annotation. The set of annotators can be identified directly (e.g., using the annotators' user identifiers) or based on a set of specified criteria. For example, the job description can specify the set of annotators based on characteristics such as expertise, experience level, location, availability, or other user data specified in the annotators' user profiles. Alternatively, annotators may be pre-assigned to one or more groups, and the job description may identify the annotators based on the groups. The job management engine 114 can then limit participation in the job to the identified set of annotators.

[0035] The content database 116 stores content that has been annotated or that may be selected for annotation by the annotation system 100. The content may include, for example, medical images, medical videos, animations (e.g., gifs), illustrations, text-based patient data, or other medical data. In some embodiments, the content database 116 may include non-medical images, videos, animations (e.g., gifs), or other data. The content in the content database 116 may include associated metadata that describes characteristics of the associated content generated at the time of capture, such as the type of content (e.g., image, video, animation, illustration, text, etc.), when the content was created, where the content was created, the entities involved in the creation, permissions associated with the content, file size, encoding format, or other metadata.

[0036] The annotation database 118 stores annotations associated with content in the content database 116. The annotations may include descriptive labels that describe various characteristics of the content provided by annotators, for example, as described in more detail below.

[0037] The machine learning engine 126 trains a machine learning model based on the content items and annotations identified for the job. The machine learning engine 126 can further apply the machine learning model to generate predictions for unannotated content items. The machine learning engine 126 can utilize techniques such as neural networks, classification, regression, or other computer-based learning techniques. Parameters (e.g., weights) associated with the machine learning model are stored in the ML model database 124.

[0038] The UI engine 112 interfaces with the annotation application 122 and the management application 132 to enable various information display, control, and presentation of content items, and processes input received from the annotation application 122 and the management application 132 as described herein. In one embodiment, the UI engine 112 may include a web server for providing the annotation application 122 and / or the management application 132 via a web interface accessible by a browser. Alternatively, the UI engine 112 may comprise an application server for interfacing with locally installed versions of the annotation application 122 and / or the management application 132. In one embodiment, the UI engine 112 may also enable direct access to the annotation database 118 to enable data scientists to manually review annotations and develop improvements to the machine learning process.

[0039] The client device 120 and the management device 130 comprise computing devices for executing the annotation application 122 and the management application 132, respectively. The annotation application 122 facilitates the submission of information and collection of user input associated with submitting a content item and obtaining labels from an annotator via the client device 120. The annotation application 122 may further track the progress of an annotator associated with an annotation job, such that a user may close the annotation application 122 and return to the same at a later time. Thus, a user need not necessarily complete an annotation job all at once, but may instead annotate several content items at a time or otherwise proceed with an annotation job accordingly. Exemplary embodiments of a user interface associated with the annotation application 122 are provided below with respect to FIGS. 3-5. The management application 132 facilitates the submission of information and collection of input from an administrator associated with creating or updating a job, as well as viewing information related to the job. Exemplary embodiments of a user interface associated with the management application 132 are provided below with respect to FIGS. 6-8.

[0040] Each of the client devices 120 and the management devices 130 may comprise, for example, a mobile phone, a tablet, a laptop or desktop computer, or other computing device. The annotation application 122 and the management application 132 may execute locally on the respective client devices 120, 130, or may comprise web applications (e.g., hosted by a remote server) accessed via a web browser. Each of the client devices 120 and the management devices 130 may include conventional computer hardware, such as a display, an input device (e.g., a touch screen), a memory, a processor, and a non-transitory computer-readable storage medium that stores instructions executed by the processor to perform the functions ascribed to the respective devices 120, 130 described in this specification.

[0041] Network 140 provides communication paths for communication between client devices 120, management device 130, annotation server 110, and other devices. Network 140 may include one or more local area networks and / or one or more wide area networks (including the Internet). Network 140 may also include one or more direct wired or wireless connections.

[0042] 2 illustrates an exemplary embodiment of job management engine 114 and its interaction with ML engine 126, which includes an ML prediction engine 232 and an ML training engine 234. In this embodiment, job management engine 114 includes a cycle manager 202, an oracle 204, and a data selector 206.

[0043] The ML prediction engine 232 receives unannotated content items 212 associated with a particular job (e.g., from the content database 116) and applies the relevant machine learning models 214 associated with the job to determine a prediction metric 220 associated with a prediction of a label for the unannotated content items 212. The prediction metric may include, for example, a confidence level associated with the prediction made by the ML prediction engine 232. Alternatively, the prediction metric may include an entropy-based score or a similarity score indicating the similarity between the content item and other content items that have already received annotations.

[0044] The data selector 206 selects a set of content items 216 selected from the content database 116 for annotating in relation to the job based on the prediction metric 220. The data selector 206 selects the content items 216 that it predicts could contribute most to improving the performance of the machine learning model 214 if included in the training set. In an example embodiment, this selection criterion may depend on the confidence level associated with the prediction. Under this framework, if the ML prediction engine 232 predicts a label for an unannotated content item 212 with a relatively low confidence level (as shown in the prediction metric 220), this indicates that the current ML model 214 is performing relatively poorly for that content item 212, and thus obtaining manual annotations for that content item 212 may provide a relatively significant improvement over the machine learning model 214 for other content items with similar characteristics. If, instead, the ML prediction engine 232 predicts a label for an unannotated content item 212 with a relatively high confidence, this indicates that the ML model 214 is already performing relatively strongly for that content item, and obtaining manual annotations for that content item 212 may not provide a significant improvement. In another embodiment, the data selector 206 may select the content items 212 for annotation based on an entropy metric. Alternatively, the predictive metric 220 may include a similarity score indicating the similarity between the content item 212 and other content items that have already been annotated. Here, the least similar content items (e.g., having a similarity score below a predetermined threshold) are most likely to contribute to improved performance of the machine learning model and may be selected by the data selector 206. In further embodiments, the data selector 206 may use other data mining strategies or combinations of techniques.

[0045] In one embodiment, the data selector 206 evaluates the predictive metric 220 of each content item 212 individually and determines to include a content item 212 in the set of selected content items 216 if the predictive metric 220 is below a threshold level. In another embodiment, the data selector 206 evaluates a batch of predictive metrics 220 for each content item 212 and then selects a predetermined number or percentage of the content items 212 in the batch for inclusion in the set of selected content items 216 that correspond to the content items 212 having the lowest relative predictive metric 220.

[0046] The oracle 204 interfaces with the annotation application 122 of the client device 120 to obtain labels 218 of selected content items 216 chosen for annotation. Based on a job description, the oracle 204 may control which annotators from a pool of annotators have access to the content items 216 selected for annotation and which labels may be assigned to the content items. The oracle 204 may also aggregate annotation results received from the client device 120 and filter out any content items flagged by annotators as containing bad data (e.g., content items not in an associated category, low quality content items, etc.). The annotations received by the oracle 204 may be stored in the annotation database 118 as described above.

[0047] The ML training model 234 trains (or updates) the machine learning model 214 associated with a particular job based on the retrieved labels 218 received from the oracle 204 and the associated selected content items 216. The ML training engine 234 may operate continuously to perform updates as newly retrieved labels 218 are received, or may perform in response to a trigger (e.g., time elapsed since the last update, a predefined number of new labels 218 being received, etc.).

[0048] The cycle manager 202 tracks the progress of each annotation job and determines when the job is complete, where the cycle manager 202 may consider the job complete when a predetermined set of completion criteria associated with the job is met, such as, for example, obtaining labels 218 for at least a predetermined number of content items 216, achieving at least a predetermined average prediction metric 220 for predictions made by the ML prediction engine 232, detecting when predictions by the ML prediction engine 232 stop improving, or satisfying other predetermined metrics associated with the job.

[0049] FIG. 3 is an exemplary sequence of user interfaces presented to an annotator via annotation application 122 to obtain annotations of surgical media content, such as video clips, images, or animations (e.g., gifs). A login screen 302 allows a user to provide credentials (e.g., username and password) to log into annotation server 110. Upon receiving and authenticating the login credentials, annotation server 110 accesses a user profile associated with the user and identifies open annotation jobs for the user. The identified annotation jobs may include annotation jobs specifically assigned to the user (e.g., based on a user identifier) ​​or may include annotation jobs matched to the user upon login. Here, annotation jobs may be matched to the user based on information in the user profile (e.g., expertise area, experience level, participation availability, etc.) and metadata associated with the jobs. A job list screen 304 presents the user with a list of selectable control elements 312 associated with different available jobs. Upon selecting a job (e.g., in this case, selecting "Surgeon Idle or Active"), the annotation application 122 presents the user with a job description screen 306 with a description 314 of the selected job and instructions. Here, for example, the job description screen 306 specifies that the user "Determine if the Surgeon is active during a clip or idle" and "Swipe right for active, swipe left for idle". The job description screen 306 also indicates that there are two possible labels to choose from for this job, either "Idle" or "Active". The user can proceed with the job by selecting the "Start Job" control element 316. The annotation application 122 then presents a series of annotation screens 308 containing content items 332 (in this case, video clips) for the annotator to review and label.In this example, the annotator can select between two labels ("Idle" or "Active") by selecting the corresponding control element 318-A, 318-B or by performing a corresponding gesture. For example, using a touch screen device, the user may swipe left on the touch screen to select "Idle" or swipe right on the touch screen to select "Active". Alternatively, the user interface may allow selection of the label by voice input or other type of input.

[0050] The user also has the option to select control element 320 “unknown” to decline to provide a label for the content item 332 and skip to the next content item. The annotation screen 308 further includes a control element for canceling the current label selection 326 or canceling all label selections 324 provided during the current session. In one embodiment, the annotation screen 308 may also provide a control element 328 that allows the user to flag the content item 332. Flagging the content item 332 may be to automatically remove the content item from the job or to flag the content item for review by a job administrator. Flagging a content item may be useful, for example, to indicate that the content item 332 is not relevant to the job, is of low quality, contains occlusions, or may not be suitable for training a machine learning model. The annotation screen 308 may also include an annotation count 330 that indicates the cumulative number of content items that the annotator has labeled during the current session.

[0051] In one embodiment, annotation application 122 can track one or more jobs that a user has started and automatically return to that job if the user exits and later resumes application 122. This allows annotators to start annotating quickly, pause when needed, and return to where they left off so that they can contribute to annotation on a convenient time schedule.

[0052] 4 illustrates an example of an annotation screen 408 associated with another job. In this example, the job is configured to request a single label from the annotator selected among four predefined labels 418. The annotator can choose a label by selecting an associated control element. Alternatively, swipe gestures may be pre-assigned to different labels to enable selection (e.g., swipe up and left to select "scissors", swipe down and left to select "grasper", swipe up and right to select "stapler", swipe down and right to select "cautery").

[0053] 5 shows another example of an annotation screen 508 associated with yet another job. In this example, the job is configured to enable multiple labels 518 (e.g., 0-4 labels) for a single content item. The annotator can select the labels using on-screen controls or by a combination of gestures as described above.

[0054] 6 illustrates an example embodiment of a user interface screen 600 for an administrator application 132 associated with creating and / or configuring an annotation job. This interface allows a job creator to provide information such as the name of the job, a description of the job, the type of job (e.g., single label selection or multi-label selection), the number of labels, the names of the labels for a given set, and the number of decisions.

[0055] 7 also shows another user interface screen 700 for the administrator application 132 that allows an administrator to assign jobs to a particular set of annotators. In the illustrated embodiment, the annotators are listed by unique identifier (e.g., username). Alternatively, the annotators can be identified using various labels that characterize the annotators, such as area of ​​expertise, availability, experience level, etc., allowing an administrator to identify a set of annotators that meet defined characteristics without necessarily identifying them individually. In further embodiments, annotators may be pre-assigned to one or more groups of annotators, and an administrator can assign jobs to one or more groups.

[0056] 8 shows an example of a user interface screen 800 of the administrator application 132 for viewing and editing a set of jobs. The screen provides a list 802 of jobs (e.g., identified by name and / or unique identifier) ​​and various parameters associated with the jobs 804. The administrator can select a job to view additional information, delete the job, or edit the parameters associated with the job.

[0057] FIG. 9 illustrates an example embodiment of a process for generating or updating a machine learning model based on annotations received via annotation application 122. Annotation server 110 configures 902 an annotation job associated with the machine learning model (e.g., based on input from management device 130) to specify a set of criteria including, for example, a pool of one or more annotators, labeling rules indicating an amount (or amount range) of labels an annotator may select, a predetermined set of selectable labels, and a target number of decisions. Annotation server 110 retrieves 904 a content item that is initially unannotated. Annotation server 110 determines 906 whether to add the content item to a set of annotations associated with the job. For example, annotation server 110 may apply the machine learning model to generate a prediction of the content item and a confidence metric associated with the prediction. Annotation server 110 then adds the content item to the set for annotation in response to the confidence metric not meeting a confidence threshold. Annotation server 110 retrieves 910 at least one selected label from a predetermined set of selectable labels via annotation application 122. For example, annotation server 110 facilitates presentation of the content item to at least one of the annotators in the pool via a user interface of annotation application 122 of client device 120. The annotation server 110 then determines 912 whether a target number of judgments has been met for the content item based on a cumulative set of labels received for the content item from the pool of annotators. In response to the target number of judgments being met, the annotation server 110 updates 914 the parameters of the machine learning model by using the obtained labels to retrain the machine learning model. The annotation server 110 then outputs 916 the machine learning model (e.g., by storing it in machine learning model database 124).

[0058] The machine learning models generated from the annotation system 100 described above can be utilized in a variety of contexts. For example, the machine learning models can be applied to pre-operative, intra-operative, or post-operative surgical images or videos to automatically classify images in a surgical context. The machine learning models may similarly be applied to other medical images or videos for purposes of diagnosing, treating, or investigating medical conditions. In other alternative embodiments, the annotation system 100 described herein may be utilized to obtain labels and train machine learning models associated with other types of content items not necessarily related to the medical field.

[0059] The described embodiments of the annotation system 100 and corresponding processes may be implemented by one or more computing systems. The one or more computing systems include at least one processor and a non-transitory computer-readable storage medium that stores instructions executable by the at least one processor to perform the processes and functions described herein. The computing systems may include distributed network-based computing systems in which the functions described herein are not necessarily performed on a single physical device. For example, some implementations may utilize cloud processing and storage technologies, virtual machines, or other technologies.

[0060] The foregoing description of the embodiments has been presented for purposes of illustration and is not intended to be exhaustive or to limit the embodiments to the precise form disclosed. Those skilled in the art will recognize that many modifications and variations are possible in light of the above disclosure.

[0061] Some portions of this description describe embodiments in terms of algorithms and symbolic representations of operations on information. These operations, while described functionally, computationally, or logically, will be understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Further, it has proven convenient at times, without loss of generality, to refer to arrangements of these operations as modules. The described operations and their associated modules may be embodied in software, firmware, hardware, or any combination thereof.

[0062] Any of the steps, operations, or processes described herein may be executed or implemented using one or more hardware or software modules, alone or in combination with other devices. The embodiments may also relate to an apparatus for performing the operations herein. The apparatus may be specially constructed for the required purposes and / or may comprise a general-purpose computing device selectively activated or reconfigured by a computer program stored in the computer. Such a computer program may be stored in a tangible non-transitory computer-readable storage medium or any type of medium suitable for storing electronic instructions and coupled to a computer system bus. Furthermore, any computing system referred to herein may include a single processor or may be an architecture employing a multiple processor design to increase computing power.

[0063] Finally, the language used herein has been selected primarily for ease of reading and instructional purposes, and may not have been selected to delineate or limit the subject matter of the invention. Accordingly, the scope is intended to be limited not by this detailed description, but rather by any claims issued on an application based thereon. Accordingly, the disclosure of the embodiments is intended to be illustrative, but not limiting, of the scope of the invention, which is set forth in the following claims.

[0064] [Embodiment] (1) A method for facilitating training of a machine learning model to predict a characteristic of a content item, including an image or video, associated with a surgical application, comprising: Configuring an annotation job associated with the machine learning model to specify at least one annotator, a predetermined set of selectable labels, and a target number of decisions; obtaining the content item; applying the machine learning model to generate a prediction associated with the content item and a confidence metric associated with the prediction; determining whether the confidence metric meets a predetermined confidence threshold; adding the content item to an annotation set for receiving annotations via an annotation application in response to the confidence metric not satisfying the confidence threshold; facilitating presentation of the content item to the at least one annotator via a user interface of the annotation application; obtaining, from the annotation application, at least one selected label from the predetermined set of selectable labels; determining whether a target number of determinations has been met for the content item; updating parameters of the machine learning model based on the at least one selected label and the content item in response to the target number of determinations being met; A method comprising: (2) obtaining the at least one selected label, Identifying a swipe gesture performed on the user interface; selecting between the predetermined set of labels based on a direction of the swipe gesture; 2. The method of embodiment 1, comprising: (3) obtaining the at least one selected label, facilitating presentation, via the user interface, of user interface elements each associated with a respective one of the predetermined set of labels; selecting between the predetermined set of labels based on a selection of one of the user interface elements; 2. The method of embodiment 1, comprising: (4) configuring the annotation job, 2. The method of embodiment 1, further comprising obtaining a labeling rule indicating a number of selectable labels that may be selected for the content item by a single annotator. (5) obtaining the at least one selected label, 2. The method of embodiment 1, comprising performing a selection of only one of the predetermined set of labels via the user interface.

[0065] (6) obtaining the at least one selected label, 2. The method of claim 1, comprising enabling selection of any number of the predetermined set of labels via the user interface. (7) The method of embodiment 1, further comprising automatically facilitating presentation of another content item by the annotation application via the user interface of the annotation application in response to obtaining the at least one selected label. (8) The method of claim 1, wherein the target number for the determination includes a number of unique annotators that provided at least one label to the content item. (9) When executed by one or more processors, the one or more processors: presenting, via a user interface of the annotation application, for an annotator, content items associated with the annotation job and a predetermined set of selectable labels; Identifying a swipe gesture performed by the annotator via the user interface; determining at least one selected label from the predetermined set of selectable labels based on the swipe gesture; storing the selected label in association with the content item; presenting, via the user interface, additional content items associated with the annotation job; A non-transitory computer-readable storage medium storing instructions to cause a computer to perform steps including: (10) determining the at least one selected label, 10. The non-transitory computer-readable storage medium of embodiment 9, further comprising selecting only a single selected label for association with the content item.

[0066] (11) determining the at least one selected label comprises: 10. The non-transitory computer-readable storage medium of embodiment 9, further comprising selecting a plurality of selected labels for association with the content item. (12) determining the at least one selected label, 10. The non-transitory computer-readable storage medium of claim 9, further comprising selecting between a first predetermined label in response to the swipe gesture being performed in a first direction and a second predetermined label in response to the swipe gesture being performed in a second direction opposite the first direction. (13) determining the at least one selected label comprises: 10. The non-transitory computer-readable storage medium of claim 9, further comprising selecting between a first predetermined label in response to the swipe gesture being performed in a first direction, a second predetermined label in response to the swipe gesture being performed in a second direction, a third predetermined label in response to the swipe gesture being performed in a third direction, and a fourth predetermined label in response to the swipe gesture being performed in the third direction. (14) The non-transitory computer-readable storage medium of embodiment 9, wherein the content item includes at least one of an image, a video, and an animation. (15) The instructions, when executed, cause the one or more processors to: presenting a control element that flags the content item for review by an administrator; storing said flag in association with said content item in response to selection of said control element; 10. The non-transitory computer-readable storage medium of embodiment 9, further comprising the steps of:

[0067] (16) The instructions, when executed, cause the one or more processors to: tracking a state of the annotation job recording a progress of annotations received for a set of content items associated with the annotation job; and in response to the annotation application closing the reopening, configuring the annotation job based on the tracked state. 10. The non-transitory computer-readable storage medium of embodiment 9, further comprising the steps of: (17) A method for facilitating training of a machine learning model to predict a characteristic of a content item, including an image or video, associated with a surgical application, comprising: configuring an annotation job associated with a machine learning model to specify a predetermined set of selectable labels for labeling a content item and labeling rules indicating a number of selectable labels that may be selected for the content item by an annotator; obtaining the content item; facilitating presentation of the content item via a user interface of an annotation application; obtaining, from the annotation application, at least one selected label from the predetermined set of selectable labels; updating parameters of the machine learning model based on the at least one selected label and the content item; and A method comprising: (18) obtaining the at least one selected label, Identifying a swipe gesture performed on the user interface; selecting between the predetermined set of labels based on a direction of the swipe gesture; 18. The method of embodiment 17, comprising: (19) Obtaining the at least one selected label comprises: facilitating presentation, via the user interface, of user interface elements each associated with a respective one of the predetermined set of labels; selecting between the predetermined set of labels based on a selection of one of the user interface elements; 18. The method of embodiment 17, comprising: (20) Configuring the annotation job includes: 18. The method of embodiment 17, further comprising identifying a set of annotators associated with the job for presenting the content item.

Claims

1. A non-temporary computer-readable storage medium for storing instructions for facilitating the collection of annotations on images or videos associated with a surgical application for training a machine learning model, wherein, when the instructions are executed by one or more processors, the instructions are sent to the one or more processors, The method involves obtaining input from the management user interface of an administrator application running on a management client device via a computer network to an annotation server, wherein the input corresponds to a set of predetermined input fields, and the input specifies at least the assignment of an annotation job to a set of annotators selectable from a predetermined list of available annotators, inclusion criteria for identifying content items to be annotated in the annotation job, a predetermined set of selectable labels for the annotation job, and the number of targets for the decision for the annotation job. Based on the aforementioned input, the annotation job associated with the machine learning model is configured by the annotation server's job management engine, Retrieving a set of unannotated content items that satisfy the aforementioned inclusion criteria from the content database based on the aforementioned input, For each of the aforementioned sets of unannotated content items, The machine learning model is applied to generate predictions associated with the content items and confidence metrics associated with the predictions. Determine whether the confidence metric satisfies a predetermined confidence threshold. In response to the confidence metric not meeting the confidence threshold, the content item is added to the annotation set associated with the annotation job, To facilitate the availability of on-demand annotation sessions for a set of annotators via each annotation application running on each annotator client device, wherein during an on-demand annotation session, the annotation application sequentially presents content items from the set of annotations, obtains selections for each label through interactions captured in the user interface of the annotation application, and transmits the selections to the annotation server. The annotation server acquires and aggregates the respective labels of the presented content items obtained from the on-demand annotation session to generate aggregated label data. Based on the aggregated label data, determine whether the target number for the determination is met for each of the content items in the annotation set, In response to the determination that the target number of a given content item in the annotation set has been met, the given content item and the aggregated label data are added to the training set. Updating the parameters of the machine learning model by training the machine learning model using the aforementioned training set, The aforementioned machine learning model is stored and A non-temporary computer-readable storage medium that causes a step including the execution of a certain procedure.

2. Obtaining the selection of each of the labels means that Identifying the swipe gesture performed on the user interface, Based on the direction of the swipe gesture, a selection is made among the predetermined set of selectable labels. A non-temporary computer-readable storage medium according to claim 1, including the following:

3. Obtaining the selection of each of the labels means To facilitate the presentation of user interface elements associated with each of the predetermined set of selectable labels via the user interface, Based on the selection of one of the user interface elements, a selection is made among the predetermined set of selectable labels. A non-temporary computer-readable storage medium according to claim 1, including the following:

4. The non-temporary computer-readable storage medium according to claim 1, wherein the input further specifies a labeling rule indicating the number of selectable labels that can be selected for each of the content items by a single annotator.

5. A computer system, One or more processors, A non-temporary computer-readable storage medium that stores instructions for facilitating the collection of image or video annotations associated with a surgical application for training a machine learning model, wherein, when the instructions are executed by the one or more processors, the one or more processors are sent to the one or more processors The method involves obtaining input from the management user interface of an administrator application running on a management client device via a computer network to an annotation server, wherein the input corresponds to a set of predetermined input fields, and the input specifies at least the assignment of an annotation job to a set of annotators selectable from a predetermined list of available annotators, inclusion criteria for identifying content items to be annotated in the annotation job, a predetermined set of selectable labels for the annotation job, and the number of targets for the decision for the annotation job. Based on the aforementioned input, the annotation job associated with the machine learning model is configured by the annotation server's job management engine, Retrieving a set of unannotated content items that satisfy the aforementioned inclusion criteria from the content database based on the aforementioned input, For each of the aforementioned sets of unannotated content items, The machine learning model is applied to generate predictions associated with the content items and confidence metrics associated with the predictions. Determine whether the confidence metric satisfies a predetermined confidence threshold. In response to the confidence metric not meeting the confidence threshold, the content item is added to the annotation set associated with the annotation job, To facilitate the availability of on-demand annotation sessions for a set of annotators via each annotation application running on each annotator client device, wherein during an on-demand annotation session, the annotation application sequentially presents content items from the set of annotations, obtains selections for each label through interactions captured in the user interface of the annotation application, and transmits the selections to the annotation server. The annotation server acquires and aggregates the respective labels of the presented content items obtained from the on-demand annotation session to generate aggregated label data. Based on the aggregated label data, determine whether the target number for the determination is met for each of the content items in the annotation set, In response to the determination that the target number of a given content item in the annotation set has been met, the given content item and the aggregated label data are added to the training set. Updating the parameters of the machine learning model by training the machine learning model using the aforementioned training set, The aforementioned machine learning model is stored and A computer system that causes a computer to perform steps including those mentioned above.

6. Obtaining the selection of each of the labels means Identifying the swipe gesture performed on the user interface, Based on the direction of the swipe gesture, a selection is made among the predetermined set of selectable labels. The computer system according to claim 5, including the computer system according to claim 5.

7. Obtaining the selection of each of the labels means To facilitate the presentation of user interface elements associated with each of the predetermined set of selectable labels via the user interface, Based on the selection of one of the user interface elements, a selection is made among the predetermined set of selectable labels. The computer system according to claim 5, including the computer system according to claim 5.

8. A method for facilitating the collection of annotations for images or videos associated with a surgical application for training a machine learning model, The method involves obtaining input from the management user interface of an administrator application running on a management client device via a computer network to an annotation server, wherein the input corresponds to a set of predetermined input fields, and the input specifies at least the assignment of an annotation job to a set of annotators selectable from a predetermined list of available annotators, inclusion criteria for identifying content items to be annotated in the annotation job, a predetermined set of selectable labels for the annotation job, and the number of targets for the decision for the annotation job. Based on the aforementioned input, the annotation job associated with the machine learning model is configured by the annotation server's job management engine, Retrieving a set of unannotated content items that satisfy the aforementioned inclusion criteria from the content database based on the aforementioned input, For each of the aforementioned sets of unannotated content items, The machine learning model is applied to generate predictions associated with the content items and confidence metrics associated with the predictions. Determine whether the confidence metric satisfies a predetermined confidence threshold. In response to the confidence metric not meeting the confidence threshold, the content item is added to the annotation set associated with the annotation job, To facilitate the availability of on-demand annotation sessions for a set of annotators via each annotation application running on each annotator client device, wherein during an on-demand annotation session, the annotation application sequentially presents content items from the set of annotations, obtains selections for each label through interactions captured in the user interface of the annotation application, and transmits the selections to the annotation server. The annotation server acquires and aggregates the respective labels of the presented content items obtained from the on-demand annotation session to generate aggregated label data. Based on the aggregated label data, determine whether the target number for the determination is met for each of the content items in the annotation set, In response to the determination that the target number of a given content item in the annotation set has been met, the given content item and the aggregated label data are added to the training set. Updating the parameters of the machine learning model by training the machine learning model using the aforementioned training set, The aforementioned machine learning model is stored and Methods that include...

9. Obtaining the aforementioned selection for each of the aforementioned labels means that Identifying the swipe gesture performed on the user interface, Based on the direction of the swipe gesture, a selection is made among the predetermined set of selectable labels. The method according to claim 8, including the method described in claim 8.

10. Obtaining the aforementioned selection for each of the aforementioned labels means that To facilitate the presentation of user interface elements associated with each of the predetermined set of selectable labels via the user interface, Based on the selection of one of the user interface elements, a selection is made among the predetermined set of selectable labels. The method according to claim 8, including the method described in claim 8.

11. The method according to claim 8, wherein the input further specifies a labeling rule indicating the number of selectable labels that can be selected for each of the content items by a single annotator.

12. The selection of each of the aforementioned labels is as follows: The method according to claim 8, comprising performing a selection of one of the predetermined set of selectable labels via the user interface.

13. The selection of each of the aforementioned labels is as follows: The method according to claim 8, comprising enabling the selection of any number of predetermined sets of selectable labels via the user interface.

14. During the aforementioned on-demand annotation session, presenting the content items sequentially is: The method according to claim 8, comprising facilitating the presentation of another content item by the annotation application via the user interface of the annotation application in response to obtaining at least one selected label for the presented content item.

15. The method according to claim 8, wherein the number of targets for the determination includes the number of unique annotators that have provided at least one label to the content item.

16. Based on the direction of the swipe gesture, the selection among the predetermined set of selectable labels is: The method according to claim 9, comprising selecting between a first predetermined label in response to the swipe gesture being performed in a first direction and a second predetermined label in response to the swipe gesture being performed in a second direction opposite to the first direction.

17. Based on the direction of the swipe gesture, the selection among the predetermined set of selectable labels is: The method according to claim 9, comprising selecting from a first predetermined label in response to the swipe gesture being performed in a first direction, a second predetermined label in response to the swipe gesture being performed in a second direction, a third predetermined label in response to the swipe gesture being performed in a third direction, and a fourth predetermined label in response to the swipe gesture being performed in a fourth direction.

18. The method according to claim 8, wherein the content item includes at least one of an image, a video, or an animation.

19. The method of claim 8, wherein the annotation application further presents a control element for flagging the content item for review by an administrator, and stores the flag associated with the content item in response to the selection of the control element.

20. To facilitate the availability of the on-demand annotation session, Tracking the status of the annotation application that records the progress of annotations by the annotator during an on-demand annotation session, In response to the annotation application being closed and then reopened, the annotation application is configured based on the tracked state. The method according to claim 8, including the method described in claim 8.