Image auditing method and device based on co-creation labeling mode, equipment and medium
By combining co-creation annotation patterns with deep learning technology, an image review method has been developed that addresses the issues of low efficiency and insufficient quality control in medical image annotation. This method enables efficient and accurate generation of annotation data, adapts to diverse needs, and improves the quality of annotation data and the scalability of the system.
Patent Information
- Application Number
- CN202511410908.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-13
AI Technical Summary
Existing technologies for medical image annotation suffer from problems such as low efficiency, long processing time, inability to allocate tasks reasonably, lack of real-time feedback and quality control, and difficulty in meeting the demand for high-quality annotated data, especially in professional team annotation and group annotation methods.
An image review method based on a co-creation annotation model is adopted. Medical image data is acquired, preprocessed, and distributed to students for annotation. A deep learning model is used for preliminary quality checks and expert review, generating annotation feedback information and dynamically adjusting annotation tasks and review strategies.
It achieves efficient, accurate and high-quality image annotation review, improves the quality and usability of annotation data, adapts to diverse annotation needs in different scenarios, and enhances the system's versatility and scalability.
Smart Images

Figure CN121330431A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of this disclosure relate to computer technology, and more specifically to image review methods, apparatus, devices, and media based on co-creation annotation patterns. Background Technology
[0002] With the rapid development of artificial intelligence technology in the medical field, there is a significant demand for high-quality labeled data in medical imaging research. Currently, the main methods for obtaining high-quality medical image labeled data include: professional annotation team annotation and group annotation. The professional team annotation method typically refers to annotation performed by a professional annotation team or company. The group annotation method typically refers to the method of distributing annotation tasks to a large number of non-professional annotators through an internet platform.
[0003] However, when using the above methods to annotate medical images, the following technical problems often arise: the above professional team annotation method may result in low efficiency and long processing time when handling large-scale annotation tasks; the above group annotation method may result in the inability to reasonably allocate annotation tasks, lack of effective real-time feedback, and the need for a lot of time and energy to review and control quality.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the present disclosure concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose image review methods, apparatuses, electronic devices, and computer-readable media based on co-creation annotation patterns to explain one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure propose an image review method based on a co-creation annotation mode. The method includes: acquiring medical image data as initial data; preprocessing the initial data to obtain a task image set; distributing the task image set to at least one student terminal; in response to detecting annotation information for the corresponding task image sent by the student terminal during the annotation process, generating annotation feedback information based on the annotation information; sending the annotation feedback information to the student terminal; in response to receiving an annotation task image submitted by the student terminal, performing a preliminary quality check on the annotation task image to obtain a preliminary quality check result; sending the preliminary quality check result to an expert terminal; and in response to receiving review information corresponding to the preliminary quality check result sent by the expert terminal, sending the review information to the student terminal, wherein the review information includes the quality check result and expert annotation data.
[0008] Secondly, some embodiments of this disclosure propose an image review device based on a co-creation annotation mode. The device includes: an acquisition unit configured to acquire medical image data as initial data; a processing unit configured to preprocess the initial data to obtain a task image set; a distribution unit configured to distribute the task image set to at least one student terminal; a generation unit configured to generate annotation feedback information based on the annotation information in response to detecting annotation information of the corresponding task image sent by the student terminal during the student annotation process; a first sending unit configured to send the annotation feedback information to the student terminal; a quality inspection unit configured to perform a preliminary quality inspection on the annotation task image submitted by the student terminal in response to receiving the annotation task image submitted by the student terminal, and obtain a preliminary quality inspection result; a second sending unit configured to send the preliminary quality inspection result to an expert terminal; and a third sending unit configured to send the review information corresponding to the preliminary quality inspection result sent by the expert terminal to the student terminal in response to receiving the review information sent by the expert terminal, wherein the review information includes the quality inspection result and expert annotation data.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0011] Fifthly, some embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0012] The above embodiments of this disclosure have the following beneficial effects: The image review method based on a co-creation annotation mode, as described in some embodiments of this disclosure, can achieve efficient, accurate, and high-quality image annotation review, improving the quality and usability of the annotation data. Specifically, traditional image annotation review methods, such as those relying on professional annotation teams, may result in low efficiency and long processing times when handling large-scale annotation tasks; if relying on a group annotation mode, problems may arise such as the inability to reasonably allocate annotation tasks, a lack of effective real-time feedback, and the need for significant time and effort to review and control quality. Therefore, the image review method based on a co-creation annotation mode, as described in some embodiments of this disclosure, firstly acquires medical image data as initial data. This provides a basic data source for subsequent annotation tasks. Then, the initial data is preprocessed to obtain a task image set. This cleans and organizes the initial data to ensure it meets annotation requirements. Next, the task image set is distributed to at least one student terminal. This completes the task allocation and initiates the annotation process. Then, in response to detecting annotation information for the corresponding task image sent by the student terminal during the annotation process, annotation feedback information is generated based on the annotation information. This allows for real-time monitoring of the annotation process, providing immediate feedback to the student terminal. Next, the aforementioned annotation feedback information is sent to the student's end. This optimizes the annotation process and improves annotation quality. Then, in response to receiving the annotation task images submitted by the student's end, a preliminary quality check is performed on the annotated task images to obtain preliminary quality check results. This allows the use of a pre-trained deep learning model to filter out annotated task images with annotation errors or requiring further manual review. Then, the preliminary quality check results are sent to the expert's end. The annotated task images that pass the preliminary quality check are then sent to the expert's end for expert review. Finally, in response to receiving the review information corresponding to the preliminary quality check results from the expert's end, the review information, including the quality check results and expert annotation data, is sent to the student's end. This sends the expert's review information on the preliminary quality check results to the corresponding student's end, improving the annotation quality in the co-creation annotation mode and providing high-quality annotation data support for the application of artificial intelligence in the medical field. Furthermore, because this method can dynamically adjust the allocation of annotation tasks and review strategies based on annotation feedback information and review results, it can adapt to diverse annotation needs in different scenarios, enhancing the system's versatility and scalability. Thus, by combining the co-creation annotation model with deep learning technology, efficient, accurate, and high-quality image annotation review was achieved, improving the overall quality of the annotation data. Attached Figure Description
[0013] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0014] Figure 1 This is a flowchart of some embodiments of the image review method based on co-creation annotation patterns according to this disclosure;
[0015] Figure 2 This is a schematic diagram of the structure of some embodiments of the image review device based on the co-creation annotation mode according to the present disclosure;
[0016] Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure.
[0017] Figure 4 These are physical product images of some embodiments of this disclosure. Detailed Implementation
[0018] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0019] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0020] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0021] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0022] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0024] Figure 1 A flow 100 of some embodiments of the image review method based on co-creation annotation patterns according to this disclosure is shown. The image review method based on co-creation annotation patterns includes the following steps:
[0025] Step 101: Obtain medical imaging data as initial data.
[0026] In some embodiments, the entity executing the image review method based on the co-creation annotation pattern (e.g., a computing device) can acquire medical image data as initial data. This medical image data may include computed tomography (CT) images, magnetic resonance (MR) images, and ultrasound images. The initial data can be medical image data obtained from a hospital image archive system or a publicly available medical image website, after format conversion (e.g., from DICOM to NIfTI, PNG, or JPG format), serving as the starting point for image annotation and review. In practice, the entity can read medical image data from a hospital image archive system or a publicly available medical image website. When reading from a hospital image archive system, since these systems typically store and transmit medical image data in DICOM format, to facilitate processing and improve compatibility, the entity can use specialized tools (e.g., MATLAB or DCMtk) to store the read medical image data in a format suitable for annotation (e.g., NIfTI, PNG, JPG). The stored medical image data in the appropriate annotation format is then used as the initial data.
[0027] Step 102: Preprocess the initial data to obtain the task image set.
[0028] In some embodiments, the executing entity may preprocess the initial data to obtain a task image set. Preprocessing may be a process of generating the task image set from the initial data. The initial data includes initial images. These initial images are obtained by format transformation of the medical video data. In practice, the executing entity may perform anonymization, preliminary processing, and image reconstruction on the initial images in the initial data to obtain the task images. Then, the task images can be integrated into a task image set.
[0029] In some optional implementations of certain embodiments, the aforementioned execution entity may preprocess the initial data through the following steps to obtain a task image set:
[0030] The first step, for each initial image in the initial data, is to perform the following steps:
[0031] The first sub-step involves anonymizing the initial image to obtain an anonymized initial image. In medical image data processing, anonymization aims to remove or replace sensitive information (such as patient name, ID number, medical record number, etc.) in the image to protect patient privacy while preserving data usability. In practice, the executing entity can use image processing tools (such as PresidioImage Redactor) to identify sensitive information in the initial image. Then, data desensitization techniques (such as substitution, randomization, or masking) can be used to encrypt or replace the sensitive information to obtain the anonymized initial image. For example, for a CT image containing a patient's name, ID number, and medical record number, the executing entity can use PresidioImage Redactor to scan the CT image and obtain its sensitive information. Then, the patient's name can be replaced with "Zhang San"; the ID number can be replaced entirely or mostly with ambiguous characters (such as asterisks); and the medical record number can be replaced entirely or mostly with ambiguous characters (such as asterisks). After replacing the sensitive information in the CT image, the anonymized CT image is obtained.
[0032] The second sub-step involves preliminary processing of the anonymized initial image to obtain a pre-processed initial image. This preliminary processing includes normalization and resizing. Normalization adjusts the pixel values to a specific range (e.g., 0 to 1 or 0 to 255). Resizing adjusts the resolution of the anonymized initial image to a specific size (e.g., 256×256). In practice, the executing entity can use range normalization to map the pixel values of the anonymized initial image to a specific range, obtaining a normalized anonymized initial image. Then, interpolation methods (e.g., bilinear interpolation or bicubic interpolation) can be used to adjust the resolution of the normalized anonymized initial image to a specific range, obtaining the pre-processed initial image. For example, an anonymized initial image could be an anonymized CT image with pixel values ranging from 0 to 255 and a resolution of 512×512. A normalization formula is applied to each pixel value of this anonymized CT image to obtain a new pixel value corresponding to each pixel value, such that the range of each new pixel value is 0 to 1. Then, each pixel value of the anonymized CT image is replaced with the corresponding new pixel value to obtain the normalized anonymized CT image. Next, the resolution of the normalized anonymized CT image can be adjusted to 256×256 using bilinear interpolation to obtain the preliminarily processed anonymized CT image.
[0033] The third sub-step involves performing frequency domain mapping on the initial image from the preliminary processing, yielding the frequency domain mapping result. Frequency domain mapping can be the process of converting a signal or image from the spatial domain (or time domain) to the frequency domain. Methods for frequency domain mapping can include Discrete Fourier Transform, Two-Dimensional Fourier Transform, and Wavelet Transform. The result of the frequency domain mapping can be a complex matrix containing the frequency domain information of the initial image from the preliminary processing. This frequency domain information can be the frequency components and their distribution in the frequency domain corresponding to the initial image from the preliminary processing. Each complex element in the complex matrix includes a real part and an imaginary part. The real part can represent the cosine component of the frequency in the frequency domain information. The imaginary part can represent the sine component of the frequency in the frequency domain information. In practice, the executing entity can use the frequency domain mapping method to convert the initial image from the preliminary processing from the spatial domain to the frequency domain to obtain the frequency domain mapping result. For example, for a grayscale image, the aforementioned execution entity can use a two-dimensional Fourier transform to convert the grayscale image from the spatial domain to the frequency domain, obtaining a complex matrix containing the frequency domain information of the grayscale image, which serves as the result of the frequency domain mapping of the grayscale image.
[0034] The fourth sub-step generates spectral intensity data based on the frequency domain mapping results described above. The frequency domain mapping results can be divided into spectral intensity data and spectral phase data. The spectral intensity data can represent the intensity distribution of different frequencies in the frequency domain information. The spectral phase data can represent the phase information of different frequencies in the frequency domain information. The spectral intensity data can be a set of amplitude values at different frequencies. In practice, the executing entity can extract the modulus of the complex number from the frequency domain information based on the frequency domain mapping results. Then, amplitude values can be generated using the modulus of the complex number. Finally, amplitude values with different probabilities can be integrated into spectral intensity data.
[0035] The fifth sub-step involves normalizing the aforementioned spectral intensity data to obtain processed spectral intensity data. This normalization operation can be a scaling operation that adjusts the amplitude values of different frequencies in the spectral intensity data to a specific range. In practice, the executing entity can use maximum value normalization to scale the amplitude values of different frequencies in the spectral intensity data to between 0 and 1, thus obtaining the processed spectral intensity data.
[0036] The sixth sub-step involves generating spectral phase data based on the frequency domain mapping results. This spectral phase data can be a set of phase values with different probabilities. In practice, the executing entity can extract complex angles from the frequency domain information obtained from the frequency domain mapping results. Then, phase values can be generated using these complex angles. Finally, the phase values with different probabilities can be integrated to obtain the spectral phase data.
[0037] The seventh sub-step involves performing symmetry processing on the aforementioned spectral phase data to obtain symmetrical phase data. Symmetry processing can be a process of adjusting or transforming the data or image to satisfy a certain symmetry condition. The symmetrical phase data can be the spectral phase data after symmetry processing. The aforementioned symmetry conditions can include rotational symmetry, conjugate symmetry, and mirror symmetry. In practice, the executing entity can explicitly set the phase of the DC component (frequency intensity of 0) in the aforementioned spectral phase data to zero. Then, the aforementioned spectral phase data can be flipped (e.g., the probability axis can be reversed) to obtain flipped phase data. Then, the flipped phase data can be negatively evaluated to obtain symmetrical phase data that is conjugate symmetrical to the aforementioned spectral phase data.
[0038] The eighth sub-step involves performing phase recovery on the aforementioned symmetrical phase data to obtain phase-recovered data. Phase recovery can be a process of recovering phase data using techniques such as iterative algorithms, optimization methods, and deep learning. The phase-recovered data can be data obtained by applying a phase recovery algorithm to the symmetrical phase data. In practice, the executing entity can use a phase recovery algorithm (such as the GS algorithm or TIE algorithm) combined with the aforementioned spectral intensity data to perform phase recovery on the aforementioned symmetrical phase data to obtain phase-recovered data. For example, for a set of symmetrical phase data, the executing entity can use the GS algorithm for phase recovery. First, the symmetrical phase data can be used as an initial phase guess. Then, the spectral intensity data corresponding to the symmetrical phase data and the initial phase guess can be used to construct a complex signal as the initial complex signal. A two-dimensional Fourier transform can be performed on the initial complex signal to obtain an initial frequency domain mapping result. The frequency intensity data in the initial frequency domain mapping result is replaced with the spectral intensity data corresponding to the symmetrical phase data to obtain a modified initial frequency domain mapping result. An inverse Fourier transform can be performed on the modified initial frequency domain mapping result to obtain an initial inverse Fourier transform result. The set of phase values from the initial inverse Fourier transform result can be extracted as a new phase guess. The above steps can be repeated until the phase guess no longer undergoes significant changes (e.g., the difference is around 10^-3) or a preset number of iterations is reached (e.g., 100 times). The phase guess that no longer undergoes significant changes or has reached the preset number of iterations is used as the phase recovery data for the symmetric phase data.
[0039] The ninth sub-step generates a frequency-phase feature map based on the processed spectral intensity data and the phase recovery data. The frequency-phase feature map is a complex data structure (such as a complex matrix) that combines the intensity and phase information of a signal or image at different frequencies in the frequency domain. In practice, the executing entity can construct a complex data structure. The modulus of each complex element in this structure is determined by the processed spectral intensity data, and the angle of each complex element is determined by the phase recovery data. The resulting complex structure is used as the frequency-phase feature map.
[0040] The tenth sub-step involves generating a task image based on the aforementioned frequency-phase feature mapping. This task image can be a pre-processed initial image. In practice, the executing entity can perform an inverse Fourier transform on the aforementioned frequency-phase feature mapping, thereby converting it from the frequency domain to the spatial domain to obtain the task image.
[0041] The second step is to integrate the generated task images into a task image set. This task image set can be a database table storing information about each task image. In practice, the executing entity can integrate the individual task images into this task image set.
[0042] The first and second steps described above are an inventive point of this disclosure, addressing the technical problem that "existing medical image processing technologies suffer from insufficient patient privacy protection, low image preprocessing efficiency, and poor preprocessing effects, making it difficult to meet the high-precision requirements of medical images." The reasons why existing technologies fail to meet the needs of accurate clinical diagnosis are as follows: existing medical image processing methods are insufficient in removing sensitive information, preserving image usability, and extracting frequency domain features, making it difficult to effectively protect patient privacy. Furthermore, they have defects in the extraction and utilization of frequency domain information, resulting in low image preprocessing efficiency and failing to fully meet the needs of accurate clinical diagnosis. Solving these problems can improve image preprocessing efficiency, protect patient privacy, and enhance the utilization of frequency domain information. To achieve this effect, this disclosure employs a medical image preprocessing method based on frequency domain mapping and phase restoration. Anonymization removes sensitive information from the image, protecting patient privacy; preliminary processing (normalization and resizing) optimizes image quality; frequency domain mapping (such as two-dimensional Fourier transform) transforms the image from the spatial domain to the frequency domain, extracting spectral intensity and phase data; phase restoration algorithms recover phase information, generating a frequency-phase feature map; finally, inverse Fourier transform generates the task image, which is then integrated into a task image set. This method effectively extracts key frequency domain features from medical images, improving the efficiency and quality of image preprocessing while protecting patient privacy and enhancing the preprocessing effect, thus meeting the high-precision requirements of medical images.
[0043] Step 103: Distribute the task image set to at least one student terminal.
[0044] In some embodiments, the executing entity may distribute the task image set to at least one student terminal. The student terminal may be a platform or device for annotators to annotate the task images. The student terminal may have the function of sending and receiving information. In practice, the executing entity may send the task image set to at least one student terminal based on the complexity of the task images and the authentication information of the student terminal.
[0045] In some optional implementations of certain embodiments, the aforementioned execution entity may distribute the aforementioned task image set to at least one student terminal through the following steps:
[0046] First, for each task image in the above task image set, perform the following steps:
[0047] The first sub-step involves generating the image information entropy of the task image. The task image can be an image to be labeled. Image information entropy serves as an indicator of image complexity. It measures the distribution and information content of pixel values in an image, specifically the diversity of grayscale and color values. Higher image information entropy indicates a more uniform distribution of pixel values, greater information content, and higher complexity. In practice, the executing entity can statistically analyze the grayscale probability distribution of the task image. Then, this grayscale probability distribution can be used to generate the image information entropy of the task image.
[0048] The second sub-step generates complexity information corresponding to the task image based on the aforementioned image information entropy and a pre-set information entropy threshold. The pre-set information entropy threshold can be a pre-defined reference value used to obtain the complexity level. The complexity information typically refers to the degree of complexity. This complexity can include low complexity, medium complexity, and high complexity. In practice, the executing entity can compare the image information entropy with the pre-set information entropy threshold. Based on the comparison of information entropy levels, the complexity information corresponding to the task image is obtained. For example, for a 256-level grayscale image, the information entropy range is typically between 0 and 8. The executing entity can set information entropy thresholds of 3 and 5, compare the grayscale image's information entropy with 3 and 5. When the grayscale image's information entropy is less than 3, the grayscale complexity information is low; when the grayscale image's information entropy is between 3 and 5, the grayscale complexity information is medium; and when the grayscale image's information entropy is greater than 5, the grayscale complexity information is high.
[0049] The second step involves dividing the task image set into task image groups based on the complexity information of each task image within the set. Each task image group within a task image group can correspond to the same complexity category. These complexity categories correspond to the aforementioned complexity information and can include: low complexity, medium complexity, and high complexity. In practice, the executing entity can group task images with the same complexity information within the task image set into task image groups corresponding to the same complexity category, thus obtaining task image group sets.
[0050] The third step involves obtaining each student's proficiency information through their authentication information. Authentication information can include name or student ID. Student proficiency information typically refers to the student's proficiency level, which can be low, medium, or high. In practice, the executing entity can retrieve the corresponding student ability information (such as past grades or study time) from the student information database based on the identified authentication information from each student's end. Then, machine learning algorithms (such as logistic regression or random forest classifiers) can be used to predict the corresponding student proficiency information based on the student's ability information. The student information database can include both the student's authentication information and ability information. For example, if the executing entity identifies a student's name and student ID, it can retrieve the student's past grades (60 points, 15 hours of study time) from the student information database and predict their proficiency level as medium using a random forest classifier.
[0051] The fourth step involves grouping the student endpoints based on the aforementioned student proficiency information, resulting in a student endpoint set. Each student endpoint group within this set can correspond to a group with the same proficiency level. These proficiency level groups can include: low proficiency group, medium proficiency group, and high proficiency group. In practice, the executing entity can group student endpoints corresponding to the same proficiency level into the same student endpoint group based on their proficiency information, thus obtaining the student endpoint set.
[0052] Fifth, for each task image group in the above task image group set, perform the following steps:
[0053] The first sub-step involves selecting student client groups from the aforementioned student client set whose proficiency levels meet the criteria, based on the complexity category corresponding to the task image groups. The criteria typically involve the proficiency level of the student client group corresponding to the complexity of the task image group (e.g., low proficiency groups corresponding to low complexity categories, or high proficiency groups corresponding to high complexity categories). In practice, the executing entity can select student client groups whose image complexity matches that of the task image groups as target student groups. For example, for task image groups with medium complexity, the executing entity can select student client groups with medium proficiency levels from the aforementioned student client set.
[0054] The second sub-step involves sending the task image group to each student terminal in the target student terminal group. In practice, the executing entity can send the task images from the task image group to each student terminal in the target student terminal group. This effectively completes the corresponding allocation of task images.
[0055] Step 104: In response to detecting the annotation information of the corresponding task image sent by the student during the annotation process, generate annotation feedback information based on the annotation information.
[0056] In some embodiments, the execution entity can generate annotation feedback information based on the annotation information received from the student during the annotation process, in response to detecting annotation information of the corresponding task image. The annotation information is typically data drawn by the annotator on the task image using annotation tools to mark a certain region. This annotation information is used to indicate lesions, organs, tissues, or other structures. The annotation tools are typically software or platforms for annotating, classifying, or commenting on data (such as images, text, audio, video, etc.). The annotation information may include annotation paths and annotation regions. The annotation path may be the boundary of a polygon, rectangle, circle, or other shape drawn by the annotator on the task image using annotation tools. The annotation region may be the area enclosed by the annotation path. The annotation feedback information may include annotation region closure information and range comparison results. The annotation region closure information may be a Boolean value indicating whether the annotation region of the annotation information is closed. The range comparison results can be obtained by comparing the boundary range of the annotation path drawn by the annotator on the task image with the image range of the annotated task image. In practice, the execution entity can receive annotation information sent by the student in real time via WebSocket. Furthermore, blockchain technology can be used to collect every step of the student's annotation process. Then, edge detection algorithms (such as Canny or Sobel) can be used to identify the annotation paths in the annotation information in real time. Thus, annotation feedback information can be generated using the coordinates of the annotation paths. Integration with third-party image annotation tools can be achieved through the following steps:
[0057] The first step is to configure the API interface for communication with third-party annotation tools. This API interface can be used to receive annotation data sent by the third-party annotation tools. This annotation data includes annotation type, annotation content, and annotation structure. The API interface can include one or more endpoints (e.g., / api / annotate). The annotation data can include annotation regions, annotation information, and annotation text. In practice, the execution entity can use common web frameworks (such as Flask, Django, or Express.js) to define one or more endpoints to receive annotation data sent by third-party annotation tools.
[0058] The second step is to import the relevant specification files for the communication protocol to define it. These specification files (such as JSON files or MODBUS-related files) can include definitions of fields such as the shape of the labeled area, coordinates of the labeled path, and labeled text. The content of these specification files is the specific definition of the communication protocol. This communication protocol is typically a set of rules and standards. It can be used to define the format, order, and control mechanisms for data transmission between computer systems, networks, and devices. In practice, the executing entity can create a JSON file. A JSON parsing library (such as Python's `json`) can be used at startup to parse this JSON file to ensure that labeled data can be received and processed according to the defined communication protocol, thus obtaining the defined communication protocol.
[0059] The third step involves processing the annotation data imported by third-party annotation tools according to the defined communication protocol. This defined communication protocol typically refers to a set of rules and standards established by importing related specification documents. It standardizes the format, order, and control mechanisms for annotation data transmission between systems. The protocol defines various fields of the annotation data, such as the shape of the annotation region, the coordinates of the annotation path, and the annotation text, and clarifies the format and meaning of these fields. Third-party annotation tools (such as LabelImg or Brat) are tools capable of generating and sending annotation data according to the defined communication protocol. In practice, the executing entity can receive annotation data imported by third-party annotation tools through the API interface. Then, it verifies the format of the received annotation data according to the communication protocol. For example, if the received annotation data does not conform to the defined communication protocol, the executing entity will return an error message. If the received annotation data conforms to the defined communication protocol, it can be organized into annotation information to generate feedback information.
[0060] The first to third steps described above are an inventive point of this disclosure, addressing the problem that "existing image annotation technologies suffer from difficulties in interfacing annotation tools, inconsistent annotation data formats, and low annotation data processing efficiency, making it difficult to meet the requirements for efficient annotation." The existing image annotation methods suffer from deficiencies in annotation tool interfacing, annotation data format standardization, and annotation data processing workflows, making it difficult to achieve efficient interfacing between different annotation tools and systems. Furthermore, they have shortcomings in the uniformity of annotation data formats and processing efficiency, resulting in low annotation data processing efficiency and failing to fully meet the requirements for efficient annotation. Solving these factors can improve the efficiency of annotation tool interfacing, standardize annotation data formats, and enhance annotation data processing efficiency. To achieve this, this disclosure employs a method for interfacing with third-party annotation tools based on a communication protocol. This involves configuring an API interface for communication with third-party annotation tools to achieve efficient interfacing between the annotation tools and the system; importing relevant specification files for the communication protocol to define the format and structure of annotation data and standardize the transmission of annotation data; processing annotation data according to the defined communication protocol to verify the data format, thereby organizing annotation information and generating feedback information. This method can effectively solve the problems of difficulty in connecting annotation tools, inconsistent annotation data formats, and low efficiency in annotation data processing, thereby improving the efficiency and quality of annotation data processing and meeting the needs of efficient annotation.
[0061] In some optional implementations of certain embodiments, the aforementioned execution entity can generate annotation feedback information through the following steps:
[0062] The first step is to perform sampling quantization on the above-mentioned annotation information to obtain sampled annotation information. This sampled annotation information can include: start-point coordinates, end-point coordinates, and sampling point coordinates. The start-point coordinates can be the quantized coordinates of the first vertex of the annotation path. The end-point coordinates can be the quantized coordinates of the last vertex of the annotation path. The sampling point coordinates can be the coordinates of points selected by the executing entity from the annotation path according to a sampling method (such as fixed-interval sampling or the Douglas-Peucker algorithm). In practice, the executing entity can identify the drawing direction of the annotation path by recognizing the movement direction of the annotation tool. Based on the annotation path and its drawing direction, the specific coordinates of the start-point and end-point can be confirmed on the annotation path. The Douglas-Peucker algorithm can be used to sample the specific coordinates of the end-point outwards along the annotation path to obtain the specific coordinates of the sampling points. Then, the specific coordinates of the start-point, end-point, and sampling points are quantized to obtain the start-point coordinates, end-point coordinates, and sampling point coordinates.
[0063] The second step is to obtain the closure information of the labeled region based on the aforementioned sampling and labeling information. In practice, the executing entity can generate the Euclidean distance between the endpoint coordinates and the starting point coordinates, as well as the Euclidean distance between each sampling point coordinate and the starting point coordinates. The generated Euclidean distances can be compared with a pre-set distance threshold. For example, if any Euclidean distance is less than or equal to the aforementioned distance threshold, the closure information of the labeled region can be considered closed. If no Euclidean distance is less than or equal to the aforementioned distance threshold, the closure information of the labeled region can be considered open.
[0064] The third step is to compare the range of the annotated area with the corresponding range of the task image to obtain a range comparison result. In practice, the executing entity can obtain the coordinates of the annotated area and the coordinates of the task image. The coordinates of the annotated area and the coordinates of the task image can be compared to obtain a range comparison result. For example, if any coordinate in the annotated area is greater than the coordinates of the task image, the range comparison result is that the annotated area is out of range. If no coordinate in any annotated area is greater than the coordinates of the task image, the range comparison result is that the annotated area is within range.
[0065] The fourth step is to integrate the above-mentioned closure information of the labeled regions and the above-mentioned range comparison results into labeling feedback information. In practice, the executing entity can combine the above-mentioned closure information of the labeled regions and the above-mentioned range comparison results to obtain labeling feedback information.
[0066] Step 105: Send the annotation feedback information to the student's device.
[0067] In some embodiments, the execution entity can send the annotation feedback information to the student's client. In practice, the execution entity can send the annotation feedback information to the student's client in real time via WebSocket. For example, when a student is annotating on their client, the execution entity can provide annotation feedback information to the student's client in real time. Thus, real-time feedback can be provided while the student is performing annotation tasks.
[0068] Step 106: In response to receiving the annotation task image submitted by the student, perform a preliminary quality check on the annotation task image and obtain the preliminary quality check result.
[0069] In some embodiments, the executing entity may, in response to receiving the annotation task image submitted by the student, perform a preliminary quality check on the annotation task image to obtain a preliminary quality check result. The annotation task image may be the completed annotation task image sent by the student to the executing entity. The preliminary quality check may be a process of extracting image markers, separating annotation regions, matching standard image features, and performing conditional review to ultimately obtain a preliminary quality check result. The preliminary quality check result may be the result of reviewing the annotation status of the annotation task image through machine learning and manual review. In practice, the executing entity can match the annotation task image with a standard image. Then, it can combine the results of the manual review to obtain the preliminary quality check result.
[0070] In some optional implementations of certain embodiments, the aforementioned execution entity may perform a preliminary quality check through the following steps:
[0071] The first step is to extract image markers from the aforementioned annotation task image. Image markers can include the image within the annotation area and label information from the image annotation process. The image within the annotation area is surrounded by the task annotation path. This task annotation path can be the boundary of a polygon, rectangle, circle, or other shape drawn by the annotator on the task image using annotation tools. The label information from the image annotation process is usually filled in or selected by the annotator during the annotation process (e.g., skull or right upper arm bone), and this label information is initially considered correct. In practice, the execution entity can extract the task annotation path using image processing tools (such as OpenCV). This allows the image within the annotation area to be obtained. Then, OCR tools (such as PaddleOCR or EasyOCR) can be used to obtain the label information from the image annotation process.
[0072] The second step is to perform the following steps on the image within the above-mentioned marked area;
[0073] The first sub-step involves obtaining image cutting range data based on the image within the aforementioned labeled area. This image cutting range data can be defined by specifying a geometric shape (such as a rectangle or polygon) within the labeled image to clearly identify the area to be extracted. In practice, the executing entity can obtain the maximum and minimum coordinates of the shape formed by the labeled path using image processing tools (such as OpenCV). Then, a rectangle can be generated based on the maximum and minimum values of the x and y coordinates. Finally, the edge path of the rectangle generated based on the maximum and minimum coordinates can be used as the cutting range data.
[0074] The second sub-step involves separating the image within the labeled area from the labeled task image based on the aforementioned image cutting range data, thereby obtaining the labeled image of the labeled task image. In practice, the executing entity can use image processing tools (such as OpenCV) to separate the image within the cutting range data from the labeled task image to obtain the labeled image of the labeled task image.
[0075] The third sub-step involves inputting the labeled image into a pre-trained feature extraction model to obtain labeled image feature data. The feature extraction model can be a neural network model (such as a CNN model or a Transformer-based model) that takes the image as input and outputs the image feature data. The image feature data can be a one-dimensional feature vector of each slice of the image. The labeled image feature data can also be a one-dimensional feature vector of each slice of the labeled image. In practice, the execution entity can use a convolutional neural network (CNN) to extract the labeled image feature data. For example, the execution entity can use a CNN model as the feature extraction model. The labeled image feature data can be extracted through the following steps:
[0076] Step one involves inputting the labeled image information into the slicing layer of the feature extraction model to obtain a predetermined number of slices. This feature extraction model may include a slicing layer, a normalization layer, a whitening layer, a convolutional layer, a pooling layer, a flattening layer, and a fully connected layer. In practice, the execution entity can use the slicing layer to evenly divide the labeled image into a predetermined number of slices (e.g., 10).
[0077] Step two involves inputting the predetermined number of slices into the normalization layer to obtain normalized image slices. The normalization layer can perform mean removal and normalization. In practice, the executing entity can use the normalization layer to adjust the feature distribution of the predetermined number of slices to have zero mean and unit variance.
[0078] Step three involves inputting the normalized image slices into the whitening layer to obtain further processed image slices. The whitening layer is used to eliminate redundancy between features, making them independent. In practice, the execution entity can use the whitening layer to decompose eigenvalues or singular values to achieve zero covariance between features.
[0079] Step four involves inputting the further processed image slices into the convolutional layer to obtain feature maps for each slice. The kernel size of this convolutional layer is typically 3x3 or 5x5. The number of kernels is typically 16 or 32. The stride is typically 1. The padding is typically the same (keeping the input and output sizes identical). The activation function is typically ReLU. In practice, a 1x1 kernel can be introduced for weight adjustment.
[0080] Step 5: Input the feature maps of each slice into the pooling layer to obtain the pooled feature maps of each slice. The pooling window size can be set to 2×2 or 3×3, etc. The stride of the pooling layer is usually the same as the pooling window size (e.g., stride of 2), but can also be set to other values to achieve overlapping pooling. In practice, the execution entity can use average pooling on the feature maps of each slice, generating an average value within the pooling window as the output to obtain the pooled feature maps of each slice.
[0081] Step six: Input the pooling feature maps of each slice into the flattening layer to obtain a one-dimensional vector for each slice. The flattening layer typically transforms a multi-dimensional vector into a one-dimensional vector through a flattening operation. In practice, the execution entity can use the flattening layer to flatten the pooling feature maps of each slice into one-dimensional vectors for input into the fully connected layer.
[0082] Step seven involves inputting the one-dimensional vectors of each slice into the fully connected layer to obtain labeled image feature data. In practice, the execution entity performs a linear transformation on the one-dimensional vectors of each slice through the fully connected layer and introduces a non-linear activation function to obtain the one-dimensional feature vectors of each slice of the labeled image. These one-dimensional feature vectors of each slice of the labeled image can then be used as the labeled image feature data to obtain the labeled image feature data.
[0083] Steps one through seven above constitute an inventive point of this disclosure, solving the technical problem that "existing medical image feature extraction technologies suffer from inaccurate feature extraction, high feature dimensionality, and insufficient feature integration, making it difficult to meet image annotation and review requirements." The reasons why existing technologies cannot meet image annotation and review requirements are as follows: existing medical image feature extraction methods are deficient in feature extraction, dimensionality reduction, and feature integration, making it difficult to effectively extract key features from images. Simultaneously, high feature dimensionality increases computational complexity, and insufficient feature integration makes it difficult for the model to learn comprehensive image features, thus failing to meet image annotation and review requirements. Solving these factors can improve the accuracy of feature extraction, reduce feature dimensionality, and enhance feature integration. To achieve this effect, this disclosure employs a neural network model based on a convolutional neural network (CNN) as the feature extraction model. The image is sliced using a slicing layer, allowing subsequent processing units to process each slice individually, improving efficiency; a normalization layer adjusts the image feature distribution; and a whitening layer eliminates redundant features, enhancing feature independence. This feature extraction model extracts multiple image features by using convolutional layers to extract local features from slices; pooling layers reduce feature dimensionality and computational cost; flattening layers flatten feature maps for easier processing by subsequent fully connected layers; and fully connected layers integrate slice features to obtain comprehensive feature data. This model improves feature extraction accuracy, reduces dimensionality, and enhances integration, meeting the requirements for image annotation and review. It enables the model to learn comprehensive and accurate feature information, providing precise evidence for image annotation and review.
[0084] Third, for the label information obtained in the above image annotation process, perform the following steps:
[0085] The first sub-step involves organizing the label information from the image annotation process into directory and classification information consistent with the standard image library, thus obtaining the organization information for the annotated task images. This organization information includes directory and classification information. Each standard image in the standard image library has corresponding organization information. Each standard image in the standard image library and its corresponding organization information can together constitute an image unit file. The basic building block of the annotated image library is the annotated image group. The grouping method for these annotated image groups can be determined by the directory and classification information. The directory information determines the first-level grouping, and the classification information determines the second-, third-, or higher-level groupings. The image unit files of the standard image library include: image files that have undergone conditional review (successful characterization), image files that have passed manual review, and image files imported from the standard knowledge base, etc.
[0086] The second sub-step involves determining the standard image group corresponding to the aforementioned organizational information in the standard image library as the first standard image group, based on the organizational information described above. The first standard image group can be the standard image group in the standard image library corresponding to the organizational information of the labeled task image. In practice, the executing entity can retrieve the corresponding standard image group from the standard image library according to the organizational information of the labeled task image. Then, this standard image group is determined as the first standard image group.
[0087] The third sub-step involves extracting standard images from the first standard image group that meet preset intra-group similarity conditions as the first target object. The preset intra-group similarity condition can be the highest sum of similarity and correlation with other images within the group. The similarity is typically the degree of similarity in certain features or attributes. The correlation is typically the strength and direction of the linear relationship between image features. The first target object can be a standard image from the first standard image group that meets the preset intra-group similarity conditions. In practice, the executing entity can use an image processing library (such as Pillow in Python) to load the standard images from the first standard image group and convert them to RGB format to obtain a uniform standard image. Then, Lanczos interpolation can be used to scale the uniform standard image to a fixed size (e.g., 64×64 pixels) to obtain a standard image of the same size. Afterward, the standard image of the same size is normalized to normalize the pixel values to the range [0, 1] to obtain a processed image. Then, the pixels of the processed image are flattened into a one-dimensional vector to obtain a one-dimensional vector of the processed image. Next, cosine similarity is generated between the one-dimensional vectors of every two processed images, resulting in cosine similarity between each processed image and all other processed images. Then, the cosine similarities between each processed image and all other processed images are summed to obtain the total similarity for each processed image. These total similarity sums are then compared to find the one with the highest similarity. Finally, the processed image corresponding to the maximum similarity sum is identified as the first target object.
[0088] The fourth sub-step involves inputting the aforementioned first target object into the pre-trained feature extraction model to obtain the feature data of the first target object. This feature data can be a one-dimensional feature vector of each slice of the first target object. In practice, the steps for extracting the feature data of the first target object can refer to the steps for extracting the feature data of the labeled image, and will not be elaborated upon here.
[0089] The fourth step involves performing similarity matching on the labeled image feature data and the first target object feature data to obtain a matching result. This matching result is a comparison between the total cosine similarity and a pre-set similarity threshold. The total cosine similarity is the sum of the cosine similarities generated by the one-dimensional feature vectors of each slice of the labeled image and the one-dimensional feature vectors of each slice of the first target object. The pre-set similarity threshold includes a maximum threshold and a minimum threshold. In practice, the executing entity can generate the total cosine similarity between the one-dimensional vectors of each slice in the labeled image feature data and the one-dimensional vectors of each slice in the first target object feature data. Then, the total cosine similarity can be compared with the pre-set similarity threshold. When the total cosine similarity is greater than or equal to the maximum threshold, the matching result is considered successful; when the total cosine similarity is less than the maximum threshold, the matching result is considered unsuccessful.
[0090] The fifth step involves conditionally reviewing the matching results to obtain preliminary quality check results. Conditional review refers to a preliminary assessment of the matching results to determine whether further manual review or other processing methods are needed. The preliminary quality check results refer to the results obtained after further processing following the conditional review. These preliminary quality check results are based on subsequent information collected during the conditional review process and feedback from the manual review. Therefore, preliminary quality check results can be obtained through conditional review.
[0091] In some optional implementations of certain embodiments, the aforementioned executing entity may perform conditional audits on the matching results through the following steps to obtain preliminary quality check results:
[0092] The first step, in response to the confirmation that the above matching result representation is successful, is to mark the above annotation task image as correct. In practice, the above executing entity can mark the above annotation task image corresponding to the above matching result as correct by using preset information (such as text, graphics, or color changes) that represents the correct annotation.
[0093] The second step involves sending the labeled task image and the matching result to a human review terminal. This human review terminal can be a platform or device used for manual inspection or verification of the labeled task image and the matching result. It should have the ability to send and receive information. In practice, the executing entity can send the labeled task image and the result indicating correct matching to the human review terminal.
[0094] The third step involves receiving the results submitted by the human reviewer and generating preliminary quality check results based on these results. The results submitted by the human reviewer can be the review results sent by the human reviewer to the executing entity. These human review results can serve as preliminary quality check results. In practice, the executing entity can store the labeled images corresponding to the correctly labeled task images in the results submitted by the human reviewer into the labeled image library according to the corresponding organizational information. The executing entity can then use the results submitted by the human reviewer as preliminary quality check results. Thus, preliminary quality check results can be obtained.
[0095] In some optional implementations of certain embodiments, the aforementioned executing entity may perform conditional audits on the matching results through the following steps to obtain preliminary quality check results:
[0096] The first step, in response to the determination that the above matching result representation failed, is to identify each standard image group in the above standard image library that differs from the above first standard image group as a second standard image group. Representation failure includes representation not being determined and representation mismatch. The condition for representation not being determined is that the total cosine similarity is less than the above maximum threshold and greater than the above minimum threshold. The condition for representation mismatch is that the total cosine similarity is not greater than the above minimum threshold. Each second standard image group can be any standard image group in the standard image library other than the above first standard image group. In practice, the executing entity can identify each standard image group in the above standard image library other than the above first standard image group as a second labeled image group.
[0097] The second step involves identifying, for each of the aforementioned second standard image groups, standard images within that group that meet the preset intra-group similarity criteria as second target objects. Each second target object can be a standard image within the aforementioned second standard image groups that meets the preset intra-group similarity criteria. In practice, the steps for determining each second target object can refer to the steps for extracting the first target object, and will not be elaborated upon here.
[0098] The third step involves inputting each identified second target object into the pre-trained feature extraction model to obtain feature data for each second target object. This feature data can be a one-dimensional feature vector of each slice of the second target object. In practice, the steps for obtaining feature data for each second target object can refer to the steps for extracting labeled image feature data, and will not be elaborated upon here.
[0099] The fourth step involves performing similarity matching on the feature data of each second target object and the feature data of the marked image, respectively, to obtain the review matching results. These review matching results include the matching results for each second target object. These matching results can be the comparison between the total cosine similarity between the feature data of each second target object and the feature data of the marked image and the pre-set similarity threshold. In practice, the steps for obtaining the matching results for each second target object can refer to the steps for performing similarity matching on the feature data of the marked image and the feature data of the first target object to obtain the matching results. Then, the matching results for each second target object can be integrated as the review matching results.
[0100] Fifth, in response to the determination that the above-mentioned review and matching result indicates that the marked image feature data and any second target object feature data have successfully matched, the marked task image and the above-mentioned review and matching result are sent to the manual review terminal. The condition for the review and matching result to indicate a successful match is that the total cosine similarity between the marked image feature data and any second target object feature data is not less than the above-mentioned maximum threshold. In practice, the executing entity can send the marked task image and the review and matching result indicating that the marked image feature data has successfully matched any second target object feature data to the manual review terminal.
[0101] Step 6: In response to the received results from the human reviewer, a preliminary quality check result is generated based on these results. In practice, the aforementioned implementing entity can store the labeled images corresponding to the correctly labeled task images in the results submitted by the human reviewer into the labeled image library according to the corresponding organizational information. The aforementioned implementing entity can use the results submitted by the human reviewer as the preliminary quality check result.
[0102] Step 7: In response to the determination that the review matching result indicates that the marked image feature data does not match the feature data of each second target object, the marked task image is removed, and the preset information indicating the labeling error is determined as the preliminary quality check result. The condition for the review matching result to indicate a mismatch is that the total cosine similarity between the marked image feature data and the feature data of each second target object is not greater than the minimum threshold mentioned above. In practice, the executing entity can remove the marked task image from the preliminary quality check and use the preset information indicating the labeling error (such as text descriptions or graphic representations) as the preliminary quality check result.
[0103] Step 8: In response to the determination that the review matching result representation cannot determine whether the marked image feature data matches any second target object feature data, the second target object corresponding to any second target object feature data is identified as the third target object. The condition for the review result representation not being able to determine whether a match is possible is that the total cosine similarity between the marked image feature data and any second target object feature data is greater than the minimum threshold and less than the maximum threshold. The third target object is the second target object corresponding to any second target object feature data among the aforementioned second target object feature data that cannot be determined by the review matching result representation of the marked image feature data. In practice, the executing entity can extract the second target object feature data that cannot be determined by the review matching result representation of the marked image feature data, and identify the second target object corresponding to the extracted second target object data as the third target object.
[0104] The ninth step involves identifying the standard images in the second standard image group corresponding to the third target object that differ from the third target object as the third standard image group. This third standard image group includes all other standard images in the second standard image group corresponding to the third target object within the standard image library, excluding the third target object itself. In practice, the executing entity can retrieve the second standard image group containing the third target object from the standard image library using the third target object's organizational information. Then, the other standard images in this second standard image group can be extracted and integrated into the third standard image group.
[0105] Step 10: Input each third standard image in the aforementioned third standard image group into the pre-trained feature extraction model to obtain feature data for each third standard image. The feature data for each third standard image can be a one-dimensional feature vector of each slice of the third standard image. In practice, the steps for obtaining the feature data of each third standard image can refer to the steps for extracting the feature data of the labeled image, and will not be repeated here.
[0106] Step 11: Perform similarity matching on the aforementioned third standard image feature data and the aforementioned marked image feature data to obtain the three-stage matching results. The three-stage matching results include the matching results of each third standard image. The matching results of each third standard image can be the comparison result of the total cosine similarity between the feature data of each third marked image and the feature data of the aforementioned marked image and the aforementioned preset similarity threshold. In practice, the steps for obtaining the matching results of each third standard image can refer to the steps for performing similarity matching on the aforementioned marked image feature data and the aforementioned first target object feature data to obtain the matching results. Then, the matching results of each third standard image can be used as the three-stage matching results.
[0107] Step 12: In response to the inability of the three-stage matching results to determine whether the labeled image feature data matches the feature data of each third-standard image, the labeled task image is sent to the manual review end. The condition for the inability to determine a match in the three-stage matching results is that the total cosine similarity between the labeled image feature data and the feature data of each third-standard image is greater than the minimum threshold and less than the maximum threshold. In practice, the executing entity can send the labeled task image to the manual review end.
[0108] Step 13: In response to receiving the results submitted by the human reviewer, a preliminary quality check result is generated based on these results. In practice, the aforementioned executing entity can store the labeled image corresponding to the correctly labeled task image represented by the results submitted by the human reviewer into the labeled image library according to the corresponding organizational information. The aforementioned executing entity can use the results submitted by the human reviewer as the preliminary quality check result. Thus, a preliminary quality check result can be obtained.
[0109] Step 107: Send the preliminary quality inspection results to the expert.
[0110] In some embodiments, the executing entity may send the preliminary quality check results to the expert. In practice, the executing entity may send the preliminary quality check results to the expert in JSON or XML format via a pre-defined network interface (such as an API interface).
[0111] Step 108: In response to receiving the review information corresponding to the preliminary quality inspection results sent by the expert, the review information is sent to the student.
[0112] In some embodiments, the executing entity may, in response to receiving review information corresponding to the preliminary quality check results sent by the expert terminal, send the review information to the student terminal. The review information includes the quality check results and expert annotation data. In practice, the executing entity may use a preset network communication protocol (e.g., HTTP / HTTPS or TCP / IP) to receive the review information from the expert terminal. The review information is encapsulated in a structured data format (e.g., JSON or XML), which includes the quality check results and expert annotation data. Subsequently, the review information can be sent to the student terminal corresponding to the annotation task image via the preset network communication protocol.
[0113] The above embodiments of this disclosure have the following beneficial effects: The image review method based on a co-creation annotation mode, as described in some embodiments of this disclosure, can achieve efficient, accurate, and high-quality image annotation review, improving the quality and usability of the annotation data. Specifically, traditional image annotation review methods, such as those relying on professional annotation teams, may result in low efficiency and long processing times when handling large-scale annotation tasks; if relying on a group annotation mode, problems may arise such as the inability to reasonably allocate annotation tasks, a lack of effective real-time feedback, and the need for significant time and effort to review and control quality. Therefore, the image review method based on a co-creation annotation mode, as described in some embodiments of this disclosure, firstly acquires medical image data as initial data. This provides a basic data source for subsequent annotation tasks. Then, the initial data is preprocessed to obtain a task image set. This cleans and organizes the initial data to ensure it meets annotation requirements. Next, the task image set is distributed to at least one student terminal. This completes the task allocation and initiates the annotation process. Then, in response to detecting annotation information for the corresponding task image sent by the student terminal during the annotation process, annotation feedback information is generated based on the annotation information. This allows for real-time monitoring of the annotation process, providing immediate feedback to the student terminal. Next, the aforementioned annotation feedback information is sent to the student's end. This optimizes the annotation process and improves annotation quality. Then, in response to receiving the annotation task images submitted by the student's end, a preliminary quality check is performed on the annotated task images to obtain preliminary quality check results. This allows the use of a pre-trained deep learning model to filter out annotated task images with annotation errors or requiring further manual review. Then, the preliminary quality check results are sent to the expert's end. The annotated task images that pass the preliminary quality check are then sent to the expert's end for expert review. Finally, in response to receiving the review information corresponding to the preliminary quality check results from the expert's end, the review information, including the quality check results and expert annotation data, is sent to the student's end. This sends the expert's review information on the preliminary quality check results to the corresponding student's end, improving the annotation quality in the co-creation annotation mode and providing high-quality annotation data support for the application of artificial intelligence in the medical field. Furthermore, because this method can dynamically adjust the allocation of annotation tasks and review strategies based on annotation feedback information and review results, it can adapt to diverse annotation needs in different scenarios, enhancing the system's versatility and scalability. Thus, by combining the co-creation annotation model with deep learning technology, efficient, accurate, and high-quality image annotation review was achieved, improving the overall quality of the annotation data.
[0114] Further reference Figure 4 The image shown is a physical illustration of a product prepared by the applicant according to an embodiment of this disclosure. Detailed descriptions can be found above.
[0115] like Figure 4 As shown, it demonstrates the practical application interface of the image review method based on the co-creation annotation mode. From Figure 4 As can be seen, Figure 4 The interface is primarily used to manage and view annotation tasks for medical image data and their review status. The top of the interface displays the current time as 10:30 and includes Wi-Fi and battery level icons, indicating that the interface may be an application running on a mobile device or electronic device. At the top of the interface are two task completion status tabs: "Incomplete" and "Completed." Users can switch between or view annotation tasks in different statuses by clicking these tabs. For example, if the currently selected tab is "Completed," it means the device is currently displaying completed annotation tasks. Next, in the middle of the interface, there are multiple category tabs, including "All," "Upper Limbs," "Lower Limbs," "Spine," "Neck," "Chest and Abdominal Wall," and "Head." These category tabs are used for categorizing and managing annotation tasks. Users can click these tabs to filter and view image annotation tasks for specific body parts. For example, if the currently selected tab is "All," it means the device is currently displaying all types of annotation tasks. Then, in the main body of the interface, several specific annotation tasks are listed. Each specific annotation task includes a task name, number of tasks, and annotation status. For example, Figure 4 The task name is "Intermuscular groove approach to brachial plexus ultrasound imaging", which indicates that the annotation task is to annotate ultrasound images of the brachial plexus via the intermuscular groove approach. Figure 4 The "6 questions in total" or "2 questions in total" indicates the number of tasks in the corresponding annotation task. The task number above shows the number of annotations included in the corresponding annotation task. Figure 4 The phrase "6 questions in total" indicates that there are 6 specific annotation sub-tasks under this annotation task. Figure 4 The annotation status for each annotation task is displayed to the right of that task. Figure 4 The "View Review" option indicates that the annotation task has been completed and reviewed, and the corresponding review results can be viewed; "Waiting for Review" indicates that the annotation task has been submitted but the review has not yet been completed, and the corresponding review results have not been obtained. Figure 4 Each specific annotation task also includes a thumbnail, which shows an example of annotation on an ultrasound image. These thumbnails help users quickly understand the specific content and results of the annotation task.
[0116] Overall, Figure 4An application interface for an example of an image review method based on a co-creation annotation pattern is presented. This interface helps users efficiently manage and review a large number of medical image annotation tasks through clear categorization and status display.
[0117] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of an image review device based on a co-creation annotation mode. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0118] like Figure 2 As shown, an image review device 200 based on a co-creation annotation mode in some embodiments includes: an acquisition unit 201, a processing unit 202, a distribution unit 203, a generation unit 204, a first sending unit 205, a quality inspection unit 206, a second sending unit 207, and a third sending unit 208. The system comprises the following components: an acquisition unit 201, configured to acquire medical image data as initial data; a processing unit 202, configured to preprocess the initial data to obtain a task image set; a distribution unit 203, configured to distribute the task image set to at least one student terminal; a generation unit 204, configured to generate annotation feedback information based on the annotation information sent by the student terminal during the annotation process, in response to detecting annotation information for the corresponding task image sent by the student terminal; a first sending unit 205, configured to send the annotation feedback information to the student terminal; a quality check unit 206, configured to perform a preliminary quality check on the annotation task image submitted by the student terminal, in response to receiving the annotation task image, and obtain a preliminary quality check result; a second sending unit 207, configured to send the preliminary quality check result to an expert terminal; and a third sending unit 208, configured to send the review information corresponding to the preliminary quality check result sent by the expert terminal to the student terminal, wherein the review information includes the quality check result and expert annotation data.
[0119] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0120] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0121] like Figure 3 As shown, the electronic device 300 may include a processing unit 301 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0122] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0123] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0124] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0125] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0126] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire medical film and television data as initial data; preprocess the initial data to obtain a task image set; distribute the task image set to at least one student terminal; in response to detecting annotation information for the corresponding task images sent by the student terminal during the annotation process, generate annotation feedback information based on the annotation information; send the annotation feedback information to the student terminal; in response to receiving the annotated task images submitted by the student terminal, perform a preliminary quality check on the annotated task images to obtain a preliminary quality check result; send the preliminary quality check result to an expert terminal; in response to receiving review information corresponding to the preliminary quality check result sent by the expert terminal, send the review information to the student terminal, wherein the review information includes the quality check result and expert annotation data.
[0127] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0129] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a processing unit, a distribution unit, a generation unit, a first sending unit, a quality inspection unit, a second sending unit, and a third sending unit. The names of these units do not necessarily limit the specific unit; for example, an acquisition unit may also be described as "a unit that acquires medical video data as initial data."
[0130] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0131] Some embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements any of the above-described image review methods based on co-creation annotation patterns.
[0132] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. An image review method based on a co-creation annotation pattern, comprising: Medical film and television data were used as initial data. The initial data is preprocessed to obtain a set of task images; Distribute the task image set to at least one student terminal; In response to detecting annotation information of the corresponding task image sent by the student during the annotation process, annotation feedback information is generated based on the annotation information; The annotation feedback information is sent to the student's terminal; In response to receiving the annotation task image submitted by the student, a preliminary quality check is performed on the annotation task image to obtain the preliminary quality check result; Send the preliminary quality inspection results to the expert's terminal; In response to receiving the review information corresponding to the preliminary quality inspection result sent by the expert terminal, the review information is sent to the student terminal, wherein the review information includes the quality inspection result and expert annotation data.
2. The method according to claim 1, wherein, The step of distributing the task image set to at least one student terminal includes: For each task image in the task image set, perform the following steps: The image information entropy of the generated task image; Based on the image information entropy and the pre-set information entropy threshold, image complexity information corresponding to the task image is generated; Based on the image complexity information of each task image in the task image set, the task image set is divided to obtain task image groups, wherein each task image group in the task image group set corresponds to the same image complexity category. Obtain each student's proficiency information through the authentication information on each student's end; Based on the proficiency information of each student, the student terminals are grouped to obtain student terminal sets, wherein each student terminal group in the student terminal set corresponds to the same proficiency level. For each task image group in the task image group set, perform the following steps: Based on the image complexity category corresponding to the task image group, select the student end group that meets the proficiency level condition from the student end group set as the target student end group; The task image group is sent to each student terminal in the target student terminal group.
3. The method according to claim 1, wherein, The step of generating annotation feedback information based on the annotation information includes: The labeled information is sampled and quantized to obtain sampled labeled information; Based on the sampling and labeling information, obtain the labeling region closure information of the labeling information; The range of the labeled area is compared with the range of the task image corresponding to the labeled information to obtain the range comparison result; The closure information of the labeled region and the range comparison result are integrated into the labeled feedback information.
4. The method according to claim 1, wherein, The preliminary quality check of the labeled task image, to obtain the preliminary quality check results, includes: Extract image tags from the labeled image, wherein the image tags include the image within the labeled area and the label information from the image labeling process; For the image within the labeled area, perform the following steps; Based on the image within the marked area, obtain the image cutting range data; Based on the image cutting range data, the image within the labeled area is separated from the labeled task image to obtain the labeled image of the labeled task image; The labeled image is input into a pre-trained feature extraction model to obtain labeled image feature data; For the label information in the image annotation process, the following steps are performed: The label information from the image annotation process is organized into catalog information and classification information consistent with the standard image library to obtain the organization information of the annotation task image, wherein the organization information includes catalog information and classification information; Based on the tissue information, the standard image group corresponding to the tissue information in the standard image library is determined as the first standard image group; Extract standard images from the first standard image group that meet the preset similarity conditions within the group as the first target object; The first target object is input into the pre-trained feature extraction model to obtain the feature data of the first target object; The marked image feature data and the first target object feature data are similarly matched to obtain a matching result; The matching results are conditionally reviewed to obtain preliminary quality check results.
5. The method according to claim 4, wherein, The conditional review of the matching results to obtain preliminary quality check results includes: In response to determining that the matching result representation is successful, the labeled task image is marked as correct; The labeled task image and the matching result are sent to the manual review terminal; In response to receiving the results submitted by the human reviewer, a preliminary quality inspection result is generated based on the results submitted by the human reviewer.
6. The method according to claim 5, wherein, The conditional review of the matching results to obtain preliminary quality check results includes: In response to determining that the matching result characterization has failed, each standard image group in the standard image library that is different from the first standard image group is identified as a second standard image group. For each of the second standard image groups, the standard image in the second standard image group that meets the preset similarity condition within the group is determined as the second target object; Each of the identified second target objects is input into the pre-trained feature extraction model to obtain feature data for each second target object; Similarity matching is performed on the feature data of each second target object and the feature data of the marked image to obtain the review matching results; In response to determining that the review matching result indicates that the labeled image feature data is successfully matched with the feature data of any second target object, the labeled task image and the review matching result are sent to the manual review terminal. In response to receiving the results submitted by the human reviewer, a preliminary quality inspection result is generated based on the results submitted by the human reviewer. In response to the determination that the review matching result indicates that the marked image feature data does not match the feature data of each second target object, the marked task image is removed, and the preset information indicating the labeling error is determined as the preliminary quality check result; In response to the determination that the review matching result indicates that it is impossible to determine whether the marked image feature data matches any second target object feature data, the second target object corresponding to any second target object feature data is determined as the third target object; Each standard image in the second standard image group that is different from the third target object is identified as the third standard image group; Each third standard image in the third standard image group is input into the pre-trained feature extraction model to obtain feature data of each third standard image. Similarity matching is performed on each of the third standard image feature data and the labeled image feature data to obtain the three-stage matching results; In response to the fact that the three-stage matching result indicates that it is impossible to determine whether the labeled image feature data matches the feature data of each third standard image, the labeled task image is sent to the manual review terminal. In response to receiving the results submitted by the human reviewer, a preliminary quality inspection result is generated based on the results submitted by the human reviewer.
7. An image review device based on a co-creation annotation mode, comprising: The acquisition unit is configured to acquire medical film and television data as initial data; The processing unit is configured to preprocess the initial data to obtain a set of task images; A distribution unit is configured to distribute the task image set to at least one student terminal; The generation unit is configured to generate annotation feedback information based on the annotation information in response to detecting annotation information of the corresponding task image sent by the student terminal during the student annotation process. The first sending unit is configured to send the annotation feedback information to the student terminal; The quality inspection unit is configured to perform a preliminary quality inspection on the labeled task image in response to receiving a labeled task image submitted by a student, and obtain a preliminary quality inspection result. The second sending unit is configured to send the preliminary quality inspection results to the expert terminal; The third sending unit is configured to send the review information corresponding to the preliminary quality inspection result to the student terminal in response to receiving the review information sent by the expert terminal, wherein the review information includes the quality inspection result and expert annotation data.
8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.
Citation Information
Cited By
Medical image intelligent labeling and auditing method based on deep learning
CN121768605A
A medical image intelligent labeling and auditing method based on deep learning
CN121768605B