A data processing method, device and product

CN122530686APending Publication Date: 2026-08-07BEIJING NORMAL UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING NORMAL UNIVERSITY
Filing Date
2026-05-21
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]然而,由于AI在复杂边界样本、跨设备图像、低质量图像等场景下的识别可能存在不稳定性

Benefits of technology

[0019]本申请实施例提供了一种数据处理方法、装置及产品,该数据处理方法包括:获取待识别图像的模型识别结果和识别自信度;其中,所述模型识别结果指示基于人工智能AI模型对所述待识别图像的识别结果,所述识别自信度指示所述模型识别结果的识别可靠程度;根据所述识别自信度和预设阈值,确定所述待识别图像的识别结果;其中,若所述识别自信度高于所述预设阈值,确定所述待识别图像的识别结果为所述模型识别结果;若所述识别自信度低于或等于所述预设阈值,获取专家对所述待识别图像的专家识别结果,将所述专家识别结果作为所述待识别图像的识别结果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122530686A_ABST
    Figure CN122530686A_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a data processing method, device and product, the data processing method takes the recognition confidence as the recognition reliability degree of the model recognition result of the to-be-recognized image, and performs dynamic switching of the AI model recognition result and the expert manual judgment based on the recognition confidence and a preset threshold. When the recognition confidence is at a high level, the recognition ability of the AI model is used to quickly give the recognition result. When the AI model recognition confidence is at a low level, the recognition task is transferred to a human expert, and the professional judgment of the expert is used to make up for the deficiency of the AI model, thereby improving the recognition accuracy of the to-be-recognized image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition, and more particularly to a data processing method, apparatus, and product. Background Technology

[0002] In the field of medical image processing and analysis, artificial intelligence (AI) has been widely used to extract features from image data and output recognition results or probability scores, thereby improving processing efficiency in resource-constrained scenarios.

[0003] However, AI recognition can be unstable in scenarios involving complex boundary samples, cross-device images, and low-quality images. For example, using AI to identify fundus images, including lesion morphology, distribution range, and fine structural features, can result in low recognition accuracy. On the other hand, providing AI output to experts for decision-making can lead to issues such as the AI ​​outputting highly confident but actually incorrect results, misleading experts, or experts making correct judgments but failing to properly reference the AI's recognition results, resulting in low recognition accuracy.

[0004] Improving the recognition accuracy of images to be identified has become a technical problem to be solved. Summary of the Invention

[0005] This application provides a data processing method, apparatus, and product for improving the recognition accuracy of images to be recognized.

[0006] In a first aspect, embodiments of this application provide a data processing method, the method comprising: Obtain the model recognition results and recognition confidence level of the image to be recognized; Wherein, the model recognition result indicates the recognition result of the image to be recognized based on artificial intelligence, and the recognition confidence level indicates the recognition reliability of the model recognition result; Based on the recognition confidence level and the preset threshold, the recognition result of the image to be recognized is determined; If the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined as the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result of the image to be recognized is obtained, and the expert recognition result is used as the recognition result of the image to be recognized.

[0007] Optionally, obtaining the model recognition result of the image to be recognized includes: The image to be identified is standardized. The standardization process includes one or more of the following: image size adjustment, pixel normalization, brightness and contrast correction, noise suppression, invalid background cropping, optic disc or macular region localization, and lesion enhancement processing. The standardized image is identified using an AI model to obtain the model's recognition result.

[0008] Optionally, the method further includes: The standardized image is identified using an AI model to obtain the recognition parameters corresponding to the model's recognition result; The identification parameters include one or more of the following: the category label corresponding to the identification result, the probability distribution, classification score, logit vector or intermediate layer features corresponding to each category; Obtaining the recognition confidence of the image to be recognized includes: The recognition confidence level is determined based on the recognition parameters.

[0009] Optionally, the method for obtaining the preset threshold includes: Obtain the training set; The training set includes high-risk samples, which indicate images that meet at least one of the following conditions: recognition confidence is lower than a preset threshold, blurry images, occluded images, abnormally exposed images or images with a field of view less than or equal to a preset field of view threshold, and inconsistent model recognition results and expert recognition results.

[0010] The threshold adjustment model is trained using the training set to obtain the preset threshold; The threshold adjustment model is used to obtain the preset threshold; the constraints of the threshold adjustment model include one or more of the following: maximizing the overall recognition accuracy, maximizing the recognition sensitivity, minimizing the specificity, minimizing the missed diagnosis rate, or minimizing the false diagnosis rate.

[0011] Optionally, if the recognition confidence level is lower than or equal to the preset threshold, the method further includes: A first interface is continuously displayed for a first preset duration; the first interface includes the image to be recognized. After continuously displaying the first preset duration, the interface switches from the first interface to the second interface and continues to display the second interface for the second preset duration; the second interface is used to obtain the expert recognition results.

[0012] The step of obtaining expert recognition results for the image to be recognized includes: In response to a trigger operation on the second interface, the expert recognition result is obtained; the trigger operation is the expert recognition result input by the expert on the second interface.

[0013] Optionally, after obtaining the expert recognition results of the image to be recognized, the method further includes: Compare the expert recognition results with the model recognition results; If the expert recognition result is inconsistent with the model recognition result, a third interface will be displayed. The third interface includes the model recognition result and the recognition confidence level; Obtain the secondary recognition result of the expert on the image to be recognized; The step of using the expert recognition result as the recognition result of the image to be recognized includes: The secondary recognition result is used as the recognition result of the image to be recognized.

[0014] Optionally, the method further includes: Obtain the expert's subjective confidence level regarding the expert identification results; The step of using the expert recognition result as the recognition result of the image to be recognized includes: Based on the subjective confidence level, the recognition result of the image to be recognized is determined.

[0015] Optionally, if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result of the image to be recognized is obtained, and the expert recognition result is used as the recognition result of the image to be recognized, including: If the recognition confidence level is lower than or equal to the threshold, the expert recognition result is used as the recognition result of the image to be recognized; a third interface is displayed; the third interface includes the model recognition result and the recognition confidence level; The expert recognition result of the image to be recognized is obtained through the third interface; the expert recognition result is used as the recognition result of the image to be recognized.

[0016] Secondly, embodiments of this application provide a data processing apparatus, the apparatus comprising: The model recognition unit is used to obtain the model recognition result and recognition confidence of the image to be recognized; Wherein, the model recognition result indicates the recognition result of the image to be recognized based on artificial intelligence, and the recognition confidence level indicates the recognition reliability of the model recognition result; The human-machine collaboration unit is used to determine the recognition result of the image to be recognized based on the recognition confidence level and a preset threshold; if the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined to be the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result is used as the recognition result of the image to be recognized. The expert recognition module is used to obtain the expert recognition results of the image to be recognized.

[0017] Thirdly, embodiments of this application provide a computer storage medium for storing a computer program; when the computer program is executed, it is used to perform a communication method as described in any of the first aspects.

[0018] Fourthly, embodiments of this application provide a computer program product containing instructions that, when run on at least one computing device, cause the at least one computing device to perform a communication method as described in any of the first aspects.

[0019] This application provides a data processing method, apparatus, and product. The data processing method includes: acquiring a model recognition result and recognition confidence level of an image to be recognized; wherein the model recognition result indicates the recognition result of the image to be recognized based on an artificial intelligence (AI) model, and the recognition confidence level indicates the reliability of the recognition result; determining the recognition result of the image to be recognized based on the recognition confidence level and a preset threshold; wherein, if the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined to be the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, acquiring an expert recognition result of the image to be recognized, and using the expert recognition result as the recognition result of the image to be recognized.

[0020] The technical solution provided in this application uses recognition confidence level as the reliability of the model's recognition result for the image to be recognized, and dynamically switches between AI model recognition results and expert judgment based on this recognition confidence level and a preset threshold. When the recognition confidence level is high, the AI ​​model's recognition capability is used to quickly provide a recognition result. When the AI ​​model's recognition confidence level is low, the recognition task is handed over to human experts, using the experts' professional judgment to compensate for the shortcomings of the AI ​​model. This avoids the possibility of errors by the AI ​​model at a high recognition confidence level misleading the expert's judgment, and also avoids the problem that the accuracy of the expert's judgment is affected by the AI ​​model's incorrect recognition result, even if the expert could have made a correct judgment. Attached Figure Description

[0021] Figure 1 A flowchart of a data processing method provided in an embodiment of this application; Figure 2 A schematic diagram illustrating a standardized process provided in an embodiment of this application; Figure 3 This is a schematic diagram of the structure of a data processing device provided in an embodiment of this application. Detailed Implementation

[0022] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. The terminology used in the following embodiments is for the purpose of describing specific embodiments only and is not intended to be a limitation of this application. As used in the specification and appended claims of this application, the singular expressions "a," "an," "the," "the," "the," and "this" are intended to also include expressions such as "one or more," unless the context clearly indicates otherwise.

[0023] It should also be understood that in the embodiments of this application, "one or more" refers to one, two, or more; "and / or" describes the correspondence between associated objects, indicating that three relationships can exist; for example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following associated objects have an "or" relationship.

[0024] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0025] The "multiple" mentioned in the embodiments of this application refers to two or more. It should be noted that in the description of the embodiments of this application, terms such as "first" and "second" are used only for the purpose of distinguishing descriptions and should not be construed as indicating or implying relative importance, nor should they be construed as indicating or implying order.

[0026] In the field of medical image processing and analysis, AI-based image recognition may produce results that are inconsistent with reality when identifying images with complex boundaries, images from different devices, and low-quality images. Specifically, when an AI model identifies an image, it may make mistakes due to feature extraction biases, thus misleading expert decisions. At the same time, experts may have their ability to make correct judgments weakened by the erroneous identification results of the AI ​​model, resulting in a decrease in recognition accuracy.

[0027] For example, in the scenario of lesion identification in fundus images, the AI ​​model may output incorrect identification results for lesion morphology containing fine structural features. If experts adopt this result and ignore the actual distribution range of the lesions during the diagnosis process, it may lead to a diagnosis error. Alternatively, experts may make a correct judgment based on image features, but adjust their conclusions due to the incorrect identification results output by the AI ​​model, which will also affect the recognition accuracy of the image to be identified.

[0028] In view of this, this application proposes a data processing method. This method first obtains the AI ​​model's recognition result and recognition confidence level of the image to be recognized; based on the recognition confidence level and a preset threshold, it determines the recognition result of the image to be recognized; if the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined as the model's recognition result; if the recognition confidence level is lower than or equal to the preset threshold, it obtains the expert recognition result of the image to be recognized and uses the expert recognition result as the recognition result of the image to be recognized. The recognition confidence level characterizes the accuracy of the AI ​​model's image recognition, and then, based on this recognition confidence level, it determines whether to use expert recognition. When the recognition confidence level is high, the AI ​​model's recognition result is directly adopted; when the confidence level is low, expert recognition is introduced. This solves the problem of model recognition errors misleading expert judgments, thus significantly improving the accuracy and reliability of image recognition.

[0029] The data processing method provided in the embodiments of this application will be described below with reference to the accompanying drawings.

[0030] It should be noted that the data processing method provided in this application can be applied to human-machine collaborative systems. A human-machine collaborative system can be a hardware device, such as a server, terminal device, computer, mobile phone, or other computing device, or a hardware module of that computing device. It can also be a software platform, such as a human-machine collaborative platform installed on a computing device; this application is not limited to any particular type.

[0031] Appendix Figure 1 This is a flowchart illustrating a data processing method provided in an embodiment of this application. Figure 1 As shown, the method includes: S10: Obtain the model recognition result and recognition confidence level of the image to be recognized.

[0032] The image to be identified refers to image data that needs to be processed and recognized. This image data can be digital images from various sources, such as medical images, industrial inspection images, or security surveillance images. For ease of explanation, the following description uses a fundus image of the retina as an example. Fundus images include the morphology, distribution range, and fine structural features of lesions, and are widely used in the screening and grading of diabetic retinopathy, age-related macular degeneration, glaucoma, and other fundus abnormalities.

[0033] The model recognition result indicates the recognition outcome of the image to be recognized based on the AI ​​model. In practical use, the model recognition result can be presented in the form of category labels, probability distributions, or scores, indicating the AI ​​model's judgment on the image content.

[0034] Recognition confidence indicates the reliability of the model's recognition results. It quantifies the certainty of the AI ​​model regarding its own recognition results. Recognition confidence can be expressed as a percentage or a probability value. For example, a higher recognition confidence value indicates stronger confidence in the AI ​​model's own recognition results. Furthermore, recognition confidence can also be output as a discrete level, such as high, medium, and low. A higher level indicates stronger confidence in the AI ​​model's own recognition results. Further, recognition confidence can be divided into more granular levels, which is not limited in the embodiments of this application.

[0035] In the embodiments of the application, the human-machine collaborative system can use a pre-trained AI model to process the image to be recognized, and obtain the model recognition result of the image to be recognized. For example, the model can be a convolutional neural network, which receives the pixel data of the original image as input and outputs one or more category labels as the model recognition result. For example, the output model recognition result is normal fundus, diabetic retinopathy, age-related macular degeneration, glaucoma, or other abnormal fundus categories.

[0036] It should be noted that the AI ​​model can also be a visual Transformer network model, a network model that combines convolution and Transformer, a classification network based on multi-scale feature extraction, or other deep learning models suitable for medical image analysis. The embodiments in this application are not limited to these.

[0037] In one example, to ensure the reliability and generalization ability of the AI ​​model's recognition results, the human-machine collaborative system can also perform standardization processing on the image to be recognized before using the AI ​​model for recognition. This standardization process is a key preliminary step in the entire recognition process. Its core lies in eliminating or significantly reducing non-essential variations in the original image introduced during the acquisition process due to factors such as differences in equipment, changes in lighting conditions, and different shooting angles through a series of image processing techniques.

[0038] In the embodiments of this application, the standardization processing may include, but is not limited to: image size adjustment to unify the image specifications of the input model; pixel normalization to eliminate differences in pixel value ranges; brightness and contrast correction to optimize the visual quality and feature recognizability of the image; noise suppression to improve image clarity; invalid background cropping to focus on the core content of the image; optic disc or macular region localization to ensure the accuracy of medical image analysis; and lesion enhancement processing to highlight key pathological features.

[0039] It should be noted that these processes are not performed in isolation, but rather one or more are selected and combined according to actual needs to form a unified image preprocessing workflow. Images that have undergone standardization have more consistent data characteristics and higher quality, thus providing stable and reliable input for subsequent AI model recognition. Based on this, the AI ​​model's recognition of the standardized images will yield more accurate and robust results.

[0040] For example, see Figure 2 As shown, this figure is a schematic diagram of a standardized processing method provided in an embodiment of this application. The human-machine collaborative system can process fundus images using the following steps: First, all input fundus images are uniformly adjusted to a size of 512×512 pixels using a bilinear interpolation algorithm to adapt to the pre-trained convolutional neural network model. Next, pixel normalization is performed on the resized images, mapping pixel values ​​from the range of 0-255 to the floating-point range of 0-1 to accelerate model training convergence. Subsequently, to compensate for brightness and contrast differences caused by shooting with different devices, an adaptive histogram equalization algorithm can be applied to correct the brightness and contrast of the images. Simultaneously, to remove speckle noise generated during image acquisition, Gaussian filtering can be used to suppress noise in the images. In some cases, to focus the model's attention on the retinal region, a semantic segmentation model can be used to locate the optic disc and macular regions in the fundus images and crop out the regions of interest containing these key areas. Finally, to make small lesions more prominent, unsharpened masking or local contrast enhancement algorithms can be applied to enhance the lesions. After the above series of standardization processes, a high-quality, highly consistent fundus image can be obtained. Inputting this image into an AI model can yield more accurate model recognition results.

[0041] Before obtaining the model's recognition results for the image to be recognized, the image is standardized to eliminate image differences caused by external factors such as equipment, lighting, and shooting conditions. This makes the data input into the AI ​​model more consistent and of higher quality, significantly improving the AI ​​model's generalization ability and recognition reliability. Therefore, it effectively reduces low confidence or misrecognition caused by image quality issues, thereby reducing interference with expert judgment, avoiding misleading results from highly confident but erroneous AI models, and preventing the negative synergy where correct expert judgments are weakened by erroneous AI model results. Ultimately, this improves the overall accuracy and stability of recognition.

[0042] Furthermore, in this embodiment, the AI ​​model can also output recognition parameters corresponding to the model recognition result. These recognition parameters include, but are not limited to, category labels, probability distributions for each category, classification scores, logit vectors, or intermediate layer features.

[0043] It's important to note that recognition parameters refer to the specific data types obtained by the AI ​​model during the recognition process. Among these, the category label is the final classification result given by the AI ​​model for the image content, such as "normal," "lesion A," and "lesion B." The probability distribution refers to the set of predicted probabilities that the AI ​​model assigns to each possible category of the input image. For example, for a binary classification problem, the AI ​​model might output [0.9, 0.1], indicating a 90% probability of belonging to the first category. The classification score is typically the raw, unnormalized score from the AI ​​model's output layer, reflecting the AI ​​model's preference strength for each category. These scores can be transformed into a probability distribution after processing. The logit vector is the AI ​​model's linear prediction for each category; its magnitude and relative relationship directly affect the final probability distribution. Intermediate layer features refer to the feature representations extracted by the AI ​​model from each hidden layer between the input and output layers in the deep network structure. These features capture abstract information at different levels of the image and can be used for more refined analysis.

[0044] Among them, the human-machine collaborative system can determine the recognition confidence level based on the recognition parameters.

[0045] For example, a human-machine collaborative system can calculate the Shannon entropy of the probability distribution. A lower entropy value indicates a more concentrated prediction by the AI ​​model for a particular category, resulting in higher recognition confidence; conversely, a higher entropy value indicates less uncertainty in the AI ​​model's predictions, leading to lower recognition confidence. Another example is that a human-machine collaborative system can determine recognition confidence by analyzing the magnitude of the logit vector or the difference between logit values ​​of different categories. For instance, a larger difference between the maximum and second-maximum logit values ​​generally indicates a more confident AI model and higher recognition confidence. Yet another example is that a human-machine collaborative system can also determine recognition confidence using intermediate layer features. For instance, by calculating the distance between intermediate layer features of the input image and features of known high-confidence samples, or by using an additional confidence prediction network with intermediate layer features as input. Furthermore, a human-machine collaborative system can also determine recognition confidence by combining multiple or all of the following: category labels, probability distributions of each category, classification scores, logit vectors, or intermediate layer features.

[0046] When using an AI model to recognize standardized images, this application simultaneously extracts various recognition parameters generated by the AI ​​model during the decision-making process. These recognition parameters contain richer and more detailed information, enabling a more comprehensive reflection of the AI ​​model's intrinsic confidence in the recognition results. Based on these internal recognition parameters, the system calculates the recognition confidence level of the image to be recognized. This confidence assessment method, based on internal AI model information rather than external heuristic rules, allows the obtained recognition confidence level to more accurately reflect the reliability of the AI ​​model's own recognition results.

[0047] In this embodiment, the human-machine collaborative system can also utilize an labeled dataset to compare the prediction results of the artificial intelligence model with gold standard labels or expert consensus labels to obtain a supervision signal indicating whether the prediction is correct or incorrect. This supervision signal is used to train the AI ​​model, ensuring that its output confidence level accurately reflects the reliability of the AI's recognition results. Furthermore, methods such as temperature scaling, ordinal-preserving regression, and / or Platt scaling can be used to post-process and calibrate the original confidence level output by the AI ​​model, thereby improving the consistency between the recognition confidence level and the actual accuracy, and providing a more reliable basis for human-machine collaborative decision-making.

[0048] S20: Determine the recognition result of the image to be recognized based on the recognition confidence level and the preset threshold.

[0049] A preset threshold is a pre-defined numerical limit used for comparison with recognition confidence. This threshold serves as the basis for decision-making, determining whether the model's recognition results are sufficiently reliable.

[0050] In this embodiment, if the recognition confidence level is higher than a preset threshold, the recognition result of the image to be recognized is determined as the model's recognition result. This means that when the AI ​​model demonstrates high reliability in its recognition results, the system directly adopts the AI ​​model's recognition result as the final recognition result. For example, if the AI ​​model's recognition confidence level for an image is 0.95, and the preset threshold is 0.80, since 0.95 is higher than 0.80, the recognition result given by the AI ​​model will be directly adopted as the final recognition result. This approach fully utilizes the efficiency advantage of AI models when processing large amounts of data.

[0051] If the AI ​​model's confidence level is lower than or equal to a preset threshold, the system obtains the expert's recognition result for the image to be recognized and uses this result as the final recognition result. This means that when the AI ​​model's confidence level is insufficient to reach the preset reliability level, the human-machine collaborative system transfers the image to a human expert for judgment, and the expert's recognition result is used as the final recognition result. For example, if the AI ​​model's confidence level for the image is 0.55, while the preset threshold is 0.80, since 0.55 is lower than 0.80, the human-machine collaborative system will prompt an expert to perform manual recognition of the image. The expert provides their professional recognition result, and the human-machine collaborative system uses this result as the final recognition result. This approach ensures that when the AI ​​model is uncertain or may err, the expertise and experience of human experts can be introduced for correction, thereby improving the accuracy of the recognition.

[0052] In this embodiment, the preset threshold can be manually set based on historical data or experience. Furthermore, the preset threshold provided in this embodiment can also be dynamically adjusted.

[0053] Furthermore, this application also provides a method for obtaining a preset threshold, the method comprising the following steps: Step 1: Obtain the training set; The training set includes high-risk samples, which are images that meet at least one of the following conditions: recognition confidence is lower than a preset threshold, blurry images, occluded images, abnormally exposed images or images with a field of view less than or equal to a preset field of view threshold, and inconsistent model recognition results with expert recognition results.

[0054] The training set includes high-risk samples to ensure that the model pays sufficient attention to situations that could lead to recognition errors or decision uncertainty during the learning process. Specifically, images with a recognition confidence level below a preset threshold indicate that the AI ​​itself is uncertain about the recognition result; blurry, occluded, abnormally exposed, or limited field of view images represent poor image quality that may affect the judgment of any subject; images where the model's recognition result is inconsistent with the expert's recognition result directly reveal the discrepancy between the AI ​​and the expert's judgment, which are points of conflict that require key attention.

[0055] It should be noted that these conditions can exist individually or in combination, together forming the basis for identifying high-risk samples.

[0056] Step 2: Use the training set to train the threshold adjustment model and obtain the preset threshold under the preset constraints.

[0057] The threshold adjustment model is used to obtain the preset threshold.

[0058] Preset constraints are key indicators guiding the threshold adjustment process of the model training. These constraints allow the acquisition of preset thresholds. These constraints may include one or more of the following: maximizing overall recognition accuracy, maximizing recognition sensitivity, minimizing specificity, minimizing the false negative rate, or minimizing the misdiagnosis rate. For example, in the recognition of fundus images, the constraint could be to maximize recognition sensitivity or minimize the false negative rate to ensure that no potential diseases are missed. In this way, the obtained preset threshold can more accurately reflect the reliability of the AI ​​model and the necessity of expert intervention under different levels of recognition confidence.

[0059] By introducing high-risk samples into the training set, the optimization process of the preset threshold can fully consider situations prone to errors, such as low recognition confidence, poor image quality, and inconsistencies between model and expert judgment. Based on this, by maximizing the overall recognition accuracy or satisfying preset constraints such as sensitivity, specificity, false negative rate, and false positive rate, the optimal threshold obtained can balance recognition performance and risk control requirements in different application scenarios. This approach helps improve the recognition accuracy and decision stability of the images to be recognized.

[0060] In one example, expert recognition results are obtained when the recognition confidence level is lower than or equal to a preset threshold to improve recognition accuracy. However, in this process, the lack of time control and interface standardization in the expert diagnosis process can lead to insufficient time for experts to observe the image or arbitrary input of recognition results, easily introducing random factors and subjective biases, affecting the consistency and repeatability of judgments. To address this, this application further proposes a first interface that continuously displays for a first preset duration if the recognition confidence level is lower than or equal to a preset threshold; the first interface includes the image to be recognized; after continuously displaying for the first preset duration, the first interface switches to a second interface, which is then continuously displayed for a second preset duration; the second interface is used to obtain expert recognition results. In response to a trigger operation on the second interface, the expert recognition results are obtained, wherein the trigger operation is the expert recognition result input by the expert on the second interface.

[0061] For example, in a lesion identification scenario for fundus images, when the human-machine collaborative system determines that expert intervention is needed based on its recognition confidence and preset thresholds, the system can present a first interface on the expert's workstation monitor. This first interface displays the fundus object to be identified and displays it continuously for 5 seconds. During this time, the expert can carefully observe the shape, size, and distribution of blood vessels, optic disc, macula, and any abnormal lesions in the image. After 5 seconds, the human-machine collaborative system automatically switches the first interface on the monitor to a second interface. This second interface may contain multiple preset recognition options, such as diabetic retinopathy, glaucoma, macular degeneration, etc., or provide a text input box for the expert to input a recognition description. Simultaneously, the second interface will continuously display a second preset duration, such as 10 seconds. In one implementation, a countdown bar can also be set on the second interface to indicate the expert's remaining decision time. Within 10 seconds, the expert needs to trigger an action, such as clicking the corresponding diagnostic option or entering the recognition result in the text box. The system uses the result triggered by the expert as the expert recognition result for the image to be identified.

[0062] The first interface, which continuously displays a first preset duration, aims to provide experts with a fixed and sufficient observation time to ensure they can fully examine the details of the image to be identified. After continuously displaying the first preset duration, it switches to a second interface, ensuring a clear separation between the observation and decision-making stages. Its function is to guide experts from a purely observational mode to a decision-making input mode. The second interface, which continuously displays a second preset duration, aims to impose time constraints on the expert's decision-making input process, reducing hesitation and external interference, and improving decision-making efficiency and consistency.

[0063] In some examples, the human-machine collaborative system can also compare the expert recognition results with the model recognition results when the recognition confidence level is lower than or equal to a preset threshold. If the expert recognition results and the model recognition results are inconsistent, indicating that the expert judgment is influenced by bias, the human-machine collaborative system displays a third interface on the expert's monitor. The third interface is used to display the model recognition results and recognition confidence level. Based on the model recognition results and recognition confidence level, combined with their own experience, the expert can perform a secondary recognition of the image to be recognized, and use the secondary recognition result as the final recognition result of the image to be recognized.

[0064] For example, still using fundus image recognition as an example, the human-machine collaborative system first acquires a fundus image to be recognized and uses an AI model to recognize it, obtaining a model recognition result, for example, mild diabetic retinopathy, with a recognition confidence level of, for example, 60%. Since this recognition confidence level is lower than a preset threshold (for example, a preset threshold of 70%), the human-machine collaborative system assigns the image to an expert for recognition. The expert, without referring to the AI ​​result, initially identifies the image as normal. At this point, the human-machine collaborative system detects a discrepancy between the expert's recognition result and the model's recognition result. The human-machine collaborative system then displays a third interface on the expert's workstation. This third interface displays the original fundus image and the model's recognition result, namely mild diabetic retinopathy and a recognition confidence level of "60%". After seeing the model's recognition result and the recognition confidence level, the expert can carefully re-examine the fundus image and, combining the model's recognition result and the recognition confidence level's indication, ultimately discover that there are indeed some tiny lesions in the image. Therefore, the expert modifies the secondary recognition result to mild diabetic retinopathy and uses the secondary recognition result as the final recognition result.

[0065] It should be noted that the third interface can also display lesion heat maps of fundus images, as well as information on areas of interest or lesion explanations, allowing experts to make a final judgment based on their knowledge of the AI's recommendations and their reliability.

[0066] By introducing comparison and secondary recognition mechanisms, the problem of experts not referring to AI model information when the results of experts and artificial intelligence are inconsistent is solved, thereby improving the accuracy of recognition.

[0067] In another example, the human-machine collaborative system can further obtain the subjective confidence level corresponding to the expert's recognition results. Subjective confidence level is used to quantify the expert's confidence level in their own judgment.

[0068] Subjective confidence can be represented using discrete levels, such as 1 to 4, 1 to 5, 1 to 7, etc., or as a continuous quantity, such as a confidence score from 0% to 100%. In one possible implementation, experts report their confidence level immediately after submitting their recognition results to reduce recall bias and additional cognitive interference. Furthermore, a confidence level input area can be set in the recognition results, or a confidence level selection interface can automatically pop up after the expert's recognition is completed. The human-machine collaborative system can determine the subjective confidence level based on the expert's triggering results on the recognition interface.

[0069] Furthermore, when the AI ​​model has low confidence in recognizing an image, the human-machine collaborative system delegates the task to an expert for recognition. The expert, while providing their recognition result, is also asked to assess their own subjective confidence level in their judgment. The human-machine collaborative system then comprehensively considers both the expert's recognition result and their subjective confidence level. For example, if the expert demonstrates high confidence in their recognition result, it will be considered the final and reliable result. Conversely, if the expert has low subjective confidence in their recognition result, the system will not blindly adopt the result but may trigger further verification processes, such as submitting the image to another expert for secondary review or marking it as a case requiring special attention. This mechanism ensures that the risk of directly adopting an expert's judgment under uncertain circumstances is mitigated when the AI ​​model's recognition is uncertain, making the human-machine collaborative decision-making process more refined and intelligent.

[0070] For example, continuing with the fundus image recognition scenario above, when the AI ​​model's confidence in identifying lesions in a fundus image is below a preset threshold, the system sends the image and its AI model recognition result to an ophthalmologist for manual review. The expert observes the image in detail on the recognition interface and inputs their recognition result, such as retinal lesions. Simultaneously with submitting the recognition result, a dialog box or a slider can pop up on the interface, asking the expert about their confidence level in their diagnosis. After receiving the expert's recognition result and subjective confidence level, the human-machine collaborative system processes the data according to a preset strategy. For example, if the expert selects a high confidence level, the human-machine collaborative system directly uses retinal lesions as the final recognition result for the fundus image. If the expert's confidence level is low, the human-machine collaborative system will mark the image as "pending review" and automatically assign it to another senior expert for secondary diagnosis, or place it in a queue requiring periodic manual review. In this way, the system can flexibly adjust the subsequent processing flow based on the expert's confidence level in their judgment, ensuring the reliability of the final diagnostic result.

[0071] Furthermore, the human-machine collaborative system can also acquire one or more of the following input learning fusion models: model recognition results, recognition confidence, subjective recognition results, subjective confidence, image quality indicators, doctor's experience level, historical individual performance parameters, and sample category information. Through the learning fusion model, the system outputs the final recognition result and the fusion probability of each result.

[0072] Among them, the learning-based fusion model can adopt logistic regression model, decision tree model, gradient boosting model, shallow neural network or other models that can achieve multi-source information fusion.

[0073] In one implementation, if the recognition confidence level is higher than a preset threshold, the recognition result of the image to be recognized is determined as the model recognition result. If the recognition confidence level is lower than or equal to the preset threshold, a third interface is displayed; the third interface includes the model recognition result and the recognition confidence level; the expert recognition result of the image to be recognized is obtained through the third interface; and the expert recognition result is used as the recognition result of the image to be recognized.

[0074] For example, still using the aforementioned fundus image recognition scenario, when the human-machine collaborative system receives an image to be recognized, it first uses an AI model to recognize it, obtaining a model recognition result, such as retinal disease, and a recognition confidence level of, for example, 0.85. At this point, the human-machine collaborative system will determine the subsequent processing procedure based on a preset threshold, for example, 0.9. If the recognition confidence level is 0.96, higher than the preset threshold of 0.9, the human-machine collaborative system will directly determine the recognition result of the fundus image as retinal disease, without expert intervention. If the recognition confidence level is 0.65, lower than the preset threshold of 0.7, the human-machine collaborative system will directly request an expert to recognize the fundus image. Further, the human-machine collaborative system will display a third interface. This third interface will display the original fundus image, the model recognition result, such as retinal disease, and the recognition confidence level, for example, 0.85. After reviewing this information, the expert can combine their own professional knowledge and experience to judge the image. For example, the expert may believe that the model recognition result is correct and confirm retinal disease; or the expert may find that the model has made a misjudgment and give a normal expert recognition result. The human-machine collaborative system ultimately uses the expert recognition results provided by the experts as the recognition result of the fundus image.

[0075] In summary, this application provides a data processing method that introduces a recognition confidence level as a measure of the reliability of the model's recognition result for the image to be recognized. Based on this recognition confidence level and a preset threshold, the method dynamically switches between the AI ​​model's recognition result and expert judgment. When the recognition confidence level is high, the AI ​​model's recognition capabilities are utilized to quickly provide a recognition result. Conversely, when the AI ​​model's recognition confidence level is low, the recognition task is transferred to human experts. The expert's professional judgment compensates for the AI ​​model's shortcomings, preventing errors that might occur when the AI ​​model's recognition confidence level is high from misleading expert judgment. It also avoids the problem that an expert's correct judgment might be affected by the AI ​​model's incorrect recognition result.

[0076] According to the method provided in the embodiments of this application, this application also provides a data processing apparatus.

[0077] Appendix Figure 3 This is a schematic diagram of a data processing apparatus provided in an embodiment of this application. The apparatus 300 includes: The model recognition unit 301 is used to obtain the model recognition result and recognition confidence of the image to be recognized; Wherein, the model recognition result indicates the recognition result of the image to be recognized based on artificial intelligence, and the recognition confidence level indicates the recognition reliability of the model recognition result; The human-machine collaboration unit 302 is used to determine the recognition result of the image to be recognized based on the recognition confidence level and a preset threshold; if the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined to be the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result is used as the recognition result of the image to be recognized. The expert recognition module 303 is used to obtain the expert recognition results of the image to be recognized.

[0078] Optionally, obtaining the model recognition result of the image to be recognized includes: The image to be identified is standardized. The standardization process includes one or more of the following: image size adjustment, pixel normalization, brightness and contrast correction, noise suppression, invalid background cropping, optic disc or macular region localization, and lesion enhancement processing. The standardized image is then identified using AI to obtain the model's recognition result.

[0079] Optionally, the model recognition unit 301 is also used for: AI is used to recognize the standardized image to obtain the recognition parameters corresponding to the model's recognition result; The identification parameters include one or more of the following: the category label corresponding to the identification result, the probability distribution, classification score, logit vector or intermediate layer features corresponding to each category; Obtaining the recognition confidence of the image to be recognized includes: The recognition confidence level is determined based on the recognition parameters.

[0080] Optionally, the method for obtaining the preset threshold includes: Obtain the training set; The training set includes high-risk samples, which indicate images that meet at least one of the following conditions: recognition confidence is lower than a preset threshold, blurry images, occluded images, abnormally exposed images or images with a field of view less than or equal to a preset field of view threshold, and inconsistent model recognition results and expert recognition results.

[0081] The threshold adjustment model is trained using the training set to obtain the preset threshold; The threshold adjustment model is used to obtain the preset threshold; the constraints of the threshold adjustment model include one or more of the following: maximizing the overall recognition accuracy, maximizing the recognition sensitivity, minimizing the specificity, minimizing the missed diagnosis rate, or minimizing the false diagnosis rate.

[0082] Optionally, if the recognition confidence level is lower than or equal to the preset threshold, the device 300 further includes a display unit, which is used to continuously display a first interface for a first preset duration; the first interface includes the image to be recognized; after continuously displaying the first preset duration, the first interface is switched to a second interface, and the second interface is continuously displayed for a second preset duration; the second interface is used to obtain the expert recognition result.

[0083] The step of obtaining expert recognition results for the image to be recognized includes: In response to a trigger operation on the second interface, the expert recognition result is obtained; the trigger operation is the expert recognition result input by the expert on the second interface.

[0084] Optionally, after obtaining the expert recognition results of the image to be recognized, the human-machine collaboration unit 302 is further configured to: compare the expert recognition results with the model recognition results; If the expert recognition result is inconsistent with the model recognition result, a third interface will be displayed. The third interface includes the model recognition result and the recognition confidence level; Obtain the secondary recognition result of the expert on the image to be recognized; The step of using the expert recognition result as the recognition result of the image to be recognized includes: The secondary recognition result is used as the recognition result of the image to be recognized.

[0085] Optionally, the human-machine collaboration unit 302 is further configured to: obtain the expert's subjective confidence level in the expert's recognition results; The step of using the expert recognition result as the recognition result of the image to be recognized includes: Based on the subjective confidence level, the recognition result of the image to be recognized is determined.

[0086] Optionally, the preset threshold includes a first threshold and a second threshold, wherein the first threshold is greater than the second threshold. If the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined as the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result of the image to be recognized is obtained, and the expert recognition result is used as the recognition result of the image to be recognized, including: If the recognition confidence level is higher than the first threshold, the recognition result of the image to be recognized is determined as the model recognition result; if the recognition confidence level is lower than or equal to the second threshold, the expert recognition result is taken as the recognition result of the image to be recognized. Optionally, the human-machine collaboration unit 302 is also used for: If the recognition confidence level is less than or equal to a preset threshold, a third interface is displayed; the third interface includes the model recognition result and the recognition confidence level; through the third interface, the expert recognition result of the expert on the image to be recognized is obtained; The expert recognition result is used as the recognition result of the image to be recognized.

[0087] This application provides a data processing apparatus that can introduce an AI model to assess the reliability of the model's recognition result for an image to be recognized based on the recognition confidence level. Based on this recognition confidence level and a preset threshold, it dynamically switches between the AI ​​model's recognition result and expert judgment. When the recognition confidence level is high, the AI ​​model's recognition capabilities are utilized to quickly provide a recognition result. Conversely, when the AI ​​model's recognition confidence level is low, the recognition task is transferred to a human expert. The expert's professional judgment compensates for the AI ​​model's shortcomings, preventing errors that might occur when the AI ​​model's recognition confidence level is high and thus avoiding the problem of expert judgments being affected by erroneous AI model results.

[0088] According to the method provided in the embodiments of this application, this application also provides a computer program product, which includes: computer program code, which, when run on a computer, causes the computer to execute the various steps or processes performed in any of the foregoing method embodiments.

[0089] According to the method provided in the embodiments of this application, this application also provides a computer-readable storage medium storing program code, which, when run on a computer, causes the computer to execute the various steps or processes performed in any of the foregoing method embodiments.

[0090] The computer-readable storage medium may be the aforementioned volatile memory or non-volatile memory, or it may include both volatile memory and non-volatile memory.

[0091] In the embodiments of this application, the terms and English abbreviations are exemplary examples given for ease of description and should not be construed as limiting the application in any way. This application does not preclude the possibility of defining other terms that can achieve the same or similar functions in existing or future agreements.

[0092] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When these computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated.

[0093] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0094] It should be understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.

[0095] In summary, the above description is merely a preferred embodiment of the technical solution of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.

Claims

1. A data processing method, characterized in that, The method includes: Obtain the model recognition results and recognition confidence level of the image to be recognized; Wherein, the model recognition result indicates the recognition result of the image to be recognized based on the artificial intelligence AI model, and the recognition confidence level indicates the recognition reliability of the model recognition result; Based on the recognition confidence level and the preset threshold, the recognition result of the image to be recognized is determined; If the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined as the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result of the image to be recognized is obtained, and the expert recognition result is used as the recognition result of the image to be recognized.

2. The method according to claim 1, characterized in that, Obtaining the model recognition result of the image to be recognized includes: The image to be identified is standardized. The standardization process includes one or more of the following: image size adjustment, pixel normalization, brightness and contrast correction, noise suppression, invalid background cropping, optic disc or macular region localization, and lesion enhancement processing. The standardized image is identified using an AI model to obtain the model's recognition result.

3. The method according to claim 2, characterized in that, The method further includes: The standardized image is identified using an AI model to obtain the recognition parameters corresponding to the model's recognition result; The identification parameters include one or more of the following: the category label corresponding to the identification result, the probability distribution, classification score, logit vector or intermediate layer features corresponding to each category; Obtaining the recognition confidence of the image to be recognized includes: The recognition confidence level is determined based on the recognition parameters.

4. The method according to claim 1, characterized in that, The method for obtaining the preset threshold includes: Obtain the training set; The training set includes high-risk samples, which indicate images that meet at least one of the following conditions: recognition confidence is lower than a preset threshold, blurry images, occluded images, abnormally exposed images or images with a field of view less than or equal to a preset field of view threshold, and inconsistent model recognition results and expert recognition results. The threshold adjustment model is trained using the training set to obtain the preset threshold; The threshold adjustment model is used to obtain the preset threshold; the constraints of the threshold adjustment model include one or more of the following: maximizing the overall recognition accuracy, maximizing the recognition sensitivity, minimizing the specificity, minimizing the missed diagnosis rate, or minimizing the false diagnosis rate.

5. The method according to claim 1, characterized in that, If the recognition confidence level is lower than or equal to the preset threshold, the method further includes: A first interface is continuously displayed for a first preset duration; the first interface includes the image to be recognized. After continuously displaying the first preset duration, the interface switches from the first interface to the second interface, and continues to display the second interface for the second preset duration; the second interface is used to obtain the expert recognition results. The step of obtaining expert recognition results for the image to be recognized includes: In response to a trigger operation on the second interface, the expert recognition result is obtained; the trigger operation is the expert recognition result input by the expert on the second interface.

6. The method according to claim 1, characterized in that, After obtaining the expert recognition results of the image to be recognized, the method further includes: Compare the expert recognition results with the model recognition results; If the expert recognition result is inconsistent with the model recognition result, a third interface will be displayed. The third interface includes the model recognition result and the recognition confidence level; Obtain the secondary recognition result of the expert on the image to be recognized; The step of using the expert recognition result as the recognition result of the image to be recognized includes: The secondary recognition result is used as the recognition result of the image to be recognized.

7. The method according to claim 1, characterized in that, The method further includes: Obtain the expert's subjective confidence level regarding the expert identification results; The step of using the expert recognition result as the recognition result of the image to be recognized includes: Based on the subjective confidence level, the recognition result of the image to be recognized is determined.

8. The method according to claim 1, characterized in that, If the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result of the image to be recognized is obtained, and the expert recognition result is used as the recognition result of the image to be recognized, including: If the recognition confidence level is lower than or equal to the threshold, a third interface is displayed; the third interface includes the model recognition result and the recognition confidence level. The expert recognition result of the image to be recognized is obtained through the third interface; the expert recognition result is used as the recognition result of the image to be recognized.

9. A data processing apparatus, characterized in that, The device includes: The model recognition unit is used to obtain the model recognition result and recognition confidence of the image to be recognized; Wherein, the model recognition result indicates the recognition result of the image to be recognized based on the artificial intelligence AI model, and the recognition confidence level indicates the recognition reliability of the model recognition result; The human-machine collaboration unit is used to determine the recognition result of the image to be recognized based on the recognition confidence level and a preset threshold; if the recognition confidence level is higher than the preset threshold, the recognition result of the image to be recognized is determined to be the model recognition result; if the recognition confidence level is lower than or equal to the preset threshold, the expert recognition result is used as the recognition result of the image to be recognized. The expert recognition module is used to obtain the expert recognition results of the image to be recognized.

10. A computer storage medium, characterized in that, Used to store computer programs; when the computer programs are executed, they are used to perform the method described in any one of claims 1-8.