Human-machine collaborative calibration method, system, device and storage medium
By using a dynamic credibility model and a hierarchical arbitration process, the problems of noise data pollution and expert resource bottlenecks caused by inconsistent quality of manual review were solved. Human-machine collaborative calibration was achieved, improving model accuracy and personnel capabilities, and building an efficient and accurate review process.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHANGHAI HAOYI INFORMATION SCI & TECH CO LTD
- Filing Date
- 2026-03-05
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies suffer from several problems: inconsistent quality of manual review leading to noise data contamination of AI models; bottlenecks in the formation of expert resources; and a lack of precision and real-time capability in traditional methods of improving personnel skills.
通过建立动态可信度模型,基于可信度分数对人工修正产生的训练数据进行加权训练,并采用分层仲裁流程和实时干预机制,实现对人机判断差异的智能分流和个性化辅导。
有效抑制噪声数据对模型的污染,缓解专家资源瓶颈,提高审核流程效率,并实现对工作人员的实时精准辅导,促进人员能力成长,构建良性循环系统。
Smart Images

Figure CN121786431B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of human-machine collaboration technology, and in particular to a human-machine collaboration calibration method, system, device and storage medium. Background Technology
[0002] In modern human-computer collaborative systems such as contact centers and content moderation platforms, a hybrid working model of "initial judgment by artificial intelligence plus manual review and confirmation" is commonly adopted. Under this model, the optimization of the AI model typically relies on treating all manually reviewed or corrected data indiscriminately as the "standard answer" and using it for model fine-tuning. However, this approach implicitly assumes that all manual corrections are completely correct. In reality, due to differences in the professional skills, proficiency, and fatigue levels of reviewers, their corrections inevitably contain errors. Treating this erroneous "noise data" the same as high-quality data systematically contaminates the model during training, causing it to learn incorrect discrimination patterns. In the long run, this can actually damage model performance, creating a vicious cycle of "the model misleading humans, and humans then misleading the model."
[0003] To ensure data quality, one approach is to submit all "disputed cases" where AI and human judgments differ to a small number of domain experts for final arbitration. However, this would quickly deplete expert resources, creating a bottleneck in high-volume scenarios. Furthermore, improving human review capabilities often relies on standardized, general training methods such as periodic exams and random checks. This approach is outdated and inefficient, failing to identify and address individual reviewers' blind spots in specific semantic understanding, leading to recurring errors and slow personnel development.
[0004] Therefore, existing technologies still have significant shortcomings in terms of how to effectively filter noisy data in manual review, how to efficiently utilize limited expert resources, and how to accurately and in real-time improve the capabilities of reviewers. Summary of the Invention
[0005] The technical problem to be solved by this application is to provide a human-machine collaborative calibration method, system, device and storage medium, which aims to solve the problems in the prior art where noise data caused by inconsistent quality of manual review contaminates artificial intelligence models, expert resources easily become a bottleneck, and traditional methods of improving personnel skills lack accuracy and real-time performance.
[0006] To address the aforementioned technical problems, this application provides a human-machine collaborative calibration method applied to a system including an artificial intelligence model and a terminal, the terminal including at least one worker terminal. The method includes: issuing a new task to the worker terminal; establishing a dynamic credibility model for the worker based on multiple preset evaluation dimensions to calculate a current credibility score based on the worker's historical work data, wherein a preset initial credibility score is assigned as the current credibility score for workers entering the system for the first time; when the worker corrects the judgment result of the artificial intelligence model, based on the worker's current credibility score, performing triage processing on the resulting discrepancies, the triage processing including: assigning the discrepancies to a preset processing path based on a comparison result between the current credibility score and at least one preset credibility threshold, the processing path including directly adopting the worker's correction result and... The discrepancies are submitted to arbitration; the training data, confirmed as correct by the aforementioned triage process and generated by manual correction, is used to train the artificial intelligence model. Based on the current credibility score of the staff member who generated the training data, a preset monotonically increasing non-linear weight mapping function maps the current credibility score to the training weights corresponding to the training data, and the training weights are used to perform weighted training on the artificial intelligence model; the error correction examples confirmed by the aforementioned triage process to be generated by the staff member are used as the staff member's personal historical error cases, and a semantic vector library of personal historical error cases is established for the staff member; and when the staff member processes the new task, the semantic similarity between the content of the new task and the personal historical error cases in the semantic vector library is calculated in real time, and when the semantic similarity exceeds a preset threshold, relevant historical error case information is pushed to the staff member's terminal.
[0007] Optionally, the multiple evaluation dimensions include at least two of the following: the worker's work accuracy within a preset time period, the standard deviation of the worker's historical work accuracy, and the reward score the worker receives for correcting a case defined as high-value by a preset rule; and / or, the push of relevant historical error case information includes: pushing non-blocking prompt information on the worker's terminal, the prompt information including the correct correction result and / or arbitration conclusion corresponding to the historical error case.
[0008] Optionally, the terminal further includes at least one expert terminal, and the preset credibility threshold includes a preset high credibility threshold and a preset low credibility threshold; the triage process includes: based on the comparison result of the worker's current credibility score with the high credibility threshold and the comparison result of the worker's current credibility score with the low credibility threshold, then: when the worker's current credibility score is higher than the high credibility threshold, the worker's correction result is directly adopted; or, when the worker's current credibility score is higher than or equal to the low credibility threshold and the worker's current credibility score is lower than or equal to the high credibility threshold, the discrepancy case is submitted for arbitration; or, when the worker's current credibility score is lower than the low credibility threshold, a hierarchical arbitration process is initiated, the hierarchical arbitration process prioritizes the discrepancy case to workers with current credibility scores higher than the high credibility threshold for cross-validation, when there is a disagreement in the cross-validation, the discrepancy case is submitted to the expert terminal for arbitration, and when the cross-validation reaches a consensus, the correction result of the discrepancy case is confirmed based on the consensus result of the cross-validation.
[0009] Optionally, the nonlinear weight mapping function is at least one of the following: the Sigmoid function, a variant of the Sigmoid function, or a piecewise linear function.
[0010] This application also provides a human-machine collaborative calibration system, including an artificial intelligence model and a terminal. The terminal includes at least one worker terminal. The system further includes: a task issuing module for issuing new tasks to the worker terminal; a credibility assessment module for establishing a dynamic credibility model for the worker based on multiple preset assessment dimensions to calculate the current credibility score based on the worker's historical work data; wherein the credibility assessment module is also used to assign a preset initial credibility score as the current credibility score to workers entering the system for the first time; and an arbitration scheduling module for triaging discrepancies when the worker corrects the judgment result of the artificial intelligence model, based on the worker's current credibility score. The triage process includes: assigning the discrepancy case to a preset processing path based on a comparison result between the current credibility score and at least one preset credibility threshold; the processing path includes directly adopting the worker's correction result and submitting the discrepancy case to arbitration. The system includes a weighted training module, which uses manually corrected training data, confirmed as correct by the arbitration and scheduling module, to train the AI model. Based on the current credibility score of the worker who generated the training data, a preset monotonically increasing non-linear weight mapping function maps the current credibility score to the corresponding training weights of the training data, and uses these training weights to perform weighted training on the AI model. A real-time intervention module is also included, which uses the erroneous correction examples confirmed by the arbitration and scheduling module as generated by the worker as the worker's personal historical error cases, and establishes a semantic vector library of these cases. When the worker processes a new task published by the task publishing module, the module calculates the semantic similarity between the content of the new task and the personal historical error cases in the semantic vector library in real time. When the semantic similarity exceeds a preset threshold, the module pushes relevant historical error case information to the worker's terminal.
[0011] Optionally, the multiple evaluation dimensions include at least two of the following: the worker's work accuracy within a preset time period, the standard deviation of the worker's historical work accuracy, and the reward score the worker receives for correcting a case defined as high-value by a preset rule; and / or, the real-time intervention module is specifically used to: push non-blocking prompts on the worker's terminal, the prompts including the correct correction results and / or arbitration conclusions corresponding to the historical error cases.
[0012] Optionally, the terminal further includes at least one expert terminal, and the preset credibility threshold includes a preset high credibility threshold and a preset low credibility threshold; the triage process includes: based on the comparison result of the worker's current credibility score with the high credibility threshold and the comparison result of the worker's current credibility score with the low credibility threshold, then: when the worker's current credibility score is higher than the high credibility threshold, the worker's correction result is directly adopted; or, when the worker's current credibility score is higher than or equal to the low credibility threshold and the worker's current credibility score is lower than or equal to the high credibility threshold, the discrepancy case is submitted for arbitration; or, when the worker's current credibility score is lower than the low credibility threshold, a hierarchical arbitration process is initiated, the hierarchical arbitration process prioritizes the discrepancy case to workers with current credibility scores higher than the high credibility threshold for cross-validation, when there is a disagreement in the cross-validation, the discrepancy case is submitted to the expert terminal for arbitration, and when the cross-validation reaches a consensus, the correction result of the discrepancy case is confirmed based on the consensus result of the cross-validation.
[0013] Optionally, the nonlinear weight mapping function is at least one of the following: the Sigmoid function, a variant of the Sigmoid function, or a piecewise linear function.
[0014] This application also provides an electronic device, including a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, it implements the human-machine collaborative calibration method described in any of the above claims.
[0015] This application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the human-machine collaborative calibration method described in any of the above claims.
[0016] This application has the following beneficial effects:
[0017] 1. By establishing a dynamic credibility model and weighting the training data generated by manual correction based on the credibility score, this application can effectively suppress the pollution of artificial intelligence models by low-quality noise data, ensure that the model optimizes in the right direction, and thus improve the accuracy and robustness of the model.
[0018] 2. By triaging the differences between human and machine judgments based on credibility scores, and especially by adopting a tiered arbitration process, most disputed cases are handed over to highly credible staff for cross-verification, while only a small number of difficult cases are submitted to experts. This effectively alleviates the bottleneck of expert resources and improves the efficiency of the overall review process.
[0019] 3. By establishing a semantic vector library of personal historical error cases for staff members, and performing real-time similarity matching and early warning prompts when handling new tasks, it is possible to accurately and in real time intervene in repetitive errors that staff members are about to make, transforming delayed post-event training into efficient in-event coaching and accelerating the growth of staff capabilities;
[0020] 4. This application quantifies human credibility and deeply couples it with machine training, while using machine intelligence to feed back into human growth, thus constructing a virtuous cycle system in which humans and machines mutually promote each other and improve quality and efficiency together. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a schematic diagram of the structure of a human-machine collaborative calibration system according to an embodiment of this application.
[0023] Figure 2 This is a flowchart illustrating a human-machine collaborative calibration method according to an embodiment of this application.
[0024] Figure 3 This is a timing diagram of real-time semantic intervention according to an embodiment of this application.
[0025] Figure 4 This is a schematic diagram illustrating the principle of nonlinear weight mapping according to an embodiment of this application.
[0026] In the diagram: 10-Human-Machine Collaborative Calibration System, 20-Task Release Module, 30-Confidence Assessment Module, 40-Arbitration Scheduling Module, 50-Weighted Training Module, 60-Artificial Intelligence Model, 70-Real-Time Intervention Module, 71-Semantic Vector Library, 80-Staff Terminal. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0028] Unless otherwise defined, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in the embodiments of this application is for the purpose of describing the embodiments of this application only and is not intended to limit this application.
[0029] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0030] (1) Dynamic credibility model: refers to a mathematical model that calculates and dynamically updates the credibility score of an employee based on multiple preset evaluation dimensions (such as the accuracy of the employee's work within a preset time period, the standard deviation of the historical work accuracy rate, and the reward score obtained by correcting high-value cases).
[0031] (2) Triage: This refers to a decision-making mechanism that automatically assigns the case of discrepancy to different processing paths (e.g., direct adoption, arbitration) when the staff corrects the judgment result of the artificial intelligence model and a difference between human and machine judgment occurs.
[0032] (3) Tiered arbitration process: This refers to a tiered arbitration mechanism for human-machine discrepancies based on personnel credibility scores. When a discrepancy case arises from a low-credibility (credibility below the threshold) staff member, it is first cross-validated by a high-credibility (credibility above the threshold) staff member. If the cross-validation reaches a consensus, the result is adopted. If the cross-validation results in a disagreement, the case is submitted to experts for final adjudication.
[0033] (4) Nonlinear weight mapping function: refers to a monotonically increasing function (such as the Sigmoid function or piecewise linear function) used to map the dynamic confidence score of the staff to the training weight of the corresponding training data, so that high confidence score corresponds to high weight and low confidence score corresponds to low weight.
[0034] (5) Semantic Vector Library: This refers to a database used to store error case information. The text content of each case is converted into a high-dimensional mathematical vector (i.e., a semantic vector) by a pre-trained language model (such as Sentence-BERT) to facilitate efficient semantic similarity calculation. Sentence-BERT is a pre-trained model based on the BERT (Bidirectional Encoder Representations from Transformers) architecture, specifically designed to map sentences or text paragraphs into high-dimensional semantic vectors for tasks such as semantic similarity calculation.
[0035] This application provides a human-machine collaborative calibration method, aiming to solve the problems in the prior art where "noise data" caused by inconsistent quality of manual review contaminates artificial intelligence models, and the lack of accuracy and real-time performance in traditional methods of improving human capabilities. Figure 2 As shown, this method is applied to a system that includes an artificial intelligence model and at least one staff terminal, constructing a complete closed loop of dynamic assessment, intelligent triage, weighted training, and real-time intervention.
[0036] This application provides a human-machine collaborative calibration method, which is applied to a system including an artificial intelligence model and a terminal, the terminal including at least one worker terminal, and the method includes steps S100 to S500.
[0037] Step S100: Issue a new task to the worker's terminal. This is the trigger point for this method, initiating the subsequent human-machine collaborative processing flow.
[0038] Step S200: Based on multiple preset evaluation dimensions, establish a dynamic credibility model for the staff member to calculate the current credibility score according to the staff member's historical work data. Specifically, assign a preset initial credibility score as the current credibility score to staff members entering the system for the first time. By establishing a dynamic credibility model, the abstract concept of staff member ability is transformed into a quantifiable dynamic indicator, thereby enabling real-time monitoring of the current judgment reliability of each staff member and providing a unified data foundation for subsequent steps. The initial credibility score setting provides a reasonable starting point for new staff members, ensuring the system's cold start capability.
[0039] Step S300: When the staff member corrects the judgment result of the artificial intelligence model and generates a human-machine judgment discrepancy, the resulting discrepancy cases are triaged based on the staff member's current credibility score. This triage process includes: assigning the discrepancy cases to a preset processing path based on a comparison between the current credibility score and at least one preset credibility threshold. This processing path includes directly adopting the staff member's correction result and submitting the discrepancy cases to arbitration. Based on the current credibility score, correction results from staff members with high current credibility scores are considered high-quality data and directly adopted, avoiding unnecessary manual review; correction results from staff members with low current credibility scores enter the arbitration process for further verification. By triaging human-machine judgment discrepancies based on credibility scores, especially with the subsequent layered arbitration process, most disputed cases are cross-validated by staff members with high credibility scores, while only a small number of difficult cases are submitted to experts. This effectively alleviates the bottleneck of expert resources, improves the efficiency of the overall review process, and achieves intelligent hierarchical processing of "human-machine discrepancy cases."
[0040] Step S400: The training data, confirmed as correct by the aforementioned diversion process and generated through manual correction, is used to train the artificial intelligence model. Based on the current credibility score of the worker who generated the training data, a preset monotonically increasing non-linear weight mapping function maps the current credibility score to the corresponding training weights for the training data. These training weights are then used to perform weighted training on the artificial intelligence model. Through the non-linear weight mapping function, data generated by workers with high credibility scores receives high weights and plays a dominant role in the training of the artificial intelligence model; data generated by workers with low credibility scores has its weights significantly suppressed. Even if a small amount of erroneous data enters the training set, the gradients generated by this erroneous data during backpropagation are minimal and cannot substantially affect the parameters of the artificial intelligence model. This mechanism achieves adaptive cleaning of training data at the mathematical level, fundamentally solving the problem of systematically polluting the model with noisy data, breaking the vicious cycle of "the model misleading people, and people misleading the model," and achieving refined control over the training data.
[0041] Step S500: Error correction examples confirmed as originating from the employee through the aforementioned triage process are designated as the employee's personal historical error cases, and a semantic vector library of these personal historical error cases is established for that employee. This is equivalent to building a personalized "error log" for each employee and converting it into semantic vectors that can be retrieved by the computer. The personal historical error case library accurately reflects each employee's skill gaps, while semantic vectorization enables the computer to understand the semantic meaning of the cases, laying a data foundation for subsequent real-time semantic matching.
[0042] Step S500 further includes calculating the semantic similarity between the content of the new task and the personal historical error cases in the semantic vector library in real time when the staff member is handling the new task, and pushing relevant historical error case information to the staff member's terminal when the semantic similarity exceeds a preset threshold. At the "critical moment" when the staff member is handling a new task, semantic similarity calculation identifies whether the current task belongs to the same scenario as a case in which the staff member has previously made a mistake. Once the preset threshold is exceeded, historical error case information is proactively pushed. This mechanism achieves a leap from lagging, extensive unified training to real-time, personalized, and precise assistance, effectively preventing the recurrence of similar errors, realizing real-time and precise assistance to staff members, and accelerating the growth of staff capabilities.
[0043] In summary, this application deeply couples human credibility metrics with the training of artificial intelligence models, while using a "mistake collection" to support human growth, thus constructing a virtuous cycle system where humans and machines mutually promote each other, and quality and efficiency are improved together. Through a complete technical chain of "credible metrics → intelligent triage → weighted training → mistake collection → real-time assistance," it systematically solves the three core problems of existing technologies: noisy data contaminating the model, bottlenecks in the formation of expert resources, and lack of accuracy and real-time capability improvement for personnel.
[0044] Specifically, in step S200, the multiple evaluation dimensions of the dynamic credibility model can be further refined into a combination of at least two of the following indicators:
[0045] The accuracy of the worker's work within a preset time period; for example, the F1 score or accuracy calculated based on a sliding window, used to reflect the worker's current business proficiency and immediate performance, ensuring that the current credibility score can keep up with changes in the worker's ability; among them, the F1 score is a commonly used indicator in the field of machine learning and data mining to evaluate the performance of classification models, and it is the harmonic mean of precision and recall.
[0046] The standard deviation of the staff member's historical work accuracy rate; the standard deviation is used to quantify the degree of fluctuation in the staff member's judgment accuracy over a longer period of time. The larger the standard deviation, the more points are deducted, thereby punishing unstable performance that fluctuates "high and low" and encouraging consistently high-quality output; and,
[0047] The staff member receives a reward score for correcting a case defined as high-value by preset rules; when the staff member's correction is confirmed to be correct, and the case corrects a high-confidence error of the artificial intelligence model, or is a marginal case, additional points are awarded; this mechanism breaks the deadlock of "low scorers never have a chance to turn things around", incentivizes all staff to actively contribute high-value knowledge, and promotes the continuous optimization of the model.
[0048] By integrating at least two of the above evaluation dimensions, the dynamic credibility model can more comprehensively and accurately characterize the judgment reliability of each staff member, providing a more reliable basis for subsequent triage and weighted training. In one specific implementation, the terminal further includes at least one expert terminal, and the preset credibility threshold includes a preset high credibility threshold and a preset low credibility threshold; the triage process includes: based on the comparison result of the worker's current credibility score with the high credibility threshold and the comparison result of the worker's current credibility score with the low credibility threshold, then: when the worker's current credibility score is higher than the high credibility threshold, the worker's correction result is directly adopted; or, when the worker's current credibility score is higher than or equal to the low credibility threshold and lower than or equal to the high credibility threshold, the discrepancy case is submitted for arbitration; or, when the worker's current credibility score is lower than the low credibility threshold, a hierarchical arbitration process is initiated, the hierarchical arbitration process prioritizes the discrepancy case to workers with current credibility scores higher than the high credibility threshold for cross-validation, when there is a disagreement in the cross-validation, the discrepancy case is submitted to the expert terminal for arbitration, and when the cross-validation reaches a consensus, the correction result of the discrepancy case is confirmed based on the consensus result of the cross-validation. This is a further refinement of the traffic splitting mechanism based on credibility scores, introducing preset high credibility thresholds and preset low credibility thresholds to construct a three-level traffic splitting structure:
[0049] 1) In the high credibility zone (where the current credibility score is higher than the high credibility threshold), the staff member is considered highly reliable, and their correction results are directly adopted without any review. This path maximizes the productivity of highly capable personnel while ensuring data quality.
[0050] 2) Medium credibility zone (current credibility score is between the low credibility threshold and the high credibility threshold), the staff member's judgment has some reference value but is not reliable enough, and the discrepancies generated should be submitted directly to arbitration; the "arbitration" here can be the regular arbitration process, such as submitting directly to experts, or other arbitration methods according to the system configuration; this layer provides necessary quality control for medium-capability personnel;
[0051] 3) Low Credibility Zone (Current Credibility Score Below Low Credibility Threshold): The reliability of this staff member's judgment is low, and discrepancies arising from their judgment will trigger a tiered arbitration process. This process first anonymizes the cases and distributes them to at least one staff member with a current credibility score above the high credibility threshold for cross-validation.
[0052] If high-scoring staff reach an agreement (or with the original corrector), a peer consensus is formed, the result is adopted, and the original staff are given positive incentives (such as bonus points) to encourage them to learn and improve.
[0053] If there are disagreements among high-scoring staff, it indicates that the case is indeed difficult and complex. In this case, the case is submitted to the expert terminal for final arbitration, and the domain expert makes the final ruling, which is regarded as the absolute truth value.
[0054] This tiered arbitration mechanism allows most disputes to be resolved among high-scoring peers, with only a very few truly complex cases requiring expert intervention. This not only effectively alleviates the bottleneck of expert resources and improves the efficiency of the overall review process, but also promotes the learning and growth of lower-scoring personnel through peer verification. Simultaneously, this tiered approach ensures that only fully validated, high-quality data is ultimately included in the training set, guaranteeing the training quality of the AI model from the outset.
[0055] In a preferred implementation, the nonlinear weight mapping function in step S400 is at least one of the following: the Sigmoid function, a variant of the Sigmoid function (e.g., a scaled and translated Sigmoid function), and a piecewise linear function. These functions all possess monotonically increasing and nonlinear characteristics, enabling a smooth mapping of staff credibility scores to the weights of training samples.
[0056] Take the Sigmoid variant function as an example:
[0057] λ(S)=(1−ϵ) / (1+e^(−k(S−S0)))+ϵ.
[0058] Where S is the worker's dynamic confidence score; S0 is the center point (e.g., 70), which can be set as the average confidence score, representing the point where the weight increases the fastest; k is the steepness factor of the curve; the larger k is, the steeper the curve is near S0, meaning the system has a higher degree of discrimination over the scores; ϵ is a very small positive value (e.g., 0.05) used to prevent the weight from being 0, ensuring that even the lowest confidence data retains a very small learning opportunity. This Sigmoid variant function can smoothly map the current confidence score to the interval (ϵ, 1), ensuring that:
[0059] 1) Data generated by staff with high credibility scores (e.g., greater than 85) has a weight close to 1, and high-quality data is fully learned;
[0060] 2) Data generated by staff with a current credibility score of medium (greater than or equal to 65 and less than or equal to 85) has a smooth transition in weight, and the data has a moderate impact on the artificial intelligence model.
[0061] 3) The weights of data generated by staff with low credibility scores (e.g., less than 65) tend to be close to the minimum value ϵ. The gradient generated by low-quality data is significantly suppressed and has almost no impact on the parameters of the artificial intelligence model, thus effectively suppressing noisy data mathematically.
[0062] Let's take a piecewise linear function as an example. For instance, we can define the following rule: when the current confidence score S is lower than the low confidence threshold T_low (e.g., 65 points), the weight λ is fixed at a very small value, such as 0.05; when S is higher than the high confidence threshold T_high (e.g., 85 points), the weight λ is fixed at a very large value, such as 0.95; when S is between T_low and T_high, the weight λ increases linearly from 0.05 to 0.95. This method is simple to calculate, logically clear, and can achieve the same non-linear mapping effect of high scores with high weights and low scores with low weights, proving that the core idea of this application does not depend on a specific function form.
[0063] By employing this type of nonlinear weight mapping function, adaptive cleaning of training data can be achieved at the mathematical level, effectively suppressing the contamination of the model by low-quality noise data, while avoiding the information loss that may result from completely discarding data. This further improves the accuracy and robustness of model training, and together with the dynamic credibility model and data splitting, constitutes a complete data quality control system.
[0064] In one specific implementation, "pushing relevant historical error case information" in step S500 specifically includes pushing non-blocking prompts to the worker's terminal. These prompts not only inform the worker that their current task is similar to a historical error, but also provide the correct correction result and / or arbitration conclusion corresponding to the historical error case. This non-blocking prompt does not interrupt the worker's normal workflow, but provides crucial reference information to help the worker make correct judgments at critical moments. By providing correct correction results and expert arbitration conclusions, workers can immediately understand where their errors lie and the correct handling methods, thereby deepening their understanding of business rules and accelerating their skill development.
[0065] In a basic embodiment, the method first executes step S100. For example, in a call center quality inspection scenario, the system retrieves a call recording to be inspected and its initial tag (such as "business consultation") generated by an artificial intelligence model from the task queue, and assigns it to an available staff member. This step ensures that the entire calibration process has a clear starting point and processing target.
[0066] Upon receiving an assigned task, the method proceeds to step S200, establishing a dynamic credibility model for the worker performing the task. This dynamic credibility model does not statically evaluate the worker but continuously calculates a dynamically changing current credibility score based on the worker's historical work data. For new workers entering the system for the first time, a preset initial credibility score is assigned, such as 60 points (out of 100). This initial credibility score represents the system's average expectation of a worker with unknown capabilities. As the worker handles more tasks, their current credibility score is dynamically adjusted based on their subsequent performance.
[0067] When staff members are processing tasks and believe that the AI model's judgment is incorrect, a "human-machine judgment discrepancy" occurs. For example, AI model 60 might label a case as "normal," while a staff member might consider it "high-risk," thus creating a discrepancy. The system will immediately capture this discrepancy and trigger subsequent triage processes. This is the core trigger point of the entire calibration mechanism, indicating that the system needs to adjudicate an inconsistent judgment.
[0068] At this point, the method enters the crucial triage process, step S300. The system obtains the current dynamic credibility score of the staff member who generated the discrepancy and determines the fate of the discrepancy case based on a comparison of this score with at least one preset credibility threshold. For example, the system can set a preset credibility threshold of 75 points. If the staff member's score is higher than 75 points, the system tends to trust their judgment; if it is lower than 75 points, the system holds reservations about their judgment. Accordingly, the system assigns the discrepancy case to different processing paths, which include at least directly adopting the staff member's corrective results and submitting the discrepancy case to arbitration. This credibility-based automated triage avoids the efficiency bottleneck caused by submitting all discrepancy cases to expert arbitration.
[0069] After the data stream is processed, the final correctness of the case is confirmed, and the process proceeds to step S400. For the data confirmed as correct by the stream process and generated through manual correction, the system does not treat it the same as all other data for training the AI model. Instead, it performs a weighted training step based on the current credibility score. Specifically, the system uses a preset monotonically increasing non-linear weight mapping function to map the current credibility score of the worker who generated the correct correction data to a corresponding training weight. The higher the current credibility score, the greater the weight. Then, the system uses this data with different training weights to perform weighted training on the AI model 60. This mechanism ensures that "high-quality signals" from high-credibility personnel (current credibility score above a preset credibility threshold) are amplified during model optimization, while "potential noise" from low-credibility personnel (current credibility score below a preset credibility threshold) is suppressed.
[0070] Simultaneously, for amendments identified as erroneous during the triage process, the system proceeds to step S500 to effectively utilize these erroneous amendments. These erroneous amendments are considered valuable learning materials for individual staff members, and the system records them as personal historical error cases, establishing a semantic vector library 71 for each staff member. This semantic vector library not only records the specific content of the errors but, more importantly, uses technical means to transform the specific content of these errors into computable semantic features.
[0071] Finally, in step S500, to achieve real-time guidance and skills enhancement for staff, when the staff member handles a new task, the system further performs a matching operation in the background in real time. The system calculates the semantic similarity between the semantic features of the new task's content and all cases in the semantic vector library 71 of the staff member's personal historical error cases. When it is found that the semantic similarity between a historical error case and the current task exceeds a preset threshold, it means that the staff member may be about to repeat the same mistake. At this time, the system will proactively push relevant historical error case information to the staff member's terminal to remind them, thereby achieving precise and real-time intervention and guidance.
[0072] To implement the above method, this application also provides a human-machine collaborative calibration system 10. For example... Figure 1 As shown, the system can have at least one server running specific software modules. The system includes an artificial intelligence model 60 and a terminal, which includes at least one worker terminal 80, as well as a series of functional modules.
[0073] Specifically, the system includes a task publishing module 20, which implements step S100. Its function is to manage the task queue and publish new tasks to the staff terminal 80 according to preset rules (such as staff busy / idle status, skill matching degree, etc.). For example, it can push a text content to be reviewed to the content reviewer's staff terminal 80.
[0074] The system also includes a credibility assessment module 30, used to implement step S200. This credibility assessment module 30 is responsible for establishing and maintaining the aforementioned dynamic credibility model for each staff member in the system based on multiple preset assessment dimensions. It continuously collects historical work data of staff members (such as the accuracy rate of corrections, the number of cases processed, etc.) and calculates each person's current credibility score based on this historical work data. The credibility assessment module 30 also assigns an initial credibility score as the current credibility score to staff members entering the system for the first time.
[0075] When a staff member corrects the judgment result of the artificial intelligence model 60, a discrepancy arises between human and machine judgments, activating the arbitration scheduling module 40 for step S300. This arbitration scheduling module 40 obtains the current credibility score of the relevant staff member from the credibility assessment module 30 and performs triage processing on the currently generated discrepancy cases. The triage processing assigns the case to a preset processing path based on a comparison between the current credibility score and at least one preset credibility threshold. The processing path includes directly adopting the staff member's corrected result and submitting the case to arbitration. It acts as the system's "traffic police," intelligently allocating disputed cases to balance efficiency and quality.
[0076] The weighted training module 50 is responsible for optimizing the artificial intelligence model 60 and is used to execute step S400. The weighted training module 50 uses training data generated by manual correction and confirmed as correct by the arbitration scheduling module 40 to train the artificial intelligence model. Specifically, based on the current credibility score of the staff member who generated this training data obtained from the credibility assessment module 30, the current credibility score is mapped to the training weight corresponding to the training data through a preset monotonically increasing non-linear weight mapping function. Then, using these training weights, the weighted training data is applied to perform weighted training on the artificial intelligence model 60.
[0077] Finally, the system also includes a real-time intervention module 70. This real-time intervention module 70 works closely with a semantic vector library 71 to execute step S500. When the arbitration scheduling module 40 confirms that the amendment example generated by the staff member is incorrect, the real-time intervention module 70 will classify the amendment example as a personal historical error case of the staff member, and call the language model to vectorize the content of the amendment example, storing it in the semantic vector library 71 of personal historical error cases established for the staff member. More importantly, when the task publishing module 20 publishes a new task to a staff member, the real-time intervention module 70 will calculate the similarity between the content of the new task and the personal historical error cases in the semantic vector library 71 of the staff member's personal historical error cases in real time. If the similarity is too high (exceeding a preset threshold), it will send a warning or prompt message to the staff member's terminal 80, such as pushing relevant historical error case information. This module plays the role of a "personal coach," helping staff members avoid repeating mistakes.
[0078] In a preferred embodiment, the "multiple evaluation dimensions" on which the above-described dynamic credibility model is based may include at least two of the following: the worker's work accuracy within a preset time period, the standard deviation of the worker's historical work accuracy rate, and the reward score the worker receives for correcting a case defined as high-value by a preset rule.
[0079] Specifically, "work accuracy" can be measured using the F1 score, which combines precision and recall and reflects staff performance on imbalanced datasets better than a single precision metric. For example, the F1 score of a staff member for all confirmed cases within the past week can be calculated. The "standard deviation of historical work accuracy" measures the stability of staff performance; a staff member with a fluctuating standard deviation, even with a decent average accuracy, should have their current confidence score reduced. The "high-value case bonus score" is an incentive mechanism to encourage staff to identify and correct errors made by the AI model with high confidence (60%), or to handle rare, poorly defined edge cases. These contributions are crucial for breakthroughs in the AI model's capabilities. An exemplary formula for calculating the current confidence score Si(t) is: Si(t) = α⋅F1_window−β⋅Var_history+γ⋅Bonus_contribution. Here, α, β, and γ are the weighting coefficients for each dimension and can be adjusted according to business needs.
[0080] Furthermore, the "submit for arbitration" path in the triage process can be designed more finely to maximize the conservation of expert resources. Preferably, the terminal also includes at least one expert terminal, and the preset credibility threshold includes a preset high credibility threshold and a preset low credibility threshold. Figure 2 As shown, when a discrepancy case is assigned to the arbitration processing path (i.e., the current credibility score S is lower than or equal to the preset high credibility threshold T_high), a tiered arbitration process is initiated. When the staff member's current credibility score S is lower than the preset low credibility threshold T_low (e.g., 65 points), the system determines that the correction result has a high risk and prioritizes anonymously submitting the discrepancy case to at least one other staff member with a current credibility score higher than the preset high credibility threshold T_high (e.g., 85 points) for cross-validation. These staff members with higher current credibility scores are equivalent to "senior auditors" in the system.
[0081] After receiving the cross-validation results, the system makes a judgment. If the cross-validation results are consistent (for example, two or more senior auditors reach the same conclusion, whether supporting or opposing the initial amendment), the system considers the conclusion sufficiently reliable and determines the final handling method for the discrepancy case based on this consistent result, thus concluding the entire process. This "peer review" mechanism can resolve the vast majority of low-credibility disputes.
[0082] The system only considers a case truly difficult to resolve when cross-validation reveals discrepancies (e.g., at least two senior auditors reach conflicting conclusions). In this case, the case is submitted to a pre-designated expert terminal for final arbitration. The expert's judgment, as the ultimate authority, will be the final conclusion of the case.
[0083] Corresponding to the handling method for low-credibility personnel (whose current credibility score S is lower than the preset low credibility threshold T_low), when the current credibility score of the personnel is already very high, that is, higher than the aforementioned preset high credibility threshold T_high, the difference cases generated by them can be directly adopted by the system for correction, which greatly improves the processing efficiency.
[0084] In one implementation, when the employee's current credibility score is higher than or equal to the low credibility threshold T_low and the current credibility score is lower than or equal to the high credibility threshold T_high, the system can directly submit the resulting discrepancy case to arbitration, for example, directly enter the expert arbitration process.
[0085] In summary, by setting high and low credibility thresholds, staff can be categorized into high, medium, and low credibility levels, enabling the system to implement a complete and efficient triage and arbitration mechanism. This mechanism directly adopts the judgments of high-credibility personnel to expedite the process, uses tiered arbitration to handle disputes from low-credibility personnel to filter noise and conserve expert resources, and provides a clear escalation path for disputes from medium-credibility personnel. This mechanism is key to overcoming the bottleneck of expert resources and improving overall review efficiency.
[0086] In the weighted training step, the "monotonically increasing nonlinear weight mapping function" can have various specific implementations. For example... Figure 4 As shown, the core idea of this function is to map the current confidence score S of the input to the training weight λ of the output, and the higher S is, the larger λ is, and this growth relationship is not linear. The non-linear design is mainly to implement different weight strategies in different score intervals, such as suppression in low score segments and saturation in high score segments.
[0087] like Figure 3 As shown, when the system detects that the semantic similarity between a new task being processed in real time and a case in the semantic vector library 71 of a person's historical error cases exceeds a preset threshold, the human-machine collaborative calibration system 10 will push a prompt message to the staff terminal 80.
[0088] Preferably, the prompt is a non-blocking prompt. This means it does not interrupt the worker's current workflow. For example, it could be a small pop-up bubble in the corner of the screen that can be closed at any time, or a wavy underline appearing below the relevant text, with detailed information only displayed when the mouse hovers over it. This design avoids interfering with the worker and ensures their work efficiency.
[0089] The content of the prompt is also crucial. It goes beyond simply warning of a potential error; it provides valuable reference information. Specifically, the prompt can include the correct corrective action and / or arbitration conclusion corresponding to the historical error case. For example, the prompt could display: "Similar Historical Errors: You previously misclassified the similar content 'Customer's emotionally charged consultation progress' as a 'complaint.' After expert arbitration, the correct label should be 'High-Priority Consultation.'" Such information allows staff to immediately recall the scenario and the correct handling method, enabling them to make the right judgment in new tasks and achieve accurate and efficient "in-process coaching."
[0090] The working process of this application is illustrated below through a specific embodiment. Assume the system's preset low credibility threshold T_low = 65, high credibility threshold T_high = 85, and semantic similarity threshold = 0.87. Staff member A is a junior auditor with a current credibility score of 52. Furthermore, their personal semantic vector library 71 contains an erroneous record: a recording of a customer speaking urgently but whose core demand was for a solution was misjudged as a "complaint threat."
[0091] A week later, staff member A received a new task on staff terminal 80, consisting of the text: "Your efficiency is too slow! I need a solution today!" The real-time intervention module 70 immediately activated in the background. It used a pre-trained language model to convert the text into a semantic vector and calculated its cosine similarity with A's personal semantic vector library 71. The calculation revealed that the semantic vector similarity between this new task and A's past incorrect question, "urgent consultation misjudged as a threat," was as high as 0.89, exceeding the preset semantic similarity threshold of 0.87.
[0092] The system immediately pushed a non-blocking prompt on A's interface: "Historical Similar Mistake Alert: The current content is highly similar to a case you previously misjudged. Please carefully assess whether it poses a threat and refer to the correct label 'High Priority Consultation'."
[0093] However, staff member A, perhaps due to fatigue or stubbornness, ignored the prompt and insisted on changing the case's label to "complaint threat," resulting in a discrepancy between human and machine judgment. At this point, the arbitration scheduling module 40 intervenes. It obtains A's credibility score of 52, below the low credibility threshold of 65. According to the conventional process, tiered arbitration should be initiated, first submitting the case to a high-credibility auditor for cross-verification. However, in this embodiment, the system can be designed more intelligently: the arbitration scheduling module 40 simultaneously learns that the real-time intervention module 70 has just issued a high-similarity warning, which was ignored. This information indicates that this error correction is a high-risk behavior of "knowingly committing a wrong." Therefore, the system dynamically increases the arbitration priority of this case and decides to skip the peer cross-verification step, directly pushing the case to experts for final arbitration to accelerate the correction of this stubborn error.
[0094] Upon receiving the case, the expert quickly determined that A's correction was incorrect, and the correct label should be "high-priority consultation." This arbitration result was recorded by the system. First, for employee A, their current credibility score will be further reduced due to this error confirmed by the expert, for example, dropping to 48 points. Second, this new error case, along with the expert's authoritative judgment, will be vectorized again and updated in A's personal semantic vector library 71, strengthening the recording of this type of error.
[0095] Meanwhile, for AI model 60, this correct data, confirmed by expert authority (i.e., "Your efficiency is too slow!..." corresponding to "high-priority consultation"), is sent to weighted training module 50. Because this data comes from expert arbitration, the system can assign it an extremely high training weight, for example, a weight close to 1.0 obtained through Sigmoid function mapping, or even the highest weight defined by a hyperparameter (such as 2.0). Weighted training module 50 uses this high-weight data to train AI model 60, thereby significantly enhancing the model's ability to distinguish between "urgent consultation" and "real threat."
[0096] Through this complete process, this application demonstrates a highly intelligent and collaborative closed-loop system. It not only successfully intercepted a low-quality manual correction, preventing contamination of the AI model, but also achieved dynamic optimization of the arbitration process through linkage with the real-time intervention module. More importantly, it transforms an error event into a dual improvement for both humans and machines: the credibility of the human (staff A) is more accurately characterized, and their individual error profile is clearer, laying the foundation for future precise guidance; the machine (AI model 60) learns from a high-quality boundary case, making it more intelligent. This is precisely the virtuous cycle of human-machine collaboration and co-evolution that this application pursues.
[0097] This application also provides an electronic device, which may be a server, a personal computer (PC), a tablet computer, or a smartphone. The computer device internally includes a processor and a memory. The memory stores a computer program, which, when executed by the processor, can implement any of the human-machine collaborative calibration methods described above. For example, the processor can execute program instructions to achieve functions such as establishing a dynamic reliability model, processing differences in human-machine judgments, training a weighted model, and performing real-time semantic intervention.
[0098] This application also provides a computer-readable storage medium storing a computer program thereon. The storage medium can be non-volatile, such as read-only memory (ROM), programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), or flash memory, or it can be volatile, such as random access memory (RAM). When the computer program is executed by a processor, it can implement any of the human-machine collaborative calibration methods described above. This storage medium enables the technical solution of this application to be released and deployed in the form of a software product.
[0099] This application has a wide range of applications, benefiting any field involving the "AI initial screening + human review" model. For example, in the intelligent quality inspection scenario of a large call center, millions of call recordings are generated daily. Using this application, the system can automatically assess the credibility of quality inspectors, quickly teaching the AI model new issues discovered by high-credibility inspectors (such as new sales misconduct scripts) with high weight, while simultaneously conducting tiered review of dispute judgments by low-credibility inspectors to prevent misjudgments from contaminating the model. When a quality inspector repeatedly makes mistakes on the boundary between "customer retention" and "customer complaints," the system can also provide real-time reminders when they encounter similar scenarios again, thereby greatly improving quality inspection efficiency and model accuracy, and accelerating the growth of quality inspectors.
[0100] Another typical application scenario is content security moderation on social media platforms. For borderline cases involving complex semantics such as culture, slang, and irony that require manual review, this application can establish a dynamic credibility model for each moderator based on their accuracy in judging different types of inappropriate content. For a moderator with a high accuracy rate in judging "slang," the cases they correct will receive high weight. Conversely, if they attempt to correct a case of "medical misinformation," which they are not familiar with, the system will initiate an arbitration process. This refined management can effectively address the challenges of the highly specialized and broad scope of content moderation, overcome the bottleneck of expert resources, and purify the online environment more quickly and accurately.
Claims
1. A human-machine collaborative calibration method, applied to a system including an artificial intelligence model and a terminal, the terminal including at least one worker terminal, characterized in that, The method includes: Issue a new task to the employee's terminal; Based on multiple preset evaluation dimensions, a dynamic credibility model is established for staff to calculate the current credibility score based on the staff's historical work data. Specifically, a preset initial credibility score is assigned as the current credibility score for staff who enter the system for the first time. When the staff member corrects the judgment result of the artificial intelligence model, the resulting discrepancy cases are triaged based on the staff member's current credibility score. The triage process includes: assigning the discrepancy cases to a preset processing path based on the comparison result of the current credibility score and at least one preset credibility threshold. The processing path includes directly adopting the staff member's correction result and submitting the discrepancy cases to arbitration. The training data that has been confirmed as correct by the aforementioned diversion process and is generated by manual correction is used to train the artificial intelligence model. Specifically, based on the current credibility score of the staff member who generated the training data, the current credibility score is mapped to the training weights corresponding to the training data through a preset monotonically increasing nonlinear weight mapping function, and the artificial intelligence model is trained using the training weights. Error amendment examples confirmed to be generated by the employee after the aforementioned traffic splitting process are designated as the employee's personal historical error cases, and a semantic vector library of personal historical error cases is established for that employee; and When the staff member processes the new task, the semantic similarity between the content of the new task and the personal historical error cases in the semantic vector library is calculated in real time. When the semantic similarity exceeds a preset threshold, the relevant historical error case information is pushed to the staff member's terminal.
2. The human-machine collaborative calibration method according to claim 1, characterized in that, The multiple evaluation dimensions include at least two of the following: the staff member's work accuracy within a preset time period, the standard deviation of the staff member's historical work accuracy rate, and the reward score the staff member receives for correcting a case defined as high-value by preset rules; the push of relevant historical error case information includes: pushing non-blocking prompt information on the staff member's terminal, the prompt information including the correct correction result and / or arbitration conclusion corresponding to the historical error case.
3. The human-machine collaborative calibration method according to claim 1, characterized in that, The terminal also includes at least one expert terminal, and the preset confidence threshold includes a preset high confidence threshold and a preset low confidence threshold; the traffic splitting process includes: Based on the comparison results of the staff member's current credibility score with the high credibility threshold and the comparison results of the staff member's current credibility score with the low credibility threshold, then: When the staff member's current credibility score is higher than the high credibility threshold, the staff member's correction result is directly adopted; or, When the employee's current credibility score is higher than or equal to the low credibility threshold and the employee's current credibility score is lower than or equal to the high credibility threshold, the discrepancy case will be submitted to arbitration; or, When the staff member's current credibility score is lower than the low credibility threshold, a tiered arbitration process is initiated. The tiered arbitration process prioritizes cross-validating the discrepancy case with staff members whose current credibility score is higher than the high credibility threshold. When there is a disagreement in the cross-validation, the discrepancy case is submitted to the expert terminal for arbitration. When the cross-validation reaches a consensus, the correction result of the discrepancy case is confirmed based on the consensus result of the cross-validation.
4. The human-machine collaborative calibration method according to any one of claims 1 to 3, characterized in that, The nonlinear weight mapping function is at least one of the following: the Sigmoid function, a variant of the Sigmoid function, or a piecewise linear function.
5. A human-machine collaborative calibration system, comprising an artificial intelligence model and a terminal, the terminal including at least one worker terminal, characterized in that, The system also includes: The task publishing module is used to publish new tasks to the worker's terminal; The credibility assessment module is used to establish a dynamic credibility model for staff based on multiple preset assessment dimensions, and to calculate the current credibility score based on the staff's historical work data. The credibility assessment module is also used to assign a preset initial credibility score as the current credibility score to staff who enter the system for the first time. The arbitration scheduling module is used to triage the resulting discrepancy cases based on the staff member's current credibility score when the staff member corrects the judgment result of the artificial intelligence model. The triage process includes: assigning the discrepancy cases to a preset processing path according to the comparison result of the current credibility score and at least one preset credibility threshold. The processing path includes directly adopting the staff member's correction result and submitting the discrepancy cases to arbitration. A weighted training module is used to train the artificial intelligence model with training data that has been corrected by manual correction and has been processed by the arbitration scheduling module. Specifically, based on the current credibility score of the worker who generated the training data, a preset monotonically increasing non-linear weight mapping function is used to map the current credibility score to training weights corresponding to the training data, and these training weights are then used to perform weighted training on the artificial intelligence model. The real-time intervention module is used to identify erroneous amendments confirmed by the arbitration scheduling module as being generated by the staff member as the staff member's personal historical error cases, and to establish a semantic vector library of personal historical error cases for the staff member; when the staff member processes a new task issued by the task issuing module, the module calculates the semantic similarity between the content of the new task and the personal historical error cases in the semantic vector library in real time; and when the semantic similarity exceeds a preset threshold, pushes relevant historical error case information to the staff member's terminal.
6. The human-machine collaborative calibration system according to claim 5, characterized in that, The multiple evaluation dimensions include at least two of the following: the staff member's work accuracy within a preset time period, the standard deviation of the staff member's historical work accuracy rate, and the reward score the staff member receives for correcting a case defined as high-value by preset rules; the real-time intervention module is specifically used to: push non-blocking prompt information on the staff member's terminal, the prompt information including the correct correction result and / or arbitration conclusion corresponding to the historical error case.
7. The human-machine collaborative calibration system according to claim 5, characterized in that, The terminal also includes at least one expert terminal, and the preset confidence threshold includes a preset high confidence threshold and a preset low confidence threshold; the traffic splitting process includes: Based on the comparison results of the staff member's current credibility score with the high credibility threshold and the comparison results of the staff member's current credibility score with the low credibility threshold, then: When the staff member's current credibility score is higher than the high credibility threshold, the staff member's correction result is directly adopted; or, When the employee's current credibility score is higher than or equal to the low credibility threshold and the employee's current credibility score is lower than or equal to the high credibility threshold, the discrepancy case will be submitted to arbitration; or, When the staff member's current credibility score is lower than the low credibility threshold, a tiered arbitration process is initiated. The tiered arbitration process prioritizes cross-validating the discrepancy case with staff members whose current credibility score is higher than the high credibility threshold. When there is a disagreement in the cross-validation, the discrepancy case is submitted to the expert terminal for arbitration. When the cross-validation reaches a consensus, the correction result of the discrepancy case is confirmed based on the consensus result of the cross-validation.
8. The human-machine collaborative calibration system according to any one of claims 5 to 7, characterized in that, The nonlinear weight mapping function is at least one of the following: the Sigmoid function, a variant of the Sigmoid function, or a piecewise linear function.
9. An electronic device, characterized in that, It includes a processor and a memory, wherein the memory stores a computer program, which, when executed by the processor, implements the human-machine collaborative calibration method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processor, implements the human-machine collaborative calibration method as described in any one of claims 1 to 4.