User AI reasoning service method based on data quality intelligent evaluation

By using data quality scoring to adjust the inference accuracy of the AI ​​model in real time, the prediction bias problem caused by input data noise in existing technologies is solved, the robustness and accuracy of the model under different quality data are improved, and efficient noise suppression and prediction reliability are achieved.

CN120706562APending Publication Date: 2025-09-26PIO CLOUD COMPUTING (SHANGHAI) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510834041.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In existing AI inference services, models are prone to noise sensitivity when processing low-quality data, leading to overfitting or misjudgment, and lack a dynamic control mechanism to cope with dynamic fluctuations in input data quality.

Method used

The inference accuracy of the AI ​​model can be regulated in real time through data quality scoring, including adjusting the confidence threshold, adaptively selecting the model complexity level, calibrating the prediction output, filtering control, and controlling the feature retention ratio, etc., to optimize the number of iterations to suppress the prediction deviation caused by input data noise.

Benefits of technology

It improves the robustness of the model under data of different qualities, balances the accuracy and inclusiveness of the model output, reduces prediction bias under high noise, and improves inference efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120706562A_ABST
    Figure CN120706562A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of AI reasoning optimization, and discloses a user AI reasoning service method based on data quality intelligent evaluation, which comprises the following steps: a target AI model receives reasoning request data input by a user and executes data format and integrity verification, and the data comprises text, image or sensor time sequence data; performing multi-dimensional quality evaluation on the inference request data passing the verification, and aggregating and generating a data quality score for judging data reliability by extracting quality indexes and analyzing data features; the reasoning precision of a target AI model is dynamically regulated and controlled in real time based on the data quality score, so that prediction deviation caused by input data noise is inhibited; and executing an inference task by using the regulated and controlled target AI model, generating a prediction result, adding quality evaluation metadata, and returning the prediction result to the user terminal. According to the method, the reasoning precision of the AI model is regulated and controlled in real time through the data quality score, the problem of prediction deviation caused by input data noise is solved, and the robustness of the model under different quality data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of AI reasoning optimization technology, and in particular to a user AI reasoning service method based on intelligent data quality evaluation. Background Art

[0002] With the in-depth application of artificial intelligence in key areas such as medical diagnosis, autonomous driving, and financial risk management, user AI inference services have become essential infrastructure. These services use pre-trained models (such as CNN and Transformer) to process user-submitted text, images, and sensor data in real time, outputting predictions (such as disease classification, obstacle location, and fraud probability).

[0003] In existing AI inference services, models typically process user input data with a fixed configuration, ignoring the dynamic fluctuations in input data quality. In real-world applications, user-submitted data often introduces noise and anomalies due to factors such as acquisition device errors, transmission noise, labeling errors, and environmental interference, leading to significant deviations in model inference. Existing technologies face challenges including noise sensitivity and a lack of dynamic control mechanisms.

[0004] Specifically, high-precision models (such as deep neural networks) tend to amplify noise under low-quality data, resulting in overfitting or misjudgment (such as misdiagnosis caused by blur in medical images); moreover, traditional methods do not provide real-time feedback of data quality scores to the inference engine, and are unable to adaptively adjust model behavior to deal with prediction bias caused by noise. Summary of the Invention

[0005] The purpose of the present invention is to address the shortcomings of the above-mentioned existing technologies and to provide a user AI reasoning service method based on intelligent data quality assessment. The method adjusts the AI ​​model reasoning accuracy in real time through data quality scoring, solves the prediction bias problem caused by input data noise, improves the robustness of the model under different quality data, dynamically reduces or increases the confidence requirements according to the data quality, and balances the accuracy and inclusiveness of the model output.

[0006] A user AI reasoning service method based on intelligent data quality assessment includes the following steps: S1: The target AI model receives inference request data input by the user and performs data format and integrity verification. The data includes text, images, or sensor time series data. S2: Perform a multi-dimensional quality assessment on the verified inference request data. By extracting quality indicators and analyzing data features, the data quality score is aggregated to evaluate the reliability of the data. S3: Dynamically control the inference accuracy of the target AI model in real time based on the data quality score. The dynamic control method includes adjusting the confidence threshold, adaptively selecting the model complexity level, calibrating the prediction output, filtering control, controlling the feature retention ratio, and optimizing the number of iterations to suppress prediction bias caused by input data noise. S4: Use the regulated target AI model to perform reasoning tasks, generate prediction results, and attach quality assessment metadata to return to the user terminal.

[0007] Furthermore, in step S1, performing data format verification and integrity check includes: S11: Calculate missing fields based on the defined set of required fields and the set of received data fields. Check field integrity by calculating the ratio of missing fields to the required set of fields. At the same time, control the size of input data accepted or rejected based on the preset threshold of the AI ​​model status. S12: Make comprehensive verification decisions based on the field integrity and data size control results, perform global verification of the data format, and log the verification results; S13: If the verification result shows that the data is missing or damaged, the error feedback mechanism is immediately triggered to prompt the user to resubmit, ensuring that the initial data flow reliably enters the next stage.

[0008] Furthermore, in step S2, the multi-dimensional quality assessment includes: S21: Acquire verified data and quantify multiple sub-indicators by analyzing data characteristics, wherein the data characteristics include noise level and completeness; S22: Weighted combination of sub-indicators to generate the final data quality score , the calculation formula is as follows: in, It is a data accuracy indicator, which is calculated by comparing the characteristic distribution deviation of the input data with the benchmark data set. is a data completeness indicator, quantified by the proportion of missing fields, is a data relevance indicator, determined by feature importance analysis. 、 and They are configurable weight coefficients corresponding to the three indicators, satisfying , .

[0009] Furthermore, in step S3, adjusting the confidence threshold includes: Setting a baseline confidence threshold and quality score lower threshold , and compare the data quality scores and the lower quality score threshold : like , the dynamic confidence threshold is calculated according to the following formula: in, is the sensitivity adjustment factor, which is used to control the response speed of the confidence threshold to quality fluctuations; like , then directly apply the preset minimum confidence threshold .

[0010] Preferably, the adjustment of the confidence threshold further introduces a noise compensation mechanism, including: Calculating the noise factor ,in, is the number of abnormal data points detected, is the total amount of input data; Generate a comprehensive quality factor based on the noise factor , , where the noise factor And the optimal value is 0; The confidence threshold is calibrated twice according to the comprehensive quality factor. The calibration formula is as follows: , if the comprehensive quality factor Less than the preset threshold, forcing the minimum confidence threshold to be enabled as the confidence threshold.

[0011] Furthermore, in step S3, the adaptive selection model complexity level includes: Calculate the complexity score of the target AI model , ,in, is the comprehensive quality factor, is the complexity coefficient; Set the complexity level threshold for each model, based on the complexity score calculated The threshold range of the value determines the target AI model complexity level, and the model of corresponding complexity quality is adaptively selected according to the target AI model complexity level; For multiple candidate models, the softmax function is used to weight the multiple models of the candidate integration. The weight calculation formula is as follows: in, and Respectively and The sensitivity parameter of the candidate model, the high and low values ​​represent the accuracy priority and robustness priority of the model respectively. If the data quality score is Less than the preset threshold, force equal weighting , is the total number of candidate models.

[0012] Preferably, in step S3, the calibration prediction output includes: Generate nonlinear calibration factors ,in, is the predicted value calibration index, like , output calibration prediction value , Indicates the direct output value of the uncalibrated AI model; like , enable realm defaults As the prediction value, to prevent the prediction distortion caused by low-quality data.

[0013] More preferably, in step S3, the filtering control includes: Calculating dynamic filter strength ,in Set the data quality critical threshold as the maximum intensity value. When the data quality is less than the critical threshold, switch to the simplified model architecture and apply The corresponding noise reduction filter.

[0014] Furthermore, in step S3, the control feature retention ratio includes: Calculate the feature preservation ratio ,in is the scaling factor, when the feature preserves the ratio If the value is less than a preset threshold, only a predefined subset of core features is used for inference, excluding low-importance noise features.

[0015] Furthermore, in step S3, the number of optimization iterations includes: Set the number of baseline iterations , and the minimum number of iterations , according to the formula Calculate the adjustment value of the number of iterations. If the calculation result , then force the setting ; By executing Model inference iterations are repeated. When the data quality is high, the iterations are increased to improve the accuracy. When the data quality is low, the iterations are reduced to avoid noise amplification.

[0016] Compared with the prior art, the present invention has the following beneficial effects: This invention uses data quality scoring to adjust the AI ​​model's inference accuracy in real time, solving the problem of prediction bias caused by input data noise and improving the model's robustness under data of varying quality. The present invention dynamically lowers or raises the confidence requirement based on data quality, balancing the accuracy and inclusiveness of the model output; This method combines data quality scores and noise levels to calculate a comprehensive quality factor, which more comprehensively reflects the overall quality of the data and reduces prediction bias under high noise conditions. It dynamically selects model complexity based on data quality, using high-precision complex models for high-quality data and switching to simpler, more robust models for low-quality data to avoid overfitting or underfitting. It optimizes model input and calculation processes by adjusting parameters such as feature retention ratio, number of inference iterations, and learning rate, reducing noise interference and improving inference efficiency and stability. The present invention dynamically adjusts the weights of multiple candidate models based on data quality scores, weights high-precision models with high-quality data, and balances the weights or biases towards robust models with low-quality data, thereby reducing integration bias and improving prediction reliability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a user AI reasoning service method based on intelligent data quality assessment of the present invention. DETAILED DESCRIPTION

[0018] To make the purpose, technical solutions, and advantages of the embodiments of this application more clear, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0019] The present invention receives inference request data input by the user, performs a quality assessment on the inference request data, and generates a data quality score; based on the data quality score, dynamically adjusts the inference accuracy of the AI ​​model; uses the adjusted AI model to perform inference tasks and output prediction results. Obtain a data quality score Q; and determine a baseline confidence threshold T0. The disclosed embodiment provides a user AI inference service method based on intelligent data quality assessment. The method uses the data quality score to adjust the inference accuracy of the AI ​​model in real time, solves the problem of prediction bias caused by input data noise, improves the robustness of the model under different quality data, dynamically reduces or increases the confidence requirement according to the data quality, and balances the accuracy and inclusiveness of the model output. For example, the threshold is relaxed for low-quality data to allow more uncertain predictions and avoid missed judgments.

[0020] The specific implementation of the present invention is described below with reference to the accompanying drawings and embodiments.

[0021] Example 1 See also Figure 1 This embodiment provides a technical solution for a user AI reasoning service method based on intelligent data quality assessment, including the following steps: S1: The target AI model receives inference request data input by the user and performs data format and integrity verification. The data includes text, images, or sensor time series data. S2: Perform a multi-dimensional quality assessment on the verified inference request data. By extracting quality indicators and analyzing data features, the data quality score is aggregated to evaluate the reliability of the data. S3: Dynamically control the inference accuracy of the target AI model in real time based on the data quality score. The dynamic control method includes adjusting the confidence threshold, adaptively selecting the model complexity level, calibrating the prediction output, filtering control, controlling the feature retention ratio, and optimizing the number of iterations to suppress prediction bias caused by input data noise. S4: Use the regulated target AI model to perform reasoning tasks, generate prediction results, and attach quality assessment metadata to return to the user terminal.

[0022] In this embodiment, when receiving the inference request data input by the user in step S1, the system obtains the user-uploaded data in real time through a standardized API interface or web service, such as deploying a RESTful API on a cloud platform, and automatically verifies the data format (such as JSON or CSV) and integrity to prevent invalid input from causing system errors. During the specific operation of performing data format verification and integrity checks, the background service will check whether the data fields are complete and whether the size exceeds the threshold, and log it for subsequent auditing. If missing or damaged data is detected, the error feedback mechanism will be immediately triggered, prompting the user to resubmit, ensuring that the initial data flow reliably enters the next stage, including: S11: Calculate missing fields based on the defined set of required fields and the set of received data fields. Check field integrity by calculating the ratio of missing fields to the required set of fields. At the same time, control the size of input data accepted or rejected based on the preset threshold of the AI ​​model status. S12: Make comprehensive verification decisions based on the field integrity and data size control results, perform global verification of the data format, and log the verification results; S13: If the verification result shows that the data is missing or damaged, the error feedback mechanism is immediately triggered to prompt the user to resubmit, ensuring that the initial data flow reliably enters the next stage.

[0023] When users upload image recognition requests through mobile applications, the system verifies the image resolution and format compliance before passing it to the quality assessment module to avoid invalid payloads occupying resources.

[0024] Next, in step S2, which assesses the quality of the inference request data and generates a data quality score, the system uses a multi-dimensional evaluation algorithm to analyze data noise, outliers, and consistency.

[0025] Specifically, the multi-dimensional quality assessment includes: S21: Acquire verified data and quantify multiple sub-indicators by analyzing data characteristics, wherein the data characteristics include noise level and completeness; S22: Weighted combination of sub-indicators to generate the final data quality score , the calculation formula is as follows: in, It is a data accuracy indicator, which is calculated by comparing the characteristic distribution deviation of the input data with the benchmark data set. is a data completeness indicator, quantified by the proportion of missing fields, is a data relevance indicator, determined by feature importance analysis. 、 and They are configurable weight coefficients corresponding to the three indicators, satisfying , .

[0026] In this example, we use pretrained neural network models or statistical methods (such as Z-score-based anomaly detection) to scan data and calculate missingness rates, distribution deviations, and noise levels. These metrics are then aggregated into a 0-100 scoring system, where high scores indicate clean and reliable data, while low scores indicate potential noise interference. Operationally, the evaluation module is integrated into the inference pipeline, running lightweight scripts in real time. For example, it calculates the PSNR (peak signal-to-noise ratio) of image data or the consistency of word frequency in text data. The generated score is then stored in a database for access by the control module. In financial risk control scenarios, when a user enters a transaction record, the system evaluates the transaction amount for outliers and missing timestamps, generating a quality score. A score below 60 indicates high data noise, potentially leading to prediction bias due to input errors.

[0027] Then, step S3 is performed to dynamically adjust the inference accuracy of the target AI model in real time based on the data quality score. The system solves the prediction bias problem caused by input data noise and adjusts the model behavior through adaptive strategies.

[0028] Specifically, we set an activation threshold and determine whether dynamic precision adjustment needs to be activated by comparing the calculated data quality score with the activation threshold. If it is close to 1, the inference result is more reliable; otherwise, Low, the system may prompt the user to optimize the input.

[0029] First, set the confidence threshold to control the confidence requirements of the model output. Specifically: Setting a baseline confidence threshold and quality score lower threshold , and compare the data quality scores and lower quality score threshold : like , the dynamic confidence threshold is calculated according to the following formula: in, is the sensitivity adjustment factor, which is used to control the response speed of the confidence threshold to quality fluctuations; like , then directly apply the preset minimum confidence threshold .

[0030] Adjusting the confidence threshold further introduces a noise compensation mechanism, including: Calculating the noise factor ,in, is the number of abnormal data points detected, is the total amount of input data; Generate a comprehensive quality factor based on the noise factor , , where the noise factor And the optimal value is 0; The confidence threshold is calibrated twice according to the comprehensive quality factor. The calibration formula is as follows: , if the comprehensive quality factor Less than the preset threshold, forcing the minimum confidence threshold to be enabled as the confidence threshold.

[0031] In this embodiment, the parameter confidence represents the certainty of the model's prediction, which ranges from [0, 1]. The parameter T is the confidence threshold, which also ranges from [0, 1]. The optimal value is determined by cross-validation or domain requirements, such as 0.7 or 0.8, to balance accuracy and coverage. The meaning of this formula is to perform a binary decision, and only when the confidence exceeds Output prediction when the prediction is correct, otherwise it is considered invalid. This setting is to minimize erroneous output in AI reasoning and improve service reliability, especially in data quality-sensitive scenarios, to avoid low-confidence predictions misleading user decisions. Related to a user AI reasoning service method based on intelligent data quality evaluation, the user submits a data set request for quality analysis. Specifically, the model evaluates data quality indicators such as completeness scores and calculates confidence; for example, Set to 0.75. If the confidence level is >= 0.75, a quality report is returned. Otherwise, a low confidence level is output and a recommendation to optimize the data source is given. This ensures that the service only provides results when reliability is high, reducing false positives caused by low-quality data.

[0032] Among them, in the noise compensation mechanism, the first step is to identify the noise elements in the data, which means to detect errors, anomalies or irrelevant information by analyzing the data set, such as finding missing values ​​or wrong labels in the AI ​​reasoning input data. The second step is to calculate the noise index using the formula ,in Indicates the number of erroneous data points, ranging from [0, T], is the total number of data points and is a positive integer. The optimal value is (noise-free); the formula means to quantify the noise level as a proportional value and normalize it to [0,1]. This setting is to provide a standardized measure to facilitate the comparison of quality across datasets. The third step is to apply the noise level assessment, which means to calculate the noise level. The value is used to adjust the weight or filtering mechanism of the AI ​​reasoning service to ensure that high-noise data is downgraded. For example, in one embodiment, a user uploads an image dataset for an object detection AI service, and the system identifies noise elements such as blurred images (erroneous data points). Assuming T = 100 images and E = 10 blurred images, then , which indicates low noise level and high data quality, and can be used directly for inference; on the contrary, if When it is close to 1, the data cleaning prompt is triggered. Specifically, this process improves the reliability of AI reasoning. , where Q represents the data quality factor, ranging from [0,1], and the optimal value is 1 (indicating the highest quality); represents the noise factor, ranging from [0, ∞), with an optimal value of 0 (no noise). The meaning of this formula is: The higher the value, the better the overall quality. Apply a penalty to the noise to ensure that when the noise increases The value decreases to prevent high-quality data from being masked by noise. The reason for this setting is that in AI inference services, noise will amplify uncertainty, and the formula is designed to be nonlinear attenuation, giving priority to ensuring the reliability of low-noise scenarios while preventing Approaching infinity becomes negative or invalid.

[0033] In this embodiment, in a user AI reasoning service method based on intelligent data quality assessment, users upload medical imaging data for AI diagnosis. For example, when quantifying the data quality factor Q, the image clarity and metadata integrity are evaluated, and Q is calculated to be 0.8; when quantifying the noise factor N, the salt and pepper noise ratio in the image is detected. The measured value is 0.5. Subsequently, the formula F = 0.8 / (1 + 0.5) ≈ 0.53 is applied. This value is lower than the threshold of 0.6, so the system rejects the data for inference to avoid incorrect diagnosis.

[0034] To further dynamically control the accuracy, we perform the adaptive selection of the model complexity level in step S3, including: Calculate the complexity score of the target AI model , ,in, is the comprehensive quality factor, is the complexity coefficient; Set the complexity level threshold for each model, based on the complexity score calculated The threshold range of the value determines the target AI model complexity level, and a model with corresponding complexity quality is adaptively selected according to the target AI model complexity level.

[0035] In this embodiment, the formula quantifies the model complexity requirement: High Increases, indicating that high-quality data requires a high-precision model; when Low Reduce, choose a simple model under noisy data to reduce overfitting bias. For example, in a user AI image classification service, if the data quality assessment is (high quality), ,but .

[0036] Secondly, the model complexity level M is determined based on S: if S > S_high (preset threshold such as 0.8), then M=3; otherwise, if S > S_medium (preset such as 0.6), then M=2; otherwise, M=1. This step divides the model level according to the S value: M=3 represents a high-complexity model (such as a deep neural network), M=2 is medium, and M=1 is a simple model (such as a decision tree) to adapt to the accuracy and robustness requirements under different data qualities. For example, in the speech recognition service, if S=0.75 > S_medium=0.6, then M=2. Then load the AI ​​model of the corresponding level M. This step dynamically selects and deploys pre-existing models to respond to real-time data quality changes and ensure efficient resource utilization. Specifically, in a user text analysis and inference service, if M=2, a medium-complexity BERT variant model is loaded. Finally, the adjusted confidence threshold is used Perform reasoning with model M. Adaptive adjustments based on data quality (e.g. Low In one embodiment, for video recognition services in noisy environments, if (low quality), Adjust to 0.6 and use a lightweight model with M=1 for inference to reduce misjudgment.

[0037] In summary, when the data is of high quality, a complex model is selected to improve accuracy, and when the data is of low quality, a simple model is selected to enhance stability.

[0038] For multiple candidate models, the softmax function is used to weight the multiple models of the candidate integration. The weight calculation formula is as follows: in, and Respectively and The sensitivity parameter of the candidate model, the high and low values ​​represent the accuracy priority and robustness priority of the model respectively. If the data quality score is Less than the preset threshold, force equal weighting , is the total number of candidate models.

[0039] We adaptively assign model weights through the formula to ensure high When Q is low, high-precision models are prioritized, and robustness is enhanced when Q is low. Use the weighted ensemble inference step to combine the outputs of multiple models for the final prediction. In the conditional weight equalization step, when Below threshold When , average weight is used to avoid the deviation caused by low-quality data. The formula is essentially a softmax function. The exponential value of is normalized so that the sum of the weights is 1. This means that when High, Zoom in high The weight of the model, giving priority to high-precision models; When low, the weights are more evenly distributed or biased towards low The purpose of this setup is to dynamically reduce ensemble bias—maximizing accuracy when the data is high-quality and preventing single model failure when the data is low-quality.

[0040] In this embodiment, we apply the user AI reasoning service method, such as medical image diagnosis: users upload X-rays for disease detection, and the data quality Automatic scoring based on image clarity and noise (e.g. Indicates high quality). Candidate models include high-precision CNN ( , sensitive to noise) and robust ResNet ( ).like (set to 0.5), then calculate the weights: W_CNN ≈0.7, W_ResNet≈0.3, and weighted integration outputs high confidence diagnosis; if , the weights are equal (0.5 each), reducing the risk of misjudgment of low-quality images.

[0041] Next, further calibration prediction output in step accuracy dynamic control is performed, including: Generate nonlinear calibration factors ,in, is the predicted value calibration index, when When the noise increases, Reduce, thereby narrowing the prediction range and reducing the deviation caused by the noise amplification effect. This design is because high noise data is prone to amplification errors. The nonlinear decay can be stably predicted.

[0042] like , output calibration prediction value , Indicates the direct output value of the uncalibrated AI model; like , enable realm defaults As the prediction value, to prevent the prediction distortion caused by low-quality data.

[0043] pass and Dynamically generate a calibration factor to adjust the sensitivity of the prediction based on data quality to cope with the impact of noise. Step 2 directly applies this factor to scale the original prediction value to achieve flexibility and adaptability of the prediction. Step 3 introduces a threshold As a safety mechanism, when data quality is too low, the default value is enabled to prevent deviation amplification and ensure service reliability.

[0044] In this embodiment, the method is applied to user AI reasoning services, such as predicting the purchasing tendency of e-commerce users. For example, the AI ​​model outputs the original prediction based on user behavior data. (e.g. purchase probability 0.7); Calculated by the data quality assessment module, if the user data is complete and accurate, (high quality), Set to 1.2 (optimal value to enhance noise suppression), then , after calibration If the user data is noisy (e.g. many missing values), , , assuming ,but , enable (For example, the default probability is 0.5) to avoid prediction distortion caused by low-quality data.

[0045] Further filtering control in dynamic regulation in S3 includes: Calculating dynamic filter strength ,in Set the data quality critical threshold as the maximum intensity value. When the data quality is less than the critical threshold, switch to the simplified model architecture and apply The corresponding noise reduction filter. When lowering Increase because The value becomes larger, which directly suppresses the noise in the input data and reduces its negative impact on model reasoning; the formula is set in this way because low-quality data contains more noise. By increasing the filtering strength, interference can be actively filtered out to improve the stability of reasoning. Applying filtering means using a digital filter (such as Gaussian filtering) to process the input data. Controls the degree of filtering (for example, higher Strength results in more blurring) to remove noise. is a preset quality threshold (e.g. 0.3), when When the number of samples is large, use a simplified model with lower computational resources (such as a lightweight neural network) to avoid overfitting noise; otherwise, use a standard model (such as a deep neural network) to maintain high-accuracy inference.

[0046] In this example, for a facial recognition system: a user uploads an image for identity verification, and the system first evaluates the data quality score. (For example, based on image blur and illumination uniformity calculation, indicates low quality). Next, calculate ,in is the maximum intensity; when Gaussian filtering is applied to the image, the intensity of 7.5 significantly blurs the noise area. (set to 0.3), the system switches to a simplified model (such as MobileNet) for inference, outputting fast but lower-precision results to avoid incorrect recognition caused by noise in the standard model. (high quality image), , filtering is lightly applied and uses standard models to ensure high accuracy.

[0047] Furthermore, the control feature retention ratio in the dynamic regulation of S3 precision is carried out, including: Calculate the feature preservation ratio ,in is the scaling factor, when the feature preserves the ratio If the feature retention ratio is less than the preset threshold, only the predefined core feature subset is used for inference, and low-importance noise features are excluded. , aiming to Dynamically adjust the feature usage ratio to prevent low-quality data from introducing noise; select the topR ratio features, which means selecting the most important features from all features The proportion is used for AI reasoning to optimize model performance; if Less than threshold , only the core feature subset is used to ensure that the model remains stable under very low data quality. When lowered, Reduce, limit the number of features, and avoid noise features causing model deviation; when High, Close to 1, retaining most features. This formula is set up this way because when data quality fluctuates, reducing low-importance features can improve model stability and inference accuracy and prevent overfitting.

[0048] In this example, a user requests an AI inference service to detect credit card fraud, and the input data includes transaction amount, timestamp, and user location. Specifically, when the data quality score When 0.6 (medium quality), set ,calculate ; The system selects the top 90% important features for inference. down to 0.3 (low quality), and ,but , but 0.45 < 0.4, so only core features such as transaction amount are used and noisy location data are ignored to ensure reliable prediction.

[0049] Next, we optimize the number of iterations and further improve the dynamic control of step S3, including: Set the number of baseline iterations , and the minimum number of iterations , according to the formula Calculate the adjustment value of the number of iterations. If the calculation result , then force the setting ; By executing Model inference iterations are repeated. When the data quality is high, the iterations are increased to improve the accuracy. When the data quality is low, the iterations are reduced to avoid noise amplification.

[0050] Aims to obtain Indicates receiving data quality scores from the intelligent evaluation module, reflecting the reliability of the input data; determining Represents the preset initial number of iterations of model inference, which serves as the basis for adjustment; calculation Aims to Dynamically adjust the number of iterations; the last step ensures that the number of iterations is not less than the minimum value , to prevent the unstable results caused by too low iterations. When the value is high, the number of iterations is increased to improve the inference accuracy (high-quality data supports deeper processing), When the quality is low, iterations are reduced to avoid noise amplification (low-quality data reduces ineffective computation), thus balancing bias (reducing errors) and computational efficiency (saving resources). This formula is set up this way because it simply and directly quantifies the impact of quality on iterations, ensuring that resource allocation matches data reliability.

[0051] In this embodiment, the user submits image data to the AI ​​inference service for target detection, and the intelligent evaluation module outputs (indicating high quality), benchmark Set to 20 times. Calculate times, since 16 is greater than (e.g. 5 times), the system uses 16 iterations to improve detection accuracy; for example, if to 0.3 (low quality), times, avoiding amplifying noise and saving computing time.

[0052] In summary, the present invention first receives inference request data input by the user. This data typically comes from the user's terminal device, such as text, images, or sensor data transmitted via an API interface or web service. This step involves data preprocessing, such as format conversion and preliminary filtering, to ensure that the input is compatible with subsequent processing modules.

[0053] Next, the inference request data is quality assessed to generate a data quality score. Specifically, a built-in quality assessment module (such as a machine learning-based classifier or rule engine) analyzes multiple dimensions of the data, including completeness (such as the proportion of missing values), consistency (such as the deviation of the data distribution from historical data), noise level (such as outlier detection), and timeliness. The assessment process uses algorithms (such as entropy methods or deep learning models) to quantify these indicators and comprehensively generate a data quality score of 0-100, where a low score indicates high data noise or poor reliability. This step relies on real-time calculations to ensure that the score reflects the immediate state of the input data.

[0054] Based on the data quality score, the AI ​​model's inference accuracy is dynamically adjusted to address prediction bias caused by input data noise. The specific adjustment mechanism includes: when the data quality score falls below a preset threshold (e.g., 60), the system automatically reduces the model's inference accuracy settings, for example by adjusting the confidence threshold (e.g., from 0.95 to 0.80) to avoid overfitting noisy data, or switching to a more robust model variant (e.g., a lightweight neural network) to reduce computational complexity and enhance noise resistance. Conversely, when the score is higher, a high-precision mode is enabled, such as increasing the number of model layers or using ensemble learning to improve prediction accuracy. This adjustment process is implemented through a feedback loop: the system monitors the deviation between the score and the model output in real time and applies adaptive algorithms (e.g., PID controllers or Bayesian optimization) to dynamically fine-tune parameters, effectively suppressing noise-induced bias. (For example, in image recognition, noisy data can lead to misclassification; adjustment reduces the model's sensitivity to reduce the false positive rate.)

[0055] Finally, the regulated AI model is used to perform inference tasks and output prediction results. Under regulation, the model (such as a convolutional neural network or Transformer) processes input data, generates predictions (such as classification labels or regression values), and returns the results through the user interface. The system also records regulation logs for continuous optimization.

[0056] Example 2 Based on Example 1, this example also discloses some other dynamic precision control methods, including model deviation control and deviation accumulation control.

[0057] Specifically, the model deviation control includes: Calculate error tolerance based on the formula ,in is the preset maximum error tolerance, which is used to adjust the strictness of the model output. The range is defined by the system (such as 0.1-0.3). The optimal value is calibrated through experiments to balance accuracy and robustness. Then, during the inference process, the error tolerance is set to , to control the deviation threshold allowed by the model. Finally, if > ( is a preset threshold), the conservative prediction mode is enabled, which prioritizes outputting more generalized or low-risk results to reduce errors.

[0058] This method aims to obtain Assess data quality to ensure that subsequent regulation is based on objective indicators; calculate Dynamically adapt to data changes, when Increase tolerance when reducing to alleviate noise impact; set Applied to inference, constrains the model output range; enabling conservative mode reduces uncertainty risk when tolerance is high.

[0059] In this example, users use AI reasoning services to classify images. For example, users upload low-quality images (such as blurry or poorly lit images), and the data quality score is Reduced to 0.3 (range 0-1). Calculation ( The default setting is 0.2). When is 0.14, the model allows output of broad categories (such as transportation rather than specific models). is 0.1, then Trigger conservative mode, outputting safer results (such as unknown objects rather than misclassifications).

[0060] The deviation accumulation control includes: Calculate dynamic learning rate based on formula ; Then, during the inference process, adjust the model learning rate to ; Finally, if , then freeze the model weights, is the attenuation coefficient ( , range 0.5-5, optimal value is about 1.0), Rate the quality.

[0061] This method aims to obtain data quality scores Indicates receiving scores from the intelligent evaluation module in real time. The value quantifies the confidence of the input data (e.g., 0-1 range, 1 is the highest quality), which is used to reflect the noise level. Calculate the dynamic learning rate This is achieved through the formula, where is the baseline learning rate (range 0.001-0.1, the optimal value is determined by model tuning); the formula means that when When lowered, decrease, resulting in This decreases, thus inhibiting the model’s adaptation to low-quality data. Reducing the learning rate when it is low can prevent the accumulation of bias and avoid the model from overfitting the noise. Adjust the learning rate Refers to applying calculated values ​​in real time in AI inference services to control the model update speed. (critical value such as 0.3), the weight update is completely stopped to ensure model stability.

[0062] In this embodiment, users upload medical images for disease diagnosis. Specifically, the system evaluates the image quality. (such as fuzziness score), if (high quality), calculated Higher (such as , hour ), the model learns normally; if (medium quality), Reduced to about 0.015, suppressing the influence of noise; if (set to 0.3), freeze the weights and stop updating to prevent misdiagnosis bias.

[0063] Finally, it should be noted that the above description is merely a preferred embodiment of the present invention, and the scope of protection of the present invention is not limited to the above embodiment. All technical solutions based on the principles of the present invention are within the scope of protection of the present invention. It should be noted that improvements and modifications that do not depart from the principles of the present invention, which are apparent to those skilled in the art, should also be considered within the scope of protection of the present invention.

[0064] The technical features of the above-mentioned embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above-mentioned embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A user AI reasoning service method based on intelligent data quality assessment, characterized in that: The steps include: S1: The target AI model receives inference request data input by the user and performs data format and integrity verification. The data includes text, images, or sensor time series data. S2: Perform a multi-dimensional quality assessment on the verified inference request data. By extracting quality indicators and analyzing data features, the data quality score is aggregated to evaluate the reliability of the data. S3: Dynamically control the inference accuracy of the target AI model in real time based on the data quality score. The dynamic control method includes adjusting the confidence threshold, adaptively selecting the model complexity level, calibrating the prediction output, filtering control, controlling the feature retention ratio, and optimizing the number of iterations to suppress prediction bias caused by input data noise. S4: Use the regulated target AI model to perform reasoning tasks, generate prediction results, and attach quality assessment metadata to return to the user terminal.

2. The user AI reasoning service method based on data quality intelligent evaluation according to claim 1 is characterized in that: In step S1, the execution of data format verification and integrity check further includes: S11: Calculate missing fields based on the defined set of required fields and the set of received data fields. Check field integrity by calculating the ratio of missing fields to the required set of fields. At the same time, control the size of input data accepted or rejected based on the preset threshold of the AI ​​model status. S12: Make comprehensive verification decisions based on the field integrity and data size control results, perform global verification of the data format, and log the verification results; S13: If the verification result shows that the data is missing or damaged, the error feedback mechanism is immediately triggered to prompt the user to resubmit, ensuring that the initial data flow reliably enters the next stage.

3. The user AI reasoning service method based on data quality intelligent evaluation according to claim 1 is characterized in that: In step S2, the multi-dimensional quality assessment includes: S21: Acquire verified data and quantify multiple sub-indicators by analyzing data characteristics, wherein the data characteristics include noise level and completeness; S22: Weighted combination of sub-indicators to generate the final data quality score , the calculation formula is as follows: in, It is a data accuracy indicator, which is calculated by comparing the characteristic distribution deviation of the input data with the benchmark data set. is a data completeness indicator, quantified by the proportion of missing fields, is a data relevance indicator, determined by feature importance analysis. 、 and They are configurable weight coefficients corresponding to the three indicators, satisfying , .

4. The user AI reasoning service method based on data quality intelligent evaluation according to claim 3 is characterized in that: In step S3, adjusting the confidence threshold includes: Setting a baseline confidence threshold and quality score lower threshold , and compare the data quality scores and lower quality score threshold : like , the dynamic confidence threshold is calculated according to the following formula: in, is the sensitivity adjustment factor, which is used to control the response speed of the confidence threshold to quality fluctuations; like , then directly apply the preset minimum confidence threshold .

5. The user AI reasoning service method based on data quality intelligent evaluation according to claim 4 is characterized in that: The adjustment of the confidence threshold further introduces a noise compensation mechanism, including: Calculating the noise factor ,in, is the number of abnormal data points detected, is the total amount of input data; Generate a comprehensive quality factor based on the noise factor , , where the noise factor And the optimal value is 0; The confidence threshold is calibrated twice according to the comprehensive quality factor. The calibration formula is as follows: , if the comprehensive quality factor Less than the preset threshold, forcing the minimum confidence threshold to be enabled as the confidence threshold.

6. The user AI reasoning service method based on data quality intelligent evaluation according to claim 5 is characterized in that: In step S3, the adaptive selection model complexity level includes: Calculate the complexity score of the target AI model , ,in, is the comprehensive quality factor, is the complexity coefficient; Set the complexity level threshold for each model, based on the complexity score calculated The threshold range of the value determines the target AI model complexity level, and the model of corresponding complexity quality is adaptively selected according to the target AI model complexity level; For multiple candidate models, the softmax function is used to weight the multiple models of the candidate integration. The weight calculation formula is as follows: in, and Respectively and The sensitivity parameter of the candidate model, the high and low values ​​represent the accuracy priority and robustness priority of the model respectively. If the data quality score is Less than the preset threshold, force equal weighting , is the total number of candidate models.

7. The user AI reasoning service method based on data quality intelligent evaluation according to claim 3 is characterized in that: In step S3, the calibration prediction output includes: Generate nonlinear calibration factors ,in, is the predicted value calibration index, like , output calibration prediction value , Indicates the direct output value of the uncalibrated AI model; like , enable realm defaults As the prediction value, to prevent the prediction distortion caused by low-quality data.

8. The user AI reasoning service method based on data quality intelligent evaluation according to claim 3 is characterized in that: In step S3, the filtering control includes: Calculating dynamic filter strength ,in Set the data quality critical threshold as the maximum intensity value. When the data quality is less than the critical threshold, switch to the simplified model architecture and apply The corresponding noise reduction filter.

9. The user AI reasoning service method based on data quality intelligent evaluation according to claim 3 is characterized in that: In step S3, the control feature retention ratio includes: Calculate the feature preservation ratio ,in is the scaling factor, when the feature preserves the proportion If the value is less than a preset threshold, only a predefined subset of core features is used for inference, excluding low-importance noise features.

10. The user AI reasoning service method based on data quality intelligent evaluation according to claim 3 is characterized in that: In step S3, the optimization iteration number includes: Set the number of baseline iterations , and the minimum number of iterations , according to the formula Calculate the adjustment value of the number of iterations. If the calculation result , then force the setting ; By executing Model inference iterations are repeated. When the data quality is high, the iterations are increased to improve the accuracy. When the data quality is low, the iterations are reduced to avoid noise amplification.

Citation Information

Cited By

  • Reservoir flood forecasting system based on data-driven model

    CN121436426A