Method, device and electronic equipment for determining data quality
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-08-11
AI Technical Summary
[0004]本申请实施例提供一种数据质量的确定方法、装置及电子设备,以解决数据质量的检测准确度较低的问题
[0019]在本申请实施例中,获取在第一时刻之前的预设时段内的历史数据流,并从预先构建的数据库中检索与所述第一时刻相关联的事件信息;将所述历史数据流和所述事件信息输入大语言模型,得到所述大语言模型的输出结果,所述输出结果包括所述第一时刻的数据流的预测值,所述预测值是所述大语言模型根据所述历史数据流和所述事件信息对所述第一时刻的数据流进行预测得到的;获取所述第一时刻的数据流的实际值,并获取所述预测值相对所述实际值的偏差值;根据所述偏差值和预先获取的偏差阈值之间的相对大小,确定所述数据流的数据质量。这样,由于考虑了历史数据流和外部事件信息带来的影响,能够提高大语言模型输出的预测值的准确性,使得预测值能够适应业务的正常波动。基于该预测值对数据质量进行确定,能够提高数据质量的检测准确度。
Smart Images

Figure CN122548183A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a method, apparatus and electronic device for determining data quality. Background Technology
[0002] In recent years, large language models (LLMs) have been increasingly applied to the field of data quality inspection due to their superior natural language understanding capabilities. In the process of using LLMs for data quality inspection, LLMs can perform logical reasoning based on the relationships between historical data, thereby predicting data flow at future moments.
[0003] However, in real-world applications, data streams may experience unusual fluctuations at certain times, such as sudden increases or decreases in data stream volume. This can lead to poor accuracy in the predicted values of the data stream, resulting in low accuracy in data quality detection. Summary of the Invention
[0004] This application provides a method, apparatus, and electronic device for determining data quality, in order to solve the problem of low accuracy in data quality detection.
[0005] To solve the above-mentioned technical problems, this application is implemented as follows:
[0006] In a first aspect, embodiments of this application provide a method for determining data quality, the method comprising:
[0007] Obtain historical data streams within a preset time period prior to the first moment, and retrieve event information associated with the first moment from a pre-built database;
[0008] The historical data stream and the event information are input into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment.
[0009] Obtain the actual value of the data stream at the first moment, and obtain the deviation value of the predicted value relative to the actual value;
[0010] The data quality of the data stream is determined based on the relative magnitude between the deviation value and a pre-acquired deviation threshold.
[0011] Secondly, embodiments of this application provide a data quality determination apparatus, the apparatus comprising:
[0012] The first acquisition module is used to acquire historical data streams within a preset time period before the first moment, and retrieve event information associated with the first moment from a pre-built database;
[0013] The input module is used to input the historical data stream and the event information into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment.
[0014] The second acquisition module is used to acquire the actual value of the data stream at the first moment and to acquire the deviation value of the predicted value relative to the actual value.
[0015] The first determining module is used to determine the data quality of the data stream based on the relative magnitude between the deviation value and a pre-acquired deviation threshold.
[0016] Thirdly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the data quality determination method described in the first aspect.
[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the data quality determination method described in the first aspect.
[0018] Fifthly, a computer program product is provided, including computer instructions that, when executed by a processor, implement the steps of the method for determining data quality as described in the first aspect.
[0019] In this embodiment, historical data streams within a preset time period prior to a first moment are acquired, and event information associated with the first moment is retrieved from a pre-built database. The historical data streams and the event information are input into a large language model to obtain the output of the large language model. The output includes a predicted value of the data stream at the first moment, which is obtained by the large language model based on the historical data streams and the event information. The actual value of the data stream at the first moment is acquired, and the deviation value of the predicted value relative to the actual value is obtained. The data quality of the data stream is determined based on the relative magnitude between the deviation value and a pre-acquired deviation threshold. Thus, by considering the influence of historical data streams and external event information, the accuracy of the predicted value output by the large language model can be improved, allowing the predicted value to adapt to normal business fluctuations. Determining data quality based on this predicted value can improve the accuracy of data quality detection. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a method for determining data quality provided in an embodiment of this application;
[0022] Figure 2 This is a second flowchart of a method for determining data quality provided in an embodiment of this application;
[0023] Figure 3 This is a schematic diagram of the structure of a data quality determination device provided in an embodiment of this application;
[0024] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0025] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0026] This application provides a method, apparatus, and electronic device for determining data quality, in order to solve the problem of low accuracy in data quality detection.
[0027] See Figure 1 , Figure 1 This is a flowchart of a data quality determination method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0028] Step 101: Obtain historical data streams within a preset time period prior to the first moment, and retrieve event information associated with the first moment from a pre-built database;
[0029] Step 102: Input the historical data stream and the event information into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment.
[0030] Step 103: Obtain the actual value of the data stream at the first moment, and obtain the deviation value of the predicted value relative to the actual value;
[0031] Step 104: Determine the data quality of the data stream based on the relative magnitude between the deviation value and the pre-acquired deviation threshold.
[0032] The data stream can be a time-series data stream that monitors the quality of the data, such as the daily active user (DAU) count of the platform or the daily playback duration of the platform.
[0033] The first moment can be the current moment or any future moment. Here, we only take the predicted value of the first moment as an example. In actual prediction, the data streams at multiple moments can be predicted in the same way to obtain the predicted value sequence of the data stream corresponding to the time series.
[0034] The preset time period can be understood as a historical time period, such as a pre-set data collection time period used to analyze the fluctuation patterns of historical data streams.
[0035] In some implementations, a real-time data stream is received, and a historical data stream (such as data from the past 60 days) is maintained in a sliding window to facilitate the large model learning the fluctuation patterns of the historical data stream.
[0036] The pre-built database can be a structured knowledge base used to store multi-source business information and event information. The event information in the database can include the event content and the time range in which the event occurred (such as start time and end time). The first moment can be any moment within the time range in which the event occurred.
[0037] In some implementations, event information and its corresponding time range or occurrence time can be pre-associated and stored in a database. During retrieval, event information corresponding to that time can be retrieved from the database. For example, retrieving event documents for "January 1, 2026" will yield the event information that "January 1, 2026 is New Year's Day".
[0038] Once event information associated with the first moment is retrieved, the historical data stream and event information are input into the large language model, which can infer the predicted value of the data stream at the first moment based on the historical data stream and event information.
[0039] Leveraging the powerful sequence modeling and pattern recognition capabilities of LLM, we can learn the inherent patterns (such as trends, periods, and seasonality) from historical data streams and generate a dynamic, intelligent expected baseline (i.e., predicted value).
[0040] Furthermore, since outliers are often triggered by external events (such as system releases, marketing campaigns, or network failures), it is essential to make the model "see" this external event information. By employing Retrieval-Augmented Generation (RAG) technology, all known event information relevant to the first moment is actively retrieved and used as contextual input to the LLM.
[0041] This is equivalent to having an analyst check the recent "operations logs" and "business calendar" before making a prediction. For example, when LLM predicts the amount of data for the day, if RAG retrieves "the S-level drama series 'AAA' is released today," the large language model can "understand" and predict a high viewership value. If the actual value is still very low, the credibility of the alert is extremely high because the model has already taken this known event into account.
[0042] To further improve prediction accuracy, prompts can be constructed based on historical data streams and retrieved event information to guide the LLM to activate its inherent reasoning and knowledge association capabilities, i.e., attribution capabilities, thus achieving a functional leap.
[0043] After acquiring historical data streams and event information, the large model performs inference based on input prompts. For example, it analyzes the patterns and trends of historical data based on the historical data stream (data from the past week shows weekly periodicity and a bimodal characteristic), and combines this with recent event information to identify events that influence the data and how these events affect the data (the release of a popular TV series leads to an increase in viewership). By combining the trends in historical data and the impact of events on the data, it obtains the predicted value of the data stream at the first moment.
[0044] In some implementations, the prompt words may include information such as the role definition, task definition, and instruction requirements of the large language model, so that the large language model can predict the data stream at the first moment based on the prompt words.
[0045] In some implementations, the instruction requirements of the prompt word may also include the confidence level of the output and the reason for obtaining the output, so that the large language model can generate the confidence level of the output and the reason for obtaining the output based on the instruction.
[0046] In some implementations, multi-model collaborative prediction can be employed, integrating the inference results of multiple large language models. For example, three roles can be set up: business analysis expert, time series expert, and root cause inference expert. Each role handles the impact of business events, historical trend patterns, and anomaly cause analysis, respectively. The analysis results are aggregated through a voting mechanism to generate the final prediction results and root cause report.
[0047] LLM allows us to obtain the predicted values of the data stream at time step one, before the first time step. We can then obtain the actual values of the data stream at time step one or after the first time step, compare the predicted and actual values to obtain the deviation value, and compare this deviation value with a deviation threshold to determine the data quality of the data stream.
[0048] In some implementations, the deviation value is calculated as follows:
[0049] deviation = |actual_value - predicted_value| / actual_value.
[0050] Here, actual_value represents the actual value, and predicted_value represents the predicted value.
[0051] In some implementations, the volatility of relative error and event sequence characteristics can be considered, and the current deviation value can be normalized using the standard deviation of historical prediction errors. The deviation value is then determined based on the normalized value.
[0052] Data quality can be reflected by the consistency between the predicted and actual values of the data stream. Data quality is determined by comparing the deviation value with a deviation threshold, which can be dynamically determined based on the confidence level of the predicted value.
[0053] Traditional data stream comparison thresholds are horizontal, fixed lines (e.g., an alarm is triggered if CPU utilization > 80%). In this embodiment, however, LLM (Limited Least Meaning) is used to predict values based on historical data streams. Since historical data streams change dynamically according to actual conditions, the predicted values obtained from them are also fluctuating curves that dynamically change with historical data patterns. This results in predicted values that more closely reflect actual conditions. The alarm threshold for the predicted values is a dynamic bandwidth around this prediction curve (e.g., an alarm is triggered if the actual value deviates from the predicted value by ±10%).
[0054] When the deviation value exceeds the deviation threshold, an alarm message is generated that includes the actual value, the predicted value, the degree of deviation, and a list of potential root causes provided by the LLM.
[0055] Although LLM can predict future data streams based on historical data, its internal knowledge base is static and general, and cannot perceive the real-time status of the system where the target is located, domain-specific knowledge (such as "the impact of the Spring Festival on playback data"), or recent events (such as "recent application version upgrades may cause some metrics to be abnormal"), resulting in a lack of context in its predictions or judgments, which can easily lead to false positives or false negatives.
[0056] This application's embodiment considers the impact of historical data streams and external event information. The predicted value can adapt to normal business fluctuations, only triggering an alarm when unexpected deviations occur, greatly reducing false alarms. Determining data quality based on this predicted value can improve the accuracy of data quality detection.
[0057] Optionally, step 102, which involves inputting the historical data stream and the event information into the large language model to obtain the output of the large language model, includes:
[0058] Based on the historical data stream and the event information, a prompt word is constructed. The prompt word includes at least one of the following: trend features of the historical data stream, inference information on the impact of the event information on the trend features, and role information of the large language model. The large language model is used to infer the output result based on the role information.
[0059] The prompt word is input into the large language model to obtain the output of the large language model.
[0060] The prompt words can include instructions that require the LLM to perform multiple tasks.
[0061] For example, outputting the specific numerical value of the predicted value, evaluating the confidence level of the predicted value, conducting preliminary root cause analysis based on the retrieved event information when a large deviation value is detected, listing possible hypotheses, outputting in JSON format to ensure machine parsing, and achieving automated closed-loop, etc.
[0062] To improve the accuracy of predictions, cue words can be designed to guide LLM inferences. For example:
[0063] External input guidance: Prompts can be used to instruct the model to process various types of historical data (e.g., "The data from the past week shows weekly periodicity and daily bimodal characteristics"), rather than just numerical values, based on the trend characteristics of historical data streams. This allows the model to learn the trend characteristics of historical data streams and make predictions based on these characteristics.
[0064] Strengthening the thought chain: This requires the model to reason step by step: "First, historical trends show...; second, the retrieved context mentions a server offline event, which may lead to...; therefore, my prediction is..." This approach guides the model to consider the impact of external event information on the trend characteristics of the data stream.
[0065] Multi-expert voting mechanism: Different roles are assigned to LLMs in the prompt words (such as "statistician", "domain operation expert" and "causal reasoning expert"). LLMs can reason from different perspectives according to the information corresponding to different roles. Finally, all opinions are combined to output a consensus prediction and root cause analysis, which improves the reliability of the results.
[0066] The constructed prompts are sent to the LLM (such as GPT-OSS, Llama 3). Due to data sensitivity, a large model that can be deployed locally is selected, and its returned JSON format results are parsed.
[0067] Here is an example of a constructed prompt word:
[0068] # Roles and Missions
[0069] You are a senior data quality analysis expert. Your task is to predict the value of the next point in time series data based on historical data and the current context, and to evaluate the data quality.
[0070] # Contextual information (provided in real-time by the RAG system)
[0071] {retrieved_context}
[0072] # Historical data (past N points in time)
[0073] {historical_data_points}
[0074] # Command
[0075] 1. Please analyze the above context information, which describes system events or business activities that may affect the data values at the next point in time.
[0076] 2. Based on historical data patterns and contextual information, infer and predict the data value at the next point in time (i.e., the first moment). Please output a single, most probable predicted value.
[0077] 3. Assess the confidence level of your prediction (high / medium / low) and briefly explain your reasoning.
[0078] 4. If the deviation between the actual value and the predicted value exceeds 10%, please list 1-3 most likely reasons based on the provided context.
[0079] # Output Format
[0080] The following JSON format must be strictly followed:
[0081] {
[0082] "predicted_value": , / / Predicted value
[0083] "certainty": "", / / Confidence level
[0084] "reasoning": "", / / Reasons for confidence level
[0085] "potential_roots": [] / / Attribution reasons
[0086] }
[0087] Currently, simple instructions (such as "predict the next value") are typically used to guide LLMs in predicting data streams, failing to fully exploit their reasoning capabilities. Furthermore, the prompts lack nuanced guidance on task context, output format, decision logic, and uncertainty estimation, limiting model performance. The approach described above guides LLMs to incorporate historical data stream trends and event information, influencing model reasoning and prediction, thus improving the accuracy of forecasts.
[0088] Optionally, the output result may also include the confidence level of the predicted value; the method may further include:
[0089] Obtain the historical deviation value of the data stream before the first time point, wherein the historical deviation value is the deviation between the predicted value and the actual value of the data stream before the first time point;
[0090] The deviation threshold is determined based on at least one of the historical deviation values and the confidence levels of the predicted values;
[0091] The deviation threshold is negatively correlated with the confidence level, and positively correlated with the historical deviation value.
[0092] Obtain historical data streams from multiple time points prior to the first time point, and then obtain the predicted and actual values of the data streams for each of these time points. Based on the predicted and actual values of the data streams at each time point, obtain the historical deviation value.
[0093] Deviation thresholds can be critical values used to determine whether a data stream is abnormal.
[0094] In some implementations, the bias threshold is dynamically adjusted based on the confidence level (uncertainty field) returned by the LLM. For example, when the confidence level of the predicted value is "low", the bias threshold is increased to 15%; when the confidence level of the predicted value is "high", the bias threshold is decreased to 7%.
[0095] In some implementations, the deviation threshold is adjusted based on historical deviation values. For example, when the historical deviation value is large, the deviation threshold is increased; when the deviation threshold is small, the deviation threshold is decreased.
[0096] In some implementations, a network model is trained and then outputs a deviation threshold. The network model outputs the deviation threshold based on information such as the input confidence level, historical deviation values, statistical characteristics of historical deviation values, the time period, external event information, and historical alarm accuracy.
[0097] Since the deviation threshold is negatively correlated with the confidence level and positively correlated with the historical deviation, it can adaptively adjust the detection sensitivity based on the confidence level and historical deviation values, thereby reducing false alarms or missed alarms.
[0098] Optionally, the method further includes:
[0099] If the deviation value is greater than the deviation threshold, an alarm message is output. The alarm message includes at least one of the actual value, predicted value, and deviation value of the data stream at the first time moment, as well as the root cause information output by the large language model.
[0100] The root cause information refers to the reason why the predicted value determined based on the event information deviates from the actual value.
[0101] When the deviation value is greater than the deviation threshold, a structured alarm message can be generated, which includes the actual value, predicted value, deviation value and root cause information of the data stream.
[0102] In practice, the actual value sequence and the predicted value sequence of the data stream can be output. The data stream can include historical data stream and data stream at the first moment.
[0103] This makes it easier for operations and maintenance personnel to obtain the trend of deviation values, and then, based on this data and root cause information, determine the reason why the deviation value is greater than the deviation threshold.
[0104] Currently, when an LLM predicts a value that deviates significantly from the actual value, it may only be a "black box" alert. Operations or data analysts struggle to quickly understand "why the prediction deviation is so large this time?" or "what potential factors caused this anomaly?", making it difficult to quickly pinpoint the root cause and prolonging troubleshooting time.
[0105] Through the above implementation method, the alarm information includes specific root cause information, which reduces the time that operation and maintenance personnel spend analyzing and hypothesizing the root cause information and can improve the efficiency of root cause information determination.
[0106] Optionally, the method further includes:
[0107] Obtain user feedback on the alarm notification information, including the user's judgment on the accuracy of the root cause information;
[0108] The feedback information is added to the database.
[0109] Users can be operations and maintenance personnel or any user. Users can receive feedback information based on the alarm prompts, that is, the judgment results obtained by users through analysis of the data in the alarm prompts. For example, "correct alarm", "false alarm", "root cause correct".
[0110] This feedback information can be used as new documents, cleaned, and then stored back into the database, enabling the system to learn itself and continuously optimize.
[0111] In practice, feedback information can be linked with alarm records to form a feedback record. After simple cleaning (removing outliers and duplicate records), this record is automatically converted into a new document and added to the knowledge base. Linking this feedback record with the original event information and the analysis results of the data stream can then be used as training data for model optimization.
[0112] Optionally, the method further includes:
[0113] Obtain the data source and construct a document fragment based on the event information in the data source;
[0114] Obtain the time information corresponding to the event information, and index the document fragment with the time information;
[0115] The step of retrieving event information associated with the first time point from a pre-built database includes:
[0116] Retrieve multiple document fragments associated with the first time point from the database according to the index;
[0117] Among the plurality of document fragments, at least one document fragment with the highest correlation to the first time point is selected.
[0118] Data sources can include various structured and unstructured data sources, such as holidays, major events (e.g., the September 3rd military parade might cause a drop in viewership), the release dates of popular dramas, operational activity schedules, business calendars, and public opinion monitoring systems. Event information is extracted from the data sources to construct document fragments. These document fragments contain the time information of the event, i.e., the timestamp of the event or a related time range. For example, [Event] Popular drama AA will be released on September 11, 2025.
[0119] Use text embedding models (such as the Beijing Academy of Artificial Intelligence General Embedding (BGE)) to convert document fragments into vectors, and index the document fragments and time information and store them in vector databases (such as Chroma and Milvus).
[0120] Using the current moment and / or the most recent time as query criteria, retrieve the top K most relevant document fragments (i.e., at least one of the above) from the vector database. For example, the query could be: "What events occurred around September 3, 2025?" Input the indexed results into the above prompt, {retrieved_context}.
[0121] Document fragments related to the first moment may include event information that occurred at the first moment, or event information that occurred before or after the first moment.
[0122] Using the above method, RAG can be used to retrieve event information related to the first moment, thereby obtaining events that affect the data stream at the first moment and improving the accuracy of the predicted values of the data stream.
[0123] To facilitate understanding of this embodiment, specific examples will be used to illustrate it below.
[0124] like Figure 2 As shown, the data quality prediction method includes the following steps:
[0125] Step 1 (Offline): Build and index a multimodal knowledge base for retrieval. The knowledge base includes information such as multi-source data access operation and maintenance logs and business calendars. Feature vectors are extracted from the multi-source data, stored in the knowledge base, and indexed.
[0126] Step Two (Online): For incoming real-time data points, the system first retrieves relevant contextual information from the knowledge base using RAG based on the time point or timestamp. For example, [Event] The popular drama AA was launched on September 11, 2025.
[0127] Step 3 (Reasoning): Combine the retrieved context information and historical time-series data obtained from the historical data buffer with the prompt word constructor to form a complete prompt word, and input it into the LLM.
[0128] Step 4 (Prediction and Judgment): Based on historical time-series data and contextual information, LLM outputs predicted values and inference information. The deviation between the predicted and actual values is calculated; if the deviation exceeds a dynamically calculated threshold, an alarm is triggered.
[0129] Step 5 (Report): The alarm information includes not only the magnitude of the deviation, but also a preliminary root cause analysis generated by LLM based on the search context, which can help operations personnel or analysts diagnose problems and generate feedback results.
[0130] Feedback results can be input into the feedback learning module, which will process the feedback results and add them to the knowledge base.
[0131] Data quality inspection aims to identify anomalies, errors, and inconsistencies in data, such as missing values, format errors, numerical offsets, and logical contradictions.
[0132] In this embodiment, the powerful sequence modeling capabilities of Large LLM (Limited Ledger Model) are utilized to learn from historical time-series data and predict the data value at the next time point. Data quality is assessed by calculating the dynamic deviation between the predicted and actual values (such as Z-score and relative error), thus achieving adaptive and intelligent anomaly detection.
[0133] By constructing a dedicated knowledge base, multidimensional information related to the time-series data to be detected (such as system change logs, business activity calendars, network status reports, historical fault records, etc.) is indexed. Before generating predictions and making anomaly judgments, RAG is first used to retrieve key information related to the current time point, and this information is injected as context into the LLM's prompts. This makes the LLM's predictions no longer "blind guesses," but rather "context-aware" reasoning based on the real-time state of the system.
[0134] In the design of cue words, structured cue words guide LLM to perform multi-step reasoning, quantify uncertainty, and provide preliminary root cause hypotheses, upgrading LLM from a simple predictor to an analysis assistant, which can improve the accuracy of the predicted values of the data stream.
[0135] See Figure 3 , Figure 3This is a schematic diagram of the structure of a data quality determination device provided in an embodiment of this application, as shown below. Figure 3 As shown, the data quality determination device 300 includes:
[0136] The first acquisition module 301 is used to acquire historical data streams within a preset time period before the first moment, and retrieve event information associated with the first moment from a pre-built database;
[0137] Input module 302 is used to input the historical data stream and the event information into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment.
[0138] The second acquisition module 303 is used to acquire the actual value of the data stream at the first moment and to acquire the deviation value of the predicted value relative to the actual value.
[0139] The first determining module 304 is used to determine the data quality of the data stream based on the relative magnitude between the deviation value and a pre-acquired deviation threshold.
[0140] Optionally, the input module includes:
[0141] A construction submodule is used to construct prompt words based on the historical data stream and the event information. The prompt words include at least one of the following: trend features of the historical data stream, inference information on the impact of the event information on the trend features, and role information of the large language model. The large language model is used to infer the output result based on the role information.
[0142] The input submodule is used to input the prompt words into the large language model and obtain the output results of the large language model.
[0143] Optionally, the output result may also include the confidence level of the predicted value; the device may further include:
[0144] The third acquisition module is used to acquire the historical deviation value of the data stream before the first moment, wherein the historical deviation value is the deviation between the predicted value and the actual value of the data stream before the first moment;
[0145] The second determining module is used to determine the deviation threshold based on at least one of the historical deviation value and the confidence level of the predicted value;
[0146] The deviation threshold is negatively correlated with the confidence level, and positively correlated with the historical deviation value.
[0147] Optionally, the device further includes:
[0148] The output module is used to output an alarm message when the deviation value is greater than the deviation threshold. The alarm message includes at least one of the actual value, predicted value, and deviation value of the data stream at the first time moment, as well as the root cause information output by the large language model.
[0149] The root cause information refers to the reason why the predicted value determined based on the event information deviates from the actual value.
[0150] Optionally, the device further includes:
[0151] The fourth acquisition module is used to acquire user feedback information on the alarm prompt information, the feedback information including the user's judgment result on the accuracy of the root cause information;
[0152] An add module is used to add the feedback information to the database.
[0153] Optionally, the device further includes:
[0154] The fourth acquisition module is used to acquire a data source and construct a document fragment based on the event information in the data source;
[0155] The fifth acquisition module is used to acquire the time information corresponding to the event information and to index the document fragment with the time information;
[0156] The first acquisition module is specifically used for:
[0157] Retrieve multiple document fragments associated with the first time point from the database according to the index;
[0158] Among the plurality of document fragments, at least one document fragment with the highest correlation to the first time point is selected.
[0159] The data quality determination device can achieve Figure 1 The various processes implemented in the method embodiments can achieve the same technical effect, and will not be described again here to avoid repetition.
[0160] like Figure 4 As shown, this application embodiment also provides an electronic device 400, including: a processor 401, a memory 402, and a program stored in the memory 402 and executable on the processor 401. When the program is executed by the processor 401, it implements the various processes of the above-described data quality determination method embodiment and can achieve the same technical effect. To avoid repetition, it will not be described again here.
[0161] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the various processes of the above-described data quality determination method embodiments and achieves the same technical effects. To avoid repetition, it will not be described again here. The computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc.
[0162] This application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the above-described... Figure 1 The various processes of the method embodiments shown can achieve the same technical effect, and will not be described again here to avoid repetition.
[0163] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0165] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.
Claims
1. A method of determining data quality, characterized by, include: Obtain historical data streams within a preset time period prior to the first moment, and retrieve event information associated with the first moment from a pre-built database; The historical data stream and the event information are input into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment. Obtain the actual value of the data stream at the first moment, and obtain the deviation value of the predicted value relative to the actual value; The data quality of the data stream is determined based on the relative magnitude between the deviation value and a pre-acquired deviation threshold.
2. The method of claim 1, wherein, The step of inputting the historical data stream and the event information into the large language model to obtain the output of the large language model includes: Based on the historical data stream and the event information, a prompt word is constructed. The prompt word includes at least one of the following: trend features of the historical data stream, inference information on the impact of the event information on the trend features, and role information of the large language model. The large language model is used to infer the output result based on the role information. The prompt word is input into the large language model to obtain the output of the large language model.
3. The method of claim 1, wherein, The output result also includes the confidence level of the predicted value; the method further includes: Obtain the historical deviation value of the data stream before the first time point, wherein the historical deviation value is the deviation between the predicted value and the actual value of the data stream before the first time point; The deviation threshold is determined based on at least one of the historical deviation values and the confidence levels of the predicted values; The deviation threshold is negatively correlated with the confidence level, and positively correlated with the historical deviation value.
4. The method according to claim 1 or 3, characterized in that, The method further includes: If the deviation value is greater than the deviation threshold, an alarm message is output. The alarm message includes at least one of the actual value, predicted value, and deviation value of the data stream at the first time moment, as well as the root cause information output by the large language model. The root cause information refers to the reason why the predicted value determined based on the event information deviates from the actual value.
5. The method according to claim 4, characterized in that, The method further includes: Obtain user feedback on the alarm notification information, including the user's judgment on the accuracy of the root cause information; The feedback information is added to the database.
6. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Obtain the data source and construct a document fragment based on the event information in the data source; Obtain the time information corresponding to the event information, and index the document fragment with the time information; The step of retrieving event information associated with the first time point from a pre-built database includes: Retrieve multiple document fragments associated with the first time point from the database according to the index; Among the plurality of document fragments, at least one document fragment with the highest correlation to the first time point is selected.
7. An apparatus for determining data quality, characterized by include: The first acquisition module is used to acquire historical data streams within a preset time period before the first moment, and retrieve event information associated with the first moment from a pre-built database; The input module is used to input the historical data stream and the event information into the large language model to obtain the output result of the large language model. The output result includes the predicted value of the data stream at the first moment. The predicted value is obtained by the large language model based on the historical data stream and the event information to predict the data stream at the first moment. The second acquisition module is used to acquire the actual value of the data stream at the first moment and to acquire the deviation value of the predicted value relative to the actual value. The first determining module is used to determine the data quality of the data stream based on the relative magnitude between the deviation value and a pre-acquired deviation threshold.
8. An electronic device, comprising: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the steps of the method for determining data quality as described in any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method for determining data quality as described in any one of claims 1 to 7.
10. A computer program product, characterised in that, It includes computer instructions that, when executed by a processor, implement the steps of the method for determining data quality as described in any one of claims 1 to 7.