Data cleaning method and system for AI intelligent database
By analyzing the impact of external events on data breakpoints, dynamically identifying and adaptively cleaning breakpoint data in AI intelligent databases, the problem of misidentifying data changes in the existing technology is solved, and the level of data quality management and decision-making support is improved.
Patent Information
- Application Number
- CN202511063154.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-31
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-07-31
AI Technical Summary
When faced with mutational data segments caused by external factors, existing AI intelligent databases are difficult to accurately identify data changes, and may mistake data that reflects real market signals as abnormal, resulting in information loss or degradation of data quality.
By analyzing the external event-driven impact of data breakpoints, obtain the driving effect coefficient of external event data on internal data, dynamically identify the breakpoint data and perform adaptive cleaning, and retain structural data changes with business significance.
It significantly improves the data quality management capabilities of intelligent databases in high dynamic environments, avoids information losses caused by mistaken cleaning, and improves data accuracy and reliability.
Smart Images

Figure CN120578656A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of data processing technology, and specifically relates to a data cleaning method and system for an AI intelligent database. Background Art
[0002] An AI intelligent database is a database system that integrates artificial intelligence technology. It can automatically manage, analyze, and optimize data. It uses machine learning and deep learning algorithms to automatically identify outliers, missing values, and noisy data, and clean and repair them. It is suitable for enterprise applications that require rapid decision-making. Data cleaning is a key step in data preprocessing, aiming to improve data accuracy, completeness, and consistency. This includes identifying noisy data such as missing values, duplicate values, and outliers, and taking appropriate action to address them. By integrating data cleaning capabilities, AI intelligent databases can automatically complete data preprocessing and optimization, providing a high-quality data foundation for data analysis and machine learning model training.
[0003] Existing data cleaning methods based on AI-powered intelligent databases typically rely on predefined rules or expert knowledge, using algorithms such as clustering, anomaly detection, and regression prediction to identify outliers or fill in missing values based on the data's global statistical characteristics. However, due to a lack of recognition and modeling of complex external features, it is often difficult to accurately capture and identify data changes when faced with sudden changes in data segments caused by external factors. Directly applying predefined cleaning strategies may mistakenly identify jumps in data that reflect real market signals and business changes as anomalies, resulting in the loss of valuable information. If left uncleaned, data quality issues will be introduced, reducing the accuracy and reliability of subsequent analysis and modeling. Summary of the Invention
[0004] To address the above issues, the present invention provides a data cleaning method and system for an AI intelligent database. By analyzing the impact of external events driving data breakpoints and confirming the rationality of data changes within the database, the breakpoint data can be dynamically identified and adaptively cleaned to effectively retain structural data changes with business significance, avoid information loss caused by incorrect cleaning, and significantly improve the data quality management capabilities and decision support level of the intelligent database in a highly dynamic environment.
[0005] According to a first aspect of an embodiment of the present application, a method for cleaning data in an AI intelligent database is provided, the method comprising: Acquiring external event data and internal data of the AI intelligent database; Obtaining a driving effect coefficient of the external event data on any breakpoint data of the internal data; Obtaining, according to the driving effect coefficient, an event-driven score of the external event data on any breakpoint data; Based on the event-driven scoring, adaptive data cleaning is performed on the internal data of the AI intelligent database.
[0006] In one embodiment, obtaining the driving effect coefficient of the external event data on any breakpoint data of the internal data includes: Obtaining the logical jump degree of the neighborhood segment data of any breakpoint data; Obtaining a linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data; Obtaining a trigger effect coefficient of the external event data on the internal data; A driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained according to the logic jump degree, the linkage coefficient, and the trigger effect coefficient.
[0007] In one embodiment, obtaining the logic jump degree of the neighborhood segment data of any breakpoint data includes: Obtaining the deviation of any breakpoint data of any type of data of the internal data; Obtain the data type and quantity of the internal data; Obtaining the correlation coefficient between any two types of data in the internal data and the neighborhood segment data of any breakpoint data; Obtain the mean value of the correlation coefficient of any two types of data in the internal data in the neighborhood segment data of all breakpoint data; The logical jump degree of the neighborhood segment data of any breakpoint data is obtained according to the deviation, the number of data types, the correlation coefficient, and the mean.
[0008] In one embodiment, obtaining the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data includes: Obtaining a heat index change sequence of the external event data; Obtaining a data change sequence of neighborhood segment data of any breakpoint data of any type of data of the internal data; Obtaining the Pearson correlation coefficient between the heat index change sequence and the data change sequence; According to the Pearson correlation coefficient, a linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data is obtained.
[0009] In one embodiment, obtaining the trigger effect coefficient of the external event data on the internal data includes: Obtaining the event coreness of any keyword in the external event data; According to the event coreness, obtaining the semantic matching degree between the external event data and the AI intelligent database; According to the semantic matching degree, a triggering effect coefficient of the external event data on the internal data is obtained.
[0010] In one embodiment, obtaining the event coreness of any keyword in the external event data includes: Clustering all keywords in the external event data based on point mutual information values between the keywords; Obtain the number of keywords in the cluster where any keyword in the external event data is located; Obtaining the total number of keywords in the external event data; Obtaining a point mutual information value between any two keywords in the external event data; The event coreness of any keyword in the external event data is obtained according to the number of keywords, the total number, and the point mutual information value.
[0011] In one embodiment, obtaining the semantic matching degree between the external event data and the AI intelligent database according to the event coreness includes: Obtaining the number of intersections between the external event data and the keywords of the internal data; Obtaining a union number of keywords of the external event data and the internal data; Obtaining a first event coreness of any keyword in the external event data; Obtaining a second event coreness of any keyword in the internal data; According to the number of intersections, the number of unions, the first event coreness, and the second event coreness, a semantic matching degree between the external event data and the AI intelligent database is obtained.
[0012] In one embodiment, obtaining the trigger effect coefficient of the external event data on the internal data according to the semantic matching degree includes: Obtaining a type semantic matching degree between the external event data and any type of data of the internal data; According to the semantic matching degree between the external event data and the AI intelligent database, and the type semantic matching degree, a triggering effect coefficient of the external event data on the internal data is obtained.
[0013] In one embodiment, obtaining, based on the driving effect coefficient, an event-driven score of the external event data on any breakpoint data includes: Obtaining a set number of data of the neighborhood segment data of any breakpoint data; Obtaining the logic jump degree of any two adjacent neighborhood segment data in the set number of neighborhood segment data; Obtaining the logic recovery degree of any breakpoint data of any type of data in the internal data according to the set data quantity and the logic jump degree of any two adjacent neighborhood segment data; An event driving score of the external event data for any breakpoint data is obtained according to the driving effect coefficient and the logic recovery degree.
[0014] According to a second aspect of an embodiment of the present application, a data cleaning system for an AI intelligent database is provided, the system including a database platform, the database platform including: a memory having a computer program stored thereon; A processor is used to execute the computer program in the memory to implement the steps of any one of the methods in the first aspect.
[0015] In summary, the embodiment of the present application provides a data cleaning method for an AI intelligent database, the method comprising: obtaining external event data and internal data of the AI intelligent database; obtaining the driving effect coefficient of the external event data on any breakpoint data of the internal data; obtaining the event-driven score of the external event data on any breakpoint data according to the driving effect coefficient; and performing adaptive data cleaning on the internal data of the AI intelligent database according to the event-driven score. The embodiment of the present application analyzes the external event-driven impact of data breakpoints to confirm the rationality of the internal data changes in the database, thereby dynamically identifying breakpoint data and adaptively cleaning the breakpoint data. It can effectively retain structural data changes with business significance, avoid information loss caused by incorrect cleaning, and significantly improve the data quality management capabilities and decision support levels of intelligent databases in highly dynamic environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the implementation scheme of the present application, the following is a brief introduction to the drawings required for use in the implementation scheme. It should be understood that the drawings only show certain implementation schemes of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on the drawings without paying any creative work.
[0017] Figure 1The present invention is a flowchart of a data cleaning method for an AI intelligent database according to an exemplary embodiment.
[0018] Figure 2 The present invention is a flowchart showing a method for obtaining a driving effect coefficient of external event data on any breakpoint data of internal data according to an exemplary embodiment.
[0019] Figure 3 The present invention is a flowchart showing a method for obtaining the logical jump degree of neighborhood segment data of any breakpoint data according to an exemplary embodiment.
[0020] Figure 4 The figure is a schematic diagram showing neighborhood segment data of breakpoint data according to an exemplary embodiment.
[0021] Figure 5 The present invention is a flowchart showing a method for obtaining a linkage coefficient between external event data and neighborhood segment data of any breakpoint data according to an exemplary embodiment.
[0022] Figure 6 The present invention is a flowchart showing a method for obtaining a trigger effect coefficient of external event data on internal data according to an exemplary embodiment.
[0023] Figure 7 The present invention is a flowchart showing a method for obtaining the event coreness of any keyword in external event data according to an exemplary embodiment.
[0024] Figure 8 It is a flowchart of a method for obtaining the semantic matching degree between external event data and an AI intelligent database based on event coreness according to an exemplary embodiment.
[0025] Figure 9 The present invention is a flowchart showing a method for obtaining a trigger effect coefficient of external event data on internal data based on semantic matching according to an exemplary embodiment.
[0026] Figure 10 The present invention is a flowchart of a method for obtaining an event-driven scoring of any breakpoint data by external event data according to a driving effect coefficient, according to an exemplary embodiment.
[0027] Figure 11 The present invention is a block diagram of a data cleaning system for an AI intelligent database according to an exemplary embodiment.
[0028] Figure 12 It is a block diagram of a database platform according to an exemplary embodiment. DETAILED DESCRIPTION
[0029] In order to clearly illustrate the technical features of this solution, this application is described in detail below through specific implementation methods and in conjunction with the accompanying drawings.
[0030] The following describes embodiments of the present application in more detail with reference to the accompanying drawings. Although certain embodiments of the present application are shown in the accompanying drawings, it should be understood that the present application can be implemented in various forms and should not be construed as limited to the embodiments described herein. Instead, these embodiments are provided to provide a more thorough and complete understanding of the present application. It should be understood that the drawings and embodiments of the present application are for illustrative purposes only and are not intended to limit the scope of protection of the present application.
[0031] It should be understood that the various steps described in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.
[0032] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0033] It should be noted that the concepts of "first" and "second" mentioned in this application are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0034] It should be noted that the modifiers "one" and "multiple" mentioned in this application are illustrative and non-restrictive. Those skilled in the art will understand that, unless the context clearly indicates otherwise, they should be understood as "one or more." In the description of this application, unless otherwise specified, "multiple" means two or more than two, and other quantifiers are similar. "At least one item (item)", "one (item) or more (items)" or similar expressions refer to any combination of these items (items), including any combination of single items (items) or plural items (items). For example, at least one item (item) a can refer to any number of a; for another example, one (item) or more (items) of a, b, and c can refer to: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural. "And / or" is a relationship that describes the association of related objects, indicating that three relationships can exist. For example, A and / or B can refer to three situations: the existence of A alone, the existence of A and B at the same time, and the existence of B alone, where A and B can be singular or plural.
[0035] Although operations or steps are described in a specific order in the drawings in the embodiments of the present application, this should not be understood as requiring that these operations or steps be performed in the specific order shown or in a serial order, or that all of the operations or steps shown be performed to obtain the desired result. In the embodiments of the present application, these operations or steps may be performed in serial; these operations or steps may be performed in parallel; or some of these operations or steps may be performed.
[0036] At the same time, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and relevant provisions.
[0037] First, the application scenario of this application is explained. In the prior art, when cleaning AI intelligent database data through predefined rules and logic between data, if the data jumps due to external events, it violates the logic and rules between data and is easily treated as abnormal data for cleaning, resulting in the loss of valuable information.
[0038] In response to the problems existing in the prior art, this application provides a data cleaning method and system for an AI intelligent database. By analyzing the correlation and driving impact of external events on current data changes before and after data breakpoints occur, the rationality of data jumps is judged, thereby achieving adaptive cleaning of breakpoint jump data in the intelligent database. This can effectively retain structural data changes with business significance, avoid information loss caused by incorrect cleaning, and significantly improve the data quality management capabilities and decision support level of the intelligent database in a highly dynamic environment.
[0039] The embodiments of the present application automatically identify mutation data points that do not conform to semantics or historical trends by establishing logical constraint relationships between data and combining them with the statistical distribution of historical intervals. This can capture potential logical contradictions between multiple fields and effectively improve the accuracy of data anomaly identification. By implementing semantic association modeling between database labels and external event text, it can quickly determine whether the current jump is likely driven by a related event, avoiding misjudging real data changes affected by macro events as data anomalies, and achieving more context-aware data cleaning. By analyzing the temporal synchronization and delayed response of the rise in external event popularity and data jumps, it constructs a causal relationship between event-driven data, and determines from a temporal perspective whether database data changes are driven by events. This significantly improves the scientific nature of jump point identification and enhances the dynamic understanding of complex background data. An event-driven scoring model is constructed to quantitatively evaluate each jump breakpoint data, intelligently distinguishing "data to be cleaned" (abnormal data) from "data to be retained" (real signal change data), and automatically determining the corresponding processing method (such as correction, marking, segmentation, etc.). This realizes personalized, adaptive, and traceable data cleaning, and significantly improves the data governance capabilities of intelligent databases in complex environments. The present application is described below with reference to specific embodiments.
[0040] Figure 1 This is a flow chart showing a data cleaning method for an AI intelligent database according to an exemplary embodiment. Figure 1 As shown, the embodiment of the present application provides a data cleaning method for an AI intelligent database, which may include the following steps: In step S10, external event data and internal data of the AI intelligent database are obtained.
[0041] In this step, external event data and internal data from the AI intelligent database are acquired. For example, data extraction and connection can be performed first, using ETL (Extract-Transform-Load) technology to extract data from distributed databases and real-time data streams through database connectors (JDBC) or API interfaces. For real-time data, a stream processing platform (Flink) is used to achieve low-latency collection. By parsing the database schema and data dictionary, the definitions of each data field, table structure information, and related business rules are collected. Next, data sources are integrated, and business-related external event data is collected through web crawlers and API interfaces (such as news and social media monitoring). Event information is then parsed, using natural language processing techniques to perform entity recognition and keyword extraction on the external text data, generating structured event records with time tags and event categories. Popularity indicators in the external data, such as news dissemination volume, search index, and social media discussion frequency, are then integrated to quantify the impact of the external event. Finally, the data is preprocessed. Since the data formats of different systems may vary, a preprocessing module is used to convert the original data into a unified structured format, record key information such as the data's timestamp and source identifier, and align the external event data with the internal data of the AI intelligent database.
[0042] In step S20 , a driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained.
[0043] In this step, the driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained. For example, the logical jump degree of the neighborhood segment data of any breakpoint data can be obtained first, followed by the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data, and then the triggering effect coefficient of the external event data on the internal data is obtained. Finally, based on the logical jump degree, linkage coefficient, and triggering effect coefficient, the driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained.
[0044] In step S30, an event-driven score of the external event data on any breakpoint data is obtained according to the driving effect coefficient.
[0045] In this step, based on the driving effect coefficient, the event-driven score of the external event data for any breakpoint data is obtained. For example, a set number of neighborhood segments for any breakpoint data can be first obtained, followed by obtaining the logical jump degree of any two adjacent neighborhood segment data within the set number of neighborhood segments. Then, based on the set number and the logical jump degree of any two adjacent neighborhood segment data, the logical recovery degree of any breakpoint data of any type of data in the internal data is obtained. Finally, based on the driving effect coefficient and the logical recovery degree, the event-driven score of the external event data for any breakpoint data is obtained.
[0046] In step S40, adaptive data cleaning is performed on the internal data of the AI intelligent database according to the event-driven scoring.
[0047] In this step, the internal data of the AI intelligent database is adaptively cleaned according to the event-driven score. For example, there may be multiple external events, and the maximum event-driven score of the data jump point is selected as the basis for judging the data jump point. If the event-driven score is high, it indicates that the data jump is driven by a real event. It is recommended to retain or perform a local repair, while retaining the breakpoint mark for subsequent analysis; retaining the real event-driven jump data can ensure that the data contains important signals reflecting changes in the external environment, avoiding the erasure of key information due to unified cleaning; at the same time, the local repair process ensures that the data maintains real changes while improving the data quality and the stability of subsequent modeling.
[0048] Specifically, when the event-driven score exceeds a first threshold (e.g., 0.7), the transition data is retained. This transition data may reflect real market changes or business adjustments. If the transition data contains minor noise or local anomalies, the data can be corrected while preserving the underlying changes (e.g., using local mean smoothing or minor corrections based on local statistical models). The data breakpoint information is recorded in the cleansing results to facilitate further business analysis and risk assessment.
[0049] When the event-driven score is less than or equal to the first score threshold (for example, 0.7) and greater than the second score threshold (for example, 0.3), it indicates that the data contains both certain external event-driven signals and a certain degree of uncertainty or noise. In this case, local repair and slight adjustments can be made. The system automatically applies local smoothing, local mean correction, or local regression models to make slight adjustments to the data to correct obvious anomalies without disrupting the overall trend, retaining the signals brought by real events while reducing deviations caused by data noise. At the same time, early warning marking and manual review are performed. The system marks the data jump as "pending review" or "uncertain" and records detailed adjustment logs and event-driven scores to provide a basis for subsequent manual review.
[0050] When the event-driven score is less than or equal to the second score threshold (e.g., 0.3), it is considered low-scoring transition data. The data transition indicates a data collection or processing error, and strong cleaning (elimination or correction of abnormal data) is performed. Strong cleaning of low-scoring transition data can remove collection errors and invalid noise, preventing them from misleading model training and analysis. Specifically, strong cleaning can include applying a forced cleaning strategy to transition data segments, directly eliminating or correcting abnormal data through high-intensity data correction methods (such as re-filling with a global statistical model). Abnormal data deletion can also be performed. If the data is completely inconsistent with the historical benchmark, this portion of the data is eliminated to ensure overall data quality. Detailed records of the cleaning process, data changes before and after cleaning, and decision-making basis are kept to facilitate subsequent audits and model feedback optimization.
[0051] In summary, the embodiment of the present application provides a data cleaning method for an AI intelligent database, the method comprising: obtaining external event data and internal data of the AI intelligent database; obtaining the driving effect coefficient of the external event data on any breakpoint data of the internal data; obtaining the event-driven score of the external event data on any breakpoint data according to the driving effect coefficient; and performing adaptive data cleaning on the internal data of the AI intelligent database according to the event-driven score. The embodiment of the present application analyzes the external event-driven impact of data breakpoints to confirm the rationality of the internal data changes in the database, thereby dynamically identifying breakpoint data and adaptively cleaning the breakpoint data. It can effectively retain structural data changes with business significance, avoid information loss caused by incorrect cleaning, and significantly improve the data quality management capabilities and decision support levels of intelligent databases in highly dynamic environments.
[0052] Figure 2 This is a flow chart showing a method for obtaining a driving effect coefficient of external event data on any breakpoint data of internal data according to an exemplary embodiment. Figure 2 As shown, obtaining the driving effect coefficient of the external event data on any breakpoint data of the internal data may include the following steps: In step S201, the logic jump degree of the neighborhood segment data of any breakpoint data is obtained.
[0053] In this step, the logical jump degree of the neighborhood segment data of any breakpoint data is obtained. For example, the deviation degree of any breakpoint data of any type of data in the internal data of the database can be first obtained, then the number of data types in the internal data can be obtained, and then the correlation coefficient of the neighborhood segment data of any two types of data in the internal data can be obtained. Then, the average of the correlation coefficients of the neighborhood segment data of any two types of data in the internal data for all breakpoint data can be obtained. Finally, based on the deviation degree, the number of data types, the correlation coefficient, and the average, the logical jump degree of the neighborhood segment data of any breakpoint data can be obtained.
[0054] In step S202, a linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data is obtained.
[0055] In this step, the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data is obtained. For example, the heat index change sequence of the external event data can be obtained first, and then the data change sequence of the neighborhood segment data of any breakpoint data of any type of internal data can be obtained. Then, the Pearson correlation coefficient between the heat index change sequence and the data change sequence of the neighborhood segment data is obtained. Finally, based on the Pearson correlation coefficient, the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data can be obtained.
[0056] In step S203, a trigger effect coefficient of the external event data on the internal data is obtained.
[0057] In this step, the trigger effect coefficient of the external event data on the internal data is obtained. For example, the event coreness of any keyword in the external event data can be obtained first. Then, based on the event coreness, the semantic matching degree between the external event data and the AI intelligent database is obtained. Finally, based on the semantic matching degree, the trigger effect coefficient of the external event data on the internal data is obtained.
[0058] In step S204 , a driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained according to the logic jump degree, the linkage coefficient, and the trigger effect coefficient.
[0059] In this step, according to the logic transition , linkage coefficient , and the trigger effect coefficient , get external event data For the database The driving effect coefficient of the neighborhood data of any breakpoint i of the class data For example, the external event data a is The driving effect coefficient of the neighborhood data of any breakpoint i of the class data It can be obtained by the following formula: Formula 1 in, Indicates the The external event data and the Class data The linkage coefficient of the neighborhood segment data of the breakpoint, Indicates the The external event data is the first The trigger effect coefficient of class data, Indicates the first Class data The logical jump degree of the neighborhood segment data of a breakpoint, Represents an exponential function with the natural number e as its base.
[0060] Indicates the The external event data is the first Class data The linkage (regulation) of the neighborhood segment data of a breakpoint and the difference between the triggering effect coefficient of the external event data on this type of data. The smaller the difference, the more likely it is that the impact of the external event on this type of data matches the linkage change of the data, and the more likely it is that the data jump is driven by the external event.
[0061] To further confirm whether a data jump is caused by an external event, the external event's influence on the breakpoint data type should be synchronized with the actual data linkage changes. The greater the influence on the breakpoint data type, the stronger the linkage changes in the actual breakpoint data. Based on the external event's trigger effect coefficient on the breakpoint data item and the linkage coefficient on the breakpoint data change, the driving effect coefficient of the external event on the breakpoint data is obtained.
[0062] Figure 3 FIG. 1 is a flow chart showing a method for obtaining the logical jump degree of the neighborhood segment data of any breakpoint data according to an exemplary embodiment. Figure 3 As shown, obtaining the logical jump degree of the neighborhood segment data of any breakpoint data may include the following steps: In step S2011, the deviation of any breakpoint data of any type of data of the internal data is obtained.
[0063] For example, when a sudden change occurs in the data within the database, the stability of the data sequence will be destroyed, causing the cumulative deviation to deviate significantly from expectations. The CUSUM (Cumulative Sum) algorithm calculates the difference between each data point and a preset target value (such as a historical mean or expected value), and accumulates these differences in chronological order. When the accumulated deviation continues to deviate from the target value, the cumulative sum will gradually increase or decrease. When it exceeds a preset threshold (such as 0.7), an alarm is triggered and marked as a breakpoint, i.e., a jump point. The CUSUM breakpoint detection algorithm is used to obtain the jump points of each type of data sequence in the database in real time, and the deviation of each jump point can also be obtained. Since the process of using the CUSUM breakpoint detection algorithm to detect each jump point and the deviation of the jump point belongs to the existing technology, it will not be described here.
[0064] When data changes at a certain moment, the database's internal logical consistency is disrupted, indicating a possible data anomaly or structural change caused by an external event. By comparing current data with historical benchmarks and utilizing logical constraints, we can accurately capture data changes and determine the logical change degree for each data breakpoint.
[0065] The fitting curve of each type of data is drawn based on the least squares principle, and the data between the breakpoint and the minimum point before the breakpoint constitute the neighborhood segment data of the breakpoint.
[0066] Figure 4 FIG. 1 is a schematic diagram showing neighborhood segment data of a breakpoint data according to an exemplary embodiment. Figure 4 As shown, the data segment between the breakpoint data B and the minimum value A is the neighborhood segment data of the breakpoint data B.
[0067] In this step, the deviation of any breakpoint data i of any type of data o in the database is obtained. .
[0068] In step S2012, the number of data types of the internal data is obtained.
[0069] In this step, the number of data types of the internal data in the database is obtained .
[0070] In step S2013, the correlation coefficient between any two types of data in the internal data and the neighborhood segment data of any breakpoint data is obtained.
[0071] In this step, the correlation coefficient between the corresponding data of any two types of data o and j in the database internal data on the neighborhood segment data time series of any breakpoint data i is obtained. For example, the correlation coefficient The Pearson correlation coefficient can be the corresponding data in the time series of the neighborhood segments of any two data types o and j in the database internal data at any breakpoint data i. It should be understood that if the lengths of the neighborhood segments of any two data types o and j at any breakpoint data i are inconsistent, the smaller neighborhood segment data can be interpolated using data interpolation to ensure that the lengths of the two neighborhood segments are consistent.
[0072] In step S2014, the mean value of the correlation coefficient of any two types of data in the internal data in the neighborhood segment data of all breakpoint data is obtained.
[0073] In this step, the mean of the correlation coefficients of any two types of data o and j in the neighborhood of all breakpoint data is obtained. .
[0074] In step S2015, the logical jump degree of the neighborhood segment data of any breakpoint data is obtained according to the deviation, the number of data types, the correlation coefficient, and the mean.
[0075] In this step, according to the deviation , number of data types , correlation coefficient , and the mean , obtain the logical jump degree of the neighborhood segment data of any breakpoint data i of any type of data o in the internal data of the database For example, the logical jump degree of the neighborhood segment data of any breakpoint data i of any type of data o in the internal data of the database is It can be obtained by the following formula: Formula 2 Wherein, j represents any type of data in the database internal data, and N is not zero.
[0076] Indicates the first The mean of the difference in the correlation between the class data and each other type of data reflects the logical relationship between the internal data of the database. The larger the value, the more likely the logical relationship of the internal data of the database is to be destroyed.
[0077] Figure 5 This is a flow chart showing a method for obtaining a linkage coefficient between external event data and neighboring segment data of any breakpoint data according to an exemplary embodiment. Figure 5 As shown, obtaining the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data may include the following steps: In step S2021, a heat index change sequence of the external event data is obtained.
[0078] In this step, the heat index change sequence of external event data a is obtained For example, the heat index of external event data a can be integrated, such as news dissemination volume, search index and social media discussion frequency, as the heat index of external event data a, and then combined with the event occurrence time parameter to obtain the heat index change sequence of external event data a. .
[0079] In step S2022, a data change sequence of the neighborhood segment data of any breakpoint data of any type of data of the internal data is obtained.
[0080] In this step, the data change sequence of the neighborhood segment data of any breakpoint data i of any type of data o in the database is obtained. For example, the data change sequence of the neighborhood data of any breakpoint data i of any type of data o in the database can be obtained according to the data point values of the neighborhood data of any breakpoint data i of any type of data o in the database, and the acquisition time corresponding to these data points. .
[0081] In step S2023, the Pearson correlation coefficient between the heat index change sequence and the data change sequence is obtained.
[0082] In this step, obtain the heat index change sequence and data change sequence Pearson correlation coefficient It should be understood that if the heat index change sequence and data change sequence If the lengths of the two change sequences are inconsistent, the data interpolation method can be used to interpolate the change sequence with the smaller length to ensure that the lengths of the two change sequences are consistent. In step S2024, the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data is obtained based on the Pearson correlation coefficient.
[0083] In this step, according to the Pearson correlation coefficient , obtain the linkage coefficient between the external event data a and the neighborhood segment data of any breakpoint data i of any type of data o in the database internal data For example, the linkage coefficient between the external event data a and the neighborhood data of any breakpoint data i of any type of data o in the database internal data is It can be obtained by the following formula: Formula 3 in, For normalization processing.
[0084] Through time series analysis, the correlation between external events and internal data changes in the database is quantified, and the driving force of external events on current data jumps is quantified.
[0085] Figure 6 FIG. 1 is a flow chart showing a method for obtaining a trigger effect coefficient of external event data on internal data according to an exemplary embodiment. Figure 6 As shown, obtaining the trigger effect coefficient of the external event data on the internal data may include the following steps: In step S2031, the event coreness of any keyword in the external event data is obtained.
[0086] In this step, the event coreness of any keyword in the external event data is obtained. For example, all keywords in the external event data can be clustered using the point mutual information values between the keywords, and then the number of keywords in the cluster where any keyword in the external event data is located is obtained. Then, the total number of keywords in the external event data is obtained, and then the point mutual information value between any two keywords in the external event data is obtained. Finally, based on the number of keywords, the total number, and the point mutual information value, the event coreness of any keyword in the external event data is obtained. In the same way, the event coreness of any keyword in the internal data of the database can also be obtained.
[0087] In step S2032, based on the event coreness, the semantic matching degree between the external event data and the AI intelligent database is obtained.
[0088] In this step, the semantic matching degree between the external event data and the AI intelligent database is obtained based on the event coreness. For example, the number of intersections of the external event data and the keywords in the database internal data can be obtained first, and then the number of unions of the keywords in the external event data and the database internal data can be obtained. Then, the first event coreness of any keyword in the external event data can be obtained, and then the second event coreness of any corresponding keyword in the database internal data can be obtained. Finally, the semantic matching degree between the external event data and the AI intelligent database can be obtained based on the number of intersections, the number of unions, the first event coreness, and the second event coreness.
[0089] In step S2033, based on the semantic matching degree, a triggering effect coefficient of the external event data on the internal data is obtained.
[0090] In this step, the trigger effect coefficient of the external event data on the internal data is obtained based on the semantic matching degree. For example, the type semantic matching degree of the external event data and any type of data in the database can be obtained first. Then, based on the semantic matching degree between the external event data and the AI intelligent database, as well as the type semantic matching degree, the trigger effect coefficient of the external event data on the internal database data can be obtained.
[0091] Figure 7 FIG. 1 is a flow chart showing a method for obtaining the event coreness of any keyword in external event data according to an exemplary embodiment. Figure 7 As shown, obtaining the event coreness of any keyword in the external event data may include the following steps: In step S20311, all keywords in the external event data are clustered based on the point mutual information values between the keywords.
[0092] In this step, all keywords in the external event data are clustered using the pointwise mutual information (PMI) between keywords. For example, NLP (Natural Language Processing) techniques can be used to extract keywords from external data such as news and announcements. Furthermore, DBSCAN clustering is performed on all keywords in each external event using the PMI (Pointwise Mutual Information) between keywords.
[0093] In order to analyze the impact of external events on data jumps, it is necessary to confirm the correlation between the external events and the content of the AI intelligent database. If the external events coincide with the research content of the database, the external events will have a certain impact on the data changes in the database, which can be obtained through the degree of matching between the external events and the keywords in the database.
[0094] In step S20312, the number of keywords in the cluster where any keyword in the external event data is located is obtained.
[0095] In this step, the number of keywords in the cluster where any keyword l in the external event data a is located is obtained .
[0096] In step S20313, the total number of keywords in the external event data is obtained.
[0097] In this step, the total number of keywords in the external event data a is obtained .
[0098] In step S20314, the point mutual information value between any two keywords in the external event data is obtained.
[0099] In this step, the point mutual information value between any keyword l and other keywords k in the cluster where any keyword l in the external event data a is located is obtained. For example, the point mutual information value between any keyword l and other keywords k in the cluster where any keyword l in the external event data a is located is It can represent the association relationship between two keywords l and k.
[0100] In step S20315, the event coreness of any keyword in the external event data is obtained according to the number of keywords, the total number, and the point mutual information value.
[0101] In this step, based on the number of keywords , total quantity , and the point mutual information value , obtain the event coreness of any keyword l in the external event data a For example, the event coreness of any keyword l in the external event data a is It can be obtained by the following formula: Formula 4 in, is not zero, Not zero.
[0102] Indicates the The first external event The ratio of the number of keywords in the cluster where the keyword is located to the total number of keywords represents the importance of the cluster where the keyword is located in the external event; Indicates the The first external event The mean of the correlation between a keyword and other keywords in its cluster. The larger the value of this formula is, the closer the relationship between the keyword and all keywords is, and the higher its coreness is.
[0103] Figure 8 This is a flow chart showing a method for obtaining semantic matching between external event data and an AI intelligent database based on event coreness according to an exemplary embodiment. Figure 8 As shown, obtaining the semantic matching degree between the external event data and the AI intelligent database according to the event coreness may include the following steps: In step S20321, the number of intersections between the external event data and the keywords of the internal data is obtained.
[0104] In this step, the number of intersections between the external event data a and the keywords of the database internal data is obtained. .
[0105] In step S20322, the union number of the keywords of the external event data and the internal data is obtained.
[0106] In this step, the union number of keywords of external event data a and database internal data is obtained. .
[0107] In step S20323, the first event coreness of any keyword in the external event data is obtained.
[0108] In this step, the first event coreness of any keyword m in the external event data a is obtained. .
[0109] In step S20324, the second event coreness of any keyword in the internal data is obtained.
[0110] In this step, the second event coreness of any keyword m in the database internal data is obtained. .
[0111] In step S20325, the semantic matching degree between the external event data and the AI intelligent database is obtained based on the number of intersections, the number of unions, the first event coreness, and the second event coreness.
[0112] In this step, according to the number of intersections , number of unions , first event coreness , and the second event coreness , obtain the semantic matching degree between external event data a and AI intelligent database For example, the semantic matching degree between external event data a and AI intelligent database It can be obtained by the following formula: Formula 5 in, is not zero, is not zero, Represents an exponential function with the natural number e as its base.
[0113] Indicates the The ratio of repeated keywords to all keywords between an external event and the AI intelligent database represents the matching degree between the external event a and the database; Indicates the The average difference between the coreness of an external event and all corresponding keywords in the database represents the impact of external event a on changes in the database's internal data. The greater the match between external event a and the database, the greater the impact of external event a on changes in the database's internal data, indicating a greater semantic match between external event a and the AI intelligent database.
[0114] The degree of match between external events and keywords in the database can be used to determine the influence of external events on data changes in the database. The closer the core degree of keywords in external events and the database, the stronger the semantic correspondence.
[0115] Figure 9 This is a flow chart showing a method for obtaining a trigger effect coefficient of external event data on internal data based on semantic matching according to an exemplary embodiment. Figure 9 As shown, obtaining the trigger effect coefficient of the external event data on the internal data according to the semantic matching degree may include the following steps: In step S20331, the type semantic matching degree between the external event data and any type of data of the internal data is obtained.
[0116] In this step, the semantic matching degree between the external event data a and any type of data o in the database is obtained in the same way as the semantic matching degree between the external event data and the AI intelligent database. .
[0117] In step S20332, based on the semantic matching degree between the external event data and the AI intelligent database, and the type semantic matching degree, the triggering effect coefficient of the external event data on the internal data is obtained.
[0118] In this step, the semantic matching between external event data and AI intelligent database is used to , and type semantic matching , obtain the trigger effect coefficient of external event data a on the oth type of data in the database internal data For example, the trigger effect coefficient of external event data a on the oth type of data in the database is It can be obtained by the following formula: Formula 6 in, Indicates the number of data types in the AI intelligent database. Not zero; Indicates the External events and AI intelligent database The type semantic matching degree between class data, represents normalization processing, Not zero.
[0119] Indicates the An external event affects the The semantic matching degree of the class data is The difference in semantic matching degree between an external event and the entire database represents the impact of external event a on the change of data o of this type.
[0120] By semantically matching the data item that caused the data jump with the external event, the system determines whether the jump caused by the event in this type of data has a driving effect. At the same time, the system compares the semantic matching degree of other types of data with the same external event. If the semantic matching of this type of data is significantly higher than that of other data categories, it indicates that the external event has a strong triggering effect on this type of data.
[0121] Figure 10 This is a flow chart of a method for obtaining an event-driven score of any breakpoint data based on external event data according to a driving effect coefficient according to an exemplary embodiment. Figure 10 As shown, obtaining the event-driven score of the external event data on any breakpoint data according to the driving effect coefficient may include the following steps: In step S301, a set data quantity of neighborhood segment data of any breakpoint data is obtained.
[0122] In this step, the set data quantity of the neighborhood segment data of any breakpoint data i of the oth type data in the database is obtained. For example, the number of set data It can be 5.
[0123] In step S302, the logic jump degree of any two adjacent neighborhood segment data in the set number of neighborhood segment data is obtained.
[0124] In this step, the logical jump degree of any two adjacent neighborhood segment data q+1 and q in the neighborhood segment data of the set data quantity is obtained. and .
[0125] In step S303, the logic recovery degree of any breakpoint data of any type of data in the internal data is obtained according to the set data quantity and the logic jump degree of any two adjacent neighborhood segments of data.
[0126] In this step, according to the set data quantity , and the logical jump degree of any two adjacent neighborhood segment data q+1 and q and , obtain the logical recovery degree of any breakpoint data i of any type of data o in the internal data of the database For example, the logical recovery degree of any breakpoint data i of any type of data o in the database internal data is It can be obtained by the following formula: Formula 7 in, -1 is not zero, For normalization processing.
[0127] Indicates the internal data of the database Class data The average of the differences in the logical transition degrees of the data points in the neighborhood of the breakpoint data i. The larger the value, the greater the logical transition recovery degree of the breakpoint data i.
[0128] Abnormal noise is a short-term deviation. If the logical consistency of other data in the neighborhood is high, there will be no data logic recovery process. If it is noise, the smaller the value of this formula will be; if it is driven by external events, there will be a gradual recovery process of data logic, and the value of this formula will be larger.
[0129] When data jumps within a database are driven by external events, this usually means that the market or business environment has undergone substantial changes, thereby changing the data generation mechanism and the internal relationships between various data items. As the external influence gradually transmits, the internal data logic will also adjust and tend to a new stable state. If the data jump is only caused by random abnormal noise, the internal logical relationships of the data items have not actually changed in essence, and the jump is only a brief deviation. In this case, the data will still maintain its original consistency after regression, and no obvious logical adjustment or recovery process is required. The internal logical recovery degree of the breakpoint data can be obtained through the logical jump of the data in the breakpoint neighborhood segment.
[0130] In step S304, an event driving score of the external event data on any breakpoint data is obtained according to the driving effect coefficient and the logic recovery degree.
[0131] In this step, according to the driving effect coefficient , and logical recovery , obtain the event-driven scoring of any breakpoint data i of any type of data o in the database internal data by external event data a For example, the external event data a is used to score the event-driven data i of any breakpoint data o of any type of data in the database internal data. It can be obtained by the following formula: Formula 8 In summary, the embodiment of the present application provides a data cleaning method for an AI intelligent database, the method comprising: obtaining external event data and internal data of the AI intelligent database; obtaining the driving effect coefficient of the external event data on any breakpoint data of the internal data; obtaining the event-driven score of the external event data on any breakpoint data according to the driving effect coefficient; and performing adaptive data cleaning on the internal data of the AI intelligent database according to the event-driven score. The embodiment of the present application analyzes the external event-driven impact of data breakpoints to confirm the rationality of the internal data changes in the database, thereby dynamically identifying breakpoint data and adaptively cleaning the breakpoint data. It can effectively retain structural data changes with business significance, avoid information loss caused by incorrect cleaning, and significantly improve the data quality management capabilities and decision support levels of intelligent databases in highly dynamic environments.
[0132] The present application also provides a computer-readable storage medium having computer program instructions stored thereon. When the program instructions are executed by a processor, the steps of the data cleaning method for the AI intelligent database provided in the present application are implemented.
[0133] Figure 11 FIG. 1 is a block diagram of a data cleaning system for an AI intelligent database according to an exemplary embodiment. Figure 11 As shown, an embodiment of the present application provides a data cleaning system 1100 for an AI intelligent database, including a database platform 1200.
[0134] Figure 12 1 is a block diagram of a database platform according to an exemplary embodiment. For example, the database platform 1200 can be provided as a server. Figure 12 The database platform 1200 includes a processing component 1222, which further includes one or more processors and a memory resource represented by a memory 1232 for storing instructions executable by the processing component 1222, such as an application. The application stored in the memory 1232 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1222 is configured to execute instructions to perform the data cleaning method of the AI intelligent database.
[0135] Database platform 1200 may also include a power component 1226 configured to perform power management of database platform 1200, a communication component 1250 configured to connect database platform 1200 to a network, and an input / output interface 1258. Database platform 1200 may operate based on an operating system stored in memory 1232.
[0136] In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a programmable electronic device, and the computer program has a code portion for executing the above-mentioned data cleaning method for the AI intelligent database when executed by the programmable electronic device.
[0137] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the scope of the present application, and such modifications and improvements are all within the scope of protection of the present application.
Claims
1. A data cleaning method for an AI intelligent database, characterized in that: The method comprises: Acquiring external event data and internal data of the AI intelligent database; Obtaining a driving effect coefficient of the external event data on any breakpoint data of the internal data; Obtaining, according to the driving effect coefficient, an event-driven score of the external event data on any breakpoint data; Based on the event-driven scoring, adaptive data cleaning is performed on the internal data of the AI intelligent database.
2. The data cleaning method of the AI intelligent database according to claim 1, characterized in that: The obtaining of the driving effect coefficient of the external event data on any breakpoint data of the internal data includes: Obtaining the logical jump degree of the neighborhood segment data of any breakpoint data; Obtaining a linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data; Obtaining a trigger effect coefficient of the external event data on the internal data; A driving effect coefficient of the external event data on any breakpoint data of the internal data is obtained according to the logic jump degree, the linkage coefficient, and the trigger effect coefficient.
3. The data cleaning method of the AI intelligent database according to claim 2, characterized in that: The obtaining of the logical jump degree of the neighborhood segment data of any breakpoint data includes: Obtaining the deviation of any breakpoint data of any type of data of the internal data; Obtain the data type and quantity of the internal data; Obtaining the correlation coefficient between any two types of data in the internal data and the neighborhood segment data of any breakpoint data; Obtain the mean value of the correlation coefficient of any two types of data in the internal data in the neighborhood segment data of all breakpoint data; The logical jump degree of the neighborhood segment data of any breakpoint data is obtained according to the deviation, the number of data types, the correlation coefficient, and the mean.
4. The data cleaning method of the AI intelligent database according to claim 2, characterized in that: The obtaining of the linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data includes: Obtaining a heat index change sequence of the external event data; Obtaining a data change sequence of neighborhood segment data of any breakpoint data of any type of data of the internal data; Obtaining the Pearson correlation coefficient between the heat index change sequence and the data change sequence; According to the Pearson correlation coefficient, a linkage coefficient between the external event data and the neighborhood segment data of any breakpoint data is obtained.
5. The data cleaning method of the AI intelligent database according to claim 2, characterized in that: The obtaining of the trigger effect coefficient of the external event data on the internal data includes: Obtaining the event coreness of any keyword in the external event data; According to the event coreness, obtaining the semantic matching degree between the external event data and the AI intelligent database; According to the semantic matching degree, a triggering effect coefficient of the external event data on the internal data is obtained.
6. The data cleaning method of the AI intelligent database according to claim 5, characterized in that: The obtaining of the event coreness of any keyword in the external event data includes: Clustering all keywords in the external event data based on point mutual information values between the keywords; Obtain the number of keywords in the cluster where any keyword in the external event data is located; Obtaining the total number of keywords in the external event data; Obtaining a point mutual information value between any two keywords in the external event data; The event coreness of any keyword in the external event data is obtained according to the number of keywords, the total number, and the point mutual information value.
7. The data cleaning method of the AI intelligent database according to claim 5, characterized in that: The obtaining of the semantic matching degree between the external event data and the AI intelligent database according to the event coreness includes: Obtaining the number of intersections between the external event data and the keywords of the internal data; Obtaining a union number of keywords of the external event data and the internal data; Obtaining a first event coreness of any keyword in the external event data; Obtaining a second event coreness of any keyword in the internal data; According to the number of intersections, the number of unions, the first event coreness, and the second event coreness, a semantic matching degree between the external event data and the AI intelligent database is obtained.
8. The data cleaning method of the AI intelligent database according to claim 5, characterized in that: The obtaining, based on the semantic matching degree, a triggering effect coefficient of the external event data on the internal data includes: Obtaining a type semantic matching degree between the external event data and any type of data of the internal data; According to the semantic matching degree between the external event data and the AI intelligent database, and the type semantic matching degree, a triggering effect coefficient of the external event data on the internal data is obtained.
9. The data cleaning method of the AI intelligent database according to claim 1, characterized in that: The step of obtaining, based on the driving effect coefficient, an event-driven score of the external event data on any breakpoint data includes: Obtaining a set number of data of the neighborhood segment data of any breakpoint data; Obtaining the logic jump degree of any two adjacent neighborhood segment data in the set number of neighborhood segment data; Obtaining the logic recovery degree of any breakpoint data of any type of data in the internal data according to the set data quantity and the logic jump degree of any two adjacent neighborhood segment data; An event driving score of the external event data for any breakpoint data is obtained according to the driving effect coefficient and the logic recovery degree.
10. A data cleaning system for an AI intelligent database, characterized in that: The system includes a database platform, and the database platform includes: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Data cleaning method, server and computer readable storage medium
CN110046151A
Dynamic threshold anomaly detection method and system, storage medium and intelligent equipment
CN110807024A
Intelligent scene big data cleaning method based on AI prediction and intelligent scene system
CN114691664A
Intelligent data cleaning method in big data environment
CN119166619A
Service knowledge base management system based on AI large model
CN119962645A