Abnormal monitoring method and system for hierarchical data of business domain under data lake
By acquiring the characteristic factors of the data lake's business domains for hierarchical matching and flow verification, the problem of inaccurate hierarchical anomaly monitoring in the data lake is solved. This enables accurate identification and dynamic accumulation of hierarchical data of business domains under the data lake, improving the accuracy and reliability of anomaly monitoring.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGXI GUIGUAN ELECTRIC POWER CO LTD
- Filing Date
- 2026-01-27
- Publication Date
- 2026-05-08
AI Technical Summary
Existing technologies lack specific methods for monitoring anomalies in layered data within data lake business domains, making it difficult to effectively identify deep-seated anomalies such as form matching and inter-layer flow. This results in untimely and inaccurate anomaly detection, impacting analytical decision-making and business application effectiveness.
By acquiring the business domain characteristic factors of the target business, data is hierarchically matched, and formal verification and inter-layer flow verification are performed. The formal matching degree and flow matching degree are calculated, and a comprehensive anomaly degree verification is performed by combining the historical abnormal business hierarchical set, and anomaly hierarchical is identified and recorded.
It enables precise location of hierarchical data in business domains under the data lake and identification of deep-level anomalies, improving the accuracy and reliability of anomaly monitoring and ensuring the consistency of data quality and the effectiveness of business applications.
Smart Images

Figure CN121996508A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data management technology, specifically to a method and system for anomaly monitoring of hierarchical data in business domains under a data lake. Background Technology
[0002] With the explosive growth of data volume and the deepening of enterprise digital transformation, data lakes, as platforms for centralized storage and management of massive amounts of heterogeneous data, have been widely used across various industries. Business domain layering is an important way for data lakes to organize and manage data. By layering data according to business logic and data characteristics, data availability and management efficiency can be improved.
[0003] However, in practical applications, data flow between different business layers under a data lake may be subject to disordered dependencies and inconsistent data quality; some business layers may be confused due to similar names and structures; at the same time, with the dynamic changes in business and the continuous access of data, new abnormal patterns are constantly emerging.
[0004] In existing technologies, anomaly monitoring of data lake data mainly focuses on basic quality dimensions such as the integrity and consistency of the data itself. There is a lack of specialized monitoring methods for the layered characteristics of business domains. It is difficult to effectively identify deep-seated anomalies in layered data such as form matching and inter-layer flow, resulting in untimely and inaccurate anomaly detection, which in turn affects the effectiveness of analysis, decision-making and business applications based on data lake data. Summary of the Invention
[0005] This application provides an anomaly monitoring method and system for hierarchical data in business domains under a data lake. This solves the technical problem in the prior art that there is a lack of specific methods for monitoring anomalies in hierarchical data in business domains of a data lake, which makes it difficult to effectively identify deep-seated anomalies such as form matching and inter-layer flow, resulting in untimely anomaly detection.
[0006] The technical solution to the above-mentioned technical problems in this application is as follows: Firstly, this application provides a method for anomaly monitoring of hierarchical data in a business domain under a data lake, the method comprising: Obtain the business domain feature factors of the target business, and perform data layer matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers; Perform formal validation on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as the target business layers, and add the matching business layers with a formal matching degree less than the formal matching degree threshold as easily confused business layers to the historical abnormal business layer set. Randomly select a first business layer from the target business layers, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of multiple first business layers; Based on the form matching degree and the flow matching degree, a comprehensive anomaly degree is obtained. Combined with the historical abnormal service layer set, the first service layer with a comprehensive anomaly degree greater than the comprehensive anomaly threshold is checked for similarity, anomaly similarity is obtained, anomaly is judged, and the first service layer with anomaly judgment result is added to the historical abnormal service layer set.
[0007] Secondly, this application provides an anomaly monitoring system for hierarchical data in business domains under a data lake, including: The information acquisition module is used to acquire the business domain feature factors of the target business, and to perform data layer matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers; The formal verification module is used to perform formal verification on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as the target business layers. The matching business layers with a formal matching degree less than the formal matching degree threshold are designated as easily confused business layers and added to the historical abnormal business layer set. The matching degree calculation module is used to randomly select a first business layer from the target business layer, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of multiple first business layers. The similarity acquisition module is used to obtain a comprehensive anomaly degree based on the form matching degree and the flow matching degree, and to perform similarity verification on the first business layer whose comprehensive anomaly degree is greater than the comprehensive anomaly threshold in combination with the historical anomaly business layer set, to obtain anomaly similarity, to perform anomaly judgment, and to add the first business layer whose anomaly judgment result is abnormal to the historical anomaly business layer set.
[0008] This application provides one or more technical solutions, which have at least the following technical effects or advantages: This application provides a method and system for anomaly monitoring of business domain layered data under a data lake. First, it acquires the business domain characteristic factors of the target business and performs data layer matching on the target business area in the data lake, obtaining multiple matching business layers. This achieves accurate positioning of the data lake's layered data from a business perspective. Second, it performs formal verification on the matching business layers to obtain the formal matching degree, and uses this to filter out the target business layer and easily confused business layers. The easily confused business layers are added to the historical abnormal business layer set, effectively solving the layer confusion problem caused by similar naming and structure. Subsequently, it randomly selects the first business layer from the target business layers for inter-layer flow verification, obtaining the flow dependency coefficient. Combined with data quality parameters, it obtains the flow matching degree, considering the data dependency relationship and update synergy between business layers, and can promptly detect abnormal situations of disordered inter-layer flow. Finally, based on the form matching degree and flow matching degree, the comprehensive anomaly degree is calculated. The first business layer with the comprehensive anomaly degree exceeding the threshold is combined with the historical abnormal business layer set for similarity verification, and the anomaly similarity is obtained and anomaly is judged. The layers judged as abnormal are added to the historical abnormal business layer set, which realizes accurate identification and dynamic accumulation of deep-level anomalies. It effectively solves the problem of untimely and inaccurate anomaly detection in the existing technology and improves the analysis, decision-making and business application effects based on data lake data.
[0009] Through the above technical solutions, this application can comprehensively monitor the business domain layered data under the data lake from multiple dimensions such as business domain characteristics, form matching, inter-layer flow dependency and historical anomaly patterns. Through multi-stage verification and dynamic learning mechanisms, it continuously optimizes the anomaly identification capability to ensure the accuracy, consistency and reliability of the business domain layered data. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 This is a flowchart illustrating an anomaly monitoring method for hierarchical data in a data lake under a business domain, provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of an anomaly monitoring system for hierarchical data in a business domain under a data lake, provided in an embodiment of this application.
[0012] The components represented by each number in the attached diagram are explained below: Information collection module 11, formal verification module 12, matching degree calculation module 13, similarity acquisition module 14. Detailed Implementation
[0013] This application provides a method and system for monitoring anomalies in layered data of business domains under a data lake. This addresses the technical problem that existing technologies lack specific methods for monitoring anomalies in layered data of business domains under data lakes, making it difficult to effectively identify deep-seated anomalies such as form matching and inter-layer flow, resulting in untimely anomaly detection.
[0014] Example 1, as Figure 1 As shown in the figure, this application embodiment provides an anomaly monitoring method for hierarchical data of business domains under a data lake, including: S10: Obtain the business domain feature factors of the target business, and perform data layer matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers; In this embodiment of the application, the business domain characteristic factor of the target business is used to describe the set of indicators of the core attributes and data requirements of the target business. Specifically, it may include business domain name, core business process identifier, data topic tag, data sensitivity level, data format requirements, data update cycle characteristics, etc.
[0015] Furthermore, based on business domain feature factors, multiple matching business layers are obtained by performing hierarchical retrieval and matching within the target business area in the data lake.
[0016] Specifically, step S10 in the method includes: Obtain the business domain characteristic factors of the target business, wherein the business domain characteristic factors include key attributes of master data, business process identifiers, and control levels; Based on the business domain feature factors, data stratification matching is performed on the target business areas in the data lake to obtain multiple matching business stratifications, wherein the target business areas are obtained based on the master data key attributes.
[0017] In this embodiment of the application, firstly, the business domain characteristic factors of the target business are obtained, wherein the key attributes of the master data include the unique identifier, classification code, basic attributes and other information of the core entity of the target business; the business process identifier is used to distinguish the data flow nodes of different business links; and the control level is determined according to the importance and sensitivity of the data, such as top secret, confidential, ordinary and other levels.
[0018] Secondly, based on business domain characteristic factors, the target business area directly related to the target business is located in the data lake according to the key attributes of the master data. For example, the data storage area containing a specific customer ID is selected from the financial data lake. Then, within the target business area, the data is scanned and the attributes are compared layer by layer according to the business process identification and control level. The hierarchical data units that match the indicators in the business domain characteristic factors to a preset ratio are determined as matching business layers, such as a preset ratio of 80%, thereby obtaining multiple matching business layers.
[0019] S20: Perform formal validation on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as target business layers, and add the matching business layers with a formal matching degree less than the formal matching degree threshold as easily confused business layers to the historical abnormal business layer set. In this embodiment of the application, formal verification is a compliance check on the formal aspects of matching business layers, such as data organization structure, naming conventions, and metadata description.
[0020] The formal matching degree is obtained by performing formal validation on multiple matching business layers. That is, different weight coefficients are set for each validation dimension, and the sum of the products of each dimension's score and its corresponding weight is taken as the formal matching degree of that matching business layer.
[0021] Furthermore, the form matching threshold can be set according to business needs. By comparing the form matching degree with the form matching threshold, the target business layer and the easily confused business layer can be filtered and obtained.
[0022] This involves performing formal validation on multiple matching business layers to obtain the formal matching degree, including: A formal validation is performed on multiple matching service layers, wherein the formal validation includes data formal validation and content formal validation; Based on the data format verification results and the content format verification results, obtain the data matching degree and the content matching degree; The data matching degree and the content matching degree are weighted and calculated to obtain the form matching degree.
[0023] In this embodiment, firstly, data format verification and content format verification are performed on multiple matching business layers, that is, the format requirements and content rationality of the data are verified. Data format verification mainly checks whether the physical storage format, field naming rules, data type definitions, table structure specifications, etc. of the layered data conform to the preset format standards of the target business. Content format verification focuses on the consistency of description at the metadata level, including whether metadata items such as the business meaning description of the layered data, source system identifier, and data owner information are complete and conform to the specifications.
[0024] Secondly, the data matching degree is calculated based on the data format verification results. First, different weights are assigned to each check item in the data format verification, such as a naming rule weight of 0.4, a data type weight of 0.3, and a table structure weight of 0.3. Each matching business is layered, and each check item is scored according to its degree of compliance, such as 100 points for complete compliance, 50 points for partial compliance, and 0 points for non-compliance. The scores for each item are multiplied by their corresponding weights, summed, and then divided by the maximum score of 100 to obtain the data matching degree.
[0025] Similarly, content format verification also uses a similar weighting system, such as a business meaning description weight of 0.5, a source system identifier weight of 0.3, and an update frequency label weight of 0.2. The completeness and standardization of each description of the metadata are scored and calculated to obtain the content matching degree.
[0026] Finally, the data matching degree and content matching degree are weighted and calculated to obtain the formal matching degree. For example, if the weight of data matching degree is set to 0.6 and the weight of content matching degree is 0.4, and the data matching degree of a certain matching business layer is 85% and the content matching degree is 90%, then its formal matching degree is 85%×0.6+90%×0.4=87%.
[0027] Furthermore, matching business layers with a form matching degree greater than or equal to the form matching degree threshold are selected as target business layers, while matching business layers with a form matching degree less than the form matching degree threshold are selected as easily confused business layers and added to the historical abnormal business layer set, including: Obtain a form matching degree threshold, wherein the form matching degree threshold is obtained based on the mapping and matching of business domain feature factors; Filter out matching business layers with a form matching degree greater than or equal to the form matching degree threshold, and use them as target business layers; The matching business layers with a form matching degree less than the form matching degree threshold are classified as easily confused business layers and added to the historical abnormal business layer set.
[0028] In this embodiment, the formal matching threshold is first obtained through dynamic mapping and matching based on business domain feature factors. Specifically, a mapping relationship library between business domain feature factors and formal matching thresholds is pre-built. For example, for business domains with extremely high data standardization requirements, when the control level of the business domain feature factors is "top secret," the mapped formal matching threshold is 95%; while for ordinary e-commerce businesses, when the control level is "ordinary," the threshold may be set to 80%. By inputting the obtained business domain feature factors of the target business into this mapping relationship library, the corresponding formal matching threshold can be automatically matched and obtained.
[0029] Secondly, the form matching degree of each matching business layer is compared with the form matching degree threshold dynamically obtained above. Matching business layers with a form matching degree greater than or equal to the threshold are selected and determined as target business layers. The target business layers have basically met the requirements of the target business in terms of form.
[0030] Finally, for matching business layers with a form matching degree less than the form matching degree threshold, they are marked as easily confused business layers. These layers may be easily confused with the target business layer due to reasons such as non-standard naming, missing metadata descriptions, or non-standard formats. They are added to the historical abnormal business layer set.
[0031] Furthermore, the historical abnormal business layer set is a collection of business layer data that has experienced abnormalities in the past. Its data structure includes information such as the unique identifier of the abnormal layer, the time of the abnormality, the type of abnormality, the abnormality characteristic parameters, and the abnormality handling records.
[0032] S30: Randomly select a first business layer from the target business layers, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of the multiple first business layers; In this embodiment, inter-layer flow verification is a systematic check of the rationality, timeliness, and completeness of data flow between business layers. It can detect inter-layer anomalies caused by asynchronous data updates, broken dependencies, etc. A first business layer is randomly selected from the target business layers. Then, using this first business layer as the core, its upstream dependent layers and downstream associated layers are identified, and an inter-layer data flow path is constructed.
[0033] Furthermore, based on the constructed inter-layer data flow paths, a flow dependency coefficient is calculated for each path. The flow dependency coefficient is a quantitative indicator that measures the tightness of data dependence between upstream and downstream layers and the degree of compliance with flow rules. Next, based on the flow dependency coefficient and data quality parameters of multiple first-level business layers, the flow matching degree is obtained. Specifically, step S30 in the method includes: Randomly select a first service layer from among the multiple target service layers; Perform inter-layer flow verification on the first business layer to obtain the proportion of data from the source layer and the update frequency; Based on the proportion and update frequency of the source layer data, calculate and obtain the data dependency coefficient and update dependency coefficient; The data dependency coefficient and the update dependency coefficient are weighted and calculated to obtain multiple flow dependency coefficients of the first business layer; Obtain the data quality parameters of the first service layer; The flow dependency coefficient is corrected and calculated based on the data quality parameters to obtain the flow matching degree of the first business layer, wherein the flow matching degree is the average of the flow matching degrees of the first business layer and multiple business layers.
[0034] In this embodiment of the application, firstly, one of the multiple target business layers is selected as the first business layer by random sampling. For example, if the target business layer includes a user registration information layer, a product browsing behavior layer, an order payment data layer, etc., then the product browsing behavior layer is randomly selected as the first business layer to be verified.
[0035] Secondly, perform inter-layer flow verification on the first business layer to obtain the proportion of source layer data and update frequency. The proportion of source layer data refers to the percentage of data in the first business layer that originates from its upstream dependent layers. For example, data in the product browsing behavior layer may partially come from the user login status layer and the product information base layer. If 80% of the total data in this layer comes from the user login status layer and 20% comes from the product information base layer, then the source layer data proportions for the two upstream layers are 80% and 20%, respectively.
[0036] Update frequency refers to whether the data update time interval between the first business layer and its upstream dependent layers conforms to the preset coordination rules. For example, if the user login status layer is preset to update once every 5 minutes, the product browsing behavior layer should complete the update within 1 minute after that. If it is actually detected that the product browsing behavior layer updates 3 minutes after the user login status layer updates, then the update frequency does not conform to the rules.
[0037] Next, based on the proportion of data from the source layer and the update frequency, the data dependency coefficient and update dependency coefficient are calculated respectively. The data dependency coefficient is determined by the sum of the products of the proportion of data from the source layer and the preset weights. The preset weights reflect the importance of each upstream layer to the first business layer. For example, if the weight of the user login status layer is 0.7 and the weight of the product information base layer is 0.3, then the data dependency coefficient is 80% × 0.7 + 20% × 0.3 = 62%. The update dependency coefficient is scored according to the degree of compliance with the update frequency. Full compliance is scored as 100%, and 20% is deducted for each minute of delay. In the example above, with a delay of 2 minutes, the update dependency coefficient is 100% - 2 × 20% = 60%.
[0038] Then, the data dependency coefficient and the update dependency coefficient are weighted and calculated. For example, the data dependency coefficient has a weight of 0.6 and the update dependency coefficient has a weight of 0.4, resulting in a flow dependency coefficient of 62%×0.6+60%×0.4=61.2%. If there are multiple upstream dependency layers in the first business layer, multiple flow dependency coefficients will be calculated accordingly.
[0039] Next, the data quality parameters of the first business layer are obtained. These parameters comprehensively consider the completeness, accuracy, consistency and timeliness of the data. For example, the data completeness rate is 95%, the accuracy rate is 98%, the consistency rate is 90%, and the timeliness is 85%. By weighted averaging, such as each indicator having a weight of 0.25, the data quality parameter is obtained as (95%+98%+90%+85%)×0.25=92%.
[0040] Finally, since the flow dependency coefficient only reflects the correlation and synchronicity of the flow, without considering the reliability of the flow data itself, a data quality parameter is introduced for correction. That is, the flow dependency coefficient is adjusted based on the data quality parameter to obtain the flow matching degree; the better the data quality, the greater the flow matching degree. The correction formula can be set as Flow Matching Degree = Flow Dependency Coefficient × Data Quality Parameter. For example, Flow Dependency Coefficient 61.2% × Data Quality Parameter 92% ≈ Flow Matching Degree 56.3%.
[0041] Furthermore, if there are multiple inter-layer flow paths in the first business layer, such as multiple upstream or downstream related layers, the flow matching degree of each path is calculated separately and the average value is taken as the final flow matching degree of the first business layer.
[0042] Specifically, the process of performing inter-layer flow verification on the first business layer to obtain the proportion and update frequency of data from the source layer includes: Perform inter-layer flow verification on the first business layer, obtain multiple data sources of the first business layer, and obtain the data proportion of the source layer data in the first business layer data as the source layer data proportion, wherein the data source includes multiple business layers; Obtain the update frequency of the first service layer, wherein the update frequency includes the update time frequency.
[0043] In this embodiment, firstly, inter-layer flow verification is performed on the first business layer. By parsing the metadata information and data lineage graph of this layer, all upstream data sources are identified, which are represented by multiple different business layers. Then, for each identified upstream business layer, the proportion of data it provides in the total data volume of the first business layer is calculated as the source layer data proportion.
[0044] Secondly, the update frequency of the first business layer is obtained. The update frequency refers to the time frequency of data updates for this business layer. Specifically, by reading the metadata management module or related scheduling logs of the data lake, the actual update time points within the most recent preset time period of the first business layer are obtained, such as the past 7 days. Then, the data update pattern is analyzed and determined, such as updating once per hour or once at 2:00 AM every day, which is used as the update time frequency of the first business layer.
[0045] Furthermore, based on the proportion and update frequency of the source layer data, the data dependency coefficient and update dependency coefficient are calculated and obtained, including: Obtain the proportion of source layer data in the first business layer as the data dependency coefficient; Obtain the update frequency of multiple data sources for the first business layer; Based on the update frequency of the first business layer and the average update frequency of multiple data sources, the update dependency coefficient is obtained.
[0046] In this embodiment of the application, firstly, the proportion of source layer data in the first business layer is directly used as the data dependency coefficient. That is, if the proportion of upstream source layer data is 80%, then the data dependency coefficient of the upstream layer on the first business layer is 80%.
[0047] Secondly, obtain multiple data sources for the first business layer, namely the update frequency of each upstream dependent layer, such as upstream layer A being updated every 10 minutes and upstream layer B being updated every 15 minutes.
[0048] Next, calculate the average update frequency of upstream data sources, such as (10+15) / 2 = 12.5 minutes / time. Then compare the update frequency of the first business layer itself with this average. The more synchronized the update frequency, the greater the proportion of data sources, and the larger the update dependency coefficient. For example, if the update frequency of the first business layer is 12 minutes / time, which is close to the average of 12.5 minutes / time, then the update dependency coefficient is relatively high; if the update frequency of the first business layer is 30 minutes / time, which is much higher than the average, then the update dependency coefficient is relatively low.
[0049] Specifically, a standard update frequency deviation range can be set, for example, the maximum allowable deviation is 20% of the average. When the actual update frequency is within this range, the update dependency coefficient is 100%. For every 10% deviation from the range, the update dependency coefficient decreases by 20%. For example, if the average is 12.5 minutes and the allowable deviation range is 10-15 minutes, if the update frequency of the first business layer is 16 minutes, exceeding the range by 1 minute, the excess ratio is (16-15) / 12.5=8%, then the update dependency coefficient is 100%-(8% / 10%)×20%=84%.
[0050] S40: Based on the form matching degree and the flow matching degree, obtain the comprehensive anomaly degree, and combine the historical abnormal service layer set to perform similarity verification on the first service layer whose comprehensive anomaly degree is greater than the comprehensive anomaly threshold, obtain the anomaly similarity, perform anomaly discrimination, and add the first service layer whose anomaly discrimination result is abnormal to the historical abnormal service layer set.
[0051] In this embodiment, the comprehensive anomaly degree is a quantitative assessment of the degree of anomaly in terms of both formal standardization and the rationality of inter-layer flow in business layering. The comprehensive anomaly degree is calculated using the form of "1-matching degree," meaning that the higher the matching degree, the lower the comprehensive anomaly degree.
[0052] Secondly, a comprehensive anomaly threshold is set, which can be dynamically adjusted based on factors such as the importance of the business domain, data sensitivity, and the frequency of historical anomalies. For the first business layer with a comprehensive anomaly score greater than the comprehensive anomaly threshold, a similarity check is performed using historical anomaly business layer sets to determine whether it belongs to a known type of anomaly or a newly emerging anomaly pattern.
[0053] Furthermore, the calculation of anomaly similarity is based on anomaly feature parameters of multiple dimensions, and then anomaly discrimination is performed. The first business layer with anomaly discrimination result is added to the historical anomaly business layer set.
[0054] Specifically, step S40 in the method includes: The form matching degree and the flow matching degree are weighted and calculated to obtain the comprehensive anomaly degree; Obtain the hierarchical set of historical abnormal business operations; When the overall anomaly degree of the first service layer is greater than the overall anomaly threshold, the first service layer is regarded as a preliminary anomaly layer; otherwise, the first service layer is randomly obtained again. Based on the business domain feature factors, flow dependency coefficients and form matching degree of the preliminary anomaly layering, similarity verification is performed on the historical anomaly business layering set to obtain multiple historical anomaly business layers whose similarity to the preliminary anomaly layering set is greater than the historical anomaly threshold, as multiple similar anomaly business layers. Based on the initial anomaly layering and the similarity of multiple historical anomalies in the multiple similar anomaly service layers, anomaly similarity is obtained and anomaly discrimination is performed; The preliminary anomaly layers that are identified as abnormal are added to the historical anomaly business layer set.
[0055] In this embodiment, the comprehensive anomaly degree is first obtained by the formula "Comprehensive Anomaly Degree = 1 - (Form Matching Degree × α + Flow Matching Degree × β)", where α and β are the weighting coefficients of the form matching degree and the flow matching degree, respectively, and α + β = 1. The value of the weighting coefficient is determined based on the degree of emphasis placed on the formal standardization and the rationality of inter-layer flow by the business domain. For example, for the business domain of financial risk control, which has extremely high requirements for the formal standardization of data, α can be 0.6 and β can be 0.4; while for real-time recommendation business, the timeliness and accuracy of data flow are more critical, so α can be 0.3 and β can be 0.7.
[0056] For example, if the form matching degree of a preliminary anomaly stratification is 85%, the flow matching degree is 60%, α=0.5, β=0.5, then the comprehensive anomaly degree = 1-(85%×0.5+60%×0.5)=1-0.725=27.5%.
[0057] Secondly, obtain the historical abnormal business layer set, which stores information on all business layers that have been judged to be abnormal in the past.
[0058] Next, the calculated comprehensive anomaly degree of the first business layer is compared with the preset comprehensive anomaly threshold. When the comprehensive anomaly degree of the first business layer is greater than the comprehensive anomaly threshold, it is marked as a preliminary anomaly layer; if the comprehensive anomaly degree is less than or equal to the comprehensive anomaly threshold, the first business layer is considered to meet the requirements in terms of both form and flow. At this time, a new first business layer needs to be randomly obtained from the target business layer, and S30 and subsequent steps are repeated until the target business layer is fully verified or the preset number of sampling verifications is reached.
[0059] Then, for the initial anomaly stratification, similarity verification is performed in the historical anomaly business stratification set based on its business domain characteristic factors, flow dependency coefficient and form matching degree.
[0060] Specifically, the process involves extracting the business domain feature factor vector, flow dependency coefficient vector, and formal matching degree value for the initial anomaly stratification. Simultaneously, vectors or values for the aforementioned three dimensions are extracted from the historical anomaly business stratification set for each historical anomaly business stratification. The similarity between the initial anomaly stratification and each historical anomaly business stratification on the business domain feature factor vector is calculated using the cosine similarity algorithm, and the similarity between the two on the flow dependency coefficient vector is calculated using the Pearson correlation coefficient. The absolute difference in formal matching degree is then normalized and used as the formal similarity. For example, formal similarity = 1 - |initial anomaly stratification formal matching degree - historical anomaly stratification formal matching degree|).
[0061] Next, a weighted average of the similarities across the three dimensions was calculated, with weights of 0.5 for business domain feature factor similarity, 0.3 for flow dependency coefficient similarity, and 0.2 for formal similarity, to obtain the historical anomaly similarity between the initial anomaly layer and the historical anomaly layer. All records in the historical anomaly business layer set were traversed, and historical anomaly business layers with a historical anomaly similarity greater than the historical anomaly threshold were selected as similar anomaly business layers to the initial anomaly layer.
[0062] Subsequently, based on the initial anomaly stratification and the historical anomaly similarities of multiple similar anomaly business stratifications, anomaly similarity is obtained and anomaly detection is performed. The anomaly similarity can be the average of the historical anomaly similarities of all similar anomaly business stratifications.
[0063] Furthermore, if the anomaly similarity is greater than or equal to the anomaly discrimination threshold, the preliminary anomaly stratification is determined to be highly similar to historically known anomaly patterns, and the anomaly discrimination result is anomaly; if the anomaly similarity is lower than the anomaly discrimination threshold, but there is at least one historical anomaly similarity of a similar anomaly business stratification that is higher than the historical anomaly threshold, it is determined to be a suspected anomaly; if there is no similar anomaly business stratification or all historical anomaly similarities are lower than the historical anomaly threshold, it is determined to be a new type of anomaly, and its discrimination result is also recorded as anomaly.
[0064] Finally, the preliminary anomaly layers that are identified as abnormal are added to the historical anomaly business layer set, including confirmed anomalies and new types of anomalies. Information such as the unique identifier of the anomaly layer, the time of occurrence, the type of anomaly, the characteristic parameters of the anomaly, and the anomaly handling records are then added. For suspected anomalies, a manual review process can be triggered, and a decision on whether to add them to the historical anomaly business layer set is made after manual confirmation.
[0065] Furthermore, based on the preliminary anomaly layering and the historical anomaly similarity of multiple similar anomaly service layers, anomaly similarity is obtained, and anomaly discrimination is performed, including: Obtain the similarity of multiple historical anomalies between the initial anomaly layer and the multiple similar anomaly service layers; Obtain the maximum historical anomaly similarity and the average historical anomaly similarity among the multiple historical anomaly similarities; Based on the deviation between the maximum historical anomaly similarity and the mean historical anomaly similarity, the historical anomaly confidence level is calculated, the mean historical anomaly similarity is corrected, and anomaly similarity is obtained for anomaly discrimination.
[0066] In this embodiment of the application, firstly, the preliminary abnormal layer and its respective historical abnormality similarity are obtained from multiple similar abnormal service layers, for example, the historical abnormality similarities are 85%, 78%, and 90%.
[0067] Secondly, the maximum value of the historical anomaly similarity, i.e., 90%, is extracted as the maximum historical anomaly similarity, and the arithmetic mean is calculated, i.e., (85%+78%+90%) / 3≈84.33%, as the mean of historical anomaly similarity.
[0068] Then, the deviation between the maximum historical anomaly similarity and the mean historical anomaly similarity is calculated, i.e., 90% - 84.33% = 5.67%. This deviation reflects the difference between the most similar case among multiple similar anomaly cases and the average similarity level. The larger the deviation, the more likely there is a historical case that is very similar to the initial anomaly stratification, but the overall average similarity may be lowered by other cases with lower similarity.
[0069] Furthermore, a historical anomaly confidence level is introduced to correct the mean historical anomaly similarity. The historical anomaly confidence level can be calculated based on the ratio of the deviation value to the maximum historical anomaly similarity. For example, historical anomaly confidence level = 1 - deviation value / maximum historical anomaly similarity = 1 - 5.67% / 90% ≈ 1 - 0.063 = 93.7%. A higher confidence level indicates a stronger representativeness of the maximum historical anomaly similarity to the overall similarity, and therefore, it should be given a higher weight during correction.
[0070] Specifically, the correction formula can be set as: Anomaly Similarity = Average Historical Anomaly Similarity × (1 - Historical Anomaly Confidence × θ) + Maximum Historical Anomaly Similarity × Historical Anomaly Confidence × θ, where θ is an adjustment coefficient used to control the influence of the maximum similarity on the correction result. The value of θ ranges from 0 to 1, for example, θ = 0.5. Substituting the example data, Anomaly Similarity = 84.33% × (1 - 93.7% × 0.5) + 90% × 93.7% × 0.5 ≈ 86.985%.
[0071] By employing the aforementioned correction method, the calculation of anomaly similarity considers both the overall similarity level and the reference value of the most similar case, thereby improving the accuracy of anomaly detection results. If the corrected anomaly similarity is greater than or equal to the anomaly detection threshold, it is determined to be an anomaly; if it is lower than the threshold but there is at least one historical anomaly with a similarity higher than the historical anomaly threshold, it is considered a suspected anomaly; otherwise, it is a new type of anomaly.
[0072] In summary, compared with existing technologies, this application extracts multi-dimensional abnormal features from business layers, quantifies inter-layer flow dependencies by combining data dependency coefficients and update dependency coefficients, obtains a comprehensive abnormality degree by weighted calculation of form matching degree and flow matching degree, introduces historical abnormal business layer sets for similarity verification, calculates abnormal similarity in multiple dimensions by combining business domain feature factors, flow dependency coefficients and form matching degree, and corrects abnormal similarity by the deviation value between the maximum historical abnormal similarity and the mean. This enables the identification and type judgment of business domain layered data anomalies under the data lake, effectively improving the comprehensiveness, accuracy and intelligence of anomaly monitoring, and better meeting the needs of data quality monitoring in complex business scenarios.
[0073] In summary, the embodiments of this application have at least the following technical effects: This application provides an anomaly monitoring method for business domain layered data under a data lake. First, it obtains the business domain characteristic factors of the target business and performs data layer matching on the target business area in the data lake, obtaining multiple matching business layers. This achieves accurate positioning of the data lake's layered data from a business perspective. Second, it performs formal verification on the matching business layers to obtain the formal matching degree, and uses this to filter out the target business layer and easily confused business layers. The easily confused business layers are added to the historical abnormal business layer set, effectively solving the layer confusion problem caused by similar naming and structure. Subsequently, it randomly selects the first business layer from the target business layers for inter-layer flow verification, obtaining the flow dependency coefficient. Combined with data quality parameters, it obtains the flow matching degree, considering the data dependency relationship and update synergy between business layers, and can promptly detect anomalies of disordered inter-layer flow. Finally, based on the form matching degree and flow matching degree, the comprehensive anomaly degree is calculated. The first business layer with the comprehensive anomaly degree exceeding the threshold is combined with the historical abnormal business layer set for similarity verification, and the anomaly similarity is obtained and anomaly is judged. The layers judged as abnormal are added to the historical abnormal business layer set, which realizes accurate identification and dynamic accumulation of deep-level anomalies. It effectively solves the problem of untimely and inaccurate anomaly detection in the existing technology and improves the analysis, decision-making and business application effects based on data lake data.
[0074] Through the above technical solutions, this application can comprehensively monitor the business domain layered data under the data lake from multiple dimensions such as business domain characteristics, form matching, inter-layer flow dependency and historical anomaly patterns. Through multi-stage verification and dynamic learning mechanisms, it continuously optimizes the anomaly identification capability to ensure the accuracy, consistency and reliability of the business domain layered data.
[0075] Example 2, as Figure 2 As shown, based on the same inventive concept as the anomaly monitoring method for hierarchical data in a business domain under a data lake provided in Embodiment 1, this application also provides an anomaly monitoring system for hierarchical data in a business domain under a data lake, including: The information acquisition module 11 is used to acquire the business domain feature factors of the target business, and perform data layer matching on the target business area in the data lake based on the business domain feature factors to acquire multiple matching business layers; The formal verification module 12 is used to perform formal verification on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as target business layers, and add the matching business layers with a formal matching degree less than the formal matching degree threshold as easily confused business layers to the historical abnormal business layer set. The matching degree calculation module 13 is used to randomly select a first business layer from the target business layer, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of multiple first business layers. The similarity acquisition module 14 is used to obtain a comprehensive anomaly degree based on the form matching degree and the flow matching degree, and to perform similarity verification on the first business layer whose comprehensive anomaly degree is greater than the comprehensive anomaly threshold in combination with the historical abnormal business layer set, to obtain anomaly similarity, to perform anomaly discrimination, and to add the first business layer whose anomaly discrimination result is abnormal to the historical abnormal business layer set.
[0076] In one embodiment, the information acquisition module 11 is specifically used for: Obtain the business domain characteristic factors of the target business, wherein the business domain characteristic factors include key attributes of master data, business process identifiers, and control levels; Based on the business domain feature factors, data stratification matching is performed on the target business areas in the data lake to obtain multiple matching business stratifications, wherein the target business areas are obtained based on the master data key attributes.
[0077] Furthermore, in one application embodiment, formal validation is performed on multiple matching service layers to obtain the formal matching degree, including: A formal validation is performed on multiple matching service layers, wherein the formal validation includes data formal validation and content formal validation; Based on the data format verification results and the content format verification results, obtain the data matching degree and the content matching degree; The data matching degree and the content matching degree are weighted and calculated to obtain the form matching degree.
[0078] Furthermore, in one embodiment, matching service layers with a form matching degree greater than or equal to a form matching degree threshold are selected as target service layers, and matching service layers with a form matching degree less than the form matching degree threshold are selected as easily confused service layers and added to the historical abnormal service layer set, including: Obtain a form matching degree threshold, wherein the form matching degree threshold is obtained based on the mapping and matching of business domain feature factors; Filter out matching business layers with a form matching degree greater than or equal to the form matching degree threshold, and use them as target business layers; The matching business layers with a form matching degree less than the form matching degree threshold are classified as easily confused business layers and added to the historical abnormal business layer set.
[0079] In one embodiment, the matching degree calculation module 13 is specifically used for: Randomly select a first service layer from among the multiple target service layers; Perform inter-layer flow verification on the first business layer to obtain the proportion of data from the source layer and the update frequency; Based on the proportion and update frequency of the source layer data, calculate and obtain the data dependency coefficient and update dependency coefficient; The data dependency coefficient and the update dependency coefficient are weighted and calculated to obtain multiple flow dependency coefficients of the first business layer; Obtain the data quality parameters of the first service layer; The flow dependency coefficient is corrected and calculated based on the data quality parameters to obtain the flow matching degree of the first business layer, wherein the flow matching degree is the average of the flow matching degrees of the first business layer and multiple business layers.
[0080] Furthermore, inter-layer flow verification is performed on the first business layer to obtain the proportion and update frequency of data from the source layer, including: Perform inter-layer flow verification on the first business layer, obtain multiple data sources of the first business layer, and obtain the data proportion of the source layer data in the first business layer data as the source layer data proportion, wherein the data source includes multiple business layers; Obtain the update frequency of the first service layer, wherein the update frequency includes the update time frequency.
[0081] Furthermore, in one embodiment, based on the proportion of data from the source layer and the update frequency, the data dependency coefficient and update dependency coefficient are calculated and obtained, including: Obtain the proportion of source layer data in the first business layer as the data dependency coefficient; Obtain the update frequency of multiple data sources for the first business layer; Based on the update frequency of the first business layer and the average update frequency of multiple data sources, the update dependency coefficient is obtained.
[0082] Furthermore, the similarity acquisition module 14 is specifically used for: The form matching degree and the flow matching degree are weighted and calculated to obtain the comprehensive anomaly degree; Obtain the hierarchical set of historical abnormal business operations; When the overall anomaly degree of the first service layer is greater than the overall anomaly threshold, the first service layer is regarded as a preliminary anomaly layer; otherwise, the first service layer is randomly obtained again. Based on the business domain feature factors, flow dependency coefficients and form matching degree of the preliminary anomaly layering, similarity verification is performed on the historical anomaly business layering set to obtain multiple historical anomaly business layers whose similarity to the preliminary anomaly layering set is greater than the historical anomaly threshold, as multiple similar anomaly business layers. Based on the initial anomaly layering and the similarity of multiple historical anomalies in the multiple similar anomaly service layers, anomaly similarity is obtained and anomaly discrimination is performed; The preliminary anomaly layers that are identified as abnormal are added to the historical anomaly business layer set.
[0083] Furthermore, based on the preliminary anomaly layering and the historical anomaly similarity of multiple similar anomaly service layers, anomaly similarity is obtained, and anomaly discrimination is performed, including: Obtain the similarity of multiple historical anomalies between the initial anomaly layer and the multiple similar anomaly service layers; Obtain the maximum historical anomaly similarity and the average historical anomaly similarity among the multiple historical anomaly similarities; Based on the deviation between the maximum historical anomaly similarity and the mean historical anomaly similarity, the historical anomaly confidence level is calculated, the mean historical anomaly similarity is corrected, and anomaly similarity is obtained for anomaly discrimination.
Claims
1. A method for anomaly monitoring of hierarchical data in a business domain under a data lake, characterized in that, include: Obtain the business domain feature factors of the target business, and perform data layer matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers; Perform formal validation on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as the target business layers, and add the matching business layers with a formal matching degree less than the formal matching degree threshold as easily confused business layers to the historical abnormal business layer set. Randomly select a first business layer from the target business layers, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of multiple first business layers; Based on the form matching degree and the flow matching degree, a comprehensive anomaly degree is obtained. Combined with the historical abnormal service layer set, the first service layer with a comprehensive anomaly degree greater than the comprehensive anomaly threshold is checked for similarity, anomaly similarity is obtained, anomaly is judged, and the first service layer with anomaly judgment result is added to the historical abnormal service layer set.
2. The anomaly monitoring method for hierarchical data in a data lake business domain according to claim 1, characterized in that, Obtain the business domain feature factors of the target business, and perform data stratification matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers, including: Obtain the business domain characteristic factors of the target business, wherein the business domain characteristic factors include key attributes of master data, business process identifiers, and control levels; Based on the business domain feature factors, data stratification matching is performed on the target business areas in the data lake to obtain multiple matching business stratifications, wherein the target business areas are obtained based on the master data key attributes.
3. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 1, characterized in that, Perform formal validation on multiple matching business layers to obtain the formal matching degree, including: A formal validation is performed on multiple matching service layers, wherein the formal validation includes data formal validation and content formal validation; Based on the data format verification results and the content format verification results, obtain the data matching degree and the content matching degree; The data matching degree and the content matching degree are weighted and calculated to obtain the form matching degree.
4. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 1, characterized in that, Select matching business layers with a form matching degree greater than or equal to the form matching degree threshold as target business layers, and add matching business layers with a form matching degree less than the form matching degree threshold as easily confused business layers to the historical abnormal business layer set, including: Obtain a form matching degree threshold, wherein the form matching degree threshold is obtained based on the mapping and matching of business domain feature factors; Filter out matching business layers with a form matching degree greater than or equal to the form matching degree threshold, and use them as target business layers; The matching business layers with a form matching degree less than the form matching degree threshold are classified as easily confused business layers and added to the historical abnormal business layer set.
5. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 1, characterized in that, A first business layer is randomly selected from the target business layers, and inter-layer flow verification is performed to obtain multiple flow dependency coefficients. Based on the flow dependency coefficients and the data quality parameters of the multiple first business layers, the flow matching degree is obtained, including: Randomly select a first service layer from among the multiple target service layers; Perform inter-layer flow verification on the first business layer to obtain the proportion of data from the source layer and the update frequency; Based on the proportion and update frequency of the source layer data, calculate and obtain the data dependency coefficient and update dependency coefficient; The data dependency coefficient and the update dependency coefficient are weighted and calculated to obtain multiple flow dependency coefficients of the first business layer; Obtain the data quality parameters of the first service layer; The flow dependency coefficient is corrected and calculated based on the data quality parameters to obtain the flow matching degree of the first business layer, wherein the flow matching degree is the average of the flow matching degrees of the first business layer and multiple business layers.
6. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 5, characterized in that, Perform inter-layer flow verification on the first business layer to obtain the proportion and update frequency of data from the source layer, including: Perform inter-layer flow verification on the first business layer, obtain multiple data sources of the first business layer, and obtain the data proportion of the source layer data in the first business layer data as the source layer data proportion, wherein the data source includes multiple business layers; Obtain the update frequency of the first service layer, wherein the update frequency includes the update time frequency.
7. The anomaly monitoring method for hierarchical data in a data lake business domain according to claim 5, characterized in that, Based on the proportion and update frequency of the source layer data, the data dependency coefficient and update dependency coefficient are calculated and obtained, including: Obtain the proportion of source layer data in the first business layer as the data dependency coefficient; Obtain the update frequency of multiple data sources for the first business layer; Based on the update frequency of the first business layer and the average update frequency of multiple data sources, the update dependency coefficient is obtained.
8. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 1, characterized in that, Based on the form matching degree and the flow matching degree, a comprehensive anomaly degree is obtained. Then, in conjunction with the historical anomaly service layer set, similarity checks are performed on the first service layer whose comprehensive anomaly degree is greater than the comprehensive anomaly threshold to obtain anomaly similarity. Anomaly discrimination is then performed, and the first service layer with anomaly discrimination result is added to the historical anomaly service layer set, including: The form matching degree and the flow matching degree are weighted and calculated to obtain the comprehensive anomaly degree; Obtain the hierarchical set of historical abnormal business operations; When the overall anomaly degree of the first service layer is greater than the overall anomaly threshold, the first service layer is regarded as a preliminary anomaly layer; otherwise, the first service layer is randomly obtained again. Based on the business domain feature factors, flow dependency coefficients and form matching degree of the preliminary anomaly layering, similarity verification is performed on the historical anomaly business layering set to obtain multiple historical anomaly business layers whose similarity to the preliminary anomaly layering set is greater than the historical anomaly threshold, as multiple similar anomaly business layers. Based on the initial anomaly layering and the similarity of multiple historical anomalies in the multiple similar anomaly service layers, anomaly similarity is obtained and anomaly discrimination is performed; The preliminary anomaly layers that are identified as abnormal are added to the historical anomaly business layer set.
9. The anomaly monitoring method for hierarchical data in a business domain under a data lake according to claim 8, characterized in that, Based on the preliminary anomaly stratification and multiple historical anomaly similarities of the multiple similar anomaly service stratifications, anomaly similarity is obtained, and anomaly discrimination is performed, including: Obtain the similarity of multiple historical anomalies between the initial anomaly layer and the multiple similar anomaly service layers; Obtain the maximum historical anomaly similarity and the average historical anomaly similarity among the multiple historical anomaly similarities; Based on the deviation between the maximum historical anomaly similarity and the mean historical anomaly similarity, the historical anomaly confidence level is calculated, the mean historical anomaly similarity is corrected, and anomaly similarity is obtained for anomaly discrimination.
10. An anomaly monitoring system for hierarchical data in a business domain under a data lake, characterized in that, A method for monitoring anomalies in hierarchical data of a business domain under a data lake, as described in any one of claims 1-9, includes: The information acquisition module is used to acquire the business domain feature factors of the target business, and to perform data layer matching on the target business area in the data lake based on the business domain feature factors to obtain multiple matching business layers; The formal verification module is used to perform formal verification on multiple matching business layers, obtain the formal matching degree, and filter out the matching business layers with a formal matching degree greater than or equal to the formal matching degree threshold as the target business layers. The matching business layers with a formal matching degree less than the formal matching degree threshold are designated as easily confused business layers and added to the historical abnormal business layer set. The matching degree calculation module is used to randomly select a first business layer from the target business layer, perform inter-layer flow verification, obtain multiple flow dependency coefficients, and obtain the flow matching degree based on the flow dependency coefficients and the data quality parameters of multiple first business layers. The similarity acquisition module is used to obtain a comprehensive anomaly degree based on the form matching degree and the flow matching degree, and to perform similarity verification on the first business layer whose comprehensive anomaly degree is greater than the comprehensive anomaly threshold in combination with the historical anomaly business layer set, to obtain anomaly similarity, to perform anomaly judgment, and to add the first business layer whose anomaly judgment result is abnormal to the historical anomaly business layer set.