Data processing method, electronic device, and computer readable storage medium

Through data lineage analysis and multi-dimensional feature extraction, the problem of low accuracy in single-source data detection is solved, multi-source anomaly detection is realized, the data quality of physical machine attributes is accurately judged and the cause of the anomaly is located, thereby improving the accuracy and robustness of detection.

WO2025196511A1PCT designated stage Publication Date: 2025-09-25CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD

Patent Information

Application Number
PCT/IB2025/050238
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-21
Filing Date
2025-01-09
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

In the existing technology, when physical machine attribute data quality detection is performed based on single-source data, the detection results have low accuracy, high probability of false detection, poor robustness, and cannot assist in the rapid investigation of the root cause of the abnormality.

Method used

By performing data lineage analysis based on the identification information of the target machine and the target attributes, multiple data objects associated with the target attributes are determined, multi-dimensional features are extracted, and multi-source anomaly detection is performed based on multiple features to obtain target detection results to assist in judging the evaluation conclusions of the target attributes on multiple data objects.

Benefits of technology

It achieves accurate data quality detection results, locates the anomaly location and determines the cause of the anomaly, improves the accuracy and robustness of detection, reduces the time for data investigation, and assists in the rapid recovery of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050238_25092025_PF_FP_ABST
    Figure IB2025050238_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computer technologies and data processing. Disclosed are a data processing method, an electronic device, and a computer readable storage medium. The method comprises: performing data lineage analysis on the basis of identification information of a target machine and a target attribute of the target machine, and determining a plurality of data objects associated with the target attribute; performing feature extraction on the plurality of data objects to obtain a plurality of features, wherein the plurality of features are used for performing multi-source anomaly detection on the plurality of data objects; and performing data quality inspection on the plurality of data objects on the basis of the plurality of features to obtain a target inspection result, wherein the target inspection result is used for assisting in determining the evaluation conclusions of the target attribute across the plurality of data objects. The present disclosure solves the technical problems in the prior art of low accuracy of detection results and poor robustness of detection schemes resulting from detection based on single-source data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] TECHNICAL FIELD The present disclosure relates to the fields of computer technology and data processing technology, and more specifically, to a data processing method, electronic device, and computer-readable storage medium. BACKGROUND In stability management scenarios, operations such as physical machine operation and maintenance, hot migration, and release management are crucial for ensuring the stability of the Elastic Compute Service (ECS) system. However, these operations often rely heavily on the physical machine's inherent attribute data. Errors in any attribute data can lead to serious ECS system vulnerabilities, making the quality of physical machine attribute data crucial. Currently, when testing the data quality of physical machine attribute data, anomaly detection is typically performed based on single-source data, that is, anomaly detection is performed based on the single data source that the physical machine attribute data relies on. Therefore, the accuracy of the detection results is limited by the distribution of this single data source. Furthermore, due to the randomness of a single data source, the probability of false detection is high, the accuracy of the detection results is low, and the robustness of the detection solution is low. Currently, no effective solution has been proposed to address the above-mentioned issues. SUMMARY OF THE INVENTION Embodiments of the present disclosure provide a data processing method, electronic device, and computer-readable storage medium to at least address the technical issues in related technologies related to detection based on single-source data, resulting in low detection accuracy and low robustness of detection solutions. According to one aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: performing data lineage relationship analysis based on identification information of a target machine and a target attribute of the target machine to determine multiple data objects associated with the target attribute; performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; and performing data quality testing on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining an evaluation conclusion of the target attribute on the multiple data objects.According to another aspect of an embodiment of the present disclosure, a data processing method is further provided, comprising: obtaining a data processing request through a first application programming interface; and returning a data processing response through a second application programming interface; wherein request data carried in the data processing request comprises: identification information of a target machine and target attributes of the target machine, and response data carried in the data processing response comprises: target detection results, the target detection results being obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine, the multiple data objects being determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes, the multiple features being obtained by performing feature extraction on the multiple data objects, the multiple features being used to perform multi-source anomaly detection on the multiple data objects, and the target detection results being used to assist in determining evaluation conclusions of the target attributes on the multiple data objects. According to another aspect of an embodiment of the present disclosure, a data processing method is provided, comprising: obtaining a currently input data processing session request; and returning a data processing session reply in response to the data processing session request. The request data carried in the data processing session request includes: identification information of a target machine and target attributes of the target machine; and the information carried in the data processing session reply includes: target detection results, obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; multiple data objects determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; multiple features obtained by performing feature extraction on multiple data objects; the multiple features used to perform multi-source anomaly detection on the multiple data objects; and the target detection results used to assist in determining evaluation conclusions of the target attributes on the multiple data objects; and displaying the target detection results in a graphical user interface. According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a memory storing an executable program; and a processor configured to execute the program, wherein when the program executes, any one of the aforementioned data processing methods is executed. According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored executable program. When the executable program is executed, the device containing the computer-readable storage medium is controlled to perform any of the aforementioned data processing methods. According to another aspect of an embodiment of the present disclosure, a computer program product is provided. The computer program includes a computer program. When executed by a processor, the computer program implements any of the aforementioned data processing methods.In the disclosed embodiments, data lineage relationship analysis is performed based on the identification information of a target machine and the target attributes of the target machine to determine multiple data objects associated with the target attributes. Feature extraction is then performed on the multiple data objects from multiple dimensions to obtain multiple features in different dimensions. Data quality detection, i.e., multi-source anomaly detection, is then performed on the multiple data objects based on the multiple features to obtain target detection results. This target detection result can then be used to assist in determining the evaluation conclusion of the target attribute on multiple data objects, thereby determining whether the target attribute has an anomaly, the location of the anomaly, and the approximate cause of the anomaly. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This addresses the technical issues in related technologies related to low detection accuracy and low robustness of detection solutions caused by single-source data detection. It should be noted that the general description above and the detailed description that follows are merely examples and explanations of the present disclosure and do not constitute limitations on the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS The drawings described herein are intended to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are intended to explain the present disclosure and do not constitute undue limitations on the present disclosure. In the accompanying drawings: Figure 1 is a hardware structure block diagram of a computer terminal (or mobile device) for implementing a data processing method according to Example 1 of the present disclosure; Figure 2 is a flow chart of a data processing method according to Example 1 of the present disclosure; Figure 3 is an overall architecture diagram of an implementation of a data processing method according to Example 1 of the present disclosure; Figure 4 is a flow chart of a data processing method according to Example 2 of the present disclosure; Figure 5 is a flow chart of a data processing method according to Example 3 of the present disclosure; Figure 6 is a flow chart of a data processing method according to Example 4 of the present disclosure; Figure 7 is a structural schematic diagram of a data processing device according to Example 5 of the present disclosure; Figure 8 is a structural schematic diagram of another data processing device according to Example 5 of the present disclosure; Figure 9 is a structural schematic diagram of yet another data processing device according to Example 5 of the present disclosure; Figure 10 is a structural schematic diagram of yet another data processing device according to Example 5 of the present disclosure; Figure 11 is a structural block diagram of a computer terminal according to an embodiment of the present disclosure.DETAILED DESCRIPTION To help those skilled in the art better understand the present disclosure, the following will provide a clear and complete description of the technical solutions in the embodiments of the present disclosure, in conjunction with the accompanying drawings. It should be noted that the described embodiments represent only a portion of the present disclosure, and are not exhaustive. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present disclosure without inventive effort should fall within the scope of protection of the present disclosure. It should be noted that the terms "first," "second," and so on, in the specification and claims of the present disclosure, and in the accompanying drawings, are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that such terms are interchangeable where appropriate, so that the embodiments of the present disclosure described herein can be implemented in an order other than that illustrated or described herein. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to the steps or units expressly listed, but may include other steps or units not expressly listed or inherent to such process, method, product, or apparatus. First, some nouns or terms used in describing the embodiments of the present disclosure are subject to the following interpretations: Machine attributes: In the embodiments of the present disclosure, machine attributes can be understood as characteristics of a physical machine (or virtual machine, etc.), such as whether it contains a local hard drive or a graphics processing unit (GPU), which are used to characterize the current state of the machine. Data lineage: In the embodiments of the present disclosure, data lineage, also known as data lineage, data provenance, or data lineage, refers to the naturally formed relationships between data throughout its entire lifecycle, from generation, processing, processing, fusion, and flow to its eventual extinction. Top-Down Algorithm for Discovering Functional Dependencies (TANE): An algorithm for discovering underlying patterns in data, specifically, an algorithm for discovering functional dependencies in relational databases. TANE is based on the concept of attribute closure and discovers functional dependencies by recursively computing attribute closures. For example, TANE first finds the functional dependencies of a single attribute, then gradually expands to the functional dependencies of multiple attributes, and finally finds all the functional dependencies.Closure Table Based Algorithm for Discovering Functional Dependencies (CTANE): This algorithm improves and extends the TANE algorithm. CTANE uses closure tables to accelerate the discovery of functional dependencies, reducing recalculation time and improving algorithm efficiency.

[0002] FDs (Functional Dependencies) refer to relationships in a database where the value of one or more attributes determines the value of another. This relationship can be represented by AB, where A and B are attribute sets. Functional dependencies are crucial in database design and optimization because they allow for a quick understanding of data relationships, leading to better database management and querying. Functional dependencies can also be used to normalize databases, reduce data redundancy, and improve data storage and query efficiency. Minimal non-trivial dependency: In a relational schema, for any given functional dependency, if no proper subset satisfies it, the dependency is considered minimal. Furthermore, a non-trivial dependency is one that cannot be derived from other dependencies. For example, a functional dependency XY is considered non-trivial if it satisfies YQX. A functional dependency is considered minimal if the set X / Y contains no additional dependencies. Minimum Approximate Non-Trivial Dependency: In a relational database, if there exists a functional dependency whose right-hand attribute is a linear combination of the left-hand attributes, and no proper subset also satisfies this condition, then this functional dependency is called a Minimum Approximate Non-Trivial Dependency. Support: In the disclosed embodiments, support refers to the degree of support, indicating the frequency with which the antecedent and consequent terms appear simultaneously in a dataset. Validity: In the disclosed embodiments, validity is used to measure the credibility of the current data. For example, the closer it is to the data source, the fewer times it has been processed, and the lower the processing risk, the higher the validity. Influence: In the disclosed embodiments, influence refers to the impact of the current data, generally referring to the business impact (such as the number of uses) and physical impact (such as the number of downstream dependencies) of the current node on downstream nodes. Physical machine operation and maintenance is a crucial element in ensuring the stability of the ECS system. Current physical machine operation and maintenance actions and rules rely heavily on physical machine attribute data. For example, when initiating a live migration operation, it is necessary to determine whether the physical machine contains the specified disk or is bare metal. Data errors can lead to misjudgments and failures, resulting in operational failures or service outages. Therefore, the quality of physical machine attribute data is crucial. Currently, anomaly detection is typically performed based on single-source data. Detection accuracy depends on the distribution of the single-source data, significantly limiting accuracy. Furthermore, due to the randomness of a single data source, the probability of false positives is high, resulting in low detection accuracy and robustness. Furthermore, current detection methods provide only "yes" or "no" results, making them ineffective for quickly identifying the root cause of anomalies.Related art anomaly detection based on single-source data suffers from the following drawbacks. Defect 1: The accuracy of the detection results is limited by the distribution of the single data source itself, resulting in a high probability of false detection, low accuracy, and low robustness of the detection scheme. Defect 2: Only a yes or no conclusion can be obtained, i.e., whether the data is anomaly or not, but the location of the anomaly cannot be located, which in turn cannot assist in quickly troubleshooting the root cause of the anomaly or facilitating rapid task recovery. Defect 3: The method cannot provide a basis for rapid judgment in task decision-making, i.e., it cannot provide relatively accurate data and provide explainability when anomaly data is detected. Prior to the present disclosure, no effective solution to the above drawbacks has been proposed. Example 1 According to an embodiment of the present disclosure, a data processing method is provided. It should be noted that the steps shown in the flowcharts of the accompanying figures can be executed in a computer system, such as a set of computer-executable instructions. Although the flowcharts illustrate a logical order, in some cases, the steps shown or described can be executed in a different order. The method embodiment provided in Example 1 of the present disclosure can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 is a hardware block diagram of a computer terminal (or mobile device) for implementing a data processing method according to Embodiment 1 of the present disclosure. As shown in Figure 1 , the computer terminal 10 (or mobile device) may include one or more processors 102 (illustrated as 102a, 102b, 102n in the figure) (processor 102 may include, but is not limited to, a processing device such as a microprocessor (MCU) or a programmable logic device (FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, the computer terminal 100 may include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS), a network interface, a power supply, and / or a camera. Those skilled in the art will appreciate that the structure shown in Figure 1 is merely illustrative and does not limit the structure of the electronic device described above. For example, the computer terminal 10 may include more or fewer components than shown in Figure 1 , or have a configuration different from that shown in Figure 1 . It should be noted that the one or more processors 102 and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be fully or partially integrated into any other component of the computer terminal 10 (or mobile device).As described in the embodiments of the present disclosure, the data processing circuit functions as a processor control (e.g., selecting a variable resistor terminal path connected to an interface). Memory 104 can be used to store application software programs and modules, such as the program instructions / data storage device corresponding to the data processing method in the embodiments of the present disclosure. Processor 102 executes the software programs and modules stored in memory 104 to perform various functional applications and data processing, thereby implementing the aforementioned data processing method. Memory 104 may include high-speed random access memory (RAM) and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, memory 104 may further include memory remotely located from processor 102, which can be connected to computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. Transmission device 106 is used to receive or transmit data via a network. Specific examples of such networks may include a wireless network provided by the communications provider of computer terminal 10. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 can be a radio frequency (RF) module for wireless communication with the Internet. The display can be, for example, a touchscreen liquid crystal display (LCD), which allows a user to interact with the user interface of the computer terminal 10 (or mobile device). In the above operating environment, the present disclosure provides a data processing method as shown in Figure 2. Figure 2 is a flow chart of a data processing method according to Example 1 of the present disclosure. As shown in Figure 2, the method may include the following steps: Step S21: performing data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine to determine multiple data objects associated with the target attribute; Step S22: performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; Step S23: performing data quality detection on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. In the embodiments of the present disclosure, the target machine may be a physical machine, a virtual machine, or other machine device, and is not limited here.Identification information can be understood as field information related to the target machine, such as the target machine's configuration information (e.g., hardware and network configuration), status information, and performance metrics, without limitation. Target attributes are machine attributes of the target machine, which can be understood as attribute information used to describe the target machine's status or performance, such as whether a local disk is present or whether a graphics processing unit (GPU) is present, characterizing the machine's current state. Data lineage refers to the naturally formed relationships between data throughout its entire lifecycle, from generation, processing, processing, integration, flow, to eventual extinction. It reflects the dependencies and propagation paths between data. For example, in a physical machine operation and maintenance scenario, data lineage may involve multiple aspects, such as the physical machine's hardware information, operation logs, and fault reports. For example, the hardware configuration data of each physical machine (such as the central processing unit (CPU) model, memory capacity, and hard drive capacity) can be traced back to hardware procurement records and configuration management databases. This data indirectly impacts server performance and stability, thus forming a data lineage relationship. During physical machine operation and maintenance management, the system generates various operation and maintenance logs, including system logs, performance logs, and error logs. These logs record information such as system operating status and abnormalities, and can be traced back to specific hardware devices and operators, forming a data lineage relationship. When a physical machine fails, a fault report and maintenance record are generated. This data can be traced back to specific hardware devices and maintenance personnel, forming a data lineage relationship (without limitation here). Data lineage relationship analysis based on the target machine's identification information and target attributes can be understood as determining, based on the identification information and target attributes, multiple data objects that have a lineage relationship with the target machine in the current scenario. Furthermore, it can be understood as determining multiple data objects associated with the target attributes in the current scenario. A data object can be understood as a data table or data node that has a data lineage relationship with the target attribute, that is, a parameter value of data that has a data lineage relationship with the target attribute, or, more specifically, as each data table or data node for the same attribute. Feature extraction from multiple data objects can be understood as extracting features from data in multiple data tables that have a data lineage relationship with the target attribute of the target machine, thereby obtaining multiple features. Multiple features can be understood as features extracted from multiple dimensions. Exemplary features include, but are not limited to, extracting features from multiple data objects based on dimensions such as influence, page views, support, and effectiveness, thereby obtaining multiple features. This is not a limitation here.Multiple features are used to perform multi-source anomaly detection on multiple data objects. This disclosure utilizes multiple features to perform data quality testing on multiple data objects, thereby obtaining target detection results. Compared to anomaly detection based on single-source data, this disclosure utilizes a multi-dimensional approach to anomaly detection, focusing on the overall data distribution and addressing the data distribution limitations inherent in single-source data. This effectively improves the accuracy of detection results and enhances the robustness of the detection scheme. Anomaly detection based on multi-source data accurately obtains target detection results, which are used to assist in determining the evaluation conclusions for target attributes across multiple data objects. These evaluation conclusions can include credibility and / or risk levels. Specifically, this helps determine whether the data for the target attribute in each data table or data node is credible and whether risks exist in each data table or data node. Furthermore, it accurately determines whether anomalies exist for the target attribute in each process step. Therefore, based on the target detection results, it is possible to determine whether an anomaly in the target attribute exists and locate the specific location of the anomaly, that is, to locate the specific process step where the anomaly occurred, such as the task processing or source collection where the error occurred. This allows the specific cause of the error to be roughly determined, thereby reducing the time required for data quality inspections and facilitating rapid recovery of data tasks. In the disclosed embodiments, data lineage relationship analysis is performed based on the identification information of the target machine and the target attribute of the target machine to determine multiple data objects associated with the target attribute. Feature extraction is then performed on these multiple data objects from multiple dimensions to obtain multiple features in different dimensions. Data quality detection is then performed on these multiple data objects based on these multiple features, i.e., multi-source anomaly detection, to obtain target detection results. These target detection results can then assist in determining the evaluation conclusion of the target attribute across multiple data objects, thereby determining whether an anomaly exists in the target attribute, the location of the anomaly, and the approximate cause of the anomaly. It can be seen that the present disclosure can achieve the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results of machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves the accuracy of detection, reduces the time required for data investigation, and improves the robustness of the detection solution.The data processing method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving data quality detection of machine attributes in fields such as e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, data quality detection of machine attributes in e-commerce services, data quality detection of machine attributes in educational services, and data quality detection of machine attributes in legal services are not limited here. According to the disclosed embodiments, data lineage relationship analysis is performed based on the identification information of a target machine and the target attributes of the target machine to determine multiple data objects associated with the target attributes. Feature extraction is then performed on the multiple data objects from multiple dimensions to obtain multiple features in different dimensions. Data quality detection, i.e., multi-source anomaly detection, is then performed on the multiple data objects based on the multiple features to obtain target detection results. This target detection result can then be used to assist in determining the evaluation conclusion of the target attribute on multiple data objects, thereby determining whether the target attribute has an anomaly, the location of the anomaly, and the approximate cause of the anomaly. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This addresses the technical issues in related technologies related to low detection accuracy and low robustness of detection solutions caused by single-source data detection. In an optional embodiment, in step S21, data lineage analysis is performed based on the identification information and the target attribute to determine multiple data objects associated with the target attribute. The method includes the following steps: Step S211: Performing data lineage analysis based on the identification information and the target attribute to match and obtain at least one target data source from multiple candidate data sources; Step S212: Retrieving multiple data objects from the at least one target data source using a preset functional dependency calculation method, wherein the preset functional dependency calculation method is used to discover a set of dependency items in the at least one target data source that satisfy a preset functional dependency relationship. In the disclosed embodiment, when performing data lineage analysis based on the identification information and the target attribute to determine multiple data objects associated with the target attribute, performing data lineage analysis based on the identification information and the target attribute to match and obtain at least one target data source from multiple candidate data sources can be understood as matching all data streams based on the predetermined fields and attributes of the target machine to obtain a data source that includes the target machine and the target attributes.The multiple candidate data sources can be understood as all data flows involved in all execution processes in this scenario, and can be represented by a network diagram, without limitation. After determining at least one target data source in all data flows that has a data lineage relationship with the target attribute based on the identification information and the target attribute, a preset functional dependency calculation method can be used to obtain multiple data objects from the at least one target data source. The preset functional dependency calculation method is used to discover a dependency item set in the at least one target data source that satisfies a preset functional dependency relationship. The preset functional dependency relationship can be a minimum approximate non-trivial dependency relationship. That is, the preset functional dependency calculation method is used to discover multiple FDs items in the at least one target data source that satisfy the minimum approximate non-trivial dependency item set, forming an FDs item set. For example, considering interpretability, the preset functional dependency calculation method can employ the TANE algorithm or the CTANE algorithm, without limitation. The TANE algorithm or the CTANE algorithm is used to calculate the at least one target data source to determine the FDs item set in the at least one target data source that satisfies the minimum approximate non-trivial dependency item set. In an optional embodiment, in step S22, feature extraction is performed on multiple data objects to obtain multiple features, including at least some of the following: Step S221: Determining influence features of the multiple data objects based on the number of downstream data objects affected by the multiple data objects; Step S222: Determining access features of the multiple data objects based on the number of queries for the multiple data objects; Step S223: Determining validity features of the multiple data objects based on the lineage distances between the multiple data objects and the target detection data, the lineage distances between the multiple data objects and the data source of the target detection data, and whether the multiple data objects are located in the lineage graph of the data source; Step S224: Determining support features of the multiple data objects based on the support of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency. In the disclosed embodiment, the multiple features are obtained by extracting features from four dimensions: influence, access, validity, and support. When extracting features from multiple data objects to obtain multiple features, influence features, access features, validity features, and support features of the multiple data objects can be extracted. When extracting influence features of the multiple data objects, the influence features of the multiple data objects can be determined based on the number of downstream data objects affected by the multiple data objects.The downstream data object can be understood as the downstream data lineage node that uses these multiple data objects, that is, the downstream data node that uses this data table. It can be understood that the downstream data node is affected by the upstream data node (i.e., the data object). If there is a parameter error in the upstream data node, the downstream data node that processes using the parameters of this upstream data node will also obtain a deviated processing result. Therefore, when determining the influence degree characteristics of multiple data objects, the influence degree characteristics of multiple data objects can be determined based on the number of downstream data objects affected by the multiple data objects. Exemplarily, the influence degree characteristic XinfiuO T 4) can be represented according to formula (1), which is used to reflect the number of downstream data lineage nodes corresponding to the multiple data objects.

[0003] Xinftu(X ->?1) = £仁 ° if(nodei G influence (%) then 1 else 0 ) Formula (1) where, influence(x) represents the downstream node, and n represents the number of all nodes in all data streams. The greater the influence degree characteristic, the higher the credibility of this data object, and therefore the greater the weight corresponding to this influence degree characteristic. When extracting the access volume characteristics of multiple data objects, the access volume characteristics of multiple data objects can be determined based on the number of query times of the multiple data objects. Among them, the number of query times can be understood as the number of times the end user queries this data table. When extracting validity features for multiple data objects, the validity features can be determined based on the kinship distances between the multiple data objects and the target detection data, the kinship distances between the multiple data objects and the data source of the target detection data, and whether the multiple data objects are located in the kinship graph of the data source. In the disclosed embodiment, the validity features are quantified by the proximity and depth of the kinship relationships with the target detection data (i.e., the target data table). Determining the validity features for the multiple data objects based on the kinship distances between the multiple data objects and the target detection data can be understood as determining the validity features for the multiple data objects based on the kinship distances between the node currently being calculated and the target data table. It is understood that the greater the kinship distance between the node currently being calculated and the target data table, the higher the comparative validity and the higher the validity feature. This is because the greater the kinship distance, the more diverse the data kinship context and the more different the data processing logic, thus reducing the probability of data consistency. Conversely, if the two data are consistent, the comparative validity is more reliable and the validity feature is higher. Determining the validity characteristics of multiple data objects based on the lineage distances between the multiple data objects and the data source of the target detection data can be understood as determining the validity characteristics of the multiple data objects based on the lineage distances between the current node to be calculated and the data source of the target data table. The data source of the target data table can be understood as the acquisition source of the target data table, i.e., the root node. It can be understood that the closer the current node to be calculated is to the data source of the target data table, the higher its credibility and the better the comparative validity. The closer it is to the root node, the fewer data processing steps there are, thus reducing the probability of data errors. Determining the validity characteristics of multiple data objects based on whether the multiple data objects are located in the lineage graph of the data source can be understood as determining the validity characteristics of the multiple data objects based on whether the current node to be calculated is located in the lineage graph of the target detection node, i.e., whether it is located in the lineage graph of the target data table. It is understandable that if the node to be calculated is not in the lineage graph of the target data table, it is likely to originate from another data source. This means that the node to be calculated is not affected by any data processing or data collection steps in the lineage graph. Therefore, the node to be calculated has the highest credibility if it is not in the lineage graph, but its validity is lower than that of the node in the lineage graph. It can be understood that if a node is not in the lineage graph of the target detection node, it is likely to originate from another data source. The more data sources, the better the overall comparative validity. For example, the validity feature x vaUd (X t 0) can be expressed according to formula (3). Formula (3) If the node to be calculated belongs to the node in the upstream lineage graph of the target detection node, its validity is its node depth. deep ). If the node to be calculated currently does not belong to the node in the upstream lineage graph of the target detection node, then its validity is the maximum depth of the upstream lineage graph of the target detection node (max (G (%) deep)) plus the inverse of the sum of the depths of the node to be calculated currently from the terminal node of the node to be calculated currently. It can be seen that the deeper the depth, the lower the weight. When extracting the support features of multiple data objects, the support features of the multiple data objects can be determined based on the support of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency. It is understandable that if the preset functional dependency calculation method allows the existence of errors, then in some approximate dependencies, there will be a situation where the support is less than 1. For example, the support feature x support (X t 4) can be expressed according to formula (4), which is used to express the support of the data item with the largest support in the minimum approximate non-trivial function dependency XA of the approximate dependency. The higher the support, the higher the credibility. The data with the highest credibility is selected as the voting item in the subsequent steps.

[0004] X support^ T 4) = P(XU 4) Formula (4) Wherein, P(XU 4) represents the support of each component corresponding to each dependency in each data table. Components can be understood as different attribute values ​​corresponding to the dependency. For example, bare metal and non-bare metal are two different components. In an optional embodiment, the data processing method further includes the following method steps: Step S225, using a preset normalization processing method to normalize the influence feature, the visit feature, and the effectiveness feature to obtain a processing result. In the embodiment of the present disclosure, after obtaining multiple features of different dimensions, normalization processing is required because the feature dimensions of the multiple features are different. The preset normalization processing method can be used to normalize the influence feature, the visit feature, and the effectiveness feature to obtain a processing result. Exemplarily, the preset normalization processing method can be using a sigmoid function for normalization, that is, using a sigmoid function to normalize the above multiple features. It is understandable that since the support feature value itself is [0, 1], no additional processing is required. The features that need to be processed are the influence feature, the effectiveness feature, and the visit volume feature. For example, the normalization processing method b(x) can be expressed according to formula (5). In an optional embodiment, in step S23, data quality testing is performed on multiple data objects based on multiple features to obtain target detection results, including the following method steps: Step S231: Data quality testing is performed on multiple data objects based on the processing results and the support features to obtain target detection results. In the disclosed embodiment, when data quality testing is performed on multiple data objects based on multiple features to obtain target detection results, since the support feature value itself is [0, 1] and does not require additional processing, data quality testing is performed on multiple data objects based on the processing results obtained by the normalization process and the support features to obtain target detection results. Table 1 When determining the target detection result of each dependency, the influence feature, effectiveness feature, and access feature of each dependency are first normalized according to Formula (5). Then, data quality detection is performed based on the normalized processing results and the support feature, thereby determining the target detection result of each dependency. In an optional embodiment, in step S231, data quality detection is performed on multiple data objects based on the processing results and the support feature to obtain the target detection result, including the following method steps: Step S2311, weighting the processing results according to a first weight calculation method to obtain first weights corresponding to the multiple data objects; Step S2312, weighting the first weights and the support feature according to a second weight calculation method to obtain second weights of the target dependency corresponding to the multiple data objects on each component of the target attribute, wherein the second weight calculation method is different from the first weight calculation method; Step S2313, data quality detection is performed on the multiple data objects according to the second weight to obtain the target detection result. In the disclosed embodiment, when data quality detection is performed on multiple data objects based on the processing results and support features to obtain target detection results, the processing results can be weighted according to a first weight calculation method to obtain first weights corresponding to the multiple data objects. This can be understood as calculating the first weight corresponding to each data object based on the normalized influence feature, access feature, and validity feature. For example, for any data object (i.e., any data node) and any minimum approximate non-trivial dependency, the first weight w can be calculated according to formula (6): node.f . „ Wherein, nodet represents a data node, / y represents a minimum approximate non-trivial dependency, and formula (6) is used to calculate the sum of any node and any minimum approximate non-trivial dependency, and to obtain the first weight corresponding to each data object. After determining the first weight, the first weight and the support feature are weighted according to the second weight calculation method to obtain the second weight of the target dependency corresponding to the multiple data objects on each component of the target attribute, that is, to obtain the second weight of any minimum approximate non-trivial dependency corresponding to the multiple data objects on each component of the target attribute, which can be understood as determining the second weight of each component of any data object (that is, any data node) and any minimum approximate non-trivial dependency. Wherein, the second weight calculation method is different from the first weight calculation method. For example, for any data object (that is, any data node) and any minimum approximate non-trivial dependency, the second weight w can be calculated according to formula (7): node .f, Vk . After determining the second weight, data quality testing is performed on multiple data objects based on the second weight to obtain a target detection result. In an optional embodiment, in step S2313, data quality testing is performed on multiple data objects based on the second weight to obtain a target detection result, including the following method steps: Step S23131: Obtaining the weights of each data object associated with the target dependency for each component of the target attribute based on the second weight to obtain a first voting result; Step S23132: Based on the first voting result, summarizing the weights of each data object associated with the target dependency for each component of the target attribute to obtain a second voting result; Step S23133: Using the second voting result, data quality testing is performed on multiple data objects to obtain a target detection result. In the embodiment of the present disclosure, considering that in actual implementation, each component of each dependency needs to be calculated and then merged into a final result, and since each dependency is associated with a data object in a different manner, the present disclosure designs a two-voting method to determine the final target detection result. When performing data quality testing on multiple data objects based on the second weights and obtaining the target detection results, the weights of each data object associated with the target dependency on each component of the target attribute are first obtained based on the second weights to obtain the first voting result. This can be understood as determining the overall voting result of the kth component of any approximately non-trivial dependency, i.e., the first voting result. For example, the first voting result, votef, can be calculated according to formula (8). Vk „ Here, n represents the number of nodes in the entire data flow corresponding to a certain minimum approximate non-trivial dependency. At this point, the voting results for each component corresponding to each minimum approximate non-trivial dependency are determined. Each minimum approximate non-trivial dependency is then matched with the target detection table to obtain each component and its corresponding weight for each machine's minimum approximate non-trivial dependency. Once all matching is complete, the weight values ​​for each component corresponding to different minimum approximate non-trivial dependencies are obtained, as shown in Table 2, which is an example of the first voting results. Table 2 After obtaining the first voting result, the weights of each data object associated with the target dependency on each component of the target attribute are summarized based on the first voting result to obtain the second voting result. For example, the second voting result total vote can be calculated according to formula (9) v , which represents the sum of the k-th component voting results of all approximately non-trivial dependencies. total_vote Vk= Y}jL1votef. Vk Formula (9) where m represents the number of all approximate non-trivial dependencies in all data streams. The second voting result can be understood as summarizing each component. Taking the data in Table 2 as an example, as shown in Table 3, Table 3 is an example of the second voting result. Table 3 It can be seen that by performing a second vote on each component in Table 2, selecting the component with the largest voting result for each dependency, Table 3 is obtained. After obtaining the first voting result, the second voting result is used to perform data quality detection on multiple data objects to obtain the target detection result. In an optional embodiment, in step S23133, using the second voting result to perform data quality detection on multiple data objects to obtain the target detection result includes the following method steps: Step S231331, calculating the vote rate corresponding to each component of the target attribute using the second voting result; Step S231332, determining the target detection result based on the maximum value of the vote rate. In the embodiments of the present disclosure, when using the second voting result to perform data quality detection on multiple data objects to obtain the target detection result, the vote rate corresponding to each component of the target attribute can be calculated using the second voting result, that is, calculating the vote rate of each component. Exemplarily, the vote rate can be calculated according to Formula (10) total vote rate v , that is, calculating a vote rate according to the weight of each vote to obtain the vote rate of the kth component. total ~ vote ~ rate. v .k Formula (10) 舄 total_vote Vk Exemplarily, the vote rate can be in percentage form, which is not limited here. Taking the data in Table 3 as an example, as shown in Table 4, Table 4 is the vote rate details. Table 4 After obtaining the vote rate corresponding to each component, determining the target detection result based on the maximum value of the vote rate can be understood as selecting the component with the highest vote rate among the vote rates corresponding to each component. This component is the correct data, that is, the multi-source detection result is the correct value. Exemplarily, the maximum value of the vote rate can be calculated according to Formula (11) ^^total_votej^atejnax a totalj^otej'atejnax = argmax(total_votej'ate Vk) Formula (11) For example, taking the data in Table 4 as an example, as shown in Table 5, Table 5 is the result value of the correct multi-source detection, that is, the component with the highest vote rate in Table 4. Table 5 Wherein, " represents correct data, S[ represents real data, and 0 represents the number of correct data obtained by calculation. In an optional embodiment, the data processing method further includes the following method steps: Step S24, marking and rendering the target detection result on the target lineage graph to obtain a rendering result, wherein the target lineage graph is a lineage graph where at least some of the multiple data objects are located, and the rendering result is used to assist in locating the abnormal state and the normal state of the target attribute in the multiple data objects, and / or to determine the abnormality category corresponding to the abnormal state. In the embodiment of the present disclosure, after determining the target detection result, the target detection result can also be marked and rendered on the target lineage graph to obtain a rendering result. It can be understood that the target detection result is marked and rendered on the lineage graph where at least some of the multiple data objects are located, for example, the abnormal part is marked in red and the normal part is marked in blue, which is not limited here. By marking and rendering on the target lineage graph, a lineage abnormality map, that is, the rendering result, can be obtained. Thus, the abnormal and normal states of target attributes can be assisted in locating them in multiple data objects through the lineage anomaly graph, for example, visually locating the abnormal and normal locations. At the same time, the anomaly category corresponding to the abnormal state can also be determined, that is, the possible cause of the anomaly can be determined, such as a data processing error or a data source acquisition error. Figure 3 is an overall architecture diagram of a data processing method according to Example 1 of the present disclosure. First, data lineage relationship collection is performed. Then, machine attribute anomaly detection is performed based on the collected lineage relationship data. The detection results of the machine attribute anomaly detection are lineage-labeled through engineering means, and the abnormal states of all lineage graphs of the machine are rendered, thereby assisting in root cause location. Among them, machine attribute anomaly detection mainly includes feature extraction and two voting parts. In the feature extraction part, in addition to the two features of support representing reference degree and visit volume representing popularity, the present disclosure also derives two features based on lineage relationship, namely influence representing popularity and validity representing reference degree. The influence is based on the principle that the more business gamma 1 is used, the lower the probability of an anomaly. Effectiveness is based on the principle that fewer processing steps reduces the probability of problems. Based on the four aforementioned characteristics, this disclosure uses the Tane algorithm to determine dependencies and employs a two-voting approach to test machine attributes from multiple dimensions, accurately determining test results and improving robustness. As can be seen, this disclosure proposes a multi-source data quality detection and repair method for machine attributes based on data lineage relationships, based on analysis of ECS data and link characteristics.The data credibility and risk level of each table (or each data node) are measured from multiple dimensions, mainly from the perspective of effectiveness, influence, access volume and support. The final mathematical model is further constructed for the credibility and risk level of the same machine attribute at each node, thereby determining the final abnormal data, possible correct data, the location of the abnormality and the possible cause of the abnormality. It is easy to understand that the beneficial effects of the data processing method provided by the present disclosure include the following points. Beneficial effect (1): Compared with detection based on single-source data, the present disclosure focuses on the overall data distribution state and performs abnormality detection based on multi-source data, which solves the inherent limitations of single-source distribution, making the detection results more accurate and the detection scheme more robust. Beneficial effect (2): The present disclosure can provide the location and possible cause of the abnormality, greatly reducing the time for data quality inspection and assisting in rapid task recovery. Beneficial effect (3): The present disclosure can not only detect abnormal data, but also provide the most likely correct data and provide explainability, which can provide a basis for rapid judgment for task decision-making. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, storage, and display, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. The collection, use, and processing of the relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or reject. Furthermore, it should be noted that for the sake of simplicity, the aforementioned method embodiments are described as a series of actions. However, those skilled in the art should understand that this disclosure is not limited by the order of the actions described, as certain steps can be performed in a different order or simultaneously according to this disclosure. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are preferred embodiments, and the actions and modules involved are not necessarily required by this disclosure. Through the above description of the embodiments, those skilled in the art will clearly understand that the methods according to the above embodiments can be implemented using software and a necessary general-purpose hardware platform, or alternatively, hardware. Based on this understanding, the technical solution of the present disclosure, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, a magnetic disk, or an optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in various embodiments of the present disclosure.Example 2 In an operating environment such as that in Example 1, the present disclosure provides a data processing method as shown in Figure 4. Figure 4 is a flowchart of a data processing method according to Example 2 of the present disclosure. As shown in Figure 4, the method includes: step S41, obtaining a data processing request through a first application programming interface; step S42, returning a data processing response through a second application programming interface; wherein, the request data carried in the data processing request includes: identification information of the target machine and target attributes of the target machine, and the response data carried in the data processing response includes: target detection results, the target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine, the multiple data objects are determined after performing data lineage relationship analysis based on the identification information of the target machine and the target attributes, the multiple features are obtained by performing feature extraction on the multiple data objects, the multiple features are used to perform multi-source anomaly detection on the multiple data objects, and the target detection results are used to assist in determining the evaluation conclusion of the target attributes on the multiple data objects. In the disclosed embodiments, a data processing request can be understood as a request initiated by calling a first application programming interface (API) to perform anomaly detection on a target attribute of a target machine. The request data carried in the data processing request includes the target machine's identification information and the target attribute of the target machine. A data processing response can be understood as a reply corresponding to the data processing request, initiated by calling a second application programming interface (API). The response data carried in the data processing response includes the target detection result. In the disclosed embodiments, a data processing request carrying the target machine's identification information and the target attribute of the target machine is obtained through the first application programming interface, and a data processing response carrying the target detection result is returned through the second application programming interface. The target detection result is obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attribute of the target machine. The multiple data objects are determined by performing data lineage relationship analysis based on the target machine's identification information and the target attribute. The multiple features are obtained by performing feature extraction on the multiple data objects. The multiple features are used to perform multi-source anomaly detection on the multiple data objects. The target detection result is used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. It can be seen that the present disclosure can achieve the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results of machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves the accuracy of detection, reduces the time required for data investigation, and improves the robustness of the detection solution.The data processing method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving data quality detection of machine attributes in fields such as e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, data quality detection of machine attributes in e-commerce services, data quality detection of machine attributes in educational services, and data quality detection of machine attributes in legal services are not limited here. According to an embodiment of the present disclosure, a data processing request carrying identification information and target attributes of a target machine is obtained through a first application programming interface (API). A data processing response carrying target detection results is then returned through a second application programming interface (API). The target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine. The multiple data objects are determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes. The multiple features are obtained by performing feature extraction on the multiple data objects. The multiple features are used to perform multi-source anomaly detection on the multiple data objects. The target detection results are used to assist in determining the evaluation conclusions of the target attributes on the multiple data objects. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing data quality anomalies, and determining the possible causes of the anomalies. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This solves the technical problem in related technologies of low detection accuracy and low robustness caused by detection based on single-source data. It should be noted that the preferred implementation of this embodiment can refer to the relevant description in Example 1 and will not be repeated here. Example 3 In the operating environment as in Example 1, the present disclosure provides a data processing method as shown in FIG5.FIG5 is a flowchart of a data processing method according to Embodiment 3 of the present disclosure. As shown in FIG5 , the method includes: step S51, obtaining a currently input data processing dialogue request; step S52, returning a data processing dialogue reply in response to the data processing dialogue request; wherein the request data carried in the data processing dialogue request includes: identification information of a target machine and target attributes of the target machine; and information carried in the data processing dialogue reply includes: target detection results, the target detection results being obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; the multiple data objects being determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; multiple features being obtained by performing feature extraction on the multiple data objects; the multiple features being used to perform multi-source anomaly detection on the multiple data objects; and the target detection results being used to assist in determining an evaluation conclusion of the target attributes on the multiple data objects; and step S53, displaying the target detection results in a graphical user interface. In the embodiments of the present disclosure, a data processing dialogue request can be understood as a data processing dialogue request initiated by a user in a conversation with an artificial intelligence. The data processing dialogue request carries identification information of a target machine and target attributes of the target machine, and is used to request anomaly detection for the target attributes of the target machine. A data processing dialogue reply can be understood as a dialogue reply from the artificial intelligence to a data processing dialogue request initiated by a user, and the data processing dialogue reply carries target detection results. In an embodiment of the present disclosure, a data processing dialogue request is obtained, and then a data processing dialogue reply is returned in response to the data processing dialogue request. The request data carried in the data processing dialogue request includes: identification information of a target machine and target attributes of the target machine. The information carried in the data processing dialogue reply includes: target detection results. The target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; the multiple data objects are determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; multiple features are obtained by performing feature extraction on the multiple data objects; the multiple features are used to perform multi-source anomaly detection on the multiple data objects; the target detection results are used to assist in determining the evaluation conclusion of the target attributes on the multiple data objects. After the target detection results are determined, the target detection results are displayed in a graphical user interface for feedback to the user. It can be seen that the present disclosure can achieve the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results of machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves the accuracy of detection, reduces the time required for data investigation, and improves the robustness of the detection solution.The data processing method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving data quality detection of machine attributes in fields such as e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, data quality detection of machine attributes in e-commerce services, data quality detection of machine attributes in educational services, and data quality detection of machine attributes in legal services are not limited here. According to an embodiment of the present disclosure, a currently input data processing session request is obtained, and then a data processing session reply is returned in response to the data processing session request. The request data carried in the data processing session request includes: identification information of a target machine and target attributes of the target machine. The information carried in the data processing session reply includes: target detection results. The target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine. The multiple data objects are determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes. The multiple features are obtained by performing feature extraction on the multiple data objects. The multiple features are used to perform multi-source anomaly detection on the multiple data objects. The target detection results are used to assist in determining the evaluation conclusion of the target attributes on the multiple data objects. After determining the target detection results, the target detection results are displayed in a graphical user interface for feedback to the user. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly, effectively improving detection accuracy, reducing data troubleshooting time, and The technical effect of improving the robustness of the detection solution solves the technical problem in related technologies of low detection accuracy and low robustness of the detection solution due to detection based on single-source data. It should be noted that the preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here. Example 4: In the operating environment of Example 1, the present disclosure provides a data processing method as shown in Figure 6.FIG6 is a flowchart of a data processing method according to Embodiment 4 of the present disclosure. As shown in FIG6 , the method includes: step S61, in response to an analysis instruction applied on an operation interface, displaying on the operation interface a plurality of data objects associated with a target attribute of a target machine, wherein the plurality of data objects are obtained by performing a data lineage relationship analysis based on identification information of the target machine and the target attribute of the target machine; step S62, in response to an extraction instruction applied on the operation interface, displaying on the operation interface a plurality of features, wherein the plurality of features are obtained by performing feature extraction on a plurality of data objects, and the plurality of features are used to perform multi-source anomaly detection on the plurality of data objects; step S63, in response to a detection instruction applied on the operation interface, displaying on the operation interface a target detection result, wherein the target detection result is obtained by performing data quality detection on a plurality of data objects based on the plurality of features, and the target detection result is used to assist in determining an evaluation conclusion of the target attribute on the plurality of data objects. In the embodiments of the present disclosure, an analysis instruction can be understood as an instruction triggered and generated by a user performing an analysis operation in an operation interface; an extraction instruction can be understood as an instruction triggered and generated by a user performing an extraction operation in an operation interface; and a detection instruction can be understood as an instruction triggered and generated by a user performing a detection operation in an operation interface. In the embodiments of the present disclosure, in response to an analysis instruction executed on the operation interface, multiple data objects associated with a target attribute of a target machine can be displayed on the operation interface. These multiple data objects are obtained by performing a data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine. Subsequently, in response to an extraction instruction executed on the operation interface, multiple features can be displayed on the operation interface. These multiple features are obtained by performing feature extraction on multiple data objects and are used to perform multi-source anomaly detection on the multiple data objects. Subsequently, in response to a detection instruction executed on the operation interface, target detection results can be displayed on the operation interface. These target detection results are obtained by performing data quality detection on the multiple data objects based on the multiple features and are used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. It can be seen that the present disclosure can achieve the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results of machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves the accuracy of detection, reduces the time required for data investigation, and improves the robustness of the detection solution.The data processing method provided in the embodiments of the present disclosure can be applied, but is not limited to, to application scenarios involving data quality detection of machine attributes in fields such as e-commerce services, educational services, legal services, medical services, conference services, social network services, financial product services, logistics services, and navigation services. For example, data quality detection of machine attributes in e-commerce services, data quality detection of machine attributes in educational services, and data quality detection of machine attributes in legal services are not limited here. According to an embodiment of the present disclosure, in response to an analysis instruction issued on an operation interface, multiple data objects associated with a target attribute of a target machine can be displayed on the operation interface. The multiple data objects are obtained by performing a data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine. Subsequently, in response to an extraction instruction issued on the operation interface, multiple features can be displayed on the operation interface. The multiple features are obtained by performing feature extraction on multiple data objects and are used to perform multi-source anomaly detection on the multiple data objects. Subsequently, in response to a detection instruction issued on the operation interface, target detection results can be displayed on the operation interface. The target detection results are obtained by performing data quality detection on multiple data objects based on the multiple features. The target detection results are used to assist in determining the evaluation conclusions of the target attributes on the multiple data objects. This achieves the purpose of performing data quality anomaly detection on physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data troubleshooting time, and enhances the robustness of the detection solution. This solves the technical problem in related technologies of low accuracy of detection results and low robustness of detection solutions caused by single-source data detection. It should be noted that the preferred implementation of this embodiment can be found in the relevant description of Example 1 and will not be repeated here. Example 5 According to an embodiment of the present disclosure, an embodiment of an apparatus for implementing the above-mentioned data processing method is also provided.FIG7 is a schematic structural diagram of a data processing device according to Embodiment 5 of the present disclosure. As shown in FIG7 , the device includes: an analysis module 701 configured to perform data lineage relationship analysis based on the identification information of a target machine and the target attribute of the target machine, and determine multiple data objects associated with the target attribute; an extraction module 702 configured to extract features from multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; and a detection module 703 configured to perform data quality detection on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. Optionally, the analysis module 701 is further configured to: perform data lineage relationship analysis based on the identification information and the target attribute to match at least one target data source from multiple candidate data sources; and obtain multiple data objects from the at least one target data source using a preset functional dependency calculation method, wherein the preset functional dependency calculation method is used to discover a set of dependency items in the at least one target data source that satisfies a preset functional dependency relationship. Optionally, the extraction module 702 is further configured to: determine influence features of the multiple data objects based on the number of downstream data objects affected by the multiple data objects; determine access features of the multiple data objects based on the number of queries for the multiple data objects; determine validity features of the multiple data objects based on the lineage distances between the multiple data objects and the target detection data, the lineage distances between the multiple data objects and the data source of the target detection data, and whether the multiple data objects are located in the lineage graph of the data source; and determine support features of the multiple data objects based on the support of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency. Optionally, the extraction module 702 further includes: a processing module configured to normalize the influence features, access features, and validity features using a preset normalization method to obtain a processing result. Optionally, the detection module 703 is further configured to: perform data quality detection on the multiple data objects based on the processing result and the support features to obtain a target detection result. Optionally, the above-mentioned detection module 703 is also configured to: perform weight calculation on the processing results according to a first weight calculation method to obtain first weights corresponding to multiple data objects; perform weight calculation on the first weight and the support feature according to a second weight calculation method to obtain second weights of target dependencies corresponding to multiple data objects on each component of the target attribute, wherein the second weight calculation method is different from the first weight calculation method; perform data quality detection on multiple data objects based on the second weight to obtain target detection results.Optionally, the detection module 703 is further configured to: obtain the weights of each data object associated with the target dependency on each component of the target attribute based on the second weight to obtain a first voting result; based on the first voting result, aggregate the weights of each data object associated with the target dependency on each component of the target attribute to obtain a second voting result; and perform data quality testing on multiple data objects using the second voting result to obtain a target detection result. Optionally, the detection module 703 is further configured to: calculate the vote percentage corresponding to each component of the target attribute using the second voting result; and determine the target detection result based on the maximum vote percentage. Optionally, the system further includes a rendering module configured to mark and render the target detection result on a target lineage graph to obtain a rendering result. The target lineage graph is a lineage graph containing at least some of the multiple data objects. The rendering result is used to assist in locating abnormal and normal states of the target attribute in the multiple data objects and / or to determine the anomaly category corresponding to the abnormal state. According to the disclosed embodiments, data lineage relationship analysis is performed based on the identification information of a target machine and the target attributes of the target machine to determine multiple data objects associated with the target attributes. Feature extraction is then performed on the multiple data objects from multiple dimensions to obtain multiple features in different dimensions. Data quality detection, i.e., multi-source anomaly detection, is then performed on the multiple data objects based on the multiple features to obtain target detection results. This target detection result can then be used to assist in determining the evaluation conclusion of the target attribute on multiple data objects, thereby determining whether the target attribute has an anomaly, the location of the anomaly, and the approximate cause of the anomaly. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This addresses the technical issues in related technologies related to low detection accuracy and low robustness of detection solutions caused by single-source data detection. It should be noted that the above analysis module 701, extraction module 702 and detection module 703 correspond to steps S21 to S23 in Example 1. The examples and application scenarios implemented by the three modules and the corresponding steps are the same, but are not limited to the contents disclosed in the above Example 1.It should be noted that the above modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, 102n). The above modules may also be part of a device and run in the computer terminal 10 provided in Example 1. According to an embodiment of the present disclosure, another device embodiment for implementing the above data processing method is also provided. FIG8 is a schematic structural diagram of another data processing device according to Embodiment 5 of the present disclosure. As shown in FIG8 , the device includes: a first acquisition module 801, configured to acquire a data processing request through a first application programming interface; and a first return module 802, configured to return a data processing response through a second application programming interface. The request data carried in the data processing request includes: identification information of a target machine and target attributes of the target machine; and the response data carried in the data processing response includes: target detection results, obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; multiple data objects determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; multiple features obtained by performing feature extraction on the multiple data objects; the multiple features are used to perform multi-source anomaly detection on the multiple data objects; and the target detection results are used to assist in determining evaluation conclusions of the target attributes on the multiple data objects. According to an embodiment of the present disclosure, a data processing request carrying identification information and target attributes of a target machine is obtained through a first application programming interface (API). A data processing response carrying target detection results is then returned through a second application programming interface (API). The target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine. The multiple data objects are determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes. The multiple features are obtained by performing feature extraction on the multiple data objects. The multiple features are used to perform multi-source anomaly detection on the multiple data objects. The target detection results are used to assist in determining the evaluation conclusions of the target attributes on the multiple data objects. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing data quality anomalies, and determining the possible causes of the anomalies. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This solves the technical problem in related technologies of low detection accuracy and low robustness caused by detection based on single-source data.It should be noted that the first acquisition module 801 and the first return module 802 correspond to steps S41 and S42 in Example 2. The examples and application scenarios implemented by these two modules and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, 102n). The above modules may also be part of an apparatus and run in the computer terminal 10 provided in Example 1. According to an embodiment of the present disclosure, another apparatus embodiment for implementing the above data processing method is also provided. FIG9 is a schematic structural diagram of another data processing device according to Embodiment 5 of the present disclosure. As shown in FIG9 , the device includes: a second acquisition module 901, configured to acquire a currently input data processing session request; a second return module 902, configured to return a data processing session reply in response to the data processing session request; wherein the request data carried in the data processing session request includes: identification information of a target machine and target attributes of the target machine; and information carried in the data processing session reply includes: target detection results, obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; multiple data objects determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; multiple features obtained by performing feature extraction on multiple data objects; the multiple features are used to perform multi-source anomaly detection on the multiple data objects; and the target detection results are used to assist in determining an evaluation conclusion of the target attributes on the multiple data objects; and a display module 903, configured to display the target detection results in a graphical user interface.According to an embodiment of the present disclosure, a currently input data processing session request is obtained, and then a data processing session reply is returned in response to the data processing session request. The request data carried in the data processing session request includes: identification information of a target machine and target attributes of the target machine. The information carried in the data processing session reply includes: target detection results. The target detection results are obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine. The multiple data objects are determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes. The multiple features are obtained by performing feature extraction on the multiple data objects. The multiple features are used to perform multi-source anomaly detection on the multiple data objects. The target detection results are used to assist in determining the evaluation conclusion of the target attributes on the multiple data objects. After determining the target detection results, the target detection results are displayed in a graphical user interface for feedback to the user. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly, effectively improving detection accuracy, reducing data troubleshooting time, and The technical effect of improving the robustness of the detection scheme solves the technical problem in related technologies of low detection accuracy and low robustness of the detection scheme due to detection based on single-source data. It should be noted that the second acquisition module 901, the second return module 902, and the display module 903 correspond to steps S51 to S53 in Example 3. The examples and application scenarios implemented by these three modules and the corresponding steps are the same, but are not limited to the content disclosed in Example 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, 102n). The above modules can also be part of a device and run in the computer terminal 10 provided in Example 1. According to the embodiments of the present disclosure, another device embodiment for implementing the above data processing method is also provided.FIG10 is a schematic structural diagram of another data processing device according to Embodiment 5 of the present disclosure. As shown in FIG10 , the device includes: a first response module 1001, configured to respond to an analysis instruction on an operation interface and display, on the operation interface, multiple data objects associated with a target attribute of a target machine, wherein the multiple data objects are obtained by performing a data lineage relationship analysis based on identification information of the target machine and the target attribute of the target machine; a second response module 1002, configured to respond to an extraction instruction on the operation interface and display, on the operation interface, multiple features, wherein the multiple features are obtained by performing feature extraction on multiple data objects, and the multiple features are used to perform multi-source anomaly detection on the multiple data objects; and a third response module 1003, configured to respond to a detection instruction on the operation interface and display, on the operation interface, a target detection result, wherein the target detection result is obtained by performing data quality detection on multiple data objects based on the multiple features, and the target detection result is used to assist in determining an evaluation conclusion of the target attribute on the multiple data objects. According to an embodiment of the present disclosure, in response to an analysis instruction issued on an operation interface, multiple data objects associated with a target attribute of a target machine can be displayed on the operation interface. The multiple data objects are obtained by performing a data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine. Subsequently, in response to an extraction instruction issued on the operation interface, multiple features can be displayed on the operation interface. The multiple features are obtained by performing feature extraction on multiple data objects and are used to perform multi-source anomaly detection on the multiple data objects. Subsequently, in response to a detection instruction issued on the operation interface, target detection results can be displayed on the operation interface. The target detection results are obtained by performing data quality detection on multiple data objects based on the multiple features. The target detection results are used to assist in determining the evaluation conclusions of the target attributes on the multiple data objects. This achieves the purpose of performing data quality anomaly detection on physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data troubleshooting time, and enhances the robustness of the detection solution. This solves the technical problem in related technologies of low accuracy of detection results and low robustness of detection solutions caused by single-source data detection. It should be noted that the first response module 1001, the second response module 1002, and the third response module 1003 correspond to steps S61 to S63 in Example 4. The examples and application scenarios implemented by these three modules and the corresponding steps are the same, but are not limited to those disclosed in Example 1.It should be noted that the above-mentioned modules or units may be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, 102n). The above-mentioned modules may also be executed as part of a device in the computer terminal 10 provided in Example 1. It should be noted that the preferred implementation schemes involved in the above-mentioned embodiments of the present disclosure are the same as the solution provided in Example 1, as well as the application scenarios and implementation processes, but are not limited to the solution provided in Example 1. Example 6 The embodiments of the present disclosure may provide a computer terminal, which may be any computer terminal device in a computer terminal group. Optionally, in this embodiment, the above-mentioned computer terminal may be replaced by a terminal device such as a mobile terminal. Optionally, in this embodiment, the above-mentioned computer terminal may be located in at least one of multiple network devices in a computer network. In this embodiment, the computer terminal can execute program code for the following steps in the data processing method: performing data lineage relationship analysis based on the identification information of the target machine and the target attributes of the target machine to determine multiple data objects associated with the target attributes; performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; performing data quality testing on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining the evaluation conclusion of the target attributes on the multiple data objects. Optionally, Figure 11 is a block diagram of a computer terminal according to an embodiment of the present disclosure. As shown in Figure 11, the computer terminal A may include: one or more (only one is shown) processors 1102, a memory 1104, a storage controller, and a peripheral interface, wherein the peripheral interface is connected to a radio frequency module, an audio module, and a display. The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the data processing method and apparatus in the embodiments of the present disclosure. The processor executes the stored software programs and modules to execute various functional applications and data processing, thereby implementing the aforementioned data processing method. The memory may include high-speed random access memory (RAM) and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory may further include memory located remotely from the processor, and such remote memory may be connected to the computer terminal A via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.The processor can call information and applications stored in the memory through the transmission device to execute the following steps: performing data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine to determine multiple data objects associated with the target attribute; performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; performing data quality detection on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. Optionally, the processor can also execute program code for the following steps: performing data lineage relationship analysis based on the identification information and the target attribute to match at least one target data source from multiple candidate data sources; and obtaining multiple data objects from the at least one target data source using a preset functional dependency calculation method, wherein the preset functional dependency calculation method is used to discover a set of dependency items in the at least one target data source that satisfies a preset functional dependency relationship. Optionally, the processor may further execute program code for the following steps: determining influence features of the multiple data objects based on the number of downstream data objects affected by the multiple data objects; determining access features of the multiple data objects based on the number of queries for the multiple data objects; determining validity features of the multiple data objects based on the lineage distances between the multiple data objects and the target detection data, the lineage distances between the multiple data objects and the data source of the target detection data, and whether the multiple data objects are located in the lineage graph of the data source; and determining support features of the multiple data objects based on the support of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency. Optionally, the processor may further execute program code for the following steps: normalizing the influence features, access features, and validity features using a preset normalization method to obtain a processing result. Optionally, the processor may further execute program code for the following steps: performing data quality testing on the multiple data objects based on the processing result and the support features to obtain a target detection result. Optionally, the processor may also execute the program code of the following steps: performing weight calculation on the processing results according to a first weight calculation method to obtain first weights corresponding to multiple data objects; performing weight calculation on the first weight and the support feature according to a second weight calculation method to obtain second weights of target dependencies corresponding to multiple data objects on each component of the target attribute, wherein the second weight calculation method is different from the first weight calculation method; performing data quality detection on multiple data objects based on the second weight to obtain target detection results.Optionally, the processor may further execute program code for the following steps: obtaining, based on the second weight, the weights of each data object associated with the target dependency on each component of the target attribute to obtain a first voting result; summarizing, based on the first voting result, the weights of each data object associated with the target dependency on each component of the target attribute to obtain a second voting result; and performing data quality testing on multiple data objects using the second voting result to obtain a target detection result. Optionally, the processor may further execute program code for the following steps: calculating, using the second voting result, the vote percentage corresponding to each component of the target attribute; and determining a target detection result based on the maximum vote percentage. Optionally, the processor may further execute program code for the following steps: marking and rendering the target detection result on a target lineage graph to obtain a rendering result, wherein the target lineage graph is a lineage graph containing at least some of the multiple data objects, and the rendering result is used to assist in locating abnormal and normal states of the target attribute in the multiple data objects and / or to determine the abnormality category corresponding to the abnormal state. According to the disclosed embodiments, data lineage relationship analysis is performed based on the identification information of a target machine and the target attributes of the target machine to determine multiple data objects associated with the target attributes. Feature extraction is then performed on the multiple data objects from multiple dimensions to obtain multiple features in different dimensions. Data quality detection, i.e., multi-source anomaly detection, is then performed on the multiple data objects based on the multiple features to obtain target detection results. This target detection result can then be used to assist in determining the evaluation conclusion of the target attribute on multiple data objects, thereby determining whether the target attribute has an anomaly, the location of the anomaly, and the approximate cause of the anomaly. This achieves the purpose of performing anomaly detection on the data quality of physical machine attributes based on multi-source data, thereby accurately obtaining data quality detection results for machine attributes, locating the location causing the data quality anomaly, and determining the possible cause of the anomaly. This effectively improves detection accuracy, reduces data investigation time, and enhances the robustness of the detection solution. This addresses the technical issues in related technologies related to low detection accuracy and low robustness of detection solutions caused by single-source data detection. Those skilled in the art will appreciate that the structure shown in FIG11 is merely illustrative, and that the computer terminal A may also be a smartphone (e.g., an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal device. FIG11 does not limit the structure of the aforementioned electronic devices.For example, computer terminal A may include more or fewer components (such as a network interface, a display device, etc.) than those shown in FIG11 , or may have a configuration different from that shown in FIG11 . Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware associated with the terminal device through a program. The program can be stored in a computer-readable storage medium, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Example 7 The present disclosure also provides a computer-readable storage medium. Optionally, in this embodiment, the computer-readable storage medium can be used to store the program code executed by the data processing method provided in Example 1. Optionally, in this embodiment, the computer-readable storage medium can be located in any computer terminal in a computer terminal group in a computer network, or in any mobile terminal in a mobile terminal group. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing data lineage relationship analysis based on the identification information of the target machine and the target attribute of the target machine to determine multiple data objects associated with the target attribute; performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; performing data quality detection on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining the evaluation conclusion of the target attribute on the multiple data objects. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing data lineage relationship analysis based on the identification information and the target attribute to match at least one target data source from multiple candidate data sources; and obtaining multiple data objects from the at least one target data source using a preset functional dependency calculation method, wherein the preset functional dependency calculation method is used to discover a set of dependency items in the at least one target data source that satisfies a preset functional dependency relationship.Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: determining influence features of the multiple data objects based on the number of downstream data objects affected by the multiple data objects; determining access features of the multiple data objects based on the number of queries for the multiple data objects; determining validity features of the multiple data objects based on the lineage distances between the multiple data objects and the target detection data, the lineage distances between the multiple data objects and the data source of the target detection data, and whether the multiple data objects are located in the lineage graph of the data source; and determining support features of the multiple data objects based on the support of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: normalizing the influence features, access features, and validity features using a preset normalization method to obtain a processing result. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing data quality testing on the multiple data objects based on the processing result and the support features to obtain a target detection result. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: performing weight calculation on the processing results according to a first weight calculation method to obtain first weights corresponding to multiple data objects; performing weight calculation on the first weights and support features according to a second weight calculation method to obtain second weights of target dependencies corresponding to the multiple data objects on each component of the target attribute, wherein the second weight calculation method is different from the first weight calculation method; performing data quality testing on the multiple data objects based on the second weights to obtain target detection results. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: obtaining weights of each data object associated with the target dependency on each component of the target attribute based on the second weights to obtain a first voting result; summarizing the weights of each data object associated with the target dependency on each component of the target attribute based on the first voting result to obtain a second voting result; and performing data quality testing on the multiple data objects using the second voting result to obtain a target detection result. Optionally, in this embodiment, the computer-readable storage medium is configured to store program code for executing the following steps: calculating the vote rate corresponding to each component of the target attribute using the second voting result; and determining the target detection result based on the maximum value of the vote rate.Optionally, in this embodiment, a computer-readable storage medium is configured to store program code for executing the following steps: marking and rendering target detection results on a target lineage graph to obtain a rendering result, wherein the target lineage graph is a lineage graph containing at least some of the multiple data objects, and the rendering result is used to assist in locating abnormal and normal states of target attributes in the multiple data objects, and / or determining the anomaly category corresponding to the abnormal state. Embodiments of the present disclosure also provide a computer program product comprising a computer program that, when executed by a processor, implements any of the aforementioned data processing methods. The serial numbers of the aforementioned embodiments of the present disclosure are for descriptive purposes only and do not represent the superiority or inferiority of any particular embodiment. In the aforementioned embodiments of the present disclosure, the descriptions of each embodiment are mutually exclusive. For portions not detailed in one embodiment, reference should be made to the relevant descriptions of other embodiments. In the several embodiments provided in this disclosure, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other divisions may be employed. For example, multiple units or components may be combined or integrated into another system, or some features may be omitted or not implemented. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be through interfaces, or indirect couplings or communication connections between units or modules, and may be electrical or other. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of these units may be selected to achieve the objectives of the present embodiments as needed. Furthermore, the functional units in the various embodiments of the present disclosure may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. These integrated units may be implemented in either hardware or software functional units. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium.Based on this understanding, the technical solution of this disclosure, or the portion that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product, stored in a storage medium, includes instructions for causing a computer device (such as a personal computer, server, or network device) to execute all or part of the steps of the methods described in various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), removable hard drives, magnetic disks, or optical disks. The above description is merely a preferred embodiment of this disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of this disclosure, and such improvements and modifications should also be considered within the scope of protection of this disclosure.

Claims

Claims 1. A data processing method, comprising: Performing data lineage relationship analysis based on identification information of the target machine and target attributes of the target machine to determine multiple data objects associated with the target attributes; Performing feature extraction on the multiple data objects to obtain multiple features, wherein the multiple features are used to perform multi-source anomaly detection on the multiple data objects; performing data quality detection on the multiple data objects based on the multiple features to obtain target detection results, wherein the target detection results are used to assist in determining an evaluation conclusion of the target attribute on the multiple data objects.

2. The data processing method according to claim 1, wherein: Performing a data lineage relationship analysis based on the identification information and the target attribute to determine the multiple data objects associated with the target attribute includes: performing a data lineage relationship analysis based on the identification information and the target attribute to match at least one target data source from multiple candidate data sources; and obtaining the multiple data objects from the at least one target data source using a preset functional dependency calculation method, wherein the preset functional dependency calculation method is used to discover a set of dependency items in the at least one target data source that satisfies a preset functional dependency relationship.

3. The data processing method according to claim 1, wherein: Extracting features from the multiple data objects to obtain the multiple features includes at least part of the following: determining influence features of the multiple data objects based on the number of downstream data objects affected by the multiple data objects; Determining access volume characteristics of the plurality of data objects based on the number of times the plurality of data objects are queried; determining validity features of the plurality of data objects based on lineage relationship distances between the plurality of data objects and target detection data, lineage relationship distances between the plurality of data objects and a data source of the target detection data, and whether the plurality of data objects are located in a lineage graph of the data source; Support features of the multiple data objects are determined based on the supports of the multiple data objects in the minimum approximate non-trivial functional dependency of the approximate dependency.

4. The data processing method according to claim 3, wherein: The data processing method further includes: A preset normalization processing method is adopted to perform normalization processing on the influence feature, the pageview feature, and the effectiveness feature to obtain a processing result.

5. The data processing method according to claim 4, wherein: Performing data quality detection on the multiple data objects according to the multiple features to obtain the target detection result includes: performing data quality detection on the multiple data objects according to the processing result and the support feature to obtain the target detection result.

6. The data processing method according to claim 5, wherein: Performing data quality detection on the multiple data objects based on the processing results and the support features to obtain the target detection results includes: performing weight calculation on the processing results according to a first weight calculation method to obtain first weights corresponding to the multiple data objects; performing weight calculation on the first weight and the support features according to a second weight calculation method to obtain second weights of target dependencies corresponding to the multiple data objects on each component of the target attribute, wherein the second weight calculation method is different from the first weight calculation method; performing data quality detection on the multiple data objects based on the second weight to obtain the target detection result.

7. The data processing method according to claim 6, wherein: Performing data quality testing on the multiple data objects based on the second weight to obtain the target detection result includes: obtaining the weights of each data object associated with the target dependency on each component of the target attribute based on the second weight to obtain a first voting result; summarizing the weights of each data object associated with the target dependency on each component of the target attribute based on the first voting result to obtain a second voting result; and performing data quality testing on the multiple data objects using the second voting result to obtain the target detection result.

8. The data processing method according to claim 7, wherein: Performing data quality detection on the multiple data objects using the second voting result to obtain the target detection result includes: calculating the vote rate corresponding to each component of the target attribute using the second voting result; and determining the target detection result based on a maximum value of the vote rates.

9. The data processing method according to claim 1, wherein: The data processing method further comprises: Marking and rendering the target detection result on a target lineage graph to obtain a rendering result, wherein the target lineage graph is a lineage graph containing at least some of the multiple data objects, and the rendering result is used to assist in locating abnormal and normal states of the target attribute in the multiple data objects and / or to determine the abnormality category corresponding to the abnormal state. A data processing method, comprising: obtaining a data processing request through a first application programming interface; A data processing response is returned through a second application programming interface; wherein, the request data carried in the data processing request includes: identification information of the target machine and target attributes of the target machine, and the response data carried in the data processing response includes: target detection results, the target detection results are obtained after data quality detection is performed on multiple data objects based on multiple features associated with the target attributes of the target machine, the multiple data objects are determined after data lineage relationship analysis is performed based on the identification information of the target machine and the target attributes, the multiple features are obtained after feature extraction on the multiple data objects, the multiple features are used to perform multi-source anomaly detection on the multiple data objects, and the target detection results are used to assist in determining the evaluation conclusions of the target attributes on the multiple data objects. , a data processing method, comprising: Get the current input data processing session request; In response to the data processing dialogue request, a data processing dialogue reply is returned; wherein the request data carried in the data processing dialogue request includes: identification information of a target machine and target attributes of the target machine; and information carried in the data processing dialogue reply includes: target detection results, the target detection results being obtained by performing data quality detection on multiple data objects based on multiple features associated with the target attributes of the target machine; the multiple data objects being determined by performing data lineage relationship analysis based on the identification information of the target machine and the target attributes; the multiple features being obtained by performing feature extraction on the multiple data objects; the multiple features being used to perform multi-source anomaly detection on the multiple data objects; and the target detection results being used to assist in determining an evaluation conclusion of the target attributes on the multiple data objects; and the target detection results being displayed in a graphical user interface. An electronic device, comprising: A memory storing an executable program; A processor is configured to run the program, wherein the program executes the data processing method according to any one of claims 1 to 11 when running.

13. A computer-readable storage medium, comprising a stored executable program, wherein when the executable program is executed, the device containing the computer-readable storage medium is controlled to execute the data processing method according to any one of claims 1 to 11.

14. A computer program product, comprising a computer program, wherein when executed by a processor, the computer program implements the data processing method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Network attack immune defense method and system adopting multi-module learning

    CN115664784A

  • Knowledge graph data fusion method and system

    CN115952862A

Cited By

  • Network technology service fault diagnosis and early warning system based on multi-source data fusion

    CN122226584A