Data analysis method and device, electronic equipment, medium and program product

By clustering differential data and analyzing based on knowledge graphs, the problem of low data analysis efficiency is solved, and the reason information of differential data is quickly obtained, which reduces the analysis threshold and duration.

CN119939156APending Publication Date: 2025-05-06CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411998095.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

In the prior art, data analysis efficiency is low, and technicians need to manually analyze the reasons for different data, which leads to excessive time and relies on experience, which increases the analysis time.

Method used

By clustering the differential data, the target data group is obtained, and the difference data representing the target data group is analyzed based on the knowledge graph to obtain the reason information of the differential data, reducing the number of analysis and improving efficiency.

Benefits of technology

There is no need to analyze all the differential data, and reason information can be obtained through clustering and knowledge graph analysis, which lowers the threshold for data analysis, reduces the analysis time, and improves data analysis efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119939156A_ABST
    Figure CN119939156A_ABST
Patent Text Reader

Abstract

The invention provides a data analysis method and device, electronic equipment, a medium and a program product, and can be applied to the field of data processing. The method comprises the steps of performing data comparison on multiple groups of data pairs to obtain multiple pieces of first difference data, and clustering the first difference data to obtain a multi-cluster target data group; obtaining first difference data representing the target data set from the first difference data of the target data set as second difference data; performing problem analysis on the second difference data based on a knowledge graph to obtain first reason information for generating the second difference data; the first reason information of the second difference data is determined as reason information for generating third difference data, and the third difference data are first difference data in a target data set where the second difference data are located. According to the method, the data analysis efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a data analysis method, device, electronic device, medium and program product. Background Art

[0002] In order to ensure the consistency of data in various data systems, data with the same data identifier is stored in multiple data systems, and it is necessary to compare the data with the same data identifier in various data systems.

[0003] After data comparison, a large amount of difference data will be generated, and the reasons for the generation of transaction data need to be analyzed. In the exemplary technology, the reasons for the generation of difference data are manually analyzed by technicians, which leads to a long time required for data analysis, and the lack of experience of technicians will also increase the time for data analysis, resulting in a problem of low data analysis efficiency. Summary of the invention

[0004] The present application provides a data analysis method, device, electronic device, medium and program product, which solve the problem of low data analysis efficiency.

[0005] In a first aspect, the present application provides a data analysis method, comprising:

[0006] Performing data comparison on multiple data pairs to obtain multiple first difference data, and clustering each of the first difference data to obtain multiple cluster target data groups, each data pair including two data with the same data identifier;

[0007] acquiring, from each of the first difference data of the target data group, first difference data representing the target data group as second difference data;

[0008] Performing problem analysis on each of the second difference data based on the knowledge graph to obtain first reason information for generating the second difference data, wherein the knowledge graph is constructed based on each of the historical difference data and the second reason information for generating the historical difference data;

[0009] The first reason information of the second difference data is determined as the reason information for generating each third difference data, and each third difference data is each first difference data in the target data group where the second difference data is located.

[0010] In some embodiments, the performing problem analysis on each of the second difference data based on the knowledge graph includes:

[0011] constructing a difference feature vector corresponding to the second difference data;

[0012] Determine a target feature vector according to the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph;

[0013] A target word vector is determined according to the target feature vector, and first reason information of the second difference data is determined based on the target word vector.

[0014] In some embodiments, determining the target feature vector according to the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph includes:

[0015] Determine the similarity between the difference feature vector and each node feature vector in the knowledge graph, and determine the graph feature vector of the knowledge graph based on each node feature vector in the knowledge graph;

[0016] A matching weight vector is constructed according to each of the similarities, and the target feature vector is determined according to the graph feature vector and the matching weight vector.

[0017] In some embodiments, determining the target feature vector according to the graph feature vector and the matching weight vector includes:

[0018] According to the association relationship between the nodes in the knowledge graph, the matching weight vector is updated to obtain a diffused matching weight vector;

[0019] The target feature vector is determined according to the product of the diffused matching weight vector and the graph feature vector.

[0020] In some embodiments, the determining the first reason information of the second difference data based on the target word vector includes:

[0021] Determine a problem description word vector matching the target feature vector as a first word vector, and determine a difference cause word vector matching the target feature vector as a second word vector, wherein the target word vector includes the first word vector and the second word vector;

[0022] Convert the first word vector into text to obtain problem description information, and convert the second word vector into text to obtain difference reason information;

[0023] The problem description information and the difference reason information are determined as the first reason information of the second difference data.

[0024] In some embodiments, constructing a difference feature vector corresponding to the second difference data includes:

[0025] Acquire attribute information of the second difference data, and construct a basic feature vector according to the attribute information;

[0026] Constructing a difference data word vector corresponding to the second difference data, and performing semantic extraction on the difference data word vector to obtain a semantic feature vector;

[0027] A difference feature vector corresponding to the second difference data is constructed according to the basic feature vector and the semantic feature vector.

[0028] In some embodiments, constructing a difference data word vector corresponding to the second difference data includes:

[0029] Determine a target field in the second difference data, where the target field includes at least one of a field corresponding to the identifier and a field corresponding to the private information;

[0030] In the second difference data, the target field is replaced with a preset field associated with the type of the target field to obtain intermediate data;

[0031] Construct a word vector corresponding to the intermediate data as a difference data word vector corresponding to the second difference data.

[0032] In some embodiments, clustering each of the first difference data to obtain multiple clusters of target data groups includes:

[0033] Performing multiple clustering on each of the first difference data, wherein the number of data groups obtained by each clustering is different;

[0034] Determining an evaluation parameter for each clustering according to a first number of first difference data in a data group obtained by each clustering and a second number of data groups obtained by each clustering;

[0035] The cluster corresponding to the largest evaluation parameter is determined as the target cluster, and each data group obtained from the target cluster is determined as a multi-cluster target data group.

[0036] In some embodiments, before performing problem analysis on each of the second difference data based on the knowledge graph, the method further includes:

[0037] Acquire each of the historical difference data and second reason information for generating the historical difference data;

[0038] Identifying entities in the historical difference data to obtain a plurality of problem entities, and determining association relationships between the problem entities according to the second cause information;

[0039] The nodes corresponding to the problem entities are configured in the configuration graph to obtain an intermediate graph, and connecting lines are configured for the nodes corresponding to the problem entities having association relationships in the intermediate graph to obtain a knowledge graph.

[0040] In some embodiments, the problem entity includes a problem description entity and a problem cause entity, wherein the problem description entity is used to describe the attributes of the historical difference data, and the problem cause entity is used to describe the reasons for generating the historical difference data.

[0041] In some embodiments, after determining the first reason information of the second difference data as the reason information for generating each third difference data, the method further includes:

[0042] determining a third amount of first difference data in each cluster of the target data set;

[0043] An analysis result report is generated according to each of the third quantities and the cause information of each of the first difference data, and the analysis result report is output.

[0044] In a second aspect, the present application provides a data analysis device, comprising:

[0045] A comparison module, used for performing data comparison on multiple data pairs to obtain multiple first difference data, and clustering each of the first difference data to obtain multiple cluster target data groups, each data pair including two data with the same data identifier;

[0046] An acquisition module, configured to acquire, from each first difference data of the target data group, first difference data representing the target data group as second difference data;

[0047] an analysis module, configured to perform problem analysis on each of the second difference data based on a knowledge graph to obtain first reason information for generating the second difference data, wherein the knowledge graph is constructed based on each of the historical difference data and the second reason information for generating the historical difference data;

[0048] The determination module is used to determine the first reason information of the second difference data as the reason information for generating each third difference data, and each third difference data is each first difference data in the target data group where the second difference data is located.

[0049] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory and a communication interface communicatively connected to the processor;

[0050] The communication interface is used to communicate with other communication devices;

[0051] The memory is used to store computer-executable instructions;

[0052] The processor is used to execute the computer-executable instructions stored in the memory to implement the data analysis method provided in the first aspect.

[0053] In a fourth aspect, the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the data analysis method provided in the first aspect is implemented.

[0054] In a fifth aspect, the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the data analysis method provided in the first aspect.

[0055] The data analysis method, device, electronic device, medium and program product provided by the present application compare multiple data pairs to obtain multiple first difference data, cluster each first difference data to obtain multiple clusters of target data groups, obtain the first difference data representing the target data group from each first difference data included in the target data group as the second difference data, analyze each second difference data through the knowledge graph, and obtain the first reason for the second difference data, so as to use the first reason information of the second difference data as the reason information of each first difference data in the target data group where the second difference data is located. In the present application, each difference data is clustered to obtain the target data group, and the difference data representing the target data group is analyzed through the knowledge graph to obtain the reason for the generation of each difference data in the target data group, without analyzing all the difference data in the target data group, reducing the number of difference data analyzed and improving the data analysis efficiency; in addition, the difference data is analyzed through the knowledge graph constructed by the historical difference data and the reason information for generating the historical difference data, without the need for experienced technicians to analyze, which not only reduces the cost of data analysis, but also reduces the analysis time of difference data, further improving the data analysis efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0057] Figure 1 This is a schematic diagram of the scenario involved in the data analysis method of this application;

[0058] Figure 2 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 1 ;

[0059] Figure 3 A schematic diagram of the knowledge graph involved in this application;

[0060] Figure 4 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 2 ;

[0061] Figure 5 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 3 ;

[0062] Figure 6 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 4 ;

[0063] Figure 7 A schematic diagram of the workflow of the analysis model of the embodiment of the present application;

[0064] Figure 8 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 5 ;

[0065] Fig. 9 The following is a schematic diagram of the steps of the data analysis method of the present application embodiment: Figure 6 ;

[0066] Fig.10 A schematic diagram of a program module of a data analysis device provided in an embodiment of the present application;

[0067] Fig.11 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0068] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0069] In order to make the purpose, technical scheme and advantages of the embodiments of the present application clearer, the technical scheme in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application. In addition, although the disclosure in the present application is introduced according to one or several exemplary examples, it should be understood that each aspect of these disclosures can also constitute a complete implementation method separately.

[0070] It should be noted that the brief description of terms in this application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their common and usual meanings.

[0071] In addition, the terms "include" and "have" and any variations thereof are intended to cover but not exclude inclusion, for example, a product or device comprising a list of components is not necessarily limited to those components expressly listed but may include other components not expressly listed or inherent to such products or devices.

[0072] The term "module" used in the embodiments of the present application refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware and / or software code that can perform the functions associated with the element.

[0073] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0074] In order to ensure the consistency of data in various data systems, data with the same data identifier is stored in multiple data systems, and it is necessary to compare the data with the same data identifier in various data systems.

[0075] After data comparison, a large amount of difference data will be generated, and the reasons for the generation of transaction data need to be analyzed.

[0076] The inventors of the present application have discovered that the reason why difference data is generated through manual analysis by technicians is that the time required for data analysis is too long, and the technicians’ insufficient experience will also increase the time for data analysis, thus resulting in the problem of low data analysis efficiency.

[0077] The inventors of the present application therefore thought of clustering each difference data to obtain a target data group, and analyzing the difference data representing the target data group through a knowledge graph, so as to obtain the cause of each difference data in the target data group, without having to analyze all the difference data in the target data group, thereby reducing the number of difference data analyzed and improving data analysis efficiency. In addition, by using a knowledge graph constructed using historical difference data and information on the causes of the historical difference data, the difference data can be analyzed without the need for analysis by experienced technicians, which not only reduces the entry cost of data analysis, but also reduces the analysis time of the difference data, further improving data analysis efficiency.

[0078] Reference Figure 1 , Figure 1 Schematic diagram of the scenario of the data analysis method of the present application. The data analysis device 100 obtains data with the same data identifier from the first system 200 and the second system 300 to form a data pair, that is, the data pair includes two data with the same data identifier. The data analysis device 100 performs data comparison on the data in each data pair, and the data pairs with differences will compare the first difference data, thereby obtaining multiple first difference data. The data analysis device 100 clusters each first difference data to obtain a multi-cluster target data group, and obtains the first difference data representing the target data group from each first difference data of the target data group as the second difference data. The data analysis device 100 analyzes the second difference data based on the knowledge graph to obtain the reason information for generating the second difference data, and then uses the reason information as the reason information of each first difference data of the target data group where the second difference data is located, and then outputs the reason information of each first difference data.

[0079] Combination Figure 1 , in order to explain the technical solution shown in the present application in detail through embodiments. It should be noted that the following embodiments can exist independently or in combination with each other, and the same or similar contents will not be repeatedly described in different embodiments.

[0080] Reference Figure 2 , Figure 2 The process diagram of the data analysis method provided in the embodiment of the present application is as follows Figure 1 , data analysis methods include:

[0081] Step S201 , performing data comparison on multiple data pairs to obtain multiple first difference data, and clustering each first difference data to obtain multiple cluster target data groups, each data pair including two data with the same data identifier.

[0082] In this embodiment, the execution subject is a data analysis device. For the convenience of description, the device is used to refer to the data analysis device. The device can be a server, or any physical device or virtual device with data comparison and analysis functions.

[0083] The device obtains multiple groups of data pairs, each group of data pairs includes two data with the same data identifier. Exemplarily, the device can obtain data with the same data identifier from two databases to form a data pair, and can also obtain data with the same data identifier from a database and a cache of a device to form a data pair, and can also obtain data with the same data identifier from two terminal devices to form a data pair. When the database and the terminal device need to perform data comparison, a data comparison request is sent to the device, and the device parses the data comparison request to obtain the two devices that need to perform data comparison, and obtains data with the same data identifier from the two devices to form a data pair. There are multiple data with the same data identifier in the two devices, and the device can obtain multiple groups of data pairs.

[0084] After obtaining the data pairs, the device compares each group of data pairs. If the data in the data pairs are different, the difference data corresponding to the data pairs is obtained, and the difference data of the data pairs is defined as the first difference data.

[0085] After obtaining each first difference data, cluster each first difference data. Exemplarily, each first difference data can be constructed as a feature vector, and the distance between each two feature vectors is calculated. If the distance is less than a preset distance, the first difference data of the two feature vectors corresponding to the distance are similar, and the two first difference data are regarded as data of the same type. In this way, similar first difference data can be regarded as data of the same cluster, thereby obtaining multiple clusters of data, and each cluster of data is defined as a target data group, that is, the target data group includes multiple similar first difference data.

[0086] Step S202 : acquiring, from each first difference data of the target data group, first difference data representing the target data group as second difference data.

[0087] The target data group includes multiple first difference data. The device obtains one of the first difference data from the target data group as a representative of the target data group, that is, the characteristics of the first difference data obtained from the target data group can be used to characterize the characteristics of all the first difference data of the target data group. The device can randomly obtain one first difference data from the target data group to characterize the target data group, and the obtained first difference data is used as the second difference data. Of course, the device can also obtain more than one first difference data from the target data group to characterize all the first difference data of the target data group.

[0088] Step S203: Perform problem analysis on each second difference data based on the knowledge graph to obtain first reason information for generating the second difference data. The knowledge graph is constructed based on each historical difference data and the second reason information for generating the historical difference data.

[0089] The device is equipped with a knowledge graph, which is built based on prior knowledge. Prior knowledge refers to the second reason information of historical difference data, which includes data migration failure, timestamp error, synchronization data failure, field garbled, field sequence error, encoding conversion error and dirty data modification. The knowledge graph also includes a detailed description of the cause, the cause and reference solution, refer to Figure 3 , Figure 3 A schematic diagram of the knowledge graph.

[0090] The device performs problem analysis on each second difference data based on the knowledge graph, thereby obtaining the first cause information of the second difference data. For example, the second difference data includes description information of the difference data, and the description information indicates that the difference field of the second difference data is due to different timestamps. The device can locate a detailed description matching the description information in the knowledge graph, and the detailed description points to a timestamp error, and the timestamp error points to a data migration failure. Therefore, the "cause" of the data migration failure in the knowledge graph is only the first cause information of the second difference data. In addition, both the data migration solution and the cause can be used as the first cause information of the second difference data, that is, the first cause information provides the cause and solution for the generation of the second difference data.

[0091] Step S204: determine the first reason information of the second difference data as reason information for generating each third difference data, where each third difference data is each first difference data in the target data group where the second difference data is located.

[0092] After determining the first reason information of each second difference data, the first reason information can be used as the reason information of each first difference data in the target data group where the second difference data is located, that is, the device determines the first reason information of the second difference data as the reason information for generating each third difference data, and each third difference data is each first difference data in the target data group where the second difference data is located. In this way, the reason information of each first difference data can be determined.

[0093] Furthermore, the device can also generate an analysis report based on the clustering of each first difference data and the cause information, and output the analysis report. Specifically, the device determines the third quantity of the second difference data in each cluster target data group, generates an analysis result report based on each third quantity and the cause information of each first difference data, and outputs the analysis result report.

[0094] The analysis result report is, for example:

[0095] This comparison compared XXX data, found XX difference data, and YY missing data. The analysis results are as follows:

[0096] There are 45 different data in category 1; Cause: The timestamp of missing data is unified on August 12, 2023. The possible reason is that the loading task on that day was not executed or failed; Recommended solution: Redo the batch loading task on that day;

[0097] There are 60 different data in category 2; Cause: There is a difference in the field rec_oper_index. The main reason is that the destination end lacks this field. The possible reason is that this field was newly added and was not loaded after it was added; Recommended solution: Reload rec_oper_index;

[0098] There are 80 discrepancy data in category 3; Cause: The fields pre_max_limit and trade_chan_mchid exist and discrepancies occur at the same time. The main reason is that there are discrepancies between the source and destination ends. The possible reason is that the data in the database or redis (cache) is not synchronized after modification; Recommended solution: Modify another part of the data according to the data in redis or database.

[0099] In addition, the device can also search the knowledge base for knowledge base content that matches the cause information of the first difference data, and the knowledge base content can be adjusted to serve as background knowledge for the analysis result report. The device adds the third quantity, the cause information of the first difference data, and the background knowledge to the context prompt information template, and inputs the context prompt information template filled with data into the large language model, and the large language model can output the analysis result report.

[0100] After the device displays the analysis result report, the user can input questions or instructions based on the report. The device searches the knowledge base for knowledge base information that matches the question or instruction and outputs the knowledge base information. In addition, the question and instruction can be used to update the large language model, that is, to generate a prompt context based on the knowledge base information, instructions and questions, and to update the generated content of the large language model based on the prompt context, so that the large language model in the application stage can generate an analysis result report that is more in line with the user's wishes, that is, to optimize the analysis result report.

[0101] In this embodiment, the device automatically implements the screening and integration of data comparison difference results, removes duplicate difference data, reduces the amount of data that needs to be processed and analyzed later, and improves the efficiency of difference data analysis. In addition, the device automatically analyzes difference data problems based on the knowledge graph, mines potential problems of data differences and analyzes the causes. In addition, the device integrates automatic analysis results with relevant knowledge to generate analysis result reports, answer user questions, and assist users in locating and processing difference data problems.

[0102] In this embodiment, multiple groups of data pairs are compared to obtain multiple first difference data, and each first difference data is clustered to obtain multiple clusters of target data groups. The first difference data representing the target data group is obtained as the second difference data from each first difference data included in the target data group, and each second difference data is analyzed through the knowledge graph to obtain the first reason for the second difference data, so that the first reason information of the second difference data is used as the reason information of each first difference data in the target data group where the second difference data is located. In this embodiment, each difference data is clustered to obtain the target data group, and the difference data representing the target data group is analyzed through the knowledge graph to obtain the reason for each difference data in the target data group. It is not necessary to analyze all the difference data in the target data group, which reduces the number of difference data analyzed and improves the data analysis efficiency. In addition, the difference data is analyzed through the knowledge graph constructed by the historical difference data and the reason information for generating the historical difference data, and there is no need for experienced technicians to analyze, which not only reduces the cost of data analysis, but also reduces the analysis time of the difference data, further improving the data analysis efficiency.

[0103] Reference Figure 4 , Figure 4 Schematic diagram of the data analysis method for this application Figure 2 ,based on Figure 2 In the illustrated embodiment, step S203 includes:

[0104] Step S401: construct a difference feature vector corresponding to the second difference data.

[0105] In this embodiment, the knowledge graph includes multiple nodes, and the device can perform problem analysis on the second difference data based on the similarity between the second difference data and the nodes.

[0106] Specifically, the device constructs a difference feature vector corresponding to the second difference data. Exemplarily, the second difference data includes multiple lines of difference fields, each line of difference fields is the difference between different lines in the data pair, and the device concatenates the multiple lines of difference fields and converts the concatenated fields into a difference feature vector.

[0107] Step S402, determining the target feature vector according to the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph.

[0108] There are multiple nodes in the knowledge graph, each node corresponds to a node feature vector, the device determines the similarity between the difference feature vector and the node feature vector corresponding to each node in the knowledge graph, and the target feature vector can be determined by each similarity. Exemplarily, the node feature vector corresponding to the maximum similarity is obtained as the target feature vector.

[0109] Step S403, determining a target word vector according to the target feature vector, and determining first reason information of the second difference data based on the target word vector.

[0110] The device stores a plurality of word vectors, each of which is obtained by converting the cause information of the historical difference data. After obtaining the target feature vector, the device determines the word vector associated with the target feature vector as the target word vector. After obtaining the target word vector, the device determines the first cause information of the second difference data based on the target word vector. Exemplarily, the device converts the target word vector into text to obtain the first cause information.

[0111] In this embodiment, the device accurately and quickly obtains the first cause information of the second difference data based on the similarity between the difference feature vector of the second difference data and the node feature vector of the node in the knowledge graph.

[0112] Reference Figure 5 , Figure 5 Schematic diagram of the data analysis method for this application Figure 3 ,based on Figure 4 In the illustrated embodiment, step S402 includes:

[0113] Step S501, determine the similarity between the difference feature vector and each node feature vector in the knowledge graph, and determine the graph feature vector of the knowledge graph based on each node feature vector in the knowledge graph.

[0114] Step S502, constructing a matching weight vector according to each similarity, and determining a target feature vector according to the graph feature vector and the matching weight vector.

[0115] In this embodiment, the device determines the matching weight vector based on the second difference data and the knowledge graph. Exemplarily, the device first determines the similarity between the difference feature vector and each node feature vector, and the similarity is:

[0116]

[0117] Among them, v d is the difference eigenvector, v n is the node feature vector, r d,n is the similarity between the node feature vector and the difference feature vector.

[0118] After the device obtains each similarity, it constructs a similarity matrix of the second difference data based on each similarity, and the similarity matrix is ​​used as a matching weight vector. Each node in the knowledge graph has a corresponding node feature vector. The integration of each node feature vector can obtain the integrated feature vector corresponding to the knowledge graph, and the integrated feature vector is defined as the graph feature vector.

[0119] After obtaining the matching weight vector, the device determines the target feature vector based on the matching weight vector and the graph feature vector.

[0120] In one example, the matching weight vector is a matrix, and the target feature vector can be obtained by multiplying the matrix with the graph feature vector.

[0121] In another example, there are multiple nodes in the knowledge graph, including problem description nodes and system problem nodes. The connection between the nodes represents the correlation between data difference phenomena or between phenomena and problems. The problem description node describes the characteristics and phenomena of the difference data, and the system problem node represents the possible causes of the difference data problem, such as data migration failure, timestamp error, synchronization failure, dirty data modification, etc. The device updates the matching weight vector based on the correlation between the nodes in the knowledge graph, that is, performs a diffusion search on the matching weight vector to obtain the diffused matching weight vector. Exemplarily, the matching weight vector is a similarity matrix, and the diffused matching weight vector refers to the update of each similarity as follows:

[0122]

[0123] in, is the updated similarity, d i,j is the parameter between node i and node j. When node i and node j have an association relationship, d i,j is 1, when node i has no association with node j, then d i,j is 0, n is the number of nodes in the knowledge graph, and i is the current node. The matrix formed by each updated similarity is the diffused matching weight vector. The device calculates the product between the matching weight vector and the image feature vector, and the product can be used as the target feature vector.

[0124] In this embodiment, the device accurately determines the target feature vector based on the similarities of the second difference data and the graph feature vector of the knowledge graph.

[0125] Reference Figure 6 , Figure 6 Schematic diagram of the data analysis method for this application Figure 4 ,based on Figure 4 In the illustrated embodiment, step S403 includes:

[0126] Step S601, determining a problem description word vector matching a target feature vector as a first word vector, and determining a difference cause word vector matching the target feature vector as a second word vector, the target word vector includes the first word vector and the second word vector.

[0127] In this embodiment, the device stores each problem description word vector and each difference reason word vector. The device determines the problem description word vector that matches the target feature vector in each problem description word vector. The device calculates the similarity between the target feature vector and each problem description word vector, and uses the problem description word vector corresponding to the maximum similarity as the first word vector, which is the problem description word vector that matches the target feature vector.

[0128] The device determines the difference cause word vector that matches the target feature vector in each difference cause word vector. The device calculates the similarity between the target feature vector and each difference cause word vector, and uses the difference cause word vector corresponding to the maximum similarity as the second word vector, and the second word vector is the difference cause word vector that matches the target feature vector. The target word vector includes the first word vector and the second word vector.

[0129] Step S602: Convert the first word vector into text to obtain problem description information, and convert the second word vector into text to obtain difference reason information.

[0130] Step S603: determine the problem description information and the difference reason information as the first reason information of the second difference data.

[0131] After obtaining the first word vector and the second word vector, the device converts the first word vector into text to obtain the problem description information, and converts the second word vector into text to obtain the difference reason information. The device uses the problem description information and the difference reason information as the first reason information for generating the second difference data.

[0132] Furthermore, the matrix formed by each updated similarity is used as a preliminary difference data feature matrix, which is a difference data feature description matrix. The target feature vector is further diffused through the analysis model, and the target feature vector is predicted by word vector probability to obtain the first word vector and the second word vector, so as to generate a problem description through the first word vector, and perform difference reason analysis based on the second word vector. The difference reason analysis is the difference reason information. For example, refer to Figure 7 , the preliminary difference data feature description matrix and the association relationship of the knowledge graph are input into the analysis network in the analysis model. The analysis network consists of multiple inference layers. Each inference layer first calculates the feature weight matrix W of the feature matrix of this round based on the feature description matrix output by the previous layer, and then updates the features based on the weight matrix. The feature weight matrix W integrates the input difference feature description matrix with the features of the difference data through the structure of the self-attention mechanism to obtain the feature weight matrix W. The feature weight matrix W is input into the difference feature node integration module in the analysis network for node feature diffusion processing.

[0133] In this embodiment, the calculation formula of the difference feature node integration module is as follows, where is the updated feature of node j after analysis, is the feature calculated last time for node i, w i,j is the inference weight from node i to node j (calculated by node diffusion weight), b i,j is the connection relationship between node i and node j. This module can be used to update node features, abstract differential data features, and activate and discover system problem nodes in the knowledge graph.

[0134]

[0135] After being processed by a multi-layer analysis network, the final analysis features are obtained, and then after pooling and dimensionality reduction processing, the feature integration is completed to obtain the target feature vector. The word vector probability prediction is performed through the target feature vector to obtain the first word vector and the second word vector, so as to generate the problem description through the first word vector, and the difference reason analysis is performed based on the second word vector. The difference willingness analysis is the difference reason information.

[0136] Example of problem description and analysis of causes of discrepancies:

[0137] 1. Problem description: The timestamp on the source side is later than that on the destination side. Cause analysis: Possible cause: The data migration of the primary key XXX data on the date 20XX-XX-XX failed.

[0138] 2. Problem description: Field A on the source side is a string with differences; Cause analysis: Binary display errors and character encoding conversion errors.

[0139] In this embodiment, the device determines the problem description word vector and the difference reason word vector that match the target word vector, thereby accurately determining the first reason information of the second difference data based on the problem description word vector and the difference reason word vector.

[0140] Reference Figure 8 , Figure 8 Schematic diagram of the data analysis method for this application Figure 5 ,based on Figures 4 to 6 In any of the embodiments shown in , step S401 includes:

[0141] Step S801: Acquire attribute information of the second difference data, and construct a basic feature vector according to the attribute information.

[0142] In this embodiment, the device extracts attribute information from the second difference data, and the attribute information includes information such as difference type, number of difference fields, name of difference field, source and destination data source type and encoding, etc. The device constructs a basic feature vector based on the attribute information, that is, the difference type, number of difference fields, name of difference field, source and destination data source type and encoding are used as basic features to construct the basic feature vector.

[0143] In addition, the second difference data contains information that needs to be hidden, such as the fields corresponding to the identifier and the fields corresponding to the private information, such as the mobile phone number, ID card number, etc. When the device performs data analysis, it is necessary to process the hidden information to avoid leakage of private data. The device determines the target field in the second difference data, and the target field includes at least one of the fields corresponding to the identifier and the fields corresponding to the private information. In the second difference data, the device replaces the target field with a preset field associated with the type of the target field to obtain intermediate data.

[0144] For example, when the target field is a mobile phone number, the preset field is<phone_num> , the target field before replacement is: "Mobile number": "150xxxxxxxx", and the target field after replacement is: "Mobile number":<phone_num_1> ; When the target field is ID number, the default field is<id_num> , the target field before replacement is: "ID number": "110xxxxxxxx4139", and the target field after replacement is: "ID number":<id_num_1> ; When the target field is the merchant number, the default field is<mchnt_cd> , the target field before replacement is: "Merchant Number": "00000000", and the target field after replacement is: "Merchant Number":<mchnt_cd_1> ; When the target field is the primary key, the default field is<primary_key> , the target field before replacement is: "ID number": "xx0000000", and the target field after replacement is: "primary key":<primary_key_1> After the target field in each difference field is replaced with the corresponding preset field, concatenation is performed to obtain a character string as intermediate data, and then a word vector corresponding to the intermediate data is constructed to serve as the difference data word vector corresponding to the second difference data.

[0145] Step S802: construct a difference data word vector corresponding to the second difference data, and perform semantic extraction on the difference data word vector to obtain a semantic feature vector.

[0146] The device constructs a difference data word vector corresponding to the second difference data. Specifically, the second difference data includes multiple multi-line difference fields, each difference field is concatenated to obtain a character string, the device performs word segmentation on the character string, and after processing, performs word vector conversion to obtain a difference data word vector. After the device obtains the difference data word vector, it performs semantic extraction on the difference data word vector to obtain a semantic feature vector. For example, each difference data word vector is input into a semantic extraction model to obtain each semantic feature vector.

[0147] Step S803: construct a difference feature vector corresponding to the second difference data according to the basic feature vector and the semantic feature vector.

[0148] After obtaining the basic feature vector and the semantic feature vector, the device concatenates the basic feature vector and the semantic feature vector to obtain the difference feature vector corresponding to the second difference data.

[0149] In this embodiment, the device constructs a basic feature vector based on the attribute information of the second difference data, and constructs a difference data word vector of the second difference data to obtain a semantic feature vector, thereby constructing a difference feature vector that can accurately represent the second difference data based on the basic feature vector and the semantic feature vector.

[0150] Reference Fig. 9 , Fig. 9 Schematic diagram of the data analysis method for this application Figure 6 ,based on Figures 2 to 8 In any of the embodiments shown in , step S201 includes:

[0151] Step S901 : performing multiple clustering on each first difference data, wherein the number of data groups obtained by clustering each time is different.

[0152] In this embodiment, the device performs multiple clustering on each first difference data, and the number of data groups obtained by each clustering is different. Exemplarily, when the device performs the first clustering on each first difference data, the number of cluster centers is set to 4, and the number of data groups obtained by clustering each first difference data is 4; when the device performs the first clustering on each first difference data, the number of cluster centers is set to 6, and the number of data groups obtained by clustering each first difference data is 6. In this way, the device clusters each first difference data with different numbers of cluster centers.

[0153] Step S902: determining an evaluation parameter for each clustering according to a first number of first difference data in a data group obtained by each clustering and a second number of data groups obtained by each clustering.

[0154] The device performs multiple clustering to obtain clusters with better clustering effects. Exemplarily, the device determines the evaluation parameters for each clustering based on the first number of first difference data of the data group obtained by each clustering and the second number of data groups obtained by each clustering. The evaluation parameter is calculated as follows:

[0155] W=∑w k =∑‖x k -c k ‖ 2

[0156] B=∑b k =∑n k ‖cc k ‖ 2

[0157]

[0158] Where K represents the number of clusters, n represents the number of all first difference data rows, and x k represents the difference data features in the kth cluster, c k represents the central feature of the kth cluster, n k represents the number of samples in the kth cluster, c represents the mean of the global difference data, and s is the evaluation parameter.

[0159] Step S903: determine the cluster corresponding to the maximum evaluation parameter as the target cluster, and determine each data group obtained by the target cluster as a multi-cluster target data group.

[0160] The larger the clustering evaluation parameter is, the better the clustering effect of each first difference data is. In this regard, the device selects the cluster corresponding to the largest evaluation parameter as the target cluster, and each data group obtained by the target cluster is determined as a multi-cluster target data group.

[0161] In this embodiment, the device performs multiple clustering on each first difference data, thereby selecting a data group obtained by clustering with the best clustering effect as the target data group, thereby improving the accuracy of data analysis.

[0162] In one embodiment, before the device analyzes each second difference data based on the knowledge graph, it is necessary to construct a knowledge graph. Specifically, the device obtains each historical difference data and the second reason information for generating the historical difference data; identifies the entities in the historical difference data to obtain multiple problem entities, and determines the association relationship between the problem entities based on the second reason information; configures the nodes corresponding to the problem entities in the configuration graph to obtain an intermediate graph, and configures connecting lines for the nodes corresponding to the problem entities with association relationships in the intermediate graph to obtain a knowledge graph. The problem entity includes a problem description entity and a problem cause entity. The problem description entity is used to describe the attributes of the historical difference data, and the problem cause entity is used to describe the reasons for generating the historical difference data.

[0163] In this embodiment, the device constructs a knowledge graph based on the historical difference data and the second reason information for generating the historical difference data, thereby conveniently realizing the analysis of the difference data based on the knowledge graph.

[0164] Based on the contents described in the above embodiments, a data analysis device is also provided in the embodiments of the present application. Fig.10 , Fig.10 Schematic diagram of a program module of a data analysis device provided in an embodiment of the present application. In some embodiments, the data analysis device 1000 includes:

[0165] A comparison module 1010 is used to compare multiple data pairs to obtain multiple first difference data, and cluster each first difference data to obtain multiple cluster target data groups, each data pair includes two data with the same data identifier;

[0166] An acquisition module 1020, configured to acquire first difference data representing the target data group from each first difference data of the target data group as second difference data;

[0167] An analysis module 1030 is used to perform problem analysis on each second difference data based on the knowledge graph to obtain first reason information for generating the second difference data, wherein the knowledge graph is constructed based on each historical difference data and the second reason information for generating the historical difference data;

[0168] The determination module 1040 is used to determine the first reason information of the second difference data as the reason information for generating each third difference data, where each third difference data is each first difference data in the target data group where the second difference data is located.

[0169] In some embodiments, the data analysis device 1000 is specifically used for:

[0170] Constructing a difference feature vector corresponding to the second difference data;

[0171] Determine the target feature vector based on the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph;

[0172] A target word vector is determined according to the target feature vector, and first reason information of the second difference data is determined based on the target word vector.

[0173] In some embodiments, the data analysis device 1000 is specifically used for:

[0174] Determine the similarity between the difference feature vector and each node feature vector in the knowledge graph, and determine the graph feature vector of the knowledge graph based on each node feature vector in the knowledge graph;

[0175] A matching weight vector is constructed according to each similarity, and a target feature vector is determined according to the graph feature vector and the matching weight vector.

[0176] In some embodiments, the data analysis device 1000 is specifically used for:

[0177] According to the association relationship between nodes in the knowledge graph, the matching weight vector is updated to obtain the diffused matching weight vector;

[0178] The target feature vector is determined according to the product of the diffused matching weight vector and the graph feature vector.

[0179] In some embodiments, the data analysis device 1000 is specifically used for:

[0180] Determine a problem description word vector that matches the target feature vector as a first word vector, and determine a difference cause word vector that matches the target feature vector as a second word vector, wherein the target word vector includes the first word vector and the second word vector;

[0181] The first word vector is converted into text to obtain the problem description information, and the second word vector is converted into text to obtain the difference reason information;

[0182] The problem description information and the difference cause information are determined as the first cause information of the second difference data.

[0183] In some embodiments, the data analysis device 1000 is specifically used for:

[0184] Acquire attribute information of the second difference data, and construct a basic feature vector according to the attribute information;

[0185] Constructing a difference data word vector corresponding to the second difference data, and performing semantic extraction on the difference data word vector to obtain a semantic feature vector;

[0186] A difference feature vector corresponding to the second difference data is constructed according to the basic feature vector and the semantic feature vector.

[0187] In some embodiments, the data analysis device 1000 is specifically used for:

[0188] Determine a target field in the second difference data, where the target field includes at least one of a field corresponding to the identifier and a field corresponding to the private information;

[0189] In the second difference data, the target field is replaced with a preset field associated with the type of the target field to obtain intermediate data;

[0190] Construct a word vector corresponding to the intermediate data as the difference data word vector corresponding to the second difference data.

[0191] In some embodiments, the data analysis device 1000 is specifically used for:

[0192] Performing multiple clustering on each first difference data, wherein the number of data groups obtained by each clustering is different;

[0193] Determining an evaluation parameter for each clustering according to a first number of first difference data in a data group obtained by each clustering and a second number of data groups obtained by each clustering;

[0194] The cluster corresponding to the maximum evaluation parameter is determined as the target cluster, and each data group obtained by the target cluster is determined as a multi-cluster target data group.

[0195] In some embodiments, the data analysis device 1000 is specifically used for:

[0196] Acquire each historical difference data and information on a second reason for generating the historical difference data;

[0197] Identify entities in the historical difference data to obtain multiple problem entities, and determine the association relationship between the problem entities based on the second cause information;

[0198] The nodes corresponding to the problem entities are configured in the configuration graph to obtain an intermediate graph, and connecting lines are configured for the nodes corresponding to the problem entities with associated relationships in the intermediate graph to obtain a knowledge graph.

[0199] In one embodiment, the problem entity includes a problem description entity and a problem cause entity. The problem description entity is used to describe the attributes of the historical difference data, and the problem cause entity is used to describe the reasons for generating the historical difference data.

[0200] In some embodiments, the data analysis device 1000 is specifically used for:

[0201] determining a third amount of first difference data in each cluster of target data sets;

[0202] An analysis result report is generated according to each third quantity and the cause information of each first difference data, and the analysis result report is output.

[0203] It should be noted that the various steps in the data analysis method executed by the data analysis device refer to the above embodiments for details and will not be described in detail here.

[0204] Furthermore, based on the contents described in the above embodiments, an electronic device is also provided in the embodiments of the present application, which includes at least one processor, and a communication interface and a memory connected to the processor; wherein the communication interface is used to communicate with other communication devices, and the memory stores computer execution instructions; the above at least one processor executes the computer execution instructions stored in the memory to implement each step in the data analysis method described in the above embodiments.

[0205] In order to better understand the embodiments of the present application, refer to Fig.11 , Fig.11 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present application.

[0206] like Fig.11 As shown, the electronic device 1100 of this embodiment includes: a processor 1101, a memory 1102, and a communication interface 1104; wherein:

[0207] Memory 1102, used to store computer-executable instructions;

[0208] The communication interface 1104 is used to communicate with other communication devices;

[0209] The processor 1101 is used to execute the computer-executable instructions stored in the memory to implement the various steps in the query optimization method described in the above embodiment.

[0210] Optionally, the memory 1102 may be independent or integrated with the processor 1101 .

[0211] When the memory 1102 is independently provided, the device further includes a bus 1103 for connecting the memory 1102 , the communication interface 1104 and the processor 1101 .

[0212] An embodiment of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, each step in the data analysis method described in the above embodiment is implemented.

[0213] An embodiment of the present application provides a computer program product, including a computer program. When the computer program is executed by a processor, the computer program implements the various steps in the data analysis method described in the above embodiment.

[0214] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0215] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0216] In addition, each functional module in each embodiment of the present application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The above-mentioned module-composed unit can be implemented in the form of hardware or in the form of hardware plus software functional units.

[0217] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the methods of various embodiments of the present application.

[0218] It should be understood that the processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in the application may be directly implemented as being executed by a hardware processor, or may be executed by a combination of hardware and software modules in the processor.

[0219] The memory may include high-speed memory, and may also include non-volatile storage, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk or an optical disk, etc.

[0220] The bus may be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus in the drawings of the present application is not limited to only one bus or one type of bus.

[0221] The storage medium may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, magnetic disk or optical disk. The storage medium may be any available medium that can be accessed by a general or special purpose computer.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A data analysis method, characterized in that: include: Performing data comparison on multiple data pairs to obtain multiple first difference data, and clustering each of the first difference data to obtain multiple cluster target data groups, each data pair including two data with the same data identifier; acquiring, from each of the first difference data of the target data group, first difference data representing the target data group as second difference data; Performing problem analysis on each of the second difference data based on the knowledge graph to obtain first reason information for generating the second difference data, wherein the knowledge graph is constructed based on each of the historical difference data and the second reason information for generating the historical difference data; The first reason information of the second difference data is determined as the reason information for generating each third difference data, and each third difference data is each first difference data in the target data group where the second difference data is located.

2. The method according to claim 1, characterized in that The performing problem analysis on each of the second difference data based on the knowledge graph includes: constructing a difference feature vector corresponding to the second difference data; Determine a target feature vector according to the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph; A target word vector is determined according to the target feature vector, and first reason information of the second difference data is determined based on the target word vector.

3. The method according to claim 2, characterized in that The determining of the target feature vector according to the similarity between the difference feature vector and the node feature vector corresponding to the node in the knowledge graph includes: Determine the similarity between the difference feature vector and each node feature vector in the knowledge graph, and determine the graph feature vector of the knowledge graph based on each node feature vector in the knowledge graph; A matching weight vector is constructed according to each of the similarities, and the target feature vector is determined according to the graph feature vector and the matching weight vector.

4. The method according to claim 3, characterized in that The determining the target feature vector according to the graph feature vector and the matching weight vector includes: According to the association relationship between the nodes in the knowledge graph, the matching weight vector is updated to obtain a diffused matching weight vector; The target feature vector is determined according to the product of the diffused matching weight vector and the graph feature vector.

5. The method according to claim 2, characterized in that: The determining the first reason information of the second difference data based on the target word vector includes: Determine a problem description word vector matching the target feature vector as a first word vector, and determine a difference cause word vector matching the target feature vector as a second word vector, wherein the target word vector includes the first word vector and the second word vector; Convert the first word vector into text to obtain problem description information, and convert the second word vector into text to obtain difference reason information; The problem description information and the difference reason information are determined as the first reason information of the second difference data.

6. The method according to claim 2, characterized in that The constructing a difference feature vector corresponding to the second difference data includes: Acquire attribute information of the second difference data, and construct a basic feature vector according to the attribute information; Constructing a difference data word vector corresponding to the second difference data, and performing semantic extraction on the difference data word vector to obtain a semantic feature vector; A difference feature vector corresponding to the second difference data is constructed according to the basic feature vector and the semantic feature vector.

7. The method according to claim 6, characterized in that The constructing the difference data word vector corresponding to the second difference data includes: Determine a target field in the second difference data, where the target field includes at least one of a field corresponding to the identifier and a field corresponding to the private information; In the second difference data, the target field is replaced with a preset field associated with the type of the target field to obtain intermediate data; Construct a word vector corresponding to the intermediate data as a difference data word vector corresponding to the second difference data.

8. The method according to claim 1, characterized in that The step of clustering each of the first difference data to obtain a plurality of cluster target data groups includes: Performing multiple clustering on each of the first difference data, wherein the number of data groups obtained by each clustering is different; Determining an evaluation parameter for each clustering according to a first number of first difference data in a data group obtained by each clustering and a second number of data groups obtained by each clustering; The cluster corresponding to the largest evaluation parameter is determined as the target cluster, and each data group obtained from the target cluster is determined as a multi-cluster target data group.

9. The data analysis method according to claim 1, characterized in that: Before performing problem analysis on each of the second difference data based on the knowledge graph, the method further includes: Acquire each of the historical difference data and second reason information for generating the historical difference data; Identifying entities in the historical difference data to obtain a plurality of problem entities, and determining association relationships between the problem entities according to the second cause information; The nodes corresponding to the problem entities are configured in the configuration graph to obtain an intermediate graph, and connecting lines are configured for the nodes corresponding to the problem entities having association relationships in the intermediate graph to obtain a knowledge graph.

10. The method according to claim 9, characterized in that The problem entity includes a problem description entity and a problem cause entity. The problem description entity is used to describe the attributes of the historical difference data, and the problem cause entity is used to describe the reason for generating the historical difference data.

11. The method according to any one of claims 1 to 10, characterized in that After determining the first reason information of the second difference data as the reason information for generating each third difference data, the method further includes: determining a third amount of first difference data in each cluster of the target data set; An analysis result report is generated according to each of the third quantities and the cause information of each of the first difference data, and the analysis result report is output.

12. A data analysis device, characterized in that: include: A comparison module, used for comparing multiple data pairs to obtain multiple first difference data, and clustering each of the first difference data to obtain multiple clusters of target data groups, each data pair including two data with the same data identifier; An acquisition module, configured to acquire, from each first difference data of the target data group, first difference data representing the target data group as second difference data; an analysis module, configured to perform problem analysis on each of the second difference data based on a knowledge graph to obtain first reason information for generating the second difference data, wherein the knowledge graph is constructed based on each of the historical difference data and the second reason information for generating the historical difference data; The determination module is used to determine the first reason information of the second difference data as the reason information for generating each third difference data, and each third difference data is each first difference data in the target data group where the second difference data is located.

13. An electronic device, characterized in that: include: A processor, and a memory and a communication interface communicatively connected to the processor; The communication interface is used to communicate with other communication devices; The memory is used to store computer-executable instructions; The processor is used to execute the computer-executable instructions stored in the memory to implement the data analysis method according to any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the data analysis method according to any one of claims 1 to 11 is implemented.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the data analysis method according to any one of claims 1 to 11 is implemented.