Customer information collection method and system based on multi-source data fusion

By establishing directed graphs and data clustering techniques, the problem of repetitive data collection in multi-source environments for enterprises was solved, improving information collection efficiency and data utilization, and enhancing data traceability.

CN120910038BActive Publication Date: 2025-12-30HANGZHOU LIUDU ENTERPRISE MANAGEMENT CONSULTING CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511439885.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-10
Publication Date
2025-12-30
Estimated Expiration
2045-10-10

AI Technical Summary

Technical Problem

In existing technologies, enterprise data collection efficiency is low, and there are problems such as waste of storage resources due to duplicate data and insufficient use of information. In particular, the lack of data integration in multi-source environments leads to low information collection efficiency and poor results.

Method used

By establishing a directed graph, we can obtain the propagation path of data and the responsibility weight of nodes, perform data clustering, filter out high-quality information, avoid duplicate data collection, and improve the traceability and usability of data.

Benefits of technology

This effectively avoids duplicate data collection, improves the efficiency of information collection and data utilization, and enhances data traceability and query efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910038B_ABST
    Figure CN120910038B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of data processing, and particularly relates to a customer information collection method and system based on multi-source data fusion, which comprises the following steps: obtaining a directed graph of several kinds of information, and obtaining several kinds of information in each time period; obtaining a responsibility interval of any kind of information according to the relative quality of data in all kinds of information of each node in each time period, and obtaining the responsibility weight of any kind of information of each node; clustering all kinds of information to obtain several information clusters; obtaining the data quality of real-time data according to the node correlation of corresponding information of real-time data and each information cluster, and the relative quality of data of the same node of real-time data, and completing information collection. The application can effectively collect and store customer information, and improve the effectiveness of information collection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology. More specifically, this invention relates to a method and system for collecting customer information based on multi-source data fusion. Background Technology

[0002] In the digital age, a vast amount of information is presented in the form of data. To effectively collect and analyze information, data collection is essential. For example, in enterprises, large amounts of data are scattered across the computers of all employees, each computer acting as a data source. This not only results in significant redundancy, incurring substantial data storage and management costs, but also leads to waste due to the dispersed nature of the data and its inability to be fully utilized. Therefore, to reduce data storage and management costs and improve data utilization, it is necessary to utilize information management platforms and other methods to organize the data. Furthermore, new data should be collected and summarized in the same way, creating a virtuous cycle.

[0003] To collect customer information from multiple data sources for enterprises, quality assessment indicators are commonly used to evaluate data quality, thereby filtering the data and obtaining high-quality information. For example, Chinese patent document CN113779150B discloses a data quality assessment method and apparatus, which discloses extracting data samples using different extraction rules and determining data quality based on the deviation of the data quality assessment values ​​of different data samples. Storing high-quality data under limited storage space can effectively reduce data storage and management costs and improve data utilization.

[0004] However, enterprises often experience significant data duplication. The more processes and personnel the same data flows through, the higher the duplication rate. Repeatedly collecting repetitive data not only reduces information gathering efficiency but also wastes storage resources. While sampling assessment can reduce resource waste, it cannot completely eliminate the problems of low information gathering efficiency and poor results caused by repetitive data. Furthermore, due to the lack of a data integration process, the problem of information not being fully utilized persists. Summary of the Invention

[0005] To address the aforementioned technical problems of insufficient efficiency in customer information collection and poor utilization of collected information, this invention provides solutions in the following aspects.

[0006] In a first aspect, the present invention provides a customer information collection method based on multi-source data fusion, comprising:

[0007] Acquire several data points from several nodes within a certain time period, and extract metadata from any data, wherein the metadata contains several data items;

[0008] Based on the transmission direction of data at different nodes and the correlation of metadata between data, all data are organized in units of time periods to obtain directed graphs of several types of information, and several types of information for each time period are obtained. The directed graphs of various types of information are traversed, and the relative quality of each data in each type of information is obtained according to the order of the traversal process. Based on the relative quality of the data of each node in all types of information in each time period, the responsibility interval of any type of information is obtained, and the responsibility weight of each node for any type of information is obtained. Based on the responsibility interval and responsibility weight of the node to which the data of each type of information belongs in each time period, all types of information are clustered to obtain several information clusters. Based on the correlation between the real-time data and the nodes of each information cluster, and the relative quality of the data of the same nodes as the real-time data, the data quality of the real-time data is obtained, and the information collection is completed.

[0009] This invention uses a directed graph to acquire information, enabling a clearer understanding of data propagation. By clustering information based on propagation nodes, it improves the efficiency of information quality assessment and provides effective reference for real-time data evaluation, thus facilitating effective collection and storage. Furthermore, by defining the responsibility intervals for information and the responsibility weights of nodes, this invention enhances data traceability. When enterprises query data, they can identify the most relevant data of the highest quality, thereby improving operational efficiency and the effectiveness of information collection.

[0010] Preferably, the directed graph that acquires several types of information acquires several types of information for each time period, including:

[0011] Take any time period as the target time period, any node as the target node, and any data of the target node as the target data. Create a directed graph of the target data's corresponding information. Obtain the relevant data of the target data of the nodes corresponding to the vertices with an out-degree greater than 0 in the directed graph of the target data's corresponding information. Together with the target data, this is called the target information. Take any data that is not target information within the target time period as new target data and create a directed graph of the type information of the new target data. Continue this process until any data within the target time period has a directed graph of its type information. This will give you several types of information for the target time period, as well as a directed graph for each type of information.

[0012] This invention establishes a directed graph of information, avoiding database redundancy caused by duplicate data and repeated information collection, thereby improving the accuracy of information collection.

[0013] Preferably, the directed graph for creating the target data corresponding to the information includes:

[0014] Get the data in other nodes within the target time period that has the same file name as the target data or the file name before modification, and record it as the relevant data of the target data. Record the corresponding node as the relevant node of the target node. Treat the target node and the relevant node of the target node as vertices of the directed graph of the information corresponding to the target data, and connect them with the target node to form the edge of the directed graph of the target data. The direction of the edge points to the vertex with the largest reception time of the corresponding data among the two vertices connected by the edge.

[0015] For each vertex in the directed graph of the target data, obtain the relevant nodes corresponding to the nodes and complete the edge connection until all the relevant nodes corresponding to the nodes of all vertices are vertices in the directed graph of the target data, thus obtaining the directed graph of the target data information.

[0016] Preferably, obtaining the relative quality of each data point in various information includes:

[0017] Perform topological sorting on the directed graph of the i-th type of information to obtain the data sequence of the i-th type of information;

[0018] The relative quality of the e-th data in the i-th type of information satisfies the expression:

[0019] ;

[0020] In the formula, This represents the relative quality of the e-th data point in the i-th type of information; This represents the index of the e-th data point of the i-th type of information within the data sequence of the i-th type of information; Represents the set of all data sequences for the i-th type of information; Represents the maximum value function; This represents the normalization function.

[0021] This invention distinguishes data of the same type of information by comparing the relative quality of all data in the same information, thus avoiding the collection of erroneous or omitted data and affecting the use of the data.

[0022] Preferably, the step of obtaining the responsibility interval for any type of information and obtaining the responsibility weight of each node for any type of information includes:

[0023] If the relative quality of the p-th data of the k-th node is greater than 0, then the type of information to which the p-th data belongs is recorded as the responsibility interval of the k-th node; the relative quality of the data of the q-th type of information in the responsibility interval of the k-th node is recorded as the responsibility weight of the k-th node for the q-th type of information.

[0024] This invention determines the responsibility range of nodes and the responsibility weight of nodes for information production, which can improve data traceability and enable quick retrieval of data from the information with the highest relative quality, thereby improving the effectiveness of data use and increasing work efficiency.

[0025] Preferably, the step of clustering all types of information to obtain several information clusters includes:

[0026] Sort all types of information across all time periods to obtain the information-responsibility weight sequence for each node; based on the numerical differences in the information-responsibility weight sequences of nodes, obtain the responsibility correlation between any two types of information.

[0027] The negative correlation normalization result of the responsibility correlation between any two pieces of information is used as the distance between the two pieces of information. All information is clustered based on the distance to obtain several information clusters.

[0028] This invention clusters information, enables batch processing of data, provides a reference for real-time data quality assessment, and allows for the management of information within the responsibility area of ​​a node, thereby improving the efficiency of information collection.

[0029] Preferably, the responsibility correlation of obtaining any two types of information includes:

[0030] Establish a responsibility correlation function for nodes with respect to two types of information, and denote the responsibility correlation function of the k-th node with respect to the q-th and r-th types of information as follows: ;

[0031] The responsibility correlation between the q-th type of information and the r-th type of information satisfies the expression:

[0032] ;

[0033] In the formula, This represents the responsibility correlation between the q-th type of information and the r-th type of information; K represents the number of nodes; The function representing the responsibility of the k-th node for the q-th and r-th types of information; This represents the normalization function.

[0034] Preferably, the responsibility correlation function of the k-th node for the q-th and r-th information satisfies the expression:

[0035] ;

[0036] In the formula, The function representing the responsibility of the k-th node for the q-th and r-th types of information; , This represents the responsibility weight of the k-th node for the q-th and r-th types of information; , This represents the set of responsibility weights for all nodes regarding the q-th and r-th types of information; Represents the maximum value function; Represents the absolute value function; This represents an exponential function with the natural constant as its base.

[0037] Preferably, the data quality of the acquired real-time data includes:

[0038] A directed graph that yields real-time data and corresponding information;

[0039] The average ratio of the intersection to the union of the node sets of the directed graph corresponding to the real-time data and the directed graph of all information of the c-th information cluster class, and the average ratio of the intersection to the union of the directed edge sets are obtained, multiplied together, and positively correlated normalized to obtain the probability that the real-time data corresponds to the c-th information cluster class.

[0040] The information cluster class with the highest probability is recorded as the information cluster class corresponding to the real-time data, and the average relative quality of the data that is the same as the real-time data node in the information cluster class corresponding to the real-time data is taken as the data quality of the real-time data.

[0041] Secondly, the present invention provides a customer information collection system based on multi-source data fusion, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned customer information collection method based on multi-source data fusion is implemented.

[0042] By adopting the above technical solution, the customer information collection method based on multi-source data fusion is generated into a computer program and stored in a memory so that it can be loaded and executed by a processor. This allows for the creation of a terminal device based on the memory and processor, making it convenient to use.

[0043] The beneficial effects of this invention are as follows:

[0044] (1) This invention filters data with the same information by establishing a directed graph of information, thereby avoiding the increase in data storage costs caused by the collection of duplicate data;

[0045] (2) This invention uses nodes as units to predict the data quality of real-time data from enterprise data sources, thereby improving the traceability of information, the convenience of using data, and the efficiency of data collection and storage.

[0046] (3) In this invention, information is clustered based on responsibility relevance, which enables batch collection of information and improves the efficiency of information collection. Attached Figure Description

[0047] Figure 1 This is a flowchart illustrating the customer information collection method based on multi-source data fusion in this invention;

[0048] Figure 2This is a schematic diagram illustrating the propagation of information A. Detailed Implementation

[0049] This invention discloses a customer information collection method based on multi-source data fusion, referring to... Figure 1 This includes steps S1-S4:

[0050] S1: Obtain several data points from several nodes within a certain time period, and extract metadata from any data point, wherein the metadata contains several data items.

[0051] Specifically, the process involves acquiring data from various nodes within an enterprise over several time periods and extracting metadata for each data point. These nodes are data sources such as computers, servers, and databases that store employee data. The data includes all data stored by any node within a given time period, and the metadata includes the data's filename, creation time, receipt time, modification time, and data format. It should be noted that the time periods are set by the implementers based on the actual implementation situation. For example, if the enterprise collects and summarizes information on a monthly basis, the time period can be set to one month.

[0052] Thus, we have obtained several data points from several nodes within several time periods, and extracted the metadata of any data.

[0053] S2: Based on the transmission direction of data at different nodes and the correlation of metadata between data, all data are organized in units of time period to obtain a directed graph of several types of information and obtain several types of information for each time period.

[0054] It should be noted that the acquisition, processing, and archiving of information are all carried out by staff. Taking information A as an example, if... Figure 2This diagram illustrates the propagation of information A. The arrows indicate the direction of information A's propagation. The data provider acquires information A through data sources such as sensors and computers, and propagates version 1 of information A. After data viewer a and data modifier receive the data, the data provider modifies version 1, creating version 2, and propagates it again. After data viewer a, data viewer b, and data modifier receive the data, the data modifier modifies version 2 and sends out version 3. After the data provider, data viewer a, and data viewer b receive the data, information A's propagation is complete. During this process, the computers of the data provider, data viewer, and data modifier all have varying degrees of data redundancy. Versions 1, 2, and 3 of information A are considered three pieces of valid information. If all three pieces of valid information from each person's computer are collected, the amount of duplicated data is 8. When the data propagation range is wider and modifications are more frequent, the amount of duplicated data increases even further, leading to a large amount of repeatedly collected information. If only the third version of archived information A is considered valid information, then the amount of information collected is too low, because there are differences between different versions of information A, and there is value in the differences in the data.

[0055] It should be noted that data is relatively stable, circulating and changing only between different nodes, while personnel are relatively variable, receiving data from different nodes and modifying it to varying degrees. Therefore, this invention tracks and records data versions to enable efficient information organization and collection.

[0056] It's important to note that for any data, the corresponding node could be a data provider, a data viewer, or a data modifier. Therefore, data providers, data viewers, and data modifiers can be viewed as vertices in a directed graph, and the direction of data propagation can be represented by the edges in the directed graph. If a vertex in the directed graph has an in-degree greater than 0 and an out-degree less than 0, then that vertex is a data viewer. Thus, any information exists within a directed graph that describes data propagation. Through this directed graph, we can obtain the number of versions of the information, and consequently, the valid data.

[0057] Specifically, taking any time period as the target time period, any node as the target node, and any data from the target node as the target data, a directed graph corresponding to the target data is created: Data with the same filename or original filename as the target data from other nodes within the target time period is obtained and recorded as related data of the target data; the corresponding nodes are recorded as related nodes of the target node. Both the target node and its related nodes are used as vertices in the directed graph corresponding to the target data, and edges are formed by connecting them to the target node, with the edges pointing towards the vertex with the longest receiving time among the two vertices connected by the edge. Related nodes are then obtained for each vertex in the directed graph corresponding to the target data, and the edges are connected sequentially until all related nodes of all vertices are vertices in the directed graph of the target data, thus obtaining the directed graph corresponding to the target data.

[0058] Preferably, the relevant data of the nodes corresponding to the vertices with an out-degree greater than 0 in the directed graph for which the target data is obtained are collectively referred to as target information along with the target data.

[0059] Preferably, any data that is not target information within the target time period is taken as new target data, and a directed graph of the type information of the new target data is created until any data within the target time period has a directed graph of the type information, thus obtaining several types of information for the target time period and a directed graph of each type of information.

[0060] Thus, we have obtained several types of information over several time periods, as well as a directed graph of various types of information.

[0061] S3: Traverse the directed graph of various information types, and obtain the relative quality of each data type in each information type according to the order of data in the traversal process; obtain the responsibility interval of any information type based on the relative quality of data of each node in all types of information in each time period, and obtain the responsibility weight of each node for any information type; cluster all types of information based on the responsibility interval and responsibility weight of the node to which the data of each information type belongs in each time period, and obtain several information clusters.

[0062] It's important to note that all information within a given time period represents all the information produced by the company during that period. However, for any given piece of information, not all versions of the data have high usability. Errors, omissions, and incompleteness exist within all versions of a single piece of information, requiring modification and supplementation at different nodes. Therefore, a depth-first traversal is performed in the directed graph of the information; the greater the depth, the higher the relative quality of the data. Furthermore, for all nodes, the responsibilities of different staff members vary, resulting in different relative quality of the data produced. By statistically analyzing the relative quality of all types of information for a node, we can demonstrate the importance of that node to any given information, thereby determining the responsibility range and weight of any given node. For any given information, data from nodes that fall within the node's responsibility range and have a higher responsibility weight possess higher data quality.

[0063] It's important to further clarify that for different time periods, there exists data with high periodic similarity, such as sales data for a product, where the only difference lies in the quantity across different time periods. There is also data that appears only within a single time period, such as a business process document. As time progresses, various new data and information will emerge. Calculating responsibility intervals and weights for each data point would impact information collection efficiency. Therefore, the information can be further organized by clustering it based on the responsibility intervals of the corresponding nodes, constructing the characteristics of large clusters. When new data appears, the characteristics of the new data are compared with those of the major clusters, thereby completing the information summarization.

[0064] Specifically, the directed graph of various information is traversed, and the relative quality of each data point in each information is obtained based on the order in which they are traversed, including:

[0065] Perform a topological sort on the directed graph containing the i-th type of information to obtain the data sequence for the i-th type of information. It should be noted that the topological sort ensures that earlier versions of the data always appear before later versions.

[0066] The relative quality of the e-th data in the i-th type of information satisfies the expression:

[0067] ;

[0068] In the formula, This represents the relative quality of the e-th data point in the i-th type of information; This represents the index of the e-th data point of the i-th type of information within the data sequence of the i-th type of information; Represents the set of all data sequences for the i-th type of information; Represents the maximum value function; This represents the normalization function.

[0069] At this point, the relative quality of various data points has been obtained.

[0070] Preferably, based on the relative quality of data from each node across all types of information in each time period, the responsibility interval for any type of information is obtained, and the responsibility weight of each node for any type of information is obtained, including:

[0071] If the relative quality of the p-th data of the k-th node is greater than 0, then the type of information to which the p-th data belongs is recorded as the responsibility interval of the k-th node; the relative quality of the data of the q-th type of information in the responsibility interval of the k-th node is recorded as the responsibility weight of the k-th node for the q-th type of information.

[0072] It should be noted that the information can be further organized, clustered according to the responsibility interval of the corresponding node, and the characteristics of the information of large clusters can be constructed. When new data appears, the characteristics of the new data are compared with those of the information of the major clusters to complete the classification of information. Then, the quality of the data is predicted according to the responsibility weight of the node to which the data belongs. If the information of the data has not been sufficiently modified, the data quality is considered to be insufficient, and therefore the data does not need to be collected.

[0073] It should be noted that the dissemination and modification of information have regional characteristics. For example, product sales data is generated, disseminated, and modified among sales-related personnel, and may be viewed by other personnel. Therefore, the responsibility area of ​​information has regional characteristics. Information that is disseminated within the same responsibility area and has a relatively similar distribution of data quality at each node has a stronger correlation and is more likely to be analyzed as information of the same cluster.

[0074] Preferably, based on the responsibility interval and responsibility weight of the nodes to which various types of information belong in each time period, all types of information are clustered to obtain several information clusters, including:

[0075] Sort all types of information across all time periods to obtain the information-responsibility weight sequence for each node.

[0076] The responsibility correlation between any two pieces of information satisfies the expression:

[0077] ;

[0078] ;

[0079] In the formula, This represents the responsibility correlation between the q-th type of information and the r-th type of information; K represents the number of nodes; The function representing the responsibility of the k-th node for the q-th and r-th types of information; , This represents the responsibility weight of the k-th node for the q-th and r-th types of information; , This represents the set of responsibility weights for all nodes regarding the q-th and r-th types of information; Represents the maximum value function; Represents the absolute value function; Represents the normalization function; This represents an exponential function with the natural constant as its base.

[0080] In the formula, The responsibility-related function is defined as follows: when the responsibility weight of the k-th node for both the q-th and r-th information is at its maximum value, that is... When the k-th node directly relates to the data quality of information q and r, it indicates a strong responsibility correlation between information q and r, denoted as 1. This avoids errors in calculating the responsibility correlation of information due to differences in the maximum responsibility weights of different information. If the responsibility weight of the k-th node for information q and r is not at its maximum, then the difference in the responsibility weight of the k-th node for information q and r is used to determine the relationship. Let q and r represent the responsibility correlation between information q and information r. The smaller the difference in responsibility weight of the k-th node for information q and information r, the greater the responsibility correlation between information q and information r. This means that the responsibility correlation of all nodes with respect to the q-th and r-th information is accumulated. The larger the responsibility correlation function value of each node with respect to the q-th and r-th information, the greater the responsibility correlation of the q-th and r-th information.

[0081] The negative correlation normalized between any two pieces of information is used as the distance between them. All information is then clustered based on this distance, resulting in several information clusters. It should be noted that the clustering method used can be the DBSCAN algorithm, with a minimum number of sample points of 2 and a radius equal to the distance between the two pieces of information equal to 0.1.

[0082] At this point, several information clusters have been obtained.

[0083] S4: Based on the correlation between the real-time data and the nodes of each information cluster, as well as the relative quality of the data of the same nodes as the real-time data, obtain the data quality of the real-time data and complete the information collection.

[0084] It should be noted that when assessing the quality of real-time data, comparing the directed graph of the real-time data with the directed graphs of all information clusters can identify the information clusters that are closest to the real-time data. This can serve as a reference for assessing the quality of real-time data and help determine whether the data needs to be collected and summarized.

[0085] Specifically, a directed graph that acquires information corresponding to real-time data.

[0086] Preferably, the probability that the real-time data corresponds to any information cluster class satisfies the expression:

[0087] ;

[0088] In the formula, This indicates the probability that the information corresponding to the real-time data belongs to the c-th information cluster class; This represents the number of information items in the c-th information cluster class; A set of vertices representing a directed graph that corresponds to real-time data information. The set of vertices of a directed graph representing the m-th type of information in the c-th information cluster class; A set of directed edges representing a directed graph corresponding to real-time data information. The set of directed edges in a directed graph representing the m-th type of information in the c-th information cluster class; Represents the intersection function. Represents the union function; This represents the normalization function.

[0089] In the formula, , The similarity between the directed graph representing the information corresponding to real-time data and the directed graph representing the m-th information of the c-th information cluster is as follows: the greater the proportion of identical nodes and the greater the proportion of identical directed edges, the more similar the propagation of the information corresponding to real-time data and the m-th information of the c-th information cluster is, and the more similar the data quality is. This represents the similarity between the real-time data corresponding information and the information of the c-th information cluster. The more identical nodes and the more identical directed edges there are on average, the higher the similarity between the real-time data corresponding information and the information of the c-th information cluster.

[0090] Preferably, the information cluster class with the highest probability is recorded as the information cluster class corresponding to the real-time data, and the average relative quality of the data that is the same as the real-time data node in the information cluster class corresponding to the real-time data is used as the data quality of the real-time data.

[0091] Preferably, data collection conditions are set, and data meeting these conditions is collected and stored in a hardware storage device belonging to the same information cluster or in the cloud at the same location. The collection conditions are those that select the data with the highest quality from the corresponding information. This data is used as the data to be collected, or the data quality is greater than [a certain value]. The data is the data that needs to be collected. The preset number of samples, To preset the quality threshold, , The settings are configured by the implementers based on the actual implementation situation, for example... The value can be set to 3. It can be set to 0.7.

[0092] This completes the real-time data quality assessment and information collection.

[0093] This invention also discloses a customer information collection system based on multi-source data fusion, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the customer information collection method based on multi-source data fusion according to this invention.

[0094] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.

[0095] While this specification has shown and described numerous embodiments of the invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Many modifications, alterations, and alternatives will occur to those skilled in the art without departing from the spirit and essence of the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in the practice of this invention.

Claims

1. A customer information collection method based on multi-source data fusion, characterized in that, The application relates to a method for collecting information, and belongs to the technical field of information collection. The method comprises the following steps: acquiring data of nodes in a plurality of time periods, extracting metadata of any data, wherein the metadata comprises a plurality of data items; based on the transmission direction of the data in different nodes and the correlation of the metadata between the data, arranging all the data in time periods to obtain a directed graph of a plurality of kinds of information, and obtaining a plurality of kinds of information in each time period; traversing the directed graph of each kind of information, and obtaining the relative quality of each data in each kind of information according to the sequence of the data in the information in the traversing process; obtaining the responsibility interval of any kind of information according to the relative quality of the data in all kinds of information of each node in each time period, and obtaining the responsibility weight of any kind of information of each node; clustering all kinds of information according to the responsibility interval and the responsibility weight of the nodes of the data of each kind of information in each time period to obtain a plurality of information clusters; The collection condition is to select the data with the highest quality from the corresponding information. This data is used as the data to be collected, or the data quality is greater than [a certain value]. The data is the data that needs to be collected; The preset number of samples, This is a preset quality threshold. 2.The customer information collection method based on multi-source data fusion according to claim 1, characterized in that, obtaining the data quality of real-time data according to the node correlation of the information corresponding to the real-time data and the information clusters, taking the information cluster with the maximum possibility as the information cluster corresponding to the real-time data, calculating the average relative quality of the data which is the same as the node of the real-time data in the information cluster corresponding to the real-time data, and obtaining the data quality of the real-time data; setting a collection condition, collecting the data meeting the collection condition, and storing the data in a hardware storage device belonging to the same information cluster or in a cloud at the same position to complete information collection. The method comprises the following steps: 3.The customer information collection method based on multi-source data fusion according to claim 2, characterized in that, taking any time period as a target time period, taking any node as a target node, taking any data of the target node as target data, and creating a directed graph of the information corresponding to the target data; obtaining the related data of the target data of the nodes corresponding to the vertices with the out-degree greater than 0 in the directed graph of the information corresponding to the target data, and taking the target data as target information; taking any data of the non-target information in the target time period as new target data, creating a directed graph of the information of the type to which the new target data belongs, and obtaining a plurality of kinds of information in the target time period and the directed graph of each kind of information until any data in the target time period has the directed graph of the information of the type to which the data belongs. The method comprises the following steps: obtaining the data of the remaining nodes in the target time period which is the same as the file name or the file name before modification of the target data, taking the data as the related data of the target data, and taking the nodes corresponding to the vertices as the related nodes of the target node; taking the target node and the related nodes of the target node as the vertices of the directed graph of the information corresponding to the target data, and connecting the target node to form the edges of the directed graph of the target data, wherein the direction of the edge points to the vertex corresponding to the data with the maximum receiving time among the two vertices connected by the edge; 4. The customer information collection method based on multi-source data fusion according to claim 1, characterized in that, sequentially obtaining the related nodes of the nodes corresponding to the vertices in the directed graph of the target data, and connecting the edges until the related nodes of all the nodes corresponding to the vertices are all the vertices in the directed graph of the target data, thereby obtaining the directed graph of the information corresponding to the target data. The method comprises the following steps: topologically sorting the directed graph of the i-th kind of information to obtain the data sequence of the i-th kind of information; The relative quality of the e-th data in the i-th information satisfies an expression: ; wherein represents the relative quality of the e-th data in the i-th information; represents the serial number of the e-th data in the i-th information in the data sequence of the i-th information; represents the set of the serial numbers of all data in the i-th information; represents the maximum function; represents the normalization function.

5. The customer information collection method based on multi-source data fusion according to claim 1, characterized in that, The responsibility interval of any information is obtained, and the responsibility weight of each node for any information is obtained, including: If the relative quality of the p-th data of the k-th node is greater than 0, the information to which the p-th data belongs is recorded as the responsibility interval of the k-th node; and the relative quality of the data belonging to the q-th information in the responsibility interval of the k-th node is recorded as the responsibility weight of the k-th node for the q-th information. 6.The customer information collection method based on multi-source data fusion according to claim 1, characterized in that, The clustering of all kinds of information is performed to obtain a plurality of information clusters, including: All kinds of information in all time periods are sorted to obtain an information-responsibility weight sequence of each node; and the responsibility correlation of any two kinds of information is obtained based on the numerical difference of the information-responsibility weight sequence of the node; The negative correlation normalization result of the responsibility correlation of any two kinds of information is taken as the distance between the two kinds of information, and all kinds of information are clustered according to the distance to obtain a plurality of information clusters.

7. The customer information collection method based on multi-source data fusion according to claim 6, characterized in that, The responsibility correlation of any two kinds of information is obtained, including: The responsibility correlation function of the kth node to the qth information and the rth information is denoted as ; The responsibility correlation of the q-th information and the r-th information satisfies an expression: ; wherein represents the responsibility correlation of the qth information and the rth information; K represents the number of nodes; represents the responsibility correlation function of the kth node to the qth information and the rth information; represents a normalization function. 8.The customer information collection method based on multi-source data fusion according to claim 7, characterized in that, The responsibility correlation function of the k-th node for the q-th information and the r-th information satisfies an expression: ; In the formula, represents the responsibility correlation function of the kth node to the qth information and the rth information; , represents the responsibility weight of the kth node to the qth information and the rth information; , represents the responsibility weight set of all nodes to the qth information and the rth information; represents the maximum function; represents the absolute value function; represents the exponential function with a natural constant as the base. 9.The customer information collection method based on multi-source data fusion according to claim 1, characterized in that, The data quality of the real-time data further includes: A directed graph of information corresponding to the real-time data is obtained; The possibility that the information corresponding to the real-time data belongs to the c-th information cluster is obtained by multiplying and positively correlating the mean value of the ratio of the intersection to the union of the node set of the directed graph of the information corresponding to the real-time data and the directed graph of all information of the c-th information cluster, and the mean value of the ratio of the intersection to the union of the directed edge set.

10. A customer information collection system based on multi-source data fusion, characterized in that, Including: A processor and a memory, the memory stores computer program instructions, when the computer program instructions are executed by the processor, the method for acquiring customer information based on multi-source data fusion according to any one of claims 1-9 is realized.

Citation Information

Patent Citations

  • A method and device for evaluating data quality

    CN113779150B

  • Electronic official document data security storage method

    CN118981803A

  • Agricultural product data real-time analysis method and system based on big data

    CN120317622A