Data circulation method and apparatus, and first platform and first node
By deploying a global data catalog on the data scheduling center platform and utilizing trusted computing nodes to process relevant information of data resources, the problem of low efficiency in point-to-point circulation is solved, and efficient sharing and secure circulation of multi-structured data resources are realized.
Patent Information
- Application Number
- PCT/CN2025/092822
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-22
- Filing Date
- 2025-05-06
- Publication Date
- 2025-11-27
AI Technical Summary
In existing technologies, data circulation mainly adopts a point-to-point approach, resulting in low circulation efficiency and an inability to effectively handle the analysis and circulation of multiple data resources with different structures.
A global data catalog is deployed on the data scheduling center platform side. Relevant information about data resources is received and processed through trusted computing nodes, including technologies such as identity verification, quality assessment, classification and grading, lineage analysis and data watermarking, to ensure the traceability and security of data. Data sharing is also carried out by generating data usage strategies according to cooperation agreements.
It enables efficient circulation of multiple data resources with different structures, improves data sharing efficiency, ensures data security and traceability, prevents duplicate uploads, and improves the quality of data resources.
Smart Images

Figure CN2025092822_27112025_PF_FP_ABST
Abstract
Description
Data flow method and device, first platform, and first node
[0001] Cross-reference to Related Applications
[0002] The present application claims priority to Chinese Patent Application No. 202410638917.8, filed on May 22, 2024, the contents of which are incorporated herein by reference in its entirety. TECHNICAL FIELD
[0003] The present disclosure relates to the technical field of communication, and particularly refers to a data flow method, device, first platform, and first node. BACKGROUND
[0004] Data element flow refers to a process of transferring data elements from a data holder to a data user according to certain rules. Data products include, but are not limited to, structured data resources, unstructured data resources, models, reports, etc. With strict data management requirements, different data product flow requirements are different, and high-level data products need to be circulated in a "data available but invisible" manner. The related technical solutions are as follows:
[0005] The data provider performs blood source analysis by using logs to associate data tables and supplement field annotation information by using a recognition rule library and similar tables. The data requester applies for permission from the data provider to obtain specific data use permission. This method is only suitable for blood source analysis of a single structured data resource and is not suitable for analysis of multiple structured data resources. In the related art, data flow is mainly in a point-to-point manner, and the actual flow process is inefficient. SUMMARY
[0006] The embodiments of the present disclosure aim to provide a data flow method, device, first platform, and first node to solve the problem of low flow efficiency caused by the point-to-point flow method in the related art.
[0007] To solve the above problem, the present disclosure provides a data flow method applied to a first platform, the method comprising:
[0008] receiving related information of a data resource sent by a first node;
[0009] adding the related information of the data resource to a global data directory;
[0010] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0011] The related information of the data resource also carries a quality evaluation result of the data resource, and before the related information of the data resource is added to the global data directory, the method further comprises:
[0012] auditing the identity information of the first node and the quality evaluation result of the data resource;
[0013] adding the related information of the data resource to the global data directory, comprising:
[0014] adding the related information of the data resource to the global data directory after the identity information and the quality evaluation result are both audited and passed.
[0015] The related information of the data resource also carries a quality evaluation result of the data resource, and before the related information of the data resource is added to the global data directory, the method further comprises:
[0016] The first platform sets different data circulation modes according to the classification and grading result of the data resource.
[0017] The classification and grading result of the data resource is obtained by the first node according to the classification and grading rules of the industry to which the data resource belongs.
[0018] The related information of the data resource is added to the global data directory, and the method further comprises:
[0019] The second node registered to the first platform is set to be able to access the basic information of the data resource; wherein the basic information includes at least one of the data resource name and the data resource application scenario;
[0020] The authorized second node is set to be able to access the detailed information of the data resource, wherein the detailed information includes at least one of the dictionary of the data resource, the statistical result, and the blood relationship analysis result;
[0021] The second node is a trusted computing node of a data demander.
[0022] After the related information of the data resource is added to the global data directory, the method further comprises at least one of the following:
[0023] auditing and confirming the source, generation mode, use purpose, access permission, use range, and use time of the data resource;
[0024] processing the data resource by using a block chain and / or data watermarking to ensure data traceability.
[0025] The method further comprises:
[0026] generating a data usage policy according to a cooperation agreement reached by the first node and the second node;
[0027] sending the data usage policy to the first node and / or the second node; starting a computing task and completing delivery according to the data usage policy by the first node.
[0028] The method further comprises:
[0029] receiving data resource application information reported by the first node;
[0030] receiving data resource evaluation information of the first node reported by the second node;
[0031] calculating an application score of the data resource according to the data resource application information and the evaluation information by the first platform;
[0032] deleting related information of the data resource in the global data directory in a case where the application score is lower than a preset threshold.
[0033] The method further comprises:
[0034] receiving a data query request sent by the second node, the data query request carrying an identifier of a first data resource;
[0035] feeding back a global data directory meeting requirements to the second node according to the data query request, the global data directory including basic information of the first data resource;
[0036] receiving encrypted content of detailed information of the first data resource sent by the first node according to an authorization request of the second node; wherein the authorization request of the second node is triggered by the second node in a case where the second node wants to obtain detailed information of the first data resource;
[0037] displaying the encrypted content of the detailed information of the first data resource to the second node.
[0038] The embodiments of the present disclosure further provide a data flow method applied to a first node, the method comprising:
[0039] After the first node accesses a data resource, the first node performs data dictionary extraction, data cleaning, and distributed analysis on the data resource to obtain related information of the data resource, and adds the related information of the data resource to local data resource management of the first node;
[0040] sending the related information of the data resource to a first platform, and adding the related information of the data resource to a global data directory by the first platform;
[0041] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0042] The distribution analysis of the data resource includes:
[0043] According to the data type, the data resource is divided into discrete type, continuous type, and date type.
[0044] The information of the discrete type includes at least one of frequency and missing value, the information of the continuous type includes at least one of missing number, missing rate, value range, mean value, standard deviation, minimum value, one-fourth score, median, three-fourth score, and maximum value, and the information of the date type includes at least one of earliest time and latest time.
[0045] Before the related information of the data resource is added to the local data resource management of the first node, the method further includes:
[0046] According to the blood relationship model, the blood graph of the data resource is obtained.
[0047] According to the blood relationship model, the blood graph is obtained, including:
[0048] The source index of the data resource is extracted, and the source index includes the name of the data resource, the name of the table corresponding to the data resource, and the field of the involved table.
[0049] The first graph data is output by using the blood relationship model with the name of the data resource as a keyword.
[0050] The second graph data is output by using the blood relationship model with the name of the table corresponding to the data resource as a keyword.
[0051] The third graph data is output by using the blood relationship model with the field of the involved table as a keyword.
[0052] The first graph data, the second graph data, and the third graph data are merged to obtain the blood graph.
[0053] Before the related information of the data resource is added to the local data resource management of the first node, the method further includes:
[0054] The similarity between the data resource and the data resource already contained in the local data resource management of the first node is calculated.
[0055] If the similarity is less than or equal to a threshold value, the related information of the data resource is added to the local data resource management, or if the similarity is greater than the threshold value, it is indicated that the data resource already exists in the local data resource management and the process is ended.
[0056] wherein the similarity of the data resource to the data resources already contained in the local data resource management is calculated, comprising:
[0057] According to the data dictionary of the data resource, the result of the distribution analysis of the data resource, and the blood relationship graph of the data resource, the similarity of the data resource to the data resources already contained in the local data resource management is calculated.
[0058] wherein before the relevant information of the data resource is added to the local data resource management, the method further comprises:
[0059] The data resource is output to a preset quality evaluation model to obtain a quality evaluation result of the data resource.
[0060] wherein before the relevant information of the data resource is added to the local data resource management, the method further comprises:
[0061] According to the classification and grading rules of the industry to which the data resource belongs, the data resource is classified and graded to obtain a classification and grading result of the data resource.
[0062] wherein the relevant information of the data resource sent to the first platform further carries a quality evaluation result of the data resource, and / or a classification and grading result of the data resource.
[0063] wherein after the relevant information of the data resource is added to the global data directory by the first platform, the method further comprises:
[0064] The data resource is sent to the first platform to request data authentication, and the first platform audits and confirms the source, generation method, use purpose, access permission, use range, and use time of the data resource; and the first platform processes the data resource by using a blockchain and / or data watermark to ensure data traceability.
[0065] wherein after the relevant information of the data resource is added to the global data directory by the first platform, the method further comprises:
[0066] The data usage strategy sent by the first platform is received; wherein the data usage strategy is generated by the first platform according to a cooperation agreement reached by the first node and the second node;
[0067] According to the data usage strategy, a computing task is started and completed.
[0068] The embodiments of the present disclosure also provide a data circulation device applied to a first platform, the device comprising:
[0069] The first receiving module is configured to receive related information of a data resource sent by a first node.
[0070] The first adding module is configured to add the related information of the data resource into a global data directory.
[0071] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0072] The first platform provided by the embodiment of the present disclosure also includes a processor and a transceiver.
[0073] The first receiving module is configured to receive related information of a data resource sent by a first node.
[0074] The first adding module is configured to add the related information of the data resource into a global data directory.
[0075] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0076] The data flow device provided by the embodiment of the present disclosure also includes a processor and a transceiver.
[0077] The second adding module is configured to, after accessing the data resource, perform data dictionary extraction, data cleaning, and distribution analysis on the data resource, obtain related information of the data resource, and add the related information of the data resource into local data resource management of the first node.
[0078] The first sending module is configured to send the related information of the data resource to a first platform, and the first platform adds the related information of the data resource into a global data directory.
[0079] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0080] The first node provided by the embodiment of the present disclosure also includes a processor and a transceiver.
[0081] After accessing the data resource, the processor is configured to perform data dictionary extraction, data cleaning, and distribution analysis on the data resource, obtain related information of the data resource, and add the related information of the data resource into local data resource management of the first node.
[0082] The first sending module is configured to send the related information of the data resource to a first platform, and the first platform adds the related information of the data resource into a global data directory.
[0083] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0084] The embodiment of the present disclosure further provides a communication device, comprising a memory, a processor and a program stored in the memory and executable on the processor, and the processor implements the data flow method as described above when executing the program.
[0085] The embodiment of the present disclosure further provides a computer readable storage medium, which stores a computer program, and the program is executed by a processor to implement the steps of the data flow method as described above.
[0086] The embodiment of the present disclosure further provides a computer program product, comprising computer instructions, and the computer instructions are executed by a processor to implement the steps of the method as described above.
[0087] The above technical solutions of the present disclosure have at least the following beneficial effects:
[0088] In the data flow method, the device, the first platform and the first node of the embodiment of the present disclosure, the first node generates a local data resource management, and sends the related information of the data resource to the first platform, and the first platform adds the related information of the data resource to the global data directory, so that the trusted computing node realizes data sharing through the global data directory of the first platform side, and improves the data flow efficiency. BRIEF DESCRIPTION OF DRAWINGS
[0089] Fig. 1 shows one of the step flowcharts of the data flow method provided by the embodiment of the present disclosure;
[0090] Fig. 2 shows an example diagram of the application scenario of the data flow method provided by the embodiment of the present disclosure;
[0091] Fig. 3 shows another of the step flowcharts of the data flow method provided by the embodiment of the present disclosure;
[0092] Fig. 4 shows one of the structural schematic diagrams of the data flow device provided by the embodiment of the present disclosure;
[0093] Fig. 5 shows the structural schematic diagram of the first platform provided by the embodiment of the present disclosure;
[0094] Fig. 6 shows another of the structural schematic diagrams of the data flow device provided by the embodiment of the present disclosure;
[0095] Fig. 7 shows the structural schematic diagram of the first node provided by the embodiment of the present disclosure. DETAILED DESCRIPTION
[0096] In order to make the technical problems, technical solutions and advantages to be solved by the present disclosure more clear, the following will be described in detail with reference to the drawings and specific embodiments.
[0097] As shown in FIG. 1, the data flow method provided by the embodiment of the present disclosure is applied to a first platform, and the method comprises the following steps:
[0098] In step 101, the first node sends the related information of a data resource to the first platform.
[0099] For example, as shown in FIG. 2, the first node sends the related information of a data resource to the first platform.
[0100] In step 102, the related information of the data resource is added to a global data directory.
[0101] For example, as shown in FIG. 2, after the global data directory module of the first platform adds the related information of the data resource to the global data directory, the first platform sends an indication of completing the addition of the directory to the first node, so as to inform the first node that the first platform has added the related information of the data resource sent by the first node to the global data directory.
[0102] In the embodiment of the present disclosure, the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0103] In an optional implementation, the data provider (or data owner) and the data demander respectively deploy trusted computing nodes. Optionally, the trusted computing node of the data provider is referred to as the first node, the trusted computing node of the data demander is referred to as the second node, and the data scheduling center platform is referred to as the first platform.
[0104] The embodiment of the present disclosure realizes the data sharing between the first node and the second node by arranging the global data directory on the first platform side, thereby improving the data flow efficiency.
[0105] As an optional embodiment, the related information of the data resource comprises the basic information of the data resource and / or the detailed information of the data resource.
[0106] In an implementation, the basic information comprises at least one of the data resource name and the data resource application scenario. In another implementation, the detailed information comprises at least one of the dictionary of the data resource, the statistical result, and the blood relationship analysis result.
[0107] In at least one embodiment of the present disclosure, the related information of the data resource further carries the quality evaluation result of the data resource, and before step 102, the method further comprises the following steps:
[0108] In step 201, the identity information of the first node and the quality evaluation result of the data resource are audited.
[0109] Correspondingly, step 102 comprises the following steps:
[0110] After the identity information and the quality evaluation result are both audited, the related information of the data resource is added to a global data directory.
[0111] For example, as shown in FIG. 2, the qualification audit module of the first platform audits the identity information of the first node. After the identity information is audited, the first platform checks the quality evaluation result of the data resource. If the quality evaluation result meets the requirements, the global data directory module of the first platform adds the related information of the data resource to the global data directory. If the quality evaluation result does not meet the requirements, the global data directory module of the first platform does not agree to add the related information of the data resource to the global data directory.
[0112] In another optional embodiment of the present disclosure, the related information of the data resource further carries a result of data resource classification and grading. Before step 102, the method further includes:
[0113] The first platform sets different data flow modes according to the result of data resource classification and grading.
[0114] The result of data resource classification and grading is obtained by the first node according to classification and grading rules of the industry to which the data resource belongs.
[0115] For example, when the result of data resource classification and grading is high grade, data flow is performed by using privacy calculation and the like. For another example, when the result of data resource classification and grading is low grade, data flow is performed by using an application programming interface (API) and the like.
[0116] As an optional embodiment, when the related information of the data resource is added to the global data directory, the method further includes:
[0117] The basic information of the data resource includes at least one of a data resource name and a data resource application scenario.
[0118] The authorized second node can access detailed information of the data resource, and the detailed information includes at least one of a dictionary of the data resource, a statistical result, and a blood relation analysis result.
[0119] The second node is a trusted computing node of a data demander.
[0120] In at least one optional embodiment of the present disclosure, after step 102, the method further includes at least one of the following:
[0121] The source, generation mode, use purpose, access permission, use range, and use time of the data resource are audited and confirmed.
[0122] The data resource is processed by using the blockchain and / or data watermarking to ensure data traceability.
[0123] Optionally, the first node sends an authentication request to the first platform, the data authentication module of the first platform is associated with a corresponding data resource directory, and data authentication is completed by using the blockchain and / or data watermarking. In this process, the source, generation method, use purpose and other information of the data resource need to be audited and confirmed to ensure the legality and compliance of the data resource; at the same time, the access permission, use range, use time and other information of the data also need to be audited and confirmed.
[0124] For example, the method of data authentication is as follows:
[0125] (1) The first platform encrypts the unique identification (ID) of the data resource, the unique ID of the first node, the detailed information of the watermarked data resource and the basic information of the data resource by using a hash function to obtain a hash value. Common hash functions include MD5, SHA-1, SHA-256, etc.
[0126] (2) The first node generates a public and private key, and encrypts the hash value by using the private key to generate a digital signature;
[0127] (3) The hash value, digital signature and public key of the first node and other information are submitted to the blockchain. This process establishes an unalterable authentication record, which proves that the ownership of the specific data belongs to the first node at a specific time point;
[0128] In at least one embodiment of the present disclosure, the method further comprises:
[0129] According to the cooperation agreement reached by the first node and the second node, a data use strategy is generated;
[0130] The data use strategy is sent to the first node and / or the second node; the first node starts a computing task and completes delivery according to the data use strategy.
[0131] Optionally, the data use strategy includes use duration, number of times, delivery method, etc., which are not limited here.
[0132] For example, as shown in FIG. 2, the strategy center module of the first platform generates a data use strategy according to the cooperation agreement of the first node and the second node, and delivers it to the first node and the second node. The data computing module of the first node starts a computing task and completes delivery according to the delivered data use strategy.
[0133] As an optional embodiment, the method further comprises:
[0134] receive the data resource application information reported by the first node;
[0135] receive the data resource evaluation information of the first node reported by the second node;
[0136] The first platform calculates the application score of the data resource according to the data resource application information and the evaluation information;
[0137] In the case where the application score is lower than the preset threshold value, the related information of the data resource is deleted in the global data directory, thereby ensuring the quality of the data resource.
[0138] In summary, in the generation of the global data directory, the embodiment of the present disclosure sets different access permissions, increases the policy center module, instructs the first node to complete data delivery according to the policy, increases the value evaluation step after the delivery is completed, and eliminates the data resources with poor credit, thereby ensuring the quality of the data resources. The data right confirmation method of the first platform adopts the technology of combining block chain and / or data watermark to confirm the right, which ensures the traceability of the data and improves the accuracy of the right confirmation.
[0139] In at least one embodiment of the present disclosure, the method further comprises:
[0140] receive the data query request sent by the second node, the data query request carrying the identifier of the first data resource;
[0141] According to the data query request, the global data directory meeting the requirements is fed back to the second node, and the global data directory includes the basic information of the first data resource;
[0142] receive the encrypted content of the detailed information of the first data resource sent by the first node according to the authorization request of the second node; wherein the authorization request of the second node is triggered by the second node when the second node wants to obtain the detailed information of the first data resource;
[0143] The encrypted content of the detailed information of the first data resource is displayed to the second node.
[0144] For example, when the second node initiates a global data query, the first platform feeds back the directory meeting the requirements to the second node according to the permissions of the data resource directory. After the second node selects the original directory of the data resource to view the basic information and clicks the detailed information, the first node sends the data dictionary and the result of data distribution analysis corresponding to the directory to the first platform with a watermark according to the authorization request information of the first node. The second node views the detailed information of the data resource in the first platform.
[0145] In summary, in the embodiment of the present disclosure, the first node generates a local data resource management, and sends the related information of the data resource to the first platform. The first platform adds the related information of the data resource to the global data directory, so that the trusted computing nodes realize data sharing through the global data directory on the first platform side, and improve the data flow efficiency.
[0146] As shown in FIG. 3, the embodiment of the present disclosure also provides a data flow method applied to a first node, and the method comprises the following steps:
[0147] In step 301, after the first node accesses the data resource, the first node performs data dictionary extraction, data cleaning, and distribution analysis on the data resource, obtains the related information of the data resource, and adds the related information of the data resource to the local data resource management of the first node.
[0148] Optionally, the related information of the data resource comprises basic information of the data resource and / or detailed information of the data resource. For example, the first node encrypts the detailed information of the data resource by using data watermarking and other encryption technologies, and the encrypted content and the basic information of the data resource constitute the local data resource management.
[0149] In step 302, the first platform is sent the related information of the data resource, and the first platform adds the related information of the data resource to the global data directory.
[0150] Optionally, the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0151] In an implementation manner, the data provider (or data owner) and the data demander respectively deploy trusted computing nodes. Optionally, the trusted computing node of the data provider is referred to as a first node, the trusted computing node of the data demander is referred to as a second node, and the data scheduling center platform is referred to as a first platform.
[0152] The embodiment of the present disclosure realizes data sharing between the first node and the second node by arranging the global data directory on the first platform side, thereby improving the data flow efficiency.
[0153] In an implementation manner, the basic information comprises at least one of a data resource name and a data resource application scenario. In another implementation manner, the detailed information comprises at least one of a dictionary of the data resource, a statistical result, and a blood relationship analysis result.
[0154] Optionally, in order to adapt to different formats of data resources, as shown in FIG. 2, the first node further comprises a data source access module, which connects different formats of data resources through an adaptation layer. The format of the data resource is not limited to Comma-Separated Values (CSV), MYSQL (a relational database management system), HIVE (a data warehouse tool based on Hadoop, used for data extraction, transformation, and loading), and the like, which are not specifically limited herein.
[0155] In order to improve the efficiency of data processing, optionally, different orders of magnitude of data resources use different computing frameworks. If the data volume of the data resource is greater than a threshold, a distributed cluster computing engine is started, and if the data volume is lower than the threshold, a single machine version computing engine is started.
[0156] After the first node completes the data resource access, the data dictionary module of the first node extracts the data dictionary of the data resource. Optionally, the data dictionary module can also correct the data dictionary, including whether ID, field name, field information, and the like.
[0157] After the first node completes the dictionary extraction, the data cleaning / conversion module of the first node is used to clean / convert the data resource by using a preset algorithm component.
[0158] Optionally, the distributed analysis of the data resource comprises:
[0159] According to the data type, the data resource is divided into discrete type, continuous type, and date type.
[0160] Among them, the information of the discrete type includes at least one of the frequency and the missing value, the information of the continuous type includes at least one of the missing number, the missing rate, the value range, the mean value, the standard deviation, the minimum value, the quarter fraction, the median, the three-quarter fraction, and the maximum value, and the information of the date type includes at least one of the earliest time and the latest time.
[0161] As an optional embodiment, before the related information of the data resource is added to the local data resource management of the first node in step 301, the method further comprises:
[0162] According to the blood relationship model, the blood graph of the data resource is obtained. Optionally, as shown in FIG. 2, the data blood analysis module of the first node obtains the blood relationship model to obtain the blood graph, wherein the scheme is not limited to the data blood analysis based on the Structured Query Language (SQL), the language model, the association graph, and the like. SQL is a database language with multiple functions such as data manipulation and data definition.
[0163] According to the blood relationship model, the blood relationship graph is obtained, including:
[0164] The source index of the data resource is extracted, and the source index includes: the name test of the data resource, the name userTable of the table corresponding to the data resource, and the fields X1, X2, and X3 related to the table.
[0165] The first graph data is output by using the blood relationship model with the name test of the data resource as a keyword; for example, the parent node test0 and the leaf nodes test1 and test3 of test.
[0166] The second graph data is output by using the blood relationship model with the name userTable of the table corresponding to the data resource as a keyword; for example, the parent node userTable0 and the leaf nodes userTable1 and userTable3 of userTable.
[0167] The third graph data is output by using the blood relationship model with the fields X1, X2, and X3 related to the table as a keyword.
[0168] The first graph data, the second graph data, and the third graph data are merged to obtain the blood relationship graph.
[0169] As an optional embodiment, before the relevant information of the data resource is added to the local data resource management of the first node in step 301, the method further includes:
[0170] The similarity between the data resource and the data resources already contained in the local data resource management of the first node is calculated.
[0171] If the similarity is less than or equal to a threshold value, the relevant information of the data resource is added to the local data resource management; or if the similarity is greater than the threshold value, it is indicated that the data resource already exists in the local data resource management, and the process is ended.
[0172] Optionally, the similarity between the data resource and the data resources already contained in the local data resource management includes:
[0173] According to the data dictionary of the data resource, the result of the distribution analysis on the data resource, and the blood relationship graph of the data resource, the similarity between the data resource and the data resources already contained in the local data resource management is calculated.
[0174] For example, as shown in FIG. 2, the data source similarity module uses the data dictionary, the result of the distribution analysis of the data resource, and the blood relationship diagram of the data resource as the similarity index to judge the similarity of the data resource with the directory saved locally in the first node. If the similarity is less than or equal to a threshold value, the relevant information of the data resource is added to the local data resource management; otherwise (greater than the threshold value), the user is prompted to check whether the data resource has been added.
[0175] For example, the data resource A has the blood relationship analysis index {database test, table name userTable, fields X1, X2, and X3 related to the table, parent source database train, table name comTable, and fields Y1 and Y2 related to the table}; the dictionary index {U1: user ID, T1: user gender, and T2: user age}; and the data analysis index {the null rate of T1, abnormal value, etc., the median, variance, and range of T2}.
[0176] The historical data resource B has the blood relationship analysis index {database test, table name userTable, fields X1, X2, and X3 related to the table, parent source database train, table name comTable, and fields Y1 and Y2 related to the table}; the dictionary index {U1: user ID, T1: user gender, and T2: user age}; and the data analysis index {the null rate of T1, abnormal value, etc., the median, variance, and range of T2}.
[0177] The blood relationship analysis index uses the character similarity function to calculate the similarity C1, the dictionary index uses the character similarity function to calculate the similarity C2, and the data analysis index uses the cosine similarity function to calculate the similarity C3. The similarity of the data resource A and the data resource B is equal to (C1+C2+C3) / 3=C4.
[0178] When C4 is greater than a threshold value g, the data resource A and the data resource B are one data source. If C4 is less than or equal to the threshold value g, the data resource A and the data resource B are two data resources.
[0179] In at least one embodiment of the present disclosure, before the relevant information of the data resource is added to the local data resource management in step 301, the method further includes:
[0180] Outputting the data resource to a preset quality evaluation model to obtain a quality evaluation result of the data resource.
[0181] For example, as shown in FIG. 2, the data quality evaluation module of the first node outputs the data resource to a preset quality evaluation model to obtain a quality evaluation result.
[0182] As an optional embodiment, before the relevant information of the data resource is added to the local data resource management in step 301, the method further includes:
[0183] According to the classification and grading rules of the industry to which the data resource belongs, the data resource is classified and graded to obtain a result of data resource classification and grading. For example, the result of data resource classification and grading is used to indicate the level of the data resource.
[0184] In the data classification and grading module of the embodiments of the present disclosure, the large model technology is used to automatically classify and grade according to the industry. For example, according to the classification and grading rules of the industry to which the data resource belongs, the data resource is classified and graded to obtain a result of data resource classification and grading, including:
[0185] 1) According to the description and industry of the data resource, load the standard;
[0186] 2) Convert the data classification and grading standard of the corresponding industry into items, and each item has a corresponding level. Load the word segmenter in the pre-trained large model, and convert the items into a digital sequence;
[0187] 3) Extract the dictionary of the data resource and convert it into a digital sequence using the large model. Calculate the semantic similarity between the data resource and the items, and output the corresponding data resource level.
[0188] As an optional embodiment, the related information of the data resource sent to the first platform also carries the quality evaluation result of the data resource and / or the result of data resource classification and grading.
[0189] Optionally, the first platform sets different data circulation modes according to the result of data resource classification and grading. For example, when the result of data resource classification and grading is a high level, the data circulation needs to be performed by using privacy calculation and the like. For another example, when the result of data resource classification and grading is a low level, the data circulation can be performed by using an application programming interface (API) and the like.
[0190] Optionally, the first platform audits the identity information of the first node and the quality evaluation result of the data resource. After the identity information and the quality evaluation result are both audited, the related information of the data resource is added to the global data directory.
[0191] In at least one embodiment of the present disclosure, after the first platform adds the related information of the data resource to the global data directory, the method further includes:
[0192] Sending a data right request to the first platform, auditing and confirming the source, generation mode, use purpose, access permission, use range, and use time of the data resource by the first platform, and processing the data resource by the first platform using a blockchain and / or data watermark to ensure data traceability.
[0193] For example, the first node sends an attestation request to the first platform, the data attestation module of the first platform associates a corresponding data resource directory, and data attestation is completed by using blockchain and data watermark calculation. In this process, the source, generation method, use purpose and other information of the data resource need to be audited and confirmed to ensure the legality and compliance of the data resource; at the same time, the access permission, use range, use time and other information of the data also need to be audited and confirmed.
[0194] In at least one embodiment of the present disclosure, after step 302, the method further comprises:
[0195] receiving a data use policy sent by the first platform; wherein the data use policy is generated by the first platform according to the cooperation agreement reached by the first node and the second node;
[0196] starting a computing task according to the data use policy and completing delivery.
[0197] For example, as shown in FIG. 2, the policy center module of the first platform generates a data use policy according to the cooperation agreement of the first node and the second node, and distributes it to the first node and the second node. The data computing module of the first node starts a computing task according to the distributed data use policy and completes delivery.
[0198] In the embodiment of the present disclosure, the first node generates a local data resource management, and sends the basic information of the data resource to the first platform to constitute a global data directory. Before generating the local data resource management, the first node adds a data dictionary module, a data cleaning / conversion module, a data analysis module, a data blood relationship analysis module, a data source similarity module, a data quality evaluation module, a data classification and grading module, and encrypts the detail information of the data resource with a data watermark to prevent subsequent leakage for tracking and ensure data quality and compliance with safety requirements and data resource consistency issues, and can solve the problem of repeated uploading of data resources.
[0199] In summary, the embodiment of the present disclosure adds a data dictionary module and a data statistical analysis module, data quality evaluation, and a data classification and grading module before data blood relationship analysis, compares the data source with the existing similar data source after data blood relationship analysis, encrypts the data details by permission, and performs data attestation and post-data value evaluation. The embodiment can prevent users from repeatedly adding the same data resource; and improves the usability of the data source through the dictionary, statistical analysis, quality evaluation and other modules; sets theme access permissions through the blood relationship and data classification and grading module to improve security; uses the data watermark technology to accurately trace the operation data user identity, job and leakage range and channel after data leakage occurs, and protect the security of the detail information of the customer data directory; adds data attestation to ensure data traceability; and adds a post-evaluation mechanism to eliminate poor resources and improve the quality of data resources in the global data directory.
[0200] As shown in FIG. 4, the embodiment of the present disclosure further provides a data flow device applied to a first platform, the device comprising:
[0201] A first receiving module 401 is configured to receive related information of a data resource sent by a first node;
[0202] A first adding module 402 is configured to add the related information of the data resource into a global data directory;
[0203] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0204] As an optional embodiment, the related information of the data resource further carries a quality evaluation result of the data resource, and the device further comprises:
[0205] An auditing module is configured to audit identity information of the first node and the quality evaluation result of the data resource;
[0206] The first adding module comprises:
[0207] A first adding submodule is configured to add the related information of the data resource into the global data directory after the identity information and the quality evaluation result are both audited.
[0208] As an optional embodiment, the related information of the data resource further carries a classification result of the data resource, and the method further comprises:
[0209] A first setting module is configured to set different data flow modes according to the classification result of the data resource;
[0210] The classification result of the data resource is obtained by the first node according to a classification rule of an industry to which the data resource belongs.
[0211] As an optional embodiment, the device further comprises:
[0212] A second setting module is configured to set that a second node registered to the first platform can access basic information of the data resource; the basic information comprises at least one of a data resource name and a data resource application scenario;
[0213] A third setting module is configured to set that an authorized second node can access detailed information of the data resource; the detailed information comprises at least one of a dictionary of the data resource, a statistical result, and a blood relation analysis result;
[0214] The second node is a trusted computing node of a data demander.
[0215] As an optional embodiment, the apparatus further comprises at least one of:
[0216] an auditing and confirming module configured to audit and confirm the source, generation mode, use purpose, access permission, use range, and use time of the data resource;
[0217] a data processing module configured to process the data resource by using a blockchain and / or a data watermark to ensure data traceability.
[0218] As an optional embodiment, the apparatus further comprises:
[0219] a policy generating module configured to generate a data use policy according to a cooperation agreement reached by the first node and the second node;
[0220] a policy sending module configured to send the data use policy to the first node and / or the second node, and start a computing task and complete delivery according to the data use policy by the first node.
[0221] As an optional embodiment, the apparatus further comprises:
[0222] a second receiving module configured to receive data resource application information reported by the first node;
[0223] a third receiving module configured to receive data resource evaluation information of the first node reported by the second node;
[0224] a computing module configured to calculate an application score of the data resource according to the data resource application information and the evaluation information;
[0225] a deleting module configured to delete related information of the data resource in the global data directory if the application score is lower than a preset threshold.
[0226] As an optional embodiment, the apparatus further comprises:
[0227] a fourth receiving module configured to receive a data query request sent by the second node, the data query request carrying an identifier of a first data resource;
[0228] a feedback module configured to feed back a global data directory meeting requirements to the second node according to the data query request, the global data directory including basic information of the first data resource;
[0229] a fifth receiving module configured to receive encrypted content of detailed information of the first data resource sent by the first node according to an authorization request of the second node; wherein the authorization request of the second node is triggered by the second node when the second node wants to obtain the detailed information of the first data resource.
[0230] a display module configured to display the encrypted content of the detailed information of the first data resource to the second node.
[0231] The data flow device provided by the embodiments of the present disclosure is a device capable of executing the data flow method, and all embodiments of the data flow method are applicable to the device and can achieve the same or similar beneficial effects, which will not be repeated here.
[0232] It should be noted that the data flow device provided by the embodiments of the present disclosure is a device capable of executing the data flow method, and all embodiments of the data flow method are applicable to the device and can achieve the same or similar beneficial effects, which will not be repeated here.
[0233] As shown in FIG. 5, the embodiments of the present disclosure also provide a first platform, comprising a processor 500 and a transceiver 510, wherein the transceiver 510 receives and sends data under the control of the processor 500, and the processor 500 is configured to perform the following operations:
[0234] receive the related information of the data resource sent by the first node;
[0235] add the related information of the data resource to the global data directory;
[0236] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0237] As an optional embodiment, the related information of the data resource further carries a quality evaluation result of the data resource, and the processor is further configured to perform the following operations:
[0238] audit the identity information of the first node and the quality evaluation result of the data resource;
[0239] add the related information of the data resource to the global data directory, comprising:
[0240] After the identity information and the quality evaluation result are both verified, the relevant information of the data resource is added to a global data catalog.
[0241] As an optional embodiment, the relevant information of the data resource also carries a result of data resource classification and grading, and the processor is further configured to perform the following operation:
[0242] Different data circulation modes are set according to the result of data resource classification and grading.
[0243] The result of data resource classification and grading is obtained by the first node according to classification and grading rules of an industry to which the data resource belongs.
[0244] As an optional embodiment, the processor is further configured to perform the following operation:
[0245] The basic information of the data resource is set to be accessible to the second node registered to the first platform, wherein the basic information includes at least one of a data resource name and a data resource application scenario.
[0246] The detailed information of the data resource is set to be accessible to the authorized second node, wherein the detailed information includes at least one of a data resource dictionary, a statistical result, and a blood relation analysis result.
[0247] The second node is a trusted computing node of a data demander.
[0248] As an optional embodiment, the processor is further configured to perform at least one of the following operations:
[0249] The source, generation mode, use purpose, access permission, use range, and use time of the data resource are verified and confirmed.
[0250] The data resource is processed by using a block chain and / or data watermarking to ensure data traceability.
[0251] As an optional embodiment, the processor is further configured to perform the following operation:
[0252] A data use strategy is generated according to a cooperation agreement reached by the first node and the second node.
[0253] The data use strategy is sent to the first node and / or the second node, and the first node starts a computing task and completes delivery according to the data use strategy.
[0254] As an optional embodiment, the processor is further configured to perform the following operation:
[0255] Data resource application information reported by the first node is received.
[0256] receiving data resource evaluation information of the first node reported by the second node;
[0257] calculating an application score of the data resource according to the data resource application information and the evaluation information;
[0258] deleting related information of the data resource in the global data directory in a case where the application score is lower than a preset threshold.
[0259] As an optional embodiment, the processor is further configured to perform the following operations:
[0260] receiving a data query request sent by the second node, the data query request carrying an identifier of the first data resource;
[0261] feeding a global data directory meeting a requirement to the second node according to the data query request, the global data directory including basic information of the first data resource;
[0262] receiving encrypted content of detailed information of the first data resource sent by the first node according to an authorization request of the second node, the authorization request of the second node being triggered by the second node in a case where the second node wants to acquire the detailed information of the first data resource;
[0263] displaying the encrypted content of the detailed information of the first data resource to the second node.
[0264] The embodiments of the present disclosure add a data dictionary module and a data statistical analysis module, data quality evaluation, a data classification and grading module before data bloodline analysis, add a comparison between a data source and an existing similar data source after data bloodline analysis, data detail permission and encryption, data right confirmation, and post-data value evaluation. The embodiments can prevent a user from repeatedly adding the same data resource, improve the usability of a data source through a dictionary, statistical analysis, quality evaluation, and the like, set theme access permissions through a bloodline and a data classification and grading module to improve security, use a data watermarking technology to accurately trace the operation data user identity, job, and leakage range and channel after data leakage occurs, protect the security of the data directory detail information of a client, increase data right confirmation to ensure data traceability, and increase a post-evaluation mechanism to eliminate poor resources and improve the quality of data resources in the global data directory.
[0265] It should be noted that the first platform provided by the embodiments of the present disclosure is a platform capable of executing the above-mentioned data circulation method, and all embodiments of the above-mentioned data circulation method are applicable to the platform and can achieve the same or similar beneficial effects, which will not be repeated here.
[0266] As shown in FIG. 6, the embodiment of the present disclosure further provides a data flow device applied to a first node, the device comprising:
[0267] a second adding module 601, configured to, after accessing a data resource, perform data dictionary extraction, data cleaning, and distribution analysis on the data resource, acquire related information of the data resource, and add the related information of the data resource to local data resource management of the first node;
[0268] a first sending module 602, configured to send the related information of the data resource to a first platform, and add the related information of the data resource to a global data directory by the first platform;
[0269] wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0270] As an optional embodiment, the second adding module is further configured to:
[0271] divide the data resource into discrete type, continuous type, and date type according to data types;
[0272] wherein the information of the discrete type includes at least one of frequency and missing value, the information of the continuous type includes at least one of missing number, missing rate, value range, mean value, standard deviation, minimum value, one-fourth fraction, median, three-fourth fraction, and maximum value, and the information of the date type includes at least one of earliest time and latest time.
[0273] As an optional embodiment, the device further comprises:
[0274] a blood relationship graph acquiring module, configured to acquire a blood relationship graph of the data resource according to a blood relationship model.
[0275] As an optional embodiment, the blood relationship graph acquiring module is further configured to:
[0276] extract a source index of the data resource, the source index including a name of the data resource, a name of a table corresponding to the data resource, and a field related to the table;
[0277] use the blood relationship model to output first graph data with the name of the data resource as a keyword;
[0278] use the blood relationship model to output second graph data with the name of the table corresponding to the data resource as a keyword;
[0279] use the blood relationship model to output third graph data with the field related to the table as a keyword;
[0280] merge the first graph data, the second graph data and the third graph data to obtain the blood relation graph.
[0281] As an optional embodiment, the apparatus further comprises:
[0282] a similarity calculation module configured to calculate a similarity between the data resource and a data resource already contained in the local data resource management of the first node;
[0283] a similarity processing module configured to, if the similarity is less than or equal to a threshold value, add relevant information of the data resource to the local data resource management; or, if the similarity is greater than the threshold value, indicate that the data resource already exists in the local data resource management and end the process.
[0284] As an optional embodiment, the similarity calculation module is further configured to:
[0285] calculate the similarity between the data resource and a data resource already contained in the local data resource management according to a data dictionary of the data resource, a result of distribution analysis on the data resource and a blood relation graph of the data resource.
[0286] As an optional embodiment, the apparatus further comprises:
[0287] a quality evaluation module configured to output the data resource to a preset quality evaluation model to obtain a quality evaluation result of the data resource.
[0288] As an optional embodiment, the apparatus further comprises:
[0289] a classification and grading module configured to classify and grade the data resource according to a classification and grading rule of an industry to which the data resource belongs, and obtain a result of classification and grading of the data resource.
[0290] As an optional embodiment, the relevant information of the data resource sent to the first platform further carries a quality evaluation result of the data resource and / or a result of classification and grading of the data resource.
[0291] As an optional embodiment, the apparatus further comprises:
[0292] a seventh sending module configured to send a data right request to the first platform, and have the first platform audit and confirm a source, a generation manner, a use purpose, an access right, a use range and a use time of the data resource, and have the first platform process the data resource by using a blockchain and / or a data watermark to ensure data traceability.
[0293] As an optional embodiment, the apparatus further comprises:
[0294] a seventh receiving module, configured to receive a data usage policy sent by the first platform, wherein the data usage policy is generated by the first platform according to a cooperation agreement reached by the first node and the second node;
[0295] a task computing module, configured to start a computing task and complete delivery according to the data usage policy.
[0296] The embodiment of the present disclosure adds a data dictionary module and a data statistical analysis module, data quality evaluation, a data classification and grading module before data bloodline analysis, adds comparison of a data source with an existing similar data source after data bloodline analysis, data details are divided by permission and encrypted, data rights are confirmed, and post-data value evaluation is performed. The embodiment can prevent a user from repeatedly adding the same data resource, and improves the usability of the data source through the dictionary, statistical analysis, quality evaluation and other modules. Through the bloodline and data classification and grading module, the theme access permission is set to improve the security. By using the data watermarking technology, the operation data user identity, job and leakage range and channel can be accurately traced after data leakage occurs, and the security of the customer data directory detail information is protected. The data rights are confirmed, and the data traceability is ensured. The post-evaluation mechanism is added to eliminate poor resources and improve the quality of data resources in the global data directory.
[0297] It should be noted that the data circulation device provided by the embodiment of the present disclosure is a device capable of executing the above-mentioned data circulation method, and all embodiments of the above-mentioned data circulation method are applicable to the device, and can achieve the same or similar beneficial effects, which will not be repeated here.
[0298] As shown in FIG. 7, the embodiment of the present disclosure also provides a first node, which comprises a processor 700 and a transceiver 710, the transceiver 710 receives and sends data under the control of the processor 700, and the processor 700 is configured to perform the following operations:
[0299] After accessing the data resource, data dictionary extraction, data cleaning, and distribution analysis on the data resource are performed to obtain related information of the data resource, and the related information of the data resource is added to the local data resource management of the first node;
[0300] The related information of the data resource is sent to the first platform, and the related information of the data resource is added to the global data directory by the first platform;
[0301] The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
[0302] As an optional embodiment, the processor is further configured to perform the following operations:
[0303] According to the data type, the data resource is divided into discrete type, continuous type and date type.
[0304] The discrete information includes at least one of a frequency and a missing value, the continuous information includes at least one of a missing number, a missing rate, a value range, a mean value, a standard deviation, a minimum value, a quarter fraction, a median, a three-quarter fraction and a maximum value, and the date information includes at least one of an earliest time and a latest time.
[0305] As an optional embodiment, the processor is further configured to perform the following operation:
[0306] According to the blood relationship model, a blood graph of the data resource is obtained.
[0307] As an optional embodiment, the processor is further configured to perform the following operation:
[0308] A source index of the data resource is extracted, and the source index includes a name of the data resource, a name of a table corresponding to the data resource, and a field related to the table.
[0309] The name of the data resource is used as a keyword, and the first graph data is output by using the blood relationship model.
[0310] The name of the table corresponding to the data resource is used as a keyword, and the second graph data is output by using the blood relationship model.
[0311] The field related to the table is used as a keyword, and the third graph data is output by using the blood relationship model.
[0312] The first graph data, the second graph data and the third graph data are merged to obtain the blood graph.
[0313] As an optional embodiment, the processor is further configured to perform the following operation:
[0314] Similarity between the data resource and a data resource already contained in a local data resource management of the first node is calculated.
[0315] If the similarity is less than or equal to a threshold value, relevant information of the data resource is added to the local data resource management, or if the similarity is greater than the threshold value, it is indicated that the data resource already exists in the local data resource management and the process is ended.
[0316] As an optional embodiment, the processor is further configured to perform the following operation:
[0317] According to a data dictionary of the data resource, a result of distributed analysis on the data resource and the blood graph of the data resource, similarity between the data resource and a data resource already contained in a local data resource management is calculated.
[0318] As an optional embodiment, the processor is further configured to perform the following operation:
[0319] outputting the data resource to a preset quality evaluation model to obtain a quality evaluation result of the data resource.
[0320] As an optional embodiment, the processor is further configured to perform the following operation:
[0321] classifying the data resource according to a classification rule of an industry to which the data resource belongs, to obtain a classification result of the data resource.
[0322] As an optional embodiment, the related information of the data resource sent to the first platform further carries the quality evaluation result of the data resource and / or the classification result of the data resource.
[0323] As an optional embodiment, the processor is further configured to perform the following operation:
[0324] sending a data right confirmation request to the first platform, and confirming, by the first platform, a source, a generation mode, a use purpose, an access right, a use range and a use time of the data resource, and processing the data resource by the first platform using a blockchain and / or a data watermark to ensure data traceability.
[0325] As an optional embodiment, the processor is further configured to perform the following operation:
[0326] receiving a data use strategy sent by the first platform, wherein the data use strategy is generated by the first platform according to a cooperation agreement reached by the first node and the second node.
[0327] starting a computing task according to the data use strategy and completing delivery.
[0328] The embodiments of the present disclosure add a data dictionary module and a data statistical analysis module, a data quality evaluation module, a data classification and grading module before data blood relationship analysis, add a comparison between a data source and an existing similar data source after data blood relationship analysis, data details are divided into permissions and encrypted, data right confirmation and post-data value evaluation are added. The embodiments can prevent a user from repeatedly adding the same data resource, improve the usability of a data source through a dictionary, statistical analysis, quality evaluation and other modules, set theme access permissions through a blood relationship and a data classification and grading module to improve security, use a data watermark technology to accurately trace the operation data user identity, job and leakage range and channel after data leakage occurs, protect the security of customer data directory detail information, add data right confirmation to ensure data traceability, and add a post-evaluation mechanism to eliminate poor resources and improve the quality of data resources in a global data directory.
[0329] It should be noted that the first node provided by the embodiments of the present disclosure is a node capable of performing the data flow method described above, and all embodiments of the data flow method are applicable to the node and can achieve the same or similar beneficial effects, which are not repeated here.
[0330] The embodiments of the present disclosure also provide a communication device, which is a first platform or a first node. The communication device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor. The processor implements each process of the data flow method embodiments described above when executing the program, and can achieve the same technical effects. To avoid repetition, they are not described here.
[0331] The embodiments of the present disclosure also provide a computer-readable storage medium having a computer program stored thereon. The program is executed by a processor to implement each process of the data flow method embodiments described above, and can achieve the same technical effects. To avoid repetition, they are not described here. The computer-readable storage medium includes, for example, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk.
[0332] The embodiments of the present disclosure also provide a computer program product comprising computer instructions. The computer instructions are executed by a processor to implement each process of the data flow method embodiments described above, and can achieve the same technical effects. To avoid repetition, they are not described here.
[0333] Those skilled in the art should understand that the embodiments of the present disclosure can be provided as a method, a system or a computer program product. Therefore, the present disclosure can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Moreover, the present disclosure can take the form of a computer program product implemented on one or more computer-readable storage media (including, but not limited to, disk storage and optical storage, etc.) containing computer usable program code.
[0334] The present disclosure is described with reference to flowcharts and / or block diagrams of the method, device (system) and computer program product according to the embodiments of the present disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of the flows and / or blocks in the flowcharts and / or block diagrams can be implemented by computer program instructions. These computer program instructions can be provided to a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing devices to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in one or more flows and / or blocks in the flowcharts and / or block diagrams.
[0335] These computer program instructions can also be stored in a computer- readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart or flowsheets and / or block or blocks of the block diagrams.
[0336] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart or flowsheets and / or block or blocks of the block diagrams.
[0337] The preferred embodiments of the present disclosure have been described above with the understanding that modifications and improvements can be made without departing from the scope of the present disclosure, and such modifications and improvements are also to be considered as being within the scope of the present disclosure.
Claims
1. A data flow method applied to a first platform, the method comprising: receiving information about a data resource sent by a first node; adding the information about the data resource to a global data directory; wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
2. The method of claim 1, wherein, The information about the data resource further carries a quality evaluation result of the data resource, and before adding the information about the data resource to the global data directory, the method further comprises: auditing identity information of the first node and the quality evaluation result of the data resource; adding the information about the data resource to the global data directory, comprising: adding the information about the data resource to the global data directory after the identity information and the quality evaluation result are both audited.
3. The method of claim 1, wherein, The information about the data resource further carries a classification result of the data resource, and before adding the information about the data resource to the global data directory, the method further comprises: setting different data flow modes by the first platform according to the classification result of the data resource; wherein the classification result of the data resource is obtained by the first node according to classification rules of an industry to which the data resource belongs.
4. The method of claim 1, wherein, When adding the information about the data resource to the global data directory, the method further comprises: setting that a second node registered to the first platform can access basic information of the data resource, wherein the basic information comprises at least one of a data resource name and a data resource application scenario; setting that an authorized second node can access detailed information of the data resource, wherein the detailed information comprises at least one of a dictionary of the data resource, a statistical result, and a blood relation analysis result; The second node is a trusted computing node of a data demander.
5. The method of claim 1, wherein, After adding the information about the data resource to the global data directory, the method further comprises at least one of: auditing and confirming a source, a generation mode, a use purpose, an access permission, a use range, and a use time of the data resource; processing the data resource by using a blockchain and / or a data watermark to ensure data traceability. 6.The method of any one of claims 1-5, further comprising: generating a data use strategy according to a cooperation agreement reached by the first node and the second node; sending the data use strategy to the first node and / or the second node; starting a computing task and completing delivery by the first node according to the data use strategy. 7.The method of claim 6, further comprising: receiving data resource application information reported by the first node; receiving data resource evaluation information of the first node reported by the second node; calculating an application score of the data resource by the first platform according to the data resource application information and the evaluation information; deleting the information about the data resource in the global data directory in a case where the application score is lower than a preset threshold. 8.The method of any one of claims 1-5, further comprising: receiving a data query request sent by a second node, the data query request carrying an identifier of a first data resource; according to the data query request, feeding back a global data directory meeting a requirement to the second node, the global data directory including basic information of the first data resource; receiving encrypted content of detailed information of the first data resource sent by the first node according to an authorization request of the second node, the authorization request of the second node being triggered by the second node in a case where the second node wants to acquire the detailed information of the first data resource; displaying the encrypted content of the detailed information of the first data resource to the second node.
9. A data flow method applied to a first node, the method comprising: after the first node accesses a data resource, performing data dictionary extraction, data cleaning, and distribution analysis on the data resource to obtain related information of the data resource, and adding the related information of the data resource to local data resource management of the first node; sending the related information of the data resource to a first platform, and adding the related information of the data resource to a global data directory by the first platform; wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
10. The method of claim 9, wherein, the distribution analysis on the data resource comprises: dividing the data resource into discrete type, continuous type, and date type according to data types; wherein the information of the discrete type includes at least one of frequency and missing value, the information of the continuous type includes at least one of missing number, missing rate, value range, mean value, standard deviation, minimum value, one-fourth fraction, median, three-fourth fraction, and maximum value, and the information of the date type includes at least one of earliest time and latest time.
11. The method of claim 9, wherein, before adding the related information of the data resource to the local data resource management of the first node, the method further comprises: obtaining a blood relationship graph of the data resource according to a blood relationship model.
12. The method of claim 11, wherein, obtaining the blood relationship graph according to the blood relationship model comprises: extracting a source index of the data resource, the source index including a name of the data resource, a name of a table corresponding to the data resource, and a field related to the table; using the name of the data resource as a keyword, outputting first graph data by using the blood relationship model; using the name of the table corresponding to the data resource as a keyword, outputting second graph data by using the blood relationship model; using the field related to the table as a keyword, outputting third graph data by using the blood relationship model; merging the first graph data, the second graph data, and the third graph data to obtain the blood relationship graph.
13. The method of claim 9, wherein, before adding the related information of the data resource to the local data resource management of the first node, the method further comprises: calculating a similarity between the data resource and a data resource already contained in the local data resource management of the first node; if the similarity is less than or equal to a threshold value, adding the related information of the data resource to the local data resource management, or if the similarity is greater than the threshold value, indicating that the data resource already exists in the local data resource management and ending the process.
14. The method of claim 13, wherein, calculating the similarity between the data resource and the data resource already contained in the local data resource management comprises: According to the data dictionary of the data resource, the result of the distribution analysis on the data resource, and the blood relationship diagram of the data resource, the similarity of the data resource and the data resource already contained in the local data resource management is calculated.
15. The method of claim 9, wherein, Before adding the related information of the data resource into the local data resource management, the method further comprises: outputting the data resource to a preset quality evaluation model to obtain a quality evaluation result of the data resource.
16. The method of claim 9, wherein, Before adding the related information of the data resource into the local data resource management, the method further comprises: According to the classification and grading rules of the industry to which the data resource belongs, the data resource is classified and graded to obtain the classification and grading result of the data resource.
17. The method of claim 15 or 16, wherein, The related information of the data resource sent to the first platform further carries the quality evaluation result of the data resource and / or the classification and grading result of the data resource.
18. The method of claim 9, wherein, After the first platform adds the related information of the data resource to the global data directory, the method further comprises: Sending a data right request to the first platform, and the first platform audits and confirms the source, generation method, use purpose, access permission, use range, and use time of the data resource, and processes the data resource by using a blockchain and / or data watermark to ensure data traceability.
19. The method of any one of claims 9-18, wherein, After the first platform adds the related information of the data resource to the global data directory, the method further comprises: Receiving a data use strategy sent by the first platform, wherein the data use strategy is generated by the first platform according to a cooperation agreement reached by the first node and the second node; Starting a computing task according to the data use strategy and completing delivery.
20. A data flow device applied to a first platform, the device comprising: a first receiving module configured to receive related information of a data resource sent by a first node; a first adding module configured to add the related information of the data resource into a global data directory; wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
21. A first platform comprising a processor and a transceiver that receives and transmits data under control of the processor, wherein, The processor is configured to perform the following operations: receive related information of a data resource sent by a first node; add the related information of the data resource into a global data directory; wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
22. A data flow device applied to a first node, the device comprising: a second adding module configured to, after accessing a data resource, perform data dictionary extraction, data cleaning, and distribution analysis on the data resource to obtain related information of the data resource, and add the related information of the data resource into a local data resource management of the first node; a first sending module configured to send the related information of the data resource to a first platform, and add the related information of the data resource into a global data directory by the first platform; wherein the first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform.
23. A first node comprising a processor and a transceiver that receives and transmits data under control of the processor, wherein, The processor is configured to perform the following operations: After accessing the data resource, data dictionary extraction, data cleaning, and distribution analysis of the data resource are performed to obtain relevant information of the data resource, and the relevant information of the data resource is added to local data resource management of the first node; The relevant information of the data resource is sent to a first platform, and the first platform adds the relevant information of the data resource to a global data directory; The first node is a trusted computing node of a data provider, and the first platform is a data scheduling center platform. 24.A communication device, comprising a memory, a processor, and a program stored in the memory and executable on the processor; the processor executes the program to implement the data flow method of any one of claims 1-8 or the data flow method of any one of claims 9-19. 25.A computer readable storage medium, having a computer program stored thereon, the program being executable by a processor to implement the steps of the data flow method of any one of claims 1-8 or the data flow method of any one of claims 9-19. 26.A computer program product, comprising computer instructions, the computer instructions being executable by a processor to implement the steps of the method of any one of claims 1-8 or the steps of the method of any one of claims 9-19.
Citation Information
Patent Citations
Electronic government affair big data processing platform
CN112765245A
Construction method of scheduling operation data base
CN115237622A
Data directory sharing method, device and system and storage medium
CN117081763A
Access method of first platform, first platform, system and storage medium
CN117081765A
Data circulation method and device, first platform and first node
CN118796798A
Cited By
A multi-constraint space-air data element circulation network construction system based on a large model
CN122366526A