Blockchain-based data analysis method and device, electronic equipment and storage medium
By dividing on-chain data into sub-data and constructing data elements for analysis, the problem of inaccurate on-chain data analysis is solved, and more efficient data mining and storage requirement analysis are achieved.
Patent Information
- Application Number
- CN202311528818.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-11-16
AI Technical Summary
In existing technologies, on-chain data analysis and mining platforms cannot effectively mine the stored information of transaction data and messages, resulting in inaccurate data analysis.
The data to be processed is divided into multiple sub-data, data information is extracted and data elements are constructed, and data analysis and processing are carried out on the blockchain based on the data information.
It improves the accuracy of data analysis, enabling more precise analysis of data storage users' storage needs.
Smart Images

Figure CN117609326B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of blockchains, and particularly relates to a data analysis method and device based on a blockchain, an electronic device, and a storage medium. BACKGROUND
[0002] Data mining is a process of processing information that is hidden, previously unknown, and has potential value from a large amount of data. With the development of blockchain technology, more and more data is stored on a blockchain.
[0003] In related technologies, after data is uploaded to a blockchain, an on-chain data analysis and mining platform performs data collection, data processing, data storage, data integration, data analysis, and other operations on the on-chain data to mine the on-chain data.
[0004] However, in related technologies, on-chain data is mostly transaction data and messages, and the structure is single and the data content is simple, which cannot effectively mine data storage information (for example, storage period, storage space, and storage demand) of a user. Moreover, in actual applications, a model obtained by mining transaction data and messages cannot accurately analyze data. SUMMARY
[0005] Embodiments of the present application provide a data analysis method and device based on a blockchain, an electronic device, and a storage medium, which can realize the diversity of on-chain data and improve the accuracy of data analysis and data mining.
[0006] In a first aspect, embodiments of the present application provide a data analysis method based on a blockchain, which includes: obtaining to-be-processed data related to data storage; dividing the to-be-processed data into multiple sub-data and extracting data information from each sub-data, wherein the data information is data required for data analysis in different data analysis scenarios; constructing a data element corresponding to each sub-data based on the data information to obtain multiple data elements corresponding to the to-be-processed data; and uploading the multiple data elements to a blockchain to perform data analysis and processing on the multiple data elements on the blockchain.
[0007] In a second aspect, embodiments of the present application provide a data analysis device based on a blockchain, which includes: a data acquisition module configured to obtain to-be-processed data related to data storage; a data extraction module configured to divide the to-be-processed data into multiple sub-data and extract data information from each sub-data, wherein the data information is data required for data analysis in different data analysis scenarios; a data element construction module configured to construct a data element corresponding to each sub-data based on the data information to obtain multiple data elements corresponding to the to-be-processed data; and a data analysis module configured to upload the multiple data elements to a blockchain to perform data analysis and processing on the multiple data elements on the blockchain.
[0008] In a third aspect, an electronic device is provided, and the electronic device includes a processor and a memory storing computer program instructions; and the processor implements the blockchain-based data analysis method according to the first aspect when executing the computer program instructions.
[0009] In a fourth aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores computer program instructions; and the computer program instructions are executed by a processor to implement the blockchain-based data analysis method according to the first aspect.
[0010] In a fifth aspect, a computer program product is provided, and instructions in the computer program product are executed by a processor of an electronic device to cause the electronic device to perform the blockchain-based data analysis method according to the first aspect.
[0011] The blockchain-based data analysis method, device, electronic device, and storage medium provided in the embodiments of the present application can realize the diversity of on-chain data and improve the accuracy of data analysis and data mining. Specifically, after obtaining the to-be-processed data related to data storage, the to-be-processed data is divided into a plurality of sub-data, and data information is extracted from each sub-data; then, a data element corresponding to each sub-data is constructed based on the data information, and a plurality of data elements corresponding to the to-be-processed data are obtained; finally, the plurality of data elements are uploaded to a blockchain, so that data analysis and processing are performed on the plurality of data elements on the blockchain. The data information is data required for data analysis in different data analysis scenarios.
[0012] As can be seen from the above, the present application divides the to-be-processed data related to data storage and extracts the related information of the divided data to construct a data element, and then performs data mining based on the data element. Since the data element is obtained by extracting the information of the divided data, and the extracted data information is data required for different data analysis scenarios, that is, in the present application, the data element contains data required for different data analysis scenarios, and the data element has diversified data information, therefore, using the data element for data analysis and mining can improve the accuracy of the data analysis model. BRIEF DESCRIPTION OF DRAWINGS
[0013] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced. Those skilled in the art can obtain other drawings according to these drawings without creating any creative labor.
[0014] Figure 1 is an architecture diagram of an on-chain data analysis and mining platform in related technologies;
[0015] Figure 2 is a flowchart of a data analysis method based on a blockchain provided by an embodiment of the present application;
[0016] Figure 3 is an architecture diagram of a data analysis system provided by an embodiment of the present application;
[0017] Figure 4 is an architecture diagram of an encrypted data processing stage provided by an embodiment of the present application;
[0018] Figure 5 is an architecture diagram of an encrypted data management stage provided by an embodiment of the present application;
[0019] Figure 6 is an architecture diagram of an encrypted data analysis stage provided by an embodiment of the present application;
[0020] Figure 7 is a data analysis flowchart based on a storage period provided by an embodiment of the present application;
[0021] Figure 8 is a schematic diagram of a coordinate system provided by an embodiment of the present application;
[0022] Figure 9 is an architecture diagram of a model verification stage provided by an embodiment of the present application;
[0023] Figure 10 is an architecture diagram of a model application stage provided by an embodiment of the present application;
[0024] Figure 11 is a structural schematic diagram of a data analysis device based on a blockchain provided by another embodiment of the present application;
[0025] Figure 12 is a hardware structural schematic diagram of an electronic device provided by yet another embodiment of the present application. DETAILED DESCRIPTION
[0026] The features and exemplary embodiments of various aspects of the present application will be described in detail below, in order to make the purposes, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in combination with the drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, but not to limit the present application. For those skilled in the art, the present application can be implemented without some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by showing examples of the present application.
[0027] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0028] Before describing the blockchain-based data analysis method provided in this application, we will first introduce the relevant knowledge in this field.
[0029] (1) Data Mining
[0030] Data mining is the process of revealing implicit, previously unknown, and potentially valuable information from massive amounts of data in databases. It is a decision support process that uses highly automated technologies such as artificial intelligence, machine learning, pattern recognition, statistics, databases, and visualization to analyze enterprise data, make inductive inferences, and uncover potential patterns to help decision-makers adjust market strategies and reduce risks. Typically, the knowledge discovery process consists of three stages: data preparation, data mining, and result presentation and interpretation. It helps enterprise users acquire, manage, process, and organize massive amounts of data within a reasonable timeframe, providing positive support for business decision-making and representing a cutting-edge technology in data storage and mining analysis.
[0031] (2) Distributed encrypted storage network
[0032] A Distributed Service Network (DSN) is a auditable, publicly verifiable, decentralized storage network with a payment system. It has evolved rapidly from advanced technologies such as distributed computing, databases, cryptography, zero-knowledge proofs, and Merkle trees. In practical applications, users typically pay encrypted storage nodes to store and retrieve data, in return for the nodes providing hard drive storage space and bandwidth. The DSN uses a blockchain-based computing scheme to periodically and randomly check the encrypted sealed sectors within the encrypted storage nodes, ensuring the integrity of the user's encrypted data.
[0033] (3) On-chain data analysis
[0034] On-chain data analytics involves indexing, processing, and storing on-chain data, then aggregating the processed on-chain data according to business problems to obtain a data model, and finally validating the data model. With the booming on-chain ecosystem, such as DeFi (Decentralized Finance) trading, lending, NFT (Non-Fungible Token) minting, and NFT trading, user behavior is directly and transparently recorded on the chain. This on-chain recorded behavioral data corresponds to the flow of on-chain value, making the analysis of this data and the insights derived from it extremely valuable.
[0035] Currently, mainstream on-chain data analytics and mining platforms primarily process the raw on-chain data retrieved through searches, then store the processed data in a data warehouse that is updated and managed by the platform. These platforms perform data mining on on-chain data according to a workflow of data collection, processing, storage, integration, and analysis. For example... Figure 1 As shown, Figure 1 This diagram illustrates the architecture of an on-chain data analysis and mining platform in a related technology, consisting of... Figure 1 It can be seen that the on-chain data analysis and mining platform mainly includes a data acquisition layer, a data processing layer, a data storage layer, a data integration layer, and a data analysis layer.
[0036] At the data acquisition layer, the on-chain data analysis and mining platform obtains raw on-chain data from blockchain nodes, or obtains on-chain data through data sources provided by third parties, or obtains on-chain data through off-chain data uploaded by users.
[0037] At the data processing layer, the on-chain data analysis and mining platform extracts, transforms, and loads raw on-chain data using either streaming or batch processing methods. In streaming processing, real-time raw blockchain data is continuously input and processed, resulting in low latency and timely analysis results. Batch processing, on the other hand, has higher latency and slightly less timely results, but is more suitable for large-volume data processing.
[0038] At the data storage layer, the on-chain data analysis and mining platform converts the processed on-chain data into a predefined format and stores it in various data tables of the dataset for later use.
[0039] At the data integration layer, the on-chain data analysis and mining platform performs aggregation calculations on the data in various tables stored in the dataset. Specifically, the on-chain data analysis and mining platform can perform aggregation calculations on the stored data based on pre-defined indicators, or it can trigger aggregation calculations in stages or according to pre-defined conditions.
[0040] At the data analysis layer, the on-chain data analysis and mining platform outputs the calculation results of the aggregated calculation in real time. Users can interact with the on-chain data analysis and mining platform at the data analysis layer. For example, the on-chain data analysis and mining platform can provide a Business Intelligence report interface on which visual charts can be displayed, and users can operate the visual charts through the API interface provided by the on-chain data analysis and mining platform.
[0041] It should be noted that in actual applications, different on-chain data analysis and mining platforms can use different schemes to build and maintain data warehouses.
[0042] However, the on-chain data used by the current on-chain data analysis and mining platform in the process of data analysis and mining is mostly transaction data and messages, and the structure of the transaction data is single, the data content is simple, and the data mining of the data storage information (such as storage period, storage space, and storage demand) of the user cannot be effectively performed. Moreover, the model obtained by performing data mining on the transaction data and messages cannot accurately analyze the data.
[0043] To solve the above problems, the embodiments of the present application provide a data analysis method based on a block chain, a device, an electronic device, and a storage medium. First, the data analysis method based on a block chain provided by the embodiments of the present application is introduced.
[0044] Figure 2 A flowchart of a data analysis method based on a block chain provided by an embodiment of the present application is shown. As shown in Figure 2 The method includes the following steps:
[0045] Step S201, obtaining data to be processed related to data storage.
[0046] In step S201, the data to be processed can be data to be uploaded by a user to a storage node in a block chain. For example, when a user needs to store data in a local device to a storage node in a block chain, a data analysis system can obtain the data uploaded by the user, and perform related processing on the data in steps S202 to S204, and then store the processed data to the storage node in the block chain.
[0047] In addition, the to-be-processed data can also be data uploaded by a data analysis user, wherein the data analysis user can obtain relevant data in a process in which a plurality of data storage users use a storage product (for example, a hard disk, a network disk, or the like), and obtain an analysis result of data analysis of the data analysis system on the to-be-processed data, so as to determine a storage requirement of the data storage user, thereby better recommending the storage product to the data storage user, or updating the storage product to meet the storage requirement of the data storage user.
[0048] In step S202, the to-be-processed data is divided into a plurality of sub-data, and data information is extracted from each sub-data.
[0049] In step S202, the data analysis system can divide the to-be-processed data into a plurality of sub-data with the same byte length. In order to ensure the integrity of the data, the data analysis system can set the byte length of each sub-data to 2K during data division of the to-be-processed data. If the byte length of the sub-data is less than 2K, a padding operation can be performed.
[0050] In addition, in step S202, the data information is data required for data analysis in different data analysis scenarios. The data analysis scenario is a scenario of analyzing different aspects of data, for example, a scenario of analyzing storage information and / or storage requirements of a data storage user based on a keyword input by the data storage user during a data storage process, a scenario of analyzing based on a storage period of data stored by the data storage user, and a scenario of analyzing based on a storage space of data stored by the data storage user. Correspondingly, the data information at least includes one of the following: a keyword related to data storage, a storage period, and a storage space.
[0051] It should be noted that, compared with the related art which only analyzes transaction information on the chain, the present application extracts a plurality of data information from the sub-data, enriches the content of the on-chain data, thereby providing diversified data information for modeling of the data analysis model, and further providing a data basis for improving the accuracy of the data analysis model.
[0052] In step S203, a data element corresponding to each sub-data is constructed based on the data information, and a plurality of data elements corresponding to the to-be-processed data are obtained.
[0053] In step S203, the data element is composed of data information, and therefore, the data element also has diversified data information, which can provide a data basis for improving the accuracy of the data analysis model, so that the data analysis model can more accurately analyze the storage requirement of the data storage user.
[0054] In step S204, the plurality of data elements are uploaded to the block chain, so as to perform data analysis and processing on the plurality of data elements on the block chain.
[0055] In step S204, after obtaining the data element, the data analysis system writes the data element into a block of the blockchain to realize storage of the data element. Meanwhile, the data element is also stored in the data pool of the blockchain, and the data analysis system can perform data analysis and mining on the data elements in the data pool to obtain a data analysis model, analyze the storage requirements of the data storage user, and better serve the customer.
[0056] It should be noted that, since the data pool of the blockchain stores data elements containing various data information, the data analysis system can analyze diversified data on the blockchain, rather than only transaction information, and thus, using the data analysis model obtained by analyzing diversified data, the storage requirements of the data storage user can be more accurately analyzed.
[0057] Based on the scheme defined in steps S201 to S204, it can be known that, after obtaining the data to be processed related to data storage, the data to be processed is divided into multiple sub-data, and data information is extracted from each sub-data; then, based on the data information, a data element corresponding to each sub-data is constructed, and multiple data elements corresponding to the data to be processed are obtained; finally, the multiple data elements are uploaded to the blockchain to perform data analysis and processing on the multiple data elements on the blockchain. The data information is data required for data analysis in different data analysis scenarios.
[0058] It is easy to note that, in the present application, the data to be processed related to data storage is split, and the relevant information of the split data is extracted to construct a data element, and then data mining is performed based on the data element. Since the data element is obtained by extracting information from the split data, and the extracted data information is data required for different data analysis scenarios, that is, in the present application, the data element contains data required for different data analysis scenarios, and the data element has diversified data information, so using the data element for data analysis and mining can improve the accuracy of the data analysis model.
[0059] In one example, Figure 3 The architecture of the data analysis system performing the method provided by the present application is shown, which includes Figure 3 It can be known that, the data analysis system realizes the scheme provided by the present application through five stages of encrypted data processing stage, encrypted data management stage, encrypted data analysis stage, model verification stage and model application stage.
[0060] In the encrypted data processing stage, the data analysis system can formulate or select a data analysis scheme according to the business requirements of the encrypted data, wherein the data analysis scheme corresponds to a data analysis scene; in the encrypted data management stage, the data analysis system constructs the data structure of the data element and encrypts the data element, and stores the encrypted data element in the block of the blockchain; in the encrypted data analysis stage, the data analysis system analyzes the data element stored in the blockchain, and analyzes and calculates the massive data element after analysis, and deduces the data analysis model to mine the model representing the storage rule of the data storage user; in the model verification stage, the data analysis system randomly selects a certain number of on-chain data (i.e. data element stored in the blockchain), and verifies the data analysis model based on the selected on-chain data; in the model application stage, the data analysis system uses the data analysis model to analyze the data of the data storage user, outputs the on-chain business analysis result, and scores the data analysis model to determine whether to update the data analysis model.
[0061] The following explains the above five stages.
[0062] In an example, Figure 4 The architecture diagram of the encrypted data processing stage is shown. In Figure 4 After the data analysis user logs in to the data analysis system, the data analysis user enters the main interface of the encrypted data processing scheme. The data analysis user can select an existing data analysis scheme through the main interface, or customize a new data analysis scheme.
[0063] It should be noted that according to different storage business scenes, the data analysis scheme can be divided into three kinds: keyword data analysis scheme, storage period analysis scheme, and storage space analysis scheme. In the keyword data analysis scheme, the data analysis system can mine a model representing the user's storage rule with the keyword according to the keyword input by the data storage user. In the storage period analysis scheme, the data analysis system can mine a model representing the user's storage rule related to the storage period according to the storage period input by the user. In the storage space analysis scheme, the data analysis system can mine a model representing the user's storage rule related to the storage space according to the storage space input by the user.
[0064] In an example, in the scene of the data analysis user creating a new data analysis scheme, the data analysis user can click the "create" button in the main interface, and the main interface pops up a creation dialog box of the data analysis scheme. The data analysis user can input the scheme name, scheme description, etc. of the data analysis scheme to be created into the dialog box, select the application category, and click the "OK" button to save the created data analysis scheme.
[0065] It should be noted that the application category represents an application scenario of a data analysis scheme created by a data analysis user, for example, a scenario of analyzing storage requirements of a data storage user.
[0066] In another example, in a scenario where a data analysis user selects a data analysis scheme from existing data analysis schemes, the data analysis user can select a target data analysis scheme from a "recent schemes" list displayed on the main interface, and then the data analysis user can edit the selected target data analysis scheme, for example, modify the scheme name, scheme description, application category, and the like of the target data analysis scheme.
[0067] In another scenario, the data analysis user can also click the "more" button of the main interface to select a data analysis scheme to be viewed or edited from a plurality of data analysis schemes popped up.
[0068] After determining the data analysis scheme, the data analysis system generates a public-private key pair, wherein the private key is the user password of the data analysis user logging into the data analysis system, and the public key is a preset point on an encryption hyperbola corresponding to the private key. As an example, the above-mentioned preset encryption hyperbola can be a BLS12-381 hyperbola, and the public key is a point on the G1 curve in the BLS12-381 hyperbola corresponding to the private key.
[0069] After generating the public-private key pair, the data analysis system constructs a data meta signature system based on the above-mentioned public-private key pair, which consists of a signature function and a verification function, wherein the signature function can be an 80K signature function, and the verification function can be an 80K verification function. The 80K signature function performs signature calculation on the split data (i.e. sub-data) to be processed by the private key, thereby obtaining 80-bit signature data and serialized data meta; the 80K verification data calculates the public key based on the signature data and the serialized data meta calculated by the 80K signature function, and verifies the data meta based on the calculated public key and the public key determined based on the encryption hyperbola.
[0070] So far, the explanation and description of the encrypted data processing phase are completed.
[0071] The encrypted data management phase is explained and described as follows.
[0072] In one example, Figure 5 The architecture diagram of the encrypted data management phase is shown. As Figure 5 shown, in the encrypted data management phase, the data analysis system first generates data meta based on the data to be processed.
[0073] Before constructing the data element, the data analysis system acquires a private key for encrypting each sub-data, determines a public key corresponding to the private key on a preset encryption hyperbolic curve, then encrypts each sub-data based on the private key to obtain encrypted sub-data, and verifies the encrypted sub-data based on the public key to obtain a verification result. Finally, in a case where the verification result represents that the encrypted sub-data is verified successfully, the data analysis system generates a random number of a preset length, and determines the random number as a data element identifier corresponding to the encrypted sub-data.
[0074] In one example, as shown in FIG. 1, before a data storage user uploads to-be-processed data to a blockchain, the data analysis system divides the to-be-processed data into sub-data of size 2k. Then, the data analysis system calls a signature function to encrypt the sub-data using a private key of a data analysis user to generate signature data (i.e., encrypted sub-data). Then, the data analysis system stores the signature data (i.e., encrypted sub-data) in a data element message pool. Figure 5
[0075] Since the data element message pool stores a large amount of data transmitted by various data storage users, in order to ensure the security and integrity of the data, the data in the data element message pool needs to be verified. The data analysis system can call a verification function to verify the signature data (i.e., encrypted sub-data) in the data element message pool using a public key. If the verification is successful, the data analysis system randomly generates a random number Nance value of a preset length (for example, 256 bits) as a data element identifier of the data element. Then, the data analysis system packs the signature data (i.e., encrypted sub-data), the data element identifier, and data information corresponding to the signature data, and writes them into a block of the data chain. Finally, the data analysis system stores the packed data in a data pool of the blockchain, so that the data analysis system analyzes the data in the data pool in the encrypted data analysis stage.
[0076] In the above process, the data analysis system uses an encryption function to encrypt each sub-data. Specifically, the data analysis system performs data conversion on target sub-data to obtain a byte array corresponding to the target sub-data, then performs data compression on the byte array to obtain compressed target sub-data, and finally encrypts the compressed target sub-data based on a private key to obtain encrypted sub-data.
[0077] It should be noted that the target sub-data is any one of the plurality of sub-data corresponding to the to-be-processed data.
[0078] In one example, the data analysis system converts the target sub-data into a byte array, and then data compresses the byte array with the compression function Blake2B.sum256 to generate 256-bit compressed data (i.e., compressed target sub-data). Then, the data analysis system encrypts the 256-bit compressed data using a private key, thereby obtaining an 80-bit data signature (i.e., encrypted sub-data).
[0079] It should be noted that the above-mentioned 80-bit data signature is a corresponding point on the G2 curve in the BLS12-381 hyperbolic curve, and the public key is a point corresponding to the private key on the G1 curve in the BLS12-381 hyperbolic curve.
[0080] Further, after completing data encryption, the data analysis system verifies the encrypted sub-data based on the public key. Specifically, the data analysis system determines the target public key based on the encrypted sub-data and the byte array; wherein, in the case that the public key and the target public key are the same, it is determined that the verification result of the encrypted sub-data is verification success; in the case that the public key and the target public key are different, it is determined that the verification result of the encrypted sub-data is verification failure.
[0081] In one example, since the 80-bit data signature is a point on the G2 curve in the BLS12-381 hyperbolic curve, the data analysis system can calculate the point on the G1 curve in the BLS12-381 hyperbolic curve through the serialized byte array of the data element, thereby obtaining the target public key. If the target public key calculated is consistent with the public key determined in the encrypted data processing stage, it indicates that the verification is successful; otherwise, it indicates that the verification fails.
[0082] Further, after completing the encryption and verification of the sub-data, the data analysis system constructs the data element corresponding to each sub-data based on the data element identifier, the encrypted sub-data, and the data information.
[0083] In one example, the data content of the data element can be represented as {key, 80-bit data signature, encryption index, data ownership, off-chain data encryption storage address, data storage time, data storage order number, Nonce value, encrypted data}. Among them, the encryption index is used for positioning the data element; the data ownership is used to determine the identifier of the user who uploads the sub-data corresponding to the data element; the off-chain data encryption storage address represents the data storage space; the data storage time represents the data storage period; the data storage order number represents the transaction information (such as payment time and payment amount) corresponding to the data storage.
[0084] At this point, the explanation and description of the encrypted data management stage are completed.
[0085] The following explains and describes the encrypted data analysis stage.
[0086] In one example, Figure 6 The architecture diagram of the encrypted data analysis stage is shown. As Figure 6 shown, in the encrypted data analysis stage, the data analysis system mines the data regularity of the data elements on the chain according to the data analysis scheme. Among them, Figure 6 The data mining process of four analysis schemes is shown, that is, the keyword data analysis scheme, the storage period analysis scheme, the storage space analysis scheme, and the collaborative scheme.
[0087] In the storage period analysis scheme, the data analysis system mines the model representing the user's storage regularity related to the storage period according to the storage period input by the data storage user.
[0088] Specifically, on the blockchain, the data analysis system compares the storage periods of multiple data elements with multiple preset storage periods, divides the multiple data elements into at least one data set; then, determines the target data element whose storage space satisfies the preset condition from the at least one data set according to the storage space corresponding to each data element, and determines the user who stores the target data element as the target user.
[0089] In one example, as Figure 6 shown, the data analysis system uses a classification prediction method to analyze the data elements based on keywords to obtain a model representing the user's storage regularity. Among them, the data analysis system can use Bayesian network, random forest, decision tree, support vector machine and other algorithm methods to analyze the data elements based on keywords.
[0090] It should be noted that the classification prediction algorithm calculates the occurrence probability of different data element attributes (for example, storage period) in the known data element set, constructs a classification tree based on judgment logic, for example, uses the logic form of IF··· THEN···, compares the input variable value to predict the output variable value, to realize the classification analysis of the data elements.
[0091] In one example, Figure 7 The data analysis flowchart based on the storage period is shown. In Figure 7The system sets five preset storage periods: storage period 1, storage period 2, storage period 3, storage period 4, and storage period 5. The data analysis system compares the storage period corresponding to the data input by the data storage user with each preset storage period to determine the data set containing the data element for that user: data set A, data set B, data set C, data set D, data set E, and data set F. After classifying data elements based on storage periods, the system can further consider the storage space corresponding to the data storage user to determine if the user corresponding to a data element in a particular data set is a large storage customer. For example, if a data element is located in data set F and its corresponding storage space is greater than 10TB, then the data storage customer corresponding to that data element is identified as a large storage customer. Additionally, the system considers the number of times a data element is stored to determine if the user corresponding to a data element is a high-frequency customer. For example, if a data element is located in data set F and its storage frequency is greater than 5 times, then the data storage customer corresponding to that data element is identified as a high-frequency customer.
[0092] In the storage space analysis scheme, the data analysis system extracts a model that represents the user's storage patterns based on the storage space input by the data storage user.
[0093] Specifically, on the blockchain, the data analysis system clusters multiple data elements based on their storage space and storage period to obtain multiple clusters; then, based on the multiple clusters, it classifies the users who store multiple data elements to obtain at least one user set, where each user set contains users with the same data storage characteristics.
[0094] In one example, such as Figure 6 As shown, the data analysis system uses clustering analysis to analyze data elements based on storage space to obtain a model representing user storage patterns. Specifically, the data analysis system can employ algorithms such as K-means, K-centroids, and hierarchical clustering to analyze data elements based on storage space. The following explanation uses the K-centroids algorithm as an example.
[0095] The data analysis system, based on the storage space input by the user, establishes a coordinate system for the data elements in the data pool, with the horizontal axis representing storage space and the vertical axis representing storage period, such as... Figure 8 As shown. The data analysis system uses this coordinate system to plot each data element as a coordinate point, and sets the location of the coordinate point set as the center point. Each center point corresponds to a classification criterion. Based on this classification criterion, the data elements are classified to obtain multiple clusters. For example, in... Figure 8In the embodiment, the classification criteria corresponding to the center points include 10M, 25M, 110M and 10G, and the data elements are divided into four clustering clusters according to the classification criteria, and the data elements included in each clustering cluster correspond to the same storage space of the storage user.
[0096] In the keyword data analysis scheme, the data analysis system mines a model representing the storage rule of the user according to the keyword input by the data storage user.
[0097] Specifically, on the blockchain, the data analysis system performs semantic recognition on the keywords of the plurality of data elements to obtain a recognition result, then classifies the plurality of data elements according to the recognition result to obtain a plurality of data element categories, and finally determines the storage requirement of the user storing the plurality of data elements based on the plurality of data element categories.
[0098] In one example, as shown in Figure 6 The data analysis system can analyze the data elements based on the keywords in the manner of association rules, wherein the data analysis system can use Apirori algorithm, Eclat algorithm (frequent item set mining algorithm), grey correlation method, FP-Growth algorithm and the like to realize the association between the keywords. Specifically, the data analysis system reads the keywords in the data elements, extracts the data elements with related keywords, and forms a data element set.
[0099] It should be noted that generally the same keyword or the keyword with the same semantic can express the same meaning, and therefore, in order to improve the accuracy of data element classification, in the process of classifying the data elements based on the keywords, the data analysis system not only classifies the data elements with the same keywords into one data element set, but also classifies the data elements with the same semantic into one data element set, so as to better analyze the storage requirement of the data storage user.
[0100] In the collaborative scheme, the data analysis system can analyze the data elements by selecting any two or three of the above-mentioned keyword data analysis scheme, storage period analysis scheme and storage space analysis scheme. In this scheme, the data analysis system can analyze the data elements in the manner of item-based collaborative filtering, can analyze the data elements in the manner of user-based collaborative filtering, or can analyze the data elements in other manners, which will not be illustrated herein.
[0101] According to the keyword input by the data storage user, a model representing the storage rule of the user is mined.
[0102] So far, the explanation and description of the encrypted data analysis stage are completed.
[0103] The model verification stage is explained and described below.
[0104] In one example, Figure 9 An architecture diagram of the model verification stage is shown. In the model verification stage, the data analysis system randomly selects a certain amount of on-chain data to form test data, and uses the test data to verify the analysis accuracy of the data model (for example, the data model corresponding to the keyword data analysis scheme) output by the encrypted data analysis stage, and optimizes the data model according to the verification result.
[0105] Specifically, the data analysis system can capture all or part of the on-chain data as test data. Then, the data analysis system sets the verification parameters of the model, for example, sets the input data, target output data, intermediate variables, etc. After setting the verification parameters, the data analysis system can input the related parameters to the model to be verified, for example, input the storage time, storage size, monthly storage frequency, etc., and obtain the output data of the model (for example, which industry and which type of enterprise have similar storage needs), and determine whether the output data of the model matches the target output data. If it matches, it is determined that the model to be verified is verified successfully; otherwise, it is determined that the model to be verified fails.
[0106] At this point, the explanation and description of the model verification stage are completed.
[0107] The following explains and describes the model application stage.
[0108] In one example, Figure 10 An architecture diagram of the model application stage is shown. In the model application stage, the data analysis system inputs test data to the data model, so that the data model outputs the on-chain business analysis result, and scores the data model to determine whether to optimize the data model. For example, the data analysis user inputs the storage time, storage size, monthly storage frequency, etc. into the data model, and the data model can infer which industry and which type of enterprise have similar storage needs, and then verify whether the storage product of this type is suitable within the next month.
[0109] At this point, the explanation and description of the entire data analysis system architecture are completed.
[0110] From the above, the application constructs a distributed encryption network based on a blockchain that can support various types of data analysis and data mining. The data on the data block contains data elements obtained by processing the data to be processed through encryption indexing, data ownership, and encrypted storage address, supporting data analysis and data mining. In addition, the blockchain and data analysis system provided by the application are designed simultaneously and complement each other. The data analysis system is an encrypted storage system based on a blockchain established for data mining. It is suitable for data mining for specific business problems to achieve data storage product analysis and recommendation, and can be extended to encrypted data storage analysis and recommendation in different industries. For example, by classifying data elements, the encrypted storage product is mined to obtain more storage users of 20M, 100M, 15G, 20G, and 1TB. The data analysis system sets storage products of 20M, 100M, 15G, 20G, and 1TB.
[0111] In addition, the data storage capacity in the data analysis platform in the related art is small, and the on-chain data is usually a string, which does not support data mining. The scheme provided by the application increases the capacity of on-chain data by designing a data element structure, which is more suitable for mining the rules of on-chain data such as keywords, storage period, and storage space.
[0112] The application also provides a data analysis device based on a blockchain, as shown in Figure 11 The device includes a data acquisition module 1101, a data extraction module 1102, a data element construction module 1103, and a data analysis module 1104.
[0113] The data acquisition module 1101 is configured to acquire data to be processed related to data storage.
[0114] The data extraction module 1102 is configured to divide the data to be processed into a plurality of sub-data, and extract data information from each sub-data, wherein the data information is data required for data analysis in different data analysis scenarios.
[0115] The data element construction module 1103 is configured to construct a data element corresponding to each sub-data based on the data information, to obtain a plurality of data elements corresponding to the data to be processed.
[0116] The data analysis module 1104 is configured to upload the plurality of data elements to a blockchain, to perform data analysis processing on the plurality of data elements on the blockchain.
[0117] In one example, the blockchain-based data analysis apparatus comprises a key acquisition module, a data encryption module, a data verification module, and an identification generation module. The key acquisition module is configured to acquire a private key for encrypting each sub-data before constructing a data element corresponding to each sub-data based on data information to obtain a plurality of data elements corresponding to the to-be-processed data, and determine a public key corresponding to the private key on a preset encryption hyperbolic curve. The data encryption module is configured to encrypt each sub-data based on the private key to obtain encrypted sub-data. The data verification module is configured to verify the encrypted sub-data based on the public key to obtain a verification result. The identification generation module is configured to generate a random number of a preset length in a case where the verification result indicates that the encrypted sub-data is verified successfully, and determine the random number as a data element identification corresponding to the encrypted sub-data.
[0118] In one example, the data element construction module comprises a construction submodule configured to construct a data element corresponding to each sub-data based on the data element identification, the encrypted sub-data, and the data information. The data information at least includes one of the following: a keyword related to data storage, a storage period, and a storage space.
[0119] In one example, the data encryption module comprises a data conversion module, a data compression module, and an encryption submodule. The data conversion module is configured to perform data conversion on target sub-data to obtain a byte array corresponding to the target sub-data, wherein the target sub-data is any one of the plurality of sub-data corresponding to the to-be-processed data. The data compression module is configured to perform data compression processing on the byte array to obtain compressed target sub-data. The encryption submodule is configured to encrypt the compressed target sub-data based on the private key to obtain the encrypted sub-data.
[0120] In one example, the data verification module comprises a public key determination module, a first verification module, and a second verification module. The public key determination module is configured to determine a target public key based on the encrypted sub-data and the byte array. The first verification module is configured to determine that the verification result of the encrypted sub-data is verification success in a case where the public key is the same as the target public key. The second verification module is configured to determine that the verification result of the encrypted sub-data is verification failure in a case where the public key is different from the target public key.
[0121] In one example, the data analysis module comprises a first analysis module configured to divide a plurality of data elements into at least one data set by comparing the storage periods of the plurality of data elements with a plurality of preset storage periods on a blockchain, determine a target data element whose storage space satisfies a preset condition from the at least one data set according to the storage space corresponding to each data element, and determine a target user who stores the target data element.
[0122] In one example, the data analysis module comprises: a second analysis module configured to cluster the plurality of data elements on the blockchain according to storage spaces of the plurality of data elements and storage periods of the plurality of data elements to obtain a plurality of clustering clusters; and classify users storing the plurality of data elements based on the plurality of clustering clusters to obtain at least one user set, wherein each user set comprises users having the same data storage characteristics.
[0123] In one example, the data analysis module comprises: a third analysis module configured to perform semantic recognition on keywords of the plurality of data elements on the blockchain to obtain a recognition result; classify the plurality of data elements according to the recognition result to obtain a plurality of data element categories; and determine storage requirements of users storing the plurality of data elements based on the plurality of data element categories.
[0124] The data analysis apparatus based on the blockchain provided by the embodiments of the present application can implement the processes implemented by the foregoing method embodiments, and thus will not be described herein again for the sake of brevity and conciseness.
[0125] Those skilled in the art can clearly understand that, for the sake of brevity and conciseness, only the division of the above functional units and modules is exemplified, and in actual applications, the above functions can be completed by different functional units and modules according to needs, that is, the internal structure of the apparatus is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiments can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the purpose of mutual distinction, and do not limit the protection scope of the present application. The specific working processes of the units and modules in the system can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.
[0126] Figure 12 A hardware structure schematic diagram of an electronic device provided by the embodiments of the present application is shown.
[0127] The electronic device can include a processor 1201 and a memory 1202 storing computer program instructions.
[0128] Specifically, the processor 1201 can include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or can be configured as one or more integrated circuits implementing the embodiments of the present application.
[0129] The memory 1202 can include mass storage for data or instructions. As an example and not by way of limitation, the memory 1202 can include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc (e.g., a compact disc (CD) or a digital versatile disc (DVD)), a solid-state drive (SSD), a USB drive, or a combination of two or more of these. Where appropriate, the memory 1202 can include removable or non-removable (or fixed) media. Where appropriate, the memory 1202 can be internal or external to the integrated gateway disaster recovery device. In particular embodiments, the memory 1202 is non-volatile, solid-state memory.
[0130] The memory can include read-only memory (ROM), random-access memory (RAM), magnetic disk storage mediums, optical storage mediums, flash memory devices, electrical, optical, or other physical / tangible memory storage devices. Thus, in general, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software that, when executed (by one or more processors), is operable to
[0131] The processor 1201 implements any one of the above-described embodiments of the blockchain-based data analysis method by reading and executing computer program instructions stored in the memory 1202.
[0132] In one example, the electronic device can further include a communication interface 1203 and a bus 1210. As shown, the processor 1201, the memory 1202, and the communication interface 1203 are connected by the bus 1210 and complete communication therebetween. Figure 12
[0133] The communication interface 1203 is mainly used to realize the communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0134] Bus 1210 includes hardware, software, or both, to couple electronic devices to each other in a manner that allows information to be passed between or among the coupled devices. Although embodiments will be described with reference to a particular bus, embodiments contemplate any suitable bus or interconnect, such as a Accelerated Graphics Port (AGP) or other graphics bus, Enhanced Industry Standard Architecture (EISA) bus, Front Side Bus (FSB), HyperTransport, Industry Standard Architecture (ISA) bus, InfiniBand, Low Pin Count (LPC) bus, memory bus, Micro Channel Architecture (MCA) bus, Peripheral Component Interconnect (PCI) bus, PCI-Express (PCI-X) bus, Serial Advanced Technology Attachment (SATA) bus, Video Electronics Standards Association local (VLB) bus, and the like, or a combination of two or more of these. Where suitable, bus 1210 can include one or more buses. Although a particular bus has been described and illustrated here, the application contemplates any suitable bus or interconnect.
[0135] The electronic device can perform the blockchain-based data analysis method in the embodiments of the present application based on the currently intercepted spam messages and the messages reported by the user, thereby realizing the blockchain-based data analysis method described above in combination with Figure 1 and Figure 2 the blockchain-based data analysis method.
[0136] In addition, in combination with the blockchain-based data analysis method in the above embodiments, the embodiments of the present application can provide a computer storage medium to realize. The computer storage medium has computer program instructions stored thereon; the computer program instructions are executed by the processor to realize any one of the blockchain-based data analysis methods in the above embodiments.
[0137] In addition, in combination with the blockchain-based data analysis method in the above embodiments, the embodiments of the present application can provide a computer program product to realize. The instructions in the computer program product are executed by the processor of the electronic device, so that the electronic device executes the blockchain-based data analysis method as any one of the above embodiments.
[0138] It needs to be clear that the present application is not limited to the specific configurations and processes described above and shown in the drawings. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between steps, after understanding the spirit of the present application.
[0139] The functions indicated in the structural block diagrams described above can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, and the like. When implemented in software, the elements of the present application are program or code segments that are used to perform the required tasks. The program or code segments can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave over a transmission medium or communication link. The "machine-readable medium" can include any medium that can store or transfer information. Examples of the machine-readable medium include an electronic circuit, a semiconductor memory device, a ROM, a flash memory, an erasable ROM (EROM), a floppy diskette, a CD-ROM, an optical disk, a hard disk, a fiber optic medium, a radio frequency (RF) link, and the like. The code segments can be downloaded via a computer network, such as the Internet, an intranet, and the like.
[0140] It is also noted that the exemplary embodiments mentioned in the present application describe some methods or systems based on a series of steps or devices. However, the present application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiments, or in an order different from the embodiments, or several steps can be performed simultaneously.
[0141] The above described aspects and / or embodiments of the present disclosure can be implemented in hardware, software, firmware or any combination thereof. Such computer program instructions are stored in a non-transitory computer readable medium or media when executed by a computer or processor cause the computer or processor to carry out a certain functionality. The non-transitory computer readable medium or media include a floppy disk, a ZIP® disk, a hard disk, a CD-ROM, a DVD-ROM, a flash memory, a computer memory, a register, a transmitter, a receiver, a semiconductor memory device, a floppy disk drive, a CD-ROM drive, a DVD- ROM drive, a flash memory drive, an RAM, a ROM, and the like. The computer program instructions can be executed by one or more processors, such as one or more digital signal processors (DSPs), general purpose microprocessors, ASICs, field programmable gate arrays (FPGAs), or other equivalent integrated or discrete logic circuitry. The processor, or processors, can be configured to perform particular functions by the implementation of one or more computer programs or code segments that are stored in a computer readable medium of the non-transitory computer readable medium or media.
[0142] The above merely describes a specific implementation of the present application. Those skilled in the art can clearly understand the specific working processes of the system, modules and units described above for the convenience and brevity of description, and can refer to the corresponding processes in the foregoing method embodiments, which will not be described herein again. It should be understood that the protection scope of the present application is not limited to this, and any person skilled in the art can easily think of various equivalent modifications or replacements within the technical range disclosed by the present application, and these modifications or replacements should be covered within the protection scope of the present application.
Claims
1. A blockchain-based data analysis method, characterized in that, The method comprises: obtaining data to be processed related to data storage; dividing the data to be processed into a plurality of sub-data, and extracting data information from each sub-data, wherein the data information is data required for data analysis in different data analysis scenarios; constructing a data element corresponding to each sub-data based on the data information, to obtain a plurality of data elements corresponding to the data to be processed; wherein the data element comprises: a keyword, data ownership, an off-chain data encrypted storage address, and a data storage time; wherein the data ownership is used to determine the identity of a user uploading the sub-data corresponding to the data element; the off-chain data encrypted storage address represents a data storage space; and the data storage time represents a data storage period; uploading the plurality of data elements to a blockchain to perform data analysis and processing on the plurality of data elements on the blockchain; performing data analysis and processing on the plurality of data elements on the blockchain, comprising: clustering the plurality of data elements on the blockchain according to the storage space of the plurality of data elements and the storage period of the plurality of data elements, to obtain a plurality of clustering clusters; classifying users storing the plurality of data elements based on the plurality of clustering clusters, to obtain at least one user set, wherein each user set contains users with the same data storage characteristics; performing data analysis and processing on the plurality of data elements on the blockchain, comprising: performing semantic recognition on the keyword of the plurality of data elements on the blockchain, to obtain a recognition result; classifying the plurality of data elements according to the recognition result, to obtain a plurality of data element categories; determining the storage requirements of users storing the plurality of data elements based on the plurality of data element categories.
2. The data analysis method of claim 1, wherein, Before constructing the data element corresponding to each sub-data based on the data information, and obtaining the plurality of data elements corresponding to the data to be processed, the method further comprises: obtaining a private key for encrypting each sub-data, and determining a public key corresponding to the private key on a preset encryption hyperbolic curve; encrypting each sub-data based on the private key to obtain encrypted sub-data; verifying the encrypted sub-data based on the public key to obtain a verification result; in a case where the verification result indicates that the encrypted sub-data is verified successfully, generating a random number of a preset length, and determining the random number as a data element identifier corresponding to the encrypted sub-data.
3. The data analysis method of claim 2, wherein, constructing the data element corresponding to each sub-data based on the data information, comprising: constructing the data element corresponding to each sub-data based on the data element identifier, the encrypted sub-data, and the data information, wherein the data information at least includes one of the following: a keyword related to the data storage, a storage period, and a storage space.
4. The data analysis method of claim 2, wherein, encrypting each sub-data based on the private key to obtain encrypted sub-data, comprising: performing data conversion on a target sub-data to obtain a byte array corresponding to the target sub-data, wherein the target sub-data is any one of the plurality of sub-data corresponding to the data to be processed; Data compression processing is performed on the byte array to obtain compressed target sub-data; The compressed target sub-data is encrypted based on the private key to obtain the encrypted sub-data.
5. The data analysis method of claim 4, wherein, The encrypted sub-data is verified based on the public key to obtain a verification result, including: A target public key is determined based on the encrypted sub-data and the byte array; In the case that the public key is the same as the target public key, it is determined that the verification result of the encrypted sub-data is verification success; In the case that the public key is different from the target public key, it is determined that the verification result of the encrypted sub-data is verification failure.
6. The data analysis method of any one of claims 1 to 3, wherein, Data analysis processing is performed on the plurality of data elements on the blockchain, including: On the blockchain, the plurality of data elements are divided into at least one data set by comparing the storage period of the plurality of data elements with a plurality of preset storage periods; A target data element is determined from the at least one data set based on the storage space corresponding to each data element, which satisfies a preset condition; A target user is determined as a user who stores the target data element. 7.A blockchain-based data analysis apparatus, characterized by comprising: Including: A data acquisition module for acquiring data to be processed related to data storage; A data extraction module for dividing the data to be processed into a plurality of sub-data and extracting data information from each sub-data, wherein the data information is required for data analysis in different data analysis scenarios; A data element construction module for constructing a data element corresponding to each sub-data based on the data information to obtain a plurality of data elements corresponding to the data to be processed; wherein the data element includes: a keyword, data ownership, an off-chain data encryption storage address, and a data storage time; wherein the data ownership is used to determine the identity of the user who uploads the sub-data corresponding to the data element; the off-chain data encryption storage address represents a data storage space; and the data storage time represents a data storage period; A data analysis module for uploading the plurality of data elements to a blockchain to perform data analysis processing on the plurality of data elements on the blockchain; The data analysis module includes: A clustering unit for clustering the plurality of data elements based on the storage space of the plurality of data elements and the storage period of the plurality of data elements on the blockchain to obtain a plurality of clustering clusters; A clustering classification unit for classifying users who store the plurality of data elements based on the plurality of clustering clusters to obtain at least one user set, wherein each user set contains users with the same data storage characteristics; The data analysis module includes: A semantic unit for performing semantic recognition on the keyword of the plurality of data elements on the blockchain to obtain a recognition result; A semantic classification unit for classifying the plurality of data elements according to the recognition result to obtain a plurality of data element categories; A storage unit for determining the storage requirements of users who store the plurality of data elements based on the plurality of data element categories.
8. An electronic device, comprising: The electronic device includes a processor and a memory storing computer program instructions; The processor implements the blockchain-based data analysis method according to any one of claims 1-6 when executing the computer program instructions.
9. A computer-readable storage medium, characterized in that, The computer program instructions are stored on the computer readable storage medium, and the computer program instructions are executed by the processor to implement the blockchain-based data analysis method according to any one of claims 1-6.
Citation Information
Patent Citations
Data processing method and device
CN104298739A
Decentralized data verification processing method, device and system and medium
CN110096542A