A multi-dimensional big data sharing method and system
By type division and distributed sharing of multivariate big data, the problems of system paralysis and inefficiency in sharing caused by centralized management are solved, and efficient data sharing and resource utilization are achieved.
Patent Information
- Application Number
- CN202411108538.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-13
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2044-08-13
AI Technical Summary
The existing multi-dimensional big data sharing methods have problems such as centralized management that can easily lead to system paralysis, inefficient sharing and waste of resources.
The data generated by the target data platform is divided into multiple types. By obtaining monitoring indicators such as historical access frequency, generation of timestamps, real-time data volume and data update frequency, the data type division coefficient is calculated, and data sharing is adopted using a distributed sharing framework and star connection method.
It improves data sharing efficiency, reduces the risk of data loss, and improves the efficiency of using shared resources.
Smart Images

Figure CN119201870B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data sharing and relates to big data technology, specifically to a multivariate big data sharing method and system. Background Art
[0002] Existing multivariate big data sharing methods have the following defects:
[0003] 1. Existing multi-dimensional big data sharing methods usually use a centralized data management system. All data is stored on a single platform for unified management and sharing. This can easily lead to system paralysis due to a single location failure.
[0004] 2. Existing multi-dimensional big data sharing methods do not differentiate between different data generated on the same platform, resulting in a lack of targeted sharing, which can easily lead to low sharing efficiency and waste of resources;
[0005] To this end, we propose a multivariate big data sharing method and system. Summary of the Invention
[0006] In response to the shortcomings of the prior art, the present invention aims to provide a multivariate big data sharing method and system. The present invention divides the data generated by a target data platform into multiple types of data, obtains historical access frequency data, generation timestamp data, real-time data volume, data period change, and data update frequency corresponding to each type of data, obtains multiple monitoring indicator data, and defines them as multivariate shared data. Based on the multivariate shared data, data type division coefficients corresponding to different types of data are calculated. A data type division coefficient threshold is obtained and numerically compared for each data type division coefficient. The different types of data are divided into first type shared data and second type shared data to obtain shared type division data. Based on the multivariate shared data, a data sharing sorting sequence corresponding to each first type of shared data is obtained. The target data platform shares the first type shared data to a data receiving terminal according to the data sharing sorting sequence, and divides the second type shared data into first terminal type shared data and second terminal type shared data according to the number of data receiving terminals. The target data platform synchronously shares the first terminal type shared data to multiple data receiving terminals through a star connection, obtains reception layer data corresponding to the second terminal type shared data, and shares the second terminal type shared data to multiple data receiving terminals according to the reception layer data.
[0007] In order to achieve the above-mentioned purpose, the present invention adopts the following technical solution: a multivariate big data sharing method includes the following specific steps:
[0008] Step F1: Divide the data generated by the target data platform into multiple types of data. Obtain the historical access frequency data, generation timestamp data, real-time data volume, data period change, and data update frequency corresponding to each type of data, obtain multiple monitoring indicator data, and define them as multivariate shared data.
[0009] Step F2: Calculate the data type division coefficients corresponding to different types of data based on the multivariate shared data, obtain the data type division coefficient threshold, perform numerical comparison on each data type division coefficient, divide the different types of data into first type shared data and second type shared data, and obtain shared type division data;
[0010] Step F3: Obtaining a data sharing ranking sequence corresponding to each first type of shared data based on the multivariate shared data analysis, and the target data platform shares the first type of shared data to the data receiving terminal based on the data sharing ranking sequence;
[0011] Step F4: Divide the second type of shared data into first terminal type shared data and second terminal type shared data according to the number of data receiving terminals. The target data platform synchronously shares the first terminal type shared data to multiple data receiving terminals through a star connection, obtains the receiving layered data corresponding to the second terminal type shared data, and the target data platform shares the second terminal type shared data to multiple data receiving terminals according to the receiving layered data.
[0012] Furthermore, the step F1 further includes the following specific steps:
[0013] Step F11: Divide various types of data generated by the target data platform into several types of data, and name them as type k1 to type kk data respectively;
[0014] Step F12: performing data generation monitoring on the k1th type of data to obtain the k1th monitoring indicator data;
[0015] Step F13: Obtain the historical access frequency data, generation timestamp data, real-time data volume, data time period change, and data update frequency corresponding to the k2 to kk types of data, and obtain the k2 to kk monitoring indicator data;
[0016] Step F14: Define the monitoring indicator data from k1 to kk as multivariate shared data.
[0017] Furthermore, the step F12 further includes the following specific steps:
[0018] In the process of monitoring the k1-th type of data, the generation time of each sub-data in the k1-th type of data is obtained to obtain multiple sub-data generation timestamps, which are named generation timestamp data;
[0019] Obtain the historical access frequency corresponding to each sub-data in the k1-th type of data, obtain multiple sub-data historical access frequencies, and collectively name them as historical access frequency data;
[0020] Within the time range of monitoring the k1th type of data, mark multiple characteristic monitoring periods, and the time interval and monitoring duration of each two consecutive characteristic monitoring periods are the same, and they are named the first to jth characteristic monitoring periods respectively, and the duration of the characteristic monitoring periods is obtained to obtain the characteristic monitoring duration value;
[0021] Obtain the cumulative number of updates of the k1th type of data corresponding to the first to jth feature monitoring periods, obtain the cumulative number of updates of the first to jth data, calculate the average of the cumulative number of updates of the first to jth data, and then calculate the ratio of the obtained average to the feature monitoring duration value to obtain the data update frequency;
[0022] Performing data volume detection on the k1th type of data at the first characteristic monitoring moment to obtain a first monitoring change data volume;
[0023] The details are as follows:
[0024] Mark the time value corresponding to the start time of the first characteristic monitoring period as the first monitoring time value, and mark the time value corresponding to the end time of the first characteristic monitoring period as the second monitoring time value;
[0025] Obtain the data volume of the k1th type data corresponding to the first monitoring time value and the second monitoring time value, and name them as the first monitoring data volume and the second monitoring data volume respectively;
[0026] Calculate the difference between the second monitoring data volume and the first monitoring data volume, and then calculate the ratio of the obtained difference to the characteristic monitoring duration value to obtain the first monitoring change data volume;
[0027] Acquire the monitoring change data amounts corresponding to the second to j-th characteristic monitoring periods respectively to obtain the second to j-th monitoring data change amounts;
[0028] Calculate the average value of the change in the first to jth monitoring data to obtain the change in the data period;
[0029] Obtain the data volume corresponding to the k1th type of data in the target data platform in real time to obtain the real-time data volume;
[0030] The historical access frequency data, generation timestamp data, real-time data volume, data period change volume and data update frequency corresponding to the k1th type of data are defined as the k1th monitoring indicator data.
[0031] Furthermore, the step F2 further includes the following specific steps:
[0032] Step F21: Acquire multi-element shared data, and acquire k1 to kkth monitoring indicator data according to the multi-element shared data;
[0033] Step F22: obtaining the real-time data volume, data period change, and data update frequency corresponding to the k1-th type of data according to the k1-th monitoring indicator data;
[0034] Step F23: Calculate the data type division coefficient corresponding to the k1th type of data by using the real-time data volume, the data time period change volume, and the data update frequency;
[0035] Calculate the data type division coefficient corresponding to the k1th type of data. The specific formula configuration is as follows:
[0036] Slx=Sjl+Sbl+Sgp*a1;
[0037] Wherein, Slx is the data type division coefficient corresponding to the k1th type of data, Sjl is the real-time data volume, Sbl is the data period change, Sgp is the data update frequency, a1 is the set proportional coefficient and a1 is greater than 0;
[0038] Step F24: obtaining data type division coefficients corresponding to the k2th to kkth type data respectively according to the k2th to kkth monitoring indicator data;
[0039] Step F25: Obtain the data type division coefficient threshold, perform numerical comparison on the data type division coefficients corresponding to the k1 to kk type data respectively with the data type division coefficient threshold to obtain shared type division data.
[0040] Furthermore, the step F25 further includes the following specific steps:
[0041] Step F251: respectively obtaining a real-time data volume threshold, a data period change threshold, and a data update frequency threshold;
[0042] Step F252: Calculating the data type division coefficient threshold by using the real-time data volume threshold, the data period change threshold, and the data update frequency threshold;
[0043] Step F253: When the data type division coefficient is greater than or equal to the data type division coefficient threshold, determining that the corresponding type of data is the first type of shared data;
[0044] Step F254: When the data type division coefficient is less than the data type division coefficient threshold, the corresponding type data is determined to be the second type shared data.
[0045] Furthermore, the step F3 further includes the following specific steps:
[0046] Step F31: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0047] Step F32: Acquire sharing type classification data;
[0048] Step F33: when the ki-th type of data is the first type of shared data, data sharing is performed on the ki-th type of data according to the ki-th monitoring indicator data;
[0049] The step F33 further includes the following specific steps:
[0050] Step F331: naming the multiple sub-data in the ki-th type of data as i1-th to ii-th sub-data respectively;
[0051] Step F332: constructing a distributed sharing framework through a distributed file management platform, and selecting multiple data nodes in the HDFS library in the distributed sharing framework;
[0052] Step F333: Obtain the i-th data sharing sorting sequence;
[0053] Step F334: Divide the i1th to iith sub-data into ii data blocks, and copy them to multiple data nodes respectively to obtain i1th to iith sub-nodes, and sort the i1th to iith sub-nodes according to the ith data shared sorting sequence;
[0054] Step F335: Acquire the APIs corresponding to the child nodes i1 to ii, and obtain the child node API data;
[0055] Step F336: A Kafka group is established through a distributed sharing framework. The target data platform publishes the sub-node API data in the Kafka group. The Kafka group pushes the sub-node API data to the data receiving terminal according to the i-th data sharing sorting sequence.
[0056] Step F337: Repeat the sharing process for the i1th sub-data, and share each sub-data in each first type shared data respectively.
[0057] Furthermore, the step F333 further includes the following specific steps:
[0058] According to the ki monitoring indicator, obtain the sub-data historical access frequency and sub-data generation timestamp corresponding to each data item in the ki type data;
[0059] Name the multiple sub-data in the ki-th type of data as the i1-th to ii-th sub-data, mark the historical access frequencies of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th access frequencies, and mark the timestamps of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th timestamps;
[0060] Get the time value corresponding to the current moment and get the current time value;
[0061] The data sharing ranking coefficient corresponding to the i1th sub-data is obtained by calculating the i1th access frequency, the i1th timestamp, and the current time value, and is named the i1th ranking coefficient;
[0062] Calculate the data sharing ranking coefficient corresponding to the i1th sub-data. The specific formula is as follows:
[0063] Spx = |Sc-Sd|*Fc;
[0064] Wherein, Spx is the data sharing ranking coefficient corresponding to the i1th sub-data, Sc is the time value corresponding to the i1th timestamp, Sd is the current time value, and Fc is the i1th access frequency;
[0065] Calculate the data sharing ranking coefficients corresponding to the i2th to iith sub-data respectively according to the i2th to iith access frequencies, the i2th to iith timestamps, and the current time value to obtain the i2th to iith ranking coefficients;
[0066] Arrange the i1th to iith sorting coefficients in descending order according to their numerical values to obtain the ith data sharing sorting sequence.
[0067] Furthermore, the step F4 further includes the following specific steps:
[0068] Step F41: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0069] Step F42: Acquire sharing type classification data;
[0070] Step F42: when the kq-th type data is the second type shared data, data sharing is performed on the ki-th type data according to the kq-th monitoring indicator data;
[0071] Step F43: Acquire the number of data receiving terminals corresponding to the kq-th type of data to obtain the number of shared terminals;
[0072] Step F44: when the number of data receiving terminals is greater than p, the kq-th type of data is classified as first terminal type shared data;
[0073] Step F45: when the number of data receiving terminals is less than or equal to p, the kq-th type of data is classified as second terminal type shared data;
[0074] Step F46: when the kq-th type data is the first terminal type shared data, data sharing is performed on the kq-th type data;
[0075] The details are as follows:
[0076] Step F461: With the target data platform as the master node and multiple data receiving terminals as branch nodes, the master node and the multiple branch nodes are connected in a star shape. The master node synchronously shares the multiple sub-data in the kq-th type data to the multiple branch nodes in the order of the timestamps of the sub-data generation;
[0077] Step F47: when the kq-th type data is the second terminal type shared data, data sharing is performed on the kq-th type data;
[0078] The details are as follows:
[0079] Step F471: Acquire received hierarchical data;
[0080] Step F472: With the target data platform as the main node and multiple data receiving terminals as branch nodes, the main node and the multiple branch nodes are hierarchically connected according to the received hierarchical data;
[0081] Step F473: The master node synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the second-layer receiving end hierarchy in the order of the sub-data generation timestamps. The branch node corresponding to the second-layer receiving end hierarchy synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the third-layer receiving end hierarchy in the order of the sub-data generation timestamps. This sharing process is repeated until the multiple sub-data in the kq-th type data are shared to the branch node corresponding to the e+1-th layer receiving end hierarchy.
[0082] Step F48: Repeat the sharing process for the kq-th type of data, and share each sub-data in each second type of shared data respectively.
[0083] Furthermore, the step F471 further includes the following specific steps:
[0084] Name the multiple data receiving terminals as first and z receiving terminals respectively;
[0085] Performing terminal monitoring on the first receiving terminal to obtain a terminal monitoring coefficient corresponding to the first receiving terminal;
[0086] The details are as follows:
[0087] The unit time length before the current moment is used as the terminal monitoring cycle;
[0088] Obtaining a total number of accesses to the kqth type of data by the first receiving terminal within a terminal monitoring period to obtain a first total number of accesses;
[0089] Acquire the number of periodic online users of the first receiving terminal in the terminal monitoring period to obtain the first number of online users;
[0090] Acquire a data reception delay of a first receiving terminal to obtain a first reception delay;
[0091] The details are as follows:
[0092] Sending test data to the first receiving terminal through the upper-level sending terminal of the first receiving terminal, and acquiring a time value corresponding to the sending moment to obtain a first time value;
[0093] The first receiving terminal receives the test data and obtains a time value corresponding to the receiving moment to obtain a second time value;
[0094] Calculate the difference between the second time value and the first time value to obtain a first receiving delay;
[0095] Obtaining a terminal monitoring coefficient corresponding to the first receiving terminal by calculating the first total number of visits, the first number of online users, and the first receiving delay;
[0096] The terminal monitoring coefficient corresponding to the first receiving terminal is calculated. The specific formula is as follows:
[0097] Zjx1=(Fws1×Sy1)+Yhs1;
[0098] Where Zjx1 is the terminal monitoring coefficient corresponding to the first receiving terminal, Fws1 is the first total number of visits, Sy1 is the first receiving delay, and Yhs1 is the first number of online users;
[0099] Obtaining the terminal monitoring coefficient corresponding to each receiving terminal respectively;
[0100] Obtain multiple terminal monitoring coefficient thresholds respectively, and arrange them in descending order as the first to e-th terminal monitoring coefficient thresholds, and perform numerical comparison on the terminal monitoring coefficient corresponding to each receiving terminal with the first to e-th terminal monitoring coefficient thresholds respectively to obtain the first to e+1th layer receiving terminal classification, and name it as receiving layered data.
[0101] A multivariate big data sharing system, comprising:
[0102] Data acquisition module: used to obtain multiple monitoring indicator data and obtain multi-dimensional shared data;
[0103] Data analysis module: used to analyze multivariate shared data, divide multiple types of data into first type shared data and second type shared data, and obtain shared type divided data;
[0104] The first sharing module is used to share the first type of shared data according to the multi-shared data and the shared type classification data.
[0105] The second sharing module is used to perform big data sharing of the second type of shared data according to the multi-dimensional shared data and the shared type classification data.
[0106] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0107] 1. The present invention divides the data to be shared into first-type shared data and second-type shared data by obtaining data type division coefficients, and obtains the data sharing sorting sequence corresponding to the first-type shared data for distributed data sharing, which can effectively ensure data sharing efficiency and reduce the risk of data loss;
[0108] 2. The present invention adopts different sharing modes for the second type of shared data according to the number of data receiving terminals, which can effectively improve the pertinence of data sharing and the efficiency of using shared resources. BRIEF DESCRIPTION OF THE DRAWINGS
[0109] To facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.
[0110] Figure 1 It is a diagram of the implementation steps of the present invention;
[0111] Figure 2 is a block diagram of the overall system of the present invention;
[0112] Figure 3 This is a star connection diagram of the present invention;
[0113] Figure 4 Schematic diagram of the hierarchical connection of the present invention. DETAILED DESCRIPTION
[0114] The technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0115] Example 1
[0116] See also Figure 2 The present invention provides a technical solution: a multivariate big data sharing system, comprising a data acquisition module, a data analysis module, a first sharing module, a second sharing module and a server, wherein the data acquisition module, the data analysis module, the first sharing module and the second sharing module are respectively connected to the server, and the server controls the data acquisition module, the data analysis module, the first sharing module and the second sharing module respectively;
[0117] The data acquisition module acquires multiple shared data;
[0118] Divide the various types of data generated by the target data platform into several types of data and name them as type k1 to type kk data respectively;
[0119] It should be noted here that:
[0120] In the present invention, k mentioned here specifically refers to the quantity value corresponding to the type data, and k is an integer greater than 0;
[0121] In the present invention, the type data needs to be specifically determined according to the specific data platform type;
[0122] For example, if the target data platform is an online shopping platform, the k1th type of data can be product price data, the k2th type of data can be product descriptions, and the k3th type of data can be sales records corresponding to the product.
[0123] Perform data generation monitoring on the k1th type of data to obtain the k1th monitoring indicator data;
[0124] In the process of monitoring the k1-th type of data, the generation time of each sub-data in the k1-th type of data is obtained to obtain multiple sub-data generation timestamps, which are named generation timestamp data;
[0125] Obtain the historical access frequency corresponding to each sub-data in the k1-th type of data, obtain multiple sub-data historical access frequencies, and collectively name them as historical access frequency data;
[0126] It should be noted here that:
[0127] In the present invention, the various data involved include existing versions and historical versions. The historical access frequency here refers to the access frequency corresponding to the historical versions of various data;
[0128] For example, if the k1 type of data is the price of a commodity, and the price of the commodity fluctuates, the historical price is the historical version of this data, and the current price is the current version of this data.
[0129] Within the time range of monitoring the k1-th type of data, mark multiple characteristic monitoring periods, and the time interval and monitoring duration of each two consecutive characteristic monitoring periods are the same, and they are named the first to j-th characteristic monitoring periods respectively;
[0130] Obtain the duration of the characteristic monitoring period to obtain the characteristic monitoring duration value;
[0131] Obtain the cumulative number of updates of the k1th type of data corresponding to the first to jth feature monitoring periods, obtain the cumulative number of updates of the first to jth data, calculate the average of the cumulative number of updates of the first to jth data, and then calculate the ratio of the obtained average to the feature monitoring duration value to obtain the data update frequency;
[0132] It should be noted here that:
[0133] In the present invention, j is the quantity value corresponding to the characteristic monitoring period, and j is an integer greater than 0;
[0134] Performing data volume detection on the k1th type of data at the first characteristic monitoring moment to obtain a first monitoring change data volume;
[0135] The details are as follows:
[0136] Mark the time value corresponding to the start time of the first characteristic monitoring period as the first monitoring time value, and mark the time value corresponding to the end time of the first characteristic monitoring period as the second monitoring time value;
[0137] Obtain the data volume of the k1th type data corresponding to the first monitoring time value and the second monitoring time value, and name them as the first monitoring data volume and the second monitoring data volume respectively;
[0138] Calculate the difference between the second monitoring data volume and the first monitoring data volume, and then calculate the ratio of the obtained difference to the characteristic monitoring duration value to obtain the first monitoring change data volume;
[0139] Acquire the monitoring change data amounts corresponding to the second to j-th characteristic monitoring periods respectively to obtain the second to j-th monitoring data change amounts;
[0140] Calculate the average value of the change in the first to jth monitoring data to obtain the change in the data period;
[0141] Obtain the data volume corresponding to the k1th type of data in the target data platform in real time to obtain the real-time data volume;
[0142] It should be noted here that:
[0143] In the present invention, the data volume involved specifically refers to the size or capacity of the data contained in the storage device or transmission pipeline;
[0144] The historical access frequency data, generation timestamp data, real-time data volume, data period change volume, and data update frequency corresponding to the k1th type of data are defined as the k1th monitoring indicator data;
[0145] Obtain the historical access frequency data, generation timestamp data, real-time data volume, data period change, and data update frequency corresponding to the k2 to kk types of data, and obtain the k2 to kk monitoring indicator data;
[0146] The monitoring indicator data from k1 to kk are defined as multivariate shared data;
[0147] The data acquisition module acquires the multi-dimensional shared data and transmits it to the data analysis module, the first sharing module and the second sharing module;
[0148] The data analysis module analyzes the multivariate shared data to obtain the shared type classification data;
[0149] Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0150] According to the k1th monitoring indicator data, obtain the real-time data volume, data period change and data update frequency corresponding to the k1th type of data;
[0151] The data type division coefficient corresponding to the k1th type of data is obtained by calculating the real-time data volume, the data period change volume and the data update frequency;
[0152] Calculate the data type division coefficient corresponding to the k1th type of data. The specific formula configuration is as follows:
[0153] Slx=Sjl+Sbl+Sgp*a1;
[0154] Wherein, Slx is the data type division coefficient corresponding to the k1th type of data, Sjl is the real-time data volume, Sbl is the data period change, Sgp is the data update frequency, a1 is the set proportional coefficient and a1 is greater than 0;
[0155] Here are some things to note:
[0156] The a1 involved here adjusts the specific value of the data update frequency to a value range that matches the real-time data volume and the change in data period through proportional adjustment. In actual application, it needs to be set accordingly based on the size relationship and weight relationship between different values;
[0157] Obtain the data type division coefficients corresponding to the k2th to kkth type data respectively according to the k2th to kkth monitoring indicator data;
[0158] Obtain a data type division coefficient threshold, compare the data type division coefficients corresponding to the k1 to kk types of data with the data type division coefficient threshold, and obtain shared type division data;
[0159] The details are as follows:
[0160] Obtain the real-time data volume threshold, data period change threshold, and data update frequency threshold respectively;
[0161] The data type division coefficient threshold is obtained by calculating the real-time data volume threshold, the data period change threshold, and the data update frequency threshold;
[0162] Calculate the data type division coefficient threshold. The specific formula configuration is as follows:
[0163] Slxy=Sjly+Sbly+Sgpy*a1;
[0164] Among them, Slxy is the data type classification coefficient threshold, Sjly is the real-time data volume threshold, Sbly is the data period change threshold, Sgpy is the data update frequency threshold, a1 is the set proportional coefficient and a1 is greater than 0;
[0165] The numerical comparison process is as follows:
[0166] When the data type division coefficient is greater than or equal to the data type division coefficient threshold, determining that the corresponding type data is the first type shared data;
[0167] When the data type division coefficient is less than the data type division coefficient threshold, determining that the corresponding type data is the second type shared data;
[0168] It should be noted here that:
[0169] In the present invention, the data update quantity, real-time data volume, and data generation speed corresponding to the first type of shared data are all greater than the data update quantity, real-time data volume, and data generation speed corresponding to the second type of shared data;
[0170] The data analysis module obtains the sharing type classification data and transmits it to the first sharing module and the second sharing module;
[0171] The first sharing module performs big data sharing on the first type of shared data according to the multi-shared data and the shared type classification data;
[0172] The first sharing module includes a target data platform and multiple data receiving terminals;
[0173] Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0174] Get sharing type classification data;
[0175] When the ki-th type of data is the first type of shared data, data sharing is performed on the ki-th type of data according to the ki-th monitoring indicator data;
[0176] The details are as follows:
[0177] Name the multiple sub-data in the ki-th type of data as the i1-th to ii-th sub-data respectively;
[0178] A distributed sharing framework is built through a distributed file management platform, and multiple data nodes are selected from the HDFS library in the distributed sharing framework.
[0179] Divide the i1th to iith sub-data into ii data blocks, and copy them to multiple data nodes respectively to obtain i1th to iith sub-nodes, and sort the i1th to iith sub-nodes according to the ith data shared sorting sequence;
[0180] It should be noted here that:
[0181] In the present invention, ii is the number value corresponding to the sub-data in the ki-th type data, and ii is an integer greater than 0;
[0182] Obtain the i-th data sharing sorting sequence;
[0183] The details are as follows:
[0184] According to the ki monitoring indicator, obtain the sub-data historical access frequency and sub-data generation timestamp corresponding to each data item in the ki type data;
[0185] Name the multiple sub-data in the ki-th type of data as the i1-th to ii-th sub-data, mark the historical access frequencies of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th access frequencies, and mark the timestamps of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th timestamps;
[0186] Get the time value corresponding to the current moment and get the current time value;
[0187] The data sharing ranking coefficient corresponding to the i1th sub-data is obtained by calculating the i1th access frequency, the i1th timestamp, and the current time value, and is named the i1th ranking coefficient;
[0188] Calculate the data sharing ranking coefficient corresponding to the i1th sub-data. The specific formula is as follows:
[0189] Spx = |Sc-Sd|*Fc;
[0190] Wherein, Spx is the data sharing ranking coefficient corresponding to the i1th sub-data, Sc is the time value corresponding to the i1th timestamp, Sd is the current time value, and Fc is the i1th access frequency;
[0191] Calculate the data sharing ranking coefficients corresponding to the i2th to iith sub-data respectively according to the i2th to iith access frequencies, the i2th to iith timestamps, and the current time value to obtain the i2th to iith ranking coefficients;
[0192] Arrange the i1th to iith sorting coefficients in descending order according to their numerical values to obtain the ith data sharing sorting sequence;
[0193] The HDFS tool obtains the API corresponding to the i1th to iith child nodes respectively, and obtains the child node API data;
[0194] A Kafka group is established through a distributed sharing framework. The target data platform is set as the "publisher" and multiple data receiving terminals are set as "consumers". The target data platform publishes the sub-node API data in the Kafka group. The Kafka group pushes the sub-node API data to the data receiving terminals according to the i-th data sharing sorting sequence.
[0195] Repeat the sharing process for the i1th sub-data, and share each sub-data in each first type of shared data;
[0196] It should be noted here that:
[0197] Kafka provides a "publish-subscribe" model where producers can publish messages to topics and consumers can subscribe to messages from these topics. In this invention, the target data platform is the "publisher" and multiple data receiving terminals are set as "consumers";
[0198] In the present invention, the storage space of each data block needs to be larger than the space for storing the corresponding sub-data;
[0199] For example: if the i1th sub-data is stored in the i1th data block, and the storage space of the i1th data block needs to be larger than the data volume of the i1th sub-data;
[0200] The second sharing module performs big data sharing on the second type of shared data according to the multi-shared data and the shared type classification data;
[0201] The second data contribution module includes a target data platform and multiple data receiving terminals;
[0202] Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0203] Get sharing type classification data;
[0204] When the kqth type of data is the second type of shared data, data sharing is performed on the kith type of data according to the kqth monitoring indicator data;
[0205] Obtain the number of data receiving terminals corresponding to the kq-th type of data to obtain the number of shared terminals;
[0206] When the number of data sharing terminals is greater than p, the kq-th type of data is classified as first terminal type shared data;
[0207] When the number of data sharing terminals is less than or equal to p, the kq-th type of data is classified as the second terminal type shared data;
[0208] It should be noted here that:
[0209] In this application, p is the division index corresponding to the number of data sharing terminals. In actual applications, the specific data of p needs to be divided according to the specific type of transmitted data;
[0210] When the kq-th type data is the first terminal type shared data;
[0211] See also Figure 3 , with the target data platform as the main node and multiple data receiving terminals as branch nodes, the main node and multiple branch nodes are connected in a star shape, and the main node synchronously shares multiple sub-data in the kq type data to multiple branch nodes in the order of the time stamps of the sub-data;
[0212] When the kq-th type data is the second terminal type shared data;
[0213] See also Figure 4 , with the target data platform as the main node and multiple data receiving terminals as branch nodes, the main node and the multiple branch nodes are hierarchically connected according to the received hierarchical data, the main node synchronously shares the multiple sub-data in the kq-th type data to the branch nodes corresponding to the second-layer receiving terminal hierarchy in the order of the time when the sub-data generated the timestamps, the branch nodes corresponding to the second-layer receiving terminal hierarchy synchronously share the multiple sub-data in the kq-th type data to the branch nodes corresponding to the third-layer receiving terminal hierarchy in the order of the time when the sub-data generated the timestamps, repeating this sharing process until the multiple sub-data in the kq-th type data are shared to the branch nodes corresponding to the e+1-th layer receiving terminal hierarchy;
[0214] Acquiring received hierarchical data;
[0215] The details are as follows:
[0216] Name the multiple data receiving terminals as first and z receiving terminals respectively;
[0217] It should be noted here that:
[0218] In this application, z is the number value corresponding to the receiving terminal, and z is an integer greater than 0;
[0219] Performing terminal monitoring on the first receiving terminal to obtain a terminal monitoring coefficient corresponding to the first receiving terminal;
[0220] The details are as follows:
[0221] The unit time length before the current moment is used as the terminal monitoring cycle;
[0222] It should be noted here that:
[0223] In this application, the unit time length is specifically 10 seconds;
[0224] Obtaining a total number of accesses to the kqth type of data by the first receiving terminal within a terminal monitoring period to obtain a first total number of accesses;
[0225] Acquire the number of periodic online users of the first receiving terminal in the terminal monitoring period to obtain the first number of online users;
[0226] Acquire a data reception delay of a first receiving terminal to obtain a first reception delay;
[0227] The details are as follows:
[0228] Sending test data to the first receiving terminal through the upper-level sending terminal of the first receiving terminal, and acquiring a time value corresponding to the sending moment to obtain a first time value;
[0229] The first receiving terminal receives the test data and obtains a time value corresponding to the receiving moment to obtain a second time value;
[0230] Calculate the difference between the second time value and the first time value to obtain a first receiving delay;
[0231] Obtaining a terminal monitoring coefficient corresponding to the first receiving terminal by calculating the first total number of visits, the first number of online users, and the first receiving delay;
[0232] The terminal monitoring coefficient corresponding to the first receiving terminal is calculated. The specific formula is as follows:
[0233] Zjx1=(Fws1×Sy1)+Yhs1;
[0234] Where Zjx1 is the terminal monitoring coefficient corresponding to the first receiving terminal, Fws1 is the first total number of visits, Sy1 is the first receiving delay, and Yhs1 is the first number of online users;
[0235] Obtaining the terminal monitoring coefficient corresponding to each receiving terminal respectively;
[0236] Obtain multiple terminal monitoring coefficient thresholds respectively and arrange them in descending order as first to e-th terminal monitoring coefficient thresholds, compare the terminal monitoring coefficient corresponding to each receiving terminal with the first to e-th terminal monitoring coefficient thresholds respectively, obtain the first to e+1th layer receiving terminal classification, and name them as receiving layer data;
[0237] It should be noted here that:
[0238] In this application, e is the numerical value corresponding to the terminal monitoring coefficient threshold, and e is an integer greater than 0;
[0239] The first terminal monitoring coefficient threshold is greater than the second terminal monitoring coefficient threshold, the second terminal monitoring coefficient threshold is greater than the third terminal monitoring coefficient threshold, and so on. The e-1th terminal monitoring coefficient threshold is greater than the eth terminal monitoring coefficient threshold;
[0240] For example:
[0241] If the terminal monitoring coefficient corresponding to the f-th receiving end is greater than the second terminal monitoring coefficient threshold and less than or equal to the first terminal monitoring coefficient threshold, then the f-th receiving end is classified as a second-tier receiving end;
[0242] In this application, the number of terminal monitoring coefficient thresholds needs to be specifically determined according to the number of receiving terminal levels, and the specific value of each specific terminal monitoring coefficient threshold needs to be determined accordingly according to the data type and data volume;
[0243] Repeat the sharing process for the kq-th type of data, and share each sub-data in each second type of shared data respectively.
[0244] In this application, if a corresponding calculation formula appears, the above calculation formula is dimensionless and its numerical calculation is performed. The weight coefficient, proportional coefficient and other coefficients in the formula are set to a result value obtained by quantifying each parameter. Regarding the size of the weight coefficient and the proportional coefficient, as long as it does not affect the proportional relationship between the parameter and the result value, it is acceptable.
[0245] Example 2
[0246] See also Figure 1 Based on another concept of the same invention, a multivariate big data sharing method is proposed, comprising the following steps:
[0247] Step F1: Obtain multi-shared data;
[0248] Step F11: Divide various types of data generated by the target data platform into several types of data, and name them as type k1 to type kk data respectively;
[0249] Step F12: performing data generation monitoring on the k1th type of data to obtain the k1th monitoring indicator data;
[0250] The step F12 further includes the following specific steps:
[0251] Step F121: in the process of monitoring the k1-th type of data, the generation time of each sub-data item in the k1-th type of data is acquired to obtain multiple sub-data generation timestamps, which are named generation timestamp data;
[0252] Step F122: Acquire the historical access frequency corresponding to each sub-data item in the k1-th type of data to obtain multiple sub-data historical access frequencies, and collectively name them as historical access frequency data;
[0253] Step F123: Within the time range for monitoring the k1-th type of data, mark multiple characteristic monitoring periods, with the time interval and monitoring duration of each two consecutive characteristic monitoring periods being the same, and name them as the first to j-th characteristic monitoring periods, respectively. Obtain the duration of the characteristic monitoring periods to obtain the characteristic monitoring duration values;
[0254] Step F124: Obtain the cumulative update times of the k1th type of data corresponding to the first to jth characteristic monitoring periods, obtain the cumulative update times of the first to jth data, calculate the average of the cumulative update times of the first to jth data, and then calculate the ratio of the obtained average to the characteristic monitoring duration value to obtain the data update frequency;
[0255] Step F125: performing data volume detection on the k1th type of data at the first characteristic monitoring moment to obtain a first monitoring change data volume;
[0256] The details are as follows:
[0257] Step F1251: Mark the time value corresponding to the start time of the first characteristic monitoring period as the first monitoring time value, and mark the time value corresponding to the end time of the first characteristic monitoring period as the second monitoring time value;
[0258] Step F1252: respectively obtaining the data amount of the k1th type data corresponding to the first monitoring time value and the second monitoring time value, and naming them as the first monitoring data amount and the second monitoring data amount respectively;
[0259] Step F1253: Calculate the difference between the second monitoring data volume and the first monitoring data volume, and then calculate the ratio of the obtained difference to the characteristic monitoring duration value to obtain the first monitoring change data volume;
[0260] Step F126: acquiring the monitoring change data amounts corresponding to the second to j-th characteristic monitoring periods respectively, to obtain the second to j-th monitoring data change amounts;
[0261] Step F127: Calculate the average value of the changes in the first to jth monitoring data to obtain the change in the data period;
[0262] Step F128: acquiring the data volume corresponding to the k1th type of data in the target data platform in real time to obtain the real-time data volume;
[0263] Step F129: defining the historical access frequency data, generation timestamp data, real-time data volume, data time period change volume, and data update frequency corresponding to the k1-th type of data as the k1-th monitoring indicator data;
[0264] Step F13: Obtain the historical access frequency data, generation timestamp data, real-time data volume, data time period change, and data update frequency corresponding to the k2 to kk types of data, and obtain the k2 to kk monitoring indicator data;
[0265] Step F14: defining the monitoring indicator data from k1 to kk as multivariate shared data;
[0266] Step F2: Analyze the multivariate sharing data to obtain sharing type classification data;
[0267] Step F21: Acquire multi-element shared data, and acquire k1 to kkth monitoring indicator data according to the multi-element shared data;
[0268] Step F22: obtaining the real-time data volume, data period change, and data update frequency corresponding to the k1-th type of data according to the k1-th monitoring indicator data;
[0269] Step F23: Calculate the data type division coefficient corresponding to the k1th type of data by using the real-time data volume, the data time period change volume, and the data update frequency;
[0270] Calculate the data type division coefficient corresponding to the k1th type of data. The specific formula configuration is as follows:
[0271] Slx=Sjl+Sbl+Sgp*a1;
[0272] Wherein, Slx is the data type division coefficient corresponding to the k1th type of data, Sjl is the real-time data volume, Sbl is the data period change, Sgp is the data update frequency, a1 is the set proportional coefficient and a1 is greater than 0;
[0273] Step F24: obtaining data type division coefficients corresponding to the k2th to kkth type data respectively according to the k2th to kkth monitoring indicator data;
[0274] Step F25: obtaining a data type division coefficient threshold, comparing the data type division coefficients corresponding to the k1-kk types of data with the data type division coefficient threshold to obtain shared type division data;
[0275] The details are as follows:
[0276] Step F251: respectively obtaining a real-time data volume threshold, a data period change threshold, and a data update frequency threshold;
[0277] Step F252: Calculating the data type division coefficient threshold by using the real-time data volume threshold, the data period change threshold, and the data update frequency threshold;
[0278] Step F253: When the data type division coefficient is greater than or equal to the data type division coefficient threshold, determining that the corresponding type of data is the first type of shared data;
[0279] Step F254: When the data type division coefficient is less than the data type division coefficient threshold, determining that the corresponding type of data is the second type of shared data;
[0280] Step F3: performing big data sharing on the first type of shared data according to the multi-dimensional shared data and the shared type classification data;
[0281] The first sharing module includes a target data platform and multiple data receiving terminals;
[0282] Step F31: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0283] Step F32: Acquire sharing type classification data;
[0284] Step F33: when the ki-th type of data is the first type of shared data, data sharing is performed on the ki-th type of data according to the ki-th monitoring indicator data;
[0285] The step F33 further includes the following specific steps:
[0286] Step F331: naming the multiple sub-data in the ki-th type of data as i1-th to ii-th sub-data respectively;
[0287] Step F332: constructing a distributed sharing framework through a distributed file management platform, and selecting multiple data nodes in the HDFS library in the distributed sharing framework;
[0288] Step F333: Obtain the i-th data sharing sorting sequence;
[0289] Step F334: Divide the i1th to iith sub-data into ii data blocks, and copy them to multiple data nodes respectively to obtain i1th to iith sub-nodes, and sort the i1th to iith sub-nodes according to the ith data shared sorting sequence;
[0290] Step F335: Acquire the APIs corresponding to the child nodes i1 to ii, and obtain the child node API data;
[0291] Step F336: A Kafka group is established through a distributed sharing framework. The target data platform publishes the sub-node API data in the Kafka group. The Kafka group pushes the sub-node API data to the data receiving terminal according to the i-th data sharing sorting sequence.
[0292] Step F337: repeating the sharing process for the i1th sub-data, and sharing each sub-data in each first type shared data;
[0293] The step F333 further includes the following specific steps:
[0294] Step F3331: Obtain the sub-data historical access frequency and sub-data generation timestamp corresponding to each item of data in the ki-th type of data according to the ki-th monitoring indicator;
[0295] Step F3332: Name the multiple sub-data in the ki-th type of data as i1-ii-th sub-data, mark the historical access frequencies of the sub-data corresponding to the i1-ii-th sub-data as i1-ii-th access frequencies, and mark the timestamps of the sub-data generation corresponding to the i1-ii-th sub-data as i1-ii-th timestamps;
[0296] Step F3333: Obtain the time value corresponding to the current moment to obtain the current time value;
[0297] Step F3334: Calculate the data sharing ranking coefficient corresponding to the i1th sub-data by using the i1th access frequency, the i1th timestamp, and the current time value, and name it the i1th ranking coefficient;
[0298] Calculate the data sharing ranking coefficient corresponding to the i1th sub-data. The specific formula is as follows:
[0299] Spx = |Sc-Sd|*Fc;
[0300] Wherein, Spx is the data sharing ranking coefficient corresponding to the i1th sub-data, Sc is the time value corresponding to the i1th timestamp, Sd is the current time value, and Fc is the i1th access frequency;
[0301] Step F3335: Calculate the data sharing ranking coefficients corresponding to the i2th to iith sub-data respectively according to the i2th to iith access frequencies, the i2th to iith timestamps, and the current time value to obtain the i2th to iith ranking coefficients;
[0302] Step F3336: Arrange the i1th to iith sorting coefficients in descending order according to their numerical values to obtain the ith data sharing sorting sequence;
[0303] Step F4: performing big data sharing on the second type of shared data according to the multi-shared data and the shared type classification data;
[0304] Step F41: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data;
[0305] Step F42: Acquire sharing type classification data;
[0306] Step F42: when the kq-th type data is the second type shared data, data sharing is performed on the ki-th type data according to the kq-th monitoring indicator data;
[0307] Step F43: Acquire the number of data receiving terminals corresponding to the kq-th type of data to obtain the number of shared terminals;
[0308] Step F44: when the number of data sharing terminals is greater than p, the kq-th type of data is classified as first terminal type shared data;
[0309] Step F45: when the number of data sharing terminals is less than or equal to p, the kq-th type of data is classified as second terminal type shared data;
[0310] Step F46: when the kq-th type data is the first terminal type shared data, data sharing is performed on the kq-th type data;
[0311] The details are as follows:
[0312] Step F461: With the target data platform as the master node and multiple data receiving terminals as branch nodes, the master node and the multiple branch nodes are connected in a star shape. The master node synchronously shares the multiple sub-data in the kq-th type data to the multiple branch nodes in the order of the timestamps of the sub-data generation;
[0313] Step F47: when the kq-th type data is the second terminal type shared data, data sharing is performed on the kq-th type data;
[0314] The details are as follows:
[0315] Step F471: Acquire received hierarchical data;
[0316] Step F472: With the target data platform as the main node and multiple data receiving terminals as branch nodes, the main node and the multiple branch nodes are hierarchically connected according to the received hierarchical data;
[0317] Step F473: The master node synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the second-layer receiving end hierarchy in the order of the sub-data generation timestamps. The branch node corresponding to the second-layer receiving end hierarchy synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the third-layer receiving end hierarchy in the order of the sub-data generation timestamps. This sharing process is repeated until the multiple sub-data in the kq-th type data are shared to the branch node corresponding to the e+1-th layer receiving end hierarchy.
[0318] The step F471 further includes the following specific steps:
[0319] Step F4711: naming multiple data receiving terminals as first and zth receiving terminals respectively;
[0320] Step F4712: Perform terminal monitoring on the first receiving terminal to obtain a terminal monitoring coefficient corresponding to the first receiving terminal;
[0321] The details are as follows:
[0322] Step F47121: taking the unit time length before the current moment as the terminal monitoring period;
[0323] Step F47122: Obtain the total number of accesses to the kq-th type of data by the first receiving terminal within the terminal monitoring period to obtain a first total number of accesses;
[0324] Step F47123: Acquire the number of periodic online users of the first receiving terminal in the terminal monitoring period to obtain the number of first online users;
[0325] Step F47124: Acquire the data receiving delay of the first receiving terminal to obtain a first receiving delay;
[0326] The details are as follows:
[0327] Step F471241: Sending test data to the first receiving terminal through the upper-level sending terminal of the first receiving terminal, and acquiring the time value corresponding to the sending moment to obtain a first time value;
[0328] Step F471242: The first receiving terminal receives the test data and obtains the time value corresponding to the receiving moment to obtain a second time value;
[0329] Step F471243: Calculate the difference between the second time value and the first time value to obtain a first receiving delay;
[0330] Step F47125: Calculate the first total number of visits, the first number of online users, and the first receiving delay to obtain a terminal monitoring coefficient corresponding to the first receiving terminal;
[0331] The terminal monitoring coefficient corresponding to the first receiving terminal is calculated. The specific formula is as follows:
[0332] Zjx1=(Fws1×Sy1)+Yhs1;
[0333] Where Zjx1 is the terminal monitoring coefficient corresponding to the first receiving terminal, Fws1 is the first total number of visits, Sy1 is the first receiving delay, and Yhs1 is the first number of online users;
[0334] Step F47126: Acquire the terminal monitoring coefficient corresponding to each receiving terminal respectively;
[0335] Step F47127: Obtain multiple terminal monitoring coefficient thresholds respectively and arrange them in descending order as the first to e-th terminal monitoring coefficient thresholds. Compare the terminal monitoring coefficient corresponding to each receiving terminal with the first to e-th terminal monitoring coefficient thresholds respectively to obtain the first to e+1th layer receiving terminal classifications, which are named receiving layer data.
[0336] Step F48: Repeat the sharing process for the kq-th type of data, and share each sub-data in each second type of shared data respectively.
[0337] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to specific embodiments. Obviously, many modifications and variations are possible based on the contents of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.
Claims
1. A multivariate big data sharing method, characterized in that: include: Step F1: Divide the data generated by the target data platform into multiple types of data. Obtain the historical access frequency data, generation timestamp data, real-time data volume, data period change, and data update frequency corresponding to each type of data, obtain multiple monitoring indicator data, and define them as multivariate shared data. Step F2: Calculating data type division coefficients corresponding to different types of data based on the multivariate shared data, obtaining a data type division coefficient threshold, numerically comparing the data type division coefficients corresponding to each type of data with the data type division coefficient threshold, dividing the different types of data into first type shared data and second type shared data, and obtaining shared type division data; Step F3: Obtaining a data sharing ranking sequence corresponding to each first type of shared data based on the multivariate shared data analysis, and the target data platform shares the first type of shared data to the data receiving terminal based on the data sharing ranking sequence; Step F4: Divide the second type of shared data into first terminal type shared data and second terminal type shared data according to the number of data receiving terminals. The target data platform synchronously shares the first terminal type shared data to multiple data receiving terminals through a star connection, obtains the receiving layered data corresponding to the second terminal type shared data, and the target data platform shares the second terminal type shared data to multiple data receiving terminals according to the receiving layered data.
2. A multivariate big data sharing method according to claim 1, characterized in that: The step F1 further includes the following specific steps: Step F11: Divide various types of data generated by the target data platform into several types of data, and name them as type k1 to type kk data respectively; Step F12: performing data generation monitoring on the k1th type of data to obtain the k1th monitoring indicator data; Step F13: Obtain the historical access frequency data, generation timestamp data, real-time data volume, data time period change, and data update frequency corresponding to the k2 to kk types of data, and obtain the k2 to kk monitoring indicator data; Step F14: Define the monitoring indicator data from k1 to kk as multivariate shared data.
3. A multivariate big data sharing method according to claim 2, characterized in that: The step F12 further includes the following specific steps: In the process of monitoring the k1-th type of data, the generation time of each sub-data in the k1-th type of data is obtained to obtain multiple sub-data generation timestamps, which are named generation timestamp data; Obtain the historical access frequency corresponding to each sub-data in the k1-th type of data to obtain multiple sub-data historical access frequencies, and name them as historical access frequency data; Within the time range of monitoring the k1th type of data, mark multiple characteristic monitoring periods, and the time interval and monitoring duration of each two consecutive characteristic monitoring periods are the same, and they are named the first to jth characteristic monitoring periods respectively, and the duration of the characteristic monitoring periods is obtained to obtain the characteristic monitoring duration value; Obtain the cumulative number of updates of the k1th type of data corresponding to the first to jth feature monitoring periods, obtain the cumulative number of updates of the first to jth data, calculate the average of the cumulative number of updates of the first to jth data, and then calculate the ratio of the obtained average to the feature monitoring duration value to obtain the data update frequency; Performing data volume detection on the k1th type of data at the first characteristic monitoring moment to obtain a first monitoring change data volume; The details are as follows: Mark the time value corresponding to the start time of the first characteristic monitoring period as the first monitoring time value, and mark the time value corresponding to the end time of the first characteristic monitoring period as the second monitoring time value; Obtain the data volume of the k1th type data corresponding to the first monitoring time value and the second monitoring time value, and name them as the first monitoring data volume and the second monitoring data volume respectively; Calculate the difference between the second monitoring data volume and the first monitoring data volume, and then calculate the ratio of the obtained difference to the characteristic monitoring duration value to obtain the first monitoring change data volume; Acquire the monitoring change data amounts corresponding to the second to j-th characteristic monitoring periods respectively to obtain the second to j-th monitoring data change amounts; Calculate the average value of the change in the first to jth monitoring data to obtain the change in the data period; Obtain the data volume corresponding to the k1th type of data in the target data platform in real time to obtain the real-time data volume; The historical access frequency data, generation timestamp data, real-time data volume, data period change volume and data update frequency corresponding to the k1th type of data are defined as the k1th monitoring indicator data.
4. The multivariate big data sharing method according to claim 1, characterized in that: The step F2 further includes the following specific steps: Step F21: Acquire multi-element shared data, and acquire k1 to kkth monitoring indicator data according to the multi-element shared data; Step F22: obtaining the real-time data volume, data period change, and data update frequency corresponding to the k1-th type of data according to the k1-th monitoring indicator data; Step F23: Calculate the data type division coefficient corresponding to the k1th type of data by using the real-time data volume, the data time period change volume, and the data update frequency; Calculate the data type division coefficient corresponding to the k1th type of data. The specific formula configuration is as follows: Slx=Sjl+Sbl+Sgp*a1; Wherein, Slx is the data type division coefficient corresponding to the k1th type of data, Sjl is the real-time data volume, Sbl is the data period change, Sgp is the data update frequency, a1 is the set proportional coefficient and a1 is greater than 0; Step F24: obtaining data type division coefficients corresponding to the k2th to kkth type data according to the k2th to kkth monitoring indicator data respectively; Step F25: Obtain the data type division coefficient threshold, perform numerical comparison on the data type division coefficients corresponding to the k1 to kk type data respectively with the data type division coefficient threshold to obtain shared type division data.
5. A multivariate big data sharing method according to claim 4, characterized in that: The step F25 further includes the following specific steps: Step F251: obtaining a real-time data volume threshold, a data period change threshold, and a data update frequency threshold; Step F252: Calculating the data type division coefficient threshold by using the real-time data volume threshold, the data period change threshold, and the data update frequency threshold; Step F253: When the data type division coefficient is greater than or equal to the data type division coefficient threshold, determining that the corresponding type of data is the first type of shared data; Step F254: When the data type division coefficient is less than the data type division coefficient threshold, the corresponding type data is determined to be the second type shared data.
6. A multivariate big data sharing method according to claim 1, characterized in that: The step F3 further includes the following specific steps: Step F31: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data; Step F32: Acquire sharing type classification data; Step F33: when the ki-th type of data is the first type of shared data, data sharing is performed on the ki-th type of data according to the ki-th monitoring indicator data; The step F33 further includes the following specific steps: Step F331: naming the multiple sub-data in the ki-th type of data as i1-th to ii-th sub-data respectively; Step F332: constructing a distributed sharing framework through a distributed file management platform, and selecting multiple data nodes in the HDFS library in the distributed sharing framework; Step F333: Obtain the i-th data sharing sorting sequence; Step F334: Divide the i1th to iith sub-data into ii data blocks, copy them to multiple data nodes respectively, obtain i1th to iith sub-nodes, and sort the i1th to iith sub-nodes according to the ith data shared sorting sequence; Step F335: Acquire the APIs corresponding to the child nodes i1 to ii, respectively, to obtain the child node API data; Step F336: A Kafka group is established through a distributed sharing framework. The target data platform publishes the sub-node API data in the Kafka group. The Kafka group pushes the sub-node API data to the data receiving terminal according to the i-th data sharing sorting sequence. Step F337: Sharing each sub-data in each first-type shared data.
7. A multivariate big data sharing method according to claim 6, characterized in that: The step F333 further includes the following specific steps: According to the ki monitoring indicator, obtain the sub-data historical access frequency and sub-data generation timestamp corresponding to each data item in the ki type data; Name the multiple sub-data in the ki-th type of data as the i1-th to ii-th sub-data, mark the historical access frequencies of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th access frequencies, and mark the timestamps of the sub-data corresponding to the i1-th to ii-th sub-data as the i1-th to ii-th timestamps; Get the time value corresponding to the current moment and get the current time value; The data sharing ranking coefficient corresponding to the i1th sub-data is obtained by calculating the i1th access frequency, the i1th timestamp, and the current time value, and is named the i1th ranking coefficient; Calculate the data sharing ranking coefficient corresponding to the i1th sub-data. The specific formula is as follows: Spx = |Sc-Sd|*Fc; Wherein, Spx is the data sharing ranking coefficient corresponding to the i1th sub-data, Sc is the time value corresponding to the i1th timestamp, Sd is the current time value, and Fc is the i1th access frequency; Calculate the data sharing ranking coefficients corresponding to the i2th to iith sub-data respectively according to the i2th to iith access frequencies, the i2th to iith timestamps, and the current time value to obtain the i2th to iith ranking coefficients; Arrange the i1th to iith sorting coefficients in descending order according to their numerical values to obtain the ith data sharing sorting sequence.
8. The multivariate big data sharing method according to claim 1, characterized in that: The step F4 further includes the following specific steps: Step F41: Acquire multivariate shared data, and acquire k1 to kk monitoring indicator data based on the multivariate shared data; Step F42: Acquire sharing type classification data; Step F42: when the kq-th type data is the second type shared data, data sharing is performed on the ki-th type data according to the kq-th monitoring indicator data; Step F43: Acquire the number of data receiving terminals corresponding to the kq-th type of data to obtain the number of shared terminals; Step F44: when the number of data receiving terminals is greater than p, the kq-th type of data is classified as first terminal type shared data; Step F45: when the number of data receiving terminals is less than or equal to p, the kq-th type of data is classified as second terminal type shared data; Step F46: when the kq-th type data is the first terminal type shared data, data sharing is performed on the kq-th type data; The details are as follows: Step F461: With the target data platform as the master node and multiple data receiving terminals as branch nodes, the master node and the multiple branch nodes are connected in a star shape. The master node synchronously shares the multiple sub-data in the kq-th type data to the multiple branch nodes in the order of the timestamps of the sub-data generation; Step F47: when the kq-th type data is the second terminal type shared data, data sharing is performed on the kq-th type data; The details are as follows: Step F471: Acquire received hierarchical data; Step F472: With the target data platform as the main node and multiple data receiving terminals as branch nodes, the main node and the multiple branch nodes are hierarchically connected according to the received hierarchical data; Step F473: The master node synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the second-layer receiving end hierarchy in the order of the sub-data generation timestamps. The branch node corresponding to the second-layer receiving end hierarchy synchronously shares the multiple sub-data in the kq-th type data to the branch node corresponding to the third-layer receiving end hierarchy in the order of the sub-data generation timestamps. This sharing process is repeated until the multiple sub-data in the kq-th type data are shared to the branch node corresponding to the e+1-th layer receiving end hierarchy. Step F48: Repeat the sharing process for the kq-th type of data, and share each sub-data in each second type of shared data respectively.
9. A multivariate big data sharing method according to claim 8, characterized in that: The step F471 further includes the following specific steps: Name the multiple data receiving terminals as first and z receiving terminals respectively; Performing terminal monitoring on the first receiving terminal to obtain a terminal monitoring coefficient corresponding to the first receiving terminal; The details are as follows: The unit time length before the current moment is used as the terminal monitoring cycle; Obtaining a total number of accesses to the kqth type of data by the first receiving terminal within a terminal monitoring period to obtain a first total number of accesses; Acquire the number of periodic online users of the first receiving terminal in the terminal monitoring period to obtain the first number of online users; Acquire a data reception delay of a first receiving terminal to obtain a first reception delay; The details are as follows: Sending test data to the first receiving terminal through the upper-level sending terminal of the first receiving terminal, and acquiring a time value corresponding to the sending moment to obtain a first time value; The first receiving terminal receives the test data and obtains a time value corresponding to the receiving moment to obtain a second time value; Calculate the difference between the second time value and the first time value to obtain a first receiving delay; Obtaining a terminal monitoring coefficient corresponding to the first receiving terminal by calculating the first total number of visits, the first number of online users, and the first receiving delay; The terminal monitoring coefficient corresponding to the first receiving terminal is calculated, and the specific formula is as follows: Zjx1=(Fws1×Sy1)+Yhs1; Where Zjx1 is the terminal monitoring coefficient corresponding to the first receiving terminal, Fws1 is the first total number of visits, Sy1 is the first receiving delay, and Yhs1 is the first number of online users; Obtaining the terminal monitoring coefficient corresponding to each receiving terminal respectively; Obtain multiple terminal monitoring coefficient thresholds respectively, and arrange them in descending order as the first to e-th terminal monitoring coefficient thresholds, and perform numerical comparison on the terminal monitoring coefficient corresponding to each receiving terminal with the first to e-th terminal monitoring coefficient thresholds respectively to obtain the first to e+1th layer receiving terminal classification, and name it as receiving layered data.
10. A multivariate big data sharing system, applicable to a multivariate big data sharing method according to any one of claims 1 to 9, characterized in that: The sharing system includes: Data acquisition module: used to obtain multiple monitoring indicator data and obtain multi-dimensional shared data; Data analysis module: used to analyze multivariate shared data, divide multiple types of data into first type shared data and second type shared data, and obtain shared type divided data; The first sharing module is used to share the first type of shared data according to the multi-shared data and the shared type classification data. The second sharing module is used to perform big data sharing of the second type of shared data according to the multi-dimensional shared data and the shared type classification data.
Citation Information
Patent Citations
Securing access to confidential data using a blockchain ledger
CN110785981A
Fusion sharing platform based on diversified data
CN111813763A