Data processing platform and method based on big data

By designing data acquisition, storage, permission management and continuous monitoring modules on the data processing platform, data dependence, system stability and data quality problems in the prior art are solved, and efficient, secure and reliable data processing and analysis are achieved.

CN120068110AActive Publication Date: 2025-05-30BEIJING QIANZE TECHNOLOGY CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510131052.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-30
Estimated Expiration
2045-02-06

AI Technical Summary

Technical Problem

The existing data processing platform has problems with data dependence, system stability and data quality, resulting in inaccurate analysis results or inability to provide effective services.

Method used

A data processing platform based on big data is designed. The data acquisition module allocates hot data, temperature data and cold data, and encrypts the temperature data in the data storage module to store cold data in a decentralized encrypted storage. The permission management module is used to judge access rights based on the user's usage frequency and behavior model, and continuously monitor the module to mark abnormal users.

Benefits of technology

The caching mechanism is optimized to improve the system response speed and stability; the security and privacy of data are ensured by dynamically adjusting the access control policy; abnormal access behaviors are discovered and identified in a timely manner, and the security and reliability of the system are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068110A_ABST
    Figure CN120068110A_ABST
Patent Text Reader

Abstract

The invention provides a data processing platform and method based on big data, and belongs to the field of data processing. The problem of low data processing efficiency is solved; the method specifically comprises the following steps: a data acquisition module acquires to-be-stored data, and the to-be-stored data is divided into hot data, temperature data and cold data; the data storage module encrypts and stores temperature data and cold data; applying access authority to the warm data and the cold data, and opening the access authority of the hot data; the authority management module is used for calculating the use frequency of each software of each user, constructing a judgment equation and judging whether the access authority of the warm data and the cold data is given to various users or not; the continuous monitoring module marks abnormal users and feeds back the abnormal users; according to the invention, the information security is improved by acquiring, analyzing and processing the related data of the user, storing the data, analyzing the use habit of the user and opening the access authority of different data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A data processing platform and method based on big data according to the present invention relates to the field of data processing. Background Art

[0002] The existing data processing platforms and methods have the following deficiencies:

[0003] Data dependence: The existing data management platforms highly depend on the usage data of users; if the data collected by the system is insufficient or of low quality, it will lead to inaccurate analysis results or the inability to provide effective services for users.

[0004] System stability: The existing data management platforms need to track and analyze the usage behavior data of users in real time, so the requirements for system stability and performance are relatively high; if the system fails or has performance bottleneck problems, it may lead to problems such as users being unable to use the system normally or data loss.

[0005] Data quality: The data collected by the existing data management platforms is easily affected by various factors (such as user misoperations, system failures, etc.), resulting in low data quality or errors; if the system does not take effective data cleaning and integration measures, it may lead to inaccurate analysis results or the inability to provide effective services for users. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, the purpose of the present invention is to provide a data processing platform and method based on big data, aiming to solve the problem of low data processing efficiency.

[0007] To achieve the above purpose, the present invention is realized through the following technical solutions: A data processing platform based on big data includes:

[0008] Data acquisition module: used to obtain the user ID numbers of the data to be stored, obtain the data to be stored corresponding to each user ID number, and divide the data to be stored into hot data, warm data, and cold data;

[0009] Data storage module: used to encrypt the warm data of all users and store it in the local server; perform decentralized encryption on the cold data of all users and store it in the cloud; impose access permissions on the warm data and cold data, and open the access permissions of the hot data;

[0010] Permission management module: used to obtain the working hours of each user for various software, calculate the usage frequency of each software corresponding to each user; based on the usage frequency of each user for each software, use the clustering algorithm to calculate the expected probability of users using various software; construct a judgment equation according to the expected probability, and judge whether to grant access permissions to the warm data and cold data of various users;

[0011] Continuous monitoring module: used to mark users who have not been granted access to warm data as abnormal users and provide feedback.

[0012] Furthermore, the specific process of the data acquisition module is as follows:

[0013] Process A1: Count the number of user ID numbers un; count the number of data to be classified fl for the users from the 1st to the unth. (1) ~fl (un) ;

[0014] Process A2: Take fl (1) as fn, and take the data to be stored corresponding to the 1st user as the target data; classify the target data into hot data, warm data, and cold data;

[0015] Process A3: Repeat the same steps of allocating the data to be classified for the 1st user (i.e., Process A2), classify the data to be classified corresponding to the user ID numbers from the 2nd to the unth into hot data, warm data, and cold data, and enter the data storage module.

[0016] Furthermore, the specific process of the data storage module is as follows:

[0017] Process B1: Encrypt all warm data and store it in the local server;

[0018] Process B2: Count the total number of cold data cc corresponding to the users from the 1st to the unth (1) ~cc (un) ; Calculate the sum of cc (1) ~cc (un) and denote it as acc;

[0019] Denote the sizes of the 1st to the cc (1) th cold data of the 1st user as: fb(1,1)~fb(1,cc (1) );

[0020] And so on, denote the sizes of the 1st to the cc (un) th cold data of the unth user as: fb(un,1)~fb(un,cc (un) );

[0021] Process B3: Calculate and extract the average value af (1) of fb(1,1)~fb(1,cc (1) ), the maximum value mf (1) , and the minimum value lf (1) ;

[0022] And so on, the average value af (un) of fb(un,1)~fb(un,cc (un) ), the maximum value mf(un) , the minimum value lf (un) ;

[0023] Process B4: Denote the total amount of cold data of the i-th user as cc (i) , denote the average value of the cold data size as af (i) , denote the maximum value as mf (i) , denote the minimum value as lf (i) ; The value range of i is: 1 to un;

[0024] Calculate the standard size of data segmentation and denote it as Δfi:

[0025]

[0026] Process B5: According to Δfi, perform decentralized encryption and data segmentation on all the cold data corresponding to the 1st to un-th users, and construct a Merkle tree; According to the node structure of the Merkle tree, store the cold data in the blockchain on the cloud, and feedback the access interface of the blockchain to the 1st to un-th users, and enter the permission management module.

[0027] Furthermore, the working process of the permission management module is as follows:

[0028] Process C1: Obtain the working time wt of the user, the current month m, and the number of working days wd of the user in the (m - 1) month;

[0029] Calculate the usage frequencies ha (1) ~ha (un) ;

[0030] The usage frequency hb of social software (1) ~hb (un) ;

[0031] The usage frequency hc of entertainment software (1) ~hc (un) ;

[0032] The usage frequency hd of other software (1) ~hd (un) ;

[0033] Process C2: Use the clustering algorithm to cluster ha (1) ~ha (un) to obtain a stable parameter cluster;

[0034] Process C3: Count the number of stable parameter clusters and denote it as cu;

[0035] Count the number of parameters in the 1st to cu-th stable parameter clusters and denote it as an (1) ~an (cu);

[0036] Denote the parameters in the 1st to the cnth stable parameter clusters as parameter ct(1, 1) to ct(cu, an (cu) );

[0037] Process C4: Extract the parameters ct(1, 1) to ct(1, an (1) ) in the 1st stable parameter cluster, and calculate the probability Pa(1).

[0038] Furthermore, the subsequent process of the said Process C4 is as follows:

[0039] Process C5: Repeat the same process of calculating Pa(1) to calculate the probabilities Pa(2) to Pa(cu) of the office software corresponding to the 2nd to the cu th stable parameter clusters;

[0040] Calculate the average value Pa of Pa(1) to Pa(cu) as the expected probability of the user using the office software;

[0041] Repeat the same process of calculating Pa to calculate the expected probability Pb of the user using the social software, the expected probability Pc of using the entertainment software, and the expected probability Pd of using other software;

[0042] Process C6: Extract the frequency ha of the 1st user using the office software (1) , the frequency hb of using the social software (1) , the frequency hc of using the entertainment software (1) , and the frequency hd of using other software (1) ;

[0043] Construct the behavior model of the 1st user, and determine whether to grant the permissions of warm data and cold data;

[0044] Process C7: Repeat the same process of determining the permissions of warm data and cold data for the 1st user; Determine whether to grant the permissions of warm data and cold data to the 2nd to the unth users.

[0045] Furthermore, the specific process of the said Process C2 is as follows:

[0046] Process C21: Perform a clustering on ha (1) ~ha (un) ;

[0047] Randomly select kn parameters in ha (1) ~ha (un) as the central parameters: haa (1) ~haa (kn) ; kn represents the number of central parameters;

[0048] Process C22: Define the calculation formula C2-1:

[0049] Among them, ha (i) represents the frequency of the i-th user using office software, haa (k) represents the kn-th central parameter, d (i-k) represents ha (i) relative to haa (k) matching coefficient;

[0050] Process C23: Using calculation formula C2-1, calculate the matching coefficient d of ha (1) ~ha (un) ; cluster ha (1-1) ~d (un-kn) ; (1) Cluster ha

[0051] Process C231: Extract the matching coefficient d (1-1) ~d (1-kn) ;

[0052] Judge whether the minimum value in d (1-1) ~d (1-kn) is unique, and enter different processes;

[0053] Process C232: If the minimum value in d (1-1) ~d (1-kn) is unique;

[0054] Extract the minimum value dl in d (1-1) ~d (1-kn) ; Divide the parameter ha (1) into the central parameter corresponding to dl;

[0055] Process C233: If the minimum value in d (1-1) ~d (1-kn) is not unique; Take the equal minimum values in d (1-1) ~d (1-kn) and denote them as dll (1) ~dll (s) ; Among them, s represents the number of equal minimum values;

[0056] Process C2331: Extract the central parameters corresponding to dll (1) ~dll (s) as the to-be-determined parameters; Denote the to-be-determined parameters as cel (1) ~cel (s) ;

[0057] Define relation C2-2:

[0058] (|cel (i) |-|dll (i) |)=|cel (i)-dll (i) |;

[0059] Among them, cel (i) represents the i-th undetermined parameter, and dll (i) represents dll (1) to dll (s) the i-th parameter in it, and the value range of i is: 1 to s;

[0060] Process C2332: Substitute cel (1) to cel (s) into the relational expression C2-2 to find the undetermined parameters that satisfy the relational expression C2-2, denoted as cey; Divide the central parameter corresponding to the parameter ha (1) into it.

[0061] Furthermore, the subsequent process of the process C23 is as follows:

[0062] Process C24: Repeatedly perform the same process of clustering the parameter ha (1) cluster the parameter ha (2) to ha (un) to obtain the 1st to the kn-th clusters, denoted as cluster c (1) to cluster c (kn) ;

[0063] Among them, ha (2) represents the frequency of the 2nd user using the office software;

[0064] Process C25: Calculate the average parameters of the 1st to the kn-th clusters to obtain ac (1) to ac (kn) ;

[0065] Define the relational expression C2-3:

[0066] (|haa (i) |-|ac (i) )×|haa (i) -ac (i) |>0;

[0067] Among them, haa (i) represents the i-th central parameter, and ac (i) represents the average parameter of the i-th cluster, and the value range of i is: 1 to kn;

[0068] Process C26: Substitute haa (1) to haa (kn) and ac (1) to ac (kn) into the relation A3 to determine the new central parameters for secondary clustering;

[0069] If haa(1) and ac (1) If it satisfies the relational expression C2-3, the new center parameter is ac (1) ; If haa (1) and ac (1) do not satisfy the relational expression C2-3, the new center parameter is haa (1) and ac (1) ;

[0070] And so on, if haa (kn) and ac (kn) satisfy the relational expression C2-3, the new center parameter is ac (kn) ; If haa (kn) and ac (kn) do not satisfy the relational expression C2-3, the new center parameter is haa (kn) and ac (kn) ;

[0071] Process C27: Using the new center parameter as the clustering benchmark, repeat the same process of clustering once to perform secondary clustering on ha (1) ~ha (un) until the center parameters of each cluster no longer change, obtaining a stable parameter cluster.

[0072] Furthermore, the specific process of the said Process C4 is as follows:

[0073] Process C41: Create an empty matrix ZO and fill it with ct(1, 1)~ct(1, an (1) ), obtaining the matrix ZE;

[0074] Process C42: Calculate the average value lx of ct(1, 1)~ct(1, an (1) );

[0075] Denote the parameter in the i-th row of the matrix ZE as x(i), and define the calculation formula C2-4:

[0076] Y(i) = x(i) - lx; where Y(i) represents the centralized value of x(i);

[0077] According to the calculation formula C2-4, calculate the center Y of the matrix ZE;

[0078] Process C43: Denote the parameter in the first stable parameter cluster as ct(1, j);

[0079] Denote the covariance matrix of the matrix ZE as E;

[0080] E = [(Y) T ×Y] / (j - 1); Construct the probability density function p(ct (1,j) ), obtaining the calculation formula C2-5:

[0081] Among them, represents the square root of matrix E;

[0082] The exponent of f(ct(1,j)) to the degree:

[0083]

[0084] Among them, I (an(1)) represents the all-ones matrix of (an (1) ×1);

[0085] Process C44: Calculate the probability density of ct(1, 1) to ct(1, an(1)): pp(ct(1, 1)) to pp(ct(1, an (1) ));

[0086] Define the relational expression C2-6:

[0087] Count the number of parameters that satisfy the relational expression 2-6, denoted as au;

[0088] Calculate the probability Pa(1) that the user uses the office software corresponding to the first stable parameter cluster, Pa(1) = au / an (1) .

[0089] Furthermore, the specific process of the process C6 is as follows:

[0090] Process C61: Calculate the ratios of ha (1) , hb (1) , hc (1) and hd (1) to obtain la, lb, lc, and ld; calculate the sum of la, lb, lc, and ld, denoted as ll;

[0091] Process C62: Construct the equation set C3;

[0092] Process C621: Let h’ represent ha (1) , hb (1) , hc (1) and hd (1) ; Let a represent the software currently used by the user; P(a - h’) represents the probability that the first user changes from software a to h , software;

[0093] If a represents the office software, construct the equation C3-1:

[0094]

[0095] If a represents the social software, construct the equation C3-2:

[0096]

[0097] If a represents entertainment software, construct equation C3-3:

[0098]

[0099] If a represents office software, construct equation C3-4:

[0100]

[0101] Equations C3-1 to C3-4, and equation set C3;

[0102] Process C622: Obtain the software currently used by the first user, and combine it with equation set C3 to calculate the frequency of the first user using office software, denoted as ya, the frequency of using social software, denoted as yb, the frequency of using entertainment software, denoted as yc, and the frequency of using other software, denoted as yd;

[0103] Process C623: Arrange office software, social software, entertainment software, and other software in descending order of ya, yb, yc, and yd to obtain an optimized sequence;

[0104] Process C624: Detect whether the subsequent software usage situation of the first user matches the optimized sequence;

[0105] If it matches, open partial warm data access permissions to the first user;

[0106] If it does not match, do not open warm data access permissions.

[0107] A data processing method based on big data includes:

[0108] Step S1: Obtain the user ID number of the data to be stored, obtain the data to be stored corresponding to each user ID number, and divide the data to be stored into hot data, warm data, and cold data;

[0109] Step S2: Encrypt the warm data of all users and store it in the local server; perform decentralized encryption on the cold data of all users and store it in the cloud; impose access permissions on the warm data and cold data, and open the access permissions of the hot data;

[0110] Step S3: Obtain the working time of each user for various software, calculate the usage frequency of each software corresponding to each user; based on the usage frequency of each user for each software, use the clustering algorithm to calculate the expected probability of the user using various software; construct a judgment equation according to the expected probability, and judge whether to grant access permissions to the warm data and cold data of each type of user;

[0111] Continuous monitoring module: used to mark users who have not been granted access to temperature data as abnormal users and provide feedback.

[0112] Step S4: Mark users who have not been granted access to temperature data as abnormal users and provide feedback.

[0113] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0114] Optimized caching: The present invention can dynamically adjust the caching mechanism, cache the data or query results frequently accessed by users locally or in memory, so as to quickly respond when users access again; reduce the response time of the system and the load pressure on the server.

[0115] Access control: The present invention can set different access control policies according to the usage habits and permission levels of users; for users who often use certain sensitive data, the system can strengthen the access control and security audit of these data to ensure the security and privacy of the data.

[0116] Anomaly detection: The present invention can timely detect and identify abnormal access behaviors, such as illegal intrusion, data leakage, etc.; this anomaly detection function can not only help the system take timely measures to prevent data loss or damage, but also provide users with a more secure and reliable data service. BRIEF DESCRIPTION OF THE DRAWINGS

[0117] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, objectives and advantages of the present invention will become more apparent:

[0118] Figure 1 It is a schematic diagram of the platform of the present invention;

[0119] Figure 2 It is a schematic diagram of the method of the present invention;

[0120] Figure 3 It is a schematic diagram of the processing logic of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0121] To make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0122] Embodiment 1:

[0123] Please refer to Figure 1 and Figure 3 , a data processing platform based on big data includes:

[0124] It should be noted that the present invention obtains data by means of "user authorization scanning";

[0125] The implementation steps of the "User Authorization Scanning" method are as follows:

[0126] Step I: The user creates a folder named "Files to be Stored" on the PC side and saves the absolute path of the "Files to be Stored (folder)".

[0127] Step II: The user deposits the user ID number and the data or files protected by the present invention into the "Files to be Stored (folder)" as data to be classified.

[0128] User ID number: used to distinguish the data source of the data to be classified.

[0129] Step III: The present invention automatically searches for the "Files to be Stored (folder)" through the absolute path, and then reads the data or files in the "Files to be Stored (folder)" to obtain the data to be classified.

[0130] Data acquisition module: used to acquire the user ID number of the data to be stored, acquire the data to be stored corresponding to each user ID number, and classify the data to be stored into hot data, warm data, and cold data.

[0131] Process A: The specific process of the data acquisition module is as follows:

[0132] Process A1: Count the number of user ID numbers, denoted as un.

[0133] Count the number of data to be classified of the users from the 1st to the unth user, denoted as: fl (1) ~fl (un) ;

[0134] Process A2: Take fl (1) as fn, and take the data to be stored corresponding to the 1st user as the target data; classify the target data into hot data, warm data, and cold data.

[0135] It should be noted that in the present invention, the "1st user" means: "the user corresponding to the 1st user ID number"; the "2nd user" means: "the user corresponding to the 2nd user ID number"; and so on, the "unth user" means: "the user corresponding to the unth user ID number".

[0136] Process A21: Obtain the sizes of the data from the 1st to the fnth, denoted as b (1) ~b (fn) ;

[0137] Load the io library; use the IOException module in the io library to obtain the access times of the data from the 1st to the fnth, denoted as: r (1) ~r (fn) ; the modification times are denoted as: w (1) ~w (fn) ;

[0138] Process A22: Calculate b (1) ~b (fn) The average value is denoted as ab; calculate r (1) ~r (fn) The average value is denoted as ar; calculate w (1) ~w (fn) The average value is denoted as aw;

[0139] Extract the maximum value in r (1) ~r (fn) and denote the maximum value as r (max) The maximum value is denoted as r (min) ;

[0140] Extract the maximum value in w (1) ~w (fn) and denote the maximum value as w (max) The maximum value is denoted as w (min) ;

[0141] Process A23: Let the size of the i-th target data be b (i) , the access times be r (i) , and the modification times be w (i) ; The value range of i is: 1 to fn;

[0142] Define relation 1: (r i +w i )×b i ≥(ar×aw) 1 / 2 ×ab;

[0143] Substitute r (1) ~r (fn) and w (1) ~w (fn) into relation 1;

[0144] Summarize the target data that does not satisfy relation 1 as cold data; count the number of cold data and denote it as cn;

[0145] Summarize the target data that satisfies relation 1 as non-cold data; count the number of non-cold data and denote it as mn;

[0146] Process A24: Let the size of the j-th non-cold data be bu (j) , the access times be ru (j) , and the modification times be wu (j) ; The value range of j is: 1 to mn;

[0147] Define relation 2: [(ru (j) +r (min) )>2×ar]∩[(w j +w(min) ) > 2 × aw];

[0148] Define relation 3: [(wu (j) + w (max) ) > 2 × aw] ∩ [(ru (j) + r (max) ) > 2 × ar];

[0149] Substitute the access times and modification times corresponding to the non-cold data into relation 2 and relation 3;

[0150] Summarize the non-cold data that only satisfies relation 2 as hot data; summarize the non-cold data that only satisfies relation 3 or simultaneously satisfies relation 2 and relation 3 as warm data;

[0151] Process A3: Repeat the same steps of allocating the to-be-classified data of the first user (i.e., Process A2), and classify the to-be-classified data corresponding to the user ID numbers from the second to the un-th as hot data, warm data, and cold data, and enter the data storage module.

[0152] Data storage module: Used to encrypt the warm data of all users and store it in the local server; perform decentralized encryption on the cold data of all users and store it in the cloud; impose access permissions on the warm data and cold data, and open the access permissions of the hot data;

[0153] Process B: The specific process of the data storage module is as follows:

[0154] Process B1: Use the RAS algorithm to encrypt the (all) warm data corresponding to the first to the un-th users and store it in the local server;

[0155] Process B11: Take the warm data corresponding to the first user as the quasi-data; use the RAS algorithm to encrypt the quasi-data;

[0156] Use the hash algorithm to process the first user ID number to obtain hn;

[0157] Use the Mersenne Twister algorithm to randomly generate two positive prime numbers, denoted as pr and pe; such that pr and pe satisfy the following conditions:

[0158] {lim[(pr + pe) / hn] ≈ 1} ∩ (pr + pe) min ;

[0159] Process B12: Calculate the Euler number of hn, denoted as φ(hn), φ(hn) = (pr - 1) × (pe - 1);

[0160] Randomly generate an integer e less than φ(hn) such that the greatest common divisor of e and φ(hn) is 1, and e is the public key;

[0161] Calculate the modular inverse d of e with respect to φ(hn), d = e (φ(hn)-2) , where d is the private key;

[0162] Process B13: Encrypt all quasi-data using the public key e to obtain the data set h (1) ; (When the user accesses the temperature data corresponding to the first user, decrypt the data set h (1) using the private key d;)

[0163] Process B14: Repeat the same steps of encrypting all the temperature data of the first user to encrypt all the temperature data corresponding to the user IDs from the 2nd to the unth user, and obtain the data sets h (2) to data set h (un) ;

[0164] Store the data sets h (1) to data set h (un) in the local server;

[0165] Process B2: Count the total number of cold data corresponding to the users from the 1st to the unth, denoted as cc (1) ~cc (un) ; Calculate the sum of cc (1) ~cc (un) , denoted as acc;

[0166] Denote the sizes of the first to the cc (1) st cold data of the first user as: fb(1,1)~fb(1,cc (1) );

[0167] Denote the sizes of the first to the cc (2) st cold data of the second user as: fb(2,1)~fb(2,cc (2) );

[0168] And so on, denote the sizes of the first to the cc (un) st cold data of the unth user as: fb(un,1)~fb(un,cc (un) );

[0169] Process B3: Calculate and extract the average value af (1) of fb(1,1)~fb(1,cc (1) ), the maximum value mf (1) , and the minimum value lf (1) ;

[0170] The average value af (2) of fb(2,1)~fb(2,cc (2) ), the maximum value mf (2), minimum value lf (2) ;

[0171] And so on, the average value af of fb(un,1) to fb(un,cc (un) ), maximum value mf (un) , minimum value lf (un) , minimum value lf (un) ;

[0172] Process B4: Denote the total number of cold data of the i-th user as cc (i) , the average value of (total) cold data size is denoted as af (i) , the maximum value is denoted as mf (i) , the minimum value is denoted as lf (i) ; The value range of i is: 1 to un;

[0173] Calculate the standard size of data segmentation and denote it as Δfi:

[0174]

[0175] Process B5: According to Δfi, perform decentralized encryption and data segmentation on all cold data corresponding to the 1st to un-th users, and construct a Merkle tree; According to the node structure of the Merkle tree, store the cold data in the blockchain in the cloud, and feedback the access interface of the blockchain to the 1st to un-th users.

[0176] Permission management module: Used to obtain the working hours of each user for various software, calculate the usage frequency of each software corresponding to each user; Based on the usage frequency of each user for each software, use the clustering algorithm to calculate the expected probability of each user using various software; Construct a judgment equation according to the expected probability, and judge whether to grant access permissions to warm data and cold data for various users;

[0177] Process C: The working process of the permission management module is as follows:

[0178] Process C1: Obtain the working hours of (all) users and denote it as wt; Obtain the current month and denote it as m; Obtain the number of working days of (all) users in the (m - 1)th month and denote it as wd;

[0179] Obtain the usage time of the 1st to un-th users using office software, social software, entertainment software and other software on working days in the (m - 1)th month, and calculate the usage frequency ha of the 1st to un-th users using office software (1) ~ha (un) ;

[0180] The usage frequency hb of using social software (1) ~hb (un) ;

[0181] Frequency hc of using entertainment software (1) ~hc (un) ;

[0182] Frequency hd of using other software (1) ~hd (un) ;

[0183] It should be noted that if m month is January of a certain year, then (m - 1) month is December of the previous year; for example: if m month is January 2023, then (m - 1) month is December 2022;

[0184] "Other software" refers to computer software that is not "office software, social software, and entertainment software", such as programming software;

[0185] Process C11: Record the time of using office software by the 1st to the unth users within the 1st to the wdth working days in (m - 1) month as ta(1, 1)~ta(un, wd);

[0186] Record the time of using social software as tb(1, 1)~tb(un, wd);

[0187] Record the time of using entertainment software as tc(1, 1)~tc(un, wd);

[0188] Record the time of using other software as td(1, 1)~td(un, wd);

[0189] Process C12: Let the time of using office software by the ith user within the jth working day in (m - 1) month be ta(i, j), the time of using social software be tb(i, j), the time of using entertainment software be tc(i, j), and the time of using other software be td(i, j); the value range of i is: 1~un, and the value range of j is: 1~wd;

[0190] Define calculation formula C1-1:

[0191] Define calculation formula C1-2:

[0192] Define calculation formula C1-3:

[0193] Define calculation formula C1-4:

[0194] Among them, ha (i) 、hb (i) 、hc (i) and hd (i), respectively representing the frequencies of the i-th user using office software, social software, entertainment software, and other software during work;

[0195] Process C13: Calculate the frequencies ha of the 1st to un-th users using office software according to calculation formulas C1-1 to C1-4 (1) ~ha (un) ;

[0196] The frequency hb of using social software (1) ~hb (un) ;

[0197] The frequency hc of using entertainment software (1) ~hc (un) ;

[0198] The frequency hd of using other software (1) ~hd (un) ;

[0199] Process C2: Use the clustering algorithm to cluster ha (1) ~ha (un) to obtain a stable parameter cluster;

[0200] Process C21: Perform a first clustering on ha (1) ~ha (un) ;

[0201] Randomly select kn parameters from ha (1) ~ha (un) as the central parameters; where, kn represents the number of central parameters, and the calculation formula for kn is: kn = (un) 1 / 2 , and kn is rounded up;

[0202] Record the 1st to kn-th central parameters as haa (1) ~haa (kn) ;

[0203] Process C22: Define calculation formula C2-1:

[0204] where, ha (i) represents the frequency of the i-th user using office software during work, and the value range of i is: 1~un;

[0205] haa (k) represents the kn-th central parameter, and the value range of k is: 1~kn;

[0206] d (i-k) represents the matching coefficient of ha (i) relative to haa (k) ;

[0207] Process C23: Calculate ha using calculation formula C2-1 (1) ~ha (un) The matching coefficient d (1-1) ~d (un-kn) , and cluster ha (1) ;

[0208] Process C231: Extract the matching coefficient d of ha (1) relative to haa (1) ~haa (kn) The matching coefficient d (1-1) ~d (1-kn) ;

[0209] Judge whether the minimum value in d (1-1) ~d (1-kn) is unique and enter different processes;

[0210] Process C232: If the minimum value in d (1-1) ~d (1-kn) is unique;

[0211] Extract the minimum value in d (1-1) ~d (1-kn) , denoted as dl; Divide the parameter ha (1) into the central parameter corresponding to dl;

[0212] Process C233: If the minimum value in d (1-1) ~d (1-kn) is not unique; Take the equal minimum values in d (1-1) ~d (1-kn) , denoted as dll (1) ~dll (s) ; Among them, s represents the number of equal minimum values in d (1-1) ~d (1-kn) , and the value range of s is: 2~kn;

[0213] Process C2331: Extract the central parameters corresponding to dll (1) ~dll (s) , as the undetermined parameters; Denote the undetermined parameters as cel (1) ~cel (s) ;

[0214] Define relation formula C2-2:

[0215] (|cel (i) |-|dll (i) |) = |cel (i) -dll (i) |;

[0216] Among them, cel(i) represents the i-th undetermined parameter, dll (i) represents dll (1) ~dll (s) the i-th parameter in dll, where the value range of i is: 1~s;

[0217] Process C2332: Substitute cel (1) ~cel (s) into the relational expression C2-2 to find the undetermined parameter that satisfies the relational expression C2-2, denoted as cey; Divide the central parameter corresponding to the parameter ha (1) into it;

[0218] Process C24: Repeatedly perform the same process of clustering the parameter ha (1) on the parameter ha (2) ~ha (un) to perform clustering, obtaining the 1st to the kn-th clusters, denoted as cluster c (1) ~cluster c (kn) (completing one clustering);

[0219] Among them, ha (2) represents the frequency of the 2nd user using office software (at work);

[0220] Process C25: Calculate the average parameters of the 1st to the kn-th clusters to obtain ac (1) ~ac (kn) ;

[0221] Define the relational expression C2-3:

[0222] (|haa (i) |-|ac (i) |)×|haa (i) -ac (i) |>0;

[0223] Among them, haa (i) represents the i-th central parameter, and ac (i) represents the average parameter of the i-th cluster, where the value range of i is: 1~kn;

[0224] Process C26: Substitute haa (1) ~haa (kn) and ac (1) ~ac (kn) into the relationship A3 to determine the new central parameters for secondary clustering;

[0225] If haa (1) and ac (1) satisfy the relational expression C2-3, then the new central parameter is ac (1) ; If haa (1) and ac(1) If the relational expression C2-3 is not satisfied, the new center parameter is haa (1) and ac (1) ;

[0226] And so on, if haa (kn) and ac (kn) satisfy the relational expression C2-3, the new center parameter is ac (kn) ; If haa (kn) and ac (kn) do not satisfy the relational expression C2-3, the new center parameter is haa (kn) and ac (kn) ;

[0227] Process C27: Using the new center parameter as the clustering benchmark, repeat the same clustering process once to perform secondary clustering on ha (1) ~ha (un) until the center parameters of each cluster no longer change, obtaining stable parameter clusters;

[0228] Process C3: Count the number of stable parameter clusters, denoted as cu;

[0229] Count the number of parameters in the 1st to the cu-th stable parameter clusters, denoted as an (1) ~an (cu) ;

[0230] Denote the parameters in the 1st to the cn-th stable parameter clusters as parameter ct(1, 1)~ct(cu, an (cu) );

[0231] Process C4: Extract the parameters ct(1, 1)~ct(1, an (1) ) in the 1st stable parameter cluster, and calculate the probability Pa(1) that the users corresponding to the 1st stable parameter cluster use office software;

[0232] Process C41: Create an empty matrix of (an (1) ×1), denoted as matrix ZO;

[0233] Fill ct(1, 1)~ct(1, an (1) ) into matrix ZO to obtain matrix ZE;

[0234] Process C42: Calculate the average value of ct(1, 1)~ct(1, an (1) ), denoted as lx;

[0235] Denote the parameter in the i-th row of matrix ZE as x(i), and define the calculation formula C2-4:

[0236] Y(i) = x(i) - lx; where, Y(i) represents the centralized value of x(i); the value range of i is from 1 to an (1) ;

[0237] According to calculation formula C2-4, calculate the centralized values corresponding to the parameters in matrix ZE to obtain the centralized matrix, denoted as matrix Y;

[0238] Process C43: Denote the parameters in the first stable parameter cluster as ct(1, j), where the value range of j is from 1 to an (1) ;

[0239] Denote the covariance matrix of matrix ZE as E;

[0240] E = [(Y) T ×Y] / (j - 1); where, T represents the transpose of the matrix, and the minimum value of (j - 1) is 1;

[0241] Construct the probability density function p(ct (1,j) ), to obtain calculation formula C2-5:

[0242] where, represents the square root of matrix E (regard matrix E as a determinant and calculate using a recursive algorithm (such as Laplace expansion));

[0243] e represents the natural constant (with a value of 2.7); (j / 2) is a positive integer, rounded up;

[0244] The exponent of f(ct(1,j)), the mathematical expression of f(ct(1,j)) is:

[0245]

[0246] where, I (an(1)) represents the all-1 matrix of (an (1) ×1), T represents the transpose of the matrix, and -1 represents the inverse of the matrix;

[0247] Process C44: Substitute the parameters ct(1, 1) to ct(1, an (1) ) into calculation formula C2-5 to obtain the probability densities of the parameters ct(1, 1) to ct(1, an(1)), denoted as pp(ct(1, 1)) to pp(ct(1, an (1) ));

[0248] Define relation formula C2-6:

[0249] Substitute pp(ct(1, 1)) to pp(ct(1, an (1)) Substitute into the relational expression 2-6, count the number of parameters that satisfy the relational expression 2-6, and denote it as au;

[0250] Calculate the probability Pa(1) that the user uses the office software corresponding to the first stable parameter cluster, Pa(1) = au / an (1) ;

[0251] Process C5: Repeat the same process of calculating Pa(1) to calculate the probabilities Pa(2) to Pa(cu) that the user uses the office software corresponding to the 2nd to cu-th stable parameter clusters;

[0252] Calculate the average value of Pa(1) to Pa(cu), and denote it as Pa; Take Pa as the expected probability that the user uses the office software;

[0253] Repeat the same process of calculating Pa to calculate the expected probability Pb that the user uses the social software, the expected probability Pc that the user uses the entertainment software, and the expected probability Pd that the user uses other software;

[0254] Process C6: Extract the frequency ha of the 1st user using the office software (1) , the frequency hb of using the social software (1) , the frequency hc of using the entertainment software (1) , the frequency hd of using other software (1) ;

[0255] Build the behavior model of the 1st user and determine whether to grant permissions for warm data and cold data;

[0256] Process C61: Calculate ha (1) , hb (1) , hc (1) and hd (1) ratios to obtain la, lb, lc, and ld; Calculate the sum of la, lb, lc, and ld, and denote it as ll;

[0257] Process C62: Build various software probability equations of the 1st user, and denote them as equation set C3;

[0258] Process C621: Let h’ represent ha (1) , hb (1) , hc (1) and hd (1) ; Let a represent the software currently used by the user; P(a - h’) represents the probability that the 1st user changes from software a to h , software;

[0259] If a represents the office software, build equation C3-1:

[0260]

[0261] If a represents a social software, construct Equation C3-2:

[0262]

[0263] If a represents an entertainment software, construct Equation C3-3:

[0264]

[0265] If a represents an office software, construct Equation C3-4:

[0266]

[0267] Combine Equations C3-1 to C3-4 into a system of equations C3;

[0268] Process C622: Obtain the software currently used by the first user, and combine it with the system of equations C3 to calculate the frequency of the first user's (subsequent) use of office software, denoted as ya, the frequency of using social software, denoted as yb, the frequency of using entertainment software, denoted as yc, and the frequency of using other software, denoted as yd;

[0269] Process C623: Arrange office software, social software, entertainment software, and other software in descending order of ya, yb, yc, and yd to obtain an optimized sequence;

[0270] Process C624: Detect whether the subsequent software usage of the first user matches the optimized sequence;

[0271] If it matches, grant the first user partial warm data access rights (i.e., 50% of the warm data access rights, and users or relevant technical personnel can adjust the access rights of warm data and cold data according to actual needs);

[0272] If it does not match, do not grant warm data access rights and conduct secondary monitoring on the first user;

[0273] If the software usage of the first user matches the optimized sequence for multiple consecutive times (i.e., more than 3 times), grant the first user full access rights to warm data and cold data;

[0274] Process C7: Repeatedly judge the same process for the warm data and cold data permissions of the first user; judge whether to grant the warm data and cold data permissions to the second to the un-th users.

[0275] Continuous monitoring module: Used to mark users who have not been granted warm data access rights as abnormal users and provide feedback.

[0276] Embodiment 2

[0277] Please refer to Figure 2, a data processing method based on big data includes:

[0278] Step S1: Obtain the user ID number of the data to be stored, obtain the data to be stored corresponding to each user ID number, and divide the data to be stored into hot data, warm data, and cold data;

[0279] Step S2: Encrypt the warm data of all users and store it in the local server; perform decentralized encryption on the cold data of all users and store it in the cloud; impose access permissions on the warm data and cold data, and open the access permissions for the hot data;

[0280] Step S3: Obtain the working hours of each user for various software, and calculate the usage frequency of each software corresponding to each user; based on the usage frequency of each user for each software, use the clustering algorithm to calculate the expected probability of each user using various software; construct a judgment equation according to the expected probability, and judge whether to grant access permissions to the warm data and cold data of each type of user;

[0281] Continuous monitoring module: used to mark the users who have not been granted access permissions to the warm data as abnormal users and give feedback;

[0282] Step S4: Mark the users who have not been granted access permissions to the warm data as abnormal users and give feedback.

[0283] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to get a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation. For example, if there are weight coefficients and proportionality coefficients, the sizes of their settings are for quantifying each parameter to obtain a specific numerical value for subsequent comparison. Regarding the sizes of the weight coefficients and proportionality coefficients, as long as they do not affect the proportional relationship between the parameters and the quantified numerical values, it is fine.

[0284] Finally, it should be noted that: the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A data processing platform and method based on big data, characterized in that: The platform includes: Data acquisition module: used to obtain the user ID number of the data to be stored, obtain the data to be stored corresponding to each user ID number, and divide the data to be stored into hot data, warm data and cold data; Data storage module: used to encrypt all users' warm data and store it in the local server; decentralize and encrypt all users' cold data and store it in the cloud; impose access rights on warm and cold data, and open access rights to hot data; Permission management module: used to obtain the working time of each user for each type of software and calculate the usage frequency of each software corresponding to each user; based on the user's usage frequency of each software, use the clustering algorithm to calculate the expected probability of the user using each type of software; construct a judgment equation based on the expected probability, and determine whether to grant each type of user access rights to warm data and cold data; Continuous monitoring module: used to mark users who are not granted warm data access rights as abnormal users and provide feedback.

2. A data processing platform based on big data according to claim 1, characterized in that: The specific process of the data acquisition module is as follows: Process A1: Count the number of user ID numbers un; Count the number of data to be classified fl for the 1st to unth users (1) ~fl (un) ; Process A2: (1) As fn, the data to be stored corresponding to the first user is taken as the target data; the target data is divided into hot data, warm data and cold data; Process A3: Repeat the same steps of allocating the data to be classified for the first user, and divide the data to be classified corresponding to the second to unth user ID numbers into hot data, warm data and cold data, and enter the data storage module.

3. A data processing platform based on big data according to claim 2, characterized in that: The specific process of the data storage module is as follows: Process B1: Encrypt all warm data and store them in the local server; Process B2: Count the total number of cold data cc corresponding to the 1st to unth users (1) ~cc (un) ; Calculate cc (1) ~cc (un) The sum of is recorded as acc; Set the first user from 1 to cc (1) The size of cold data is recorded as: fb(1,1)~fb(1,cc (1) ); and so on, the unth user 1 to cc (un) The size of cold data is recorded as: fb(un,1)~fb(un,cc (un) ); Process B3: Calculate and extract fb(1,1)~fb(1,cc (1) ) the average value af (1) , maximum value mf (1) , minimum value lf (1) ; and so on, fb(un,1)~fb(un,cc (un) ) the average value af (un) , maximum value mf (un) , minimum value lf (un) ; Process B4: The total amount of cold data for the i-th user is recorded as cc (i) , the average value of cold data size is recorded as af (i) , the maximum value is recorded as mf (i) , the minimum value is denoted as lf (i) ; The value range of i is: 1~un; The standard size of data segmentation is calculated and recorded as Δfi: Process B5: According to Δfi, all cold data corresponding to the 1st to unth users are decentralized encrypted and split, and a Merkle tree is constructed; according to the node structure of the Merkle tree, the cold data is stored in the blockchain in the cloud, and the access interface of the blockchain is fed back to the 1st to unth users to enter the permission management module.

4. A data processing platform based on big data according to claim 3, characterized in that: The workflow of the rights management module is as follows: Process C1: Get the user's working time wt, the current month m, and the number of working days wd of the user in month (m-1); Calculate the frequency of office software used by the 1st to unth users ha (1) ~ha (un) ; Frequency of using social software hb (1) ~hb (un) ; Frequency of using entertainment software hc (1) ~hc (un) ; Frequency of using other software hd (1) ~hd (un) ; Process C2: ha (1) ~ha (un) Perform clustering processing to obtain stable parameter clusters; Process C3: Count the number of stable parameter clusters, denoted as cu; count the parameters an in the 1st to cuth stable parameter clusters (1) ~an (cu) ; The parameters in the 1st to cnth stable parameter clusters are recorded as parameters ct(1,1)~ct(cu,an (cu) ); Process C4: Extract the parameters ct(1,1)~ct(1,an in the first stable parameter cluster (1) ), calculate the probability Pa(1).

5. A data processing platform based on big data according to claim 4, characterized in that: The subsequent process of process C4 is as follows: Process C5: Repeat the same process of calculating Pa(1) to calculate the probabilities Pa(2)~Pa(cu) of the office software corresponding to the 2nd to cuth stable parameter clusters; Calculate the average value Pa of Pa(1)~Pa(cu) as the expected probability of the user using office software; Repeat the same process of calculating Pa to calculate the expected probability Pb ​​of the user using social software, the expected probability Pc of using entertainment software, and the expected probability Pd of using other software; Process C6: Extract the frequency of the first user using office software ha (1) , frequency of using social software hb (1) , frequency of using entertainment software hc (1) , use other software frequency hd (1) ; Build the behavior model of the first user to determine whether to grant permissions for warm and cold data; Process C7: Repeat the same process of determining the warm data and cold data permissions of the first user; determine whether to grant the warm data and cold data permissions to the second to unth users.

6. A data processing platform based on big data according to claim 4, characterized in that: The specific process of process C2 is as follows: Process C21: ha (1) ~ha (un) Perform a clustering; randomly select kn parameters as the central parameters: haa (1) ~haa (kn) ; Process C22: Define calculation formula C2-1: Among them, ha (i) represents the frequency of the i-th user using office software, haa (k) represents the knth center parameter, d (i-k) Indicates ha (i) Relative haa (k) The matching coefficient of Process C23: Calculate ha using formula C2-1 (1) ~ha (un) The matching coefficient d (1-1) ~d (un-kn) , yes (1) Perform clustering; Process C231: Extracting matching coefficient d (1-1) ~d (1-kn) ; Judge d (1-1) ~d (1-kn) Whether the minimum value in is unique, enter different processes; Process C232: If d (1-1) ~d (1-kn) The minimum value in is unique; extract d (1-1) ~d (1-kn) The minimum value dl in the (1) Divide the center parameters corresponding to dl; Process C233: If d (1-1) ~d (1-kn) The minimum value in is not unique; (1-1) ~d (1-kn) The minimum value of the tie is denoted as dll (1) ~dll (s) ; Process C2331: Extract dll (1) ~dll (s) The corresponding central parameter is taken as the undetermined parameter; the undetermined parameter is denoted as cel (1) ~cel (s) ; Define the relationship C2-2: (|the (i) |-|dll (i) |)=|the (i) -dll (i) |; Among them, cel (i) Indicates the i-th pending parameter, dll (i) Indicates dll (1) ~dll (s) The i-th parameter in; Process C2332: cel (1) ~cel (s) Substitute into the relation C2-2, and the undetermined parameter that satisfies the relation C2-2 is denoted as cey; replace the parameter ha (1) Divide cey into the center parameters corresponding to it.

7. A data processing platform based on big data according to claim 6, characterized in that: The subsequent process of process C23 is as follows: Process C24: Repeat for parameter ha (1) The same process of clustering is carried out, and the parameter ha (2) ~ha (un) Perform clustering and get cluster c (1) ~Cluster c (kn) ; Process C25: Calculate the average parameters of the 1st to knth clusters and obtain ac (1) ~ac (kn) ; Define the relationship C2-3: (|haa (i) |-|ac (i) |)×|haa (i) -ac (i) |>0; Among them, haa (i) represents the i-th center parameter, ac (i) represents the average parameter of the i-th cluster; Process C26: Haa (1) ~haa (kn) and ac (1) ~ac (kn) Substitute into relation A3 to determine the new center parameter of the secondary clustering; If haa (1) and ac (1) If the relationship C2-3 is satisfied, the new center parameter is ac (1) ; If the relationship C2-3 is not satisfied, the new center parameter is haa (1) and ac (1) ; By analogy, if haa (kn) and ac (kn) If the relationship C2-3 is satisfied, the new center parameter is ac (kn) ; If the relationship C2-3 is not satisfied, the new center parameter is haa (kn) and ac (kn) ; Process C27: Using the new center parameter as the clustering basis, repeat the same clustering process once, and (1) ~ha (un) Perform secondary clustering until the central parameter of each cluster no longer changes, and obtain a stable parameter cluster.

8. The data processing platform based on big data according to claim 4, characterized in that: The specific process of process C4 is as follows: Process C41: Create an empty matrix ZO and fill it with ct(1,1)~ct(1,an (1) ), and get the matrix ZE; Process C42: Calculate ct(1,1)~ct(1,an (1) ) is the average value lx; let the parameter of the i-th row in the matrix ZE be x(i), and define the calculation formula C2-4: Y(i)=x(i)-lx;where Y(i) represents the central value of x(i); According to formula C2-4, calculate the center Y of matrix ZE; Process C43: The parameters in the first stable parameter cluster are recorded as ct(1, j); the covariance matrix of the matrix ZE is recorded as E; E = [(Y) T ×Y] / (j-1); construct the probability density function p(ct (1,j) ), and obtain formula C2-5: in, represents the square root of the matrix E; The exponential order of f(ct(1,j)): Among them, I (an(1)) indicates (1) ×1) represents a matrix of all 1s; Process C44: Calculate the probability density of ct(1,1)~ct(1,an(1)): pp(ct(1,1))~pp(ct(1,an (1) )); Define the relationship C2-6: Count the number of parameters that satisfy equations 2-6, denoted as au; Calculate the probability Pa(1) that the user corresponding to the first stable parameter cluster uses office software, Pa(1) = au / an (1) .

9. The data processing platform based on big data according to claim 5, characterized in that: The specific process of process C6 is as follows: Process C61: Calculate ha (1) ,hb (1) 、hc (1) and hd (1) The ratio of , we get la, lb, lc and ld; sum them up to get ll; Process C62: Construct equation group C3; let h , Indicates ha (1) ,hb (1) 、hc (1) and hd (1) ; Let a represent the software currently used by the user; P(ah , ) indicates that the first user changes from software a to h , The probability of software; if a represents office software, construct equation C3-1: If a represents social software, construct equation C3-2: If a represents entertainment software, construct equation C3-3: If a represents office software, construct equation C3-4: Equations C3-1 to C3-4, equation group C3; Get the software currently used by the first user, and combine it with equation group C3 to calculate the frequency of the first user using office software, recorded as ya, the frequency of using social software, recorded as yb, the frequency of using entertainment software, recorded as yc, and the frequency of using other software, recorded as yd; Arrange office software, social software, entertainment software and other software in descending order of ya, yb, yc and yd to obtain a preferred sequence; detect whether the subsequent software usage of the first user matches the preferred sequence; If they match, partial warm data access rights are opened to the first user; If they do not match, the warm data access permission will not be granted.

10. A data processing method based on big data, applicable to a data processing platform based on big data as claimed in any one of claims 1 to 9, characterized in that: The method comprises: Step S1: Obtain user ID numbers of data to be stored, obtain the data to be stored corresponding to each user ID number, and divide the data to be stored into hot data, warm data, and cold data; Step S2: Encrypt all users’ warm data and store them in the local server; decentralize and encrypt all users’ cold data and store them in the cloud; impose access rights on warm data and cold data, and open access rights to hot data; Step S3: Obtain the working time of each user for each type of software, and calculate the usage frequency of each software corresponding to each user; based on the usage frequency of each software by the user, use the clustering algorithm to calculate the expected probability of the user using each type of software; construct a judgment equation based on the expected probability, and judge whether to grant each type of user access rights to warm data and cold data; Continuous monitoring module: used to mark users who are not granted warm data access rights as abnormal users and provide feedback; Step S4: Mark users who are not granted warm data access rights as abnormal users and provide feedback.

Citation Information

Patent Citations

  • Access authority setting method and device, server, and storage medium

    CN106911697A

  • GIS data management and processing method based on cloud platform

    CN110765192A

  • Database authority management system and method

    CN117874826A

  • Enterprise sensitive data security access management method and system

    CN118656870A

  • Multi-element big data storage method

    CN118796137A