Data integrated management method based on data medium station
Through the integrated data management method based on the data middle platform, user permissions and backup plans are dynamically adjusted, data leakage risks caused by lagging permission adjustments in the existing technology are solved, real-time monitoring and refined management of user behavior are realized, and data security and backup efficiency are improved.
Patent Information
- Application Number
- CN202510398025.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, data integrated management methods based on data middle platform are difficult to dynamically adjust permissions based on user real-time access behavior, resulting in difficult to detect and process abnormal user behaviors as soon as possible, and there is a risk of data leakage.
By authenticating users, dynamically adjusting permissions in combination with access control lists and user role attributes, monitoring data access behavior in real time, generating behavior evaluation results, and adjusting permissions according to behavior pattern deviation values, setting backup time intervals and backup plans, monitoring the backup process in real time, ensuring data security.
It realizes precise control of user permissions, reduces the risk of data leakage, improves the flexibility and timeliness of backup plans, and ensures the stability and security of data.
Smart Images

Figure CN120337250A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data protection, and particularly to a data integration management method based on a data middle platform. Background Art
[0002] The data integration management method based on a data middle platform is a data protection technology for realizing secure data access, real-time monitoring of user access behaviors, dynamically adjusting permissions, and optimizing backup strategies, and can achieve efficient management of sensitive data and prevent potential security risks.
[0003] In the prior art, only relying on a preset access control list to perform data permission management, the adjustment of user permissions depends on a fixed periodic review, and it is difficult to dynamically adjust permissions according to the real-time access behaviors of users, making it difficult to discover and handle abnormal user behaviors in a timely manner. For example, when high-risk data calls occur frequently, the permissions remain in the original state, easily resulting in the risk of data leakage. Therefore, improvements are needed. Summary of the Invention
[0004] The purpose of the present invention is to solve the deficiencies existing in the prior art and propose a data integration management method based on a data middle platform.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions. The data integration management method based on a data middle platform includes the following steps:
[0006] Authenticate the user, and determine the data access permissions according to the user's role attributes in cooperation with the preset access control list in the database to obtain the user access permission configuration; monitor the user's real-time data access behaviors, evaluate potential security risks by comparing the deviation between the behavior pattern and the normal pattern, and generate a user behavior evaluation result;
[0007] Based on the user access permission configuration, match the permission verification rules for each data access request, perform permission verification, and record the verification results to obtain the permission verification record; use the user behavior evaluation result to adjust the user's real-time access permissions by comparing with abnormal behaviors, restrict or expand the data access scope, and obtain the adjusted user permissions;
[0008] Set the backup time interval according to the data access frequency in the permission verification record to obtain the data backup plan; execute data backup according to the data backup plan, and real-time monitor the backup completion degree and errors during the backup process to generate a backup status result;
[0009] Based on the backup status result, evaluate the effectiveness of the backup operation, and re-back up the data items that have not been successfully backed up to obtain a backup effect evaluation result.
[0010] Preferably, the step of obtaining the user access permission configuration is as follows:
[0011] Compare item by item the role attributes in the user identity information with the preset access control list in the database, screen item by item based on the permission limitation rules of the role attributes, and confirm the permission level in combination with the identity authentication status corresponding to the user identity information to generate the user role permission level;
[0012] Based on the user role permission level, match the data access permission entries, screen and combine the contents of the permission entries, and perform validity checks on the permission entries to eliminate redundant or invalid permission entries to generate the user access permission configuration.
[0013] Preferably, the step of obtaining the user behavior evaluation result is as follows:
[0014] Invoke the user's real-time data access behavior, extract and record item by item the time interval of the user access behavior, the access request response delay, and the data call quantity based on the target data entry accessed by the user, the access duration, and the frequency of the data access request to generate the user real-time access behavior characteristic data;
[0015] Based on the user real-time access behavior characteristic data, calculate the behavior pattern deviation value, and the calculation formula is:
[0016]
[0017] where Y represents the behavior pattern deviation value, T a represents the time interval of the user's current access behavior, T n represents the average time interval of access behavior in the normal mode, F r represents the user's current data call quantity, F n represents the average value of the data call quantity in the normal mode, D a represents the user's current access request response delay, D n represents the average value of the access request response delay in the normal mode;
[0018] Based on the behavior pattern deviation value, determine the security risk level of the user's real-time data access behavior, and determine whether the behavior risk exceeds the security standard in combination with the security risk level to generate the user behavior evaluation result.
[0019] Preferably, the step of obtaining the permission verification record is as follows:
[0020] Based on the user access permission configuration, for each data access request submitted by the user, retrieve the data access control list item by item, perform permission verification rule matching for the access request data entries, and complete the validity check of the permission matching rules to form the permission verification rules;
[0021] Based on the permission verification rules, item-by-item comparison and verification of the permissions of data entries in each data access request is performed to verify whether the data access request complies with the permission verification rules, and the execution results of the permission verification of each data access request are recorded item by item to obtain a permission verification record.
[0022] Preferably, the step of obtaining the adjusted user permissions is as follows:
[0023] According to the security risk level in the user behavior evaluation result, the response delay duration of the access data corresponding to the abnormal behavior, the recurrence times of the abnormal requests, and the data call quantity of a single request are extracted and recorded item by item to generate abnormal behavior characteristic data;
[0024] Based on the abnormal behavior characteristic data, a behavior risk index is calculated, and the calculation formula is:
[0025]
[0026] Among them, R represents the behavior risk index, f a represents the recurrence request frequency of the user's current abnormal request, d a represents the total amount of data calls associated with the abnormal behavior, q a represents the data sensitivity level value of the current access request, u n represents the data access permission level corresponding to the user's current request, t a represents the continuous duration of the occurrence of abnormal request behaviors;
[0027] Based on the behavior risk index, a numerical comparison is made between the behavior risk index and the permission risk threshold, and the deviation direction between the behavior risk index and the permission risk threshold is judged to determine whether the user's real-time access permission needs to be extended or restricted, a permission adjustment determination result is generated, and according to the permission adjustment determination result, a permission update operation of the data access control list is performed to generate adjusted user permissions.
[0028] Preferably, the step of obtaining the data backup plan is as follows:
[0029] According to the cumulative access times and access time intervals of each data access entry in the permission verification record, the access times of each data access entry in the current consecutive five data access requests are counted, and the average access response delay of the access requests is calculated to generate the consecutive access times of the data access entry and the average response delay of a single access request;
[0030] Based on the consecutive access times of the data access entry and the average response delay of a single access request, a backup priority index is calculated, and the calculation formula is:
[0031]
[0032] Among them, G is the backup priority index, and C a is the cumulative number of accesses to the data access entry within the most recent hour, EF r is the average response latency of a single access request to the data access entry, Q a is the number of data items called by a single request to the data access entry, T r is the average time interval between two consecutive access requests to the data access entry;
[0033] Based on the backup priority index, confirm the backup execution time interval corresponding to each data access entry, generate the backup time interval, and obtain the data backup plan.
[0034] Preferably, the step of obtaining the backup status result is as follows:
[0035] According to the data entries in the data backup plan, collect the backup duration, data transfer rate, and the number of errors during the backup of each data entry item by item to form a record of the data backup execution process;
[0036] Based on the data backup execution process record, calculate the backup stability index, and the calculation formula is:
[0037]
[0038] Among them, K represents the backup stability index, and E d represents the cumulative number of errors that occurred during the backup, R b represents the average data transfer rate of this backup, L s represents the total number of data entries associated with a single backup, H b is the duration of a single data access entry backup;
[0039] Based on the backup stability index, verify the completion degree of the backup, identify the entries that were not successfully backed up and the reasons, and generate the backup status result.
[0040] Preferably, the step of obtaining the backup effect evaluation result is as follows:
[0041] According to the backup completion status of each data access entry recorded in the backup status result, screen and extract the data access entries that failed to be backed up item by item, and record the reasons for the backup failure, the data transfer rate at the time of failure, and the error type to generate a list of data access entries that failed to be backed up;
[0042] Based on the list of data access entries that failed to be backed up, perform the re-backup operation on the data access entries item by item, and monitor the data transfer rate and backup error situation during the re-backup process in real time to generate the backup effect evaluation result.
[0043] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0044] In the present invention, by screening the permission control list item by item according to the user role attributes, dynamically adjusting the user permission scope, and matching the permission verification rules item by item in real time according to the user access request for verification, the precise control of data access is strengthened, and the risk of permission error caused by traditional static permission configuration is avoided; further, according to the difference between the user's real-time data access behavior characteristics and the normal behavior pattern, the behavior pattern deviation value is dynamically calculated, and the user permission is adjusted in real time according to the behavior pattern deviation value to achieve refined management of access permissions. At the same time, the backup priority index is calculated through characteristic indicators such as data access frequency and access response delay, and the data backup time interval is adjusted in real time to improve the flexibility and timeliness of the backup plan; during the data backup process, the backup stability index is calculated in real time, the backup exception is quickly located and the re-backup operation is carried out in time to ensure the effectiveness of the data backup operation, reduce the risk of data loss and damage, and improve the overall stability and security of sensitive data protection. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a step schematic diagram of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0046] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0047] Please refer to Figure 1 , the present invention provides a technical solution, a data integration management method based on a data middle platform, including the following steps:
[0048] Authenticate the user, and according to the user's role attributes and the preset access control list in the database, determine the data access permission to obtain the user access permission configuration; monitor the user's real-time data access behavior, and evaluate the potential security risks by comparing the deviation between the behavior pattern and the normal pattern to generate the user behavior evaluation result;
[0049] Based on the user access permission configuration, match the permission verification rules for each data access request, perform permission verification, and record the verification result to obtain the permission verification record; use the user behavior evaluation result to adjust the user's real-time access permission by comparing with abnormal behaviors, restrict or expand the data access scope to obtain the adjusted user permission;
[0050] Set the backup time interval according to the data access frequency in the permission verification record to obtain a data backup plan; execute data backup according to the data backup plan, and monitor the backup completion degree and errors during the backup process in real time to generate a backup status result;
[0051] Based on the backup status result, evaluate the effectiveness of the backup operation, re-back up the data items that have not been successfully backed up to obtain a backup effect evaluation result.
[0052] The steps for obtaining the user access permission configuration are as follows:
[0053] Compare each item of the role attribute in the user identity information with the preset access control list in the database one by one, screen item by item based on the permission limitation rules of the role attribute, and confirm the permission level in combination with the identity authentication status corresponding to the user identity information to generate the user role permission level;
[0054] Based on the user role permission level, match the data access permission entries, screen and combine the contents of the permission entries, and check the effectiveness of the permission entries to eliminate redundant or invalid permission entries to generate the user access permission configuration.
[0055] Specifically, based on the obtained user identity information and its role attributes, first establish an access control list index in the local database, which contains all possible role names and their corresponding access permission ranges. Then, compare each role attribute carried by the user one by one with the access control list index, and record the judgment parameters required for each comparison result during the matching process. For example, when comparing the user's role name, a threshold for character matching degree can be referred to. For example, set the threshold to 0.8 and determine whether it exceeds the threshold by calculating the character similarity of the role name. Here, the threshold can be obtained by counting at least 200 known role names, calculating the average value of their character matching rates, and then adjusting it downward by 0.05 to get the final value of 0.8. Then, perform the same comparison on other attributes in the identity information. If a role attribute is missing or has an abnormal format during the comparison process, use a fixed padding value for placeholder to continue the comparison, but do not perform any additional screening or transformation. Just to keep the data structure consistent. Then, pair these role attributes with the role permission limitation rules in the access control list one by one and perform screening. If the matching rate between the permission range defined by a certain rule and the user's role attribute is lower than the set threshold, it is regarded as not meeting the conditions and excluded. The matching rate can be judged whether to exclude by dividing the number of matching attributes by the total number of attributes participating in the comparison and comparing it with 0.8. If it cannot be excluded, record this rule. Finally, further confirm the identity authentication status for all the rules retained after screening. Compare the user's current identity authentication status with the verification reference value inside the system. For example, set a reference value of 2, indicating that the user must pass both fingerprint recognition and password verification to be judged as authentication completed. This 2 is set by researching the two most commonly used methods of multi-factor authentication. If the record of the user's identity authentication status shows that all verification requirements meeting the reference value of 2 have been completed, confirm that the user's permission level is high, otherwise it is low. Finally, combine each rule and summarize the permission ranges of the rules uniformly to obtain the user's role permission level.
[0056] Based on the previously determined user role permission levels, first retrieve the set of data access permission entries preset in the database for this level, and then group and summarize these permission entries by functional category. For example, they are classified into two categories: "general information access" and "sensitive information access". Then, perform content screening on each permission entry. Determine whether to retain or eliminate it by checking whether the actual usage threshold of the corresponding entry matches the user role permission level by more than 0.75. Here, 0.75 can be obtained by calculating the average value of 0.70 of the successful matches between the role permission level and the access entry after collecting 100 historical records and then increasing it by 0.05. If the actual usage threshold only allows users with a permission level greater than or equal to 3 to access, then the specific level of the user should be used as a reference during the comparison process. If it does not meet the requirement, it will be eliminated; if it meets the requirement, it will be retained. After combining all the retained entries, confirm their status one by one according to the validity check process. For example, set a benchmark for the number of error calls as 3. This benchmark value is obtained by rounding up the average value of 2.4 statistically calculated from the preliminary test data of 10 users to 3. If the number of error calls of a certain entry exceeds 3 and does not recover within 24 hours, it is regarded as an invalid entry and directly eliminated. The remaining entries, after being sorted out, form a final permission configuration list that can be accessed, and finally output this list to obtain the user access permission configuration.
[0057] The steps to obtain the user behavior evaluation result are as follows:
[0058] Invoke the user's real-time data access behavior. According to the target data entry accessed by the user, the access duration, and the frequency of data access requests, extract and record the time interval of the user's access behavior, the access request response delay, and the data call quantity item by item, and generate the user real-time access behavior characteristic data.
[0059] Based on the user real-time access behavior characteristic data, calculate the behavior pattern deviation value. The calculation formula is:
[0060]
[0061] Among them, Y represents the behavior pattern deviation value, T a represents the time interval of the user's current access behavior, T n represents the average time interval of access behavior in the normal mode, F r represents the user's current data call quantity, F n represents the average value of the data call quantity in the normal mode, D a represents the user's current access request response delay, D n represents the average value of the access request response delay in the normal mode;
[0062] Based on the behavior pattern deviation value, determine the security risk level of the user's real-time data access behavior, and combine the security risk level to determine whether the behavior risk exceeds the security standard, and generate the user behavior evaluation result.
[0063] Specifically, based on the monitoring data of the target data items accessed by the user, the start and end time points of each request and the access content identifier corresponding to the request are first extracted from the user's recent access log, and the time difference between adjacent access requests is recorded as the time interval and abnormal timeouts are excluded. For example, when the time interval exceeds sixty minutes, it is necessary to confirm whether there is a disconnection or long-term idleness based on the actual situation. Then, according to the distribution of the access duration in the range of ten seconds to three thousand seconds, the record is archived and it is determined whether it needs to be classified as continuous access or scattered access. If it is found that the time interval is concentrated less than thirty seconds and the number of data requests is more than twenty times, it can be noted as high-frequency access in the record of this stage, which can be compared with the security policy later. In addition, it is necessary to collect the response delay of the access request. The specific method is to measure the time spent waiting for the system to return a preliminary response after sending the request. If it takes a long time, if it exceeds ten seconds, it is recorded as a high latency state, and if it is less than one second, it is recorded as a low latency state. The middle interval is one to ten seconds as the normal response range. The statistics of the number of data calls are accumulated and recorded through real-time monitoring of the number of target data entries involved in a single request. For example, when a request reads five data entries at the same time, it can be recorded as the number of calls for this request equals five. After each access, all time intervals, response delays, and data call quantity entries are summarized and compared with the reference range. For example, the time interval is compared with the range of one second to three thousand seconds, the response delay is compared with the range of zero seconds to twenty seconds, and the number of data calls is compared with the range of zero to one thousand. After comparison, the legal records are stored uniformly. Finally, the user's real-time access behavior characteristic data is obtained by comprehensively analyzing the time interval, response delay, and number of data calls and performing simple marking and grouping classification.
[0064] The benefit of the formula is that it comprehensively considers the access time difference, the difference in the number of calls, and the difference in response delay, and presents the degree of deviation between the user's current behavior and normal behavior by incorporating these three quantitative indicators at the same time.
[0065] T a The steps for obtaining the parameters are as follows: first, read the start and end time from the actual operation log of a single access request of the user, count the difference between the end time and the start time as the duration of one access, then continuously extract the duration values of the most recent 20 accesses and arrange them in order, select the duration value corresponding to the latest access and record it as T a The time consumption includes the entire period from when the user initiates a request to the system until the request is processed and the final response is returned. In order to collect accurate time consumption, it is necessary to generate time stamps at the moment when the user request is initiated and the moment when the system completes processing, and then subtract the two to obtain millisecond accuracy. For example, when the user initiates a request at 10:00:05.350 and completes the response record at 10:00:13.570, it takes 8.22 seconds, which is recorded as Ta =8.22 seconds.
[0066] T n The steps to obtain the parameters are to collect the time values of at least 500 past access requests under the same role or the same operation type, calculate the arithmetic mean of this batch of time data as the average time interval of access behavior in normal mode, first subtract the start time and end time of each access to obtain a series of time costs, and then accumulate these time costs and divide them by the total number of entries to get T n , if the total time consumed by the collected 500 accesses is 3,250 seconds, then
[0067] F r The parameter acquisition step is to count the total number of data items read or retrieved in a single user request, and then accumulate all the accompanying call operations of the user's visit, and use the merge summation method to obtain the sum of the number of calls for this visit as F r For example, if this access retrieves ninety different record entries and also calls thirty other reference data blocks, then F can be recorded. r =120.
[0068] F n The steps to obtain the parameters are as follows: select one thousand normal historical accesses of the same access type, count the total number of data calls one by one, and then calculate the average value, which is used as F n For example, if these accesses cumulatively call 102,000 data entries, then
[0069] D a The parameter acquisition step is to read the time interval from the user's current access request to the system's first feedback. This does not include the completion time of all subsequent data transmissions. Only the time taken for the initial response to arrive is counted. The same method is used to measure the millisecond level. If the time is 0.52 seconds, record D a =0.52.
[0070] D n The parameter acquisition step is to extract at least 300 response delay values from the historical records of the same access operation, and calculate their average value as the average response delay value in normal mode. Assuming that the sum of these 300 response delays is 150 seconds, then
[0071] Calculation process: The first step is to bring in the parameter values obtained above and set T a =8.22, T n =6.50, F r =120, F n =102, D a =0.52, Dn =0.50; Step 2, calculate the numerator part first |T a -T n |×|F r -F n |:
[0072] |8.22-6.50|×|120-102|=1.72×18=30.96;
[0073] Step 3: Calculate the denominator
[0074]
[0075] 1+0.1414=1.1414;
[0076] Step 4: Divide the numerator by the denominator and take the cube root:
[0077]
[0078]
[0079] The result shows that when the value of Y is about 3.00, it means that the overall differences between the current user's access time interval and the average value, the number of calls and the average value, and the response delay and the average value are obvious. Subsequent steps are needed to judge the deviation value and compare it with the security standard to determine the user behavior risk level.
[0080] Based on the behavior pattern deviation value records obtained above and combined with the security standard definition of user access operations, an interval is first set to identify different security risk levels. For example, if the deviation value is between one and two, it is marked as a lower risk, if it is between two and four, it is marked as a medium risk, and if it is greater than four, it is marked as a high risk. This interval definition needs to be summarized with historical deviation values and statistical data. For example, one thousand access deviation value records are used to determine the percentile distribution. In this distribution, the deviation values within the initial 30% range can be classified as lower risk, the middle 40% as medium risk, and the last 30% as high risk. Then the current deviation value is compared with the corresponding interval boundary. If it is found that the value is greater than four, it is judged as high risk, if the value is between two and four, it is judged as medium risk, and if it is less than two, it is judged as low risk. The risk level is marked and checked with the security standard. The security standard itself can be set through the company's internal control requirements or industry regulations. For example, a deviation value greater than four is considered to exceed the security standard. If it is confirmed that the deviation value exceeds the standard, it is classified as abnormal behavior. Finally, after all judgments are completed, the overall access situation is recorded and compared to generate a user behavior evaluation result.
[0081] The steps to obtain the permission verification record are:
[0082] Based on the user access permission configuration, for each data access request submitted by the user, retrieve the data access control list item by item, perform permission verification rule matching for the access request data entries, complete the validity check of the permission matching rules, and form permission verification rules.
[0083] Based on the permission verification rules, compare and verify the permissions of the data entries in each data access request item by item, verify whether the data access request conforms to the permission verification rules, record the execution results of the permission verification of each data access request item by item, and obtain permission verification records.
[0084] Specifically, based on the user access permission configuration obtained previously, first read the user identifier, request data entry number, and request initiation time included in each newly submitted data access request. Compare item by item with the permission scope corresponding to the user role or identity identifier in the data access control list to check whether there is a prohibited item record for this data entry. If it is detected that it is marked as not allowed to access in the control list, immediately record the verification result and end this comparison. Otherwise, continue to search for whether an access condition restriction item is matched. For example, there is a restriction that the number of accesses to the same entry or similar keywords should not exceed five times within one minute. When it is found that the user's access count has reached six times, it is regarded as an over-limit access, and the over-limit entry is recorded. If the access count is less than or equal to five times, it is regarded as meeting the conditions, and continue to evaluate the next possible restriction item. During the process, separate checks can be combined with specific advanced requirements. For example, set a request frequency benchmark value of ten times per minute. This value is determined by summing the means after statistically analyzing the regular operations of one hundred users. On the basis of this mean, it is increased by two times as a safety factor to obtain the benchmark value of ten times per minute. If the request frequency submitted by the user has exceeded this benchmark value, it is judged that the access is too concentrated. At this time, an exception identifier needs to be registered and enter the subsequent more refined verification process. If it has not been exceeded, continue to compare other restriction rules. A default query also needs to be performed on the entries that do not appear in the data access control list to confirm whether there are global-level restriction items. This global restriction can be manually set by the configuration manager in a list and matched here. When all possible restriction items have been compared, if no prohibited access or over-limit access conditions are met, it indicates that the permission verification rule for the current data entry can remain in the allowed state. Finally, the accessible or intercepted records corresponding to each request are statistically analyzed in segments and the correctness is checked. If it is found that there are duplicate markings or data reference confusions during the statistics, a re-comparison needs to be performed according to the request initiation time sequence. After excluding duplicates, summarize and form permission verification rules.
[0085] Based on the previously formed permission verification rules, perform matching comparisons on each data access request submitted by the user in sequence. First, identify the target data entry number and access type from the request content, such as read, write, or delete, etc. Then, retrieve the corresponding restrictive conditions in the rules, such as whether a read-type access requires only one occurrence within thirty seconds or a delete-type access requires confirmation that the user has full access level. If a restrictive condition is matched, check whether the user identity meets the condition. If not, consider this access as rejected and record the reason for rejection, and record the user identifier, timestamp, and the number of the rejected entry at the corresponding position. If it meets the condition, record the pass identifier and continue to review the next data access request. For all requests that meet the conditions, it is necessary to further confirm whether there are conflicts with other rules, such as determining whether the frequency of this request is mutually exclusive with the shared rules of other roles. If there is a mutual exclusion, record the mutual exclusion result in the special conflict list. For each request in the conflict list, re-compare the access type and permission level, and combine the user role attributes and the previously set benchmark value of three times per minute. This benchmark is obtained by analyzing the average request value of five hundred access logs, which is two times per minute, and then increasing it by one to get three times per minute. If it exceeds this standard, mark it as multiple repeated accesses in the additional record. After traversing all requests, a collection of verification results item by item will be formed, including allowed, rejected, or mutually exclusive status, and attach the corresponding matching basis after each result item. Finally, integrate this set to obtain the permission verification record.
[0086] The steps to obtain the adjusted user permissions are as follows:
[0087] According to the security risk level in the user behavior assessment results, extract and record item by item the response delay duration of the access data corresponding to the abnormal behavior, the repeated occurrence times of the abnormal requests, and the data call quantity of a single request, and generate abnormal behavior characteristic data;
[0088] Based on the abnormal behavior characteristic data, calculate the behavior risk index, and the calculation formula is:
[0089]
[0090] Among them, R represents the behavior risk index, f a represents the repeated request frequency of the user's current abnormal request, d a represents the total amount of data calls associated with the abnormal behavior, q a represents the data sensitivity level value of the current access request, u n represents the data access permission level corresponding to the user's current request, t a represents the continuous duration of the cumulative occurrence of abnormal request behaviors;
[0091] Based on the behavioral risk index, numerically compare the behavioral risk index with the permission risk threshold, and determine the deviation direction between the behavioral risk index and the permission risk threshold. Determine whether the user's real-time access permission needs to be expanded or restricted, generate a permission adjustment determination result. According to the permission adjustment determination result, perform a permission update operation on the data access control list to generate the adjusted user permission.
[0092] Specifically, according to the security risk level in the user behavior assessment result, first read each request record marked as abnormal in the previous step, and check the access data response delay duration, the repeated occurrence times of abnormal requests, and the number of data calls in a single request item by item. Separate the specific values of this information and record and confirm the sources. For example, the response delay duration is calculated by comparing the request initiation time with the system's preliminary response time and retaining millisecond-level precision. The repeated occurrence times of abnormal requests are obtained from the actual number of times the same request appears within one hour. The number of data calls in a single request is obtained by counting the target data entries involved in the same request and cumulatively summarizing them. Store all entries in different regions for unified extraction and comparison in the next step. If it is detected that the response delay duration exceeds two seconds, use the range from one second to ten seconds as the upper limit of the normal interval to judge the difference. If the difference is higher than eight seconds, mark it as extreme delay. The repeated occurrence times of abnormal requests can be compared with the threshold of more than five times within five minutes. This threshold is obtained by statistically analyzing the historical records of two hundred users, taking the average repeated value four times, and then increasing it once to get five times. If the current record's repeated occurrence times have reached six times, it is recorded as exceeding the threshold. Compare the number of data calls in a single request with the previous data. If it exceeds two hundred and fifty target data entries, it is marked as large-scale access. The limit standard of two hundred and fifty entries mainly comes from the maximum value of two hundred entries for the combined query of multiple related data in ordinary requests in the daily environment and an additional reserve of fifty entries. After completing the above comparison, integrate and save the delay duration, repeated occurrence times, and number of data calls of each abnormal access as abnormal behavior characteristic data.
[0093] The benefit of the formula is that it simultaneously incorporates key factors such as the repetition frequency of user abnormal requests, data call scale, data sensitivity level, permission level, and the cumulative duration of abnormal behavior. It reflects the overall risk level of abnormal behavior through the comprehensive calculation of multi-dimensional parameters.
[0094] f aThe steps to obtain the parameter are as follows: First, record the unique number of each request and its specific triggering time in the monitoring system. Then, select the requests marked as abnormal from these records, merge them according to the same request content and similar resource pointers, and count the number of occurrences of the same abnormal request within a fixed time duration. Consider this count as the repetition request frequency of the user's current abnormal request. To accurately quantify this frequency, it is necessary to extract the occurrence information of the same request from the historical records of at least thirty days and divide the time interval. Count the total number of abnormal requests per hour, divide the total number per hour by the total number of all requests within that hour to obtain the frequency score, and finally divide the sum of all frequency scores by the number of hour intervals to obtain the concentration value, thereby obtaining the value of f a The numerical value of, for example, after counting the logs of the user for two hundred and forty hours, it is found that there are a total of sixty abnormal requests, and the total number of all requests is one thousand two hundred. Set the average total number of requests per hour to five, and obtain the average frequency score as Multiply by a hundred times to expand and obtain the final f a = 25.
[0095] d a The steps to obtain the parameter are as follows: Extract all data call entries marked with abnormal behaviors and accumulate them one by one to form a total. Here, the data call entries include the main access target and several associated queries. It is necessary to confirm each entry involved in the associated query and include it in the statistics. To ensure the accuracy of the statistical results, it is necessary to count the data entry IDs for each abnormal access call at the log level, and finally integrate these data entry ID numbers together to obtain the corresponding total. If it is confirmed during the screening process that a certain access calls 327 associated records, then directly include 327 in the summary value of d a After completing the merging of all abnormal accesses, obtain d a , for example, during a monitoring period, it is found that this abnormal behavior has called a total of 1200 data records, then d a = 1200.
[0096] q a The steps to obtain the parameter are as follows: Quantify the numerical value of the data sensitivity level involved in the current access request. A pre-developed classification standard table is required, in which the data is divided into several sensitivity intervals, and the data classification codes associated with the access request are used for matching. After reading the sensitivity level values of each data item in the matched items, perform a normalization operation and summarize. Finally, use the value obtained by superimposing the normalization results as the data sensitivity level value q of this request a, The specific approach is to read the confidentiality classification marks of each data item from the database. The confidentiality classification marks range from 1 to 10 numerically. Then, if there are multiple data items in this request, the average or weighted value of these classification marks is used as the overall sensitivity level value of this request. To ensure accuracy, each classification mark needs to be numerically combined and then divided by the number of items. If a request accesses three data items with classification marks of 7, 8, and 9 respectively, then we can get
[0097] u n The steps to obtain the u parameter are as follows: Check the permission level corresponding to the user's current request. The permission level is generally given by the system after comprehensive evaluation based on the user's role, authentication method, and past access records. This comprehensive evaluation is divided into five levels from 1 to 5, and the larger the number, the higher the permission. It is necessary to read from the user's current authorization status which permission level this request is at, and then directly use this level as u n , For example, if the permission level shown by the user role assessment result is 3, then record u n = 3.
[0098] t a The steps to obtain the t parameter are as follows: Calculate the cumulative duration from the start of the abnormal request behavior to the current moment. In the system's activity timeline, locate the moment when the first abnormal request occurred and the current moment, and calculate the difference to obtain the duration. If the first abnormal request occurred at 08:00 and the current recording moment is 10:30, then the duration is two and a half hours, which is 150 minutes. We can record t a = 150.
[0099] Calculation process: First step, substitute the parameter values given above: f a = 25, d a = 1200, q a = 8, u n = 3, t a = 150; Second step, first calculate Then multiply it by f a to get 25 × 1,728,000,000 = 43,200,000,000; Third step, calculate Add the two to get 64 + 9 = 73; Fourth step, multiply the results of the first two steps: 43,200,000,000 × 73 = 3,153,600,000,000; Fifth step, take the fourth root of this result: First find the square root and then take the square root again. The process is as follows:
[0100]
[0101] Therefore, the fourth root is approximately 1331.0; Step 6, calculate the denominator ln(t a +1)=ln(150+1)=ln(151)≈5.0173; Step 7: Divide the numerator by the denominator:
[0102]
[0103] The result shows that when R≈265.33, it means that the current abnormal request has a high behavioral risk index due to the combined effects of repetition frequency, data call scale, data sensitivity level, permission level and duration. It is necessary to further compare it with the permission risk threshold defined previously. Only in the subsequent steps can we confirm whether to expand or restrict user permissions based on the comparison results.
[0104] Based on the behavior risk index, all R values are first retrieved from the risk records of the current abnormal behavior and compared with the corresponding values of the permission risk threshold. The threshold can be set in combination with industry standards and local historical data. For example, after counting one thousand abnormal records, the average R value is calculated to be one hundred and eighty and increased by twenty to get two hundred as the threshold. Then, the direction and degree of deviation between the R value and the threshold are judged according to the comparison process. If the R value exceeds two hundred, it is marked as exceeding the threshold and the permission operation is restricted. If the R value does not exceed the threshold, it continues to be marked as normal range permission. Then, the judgment result is matched with the user's actual access permission record one by one. If it is found If there are records that require permission restrictions, the specific access data entries will be adjusted in a contraction manner, reducing the permission level from four to three or from three to two, and the new permission status will be recorded again after matching. If the R value is found to be lower than the threshold during the deviation comparison process, the permission will be expanded but not exceeding the highest level allowed by the user role. It will be extended in sequence from low to high according to the five levels set previously. When all records have completed the R value comparison and the corresponding update of the permission level, check the users whose permissions have changed in the final list and form an update summary table. Finally, the permission annotations after the update are organized into one record and output as the adjusted user permissions.
[0105] The steps to obtain a data backup plan are:
[0106] According to the cumulative number of accesses and the access time interval of each data access entry in the permission verification record, the number of accesses of each data access entry in the current five consecutive data access requests is counted, and the average access response delay of the access requests is calculated to generate the number of consecutive accesses of the data access entry and the average response delay of a single access request;
[0107] The backup priority index is calculated based on the number of consecutive accesses to the data access entry and the average response delay of a single access request. The calculation formula is:
[0108]
[0109] Among them, G is the backup priority index, C a is the cumulative number of accesses to the data access entry within the most recent hour, EF r is the average response latency of a single access request for this data access entry, Q a is the amount of data called by a single request for this data access entry, T r is the average time interval between two consecutive access requests for this data access entry;
[0110] Based on the backup priority index, confirm the backup execution time interval corresponding to each data access entry, generate the backup time interval, and obtain the data backup plan.
[0111] Specifically, according to the cumulative number of accesses and the access time interval of each data access entry in the permission verification record, first read the access behaviors of all data access entries within the most recent operation cycle from the log, group them according to the entry number, and extract the access times of each group within the specified time period. Then, select the timestamps corresponding to the current five consecutive data access requests, identify the continuity by calculating the differences between adjacent request time points, and exclude the entries that have not been accessed again for more than 24 hours. If the five consecutive accesses all occur within one hour and there is no obvious time gap, include them in the statistical scope. Then, count the number of accesses within these five access requests and record it. For requests with an access interval lower than 10 seconds, high-frequency marking can also be performed. The judgment criterion of 10 seconds comes from the reference median of about 7 seconds obtained for 100 average access intervals and is increased by 3 seconds to cover unexpected jitters. If this mark appears more than three times, it indicates a relatively high access demand intensity. By recording these marks, high-load entries can be distinguished for subsequent troubleshooting. Subsequently, measure the system response time of each access request, regard the difference between the request sending time and the system initial return time as the response latency and summarize the records. If it is found that the latency exceeds 2 seconds, classify it into the long response queue. The 2-second threshold is obtained by averaging about 1.3 seconds of the 500 request response latency measurements and then increasing it by 0.7 seconds. For those values between 0.2 seconds and 2 seconds, they are regarded as normal responses. Calculate the average of these latency values to obtain the average response latency of a single access request, and then combine it with the previously obtained consecutive access count to form a set of key information. Finally, integrate and map this set of information to the number of each data access entry to generate the consecutive access count and the average response latency of a single access request for the data access entry.
[0112] The benefit of the formula is to comprehensively quantify the intensity of recent accesses (C a ), access response performance (EF r ), the scale of a single call (Q a ), and the access gap of the entry (T r ), and measure the urgency of the entry for backup through a three-dimensional operation.
[0113] C a The steps to obtain the parameter are as follows: First, count the access situation of the specified entry within the last hour, retrieve one by one from the access log whether it belongs to this entry. If retrieved, record the number of access times once and accumulate to get the final total. To ensure data accuracy, all request records within this hour need to be scanned, and the timestamps and access entry numbers of each record are matched. Then, add up the matching results to get C a , if the access entry numbers match 200 times within sixty minutes, then C can be recorded a = 200
[0114] EF r The steps to obtain the parameter are as follows: Summarize all single-request response delays of this data access entry. It is required to record the difference between the time when each request is sent and the time when the system returns the initial information, and then divide the sum of all delays by the number of requests within the statistical period to obtain the average access response delay. If the statistical period is selected as one day and a total of 1,000 requests for this entry are collected, and the total delay of these requests is 1,100 seconds, then
[0115] Q a The steps to obtain the parameter are as follows: Calculate the number of data calls in each access request for the entry. It is necessary to scan how many independent data records or data blocks are involved in this access during the request processing stage and accumulate them. For example, if one access contains 50 related records, the number of calls for this time can be registered as 50. When multiple access records are aggregated, the final number of calls per single request can be obtained by weighted or simple average methods. If a total of 1,000 data are called in 10 accesses are found within the same statistical period, then
[0116] T r The steps to obtain the parameter are as follows: Accumulate the time intervals between two adjacent accesses for this entry and then divide by the number of adjacent access pairs to get the average value. For example, within one hour, this entry has 6 accesses, which occur at 08:00, 08:10, 08:22, 08:35, 08:40, and 08:50. The difference between 08:10 and 08:00 can be calculated as 10 minutes, the difference between 08:22 and 08:10 is 12 minutes, the difference between 08:35 and 08:22 is 13 minutes, the difference between 08:40 and 08:35 is 5 minutes, and the difference between 08:50 and 08:40 is 10 minutes. Then, divide (10 + 12 + 13 + 5 + 10 = 50 minutes) by 5 time difference pairs to get 10 minutes, and convert it to the value 10 and record it as T r = 10
[0117] Calculation process: First step, substitute the numerical values of each parameter: C a= 200, EF r = 1.1, Q a = 100, T r = 10; In the second step, calculate the numerator: where 200 2
[0118] = 40000, then multiply it by 1.1 × 100 = 110 to get 40000 × 110 = 4400000, and then find the cube root: In the third step, calculate the denominator:
[0119] In the fourth step, perform the division operation:
[0120] This result indicates that when G ≈ 39.17, the backup priority index composed of the cumulative access intensity, average response latency, call data volume, and access time interval of this entry is relatively high. This value can be compared with the corresponding indices of other entries, and the subsequent steps will determine the backup order and time interval based on this.
[0121] Based on the backup priority index calculated previously, first sort all data access entries numerically from largest to smallest and mark the entry numbers and their corresponding priority values in the index. If the corresponding index of an entry exceeds forty, it can be considered within the scope of urgent attention. The threshold of forty is obtained by taking the average value of thirty-five of the backup priority index data sequences of one hundred high-traffic entries and increasing it by five. Then, perform a one-hour-level backup plan for the entries within the scope of urgent attention. For entries with an index between twenty and forty, set a backup every two hours as the main execution cycle. The two-hour interval is specified by comprehensively comparing the tolerable access pressure of these entries in the previous index. For entries with an index lower than twenty, temporarily set a backup once a day to allocate resources to higher-priority entries. Then, form a one-to-one corresponding list of these time intervals, entry numbers, and their priority values and record it. To determine the backup execution time point, it is necessary to coordinate the scheduling information of other jobs in the same storage service. For example, check whether there are a large number of data import operations occurring at the whole hour, and then fine-tune the backup start time according to this scheduling information. If a resource conflict occurs, push it back by several minutes and re-detect. For example, if the load of a server does not exceed 70% per hour, it is considered executable for backup. If the load has reached 80%, wait for another ten minutes and then evaluate again. After arranging the backup time intervals for all entries, summarize them uniformly to generate the final backup time interval information and obtain the data backup plan.
[0122] The steps to obtain the backup status result are as follows:
[0123] According to the data entries in the data backup plan, collect the backup duration, data transfer rate, and the number of errors during the backup process for each data entry item by item, and form a record of the data backup execution process;
[0124] Based on the record of the data backup execution process, calculate the backup stability index. The calculation formula is:
[0125]
[0126] Among them, K represents the backup stability index, E d represents the cumulative number of errors that occurred during the backup process, R b represents the average data transfer rate of this backup, L s represents the total number of data entries associated with a single backup, H b is the duration of the backup for a single data access entry;
[0127] Based on the backup stability index, verify the completion degree of the backup, identify the entries that were not successfully backed up and the reasons, and generate the backup status result.
[0128] Specifically, according to the data entries in the data backup plan, first record the data numbers that need to be backed up one by one, and extract the time consumed, the transfer rate, and the number of errors that occurred during the actual backup process for each entry. Obtain the backup duration by recording the total time consumed from the start time to the end time of the backup, and conduct unified statistics in seconds or minutes. Then, combine the values obtained from the real-time bandwidth monitoring during the transmission process to accumulate the speeds for each period and calculate the average value to obtain the data transfer rate. For the measurement of the number of errors, increment the count each time a fault prompt or interruption occurs, and organize it corresponding to the corresponding data number. If the number of errors exceeds three times, it is marked as a multiple-fault occurrence. Here, the judgment of three times comes from twice the average value of the errors that occurred in previous backups and then increased by one time to obtain a margin. Subsequently, form a comparison table of the backup durations and the number of errors of all entries in the order of the numbers, and compare the transfer rate and the duration of each item in the comparison table. If the transfer rate is lower than two megabits per second or the duration exceeds one hour, it is additionally marked so that it can be classified as an abnormal status entry later. If the number of errors reaches five times, it is included in the priority troubleshooting list and the reasons for the abnormality are recorded, such as network jitter or unstable storage device ports, etc. To judge these reasons, detailed fault prompts need to be recorded when collecting relevant logs and compared with the running status of the server. For example, match the abnormal port moment with the transmission interruption moment to confirm the fault point, and then summarize all the records into a record of the data backup execution process.
[0129] The usefulness of the formula lies in that it comprehensively considers four factors: error frequency during the backup process, data transmission efficiency, total number of backup items and backup duration, and measures the stability of the backup execution through the synergy of multiple parameters.
[0130] E d The steps to obtain the parameters are as follows: register the errors that have accumulated during this backup process one by one, and record the specific time of occurrence, error type, and the resulting termination or suspension for each error event. Then, summarize the number of occurrences to obtain the number of errors generated by this data or this backup. If eight of the records are error events after checking the backup operation logs for 100 times, count these eight times as error events. d =8.
[0131] R b The steps for obtaining the parameters are as follows: measure the rate during the data transmission process, read the amount of data bytes transmitted at regular intervals (e.g., ten seconds) and convert it into megabits per second and record it, then accumulate the recorded rate values during the entire backup time and divide it by the number of acquisitions to get the average value. If the rate values collected thirty times total 1,200 megabits per second, then
[0132] L s The steps to obtain the parameters are: count the total number of data items associated with a single backup, read the list of data numbers that need to be backed up before starting the backup, and count them one by one. For example, if a backup contains 300 data items, then register this number as L s =300.
[0133] H b The steps to obtain the parameters are as follows: measure the duration of a single data access entry from the start of backup to the completion of backup. If a data entry is backed up from 09:00:00 and completed at 09:00:20, the backup duration of the entry is 20 seconds. Then take the average value of all entries or weight them according to actual needs and confirm H. b For example, if the duration of twenty items is sampled and the sum of these durations is divided by twenty to get thirty seconds, then record H b =30.
[0134] Calculation process:
[0135] The first step is to bring in the parameters: E d =8, R b =40, L s =300, H b =30;
[0136] The second step is to calculate the numerator first. Then with R b×L s
[0137] = 40 × 300 = 12000. Multiply to get 64 × 12000 = 768000, and then take the cube root
[0138] Step 3: Calculate the denominator
[0139] Step 4: Divide the numerator by the denominator:
[0140] This result indicates that when K ≈ 16.43, the backup stability index generated by the combination of factors such as the number of error events, transmission speed, number of horizontal entries, and entry backup duration in this backup reaches 16.43. In subsequent steps, this value can be compared with the stability indices of other backup tasks within the system. If it is greater than a certain threshold, it is determined that there is a relative fluctuation in backup stability; otherwise, it is considered to be within the normal range.
[0141] Based on the backup stability index obtained previously, first compare the K values of all entries one by one and match them with the pre-defined grade range. For example, set a threshold of ten, which is obtained by taking the average value of eight of the K values of one hundred successfully backed-up records and increasing it by two. If the K value of the current entry is higher than ten, it is assigned to the priority viewing category, and continue to detect whether the number of its errors exceeds the three-time threshold or whether the transmission rate is lower than twenty megabits per second. If either condition is met, it is marked as potentially abnormal and the cause of the error is retrieved. If it is detected that there is an idle disk failure prompt or a timeout disconnection event during the backup of the entry, record the details. Finally, after completing the comparison of all entries, separately summarize the entries that have not been successfully backed up, summarize their performance in terms of error events, network bandwidth, and disk write load, and assemble and output this information to form the backup status result.
[0142] The steps to obtain the backup effect evaluation result are as follows:
[0143] According to the backup completion status of each data access entry recorded in the backup status result, screen and extract the data access entries that failed to be backed up item by item, and record the reasons for backup failure, the data transmission rate at the time of failure, and the error type to generate a list of data access entries that failed to be backed up;
[0144] Based on the list of data access entries that failed to be backed up, perform the re-backup operation of the data access entries item by item, and monitor the data transmission rate and backup error conditions during the re-backup process in real time to generate the backup effect evaluation result.
[0145] Specifically, according to the backup completion status of each data access entry presented in the previous backup status result, first compare the access numbers of all entries and record whether the current backup is successful. Then, screen item by item according to records such as whether there are interruption prompts, network anomalies, or insufficient storage space. Classify the entries determined to be incompletely backed up as failure cases. For each failed item, it is necessary to further extract the specific transmission rate and error type when the failure occurred. For example, when the transmission rate is lower than two megabits per second, it can be marked as insufficient rate and an appropriate anomaly identifier can be attached. The threshold of two megabits per second is obtained by averaging the transmission speed of one megabit and eight hundred kilobits per second for one hundred completed backups and then increasing it by two hundred kilobits per second. If the error type shows disconnection or read-write conflict, it is necessary to check the server log to find out what kind of failure occurred during the corresponding time period and confirm the failure point. For example, if a port occupancy conflict is detected at 10:30 and this conflict occurs within ten seconds after the backup request is issued, the conflict can be recorded in the failure reason of this data item. If the error type is marked as checksum mismatch, it is necessary to compare the backup checksum with the source data checksum and record the final result. Then, compare it with the data transmission rate collected at the time when the error occurred. If the rate is lower than 1.5 megabits per second, it is recorded as a deep anomaly. 1.5 megabits per second is obtained by averaging the rate of 1.2 megabits per second for fifty worst backup records and then increasing it by 0.3 megabits per second. If the error reason is that the device capacity is exceeded, the current remaining capacity of the device can be further queried. For example, a capacity range of 8GB to 16GB is set as the normal remaining range. If the detected value is lower than 8GB, it is marked as insufficient capacity. List all these failure reasons and error types separately, along with the specific time and transmission speed when they occurred, and merge them item by item under the information of the same item. Ensure the accuracy of the records and exclude duplicate false alarms through multiple comparisons and log verifications. Then, summarize the item numbers corresponding to each record, their failure reasons, error types, and data transmission rates, and finally organize them into a list of backup failure data entries.
[0146] Based on the list of backup failure data entries generated previously, first assign a new re-backup period for each failure entry and record the required network bandwidth threshold and available storage capacity range. For example, set the network bandwidth to be no less than two megabits per second and the remaining storage space to be no less than 10 GB. The bandwidth value of two megabits per second is obtained by adding 200 kilobits per second to the known average rate of 1.8 megabits per second from the investigation to ensure redundancy. At the same time, the 10 GB space threshold is obtained by basing on the maximum storage amount of 8 GB required for some entries detected previously and reserving 2 GB. Then, start executing the re-backup operation item by item and monitor the transmission rate during this process. If the transmission rate drops below two megabits per second, mark the speed reduction in the monitoring record and track the duration of this status. If the number of errors accumulates more than twice during the re-backup process, register it in the failure statistics of the corresponding entry and consult the corresponding backup timestamp to check whether there is a port conflict or network switch. If it is a port conflict, compare the server working status at that moment and view the hardware monitoring record to confirm the specific conflict reason. If it is a network switch, check whether there is a primary and standby line switch operation in the routing forwarding table and record the switching duration. After completing the re-backup of all entries, compare the number of errors and average rate observed during the re-backup for each entry again. Aggregate the recorded data of each entry for subsequent analysis. Finally, obtain the final status and statistical information of each entry during the overall re-backup process and compile them into the backup effect evaluation result.
[0147] The above is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes and apply them to other fields. However, as long as it does not depart from the technical solution content of the present invention, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention still fall within the protection scope of the technical solution of the present invention.
Claims
1. A data integration management method based on a data middle platform, characterized in that, It includes the following steps: Authenticate the user, and based on the user's role attributes and the preset access control list in the database, determine the data access permissions to obtain the user access permission configuration; monitor the user's real-time data access behavior, evaluate the potential security risks by comparing the deviation between the behavior pattern and the normal pattern, and generate the user behavior evaluation result; Based on the user access permission configuration, match the permission verification rules for each data access request, perform permission verification, and record the verification results to obtain the permission verification record; Utilize the user behavior evaluation result, compare with the abnormal behavior, adjust the user's real-time access permissions, restrict or expand the data access scope to obtain the adjusted user permissions; Set the backup time interval according to the data access frequency in the permission verification record to obtain the data backup plan; Execute data backup according to the data backup plan, and monitor the backup completion degree and errors in the backup process in real time to generate the backup status result; Based on the backup status result, evaluate the effectiveness of the backup operation, and re-back up the data items that have not been successfully backed up to obtain the backup effect evaluation result.
2. The data integration management method based on the data middle platform according to claim 1, wherein The steps for obtaining the user access permission configuration are as follows: Compare the role attributes in the user identity information item by item with the preset access control list in the database, perform item-by-item screening based on the permission limitation rules of the role attributes, and confirm the permission level in combination with the identity authentication status corresponding to the user identity information to generate the user role permission level; Based on the user role permission level, match the data access permission entries, perform screening and combination of the content of the permission entries, and perform validity checks on the permission entries to eliminate redundant or invalid permission entries to generate the user access permission configuration.
3. The data integration management method based on the data middle platform according to claim 1, wherein, The steps for obtaining the user behavior evaluation result are as follows: Call the user's real-time data access behavior, extract and record the time interval of the user access behavior, the access request response delay, and the data call quantity item by item according to the target data entry accessed by the user, the access duration, and the frequency of the data access request to generate the user real-time access behavior characteristic data; Based on the user real-time access behavior characteristic data, calculate the behavior pattern deviation value, and the calculation formula is: Among them, Y represents the deviation value of the behavior pattern, and T a represents the time interval of the user's current access behavior, and T n represents the average time interval of access behavior in the normal mode, and F r represents the current data call quantity of the user, and F n represents the average value of the data call quantity in the normal mode, and D a represents the response delay of the user's current access request, and D n represents the average value of the response delay of access requests in the normal mode; Based on the behavior pattern deviation value, determine the security risk level of the user's real-time data access behavior, and combine the security risk level to determine whether the behavior risk exceeds the security standard to generate the user behavior evaluation result.
4. The data integration management method based on the data middle platform according to claim 1, characterized in that The steps for obtaining the permission verification record are as follows: Based on the user access permission configuration, for each data access request submitted by the user, retrieve the data access control list item by item, perform permission verification rule matching for the data entries of the access request, and complete the validity check of the permission matching rules to form the permission verification rules; Based on the permission verification rules, compare and verify the permissions of the data entries in each data access request item by item to verify whether the data access request conforms to the permission verification rules, and record the execution results of the permission verification of each data access request item by item to obtain the permission verification record.
5. The data integration management method based on the data middle platform according to claim 1, wherein The steps for obtaining the adjusted user permissions are as follows: Extract and record item by item the response delay duration of the access data corresponding to the abnormal behavior, the recurrence times of the abnormal requests, and the data call quantity of a single request according to the security risk level in the user behavior evaluation result, and generate abnormal behavior feature data; Calculate a behavior risk index based on the abnormal behavior feature data, and the calculation formula is: Among them, R represents the behavioral risk index, f a represents the repetition request frequency of the user's current abnormal request, d a represents the total amount of data calls associated with abnormal behaviors, q a represents the data sensitivity level value of the current access request, u n represents the data access permission level corresponding to the user's current request, t a represents the continuous duration during which abnormal request behaviors have occurred cumulatively; Based on the behavior risk index, numerically compare the behavior risk index with the permission risk threshold, judge the deviation direction between the behavior risk index and the permission risk threshold, determine whether the real-time access permission of the user needs to be expanded or restricted, generate a permission adjustment determination result, and perform a permission update operation on the data access control list according to the permission adjustment determination result to generate an adjusted user permission.
6. The data integration management method based on a data middle platform according to claim 1, wherein The steps for obtaining the data backup plan are as follows: According to the cumulative access times and access time intervals of each data access entry in the permission verification record, count the access times of each data access entry in the current consecutive five data access requests, and calculate the average access response delay of the access requests to generate the consecutive access times of the data access entry and the average response delay of a single access request; Calculate a backup priority index based on the consecutive access times of the data access entry and the average response delay of a single access request, and the calculation formula is: Among them, G is the backup priority index, C a is the cumulative access times of the data access entry within the most recent hour, EF r is the average response latency of a single access request for this data access entry, Q a is the amount of data called by a single request for this data access entry, T r is the average time interval between two consecutive access requests for this data access entry; Based on the backup priority index, confirm the backup execution time interval corresponding to each data access entry to generate a backup time interval and obtain a data backup plan.
7. The data integration management method based on a data middle platform according to claim 1, characterized in that The steps for obtaining the backup status result are as follows: According to the data entries in the data backup plan, collect item by item the backup duration, data transmission rate, and the number of errors during the backup process of the data entries to form a record of the data backup execution process; Calculate a backup stability index based on the data backup execution process record, and the calculation formula is: Among them, K represents the backup stability index, E d represents the cumulative number of errors that occurred during the backup process, R b represents the average data transfer rate of this backup, L s represents the total number of data entries associated with a single backup, H b is the duration of the backup for a single data access entry; Based on the backup stability index, verify the completion degree of the backup, identify the entries that have not been successfully backed up and the reasons, and generate a backup status result.
8. The data integration management method based on a data middle platform according to claim 1, wherein, The steps for obtaining the backup effect evaluation result are as follows: According to the backup completion status of each data access entry recorded in the backup status result, screen and extract item by item the data access entries that have failed to be backed up, and record the reasons for the backup failure, the data transmission rate at the time of failure, and the error type to generate a list of data access entries that have failed to be backed up; Based on the list of data access entries that have failed to be backed up, perform a re-backup operation on each data access entry item by item, and monitor in real time the data transmission rate and backup error conditions during the re-backup process to generate a backup effect evaluation result.