A cloud computing-based file collaborative processing method and system
By introducing asymmetric risk assessment and nonlinear ranking amplification optimization factors into the file collaborative processing system, a risk-aware asymmetric distance is constructed, which solves the problem that the Euclidean distance cannot distinguish the deviation direction, and achieves more accurate abnormal behavior identification and security monitoring.
Patent Information
- Application Number
- CN202511057374.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-07-30
AI Technical Summary
The existing Euclidean distance cannot distinguish the directionality of operation deviations in different feature dimensions in the file collaborative processing system, resulting in difficulty in accurately identifying high-risk deviations and reducing the accuracy of clustering algorithms in identifying abnormal behaviors.
By performing asymmetric risk assessment and nonlinear ranking amplification on multidimensional feature vectors, introducing asymmetric direction optimization factors and nonlinear ranking amplification optimization factors, a risk-aware asymmetric distance is constructed for identifying and processing abnormal behaviors in the clustering process.
It improves the accuracy of abnormal behavior identification, reduces the false alarm rate, can timely identify sparse behavior clusters and locate highly correlated users, and improves the data security level of the collaborative platform.
Smart Images

Figure CN120561626B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of file collaboration, and in particular to a file collaborative processing method and system based on cloud computing. Background Art
[0002] With the development of informatization and networking, collaborative file processing has become an indispensable part of daily work for businesses and teams. Cloud-based collaborative platforms, such as online documents, collaborative design tools, and cloud-based integrated development environments, allow multiple geographically dispersed users to simultaneously preview, edit, and save the same shared file, significantly improving work efficiency.
[0003] In the file collaborative processing system based on cloud computing, a common technology for detecting abnormal behavior is to use The clustering algorithm processes massive amounts of user operation logs. This method abstracts each user operation into a multidimensional feature vector based on its basic characteristics such as operation time, operation type, and operation data size, and then uses The algorithm clusters these feature vectors.
[0004] However The standard Euclidean distance metric used in the algorithm presents a problem due to its inherent mathematical properties when applied to collaborative operation logs composed of the aforementioned basic features. Standard Euclidean distance is symmetric when calculating the distance between an operation data point and a cluster centroid. Its calculation process shows that the distance depends solely on the sum of the squares of the differences between the two feature vectors in each dimension. Therefore, it only measures the absolute numerical difference between the two in the feature space and cannot reflect the directionality of this difference. In the security risk assessment of collaborative systems, changes in different feature dimensions often have asymmetric risk implications. For example, the discrete feature of operation type is encoded as a set of integers for calculation. Although the deviation from a normal operation type (update) to a higher-risk operation type (delete) and the deviation from a lower-risk operation type (read) have the same numerical difference, the security risk levels they represent are completely different. The symmetry of the standard Euclidean distance prevents it from distinguishing risk differences caused by deviations in direction. A deviation toward higher risk and a deviation of the same magnitude toward lower risk are considered equally abnormal. Furthermore, for continuous features like operation data size, the risk increase is not linear. An operation with a large value far outside the normal range will increase risk much more than a small value slightly below the normal range. The symmetry of the standard Euclidean distance also fails to capture this asymmetric risk variation.
[0005] Due to the symmetry of the standard Euclidean distance, it cannot differentiate the risk according to the direction of the numerical deviation in different feature dimensions, thereby reducing the accuracy of the clustering algorithm in identifying specific types of high-risk abnormal behaviors when processing the collaborative operation log the accuracy of the clustering algorithm in identifying specific types of high-risk abnormal behaviors. SUMMARY
[0006] Therefore, the embodiments of the present application provide a file collaborative processing method based on cloud computing to solve the problem that the symmetry of the existing Euclidean distance makes it difficult to accurately identify high-risk deviations.
[0007] To achieve the above object, the technical scheme of the present application is as follows:
[0008] In a first aspect, the present application provides a file collaborative processing method based on cloud computing, which comprises the following steps:
[0009] Step S1: obtaining a multi-dimensional feature vector by performing feature mapping on the collaborative operation log in the background of the cloud collaborative processing system server;
[0010] Step S2: obtaining an asymmetric direction optimization factor by performing asymmetric risk assessment on the dimension distribution of the multi-dimensional feature vector;
[0011] Step S3: obtaining a nonlinear ranking amplification optimization factor by performing nonlinear amplification on the intra-cluster deviation ranking of the multi-dimensional feature vector;
[0012] Step S4: obtaining a risk-aware asymmetric distance by performing fusion assessment on the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor;
[0013] Step S5: obtaining an optimized clustering result by performing a clustering process based on the risk-aware asymmetric distance, and identifying and processing abnormal behaviors based on the optimized clustering result.
[0014] Preferably, the multi-dimensional feature vector is obtained by performing feature mapping on the collaborative operation log in the background of the cloud collaborative processing system server, which comprises:
[0015] Data is collected in the background of the cloud collaborative processing server to obtain log data recording user historical operations; for each independent operation in the log, a multi-dimensional feature vector is generated; the first dimension of the multi-dimensional feature vector is a log data time sine feature; the second dimension is a log data time cosine feature; the third dimension is a log data operation code; and the fourth dimension is the number of bytes affected by the corresponding operation of the log data.
[0016] Preferably, the asymmetric direction optimization factor is obtained by performing asymmetric risk assessment on the dimension distribution of the multi-dimensional feature vector, which comprises:
[0017] The asymmetric risk coefficient is obtained by performing random difference sampling analysis on the multidimensional feature vectors of historical operation logs. The direction deviation assessment is obtained by performing directional discrimination processing on the difference between the eigenvalues and the cluster centroid in the multi-bit feature vector. The asymmetric direction optimization factor is obtained by performing a fusion evaluation of the asymmetric risk coefficient and the direction deviation assessment.
[0018] Preferably, the step of obtaining an asymmetric risk coefficient by performing random difference sampling analysis on the multi-dimensional feature vectors of historical operation logs includes:
[0019] In the historical operation log, two historical operation data are randomly selected with replacement. The value of the subtraction of the eigenvalues of the eigenvectors of the two historical operation data in any dimension compared with 0 is used as the first dimension difference of the dimension; the absolute value of the subtraction of the eigenvalues of the eigenvectors of the two historical operation data in any dimension is used as the second dimension difference of the dimension;
[0020] Set the number of random difference samplings; for any dimension, use the result of adding the first dimension differences of all random difference samplings as the numerator, and the result of adding the second dimension differences of all random difference samplings as the denominator, and use the result of the corresponding fraction as the asymmetric risk coefficient of the dimension.
[0021] Preferably, the step of performing direction discrimination processing on the difference between the eigenvalues in the multi-bit eigenvector and the cluster centroid to obtain the direction deviation assessment includes:
[0022] Set the unit step function. When the independent variable of the unit step function is greater than 0, the output value of the unit step function is 1. When the independent variable of the unit step function is less than or equal to 0, the output value of the unit step function is 0.
[0023] Obtain all cluster centroids during the clustering process; for any dimension of the feature vector of any historical operation data, the result of subtracting the feature value of that dimension from the centroid of any cluster of that dimension is used as the first centroid difference evaluation;
[0024] The first center of mass difference assessment is mapped through the unit step function, and the corresponding mapping result is used as the direction deviation assessment.
[0025] Preferably, the step of obtaining the asymmetric direction optimization factor by fusing the asymmetric risk coefficient with the direction deviation assessment includes:
[0026] For any dimension of the characteristic vector of any historical operation data, the larger value between the settlement result of subtracting the asymmetric risk coefficient of the dimension from one half and the constant 0 is multiplied by the constant 2, and the corresponding calculation result is used as the first evaluation of direction optimization; the calculation result of multiplying the direction deviation evaluation by the first evaluation of direction optimization and adding the constant 1 is used as the asymmetric direction optimization factor.
[0027] Preferably, the step of obtaining a nonlinear ranking amplification optimization factor by nonlinearly amplifying the intra-cluster deviation ranking of the multi-bit feature vectors includes:
[0028] In each cluster assignment step of the iterative process of the clustering algorithm, for any target cluster, all member data points belonging to the target cluster are obtained, and a deviation value set is formed by the absolute deviation values of the member data points relative to the cluster centroid in different dimensions; the absolute deviation value is the absolute value of the difference between the eigenvalue and the cluster centroid; for any member data point, the normalized deviation ranking of the member data point is obtained by calculating the percentile ranking of the deviation value of the member data point in the deviation value set; the square of the calculation result of adding the normalized deviation ranking to the constant 1 is used as the nonlinear ranking amplification weight factor.
[0029] Preferably, the step of obtaining the risk-aware asymmetric distance by fusion evaluation of the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor includes:
[0030] For any dimension of any operation log data and any dimension of any cluster centroid; the square of the calculation result of subtracting the eigenvalue of the target dimension of the eigenvector of the log data from the target dimension of the cluster centroid is used as the first distance evaluation between the operation log data in the target dimension and the corresponding dimension of the cluster centroid; the calculation result of multiplying the asymmetric direction optimization factor between the operation log data in the target dimension and the corresponding dimension of the cluster centroid by the nonlinear ranking amplification weight factor is used as the distance optimization factor; the calculation result of multiplying the distance optimization factor by the first distance evaluation is used as the corresponding optimized distance; the calculation result of adding the optimized distances of all dimensions is used as the risk-aware asymmetric distance between the operation log data and the cluster centroid.
[0031] Preferably, obtaining an optimized clustering result through a clustering process based on risk-aware asymmetric distance, and performing abnormal behavior identification based on the optimized clustering result, includes:
[0032] Set the number of clusters for the clustering algorithm; perform clustering during the clustering iteration process using the risk-aware asymmetric distance between the operation log data and the cluster centroids; complete the iteration process and obtain the optimized clustering results;
[0033] The abnormal behavior evaluation is performed through two stages; the first stage is to evaluate the cluster data point quantity in the optimized clustering result, and determine the sparse cluster as an abnormal behavior mode; the second stage is to obtain the user highly associated with the abnormal behavior mode through the associated abnormal user identification based on the user contribution degree.
[0034] In a second aspect, the present application provides a cloud computing-based file collaborative processing system, comprising a processor and a memory, the memory storing computer program instructions, when the computer program instructions are executed by the processor, a cloud computing-based file collaborative processing method is realized.
[0035] Compared with the prior art, the embodiment of the present application has the following beneficial effects:
[0036] The present application introduces an asymmetric direction optimization factor and a nonlinear ranking amplification optimization factor, so that the distance metric can impose greater punishment on operations that deviate towards high-risk directions, and exponentially amplify extreme deviations, so that high-risk behaviors such as download surge and batch deletion are significantly pushed away from the normal behavior cluster area in the feature space. Combined with the endogenous asymmetric risk coefficient calculated by the sliding window, the model can continuously track the latest business mode and reduce the missed detection rate of new risk patterns; at the same time, the contour coefficient is used to automatically select the optimal clustering number, improve the clustering boundary definition, and make the abnormal and normal behavior more accurate. In the actual application of cloud file collaboration, this risk-aware distance metric can identify sparse behavior clusters that only account for a very small proportion but have extremely high risk, and automatically locate highly associated users for subsequent operation to perform enhanced monitoring or multi-factor verification. Compared with the traditional symmetric distance method, this scheme greatly reduces false positives without increasing computational complexity, improves the credibility of security alerts, helps enterprises to discover potential leakage events earlier under the same log size, and improves the overall data security level of the collaborative platform. BRIEF DESCRIPTION OF DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0038] Figure 1 It is a method flowchart of a cloud computing-based file collaborative processing method provided by the first embodiment of the present application. DETAILED DESCRIPTION
[0039] The embodiments of the present disclosure are described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to be used to explain the present disclosure, but should not be understood as limiting the present disclosure.
[0040] In order to illustrate the technical solution of the present invention, specific embodiments are provided below.
[0041] See also Figure 1 , is a method flow chart of a file collaborative processing method based on cloud computing provided by the first embodiment of the present invention, such as Figure 1 As shown, the method may include:
[0042] Step S1 : obtaining a multi-dimensional feature vector by performing feature mapping on the collaborative operation log of the backend of the cloud collaborative processing system server.
[0043] First, data collection is performed on the backend of a cloud-based collaborative processing server to obtain log data recording historical user operations. For each independent operation in the log, a multidimensional feature vector is generated. The first dimension of the multidimensional feature vector is the temporal sine feature of the log data; the second dimension is the temporal cosine feature of the log data; the third dimension is the operation encoding of the log data; and the fourth dimension is the number of bytes affected by the corresponding operation in the log data.
[0044] Specifically, in order to accurately represent the periodic position of the operation time within 24 hours in the first dimension and the second dimension and avoid the distance calculation distortion caused by the numerical boundary, the present invention rounds down the number of minutes from the timestamp information recorded in the operation log to calculate the total number of minutes from the time the operation occurred to midnight on the same day. The range of this value is The number of minutes is then transformed by trigonometric functions and mapped to a feature dimension representing its position on the unit circle. For the time sine feature The calculation formula is ; For the time cosine feature The calculation formula is .
[0045] In the third dimension, according to the operation type identifier recorded in the operation log, it is mapped to a preset integer code, which is used to convert the character-type operation type into a feature that can be used for numerical calculation. In this embodiment, the mapping rule is defined as follows: "Read" operation is coded as 0, "Update" operation is coded as 1, "Delete" operation is coded as 2, "Share" operation is coded as 3, and "Download" operation is coded as 4.
[0046] In the fourth dimension, the number of bytes of the data affected by the operation is obtained, and the number of bytes is logarithmically transformed, thereby smoothing the data with a large data level difference and causing uneven numerical distribution. In the embodiment of the present application, the logarithmic transformation is performed by a logarithmic function The number of bytes is added to the constant 1 to avoid mathematical nonsense results when taking the logarithm of 0 bytes.
[0047] At this point, the multi-dimensional feature vector is obtained by feature mapping the collaborative operation log in the background of the cloud collaborative processing system server.
[0048] In step S2, an asymmetric direction optimization factor is obtained by asymmetric risk assessment of the multi-dimensional feature vector distribution.
[0049] In In the implementation scenario of the clustering algorithm applied to collaborative operation log anomaly behavior detection, the calculation formula of the standard Euclidean distance is on a single feature dimension The distance contribution is calculated as the square of the difference between the data point and the cluster center , that is . Regardless of the positive or negative of the difference, the square result is positive, so the distance measure is symmetric. It can only measure the size of the difference between two data points in numerical value, and completely ignores the direction of the difference. However, in the specific technical problem of security risk assessment of the file collaborative processing system, the risk degree of an operation behavior is not only related to the magnitude of its feature value deviating from the normal mode, but also closely related to the direction of its deviation. Taking the single download data size feature as an example, a download operation much larger than the normal average value indicates a much higher data leakage risk than a download operation smaller than the normal average value. The standard Euclidean distance cannot distinguish between the two directions of deviation with completely different properties due to its inherent symmetry, and will incorrectly judge that a high-risk upward deviation and a generally risk-free downward deviation have the same degree of abnormality under the same numerical difference, thereby reducing the recognition sensitivity of the anomaly detection model to specific risk patterns.
[0050] In order to solve the problem that the distance metric function in the prior art cannot accurately evaluate the risk directionality due to symmetry, the first step of the present application is to obtain an asymmetric risk coefficient by randomly sampling and analyzing the multi-dimensional feature vectors of the historical operation log, to obtain a direction deviation evaluation by performing direction discrimination processing on the difference between the feature values in the multi-dimensional feature vector and the cluster centroid, and to obtain an asymmetric direction optimization factor by performing fusion evaluation on the asymmetric risk coefficient and the direction deviation evaluation. A multiplication adjustment term related to the deviated direction is introduced when calculating the distance. When the value of an operation data point in a certain feature dimension is higher than the cluster centroid compared with it (i.e. upward deviation occurs, and this direction is usually related to risk increase in safety evaluation), the weight factor is an amplification coefficient greater than 1, so as to increase the distance penalty in this dimension; and when the value is lower than or equal to the cluster centroid (i.e. downward deviation or no deviation occurs), the weight factor should be 1, i.e. no amplification is performed on the distance, so as to reflect the low risk property of this direction. The present application proposes an amplification coefficient for upward deviation, and the size of the amplification coefficient is adaptively determined by the overall distribution form of the feature dimension in the historical operation log data. The higher the internal risk of upward deviation of a feature dimension is, the more the data distribution of the feature dimension presents a positive skewness with long tail on the right (i.e. normal values are concentrated in the low area, and high-value anomalies are sparse), and the greater the amplification coefficient corresponding to the feature dimension should be. By constructing such a non-symmetrical weight factor which can convert the objective non-symmetry of data distribution into a calculable non-symmetrical weight factor, the standard Euclidean distance is optimized.
[0051] Specifically, in the historical operation log, two historical operation data are randomly extracted in a replacement manner, a calculation result of subtracting the feature value of the feature vector of the two historical operation data in any dimension is compared with 0, and a value greater than 0 is taken as a first dimension difference of the dimension; an absolute value of the calculation result of subtracting the feature value of the feature vector of the two historical operation data in any dimension is taken as a second dimension difference of the dimension; a random difference sampling number is set; for any dimension, a calculation result of adding all the first dimension differences of the random difference sampling is taken as a numerator, a calculation result of adding all the second dimension differences of the random difference sampling is taken as a denominator, and a calculation result of the corresponding fraction is taken as an asymmetric risk coefficient of the dimension.
[0052] In an embodiment, it is assumed that the value of the first historical operation data in the first feature dimension in the first random sampling is a, the value of the second historical operation data in the first feature dimension in the first random sampling is b, the value of the first historical operation data in the second feature dimension in the first random sampling is c, the value of the second historical operation data in the second feature dimension in the first random sampling is d, the random sampling number is n, and the asymmetric risk coefficient of the first feature dimension is calculated as follows:
[0053]
[0054] in, Indicates the Asymmetric risk coefficients for each characteristic dimension; Indicates the number of random draws; Indicates the The first historical operation data in the random sampling is The numerical value of the feature dimension; Indicates the The second historical operation data in the random sampling is The numerical value of the feature dimension; represents the maximum value function; Indicates absolute value calculation.
[0055] It should be noted that the historical operation log data set described in this embodiment is dynamically determined by a mechanism based on a sliding time window, and contains all operation logs within the past 24 hours from the current moment. The design of using the data of the last 24 hours in this invention is intended to give priority to ensuring that the model is sensitive to business changes and adapts quickly. It sacrifices a certain amount of memory of ultra-long-term historical patterns in exchange for the accurate capture of recent patterns, thereby providing more timely abnormal behavior detection in a dynamically changing collaborative environment. The total number of random comparisons The value of is independent of the total number of data points in the historical operation log dataset; it represents the number of independent sampling comparisons performed to obtain a robust statistical estimate. In this embodiment of the present invention, regardless of the number of data points in the historical dataset, the number of random sampling comparisons can be set to a fixed, large value. In this embodiment, the initial value is 10,000 to balance computational efficiency and statistical accuracy.
[0056] After obtaining the asymmetric risk coefficient, continue to perform directional discrimination processing on the difference between the eigenvalues in the multi-bit eigenvector and the cluster centroid to obtain a directional deviation assessment. Specifically, set a unit step function. When the independent variable of the unit step function is greater than 0, the output value of the unit step function is 1, and when the independent variable of the unit step function is less than or equal to 0, the output value of the unit step function is 0; obtain all cluster centroids in the clustering process; for any dimension of the eigenvector of any historical operation data, subtract the eigenvalue of the dimension from the centroid of any cluster of the dimension as the first centroid difference assessment; map the first centroid difference assessment through the unit step function, and use the corresponding mapping result as the directional deviation assessment. For any dimension of the eigenvector of any historical operation data, multiply the larger value between the settlement result of subtracting one-half of the asymmetric risk coefficient of the dimension and the constant 0 by a constant 2, and use the corresponding calculation result as the first directional optimization assessment; multiply the directional deviation assessment by the first directional optimization assessment and add the constant 1 to calculate the result as the asymmetric directional optimization factor.
[0057] In one embodiment, the unit step function is assumed to be ;No. Operation log data The characteristic data of the dimension is ;No. The centroid of the cluster is The characteristic data of the dimension is , then The operation log data and The centroid of the cluster is The calculation expression of the asymmetric direction optimization factor of the dimension is:
[0058]
[0059] in, Indicates the The operation log data and The centroid of the cluster is Asymmetric directional optimization factors in dimensions; Indicates the Operation log data Feature data of dimensions; Indicates the The centroid of the cluster is Feature data of dimensions; represents the unit step function; Indicates the Asymmetric risk coefficients for each characteristic dimension; Represents the maximum function.
[0060] It should be noted that the core of the present application is to solve the symmetry defect of the standard Euclidean distance, so that different risk weights can be given to deviations in different directions. First, the formula in the This part is the switch structure designed by the present application to introduce directional judgment. In the context of collaborative security assessment, upward deviation (i.e. greater than the normal value) is usually a risk signal that needs to be focused on. The role is to open the subsequent risk amplification calculation when the data point The value of is greater than the cluster center , the independent variable is greater than 0, and the function result is 1, so as to open the subsequent risk amplification calculation; and when the value of is less than or equal to , the independent variable is less than or equal to 0, and the function result is 0, so as to close the risk amplification calculation, so that the final weight is equal to its baseline value 1, ensuring that the amplification effect is only activated when upward deviation occurs in this high-risk direction. Secondly, the asymmetric risk coefficient decides the strength of risk amplification. This coefficient calculates the proportion of the total amount of upward difference (positive difference) in all randomly generated differences to objectively measure the skew direction and degree of a data distribution. In a typical characteristic dimension with security risk (such as single-day download volume), the data distribution is almost certainly positively skewed, i.e. the vast majority of normal operation values are concentrated in a very low area, while a few abnormal operation values are very large. In this right-tailed distribution, the probability and total amount of generating a large positive difference between two randomly selected points are much greater than generating a large negative difference. Therefore, for such risk dimensions, the calculated value of will be significantly greater than 0.5. Conversely, for a symmetrically distributed dimension, the value will be approximately equal to 0.5. Finally, substitute into . Compare with a representative baseline point of 0.5. Only when is greater than 0.5, i.e. the data distribution indeed exhibits a positive skew that needs attention, the result of the maximum value calculation is a positive number greater than 0, thus producing a final weight greater than 1. This ensures that only those dimensions that have proven to have asymmetric risk will be punished for upward deviation, while for dimensions that are inherently symmetrically distributed, their weight will only be 1 and will not be unfairly punished even if upward deviation occurs.
[0061] Thus, the non-symmetric directional optimization factor is obtained by performing non-symmetric risk assessment on the distribution of the multi-bit feature vector dimension.
[0062] Step S3, obtain a non-linear ranking amplification optimization factor by non-linear amplification of the intra-cluster deviation ranking of the multi-bit feature vector.
[0063] In step S2, the present application solves the symmetry problem that the standard Euclidean distance cannot distinguish the risk deviation direction by constructing an asymmetric direction optimization factor. However, in the collaborative security risk assessment, the growth of risk is nonlinear. The greater the deviation of an operation, the greater the risk growth is not usually in a linear relationship proportional to the square of the distance. For example, for the feature of single download data size, an operation deviates from the normal value of 1 MB to 10 MB, and another operation deviates from the normal value of 1 GB to 1.01 GB. Although the former is much larger than the latter in numerical value and relative amplitude, the behavior of the latter (downloading a GB-level file) itself is in a risk interval that needs to be highly concerned. The standard Euclidean distance, even after the weighting in step a, still mainly depends on the square of the numerical difference, and its risk measurement is relatively smooth, and cannot amplify the risk of operations with small numerical deviation.
[0064] In order to solve this problem, that is, to make the distance measurement more sensitive to capture and amplify the risk of extreme deviation, in each cluster assignment step of the iteration process of the clustering algorithm, for any target cluster, all member data points belonging to the target cluster are obtained, and a deviation value set is formed by the absolute deviation values of the member data points in different dimensions relative to the cluster centroid. The absolute deviation value is the absolute value of the feature value minus the cluster centroid. For any member data point, the normalized deviation ranking of the member data point is obtained by calculating the percentile ranking of the deviation value of the member data point in the deviation value set. The square of the calculation result of adding the normalized deviation ranking to a constant 1 is taken as a non-linear ranking amplification weight factor. The core design idea of this factor is that the risk amplification degree of an operation should not be determined only by the absolute or relative numerical value of its deviation, but more by the relative position of its deviation value in all member deviation degrees in its belonging cluster. Even if its numerical value is not large, if it is already the fewest in its cluster, it should be punished by an additional non-linear amplification.
[0065] In an embodiment, assuming that the normalized deviation ranking of the jth operation log data in the ith dimension for the ith cluster is , then the calculation expression of the non-linear ranking amplification weight factor of the jth operation log data in the ith dimension for the ith cluster is:
[0066]
[0067] wherein, represents the th operation log data for the th cluster class in the th dimension; represents the th operation log data for the th cluster class in the th dimension.
[0068] It is to be noted that the present application aims to solve the linear problem of the existing distance metric method in step S3, that is, it can impose a more severe punishment on extreme deviation. For this purpose, the present application designs a non-linear amplification weight based on relative ranking . According to the relative ranking, all metrics that depend on specific numerical distribution assumptions or statistical parameters are replaced. Whether a data point is extremely deviated is no longer determined by its absolute value, but by its relative position in the local group to which it belongs. By mapping the ranking in the range of to the interval of , it is ensured that the basic amplification factor is at least 1. Squaring it introduces a non-linear transformation. A deviation value at the point of the median , its amplification weight is times. And a deviation value at the point of the 99th percentile , its amplification weight is times. This design achieves the effect that the more extreme the deviation, the more exponential growth of the amplification penalty it receives. It makes those deviated from the group abnormally, whose calculated distance is disproportionately enlarged, so as to be pushed further in the feature space and more easily identified as an abnormal point. This step is a secondary adjustment of distance calculation after judging the risk direction in step S2. The of step S2 solves the problem of whether to punish and the strength of the basic punishment, while the of step S3 further non-linearly increases the punishment according to the degree of extreme deviation. The two work together to form a correction system that can identify the risk direction and amplify extreme risks.
[0069] At this point, the non-linear ranking amplification optimization factor is obtained by non-linearly amplifying the intra-cluster deviation ranking of the multi-dimensional feature vector.
[0070] Step S4, the risk-aware asymmetric distance is obtained by fusing and evaluating the asymmetric direction optimization factor and the non-linear ranking amplification optimization factor.
[0071] After obtaining the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor, this step is for any dimension of any operation log data and any dimension of any cluster centroid; the square of the calculation result of subtracting the eigenvalue of the target dimension of the eigenvector of the log data from the target dimension of the cluster centroid is used as the first distance evaluation between the operation log data in the target dimension and the corresponding dimension of the cluster centroid; the calculation result of multiplying the asymmetric direction optimization factor and the nonlinear ranking amplification weight factor between the operation log data in the target dimension and the corresponding dimension of the cluster centroid is used as the distance optimization factor; the calculation result of multiplying the distance optimization factor and the first distance evaluation is used as the corresponding optimized distance; the calculation result of adding the optimized distances of all dimensions is used as the risk-aware asymmetric distance between the operation log data and the cluster centroid.
[0072] In one embodiment, the Operation log data and The calculation expression of the risk perception asymmetric distance between cluster centroids is:
[0073]
[0074] in, Indicates the Operation log data and The risk perception asymmetric distance between cluster centroids; The total number of dimensions of the feature vector representing the operation log data; Indicates the The operation log data and The centroid of the cluster is Asymmetric directional optimization factors in dimensions; Indicates the Operation log data for the The cluster class The nonlinear ranking amplification weight factor of each dimension; Indicates the Operation log data Feature data of dimensions; Indicates the The centroid of the cluster is feature data of dimensions.
[0075] At this point, the risk-aware asymmetric distance is obtained by integrating the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor.
[0076] Step S5: Obtain optimized clustering results through a clustering process based on risk-aware asymmetric distance, and perform abnormal behavior identification and processing based on the optimized clustering results.
[0077] After obtaining the risk perception asymmetric distance, continue to set the number of clusters of the clustering algorithm. In the embodiment of the present invention, the candidate range of the number of clusters is set to , the optimal number of clusters is selected in the candidate range by the silhouette coefficient. After obtaining the number of clusters of the clustering algorithm, the risk-aware asymmetric distance between the operation log data and the cluster centroid is used to perform cluster division in the clustering iteration process; the iterative process is completed and the optimized clustering result is obtained. Specifically, after determining the number of clusters ( value), and use this number of clusters to perform the final Clustering. The algorithm first randomly initializes The centroid of the cluster. Then, enter the iterative process: in the assignment step, for each data point , calculate the distance to all risk perceptions according to the above asymmetric distance formula The distance between the cluster centroids and the data point Assign it to the cluster with the smallest distance; in the update step, the centroid of each cluster is recalculated according to the new cluster member division and the calculation is prepared for the next iteration. The required set of intra-cluster deviation values. Repeat the allocation and update steps until the cluster allocation result no longer changes or the preset maximum number of iterations is reached, the algorithm converges, and clustering is completed;
[0078] After obtaining the optimized clustering results, abnormal behavior evaluation is continued in two stages. The first stage evaluates the number of cluster data points in the optimized clustering results and determines that sparse clusters are abnormal behavior patterns. The second stage identifies associated abnormal users based on user contribution and obtains users who are highly associated with abnormal behavior patterns.
[0079] Specifically, the first stage involves identifying abnormal behavior patterns based on cluster size. First, the present invention performs a scale analysis on all generated clusters. Abnormal events are sparse in frequency, while normal behavior constitutes the majority of the data. Therefore, an outlier cluster containing only a very small number of data points, representing a behavior pattern that is inherently highly suspected of being abnormal, is rare.
[0080] The second stage is the identification of highly associated abnormal users based on their contribution to the abnormal behavior patterns. After the sparse clusters representing abnormal behavior patterns are identified, the system further analyzes the composition of these clusters to locate the users that are highly associated with these abnormal behaviors. For each sparse cluster that is determined to represent an abnormal behavior pattern, the system counts the number of users whose operation log data points are in this cluster. Then, the system calculates the contribution of each user to this abnormal cluster. The contribution of a user U to a cluster C is quantified as the proportion of the total number of operations of user U that are in cluster C. Through this contribution analysis, the system can identify the users that frequently appear in one or more abnormal behavior pattern clusters. A user, if a large number of his / her operations are classified into different sparse clusters by the improved clustering algorithm of the present application, objectively indicates that the behavior pattern of this user is significantly deviated from all known normal behavior patterns.
[0081] Finally, the system identifies the users that are highly associated with the abnormal behavior patterns as the highest priority objects for security audit. The executable processes include but are not limited to: pushing the detailed abnormal behavior logs of these users to the security administrator for manual review; increasing the monitoring and auditing level of the subsequent operations of these users in the system; triggering stronger security measures such as secondary authentication when they perform certain critical operations.
[0082] The above embodiments are only used to illustrate the technical solutions of the present application, but not limit the present application; although the present application is described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement to part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A file collaborative processing method based on cloud computing, characterized in that: The cloud computing-based collaborative file processing method includes: Step S1: obtaining a multi-dimensional feature vector by performing feature mapping on the collaborative operation log of the backend of the cloud collaborative processing system server; Step S2: Obtain an asymmetric direction optimization factor by performing an asymmetric risk assessment on the dimensional distribution of multiple feature vectors; Step S3: Obtaining a nonlinear ranking amplification optimization factor by nonlinearly amplifying the intra-cluster deviation ranking of the multi-bit feature vector; Step S4: Obtaining the risk-aware asymmetric distance by fusion evaluation of the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor; Step S5: Obtaining optimized clustering results through a clustering process based on risk-aware asymmetric distance, and performing abnormal behavior identification and processing based on the optimized clustering results; The method of obtaining an asymmetric directional optimization factor by performing an asymmetric risk assessment on the dimensional distribution of a multi-bit feature vector includes: obtaining an asymmetric risk coefficient by performing random difference sampling analysis on the multi-dimensional feature vectors of the historical operation log; obtaining a directional deviation assessment by performing directional discrimination processing on the difference between the eigenvalues in the multi-bit feature vector and the cluster centroid; and obtaining an asymmetric directional optimization factor by performing a fusion assessment on the asymmetric risk coefficient and the directional deviation assessment. The method of obtaining a nonlinear ranking amplification optimization factor by nonlinearly amplifying the intra-cluster deviation ranking of a multi-bit eigenvector comprises: in each cluster assignment step of an iterative process of a clustering algorithm, for any target cluster, obtaining all member data points belonging to the target cluster, and forming a deviation value set by the absolute deviation values of the member data points relative to the cluster centroid in different dimensions; the absolute deviation value is the absolute value of the difference between the eigenvalue and the cluster centroid; for any member data point, obtaining a normalized deviation ranking of the member data point by calculating the percentile ranking of the deviation value of the member data point in the deviation value set; and using the square of the calculation result of adding the normalized deviation ranking to a constant 1 as a nonlinear ranking amplification weight factor; The risk-aware asymmetric distance is obtained by fusing and evaluating the asymmetric direction optimization factor and the nonlinear ranking amplification optimization factor, including: for any dimension of any operation log data and any dimension of any cluster centroid; the square of the calculation result of subtracting the eigenvalue of the target dimension of the eigenvector of the log data from the target dimension of the cluster centroid is used as the first distance evaluation between the operation log data and the dimension corresponding to the cluster centroid in the target dimension; the calculation result of multiplying the asymmetric direction optimization factor between the operation log data and the dimension corresponding to the cluster centroid in the target dimension and the nonlinear ranking amplification weight factor is used as the distance optimization factor; the calculation result of multiplying the distance optimization factor and the first distance evaluation is used as the corresponding optimized distance; and the calculation result of adding the optimized distances of all dimensions is used as the risk-aware asymmetric distance between the operation log data and the cluster centroid.
2. The method for collaborative file processing based on cloud computing according to claim 1, characterized in that: The method of obtaining a multi-dimensional feature vector by performing feature mapping on the collaborative operation log of the backend of the cloud collaborative processing system server includes: Data collection is performed on the backend of a server for collaborative processing in the cloud to obtain log data recording historical user operations; for each independent operation in the log, a multidimensional feature vector is generated; the first dimension of the multidimensional feature vector is the temporal sine feature of the log data; the second dimension is the temporal cosine feature of the log data; the third dimension is the operation encoding of the log data; and the fourth dimension is the number of bytes affected by the corresponding operation in the log data.
3. The method for collaborative file processing based on cloud computing according to claim 1, characterized in that: The asymmetric risk coefficient is obtained by performing random difference sampling analysis on the multi-dimensional feature vectors of historical operation logs, including: In the historical operation log, two historical operation data are randomly selected with replacement. The value of the subtraction of the eigenvalues of the eigenvectors of the two historical operation data in any dimension compared with 0 is used as the first dimension difference of the dimension; the absolute value of the subtraction of the eigenvalues of the eigenvectors of the two historical operation data in any dimension is used as the second dimension difference of the dimension; Set the number of random difference samplings; for any dimension, use the result of adding the first dimension differences of all random difference samplings as the numerator, and the result of adding the second dimension differences of all random difference samplings as the denominator, and use the result of the corresponding fraction as the asymmetric risk coefficient of the dimension.
4. The method for collaborative file processing based on cloud computing according to claim 1, characterized in that: The direction deviation assessment is obtained by performing direction discrimination processing on the difference between the eigenvalues in the multi-bit eigenvector and the cluster centroid, including: Set the unit step function. When the independent variable of the unit step function is greater than 0, the output value of the unit step function is 1. When the independent variable of the unit step function is less than or equal to 0, the output value of the unit step function is 0. Obtain all cluster centroids during the clustering process; for any dimension of the feature vector of any historical operation data, the result of subtracting the feature value of that dimension from the centroid of any cluster of that dimension is used as the first centroid difference evaluation; The first center of mass difference assessment is mapped through the unit step function, and the corresponding mapping result is used as the direction deviation assessment.
5. The method for collaborative file processing based on cloud computing according to claim 1, characterized in that: The asymmetric direction optimization factor is obtained by fusing the asymmetric risk coefficient with the direction deviation assessment, including: For any dimension of the characteristic vector of any historical operation data, the larger value between the settlement result of subtracting the asymmetric risk coefficient of the dimension from one half and the constant 0 is multiplied by the constant 2, and the corresponding calculation result is used as the first evaluation of direction optimization; the calculation result of multiplying the direction deviation evaluation by the first evaluation of direction optimization and adding the constant 1 is used as the asymmetric direction optimization factor.
6. The method for collaborative file processing based on cloud computing according to claim 1, characterized in that: The method of obtaining an optimized clustering result through a clustering process based on risk-aware asymmetric distance and identifying abnormal behavior based on the optimized clustering result includes: Set the number of clusters for the clustering algorithm; perform clustering during the clustering iteration process using the risk-aware asymmetric distance between the operation log data and the cluster centroids; complete the iteration process and obtain the optimized clustering results; Abnormal behavior evaluation is performed in two stages. The first stage evaluates by optimizing the number of cluster data points in the clustering results, and determines that sparse clusters are abnormal behavior patterns. The second stage identifies associated abnormal users based on user contribution, and obtains users who are highly associated with abnormal behavior patterns.
7. A file collaborative processing system based on cloud computing, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a cloud computing-based file collaborative processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Behavior risk detection method, clustering model construction method and device
CN114741673A
Product quantification method based on fuzzy clustering and asymmetric distance calculation
CN114780781A