Audit informatization processing method and system for enterprise data assets
By improving clustering algorithms and deep learning technology, abnormal analysis and risk assessment of user behavior data is solved, the problem of inefficiency of traditional audit methods is achieved, real-time monitoring and rapid response are achieved, and the efficiency and accuracy of data security control are improved.
Patent Information
- Application Number
- CN202510519919.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-04-24
AI Technical Summary
Traditional user behavior audit methods rely on manual monitoring and manual recording, which are inefficient and prone to omissions, making it difficult to meet the needs of modern enterprises for real-time monitoring and rapid response.
The improved clustering algorithm is used to cluster user behavior data, build a benchmark behavior model, and build a behavior recognition model based on deep learning technology, integrate the benchmark behavior model and behavior recognition model, perform abnormal analysis of user behavior data, identify user abnormal behavior, and generate personalized response strategies through multi-dimensional analysis and optimization algorithm.
It realizes accurate identification and risk assessment of user abnormal behavior, automatically triggers alarms and generates targeted response strategies, improves the efficiency of enterprise data security management and accuracy in dealing with data risks, and overcomes the limitations of traditional relying on manual monitoring.
Smart Images

Figure CN120046992A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data auditing, and more specifically, to an information-based processing method and system for auditing enterprise data assets. Background Art
[0002] With the deepening of digital transformation, enterprises generate a large amount of data assets in their daily operations. These data are not only an important basis for enterprise decision-making but also a core manifestation of its competitiveness. However, with the rapid increase in the amount of data, problems such as data leakage, misuse, and compliance risks have become increasingly prominent. To effectively protect enterprise data assets and ensure data security and compliance, user behavior auditing has become a crucial task. User behavior auditing mainly focuses on the operation records of users in the system, including behaviors such as accessing, modifying, and deleting data. By monitoring and analyzing user behaviors, enterprises can timely discover abnormal operations, identify potential security threats, and track the compliance of data usage. However, traditional user behavior auditing methods often rely on manual monitoring and manual recording, which are inefficient and prone to omissions, and it is difficult to meet the needs of modern enterprises for real-time monitoring and rapid response. Therefore, there is an urgent need to develop an information-based user behavior auditing processing method.
[0003] The patent application with the publication number CN112685768A discloses a data leakage prevention method and device based on software asset auditing. The method includes: in the acquisition management stage, configuring a data acquisition strategy; in the data acquisition stage, defining a data extraction strategy according to the data acquisition strategy and acquiring data; analyzing the access behaviors of users based on the acquired data and auditing security events of users; auditing data assets and analyzing the acquired data; customizing the report structure by a report engine and generating corresponding reports for various analyzed indicators. This invention can timely locate the source of potential hazards and perform alarm responses, greatly avoiding potential information security hazards that may occur in enterprises and ensuring the information security of enterprises.
[0004] However, although the above technology involves user behavior auditing, it does not elaborate on specific auditing methods in detail, lacking clear operation guidelines and monitoring frameworks, which makes it face great complexity and uncertainty when enterprises implement user behavior auditing methods. This lack of systematic and standardized implementation methods cannot effectively prevent potential security threats, resulting in the formation of information security blind spots and reducing the accuracy and efficiency of abnormal behavior identification.
[0005] In view of this, the present invention proposes an information-based processing method and system for auditing enterprise data assets to solve the above problems. Summary of the Invention
[0006] To overcome the above defects of the prior art and achieve the above object, the present invention provides the following technical solutions: An audit informatization processing method for enterprise data assets, including: Collect user behavior data; Use the improved clustering algorithm to cluster the user behavior data, cluster the user behavior data into multiple benchmark sets, and construct corresponding benchmark behavior models according to the user behavior data in each benchmark set; Construct a behavior recognition model based on deep learning technology, fuse the benchmark behavior model and the behavior recognition model, perform anomaly analysis on the user behavior data, and identify user abnormal behaviors; Conduct multi-dimensional analysis on user abnormal behaviors, calculate the corresponding anomaly scores, and evaluate the corresponding risk levels based on a preset level mapping table; Automatically trigger an alarm according to the risk level, and use an optimization algorithm to iteratively analyze user abnormal behaviors and anomaly scores, quantify the response values corresponding to different response strategies during each iterative analysis process, and generate and execute personalized response strategies based on the iterative analysis results.
[0007] Further, the user behavior data includes historical behavior data and real-time behavior data; the historical behavior data is user operation data collected at historical moments, and the real-time behavior data is user operation data collected in real time; the user operation data includes user accounts, operation objects, operation types, and operation times; The steps of constructing a benchmark behavior model include: Step S101: Take the user operation data corresponding to the same user account in the historical behavior data as a group of user sets, and the user sets correspond one-to-one with the user accounts; Step S102: Set different digital labels for different operation objects and mark them as object labels, and set different digital labels for different operation types and mark them as type labels; replace the operation objects in each group of user sets with the corresponding object labels, and replace the operation types in each group of user sets with the corresponding type labels; Step S103: According to the operation time in each group of user sets, divide each group of user sets into three groups of time sets; divide the data of the same type in each group of time sets into a group of classification sets; Step S104: Set the number of classifications ; Step S105: Perform clustering processing on each group of classification sets in turn according to the number of classifications to obtain the corresponding clustering sets; where each clustering set includes groups of data sets; Step S106: Let , , The number of data for user operations; Step S107: Loop through steps S105 to S106 until the loop ends and proceed to step S108; Step S108: Calculate the average distance corresponding to each group of clustering sets, and among all the clustering sets corresponding to each group of time sets, select the clustering set with the largest average distance as the optimal set; Step S109: Count the number of data within each data set in each group of optimal sets and mark it as the benchmark quantity; Based on the benchmark quantity, filter out the benchmark sets from the data sets within each group of optimal sets; Step S110: Construct a benchmark behavior model for the corresponding user according to the benchmark set corresponding to each user account.
[0008] Furthermore, in the step S103, the method of dividing each group of user sets into three groups of time sets is as follows: Obtain the operation time in each group of user sets, and divide each group of user sets into three groups of time sets according to the dates corresponding to the operation time. The three groups of time sets respectively correspond to weekdays, weekends, and holidays; In the step S105, the methods of performing clustering processing on each group of classification sets are all the same; Among them, the steps of performing clustering processing on a classification set corresponding to an object label include: Step S201: Take each object label in the classification set as a sample point, and the sample points correspond one-to-one with the object labels; Step S202: Randomly select sample points as the center points, and sequentially incrementally label each center point as , ; Mark the sample points that are not used as center points as assignment points, and sequentially incrementally label each assignment point as , , is the number of object labels in the classification set; Step S203: Establish corresponding groups of data sets based on the center points, and sequentially calculate the point distances from each assignment point to each center point; Step S204: Compare the point distances from the assignment point to each center point, and assign the assignment point to the data set corresponding to the center point with the smallest point distance; Step S205: Let ; Step S206: Loop through steps S204 to S205 until the loop ends and proceed to step S207; Step S207: Recalculate the new center point corresponding to each data set; Step S208: Loop through steps S203 to S207 until the new center points of each data set recalculated in step S207 are the same as the new center points calculated in the previous loop. Then the loop ends, and obtain the data sets and their corresponding assigned points, which are used as the clustering sets.
[0009] Further, in step S203, the expression for the point distance is: ; where is the point distance from the assigned point to the center point ; In step S207, the calculation method for the new center point of each data set includes: ; where is the new center point corresponding to the th data set, is the th assigned point in the th data set, is the number of assigned points in the th data set, ; In step S108, the method for calculating the average distance corresponding to each clustering set includes: Calculate the distance coefficient corresponding to each sample point in each clustering set, add up the distance coefficients corresponding to each clustering set in sequence, and then divide by to obtain the average distance corresponding to each clustering set; the expression for the distance coefficient is: ; where is the distance coefficient of the th sample point, is the external distance of the th sample point, is the internal distance of the th sample point, is the maximum value function, ; The calculation method for the internal distance of the th sample point is: Mark the data set corresponding to the th sample point as the current set, and mark all sample points in the current set except the th sample point as other points; Calculate the point distance from the th sample point to each other point and mark it as the within-cluster distance; Add up each within-cluster distance in sequence and then divide by the number of within-cluster distances to obtain the Internal distance of a sample point; The calculation method of the external distance of the th sample point is as follows: calculate the point distance from the th sample point to each center point, and mark it as the center distance; sort each center point in descending order according to the corresponding center distance, mark the data set corresponding to the second-ranked center point as the nearest set, and mark all the sample points in the nearest set as neighbor points; calculate the point distance from the th sample point to each neighbor point, and mark it as the out-of-cluster distance; add up each out-of-cluster distance in turn, and then divide by the number of out-of-cluster distances to obtain the external distance of the
[0010] th sample point. Further, in the step S109, the method for screening the reference set from the data sets in each optimal set is as follows: preset a quantity threshold, and compare the reference quantity corresponding to each data set with the quantity threshold respectively; if the reference quantity is greater than the quantity threshold, mark the corresponding data set as the reference set; if the reference quantity is less than or equal to the quantity threshold, do not mark the corresponding data set; In the step S110, the method for constructing a reference behavior model for the user includes:
[0011] Further, the method for identifying the abnormal behavior of the user includes: Mark the user accounts in the real-time behavior data as identified accounts, and obtain the behavior data of the identified accounts for the current day. The behavior data for the current day is the user operation data collected at historical moments within the current day. Take the user operation data corresponding to the same identified account in the behavior data for the current day and the real-time behavior data as an analysis set. Each analysis set corresponds to one identified account. Count the quantity of each operation object in the analysis set and mark it as the object quantity. Count the quantity of each operation type in the analysis set and mark it as the type quantity. Compare the operation times in the analysis set, mark the earliest operation time as the extremely early time of the current day, mark the latest operation time as the extremely late time of the current day, and take the extremely early time and the extremely late time of the current day as the operation period. Take the baseline behavior model of the user corresponding to the identified account as the analysis behavior model, and take the baseline period, baseline object frequency, and baseline type frequency in the analysis behavior model as baseline data. Take the object quantity, type quantity, operation period, and baseline data as analysis data, and input the analysis data into the trained behavior recognition model to identify the corresponding behavior label. The behavior label is the digital label corresponding to the user's abnormal behavior, and different digital labels correspond to different user abnormal behaviors. Obtain the corresponding user abnormal behavior according to the behavior label. The behavior recognition model is a deep neural network model.
[0012] Further, the method for evaluating the risk level includes: Take the object quantity, type quantity, and operation period as evaluation data. According to the user's abnormal behavior, obtain the corresponding data in the evaluation data and mark it as abnormal data. According to the abnormal data, obtain the corresponding data in the baseline data and mark it as standard data. If the abnormal data is the object quantity, subtract the corresponding baseline object frequency in the baseline data from the corresponding object quantity and take the absolute value to obtain the object difference. If the abnormal data is the type quantity, subtract the corresponding baseline type frequency in the baseline data from the corresponding type quantity and take the absolute value to obtain the type difference. If the abnormal data is the operation period, subtract the extremely early time of the corresponding baseline period from the extremely early time of the current day of the corresponding operation period and take the absolute value to obtain the extremely early difference. Subtract the extremely late time of the corresponding baseline period from the extremely late time of the current day of the corresponding operation period and take the absolute value to obtain the extremely late difference. Add the extremely early difference and the extremely late difference to obtain the period difference. A preset weight set, which includes a first set, a second set, and a third set; the first set includes weight coefficients corresponding to different operation objects, the second set includes weight coefficients corresponding to different operation types, and the third set includes weight coefficients corresponding to operation time; according to the operation object corresponding to the object difference, obtain the corresponding weight coefficient from the first set and mark it as the first coefficient; according to the operation type corresponding to the type difference, obtain the corresponding weight coefficient from the second set and mark it as the second coefficient; according to the time period difference, obtain the weight coefficient in the third set and mark it as the third coefficient; multiply each object difference by the corresponding first coefficient to obtain the first score; multiply each type difference by the corresponding second coefficient to obtain the second score; multiply the time period difference by the third coefficient to obtain the third score; add up all the first scores, second scores, and third scores in sequence to obtain the anomaly score. A preset level mapping table, which includes anomaly score ranges corresponding to different risk levels; the risk levels include high risk, medium risk, and low risk; mark the calculated anomaly score as the calculated score, and compare the calculated score with the anomaly score ranges corresponding to different risk levels in the level mapping table respectively to obtain the risk level corresponding to the anomaly score range where the calculated score is located.
[0013] Furthermore, the steps of generating a response strategy include: Step S301: Set different digital tags for different response strategies and mark them as strategy tags; Step S302: Initialize the population, where the population includes n individuals, the individual positions correspond one-to-one with the strategy tags, and the initial iteration number is 0; Step S303: Define a response value function and an iteration threshold ; Step S304: Determine the optimal individual in the population; Step S305: Determine the update method for each individual according to the optimal individual and perform the update; Step S306: Judge whether ; if so, let , and return to Step S304; if not, enter Step S307; Step S307: Obtain the response strategy corresponding to the strategy tag of the optimal individual; In the said Step S302, the expression of the individual position is: ; in the formula, is the position of the i-th individual, is the random coefficient of the i-th individual, , , is the number of strategy tags; Define the speed range value ; Define the corresponding individual speed for each individual. The expression of the individual speed is: ; In the step S303, the expression of the response value function is: ; In the formula, is the response value, is the security; The calculation method of the security is: Take the behavior label, the anomaly score, and the policy label corresponding to the individual as the calculation data, input the calculation data into the trained security analysis model, and calculate the corresponding security; The security analysis model is a deep neural network model; In the step S304, the method for determining the optimal individual in the population is: Calculate the response value of each individual in the population, sort all the response values from large to small, and take the individual corresponding to the first-ranked response value as the optimal individual.
[0014] Further, in the step S305, the method for determining the update method of each individual includes: Preset a response threshold, subtract the response value of the optimal individual from the response value of each individual respectively, and take the absolute value to obtain the response difference; Compare each response difference with the response threshold respectively, mark the individual corresponding to the response difference greater than or equal to the response threshold as the exploration individual, and mark the individual corresponding to the response difference less than the response threshold as the pursuit individual; Among them, the update method corresponding to the exploration individual is exploration update, and the update method corresponding to the pursuit individual is pursuit update; The method for each individual to be updated includes: When the exploration individual is updated, the calculation method for updating the individual speed includes: ; In the formula, is the individual speed after the i-th individual is updated, is the individual speed before the i-th individual is updated, is the position of the i-th individual, is the position of the optimal individual, is a random number in the standard normal distribution, are all random numbers in; The calculation method for updating the individual position includes: ; In the formula, is the position after the i-th individual is updated, is the position of the global optimal individual, and the global optimal individual is the individual with the largest response value in the entire iteration process; When the pursuit individual is updated, the calculation method for updating the individual speed includes: ; In the formula, is a random number in The calculation method for updating the individual position includes: ; In the formula, is the position of the i-th individual before update in the previous iteration process, is a random number in
[0015] An audit information processing system for enterprise data assets, implementing the described audit information processing method for enterprise data assets, includes: A data collection module for collecting user behavior data; A benchmark modeling module that uses an improved clustering algorithm to cluster user behavior data, clusters the user behavior data into multiple benchmark sets, and constructs corresponding benchmark behavior models based on the user behavior data in each benchmark set; A behavior analysis module that constructs a behavior recognition model based on deep learning technology, fuses the benchmark behavior model and the behavior recognition model, performs anomaly analysis on user behavior data, and identifies user abnormal behaviors; A risk assessment module for performing multi-dimensional analysis on user abnormal behaviors, calculating corresponding anomaly scores, and evaluating the corresponding risk levels based on a preset level mapping table; An alarm response module that automatically triggers an alarm according to the risk level, and uses an optimization algorithm to perform iterative analysis on user abnormal behaviors and anomaly scores, quantifies the response values corresponding to different response strategies during each iterative analysis process, and generates and executes personalized response strategies based on the iterative analysis results.
[0016] The technical effects and advantages of the audit information processing method and system for enterprise data assets of the present invention: By collecting user behavior data and constructing a benchmark behavior model, it can accurately identify user abnormal behaviors and timely discover potential data security risks; moreover, it can perform risk assessment on the identified user abnormal behaviors, automatically trigger an alarm and generate targeted response strategies, which not only improves the efficiency of enterprise data security management and control, but also enhances the accuracy of dealing with data risks, realizing rapid response and precise prevention and control; using intelligent algorithms to achieve automated and intelligent data auditing, overcoming the limitations of traditional manual monitoring, can effectively protect the security of enterprise data assets, and has important practical significance and application value for enterprise digital transformation and information management. Description of the Drawings
[0017] Figure 1Schematic diagram of an audit information processing system for enterprise data assets in Embodiment 1 of the present invention; Figure 2 Flowchart of the method for constructing a benchmark behavior model in Embodiment 1 of the present invention; Figure 3 Flowchart of an audit information processing method for enterprise data assets in Embodiment 2 of the present invention. Detailed implementation manners
[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0019] Embodiment 1 Please refer to Figure 1 As shown, an audit information processing system for enterprise data assets in this embodiment includes a data collection module, a benchmark modeling module, a behavior analysis module, a risk assessment module, and an alarm response module; each module is connected by wired and / or wireless means to realize data transmission between modules.
[0020] The data collection module is used to collect user behavior data.
[0021] The user behavior data includes historical behavior data and real-time behavior data; the historical behavior data is the user operation data collected at historical moments, and the real-time behavior data is the user operation data collected in real time; the user operation data includes user accounts, operation objects, operation types, and operation times; the user behavior data is obtained through the enterprise's database management system; The user account is a credential for uniquely identifying each user, usually composed of a user name or a user ID, and is used to track all operation behaviors of a specific user in the system, so as to ensure clear responsibilities; the operation object is the data in the enterprise data assets (such as customer data, financial data, employee information, etc.) operated by the user, which helps to analyze the usage of enterprise data assets and identify potential security risks; the operation type is the operation behavior executed by the user, such as viewing, modifying, deleting, uploading, downloading, etc., which helps to monitor and manage user behaviors to identify abnormal behaviors; the operation time is the time when the user performs each operation, represented in the form of a time stamp, which is used to track the occurrence time of the operation and facilitate audit traceability and behavior analysis.
[0022] The benchmark modeling module performs clustering processing on the user behavior data by using an improved clustering algorithm, clusters the user behavior data into multiple benchmark sets, and constructs corresponding benchmark behavior models according to the user behavior data in each benchmark set.
[0023] As Figure 2 shown, the steps to build a benchmark behavior model include: Step S101: Take the user operation data corresponding to the same user account in the historical behavior data as a set of user collections, and the user collections correspond one-to-one with the user accounts; Step S102: Set different digital tags for different operation objects and mark them as object tags, and set different digital tags for different operation types and mark them as type tags; Replace the operation objects in each group of user collections with the corresponding object tags, and replace the operation types in each group of user collections with the corresponding type tags; Step S103: According to the operation time in each group of user collections, divide each group of user collections into three groups of time collections; Divide the data of the same category in each group of time collections into a group of classification collections, and the classification collections correspond one-to-one with the data in the corresponding time collections; Step S104: Set the number of classifications ; Step S105: Perform clustering processing on each group of classification collections in turn according to the number of classifications to obtain the corresponding clustering collections, and the clustering collections correspond one-to-one with the classification collections; Among them, each clustering collection includes groups of data collections; Step S106: Let , , be the number of user operation data; Step S107: Loop steps S105 to S106 until the loop ends and proceeds to step S108; Step S108: Calculate the average distance corresponding to each group of clustering collections, and take the clustering collection with the largest average distance among all the clustering collections corresponding to each group of time collections as the optimal collection; Step S109: Count the number of data in each group of data collections in each group of optimal collections and mark it as the benchmark quantity; According to the benchmark quantity, screen out the benchmark collections from the data collections in each group of optimal collections; Step S110: Build a benchmark behavior model for the corresponding user according to the benchmark collection corresponding to each user account.
[0024] In the above step S103, the method of dividing each group of user collections into three groups of time collections includes: Obtain the operation time in each group of user collections, and divide each group of user collections into three groups of time collections according to the date corresponding to the operation time. The three groups of time collections correspond to weekdays, weekends, and holidays respectively.
[0025] In the above step S105, the method for clustering each group of classification sets is the same; among them, the steps for clustering a classification set corresponding to an object label include: Step S201: Each object label in the classification set is used as a sample point, and the sample points correspond one-to-one with the object labels; Step S202: Randomly select sample points as the center points, and sequentially incrementally label each center point as , ; that is, label the first center point as , label the second center point as , and label the th center point as ; label the sample points that are not used as center points as distribution points, and sequentially incrementally label each distribution point as , , is the number of object labels in the classification set; that is, label the first distribution point as , label the second distribution point as , and label the th center point as ; Step S203: Establish corresponding groups of data sets according to the center points, and sequentially calculate the point distance from each distribution point to each center point; Step S204: Compare the point distances from the distribution point to each center point, and assign the distribution point to the data set corresponding to the center point with the smallest point distance; Step S205: Let ; Step S206: Loop steps S204 to S205 until when the loop ends, and enter step S207; Step S207: Recalculate the new center point corresponding to each data set; Step S208: Loop steps S203 to S207 until the new center points of each data set recalculated in step S207 are the same as the new center points calculated in the previous loop, then the loop ends, and obtain groups of data sets and the corresponding distribution points, and use them as the clustering set.
[0026] In the above step S203, the expression of the point distance is: ; in the formula, is the distribution point to the center point Point distance.
[0027] In the above step S207, the calculation method of the new center point of each data set includes: ; In the formula, is the new center point corresponding to the th data set, is the th assignment point in the th data set, is the number of assignment points in the th data set, .
[0028] In the above step S108, the method for calculating the average distance corresponding to each clustering set includes: Calculate the distance coefficient corresponding to each sample point in each clustering set, add up the distance coefficients corresponding to each clustering set in sequence, and then divide by to obtain the average distance corresponding to each clustering set; the expression of the distance coefficient is: ; in the formula, is the distance coefficient of the th sample point, is the external distance of the th sample point, is the internal distance of the th sample point, is the maximum value function, .
[0029] The calculation method of the internal distance of the th sample point is: Mark the data set corresponding to the th sample point as the current set, and mark all sample points in the current set except the th sample point as other points; calculate the point distance from the th sample point to each other point and mark it as the intra-cluster distance; add up each intra-cluster distance in sequence, and then divide by the number of intra-cluster distances to obtain the internal distance of the th sample point.
[0030] The calculation method of the external distance of the th sample point is: Calculate the point distance from the th sample point to each center point and mark it as the center distance; sort each center point in descending order according to the corresponding center distance, mark the data set corresponding to the center point ranked second as the nearest set, and mark all sample points in the nearest set as neighbor points; calculate the The point distances from each sample point to each neighbor point are marked as out-of-cluster distances; the out-of-cluster distances are added up in sequence and then divided by the number of out-of-cluster distances to obtain the external distance of the th sample point.
[0031] In the above step S109, the method for screening out the reference set from the data sets within each group of optimal sets is as follows: a preset quantity threshold is set, and the quantity threshold is preset by those skilled in the art according to the actual situation; the reference quantities corresponding to each data set are respectively compared with the quantity threshold; if the reference quantity is greater than the quantity threshold, the corresponding data set is marked as the reference set; if the reference quantity is less than or equal to the quantity threshold, the corresponding data set is not marked.
[0032] In the above step S110, the method for the user to construct the reference behavior model includes: Mark the reference set corresponding to the object label as the object set, mark the reference set corresponding to the type label as the type set, and mark the reference set corresponding to the operation time as the time period set; use the operation object corresponding to the object label in each object set as the reference object of the corresponding time set; use the operation type corresponding to the type label in each type set as the reference type of the corresponding time set; compare the operation times in each time period set, mark the earliest operation time as the extremely early time, and mark the latest operation time as the extremely late time; use the extremely early time and the extremely late time of each time period set as the reference time period of the corresponding time set; use the reference quantity corresponding to each object set as the reference object frequency of the corresponding time set; use the reference quantity corresponding to each type set as the reference type frequency of the corresponding time set; construct the reference behavior model of the corresponding user according to the reference time period, reference object frequency and reference type frequency corresponding to each user's corresponding time set.
[0033] Exemplarily, the reference behavior model of user A is: Time set: weekdays; Reference time period: 9:00 - 17:00; Reference object frequency: Data A: 10 times a day, Data B: 15 times a day; Reference type frequency: Upload: 8 times a day, Download: 20 times a day; Time set: weekends; Reference time period: 14:00 - 16:00; Reference object frequency: Data C: 5 times a day, Data D: 3 times a day; Reference type frequency: View: 7 times a day, Delete: 2 times a day.
[0034] The behavior analysis module constructs a behavior recognition model based on deep learning technology, integrates the benchmark behavior model and the behavior recognition model, performs anomaly analysis on user behavior data, and identifies user abnormal behaviors.
[0035] The methods for identifying user abnormal behaviors include: Mark the user account in the real-time behavior data as the identified account, and obtain the behavior data of the identified account on the same day. The behavior data on the same day is the user operation data collected at historical moments within the same day; regard the user operation data corresponding to the same identified account in the behavior data on the same day and the real-time behavior data as a set of analysis data, and each set of analysis data corresponds one-to-one to the identified account; count the quantity corresponding to each operation object in the set of analysis data and mark it as the object quantity; count the quantity corresponding to each operation type in the set of analysis data and mark it as the type quantity; compare the operation times in the set of analysis data, mark the earliest operation time as the extremely early time of the day, mark the latest operation time as the extremely late time of the day, and use the extremely early time and the extremely late time of the day as the operation period; Regard the benchmark behavior model of the user corresponding to the identified account as the analysis behavior model, and use the benchmark period, benchmark object frequency, and benchmark type frequency in the analysis behavior model as the benchmark data; use the object quantity, type quantity, operation period, and benchmark data as the analysis data, input the analysis data into the trained behavior recognition model, and identify the corresponding behavior label; the behavior label is the digital label corresponding to the user abnormal behavior, and different user abnormal behaviors correspond to different digital labels. User abnormal behaviors include, for example, large-scale data download, excessive operation frequency of Data A, accessing Data B during non-benchmark periods, etc.; according to the behavior label, obtain the corresponding user abnormal behavior; The training process of the behavior recognition model includes: Pre-collect a set of analysis data, for each set of analysis data, set the corresponding behavior label, n is an integer greater than 1. Convert the analysis data and the corresponding behavior label into a corresponding set of feature vectors; the behavior label corresponding to the analysis data is collected by those skilled in the art during the process of historically identifying user abnormal behaviors. Pre-collect n sets of analysis data, analyze each set of analysis data in turn in combination with the actual situation, determine the user abnormal behavior corresponding to each set of analysis data, and set the corresponding behavior label for each set of analysis data in turn; Take each set of feature vectors as the input of the behavior recognition model. The behavior recognition model takes a set of predicted behavior labels corresponding to each set of analysis data as the output, uses the actual behavior label corresponding to each set of analysis data as the prediction target, and the actual behavior label is the behavior label preset corresponding to the analysis data; use minimizing the sum of the prediction errors of all analysis data as the training target; among them, the calculation formula for the prediction error is , where is the prediction error, is the group number of the eigenvector corresponding to the analysis data, is the prediction behavior label corresponding to the th group of analysis data, and is the actual behavior label corresponding to the th group of analysis data; the behavior recognition model is trained until the sum of the prediction errors converges and then the training stops.
[0036] Specifically, the above behavior recognition model is a deep neural network model; it includes an input layer, a hidden layer, and an output layer; each hidden layer includes multiple neurons, and there are connections between each neuron and the neurons in the next layer. The connections contain weights that determine the importance and influence of data transmission in the neural network; an activation function is applied to each neuron between the hidden layer and the output layer. The activation function introduces non-linearity and allows the network to learn more complex patterns and features.
[0037] A risk assessment module is used to perform multi-dimensional analysis on user abnormal behaviors, calculate corresponding abnormal scores, and evaluate corresponding risk levels based on a preset level mapping table.
[0038] The methods for evaluating risk levels include: Taking the number of objects, the number of types, and the operation period as evaluation data; according to the user's abnormal behavior, obtaining the corresponding data in the evaluation data and marking it as abnormal data; for example, if the user's abnormal behavior is large-scale data download, obtaining the number of types corresponding to the operation type of download, and if the user's abnormal behavior is excessive operation frequency of data A, obtaining the number of objects corresponding to the operation object of data A; according to the abnormal data, obtaining the corresponding data in the reference data and marking it as standard data; if the abnormal data is the number of objects, subtracting the reference object frequency corresponding to it in the reference data and taking the absolute value to obtain the object difference; if the abnormal data is the number of types, subtracting the reference type frequency corresponding to it in the reference data and taking the absolute value to obtain the type difference; if the abnormal data is the operation period, subtracting the extremely early time of the corresponding day of the operation period from the extremely early time of the corresponding reference period and taking the absolute value to obtain the extremely early difference; subtracting the extremely late time of the corresponding day of the operation period from the extremely late time of the corresponding reference period and taking the absolute value to obtain the extremely late difference; adding the extremely early difference and the extremely late difference to obtain the period difference.
[0039] A preset weight set, the weight set includes a first set, a second set and a third set, and the weight set is preset by those skilled in the art according to the actual situation; the first set includes weight coefficients corresponding to different operation objects, the second set includes weight coefficients corresponding to different operation types, and the third set includes weight coefficients corresponding to operation time; according to the operation object corresponding to the object difference, obtain the corresponding weight coefficient from the first set and mark it as the first coefficient; according to the operation type corresponding to the type difference, obtain the corresponding weight coefficient from the second set and mark it as the second coefficient; according to the time period difference, obtain the weight coefficient in the third set and mark it as the third coefficient; multiply each object difference by the corresponding first coefficient to obtain the first score; multiply each type difference by the corresponding second coefficient to obtain the second score; multiply the time period difference by the third coefficient to obtain the third score; add all the first scores, second scores and third scores in sequence to obtain the anomaly score.
[0040] A preset level mapping table, the level mapping table includes anomaly score ranges corresponding to different risk levels, and the level mapping table is preset by those skilled in the art according to the actual situation; the risk levels include high risk, medium risk and low risk; mark the calculated anomaly score as the calculated score, and compare the calculated score with the anomaly score ranges corresponding to different risk levels in the level mapping table respectively to obtain the risk level corresponding to the anomaly score range where the calculated score is located.
[0041] An alarm response module, which automatically triggers an alarm according to the risk level, and uses an optimization algorithm to iteratively analyze the user's abnormal behavior and the anomaly score. During each iterative analysis process, quantify the response values corresponding to different response strategies, and generate and execute personalized response strategies based on the iterative analysis results.
[0042] When the alarm is triggered, the system automatically generates alarm information, which includes the identified account, the user's abnormal behavior and the risk level, and sends the alarm information to the mobile devices of relevant enterprise personnel in the form of email or text message; the response strategies are, for example, temporarily disabling the account, reducing the access permission, freezing the user operation, isolating the network access, etc.
[0043] The steps of generating the response strategy include: Step S301: Set different digital tags for different response strategies and mark them as strategy tags; Step S302: Initialize the population, the population includes n individuals, the individual positions correspond one-to-one with the strategy tags, and the initial iteration number is 0; Step S303: Define the response value function and the iteration threshold ; Step S304: Determine the optimal individual in the population; Step S305: Determine the update method for each individual according to the optimal individual and perform the update; Step S306: Judge whether ; if so, let , and return to Step S304; if not, proceed to Step S307; Step S307: Obtain the response strategy corresponding to the policy label of the optimal individual.
[0044] In the above Step S302, the expression for the individual position is: ; in the formula, is the position of the i-th individual, is the random coefficient of the i-th individual, , , is the number of policy labels; Define the speed range value , and the speed range value is preset by those skilled in the art according to the actual situation; define the corresponding individual speed for each individual, and the expression for the individual speed is: .
[0045] In the above Step S303, the expression for the response value function is: ; in the formula, is the response value, is the security; the calculation method of security is: use the behavior label, anomaly score, and the policy label corresponding to the individual as calculation data, input the calculation data into the trained security analysis model, and calculate the corresponding security; the training process of the security analysis model is the same as that of the behavior recognition model, and both are deep neural network models; the security corresponding to the calculation data is collected by those skilled in the art when generating response strategies historically groups of calculation data, under the conditions of the behavior label and anomaly score in each group of calculation data, adopt the response strategy corresponding to the policy label in the corresponding calculation data for risk management, and analyze the security after risk management, and groups of calculation data are sequentially set with corresponding security; the iteration threshold is preset by those skilled in the art according to the actual situation.
[0046] In the above Step S304, the method for determining the optimal individual in the population is: calculate the response value of each individual in the population, sort all the response values from largest to smallest, and take the individual corresponding to the first-ranked response value as the optimal individual.
[0047] In the above Step S305, the method for determining the update method for each individual includes: A preset response threshold, which is preset by those skilled in the art according to the actual situation; subtract the response value of each individual from the response value of the optimal individual respectively, and take the absolute value to obtain the response difference; compare each response difference with the response threshold respectively, mark the individuals corresponding to the response difference greater than or equal to the response threshold as exploration individuals, and mark the individuals corresponding to the response difference less than the response threshold as pursuit individuals; among them, the update method corresponding to the exploration individuals is exploration update, and the update method corresponding to the pursuit individuals is pursuit update.
[0048] The method for updating each individual includes: When updating the exploration individuals, the calculation method for updating the individual velocity includes: ; In the formula, is the individual velocity of the i-th individual after update, is the individual velocity of the i-th individual before update, is the position of the i-th individual, is the position of the optimal individual, is a random number in the standard normal distribution, are all random numbers in.
[0049] The calculation method for updating the individual position includes: ; In the formula, is the position of the i-th individual after update, is the position of the global optimal individual, and the global optimal individual is the individual with the largest response value in the whole iteration process.
[0050] When updating the pursuit individuals, the calculation method for updating the individual velocity includes: ; In the formula, is a random number in.
[0051] The calculation method for updating the individual position includes: ; In the formula, is the position of the i-th individual before update in the previous iteration process, is a random number in.
[0052] In this embodiment, by collecting user behavior data and constructing a baseline behavior model, it is possible to accurately identify abnormal user behaviors and timely discover potential data security risks. Moreover, it can conduct risk assessments on the identified abnormal user behaviors, automatically trigger alarms, and generate targeted response strategies, which not only improves the efficiency of enterprise data security management and control but also enhances the accuracy in dealing with data risks, achieving rapid response and precise prevention and control. The use of intelligent algorithms to achieve automated and intelligent data auditing overcomes the limitations of traditional manual monitoring and can effectively protect the security of enterprise data assets, having important practical significance and application value for the digital transformation and information management of enterprises.
[0053] Embodiment 2 Please refer to Figure 3 As shown, for the parts not described in detail in this embodiment, refer to the description in Embodiment 1. A method for auditing information processing of enterprise data assets is provided, and the method includes: Collect user behavior data; Use the improved clustering algorithm to cluster the user behavior data, cluster the user behavior data into multiple baseline sets, and construct corresponding baseline behavior models according to the user behavior data in each baseline set; Based on deep learning technology, construct a behavior recognition model, fuse the baseline behavior model and the behavior recognition model, conduct anomaly analysis on the user behavior data, and identify abnormal user behaviors; Conduct multi-dimensional analysis on the abnormal user behaviors, calculate the corresponding anomaly scores, and evaluate the corresponding risk levels based on a preset level mapping table; Automatically trigger an alarm according to the risk level, and use an optimization algorithm to conduct iterative analysis on the abnormal user behaviors and anomaly scores. During each iterative analysis process, quantify the response values corresponding to different response strategies, and generate and execute personalized response strategies based on the iterative analysis results.
[0054] Embodiment 3 This application also provides an electronic device. The electronic device may include one or more processors and one or more memories. Among them, computer-readable code is stored in the memory, and when the computer-readable code is run by one or more processors, it can execute a method for auditing information processing of enterprise data assets as described above.
[0055] The method or system according to the embodiments of the present application can also be implemented by means of the architecture of the electronic device shown in the present application. The electronic device may include a bus, one or more CPUs, a ROM, a RAM, a communication port connected to a network, an input / output, a hard disk, etc. The storage device in the electronic device, such as a ROM or a hard disk, can store a method for auditing informatization processing of an enterprise data asset provided by the present application. Further, the electronic device may further include a user interface. Of course, the architecture shown in the present application is only exemplary. When implementing different devices, one or more components shown in the electronic device of the present application can be omitted according to actual needs.
[0056] Embodiment 4 One embodiment of the present application discloses a computer-readable storage medium. Computer-readable instructions are stored on the computer-readable storage medium. When the computer-readable instructions are run by a processor, a method for auditing informatization processing of an enterprise data asset according to the embodiments of the present application as described with reference to the above drawings can be executed. The storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and cache memory, etc. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc.
[0057] In addition, according to the embodiments of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the present application provides a non-transitory machine-readable storage medium storing machine-readable instructions that can be run by a processor to execute instructions corresponding to the method steps provided by the present application, such as: a method for auditing informatization processing of an enterprise data asset. When the computer program is executed by a central processing unit (CPU), the above functions defined in the method of the present application are executed.
[0058] As described above, the above are only specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
[0059] Finally: The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should all be included within the protection scope of the present invention.
Claims
1. A method for processing enterprise data assets through audit informatization, characterized in that: include: Collect user behavior data; The improved clustering algorithm is used to cluster the user behavior data, the user behavior data is clustered into multiple benchmark sets, and a corresponding benchmark behavior model is constructed according to the user behavior data in each benchmark set; Build a behavior recognition model based on deep learning technology, integrate the baseline behavior model and the behavior recognition model, perform abnormal analysis on user behavior data, and identify abnormal user behavior; Perform multi-dimensional analysis on abnormal user behaviors, calculate corresponding abnormal scores, and evaluate corresponding risk levels based on a preset level mapping table; Alarms are automatically triggered based on risk levels, and optimization algorithms are used to iteratively analyze abnormal user behaviors and abnormal scores. During each iterative analysis, the response values corresponding to different response strategies are quantified, and personalized response strategies are generated and executed based on the iterative analysis results.
2. The audit information processing method for enterprise data assets according to claim 1 is characterized in that: The user behavior data includes historical behavior data and real-time behavior data; historical behavior data is user operation data collected at historical moments, and real-time behavior data is user operation data collected in real time; User operation data includes user account, operation object, operation type, and operation time; The steps to build a baseline behavioral model include: Step S101: user operation data corresponding to the same user account in the historical behavior data is taken as a group of user sets, where the user sets correspond to the user accounts one-to-one; Step S102: different digital labels are set for different operation objects and marked as object labels, and different digital labels are set for different operation types and marked as type labels; the operation objects in each user set are replaced with corresponding object labels, and the operation types in each user set are replaced with corresponding type labels; Step S103: Divide each user set into three time sets according to the operation time in each user set; divide the same type of data in each time set into a classification set; Step S104: Setting the number of categories ; Step S105: cluster each group of classification sets in turn according to the number of classifications to obtain the corresponding cluster sets; wherein each cluster set includes Group data collection; Step S106: , , The amount of data operated for the user; Step S107: loop through steps S105 to S106 until When the loop ends, the process goes to step S108; Step S108: Calculate the average distance corresponding to each group of cluster sets, and take the cluster set with the largest average distance among all cluster sets corresponding to each group of time sets as the optimal set; Step S109: Count the number of data in each data set in each optimal set and mark it as a benchmark number; and select a benchmark set from the data sets in each optimal set according to the benchmark number; Step S110: constructing a benchmark behavior model for each user account according to the benchmark set corresponding to the user account.
3. The audit information processing method for enterprise data assets according to claim 2 is characterized in that: In step S103, the method of dividing each user set into three time sets is: obtaining the operation time in each user set, and dividing each user set into three time sets according to the date corresponding to the operation time, wherein the three time sets correspond to working days, weekends and holidays respectively; In step S105, the method of clustering each group of classification sets is consistent; wherein the step of clustering a classification set corresponding to an object label includes: Step S201: each object label in the classification set is used as a sample point, and the sample point corresponds to the object label one by one; Step S202: Random selection sample points as the center point, and mark each center point in ascending order. , ; Mark the sample points that are not the center points as allocation points, and mark each allocation point as , , is the number of object labels in the classification set; Step S203: According to The corresponding center point is established Group the data sets and calculate the point distance from each assigned point to each center point in turn; Step S204: Assign points The distance to each center point is compared and the points are assigned Assign to the data set corresponding to the center point with the minimum point distance; Step S205: ; Step S206: loop through steps S204 to S205 until When the loop ends, the process goes to step S207; Step S207: recalculate the new center point corresponding to each data set; Step S208: loop through steps S203 to S207 until the new center point of each data set recalculated in step S207 is consistent with the new center point calculated in the previous loop, the loop ends, and the A data set and the corresponding distribution points are used as cluster sets.
4. The audit information processing method for enterprise data assets according to claim 3 is characterized in that: In step S203, the expression of point distance is: ; In the formula, For allocation point To center point The point distance; In step S207, the method for calculating the new center point of each data set includes: ; In the formula, For the The data set corresponds to the new center point, For the In the data set distribution points, For the The number of distribution points in the data set, ; In step S108, the method for calculating the average distance corresponding to each cluster set includes: Calculate the distance coefficient corresponding to each sample point in each cluster set, add the distance coefficients corresponding to each cluster set in turn, and then divide by , get the average distance corresponding to each cluster set; the expression of the distance coefficient is: ; In the formula, For the The distance coefficient of the sample points, For the The external distance of the sample points, For the The internal distance of the sample points, is the maximum value function, ; No. The calculation method of the internal distance of the sample points is: The data set corresponding to the sample point is marked as the current set, and the data set except the All sample points other than the sample points are marked as other points; calculate the The point distance from the sample point to each other point is marked as the intra-cluster distance; each intra-cluster distance is added in turn, and then divided by the number of intra-cluster distances to obtain the The internal distance of sample points; No. The calculation method of the external distance of the sample points is: The point distance from each sample point to each center point is marked as the center distance; each center point is sorted from large to small according to the corresponding center distance, and the data set corresponding to the second center point is marked as the nearest set, and the sample points in the nearest set are marked as neighbor points; calculate the The point distance from each sample point to each neighbor point is marked as the out-of-cluster distance; each out-of-cluster distance is added in turn, and then divided by the number of out-of-cluster distances to obtain the The outer distance of the sample points.
5. The audit information processing method for enterprise data assets according to claim 4 is characterized in that: In step S109, the method of selecting the reference set from the data sets in each optimal set is: presetting a quantity threshold, and comparing the reference quantity corresponding to each data set with the quantity threshold respectively; If the number of benchmarks is greater than the number threshold, the corresponding data set is marked as a benchmark set; If the benchmark quantity is less than or equal to the quantity threshold, the corresponding data set is not marked; In step S110, the method for constructing a baseline behavior model for the user includes: Mark the benchmark set corresponding to the object label as the object set, mark the benchmark set corresponding to the type label as the type set, and mark the benchmark set corresponding to the operation time as the time period set; use the operation object corresponding to the object label in each object set as the benchmark object of the corresponding time set; use the operation type corresponding to the type label in each type set as the benchmark type of the corresponding time set; compare the operation time in each time period set, mark the earliest operation time as the extremely early time, and mark the latest operation time as the extremely late time; use the extremely early time and the extremely late time of each time period set as the benchmark time period of the corresponding time set; use the benchmark quantity corresponding to each object set as the benchmark object frequency of the corresponding time set; use the benchmark quantity corresponding to each type set as the benchmark type frequency of the corresponding time set; construct the benchmark behavior model of the corresponding user according to the benchmark time period, benchmark object frequency and benchmark type frequency corresponding to the corresponding time set of each user.
6. The audit information processing method for enterprise data assets according to claim 5 is characterized in that: The method for identifying abnormal user behavior includes: Mark the user account in the real-time behavior data as the identification account, and obtain the behavior data of the day corresponding to the identification account, where the behavior data of the day is the user operation data collected at historical moments within the day; take the user operation data corresponding to the same identification account in the behavior data of the day and the real-time behavior data as a set of analysis sets, and the analysis sets correspond to the identification accounts one by one; count the number of each operation object in the analysis set and mark it as the number of objects; count the number of each operation type in the analysis set and mark it as the number of types; compare the operation times in the analysis set, mark the earliest operation time as the very early time of the day, mark the latest operation time as the very late time of the day, and take the very early time of the day and the very late time of the day as the operation time period; The baseline behavior model of the user corresponding to the identified account is used as the analysis behavior model, and the baseline time period, baseline object frequency and baseline type frequency in the analysis behavior model are used as baseline data; the number of objects, number of types, operation time period and baseline data are used as analysis data, and the analysis data is input into the trained behavior recognition model to identify the corresponding behavior label; the behavior label is a digital label corresponding to the user's abnormal behavior, and different user abnormal behaviors have different digital labels; according to the behavior label, the corresponding user abnormal behavior is obtained; the behavior recognition model is a deep neural network model.
7. The audit information processing method for enterprise data assets according to claim 6 is characterized in that: Methods for assessing risk levels include: The number of objects, the number of types and the operation period are used as evaluation data; according to the abnormal behavior of the user, the corresponding data in the evaluation data is obtained and marked as abnormal data; according to the abnormal data, the corresponding data in the benchmark data is obtained and marked as standard data; if the abnormal data is the number of objects, the corresponding number of objects is subtracted from the corresponding benchmark object frequency in the benchmark data, and the absolute value is taken to obtain the object difference; if the abnormal data is the number of types, the corresponding number of types is subtracted from the corresponding benchmark type frequency in the benchmark data, and the absolute value is taken to obtain the type difference; if the abnormal data is the operation period, the extremely early time of the corresponding operation period is subtracted from the extremely early time of the corresponding benchmark period, and the absolute value is taken to obtain the extremely early difference; the extremely late time of the corresponding operation period is subtracted from the extremely late time of the corresponding benchmark period, and the absolute value is taken to obtain the extremely late difference; the extremely early difference is added to the extremely late difference to obtain the period difference; A preset weight set, the weight set includes a first set, a second set and a third set; the first set includes weight coefficients corresponding to different operation objects, the second set includes weight coefficients corresponding to different operation types, and the third set includes weight coefficients corresponding to operation time; according to the operation object corresponding to the object difference, the corresponding weight coefficient is obtained from the first set and marked as the first coefficient; according to the operation type corresponding to the type difference, the corresponding weight coefficient is obtained from the second set and marked as the second coefficient; according to the time period difference, the weight coefficient in the third set is obtained and marked as the third coefficient; each object difference is multiplied by the corresponding first coefficient to obtain a first score; each type difference is multiplied by the corresponding second coefficient to obtain a second score; the time period difference is multiplied by the third coefficient to obtain a third score; all first scores, second scores and third scores are added in sequence to obtain an abnormal score; A preset level mapping table includes abnormal score segments corresponding to different risk levels; the risk levels include high risk, medium risk and low risk; the calculated abnormal score is marked as the calculated score, and the calculated score is compared with the abnormal score segments corresponding to different risk levels in the level mapping table to obtain the risk level corresponding to the abnormal score segment in which the calculated score is located.
8. The audit information processing method for enterprise data assets according to claim 7 is characterized in that: The steps to generate a response strategy include: Step S301: setting different digital labels for different response strategies and marking them as strategy labels; Step S302: Initialize the population, which includes n individuals, where the individual positions correspond to the strategy labels one by one, and the initial number of iterations is 0; Step S303: Define the response value function and iteration threshold ; Step S304: determine the best individual in the population; Step S305: According to the optimal individual, determine the update method of each individual and perform the update; Step S306: Determine whether If so, then , and return to step S304; if not, proceed to step S307; Step S307: Obtain the response strategy corresponding to the optimal individual corresponding strategy label; In step S302, the expression of the individual position is: ; In the formula, is the position of the ith individual, is the random coefficient of the ith individual, , , is the number of strategy tags; Define speed range value ; The corresponding individual speed is defined for each individual, and the expression of individual speed is: ; In step S303, the expression of the response value function is: ; In the formula, is the response value, The security is calculated by taking the behavior label, the anomaly score and the strategy label corresponding to the individual as the calculation data, inputting the calculation data into the trained security analysis model, and calculating the corresponding security; the security analysis model is a deep neural network model; In step S304, the method for determining the best individual in the population is: calculating the response value of each individual in the population, sorting all the response values from large to small, and taking the individual corresponding to the first response value as the best individual.
9. The audit information processing method for enterprise data assets according to claim 8 is characterized in that: In step S305, the method for determining the update mode of each individual includes: A response threshold is preset, and the response value of each individual is subtracted from the response value of the optimal individual, and the absolute value is taken to obtain the response difference; each response difference is compared with the response threshold, and the individuals corresponding to the response difference greater than or equal to the response threshold are marked as exploration individuals, and the individuals corresponding to the response difference less than the response threshold are marked as pursuit individuals; among which, the update mode corresponding to the exploration individual is the exploration update, and the update mode corresponding to the pursuit individual is the pursuit update; The methods for each individual to update include: When the exploration individual is updated, the calculation method of the update individual speed includes: ; In the formula, is the updated individual speed of the ith individual, is the individual speed before the update of the ith individual, is the position of the ith individual, is the position of the optimal individual, is a random number from a standard normal distribution, Both Random numbers in ; The calculation method for updating individual positions includes: ; In the formula, is the updated position of the i-th individual, is the position of the global optimal individual, which is the individual with the largest response value in the entire iteration process; When the pursuit individual is updated, the calculation method of the update individual speed includes: ; In the formula, for Random numbers in ; The calculation method for updating individual positions includes: ; In the formula, is the position of the i-th individual before update in the last iteration, for The random number in .
10. An enterprise data asset audit information processing system, used to implement the enterprise data asset audit information processing method according to any one of claims 1 to 9, characterized in that: include: Data collection module, used to collect user behavior data; The benchmark modeling module uses an improved clustering algorithm to cluster the user behavior data, cluster the user behavior data into multiple benchmark sets, and build a corresponding benchmark behavior model based on the user behavior data in each benchmark set; The behavior analysis module builds a behavior recognition model based on deep learning technology, integrates the baseline behavior model and the behavior recognition model, performs abnormal analysis on user behavior data, and identifies abnormal user behavior; The risk assessment module is used to perform multi-dimensional analysis on abnormal user behaviors, calculate the corresponding abnormal scores, and assess the corresponding risk levels based on a preset level mapping table; The alarm response module automatically triggers alarms according to the risk level, and uses optimization algorithms to iteratively analyze user abnormal behaviors and abnormal scores. In each iterative analysis process, the response values corresponding to different response strategies are quantified, and personalized response strategies are generated and executed based on the iterative analysis results.
Citation Information
Patent Citations
Data leakage prevention method and device based on software asset auditing
CN112685768A
Intrusion detection system and method based on intelligent network
CN118413406A
Authentication analysis system and method based on big data
CN118608167A
Server security access monitoring method based on Internet of Things
CN119071049A
Task package permission allocation method and device, computer equipment and storage medium
CN119128875A
Cited By
Digital asset security assessment method based on block chain
CN120781390A
A blockchain-based digital asset security assessment method
CN120781390B