Account sharing violation quantity analysis method and device, equipment and medium
By training the model based on machine learning methods and predicting the number of future account sharing violations in enterprises, the problem of enterprises' lack of accurate analysis on account sharing issues is solved, and accurate prediction and effective control of account sharing violation risks are achieved, thereby reducing the losses of enterprises.
Patent Information
- Application Number
- CN202510661255.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-26
AI Technical Summary
In existing technologies, companies lack accurate data analysis and prediction on account sharing issues, resulting in serious risks of account sharing violations, information leakage and network vulnerabilities. Companies also do not pay enough attention to this, resulting in user privacy data being exposed and losing user trust.
Through machine learning-based methods, we obtain the target organization's account sharing violation control data, train the machine learning model, predict the number of account sharing violations in a certain period in the future, and use the machine learning model's training sample data set to optimize the model until the test error reaches the preset value to achieve accurate risk prediction.
It achieves accurate prediction of corporate account sharing violation risks, helps companies avoid or reduce unnecessary losses in advance, and improves the effectiveness of security management measures.
Smart Images

Figure CN120705829A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the fields of artificial intelligence and financial technology, and in particular to a method, apparatus, device and storage medium for analyzing the number of account sharing violations based on machine learning. Background Art
[0002] Account sharing is a serious and difficult issue for businesses to resolve completely, leading to a range of issues such as unauthorized access, information leaks, and network vulnerabilities. With regulators increasingly protecting sensitive personal user information, the annual costs incurred by various industries are substantial. For example, in the field of fintech, account sharing is a common issue across many sectors, including but not limited to banking, insurance, and government agencies, encompassing interrelated departments, institutions, and agencies. Account sharing can be encountered in numerous application scenarios, including but not limited to payment, authentication, notarization, and banking transactions. These scenarios all raise issues of image privacy and security. These application scenarios are typically implemented by corresponding business systems, including but not limited to shopping, insurance, banking, transaction, and order systems.
[0003] Currently, many companies or departments only implement localized controls on user monitoring, but lack accurate data analysis and forecasting for related departments, institutions, and agents. Instead, they simply impose post-event penalties, treating the symptoms rather than the root cause. Many companies even neglect account sharing violations, potentially leading to significant risks. In today's increasingly internet-enabled world, this means that the privacy of internet users will be exposed online, leading to the loss of trust between these companies and the government and their users. Summary of the Invention
[0004] The present application provides a method, apparatus, device and storage medium for analyzing the number of account sharing violations based on machine learning, which is used to solve the technical problem that traditional enterprises may have a large number of account sharing violation risks.
[0005] A method for analyzing the number of account sharing violations based on machine learning, the method comprising: Obtaining account sharing violation control data of a target organization, the control data including a plurality of different attributes describing account sharing violation control measures of the target organization; Inputting the control data into a pre-trained machine learning model to predict the number of account sharing violations within a certain period of time within the target organization; The machine learning model is pre-trained in the following way: Obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each of the control training sample data including multiple different attributes describing account sharing violation control measures of a sample organization; The initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
[0006] A device for analyzing the number of account sharing violations based on machine learning, the device comprising: an acquisition module, configured to acquire account sharing violation control data of a target organization, wherein the control data includes a plurality of different attributes describing account sharing violation control measures of the target organization; A prediction module, configured to input the control data into a pre-trained machine learning model to predict the number of account sharing violations committed by the target organization within a certain period of time in the future; Among them, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
[0007] A computer device comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the steps of analyzing the number of account sharing violations based on machine learning as described in any of the above items are implemented.
[0008] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of analyzing the number of account sharing violations based on machine learning as described in any of the above items.
[0009] It can be seen that a solution for analyzing the number of account sharing violations based on machine learning is provided. In this solution, a control training sample data set is first obtained, and the control training sample data set includes multiple control training sample data, and each control training sample data includes a variety of different attributes that describe the account sharing violation control measures of the sample organization; the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, and a machine learning model for predicting the number of account sharing violations in a certain period in the future is obtained, and the control data of the target organization's account sharing violations is obtained, and the control data includes a variety of different attributes that describe the account sharing violation control measures of the target organization; the control data is input into the pre-trained machine learning model to predict the number of account sharing violations of the target organization in a certain period in the future. It can be seen that the purpose of the embodiment of the present application is to conduct model training based on historical data such as security control measures for user account violation risks. By learning from historical data, the number of account sharing violations in the target organization in a certain period of time in the future can be predicted, thereby predicting what measures in the company's long-term security control measures will be executed within a certain period to obtain the number of account risk violations that the company wants, thereby effectively reducing unnecessary losses to the company due to account risk violations. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0011] Figure 1 This is a schematic diagram of a processing framework of a method for analyzing the number of account sharing violations based on machine learning in one embodiment of the present application; Figure 2 This is a flow chart of a device for analyzing the number of account sharing violations based on machine learning in one embodiment of the present application; Figure 3 This is a schematic diagram of a comparison result between the output value of a linear regression model and the actual value in a method for analyzing the number of account sharing violations based on machine learning in one embodiment of the present application; Figure 4 This is a schematic diagram of a training process of a model in a method for analyzing the number of account sharing violations based on machine learning in one embodiment of the present application; Figure 5 This is a structural diagram of a device for analyzing the number of account sharing violations based on machine learning in one embodiment of the present application; Figure 6 It is a structural diagram of a computer device in one embodiment of the present application. DETAILED DESCRIPTION
[0012] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0013] In the field of financial technology, many application scenarios involve account sharing. These include, but are not limited to, payment, authentication, and notarization, all of which may involve account identification and verification, leading to account sharing. Banking and insurance transactions may also involve account identification and verification. All of these application scenarios raise privacy and security issues related to account sharing. These application scenarios are typically implemented by corresponding business systems, including but not limited to shopping, insurance, banking, transaction, and order systems. Account sharing is a common technology for convenience. However, account sharing is a serious and difficult issue for businesses to fully resolve. Account sharing can lead to a range of issues, including unauthorized access, information leakage, and network vulnerabilities. Currently, many businesses only implement limited user monitoring and control measures, lacking accurate data analysis and prediction for subordinate departments, agencies, and agents, resulting in significant information leakage.
[0014] like Figure 1 As shown, Figure 1This is a framework diagram of a method for analyzing the number of account sharing violations based on machine learning provided by the present application. A control training sample data set will be obtained first, and the control training sample data set will include multiple control training sample data, and each control training sample data will include multiple different attributes describing the account sharing violation control measures of the sample organization; based on the control training sample data, an initial machine learning model will be trained until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, and a machine learning model for predicting the number of account sharing violations in a certain period in the future is obtained, and the control data of the account sharing violations of the target organization is obtained, and the control data includes multiple different attributes describing the account sharing violation control measures of the target organization; the control data is input into the pre-trained machine learning model to predict the number of account sharing violations of the target organization in a certain period in the future. It can be seen that the purpose of the embodiment of the present application is mainly to perform model training based on historical data such as security control measures for user account violation risks. By learning from historical data, the number of account sharing violations in the target organization in a certain period of time in the future can be predicted, thereby predicting which measures in the company's long-term security control measures will be executed within a certain period to obtain the number of account risk violations that the company wants, thereby avoiding or reducing unnecessary losses to the company due to future account risk violations in advance.
[0015] It should be noted that the above method may be implemented through a server, and the server may be implemented by an independent server or a server cluster composed of multiple servers, without specific limitation.
[0016] In one embodiment, taking application to a server as an example, Figure 2 As shown, a method for analyzing the number of account sharing violations based on machine learning is provided, and the method mainly includes the following steps: S10. Acquire account sharing violation control data of the target organization, where the control data includes multiple different attributes describing the account sharing violation control measures of the target organization; S20. Inputting the control data into a pre-trained machine learning model to predict the number of account sharing violations in the target organization within a certain period of time in the future; In this embodiment, it is first necessary to train a machine learning model that can be used to predict the number of account sharing violations in the target organization within a certain period in the future. The machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each of the control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value. The above-mentioned preset error level or the number of training rounds reaching the preset value can be set based on experience.
[0017] The attributes of the account sharing violation control measures represent a certain control measure for account sharing violations, and the various different attributes of the account sharing violation control measures represent multiple different control measures for account sharing violations. For example, in Enterprise D, there are 9 values related to the control measures for account sharing violation risks. The first 8 are used to describe the various attributes related to the control measures, i.e., xi in the machine learning model; the last value is the median number of people of this type of violation to be predicted, i.e., yi in the model. Therefore, in the embodiment of the present application, the machine learning model can be represented as Y, where Y represents the prediction result of the machine learning model. After the machine learning model is established, an optimization target needs to be given to the machine learning model so that the learned model parameters can make the predicted value as close to the true value as possible.
[0018] After obtaining the trained machine learning model, the control data can be input into the pre-trained machine learning model to predict the number of account sharing violations of the target organization in a certain period in the future. The target organization can include any organizational unit such as an enterprise, institution or government unit. For example, in order to realize the number of account sharing violations of the target organization (such as an upstream organization or a subordinate department, agency or agency of a superior organization) in a certain period in the future, for example, the number of account sharing violations can be the number of account sharing violations per month, the number of account sharing violations per quarter or the number of account sharing violations per year, etc., depending on the training process, and there is no specific limitation. This embodiment obtains the control data of the target organization's account sharing violations and inputs it into the trained machine learning model for prediction.
[0019] It can be seen that the embodiment of the present application provides a method for analyzing the number of account sharing violations based on machine learning, which will perform model training based on historical data such as security control measures for user account violation risks. By learning from historical data, the number of account sharing violations in the target organization in a certain period of time in the future can be predicted, thereby predicting what measures in the company's long-term security control measures will be executed within a certain period to obtain the number of account risk violations that the company wants, thereby effectively giving the company an accurate prediction of the risks that may be discovered, so that the company can make rectifications, which is conducive to effectively reducing unnecessary losses to the company due to account risk violations.
[0020] In one embodiment, the various attributes of the account sharing violation control measures include: The first attribute: whether a comprehensive account security management system has been established; Second attribute: whether the account security management system has been promoted during the target period; The third attribute is whether safety system education is provided to all new employees; The fourth attribute: whether the account system intercepts any illegal account sharing activities; The fifth attribute is whether account sharing violations are monitored after the violation occurs. The sixth attribute is whether the current account monitoring model is regularly optimized. An account monitoring model refers to the management model used by an enterprise or organization to monitor accounts, and can be an artificial intelligence model. Are there any follow-up actions, including warnings, dismissal, or penalties, for individuals who violate account sharing regulations? The seventh attribute: Is there any follow-up action for those who violate the account sharing rules? Attribute 8: Whether complaints and false positives regarding account sharing violations are regularly removed to optimize the rules for monitoring account sharing violations.
[0021] It should be noted that the above target period can be one month, one quarter, one year or other time periods, without specific limitation.
[0022] That is, when training a machine learning model, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including the above eight attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value. When the model is applied, the control data of the target organization's account sharing violation is obtained, the control data including the above eight different attributes describing the account sharing violation control measures of the target organization; and then inputting the control data into the pre-trained machine learning model to predict the number of account sharing violations of the target organization in a certain period in the future.
[0023] For example, the above eight attributes can be shown in the following table: Table 1 As can be seen from Table 1, in this embodiment, the above eight attributes are all descriptions of yes or no, so the corresponding field types are discrete values, and illustratively, can be represented as discrete values (0, 1).
[0024] In this embodiment, the various different attributes of the account sharing violation control measures are clarified to ensure the feasibility of the solution. Moreover, since they are discrete values, they can be uniformly normalized to discrete values (0, 1), that is, the training data is preprocessed to reduce the data processing pressure of training and improve the timeliness of model training.
[0025] In one embodiment, in combination with the above embodiment, in order to make the prediction of the machine learning model more accurate, the various attributes of the account sharing violation control measures include, in addition to the above eight attributes, the following: Ninth attribute: the ratio of the number of people promoted during the target period to the total number of monitored people; Tenth attribute: The ratio of the number of people who violated account sharing rules during the target period to the total number of people monitored. Eleventh attribute: the ratio of the number of people who filed complaints about account sharing violations to the total number of monitored people during the target period.
[0026] That is, when training a machine learning model, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including the above eleven attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value. When the model is applied, the control data of the target organization's account sharing violation is obtained, the control data including the above eleven different attributes describing the account sharing violation control measures of the target organization; and then inputting the control data into the pre-trained machine learning model to predict the number of account sharing violations of the target organization in a certain period in the future.
[0027] For example, the above eleven attributes may be as shown in the following table: Table 2 An examination of the data in Table 2 reveals that among the eleven attributes in Table 2, three are continuous values and eight are discrete values. Although discrete values are often represented using discrete numbers like 0, 1, and 2, their meaning differs from continuous values because the difference here has no additional meaning. For example, if 0, 1, and 2 represent red, green, and blue, respectively, this does not mean that "blue" and "red" are farther apart than "green" and "red." Therefore, for a discrete attribute with d possible values, this embodiment typically converts them into d binary attributes, each taking the value 0 or 1, or maps each possible value into a multidimensional vector. However, in this case, since discrete values are inherently binary attributes, this effort is omitted. Furthermore, for non-discrete attributes, the value ranges of the various attributes can vary significantly. For example, the value ranges of the tenth and eleventh attributes can differ significantly. Therefore, in this embodiment, the ninth, tenth, and eleventh attributes are attribute data that has undergone normalization. The goal of normalization is to reduce the value range of each attribute to a similar range. For example, as an example, a common operation method is to subtract the mean and then divide by the original value range.
[0028] This embodiment clearly defines the various attributes of the account-sharing violation control measures, encompassing a total of eleven attribute data types, ensuring the feasibility of the solution. Furthermore, since both discrete and continuous values are included, they can be normalized to discrete values, preprocessing the training data to reduce data processing pressure and improve training timeliness. Furthermore, the inclusion of multiple additional attributes that characterize the risk of account-sharing violations can improve the model's prediction of the number of account-sharing violations within the target organization within a certain period of time, thereby enhancing prediction accuracy.
[0029] It should be noted that in the embodiments of the present application, the machine learning model may include multiple types. As an example, the machine learning model includes a linear regression model. In this embodiment, a control training sample data set of size n is given, where each control training sample data i has d attributes, respectively denoted as xi1, xi2, ..., xid, and yi is the target variable to be predicted for the control training sample data, that is, the number of account sharing violations within a certain period in the future. Incorporating it into the linear regression model, it is assumed that the target yi can be described by a linear combination of attributes, xij is the various attributes describing the control measure i, and yi is the number of violations.
[0030] like Figure 3 As shown, Figure 3 This is one of the methods in the present application to control the training sample data set to perform a linear regression model. type A schematic diagram comparing the predicted value and the actual value, wherein the horizontal axis of each point represents the median number of account-sharing violations, the vertical axis represents the predicted value of the linear regression model, and the horizontal axis represents the actual value, that is, the actual number of account-sharing violations. When the two values are completely equal, they will fall on the dotted line. Therefore, the more accurate the linear regression model's prediction, the closer the point is to the dotted line. Therefore, a machine learning model in the embodiment of the present application can use a linear regression model.
[0031] In one embodiment, if Figure 4 As shown, taking the machine learning model as a linear regression model as an example, the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including the following steps: a1. Initialize the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b1. Training the initialized linear regression model based on the control training sample data to calculate the training error of the initialized linear regression model. In this embodiment, a fixed loss function is used to calculate the training error, including but not limited to a mean squared error (MSE) loss function. The specific type of loss function is not specifically limited. c1. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d1. Determine whether the test error of the linear regression model reaches the preset error level or the number of training rounds reaches the preset value. Repeat steps b1 to c1 until the test error of the initialized linear regression model reaches the preset error level or the number of training rounds reaches the preset value.
[0032] In this embodiment, the control training sample data in the control training sample dataset is historical data. In this embodiment, assume a dataset of size n, where each sample i has d attributes, denoted as xi1, xi2, ..., xid, and yi is the target variable to be predicted for that sample. The linear regression model assumes that the target yi can be described by a linear combination of attributes. Specifically, in the account sharing violation prediction problem, xij represents the various attributes describing control measure i, and yi represents the number of violators (which can be the number of account risk violations per month, quarter, or year, etc.). First, a control training sample data set is obtained. The control training sample data set includes multiple control training sample data. Each control training sample data includes multiple different attributes describing the account sharing violation control measures of the sample organization, such as the eleven attributes mentioned in Table 2. An initialized linear regression model is constructed, and the model network parameters of the initial linear regression model are initialized using a random initialization method. The model network parameters include weights w and bias b. Exemplarily, the weights w and bias b can be initialized using a random initialization method, for example, using a normal distribution with a mean of 0 and a variance of 1 for initialization, without specific limitation. Then, the initialized linear regression model is trained based on the control training sample data to calculate the training error of the initialized linear regression model. The initialized linear regression model is back-propagated based on the training error to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model, and the model network parameters are updated. The training and parameter transfer process is repeated until the test error of the initialized linear regression model reaches a preset error level or the number of training rounds reaches a preset value.
[0033] It is worth noting that, in one embodiment, a control sample data set is obtained, and the control sample data set includes multiple control sample data, each of which includes multiple different attributes describing the account sharing violation control measures of the sample organization; the control sample data set is then divided into a control training sample data set and a control test sample data set. The control training sample data set is used to adjust the parameters of the model, that is, to train the linear regression model, and the error of the linear regression model on this data set is called the training error; the control test sample data set is used for testing, and the error of the linear regression model on this data set is called the test error. Exemplarily, the split ratio of the control training sample data set and the control test sample data set can be a ratio of 2:8, which is not specifically limited.
[0034] The purpose of training a linear regression model is to predict unknown new data by finding patterns in the training data. Therefore, test error is a more effective indicator of model performance. The test error in step d1 is calculated by inputting the controlled test sample data set into the linear regression model. This is done until the test error of the initialized linear regression model reaches the preset error level or the number of training rounds reaches the preset value. The same principle applies to other machine learning models.
[0035] In another embodiment, the machine learning model includes a linear regression model, and the initial machine learning model is trained based on the controlled training sample data until a test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including: a2. Initializing the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b2. Training the initialized linear regression model based on the control training sample data and the first loss function to calculate the training error of the initialized linear regression model; c2. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d2. Determine, based on the training error, whether the data aggregation degree output by the linear regression model deviates from a preset aggregation degree; e2. When the answer is yes, determine an adapted second loss function based on the cause of the data aggregation deviation, and replace the first loss function in step b2. Repeat steps b2 to c2 after replacement until the test error of the initialized linear regression model reaches a preset error level or the number of training rounds reaches a preset value.
[0036] In this embodiment, similarly, a control training sample data set is first obtained, and the control training sample data set includes multiple control training sample data. Each control training sample data includes multiple different attributes that describe the account sharing violation control measures of the sample organization, such as the eleven attributes mentioned in Table 2, and an initialized linear regression model is constructed. The model network parameters of the initial linear regression model are initialized using a random initialization method. The model network parameters include weights w and bias b. Exemplarily, the weights w and bias b can be initialized using a random initialization method, for example, using a normal distribution with a mean of 0 and a variance of 1 for initialization, without specific limitation. Then, the initialized linear regression model is trained based on the control training sample data, and the training error of the initialized linear regression model is calculated using the first loss function. The initialized linear regression model is subjected to reverse error propagation based on the training error, so that the training error of the initialized linear regression model is sequentially transferred to the input layer of the initialized linear regression model, and the model network parameters are updated, and the training and parameter transfer processes are repeated continuously. Different from the previous embodiment, in this embodiment, during the training process, it is determined based on the training error whether the data aggregation degree (or data aggregation property) of the model output of the linear regression model deviates from the preset aggregation degree, wherein the data aggregation degree refers to the trend or distribution characteristics after integrating multiple prediction results. If yes, an adapted second loss function is determined based on the cause of the deviation of the data aggregation degree, and the first loss function in step b2 is replaced. The replaced steps b2 to c2 are repeated until the test error of the initialized linear regression model reaches a preset error level or the number of training rounds reaches a preset value.
[0037] Similarly, in one embodiment, a control sample data set is obtained, and the control sample data set includes multiple control sample data, each of which includes multiple different attributes describing the account sharing violation control measures of the sample organization; the control sample data set is then divided into a control training sample data set and a control test sample data set. The control training sample data set is used to adjust the parameters of the model, that is, to train the linear regression model, and the error of the linear regression model on this data set is called the training error; the control test sample data set is used for testing, and the error of the linear regression model on this data set is called the test error. Exemplarily, the split ratio of the control training sample data set and the control test sample data set can be a ratio of 2:8, which is not specifically limited.
[0038] It can be seen that a second training method is provided in the embodiment of the present application. The first training method does not require fine-tuning of the loss function, reduces training complexity, and can improve the model training speed when the data is good enough. Compared with the first training method, during the training process of the second training method, based on the training error, when judging whether the data aggregation degree of the model output of the linear regression model deviates from the preset aggregation degree, an adapted second loss function will be determined according to the cause of the deviation of the data aggregation degree, and the loss function of the training process will be replaced, that is: during the model training process, if the data aggregation deviates slightly from expectations, the model can be fine-tuned by adjusting the loss function to make it better adapt to the data characteristics of the training data, thereby accelerating model convergence and improving model accuracy.
[0039] In one embodiment, determining the adapted second loss function according to the cause of the data aggregation deviation includes: When the reason for the deviation of the data aggregation degree is the presence of outliers in the control training sample data, determining the weighted loss function or the Huber loss function as the adapted second loss function; When the reason for the deviation of the data aggregation degree is the uneven distribution of the control training sample data, the quantile loss function is determined as the adapted second loss function.
[0040] Depending on the specific circumstances of data aggregation deviating from expectations, you can choose an appropriate loss function fine-tuning method. For outlier issues, the Huber loss function can be used. For uneven data distribution, the quantile loss function can be used. For dynamic adjustment of the loss function, an adaptive loss function can be used. It should be noted that in practical applications, the effectiveness of different loss functions can be verified experimentally and the optimal solution can be selected based on model evaluation metrics (such as mean squared error (MSE), mean absolute error (MAE), and coefficient of determination (R²).
[0041] It should be noted that in addition to the weighted loss function, Huber loss function, or quantile loss function, the loss function can also be fine-tuned according to other specific situations where the data aggregation deviates from the expectation, and there is no specific limitation.
[0042] For example, in other embodiments, if the deviation of data aggregation from expectations is due to the asymmetry or skewness of the data distribution, it is also possible to consider using a loss function based on the data distribution. For example, negative log-likelihood loss (Negative Log-Likelihood Loss) is used for replacement training. For another example, in other embodiments, an adaptive loss function can be directly used to dynamically adjust the learning parameters of the adaptive loss function according to the degree of deviation of the data aggregation. For example, a learnable parameter can be introduced to adjust the loss function. The learnable parameter is mainly used to balance the weights of the mean square error (MSE) and the absolute value loss. When the data aggregation deviates from expectations and is difficult to adjust through a fixed loss function, an adaptive parameter can be introduced to dynamically adjust the loss function. This ensures the smooth progress of training.
[0043] It can be seen that in this embodiment, during the model training process, a suitable loss function fine-tuning method can be selected according to the specific situation in which the data aggregation deviates from the expectation, thereby ensuring the quality and speed of training.
[0044] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0045] In one embodiment, an image processing device is provided, which corresponds one-to-one to an image processing device method in the above embodiment. Figure 5 As shown, the image processing device includes an acquisition module 101 and a prediction module 102. The functional modules are described in detail as follows: An acquisition module 101 is configured to acquire account sharing violation control data of a target organization, wherein the control data includes a plurality of different attributes describing account sharing violation control measures of the target organization; Prediction module 102, configured to input the control data into a pre-trained machine learning model to predict the number of account sharing violations committed by the target organization within a certain period of time in the future; Among them, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
[0046] In conjunction with the above device embodiment, in one embodiment, the various attributes of the account sharing violation control measures include: The first attribute: whether a comprehensive account security management system has been established; Second attribute: whether the account security management system has been promoted during the target period; The third attribute is whether safety system education is provided to all new employees; The fourth attribute: whether the account system intercepts any illegal account sharing activities; The fifth attribute is whether account sharing violations are monitored after the violation occurs. The sixth attribute: whether the current account monitoring model is regularly optimized; The seventh attribute: Is there any follow-up action for those who violate the account sharing rules? Attribute 8: Whether complaints and false positives regarding account sharing violations are regularly removed to optimize the rules for monitoring account sharing violations.
[0047] In conjunction with the above device embodiment, in one embodiment, the multiple different attributes of the account sharing violation control measures further include: Ninth attribute: the ratio of the number of people promoted during the target period to the total number of monitored people; Tenth attribute: the ratio of the number of people who violated account sharing rules during the target period to the total number of monitored people; Eleventh attribute: the ratio of the number of people who filed complaints about account sharing violations to the total number of monitored people during the target period.
[0048] In combination with the above device embodiment, in one embodiment, the ninth attribute, the tenth attribute, and the eleventh attribute are attribute data that have been normalized.
[0049] In combination with the above-mentioned device embodiment, in one embodiment, the machine learning model includes a linear regression model, and the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including: a1. Initialize the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b1. Training the initialized linear regression model based on the control training sample data to calculate the training error of the initialized linear regression model; c1. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d1. Repeat steps b1 to c1 until the test error of the initialized linear regression model reaches the preset error level or the number of training rounds reaches the preset value.
[0050] In combination with the above-mentioned device embodiment, in one embodiment, the machine learning model includes a linear regression model, and the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including: a2. Initializing the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b2. Training the initialized linear regression model based on the control training sample data and the first loss function to calculate the training error of the initialized linear regression model; c2. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d2. Determine, based on the training error, whether the data aggregation degree output by the linear regression model deviates from a preset aggregation degree; e2. When the answer is yes, determine an adapted second loss function based on the cause of the data aggregation deviation, and replace the first loss function in step b2. Repeat steps b2 to c2 after replacement until the test error of the initialized linear regression model reaches a preset error level or the number of training rounds reaches a preset value.
[0051] In combination with the above device embodiment, in one embodiment, determining the adapted second loss function according to the cause of the data aggregation deviation includes: When the reason for the deviation of the data aggregation degree is the presence of outliers in the control training sample data, determining the weighted loss function or the Huber loss function as the adapted second loss function; When the reason for the deviation of the data aggregation degree is the uneven distribution of the control training sample data, the quantile loss function is determined as the adapted second loss function.
[0052] It can be seen that a device for analyzing the number of account sharing violations based on machine learning is provided, which will perform model training based on historical data such as security control measures for user account violation risks. By learning from historical data, the number of account sharing violations in the target organization in a certain period of time in the future can be predicted, thereby predicting which measures in the company's long-term security control measures, when executed within a certain period, will result in the number of account risk violations that the company wants, thereby effectively reducing unnecessary losses to the company due to account risk violations.
[0053] Regarding the specific definition of a device for analyzing the number of account sharing violations based on machine learning, please refer to the definition of the method for analyzing the number of account sharing violations based on machine learning above, which will not be repeated here. Each module in the above-mentioned device for analyzing the number of account sharing violations based on machine learning can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0054] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 6 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. When the computer program is executed by the processor, it implements the functions of a device for analyzing the number of account sharing violations based on machine learning, or when executed, it implements the steps of a method for analyzing the number of account sharing violations based on machine learning.
[0055] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the functions of an apparatus for analyzing the number of account sharing violations based on machine learning, or when executed, it implements the steps of a method for analyzing the number of account sharing violations based on machine learning.
[0056] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: Obtaining account sharing violation control data of a target organization, the control data including a plurality of different attributes describing account sharing violation control measures of the target organization; Inputting the control data into a pre-trained machine learning model to predict the number of account sharing violations within a certain period of time within the target organization; Among them, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
[0057] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the functions of an image processing device, or implements the steps of an image processing method.
[0058] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements the following steps: Obtaining account sharing violation control data of a target organization, the control data including a plurality of different attributes describing account sharing violation control measures of the target organization; Inputting the control data into a pre-trained machine learning model to predict the number of account sharing violations within a certain period of time within the target organization; Among them, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
[0059] For the specific definition of a computer-readable storage medium or a computer device, please refer to the above definition of the steps of a method for analyzing the number of account sharing violations based on machine learning, which will not be repeated here.
[0060] Those skilled in the art will understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), Synchronous Link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0061] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0062] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.
Claims
1. A method for analyzing the number of account sharing violations based on machine learning, characterized in that: The method comprises: Obtaining account sharing violation control data of a target organization, the control data including a plurality of different attributes describing account sharing violation control measures of the target organization; Inputting the control data into a pre-trained machine learning model to predict the number of account sharing violations within a certain period of time within the target organization; The machine learning model is pre-trained in the following way: Obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each of the control training sample data including multiple different attributes describing account sharing violation control measures of a sample organization; The initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
2. The method for analyzing the number of account sharing violations based on machine learning according to claim 1 is characterized in that: The various attributes of the account sharing violation control measures include: The first attribute: whether a comprehensive account security management system has been established; Second attribute: whether the account security management system has been promoted during the target period; The third attribute is whether safety system education is provided to all new employees; The fourth attribute: whether the account system intercepts any illegal account sharing activities; The fifth attribute is whether account sharing violations are monitored after the violation occurs. The sixth attribute: whether the current account monitoring model is regularly optimized; The seventh attribute: Is there any follow-up action for those who violate the account sharing rules? Attribute 8: Whether complaints and false positives regarding account sharing violations are regularly removed to optimize the rules for monitoring account sharing violations.
3. The method for analyzing the number of account sharing violations based on machine learning according to claim 2 is characterized in that: The various attributes of the account sharing violation control measures also include: Ninth attribute: the ratio of the number of people promoted during the target period to the total number of monitored people; Tenth attribute: the ratio of the number of people who violated account sharing rules during the target period to the total number of monitored people; Eleventh attribute: the ratio of the number of people who filed complaints about account sharing violations to the total number of monitored people during the target period.
4. The method for analyzing the number of account sharing violations based on machine learning according to claim 3 is characterized in that: The ninth attribute, the tenth attribute, and the eleventh attribute are attribute data that have been normalized.
5. The method for analyzing the number of account sharing violations based on machine learning according to claim 2 is characterized in that: The machine learning model includes a linear regression model, and the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including: a1. Initialize the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b1. Training the initialized linear regression model based on the control training sample data to calculate the training error of the initialized linear regression model; c1. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d1. Repeat steps b1 to c1 until the test error of the initialized linear regression model reaches the preset error level or the number of training rounds reaches the preset value.
6. The method for analyzing the number of account sharing violations based on machine learning according to claim 2, characterized in that: The machine learning model includes a linear regression model, and the initial machine learning model is trained based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value, including: a2. Initializing the model network parameters of the initial linear regression model using a random initialization method, wherein the model network parameters include weights and biases; b2. Training the initialized linear regression model based on the control training sample data and the first loss function to calculate the training error of the initialized linear regression model; c2. Performing reverse error propagation on the initialized linear regression model according to the training error, so as to sequentially transfer the training error of the initialized linear regression model to the input layer of the initialized linear regression model and update the model network parameters; d2. Determine, based on the training error, whether the data aggregation degree output by the linear regression model deviates from a preset aggregation degree; e2. When the answer is yes, determine an adapted second loss function based on the cause of the data aggregation deviation, and replace the first loss function in step b2. Repeat steps b2 to c2 after replacement until the test error of the initialized linear regression model reaches a preset error level or the number of training rounds reaches a preset value.
7. The method for analyzing the number of account sharing violations based on machine learning according to claim 5 is characterized in that: The determining of an adapted second loss function according to a cause of the data aggregation deviation includes: When the reason for the deviation of the data aggregation degree is the presence of outliers in the control training sample data, determining the weighted loss function or the Huber loss function as the adapted second loss function; When the reason for the deviation of the data aggregation degree is the uneven distribution of the control training sample data, the quantile loss function is determined as the adapted second loss function.
8. A device for analyzing the number of account sharing violations based on machine learning, characterized in that: The device comprises: an acquisition module, configured to acquire account sharing violation control data of a target organization, wherein the control data includes a plurality of different attributes describing account sharing violation control measures of the target organization; A prediction module, configured to input the control data into a pre-trained machine learning model to predict the number of account sharing violations committed by the target organization within a certain period of time in the future; Among them, the machine learning model is pre-trained in the following manner: obtaining a control training sample data set, the control training sample data set including multiple control training sample data, each control training sample data including multiple different attributes describing the account sharing violation control measures of the sample organization; training the initial machine learning model based on the control training sample data until the test error of the initial machine learning model reaches a preset error level or the number of training rounds reaches a preset value.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method for analyzing the number of account sharing violations based on machine learning are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for analyzing the number of account sharing violations based on machine learning are implemented.