Risk prediction method and device for preventing information leakage

By obtaining the historical behavior data of users for third-party applications and using deep neural network models, predicting the confidence score of user attribute information, determining users with information leakage risks, solving the problem of predicting user information leakage risks, and realizing the effectiveness of information security protection.

CN120217442APending Publication Date: 2025-06-27ANT BLOCKCHAIN TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510370352.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

When users use third-party merchant application products, there is a risk of information leakage, especially for important and sensitive user information. The leakage not only has an adverse impact on users, but may also cause economic losses to the digital service platform.

Method used

By acquiring multiple first historical behavior data, including user access data and operation data to third-party applications, using a pre-constructed deep neural network model, each user's attribute information prediction confidence score is obtained, and users with a risk of information leakage are determined to perform information security protection.

Benefits of technology

Effectively predict the security of user behavior, provide a basis and basis for information security protection, help ensure system security, and avoid user information leakage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217442A_ABST
    Figure CN120217442A_ABST
Patent Text Reader

Abstract

One or more embodiments of the invention provide a risk prediction method and device for preventing information leakage, and the method comprises the steps: carrying out the reasoning of the access data and operation data of a plurality of first users for a third-party application through a pre-constructed deep neural network model, and predicting a corresponding confidence score for each piece of attribute information of each first user. The attribute information represents relatively important and sensitive user information of the corresponding first user, so that the confidence score of each piece of attribute information of each first user reflects the possibility that the operation behavior of the corresponding first user has an information leakage risk. Based on this, according to the confidence score predicted by the deep neural network model for each kind of attribute information of each first user, whether the corresponding first user has an information leakage risk or not can be determined, whether the user behavior is safe or not can be effectively predicted, basis and foundation are provided for necessary information safety protection, the system safety is guaranteed, and the user experience is improved. And user information leakage is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the field of information security technology, and in particular, to a risk prediction method and device for preventing information leakage. Background Art

[0002] With the development of Internet technology, network services have spread to all fields, and most of people's daily life and work needs can be completed by using the network services provided by digital service platforms. To meet people's diverse usage needs, the forms of network services are constantly updated. With the emergence of mini-programs, many digital service platforms will embed the application products of third-party merchants in the form of mini-programs in their own platforms, so as to promote cooperation with third-party merchants while providing users with richer network services.

[0003] However, in the process of users using the application products of third-party merchants, there may be a risk of information leakage. Especially for important and sensitive user information, once leaked, it will not only have an adverse impact on users, but may even cause serious economic losses to digital service platforms.

[0004] Therefore, for the situation where users use the application products of third-party merchants, how to accurately predict users at risk of information leakage is the key to providing information security protection for corresponding users. Summary of the Invention

[0005] To ensure system and user information security and prevent information leakage, one or more embodiments of this specification provide a risk prediction method and device for preventing information leakage.

[0006] In a first aspect, one or more embodiments of this specification provide a risk prediction method for preventing information leakage, including: obtaining a plurality of first historical behavior data; the plurality of first historical behavior data includes historical behavior data of a plurality of first users within a preset historical period; the historical behavior data includes access data and operation data for at least one third-party application; based on the plurality of first historical behavior data and a pre-constructed deep neural network model, obtaining a confidence score for predicting each type of attribute information of each first user; the deep neural network model is trained based on a plurality of second historical behavior data and an attribute label corresponding to each second historical behavior data; the plurality of second historical behavior data includes historical behavior data of a plurality of second users; different types of attribute information of the first user respectively correspond to different attributes of user information; according to the confidence scores of each type of attribute information of each first user, determining third users among all first users who are at risk of information leakage, so as to perform information security protection on the user information of the third users.

[0007] In an alternative embodiment, obtaining a plurality of first historical behavior data includes: obtaining access data of a plurality of first users to at least one third-party application during a first preset historical period; obtaining operation data of the plurality of first users on at least one third-party application during a second preset historical period; based on the access data and operation data of the plurality of first users, obtaining a plurality of first historical behavior data; each first historical behavior data includes the access data and operation data of a first user.

[0008] In an alternative embodiment, obtaining access data of a plurality of first users to at least one third-party application during a first preset historical period includes: according to the user identifiers of the plurality of first users, obtaining the application identifiers and application types of at least one third-party application accessed by each first user during the first preset historical period as the access data.

[0009] In an alternative embodiment, obtaining operation data of the plurality of first users on at least one third-party application during a second preset historical period includes: according to the user identifiers of the plurality of first users, obtaining the element identifiers corresponding to the triggering operations performed by each first user on the page elements of at least one third-party application at multiple moments during the second preset historical period as the operation data.

[0010] In an alternative embodiment, based on the plurality of first historical behavior data and a pre-constructed deep neural network model, obtaining the confidence scores for predicting each type of attribute information of each first user includes: calibrating the deep neural network model according to preset calibration parameters and a validation data set to generate a calibrated deep neural network model; inputting the plurality of first historical behavior data into the calibrated deep neural network model to obtain the confidence scores for predicting each type of attribute information of each first user output by the calibrated deep neural network model.

[0011] In an alternative embodiment, according to the confidence scores of each type of attribute information of each first user, determining third users among all first users who have a risk of information leakage includes: according to the preset confidence thresholds corresponding to each type of attribute information, determining first users among the plurality of first users whose confidence scores of attribute information are greater than the corresponding preset confidence thresholds as third users who have a risk of information leakage.

[0012] Second aspect, one or more embodiments of this specification provide a risk prediction device for preventing information leakage, and the device includes: a data collection module, a risk prediction module, and a risk assessment module; the data collection module is configured to obtain a plurality of first historical behavior data; the plurality of first historical behavior data includes historical behavior data of a plurality of first users within a preset historical period, and the historical behavior data includes access data and operation data for at least one third-party application; the risk prediction module is configured to obtain a confidence score for predicting each type of attribute information of each first user based on the plurality of first historical behavior data and a pre-constructed deep neural network model; the deep neural network model is trained based on a plurality of second historical behavior data and an attribute label corresponding to each second historical behavior data; the plurality of second historical behavior data includes historical behavior data of a plurality of second users; different types of attribute information of the first user respectively correspond to different attributes of user information; the risk assessment module is configured to determine third users among all first users who have a risk of information leakage according to the confidence score of each type of attribute information of each first user, so as to perform information security protection on the user information of the third users.

[0013] In an alternative embodiment, when obtaining the plurality of first historical behavior data, the data collection module is configured to: obtain access data of a plurality of first users to at least one third-party application within a first preset historical period; obtain operation data of the plurality of first users on at least one third-party application within a second preset historical period; based on the access data and operation data of the plurality of first users, obtain the plurality of first historical behavior data; each first historical behavior data includes access data and operation data of one first user.

[0014] In an alternative embodiment, when obtaining the access data of a plurality of first users to at least one third-party application within a first preset historical period, the data collection module is configured to: according to the user identifiers of the plurality of first users, obtain the application identifier and application type of at least one third-party application accessed by each first user within the first preset historical period as the access data.

[0015] In an alternative embodiment, when obtaining the operation data of the plurality of first users on at least one third-party application within a second preset historical period, the data collection module is configured to: according to the user identifiers of the plurality of first users, obtain the element identifier corresponding to the triggering operation performed by each first user on the page elements of at least one third-party application at multiple moments within the second preset historical period as the operation data.

[0016] In an alternative embodiment, when the risk prediction module obtains the confidence scores for the prediction of each type of attribute information of each first user based on the multiple first historical behavior data and a pre-constructed deep neural network model, it is configured to: calibrate the deep neural network model according to preset calibration parameters and a validation data set to generate a calibrated deep neural network model; input the multiple first historical behavior data into the calibrated deep neural network model to obtain the confidence scores for the prediction of each type of attribute information of each first user output by the calibrated deep neural network model.

[0017] In an alternative embodiment, when the risk assessment module determines the third users with information leakage risks among all the first users according to the confidence scores of each type of attribute information of each first user, it is configured to: determine, according to a preset confidence threshold corresponding to each type of attribute information, the first users among the multiple first users whose confidence scores of the attribute information are greater than the corresponding preset confidence threshold as the third users with information leakage risks.

[0018] In a third aspect, one or more embodiments of this specification provide an electronic device, which includes: a memory for storing a computer program product; a processor for executing the computer program product stored in the memory, and when the computer program product is executed, implementing the above-mentioned risk prediction method for preventing information leakage.

[0019] In a fourth aspect, one or more embodiments of this specification provide a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed, implementing the above-mentioned risk prediction method for preventing information leakage.

[0020] In summary, the risk prediction method for preventing information leakage provided by one or more embodiments of this specification can infer the access data and operation data of multiple first users to third-party applications within a preset historical period through a pre-constructed deep neural network model to determine the association relationship between the operation behavior and attribute information of each first user, and calculate the confidence scores corresponding to the prediction of each type of attribute information of each first user. Since these attribute information all represent important and sensitive user information, the confidence score of each type of attribute information of each first user reflects the possibility that the operation behavior of the corresponding first user has a risk of user information leakage. Based on this, according to the confidence scores of the prediction of each type of attribute information of each first user by the deep neural network model, it can be determined whether the corresponding first user has an information leakage risk. Through this method, the security of user behavior can be effectively predicted, providing a basis and foundation for necessary information security protection, helping to ensure system security and avoid user information leakage. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] To more clearly illustrate the technical solutions of one or more embodiments of this specification, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 It is a flowchart of a risk prediction method for preventing information leakage provided by an embodiment of this specification.

[0023] Figure 2 It is a schematic structural diagram of a risk prediction device for preventing information leakage provided by an embodiment of this specification.

[0024] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Specific embodiments

[0025] The following will further elaborate on one or more embodiments of this specification through the accompanying drawings and embodiments. Through these descriptions, the features and advantages of one or more embodiments of this specification will become clearer and more distinct.

[0026] The special term "exemplary" here means "serving as an example, embodiment, or illustrative". Any embodiment described as "exemplary" here does not have to be construed as superior to or better than other embodiments. Although various aspects of the embodiments are shown in the drawings, unless otherwise specified, the drawings do not have to be drawn to scale.

[0027] In addition, the technical features involved in different embodiments of this specification described below can be combined with each other as long as they do not conflict with each other.

[0028] With the development of Internet technology, various digital service platforms have emerged. These digital service platforms can provide users with various online services, bringing great convenience to users in life and work. However, with the development of the economy, people's living standards have been continuously improved, and their material and spiritual needs have become more abundant. To meet the more diversified service needs of users and promote the cooperation between the platform and third-party merchants, digital service platforms usually introduce various application products of third-party merchants into the platform for users within the platform to use, so as to improve the user experience. Through this mode, users can use more diversified and personalized online services, which not only enhances the user experience but also helps with the acquisition, retention, recall, and conversion of platform users. At the same time, for third-party merchants, promoting products and services through digital service platforms greatly reduces the operating costs.

[0029] Since mini programs are lightweight applications that do not require installation and are ready to go, they not only save processing resources but are also more convenient for users to use. Therefore, in actual applications, digital service platforms usually embed various mini programs in the platform to introduce various application products from third-party merchants. For digital platforms with a wide range of services and large scale, there are many types of embedded mini programs. Users will generate different behavioral data when using different types of mini programs. These behavioral data may involve very important user information. If it is not identified, once information leakage occurs, it may cause serious consequences for both users and platforms.

[0030] Therefore, for the above application scenarios, one or more embodiments of this specification provide a risk prediction method for preventing information leakage, which can infer the behavior habits of the corresponding user based on the historical behavior data generated when the user uses various mini-programs, determine whether the user's operation behavior involves important user information, and predict whether the corresponding user has information leakage risks. Furthermore, for users who are determined to have information leakage risks, they are identified, and information security protection can be performed on the user information of the corresponding user.

[0031] Before explaining the various steps of the above-mentioned risk prediction method, first, the basic premise for the application of the risk prediction method is explained.

[0032] In one or more embodiments of the present specification, a deep neural network model is pre-constructed, and the model is trained based on multiple historical behavior data generated by each user in the digital service platform when using various mini-programs, and the attribute labels corresponding to each historical behavior data. Among them, the attribute label corresponding to each historical behavior data is used to identify the attribute information of the user involved in the corresponding historical behavior data. Among them, the attribute information of the user is used to characterize the attributes of the user information involved. Different types of attribute information correspond to different attributes of user information. Exemplarily, the attributes of user information may include, but are not limited to, attributes such as personal information, address, location, property ownership, vehicle ownership, marital status, and parental status.

[0033] In an alternative embodiment of this specification, the attribute information of the user involved in the above historical behavior data may include the names of the attributes of the user information. Correspondingly, the content of the attribute label corresponding to the above historical behavior data may include the names of the corresponding attributes. It should be understood that there may be multiple attribute labels corresponding to the above historical behavior data, that is to say, the corresponding historical behavior data may involve multiple attributes of the user information. In this case, multiple attribute labels can be added to the corresponding historical behavior data, and the content of each attribute label includes the name of the corresponding attribute. For example, the attribute labels added to a certain historical behavior data can be two, where the content of one attribute label is personal information A, and the content of the other attribute label is personal information B.

[0034] In an alternative embodiment of this specification, the attribute information of the user involved in the above historical behavior data may include the specific content of the attributes of the user information. Correspondingly, the content of the attribute label corresponding to the above historical behavior data may include the specific content of the corresponding attributes. It should be understood that in this scenario, multiple attribute labels can also be added to the historical behavior data according to the multiple attributes involved in the historical behavior data, and the content of each attribute label includes the specific content of the corresponding attribute. For example, a certain historical behavior data can be added with two attribute labels, where the content of one attribute label may include the following content: User A is male; the content of the other attribute label may include the following content: User A is 30 years old, etc.

[0035] Based on this, the trained deep neural network model has the ability to infer the correlation between historical behavior data and attribute information, and using this model can identify whether important and sensitive user information is involved in the historical behavior data, so as to be used to predict users at risk of information leakage.

[0036] In one or more embodiments of this specification, the specific type of the above deep neural network model is not limited. Optionally, it includes but is not limited to convolutional neural network models (Convolutional Neural Networks, CNN), recurrent neural network models (Recurrent Neural Network, RNN), long short-term memory network models (Long Short-Term Memory, LSTM), etc.; correspondingly, the method for model training is not limited either. Optionally, the cross-entropy loss function (Cross Entropy Loss) can be used to train the model. For example, the binary cross-entropy loss function (Binary Cross Entropy Loss) or the weighted cross-entropy loss function (Weighted Cross-Entropy Loss) can be used. Of course, the specific model type and training method are not limited to this, and can be selected according to actual needs, which will not be elaborated here.

[0037] For the sake of easy distinction, in one or more embodiments of this specification, various applets embedded in the digital service platform are referred to as third-party application programs, users who need to predict whether there is a risk of information leakage are referred to as first users, and users predicted to have a risk of information leakage among the first users are referred to as third users; correspondingly, the historical behavior data generated when the first users use the third-party application programs is referred to as first historical behavior data, the historical behavior data used to train the deep neural network model is referred to as second historical behavior data, and the users who generate the second historical behavior data by using the third-party application programs are referred to as second users.

[0038] Based on the above, the following will describe each step of the risk prediction method for preventing information leakage provided in one or more embodiments of this specification with reference to the accompanying drawings.

[0039] Figure 1 The flowchart of a risk prediction method for preventing information leakage provided in an embodiment of this specification is as Figure 1 shown, and the method may include the following steps:

[0040] S102. Obtain a plurality of first historical behavior data; the plurality of first historical behavior data includes the historical behavior data of a plurality of first users within a preset historical period; the historical behavior data includes access data and operation data for at least one third-party application program;

[0041] S104. Based on the plurality of first historical behavior data and a pre-constructed deep neural network model, obtain the confidence score predicted for each type of attribute information of each first user; the deep neural network model is trained based on the plurality of second historical behavior data and the attribute labels corresponding to each second historical behavior data; the plurality of second historical behavior data includes the historical behavior data of a plurality of second users; different types of attribute information of the first user respectively correspond to different attributes of the user information.

[0042] S106. Determine the third users with information leakage risks among all the first users according to the confidence scores of each type of attribute information of each first user, so as to perform information security protection on the user information of the third users.

[0043] In the embodiments of this specification, the application manner of the above risk prediction method is not limited, and according to different usage requirements, the application manner of the above risk prediction method may also be different. For example, the above risk prediction method may be executed by a cloud server to provide remote risk prediction services. Another example is that the above risk prediction method may also be executed by an offline application server to provide offline risk prediction services. The specific application manner adopted is not limited herein.

[0044] In the embodiments of this specification, the access data and operation data of each first user for at least one third-party application within a preset historical period are referred to as a first historical behavior data. Based on this, in order to determine whether there is a third user at risk of information leakage among multiple first users, it is necessary to first obtain multiple first historical behavior data to predict the third user at risk of information leakage from multiple first users based on the multiple first historical behavior data. Among them, there is no limitation on the method of obtaining multiple first historical behavior data. Optionally, it can be obtained regularly or in real time. According to the different execution times of the above risk prediction method, the acquisition method can also be different.

[0045] As can be seen from the above description, the pre-constructed deep neural network model is trained based on multiple second historical behavior data and the attribute labels corresponding to each second historical behavior data. The trained deep neural network model has the ability to predict the correlation between each first historical behavior data and the attribute information of the first user. Among them, the prediction accuracy of the deep neural network model for the correlation between each first historical behavior data and the attribute information of the first user is measured by a confidence score. The higher the confidence score predicted for each attribute information, the stronger the correlation between the corresponding attribute information and the corresponding first historical behavior data, that is, the greater the possibility that the first user who generates the corresponding first historical behavior data is at risk of information leakage.

[0046] Based on this, after obtaining multiple first historical behavior data, the confidence scores predicted for each type of attribute information of each first user can be obtained based on the multiple first historical behavior data and the pre-constructed deep neural network model. Optionally, after obtaining multiple first historical behavior data, preprocessing operations such as data cleaning, normalization, and standardization can be performed on the multiple first historical behavior data to convert the multiple first historical behavior data into a data format that the deep neural network model can directly recognize. Among them, there is no limitation on the specific form of the converted data format. According to different actual requirements, the selected data format can also be different, which will not be elaborated here.

[0047] Furthermore, after performing the above preprocessing operations, the preprocessed multiple first historical behavior data can be input into the deep neural network model for the deep neural network model to predict the attribute information of the first user involved according to each first historical behavior data, and output the predicted attribute information, as well as the confidence scores corresponding to the prediction of each type of attribute information of each first user. Based on this, the prediction results output by the deep neural network model are obtained. According to the confidence scores of each type of attribute information of each first user among them, the third user at risk of information leakage among all first users can be determined for information security protection of the user information of the third user.

[0048] In the embodiments of this specification, the specific duration and start and end times of the preset historical period are not limited, and according to different actual requirements, the specific duration and start and end times of the preset historical period can also be different. To improve the accuracy of the prediction of the deep neural network model, in the embodiments of this specification, a plurality of first historical behavior data are obtained by obtaining access data and operation data from two different preset historical periods respectively. For the sake of distinction, the preset historical period for obtaining access data is called the first preset historical period, and the preset historical period for obtaining operation data is called the second preset historical period; among them, the first preset historical period can be a relatively long historical period, for example, the past 30 days, for inferring the user's access habits from a long-term dimension; the second preset historical period can be a relatively short historical period, for example, the past 3 days, for measuring the user's operation habits from a short-term dimension. Of course, the above selection method for the preset historical period is only an exemplary illustration and is not limited thereto.

[0049] Based on this, according to the first preset historical period, access data of a plurality of first users to at least one third-party application within the first preset historical period can be obtained, and according to the second preset historical period, operation data of a plurality of first users to at least one third-party application within the second preset historical period can be obtained. Furthermore, based on the access data and operation data obtained for each first user, it is determined as one first historical behavior data obtained.

[0050] In the embodiments of this specification, the method for obtaining access data within the first preset historical period and operation data within the second preset historical period is not limited. Considering that the action of the first user accessing the third-party application is an irregular action, therefore, when obtaining access data, for each first user, all access data within the first preset historical period can be directly obtained, so as to more accurately predict the access habits of the corresponding first user. However, after successfully accessing the third-party application, the first user's various trigger operations on the page elements on the application page may be a series of continuous actions, generating a large amount of operation data. Therefore, when obtaining operation data, for each first user, multiple operation data within the second preset historical period can be obtained in the form of a preset time interval within the second preset historical period, so as to reduce the processing pressure while realizing the prediction of the operation habits of the corresponding first user.

[0051] Furthermore, this specification embodiment does not limit the specific content of accessing data and operating data. For the sake of simplifying the processing, the application identifier and application type of a third-party application are used as the accessed data, and the element identifier corresponding to the triggered page element is used as the operating data for illustration. Of course, in actual applications, according to specific requirements, other content may also be included, which will not be elaborated here. Among them, the application type of the third-party application can be understood as the type of application service provided by the third-party application. For example, it includes but is not limited to finance, life, shopping, retail, catering, life services, education, tourism, entertainment, etc.; the triggered page elements include but are not limited to radio buttons, checkboxes, dropdown lists, scroll bars, text boxes, graphic labels, etc. Of course, the above types of application types and page elements are only for illustrative purposes, and the specific types are subject to actual applications.

[0052] Based on this, when obtaining the access data of multiple first users to at least one third-party application during a first preset historical period, the application identifier and application type of at least one third-party application accessed by each first user during the first preset historical period can be obtained according to the user identifiers of the multiple first users, and used as the access data of the corresponding first user. When obtaining the operation data of multiple first users to at least one third-party application during a second preset historical period, the element identifier corresponding to the trigger operation performed on the page element of at least one third-party application by each first user at multiple moments during the second preset historical period can be obtained according to the user identifiers of the multiple first users, and used as the operation data of the corresponding first user. Among them, the size of the preset time interval between any two adjacent moments is not limited, and according to actual needs, the value of the preset time interval can be flexibly selected. For example, 500 milliseconds.

[0053] In the embodiments of this specification, the specific structure of the deep neural network model is not limited. In an alternative manner, the deep neural network model is a general multi-layer structure. Correspondingly, in the above embodiments, obtaining the confidence score for predicting each type of attribute information of each first user refers to obtaining the prediction result directly output by the output layer of the deep neural network model. In another alternative manner, in order to improve the accuracy of the prediction result, a calibration module can be added to the output layer of the deep neural network model to form a calibrated deep neural network model. Among them, the calibration module can calibrate the prediction result output by the deep neural network model to adjust the gap between the confidence score for predicting each type of attribute information of each first user and the actual accuracy.

[0054] Based on this, obtaining the confidence score for the prediction of each type of attribute information of each first user based on multiple first historical behavior data and a pre-constructed deep neural network model means that first, the deep neural network model is calibrated according to preset calibration parameters and a validation data set to generate a calibrated deep neural network model; then, the multiple first historical behavior data are input into the calibrated deep neural network model to obtain the confidence score for the prediction of each type of attribute information of each first user output by the calibrated deep neural network model. It should be understood that the calibrated deep neural network model will also output the predicted attribute information for each first user. The specific content of the predicted attribute information for each first user can refer to the content of the attribute information corresponding to the second historical data in the aforementioned model training, which will not be elaborated here. Among them, the validation data set can be understood as other historical behavior data different from the multiple second historical behavior data used to train the deep neural network model. For the sake of distinction, it is called multiple third historical behavior data.

[0055] In this embodiment, the specific form of the preset calibration parameters is not limited. Depending on the calibration method, the form of the preset calibration parameters can also be different. Optionally, this embodiment takes the temperature calibration method for calibrating the deep neural network model as an example. Among them, the preset calibration parameter can be a negative logarithm with a temperature parameter. By adjusting the size of the negative logarithm, the confidence score for the prediction of each type of attribute information of each first user by the deep neural network model is adjusted to narrow the gap between the confidence score for the prediction of each type of attribute information and the actual accuracy.

[0056] Based on the above, after obtaining the confidence score for each type of attribute information of each first user, it is possible to determine whether there is a third user with information leakage risk among all first users according to the confidence score for each type of attribute information of each first user. Optionally, a corresponding preset confidence threshold can be set for each type of attribute of the user information. Based on this, according to the preset confidence threshold corresponding to each type of attribute information of the first user, the first users whose confidence score of the attribute information is greater than the corresponding preset confidence threshold are determined from the multiple first users as the third users with information leakage risk. Correspondingly, for the first users whose confidence score of the attribute information is less than or equal to the corresponding preset confidence threshold, it is considered that there is no information leakage risk.

[0057] That is to say, for any first user, if there is one or more pieces of attribute information with confidence scores greater than the corresponding preset confidence threshold among all the predicted attribute information corresponding to the user, it is considered that there is a risk of leakage of the user's information. On the contrary, for any first user, if the confidence scores of all the predicted attribute information corresponding to the user are less than or equal to the corresponding preset confidence threshold, it is considered that there is no risk of leakage of the user's information for the corresponding first user.

[0058] In the embodiments of this specification, the specific values of the preset confidence thresholds for each attribute of the user information are not limited. Depending on the application scenario and the specific attribute, the values of the corresponding preset confidence thresholds can also be different.

[0059] Exemplarily, the preset confidence threshold can be set to a specific value. For example, the preset confidence threshold for a certain attribute can be 0.9. Alternatively, the preset confidence threshold can also be any value selected from a numerical range. For example, the numerical range can be set to a range where the error from the actual accuracy does not exceed 10%.

[0060] It should be noted that in the above embodiments, the setting method of the preset confidence threshold for each attribute and the method of determining the third user from the first users based on the corresponding preset confidence threshold are only exemplary illustrations. The above examples are intended to illustrate the implementation principle of the risk prediction method provided by the embodiments of this specification and are not limited thereto. In the embodiments of this specification, as long as the confidence score of the attribute information of the first user is greater than the preset confidence threshold of the corresponding attribute, the first user is determined as the third user with the risk of information leakage. However, in actual applications, the determination method is not limited thereto. For example, other dimensions such as the type of attribute information and the number of pieces of attribute information with confidence scores greater than the corresponding preset confidence threshold can also be combined to determine the third user. The specific method can be determined according to the actual application requirements and will not be elaborated here.

[0061] Furthermore, it should be noted that for the setting method of the preset confidence threshold for other attributes and the method of determining the third user from the first users based on the corresponding preset confidence threshold, reference can be made to the above embodiments and will not be elaborated here.

[0062] Based on the above, after determining the third users with the risk of information leakage from multiple first users, the corresponding third users can be identified according to the actual application requirements, so as to perform information security protection on the user information of the corresponding third users through the preset information security protection function. For example, before the third user performs subsequent operations, a prompt message can be sent to remind the corresponding third user to pay attention that the next operation may leak personal information and avoid causing adverse consequences.

[0063] In summary, the risk prediction method for preventing information leakage provided by one or more embodiments of this specification can infer the access data and operation data of multiple first users to third-party applications within a preset historical period through a pre-constructed deep neural network model, so as to determine the association relationship between the operation behavior and attribute information of each first user, and then predict the corresponding confidence score for each type of attribute information of each first user. Since these attribute information involve relatively important and sensitive user information of the corresponding first user, therefore, the confidence score of each type of attribute information of each first user reflects the possibility that the operation behavior of the corresponding first user has the risk of information leakage. Based on this, according to the confidence score predicted by the deep neural network model for each type of attribute information of each first user, it can be determined whether the corresponding first user has the risk of information leakage. Through this method, the security of user behavior can be effectively predicted, providing a basis and foundation for necessary information security protection, helping to ensure system security and avoid user information leakage.

[0064] It can be understood that the execution subject of each step in the above risk prediction method for preventing information leakage can be the same device, or the method can also be executed by different devices as the execution subject. In addition, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations can be executed not in the order in which they appear in this article or in parallel. The operation numbers such as S102, S104, etc. are only used to distinguish different operations, and the numbers themselves do not limit the execution order. In addition, these processes can include more or fewer operations, and these operations can be executed in sequence or in parallel.

[0065] It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different information, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types. The above embodiments are only examples. When actually implemented, the above embodiments can be deformed. Those skilled in the art can understand that the deformation methods of the above embodiments that do not require creative labor all fall within the protection scope of this specification and will not be elaborated in the embodiments. All the above optional technical solutions can be borrowed from or combined with each other to form optional embodiments of this specification, which will not be elaborated one by one here.

[0066] Based on the same inventive concept, one or more embodiments of this specification also provide a risk prediction device for preventing information leakage. Since the principle of the problem solved by this risk prediction device is similar to that of the foregoing risk prediction method, therefore, the implementation of the risk prediction device can refer to the implementation of the foregoing risk prediction method, and the repeated parts will not be elaborated.

[0067] See Figure 2 , Figure 2The block diagram of a risk prediction device for preventing information leakage provided by the embodiments of this specification. As Figure 2 shown, the risk prediction device 200 may include: a data collection module 201, a risk prediction module 202, and a risk assessment module 203, where:

[0068] The data collection module 201 is used to obtain a plurality of first historical behavior data; among them, the plurality of first historical behavior data includes the historical behavior data of a plurality of first users within a preset historical period, and the historical behavior data includes access data and operation data for at least one third-party application; the risk prediction module 202 is used to obtain the confidence score predicted for each type of attribute information of each first user based on the plurality of first historical behavior data and a pre-constructed deep neural network model; among them, the deep neural network model is trained based on a plurality of second historical behavior data and the corresponding attribute labels of each second historical behavior data; the plurality of second historical behavior data includes the historical behavior data of a plurality of second users; different types of attribute information of the first user respectively correspond to different attributes of user information; the risk assessment module 203 is used to determine the third users with information leakage risks among all the first users according to the confidence scores of each type of attribute information of each first user, so as to perform information security protection on the user information of the third users.

[0069] In an optional embodiment, when obtaining a plurality of first historical behavior data, the data collection module 201 is used to: obtain the access data of a plurality of first users to at least one third-party application within a first preset historical period; obtain the operation data of a plurality of first users to at least one third-party application within a second preset historical period; based on the access data and operation data of the plurality of first users, obtain a plurality of first historical behavior data; each first historical behavior data includes the access data and operation data of a first user.

[0070] In an optional embodiment, when obtaining the access data of a plurality of first users to at least one third-party application within a first preset historical period, the data collection module 201 is used to: according to the user identifiers of the plurality of first users, obtain the application identifiers and application types of at least one third-party application accessed by each first user within the first preset historical period as the access data.

[0071] In an optional embodiment, when obtaining the operation data of a plurality of first users to at least one third-party application within a second preset historical period, the data collection module 201 is used to: according to the user identifiers of the plurality of first users, obtain the element identifiers corresponding to the trigger operations performed on the page elements of at least one third-party application by each first user at multiple moments within the second preset historical period as the operation data.

[0072] In an alternative embodiment, when the risk prediction module 202 obtains the confidence scores for predicting each type of attribute information of each first user based on multiple pieces of first historical behavior data and a pre-constructed deep neural network model, it is configured to: calibrate the deep neural network model according to preset calibration parameters and a validation data set to generate a calibrated deep neural network model; input the multiple pieces of first historical behavior data into the calibrated deep neural network model to obtain the confidence scores for predicting each type of attribute information of each first user output by the calibrated deep neural network model.

[0073] In an alternative embodiment, when the risk assessment module 203 determines third users with information leakage risks among all first users according to the confidence scores of each type of attribute information of each first user, it is configured to: determine, according to a preset confidence threshold corresponding to each type of attribute information, first users among the multiple first users whose confidence scores of the attribute information are greater than the corresponding preset confidence threshold as the third users with information leakage risks.

[0074] It should be noted that each module of the above risk prediction device has the same implementation principle as the corresponding steps of the above risk prediction method. For specific content, please refer to the description of the corresponding part in the foregoing embodiments, and details are not described herein again.

[0075] Based on the same inventive concept, one or more embodiments of this specification further provide an electronic device. Refer to Figure 3 , Figure 3 which is a structural block diagram of an electronic device provided in an embodiment of this specification.

[0076] As Figure 3 shown, the electronic device 300 may include a processor 301, a memory 302, and a program or instruction stored on the memory 302 and executable on the processor 301. When the program or instruction is executed by the processor 301, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, details are not described herein again.

[0077] It should be noted that the electronic devices in the embodiments of this specification include mobile electronic devices and non-mobile electronic devices.

[0078] One or more embodiments of this specification further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements each process of the above method embodiment and can achieve the same technical effect. To avoid repetition, details are not described herein again.

[0079] Among them, the processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes computer-readable storage media such as computer read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc.

[0080] One or more embodiments of this specification are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to one or more embodiments of this specification. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0081] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0082] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0083] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the corresponding descriptions in the method embodiments. In this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. For those of ordinary skill in the art, the specific meanings of the above terms in the embodiments of this specification can be understood according to specific circumstances.

[0084] It should be noted that, without conflict, the embodiments in this specification and the features in the embodiments can be combined with each other. This specification is not limited to any single aspect, nor to any single embodiment, nor to any arbitrary combination and / or permutation of these aspects and / or embodiments. Moreover, each aspect and / or embodiment of this specification can be used alone or in combination with one or more other aspects and / or their embodiments.

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this specification, and not to limit them; although this specification has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this specification, and they should all be covered within the scope of this specification.

Claims

1. A risk prediction method for preventing information leakage, characterized in that: include: Obtaining a plurality of first historical behavior data; The plurality of first historical behavior data include historical behavior data of a plurality of first users within a preset historical period; the historical behavior data include access data and operation data of at least one third-party application; Based on the multiple first historical behavior data and the pre-built deep neural network model, a confidence score for each attribute information prediction of each first user is obtained; the deep neural network model is trained based on the multiple second historical behavior data and the attribute label corresponding to each second historical behavior data; the multiple second historical behavior data include historical behavior data of multiple second users; different attribute information of the first user respectively corresponds to different attributes of user information; According to the confidence score of each type of attribute information of each first user, a third user with information leakage risk among all the first users is determined, so as to perform information security protection on the user information of the third user.

2. The method according to claim 1, characterized in that Get multiple first historical behavior data, including: Acquire access data of multiple first users to at least one third-party application within a first preset historical period; Acquire operation data of the plurality of first users on at least one third-party application within a second preset historical period; Based on the access data and operation data of the multiple first users, multiple first historical behavior data are acquired; each first historical behavior data includes the access data and operation data of one first user.

3. The method according to claim 2, characterized in that Acquiring access data of multiple first users to at least one third-party application within a first preset historical period includes: According to the user identifiers of the plurality of first users, an application identifier and an application type of at least one third-party application program accessed by each first user within a first preset historical period are obtained as access data.

4. The method according to claim 2, characterized in that: Acquiring operation data of the plurality of first users on at least one third-party application within a second preset historical period includes: According to the user identifiers of the multiple first users, element identifiers corresponding to trigger operations performed by each first user on a page element of at least one third-party application at multiple moments within the second preset historical period are obtained as operation data.

5. The method according to any one of claims 1 to 4, characterized in that: Based on the plurality of first historical behavior data and the pre-built deep neural network model, obtaining a confidence score for each attribute information prediction of each first user includes: Calibrate the deep neural network model according to preset calibration parameters and a verification data set to generate a calibrated deep neural network model; The multiple first historical behavior data are input into the calibrated deep neural network model, and a confidence score for each attribute information prediction of each first user output by the calibrated deep neural network model is obtained.

6. The method according to claim 5, characterized in that Determining third users with information leakage risks among all first users according to the confidence score of each type of attribute information of each first user includes: According to the preset confidence threshold corresponding to each type of attribute information, a first user whose attribute information confidence score is greater than the corresponding preset confidence threshold is determined from the multiple first users as a third user with information leakage risk.

7. A risk prediction device for preventing information leakage, characterized in that: The device comprises: a data collection module, a risk prediction module and a risk assessment module; The data collection module is used to obtain a plurality of first historical behavior data; the plurality of first historical behavior data include historical behavior data of a plurality of first users within a preset historical period, and the historical behavior data include access data and operation data of at least one third-party application; The risk prediction module is used to obtain a confidence score for each attribute information prediction of each first user based on the multiple first historical behavior data and a pre-built deep neural network model; the deep neural network model is trained based on multiple second historical behavior data and attribute labels corresponding to each second historical behavior data; the multiple second historical behavior data include historical behavior data of multiple second users; different attribute information of the first user corresponds to different attributes of user information respectively; The risk assessment module is used to determine third users with information leakage risks among all first users based on the confidence scores of each type of attribute information of each first user, so as to perform information security protection on user information of the third users.

8. The device according to claim 7, characterized in that When acquiring a plurality of first historical behavior data, the data collection module is used to: Acquire access data of multiple first users to at least one third-party application within a first preset historical period; Acquire operation data of the plurality of first users on at least one third-party application within a second preset historical period; Based on the access data and operation data of the multiple first users, multiple first historical behavior data are acquired; each first historical behavior data includes the access data and operation data of one first user.

9. The device according to claim 8, characterized in that When acquiring access data of multiple first users to at least one third-party application within a first preset historical period, the data collection module is used to: According to the user identifiers of the plurality of first users, an application identifier and an application type of at least one third-party application program accessed by each first user within a first preset historical period are obtained as access data.

10. The device according to claim 8, characterized in that When acquiring the operation data of the plurality of first users on at least one third-party application within the second preset historical period, the data collection module is used to: According to the user identifiers of the multiple first users, element identifiers corresponding to trigger operations performed by each first user on a page element of at least one third-party application at multiple moments within the second preset historical period are obtained as operation data.

11. The device according to any one of claims 7 to 10, characterized in that: When the risk prediction module obtains the confidence score of each attribute information prediction of each first user based on the plurality of first historical behavior data and the pre-built deep neural network model, it is used to: Calibrate the deep neural network model according to preset calibration parameters and a verification data set to generate a calibrated deep neural network model; The multiple first historical behavior data are input into the calibrated deep neural network model, and a confidence score for each attribute information prediction of each first user output by the calibrated deep neural network model is obtained.

12. The device according to claim 11, characterized in that The risk assessment module is used to determine, based on the confidence score of each attribute information of each first user, a third user with information leakage risk among all first users: According to the preset confidence threshold corresponding to each type of attribute information, a first user whose attribute information confidence score is greater than the corresponding preset confidence threshold is determined from the multiple first users as a third user with information leakage risk.

13. An electronic device, characterized in that: The electronic device comprises: a memory for storing a computer program product; A processor is used to execute the computer program product stored in the memory, and when the computer program product is executed, the method described in any one of claims 1 to 6 is implemented.

14. A computer-readable storage medium storing a computer program, characterized in that: The computer program is configured to implement the method of any one of claims 1-6 when executed by a processor.