Risk assessment method and equipment for target data

Through distributed collaboration and model fusion methods, the problem of inaccurate data security evaluation results in the prior art is solved, and higher accuracy and reliability of evaluation results are achieved.

CN120068151APending Publication Date: 2025-05-30CHINA EVERBRIGHT BANK
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510146428.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, data security assessment methods rely on manual audits or are based on classical statistical models, resulting in low accuracy of assessment results, especially in cross-departmental collaboration, which leads to poor data sharing, affecting the accuracy of assessment results.

Method used

The distributed collaboration model training method is adopted to divide the data set into multiple sub-data sets and allocate them to different clients for local model training. The global model is used for model fusion and update, and the convergence of the global model is determined through the model comparison loss function, thereby achieving risk assessment of the target data.

Benefits of technology

Through distributed collaboration and model fusion, information barriers are broken, the accuracy and reliability of evaluation results are improved, and the problem of inaccurate evaluation results in traditional methods is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068151A_ABST
    Figure CN120068151A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a risk assessment method and equipment for target data. The method comprises the following steps: segmenting a data set into a plurality of sub-data sets; sending one sub-data set to a client to enable the client to train a local model based on the sub-data set to obtain a local model of a first stage; obtaining a global model based on the plurality of local models of the first stage, and sending the global model to each client, so that the client updates the local model of the first stage based on the global model to obtain a local model of a second stage, the client determines a model comparison loss function based on the local model of the first stage, the global model and the local model of the second stage to obtain a loss value, and determines global model convergence based on the loss value; and sending the converged global model to each client, so that the client performs risk assessment on the target data based on the converged global model to obtain an assessment result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the technical field of data processing, and in particular, to a method and device for risk assessment of target data. Background Art

[0002] In today's digital age, data has become one of the core assets of enterprises, organizations and even countries. Data security is not only related to personal privacy protection, but also involves business secrets, national security and the stable operation of society. Data leakage, tampering or abuse may lead to serious economic, legal and reputational losses, so it is crucial to ensure the security of data.

[0003] With the rapid development of big data technology, the scale and complexity of data are constantly increasing, and traditional data security detection methods can no longer meet modern needs. In recent years, data security detection technology has evolved from simple firewalls and intrusion detection systems (IDS) to intelligent detection systems based on machine learning and deep learning. For example, intrusion detection systems based on deep neural networks (DNN) can process and analyze large-scale data in real time and effectively identify unforeseen network attacks. In addition, data security assessment tools are also constantly evolving, from single data encryption and access control to comprehensive data classification and grading, vulnerability scanning and compliance detection.

[0004] Although data security detection technology has made significant progress, there are still some evaluation methods that rely on manual review or are based on classical statistical models. These methods have many drawbacks in practical applications. First, manual review is inefficient and inaccurate, making it difficult to cope with the real-time monitoring needs of massive data. Second, classical statistical models face the problem of insufficient accuracy in data security assessment, especially in cross-departmental collaboration, where information barriers may lead to poor data sharing, further affecting the accuracy of the evaluation results. Therefore, there is a problem of low accuracy in the evaluation results. Summary of the invention

[0005] The embodiments of the present invention provide a method and device for risk assessment of target data, so as to at least solve the problem of low accuracy of assessment results existing in the related art.

[0006] According to an embodiment of the present invention, a risk assessment method for target data is provided, including: splitting a data set into multiple sub-data sets; sending one of the sub-data sets to a client, so that the client trains a local model based on the sub-data set to obtain a local model in the first stage; obtaining a global model based on multiple local models in the first stage, and sending the global model to each client, so that the client updates the local model in the first stage based on the global model to obtain a local model in the second stage, and so that the client determines a model comparison loss function based on the local model in the first stage, the global model, and the local model in the second stage to obtain a loss value, and determines the convergence of the global model based on the loss value; sending the converged global model to each client, so that the client performs risk assessment on the target data based on the converged global model to obtain an assessment result.

[0007] According to another embodiment of the present invention, a risk assessment method for target data is further provided, including: training a local model based on a sub-data set to obtain a local model in the first stage; wherein, the sub-data set is any one of multiple sub-data sets split by a server based on a data set; sending the local model in the first stage to the server, so that the server obtains a global model based on multiple local models in the first stage; updating the local model in the first stage based on the global model to obtain a local model in the second stage; determining a model comparison loss function based on the local model in the first stage, the global model, and the local model in the second stage to obtain a loss value, and determining the convergence of the global model based on the loss value; performing risk assessment on the target data based on the converged global model to obtain an assessment result.

[0008] According to another embodiment of the present invention, a computer-readable storage medium is further provided, in which a computer program is stored, wherein the computer program, when run by a processor, executes the steps in any one of the above method embodiments.

[0009] According to another embodiment of the present invention, an electronic device is further provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0010] According to another embodiment of the present invention, a computer program product is further provided, including computer instructions, and the computer instructions, when executed by a processor, implement the steps in any one of the above method embodiments.

[0011] According to one embodiment of the present invention, since a distributed collaborative model training method is adopted, by splitting a data set into multiple sub-data sets and distributing them to different clients for local model training, and at the same time using a global model for model fusion and update to obtain the characteristics of the target data of multiple clients, the information barrier is broken, and the convergence of the global model is determined through a model comparison loss function. Therefore, the problem of low accuracy of evaluation results in the related art is solved, and the effect of improving the accuracy of the evaluation results of the target data is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present invention, and constitute a part of the present invention. The illustrative embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0013] Figure 1 is a schematic structural diagram of a computer terminal for executing the risk assessment method of target data in an embodiment of the present invention;

[0014] Figure 2 is the flow of the risk assessment method of target data according to an embodiment of the present invention Figure 1 ;

[0015] Figure 3 is a flowchart of a method for splitting a data set into multiple sub-data sets according to an embodiment of the present invention;

[0016] Figure 4 is the flow of the risk assessment method of target data according to an embodiment of the present invention Figure 2 ;

[0017] Figure 5 is a schematic structural diagram of a local model of a client according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] The present invention will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0019] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. The embodiments of the present invention can be applied to scenarios for risk assessment of credit status data of users (such as borrowers). The following takes the scenario of a bank's risk assessment of credit status data of users (such as borrowers) as an example for explanation. Of course, the scenarios of the embodiments of the present invention are not limited to the following scenarios.

[0020] The method embodiments provided by the embodiments of the present invention can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 is a schematic structural diagram of a computer terminal for the risk assessment method of execution target data according to an embodiment of the present invention. As Figure 1 shown, the computer terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a field programmable gate array FPGA) and a memory 104 for storing data. Optionally, the above computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above computer terminal. For example, the computer terminal may further include more or fewer components than those Figure 1 shown in the figure, or have a different configuration from that Figure 1 shown.

[0021] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the risk assessment method of target data in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0022] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computer terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a gateway and thus can communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.

[0023] In this embodiment, a risk assessment method for target data is provided. Figure 2 is the flow of the risk assessment method for target data according to an embodiment of the present inventionFigure 1 , as Figure 2 shown, this method is applied to a server, and the method flow includes the following steps:

[0024] Step S201: Split the data set into multiple sub - data sets;

[0025] Step S202: Send a sub - data set to a client so that the client trains a local model based on the sub - data set to obtain a local model in the first stage;

[0026] In an exemplary implementation, for example, the server is located in a branch. The server of the branch uses personal customer data, classifies and organizes it into a data set containing features such as customer basic information, transaction behavior, wealth management purchase behavior, and asset size, extracts historical data of the past year as a sample set (i.e., the data set), splits the sample set (i.e., the data set) into multiple sub - data sets for model (such as a local model) training, and combines the customer credit repayment data of the sub - branches to train a local model (such as a local model in the first stage). There can be multiple clients, and the multiple clients are respectively located in multiple sub - branches under the branch. It should be noted that the above "one client" and "multiple clients" do not limit the number of clients, but limit the mapping relationship between the clients and the sub - branches. There can be multiple clients in one sub - branch, and these multiple clients can be the "one client" that receives a sub - data set mentioned above.

[0027] Therefore, by splitting the data set into multiple sub - data sets and distributing the sub - data sets to the clients (such as Client 1, Client 2, Client 3) of multiple sub - branches under the branch for local model training, centralized processing of data is avoided, thus effectively protecting data privacy. The client of each sub - branch only processes local data and does not need to share the original data, reducing the risk of data leakage.

[0028] Step S203: Obtain a global model based on multiple local models in the first stage, send the global model to each client so that the client updates the local model in the first stage based on the global model to obtain a local model in the second stage, and so that the client determines a model comparison loss function based on the local model in the first stage, the global model, and the local model in the second stage to obtain a loss value, and determine the convergence of the global model based on the loss value;

[0029] In an exemplary implementation, for example, Branch A has Sub - branch A 1 , Sub - branch A 2 , Sub - branch A 3 , and Branch A has a server. Among them, Sub - branch A 1 has Client A 1 , and in Client A 1 a local model A in the first stage is trained1 Sub-branch A 2 is equipped with Client A 2 and Client A 2 has a locally-trained model A for the first stage 2 Sub-branch A 3 is equipped with Client A 3 and Client A 3 has a locally-trained model A for the first stage 3 The server receives the locally-trained model A for the first stage 1 the locally-trained model A for the first stage 2 the locally-trained model A for the first stage 3 Based on the locally-trained model A for the first stage 1 the locally-trained model A for the first stage 2 the locally-trained model A for the first stage 3 a global model is obtained and sent to Client A 1 Client A 2 Client A 3 so that Client A 1 updates the locally-trained model A for the first stage based on the global model 1 to obtain the locally-trained model A for the second stage 1 so that Client A 2 updates the locally-trained model A for the first stage based on the global model 2 to obtain the locally-trained model A for the second stage 2 so that Client A 3 updates the locally-trained model A for the first stage based on the global model 3 to obtain the locally-trained model A for the second stage 3 Moreover, for example, Client A 1 determines a model comparison loss function based on the locally-trained model A for the first stage 1 the global model, and the locally-trained model A for the second stage 1 to obtain a loss value and determine the convergence of the global model based on the loss value. Or, Client A 2 determines a model comparison loss function based on the locally-trained model A for the first stage 2 the global model, and the locally-trained model A for the second stage 2 to obtain a loss value and determine the convergence of the global model based on the loss value. Or, Client A 3 determines a model comparison loss function based on the locally-trained model A for the first stage 3 the global model, and the locally-trained model A for the second stage 3 to obtain a loss value and determine the convergence of the global model based on the loss value.

[0030] Therefore, by collecting the first-stage local models trained by each branch's client and constructing a global model based on these local models, the fusion and optimization of the models are achieved. The construction of the global model synthesizes the characteristics of the local models of each branch and can more comprehensively reflect the overall risk characteristics of the data. For example, Client A 1 Client A 2 and Client A 3 The global model obtained after the fusion of the local models trained respectively can make up for the limitations of a single local model and improve the generalization ability and accuracy of the model. Moreover, by updating the local models of each branch's client with the global model, the second-stage local models are generated, and the global model is dynamically adjusted through the model comparison loss function until the global model converges. This process ensures the continuous optimization and dynamic adjustment of the model. For example, Client A 1 Client A 2 and Client A 3 After updating the local model, calculate the loss value based on the comparison loss function between the global model and the local model, and dynamically adjust the global model to finally achieve the convergence of the global model, thus ensuring the stability and accuracy of the model.

[0031] Step S204: Send the converged global model to each client so that the client can perform a risk assessment on the target data based on the converged global model to obtain an assessment result.

[0032] In an exemplary implementation, when the server learns that the global model has converged, send the converged global model to Client A 1 Client A 2 Client A 3 . So that Client A 1 can perform a risk assessment on the target data of this branch based on the global model to obtain an assessment result. Or, so that Client A 2 can perform a risk assessment on the target data of this branch based on the global model to obtain an assessment result. Or, so that Client A 3 can perform a risk assessment on the target data of this branch based on the global model to obtain an assessment result.

[0033] Therefore, after the global model converges, send the converged global model to the clients of each branch for risk assessment of the target data. Since the global model integrates the characteristics of the local models of multiple branches and has undergone dynamic optimization and convergence verification, it can more accurately assess the risk of the target data. For example, Client A 1 Client A 2 Client A 3Risk assessment of the target data of each branch based on a converged global model can obtain more accurate and valuable evaluation results, effectively solving the problem of inaccurate evaluation results caused by information barriers in traditional methods.

[0034] Through the above steps S201 to S204, since a distributed collaborative model training method is adopted, by splitting the dataset into multiple sub-datasets and distributing them to different clients for local model training, and at the same time using the global model for model fusion and update to obtain the characteristics of the target data of multiple clients, the information barrier is broken, and the convergence of the global model is determined through the model comparison loss function. Therefore, the problem of low accuracy of evaluation results existing in the related technology is solved, and the effect of improving the accuracy of the evaluation results of the target data is achieved.

[0035] In one implementation manner, Figure 3 is a flowchart of a method for splitting a dataset into multiple sub-datasets according to an embodiment of the present invention. As Figure 3 shown, splitting the dataset into multiple sub-datasets includes:

[0036] Step S301, performing standard normalization and data augmentation on the dataset to obtain a processed dataset;

[0037] In an exemplary implementation manner, for example, the server performs standard normalization on the dataset D, sets the data specification to be between 0 and 1, and the calculation formula for normalization can be: where x is any element in the dataset D, and then uses methods such as rotation, scaling, and cropping to perform data augmentation on the normalized dataset.

[0038] Step S302, splitting the processed dataset into multiple sub-datasets based on the non-independent and identically distributed algorithm.

[0039] In one implementation manner, the non-independent and identically distributed algorithm is the Dirichlet distribution algorithm.

[0040] In an exemplary implementation manner, for example, the server uses the Dirichlet distribution algorithm to split the dataset D into sub-dataset D 1 、sub-dataset D 2 、sub-dataset D 3 . And send sub-dataset D 1 to client A 1 , send sub-dataset D 2 to client A 2 , send sub-dataset D 3 to client A3.

[0041] For example, the server has a dataset D containing customer transaction data, and dataset D has the following features: 1. Basic customer information (such as age, gender); 2. Transaction behavior (such as transaction amount, transaction frequency); 3. Wealth management purchase behavior (such as types of wealth management products purchased, amount); 4. Asset size (such as account balance, asset category). Among them, dataset D contains 1000 records, and each record represents the transaction behavior and related information of a customer.

[0042] Based on the server, the parameter α of the Dirichlet distribution is selected. The parameter α determines the concentration degree of the data distribution. A smaller α value will result in a more uneven data distribution, while a larger α value makes the data distribution closer to a uniform distribution. For example, selecting α = [0.5, 0.5, 0.5] means that it is expected that the generated sub-datasets have a high degree of heterogeneity.

[0043] According to the Dirichlet distribution, generate the data distribution probabilities for each client. For example, for three clients (Client A 1 、Client A 2 、Client A 3 ), the generated probability distributions may be:

[0044] Client A 1 : p 1 = 0.4;

[0045] Client A 2 : p 2 = 0.3;

[0046] Client A 3 : p 3 = 0.3.

[0047] These probability values represent the proportion of the data volume that each client will obtain in the total dataset.

[0048] According to the above probability distribution, dataset D is sliced into three sub-datasets:

[0049] Sub-dataset D 1 : Contains 40% of the data in dataset D, that is, 400 records.

[0050] Sub-dataset D 2 : Contains 30% of the data in dataset D, that is, 300 records.

[0051] Sub-dataset D 3 : Contains 30% of the data in dataset D, that is, 300 records.

[0052] Allocate the sliced sub-datasets to the corresponding clients:

[0053] Allocate sub-dataset D1 Sent to client A 1 ;

[0054] Send the sub - dataset D 2 Sent to client A 2 ;

[0055] Send the sub - dataset D 3 Sent to client A 3 .

[0056] The above process will be explained with examples as follows:

[0057] For example, the customer transaction data in dataset D is distributed as follows according to the transaction amount:

[0058] Low - amount transactions (< 100 yuan): 300 records;

[0059] Medium - amount transactions (100 - 500 yuan): 400 records;

[0060] High - amount transactions (> 500 yuan): 300 records.

[0061] Through the Dirichlet distribution algorithm, the server divides dataset D into three sub - datasets and assigns them to client A 1 , client A 2 and client A 3 . The sub - datasets after splitting may have the following characteristics:

[0062] Sub - dataset D 1 (client A 1 ):

[0063] Low - amount transactions: 120 records;

[0064] Medium - amount transactions: 180 records;

[0065] High - amount transactions: 100 records.

[0066] Sub - dataset D 2 (client A 2 ):

[0067] Low - amount transactions: 90 records;

[0068] Medium - amount transactions: 120 records;

[0069] High - amount transactions: 90 records.

[0070] Sub - dataset D 3 (client A 3 ):

[0071] Low - amount transactions: 90 records;

[0072] Medium - amount transactions: 100 records;

[0073] High - amount transactions: 110 records.

[0074] Therefore, through the Dirichlet distribution algorithm, the server can split the dataset D into sub - datasets with heterogeneous characteristics and allocate these sub - datasets to different clients. This splitting method not only ensures the non - independent and identically - distributed characteristics of the data but also can simulate the data distribution differences among different clients in the real - world scenario, thus providing a more practical data basis for subsequent distributed model training and risk assessment.

[0075] In one implementation manner, obtaining a global model based on multiple local models in the first stage includes:

[0076] Performing weighted - average aggregation on multiple local models in the first stage to obtain a global model.

[0077] In one implementation manner, the model contrast loss function is:

[0078]

[0079] where z prev is the local model in the first stage, z glob is the global model, z is the local model in the second stage, and sim() is a function for calculating the similarity between two features.

[0080] In an exemplary implementation manner,

[0081] In one implementation manner, z prev = R wt-1 (x); z glob = R wt (x); z = R wt+1 (x). Where x is an element of the sub - dataset.

[0082] According to another embodiment of the present invention, a risk assessment method for target data is further provided. Figure 4 It is the process of the risk assessment method for target data according to the embodiment of the present invention. Figure 2 , as Figure 4 shown, this method is applied to the client, and the method process includes the following steps:

[0083] Step S401, training a local model based on the sub - dataset to obtain a local model in the first stage; where the sub - dataset is any one of the multiple sub - datasets split by the server based on the dataset;

[0084] Step S402, sending the local model of the first stage to the server, so that the server obtains a global model based on multiple local models of the first stage;

[0085] In an exemplary embodiment, the client trains a local model based on a sub-data set in an automated manner (step S401), without manual intervention, which greatly improves the efficiency of data processing. The client can quickly generate a local model of the first stage and send it to the server (step S402), thereby achieving efficient data processing and model training.

[0086] Step S403, updating the local model of the first stage based on the global model to obtain the local model of the second stage;

[0087] Step S404, determining a model comparison loss function based on the local model of the first stage, the global model, and the local model of the second stage, obtaining a loss value, and determining the convergence of the global model based on the loss value;

[0088] In an exemplary embodiment, the traditional classical statistical model has the problem of insufficient accuracy in data security assessment, especially when dealing with complex and large-scale data. The present invention realizes dynamic optimization and updating of the model through the collaboration between the client and the server. The client updates the local model of the first stage based on the global model, generates the local model of the second stage (step S403), and dynamically adjusts the global model through the model comparison loss function (step S404). This process not only utilizes the characteristics of local data, but also integrates the advantages of the global model, makes up for the limitations of a single model, and thus significantly improves the accuracy and generalization ability of the model.

[0089] Secondly, it is mentioned in the background technology that information barriers in cross-departmental collaboration will lead to poor data sharing, further affecting the accuracy of the evaluation results. The present invention divides the data set into multiple sub-datasets and distributes them to different clients (such as branches under a branch) in a distributed collaborative manner. Each client only processes local data without sharing the original data. The server builds a global model based on the first-stage local models of multiple clients (step S402) and sends it back to the client for updating the local model (step S403). This distributed collaboration method not only protects data privacy, but also breaks down information barriers, so that the models of different clients can be integrated and optimized, thereby improving the accuracy of the evaluation results.

[0090] Moreover, in the traditional method, the training of the model is often static and difficult to adapt to the changes in data. Through the interaction between the client and the server, the present invention realizes the dynamic optimization and global convergence of the model. The client calculates the model comparison loss function based on the local model in the first stage, the global model, and the local model in the second stage to obtain a loss value (step S404), and dynamically adjusts the global model according to the loss value until the global model converges. This process ensures the continuous optimization and stability of the model, making the final evaluation result more accurate.

[0091] Step S405: Perform a risk assessment on the target data based on the converged global model to obtain an evaluation result.

[0092] In an exemplary embodiment, after the global model converges, the client performs a risk assessment on the target data based on the converged global model (step S405). Since the global model integrates the local model features of multiple clients and has undergone dynamic optimization and convergence verification, it can more comprehensively and accurately evaluate the risk of the target data. For example, Client A 1 、Client A 2 and Client A 3 respectively perform a risk assessment on the target data of their respective branches based on the converged global model, and can obtain more accurate and valuable evaluation results, effectively solving the problem of inaccurate evaluation results caused by information barriers in the traditional method.

[0093] Through the above steps S401 to S405, the embodiments of the present invention are based on the clients of each branch, and through technical means such as distributed cooperation, model fusion, and dynamic optimization, solve the problems of low efficiency of manual review, insufficient accuracy of classical statistical models, and information barriers in cross-departmental cooperation in the traditional method. The collaborative work between the client and the server not only improves the efficiency of data processing, but also significantly improves the accuracy and reliability of the evaluation results, providing an efficient and intelligent solution for data security assessment.

[0094] In one embodiment, the model comparison loss function is:

[0095]

[0096] where z prev is the local model in the first stage, z glob is the global model, z is the local model in the second stage, and sim() is a function for calculating the similarity between two features.

[0097] In one embodiment, z prev = R wt-1 (x); z glob = R wt (x); z = Rwt+1 (x). Where x is a sub-dataset.

[0098] In one embodiment, the global model, the local model in the first stage, and the local model in the second stage can be neural network models.

[0099] In an exemplary embodiment, Figure 5 is a schematic structural diagram of the local model of the client according to an embodiment of the present invention, as Figure 5 shown, the neural network model includes: a first convolutional layer (Conv2d(3, 6, 5)), a first max pooling layer (MaxPool2d(2, 2)), an activation function layer (ReLU()), a second convolutional layer (Conv2d(6, 16, 5)), a second max pooling layer (MaxPool2d(2, 2)), a first fully connected layer (FC(16, 120)), a batch normalization layer (BatchNorm1d(120)), a second fully connected layer (FC(120, 84)), and a third fully connected layer (FC(84, 10)).

[0100] Among them, the first convolutional layer (Conv2d(3, 6, 5)), the first max pooling layer (MaxPool2d(2, 2)), the activation function layer (ReLU()), the second convolutional layer (Conv2d(6, 16, 5)), the second max pooling layer (MaxPool2d(2, 2)), the first fully connected layer (FC(16, 120)), the batch normalization layer (BatchNorm1d(120)), the second fully connected layer (FC(120, 84)), and the third fully connected layer (FC(84, 10)) are connected in sequence.

[0101] For example, the functions of each layer are as follows:

[0102] The first convolutional layer (Conv2d(3, 6, 5)): Applies 6 5x5 convolutional kernels to the 3-channel input data to extract preliminary features.

[0103] The first max pooling layer (MaxPool2d(2, 2)): Performs downsampling through a 2x2 pooling window to reduce the size of the features.

[0104] The activation function layer (ReLU()): Applies the ReLU activation function to introduce non-linearity.

[0105] The second convolutional layer (Conv2d(6, 16, 5)): Applies 16 5x5 convolutional kernels to the output of the previous layer to further extract features.

[0106] The second max pooling layer (MaxPool2d(2, 2)): Performs downsampling again through a 2x2 pooling window.

[0107] The first fully connected layer (FC(16,120)): Flattens the output of the previous layer and maps it to 120 neurons to integrate the extracted features.

[0108] The batch normalization layer (BatchNorm1d(120)): Performs batch normalization on the output of 120 neurons to improve the stability and speed of training.

[0109] The second fully connected layer (FC(120,84))): Maps 120 neurons to 84 neurons to further integrate features.

[0110] The third fully connected layer (FC(84,10)): Maps 84 neurons to 10 outputs, which may represent different risk categories or evaluation results.

[0111] By adopting this neural network model, the client can efficiently process and analyze data to support risk assessment tasks.

[0112] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software adding the necessary general hardware platform, and of course, it can also be implemented by hardware. However, in many cases, the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0113] An embodiment of the present invention also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run by a processor, it executes the steps in any one of the above method embodiments.

[0114] In one embodiment, in this embodiment, the above storage medium may include, but is not limited to: USB flash drive, read-only memory (abbreviated as ROM), random access memory (abbreviated as RAM), mobile hard disk, magnetic disk, or optical disk, etc., various media that can store computer programs.

[0115] An embodiment of the present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0116] An embodiment of the present invention further provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the methods described in the various embodiments of the present invention.

[0117] In one implementation manner, the specific examples in this embodiment may refer to the examples described in the above embodiments and optional implementation manners, and will not be elaborated herein.

[0118] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a sequence different from that here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.

[0119] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A target data risk assessment method, characterized in that: include: Split the data set into multiple sub-data sets; Sending one of the sub-datasets to a client, so that the client trains a local model based on the sub-dataset to obtain a local model of the first stage; A global model is obtained based on multiple local models of the first stage, and the global model is sent to each of the clients, so that the client updates the local model of the first stage based on the global model to obtain the local model of the second stage, and the client determines a model comparison loss function based on the local model of the first stage, the global model, and the local model of the second stage to obtain a loss value, so as to determine the convergence of the global model based on the loss value; The converged global model is sent to each of the clients, so that the client performs risk assessment on the target data based on the converged global model to obtain an assessment result.

2. The method according to claim 1, characterized in that A global model is obtained based on multiple local models of the first stage, including: A plurality of local models of the first stage are weighted averaged and aggregated to obtain the global model.

3. The method according to claim 1, characterized in that The model contrast loss function is: Among them, z prev is the local model of the first stage, z glob is the global model, z is the local model of the second stage, and sim() is a function for calculating the similarity between two features.

4. The method according to claim 1, characterized in that: The dataset is divided into multiple sub-datasets, including: Performing standard normalization and data enhancement processing on the data set to obtain a processed data set; The processed data set is divided into the multiple sub-data sets based on a non-independent and identically distributed algorithm.

5. The method according to claim 4, characterized in that The non-independent and identically distributed algorithm is a Dirichlet distribution algorithm.

6. A method for risk assessment of target data, characterized in that: include: Training a local model based on a sub-dataset to obtain a local model of the first stage; wherein the sub-dataset is any one of a plurality of sub-datasets divided by the server based on the data set; Sending the local model of the first stage to the server, so that the server obtains a global model based on multiple local models of the first stage; Updating the local model of the first stage based on the global model to obtain the local model of the second stage; Determine a model comparison loss function based on the local model of the first stage, the global model, and the local model of the second stage to obtain a loss value, and determine whether the global model converges based on the loss value; The target data is risk assessed based on the converged global model to obtain the assessment results.

7. The method according to claim 6, characterized in that The model contrast loss function is: Among them, z prev is the local model of the first stage, z glob is the global model, z is the local model of the second stage, and sim() is a function for calculating the similarity between two features.

8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program executes the steps of the method described in any one of claims 1 to 5, or executes the steps of the method described in claim 6 when executed by a processor.

9. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to execute the steps of the method according to any one of claims 1 to 5, or to execute the steps of the method according to claim 6.

10. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented, or the steps of the method according to claim 6 are performed.