Data risk assessment method and device, equipment, medium and product

By using the transformer network model to perform risk assessment in financial data, the problems of low risk assessment accuracy and low efficiency in existing technologies are solved, efficient and accurate risk assessment and prompts are achieved, and risk management of financial institutions is supported.

CN120634735APending Publication Date: 2025-09-12INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510716755.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing data risk assessment methods have the problems of low accuracy of risk assessment results and low operational efficiency.

Method used

By obtaining the target data features input by the user and inputting them into the pre-built risk assessment model, the risk assessment results and prompt information are generated for risk warning. The transformer network model is used for training, including encoder and decoder, and the multi-head attention mechanism and normalization layer are used to improve the model's expressiveness and training stability.

Benefits of technology

It improves the accuracy and operational efficiency of risk assessment results, realizes efficient risk assessment of data, and provides support for dynamic risk monitoring and precise intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634735A_ABST
    Figure CN120634735A_ABST
Patent Text Reader

Abstract

The invention discloses a data risk assessment method, device and equipment, a medium and a product. The method comprises the steps of obtaining target data features input by a user; inputting the target data feature into a pre-constructed risk assessment model to obtain a risk assessment result matched with the target data feature; and generating prompt information based on the risk assessment result so as to carry out risk prompt on a user. By means of the technical scheme, risk assessment can be conducted on the data, the risk assessment result of the data is obtained, the accuracy of the risk assessment result is improved, and meanwhile the efficiency of risk assessment operation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a data risk assessment method, device, equipment, medium and product. Background Art

[0002] Amidst the rapid development of the financial industry, the scale of financial transactions continues to snowball. With this growth, the complexity and diversity of data are becoming increasingly prominent. Financial risk control, a key bulwark in safeguarding the stability of the financial system, is tasked with rapidly and efficiently analyzing large amounts of financial data from diverse sources and structures. Its purpose is to discern potential risk events and safeguard the sound operations of financial institutions.

[0003] Traditional financial risk control methods largely rely on rule engines or relatively simple statistical models. Specifically, rule engines operate based on pre-set static rules and lack sufficient adaptability. Once new trading patterns or risk profiles emerge in the financial market, established rules often fail to respond promptly, leading to loopholes in risk monitoring. While traditional statistical models can analyze data to a certain extent, they struggle to accurately capture the complex interplay of relationships within high-dimensional data. This significantly reduces the accuracy of risk identification. Furthermore, with the emergence of a wide variety of innovative financial products, the data underlying trading behavior is becoming increasingly complex. Against this backdrop, relying solely on traditional models can no longer meet the efficiency and accuracy requirements of financial risk control systems.

[0004] In summary, the existing data risk assessment methods have the problems of low accuracy of risk assessment results and low efficiency of risk assessment operations. Summary of the Invention

[0005] The present invention provides a data risk assessment method, device, equipment, medium and product, which can solve the problems of low accuracy of risk assessment results and low efficiency of risk assessment operations in existing data risk assessment methods.

[0006] In a first aspect, an embodiment of the present invention provides a data risk assessment method, the method comprising:

[0007] Obtain target data features input by the user;

[0008] Inputting the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features;

[0009] Prompt information is generated based on the risk assessment result to provide risk prompts to users.

[0010] In a second aspect, an embodiment of the present invention provides a data risk assessment device, the device comprising:

[0011] A data acquisition module is used to obtain target data features input by the user;

[0012] A result acquisition module is used to input the target data characteristics into a pre-built risk assessment model to obtain a risk assessment result that matches the target data characteristics;

[0013] The information generation module is used to generate prompt information based on the risk assessment result to provide risk prompts to users.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, comprising:

[0015] at least one processor; and

[0016] a memory communicatively connected to the at least one processor; wherein,

[0017] The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform a data risk assessment method according to any embodiment of the present invention.

[0018] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a data risk assessment method described in any embodiment of the present invention when executed.

[0019] In a fifth aspect, an embodiment of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements a data risk assessment method described in any embodiment of the present invention.

[0020] The technical solution of the embodiment of the present invention obtains the target data features input by the user, then inputs the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features, and finally generates prompt information based on the risk assessment result to provide risk prompts to the user. This solves the problems of low accuracy of risk assessment results and low efficiency of risk assessment operations in existing data risk assessment methods, realizes risk assessment of data, obtains risk assessment results of data, improves the accuracy of risk assessment results, and improves the efficiency of risk assessment operations.

[0021] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0023] Figure 1 This is a flow chart of a data risk assessment method provided in accordance with the first embodiment of the present invention;

[0024] Figure 2 is a flow chart of a method for constructing a risk assessment model according to a second embodiment of the present invention;

[0025] Figure 3 This is a schematic diagram of the structure of a data risk assessment device provided in accordance with a third embodiment of the present invention;

[0026] Figure 4 The present invention is a schematic diagram of an electronic device for implementing a data risk assessment method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0027] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0028] It should be noted that the terms "first," "second," and the like in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, any variations of the terms "including" and "having" are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units that are explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0029] Example 1

[0030] Figure 1 This is a flowchart of a data risk assessment method provided in Example 1 of the present invention. This embodiment is applicable to situations where data risk assessment is performed. The method can be performed by a data risk assessment device. The data risk assessment device can be implemented in the form of hardware and / or software. The data risk assessment device can be configured in a terminal or server with data risk assessment function.

[0031] like Figure 1 As shown, the method includes:

[0032] S110: Obtain target data features input by the user.

[0033] In this embodiment, the target data features include: at least one of a user identity identifier, a transaction amount, a transaction type, a transaction time, a transaction counterparty identity identifier, a transaction location, a financial product category, a market volatility index, a user credit score, and a macroeconomic indicator.

[0034] Specifically, the user identity identifier is a code used to uniquely identify a user's financial account, such as a bank account number, securities account number, etc., and is the core identifier of the multi-dimensional transaction data of the associated user; the transaction amount is used to represent the monetary value of a specific financial transaction, such as the amount of a transfer or payment, and its size and fluctuation can reflect the risk of abnormal transactions; the transaction type is used to describe the classification of the nature of the transaction, including transfers, payments, deposits, loan applications, etc.; the transaction time records the specific time when the transaction occurs, accurate to the minute, and can be used to analyze transaction time patterns and abnormal period transactions; the counterparty identity identifier is the account identifier of the counterparty, such as the counterparty's bank account number, investor account number, etc., which is used to track the flow of funds and related transaction networks. The transaction location is the physical or online location where the transaction occurs, such as a bank branch, ATM location, online trading platform, etc. The financial product category is the type of financial product in which the user participates, including loans, credit cards, securities, futures, etc. The market volatility index is an indicator used to measure the volatility of financial markets, such as the volatility of the stock market and the interest rate fluctuations in the bond market, reflecting the impact of the macro market environment on risk. The user credit score is a quantitative score calculated based on the user's historical credit record, used to assess the debt repayment ability and credit risk of an individual or enterprise, such as the credit score generated by a bank based on lending and borrowing records. Macroeconomic indicators reflect the overall economic performance, including GDP growth rate, interest rate, inflation rate, etc.

[0035] S120: Input the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features.

[0036] The risk assessment results include: high risk, medium risk and low risk.

[0037] S130: Generate prompt information based on the risk assessment result to provide risk prompts to the user.

[0038] Based on the above steps, the prompt information can be designed in different forms according to the risk level: for high-risk transactions, the system can trigger a real-time SMS alert, prompting the user to confirm the authenticity of the transaction; for medium-risk scenarios, a risk reminder can be displayed on the user's app, suggesting that the counterparty information be verified; for low-risk scenarios, the user can be briefly informed of the transaction compliance. This type of prompt information is intended to help users identify potential risks in a timely manner, while also providing financial institutions with a basis for decision-making on risk prevention and control, enabling dynamic risk monitoring and precise intervention.

[0039] The technical solution of the embodiment of the present invention obtains the target data features input by the user, then inputs the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features, and finally generates prompt information based on the risk assessment result to provide risk prompts to the user. This solves the problems of low accuracy of risk assessment results and low efficiency of risk assessment operations in existing data risk assessment methods, realizes risk assessment of data, obtains risk assessment results of data, improves the accuracy of risk assessment results, and improves the efficiency of risk assessment operations.

[0040] Example 2

[0041] Figure 2 This is a flow chart of a method for constructing a risk assessment model provided in the second embodiment of the present invention. This embodiment is supplemented based on the above embodiment, and specifically supplements the construction process of the risk assessment model.

[0042] like Figure 2 As shown, the method includes:

[0043] S210: Obtain a training sample set.

[0044] The training sample includes at least one sample data and a risk level matching the sample data, and the risk label is used as the annotation data in the training sample.

[0045] Furthermore, the sample data is derived from multi-source heterogeneous financial datasets such as bank transaction records, securities market dynamics, loan approval data, and consumer credit records. For example, bank transaction data includes information such as account access, transfers, and payments (e.g., a user's single transfer amount is 500,000 yuan and the transaction time is 2 a.m.); securities market data covers price fluctuations and trading volumes of stocks, futures, and bonds (e.g., a certain stock's daily volatility reaches 8%); consumer credit data includes loan application records and historical repayment status (e.g., a company's loan is overdue three times); and macroeconomic data involves GDP growth rate, interest rate, inflation rate, etc. (e.g., GDP grew by 4.5% year-on-year in a certain quarter and interest rates increased by 0.5 percentage points). These data are collected through real-time data capture, historical data import, and third-party interface access, and contain dozens or even hundreds of attributes such as user identity identifiers, transaction amounts, and transaction types, which can comprehensively characterize the risk characteristics in financial scenarios.

[0046] Furthermore, the risk level labels are manually labeled to assign risk categories to the sample data, including high risk, medium risk, and low risk.

[0047] S220 : Using the training sample set to train a pre-configured converter network model to obtain a risk assessment model.

[0048] In which, the preconfigured converter network model includes an encoder and a decoder connected in sequence, the encoder includes an encoding feedforward network layer, a normalization layer and an encoding multi-head attention mechanism layer connected in sequence, and the decoder further includes a first normalization layer, a decoding feedforward network layer, a second normalization layer, a decoding multi-head attention mechanism layer, a third normalization layer, a mask multi-head attention layer and a decoding output layer connected in sequence.

[0049] Specifically, in this embodiment, the encoding feedforward network layer is a neural network structure consisting of two fully connected layers and a ReLU activation function, which is used to perform nonlinear transformations on the input financial data features. For example, numerical features such as transaction amounts and credit scores are processed through linear transformations and activation functions to enhance the model's ability to express complex patterns; further, the normalization layer is used to perform layer normalization (LayerNormalization) on the feedforward network output, and to avoid gradient vanishing by adjusting the data distribution to ensure training stability; further, the encoding multi-head attention mechanism layer: the core module captures the importance relationship of each position in the input sequence through a multi-head self-attention mechanism. For example, when analyzing a loan application, the mechanism can automatically calculate the associated weights of features such as transaction time, applicant credit score, and loan amount, generate a global context representation, and identify potential risk patterns (such as the combination of large loans late at night and low credit scores may indicate default risk). The multi-head structure divides the features into multiple subspaces for parallel processing, improving the model's ability to capture data diversity.

[0050] On the basis of the above steps, the first normalization layer, the second normalization layer and the third normalization layer normalize the features in the input and decoding process respectively to maintain the stability of the data distribution; further, the decoding feedforward network layer is similar to the feedforward network in the encoder, and further refines the features through nonlinear transformation to prepare for subsequent attention calculations; the decoding multi-head attention mechanism layer adopts a multi-head cross-attention mechanism, which uses the global features output by the encoder to interact with the current input of the decoder. For example, when predicting the risk level, the mechanism combines the global transaction features extracted by the encoder (such as capital flow, market fluctuations) and the local features of the decoder (such as the current transaction type) to generate a more accurate risk probability distribution; the masked multi-head attention layer masks future information when processing sequence data to avoid information leakage during training and ensure that the model only relies on current and historical data for prediction; the decoding output layer converts the decoder output into a probability distribution of risk level through linear transformation and Softmax function, and finally determines the risk assessment result of the sample.

[0051] Specifically, the pre-configured converter network model is trained using the training sample set to obtain a risk assessment model, including: obtaining a target training sample from the training sample set and inputting the target training sample into the converter network model; performing a parameter adjustment operation on the encoding feedforward network layer in the encoder based on the sample data in the target training sample to adjust the parameters of the encoder; using the adjusted encoder to encode the sample data to obtain a sample vector matching the sample data; performing a parameter adjustment operation on the decoding feedforward network layer in the decoder based on the sample data in the target training sample to adjust the parameters of the decoder; using the adjusted decoder to decode the sample vector to obtain a sample feature matching the sample data; using the encoding output layer to perform a cross-attention calculation on the sample feature to obtain a sample test result matching the sample feature; adjusting the parameters of the converter network model based on the sample test result and the sample label in the target training sample; returning to execute the operation of obtaining the target training sample in the training sample set until the training end condition is met, and determining the trained converter network model as the risk assessment model.

[0052] Specifically, after the target training sample is input into the transformer network (Transformer) model, the parameters of the encoding feedforward network layer in the encoder are first adjusted. The encoding feedforward network layer consists of two fully connected layers and a ReLU activation function. Its function is to perform nonlinear transformation on the input financial data features and enhance the model's ability to express complex patterns. After parameter adjustment, the encoder encodes the sample data through a multi-head self-attention mechanism. The encoder is stacked by multiple encoding layers, each encoding layer containing a multi-head attention sublayer and a feedforward network sublayer, and stabilizes the training process through residual connections and layer normalization. The multi-head attention mechanism decomposes the input features into multiple subspaces for parallel processing. Each head generates attention weights by calculating the dot product of the query, key, and value vectors, thereby capturing the association weights between different features. Subsequently, the decoding feedforward network layer in the decoder is adjusted. It should be noted that in this embodiment, the method for adjusting the decoding feedforward network layer is the same as the method for adjusting the encoding feedforward network layer. After parameter adjustment, the decoder uses a multi-headed cross-attention mechanism to decode the sample vectors. The first multi-headed attention sublayer processes the self-attention relationship of the decoder input, while the second multi-headed attention sublayer fuses the sample vectors output by the encoder with the decoder's current features through a cross-attention mechanism to generate a more accurate representation of risk-related features. For example, the decoder can combine features such as "large transactions" and "low credit scores" extracted by the encoder to further refine sample features indicating "high default probability." The encoder output layer then performs cross-attention calculations on the sample features, generating a probability distribution of risk levels through linear transformations and softmax functions. The cross-attention mechanism allows the decoder to dynamically focus on key features of the encoder output when generating output. For example, when predicting the risk level of a transaction, the encoder output layer calculates the probability of the sample being high, medium, or low risk based on the attention weights of features such as transaction time and market volatility index, i.e., the sample test result.

[0053] Based on the above steps, the loss function is calculated using the backpropagation algorithm based on the sample test results and the true labels of the target training samples, and the parameters of the transformer network model are globally adjusted to update the weights and biases in the encoder and decoder. During training, the process of "obtaining samples → adjusting parameters → encoding and decoding → evaluating losses → updating parameters" is executed repeatedly until the preset end conditions are met (such as reaching the maximum number of iterations of 1000, the loss function converges, or the accuracy of the validation set no longer improves). At this point, the trained transformer network model is determined to be a risk assessment model and can be used for risk identification of real-time financial data, such as generating a risk probability distribution for new transaction data, and assisting the risk control system in automated risk warning and decision-making.

[0054] Among them, the parameter adjustment operation of the encoding feedforward network layer in the encoder is performed based on the sample data in the target training sample, including: initializing the encoding feedforward network layer to obtain the initial weight and initial bias of at least one encoding parameter in the encoding feedforward network layer; obtaining the data attributes of the target training sample, and according to the target training sample, data attributes, and initial loss function respectively matched with each encoding parameter; based on random sampling of normal distribution and the initial weight and initial bias of each encoding parameter, determining the exploration weight and exploration bias value respectively matched with each encoding parameter, and according to the initial weight, initial bias and initial weight matched with each encoding parameter, determining the exploration weight and exploration bias value respectively matched with each encoding parameter, and determining the exploration weight and exploration bias value respectively matched with each encoding parameter, and determining the exploration weight and exploration bias value respectively matched with each encoding parameter, and determining the exploration weight and initial ... The initial bias, exploration weight, exploration bias value, initial loss function and preset migration adjustment coefficient are used to calculate the position change amount that matches each coding parameter respectively; the preset competition intensity adjustment parameter, the parameter for controlling distance sensitivity, and the adaptive parameter influencing factor are obtained, and based on the competition intensity adjustment parameter, the parameter for controlling distance sensitivity, the adaptive parameter influencing factor, the initial weights that match each coding parameter, the initial bias and the position change amount, the parameter adjustment results that match each coding parameter are calculated respectively; based on the parameter adjustment results, each coding parameter is adjusted to perform parameter adjustment on the coding feedforward network layer.

[0055] In a specific implementation scenario of this embodiment, the specific steps of using the training sample set to train the pre-configured converter network model to obtain the risk assessment model may be:

[0056] 1) Initialize the coding feedforward network layer to obtain at least one coding parameter θ in the coding feedforward network layer p The initial weight and initial bias Where n is the serial number of the encoding parameter, and the definition Encoding parameters Parameter position.

[0057] 2) Obtain the data attributes of the target training sample, and generate an initial loss function that matches the target training sample, data attributes, and encoding parameters: Based on the formula The initial loss function of the encoding parameters is calculated, where L i () is the loss function value of the target training sample described in the data attributes; y is the sample label of the target training sample, f p (x i ;W p , b p ) is the processing result of the target training sample processed by the current encoding feedforward network layer; n rer It is the combing of target training samples described in the data attributes; x i is the current training round in the data attributes; L(W p, b p ) is the initial loss function.

[0058] 3) Based on random sampling of the normal distribution and the initial weights and initial biases of the encoding parameters, determining exploration weights and exploration bias values ​​that match the encoding parameters: the exploration weights and exploration bias values ​​of the encoding parameters are based on the formula: Calculated, where the is the exploration position of the encoding parameter, η p is the exploration step size, randn(μ p ,σ p ) is a random sampling based on normal distribution. From step 1), It can be expressed as That is, exploration weight and explore bias values

[0059] Based on the above steps, η p The calculation method is: Based on the formula Calculated, where η0 is the initial step size preset in the migration adjustment coefficient, β p is the step attenuation coefficient in the migration adjustment coefficient, x i is the current training round in the data attributes.

[0060] 4) The position change amount matching the encoding parameters is calculated based on the initial weight, initial bias, exploration weight, exploration bias value, initial loss function and preset migration adjustment coefficient matching the encoding parameters: Based on the formula Calculate the position change Δθ p ; Among them, γ p is the learning rate of the migration decision in the migration adjustment coefficient, α p is the migration intensity adjustment coefficient in the migration adjustment coefficient, is the ecological position gradient.

[0061] Furthermore, the The calculation method is:

[0062]

[0063] 5) Obtaining a preset competition intensity adjustment parameter, a parameter for controlling distance sensitivity, and an adaptive parameter influencing factor, and calculating the parameter adjustment results that match each encoding parameter based on the competition intensity adjustment parameter, the parameter for controlling distance sensitivity, the adaptive parameter influencing factor, the initial weights, initial biases, and position changes that match each encoding parameter: First, calculating a distance calculation function that matches the encoding parameter, the distance calculation function is calculated as follows:

[0064] Among them, σ pes It is a preset parameter for controlling distance sensitivity and can be set by the user.

[0065] Afterwards, based on the formula Calculate the similarity function that matches the encoding parameters and based on the formula

[0066]

[0067] The calculated ecological position change that matches the encoding parameters Among them, λ pas is the competition intensity adjustment parameter.

[0068] Afterwards, the parameter adjustment result matching the encoding parameters is calculated as follows:

[0069] Among them, I p is the adaptive parameter influencing factor, The result of adjusting the encoding parameters.

[0070] 6) Based on the parameter adjustment results, each encoding parameter is adjusted to adjust the encoding feedforward network layer: Easy to get That is, once the adjustment result of the coding parameter is confirmed, the adjusted weight and the adjusted deviation value of the coding parameter can be confirmed. In summary, each coding parameter is adjusted based on the adjusted weight and the adjusted deviation value that match each coding parameter, so as to perform a parameter adjustment operation on the coding feedforward network layer.

[0071] Furthermore, after performing parameter adjustment operations on each encoding parameter based on each parameter adjustment result to adjust the encoding feedforward network layer, it also includes: after detecting that the preset adjustment rules are met, obtaining the preset first sensitivity parameter, second sensitivity parameter, environmental sensitivity, and the current initial weight and initial bias of the decoding feedforward network layer; updating the parameters of the control distance sensitivity of the encoding feedforward network layer according to the first sensitivity parameter, second sensitivity parameter, environmental sensitivity, and the initial loss function, initial weight and initial bias of the encoding feedforward network layer.

[0072] On the basis of the above steps, when the parameter type of the encoding parameter is a parameter for controlling distance sensitivity, due to the changes in the weight and bias values ​​when adjusting the weight and bias values, it can be seen from the calculation method of the initial loss function in the above step 2) that the initial loss function of the parameter for controlling distance sensitivity will change accordingly. In actual applications, in order to adapt to the training progress of the model, the algorithm periodically adjusts the parameter for controlling distance sensitivity, that is, after detecting that the preset adjustment rules are met, based on the formula Update the parameters that control distance sensitivity, where σ pes is the current parameter for controlling distance sensitivity. is the updated parameter for controlling distance sensitivity, β pes is the preset environmental sensitivity, is the rate of change of the feedforward network loss to the environmental parameters.

[0073] Furthermore, the calculation method of the rate of change of the feedforward network loss to the environmental parameters is:

[0074]

[0075] Optionally, the encoding output layer is used to perform cross-attention calculation on the sample features to obtain a sample test result that matches the sample features, including: obtaining the output weight value and output bias value of the encoding output layer; using a preset normalized exponential function to process the sample features, output weight value and output bias value to obtain a sample test result that matches the sample features.

[0076] Specifically, the encoding output layer performs cross attention calculation on the sample features, and the formula P risk =Softmax(WH+b) to obtain the sample test result that matches the sample feature, where H is the sample feature, W and b are the output weight value and output bias value of the encoding output layer, Softmax() is the Softmax function, P risk The results of sample tests are shown in Figure 2.

[0077] The technical solution of the embodiment of the present invention obtains a training sample set and then uses the training sample set to train a pre-configured converter network model to obtain a risk assessment model, thereby realizing the construction of the risk assessment model, improving the accuracy and efficiency of the risk assessment results obtained through the risk assessment model, and at the same time improving the construction efficiency of the risk assessment model.

[0078] Example 3

[0079] Figure 3 This is a structural diagram of a data risk assessment device provided in Example 3 of the present invention.

[0080] like Figure 3 As shown, the device includes:

[0081] The data acquisition module 310 is used to acquire target data features input by the user;

[0082] A result acquisition module 320 is used to input the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features;

[0083] The information generation module 330 is used to generate prompt information based on the risk assessment result to provide risk prompts to the user.

[0084] The technical solution of the embodiment of the present invention obtains the target data features input by the user, then inputs the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features, and finally generates prompt information based on the risk assessment result to provide risk prompts to the user. This solves the problems of low accuracy of risk assessment results and low efficiency of risk assessment operations in existing data risk assessment methods, realizes risk assessment of data, obtains risk assessment results of data, improves the accuracy of risk assessment results, and improves the efficiency of risk assessment operations.

[0085] Based on the above embodiment, the data risk assessment device further includes:

[0086] A model training module is used to obtain a training sample set, wherein the training sample includes at least one sample data and a risk level matching the sample data, and the risk label is used as annotation data in the training sample; the pre-configured transformer network model is trained using the training sample set to obtain a risk assessment model, wherein the pre-configured transformer network model includes an encoder and a decoder connected in sequence, the encoder includes an encoding feedforward network layer, a normalization layer, and an encoding multi-head attention mechanism layer connected in sequence, and the decoder further includes a first normalization layer, a decoding feedforward network layer, a second normalization layer, a decoding multi-head attention mechanism layer, a third normalization layer, a mask multi-head attention layer, and a decoding output layer connected in sequence.

[0087] Based on the above embodiment, the model training module specifically includes:

[0088] A sample acquisition unit, configured to acquire a target training sample from the training sample set and input the target training sample into the converter network model;

[0089] An encoding parameter adjustment unit, configured to perform a parameter adjustment operation on an encoding feedforward network layer in the encoder based on sample data in the target training sample, so as to adjust the parameters of the encoder;

[0090] An encoding unit, configured to perform an encoding operation on the sample data using the adjusted encoder to obtain a sample vector matching the sample data;

[0091] A decoding parameter adjustment unit, configured to perform a parameter adjustment operation on a decoding feedforward network layer in the decoder based on the sample data in the target training sample, so as to adjust the parameters of the decoder;

[0092] A decoding unit, configured to perform a decoding operation on the sample vector using a decoder after parameter adjustment to obtain a sample feature that matches the sample data;

[0093] An encoding output unit, configured to perform a cross-attention calculation on the sample features using the encoding output layer to obtain a sample test result matching the sample features;

[0094] a parameter adjustment unit, configured to adjust parameters of the converter network model based on the sample test results and the sample labels in the target training samples;

[0095] The return execution unit is used to return to execute the operation of obtaining the target training sample in the training sample set until the training end condition is met, and the trained converter network model is determined as the risk assessment model.

[0096] Based on the above embodiment, the encoding parameter adjustment unit includes:

[0097] An initialization unit, configured to initialize the coding feedforward network layer and obtain an initial weight and an initial bias of at least one coding parameter in the coding feedforward network layer;

[0098] A loss function calculation unit, configured to calculate initial loss functions for each encoding parameter based on the target training sample, data attributes, and initial weights and initial biases that match each encoding parameter;

[0099] A random sampling unit, configured to determine exploration weights and exploration bias values ​​that match each encoding parameter based on random sampling from a normal distribution and initial weights and initial biases of each encoding parameter, and to calculate position changes that match each encoding parameter based on the initial weights, initial biases, exploration weights, exploration bias values, initial loss functions, and preset migration adjustment coefficients that match each encoding parameter;

[0100] A competition unit, configured to obtain a preset competition intensity adjustment parameter, a parameter for controlling distance sensitivity, and an adaptive parameter influencing factor, and to calculate, based on the competition intensity adjustment parameter, the parameter for controlling distance sensitivity, the adaptive parameter influencing factor, the initial weights, initial biases, and position changes that match each encoding parameter, to obtain parameter adjustment results that match each encoding parameter;

[0101] The coding feedforward network layer parameter adjustment unit is used to adjust the coding parameters based on the parameter adjustment results to adjust the coding parameters of the coding feedforward network layer.

[0102] On the basis of the above embodiment, the encoding feedforward network layer parameter adjustment unit is also used to: perform parameter adjustment operations on each encoding parameter based on each parameter adjustment result, so as to adjust the encoding feedforward network layer, and after detecting that the preset adjustment rules are met, obtain the preset first sensitivity parameter, second sensitivity parameter, environmental sensitivity, and the current initial weight and initial bias of the decoding feedforward network layer; update the parameters of the control distance sensitivity of the encoding feedforward network layer according to the first sensitivity parameter, second sensitivity parameter, environmental sensitivity, and the initial loss function, initial weight and initial bias of the encoding feedforward network layer.

[0103] Based on the above embodiment, the encoding output unit includes:

[0104] An output layer parameter acquisition unit, configured to acquire an output weight value and an output bias value of the encoding output layer;

[0105] The normalization unit is used to process the sample features, output weight values ​​and output bias values ​​using a preset normalization exponential function to obtain a sample test result that matches the sample features.

[0106] A data risk assessment device provided by an embodiment of the present invention can execute a data risk assessment method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects of the execution method.

[0107] Example 4

[0108] Figure 4A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0109] like Figure 4 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12 and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0110] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0111] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any other suitable processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as a data risk assessment method.

[0112] Accordingly, the method includes:

[0113] Obtain target data features input by the user;

[0114] Inputting the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features;

[0115] Prompt information is generated based on the risk assessment result to provide risk prompts to users.

[0116] In some embodiments, a data risk assessment method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data risk assessment method described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform a data risk assessment method in any other appropriate manner (e.g., by means of firmware).

[0117] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0118] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0119] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0120] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0121] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0122] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0123] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

Claims

1. A data risk assessment method, characterized in that: include: Obtain target data features input by the user; Inputting the target data features into a pre-built risk assessment model to obtain a risk assessment result that matches the target data features; Prompt information is generated based on the risk assessment result to provide risk prompts to users.

2. The method according to claim 1, characterized in that The process of constructing the risk assessment model includes: Acquire a training sample set, wherein the training sample includes at least one sample data and a risk level matching the sample data, and the risk label is used as labeled data in the training sample; The preconfigured transformer network model is trained using the training sample set to obtain a risk assessment model, wherein the preconfigured transformer network model includes an encoder and a decoder connected in sequence, the encoder includes an encoding feedforward network layer, a normalization layer, and an encoding multi-head attention mechanism layer connected in sequence, and the decoder further includes a first normalization layer, a decoding feedforward network layer, a second normalization layer, a decoding multi-head attention mechanism layer, a third normalization layer, a mask multi-head attention layer, and a decoding output layer connected in sequence.

3. The method according to claim 2, characterized in that The pre-configured converter network model is trained using the training sample set to obtain a risk assessment model, including: Obtaining a target training sample from the training sample set, and inputting the target training sample into the converter network model; Performing a parameter adjustment operation on the encoding feedforward network layer in the encoder based on the sample data in the target training sample to adjust the parameters of the encoder; Using the adjusted encoder to encode the sample data, a sample vector matching the sample data is obtained; Performing a parameter adjustment operation on a decoding feedforward network layer in the decoder based on the sample data in the target training sample to adjust the parameters of the decoder; Using the decoder after parameter adjustment to perform a decoding operation on the sample vector to obtain a sample feature that matches the sample data; Using the encoding output layer to perform cross attention calculation on the sample features to obtain a sample test result matching the sample features; Adjusting parameters of the converter network model based on the sample test results and the sample labels in the target training samples; Return to execute the operation of obtaining the target training sample in the training sample set until the training end condition is met, and determine the transformer network model obtained by training as the risk assessment model.

4. The method according to claim 3, characterized in that Performing a parameter adjustment operation on the encoding feedforward network layer in the encoder based on the sample data in the target training sample includes: Initializing the encoding feedforward network layer to obtain an initial weight and an initial bias of at least one encoding parameter in the encoding feedforward network layer; Obtaining data attributes of the target training sample, and matching initial loss functions with the target training sample, data attributes, and encoding parameters; Based on random sampling of the normal distribution and the initial weights and initial biases of each encoding parameter, the exploration weights and exploration bias values ​​that match each encoding parameter are determined, and the position changes that match each encoding parameter are calculated based on the initial weights, initial biases, exploration weights, exploration bias values, initial loss functions, and preset migration adjustment coefficients that match each encoding parameter; Obtaining a preset competition intensity adjustment parameter, a parameter for controlling distance sensitivity, and an adaptive parameter influencing factor, and calculating, based on the competition intensity adjustment parameter, the parameter for controlling distance sensitivity, the adaptive parameter influencing factor, and the initial weight, initial bias, and position change that match each encoding parameter, to obtain parameter adjustment results that match each encoding parameter; Based on the parameter adjustment results, each encoding parameter is adjusted to adjust the parameters of the encoding feedforward network layer.

5. The method according to claim 4, characterized in that After performing parameter adjustment operations on the encoding parameters based on the parameter adjustment results to adjust the parameters of the encoding feedforward network layer, the method further includes: After detecting that a preset adjustment rule is satisfied, obtaining a preset first sensitivity parameter, a second sensitivity parameter, an environmental sensitivity, and a current initial weight and an initial bias of the decoding feedforward network layer; The parameters controlling the distance sensitivity of the coding feedforward network layer are updated according to the first sensitivity parameter, the second sensitivity parameter, the environmental sensitivity, and the initial loss function, the initial weight and the initial bias of the coding feedforward network layer.

6. The method according to claim 3, characterized in that Using the encoding output layer to perform cross attention calculation on the sample features, obtaining a sample test result matching the sample features, including: Obtaining the output weight value and output bias value of the encoding output layer; The sample features, output weight values, and output bias values ​​are processed using a preset normalized exponential function to obtain a sample test result that matches the sample features.

7. A data risk assessment device, characterized in that: include: A data acquisition module is used to obtain target data features input by the user; A result acquisition module is used to input the target data characteristics into a pre-built risk assessment model to obtain a risk assessment result that matches the target data characteristics; The information generation module is used to generate prompt information based on the risk assessment result to provide risk prompts to users.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor. The computer program is executed by the at least one processor so that the at least one processor can perform a data risk assessment method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement a data risk assessment method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements a data risk assessment method according to any one of claims 1 to 6.