A method, device, storage medium and equipment for training a risk identification model

By training the risk identification model to extract and fusion the multimodal data, the problem of low efficiency of multimodal data fusion in the prior art is solved, and more efficient and accurate data fusion and risk identification are achieved.

CN116188023BActive Publication Date: 2025-06-27ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310202513.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-23
Publication Date
2025-06-27
Estimated Expiration
2043-02-23

AI Technical Summary

Technical Problem

In the fusion of multimodal data, the prior art has low efficiency and cannot fuse all modal data at one time, resulting in poor fusion effect.

Method used

By training the risk identification model, historical data of different modes are obtained as training samples, feature extraction and fusion is used using coding subnets, two-dimensional convolutions and three-dimensional convolutions, feature fusion matrix is ​​generated for prediction results, and the model is trained.

Benefits of technology

The one-time full fusion of multimodal data is achieved, the effect and efficiency of data fusion is improved, and the risks of the business to be executed can be more accurately identified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188023B_ABST
    Figure CN116188023B_ABST
Patent Text Reader

Abstract

In the method for training a risk identification model provided in this specification, training samples of multi-modal data are obtained, the annotations of the training samples are determined, and through the encoding subnet of the risk identification model, the data of each modality in the training samples are respectively converted into matrices to obtain respective first feature matrices. A multi-channel second feature matrix is determined based on the respective first feature matrices, and the second feature matrix is input into the three-dimensional convolutional subnet of the model for feature fusion. Based on the feature fusion matrix obtained after fusion, a prediction result is determined, and the model is trained based on the prediction result and the annotations of the training samples. As can be seen from the above method, by converting data of different modalities into matrices and then performing fusion, it facilitates the feature fusion of multi-modal data. The feature fusion matrix obtained after fusion does not solely depend on the data of any one modality, but fully fuses the data of each modality at once, improving the effect and efficiency of data fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computers, and particularly to a method, device, storage medium, and equipment for training a risk recognition model. Background Art

[0002] With the development of Internet technology, more and more attention has been paid to the protection of privacy data. Moreover, by using multi-modal data fusion technology, different modal data are fused to draw on the advantages of different modal data and achieve complementarity between data, which has been widely applied in various fields.

[0003] When performing data fusion in the prior art, the attention mechanism is usually adopted to fuse multi-modal data. However, this method can only fuse different modal data pairwise and cannot fuse all modal data at once, resulting in low efficiency of multi-modal data fusion. And when performing data fusion based on the attention mechanism, the features of one modal data are weighted according to the similarity between the two modal data. The other modal data is not fully utilized. It can be seen that the effect of fusing multiple modal data by this method is not very good.

[0004] Therefore, the present application proposes a method for realizing one-time and full fusion of multi-modal data by training a model. Summary of the Invention

[0005] This specification provides a method, device, storage medium, and electronic device for training a risk recognition model to at least partially solve the above problems.

[0006] This specification adopts the following technical solutions:

[0007] This specification provides a method for training a risk recognition model, and the method includes:

[0008] Obtain historical data of different modalities of historical risk services as training samples, and determine the execution result of the historical risk service as the annotation of the training samples;

[0009] For each modality, use the historical data of this modality as input and input it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality;

[0010] According to the first feature matrices of different modalities, determine a multi-channel second feature matrix, and input the second feature matrix into the three-dimensional convolution subnet of the risk recognition model to perform three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix;

[0011] Input the feature fusion matrix into the prediction subnet of the risk recognition model to obtain a prediction result. According to the prediction result and the annotation corresponding to the training sample, train the risk recognition model. The trained risk recognition model is used to identify whether there is a risk in the to-be-executed service based on data of different modalities of the to-be-executed service.

[0012] Optionally, for each modality, use the historical data of this modality as input and input it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality, specifically including:

[0013] For the text data in the training sample, determine the input data corresponding to each character according to each character in the text data and the position of each character in the text data;

[0014] Input the determined input data into the encoding subnet corresponding to the text modality in the risk recognition model to perform two-dimensional convolution on the text data according to the character and the position;

[0015] Take the matrix obtained after two-dimensional convolution as the first feature matrix of the text data in the training sample.

[0016] Optionally, for each modality, use the historical data of this modality as input and input it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality, specifically including:

[0017] For the structured data in the training sample, determine the input data corresponding to each group of key-value pairs according to the value of each group of key-value pairs in the structured data and the position of each group of key-value pairs in the structured data;

[0018] Input the determined input data into the encoding subnet corresponding to the structured data in the risk recognition model to perform two-dimensional convolution on the structured data according to the value of the key-value pair and the position;

[0019] Take the matrix obtained after two-dimensional convolution as the first feature matrix of the structured data in the training sample.

[0020] Optionally, for each modality, use the historical data of this modality as input and input it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality, specifically including:

[0021] For the image data in the training samples, according to the positions of the pixel points in the image data, input the image data into the encoding subnet corresponding to the image modality in the risk recognition model, and convert the image data into a matrix as the first feature matrix of the image data in the training samples.

[0022] Optionally, determine a multi-channel second feature matrix according to the first feature matrices of different modalities, specifically including:

[0023] For each modality, use the first feature matrix of this modality as the input, input it into the two-dimensional convolutional subnet corresponding to this modality in the risk recognition model, and perform two-dimensional convolution to obtain the third feature matrix of this modality. The sizes of the third feature matrices of each modality are the same;

[0024] Use the obtained third feature matrices of different modalities as data for different channels, and stack the third feature matrices to obtain a multi-channel second feature matrix.

[0025] Optionally, the size of the convolutional matrix in the two-dimensional convolutional subnet of each modality in the risk recognition model is positively correlated with the size of the first feature matrix of the corresponding modality input into the two-dimensional convolutional subnet of each modality.

[0026] Optionally, the trained risk recognition model is used to identify whether there is a risk in the to-be-executed service according to the data of different modalities of the to-be-executed service, specifically including:

[0027] In response to the to-be-executed service, determine the image and text input by the user, where the number of words of the received text input by the user does not exceed the preset number;

[0028] According to the user identification of the user, obtain the structured data required to execute the to-be-executed service;

[0029] According to the image input by the user, the text input by the user, and the structured data, determine the risk recognition result through the risk recognition model;

[0030] Execute the risk control service according to the risk recognition result.

[0031] Optionally, the number of channels of the convolutional matrix in the three-dimensional convolutional subnet of the risk recognition model is equal to the number of channels of the multi-channel second feature matrix.

[0032] This specification provides a device for training a risk recognition model. The device includes:

[0033] An acquisition module, configured to acquire historical data of different modalities of historical risk services as training samples, and determine the execution result of the historical risk services as the annotation of the training samples;

[0034] A two-dimensional convolution module, for each modality, takes the historical data of this modality as input and inputs it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality;

[0035] A three-dimensional convolution module, which is used to determine a multi-channel second feature matrix according to the first feature matrices of different modalities, inputs the second feature matrix into the three-dimensional convolution subnet of the risk recognition model, and performs three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix

[0036] A prediction module, which is used to input the feature fusion matrix into the prediction subnet of the risk recognition model to obtain a prediction result, and train the risk recognition model according to the prediction result and the annotation corresponding to the training sample. The trained risk recognition model is used to identify whether there is a risk in the to-be-executed business according to the data of different modalities of the to-be-executed business.

[0037] This specification provides a computer-readable storage medium, and the storage medium stores a computer program, and when the computer program is executed by a processor, the method for training the above-mentioned risk recognition model is implemented.

[0038] This specification provides an electronic device, including a storage, a processor, and a computer program stored on the storage and executable on the processor. When the processor executes the program, the method for training the above-mentioned risk recognition model is implemented.

[0039] At least one of the above technical solutions adopted in this specification can achieve the following beneficial effects:

[0040] In the method for training the risk recognition model provided in this specification, training samples of multi-modal data are obtained, the annotations of the training samples are determined, and through the encoding subnet of the risk recognition model, the data of each modality in the training samples are respectively converted into matrices to obtain the first feature matrices of each modality. A multi-channel second feature matrix is determined according to the first feature matrices of each modality, the second feature matrix is input into the three-dimensional convolution subnet of the model for feature fusion, a feature fusion matrix is obtained according to the fusion, a prediction result is determined, and the model is trained according to the prediction result and the annotation of the training sample.

[0041] As can be seen from the above method, after converting the data of different modalities into matrices and then performing fusion, it is convenient for the feature fusion of multi-modal data. The obtained feature fusion matrix after fusion does not simply depend on the data of any one of the modalities, but fully fuses the data of each modality at one time, improving the effect and efficiency of data fusion. Brief Description of the Drawings

[0042] The accompanying drawings described herein are used to provide a further understanding of the present specification and form a part of the present specification. The schematic embodiments and descriptions thereof in the present specification are used to explain the present specification and do not constitute an improper limitation to the present specification. In the drawings:

[0043] Figure 1 It is a schematic flowchart of the training process of a risk identification model in the present specification;

[0044] Figure 2 It is a schematic diagram of an encoding subnet provided in the present specification;

[0045] Figure 3 It is a schematic diagram of a two-dimensional convolutional subnet provided in the present specification;

[0046] Figure 4 It is a schematic diagram of a three-dimensional convolutional subnet provided in the present specification;

[0047] Figure 5 It is a schematic diagram of the encoding process of text data provided in the present specification;

[0048] Figure 6 It is a schematic diagram of the encoding process of structured data provided in the present specification;

[0049] Figure 7 It is a schematic diagram of the whole process of fusing data of a risk identification model provided in the present specification;

[0050] Figure 8 It is a schematic diagram of a risk identification model training device provided in the present specification;

[0051] Figure 9 corresponding to that provided in the present specification Figure 1 schematic diagram of an electronic device. Specific Embodiments

[0052] To make the objectives, technical solutions and advantages of the present specification clearer, the technical solutions of the present specification will be clearly and completely described below in conjunction with the specific embodiments of the present specification and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present specification, rather than all the embodiments. Based on the embodiments in the present specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0053] The technical solutions provided in each embodiment of the present specification will be described in detail below in conjunction with the drawings.

[0054] Figure 1 It is a schematic flowchart of a method for training a risk identification model provided in the present specification, which specifically includes the following steps:

[0055] S100: Obtain historical data of different modalities of historical risk services as training samples, and determine the execution result of the historical risk service as the annotation of the training sample.

[0056] The execution entity of the model training provided in this specification can be a server or an electronic device such as a personal computer (PC). For the convenience of description, only the server is used as the execution entity to illustrate the model training method provided in this specification below.

[0057] In the embodiments of this specification, the trained risk identification model is used to determine the risk identification result of the service to be executed based on the data of different modalities of the service to be executed input, so that the service of risk control can be performed on the service to be executed according to the risk identification result. Therefore, when training the risk identification model, it is necessary to first determine the historical risk services of the services that have undergone risk control in history, obtain the historical data of different modalities of the historical risk services as training samples, and use the execution result of performing risk control on the historical risk services in history as the annotation of the training sample. By inputting the training sample into the risk identification model, obtaining the output result of the model, determining the loss function according to the output result and the annotation of the training sample, and training the risk identification model with the goal of minimizing the loss.

[0058] Among them, the historical data of different modalities can include image data, text data, structured data, etc. The model structure of the risk identification model is an encoder-decoder structure. Through this risk identification model, feature fusion can be performed on data of different modalities first, and the risk identification result can be determined based on the fused features. Of course, in the embodiments of this specification, the number of modalities of the input data is not limited, and can be set according to needs, and the corresponding model structure is determined when establishing the risk identification model.

[0059] In the embodiments of this specification, the service to be executed is a service that needs to perform risk identification to determine whether to execute the risk control service. For example, when a transaction service is executed between users and risk identification is required, then the transaction service is the service to be executed. Or, if a user's account is suspected of having risks, then risk identification needs to be performed on the user's account, and the server can automatically initiate a risk identification service for the user's account, and this risk identification service is the service to be executed. Or, when a user's account is frozen due to meeting the risk control rules, or a service initiated by the user is rejected from execution due to meeting the risk control rules, and the user can file an appeal for the frozen account or the service that has been rejected from execution, then the appeal initiated by the user at this time is the service to be executed. Similarly, the historical risk service is the historical service that has undergone risk control for different reasons in history.

[0060] In the embodiments of this specification, taking the historical risk business as a transaction business between users as an example, the server obtains a screenshot provided by the user to prove that the transaction is risk-free as the image data of the historical risk business; obtains the reason input by the user to prove that the current transaction is risk-free as the text data of the historical risk business; and obtains the structured data of the transaction business stored in the server according to the user identification as the structured data of the historical risk business, where the types of structured data are generally preset in advance. Common structured data may include, for example, user level information, user identity information, item category information, etc.

[0061] In the embodiments of this specification, the execution result of the historical risk business is determined based on manual experience according to the historical data of different modalities of the risk business. The server can execute corresponding operations according to the determined execution result of the historical risk business. When the execution result of the historical risk business is that there is a risk, the server can continue to restrict the user's account; when the execution result of the historical risk business is that there is no risk, the server can lift the restriction on the user's account. Using the execution result of the risk business as the annotation of the training sample can better train the risk recognition model. The trained risk recognition model can replace manual prediction of the results of risk operations, greatly improving the efficiency of risk recognition.

[0062] S101: For each modality, use the historical data of this modality as the input and input it into the encoding subnet corresponding to this modality in the risk recognition model to perform two-dimensional convolution to obtain the first feature matrix of this modality.

[0063] In the embodiments of this specification, the risk recognition model is provided with an encoding subnet corresponding to the data of each modality. The first feature matrix of each modality is subjected to two-dimensional convolution through the encoding subnet of each modality to realize feature extraction of the data of this modality. Since the encoding subnet of this modality is used for feature extraction of the data of this modality and does not need to be fused with the data of other modalities, only two-dimensional convolution is required. Of course, for the convenience of subsequent fusion of the features of each modality, the encoding subnet of each modality outputs the result in the form of a feature matrix, that is, the first feature matrix of this modality can be determined through the encoding subnet of this modality. Of course, since the data volumes of different modality data are different, the sizes of the first feature matrices of each modality obtained are generally not exactly the same. In this specification, the sizes of the first feature matrices of each modality are not restricted. By configuring the encoding subnet and restricting the size of the input data, the sizes of the first feature matrices of each modality can be set.

[0064] Specifically, for the convenience of description, in the following, it is taken as an example that each modality includes: image data, text data, and structured data.

[0065] Taking the training samples including text data, image data, and structured data as an example, the method for model training provided in this specification will be described. As Figure 2 shown, after obtaining the training samples, the historical data of each modality in the training samples are respectively input into the encoding subnets of the corresponding modalities in the model. For example, the text data in the training samples is input into the encoding subnet of the corresponding text modality in the model, and this encoding subnet can convert the text data into a matrix. Similarly, the encoding subnets corresponding to other modalities in the model can also convert the historical data of the corresponding modalities in the training samples into matrices. Through the encoding subnets corresponding to each modality in the model, the historical data of different modalities are converted into the form of feature matrices to facilitate the fusion of data of different modalities.

[0066] S102: Determine a multi-channel second feature matrix according to the first feature matrices of different modalities, input the second feature matrix into the three-dimensional convolutional subnet of the risk recognition model, and perform three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix.

[0067] In the embodiments of this specification, after determining the first feature matrices of each modality, the first feature matrices of different modalities can also be stacked as data of different channels to determine a multi-channel second feature matrix, so as to perform three-dimensional convolution on the multi-channel second feature matrix through the three-dimensional convolutional subnet to fuse the features of different modalities and obtain a feature fusion matrix.

[0068] Of course, since each first feature matrix is converted from the historical data of different modalities, the sizes of these first feature matrices may be different. Therefore, after the server stacks the first feature matrices of different sizes and different modalities to obtain the second feature matrix, it can determine the two-dimensional size of the second feature matrix. For example, determine the maximum length and maximum width of the second feature matrix, and then fill the features of each channel of the second feature matrix. For example, assume that the two-dimensional size of the second feature matrix is 100×100, and the size of a first feature matrix as a certain channel is 50×50. Numerical filling is performed on the first feature matrix of this channel to obtain a 100×100 feature matrix as the feature of this channel in the second feature matrix.

[0069] Then the sizes of each channel of the obtained second feature matrix are unified, and three-dimensional convolution is performed through the three-dimensional convolutional subnet.

[0070] However, by filling in the specified value, although the size of the feature matrix of each channel of the second feature matrix can be unified, the specified value filled in cannot represent any feature, resulting in the sparsity of the features of the second feature matrix. Moreover, when performing three-dimensional convolution, it is also difficult to ensure that the features of each modality can be averaged and fused.

[0071] Therefore, in order to obtain a more accurate fused feature matrix during three-dimensional convolution, the server can also convert each first feature matrix of different sizes into each third feature matrix of the same size through the two-dimensional convolution subnets corresponding to each modality in the model. Therefore, when the convolution matrix of the two-dimensional convolution subnet of the model is convolved with the first feature matrix, the convolution mode adopted is a convolution mode that can change the matrix size before and after convolution. The specific convolution mode is not limited in this specification.

[0072] Specifically, after the encoding subnets corresponding to each modality output each matrix, the matrices are used as the first feature matrices of the historical data of each modality, and each first feature matrix is respectively input into the two-dimensional convolution subnet of the corresponding modality in the model, as Figure 3 shown, Figure 3 the part surrounded by the dashed box in the figure is each third feature matrix obtained through the two-dimensional convolution subnet. For example, the first feature matrix of the image modality is input into the two-dimensional convolution subnet of the corresponding image modality in the model, and the convolution matrix in the two-dimensional convolution subnet is convolved with the first feature matrix to obtain the third feature matrix of the image modality.

[0073] Similarly, the two-dimensional convolution subnets corresponding to other modalities in the model can also convert the first feature matrices of other modalities into third feature matrices. The sizes of the convolution matrices of each modality are different. Since the sizes of the first feature matrices of different modalities can be determined according to the input data, such as the upper limit of the character length, the maximum resolution of the picture, and the number of key-value pairs in the structured information. And the size of the third feature matrix can be determined in advance. Therefore, for each modality, the size of the convolution matrix of the modality can be determined according to the size of the first feature matrix of the modality and the preset size of the third feature matrix. Among them, the size of the convolution matrix is positively correlated with the size of the first feature matrix, and the sizes of the third feature matrices are the same.

[0074] In the embodiments of this specification, after obtaining the third feature matrices of each modality with the same size, the third feature matrices can be fused into a feature fusion matrix, and the feature fusion matrix contains the features of the historical data of each modality input into the risk recognition model.

[0075] As Figure 4As shown, since the third feature matrices of each modality need to be fused and the sizes of the third feature matrices of each modality are the same, the server can stack the third feature matrices of each modality as data of different channels to obtain a second feature matrix, and input the second feature matrix into the three-dimensional convolutional subnet of the model for three-dimensional convolution to obtain a feature fusion matrix, which includes both the fusion of data of the same modality by the three-dimensional convolutional subnet and the fusion of data of different modalities by the three-dimensional convolutional subnet.

[0076] Of course, the embodiments of this specification are described by taking the data of three modalities as an example. Therefore, the convolutional matrix of the three-dimensional convolutional subnet is three-channel. When the number of modalities corresponding to the input data increases, the number of channels of the convolutional matrix of the three-dimensional convolutional subnet can be set correspondingly. That is to say, the number of channels of the convolutional matrix of the three-dimensional convolutional subnet is equal to the number of channels of the multi-channel second feature matrix.

[0077] S103: Input the feature fusion matrix into the prediction subnet of the risk recognition model to obtain a prediction result, and train the risk recognition model according to the prediction result and the annotation corresponding to the training sample. The trained risk recognition model is used to identify whether there is a risk in the to-be-executed service according to the data of different modalities of the to-be-executed service.

[0078] In the embodiments of this specification, after obtaining the feature fusion matrix, the server can input the feature fusion matrix into the prediction subnet of the risk recognition model. The prediction subnet can output a prediction result according to the feature fusion matrix to predict whether there is a risk in the historical risk service, compare the prediction result with the annotation of the training sample to determine the difference between the two, that is, determine the loss function, and train the risk recognition model with the goal of minimizing the loss function.

[0079] The trained risk identification model can be used to determine whether there is a risk in the business to be executed. When the server responds to the business to be executed by the user, for example, when the server determines that there is a risk in the transaction executed by the user account and restricts the subsequent transaction activities of the user account, the user can send an appeal request to the server through the terminal to request the lifting of the restriction on the user account. Then, the server can send an appeal interface to the user's terminal, and the user can fill in the appeal information on the appeal interface. The appeal information includes data in different modalities. For example: text data, that is, the reason for the user's appeal; image data, that is, the screenshots provided by the user, etc. Then, according to the user identification of the user or the business identification of this type of business to be executed, the structured data required for executing this business to be executed is obtained. The obtained structured data and the appeal information are used as inputs and input into the trained risk identification model. According to the prediction result output by the risk identification model, it is determined whether there is a risk in the user account. If there is a risk, the subsequent transaction activities of the user account are continuously restricted; if there is no risk, the restriction on the user account is cancelled.

[0080] Based on Figure 1 In the risk identification model training method shown in the foregoing, in the embodiments of this specification, by using historical data in different modalities of historical risk identification services, a risk identification model is trained. The risk identification model first converts the historical data in different modalities into matrices in a unified form, then performs convolution on each matrix to unify the sizes of the matrices, then fuses the matrices in different modalities with the same size obtained, makes a prediction on the feature fusion matrix obtained after fusion, and finally continuously adjusts the parameters of the risk identification model according to the loss between the prediction result and the annotation. The trained risk identification model can accurately determine whether there is a risk in the risk identification service to be executed. The risk identification model realizes the full fusion of data in different modalities at one time, improving the effect and efficiency of data fusion.

[0081] In step S101, the text data in the training sample is input into the encoding subnet corresponding to the text modality in the model. Specifically, according to each character in the text data, the position of each character in the text data, and the data type identifier of the text data, the input data corresponding to each character is determined. According to the order of each character in the text data, the determined input data is concatenated to obtain an input sequence. The input sequence is input into the encoding subnet corresponding to the text modality in the risk identification model to encode the text data according to each character and the position of each character. The matrix obtained after encoding is used as the first feature matrix of the text data in the training sample.

[0082] For example, as Figure 5The text data shown is "The transaction party is a new customer", the data type identifier of the text data is "1", there are seven characters in total, the identifier of the word "交" in the text data is f11, and the position of this character in the text data is 1, the identifier of the word "易" in the text data is f12, and the position of this character in the text data is 2, the identifier of the word "方" in the text data is f13, and the position of this character in the text data is 3, and so on, the character identifier and position identifier of each character in the text data are determined, and the character identifier, position identifier and data type identifier of each character are used as input data. According to the order of each character in the text data, the determined input data are spliced ​​to obtain the input sequence [(f11, 1, 1), (f12, 1, 2), (f13, 1, 3), (f14, 1, 4), (f15, 1, 5), (f16, 1, 6), (f17, 1, 7)].

[0083] Similarly, the structured data in the training sample is input into the encoding subnet corresponding to the structured data in the model. Specifically, the input data corresponding to each group of key-value pairs can be determined according to the value of each group of key-value pairs in the structured data, the position of each group of key-value pairs in the structured data, and the data type identifier of the structured data. The determined input data is input into the encoding subnet corresponding to the structured data in the risk identification model to encode the structured data according to the value of the key-value pair and the position, and the matrix obtained after encoding is used as the first feature matrix of the structured data in the training sample.

[0084] In the embodiment of this specification, the structured data obtained by the server may be structured data corresponding to the user identifier or structured data corresponding to the service identifier. The following is an example of obtaining structured data corresponding to the user identifier. Among the structured data corresponding to the user identifier stored on the server, the structured data corresponding to the user identifier of the user of the historical risk identification service is determined. For example, the user who performs the historical risk identification service is user 2. Figure 6 As shown, the server obtains the structured data of user 2, the data type of the structured data is identified as "0", and determines the input data corresponding to each key-value pair according to the value of each key-value pair in the structured data, the position of each key-value pair in the structured data, and the data type identifier of the structured data.

[0085] Similarly, the image data in the training sample is input into the encoding subnet of the corresponding image modality in the model. Specifically, according to the position of each pixel in the image data, the image data can be input into the encoding subnet of the corresponding image modality in the risk identification model, and the image data can be converted into a matrix as the first feature matrix of the image data in the training sample.

[0086] In the embodiments of this specification, for each modality, in order to avoid the difficulty in converting the first feature matrices of different modalities into third feature matrices of a unified size due to different data sizes corresponding to different services for the same modality, it is necessary to limit the maximum size of the data for each modality. For example, the resolution of image data is limited to 256×256, and the text data is limited to 400 words. Of course, since the key-value pairs included in the structured data are all preset, when setting the key-value pairs included in the structured data, the size of the structured data is already limited.

[0087] In addition, since the data of each modality provided by the user during the application process may not match the maximum size of the data of each modality restricted, the data obtained for the risk identification service to be executed can be cropped, extracted, or padded before being input into the risk identification model. For example, the resolution of image data is limited to 256×256, while the resolution of the screenshot provided by the user is 200×200. In this case, the screenshot needs to be padded to make its resolution 256×256. There are various ways to pad the screenshot, which are not specifically limited in this specification. However, when cropping, extracting, or padding the data before it is input into the risk identification model, it may occur that the key information in the data is cropped off or the key information in the data is not extracted. Therefore, to ensure the integrity of the data and thus achieve a better data fusion effect, the size of the appeal information filled in by the user can be restricted in the appeal interface. For example, the number of words that the user can input in the appeal interface does not exceed 200 words, and the resolution of the screenshot provided by the user in the appeal interface is 256×256, etc.

[0088] The traditional method of fusing data using the attention mechanism only weights the data of one of the input modalities, where the weights are determined based on the correlation between the data of each input modality. In fact, it is still a transformation of the data of one of the modalities and does not fully utilize the data of each modality for feature fusion. However, the method provided in this specification, as Figure 7 shown, is to convert the data of each modality into matrices of the same size, and then regard each matrix as data of different channels for fusion, which can fully fuse the data of each modality and is not restricted by the limitation in the traditional method that only two different modalities can be fused pairwise, thus improving the efficiency of data fusion.

[0089] The above is the method for training the risk identification model provided by one or more embodiments of this specification. Based on the same idea, this specification also provides a corresponding risk identification model training device, as Figure 8 shown.

[0090] Figure 8 It is a schematic diagram of a risk identification model training device provided by this specification, specifically including:

[0091] An acquisition module, configured to acquire historical data of different modalities of historical risky services as training samples, and determine the execution results of the historical risky services as the annotations of the training samples;

[0092] A two-dimensional convolution module, configured to, for each modality, use the historical data of this modality as input, input it into the encoding subnet corresponding to this modality in the risk recognition model, and perform two-dimensional convolution to obtain the first feature matrix of this modality;

[0093] A three-dimensional convolution module, configured to determine a multi-channel second feature matrix according to the first feature matrices of different modalities, input the second feature matrix into the three-dimensional convolution subnet of the risk recognition model, and perform three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix

[0094] A prediction module, configured to input the feature fusion matrix into the prediction subnet of the risk recognition model to obtain a prediction result, and train the risk recognition model according to the prediction result and the annotation corresponding to the training sample. The trained risk recognition model is used to identify whether there is a risk in the to-be-executed service according to the data of different modalities of the to-be-executed service.

[0095] Optionally, the two-dimensional convolution module 802 is specifically configured to, for the text data in the training sample, determine the input data corresponding to each character according to each character in the text data and the position of each character in the text data; input the determined input data into the encoding subnet corresponding to the text modality in the risk recognition model to perform two-dimensional convolution on the text data according to the character and the position; and use the matrix obtained after two-dimensional convolution as the first feature matrix of the text data in the training sample.

[0096] Optionally, the two-dimensional convolution module 802 is specifically configured to, for the structured data in the training sample, determine the input data corresponding to each group of key-value pairs according to the value of each group of key-value pairs in the structured data and the position of each group of key-value pairs in the structured data; input the determined input data into the encoding subnet corresponding to the structured data in the risk recognition model to perform two-dimensional convolution on the structured data according to the value of the key-value pair and the position; and use the matrix obtained after two-dimensional convolution as the first feature matrix of the structured data in the training sample.

[0097] Optionally, the two-dimensional convolution module 802 is specifically configured to, for the image data in the training sample, according to the positions of the pixel points in the image data, input the image data into the encoding subnet corresponding to the image modality in the risk recognition model, and convert the image data into a matrix as the first feature matrix of the image data in the training sample.

[0098] Optionally, the three-dimensional convolution module 803 is specifically configured to, for each modality, use the first feature matrix of that modality as the input, input it into the two-dimensional convolution subnet corresponding to that modality in the risk recognition model, perform two-dimensional convolution to obtain the third feature matrix of that modality, and the sizes of the third feature matrices of each modality are the same; use the obtained third feature matrices of different modalities as data of different channels, and stack the third feature matrices to obtain a multi-channel second feature matrix third feature matrix second feature matrix.

[0099] Optionally, the size of the convolution matrix in the two-dimensional convolution subnet of each modality in the risk recognition model is positively correlated with the size of the first feature matrix of the corresponding modality input into the two-dimensional convolution subnet of each modality.

[0100] Optionally, the prediction module 804 is specifically configured to, in response to the business to be executed, determine the image and text input by the user, where the number of words input by the user received does not exceed a preset number; obtain the structured data required to execute the business to be executed according to the user identification of the user; determine the risk recognition result through the risk recognition model according to the image input by the user, the text input by the user, and the structured data; and execute the risk control business according to the risk recognition result.

[0101] Optionally, the number of channels of the convolution matrix of the three-dimensional convolution subnet of the risk recognition model is equal to the number of channels of the multi-channel second feature matrix.

[0102] This specification also provides a computer-readable storage medium, which stores a computer program that can be used to execute the above Figure 1 Provided risk recognition model training method.

[0103] This specification also provides Figure 9 The structural schematic diagram of the electronic device shown. As Figure 9 Described, at the hardware level, the interface matching device includes a processor, an internal bus, a network interface, a memory, and a non-volatile memory. Of course, it may also include other hardware required for other services. The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it to implement the above Figure 1The described risk identification model training method. Of course, in addition to the software implementation method, this specification does not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.

[0104] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to circuit structures such as diodes, transistors, switches, etc.) or software improvements (improvements to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Designers almost always obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented using a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is an integrated circuit whose logic function is determined by the user programming the device. Designers can program themselves to "integrate" a digital system onto a single PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compilers used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called a Hardware Description Language (HDL), and there is not just one type of HDL, but many, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. The most commonly used ones currently are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be aware that by simply performing a little logical programming on the method flow using the above-mentioned several hardware description languages and programming it into an integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.

[0105] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0106] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0107] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.

[0108] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.

[0109] The present invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block of the flowchart illustrations and / or block diagrams, and combinations of flows and / or blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general purpose computer, special purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing apparatus create means for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks for implementing the specified functions.

[0110] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer-readable memory produce an article of manufacture including instruction means that implement the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks for implementing the specified functions.

[0111] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in the flow Figure 1 one or more flows and / or blocks Figure 1 one or more blocks for implementing the specified functions.

[0112] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0113] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read only memory (ROM) or flash memory (flash RAM). Memory is an example of computer-readable media.

[0114] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0115] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0116] Those skilled in the art will appreciate that the embodiments of this specification may be provided as methods, systems or computer program products. Therefore, this specification may take the form of a complete hardware embodiment, a complete software embodiment or an embodiment combining software and hardware. Moreover, this specification may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0117] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0118] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiment.

[0119] The above is only the embodiment of this specification and is not used to limit this specification. For those skilled in the art, various modifications and changes can be made to this specification. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification shall be included in the scope of the claims of this application.

Claims

1. A method for training a risk identification model, the method comprising: Obtaining historical data of different modalities of historical risk services as training samples, and determining the execution results of the historical risk services as the annotations of the training samples; For each modality, taking the historical data of this modality as input and inputting it into the encoding subnet corresponding to this modality in the risk identification model to perform two-dimensional convolution to obtain the first feature matrix of this modality; According to the first feature matrices of different modalities, determining a multi-channel second feature matrix, inputting the second feature matrix into the three-dimensional convolution subnet of the risk identification model, and performing three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix; The number of channels of the convolution matrix of the three-dimensional convolution subnet of the risk identification model is equal to the number of channels of the multi-channel second feature matrix; Inputting the feature fusion matrix into the prediction subnet of the risk identification model to obtain a prediction result, and training the risk identification model according to the prediction result and the annotation corresponding to the training sample. The trained risk identification model is used to identify whether there is a risk in the service to be executed according to the data of different modalities of the service to be executed; Wherein, the determining the multi-channel second feature matrix according to the first feature matrices of different modalities specifically includes: For each modality, taking the first feature matrix of this modality as input and inputting it into the two-dimensional convolution subnet corresponding to this modality in the risk identification model to perform two-dimensional convolution to obtain the third feature matrix of this modality. The sizes of the third feature matrices of each modality are the same; taking the obtained third feature matrices of different modalities as data of different channels, and superimposing each third feature matrix to obtain a multi-channel second feature matrix; Alternatively, after superimposing the first feature matrices of different sizes and different modalities to obtain the second feature matrix, determining the two-dimensional size of the second feature matrix, and padding the feature matrix of each channel of the second feature matrix according to the two-dimensional size to make the sizes of the feature matrices of each channel of the second feature matrix uniform.

2. The method according to claim 1, for each modality, taking the historical data of this modality as input and inputting it into the encoding subnet corresponding to this modality in the risk identification model to perform two-dimensional convolution to obtain the first feature matrix of this modality, specifically including: For the text data in the training sample, determining the input data corresponding to each character according to each character in the text data and the position of each character in the text data; Inputting the determined input data into the encoding subnet corresponding to the text modality in the risk identification model to perform two-dimensional convolution on the text data according to the character and the position; Taking the matrix obtained after two-dimensional convolution as the first feature matrix of the text data in the training sample.

3. The method according to claim 1, for each modality, taking the historical data of this modality as input and inputting it into the encoding subnet corresponding to this modality in the risk identification model to perform two-dimensional convolution to obtain the first feature matrix of this modality, specifically including: For the structured data in the training sample, determine the input data corresponding to each key-value pair according to the value of each key-value pair in the structured data and the position of each key-value pair in the structured data. Input the determined input data into the encoding subnet corresponding to the structured data in the risk identification model to perform two-dimensional convolution on the structured data according to the value and position of the key-value pair. Use the matrix obtained after two-dimensional convolution as the first feature matrix of the structured data in the training sample.

4. The method according to claim 1, for each modality, use the historical data of the modality as input, input it into the encoding subnet corresponding to the modality in the risk identification model, and perform two-dimensional convolution to obtain the first feature matrix of the modality, specifically including: For the image data in the training sample, according to the positions of the pixels in the image data, input the image data into the encoding subnet corresponding to the image modality in the risk identification model, and convert the image data into a matrix as the first feature matrix of the image data in the training sample.

5. The method according to claim 1, the size of the convolution matrix in the two-dimensional convolution subnet of each modality in the risk identification model is positively correlated with the size of the first feature matrix of the corresponding modality input into the two-dimensional convolution subnet of each modality.

6. The method according to claim 2, the trained risk identification model is used to identify whether there is a risk in the to-be-executed business according to the data of different modalities of the to-be-executed business, specifically including: In response to the to-be-executed business, determine the image and text input by the user, where the number of the received text input by the user does not exceed the preset number. According to the user identification of the user, obtain the structured data required to execute the to-be-executed business. Determine the risk identification result through the risk identification model according to the image input by the user, the text input by the user, and the structured data. Execute the risk control business according to the risk identification result.

7. An apparatus for training a risk identification model, the apparatus includes: An acquisition module, configured to acquire the historical data of different modalities of historical risk services as training samples, and determine the execution result of the historical risk services as the annotation of the training samples. A two-dimensional convolution module, configured to, for each modality, use the historical data of the modality as input, input it into the encoding subnet corresponding to the modality in the risk identification model, and perform two-dimensional convolution to obtain the first feature matrix of the modality. A three-dimensional convolution module, configured to determine a multi-channel second feature matrix according to the first feature matrices of different modalities, input the second feature matrix into the three-dimensional convolution subnet of the risk identification model, and perform three-dimensional convolution on the second feature matrix to obtain a feature fusion matrix. The number of channels of the convolution matrix in the three-dimensional convolution subnet of the risk identification model is equal to the number of channels of the multi-channel second feature matrix. A prediction module, configured to input the feature fusion matrix into a prediction subnet of the risk identification model to obtain a prediction result, and train the risk identification model according to the prediction result and the annotation corresponding to the training sample. The trained risk identification model is used to identify whether there is a risk in the to-be-executed service according to data of different modalities of the to-be-executed service; Wherein, determining the multi-channel second feature matrix according to the first feature matrices of different modalities specifically includes: For each modality, taking the first feature matrix of this modality as an input, inputting it into the two-dimensional convolutional subnet corresponding to this modality in the risk identification model, performing two-dimensional convolution to obtain the third feature matrix of this modality, and the sizes of the third feature matrices of each modality are the same; taking the obtained third feature matrices of different modalities as data of different channels, and superimposing each third feature matrix to obtain a multi-channel second feature matrix; Alternatively, after superimposing the first feature matrices of different sizes and different modalities to obtain the second feature matrix, determining the two-dimensional size of the second feature matrix, and padding the feature matrix of each channel of the second feature matrix according to this two-dimensional size to make the sizes of the feature matrices of each channel of the second feature matrix uniform.

8. A computer-readable storage medium, storing a computer program, which when executed by a processor, implements the method according to any one of claims 1 to 6 above.

9. An electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method according to any one of claims 1 to 6 above when executing the program.

Citation Information

Patent Citations

  • Multi-mode complex activity recognition method based on deep learning model

    CN108960337A

  • Risk assessment method and device

    CN115222405A