User tag determination method, device and server

By acquiring and analyzing transaction data and market data, and using hierarchical clustering and Marshall distance improvement models for adversarial training, the problems of high cost, low efficiency and low accuracy of model training in the existing technology are solved, and accurate user tag determination is achieved in the transaction data processing scenario.

CN112836743BActive Publication Date: 2025-05-09INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110141802.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-02
Publication Date
2025-05-09
Estimated Expiration
2041-02-02

AI Technical Summary

Technical Problem

The prior art requires manual labeling of large amounts of sample data when training prediction models, resulting in high cost and low efficiency. The labeling process is subjectively affected by people, resulting in low model accuracy and inability to be applicable to real transaction data processing scenarios.

Method used

By obtaining target transaction data and associated target market data, using hierarchical clustering to obtain preset market categories, further using improved model based on Mahjong distance for confrontation training, obtaining the preset trading behavior characteristic prediction model, and then determining the user tag of the target user.

Benefits of technology

It realizes accurate determination of the user tag of the target user in the transaction data processing scenario, reduces the cost of model training, improves training efficiency, and reduces the error of user tag determination.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112836743B_ABST
    Figure CN112836743B_ABST
Patent Text Reader

Abstract

This specification provides a method, device and server for determining user tags. Based on this method, when determining the user tag of a target user, first obtain the target transaction data corresponding to the target user and the target market data associated with the target transaction data; then determine the matching target market category from multiple preset market categories obtained in advance through hierarchical clustering based on the target market data; then determine the matching target transaction behavior feature prediction model from multiple preset transaction behavior feature prediction models obtained in advance through adversarial training using an improved model based on Mahalanobis distance; then process the target transaction data by calling the target transaction behavior feature prediction model, and determine the user tag of the target user based on the determined transaction behavior features. In this way, the transaction behavior features can be accurately obtained and used to more accurately determine the user tag of the target user in the transaction data processing scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification belongs to the field of artificial intelligence technology, and in particular to a method, device and server for determining user tags. Background Art

[0002] In many transaction data processing scenarios (for example, stock trading scenarios or precious metal trading scenarios, etc.), it is often necessary to use a pre-trained prediction model to predict user tags, and then perform subsequent specific data processing based on the user tags.

[0003] When training the above prediction model based on the existing model training method, it is often necessary to manually label the acquired sample data with a large amount of data first, and then use the labeled sample data for specific model training. This will inevitably make the training process of the prediction model more costly and less efficient. In addition, when training the prediction model based on the above method, the sample data is often labeled according to artificially set rules, which will inevitably make the labeling process subject to human subjective influence, resulting in the labeling process itself being not comprehensive and objective, and prone to omissions. The prediction model obtained based on the above labeled sample data training has low accuracy and cannot be well applied to real transaction data processing scenarios. Furthermore, if the above prediction model is used to determine user labels, technical problems such as large errors and inaccuracies in determining user labels in transaction data processing scenarios are bound to occur.

[0004] Currently, no effective solution has been proposed to the above problems. Summary of the invention

[0005] This specification provides a method, device and server for determining a user tag, so as to more precisely and accurately determine the user tag of a target user in a transaction data processing scenario.

[0006] This specification provides a method for determining a user tag, including:

[0007] Acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user;

[0008] According to the target market data, a matching target market category is determined from a plurality of preset market categories; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data;

[0009] Determining a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training sample data of the corresponding market category using an improved model based on Mahalanobis distance;

[0010] Calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data;

[0011] A user tag of a target user is determined according to the transaction behavior characteristics of the target transaction data.

[0012] In one embodiment, after determining the user tag of the target user, the method further includes:

[0013] Generate risk warning information matching the target user based on the target user's user tag;

[0014] The risk warning information is sent to the target user.

[0015] In one embodiment, the method further comprises:

[0016] Acquire historical transaction data as the second type of sample data; and acquire historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label;

[0017] According to the association relationship between the historical transaction data and the historical market data, a correspondence relationship between the first type of sample data and the second type of sample data is established; wherein the first type of sample data corresponds to one or more second type of sample data;

[0018] Performing hierarchical clustering on the first category sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first category sample data;

[0019] Determine the second type of sample data corresponding to each market category according to the first type of sample data included in the plurality of category groups and the correspondence between the first type of sample data and the second type of sample data, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively;

[0020] An improved model based on Mahalanobis distance was constructed as the initial model;

[0021] The initial model is respectively subjected to preset virtual adversarial training using the multiple training data sets to obtain multiple preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category.

[0022] In one embodiment, the market category includes at least one of the following: rising market, falling market, and volatile market.

[0023] In one embodiment, the initial model at least includes an improved loss function based on Mahalanobis distance.

[0024] In one embodiment, an improved model based on Mahalanobis distance is constructed as an initial model, including:

[0025] Using the improved loss function based on Mahalanobis distance, the corresponding objective function is constructed to obtain the initial model;

[0026] The improved loss function based on Mahalanobis distance is constructed as follows:

[0027]

[0028] in, is the improved loss function based on Mahalanobis distance in the initial model, D l D is the sample data set composed of the second type of sample data with transaction behavior feature labels. ul N is the sample data set composed of the second type of sample data without transaction behavior feature labels. l D l The total number of the second type of sample data contained in, N ul D ul The total number of the second type of sample data contained in , x is the second type of sample data, θ is the independent variable of the loss function, and LMS is the Mahalanobis smoothness.

[0029] In one embodiment, Marano smoothness is determined according to the following formula:

[0030]

[0031]

[0032] Among them, x i is the transaction behavior feature label of the second type of sample data numbered i, is the specific data value in the independent variable of the loss function, r MDVAT is the perturbation term based on Mahalanobis distance, is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set, p(x i +rMDVAT ,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after being disturbed by the disturbance term based on the Mahalanobis distance, md[] represents the calculation function of the Mahalanobis distance, r is the random disturbance, p(x i +r,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after random perturbation.

[0033] In one embodiment, the initial model is subjected to preset virtual adversarial training using the multiple training data sets to obtain multiple preset transaction behavior feature prediction models, including:

[0034] The initial model is subjected to a preset virtual adversarial training using a current training data set among the multiple training data sets in the following manner to obtain a preset transaction behavior feature prediction model corresponding to the current market category:

[0035] Using the initial model to process the second type of sample data in the current data set that does not carry the transaction behavior feature label, to obtain pseudo labels corresponding to the second type of sample data that does not carry the transaction behavior feature label;

[0036] Combining the second type of sample data carrying the transaction behavior feature label in the current training data set with the second type of sample data carrying the pseudo label to obtain combined training data;

[0037] The combined training data is used to train the initial model to obtain a preset transaction behavior feature prediction model corresponding to the current market category.

[0038] In one embodiment, the transaction behavior characteristic label includes at least one of the following: an aggressive label, a conservative label, and a risk-averse label.

[0039] In one embodiment, the target transaction data includes multiple target transaction data corresponding to the same target user; accordingly, determining the user tag of the target user according to the transaction behavior characteristics of the target transaction data includes:

[0040] Combining the transaction behavior features of the plurality of target transaction data to obtain a combined transaction behavior feature;

[0041] A user tag of the target user is determined according to the combined transaction behavior characteristics.

[0042] In one embodiment, the target transaction data includes at least one of the following: stock transaction data, precious metal transaction data, and crude oil transaction data.

[0043] This specification also provides a method for training a preset transaction behavior feature prediction model, including:

[0044] Acquire historical transaction data as the second type of sample data; and acquire historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label;

[0045] According to the association relationship between the historical transaction data and the historical market data, a correspondence relationship between the first type of sample data and the second type of sample data is established; wherein the first type of sample data corresponds to one or more second type of sample data;

[0046] Performing hierarchical clustering on the first category sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first category sample data;

[0047] Determine the second type of sample data corresponding to each market category according to the first type of sample data included in the plurality of category groups and the correspondence between the first type of sample data and the second type of sample data, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively;

[0048] An improved model based on Mahalanobis distance was constructed as the initial model;

[0049] The initial model is respectively subjected to preset virtual adversarial training using the multiple training data sets to obtain multiple preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category.

[0050] This specification also provides a device for determining a user tag, including:

[0051] An acquisition module, used to acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user;

[0052] A first determination module is used to determine a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data;

[0053] A second determination module is used to determine a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training sample data of the corresponding market category using an improved model based on Mahalanobis distance;

[0054] A calling module, used for calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data;

[0055] The third determination module is used to determine the user tag of the target user according to the transaction behavior characteristics of the target transaction data.

[0056] The present specification also provides a server, including a processor and a memory for storing processor executable instructions, wherein the processor implements the relevant steps of the method for determining the user tag when executing the instructions.

[0057] The present specification also provides a computer-readable storage medium on which computer instructions are stored. When the instructions are executed, the relevant steps of the method for determining the user tag are implemented.

[0058] The present specification provides a method, device and server for determining a user label. When determining the user label of a target user, the target transaction data corresponding to the target user and the target market data associated with the target transaction data are obtained; then, according to the target market data, a matching target market category is determined from a plurality of preset market categories obtained in advance through hierarchical clustering; then, a matching target transaction behavior feature prediction model can be determined from a plurality of preset transaction behavior feature prediction models obtained in advance through adversarial training using an improved model based on Mahalanobis distance; then, the target transaction data is processed by calling the above target transaction behavior feature prediction model, and the user label of the target user is determined according to the determined transaction behavior features. Thus, the user's transaction behavior features can be accurately obtained and used, and the user label of the target user in the transaction data processing scenario can be more accurately determined. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the embodiments of this specification, the drawings required for use in the embodiments will be briefly introduced below. The drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0060] Figure 1 It is a schematic diagram of an embodiment of the structural composition of a system applying the method for determining user tags provided in the embodiments of this specification;

[0061] Figure 2 is a flowchart of a method for determining a user tag provided by an embodiment of this specification;

[0062] Figure 3It is a flowchart of a training method of a preset transaction behavior feature prediction model provided by an embodiment of this specification;

[0063] Figure 4 It is a schematic diagram of the structure of a server provided by an embodiment of this specification;

[0064] Figure 5 It is a schematic diagram of the structure of a device for determining a user tag provided by an embodiment of this specification;

[0065] Figure 6 It is a schematic diagram of the structure of a training device for a preset transaction behavior feature prediction model provided by an embodiment of this specification;

[0066] Figure 7 is a schematic diagram of an embodiment of a method for determining a user tag provided by an embodiment of this specification, in a scenario example;

[0067] Figure 8 It is a schematic diagram of an embodiment of a method for determining a user tag provided by an embodiment of this specification, in a scenario example. DETAILED DESCRIPTION

[0068] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0069] Considering that when training a prediction model for predicting user tags based on existing methods, it is often necessary to first label a large amount of sample data according to artificially set rules and set corresponding labels; then use the labeled sample data to train the model to obtain the corresponding prediction model. First of all, the above model training process is bound to consume a lot of human resources and a lot of processing time to complete the labeling of sample data. This results in high cost of model training and low training efficiency.

[0070] Secondly, the prediction model trained based on the above method is prone to being subjective and unreasonable due to human factors, as the labeling process is based on artificially set rules to set labels. In addition, labeling sample data based on artificially set rules cannot take into account the specific circumstances in the actual transaction data processing scenario, and the label types used in labeling are often incomplete, have omissions, or are not suitable for the actual transaction data processing scenario currently faced. As a result, the labeling effects of the labeled sample data obtained based on the artificially set rules are relatively different, affecting the accuracy of the prediction model finally trained.

[0071] As a result, when the above prediction model is used to determine user tags in the future, technical problems such as large errors and inaccuracies in determining user tags in transaction data processing scenarios are bound to occur.

[0072] In view of the root cause of the above problems, this specification considers that before specific implementation, while obtaining a large amount of historical transaction data as the second type of sample data, historical market data associated with the above historical transaction data is obtained as the first type of sample data. Among them, only part of the sample data in the first type of sample data is required to carry the corresponding transaction behavior feature label, and only part of the sample data in the second type of sample data is required to carry the market category label. This can effectively reduce the workload of the labeling process, reduce the cost of model training, and improve the overall training efficiency.

[0073] Furthermore, hierarchical clustering is first performed on the first type of sample data to find the common features of multiple historical market data in the transaction data processing scenario based on the characteristics of the first type of sample data itself, rather than artificially set rules, and automatic clustering is performed, so that multiple category groups corresponding to multiple market types can be obtained.

[0074] Then, based on the first type of sample data contained in multiple category groups and the correlation between historical transaction data and historical market data, the second type of sample data corresponding to each market category is determined to obtain multiple training data sets corresponding to the multiple market categories.

[0075] Furthermore, an improved model based on Mahalanobis distance instead of Euclidean distance can be constructed as the initial model. Then, the initial model is subjected to preset virtual adversarial training using multiple training data sets, thereby obtaining multiple preset transaction behavior feature prediction models that can predict the transaction features of the transaction data of the corresponding market category.

[0076] Therefore, preset transaction behavior feature prediction models corresponding to multiple market categories in transaction processing scenarios with higher accuracy and better effects can be efficiently trained at lower training costs.

[0077] In specific implementation, when it is necessary to determine the user label of the target user, the target market data associated with the target transaction can be obtained while obtaining the target transaction data corresponding to the target user. The matching target market category is first determined based on the target market data. Then, based on the target market category, a matching target transaction behavior feature prediction model is determined from multiple preset transaction behavior feature prediction models trained previously.

[0078] Furthermore, the target transaction behavior feature prediction model can be called to process the target transaction data to accurately determine the transaction behavior features of the target transaction data. Then, based on the transaction behavior features, the user label of the target user in the transaction data processing scenario can be accurately determined.

[0079] Therefore, based on the transaction behavior characteristics of the target transaction data corresponding to the target user, the user label in the transaction data processing scenario of the target user can be accurately determined.

[0080] The embodiment of this specification provides a method for determining a user tag, which can be specifically applied to a system including a server and a user terminal. Figure 1 The server and the user terminal can be connected via wired or wireless means and perform specific data interaction.

[0081] In this embodiment, the server may specifically include a background server applied to a transaction data processing platform and capable of realizing functions such as data transmission and data processing. Specifically, the server may be, for example, an electronic device having data calculation, storage and network interaction functions. Alternatively, the server may also be a software program running in the electronic device to provide support for data processing, storage and network interaction. In this embodiment, the number of the servers is not specifically limited. The server may specifically be one server, or several servers, or a server cluster formed by several servers.

[0082] In this embodiment, the user terminal may specifically include a front-end electronic device applied to the user side and capable of realizing functions such as data collection and data transmission. Specifically, the user terminal may be, for example, a desktop computer, a tablet computer, a laptop computer, a smart phone, a computer, etc. Alternatively, the user terminal may also be a software application that can be run in the above electronic devices. For example, it may be a transaction APP installed and run on a smart phone.

[0083] In this embodiment, the user can perform specific transaction operations, such as buying operations, selling operations, etc., on a transaction data processing platform (eg, a precious metals trading platform, etc.) through the user terminal held by the user.

[0084] Correspondingly, the user terminal can respond to the above transaction operation of the user, generate corresponding target transaction data, and send the above target transaction data to the server of the transaction data processing platform for processing.

[0085] While receiving and processing the above target transaction data, the server also determines the user tag of the user in the current transaction data processing scenario based on the target transaction data.

[0086] Specifically, the server may first obtain the target market data associated with the target transaction data. For example, the server may collect the market data when the user sends the target transaction data as the target market data associated with the target transaction data.

[0087] Next, the server can determine the matching target market category from multiple preset market categories obtained in advance through hierarchical clustering according to the target market data, and determine the matching target transaction behavior feature prediction model from multiple preset transaction behavior feature prediction models obtained by pre-training the second type of sample data of the corresponding market category using an improved model based on Mahalanobis distance with good generalization.

[0088] Furthermore, the target transaction behavior feature prediction model can be called to process the target transaction data to determine the behavior features of the target transaction data. And based on the transaction behavior features of the target transaction data, the user tags in the current transaction data processing scenario can be more accurately determined.

[0089] Furthermore, the server can generate risk warning information matching the user in a targeted manner according to the user tag, and feed back the risk warning information to the user terminal. The user terminal displays the risk warning information to the user.

[0090] Through the above embodiments, users can obtain risk warning information that is closely related to them and has high reference value in a timely manner. Accordingly, users are more willing to adjust their trading operations in a targeted manner according to the above risk warning information. This can effectively reduce the risk of users suffering losses when performing trading operations on the trading data processing platform, improve the user experience, and reduce the loss of users of the trading data processing platform.

[0091] See also Figure 2 As shown, the embodiment of this specification provides a method for determining a user tag. The method is specifically applied to the server side. When implemented specifically, the method may include the following contents:

[0092] S201: Acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user;

[0093] S202: According to the target market data, determine a matching target market category from a plurality of preset market categories; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data;

[0094] S203: determining a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training sample data of the corresponding market category using an improved model based on Mahalanobis distance;

[0095] S204: calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data;

[0096] S205: Determine a user tag of a target user according to the transaction behavior characteristics of the target transaction data.

[0097] Through the above embodiments, the target transaction data can be processed in a targeted manner by calling the target transaction behavior feature prediction model corresponding to the target market category matching the current market conditions, so as to accurately determine the transaction behavior features of the target transaction data; and then based on the above transaction behavior features, the user label of the target user corresponding to the target transaction data in the transaction data processing scenario can be more accurately determined, thereby reducing the error in determining the user label.

[0098] In this embodiment, the target transaction data can be specifically understood as request data initiated by a target user to instruct a transaction data processing in a transaction data processing scenario, or parameter data used to characterize a certain type of attribute characteristics (e.g., revenue characteristics, etc.) of the target user in the transaction data processing scenario. The target transaction data corresponds to the target user.

[0099] In this embodiment, the target transaction data may carry a user identifier of a corresponding target user, such as the target user's account name, the target user's name, or the target user's identity number and other information.

[0100] In one embodiment, the target transaction data includes at least one of the following: stock transaction data, precious metal transaction data, crude oil transaction data, etc.

[0101] Through the above embodiments, the method provided in this specification can be applied to obtain and process various types of target transaction data in various different transaction data processing scenarios.

[0102] In this embodiment, for different transaction data processing scenarios, the target transaction data may be transaction data of different types and contents. Specifically, for example, in a stock transaction scenario, the target transaction data may be stock transaction data, such as a target user's request to buy a certain stock, or a target user's request to sell a certain stock, etc. In a precious metal transaction scenario, the target transaction data may be precious metal transaction data, such as the closing price of a certain precious metal purchased by the target user, or the yield of a certain precious metal purchased by the target user, etc.

[0103] Of course, the various transaction data processing scenarios and target transaction data listed above are only schematic illustrations. In specific implementation, according to specific application scenarios and processing requirements, the above method can also be extended to other types of transaction data processing scenarios. Accordingly, the above target transaction data can also include other types of transaction data in other types of transaction data processing scenarios. This specification does not limit this.

[0104] In this embodiment, the above-mentioned target market data can be specifically understood as parameter data associated with the target transaction data, which is used to reflect, from a macro perspective, a certain type of attribute characteristics of the overall transaction data processing scenario when the target transaction data is generated.

[0105] Similar to the target transaction data, the target market data may also be market data of different types and contents for different transaction data processing scenarios. Specifically, for example, in a stock trading scenario, the target market data may be the opening price of the market, the rise and fall of the market, or the closing price of the market.

[0106] In this embodiment, when the server obtains the target transaction data, it can actively query and collect the market data in the transaction data processing scenario when the target transaction data is generated as the target market data associated with the target transaction data.

[0107] In this embodiment, the above-mentioned preset market category can be specifically understood as a predetermined category data for describing the scenario type of the transaction data processing scenario in which the transaction data is located from a macro dimension.

[0108] The preset market category can be obtained by pre-acquiring historical market data as the first type of sample data and performing hierarchical clustering on the first type of sample data. The specific process of obtaining the preset market category through hierarchical clustering will be described later.

[0109] In one embodiment, specifically, the above market categories may include at least one of the following: rising market, falling market, oscillating market, etc. Of course, the above listed market categories are only a schematic illustration. In specific implementation, other types of market categories may also be included according to specific application scenarios and processing requirements. This specification does not limit this.

[0110] Through the above embodiments, it is possible to finely distinguish a variety of specific market conditions that exist in real transaction data processing scenarios, so that targeted prediction processing can be performed on transaction data under different market conditions in the future to obtain relatively more accurate prediction results.

[0111] In one embodiment, during implementation, the server may first calculate the similarity between the target market data and each of the multiple preset market categories; and then find the preset market category with the highest similarity from the multiple preset market categories as the matching target market category.

[0112] In this embodiment, the preset transaction behavior feature prediction model corresponds to a market category. The preset transaction behavior feature prediction model can be specifically understood as a prediction model obtained by pre-training, which can predict the corresponding transaction behavior feature based on the transaction data of the corresponding market category.

[0113] The preset trading behavior feature prediction model is obtained by pre-using an improved model based on Mahalanobis distance as an initial model, using the historical trading data of the corresponding market category as the second type of sample data, and using the second type of sample data to perform adversarial training on the initial model. The specific training process of the preset trading behavior feature prediction model will be described later.

[0114] In one embodiment, when the target transaction behavior feature prediction model is called to specifically process the target transaction data, the target transaction data can be used as a model input, input into the target transaction behavior feature prediction model, and the model can be run to obtain the corresponding model output; and then the transaction behavior features of the target transaction data can be determined based on the model output. The above-mentioned transaction behavior features can specifically include at least one of the following: aggressive behavior, conservative behavior, risk-averse behavior, etc.

[0115] In one embodiment, during specific implementation, the user tag of the target user in the transaction data processing scenario may be determined according to the transaction behavior characteristics of the target transaction data.

[0116] Specifically, for example, the transaction behavior characteristics of the target transaction data may be used as a user tag of a target user corresponding to the target transaction data in a transaction data processing scenario.

[0117] In one embodiment, corresponding to the transaction behavior characteristics, the user label may specifically include at least one of the following: an aggressive label, a conservative label, a risk-averse label, etc.

[0118] Through the above embodiments, the transaction behavior characteristics can be used to determine the user tags of the target users of the transaction data processing scenario, so that the obtained user tags can be closer to the corresponding real transaction data processing scenario and have higher reference value.

[0119] In one embodiment, the target transaction data may specifically include multiple target transaction data corresponding to the same target user. Accordingly, when the user tag of the target user is determined based on the transaction behavior characteristics of the target transaction data, it may include: combining the transaction behavior characteristics of the multiple target transaction data to obtain a combined transaction behavior characteristic; and determining the user tag of the target user based on the combined transaction behavior characteristic.

[0120] Through the above embodiments, one or more transaction behavior characteristics obtained based on multiple different target transaction data of the same target user can be comprehensively utilized to more accurately determine the user label in the real transaction data processing scenario for the target user, thereby further reducing the error in determining the user label.

[0121] In one embodiment, during specific implementation, the transaction behavior features of multiple target transaction data may be weighted and summed to obtain a combined transaction behavior feature, and then the combined transaction behavior feature may be used as the user tag of the target user. It is also possible to combine the specific transaction data processing scenario with other associated information of the target user, and select a transaction behavior feature with the highest matching degree with the target user from the transaction behavior features of multiple target transaction data as the user tag of the target user.

[0122] In one embodiment, after determining the user tag of the target user, the method may further include the following when implemented: generating risk warning information matching the target user according to the user tag of the target user; and sending the risk warning information to the target user.

[0123] In this embodiment, the above-mentioned risk warning information can be specifically understood as a warning information containing transaction risks that may be caused by the above-mentioned transaction behavior characteristics, which is obtained by the server after analyzing and simulating the target user's usual transaction behavior characteristics based on the target user's user tag.

[0124] Furthermore, the above risk warning information may also include adjustment strategies and suggestions for reducing the above transaction risks.

[0125] Through the above embodiment, the server can generate risk warning information that is more closely related to the target user, has higher reference value and is more willing to be accepted by the target user according to the user tag of the target user; and timely provide the above risk warning information to the target user. For the target user, since he is relatively more willing to accept the above risk warning information and is willing to refer to the specific content of the above risk warning information; therefore, the target user can adjust his investment trading plan or strategy more timely, effectively avoid the loss of income caused by his own habitual trading behavior characteristics, so that the user can get a better user experience.

[0126] In this embodiment, when determining the user tag of the target user, first obtain the target transaction data corresponding to the target user and the target market data associated with the target transaction data; then determine the matching target market category from multiple preset market categories obtained in advance through hierarchical clustering based on the target market data; and then determine the matching target transaction behavior feature prediction model from multiple preset transaction behavior feature prediction models obtained in advance through adversarial training using an improved model based on Mahalanobis distance; then call the above target transaction behavior feature prediction model to process the target transaction data, and determine the user tag of the target user based on the determined transaction behavior feature. In this way, the user's transaction behavior features can be accurately obtained and utilized, and the user tag of the target user in the transaction data processing scenario can be determined more accurately. Furthermore, based on the above user tags, risk warning information with strong pertinence and high value can be generated and sent to the target user, so that the target user can obtain a better user experience and reduce customer churn of the transaction data processing platform.

[0127] In one embodiment, before implementation, historical transaction data within a historical time period may be obtained as first-category sample data; at the same time, historical market data associated with the historical transaction data within the same historical time period may be obtained as second-category sample data. Multiple preset market categories and multiple preset transaction behavior feature prediction models corresponding to the multiple preset market categories may then be determined based on the first-category sample data and the second-category sample data.

[0128] In one embodiment, before implementation, the method may further include the following contents:

[0129] S1: Acquire historical transaction data as the second type of sample data; and acquire historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label;

[0130] S2: establishing a correspondence between the first type of sample data and the second type of sample data according to the association relationship between the historical transaction data and the historical market data; wherein the first type of sample data corresponds to one or more second type of sample data;

[0131] S3: performing hierarchical clustering on the first category sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first category sample data;

[0132] S4: determining the second type of sample data corresponding to each market category according to the first type of sample data included in the plurality of category groups and the correspondence between the first type of sample data and the second type of sample data, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively;

[0133] S5: Construct an improved model based on Mahalanobis distance as the initial model;

[0134] S6: Performing preset virtual adversarial training on the initial model using the multiple training data sets respectively to obtain multiple preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category.

[0135] Through the above embodiment, the first type of sample data containing only part of the labels can be used to accurately determine multiple preset market categories in the real transaction data processing scenario through hierarchical clustering; at the same time, the initial model based on the Mahalanobis distance improvement can be introduced, and the second type of sample data containing only part of the labels can be used for adversarial training to obtain multiple relatively accurate preset transaction behavior feature prediction models that are suitable for the real transaction data processing scenario and correspond to multiple preset market categories. In addition, through the above embodiment, during the training process, since all sample data do not need to be manually labeled, the overall processing cost can be effectively reduced and the overall processing efficiency can be improved.

[0136] In one embodiment, during specific implementation, the server may obtain historical transaction data that appears on the transaction data processing platform within a certain historical time period (e.g., the past year, etc.), and set corresponding transaction behavior feature tags for only part of the above historical transaction data, and obtain the second type of sample data in which at least one sample data carries the transaction behavior feature tag. While obtaining the historical transaction data, the server also obtains the historical market data associated with the historical transaction data within the same historical time period, and similarly sets corresponding market category tags for only part of the above historical market data, and obtains the first type of sample data in which at least one sample data carries the market category tag.

[0137] Furthermore, the server can establish a correspondence between the first type of sample data and the second type of sample data according to the association between the historical business data and the historical market data. Each first type of sample data can correspond to one or more second type of sample data. For example, multiple different second type sample data at the same time point can correspond to one first type sample data object at the same time point.

[0138] In one embodiment, during specific implementation, the server can perform hierarchical clustering on the first-category sample data, find multiple first-category sample data with common features based on the data of the first-category sample data itself, and cluster them to obtain category groups of multiple market categories in the respective transaction data processing scenarios. Each of the above multiple market category groups contains one or more first-category sample data.

[0139] In one embodiment, when hierarchical clustering is performed on the first type of sample data, the first type of sample data can be segmented layer by layer multiple times in a serial or parallel manner to efficiently find multiple categories (or multilateral factors), and then multiple category groups can be obtained. In each segmentation, the average value of the mutual distances between multiple nodes corresponding to the multiple first type sample data can be used for calculation, and a category with the largest range can be selected for segmentation. When segmenting the category, the following steps can be performed:

[0140] S1: Let the category to be segmented be recorded as U, find a node in U with the longest average distance to other nodes, recorded as d, and form a new category Ui.

[0141] S2: Select a node d' in Ui, where the node d' satisfies the following characteristics: the difference between the average distance to other nodes in Ui and the average distance from d' to all nodes in Ui is the largest, and assign the node d' to U.

[0142] S3: Repeat step S2 until the calculated difference is negative.

[0143] By using the above-mentioned hierarchical clustering, we can better combine the scenario characteristics in the transaction data processing scenario, and perform hierarchical segmentation and clustering on the market categories with relatively small number of categories in the transaction data processing scenario, so as to quickly find the category set U composed of multiple first-category sample data with common characteristics, and obtain multiple category groups corresponding to multiple market categories.

[0144] In one embodiment, before performing hierarchical clustering, the first type of sample data may be cleaned first; then the cleaned first type of sample data may be one-hot encoded to obtain encoded first type of sample data. Then, hierarchical clustering may be performed on the encoded first type of sample data, so that multiple preset market categories in the transaction data processing scenario may be determined more efficiently and accurately.

[0145] In one embodiment, based on the first type of sample data included in each category group and the previously determined correspondence between the first type of sample data and the second type of sample data, the second type of sample data corresponding to the first type of sample data included in the category group can be found as the second type of sample data corresponding to the market category; and the second type of sample data corresponding to the market category can be determined as the training data set corresponding to the market category. Thus, multiple training data sets corresponding to multiple market categories can be obtained.

[0146] In one embodiment, considering that the initial models commonly used for training are mostly constructed based on Euclidean distance, when performing model training based on such initial models, it is often impossible to eliminate the interference caused by the correlation between different attribute features in the training data. At the same time, the trained model will be affected by the data units of the training data, which leads to relatively poor generalization and unsatisfactory results of the trained model.

[0147] It is precisely because of the above problem that an improved model based on Mahalanobis distance with good generalization is first constructed as an initial model in this embodiment. Then, the initial model based on Mahalanobis distance can be used for specific model training to obtain a preset transaction behavior feature prediction model with high generalization and good accuracy.

[0148] In one embodiment, the initial model may at least include an improved loss function based on Mahalanobis distance. The Mahalanobis distance (MD) may be used to represent the distance between a point and a distribution. Specifically, unlike the Euclidean distance, the Mahalanobis distance takes into account the connection between various attribute characteristics and is a distance that is independent of the scale unit.

[0149] In this embodiment, when implementing, the characteristics of Mahalanobis distance can be used. When constructing the initial model, the Mahalanobis distance can be introduced in a targeted manner to improve the loss function used in the initial model, so as to obtain a loss function with better effect. In this way, when training the initial model later, the characteristics of Mahalanobis distance can be used to enhance the generalization performance of virtual adversarial training, forming a Mahalanobis-virtual adversarial training model (which can be recorded as MD-VAT), so that a preset transaction behavior feature prediction model with high generalization and good effect can be obtained later.

[0150] In one embodiment, the above-mentioned construction of the improved model based on the Mahalanobis distance as the initial model may specifically include: using the improved loss function based on the Mahalanobis distance to construct a corresponding objective function to obtain the initial model.

[0151] The improved loss function based on Mahalanobis distance can be constructed in the following way:

[0152]

[0153] in, is the improved loss function based on Mahalanobis distance in the initial model, D l D is the sample data set composed of the second type of sample data with transaction behavior feature labels. ul N is the sample data set composed of the second type of sample data without transaction behavior feature labels. l D l The total number of the second type of sample data contained in, N ul D ul The total number of the second type of sample data contained in , x is the second type of sample data, θ is the independent variable of the loss function, and LMS is the Mahalanobis smoothness.

[0154] Through the above embodiments, a corresponding improved loss function based on the Mahalanobis distance can be constructed, and the above loss function can be used to construct the objective function of the model, so that an improved initial model based on the Mahalanobis distance with better effect can be established.

[0155] Furthermore, the above It is obtained based on the Mahalanobis smoothness (LMS), which can also be called the sum of the loss functions of the Mahalanobis smoothness.

[0156] Correspondingly, the objective function of the initial model can be established in the following way using the improved loss function based on the Mahalanobis distance mentioned above:

[0157]

[0158] in, Specifically, it can be expressed as the objective function in the virtual adversarial network, l(D l , θ) can be specifically expressed as the loss function of the sample data set carrying transaction behavior feature labels (which can be referred to as the labeled loss function for short), and α is a constant.

[0159] Through the above embodiments, the Mahalanobis distance can be used to specifically improve the loss function in the objective function used by the model, so as to obtain an initial model with higher generalization and better accuracy.

[0160] In one embodiment, the Martensitic smoothness can be determined according to the following formula:

[0161]

[0162]

[0163] Among them, x i is the transaction behavior feature label of the second type of sample data numbered i, is the specific data value in the independent variable of the loss function, r MDVAT is the perturbation term based on Mahalanobis distance, is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set, p(x i +r MDVAT ,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after being disturbed by the disturbance term based on the Mahalanobis distance, md[] represents the calculation function of the Mahalanobis distance, r is the random disturbance, p(x i +r,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after random perturbation.

[0164] Through the above embodiments, Mahalanobis smoothness can be introduced and utilized to improve the loss function of the model based on the Mahalanobis distance, thereby obtaining an initial model with higher generalization ability.

[0165] In this embodiment, the specific calculation process of the calculation function of the Mahalanobis distance can be implemented by referring to the following formula:

[0166]

[0167] Among them, x and y are different schematic variables, and σ is the standard deviation of variable x.

[0168] In one embodiment, when conducting model training, different market categories can be distinguished, and multiple training data sets corresponding to multiple different market categories can be used to perform virtual adversarial training based on the initial model, so as to obtain multiple different preset transaction behavior feature prediction models. Each preset transaction behavior feature prediction model corresponds to a market category and is used to predict the transaction behavior features of the transaction data of the corresponding market category.

[0169] In one embodiment, taking the training of a preset trading behavior feature prediction model corresponding to the current market category among multiple preset trading behavior feature prediction models as an example, the above-mentioned using the multiple training data sets to perform preset virtual adversarial training on the initial model to obtain multiple preset trading behavior feature prediction models, when specifically implemented, may include: using the current training data set among the multiple training data sets to perform preset virtual adversarial training on the initial model in the following manner to obtain the preset trading behavior feature prediction model corresponding to the current market category:

[0170] S1: using the initial model to process the second type of sample data in the current data set that does not carry the transaction behavior feature label, and obtain pseudo labels corresponding to the second type of sample data that does not carry the transaction behavior feature label;

[0171] S2: combining the second type of sample data carrying the transaction behavior feature label in the current training data set with the second type of sample data carrying the pseudo label to obtain combined training data;

[0172] S3: Using the combined training data, the initial model is trained to obtain a preset transaction behavior feature prediction model corresponding to the current market category.

[0173] By repeating the above method to perform virtual adversarial training, multiple preset trading behavior feature prediction models corresponding to multiple market categories can be obtained.

[0174] Through the above-mentioned embodiments, only a relatively low processing cost is required to annotate part of the second category sample data in the training data set, and virtual adversarial training based on semi-supervised learning can be used to learn the second category sample data carrying transaction behavior feature labels in the training data set, so as to automatically annotate the second category sample data that does not carry transaction behavior feature labels in the original training data set with corresponding pseudo-labels; further, the second category sample data carrying transaction behavior feature labels and the second category sample data carrying pseudo-labels are simultaneously integrated and utilized for learning and training to obtain a preset transaction behavior feature prediction model with higher generalization and better accuracy.

[0175] In one embodiment, before training the preset transaction behavior feature prediction model, the method further includes: performing data cleaning on the second type of sample data; and then one-hot encoding the cleaned second type of sample data. The encoded second type of sample data can then be used to perform virtual adversarial training to obtain the corresponding preset transaction behavior feature prediction model.

[0176] As can be seen from the above, the method for determining user tags provided in the embodiment of this specification, when determining the user tag of the target user, first obtains the target transaction data corresponding to the target user, and the target market data associated with the target transaction data; then, according to the target market data, determines the matching target market category from multiple preset market categories obtained in advance through hierarchical clustering; and then determines the matching target transaction behavior feature prediction model from multiple preset transaction behavior feature prediction models obtained in advance through adversarial training using an improved model based on Mahalanobis distance; then, by calling the above target transaction behavior feature prediction model to process the target transaction data, the user tag of the target user is determined according to the determined transaction behavior features. Thus, the user's transaction behavior features can be accurately obtained and utilized, and the user tag of the target user in the transaction data processing scenario can be determined more accurately. Furthermore, according to the above user tags, risk warning information with strong pertinence and high value can be generated and sent to the target user, so that the target user can obtain a better user experience and reduce customer churn of the transaction data processing platform. By first performing hierarchical clustering on the first type of sample data containing multiple historical market data, multiple market categories in the transaction scenario are determined; then, based on the correlation between the historical transaction data and the historical market data, the second type of sample data corresponding to each market category is determined to obtain multiple training data sets corresponding to multiple market categories, and then the multiple different market categories existing in the transaction data processing scenario can be finely distinguished, and the corresponding training data sets are used for model training, thereby improving the accuracy of model training. When the model training is specifically performed, by first introducing the improved loss function based on the Mahalanobis distance into the objective function of the initial model, an improved model based on the Mahalanobis distance with stronger generalization and better effect is obtained as the initial model; further, the training data set in which only part of the sample data carries labels can be used to perform virtual adversarial training on the above-mentioned improved model based on the Mahalanobis distance, so that a preset transaction behavior feature prediction model corresponding to the corresponding market category with high generalization, good accuracy, small error, and suitable for the transaction data processing scenario can be obtained. In addition, since the above model training process only needs to label part of the sample data, it can effectively reduce the amount of data processing involved in the model training process, reduce the training cost, and improve the efficiency of model training.

[0177] See also Figure 3 As shown, the embodiment of this specification also provides a training method for a preset transaction behavior feature prediction model, and when the method is implemented specifically, it may include the following contents.

[0178] S301: Acquire historical transaction data as the second type of sample data; and acquire historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label;

[0179] S302: establishing a correspondence between the first type of sample data and the second type of sample data according to the association relationship between the historical transaction data and the historical market data; wherein the first type of sample data corresponds to one or more second type of sample data;

[0180] S303: performing hierarchical clustering on the first category sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first category sample data;

[0181] S304: Determine the second type of sample data corresponding to each market category according to the first type of sample data included in the plurality of category groups and the correspondence between the first type of sample data and the second type of sample data, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively;

[0182] S305: constructing an improved model based on Mahalanobis distance as an initial model;

[0183] S306: Performing preset virtual adversarial training on the initial model using the multiple training data sets respectively to obtain multiple preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category.

[0184] As can be seen from the above, the training method of the preset transaction behavior feature prediction model provided in the embodiment of this specification firstly performs hierarchical clustering on the first type of sample data containing multiple historical market data to determine multiple market categories in the transaction scenario; then, according to the correlation between the historical transaction data and the historical market data, the second type of sample data corresponding to each market category is determined to obtain multiple training data sets corresponding to multiple market categories, and then the multiple different market categories existing in the transaction data processing scenario can be finely distinguished, and the corresponding training data sets are used for model training, thereby improving the accuracy of model training. Further, when conducting adversarial training, the improved loss function based on Mahalanobis distance is first introduced into the objective function of the initial model to obtain an improved model based on Mahalanobis distance with stronger generalization and better effect as the initial model; and then, the training data set in which only part of the sample data carries labels is used to perform virtual adversarial training on the improved model based on Mahalanobis distance. Thus, a preset transaction behavior feature prediction model with high generalization, good accuracy, small error, and applicable to real transaction data processing scenarios and corresponding to each market category can be obtained.

[0185] The embodiment of the present specification also provides a server, including a processor and a memory for storing processor executable instructions, wherein the processor can perform the following steps according to the instructions when implemented: obtaining target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user; determining a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by hierarchical clustering of sample data; determining a target transaction behavior feature prediction model matching the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained in advance by adversarially training sample data of the corresponding market category using an improved model based on Mahalanobis distance; calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data; and determining a user tag of the target user according to the transaction behavior features of the target transaction data.

[0186] In order to complete the above instructions more accurately, refer to Figure 4 As shown, the embodiment of this specification also provides another specific server, wherein the server includes a network communication port 401, a processor 402 and a memory 403, and the above structures are connected through internal cables so that each structure can perform specific data interaction.

[0187] The network communication port 401 can be specifically used to obtain target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user.

[0188] The processor 402 can be specifically used to determine a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on sample data; determine a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained in advance by performing adversarial training on sample data of the corresponding market categories using an improved model based on Mahalanobis distance; call the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data; and determine a user tag of a target user according to the transaction behavior features of the target transaction data.

[0189] The memory 403 may be specifically used to store corresponding instruction programs.

[0190] In this embodiment, the network communication port 401 can be a virtual port that is bound to different communication protocols so that different data can be sent or received. For example, the network communication port can be a port responsible for web data communication, a port responsible for FTP data communication, or a port responsible for email data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.

[0191] In this embodiment, the processor 402 may be implemented in any appropriate manner. For example, the processor may take the form of a microprocessor or processor and a computer-readable medium storing a computer-readable program code (such as software or firmware) executable by the (micro)processor, a logic gate, a switch, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller, etc. This specification does not limit this.

[0192] In this embodiment, the memory 403 may include multiple levels. In a digital system, anything that can store binary data can be a memory; in an integrated circuit, a circuit with a storage function that has no physical form is also called a memory, such as RAM, FIFO, etc.; in a system, a storage device with a physical form is also called a memory, such as a memory stick, TF card, etc.

[0193] The embodiment of the present specification also provides a computer storage medium based on the above-mentioned user label determination method, wherein the computer storage medium stores computer program instructions, and when the computer program instructions are executed, the following steps are implemented: obtaining target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user; determining a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by hierarchical clustering of sample data; determining a target transaction behavior feature prediction model matching the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained in advance by adversarially training sample data of the corresponding market category using an improved model based on Mahalanobis distance; calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data; and determining the user label of the target user according to the transaction behavior features of the target transaction data.

[0194] In this embodiment, the storage medium includes, but is not limited to, a random access memory (RAM), a read-only memory (ROM), a cache, a hard disk (HDD), or a memory card. The memory may be used to store computer program instructions. The network communication unit may be an interface for network connection communication set in accordance with the standard specified by the communication protocol.

[0195] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained in comparison with other implementations and will not be described in detail here.

[0196] The present specification also provides a computer storage medium for a training method based on the above-mentioned preset transaction behavior feature prediction model, wherein the computer storage medium stores computer program instructions, and when the computer program instructions are executed, the following are achieved: obtaining historical transaction data as the second type of sample data; and obtaining historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label; and according to the association relationship between the historical transaction data and the historical market data, establishing a correspondence between the first type of sample data and the second type of sample data; wherein the first type of sample data corresponds to one or more A plurality of second-category sample data; hierarchically clustering the first-category sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first-category sample data; according to the first-category sample data included in the plurality of category groups, and the correspondence between the first-category sample data and the second-category sample data, determining the second-category sample data corresponding to each market category, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively; constructing an improved model based on Mahalanobis distance as an initial model; performing preset virtual adversarial training on the initial model using the plurality of training data sets respectively, so as to obtain a plurality of preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category.

[0197] See also Figure 5 As shown, at the software level, the embodiment of this specification also provides a device for determining a user tag, which may specifically include the following structural modules:

[0198] The acquisition module 501 may be specifically used to acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user;

[0199] The first determination module 502 may be specifically configured to determine a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data;

[0200] The second determination module 503 may be specifically used to determine a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training the sample data of the corresponding market category using an improved model based on Mahalanobis distance;

[0201] The calling module 504 may be specifically used to call the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data;

[0202] The third determination module 505 may be specifically configured to determine a user tag of a target user according to the transaction behavior characteristics of the target transaction data.

[0203] It should be noted that the units, devices or modules described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. For the convenience of description, the above devices are described separately by functions divided into various modules. Of course, when implementing this specification, the functions of each module can be implemented in the same or more software and / or hardware, or the modules that implement the same function can be implemented by a combination of multiple sub-modules or sub-units. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0204] It can be seen from the above that the user tag determination device provided in the embodiments of this specification can accurately obtain and utilize the transaction behavior characteristics of the user, and more accurately determine the user tag of the target user in the transaction data processing scenario.

[0205] See also Figure 6 As shown, at the software level, the embodiment of this specification also provides a training device for a preset transaction behavior feature prediction model, which may specifically include the following structural modules:

[0206] The acquisition module 601 may be specifically used to acquire historical transaction data as the second type of sample data; and acquire historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label;

[0207] Establishing module 602, which can be specifically used to establish a correspondence between the first type of sample data and the second type of sample data according to the association relationship between the historical transaction data and the historical market data; wherein the first type of sample data corresponds to one or more second type of sample data;

[0208] The clustering module 603 may be specifically used to perform hierarchical clustering on the first type of sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first type of sample data;

[0209] The determination module 604 may be specifically configured to determine the second-category sample data corresponding to each market category according to the first-category sample data included in the plurality of category groups and the correspondence between the first-category sample data and the second-category sample data, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively;

[0210] A construction module 605 may be specifically used to construct an improved model based on Mahalanobis distance as an initial model;

[0211] The training module 606 can be specifically used to perform preset virtual adversarial training on the initial model using the multiple training data sets respectively to obtain multiple preset trading behavior feature prediction models; wherein the preset trading behavior feature prediction model corresponds to a market category.

[0212] It can be seen from the above that the training device of the preset transaction behavior feature prediction model provided in the embodiment of this specification can effectively reduce the amount of data processing involved in the model training process, reduce the training cost, and improve the efficiency of model training. It can efficiently train to obtain a preset transaction behavior feature prediction model corresponding to the corresponding market category that has high generalization, good accuracy, and small error and is suitable for transaction data processing scenarios.

[0213] In a specific scenario example, the method provided in this specification can be applied to the precious metal trading scenario to train a model for predicting user tags. For the specific implementation process, please refer to the following content.

[0214] In this scenario example, a semi-supervised learning method and system based on Markov virtual adversarial training is proposed to solve the problem of inaccurate and erroneous identification of customer trading behaviors in precious metal trading scenarios.

[0215] For specific implementation, see Figure 7 As shown, you can follow the steps below:

[0216] S701: Obtain market data and transaction data (a small amount of labeled data and a large amount of unlabeled data), and clean and organize them.

[0217] S702: Use hierarchical clustering to find various multilateral factors (eg, market categories).

[0218] S703: Under each market category, perform Markov virtual adversarial training on unlabeled data, and store the training results.

[0219] S704: Match the data with transaction type labels with user dimensions, and construct data with transaction feature label user dimensions (for example, determine the user label of the corresponding target user according to the transaction behavior characteristics of the transaction data).

[0220] In this scenario example, the above-mentioned multilateral factors can be specifically generated based on various changes in the precious metal trading scenario. Specifically, each segmentation of hierarchical clustering uses the average value of the mutual distances of all nodes in the cluster to calculate, and directly selects the category with the largest range for segmentation. The segmentation method after clustering is as follows: 1. The category to be segmented is recorded as U, and a point d with the longest average distance to other points is taken out from U to form a new category u i ; 2. in u i Select such a point d', d' to u i The average distance of the other points in minus d' to u i The average distance of all points in the is the largest, and it is included in U; ​​3. Repeat the previous step until the difference is negative. This method is suitable for data with fewer categories. After segmentation, the category set U is obtained. Among them,

[0221] In this scenario example, the above virtual adversarial training is an advancement of supervised learning anti-training. The meaning of virtual adversarial training is to generate "virtual" labels (for example, pseudo labels) from unlabeled data. Then the newly generated virtual labels are added to the set of labeled samples to enhance the training accuracy of the model. This technology can be well applied in semi-supervised learning and unsupervised learning, and has the characteristics of enhancing data efficiency and improving generalization.

[0222] In this scenario example, improvements are made to the virtual adversarial training objective function and loss function, aiming to use the Mahalanobis (MD) distance to enhance the generalization of virtual adversarial training to form the Markov-Virtual Adversarial Training model (MD-VAT). Compared with many existing semi-supervised learning techniques, using the MD-VAT model, the input unlabeled data does not have to come from the same group of classes as the labeled data. It is more in line with the scenario requirements of this scenario of transaction data processing, making the model more suitable for applications in the field of precious metals trading. The modified objective function can be specifically expressed as:

[0223] Among them, the loss function used in the objective function can be specifically expressed as:

[0224]

[0225]

[0226]

[0227]

[0228] Among them, D l is the labeled sample set, D ul is the unlabeled sample set, is the objective function of the virtual adversarial network, l(D l ,θ) is the labeled loss function, is the sum of the LMS loss functions, LMS is Mahalanobis smoothness, r MDVAT is a disturbance term evaluated using the Mahalanobis distance, md is the Mahalanobis distance calculation formula, and the independent variable in the θ loss function. is a specific value in θ, x i is the label of the specific training sample, α is a constant, N l D l The number of labeled samples, N ul D ul The number of unlabeled samples in is the true distribution of the sample set.

[0229] In the md Mahalanobis distance calculation formula, x and y are different schematic variables, and σ is the standard deviation of variable x. MDVAT In the equation, r is x i Random perturbations of the variables.

[0230] In this scenario example, during specific training, the overall generalization ability of the model is improved by modifying the loss function in the virtual adversarial training. MDVAT is an adversarial perturbation that maximally changes the input variable x, as evaluated by the Mahalanobis distance.

[0231] In the LMS formula, is the current setting of θ at a particular moment in the optimization process, i.e., it is considered a constant. θ and The difference between the two is that the gradient of LMS(x, θ) recovers the perturbation only by the value generated by the input. In simple terms, LMS represents the sensitivity to the input X.

[0232] Moreover, by introducing multilateral factors, it is possible to obtain more diverse and economically logical transaction characteristics (for example, transaction behavior characteristics) for users under different market conditions.

[0233] In this scenario example, see Figure 8 As shown, during specific implementation, periodic data collection 1 can be performed based on the acquired transaction and market data 1, and data cleaning 2 can be performed on the collected data. One-hot encoding 7 can also be performed on these data to facilitate subsequent modeling.

[0234] Then, the market data is classified through hierarchical clustering to obtain various market factors 3. The following model training 4 must be carried out when the market factors are limited to the same type.

[0235] Then the processed transaction data is fed into the MD-VAT model training 4, the label data predicted by the model is output 5 and the result data is pre-stored 8. The required customer label is obtained.

[0236] The above scenario examples verify that based on the method provided in this specification, the acquired market data and transaction data are first collected and the data is one-hot processed to facilitate subsequent model training. Then the Mahalanobis virtual adversarial training model is used for training. The virtual adversarial training model used is improved by adding Mahalanobis distance to the model so that the input unlabeled data does not have to come from the same group of classes as the labeled data, so as to generalize the labels and improve the generalization of the model. Finally, by matching the labeled transaction data with the user data, transaction risk data with transaction feature labels in the user dimension is formed. In addition, in the case of large fluctuations in the precious metal market, while controlling the risk of customers, the probability of liquidity risk in member banks can be reduced, thereby improving the customer experience.

[0237] Although the present specification provides method operation steps as described in the embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way of executing the order of many steps, and does not represent a unique execution order. When the device or client product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiments or the drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed data processing environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or device including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such a process, method, product or device. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or device including the elements. The first, second, etc. words are used to represent the name, and do not represent any particular order.

[0238] Those skilled in the art also know that, in addition to implementing the controller in a purely computer-readable program code, the controller can be made to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered as a hardware component, and the devices for implementing various functions included therein can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules for implementing the method and structures within the hardware component.

[0239] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0240] Through the description of the above embodiments, it can be known that those skilled in the art can clearly understand that the present specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present specification can essentially be embodied in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a mobile terminal, a server, or a network device, etc.) to execute the methods described in each embodiment of the present specification or some parts of the embodiments.

[0241] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. This specification can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0242] Although the present specification is described through embodiments, those skilled in the art will appreciate that there are many modifications and changes to the present specification without departing from the spirit of the present specification, and it is intended that the appended claims include these modifications and changes without departing from the spirit of the present specification.

Claims

1. A method for determining a user tag, characterized in that: include: Acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user; According to the target market data, a matching target market category is determined from a plurality of preset market categories; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data; Determining a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training sample data of the corresponding market category using an improved model based on Mahalanobis distance; Calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data; Determining a user tag of a target user according to transaction behavior characteristics of the target transaction data; The method further comprises: obtaining historical transaction data as the second type of sample data; and obtaining historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label; establishing a correspondence between the first type of sample data and the second type of sample data according to the association between the historical transaction data and the historical market data; wherein the first type of sample data corresponds to one or more second type of sample data; performing hierarchical clustering on the first type of sample data to obtain a plurality of category groups; wherein the category group corresponds to a market category, and the category group includes one or more first type of sample data; according to the first type of sample data included in the plurality of category groups, and the correspondence between the first type of sample data and the second type of sample data, determining the second type of sample data corresponding to each market category, so as to obtain a plurality of training data sets corresponding to the plurality of market categories respectively; constructing an improved model based on Mahalanobis distance as an initial model; performing preset virtual adversarial training on the initial model using the plurality of training data sets respectively, so as to obtain a plurality of preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category; The initial model at least includes an improved loss function based on Mahalanobis distance, and constructing an improved model based on Mahalanobis distance as the initial model includes: using the improved loss function based on Mahalanobis distance to construct a corresponding objective function to obtain the initial model; the improved loss function based on Mahalanobis distance is constructed in the following manner: in, is the improved loss function based on Mahalanobis distance in the initial model, D l D is the sample data set composed of the second type of sample data with transaction behavior feature labels. ul N is the sample data set composed of the second type of sample data without transaction behavior feature labels. l D l The total number of the second type of sample data contained in, N ul D ul The total number of the second type of sample data contained in , x is the second type of sample data, θ is the independent variable of the loss function, and LMS is the Mahalanobis smoothness.

2. The method according to claim 1, characterized in that After determining the user tag of the target user, the method further includes: Generate risk warning information matching the target user based on the target user's user tag; The risk warning information is sent to the target user.

3. The method according to claim 1, characterized in that The market conditions include at least one of the following: rising market conditions, falling market conditions, and volatile market conditions.

4. The method according to claim 1, characterized in that: The Marano smoothness is determined by the following formula: Among them, x i is the transaction behavior feature label of the second type of sample data numbered i, is the specific data value in the independent variable of the loss function, r MDVAT is the perturbation term based on Mahalanobis distance, is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set, p(x i +r MDVAT ,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after being disturbed by the disturbance term based on the Mahalanobis distance, md[] represents the calculation function of the Mahalanobis distance, r is the random disturbance, p(x i +r,θ) is the data distribution of the transaction behavior feature label numbered i in the corresponding training data set after random perturbation.

5. The method according to claim 4, characterized in that The initial model is respectively subjected to preset virtual adversarial training using the multiple training data sets to obtain multiple preset transaction behavior feature prediction models, including: The initial model is subjected to a preset virtual adversarial training using a current training data set among the multiple training data sets in the following manner to obtain a preset transaction behavior feature prediction model corresponding to the current market category: Using the initial model to process the second type of sample data in the current data set that does not carry the transaction behavior feature label, to obtain pseudo labels corresponding to the second type of sample data that does not carry the transaction behavior feature label; Combining the second type of sample data carrying the transaction behavior feature label in the current training data set with the second type of sample data carrying the pseudo label to obtain combined training data; The combined training data is used to train the initial model to obtain a preset transaction behavior feature prediction model corresponding to the current market category.

6. The method according to claim 1, characterized in that The transaction behavior characteristic label includes at least one of the following: an aggressive label, a conservative label, and a risk-averse label.

7. The method according to claim 1, characterized in that The target transaction data includes a plurality of target transaction data corresponding to the same target user; accordingly, according to the transaction behavior characteristics of the target transaction data, determining the user tag of the target user includes: Combining the transaction behavior features of the plurality of target transaction data to obtain a combined transaction behavior feature; A user tag of the target user is determined according to the combined transaction behavior characteristics.

8. The method according to claim 1, characterized in that The target transaction data includes at least one of the following: stock transaction data, precious metal transaction data, and crude oil transaction data.

9. A device for determining a user tag, characterized in that: include: An acquisition module, used to acquire target transaction data and target market data associated with the target transaction data; wherein the target transaction data corresponds to a target user; A first determination module is used to determine a matching target market category from a plurality of preset market categories according to the target market data; wherein the plurality of preset market categories are obtained in advance by performing hierarchical clustering on the sample data; A second determination module is used to determine a target transaction behavior feature prediction model that matches the target market category from a plurality of preset transaction behavior feature prediction models; wherein the plurality of preset transaction behavior feature prediction models are obtained by pre-training sample data of the corresponding market category using an improved model based on Mahalanobis distance; A calling module, used for calling the target transaction behavior feature prediction model to process the target transaction data to determine the transaction behavior features of the target transaction data; A third determination module is used to determine a user tag of a target user according to the transaction behavior characteristics of the target transaction data; The device is also used to obtain historical transaction data as the second type of sample data; and obtain historical market data associated with the historical transaction data as the first type of sample data; wherein at least one sample data in the first type of sample data carries a market category label; and at least one sample data in the second type of sample data carries a transaction behavior feature label; according to the association relationship between the historical transaction data and the historical market data, a correspondence relationship between the first type of sample data and the second type of sample data is established; wherein the first type of sample data corresponds to one or more second type of sample data; hierarchical clustering is performed on the first type of sample data to obtain multiple category groups; wherein the category group corresponds to a market category, and the category group includes one or more first type of sample data; according to the first type of sample data included in the multiple category groups, and the correspondence relationship between the first type of sample data and the second type of sample data, the second type of sample data corresponding to each market category is determined to obtain multiple training data sets corresponding to multiple market categories respectively; an improved model based on Mahalanobis distance is constructed as an initial model; and the initial model is respectively subjected to preset virtual adversarial training using the multiple training data sets to obtain multiple preset transaction behavior feature prediction models; wherein the preset transaction behavior feature prediction model corresponds to a market category; The initial model at least includes an improved loss function based on Mahalanobis distance, and constructing an improved model based on Mahalanobis distance as the initial model includes: using the improved loss function based on Mahalanobis distance to construct a corresponding objective function to obtain the initial model; the improved loss function based on Mahalanobis distance is constructed in the following manner: in, is the improved loss function based on Mahalanobis distance in the initial model, D l The sample data set D is composed of the second type of sample data with transaction behavior feature labels. ul N is the sample data set composed of the second type of sample data without transaction behavior feature labels. l D l The total number of the second type of sample data contained in, N ul D ul The total number of the second type of sample data contained in , x is the second type of sample data, θ is the independent variable of the loss function, and LMS is the Mahalanobis smoothness.

10. A server, characterized in that: The method comprises a processor and a memory for storing processor-executable instructions, wherein the processor implements the steps of the method according to any one of claims 1 to 8 when executing the instructions.

11. A computer-readable storage medium, characterized in that: Computer instructions are stored thereon, and when the instructions are executed, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Modeling method, marketing method and device for predicting fund potential customers

    CN109509040A

  • Trransaction feature generation model trainingaining method and device of transaction feature generation model, and transaction feature generation method and device

    CN110263821A