Text generation method and related equipment

By comparing a target transaction account's features with the most similar normal accounts and using isolated forests and language models, the method generates accurate risk descriptions, addressing the inaccuracies in existing methods and enhancing the reliability of risk identification.

CN120317875APending Publication Date: 2025-07-15TENPAY PAID TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510373926.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

In the prior art, there is hallucination during the generation of risk description text, and the generated text cannot accurately describe the transaction account as the reason for the risk, resulting in inaccurate description.

Method used

By obtaining the transaction feature information of the target transaction account and performing abnormal identification, the most similar K transaction feature information are filtered out from the transaction feature information of multiple normal transaction accounts, and the abnormal cause description text is generated using the isolated forest algorithm and large language model to avoid using normal transaction features to interfere with the abnormal cause description.

Benefits of technology

It improves the accuracy of the abnormal cause description text, ensures the accuracy of abnormal transaction feature recognition, reduces the model hallucination phenomenon, and improves the accuracy of generated text.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317875A_ABST
    Figure CN120317875A_ABST
Patent Text Reader

Abstract

The invention discloses a text generation method and related equipment. The text generation method comprises the steps of obtaining transaction feature information of a target transaction account; performing anomaly identification according to the transaction feature information of the target transaction account to obtain an identification tag of the target transaction account; if the identification tag indicates that the target transaction account is abnormal, determining K pieces of transaction feature information most similar to the transaction feature information of the target transaction account in the transaction feature information of the plurality of normal transaction accounts; the normal transaction account refers to the transaction account indicated to be normal by the identification tag; according to the K pieces of transaction feature information, abnormal transaction feature recognition is carried out on the transaction feature information of the target transaction account, and abnormal transaction features in the transaction feature information of the target transaction account are determined; and performing text generation according to a target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account to obtain an abnormal reason description text of the target transaction account. The accuracy of the abnormal reason description text can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a text generation method and related devices. Background Art

[0002] To ensure the security of financial transactions, it is necessary to identify the risks of trading accounts to screen out trading accounts that may have risks such as financial fraud and black production transactions, which can also be called abnormal trading accounts. In related technologies, usually after identifying that a trading account has risks, the full amount of transaction characteristics of the trading account is input into a text generation model, and the text generation model generates text based on the full amount of transaction characteristics of the trading account to obtain a risk description text of the trading account. This risk description text is used to describe the reasons why the trading account is identified as having risks. However, in this way, there is a hallucination phenomenon, and the generated risk description text has the situation of misattribution and fabrication out of thin air, resulting in the risk description text not being able to accurately describe the reasons for identifying the trading account as having risks. Summary of the Invention

[0003] In view of the above problems, embodiments of this application propose a text generation method and related devices to solve the hallucination problem existing in the process of generating risk description texts in related technologies.

[0004] According to one aspect of the embodiments of this application, a text generation method is provided, including:

[0005] Obtain the transaction characteristic information of a target trading account;

[0006] Perform anomaly identification based on the transaction characteristic information of the target trading account to obtain an identification label of the target trading account;

[0007] If the identification label indicates that the target trading account is abnormal, determine the K transaction characteristic information that is most similar to the transaction characteristic information of the target trading account among the transaction characteristic information of multiple normal trading accounts; the normal trading account refers to a trading account whose identification label indicates normal; K is an integer greater than 1;

[0008] Perform abnormal transaction characteristic identification on the transaction characteristic information of the target trading account according to the K transaction characteristic information, and determine the abnormal transaction characteristics in the transaction characteristic information of the target trading account;

[0009] Generate a text description of the abnormal reason of the target trading account according to the target characteristic value of the abnormal transaction characteristic in the transaction characteristic information of the target trading account.

[0010] According to one aspect of the embodiments of this application, a text generation device is provided, including:

[0011] An acquisition module, configured to acquire transaction feature information of a target transaction account;

[0012] An anomaly recognition module, configured to perform anomaly recognition based on the transaction feature information of the target transaction account to obtain an identification label of the target transaction account;

[0013] A processing module, configured to, if the identification label indicates that the target transaction account is abnormal, determine K transaction feature information that is most similar to the transaction feature information of the target transaction account from the transaction feature information of multiple normal transaction accounts; the normal transaction account refers to a transaction account whose identification label indicates normal; K is an integer greater than 1;

[0014] An abnormal transaction feature determination module, configured to perform abnormal transaction feature recognition on the transaction feature information of the target transaction account according to the K transaction feature information, and determine abnormal transaction features in the transaction feature information of the target transaction account;

[0015] A text generation module, configured to generate a description text of the abnormal reason of the target transaction account according to a target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account.

[0016] In some embodiments, the transaction feature information includes feature values of multiple transaction features; the abnormal transaction feature determination module is configured to:

[0017] Perform the following processing for each of the transaction features:

[0018] Locate the position of the target feature value of the transaction feature in each binary tree of the target isolation forest in the transaction feature information of the target transaction account, and determine multiple first path lengths of the target feature value in the target isolation forest; the target isolation forest refers to an isolation forest applicable to anomaly detection of the transaction feature;

[0019] Locate the position of the reference feature value of the transaction feature in each binary tree of the target isolation forest in the K transaction feature information, and determine multiple second path lengths of each reference feature value in the target isolation forest;

[0020] Calculate the mean of the multiple first path lengths of the target feature value in the target isolation forest to obtain a first average path length;

[0021] Calculate the mean of the multiple second path lengths of each reference feature value in the target isolation forest to obtain the second average path length of each reference feature value in the target isolation forest;

[0022] Based on the first average path length and the second average path lengths of the reference eigenvalues in the target isolation forest, perform anomaly feature recognition on the transaction features in the transaction feature information of the target transaction account.

[0023] In some embodiments, in the step of performing anomaly feature recognition on the transaction features in the transaction feature information of the target transaction account based on the first average path length and the second average path lengths of the reference eigenvalues in the target isolation forest, the anomaly transaction feature determination module is further configured to: determine a path length threshold according to the second average path lengths of the K reference eigenvalues in the target isolation forest;

[0024] If the first average path length is less than the path length threshold, determine that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature.

[0025] In some embodiments, the text generation device further includes:

[0026] A first acquisition module, configured to acquire the first transaction feature information of the first sample abnormal transaction account and the second transaction feature information of multiple first sample normal transaction accounts that are most similar to the first transaction feature information;

[0027] An eigenvalue set construction module, configured to construct an eigenvalue set corresponding to each transaction feature according to the eigenvalue of the same transaction feature in the first transaction feature information and the eigenvalue in the second transaction feature information;

[0028] A binary tree construction module, configured to construct T binary trees corresponding to each transaction feature according to the eigenvalue set corresponding to each transaction feature; T is an integer greater than 1;

[0029] An integration module, configured to integrate the T binary trees corresponding to the same transaction feature to obtain a target isolation forest applicable to anomaly detection of the corresponding transaction feature.

[0030] In other embodiments, the transaction feature information includes eigenvalues of multiple transaction features; the abnormal transaction feature determination module is configured to:

[0031] Perform the following processing for each transaction feature:

[0032] Perform distribution fitting on the reference eigenvalue of the transaction feature in the K transaction feature information and the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account to determine the target feature distribution corresponding to the transaction feature;

[0033] Perform range delineation based on the target feature distribution corresponding to the transaction feature, and determine the target normal feature value range corresponding to the target feature distribution;

[0034] If the target feature value exceeds the target normal feature value range, determine that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature.

[0035] In some embodiments, the processing module includes:

[0036] A feature similarity calculation unit for calculating the feature similarity between the transaction feature information of the target transaction account and the transaction feature information of each normal transaction account;

[0037] A determination unit for using the transaction feature information of the K normal transaction accounts with the largest feature similarity as the K transaction feature information.

[0038] In some embodiments, the text generation device further includes:

[0039] A distribution determination module for performing distribution fitting based on the feature values of the risk transaction feature in the K transaction feature information and the corresponding target feature values to obtain the target feature distribution corresponding to the risk transaction feature;

[0040] Correspondingly, the text generation module is configured to:

[0041] The text generation model generates a text description of the abnormal reason of the target transaction account according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account and the target feature distribution corresponding to the abnormal transaction feature.

[0042] In some embodiments, the text generation device further includes:

[0043] A second acquisition module for acquiring a plurality of training samples, where one training sample includes a sample abnormal transaction feature, the sample feature distribution corresponding to the sample abnormal transaction feature, and the sample text description of the abnormal reason corresponding to the sample abnormal transaction feature;

[0044] A first text generation module for the text generation model to generate a predicted text description of the abnormal reason according to the sample abnormal transaction feature and the sample feature distribution corresponding to the sample abnormal transaction feature;

[0045] A text generation loss calculation module for calculating the text generation loss according to the sample text description of the abnormal reason and the predicted text description of the abnormal reason;

[0046] The first adjustment module is used to generate a loss based on the text and adjust the parameters of the text generation model until a first training end condition is reached.

[0047] In some embodiments, the text generation model includes a large language model and a fine-tuning network; the first adjustment module is configured to:

[0048] Fix the parameters of the large language model and adjust the parameters of the fine-tuning network according to the text generation loss until a first training end condition is reached.

[0049] In some embodiments, the anomaly recognition module is configured to: perform anomaly recognition on the transaction feature information of the target transaction account through an anomaly classification model to obtain the recognition label of the target transaction account;

[0050] Correspondingly, the text generation device further includes:

[0051] The third acquisition module is used to acquire the sample transaction feature information of multiple sample transaction accounts and the labeled recognition label of the sample transaction account;

[0052] The risk classification module is used to perform anomaly classification on each sample transaction account according to the sample transaction feature information of each sample transaction account through the anomaly classification model to obtain the predicted category of each sample transaction account;

[0053] The anomaly classification loss calculation module is used to calculate the anomaly classification loss according to the predicted category of each sample transaction account and the labeled recognition label of the sample transaction account;

[0054] The second adjustment module is used to adjust the parameters of the anomaly classification model according to the anomaly classification loss until a second training end condition is reached.

[0055] According to one aspect of the embodiments of the present application, an electronic device is provided, including: a processor; a memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the above text generation method is implemented.

[0056] According to one aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by a processor, the above text generation method is implemented.

[0057] According to one aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, and when the computer instructions are executed by a processor, the above text generation method is implemented.

[0058] In this application, after identifying that the target trading account is abnormal, among the trading characteristic information of multiple normal trading accounts, the K trading characteristic information that is most similar to the trading characteristic information of the target trading account is filtered out, and the K trading characteristic information filtered out is compared with the trading characteristic information of the target trading account to identify the abnormal trading characteristics in the trading characteristic information of the target trading account, and the identified abnormal trading characteristics are used to generate the abnormal reason description text of the target trading account, without using the normal trading characteristics in the trading characteristic information of the target trading account to generate the abnormal reason description text. In this way, it is possible to avoid the interference of normal trading characteristics on the generation process of the abnormal reason description text, effectively alleviate the problem of model hallucination, and improve the accuracy of the generated abnormal reason description text.

[0059] Moreover, the K trading characteristic information that is most similar to the trading characteristic information of the target trading account in the normal trading accounts is used as a reference, rather than using the trading characteristic information of the normal trading accounts that has a large difference from the trading characteristic information of the target trading account as a reference. In this way, it is possible to accurately identify the boundary between normal and abnormal trading characteristics, thereby ensuring the accuracy of abnormal trading characteristic identification, and further ensuring the accuracy of the abnormal reason description text generated based on the abnormal trading characteristics. Brief Description of the Drawings

[0060] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application, and those of ordinary skill in the art can obtain other accompanying drawings based on these drawings without creative efforts.

[0061] Figure 1 is a schematic diagram of the application scenario of this application shown according to an embodiment of this application.

[0062] Figure 2 is a flowchart of a text generation method shown according to an embodiment of this application.

[0063] Figure 3 is a flowchart of a text generation method shown according to another embodiment of this application.

[0064] Figure 4 is a flowchart of a text generation method shown according to another embodiment of this application.

[0065] Figure 5 is a schematic structural diagram of a text generation model shown according to an embodiment of this application.

[0066] Figure 6It is a flowchart of a text generation method shown according to another embodiment of the present application.

[0067] Figure 7 It is a flowchart of constructing a target isolation forest shown according to another embodiment of the present application.

[0068] Figure 8 It is a flowchart of a text generation method shown according to another embodiment of the present application.

[0069] Figure 9 An exemplary flowchart of training a model participating in implementing the method of the present application is shown.

[0070] Figure 10 It is a flowchart of a text generation method shown according to another embodiment of the present application.

[0071] Figure 11 It is a block diagram of a text generation device shown according to an embodiment of the present application.

[0072] Figure 12 A schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application is shown. Detailed implementation manners

[0073] The following details the implementation manners of the present application. Examples of the implementation manners are shown in the drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below with reference to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application.

[0074] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0075] In the following description, the terms "first / second" etc. involved are only used to distinguish similar objects and do not represent a specific order for the objects. It can be understood that "first / second" can be interchanged with a specific order or sequence when allowed, so that the embodiments of the present application described here can be implemented in an order other than that illustrated or described here.

[0076] As used herein, "a plurality of" means two or more. " / or" describes the relationship between associated objects and indicates that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In the following description, when referring to "some embodiments or some implementation manners", it describes a subset of all possible embodiments. However, it can be understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and they can be combined with each other without conflict.

[0077] Figure 1 is a schematic diagram of the application scenario of the present application shown according to an embodiment of the present application. As Figure 1 shown, the application scenario includes an electronic device 120, and the electronic device 120 can obtain the transaction feature information of the target transaction account to be detected from the transaction data set 110. The transaction feature information of each transaction account in the transaction data set 110 can be obtained by extracting transaction features based on the transaction records of the transaction account, and the transaction feature information of a transaction account can include the feature values of multiple transaction features.

[0078] Then, the electronic device 120 performs anomaly recognition based on the transaction feature information of the target transaction account to obtain the recognition label of the target transaction account; after the electronic device 120 identifies that the target transaction account has a risk, that is, the recognition label of the target transaction account indicates that the target transaction account is an abnormal transaction account, among the transaction feature information of multiple normal transaction accounts, it determines the K transaction feature information that is most similar to the transaction feature information of the target transaction account; a normal transaction account refers to a transaction account whose recognition label indicates normal; K is an integer greater than 1; and based on the K transaction feature information, it performs anomaly transaction feature recognition on the transaction feature information of the target transaction account to determine the anomaly transaction features in the transaction feature information of the target transaction account; finally, the electronic device 120 can generate text according to the target feature value of the anomaly transaction feature in the transaction feature information of the target transaction account to obtain the anomaly reason description text of the target transaction account. This anomaly reason description text is used to explain the reason for identifying the target transaction account as having a risk or identifying the target transaction account as an abnormal transaction account.

[0079] Among them, the electronic device 120 can be a terminal, an edge computing device, a vehicle-mounted device, a server, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0080] The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below:

[0081] Figure 2 It is a flowchart of a text generation method shown according to an embodiment of the present application. This method can be executed by an electronic device.

[0082] Refer to Figure 2 As shown, this method at least includes steps 210 to 250, which are introduced in detail as follows:

[0083] Step 210, obtain the transaction feature information of the target transaction account.

[0084] The target transaction account refers to the transaction account to be currently detected. A transaction account refers to an account participating in a transaction. It can be an account in an application with transaction functions. Applications with transaction functions can be instant messaging applications, shopping applications, video applications, live broadcast applications, music applications, etc., which are not specifically limited here. In some embodiments, the transaction account can also be a financial account registered with a financial institution.

[0085] The transaction feature information of a transaction account can include multiple transaction features extracted from the transaction data of the transaction account within a preset time period. The duration of the preset time period can be set according to actual needs, such as one week, 10 days, one month, three months, five months, etc. In some embodiments, the transaction feature information of a transaction account can also include the account features of the transaction account, such as the gender of the object to which the transaction account belongs, the geographical location, age, and registration time of the transaction account.

[0086] The transaction data of a transaction account within a preset time period includes the transaction-related information of each transaction participated in by the transaction account within the preset time period. The transaction-related information includes, for example, the transaction initiator, the transaction recipient, the transaction amount, the transaction time, the transaction status, and the payment method (such as the payment application called).

[0087] In some embodiments, the transaction features may include transaction statistical features. The transaction statistical features refer to the features that reflect the overall situation of a transaction account within a preset time period by statistically analyzing the transaction basic features of each transaction of the transaction account within the preset time period. The transaction statistical features may be obtained by statistical analysis in different statistical dimensions. For example, the statistical dimensions are as follows: ① The transaction frequency of the transaction account within the preset time period (such as transaction frequency and number of transactions); ② The average transaction time interval between two adjacent transactions of the transaction account within the preset time period; ③ The total transaction amount and / or number of transactions of the transaction account within the preset time period (or in each time segment within the preset time period); ④ The total transaction amount and / or average transaction amount between the transaction account and the same transaction account within the preset time period; ⑤ The total number of transactions and / or total transaction amount of the transaction account with each counterparty transaction account within the preset time period (or in each time segment within the preset time period), etc. The above-listed statistical dimensions are only exemplary examples and should not be considered as a limitation on the scope of use of this application.

[0088] In some embodiments, the transaction features may also include the transaction basic features of each transaction, such as the transaction initiator, transaction recipient, transaction amount, transaction time, transaction status, and payment method listed above.

[0089] Step 220: Perform anomaly identification based on the transaction feature information of the target transaction account to obtain the identification label of the target transaction account.

[0090] The identification label of the target transaction account is used to indicate whether the target transaction account is a normal transaction account or an abnormal transaction account. An abnormal transaction account refers to a transaction account with transaction risks, such as financial fraud risks and black production transaction risks.

[0091] Anomaly identification can be performed on the transaction feature information of the target transaction account through an anomaly classification model to obtain the identification label of the target transaction account. The identification label of the target transaction account is used to indicate whether the target transaction account is abnormal. In some embodiments, the identification label may include a first identification label indicating that there is no risk for the transaction account (i.e., the label indicating that the transaction account is normal) and a second identification label indicating that there is a risk for the transaction account (i.e., the label indicating that the transaction account is an abnormal transaction account). In this case, the anomaly classification model is used for binary classification, and the identification label output for a transaction account is one of the first identification label and the second identification label.

[0092] In some other embodiments, the situation indicating that there is a risk for the transaction account may also be divided into multiple anomaly types, and each anomaly type corresponds to an identification label. In this case, the anomaly classification model is used for multi-classification.

[0093] The anomaly classification model can be a deep neural network model or a tree model (such as a decision tree model, a random forest model, a gradient boosting tree model, etc.). In some embodiments, if the anomaly classification model is used for binary classification, the anomaly classification model can be a decision tree model based on the XGBoost (eXtreme Gradient Boosting) algorithm. During the training process, multiple decision tree models can be iteratively trained and combined into a powerful ensemble model as the anomaly classification model. In each round of iteration, the anomaly classification model can adjust the weights of the samples according to the previous prediction results so that the model can pay more attention to the samples with incorrect predictions, thereby improving the performance of the overall model.

[0094] Step 230, if the identification label indicates that the target trading account is abnormal, among the trading feature information of multiple normal trading accounts, determine the K trading feature information that is most similar to the trading feature information of the target trading account; a normal trading account refers to a trading account whose identification label indicates normal; K is an integer greater than 1.

[0095] If through step 220, it is identified that the target trading account is a normal trading account, then the processes of steps 230 - 250 do not need to be executed. Only when it is identified that the target trading account is abnormal (i.e., there is a risk), is it further necessary to generate the text description of the anomaly reason.

[0096] Among them, the time period corresponding to the trading feature information of multiple normal trading accounts can be the same as the time period to which the trading feature information of the target trading account belongs. For example, if the trading feature information of the target trading account is extracted from the trading data of the target trading account in the recent three months, then the trading feature information of the multiple normal trading accounts can also be extracted from the trading data of the corresponding trading accounts in the recent three months. A normal trading account can also be identified by the above-mentioned anomaly classification model according to the corresponding trading feature information.

[0097] In some embodiments, the feature similarity between the trading feature information of the target trading account and the trading feature information of each normal trading account can be calculated, and then, according to the calculated feature similarity, among the trading feature information of multiple normal trading accounts, select the K trading feature information that is most similar to the trading feature information of the target trading account. The feature similarity can be cosine similarity, Euclidean distance, Manhattan distance, etc., which are not specifically limited herein.

[0098] In some embodiments, considering that there are significant differences in the transaction basic characteristics of each transaction involved in different transaction accounts, in order to avoid using the transaction basic characteristics for calculating feature similarity, the resulting results may not accurately reflect the feature similarity between different transaction accounts. Therefore, it is possible to calculate the feature similarity between the transaction statistical characteristics in the transaction feature information of the target transaction account and the transaction statistical characteristics in the transaction feature information of each normal transaction account, while the transaction basic characteristics do not participate in the calculation of feature similarity.

[0099] Step 240: Based on the K transaction feature information, identify the abnormal transaction features in the transaction feature information of the target transaction account, and determine the abnormal transaction features in the transaction feature information of the target transaction account.

[0100] Each transaction feature information includes the feature values of multiple transaction features. For example, the feature values of the above-mentioned transaction statistical features, such as the transaction frequency of the transaction account within the preset time period, which is a transaction statistical feature, and the numerical value of the counted transaction frequency is the corresponding feature value. Since the K transaction feature information is from normal transaction accounts, the probability that the feature values of each transaction feature in the K transaction feature information are normal is very high. For the target transaction account, identifying the target transaction account as having risks may be caused by the abnormal feature values of some or all transaction features in the transaction feature information of the target transaction account. That is to say, although there must be transaction features with abnormal feature values in the transaction feature information of the target transaction account, there may also be transaction features with normal feature values.

[0101] Therefore, in this application, using the feature values of each transaction feature in the determined transaction feature information of the K normal transaction accounts as a reference, to evaluate whether the feature values of each transaction feature in the transaction feature information of the target transaction account are normal. If the feature value of a transaction feature in the transaction feature information of the target transaction account is abnormal, this transaction feature can be determined as the abnormal transaction feature (also called the risk transaction feature) in the transaction feature information of the target transaction account.

[0102] In some embodiments, for each transaction feature in the transaction feature information, the feature value of the transaction feature in each piece of transaction feature information among the K pieces of transaction feature information is added to the normal feature value set corresponding to the transaction feature. After that, according to the normal feature value set corresponding to each transaction feature (taking transaction feature A as an example), the normal feature value range of transaction feature A is constructed. If the feature value of transaction feature A in the transaction feature information of the target transaction account is within the normal feature value range of transaction feature A, it is determined that transaction feature A in the transaction feature information of the target transaction account is not an abnormal transaction feature; otherwise, if the feature value of transaction feature A in the transaction feature information of the target transaction account exceeds the normal feature value range of transaction feature A, it is determined that transaction feature A in the transaction feature information of the target transaction account is an abnormal transaction feature.

[0103] In some embodiments, the minimum feature value in the normal feature value set corresponding to transaction feature A may be used as the lower limit feature value in the normal feature value range of transaction feature A, and the maximum feature value in the normal feature value set corresponding to transaction feature A may be used as the upper limit feature value in the normal feature value range of transaction feature A.

[0104] In some other embodiments, the difference between the minimum feature value in the normal feature value set corresponding to transaction feature A and the specified surplus amount may be used as the lower limit feature value in the normal feature value range of transaction feature A, and the difference between the maximum feature value in the normal feature value set corresponding to transaction feature A and the specified surplus amount may be used as the upper limit feature value in the normal feature value range of transaction feature A. The specified surplus amount may be an integer.

[0105] In some other embodiments, the feature values in the normal feature value set corresponding to transaction feature A may be averaged to obtain the average feature value corresponding to transaction feature A. After that, the difference between the average feature value corresponding to transaction feature A and the first offset may be used as the lower limit feature value in the normal feature value range of transaction feature A, and the sum of the average feature value corresponding to transaction feature A and the first offset may be used as the upper limit feature value in the normal feature value range of transaction feature A. The first offset may be the average distance obtained by averaging the distances between the feature values in the normal feature value set corresponding to transaction feature A and the average feature value corresponding to transaction feature A.

[0106] In some other embodiments, in step 240, each transaction feature may be processed according to the following Figure 3 procedure as shown, such as Figure 3 shown, including step 310 - step 330:

[0107] Step 310: Perform distribution fitting on the reference eigenvalue of the transaction feature in the K transaction feature information and the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account to determine the target feature distribution corresponding to the transaction feature.

[0108] The reference eigenvalue refers to the eigenvalue of the current transaction feature in each transaction feature information among the determined K transaction feature information. The distribution followed by the eigenvalues of a transaction feature is called the target feature distribution corresponding to the transaction feature. Considering that a distribution can be described by distribution parameters, therefore, determining the target feature distribution corresponding to the transaction feature is equivalent to determining the values of the distribution parameters used to describe the target feature distribution. For example, if the target feature distribution is a normal distribution (also known as a Gaussian distribution), this target feature distribution can be described by two distribution parameters, the mean and the variance. On this basis, the mean and variance of the K reference eigenvalues corresponding to the transaction feature can be calculated based on the reference eigenvalue of the transaction feature in the K transaction feature information and the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account to determine the target feature distribution.

[0109] Step 320: Determine the target normal eigenvalue range corresponding to the target feature distribution according to the target feature distribution corresponding to the transaction feature.

[0110] For a normal distribution, assume that the mean of a normal distribution is μ and the variance is σ 2 , the normal value range of a variable following a normal distribution can be any one of: [μ - σ, μ + σ], [μ - 2σ, μ + 2σ], and [μ - 3σ, μ + 3σ]. Through the above Step 310, the mean μ and variance σ used to describe the target feature distribution can be determined 2 After that, the target normal eigenvalue range corresponding to the target feature distribution can be determined accordingly. For example, if it is determined that the normal value range of a variable following a normal distribution is [μ - σ, μ + σ], then the mean μ and variance σ of the target feature distribution 2 are substituted for calculation, and the target normal eigenvalue range corresponding to the target feature distribution can be obtained.

[0111] In some embodiments, considering that different transaction features are different variables, and the normal value ranges applicable to different variables may be different. For a variable following a normal distribution, the normal value range of this variable can be: [μ - i * σ, μ + i * σ], where i = 1, 2, 3, etc. In some embodiments, based on the reference eigenvalue of the transaction feature in the K transaction feature information, and the mean μ1 and variance σ determined for the target feature distribution 21. Among the three candidate normal value ranges of [μ1 - σ1, μ1 + σ1], [μ1 - 2σ1, μ1 + 2σ1], and [μ1 - 3σ1, μ1 + 3σ1], determine the reference normal value range corresponding to the transaction feature. The reference normal value range corresponding to the transaction feature includes the feature values that exceed a specified proportion among the reference feature values of the transaction feature in the K transaction feature information, and use the determined reference normal value range as the target normal feature value range corresponding to the target feature distribution applicable to the current transaction feature. Among them, the specified proportion can be set as needed. For example, the specified proportion can be 100%, 98%, 95%, 90%, etc.

[0112] In some embodiments, considering that different transaction features are different variables, and the range constraint parameters of the normal value ranges applicable to different variables are different. For variables that follow a normal distribution, the normal value range of the variable can be described as: [μ - i*σ, μ + i*σ], where i is a positive integer, such as i = 1, 2, 3, etc., and i is the range constraint parameter of the normal value range of the variable. The range constraint parameters of the normal value ranges for different transaction features can be determined through the first training data.

[0113] The first training data may include the third transaction feature information of multiple second sample target transaction accounts, and the fourth transaction feature information of multiple second sample normal transaction accounts that are most similar to each third transaction feature information. The second sample target transaction account refers to a transaction account identified as having risks (i.e., an abnormal transaction account), and the third transaction feature information refers to the transaction feature information of the second sample target transaction account; the second sample normal transaction account refers to any one of the N normal transaction accounts with the highest similarity between the transaction feature information and the third transaction feature information of the second sample target transaction account. N can be the same as K in the above text or different, and N is a positive integer greater than 1.

[0114] Based on the third transaction feature information of multiple second sample target transaction accounts and the fourth transaction feature information of multiple second sample normal transaction accounts, for each transaction feature (assumed to be transaction feature A), the feature values of transaction feature A in the third transaction feature information of each second sample target transaction account can be added to the first feature value set corresponding to transaction feature A, and the feature values of transaction feature A in the fourth transaction feature information of each second sample normal transaction account can be added to the second feature value set corresponding to transaction feature A; then, the feature values in the first feature value set corresponding to transaction feature A and the feature values in the second feature value set corresponding to it are subjected to distribution fitting to determine the distribution parameters of the first feature distribution followed by transaction feature A, that is, the mean and variance. Assume the mean is μ2 and the variance is

[0115] After that, among the three candidate normal value ranges of [μ2 - σ2, μ2 + σ2], [μ2 - 2σ2, μ2 + 2σ2] and [μ2 - 3σ2, μ2 + 3σ2], determine the smallest candidate normal value range that contains the eigenvalues of the transaction feature A exceeding the specified proportion in the second eigenvalue set. After that, take the value of i corresponding to the determined smallest candidate normal value range as the range constraint parameter applicable to the transaction feature A. For example, if the determined smallest candidate normal value range is [μ2 - 2σ2, μ2 + 2σ2], then the range constraint parameter i applicable to the transaction feature A is 2.

[0116] On this basis, based on the mean and variance corresponding to the target feature distribution determined for the transaction feature A in the above text, and the range constraint parameter applicable to the transaction feature A, the target normal eigenvalue range corresponding to the target feature distribution can be correspondingly determined. For example, if the mean determined for the target feature distribution followed by the transaction feature A is μ1 and the variance is σ 2 1, and the determined range constraint parameter for the transaction feature A is i1, then the target normal eigenvalue range corresponding to the target feature distribution followed by the transaction feature A is [μ1 - i1 * σ1, μ1 + i1 * σ1].

[0117] In the above embodiment, the range constraint parameter applicable to each transaction feature is determined by means of the first training data. Since there are more normal eigenvalues for each transaction feature in the first training data, the determined range constraint parameter is more accurate, thus ensuring the accuracy of the subsequent anomaly judgment.

[0118] Step 330, if the target eigenvalue exceeds the target normal eigenvalue range, determine that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature.

[0119] If the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account is within the target normal eigenvalue range, it can be determined that the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account is normal. Therefore, it can be determined that the transaction feature in the transaction feature information of the target transaction account is normal; on the contrary, if the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account exceeds the target normal eigenvalue range, it can be determined that the target eigenvalue of the transaction feature in the transaction feature information of the target transaction account is abnormal. Therefore, it can be determined that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature. That is to say, whether a transaction feature is a normal transaction feature or an abnormal transaction feature actually refers to whether the eigenvalue of the transaction feature is normal or abnormal.

[0120] Step 250: Generate text based on the target feature value in the transaction feature information of the target transaction account according to the abnormal transaction feature, and obtain the text description of the reason for the abnormality of the target transaction account.

[0121] The text description of the reason for the abnormality of the target transaction account refers to the text used to describe the reason for identifying the target transaction account as abnormal. Thus, through the text description of the reason for the abnormality of the target transaction account, it can be directly known in which aspects the transaction account is abnormal.

[0122] In some embodiments, the name of the abnormal transaction feature, the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account, and the risk description prompt text can be input into the large language model. By describing the prompt text, the large language model is guided to generate the risk description text of the target transaction account according to the name of the abnormal transaction feature and the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account. Among them, the name of the abnormal transaction feature is used for the large language model to understand the physical meaning of the abnormal transaction feature. The description prompt text includes instruction text for instructing the large language model to generate the text description of the reason for the abnormality. The instruction text is, for example: Please generate a text explaining the reason why the corresponding transaction account is identified as abnormal according to the name of the input transaction feature and the feature value of the transaction feature.

[0123] In some embodiments, the risk description prompt text may further include description examples. The description examples include the name of the given sample abnormal transaction feature, the abnormal feature value of the sample abnormal transaction feature, and the corresponding sample text description of the reason for the abnormality. The description examples serve as examples for the large language model to generate the text description of the reason for the abnormality.

[0124] In other embodiments, a text generation model can be used to generate the risk description text of the target transaction account according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account and the name of the abnormal transaction feature. For example, the name of the abnormal transaction feature and the feature value of the abnormal transaction feature in the transaction feature information of the target transaction account are input into the text generation model, and the text generation model performs text generation to output the text description of the reason for the abnormality of the target transaction account.

[0125] The text generation model is a text generation model constructed by a neural network. The text generation model can be trained with training data so that the text generation model has the ability to generate the text description of the reason for the abnormality according to the input name of the abnormal transaction feature and the abnormal transaction feature. The training process of the text generation model can be referred to the description below.

[0126] In some embodiments, before step 240, the method further includes: performing distribution fitting on the feature values of the abnormal transaction features in the K transaction feature information and the corresponding target feature values to obtain the target feature distribution corresponding to the abnormal transaction features; correspondingly, step 240 includes: using a text generation model to generate a text description of the abnormal reason of the target transaction account according to the target feature values of the abnormal transaction features in the transaction feature information of the target transaction account and the target feature distribution corresponding to the abnormal transaction features.

[0127] In this embodiment, the target feature values of the abnormal transaction features in the transaction feature information of the target transaction account, the names of the abnormal transaction features, and the target feature distribution corresponding to the abnormal transaction features can be input into the text generation model. The text generation model understands and generates text, and outputs a text description of the abnormal reason of the target transaction account. The target feature distribution corresponding to the input abnormal transaction feature is used for the text generation model to understand the position of the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account in the target feature distribution, so as to more accurately understand the reason why the target transaction account is identified as having a risk, or more accurately understand the reason why the feature value of the abnormal transaction feature in the transaction feature information of the target transaction account is identified as abnormal.

[0128] In some embodiments, if there are multiple abnormal transaction features in the transaction feature information of the target transaction account, the multiple abnormal transaction features in the transaction feature information of the target transaction account can be input into the text generation model together, and the text generation model outputs a text description of the abnormal reason. This text description of the abnormal reason can explain the reason why the target transaction account is identified as abnormal caused by multiple abnormal transaction features.

[0129] In other embodiments, if there are multiple abnormal transaction features in the transaction feature information of the target transaction account, each abnormal transaction feature in the transaction feature information of the target transaction account can be input into the text generation model separately, and the text generation model outputs a text description of the abnormal reason. This text description of the abnormal reason can explain the reason why the target transaction account is identified as having a risk caused by the corresponding abnormal transaction feature. Finally, the text descriptions of the abnormal reasons generated for multiple abnormal transaction features can be combined as the comprehensive text description of the abnormal reason of the target transaction account.

[0130] In some embodiments, after step 250, the text description of the abnormal reason of the target transaction account and the identification label of the target transaction account can be sent to the audit side, and the auditors on the audit side perform manual review on this transaction risk account. In some embodiments, the feature values of the abnormal transaction features in the transaction feature information of the target transaction account can also be sent to the audit side as reference information for manual review.

[0131] In this application, after identifying that the target trading account is abnormal, among the trading characteristic information of multiple normal trading accounts, the K trading characteristic information that is most similar to the trading characteristic information of the target trading account is screened out. The K trading characteristic information screened out and the trading characteristic information of the target trading account are used to identify the abnormal trading characteristics in the trading characteristic information of the target trading account, and the identified abnormal trading characteristics are used to generate the text description of the abnormal reason of the target trading account, without using the normal trading characteristics in the trading characteristic information of the target trading account to generate the text description of the abnormal reason. In this way, the interference of normal trading characteristics in the generation process of the text description of the abnormal reason can be avoided, the problem of model hallucination can be effectively alleviated, and the accuracy of the generated text description of the abnormal reason can be improved.

[0132] Moreover, the K trading characteristic information that is most similar to the trading characteristic information of the target trading account in the normal trading accounts is used as a reference, rather than using the trading characteristic information of the normal trading accounts with a large difference from the trading characteristic information of the target trading account as a reference. In this way, the boundary between normal and abnormal trading characteristics can be accurately identified, thereby ensuring the accuracy of the identification of abnormal trading characteristics, and further ensuring the accuracy of the text description of the abnormal reason generated according to the abnormal trading characteristics.

[0133] In some embodiments, the text generation model can be trained according to the Figure 4 process shown, as Figure 4 shown, including:

[0134] Step 410, obtain a plurality of training samples. A training sample includes sample abnormal trading characteristics, the sample characteristic distribution corresponding to the sample abnormal trading characteristics, and the sample text description of the abnormal reason corresponding to the sample abnormal trading characteristics.

[0135] The sample abnormal trading characteristics are for the trading characteristic information of a sample trading account determined to be abnormal (at risk). In other words, the sample abnormal trading characteristics refer to the trading characteristics with abnormal characteristic values in the trading characteristic information of the sample trading account determined to be at risk. The sample abnormal trading characteristics in a training sample can include the name of the sample abnormal trading characteristics and the characteristic value determined to be abnormal for the sample abnormal trading characteristics. In some embodiments, the sample abnormal trading characteristics can be the trading characteristics marked with abnormal characteristic values in the trading characteristic information of the sample trading account indicated by the identification label as being at risk.

[0136] In some other embodiments, in a similar manner as above, transaction feature information of a sample transaction account with a risk indicated by an identification tag can be obtained, and transaction feature information of P sample normal transaction accounts with the highest similarity to the transaction feature information between the transaction feature information and that of the sample transaction account can be obtained. Based on the transaction feature information of the P sample normal transaction accounts, abnormal transaction features in the transaction feature information of the sample transaction account marked as having a risk are determined as sample abnormal transaction features. Correspondingly, in combination with the eigenvalue of the sample abnormal transaction feature in the transaction feature information of the sample transaction account and the eigenvalue in the transaction feature information of the corresponding P sample normal transaction accounts, a sample feature distribution corresponding to the sample abnormal transaction feature is determined. The specific process is similar to the above and will not be elaborated here.

[0137] The sample abnormal reason description text corresponding to the sample abnormal transaction feature can be manually written, which describes the reason for the risk (abnormality) of the sample transaction account from which the sample abnormal transaction feature is derived. Or rather, the sample abnormal reason description text corresponding to the sample abnormal transaction feature can be a text generated by manually combining the eigenvalue of the sample abnormal transaction feature in the transaction feature information of the corresponding sample transaction account and the sample feature distribution corresponding to the sample abnormal transaction feature for abnormal reason description.

[0138] Step 420: The text generation model generates text based on the sample abnormal transaction feature and the sample feature distribution corresponding to the sample abnormal transaction feature to obtain a predicted abnormal reason description text.

[0139] The text generation model can jointly encode the name of the sample abnormal feature, the eigenvalue of the sample abnormal feature, and the sample feature distribution corresponding to the sample abnormal transaction feature to obtain a jointly encoded feature. Then, the jointly encoded feature is decoded to output a predicted abnormal reason description text.

[0140] Step 430: Calculate a text generation loss based on the sample abnormal reason description text and the predicted abnormal reason description text.

[0141] In some embodiments, according to a first loss function, the difference between the sample abnormal reason description text and the predicted abnormal reason description text is calculated as the text generation loss. The first loss function can be a cross-entropy loss function, a mean squared error loss function, a cosine loss function, etc., and no specific limitation is made here.

[0142] Step 440: Adjust the parameters of the text generation model according to the text generation loss until a first training end condition is reached.

[0143] The gradient of the text generation loss with respect to the parameters of the text generation model can be calculated, and then, according to the gradient descent algorithm, the parameters of the text generation model can be adjusted until the first training end condition is reached. The first training end condition can be that the number of iterations of the text generation model reaches the first number threshold, or the text generation loss converges, which is not specifically limited herein. The parameters to be adjusted can be all the parameters in the text generation model or some of them.

[0144] In some embodiments, the text generation model includes a large language model and a fine-tuning network; correspondingly, step 440 includes: fixing the parameters of the large language model and adjusting the parameters of the fine-tuning network according to the text generation loss until the first training end condition is reached.

[0145] In this embodiment, the parameters of the large language model are fixed, the gradient of the text generation loss with respect to the parameters of the fine-tuning network is calculated, and the parameters of the text generation model are adjusted according to the gradient descent algorithm until the first training end condition is reached.

[0146] Figure 5 An exemplary structural schematic diagram of the text generation model is shown, such as Figure 5 shown. The large language model is on the left. During training, the parameter matrix W of the large language model is fixed, and the dimension of the parameter matrix of the large language model is d×k. On the right is the fine-tuning network. The matrix formed by the parameters of the fine-tuning network is a low-rank separable matrix. In other words, the fine-tuning network includes a first neural network and a second neural network connected in series. The parameter matrix of the first neural network is matrix A, and matrix A can be initialized with a random Gaussian distribution Ν(0,σ 2 )). The parameter matrix of the second neural network is matrix B, and matrix B can be initialized with a 0 matrix. Correspondingly, during the training stage, the parameters in matrix A and matrix B are adjusted. The input dimension (d) of the first neural network is greater than the output dimension (r) of the first neural network, the input dimension (r) of the second neural network is less than the output dimension of the second neural network, and the input dimension of the first neural network is the same as the output dimension of the second neural network. That is to say, the first neural network is responsible for dimensionality reduction, and the second neural network is responsible for dimensionality increase. The two low-rank matrices, matrix A and matrix B, can greatly reduce the number of parameters.

[0147] Since large language models have been pre-trained with a large amount of text data, they have relatively strong text semantic understanding capabilities. Therefore, in this embodiment, by leveraging the semantic understanding capabilities of large language models and adding a fine-tuning network on top of the large language model, the trained model can be made suitable for the current risk description text generation task, ensuring compatibility with the current task of generating abnormal cause description texts. Moreover, by fixing the parameters of the large language model and only adjusting the parameters of the fine-tuning network, the number of parameters to be adjusted can be significantly reduced, greatly reducing the time required to train the text generation model. Additionally, by fixing the parameters of the large language model during the training phase and adjusting the parameters of the fine-tuning network, the generalization ability of the trained text generation model in the task of generating abnormal cause description texts can be improved, ensuring compatibility with this task.

[0148] Furthermore, by performing supervised training on the text generation model using sample abnormal cause description texts, the text generation model can learn the topic content to be described under different abnormal feature values of different transaction characteristics. This can effectively alleviate the hallucination problem and reduce the situation where the generated abnormal cause description text does not match the actual situation or does not meet the actual needs.

[0149] In some other embodiments, the transaction feature information includes the feature values of multiple transaction features; in step 240, the process shown in steps 610 - 660 can be performed for each transaction feature as follows: Figure 6 The process shown in steps 610 - 660 is performed for each transaction feature:

[0150] Step 610: Locate the position of the target feature value of the transaction feature in the transaction feature information of the target transaction account in each binary tree of the target isolation forest, and determine multiple first path lengths of the target feature value in the target isolation forest; the target isolation forest refers to an isolation forest suitable for anomaly detection of transaction features.

[0151] For ease of distinction, for a transaction feature, the feature value of the transaction feature in the transaction feature information of the target transaction account is referred to as the target feature value; the feature value of the transaction feature in the K transaction feature information is referred to as the reference feature value.

[0152] An isolation forest can be constructed in advance for each transaction feature to obtain an isolation forest suitable for anomaly detection of each transaction feature. For each transaction feature, the isolation forest suitable for anomaly detection of the transaction feature is referred to as the target isolation forest, and the target isolation forests corresponding to different transaction features can be different. The target isolation forest corresponding to a transaction feature includes multiple binary trees.

[0153] For each transaction feature, before step 610, it can be performed according to Figure 7The process shown constructs the target isolation forests corresponding to each transaction feature, as Figure 7 shown, including:

[0154] Step 710: Obtain the first transaction feature information of the first sample abnormal transaction account and the second transaction feature information of multiple first sample normal transaction accounts that are most similar to the first transaction feature information.

[0155] The abnormal transaction accounts used to construct the target isolation forests of each transaction feature are called the first sample abnormal transaction accounts, and the normal transaction accounts used to construct the target isolation forests of each transaction feature are called the first sample normal transaction accounts. Among them, the transaction features involved in the first transaction feature information and the second transaction feature information are the same as those involved in the transaction feature information in step 210 above, except that the feature values of the transaction features may be different.

[0156] Step 720: Construct the eigenvalue set corresponding to each transaction feature according to the eigenvalue of the same transaction feature in the first transaction feature information and the eigenvalue in the second transaction feature information.

[0157] For each transaction feature, obtain the eigenvalue of this transaction feature from the first transaction feature information, and obtain the eigenvalue of this transaction feature from each second transaction feature information, and add the eigenvalue of the same transaction feature in the first transaction feature information and the eigenvalues in multiple second transaction feature information to the same eigenvalue set as the eigenvalue set corresponding to this transaction feature.

[0158] Step 730: Construct T binary trees corresponding to each transaction feature; T is an integer greater than 1.

[0159] For each transaction feature, in the set of feature values corresponding to the transaction feature, a feature value is randomly selected between the maximum feature value and the minimum feature value of the transaction feature. The selected feature value is used as the root node in a binary tree, and the feature value represented by the root node is used as a cut-off point. After that, in the set of feature values corresponding to the transaction feature, the feature values less than the cut-off point can be placed on the left branch node of the root node, and the feature values greater than or equal to the cut-off point can be placed on the right branch node of the root node. Subsequently, for the left branch node and the right branch node of the root node, one node is selected as the target node, and a feature value is randomly selected from the feature values contained in the target node as the split point corresponding to the target node. Then, among the feature values contained in the target node, the feature values less than the split point corresponding to the target node are placed on the left branch node of the target node, and the feature values greater than or equal to the split point corresponding to the target node are placed on the right branch node of the target node. And so on, new nodes are continuously constructed until there is only one feature value on the leaf node (indicating that no further splitting is possible), or the binary tree has grown to the set height. In this way, the construction of a binary tree can be completed. After that, in a similar manner, other binary trees are constructed until the number of constructed binary trees reaches T.

[0160] Step 740, integrate the T binary trees corresponding to the same transaction feature to obtain a target isolation forest applicable to anomaly detection for the corresponding transaction feature.

[0161] The T binary trees corresponding to the same transaction feature respectively serve as a target isolation forest applicable to anomaly detection for the corresponding transaction feature. For each transaction feature, the process similar to that shown Figure 7 can be used to construct a target isolation forest applicable to anomaly detection for the corresponding transaction feature.

[0162] In each binary tree of the target isolation forest corresponding to a transaction feature, if a feature value of the transaction feature is given, on this binary tree, the node corresponding to the location of the feature value on the binary tree can be located (that is, the node whose contained feature values include the feature value of the transaction feature). Then, calculate the path length from the root node of this binary tree to the node corresponding to the location of the feature value on the binary tree, that is, the number of nodes on the path from the root node of this binary tree to the node corresponding to the location of the feature value on the binary tree. Generally, a shorter path length is considered an abnormal feature value, while a longer path length is considered a normal feature value.

[0163] After constructing the target isolation forests corresponding to each transaction feature, for each binary tree in the target isolation forest applicable to a transaction feature, the node where the target feature value of the transaction feature is located in the binary tree can be determined, and the path length from the root node of the binary tree to the node where the target feature value of the transaction feature is located in the binary tree is calculated correspondingly, obtaining the first path length of the transaction feature in the binary tree.

[0164] Step 620, locate the positions of the reference feature values of the transaction feature in the K transaction feature information in each binary tree of the target isolation forest, and determine multiple second path lengths of each reference feature value in the target isolation forest.

[0165] Similar to step 610, after constructing the target isolation forests corresponding to each transaction feature, for each reference feature value of the current transaction feature, in each binary tree of the target isolation forest applicable to the transaction feature, the node where the reference feature value of the transaction feature is located in the binary tree can be determined, and the path length from the root node of the binary tree to the node where the reference feature value of the transaction feature is located in the binary tree is calculated correspondingly, obtaining the second path length of the transaction feature in the binary tree.

[0166] Since the binary trees in the target isolation forest applicable to the current transaction feature are different (that is, the feature values represented by the nodes at different levels or the feature values included are different), therefore, for each reference feature value, the second path length of the reference feature value in each binary tree is determined in a similar manner. Since the target isolation forest applicable to the current transaction feature includes T binary trees, therefore, for a reference feature value of the transaction feature, the path lengths on T binary trees can be correspondingly determined, obtaining T second path lengths.

[0167] Step 630, calculate the mean of the multiple first path lengths of the target feature value in the target isolation forest, obtaining the first average path length.

[0168] By calculating the mean of the first path lengths corresponding to the T binary trees in the target isolation forest for the target feature value, the first average path length can be obtained. This first average path length reflects the average path length distribution of the target feature value in the target isolation forest.

[0169] Step 640, calculate the mean of the multiple second path lengths of each reference feature value in the target isolation forest, obtaining the second average path length of each reference feature value in the target isolation forest.

[0170] For each reference eigenvalue, calculate the average of the second path lengths corresponding to the reference eigenvalue on the T binary trees in the target isolation forest, and the second average path length of the reference eigenvalue in the target isolation forest can be obtained. The second average path length of a reference eigenvalue in the target isolation forest reflects the average path length distribution of the corresponding reference eigenvalue in the corresponding target isolation forest.

[0171] Step 650, based on the first average path length and the second average path lengths of each reference eigenvalue in the target isolation forest, identify abnormal features among the trading features in the trading feature information of the target trading account.

[0172] In some embodiments, step 650 includes: determining a path length threshold according to the second average path lengths of the K reference eigenvalues in the target isolation forest; if the first average path length is less than the path length threshold, determining that the trading feature in the trading feature information of the target trading account is an abnormal trading feature.

[0173] Since the K reference eigenvalues are from the same trading feature in the trading feature information of normal trading accounts, the second average path lengths of the K reference eigenvalues in the target isolation forest can reflect the path length distribution of the normal eigenvalue of this trading feature in the target isolation forest corresponding to this trading feature.

[0174] In some embodiments, the minimum second average path length among the second average path lengths of the K reference eigenvalues in the target isolation forest can be used as the path length threshold.

[0175] In other embodiments, the sum of the minimum second average path length among the second average path lengths of the K reference eigenvalues in the target isolation forest and a preset length surplus can be used as the path length threshold.

[0176] As described above, in the isolation forest, a shorter path length is considered an abnormal eigenvalue, while a longer path length is considered a normal eigenvalue. Therefore, if the first average path length of the target eigenvalue of a trading feature in the corresponding target isolation forest is less than the path length threshold, it indicates that the probability of the target eigenvalue of this trading feature being abnormal is relatively high. Therefore, it can be determined that the current trading feature in the trading feature information of the target trading account is a risky trading feature, that is, it indicates that the target eigenvalue is an abnormal eigenvalue.

[0177] Conversely, if the target feature value of a transaction feature has a first average path length in the corresponding target isolation forest that is not less than the path length threshold, it indicates that the probability of the target feature value of this transaction feature being abnormal is low. Therefore, it can be determined that the current transaction feature in the transaction feature information of the target transaction account is not a risk transaction feature, that is, it indicates that the target feature value is a normal feature value.

[0178] In the above embodiments, for different transaction features, a target isolation forest suitable for anomaly detection of different transaction features is constructed. Moreover, in the process of anomaly detection for a transaction feature, the reference feature value of this transaction feature in the transaction feature information of the determined K normal transaction accounts (i.e., the K normal transaction accounts most similar to the transaction feature information of the current target transaction account) is used to specifically determine the path length threshold suitable for anomaly detection of the target feature value under the current single transaction feature, which can ensure the adaptability of the determined path length threshold to the target feature value under the current transaction feature and ensure the accuracy of the anomaly detection. This avoids problems such as inaccurate anomaly detection results caused by an unreasonable determined path length threshold, such as being too large or too small.

[0179] In some embodiments, before step 220, the anomaly classification model can be trained according to the Figure 8 process shown, which specifically includes:

[0180] Step 810, obtain the sample transaction feature information of multiple sample transaction accounts and the labeled recognition labels of the sample transaction accounts.

[0181] The recognition label of a sample transaction account is used to indicate whether the corresponding sample transaction account has risks. In some embodiments, the recognition label of the sample risk transaction account can be manually labeled, that is, by the reviewer according to the sample transaction feature information of the sample transaction account, the recognition label of the sample transaction account is manually labeled, and this labeled recognition label is used to indicate whether the corresponding sample transaction account is normal or abnormal. In other embodiments, abnormal feedback record data can be obtained. The abnormal feedback record data includes transaction abnormal feedback information submitted by participating users for transaction records. In other words, when a transaction participating user determines that a certain transaction is abnormal, the transaction abnormal feedback information is submitted. After that, the transaction accounts involved in the abnormal transaction in the abnormal feedback record data can be identified as sample transaction accounts with risks (abnormalities), and the transaction feature information of this sample transaction account in the time period where the target time is located (assuming it is the target time) can be obtained correspondingly.

[0182] Step 820: The anomaly classification model performs anomaly classification based on the sample transaction feature information of each sample transaction account to obtain the predicted category of each sample transaction account.

[0183] Step 830: Calculate the anomaly classification loss based on the predicted category of each sample transaction account and the labeled recognition label of the sample transaction account.

[0184] The anomaly classification loss is used to reflect the difference between the predicted category of the predicted output for the sample transaction account and the category indicated by the labeled recognition label of the sample transaction account. The anomaly classification loss can be determined by combining the predicted category of each sample transaction account and the labeled recognition label of the sample transaction account through a second loss function. The second loss function can be a mean squared error loss function, a cross-entropy loss function, an absolute value loss function, etc., and is not specifically limited here.

[0185] Step 840: Adjust the parameters of the anomaly classification model according to the anomaly classification loss until the second training end condition is reached.

[0186] The gradient of the anomaly classification loss with respect to the parameters of the risk recognition model can be calculated, and then the parameters of the anomaly classification model can be adjusted according to the gradient descent algorithm until the second training end condition is reached. The second training end condition can be that the number of iterations of the anomaly classification model reaches the second number threshold, or the anomaly classification loss converges, and is not specifically limited here.

[0187] Through the above training process, the anomaly classification model can learn the correlation between the transaction feature information of the transaction account and the category indicating whether it is abnormal. Then, after training, the anomaly classification model can accurately predict the category of the transaction account according to the input transaction feature information, realizing the accurate identification of abnormal transaction accounts.

[0188] Next, a specific embodiment is used to illustrate the method of the present application. In this embodiment, the anomaly classification model is a binary classification decision tree model based on the XGBoost (eXtreme Gradient Boosting) algorithm. This binary classification decision tree model has the characteristics of high performance (parallel processing, feature column storage, and cache optimization, etc., making the model training and prediction faster), high prediction accuracy (able to handle high-dimensional sparse data, and having strong generalization ability, and can obtain good prediction results on complex data sets), and interpretability (having feature importance analysis and model interpretation functions, which can help understand the prediction process and influencing factors of the model). Through the KNN (K-NearestNeighbors) model, the K trading feature information most similar to the trading feature information of the risky trading account is identified. Through the anomaly detection model based on the isolation forest algorithm, the abnormal trading features in the trading feature information of the risky trading account are identified; and through the text generation model formed by stacking and fine-tuning the network on the large language model, according to the abnormal trading features, the description text of the abnormal reason of the risky trading account is generated.

[0189] Figure 9 Exemplarily shows the training flowcharts of the binary classification decision tree model, the KNN model, the anomaly detection model, and the text generation model. As Figure 9 shown, a training set can be pre-constructed, and the training set includes the trading feature information of the trading accounts marked as normal and the trading feature information of the trading accounts marked as risky. In some embodiments, after obtaining the trading data of each trading account, data cleaning, feature extraction, and data standardization processing can be performed to obtain the trading feature information of each trading account, and the risk labels (normal or risky) of each trading account are marked.

[0190] In this embodiment, the trading feature information of the trading accounts marked as normal is added to the normal sample set, and the trading feature information of one trading account marked as normal in the normal sample set is called a normal sample; similarly, the trading feature information of the trading accounts marked as risky is added to the risk sample set, and the trading feature information of one trading account marked as risky in the risk sample set is called a risk sample.

[0191] The normal sample set and the risk sample set can be used to train the binary classification decision tree model, and the specific training process can refer to the process of Figure 9 In addition, before training the binary classification decision tree model, the hyperparameters of the binary classification decision tree model are correspondingly set, and the hyperparameters are, for example, the learning rate, the tree depth of the binary classification decision tree model, and the regularization parameter, etc.

[0192] After training the binary classification decision tree model for a period of time, a test data set can be used. The test data set is independent of the training set in the above text. The test data set also includes the transaction feature information of the transaction accounts marked as normal and the transaction feature information of the transaction accounts marked as risky. The trained binary classification decision tree model is evaluated for performance through the test data set. For example, the accuracy, precision, recall, etc. of the trained binary classification decision tree model are evaluated.

[0193] After that, according to the performance evaluation results of the trained binary classification decision tree model, if the performance evaluation results indicate that the requirements are not met, the hyperparameters of the binary classification decision tree model can be adjusted. Then, the binary classification decision tree model with adjusted hyperparameters is continuously trained to further improve the performance of the binary classification decision tree model.

[0194] In addition, based on the normal sample set and the risk sample set, a single risk sample can be taken out from the risk sample set each time. Assume that the taken-out risk sample is A. And through the KNN model, according to the risk sample A, K normal samples most similar to A are determined from the normal sample set. In some embodiments, since the risk samples and the normal samples include the feature values of multiple transaction features, and some transaction features may not be applicable to screening similar samples, therefore, according to the task requirements and domain knowledge, the feature values of some transaction features are selected from each risk sample and each normal sample for training the KNN model. During the training process, the similarity (such as Euclidean distance, Manhattan distance, etc.) between the risk sample A and each normal sample in the normal sample set can be calculated to evaluate the distance between the risk sample A and each normal sample in the normal sample set, and according to the similarity calculation results, K normal samples most similar to the risk sample A are determined. During the training process, the distance metric method can be selected to be changed (such as adjusting the Euclidean distance to calculate the Manhattan distance) to improve the performance of the KNN model.

[0195] After completing the training of the KNN model, the K normal samples most similar to each risk sample A can be determined from the normal sample set through the trained KNN model for training the anomaly detection model, that is, for constructing the target isolation forest applicable to each transaction feature. The process of constructing the target isolation forest applicable to anomaly detection of each transaction feature can be referred to the above description and will not be elaborated here.

[0196] After completing the training of the anomaly detection model, the trained anomaly detection model can be used to identify anomalies in the transaction features of each risk sample A based on the K most similar normal samples determined for each risk sample A, determine the abnormal transaction features in the risk sample A, and combine the K most similar normal samples and the risk sample A to determine the distribution of the abnormal transaction features, that is, the feature distribution it follows. In addition, it is also possible to manually write a sample abnormal reason description text for the transaction account to which the risk sample A belongs according to the eigenvalue of the abnormal transaction feature in the risk sample A, and use the abnormal transaction feature and its distribution in the risk sample A, as well as the sample abnormal reason description text for the abnormal transaction feature in the risk sample A to train the text generation model. During the process of training the text generation model, the parameters of the large language model are fixed, and the parameters in the fine-tuning network are fine-tuned to obtain the trained text generation model.

[0197] After completing the above training process, it can be generated according to the Figure 10 process shown. As Figure 10 shown, obtain the transaction feature information of multiple transaction accounts to be detected (each transaction account can be regarded as the target transaction account in this application), and through the binary classification decision tree model, identify the risks of each transaction account according to the transaction feature information of each transaction account to obtain the risk labels of each transaction account. On this basis, according to the risk labels of multiple transaction accounts detected in the same batch, the multiple transaction accounts detected in the same batch can be correspondingly divided into risk transaction accounts and normal transaction accounts. Assume that the set of transaction feature information of normal transaction accounts among the multiple transaction accounts detected in the same batch is the target transaction feature information set.

[0198] Based on the risk identification result, obtain the transaction feature information of a single risk transaction account, such as transaction feature information B, and through the trained KNN model, screen out the transaction feature information of the K normal transaction accounts most similar to B in the target transaction feature information set, and through the trained feature anomaly detection model, identify the abnormal transaction features in B according to the transaction feature information of the K normal transaction accounts most similar to B, and determine the distribution of the abnormal transaction features in B according to B and the transaction feature information of the K normal transaction accounts most similar to B.

[0199] Then, the text generation model generates a risk description generation text according to the abnormal transaction features and the distribution of the abnormal transaction features in B. Finally, the abnormal reason description text of the risk transaction account and the risk label of the risk transaction account can be output to the management side, so that the reviewers on the management side can review the risk transaction account.

[0200] In the related art, usually after identifying that a trading account has risks, all trading characteristics of the trading account are input into a text generation model, and the text generation model generates text according to all trading characteristics of the trading account to obtain a text describing the reason for the abnormality of the trading account. However, this method has a hallucination phenomenon, that is, there will be a phenomenon of misattribution and fabrication out of thin air, resulting in the text describing the reason for the abnormality not being able to accurately describe the reason for identifying the trading account as having risks.

[0201] In this application, after identifying that the target trading account has risks, among the trading characteristic information of multiple normal trading accounts, the K trading characteristic information most similar to the trading characteristic information of the target trading account is selected, and the selected K trading characteristic information and the trading characteristic information of the target trading account form a group, and the group is analyzed through the isolation forest algorithm to identify the abnormal trading characteristics in the trading characteristic information of the target trading account, and the identified abnormal trading characteristics are used to generate the text describing the reason for the abnormality of the target trading account, without using the normal trading characteristics in the trading characteristic information of the target trading account to generate the text describing the reason for the abnormality. In this way, the problem of model hallucination can be effectively alleviated, and the accuracy of the generated text describing the reason for the abnormality can be improved. Moreover, using the K trading characteristic information most similar to the trading characteristic information of the target trading account as a reference to identify the abnormal trading characteristics in the trading characteristic information of the target trading account can ensure the accuracy of the identification of abnormal trading characteristics, and thus ensure the accuracy of the subsequent text describing the reason for the abnormality generated according to the abnormal trading characteristics. In addition, in this application, by adding a fine-tuning network to the large language model and training the fine-tuning network, the generalization ability of the large language model in the task of generating text describing the reason for the abnormality can be improved.

[0202] The following introduces the device embodiments of this application, which can be used to execute the methods in the above embodiments of this application. For the details not disclosed in the device embodiments of this application, please refer to the above method embodiments of this application.

[0203] Figure 11 is a block diagram of a text generation device shown according to an embodiment of this application. As Figure 11 shown, the text generation device includes:

[0204] An acquisition module 1110, configured to acquire the trading characteristic information of the target trading account;

[0205] An anomaly recognition module 1120, configured to perform anomaly recognition according to the trading characteristic information of the target trading account to obtain an identification label of the target trading account;

[0206] A processing module 1130, configured to, if the identification tag indicates that the target trading account is abnormal, determine the K trading feature information that is most similar to the trading feature information of the target trading account from the trading feature information of multiple normal trading accounts; a normal trading account refers to a trading account whose identification tag indicates normal; K is an integer greater than 1;

[0207] An abnormal trading feature determination module 1140, configured to identify abnormal trading features in the trading feature information of the target trading account according to the K trading feature information, and determine the abnormal trading features in the trading feature information of the target trading account;

[0208] A text generation module 1150, configured to generate a description text of the abnormal reason of the target trading account according to the target feature value of the abnormal trading feature in the trading feature information of the target trading account.

[0209] In some embodiments, the trading feature information includes the feature values of multiple trading features; the abnormal trading feature determination module 1140 is configured to:

[0210] Perform the following processing for each trading feature:

[0211] Locate the position of the target feature value of the trading feature in the trading feature information of the target trading account in each binary tree of the target isolation forest, and determine multiple first path lengths of the target feature value in the target isolation forest; the target isolation forest refers to an isolation forest applicable to abnormal detection of trading features;

[0212] Locate the position of the reference feature value of the trading feature in the K trading feature information in each binary tree of the target isolation forest, and determine multiple second path lengths of each reference feature value in the target isolation forest;

[0213] Calculate the mean value of the multiple first path lengths of the target feature value in the target isolation forest to obtain the first average path length;

[0214] Calculate the mean value of the multiple second path lengths of each reference feature value in the target isolation forest to obtain the second average path length of each reference feature value in the target isolation forest;

[0215] Identify abnormal features of the trading features in the trading feature information of the target trading account according to the first average path length and the second average path length of each reference feature value in the target isolation forest.

[0216] In some embodiments, in the step of identifying abnormal features of the transaction features in the transaction feature information of the target trading account according to the first average path length and the second average path length of each reference feature value in the target isolation forest, the abnormal transaction feature determination module is further configured to: determine a path length threshold according to the second average path length of the K reference feature values in the target isolation forest;

[0217] If the first average path length is less than the path length threshold, determine that the transaction feature in the transaction feature information of the target trading account is an abnormal transaction feature.

[0218] In some embodiments, the text generation device further includes:

[0219] A first acquisition module, configured to acquire the first transaction feature information of the first sample abnormal trading account and the second transaction feature information of multiple first sample normal trading accounts that are most similar to the first transaction feature information;

[0220] A feature value set construction module, configured to construct a feature value set corresponding to each transaction feature according to the feature value of the same transaction feature in the first transaction feature information and the feature value in the second transaction feature information;

[0221] A binary tree construction module, configured to construct T binary trees corresponding to each transaction feature according to the feature value set corresponding to each transaction feature; T is an integer greater than 1;

[0222] An integration module, configured to integrate the T binary trees corresponding to the same transaction feature to obtain a target isolation forest applicable to detecting abnormalities of the corresponding transaction feature.

[0223] In other embodiments, the transaction feature information includes feature values of multiple transaction features; the abnormal transaction feature determination module 1140 is configured to:

[0224] Perform the following processing for each transaction feature:

[0225] Perform distribution fitting on the reference feature value of the transaction feature in the K transaction feature information and the target feature value of the transaction feature in the transaction feature information of the target trading account to determine the target feature distribution corresponding to the transaction feature;

[0226] Perform range delineation according to the target feature distribution corresponding to the transaction feature to determine the target normal feature value range corresponding to the target feature distribution;

[0227] If the target feature value exceeds the target normal feature value range, determine that the transaction feature in the transaction feature information of the target trading account is an abnormal transaction feature.

[0228] In some embodiments, the processing module 1130 includes:

[0229] A feature similarity calculation unit, configured to calculate the feature similarity between the transaction feature information of the target transaction account and the transaction feature information of each normal transaction account respectively;

[0230] A determination unit, configured to use the transaction feature information of the K normal transaction accounts with the largest feature similarity as the K transaction feature information.

[0231] In some embodiments, the text generation device further includes:

[0232] A distribution determination module, configured to perform distribution fitting based on the feature values of the risk transaction features in the K transaction feature information and the corresponding target feature values, to obtain the target feature distribution corresponding to the risk transaction features;

[0233] Correspondingly, the text generation module 1150 is configured as:

[0234] The text generation model generates a description text of the abnormal reason of the target transaction account according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account and the target feature distribution corresponding to the abnormal transaction feature.

[0235] In some embodiments, the text generation device further includes:

[0236] A second acquisition module, configured to acquire a plurality of training samples, where one training sample includes a sample abnormal transaction feature, a sample feature distribution corresponding to the sample abnormal transaction feature, and a sample description text of the sample abnormal reason;

[0237] A first text generation module, configured to generate a predicted description text of the abnormal reason by the text generation model according to the sample abnormal transaction feature and the sample feature distribution corresponding to the sample abnormal transaction feature;

[0238] A text generation loss calculation module, configured to calculate the text generation loss according to the sample description text of the abnormal reason and the predicted description text of the abnormal reason;

[0239] A first adjustment module, configured to adjust the parameters of the text generation model according to the text generation loss until a first training end condition is reached.

[0240] In some embodiments, the text generation model includes a large language model and a fine-tuning network; the first adjustment module is configured as:

[0241] Fix the parameters of the large language model, and adjust the parameters of the fine-tuning network according to the text generation loss until a first training end condition is reached.

[0242] In some embodiments, the anomaly recognition module 1120 is configured to: perform anomaly recognition on the basis of the transaction feature information of the target transaction account through an anomaly classification model, and obtain the recognition label of the target transaction account;

[0243] Correspondingly, the text generation device further includes:

[0244] A third acquisition module, configured to acquire the sample transaction feature information of multiple sample transaction accounts and the labeled recognition labels of the sample transaction accounts;

[0245] A risk classification module, configured to perform anomaly classification on the basis of the sample transaction feature information of each sample transaction account through the anomaly classification model, and obtain the predicted category of each sample transaction account;

[0246] An anomaly classification loss calculation module, configured to calculate the anomaly classification loss according to the predicted category of each sample transaction account and the labeled recognition label of the sample transaction account;

[0247] A second adjustment module, configured to adjust the parameters of the anomaly classification model according to the anomaly classification loss until a second training end condition is met.

[0248] Figure 12 FIG. shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that, Figure 12 The computer system 1200 of the electronic device shown is only an example, and should not impose any limitation on the functions and usage scope of the embodiments of the present application. The electronic device can be used to execute the text generation method provided by the application.

[0249] As Figure 12 shown, the computer system 1200 includes a central processing unit (CPU) 1201, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1202 or the program loaded from the storage section 1208 into the random access memory (RAM) 1203, such as executing the method in the above embodiments. In the RAM 1203, various programs and data required for system operation are also stored. The CPU 1201, the ROM 1202, and the RAM 1203 are connected to each other through a bus 1204. The input / output (I / O) interface 1205 is also connected to the bus 1204.

[0250] The following components are connected to the I / O interface 1205: an input section 1206 including a keyboard, a mouse, etc.; an output section 1207 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 1208 including a hard disk, etc.; and a communication section 1209 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 1209 performs communication processing via a network such as the Internet. A drive 1210 is also connected to the I / O interface 1205 as required. A removable medium 1211 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is mounted on the drive 1210 as required so that a computer program read therefrom is installed into the storage section 1208 as required.

[0251] Specifically, according to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including program code for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication section 1209, and / or installed from the removable medium 1211. When the computer program is executed by a central processing unit (CPU) 1201, various functions defined in the system of the present application are executed.

[0252] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0253] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0254] The units involved in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not constitute a limitation to the units themselves in certain cases.

[0255] As another aspect, the present application also provides a computer-readable storage medium, which can be included in the electronic device described in the above embodiments; or can exist alone without being assembled into the electronic device. The above computer-readable storage medium carries computer-readable instructions, and when the computer-readable storage instructions are executed by a processor, the method in any of the above embodiments is implemented.

[0256] In the embodiments of the present application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be an overall module or a part of the unit of the function of the module or unit.

[0257] According to one aspect of the embodiments of the present application, a computer program product is provided, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method in any of the above embodiments.

[0258] It should be noted that although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0259] From the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, and includes several instructions to enable a computing device (such as a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of the present application.

[0260] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.

[0261] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.

Claims

1. A text generation method, characterized in that, Including: Obtain the transaction feature information of the target transaction account; Perform anomaly recognition based on the transaction feature information of the target transaction account to obtain the recognition label of the target transaction account; If the recognition label indicates that the target transaction account is abnormal, among the transaction feature information of multiple normal transaction accounts, determine the K transaction feature information that is most similar to the transaction feature information of the target transaction account; the normal transaction account refers to a transaction account whose recognition label indicates normal; K is an integer greater than 1; Based on the K transaction feature information, perform abnormal transaction feature recognition on the transaction feature information of the target transaction account to determine the abnormal transaction features in the transaction feature information of the target transaction account; Generate text according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account to obtain the abnormal reason description text of the target transaction account.

2. The method according to claim 1, wherein The transaction feature information includes the feature values of multiple transaction features; the performing abnormal transaction feature recognition on the transaction feature information of the target transaction account based on the K transaction feature information to determine the abnormal transaction features in the transaction feature information of the target transaction account includes: Perform the following processing for each of the transaction features: Locate the position of the target feature value of the transaction feature in the transaction feature information of the target transaction account in each binary tree of the target isolation forest, and determine multiple first path lengths of the target feature value in the target isolation forest; the target isolation forest refers to an isolation forest applicable to anomaly detection of the transaction feature; Locate the position of the reference feature value of the transaction feature in the K transaction feature information in each binary tree of the target isolation forest, and determine multiple second path lengths of each reference feature value in the target isolation forest; Calculate the mean of the multiple first path lengths of the target feature value in the target isolation forest to obtain the first average path length; Calculate the mean of the multiple second path lengths of each reference feature value in the target isolation forest to obtain the second average path length of each reference feature value in the target isolation forest; Based on the first average path length and the second average path lengths of each reference feature value in the target isolation forest, perform abnormal feature recognition on the transaction feature in the transaction feature information of the target transaction account.

3. The method according to claim 2, characterized in that, The performing abnormal feature recognition on the transaction feature in the transaction feature information of the target transaction account based on the first average path length and the second average path lengths of each reference feature value in the target isolation forest includes: Determine the path length threshold according to the second average path lengths of the K reference feature values in the target isolation forest; If the first average path length is less than the path length threshold, determine that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature.

4. The method according to claim 2, wherein Before locating the position of the target feature value of the transaction feature in the transaction feature information of the target transaction account in each binary tree of the target isolation forest and determining multiple first path lengths of the target feature value in the target isolation forest, the method further includes: Obtain the first transaction feature information of the first sample abnormal transaction account and the second transaction feature information of multiple first sample normal transaction accounts that are most similar to the first transaction feature information; Construct a feature value set corresponding to each of the transaction features according to the feature value of the same transaction feature in the first transaction feature information and the feature value in the second transaction feature information; Construct T binary trees corresponding to each of the transaction features according to the feature value sets corresponding to each of the transaction features; T is an integer greater than 1; Integrate the T binary trees corresponding to the same transaction feature to obtain a target isolation forest suitable for abnormal detection of the corresponding transaction feature.

5. The method according to claim 1, wherein The transaction feature information includes feature values of multiple transaction features; the identifying abnormal transaction features in the transaction feature information of the target transaction account according to the K transaction feature information includes: Perform the following processing for each of the transaction features: Perform distribution fitting on the reference feature value of the transaction feature in the K transaction feature information and the target feature value of the transaction feature in the transaction feature information of the target transaction account to determine the target feature distribution corresponding to the transaction feature; Determine the target normal feature value range corresponding to the target feature distribution according to the range delineation of the target feature distribution corresponding to the transaction feature; If the target feature value exceeds the target normal feature value range, determine that the transaction feature in the transaction feature information of the target transaction account is an abnormal transaction feature.

6. The method according to claim 1, wherein The determining, in the transaction feature information of multiple normal transaction accounts, the K transaction feature information that is most similar to the transaction feature information of the target transaction account if the identification label indicates that the target transaction account is abnormal includes: Calculate the feature similarity between the transaction feature information of the target transaction account and the transaction feature information of each of the normal transaction accounts; Use the transaction feature information of the K normal transaction accounts with the largest feature similarity as the K transaction feature information.

7. The method according to claim 1, characterized in that Before generating a text of the abnormal reason description of the target transaction account according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account, the method further includes: Perform distribution fitting according to the feature value of the risk transaction feature in the K transaction feature information and the corresponding target feature value to obtain the target feature distribution corresponding to the risk transaction feature; The generating a text of the abnormal reason description of the target transaction account according to the target feature value of the abnormal transaction feature in the transaction feature information of the target transaction account includes: The text generation model generates a description text of the abnormal reason of the target trading account according to the target feature value of the abnormal trading feature in the trading feature information of the target trading account and the target feature distribution corresponding to the abnormal trading feature.

8. The method according to claim 7, characterized in that Before the text generation model generates the description text of the abnormal reason of the target trading account according to the target feature value of the abnormal trading feature in the trading feature information of the target trading account and the target feature distribution corresponding to the abnormal trading feature, the method further includes: Obtain a plurality of training samples, where one training sample includes a sample abnormal trading feature, a sample feature distribution corresponding to the sample abnormal trading feature, and a sample description text of the sample abnormal reason corresponding to the sample abnormal trading feature; The text generation model performs text generation according to the sample abnormal trading feature and the sample feature distribution corresponding to the sample abnormal trading feature to obtain a predicted description text of the abnormal reason; Calculate the text generation loss according to the sample description text of the abnormal reason and the predicted description text of the abnormal reason; Adjust the parameters of the text generation model according to the text generation loss until a first training end condition is reached.

9. The method according to claim 8, wherein The text generation model includes a large language model and a fine-tuning network; The adjusting the parameters of the text generation model according to the text generation loss until a first training end condition is reached includes: Fix the parameters of the large language model, and adjust the parameters of the fine-tuning network according to the text generation loss until a first training end condition is reached.

10. The method according to any one of claims 1 to 9, characterized in that, The performing abnormal identification according to the trading feature information of the target trading account to obtain the identification label of the target trading account includes: An abnormal classification model performs abnormal identification according to the trading feature information of the target trading account to obtain the identification label of the target trading account; Before the abnormal classification model performs abnormal identification according to the trading feature information of the target trading account to obtain the identification label of the target trading account, the method further includes: Obtain the sample trading feature information of a plurality of sample trading accounts and the labeled identification labels of the sample trading accounts; The abnormal classification model performs abnormal classification according to the sample trading feature information of each sample trading account to obtain the predicted category of each sample trading account; Calculate the abnormal classification loss according to the predicted category of each sample trading account and the labeled identification label of the sample trading account; Adjust the parameters of the abnormal classification model according to the abnormal classification loss until a second training end condition is reached.

11. A text generation device, characterized in that, including: An acquisition module, configured to acquire the trading feature information of the target trading account; An abnormal identification module, configured to perform abnormal identification according to the trading feature information of the target trading account to obtain the identification label of the target trading account; A processing module, configured to, if the identification label indicates that the target trading account is abnormal, determine K trading feature information that is most similar to the trading feature information of the target trading account from the trading feature information of multiple normal trading accounts; the normal trading account refers to a trading account whose identification label indicates normal; K is an integer greater than 1; An abnormal trading feature determination module, configured to identify abnormal trading features in the trading feature information of the target trading account according to the K trading feature information; A text generation module, configured to generate a description text of the abnormal reason of the target trading account according to the target feature value of the abnormal trading feature in the trading feature information of the target trading account.

12. An electronic device, characterized in that, Comprising: A processor; A memory, on which computer-readable instructions are stored, and when the computer-readable instructions are executed by the processor, the method described in any one of claims 1-10 is implemented.

13. A computer-readable storage medium having computer-readable instructions stored thereon, characterized in that, When the computer-readable instructions are executed by the processor, the method described in any one of claims 1-10 is implemented.

14. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method described in any one of claims 1-10 is implemented.