Network transaction risk assessment method and system for commodities

By constructing a distribution prediction model based on the self-attention layer and the LSTM layer, the trend prediction of the number of abnormal behaviors of commodities in online transactions is solved, and the problems of unclear definition of network transaction risk assessment, incomplete evaluation indicators and data processing in the existing technology are solved, and more efficient and accurate risk assessment is achieved.

CN120047211APending Publication Date: 2025-05-27SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510047653.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing technology has problems such as unclear definitions, incomplete evaluation indicators, difficulty in dealing with massive data and lack of system module integration in online transaction risk assessment, resulting in insufficient evaluation accuracy and stability.

Method used

A method based on the definition and evaluation index system of online transaction risk is proposed. Through the distribution prediction model constructed by the self-attention layer and the LSTM layer, the trend prediction number of commodity abnormal behavior is predicted, and the probability distribution is used to evaluate the probability of online transaction risk occurrence.

Benefits of technology

It improves the efficiency and accuracy of online transaction risk assessment, can effectively process massive data and provide trend forecasts and probability forecasts to meet the needs of regulators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047211A_ABST
    Figure CN120047211A_ABST
Patent Text Reader

Abstract

The invention provides a network transaction risk assessment method and system for commodities, and relates to the field of network transaction risk assessment, and the method comprises the steps: collecting the historical index data of a to-be-assessed commodity according to a network transaction risk rating index system, and processing the historical index data into a time sequence, the indexes comprising the number of abnormal behaviors of the commodity; inputting the time sequence into a trained distribution prediction model to obtain probability distribution of the number of abnormal behaviors of the commodity; on the basis of probability distribution, predicting the number of commodity abnormal behaviors in a period of time in the future; utilizing a cumulative distribution function constructed based on probability distribution to predict the probability that the commodity abnormal behavior quantity prediction value exceeds a threshold value to obtain a network transaction risk assessment result; on the basis of the network transaction risk definition and the network transaction risk evaluation index system, trend prediction is carried out on the number of abnormal behaviors of commodities, the probability of occurrence of the network transaction risk is predicted and evaluated by utilizing the probability, and the efficiency of network transaction risk evaluation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of network transaction risk assessment, and in particular to a network transaction risk assessment method and system for commodities. Background Art

[0002] With the joint development of e-commerce and logistics industries, online transactions have emerged to meet people's needs for more convenient consumption experience and more efficient consumption patterns. Through online transactions, e-commerce platforms provide consumers with more choices, allowing them to complete various consumption activities without leaving home, significantly reducing consumers' economic and time costs.

[0003] In order to enhance competitiveness and attract more consumers, different businesses have formulated different competition strategies and service models based on e-commerce platforms, which has promoted the healthy development of online transactions. However, it has also provided opportunities for abnormal behaviors and brought new challenges and risks to consumers. For example, the most representative abnormal behaviors such as counterfeiting, sales of prohibited or restricted goods, information asymmetry, contract anomalies and false reviews have seriously threatened the lives, health and property safety of consumers, and have gradually attracted widespread attention from all walks of life. Assessing the risks of online transactions has become a new industry demand.

[0004] Since the research on online transaction risks is still in its infancy, online transaction risk assessment mainly faces the following technical difficulties and problems:

[0005] 1. The definition of online transaction risk is not yet clear. Risk itself is uncertain, and determining the prediction object and evaluation indicators becomes a key factor. The definition of online transaction risk may vary depending on the business scenario. For example, for regulatory authorities to carry out necessary regulatory business, the definition of online transaction risk needs to comprehensively cover all aspects such as merchants, customers, platforms and transaction order, while for e-commerce platforms, it may only focus on a few aspects.

[0006] 2. There are many factors that trigger network transaction risks, and there is currently a lack of support for a comprehensive and reasonable evaluation index system. The factors that trigger network transaction risks are different, and the evaluation indicators will also be different. For example, abnormal behavior is the main factor that triggers network transaction risks, so the number of abnormal behaviors is the object of trend prediction, and the evaluation indicators need to be designed around abnormal behaviors. Other potential factors such as network facilities, logistics distribution, and hot events may also trigger network transaction risks, so the evaluation indicators need to be appropriately adjusted and supplemented.

[0007] 3. The huge user scale and wide variety of goods have given rise to a huge amount of online transaction records. Most traditional time series prediction models are unable to efficiently process massive data and are unable to provide trend predictions and probability predictions.

[0008] 4. Regulators do not need to understand the principles of the model when assessing online transaction risks, but rather hope to call the model through a user-friendly system interface. There is currently a lack of clear support for system module integration.

[0009] Therefore, the existing network transaction risk assessment technology has deficiencies in accuracy and stability. Summary of the invention

[0010] In order to solve the above problems, the present invention proposes a method and system for evaluating the risks of online transaction of commodities. Based on the definition of online transaction risks and the network transaction risk evaluation index system, the trend of the number of abnormal behaviors of commodities is predicted, and then the probability of occurrence of network transaction risks is evaluated by probability prediction, so as to improve the efficiency of network transaction risk assessment.

[0011] According to some embodiments, the present disclosure adopts the following technical solutions:

[0012] A method for assessing the risks of online transactions of commodities, comprising:

[0013] According to the network transaction risk rating index system, historical index data of the commodity to be evaluated is collected and processed into a time series, wherein the index of the index system includes the number of abnormal behaviors of the commodity;

[0014] Input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the product;

[0015] Based on probability distribution, predict the number of abnormal behaviors of commodities in the future period;

[0016] Using the cumulative distribution function constructed based on probability distribution, the probability that the predicted value of the number of abnormal behaviors of commodities exceeds the threshold is predicted to obtain the network transaction risk assessment result;

[0017] Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

[0018] According to some embodiments, the present disclosure adopts the following technical solutions:

[0019] A network transaction risk assessment system for commodities, comprising:

[0020] The sequence building module is configured to: collect historical indicator data of the commodity to be evaluated according to the network transaction risk rating indicator system, and process it into a time series, wherein the indicators of the indicator system include the number of abnormal behaviors of the commodity;

[0021] The distribution prediction module is configured to: input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the commodity;

[0022] The quantity prediction module is configured to: predict the quantity of abnormal behaviors of commodities in a future period of time based on probability distribution;

[0023] The risk assessment module is configured to: use a cumulative distribution function constructed based on probability distribution to predict the probability that the predicted value of the number of abnormal behaviors of commodities exceeds a threshold value, and obtain a network transaction risk assessment result;

[0024] Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

[0025] According to some embodiments, the present disclosure adopts the following technical solutions:

[0026] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, a network transaction risk assessment method for commodities is implemented.

[0027] According to some embodiments, the present disclosure adopts the following technical solutions:

[0028] An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device implements the network transaction risk assessment method for commodities.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention is based on the definition of network transaction risk and the network transaction risk evaluation index system, performs trend prediction on the number of abnormal behaviors of commodities, and then uses probability prediction to evaluate the probability of network transaction risk occurrence, thereby improving the efficiency of network transaction risk evaluation.

[0031] The present invention effectively utilizes the characteristics of the self-attention mechanism and deep autoregressive recurrent network, and designs a distribution prediction model based on Self-Attention DeepAR to perform trend prediction on the number of abnormal behaviors of commodities, thereby improving the model's ability to learn long-term dependencies. At the same time, it solves the problem that traditional time series prediction models are difficult to process data efficiently and provide trend prediction and probability prediction.

[0032] The present invention defines network transaction risks, and to a certain extent solves the problem that the definition of network transaction risks is still unclear. The present invention mainly focuses on the network transaction risks caused by abnormal network transaction behaviors, and provides an example for conducting network transaction risk supervision for specific business scenarios. The present invention takes the number of abnormal behaviors of commodities as the object of trend prediction, and at the same time requires the calculation of the possibility of network transaction risks, which conforms to the definition of network transaction risks, clarifies the core requirements, and guides the design of time series prediction and risk assessment models.

[0033] The present invention designs a relatively comprehensive and reasonable network transaction risk evaluation index system, which to a certain extent solves the current problem of lack of support for the network transaction risk evaluation index system; the network transaction risk evaluation index system is constructed in accordance with the principles of comprehensiveness, systematicness, scientific rationality, objectivity, practicality, safety and reliability, and divides and explains the evaluation indicators based on the four basic dimensions of merchants, platforms, commodities and environment, providing important guidance and reference basis for conducting network transaction risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings constituting a part of the present disclosure are used to provide a further understanding of the present disclosure. The illustrative embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation on the present disclosure.

[0035] Figure 1 This is a schematic diagram of the elements constituting network transaction risks in Example 1;

[0036] Figure 2 This is a diagram showing the framework structure of the network transaction risk assessment indicator system in Example 2;

[0037] Figure 3 This is a flow chart of the method of Example 3. DETAILED DESCRIPTION

[0038] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.

[0039] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present disclosure belongs.

[0040] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0041] The overall idea proposed in the present invention is as follows: defining network transaction risks for network transaction risk supervision scenarios and designing a network transaction risk evaluation index system; focusing on network transaction risks caused by abnormal network transaction behaviors, collecting required data based on the network transaction risk evaluation index system, designing covariates and performing standardization processing, extracting the number of abnormal commodity behaviors as the prediction object, constructing time series and processing based on scale factors; applying the network transaction risk assessment method based on Self-Attention DeepAR to predict the trend of the number of abnormal commodity behaviors and assess the possibility of network transaction risks; applying the network transaction risk intelligent monitoring and assessment system integration functional module, providing a model calling interface, and displaying network transaction risk trends and assessment results.

[0042] Example 1

[0043] This embodiment discloses a network transaction risk definition focusing on abnormal network transaction behavior, and its constituent elements are as follows: Figure 1 As shown, the components include: risk sources, potential events, consequences, and possibilities.

[0044] The definition of network transaction risk focusing on abnormal network transaction behavior is mainly to focus on abnormal network transaction behavior in the risk source dimension. The specific explanation is: in network transactions, due to the occurrence and development of abnormal network transaction behavior, the consequences and possibility of affecting the goals (including property, safety, reputation, network transaction order, etc.) of all parties involved in the network transaction (including customers, merchants, platforms, etc.) (usually refers to negative impact, such as damage and destruction).

[0045] Based on specific business needs, the definition of network transaction risk can be further focused on:

[0046] For the scenario where we are concerned about the number of abnormal behaviors of products, we select false reviews as the representative of abnormal behaviors of products and define the risk type of large-scale false review risk: in online transactions, when the number of false reviews exceeds a certain threshold, the consequences and possibility of causing damage to the goals (including property, safety, reputation, online transaction order, etc.) of all parties involved in the online transaction (including customers, merchants, platforms, etc.).

[0047] Example 2

[0048] This embodiment discloses a network transaction risk assessment index system, and its framework structure is as follows: Figure 2 As shown, the following indicators are included:

[0049] 1. Merchant abnormal situation, which is used to evaluate the status and trend of merchants' abnormal behavior. It has five secondary indicators, namely the number of abnormal merchants, the online sales rate of merchants' abnormal goods, the time trend of the number of abnormal merchants, the spatial trend of the number of abnormal merchants, and the degree of abnormal correlation between merchants;

[0050] 2. Business operating status: used to evaluate the business operating status of the business, with 7 secondary indicators, namely business qualifications, administrative licenses, penalty records, business hours, transaction activity, credit score, and after-sales service;

[0051] 3. Merchant information disclosure, which is used to evaluate the completeness of merchant information disclosure. It consists of four secondary indicators, namely, product attribute description, product price indication, product logistics information, and after-sales charge indication;

[0052] 4. Platform abnormal situation, which is used to evaluate the status and trend of abnormal behavior on the platform. It has three secondary indicators, namely, platform coverage rate of abnormal commodity behavior, platform abnormal commodity sales rate, and platform-platform abnormal correlation degree;

[0053] 5. Platform supervision system, which is used to evaluate the completeness of the platform supervision system construction. It has three secondary indicators, namely, merchant entry threshold, supervision system formulation, and supervision quality feedback;

[0054] 6. Product abnormality situation, which is used to evaluate the status and trend of abnormal behavior of products. It has five secondary indicators, namely the number of abnormal behaviors of products, the incidence rate of abnormal behaviors of products, the time trend of the number of abnormal behaviors of products, the spatial trend of the number of abnormal behaviors of products, and the degree of abnormal correlation between products.

[0055] 7. Product quality status, which is used to evaluate product quality. It has four secondary indicators, namely, quality inspection pass rate, return and exchange rate, customer rating, and complaint feedback;

[0056] 8. Potential environmental impact is used to evaluate potential environmental impact factors. It has four secondary indicators, namely network facility impact, logistics distribution impact, hot event impact, and relevant policy impact.

[0057] Furthermore, with respect to the large-scale false review risk of Example 1, the abnormal behavior-related index items in the network transaction risk assessment index system need to focus the abnormal behavior types on false reviews.

[0058] Example 3

[0059] This embodiment discloses a method for assessing the risk of online transactions of commodities, including:

[0060] Step S1: According to the network transaction risk rating index system, historical index data of the commodity to be evaluated is collected and processed into a time series, wherein the index of the index system includes the number of abnormal behaviors of the commodity;

[0061] Step S2: input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the commodity;

[0062] Step S3: Based on the probability distribution, predict the number of abnormal behaviors of commodities in the future period;

[0063] Step S4: using the cumulative distribution function constructed based on the probability distribution, predicting the probability that the predicted value of the number of abnormal behaviors of commodities exceeds the threshold, and obtaining the network transaction risk assessment result;

[0064] Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

[0065] As an embodiment, the present disclosure provides a method for assessing the risk of online transactions for commodities. Based on the definition of online transaction risk in embodiment 1 and the evaluation index system of online transaction risk in embodiment 2, the method predicts the trend of the number of abnormal behaviors of commodities, and then uses probability prediction to assess the probability of occurrence of online transaction risk, thereby improving the efficiency of online transaction risk assessment. Figure 3 The specific implementation process is as follows:

[0066] [Step (1)] Data processing: Based on data sources such as laws and regulations at all levels related to the e-commerce industry, e-commerce platform data interfaces, and relevant penalty documents, collect and process the required data to make it meet the model input requirements. The specific steps are as follows:

[0067] [Step (1-1)] Collect relevant data: According to the online transaction risk rating index system, collect index data, and extract the number of abnormal behaviors of commodities as the prediction object separately, and perform preprocessing steps such as supplementing missing values;

[0068] [Step (1-2)] Design covariates: In order to make the prediction of the number of abnormal behaviors of products more accurate, add time information, product inherent information, indicator information and other data as covariates, and perform standardization processing to assist prediction;

[0069] [Step (1-3)] Build a time series: Set the appropriate time window length and each sliding length, build a time series for the number of abnormal behaviors of products, and add the designed covariates;

[0070] [Step (1-4)] Calculate the scale factor: Calculate the scale factor v for the constructed time series i When data is input into the model, the scale factor is used to unify the scales of all time series to a similar range. When data is output from the model, the scale factor is used to restore the data to its original scale. The scale factor v i The calculation formula is as follows:

[0071]

[0072] Among them, i represents the serial number of the abnormal behavior quantity sequence of the commodity, t 0 Indicates the sequence number of the starting time point of the prediction sequence, z i,t Represents the observed value of the abnormal behavior quantity sequence i of the commodity at time step t.

[0073] Because we need to predict the time series of multiple commodities, and the time series of each commodity will be further divided into multiple small time series according to the length of the time window, there are multiple sequences here, and the sequence number i is used to distinguish different sequences.

[0074] Taking false product reviews as an example, in view of the large-scale false review risk, the number of false product reviews is used as the prediction object. At the same time, based on the network transaction risk rating index system, index data is collected. According to the time information such as year, month, week, day, inherent product information such as product type, and index information such as the incidence rate of false product reviews, covariates are designed and standardized. Based on the number of false product reviews, a time series is constructed and covariates are added to calculate the scale factor v. i Used to unify the time series scale.

[0075] [Step (2)] Model construction: Build the Self-Attention DeepAR model architecture.

[0076] The distribution prediction model uses the Self-Attention DeepAR model, which is a sequence-to-sequence (Seq2Seq) framework. From the perspective of the Seq2Seq framework, the Self-Attention DeepAR model uses an autoregressive recurrent network (DeepAR) for trend prediction and probability prediction, and is an encoder-decoder (Encoder-Decoder) structure. The known sequence is encoded as the input of the encoder, and the output is the hidden state of the model. The encoder output corresponding to the last time step of the known sequence is As the initialized hidden state of the decoder, at the first time step of the prediction sequence, the initialized hidden state, the observed value of the previous time step and the covariate of the current time step are used as input for forward propagation to obtain the updated hidden state, which is then parameterized to obtain the parameters of the probability distribution, thereby constructing the probability distribution.

[0077] Furthermore, the Self-Attention DeepAR model does not directly predict the value of the time series, but predicts the parameters of the probability distribution that the value of the time series conforms to. The model architecture includes input layer, self-attention layer, LSTM layer, linear layer and activation function layer.

[0078] The input layer is used to input time series data and perform some possible processing on the input data, such as embedding discrete data, that is, mapping high-dimensional discrete input to low-dimensional dense vector representation as the input of the neural network.

[0079] The self-attention layer is used to calculate the similarities between different time steps in the time series processed by the input layer, generate attention weights, and generate new time step vectors based on the attention weights, giving the model the ability to capture long-term dependencies in the time series.

[0080] The LSTM layer is used to learn long-term dependencies from the time series data processed by the self-attention layer, memorize the state of the data over time, and use it for further prediction. For the LSTM layer, it can be optimized by adjusting its number of layers, Dropout parameters, and other configuration parameters.

[0081] The linear layer is actually a fully connected layer, which maps the hidden state of the LSTM layer output to the pre-parameters of the probability distribution through a series of transformations. The reason why we obtain pre-parameters instead of parameters here is that the transformation of the linear layer cannot ensure that the output conforms to the reasonable range of the probability distribution parameters. For example, the standard deviation of the normal distribution cannot be negative.

[0082] The activation function layer is used to perform nonlinear transformation on the output of the linear layer. The activation function used is the Softplus function, which ensures that the output is non-negative and has a smooth derivative. The calculation formula of the Softplus function is as follows:

[0083] Softplus(x)=log(1+x)

[0084] Here, x represents the input of the Softplus function.

[0085] [Step (3)] Algorithm training: Design a loss function for training and optimize model parameters. Specifically:

[0086] (1) Design loss function:

[0087] The core of the Self-Attention DeepAR model lies in the design of the loss function, which is actually a variant of the likelihood function. This is different from the traditional time series prediction algorithm in concept. The fundamental difference is that the Self-Attention DeepAR model introduces relevant ideas of probability theory in order to achieve probabilistic prediction.

[0088] The significance of the loss function is to calculate the gap between the prediction results of the current model and the actual situation, so as to guide the optimization of the model parameters and make the prediction effect of the model more and more accurate. It is inappropriate to simply use the predicted parameters to construct a probability distribution, and then take random values ​​and actual values ​​on the probability distribution to calculate the error, because it has too much randomness and contingency, which can easily lead to instability in model training and unsatisfactory final results. Therefore, it is necessary to change the thinking to measure the gap between the prediction results and the actual situation.

[0089] Since the parameters of the predicted probability distribution have been obtained, and the actual value can also be obtained through the label, the probability of the actual value in the probability distribution constructed by the current predicted parameters can be calculated. The larger the probability, the higher the possibility of obtaining the true value under the current distribution, and the smaller the gap with the actual situation. Through multiple rounds of training, the loss calculated by the loss function should become smaller and smaller, so the gap between the predicted result and the actual value should also become smaller and smaller. It can be further deduced that the probability is getting larger and larger. It can be found that the change direction of loss and probability is opposite. If the change direction is consistent through a certain transformation, the probability can be used as a loss. The solution is to take a negative value for the probability, and the probability is calculated by the likelihood function, so the likelihood function can be transformed to serve as a loss function.

[0090] The likelihood function is a function of parameters, which indicates how likely the parameter value is to make the observed data appear under the condition of given observed data; the log-likelihood function is the logarithm of the likelihood function, that is, the function obtained by taking the natural logarithm of the likelihood function; using the log-likelihood function, multiplication can be converted into addition, thereby simplifying the calculation. When the likelihood function contains many multiplications, the log-likelihood function can avoid the problem of floating-point underflow, and the log-likelihood function is a convex function, which is easier to optimize in optimization problems; therefore, the likelihood function is first converted into a log-likelihood function, and then the inverse is taken as the model loss function. The calculation formula is as follows:

[0091]

[0092] Where N represents the number of abnormal behavior sequences of commodities, T represents the length of abnormal behavior sequence of commodities, function l represents the likelihood function, and z i,trepresents the observed value of the abnormal behavior quantity sequence i of the commodity at time step t, and θ represents the hidden state h of the LSTM output i,t The transformation function of , used to generate the parameters of the probability distribution.

[0093] Taking false product reviews as an example, for the large-scale false review risk, the number of false product reviews as the prediction object belongs to discrete non-negative counting data, and the probability distribution that meets the requirements is the negative binomial distribution. The negative binomial distribution is an important probability distribution in probability theory and statistics, and is defined as follows:

[0094] If each Bernoulli trial has two results, success or failure, and in each trial, the probability of success is p and the probability of failure is 1-p. Repeat the Bernoulli trial until the rth success occurs. At this time, the distribution of the number of trial failures X is the negative binomial distribution (Pascal distribution).

[0095] For example, in a box containing 6 red balls, 3 yellow balls and 1 white ball, we define touching a red ball as success and touching a non-red ball as failure. At this point, we conduct a series of ball-touching tests with replacement, and stop the test when we touch 5 red balls. At this time, the probability distribution of the number of times we touch a non-red ball is the negative binomial distribution. When r is an integer, the negative binomial distribution is also called the Pascal distribution, and its probability mass function formula is as follows:

[0096]

[0097] Here, k is the number of failures, r is the number of successes, and p is the probability of success of the event.

[0098] Negative binomial distribution and Poisson distribution are both discrete distributions and have certain similarities. However, the mean and variance of Poisson distribution are equal, while negative binomial distribution does not have this problem. Therefore, negative binomial distribution is more suitable for scenarios with high discreteness and variance greater than the mean.

[0099] Since the parameters of the predicted negative binomial distribution have been obtained, and the actual value can also be obtained through the label, the probability of the actual value in the negative binomial distribution constructed by the current predicted parameters can be calculated. The larger the probability, the higher the possibility of obtaining the true value under the current distribution, and the smaller the gap with the actual situation. Through multiple rounds of training, the loss calculated by the loss function should become smaller and smaller, so the gap between the predicted result and the actual value should also become smaller and smaller. It can be further deduced that the probability is getting larger and larger. It can be found that the change direction of loss and probability is opposite. If the change direction is consistent, the probability can be used as loss. The solution is to take a negative value for the probability, and the probability is calculated by the likelihood function, so the likelihood function can be transformed to serve as a loss function; therefore, first convert the likelihood function into a log-likelihood function, and then take the opposite number as the model loss function. The calculation formula is as follows:

[0100]

[0101] Where N is the number of false product review sequences, T is the length of the false product review sequence, function l is the likelihood function of the negative binomial distribution, and z is i,t represents the observed value of the number of false reviews of the product sequence i at time step t, and θ represents the hidden state h of the LSTM output i,t Transformation function for generating the parameters of the negative binomial distribution, including mean μ and dispersion parameter α.

[0102] The likelihood function for the negative binomial distribution is calculated as follows:

[0103]

[0104] Among them, x is the actual value, the mean μ and the discrete parameter α are the hidden state h output by the LSTM layer i,t It is obtained through transformation of the linear layer and the activation function layer. In this process, some parameters of the model need to be used, such as the matrix and bias for linear transformation. The calculation formula of the transformation process is as follows:

[0105]

[0106] in, and b μ It is used to process h i,t Get the matrix and bias of μ, and b α It is used to process h i,t Get the matrix and bias of α.

[0107] (2) Model training optimization:

[0108] Set the observation value of the abnormal behavior quantity sequence of the i-th commodity at time step t to zi,t , with t 0 As the number of the starting time point of the prediction sequence in the prediction phase time window, the goal of the model is to and the covariates x known at all time points i,1:T , for the prediction sequence Modeling to obtain conditional probability distribution Next, the conditional probability distribution is converted into a probability product, and then the likelihood function is used to calculate the probability. The formula is as follows:

[0109]

[0110] h i,t =h(h i,t-1 ,z i,t-1 ,x i,t ,Θ)

[0111] Among them, z i,t-1 represents the hidden state of the previous time step, x i,t represents the covariate input at the current time step, Θ represents the parameters of the model, and the function h represents the internal operation of LSTM. The hidden state of the current time step is obtained and used together with the model parameters Θ as the input of the function θ. The function θ is used to convert the hidden state into the parameters of the probability distribution, so that the likelihood function l(z i,t |θ(h i,t ,Θ)).

[0112] In the training phase, all data are known, including the prediction sequence, and can be directly used as input without the need for recursive prediction. The input to the network for the abnormal behavior quantity sequence of the i-th commodity at time step t includes the observation value z of the previous time step i,t-1 , the covariate x at the current time step i,t and the hidden state h at the previous time step i,t-1 In this embodiment, the network adopts the long short-term memory network LSTM, and variants such as recurrent neural network RNN ​​and GRU can also be selected according to actual requirements. The output of the network has two flows: one is used to transfer the hidden state of the current time step to the next time step to maintain memory, and the other is to obtain the parameters of the probability distribution through calculation, which are also the parameters of its likelihood function, and then use the loss function to maximize the likelihood function to optimize the model parameters.

[0113] [Step (4)] Model prediction and evaluation: Make predictions based on the trained Self-Attention DeepAR model and select appropriate indicators to evaluate the model effect;

[0114]

Step (4-1)

[0115] The encoder structure is the same as that of the training phase, but the decoder is somewhat different. In the prediction phase, the prediction sequence is unknown, so the observation value of the previous moment cannot be directly input as in the training phase; the prediction phase uses the Monte Carlo sampling method to generate the prediction value, and uses it as the input of the next time step for recursive prediction, and continuously iterates to obtain the prediction result of the prediction sequence.

[0116] The application of Monte Carlo sampling is reflected in the fact that after using the predicted parameters to construct a probability distribution, multiple sampling is performed on the distribution, and the median of the sampling results is used as the predicted value. The median is chosen here instead of the mean because the mean is more easily affected by extreme data; at the same time, the upper and lower bounds of the corresponding confidence interval of the sampling result are calculated according to the specified confidence interval parameters to achieve trend prediction.

[0117] The application of the cumulative distribution function is reflected in directly substituting the predicted parameters into the cumulative distribution function of the corresponding probability distribution, thereby calculating the probability that the predicted value exceeds a certain threshold, realizing probabilistic prediction, and meeting the requirements for possibility in the definition of network transaction risk.

[0118] In the prediction process, the scale factor needs to be applied for scale processing; first, the parameters predicted by the model need to be restored to the original scale according to the scale factor; second, when the predicted value of the current time step is used as the input of the next time step for recursive prediction, the predicted value needs to be unified to a similar scale according to the scale factor. The calculation formula is as follows:

[0119]

[0120] in, represents the predicted value of the abnormal behavior quantity sequence i of the commodity at time step t, v i Represents the scale factor.

[0121] In the prediction process, the scale factor needs to be applied for scale processing. The parameters predicted by the model need to be restored to the original scale according to the scale factor. The calculation formula is as follows:

[0122] μ=v i log(1+exp(o μ ))

[0123]

[0124] Where μ represents the mean of the negative binomial distribution, α represents the dispersion parameter of the negative binomial distribution, and v i The scale factor representing the number of abnormal behaviors of commodities i, o μ and α Represents the predicted values ​​of the mean and dispersion parameters of the model output.

[0125] [Step (4-2)] Model evaluation: Model evaluation indicators include MAE, MAPE, MSE, RMSE, ND, NRMSE, and ρ-risk.

[0126] MAE (Mean Absolute Error) represents the mean absolute error, and the calculation formula is as follows:

[0127]

[0128] MAPE (Mean Absolute Percentage Error) represents the mean absolute percentage error, and the calculation formula is as follows:

[0129]

[0130] MSE (Mean Squared Error) represents the mean square error, and the calculation formula is as follows:

[0131]

[0132] RMSE (Root Mean Square Error) represents the root mean square error, and the calculation formula is as follows:

[0133]

[0134] ND (Normalized Deviation) represents the standardized deviation, and the calculation formula is as follows:

[0135]

[0136] NRMSE (Normalized Root Mean Square Error) represents the standardized root mean square error, and the calculation formula is as follows:

[0137]

[0138] Among them, z i,t represents the true value of the abnormal behavior quantity sequence i of the commodity at time step t, represents the predicted value of the product abnormal behavior quantity sequence i at time step t. N represents the number of product abnormal behavior quantity sequences, Tt 0 Indicates the number of time steps for the prediction.

[0139] Use (L,S) to represent the time step t of the division between the known sequence and the predicted sequence 0 In the subsequent time span [L, L+S], the aggregate target value of the abnormal behavior quantity sequence i of the product in this span is expressed as For the quantile ρ∈(0,1), we predict Z i The ρ-quantile of (L,S) is expressed as In order to obtain the predicted quantile from multiple paths of the sampling results, it is necessary to first calculate the sum of each sample path within the specified span, and the sum of these sample paths is used as Z i The distribution estimation of (L,S) can obtain the corresponding quantile from the empirical distribution. The corresponding quantile loss calculation formula is as follows:

[0140]

[0141] ρ-risk represents the normalized quantile loss and is calculated as follows:

[0142]

[0143] Example 4

[0144] In one embodiment of the present disclosure, a network transaction risk assessment system for commodities is provided, comprising:

[0145] The sequence building module is configured to: collect historical indicator data of the commodity to be evaluated according to the network transaction risk rating indicator system, and process it into a time series, wherein the indicators of the indicator system include the number of abnormal behaviors of the commodity;

[0146] The distribution prediction module is configured to: input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the commodity;

[0147] The quantity prediction module is configured to: predict the quantity of abnormal behaviors of commodities in a future period of time based on probability distribution;

[0148] The risk assessment module is configured to: use a cumulative distribution function constructed based on probability distribution to predict the probability that the predicted value of the number of abnormal behaviors of commodities exceeds a threshold value, and obtain a network transaction risk assessment result;

[0149] Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

[0150] Example 5

[0151] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, the method for assessing the risks of online transactions of commodities is implemented.

[0152] Example 6

[0153] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes and implements the network transaction risk assessment method for commodities.

[0154] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0156] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A method for assessing the risk of online transactions of commodities, characterized in that: include: According to the network transaction risk rating index system, historical index data of the commodity to be evaluated is collected and processed into a time series, wherein the index of the index system includes the number of abnormal behaviors of the commodity; Input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the product; Based on probability distribution, predict the number of abnormal behaviors of commodities in the future period; Using the cumulative distribution function constructed based on probability distribution, the probability that the predicted value of the number of abnormal behaviors of commodities exceeds the threshold is predicted to obtain the risk assessment result of network transactions; Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

2. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: The network transaction risk rating index system is a two-level index. The first-level index includes merchant abnormal situation, merchant operating conditions, merchant information disclosure, platform abnormal situation, platform supervision system, commodity abnormal situation, commodity quality status, and potential environmental impact; Among them, the abnormal situation of goods is used to evaluate the status and trend of abnormal behavior of goods, and it has 5 secondary indicators, namely the number of abnormal behaviors of goods, the incidence rate of abnormal behaviors of goods, the time change trend of the number of abnormal behaviors of goods, the spatial change trend of the number of abnormal behaviors of goods, and the degree of correlation between product-product abnormalities.

3. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: The time series is composed of a plurality of historical time step vectors arranged in chronological order. The time step vector is the number of abnormal behaviors of commodities in the current time step, with time information, commodity inherent information and other indicator information added as covariates.

4. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: The distribution prediction model is constructed based on Self-Attention DeepAR, including input layer, self-attention layer, LSTM layer, linear layer and activation function layer; The linear layer is used to map the hidden state output by the LSTM layer into pre-parameters of the probability distribution; The activation function layer is used to perform nonlinear transformation on the pre-parameters output by the linear layer to obtain parameters that meet the requirements of the probability distribution.

5. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: The probability distribution adopts negative binomial distribution, and the parameters of the probability distribution include mean and dispersion parameter; Multiple Monte Carlo sampling is performed on the negative binomial distribution, and the median of the sampling results is used as the predicted value of the number of abnormal behaviors of commodities in the future time step.

6. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: The threshold is set in the network transaction risk definition. When the probability that the number of abnormal behaviors of commodities exceeds the threshold meets the conditions preset in the network transaction risk definition, it is considered that a network transaction risk occurs.

7. A method for assessing the risk of online transactions of commodities as claimed in claim 1, characterized in that: Also includes: After constructing the time series, the scale factor is calculated for the constructed time series, and the scales of all time series are unified through the scale factor. When the predicted value of the number of abnormal behaviors of commodities is obtained, the predicted value is restored to the original scale through the scale factor.

8. A network transaction risk assessment system for commodities, characterized in that: include: The sequence building module is configured to: collect historical indicator data of the commodity to be evaluated according to the network transaction risk rating indicator system, and process it into a time series, wherein the indicators of the indicator system include the number of abnormal behaviors of the commodity; The distribution prediction module is configured to: input the time series into the trained distribution prediction model to obtain the probability distribution of the number of abnormal behaviors of the commodity; The quantity prediction module is configured to: predict the quantity of abnormal behaviors of commodities in a future period of time based on probability distribution; The risk assessment module is configured to: use a cumulative distribution function constructed based on probability distribution to predict the probability that the predicted value of the number of abnormal behaviors of commodities exceeds a threshold value, and obtain a network transaction risk assessment result; Among them, the distribution prediction model first uses the self-attention layer to calculate the similarity between different time steps in the input time series, generates attention weights, and generates new time step vectors based on the attention weights. Then, the LSTM layer is used to learn long-term dependencies from the time series composed of new time step vectors, memorize the state of data changing over time, and finally, based on the hidden state output by the LSTM layer, the parameters of the probability distribution are generated.

9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, a method for assessing the risk of online transactions for commodities as described in any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a network transaction risk assessment method for commodities as described in any one of claims 1 to 7.