Data dynamic transaction value evaluation method
Through a hybrid model of graph attention neural network and long and short-term memory network, combined with sparse matrix acceleration and Monte Carlo simulation, the problem of capturing network relationships and dynamic characteristics in data transaction value assessment is solved, and efficient and accurate data transaction value assessment and risk monitoring are achieved.
Patent Information
- Application Number
- CN202510740697.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-07-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to capture the network relationship between data and dynamic characteristics that change over time in the evaluation of data transaction value, resulting in low accuracy of evaluation results and lack of comprehensive consideration of risk factors, making it difficult to adapt to a rapidly changing environment.
A hybrid model of graph attention neural network and long and short-term memory network is used, combining sparse matrix acceleration and Monte Carlo simulation, the intrinsic value, network value, time value and dynamic risk of data is calculated, and the evaluation results are pushed through WebSocket and risk is monitored in real time.
It improves the accuracy of data transaction value evaluation, reduces calculation time and memory usage, improves the accuracy and response speed of risk monitoring, and is suitable for high-frequency trading scenarios such as finance and e-commerce.
Smart Images

Figure CN120258638A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method for evaluating the dynamic trading value of data. Background Art
[0002] The evaluation of data trading value is a core technology in fields such as finance, e-commerce, and healthcare, aiming to quantify the economic value of data to support trading decisions. In the prior art, data value evaluation usually adopts methods based on static rules or single dimensions. For example, the accuracy or integrity of data is calculated through statistical analysis, or linear regression prediction is performed based on historical transaction prices. These methods often use simple machine learning models (such as decision trees or support vector machines) or fixed weight formulas, combined with a database management system (such as MySQL) and a conventional computing framework (such as Apache Spark), to achieve a certain degree of automated evaluation on medium and small-scale data sets.
[0003] However, the prior art has significant deficiencies in dealing with complex data associations and dynamic time-series characteristics. Static rules or single-dimensional methods are difficult to capture the network relationships between data (such as node dependencies in a knowledge graph) and the dynamic characteristics that change over time (such as the attenuation of data timeliness), resulting in relatively low accuracy of evaluation results. For example, in the evaluation of financial transaction data, the mean absolute error (MAE) of a linear regression-based model is usually around 0.15, which is difficult to meet the requirements of high-dynamic scenarios. In addition, the existing methods lack comprehensive consideration of risk factors and are difficult to adapt to the rapidly changing environment in data trading, thus limiting the reliability and real-time nature of the evaluation results. To address the above problems, an evaluation method that can comprehensively consider data correlation, time-series characteristics, and dynamic risks is needed to improve the accuracy of data trading value evaluation. Summary of the Invention
[0004] Aiming at the deficiencies of the prior art, the present invention provides a method for evaluating the dynamic trading value of data, which solves the problem that the existing static rules or single-dimensional methods are difficult to capture the network relationships between data (such as node dependencies in a knowledge graph) and the dynamic characteristics that change over time (such as the attenuation of data timeliness), resulting in relatively low accuracy of evaluation results.
[0005] To achieve the above objectives, the present invention is realized through the following technical solutions: A method for evaluating the dynamic trading value of data, applied to a server, includes: S1. Receiving a data set from a data source, where the data set includes data attributes, relationships between data, and historical operation records; S2. Based on the received data set, using a hybrid model of a graph attention neural network and a long short-term memory network to calculate the total trading value at time t under the operation of a user , the total transaction value is calculated by the following formula:
[0006] where represents the data under the operation of the user , the total value at time t, is the intrinsic value of the data , is the value of the data in the network environment, is the time value of the data , is the data under the operation of the user at time t, the dynamic risk faced; S3. Output the total transaction value to the user terminal, and perform real-time risk monitoring according to the dynamic risk. When the risk exceeds the preset threshold, send a warning message to the user terminal.
[0007] Through the above technical solution, in step S1, the server receives a JSON format data set from a data source (such as cloud storage or enterprise database) through an API, including data attributes (such as type, source), association relationships (such as knowledge graph edges), and historical operation records (such as access logs). Data preprocessing uses the Pandas library for deduplication and format standardization, taking about milliseconds. Step S2 calculates the total transaction value through a GAT-LSTM hybrid model. The model is accelerated based on GPU (NVIDIA CUDA), trained for 1000 rounds, and the MAE is lower than 0.087, suitable for financial and e-commerce scenarios. Step S3 pushes the total value to the user terminal (such as APP) through WebSocket, and uses a rule engine to monitor risks in real time. When the risk exceeds the threshold (0.7), send a warning through SMS or APP, and the response time is less than 100 milliseconds.
[0008] Preferably, the calculation of the intrinsic value includes: S101. Obtain multiple quality dimension scores of the data , and the quality dimensions at least include accuracy, integrity, and timeliness; S102. Calculate the intrinsic value based on the following formula:
[0009] Q where Q m represents the score of the m th quality dimension, and the value range is [0,1], after being standardized; is the dynamic weight coefficient, satisfying .
[0010] Preferably, the calculation of the quality dimension includes: S103. Calculate the accuracy, and determine the deviation ratio between the data and the true value through sampling inspection; S104. Calculate the integrity, and determine the proportion of non-missing data; S105. Calculate the timeliness, using introduce the influence of the logarithmic function to smooth the update frequency or through the formula , is the number of days the data exists.
[0011] Preferably, the calculation of the network value includes: S201. Based on the graph attention network, calculate the network value of the data , and the formula is:
[0012] where represents the set of direct associated nodes of the data , is the associated edge weight, and the value range is [0, 1], and after being normalized, is the Sigmoid activation function, used to strengthen the key connections, is the total transaction value of the data .
[0013] Preferably, the network value calculation uses a sparse matrix to accelerate, specifically including: S202. Calculate the network value based on the following formula:
[0014] where A is the adjacency matrix, , or is the dynamic decay factor, where is the contribution degree of the network value to the total value at time t = 0, and the value range is [0, 1], is the decay rate of the network value over time, is the total transaction value at the previous moment of time t.
[0015] Preferably, the calculation of the time value includes: S203. Calculate the time value based on the following formula:
[0016] Among them, is the maximum effective life cycle of the data.
[0017] Preferably, the calculation of the dynamic risk includes: S401. Based on the risk propagation model, calculate the dynamic risk, and the formula is:
[0018] Among them, represents the occurrence probability of the r-th type of risk in the range [0, 1], which is statistically obtained from historical events. represents the corresponding risk loss amount. represents the risk time distribution model, where represents the risk probability density value corresponding to time t, that is, the relative possibility of the risk occurring near time t. t represents the time variable, which can be absolute time or relative time. represents the average time point of risk occurrence, that is, the central position of the risk distribution. represents the dispersion degree of the risk distribution, reflecting the uncertainty of the risk time. η ∈ [0, 1] represents the risk propagation attenuation coefficient, and λ ij ∈ [0, 1] represents the propagation coefficient = the proportion of shared sensitive fields × the access frequency. represents the data of the direct associated node set.
[0019] Preferably, the dynamic risk is optimized by the Monte Carlo simulation algorithm, which specifically includes: S3011. Initialize the risk seed vector; S3012. Iteratively propagate the risk based on the adjacency matrix and perform a predetermined number of simulations; S3013. Average the risk matrices to estimate the propagated risk.
[0020] Preferably, the construction of the hybrid model includes: S204. Use the graph attention network to capture the correlation between data and calculate the network value; S205. Use the long short-term memory network to encode the temporal features and calculate the time value; S206. Through the fusion layer, combine the outputs of the graph attention network and the long short-term memory network to calculate the total transaction value.
[0021] Preferably, a data dynamic transaction value evaluation system includes: A data receiving module, used to execute step S1 to receive a data set from a data source; A value calculation module, used to execute step S2 to calculate the total transaction value; A risk monitoring module for performing step S3 to conduct real-time risk monitoring and send warning messages. An output module for performing step S3 to send the total transaction value and warning messages to the user terminal.
[0022] The present invention provides a method for evaluating the dynamic trading value of data, having the following beneficial effects: 1. By adopting a hybrid model of graph attention neural network (GAT) and long short-term memory network (LSTM), the present invention comprehensively evaluates the intrinsic value, network value, time value, and dynamic risks of data. Compared with the existing methods based on static rules, it has higher accuracy. In the evaluation of financial trading data, the mean absolute error (MAE) is reduced to 0.087, which is about 42% higher than that of the traditional regression model (MAE is about 0.15), supporting real-time evaluation in a dynamic data environment.
[0023] 2. The present invention accelerates the calculation of network value through a sparse matrix, stores the adjacency matrix in CSR format, reducing the memory occupancy by about 70%. In a financial trading network with millions of nodes, the calculation time is shortened from 1.2 seconds to 0.3 seconds, with the speed increased by 8.7 times. It is more efficient than the dense matrix method and is applicable to high-frequency trading scenarios.
[0024] 3. The present invention realizes dynamic risk monitoring through a risk propagation model and Monte Carlo simulation, with a warning accuracy rate of 95% and a response time of less than 50 milliseconds. Compared with the static threshold method, it can identify risks more accurately and is applicable to high-security scenarios such as bank data transactions.
[0025] 4. The present invention evaluates the intrinsic value of data from multiple dimensions such as accuracy, integrity, and timeliness, and improves the reliability through dynamic weight weighting. In the e-commerce scenario, the accuracy weight can be adjusted to 0.5, and the evaluation time is about 3 seconds, which is more comprehensive than the single-dimensional method. Description of the Drawings
[0026] Figure 1 It is a flowchart of the method for evaluating the dynamic trading value of data according to the present invention; Figure 2 It is a schematic diagram of the streaming data processing pipeline architecture according to the present invention. Detailed Embodiments
[0027] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0028] Please refer to the attached Figure 1 - attached Figure 2, an embodiment of the present invention provides a method for evaluating the dynamic trading value of data, which is applied to a server and includes the following steps: S1. Receive a data set from a data source, where the data set includes data attributes, relationships between data, and historical operation records; S2. Based on the received data set, use a hybrid model of a graph attention neural network and a long short-term memory network to calculate the total trading value at time t under the operation of the user, and the total trading value is calculated by the following formula:
[0029] where, represents the total value of the data under the operation of the user at time t, is the intrinsic value of the data , is the value of the data in the network environment, is the time value of the data , is the data under the operation of the user, and the dynamic risk faced at time t; S3. Output the total trading value to the user terminal, and perform real-time risk monitoring according to the dynamic risk. When the risk exceeds the preset threshold, send a warning message to the user terminal.
[0030] Specifically, the method for evaluating the dynamic trading value of data is deployed on a high-performance server cluster. The server adopts a distributed architecture and supports multi-threaded parallel computing to meet the requirements of large-scale data processing. In step S1, the server receives the data set from the data source (such as an enterprise database, cloud storage, or third-party data platform) through the API gateway. The data set is transmitted in JSON or Parquet format and includes data attributes (such as data type, source, generation time), relationships between data (such as the edge and node relationships in a knowledge graph), and historical operation records (such as access logs, modification records). To ensure data integrity, the server uses a data verification module to preprocess the received data, including deduplication, missing value filling, and format standardization. The preprocessing process is implemented using the Pandas library, and the time consumption is controlled within milliseconds.
[0031] In step S2, a hybrid model of a graph attention neural network (GAT) and a long short-term memory network (LSTM) is used to achieve efficient computing through GPU acceleration (such as NVIDIA CUDA). The model training is based on a historical dataset, using the Adam optimizer with a learning rate of 0.001 and 1000 training epochs. The validation set error (MAE) is lower than 0.087. During the calculation process, the server dynamically allocates memory, preferentially processes high-value data nodes, and avoids resource waste. The calculation of the total transaction value takes into account the comprehensive effects of the intrinsic value of the data, network value, time value, and dynamic risks, and is applicable to scenarios such as finance, e-commerce, and healthcare. For example, in the financial scenario, this method can evaluate the real-time value of transaction data and assist in quantitative investment decisions.
[0032] In step S3, the server pushes the total transaction value to the user terminal (such as a mobile APP or a web interface) through the WebSocket protocol, supporting real-time visual display (such as line charts and heat maps). The risk monitoring module analyzes dynamic risks in real time based on a rule engine and an anomaly detection algorithm (such as IsolationForest). When the risk value exceeds a threshold (such as 0.7), an early warning mechanism is triggered, and the user is notified via SMS, email, or the APP. The warning information includes the risk type, scope of influence, and recommended measures to ensure that users can take corresponding strategies in a timely manner. The response time of the entire process is controlled within 100 milliseconds to meet the requirements of high-concurrency scenarios.
[0033] Intrinsic value The calculation includes: S101. Obtain the quality dimension scores of multiple data, and the quality dimensions at least include accuracy, integrity, and timeliness; S102. Calculate the intrinsic value based on the following formula:
[0034] where Q m represents the score of the m th quality dimension, with a value range of [0, 1] and after being normalized; is a dynamic weight coefficient, satisfying .
[0035] Specifically, the calculation of the intrinsic value is implemented in the feature extraction module of the server. In step S101, multiple quality dimension scores are obtained through the data quality assessment framework. The server first extracts metadata from the dataset (such as field completeness rate, update frequency), and then calls the quality assessment algorithm library (such as TensorFlow Data Validation) to quantify accuracy, integrity, and timeliness. The accuracy assessment is based on random sampling with a sampling ratio of 5%, and the deviation is calculated by comparing with the standard dataset; the integrity assessment is performed by scanning the dataset fields and counting the non-null value ratio; the timeliness assessment combines the data generation timestamp and uses a sliding time window (default 24 hours) to analyze the update frequency.
[0036] In step S102, the dynamic weight coefficient is generated through XGBoost model training. The training data includes historical quality dimension scores and user feedback. The weight update period is 1 hour to ensure adaptation to data changes. The server uses Redis cache to store real-time weights to reduce query latency. The intrinsic value calculation result is stored in floating-point form with 4 decimal places reserved for subsequent total value calculation. This embodiment performs excellently in the e-commerce scenario. For example, when evaluating the intrinsic value of product review data, the accuracy weight is relatively high (about 0.5), significantly improving the reliability of value evaluation.
[0037] The calculation of quality dimensions includes: S103. Calculate the accuracy by determining the deviation ratio between the data and the true value through sampling inspection; S104. Calculate the integrity by determining the proportion of non-missing data; S105. Calculate the timeliness by using introducing a logarithmic function to smooth the impact of the update frequency or through the formula , where
[0038] is the number of days the data exists.
[0039] The timeliness calculation of step S105 is based on timestamp analysis. The server uses a time series database (such as InfluxDB) to store data generation and update records. The logarithmic smoothing function is implemented through NumPy, and the parameter λ is dynamically adjusted according to the data type (for example, λ = 0.05 for financial data and λ = 0.1 for social data). When applied in the evaluation of medical data in this embodiment, expired diagnostic data is identified through timeliness analysis, significantly reducing the risk of misdiagnosis.
[0040] Value in the network environment The calculation includes: S201. Calculate the network value of the data based on the graph attention network, and the formula is: The formula is:
[0041] Among them, represents the set of direct associated nodes of the data of, is the weight of the associated edge, and the value range is [0, 1], and after normalization processing, is the Sigmoid activation function, which is used to strengthen the key connection.
[0042] Specifically, the calculation of the network value is executed in the graph computing engine of the server. Step S201 stores the association relationship between data based on the Neo4j graph database. The nodes represent data entities (such as users, transaction records), and the edges represent relationships (such as references, dependencies). The graph attention network (GAT) model is implemented through PyTorchGeometric. During training, Dropout (ratio 0.2) is used to prevent overfitting. The server supports dynamic graph updates. When new nodes or edges are added, the network value is adjusted through incremental calculation, and the update frequency is once a minute.
[0043] To improve the calculation efficiency, the server uses the distributed graph partitioning technology to split the large-scale graph into subgraphs and allocate them to multiple nodes for processing. Each node is equipped with 32GB of memory and 8-core CPUs. In the evaluation of social network data in this embodiment, high-influence user nodes (such as KOLs) are successfully identified, and the network value weight is close to 0.9, significantly improving the accuracy of advertising placement.
[0044] The network value calculation uses sparse matrix acceleration, specifically including: S202. Calculate the network value based on the following formula:
[0045] Among them, A is the adjacency matrix, , or is the dynamic decay factor, among which, is the contribution degree of the network value to the total value at time t = 0, and its value range is [0, 1]. is the decay rate of the network value over time. is the total transaction value at the previous moment of time t. When = 0, the network value makes no contribution to the total value.
[0046] When = 1, the contribution of the network value to the total value is the largest at the initial moment.
[0047] The average association strength (such as the mean weight of edges) between data nodes can be calculated based on data statistics and normalized to the interval [0, 1] as the initial value.
[0048] For example, if the mean weight of edges is 0.8, then can be set to 0.8. The physical meaning of μ is that μ represents the decay rate of the network value over time: When μ = 0, the network value does not decay over time ( γ ( t ) = ).
[0049] When μ > 0, the network value decays exponentially over time, μ the larger, the faster the decay.
[0050] The method of obtaining the value of μ: (1) Based on the data life cycle If the data life cycle is short (such as news data, social media data), a larger μ can be set. For example μ ∈[0.1, 0.5].
[0051] If the data life cycle is long (such as historical data, knowledge graph data), a smaller μ can be set. For example μ ∈[0.01, 0.1].
[0052] (2) Based on business requirements If the business pays more attention to the value of recent data (such as real-time recommendation systems), a larger μ can be set to quickly decay the value of old data.
[0053] If the business needs to retain the value of data for a long time (such as historical data analysis), a smaller μ can be set to slow down the decay rate.
[0054] (3) Based on statistical characteristics Analyze the variation law of data value over time, fit an exponential decay curve, and obtain μ the estimated value of
[0055] For example, if the data value decays to half of the initial value within 10 days, then it can be calculated through the formula to obtain μ ≈0.069.
[0056] (4)Based on model optimization Take μ as the trainable parameter and optimize its value through a machine learning model (such as the gradient descent method).
[0057] During the training process, μ will be automatically adjusted according to the loss function to achieve the best evaluation effect.
[0058] Suppose in a news dataset, the data value decays to half of the initial value within 7 days, then μ can be calculated through the following formula: Solve to get:
[0059] If the business pays more attention to recent data, μ can be appropriately increased, for example μ =0.15.
[0060] Sparse matrix acceleration is implemented through the SciPy library in step S202. The adjacency matrix A is stored in the CSR (Compressed Sparse Row) format, only recording non-zero elements, reducing memory occupancy by about 70%. The node feature matrix is generated through feature engineering, including features such as data popularity and access frequency, with a dimension of 128. The server uses multi-threaded matrix operations, and the single calculation time is controlled within 50 milliseconds.
[0061] In the evaluation of financial transaction data, sparse matrix acceleration reduces the network value calculation time from 1.2 seconds to 0.3 seconds, and the inference speed is increased by 8.7 times, which is applicable to high-frequency trading scenarios. The server supports dynamic sparsity adjustment. When the graph density is lower than 0.1, it automatically switches to the sparse algorithm to further optimize performance.
[0062] Time value The calculation of S203. Calculate the time value based on the following formula:
[0063] Among them, is the maximum effective life cycle of the data.
[0064] Specifically, the calculation of the time value is implemented by the time series analysis module of the server in step S203, and the maximum effective life cycle of the data is preset according to the data type (for example, 7 days for news data and 30 days for financial data). The server uses the Kafka stream processing platform to collect data in real time to generate timestamps, and the time interval t is calculated in seconds. The base coefficient α and the decay rate β are determined through historical data regression analysis and stored in the MySQL database, with an update cycle of once a day.
[0065] In this embodiment, in the news recommendation system, the latest news is preferentially pushed through the calculation of the time value, and the click-through rate is increased by 12%. The server supports a multi-tenant architecture, and different users can customize , meeting personalized needs.
[0066] The calculation of the dynamic risk includes: S301. Based on the risk propagation model, calculate the dynamic risk, and the formula is:
[0067] Among them, represents the probability of the r-th type of risk occurring in the range [0,1], which is statistically analyzed by historical events, represents the corresponding risk loss amount, represents the risk time distribution model, where, represents the risk probability density value corresponding to time t, that is, the relative possibility of the risk occurring near time t. t represents the time variable, which can be absolute time or relative time, represents the average time point of the risk occurrence, that is, the central position of the risk distribution, represents the dispersion degree of the risk distribution, reflecting the uncertainty of the risk time. η ∈ [0,1] represents the risk propagation attenuation coefficient, and λ ij ∈ [0,1] represents the propagation coefficient = the proportion of shared sensitive fields × the access frequency the proportion of shared sensitive fields × the access frequency, represents the data 's set of direct associated nodes.
[0068] Specifically, the calculation of the dynamic risk is executed by the risk analysis engine of the server in step S301. The risk propagation model stores the risk propagation path based on the DGraph database. The risk occurrence probability p r is statistically analyzed through historical event logs (such as data leakage records), and the probability distribution is updated using a Bayesian network. The risk loss amount L rCalculated through a monetization assessment model, combined with industry standards (such as GDPR fine rules). The risk time distribution model uses Gaussian distribution fitting, and the parameters are determined by maximum likelihood estimation. The server supports multi-dimensional risk analysis, including data leakage, tampering, and access anomalies, with a risk propagation coefficient λ ij Calculated through shared field analysis (based on SQL JOIN) and access log statistics (based on Elasticsearch). In this embodiment, high-risk transaction nodes are successfully detected in bank data assessment, and the early warning accuracy rate reaches 95%.
[0069] The dynamic risk is optimized by the Monte Carlo simulation algorithm, specifically including: S3011. Initialize the risk seed vector; S3012. Iteratively propagate the risk based on the adjacency matrix and perform a predetermined number of simulations; S3013. Average the risk matrix to estimate the propagated risk.
[0070] Specifically, the Monte Carlo simulation is implemented through the simulation module of the server in steps S3011 - S3013. The risk seed vector is randomly initialized based on historical risk events in step S3011, with the number of seeds being 1000. Step S3012 uses MPI (Message Passing Interface) to perform iterative propagation in parallel, updating the risk values in the adjacency matrix for each iteration, and the number of iterations is 500 times. Step S3013 averages the risk matrix through the MapReduce framework to generate the final risk estimate value, with the accuracy controlled within 0.01.
[0071] In the supply chain data assessment, potential risk paths (such as supplier data leakage) are identified through the Monte Carlo simulation. The risk assessment time is shortened from 10 seconds to 2 seconds, and the efficiency is increased by 5 times. The server supports dynamically adjusting the simulation scale to adapt to different data scales.
[0072] The construction of the hybrid model includes: S204. Adopt a graph attention network to capture the correlation between data and calculate the network value; S205. Adopt a long short-term memory network to encode the temporal features and calculate the time value; S206. Combine the outputs of the graph attention network and the long short-term memory network through a fusion layer to calculate the total transaction value.
[0073] Specifically, the construction of the hybrid model is implemented through the deep learning framework of the server in steps S204 - S206. The GAT model in step S204 is based on DGL (DeepGraphLibrary), with a node embedding dimension of 64 and 8 attention heads. The LSTM model in step S205 adopts a two - layer structure, with a hidden layer dimension of 128. During training, gradient clipping (threshold 1.0) is used to prevent gradient explosion. The fusion layer in step S206 is implemented through a fully - connected layer, with the activation function ReLU, and outputs the total transaction value vector.
[0074] The server uses the Kubernetes cluster management model for deployment, supports model hot - updating, and the update period is 1 week. In the data evaluation of the e - commerce platform in this embodiment, the hybrid model reduces the MAE to 0.087, which is better than the traditional regression model (MAE 0.15), significantly improving the value prediction accuracy.
[0075] A data dynamic transaction value evaluation system includes: A data receiving module, used to execute step S1 to receive a data set from a data source; A value calculation module, used to execute step S2 to calculate the total transaction value; A risk monitoring module, used to execute step S3 to perform real - time risk monitoring and send early warning information; An output module, used to execute step S3 to send the total transaction value and early warning information to the user terminal.
[0076] Specifically, the data dynamic transaction value evaluation system is deployed in the cloud, adopts a microservices architecture, and includes the following modules: Data receiving module: Implemented based on SpringCloudGateway, receives HTTP / HTTPS requests, and supports 100,000 QPS (queries per second). The data is stored in the TiDB distributed database to ensure high availability.
[0077] Value calculation module: Integrates the GAT - LSTM model, is deployed using TensorFlowServing, supports batch inference, and it takes about 200 milliseconds to process 1000 pieces of data in a single batch.
[0078] Risk monitoring module: Based on the Prometheus monitoring system, real - time collects risk metrics, combines with Grafana to visualize the risk trend, and the abnormal detection response time is less than 50 milliseconds.
[0079] Output module: Pushes the results through Nginx load balancing, supports multi - end synchronization (Web, mobile), and the results are returned in JSON format, including value, risk, and timestamp.
[0080] The following is an introduction in combination with specific implementation examples: 1. Data Preparation Phase (1) Data Collection: Collect relevant data from various data sources, including the attributes of the data itself, association relationships, and historical operation records, etc.
[0081] (2) Data Cleaning: Clean the collected data, remove noisy data, fill in missing values, etc.
[0082] (3) Data Preprocessing: Standardize and normalize the cleaned data so that the subsequent model can process it better.
[0083] 2. Model Construction (1) Feature Extraction: Extract relevant features according to the characteristics and requirements of the data, including intrinsic value features, network value features, and time value features, etc.
[0084] (2) Model Selection: Select a suitable machine learning or deep learning model. Here, a GAT - LSTM hybrid model is used to construct a data value evaluation model.
[0085] (3) Model Training: Use the prepared dataset to train the model and adjust the model parameters to improve the evaluation accuracy.
[0086] (4) Model Construction Technical Innovation: Graph Neural Network Architecture Innovation Based on Heterogeneous Data Fusion (GAT - LSTM Hybrid Model) A. Graph Attention Network Topology Optimization Propose a multi - hop attention propagation mechanism to reduce the computational complexity through sparse processing of the adjacency matrix: # Adjacency Matrix Compression Algorithm Implementation (PyTorch Example)
[0087] Design a dynamic edge weight update algorithm, combined with node degree centrality features:
[0088] B. Temporal Feature Encoding Optimization Construct a bidirectional LSTM temporal encoder, introducing a time decay gating mechanism:
[0089] Develop a sliding window sampling algorithm to achieve efficient storage of historical value trajectories. The pseudo - code is as follows: class CircularBuffer: def __init__(self, capacity): self.buffer = np.zeros(capacity) self.index = 0 def add(self, value): self.buffer[self.index % len(self.buffer)] = value self.index += 1 3. Dynamic Evaluation System Setup (1) Real-time Data Access: Establish a real-time data access module to ensure timely acquisition of the latest data and operation records.
[0090] (2) Value Calculation and Update: Dynamically calculate the data value based on real-time data and the trained model, and update the evaluation results in a timely manner.
[0091] (3) Risk Monitoring and Early Warning: Monitor the risks faced by the data in real time, and send early warning messages in a timely manner when the risk exceeds the set threshold.
[0092] (4) Technological Innovation in Value Evaluation Model Training (A) Risk Propagation Path Optimization Algorithm Propose a risk path search method based on Monte Carlo simulation
[0093] (B) Real-time Risk Early Warning Engine Design a streaming data processing pipeline to achieve sub-second risk detection response; (C) Technological Innovation in Value Evaluation Model Training Mixed Precision Training Acceleration: Adopt the FP16 / FP32 mixed precision training strategy to reduce the video memory occupancy by 40%: scaler = torch.cuda.amp.GradScaler() with torch.cuda.amp.autocast(): outputs = model(inputs) loss = criterion(outputs, labels) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update() Distributed Model Training Optimization Develop a parameter server architecture to support the training of models with hundreds of billions of parameters (5) Key Technologies for Engineering Implementation Real-time Computing Engine Optimization: Achieved acceleration of kernel functions based on CUDA, reducing the value calculation time to the millisecond level:
[0094] High-concurrency Service Architecture: Designed a microservices-based deployment solution to support tens of thousands of concurrent evaluation requests per second: APIGateway → Load Balancing → Evaluation Service Cluster (Docker Containers) ↓ Redis Cache Layer ↓ TiDB Distributed Database.
[0095] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A method for evaluating the dynamic trading value of data, which is applied to a server, is characterized in that, Including the following steps: S1. Receive a data set from a data source, where the data set includes data attributes, relationships between data, and historical operation records; S2. Based on the received data set, use a hybrid model of graph attention neural network and long short-term memory network to calculate the data Under the operation of the user The total transaction value at time t , and the total transaction value is calculated by the following formula: Among them, represents data under the operation of the user at time t, the total value, is the intrinsic value of the data ; is the value of the data in the network environment, is the time value of the data ; is the dynamic risk faced by the data under the operation of the user at time t; S3. Output the total transaction value to the user terminal, and perform real-time risk monitoring based on the dynamic risk. When the risk exceeds the preset threshold, send a warning message to the user terminal.
2. The data dynamic transaction value evaluation method according to claim 1, characterized in that The intrinsic value is calculated as follows: S101. Obtain data scores of multiple quality dimensions , where the quality dimensions at least include accuracy, integrity, and timeliness; S102. Calculate the intrinsic value based on the following formula: Among them, Q m represents the score of the m th quality dimension, with a value range of [0, 1] and is standardized. is the dynamic weight coefficient, satisfying .
3. The data dynamic trading value evaluation method according to claim 2, characterized in that The calculation of the quality dimension includes: S103. Calculate the accuracy by determining the deviation ratio between the data and the true value through sampling inspection; S104. Calculate the integrity by determining the proportion of non-missing data; S105. Calculate timeliness, and adopt introduce the influence of the logarithmic function to smooth the update frequency or through the formula , where is the number of days the data exists.
4. The data dynamic trading value evaluation method according to claim 1, wherein The value in the described network environment is calculated as follows: S201. Calculate the network value of the data based on the graph attention network, and the formula is: Among them, is the set of direct associated nodes of the data , is the weight of the associated edge, and its value range is [0, 1]. After normalization processing, is the Sigmoid activation function, which is used to strengthen the key connections, is the total transaction value of the data .
5. The data dynamic trading value evaluation method according to claim 4, wherein The calculation of the network value uses a sparse matrix to accelerate, specifically including: S202. Calculate the network value based on the following formula: Among them, A is the adjacency matrix, , or is the dynamic decay factor, where is the contribution degree of the network value to the total value at time t = 0, and its value range is [0, 1], is the decay rate of the network value over time, is the total transaction value at the previous moment of time t.
6. The data dynamic transaction value evaluation method according to claim 1, wherein The time value is calculated as follows: S203. Calculate the time value based on the following formula: Among them, is the maximum effective life cycle of the data, is the base coefficient, and its value range is [0.5, 1], is the decay rate, and its value range is (0, 0.1].
7. The data dynamic trading value evaluation method according to claim 1, wherein The calculation of the dynamic risk includes: S301. Calculate the dynamic risk based on the risk propagation model, and the formula is: Among them, represents the occurrence probability of the r-th type of risk in the range [0,1], which is statistically obtained from historical events. represents the corresponding risk loss amount. represents the risk time distribution model, where represents the risk probability density value corresponding to time t, that is, the relative possibility of the risk occurring near time t. t represents the time variable, which can be absolute time or relative time. represents the average time point of risk occurrence, that is, the central position of the risk distribution. represents the dispersion degree of the risk distribution, reflecting the uncertainty of the risk time. η ∈ [0,1] represents the risk propagation attenuation coefficient, and λ ij ∈ [0,1] represents the propagation coefficient = the proportion of shared sensitive fields × access frequency / the proportion of shared sensitive fields × access frequency. represents the data of the direct associated node set.
8. The data dynamic transaction value evaluation method according to claim 1, wherein The dynamic risk is optimized by the Monte Carlo simulation algorithm, specifically including: S3011. Initialize the risk seed vector; S3012. Iteratively propagate the risk based on the adjacency matrix and perform a predetermined number of simulations; S3013. Average the risk matrix to estimate the propagated risk.
9. The data dynamic trading value evaluation method according to claim 1, characterized in that The construction of the hybrid model includes: S204. Use a graph attention network to capture the relevance between data and calculate the network value; S205. Use a long short-term memory network to encode the temporal features and calculate the time value; S206. Calculate the total transaction value by combining the outputs of the graph attention network and the long short-term memory network through a fusion layer.
10. A data dynamic trading value evaluation system, characterized in that, For the data dynamic transaction value evaluation method according to any one of claims 1-9, including: A data reception module for performing step S1 to receive a data set from a data source; A value calculation module for performing step S2 to calculate the total transaction value; A risk monitoring module for performing step S3 to perform real-time risk monitoring and send a warning message; An output module for performing step S3 to send the total transaction value and the warning message to the user terminal.
Citation Information
Cited By
Intelligent insurance pricing method based on public network information
CN120634657A
Data quality evaluation method and system based on rule engine and machine learning
CN121030265A